الملفات
ghaymah-exam-sre-abdelrahman/q2-postmortem/postmortem-report.md

4.2 KiB

Incident Postmortem Report: Memory Outage & CrashLoopBackOff

1. Incident Overview & Summary

  • Service Name: Ghaymah Core Application
  • Outage Duration: 45 Minutes
  • Severity: High (P1 - Service Outage)
  • Status: Resolved
  • Impact: HTTP 502 Bad Gateway errors observed during traffic peak due to repeated pod failures.
  • Root Cause: Container memory usage continuously exceeded its hard limit (512MB), triggering Kernel OOM Killer (Exit Code 137) and Kubernetes CrashLoopBackOff.

2. Incident Timeline

  • 12:00 UTC: Traffic surge initiated on application endpoints.
  • 12:10 UTC: Memory utilization reached 95% of defined container resource limits.
  • 12:15 UTC: Linux Kernel OOM Killer terminated the primary pod container (OOMKilled).
  • 12:16 UTC: Kubernetes entered a CrashLoopBackOff restart cycle; application became unreachable (502 Bad Gateway).
  • 12:45 UTC: SRE team identified resource constraints, updated memory limits to 2Gi, applied HPA auto-scaling, and restored normal service operations.

3. Root Cause Analysis (RCA)

  1. Inadequate Resource Allocations: Memory limits were set statically at 512MB, which was insufficient for processing concurrent request spikes.
  2. Absence of Dynamic Scaling: The deployment lacked auto-scaling rules, preventing horizontal pod expansion under load.

4. Remediation & Preventive Recommendations

  • Increased container memory resource request to 1Gi and hard limit to 2Gi.
  • Implemented Horizontal Pod Autoscaler (HPA) targeting memory and CPU thresholds.
  • Implement proactive alerting via Prometheus Alertmanager for high memory consumption.

5. Ghaymah Auto-Scaling Policy Configuration

To prevent future memory exhaustion outages, the following Kubernetes HorizontalPodAutoscaler (HPA) manifest was created and applied to scale pods dynamically:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: ghaymah-core-app-hpa
  namespace: production
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: ghaymah-core-app
  minReplicas: 3
  maxReplicas: 10
  metrics:
  - type: Resource
    resource:
      name: memory
      target:
        type: Utilization
        averageUtilization: 75
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 80
6. Early Detection via Ghaymah Monitoring Tools
To detect and mitigate memory leakage or depletion early before service degradation occurs, we leverage Ghaymah Cloud Monitoring (Prometheus & Grafana ecosystem) as follows:

A. Key Metrics to Track:
container_memory_working_set_bytes: Measures actual memory used by the container excluding cached pages.

kube_pod_container_status_restarts_total: Tracks container restart loops triggered by OOMKilled events.

container_spec_memory_limit_bytes: Monitors usage percentage against defined resource boundaries.

B. Prometheus Alerting Rules Configuration:
We configure proactive alert rules in Prometheus/Alertmanager:

YAML
groups:
- name: ghaymah_memory_alerts
  rules:
  # 1. Early Warning Alert (Memory > 85% for 3 minutes)
  - alert: HighMemoryUsageWarning
    expr: (container_memory_working_set_bytes{container="ghaymah-core-app"} / container_spec_memory_limit_bytes{container="ghaymah-core-app"}) * 100 > 85
    for: 3m
    labels:
      severity: warning
    annotations:
      summary: "High Memory Utilization on Ghaymah Pod"
      description: "Pod {{ $labels.pod }} memory usage is above 85% for more than 3 minutes."

  # 2. Critical Alert (OOM Kill Detected)
  - alert: ContainerOOMKilledCritical
    expr: increase(kube_pod_container_status_restarts_total{container="ghaymah-core-app"}[5m]) > 0
    for: 0m
    labels:
      severity: critical
    annotations:
      summary: "Container Restarted due to OOMKilled"
      description: "Pod {{ $labels.pod }} was killed by Linux OOM Killer."
C. Grafana Visual Dashboard:
Setup a real-time Memory Threshold gauge with visual color indicators (Yellow at 75%, Red at 90%).

Enable automated PagerDuty / Slack notifications when the HighMemoryUsageWarning fires, allowing SRE engineers to intervene before pod crashes occur.