# πŸ”΄ Incident Postmortem Report β€” OOMKilled Application Crash **Incident ID:** INC-2026-0726 **Severity:** P1 β€” Critical **Duration:** 45 minutes **Date:** July 26, 2026 **Author:** DevOps Team **Status:** Resolved --- ## 1. Executive Summary On July 26, 2026, the **Ghayma REST API** application hosted on a cloud container platform experienced a **45-minute outage** caused by repeated **OOMKilled** (Out of Memory Killed) events. The container runtime terminated the application process multiple times after it exceeded its allocated memory limit. Each restart triggered the same memory spike, creating a crash loop that prevented the service from recovering. **Impact:** - 100% service unavailability for 45 minutes - All API consumers (frontend clients, monitoring, health probes) received connection errors - Approximately **2,700 failed requests** during the outage window (estimated ~60 req/min baseline) - Container orchestrator marked the pod/container as `CrashLoopBackOff` after repeated restart failures --- ## 2. Incident Timeline | Time (UTC) | Event | | :--- | :--- | | **11:00** | 🟒 Application running normally. Memory usage stable at ~120MB (limit: 256MB) | | **11:12** | πŸ“ˆ Traffic spike begins β€” external batch job sends large payload requests to `POST /api/v1/items` | | **11:15** | ⚠️ Memory usage crosses **200MB** (78% of limit). No alerts triggered | | **11:18** | πŸ”΄ Memory hits **256MB** limit. Kernel OOM killer terminates the container process (`OOMKilled`) | | **11:18** | πŸ”„ Container runtime automatically restarts the container (Restart #1) | | **11:20** | πŸ”΄ Application starts, loads cached data into memory, immediately OOMKilled again (Restart #2) | | **11:20–11:45** | πŸ” **Crash loop** β€” container restarts 12 times. Orchestrator applies exponential backoff (`CrashLoopBackOff`) | | **11:32** | πŸ“Ÿ On-call engineer alerted via PagerDuty after health check failures exceed 10 minutes | | **11:38** | πŸ” Engineer identifies OOMKilled events in container logs and platform event stream | | **11:45** | πŸ› οΈ Engineer increases memory limit from **256MB β†’ 512MB** and deploys hotfix | | **11:48** | πŸ”„ Application restarts successfully. Memory stabilizes at ~180MB | | **11:50** | πŸ“‰ Batch job completes. Traffic returns to normal levels | | **12:03** | 🟒 Service fully confirmed stable. Incident resolved | **Total downtime:** 45 minutes (11:18 β€” 12:03 UTC) --- ## 3. Root Cause Analysis ### 3.1 Direct Cause The container was configured with a **memory limit of 256MB**, which was insufficient to handle traffic spikes. When an external batch job sent a burst of large `POST` requests with sizable JSON payloads, the Node.js process accumulated in-memory data (parsed request bodies, in-memory item storage, response buffers) beyond the container's limit. The Linux kernel's **OOM Killer** terminated the process when it attempted to allocate memory beyond the `256MB` cgroup limit. ### 3.2 Contributing Factors ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ ROOT CAUSE BREAKDOWN β”‚ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”‚ β”‚ β”‚ 1. INSUFFICIENT MEMORY LIMIT β”‚ β”‚ └─ Container limited to 256MB (too tight for Node.js) β”‚ β”‚ β”‚ β”‚ 2. NO MEMORY-AWARE AUTO-SCALING β”‚ β”‚ └─ Only 1 replica running, no HPA/scaling policy β”‚ β”‚ β”‚ β”‚ 3. UNBOUNDED IN-MEMORY DATA STORE β”‚ β”‚ └─ Items array grows indefinitely with no cap β”‚ β”‚ β”‚ β”‚ 4. NO MEMORY ALERTS / EARLY WARNING β”‚ β”‚ └─ Alert only triggered after 10 min of health failures β”‚ β”‚ β”‚ β”‚ 5. NO REQUEST PAYLOAD SIZE LIMIT β”‚ β”‚ └─ Express accepted arbitrarily large JSON bodies β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` ### 3.3 Why Did Restarts Fail to Recover? Each restart reloaded the same conditions: 1. Container starts β†’ Node.js initializes (~80MB baseline) 2. Pending queued requests immediately hit the API 3. Large payloads + in-memory storage push memory past the limit 4. OOMKilled again within seconds 5. Orchestrator enters **CrashLoopBackOff** with increasing delays between restarts --- ## 4. Remediation Actions Taken | Action | Status | | :--- | :--- | | Increased container memory limit to 512MB | βœ… Done (hotfix) | | Restarted container with new limits | βœ… Done | | Confirmed service recovery and stability | βœ… Done | --- ## 5. Recommendations & Prevention ### 5.1 Application-Level Fixes #### A. Limit Request Payload Size ```javascript // In src/index.js β€” restrict incoming JSON body size app.use(express.json({ limit: '1mb' })); ``` > Prevents a single large request from consuming excessive memory. #### B. Cap In-Memory Data Store ```javascript // Limit the items array to prevent unbounded growth const MAX_ITEMS = 1000; app.post('/api/v1/items', (req, res) => { if (items.length >= MAX_ITEMS) { return res.status(429).json({ success: false, error: 'Maximum item limit reached' }); } // ... rest of handler }); ``` #### C. Set Node.js Memory Ceiling ```dockerfile # In Dockerfile β€” set explicit V8 heap limit CMD ["node", "--max-old-space-size=384", "src/index.js"] ``` > Ensures Node.js garbage collector runs more aggressively before hitting the container limit. #### D. Use External Storage for Production Replace the in-memory `items` array with a proper database (Redis, PostgreSQL, MongoDB) so application memory remains constant regardless of data volume. --- ### 5.2 Container & Infrastructure Fixes #### A. Set Proper Resource Requests & Limits ```yaml # Kubernetes Deployment example resources: requests: memory: "256Mi" # Guaranteed minimum cpu: "100m" limits: memory: "512Mi" # Hard ceiling cpu: "500m" ``` > [!IMPORTANT] > **Rule of thumb:** Set `limits.memory` to at least **2x** the average working set. Set `requests.memory` to the **steady-state average**. #### B. Add Graceful Shutdown Handling ```javascript process.on('SIGTERM', () => { console.log('SIGTERM received. Shutting down gracefully...'); server.close(() => { process.exit(0); }); }); ``` --- ## 6. Auto-Scaling Policy Design ### 6.1 Horizontal Pod Autoscaler (HPA) Policy The following auto-scaling policy prevents a single container from being overwhelmed by distributing load across multiple replicas: ```yaml apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: ghayma-api-hpa namespace: production spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: ghayma-api # Replica bounds minReplicas: 2 # Always run at least 2 for high availability maxReplicas: 10 # Cap to control costs # Scaling metrics metrics: # Scale based on MEMORY usage (primary - prevents OOMKill) - type: Resource resource: name: memory target: type: Utilization averageUtilization: 70 # Scale up when memory > 70% of limit # Scale based on CPU usage (secondary) - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 75 # Scale up when CPU > 75% # Scaling behavior (prevents flapping) behavior: scaleUp: stabilizationWindowSeconds: 30 # React quickly to spikes policies: - type: Pods value: 2 # Add up to 2 pods at a time periodSeconds: 60 scaleDown: stabilizationWindowSeconds: 300 # Wait 5 min before scaling down policies: - type: Pods value: 1 # Remove 1 pod at a time periodSeconds: 120 ``` ### 6.2 How This Policy Prevents Duplication of the Incident ``` Traffic Spike Detected β”‚ β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Memory usage > 70% │──── YES ──▢ HPA adds new replicas (up to 10) β”‚ on existing pods? β”‚ Load balancer distributes traffic β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ Memory per-pod stays under limit β”‚ NO β”‚ β–Ό β–Ό Normal operation Spike absorbed across multiple pods No single pod reaches OOMKill threshold ``` ### 6.3 Cloud Platform Auto-Scaling (Non-Kubernetes) For managed container platforms (e.g., Ghayma Cloud, AWS ECS, Google Cloud Run): | Setting | Value | Rationale | | :--- | :--- | :--- | | **Min instances** | 2 | Avoid cold starts & single point of failure | | **Max instances** | 10 | Cost ceiling | | **Scale-up trigger** | Memory > 70% OR CPU > 75% OR Concurrent requests > 50 | Multi-signal scaling | | **Scale-up cooldown** | 30 seconds | React quickly to spikes | | **Scale-down cooldown** | 5 minutes | Prevent flapping during variable traffic | | **Memory per instance** | 512MB | 2x steady-state working set | --- ## 7. Early Detection β€” Monitoring & Alerting Strategy ### 7.1 Key Metrics to Monitor | Metric | Source | Warning Threshold | Critical Threshold | | :--- | :--- | :--- | :--- | | **Container Memory Usage** | cAdvisor / Platform metrics | > 70% of limit | > 85% of limit | | **Container Restart Count** | Kubelet / Platform events | β‰₯ 1 restart in 5 min | β‰₯ 3 restarts in 10 min | | **OOMKilled Events** | Kernel / Container runtime | Any occurrence | N/A (always critical) | | **Response Time (p95)** | Application `/metrics` endpoint | > 500ms | > 2000ms | | **Error Rate (5xx)** | Load balancer / Application | > 1% | > 5% | | **Health Check Failures** | Platform health probe | 1 consecutive failure | 3 consecutive failures | | **Request Queue Depth** | Load balancer | > 50 pending | > 200 pending | ### 7.2 Alerting Rules (Prometheus / Grafana Example) ```yaml # Alert: Memory approaching container limit - alert: HighMemoryUsage expr: | (container_memory_usage_bytes{container="ghayma-api"} / container_spec_memory_limit_bytes{container="ghayma-api"}) > 0.70 for: 2m labels: severity: warning annotations: summary: "Ghayma API memory usage above 70%" description: "Container {{ $labels.pod }} memory at {{ $value | humanizePercentage }}" # Alert: OOMKill detected (CRITICAL β€” immediate page) - alert: OOMKillDetected expr: | increase(kube_pod_container_status_restarts_total{container="ghayma-api"}[5m]) > 0 and kube_pod_container_status_last_terminated_reason{container="ghayma-api"} == "OOMKilled" for: 0m labels: severity: critical annotations: summary: "OOMKill detected on Ghayma API" description: "Pod {{ $labels.pod }} was OOMKilled. Immediate investigation required." # Alert: High error rate - alert: HighErrorRate expr: | rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) > 0.05 for: 3m labels: severity: critical annotations: summary: "Ghayma API error rate exceeds 5%" ``` ### 7.3 Monitoring Architecture ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ MONITORING STACK β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ Ghayma │───▢│ Prometheus │───▢│ Grafana Dashboard β”‚ β”‚ β”‚ β”‚ API β”‚ β”‚ (scrapes β”‚ β”‚ (visualization + β”‚ β”‚ β”‚ β”‚ /metrics β”‚ β”‚ /metrics) β”‚ β”‚ alerting rules) β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β”‚ β”‚ β–Ό β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ Alert Manager β”‚ β”‚ β”‚ β”‚ (routing & β”‚ β”‚ β”‚ β”‚ deduplication)β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β–Ό β–Ό β–Ό β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ Slack β”‚ β”‚ PagerDutyβ”‚ β”‚ Email β”‚ β”‚ β”‚ β”‚ #alerts β”‚ β”‚ on-call β”‚ β”‚ team β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` ### 7.4 Cloud Platform Native Tools | Cloud Platform | Monitoring Service | Key Feature for OOM Detection | | :--- | :--- | :--- | | **AWS ECS / Fargate** | CloudWatch Container Insights | `MemoryUtilization` metric + CloudWatch Alarms | | **Google Cloud Run** | Cloud Monitoring | `container/memory/utilization` + Alert Policies | | **Azure Container Apps** | Azure Monitor | `MemoryWorkingSet` + Metric Alerts | | **Kubernetes (any)** | Prometheus + Grafana | `container_memory_usage_bytes` + `kube_pod_container_status_last_terminated_reason` | ### 7.5 Recommended Grafana Dashboard Panels 1. **Memory Usage vs Limit** β€” Time series showing memory consumption relative to the container limit (with 70% and 85% threshold lines) 2. **Container Restarts** β€” Counter panel showing restart events over time 3. **Response Time Heatmap** β€” p50 / p95 / p99 latency distribution 4. **Request Rate & Error Rate** β€” Dual-axis graph (total requests + 5xx rate) 5. **Pod/Replica Count** β€” Shows HPA scaling activity over time --- ## 8. Lessons Learned | # | Lesson | Action Item | | :--- | :--- | :--- | | 1 | Default memory limits were set too low without load testing | Conduct load tests before production deployment | | 2 | Single-replica deployment has no resilience | Always run **β‰₯ 2 replicas** for production services | | 3 | Alert threshold (10 min) was too slow for a P1 outage | Reduce critical alert threshold to **2 minutes** | | 4 | In-memory data stores are dangerous without bounds | Use external databases or enforce collection size limits | | 5 | No runbook existed for OOMKilled incidents | Create and publish an OOMKilled response runbook | --- ## 9. Action Items Tracker | # | Action Item | Owner | Priority | Deadline | Status | | :--- | :--- | :--- | :--- | :--- | :--- | | 1 | Increase container memory limit to 512MB | DevOps | P0 | Done | βœ… | | 2 | Add `express.json({ limit: '1mb' })` payload limit | Backend | P1 | +1 day | ⬜ | | 3 | Implement HPA with memory-based scaling | DevOps | P1 | +3 days | ⬜ | | 4 | Set up Prometheus memory alerts (70% / 85%) | DevOps | P1 | +3 days | ⬜ | | 5 | Add OOMKilled alert rule (immediate paging) | DevOps | P0 | +1 day | ⬜ | | 6 | Run minimum 2 replicas in production | DevOps | P1 | +1 day | ⬜ | | 7 | Replace in-memory store with Redis/DB | Backend | P2 | +1 week | ⬜ | | 8 | Conduct load testing with realistic traffic | QA | P2 | +2 weeks | ⬜ | | 9 | Create OOMKilled incident runbook | DevOps | P2 | +1 week | ⬜ | | 10 | Build Grafana dashboard for container metrics | DevOps | P2 | +1 week | ⬜ | --- > **Report prepared by:** DevOps Team > **Review date:** July 26, 2026 > **Next review:** August 2, 2026 (verify all P0/P1 items completed)