feat: Ghayma Assessment - complete submission
هذا الالتزام موجود في:
387
q2-postmortem/postmortem-report.md
Normal file
387
q2-postmortem/postmortem-report.md
Normal file
@@ -0,0 +1,387 @@
|
||||
# 🔴 Incident Postmortem Report — OOMKilled Application Crash
|
||||
|
||||
**Incident ID:** INC-2026-0726
|
||||
**Severity:** P1 — Critical
|
||||
**Duration:** 45 minutes
|
||||
**Date:** July 26, 2026
|
||||
**Author:** DevOps Team
|
||||
**Status:** Resolved
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive Summary
|
||||
|
||||
On July 26, 2026, the **Ghayma REST API** application hosted on a cloud container platform experienced a **45-minute outage** caused by repeated **OOMKilled** (Out of Memory Killed) events. The container runtime terminated the application process multiple times after it exceeded its allocated memory limit. Each restart triggered the same memory spike, creating a crash loop that prevented the service from recovering.
|
||||
|
||||
**Impact:**
|
||||
- 100% service unavailability for 45 minutes
|
||||
- All API consumers (frontend clients, monitoring, health probes) received connection errors
|
||||
- Approximately **2,700 failed requests** during the outage window (estimated ~60 req/min baseline)
|
||||
- Container orchestrator marked the pod/container as `CrashLoopBackOff` after repeated restart failures
|
||||
|
||||
---
|
||||
|
||||
## 2. Incident Timeline
|
||||
|
||||
| Time (UTC) | Event |
|
||||
| :--- | :--- |
|
||||
| **11:00** | 🟢 Application running normally. Memory usage stable at ~120MB (limit: 256MB) |
|
||||
| **11:12** | 📈 Traffic spike begins — external batch job sends large payload requests to `POST /api/v1/items` |
|
||||
| **11:15** | ⚠️ Memory usage crosses **200MB** (78% of limit). No alerts triggered |
|
||||
| **11:18** | 🔴 Memory hits **256MB** limit. Kernel OOM killer terminates the container process (`OOMKilled`) |
|
||||
| **11:18** | 🔄 Container runtime automatically restarts the container (Restart #1) |
|
||||
| **11:20** | 🔴 Application starts, loads cached data into memory, immediately OOMKilled again (Restart #2) |
|
||||
| **11:20–11:45** | 🔁 **Crash loop** — container restarts 12 times. Orchestrator applies exponential backoff (`CrashLoopBackOff`) |
|
||||
| **11:32** | 📟 On-call engineer alerted via PagerDuty after health check failures exceed 10 minutes |
|
||||
| **11:38** | 🔍 Engineer identifies OOMKilled events in container logs and platform event stream |
|
||||
| **11:45** | 🛠️ Engineer increases memory limit from **256MB → 512MB** and deploys hotfix |
|
||||
| **11:48** | 🔄 Application restarts successfully. Memory stabilizes at ~180MB |
|
||||
| **11:50** | 📉 Batch job completes. Traffic returns to normal levels |
|
||||
| **12:03** | 🟢 Service fully confirmed stable. Incident resolved |
|
||||
|
||||
**Total downtime:** 45 minutes (11:18 — 12:03 UTC)
|
||||
|
||||
---
|
||||
|
||||
## 3. Root Cause Analysis
|
||||
|
||||
### 3.1 Direct Cause
|
||||
|
||||
The container was configured with a **memory limit of 256MB**, which was insufficient to handle traffic spikes. When an external batch job sent a burst of large `POST` requests with sizable JSON payloads, the Node.js process accumulated in-memory data (parsed request bodies, in-memory item storage, response buffers) beyond the container's limit.
|
||||
|
||||
The Linux kernel's **OOM Killer** terminated the process when it attempted to allocate memory beyond the `256MB` cgroup limit.
|
||||
|
||||
### 3.2 Contributing Factors
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────────────────┐
|
||||
│ ROOT CAUSE BREAKDOWN │
|
||||
├──────────────────────────────────────────────────────────────────┤
|
||||
│ │
|
||||
│ 1. INSUFFICIENT MEMORY LIMIT │
|
||||
│ └─ Container limited to 256MB (too tight for Node.js) │
|
||||
│ │
|
||||
│ 2. NO MEMORY-AWARE AUTO-SCALING │
|
||||
│ └─ Only 1 replica running, no HPA/scaling policy │
|
||||
│ │
|
||||
│ 3. UNBOUNDED IN-MEMORY DATA STORE │
|
||||
│ └─ Items array grows indefinitely with no cap │
|
||||
│ │
|
||||
│ 4. NO MEMORY ALERTS / EARLY WARNING │
|
||||
│ └─ Alert only triggered after 10 min of health failures │
|
||||
│ │
|
||||
│ 5. NO REQUEST PAYLOAD SIZE LIMIT │
|
||||
│ └─ Express accepted arbitrarily large JSON bodies │
|
||||
│ │
|
||||
└──────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### 3.3 Why Did Restarts Fail to Recover?
|
||||
|
||||
Each restart reloaded the same conditions:
|
||||
1. Container starts → Node.js initializes (~80MB baseline)
|
||||
2. Pending queued requests immediately hit the API
|
||||
3. Large payloads + in-memory storage push memory past the limit
|
||||
4. OOMKilled again within seconds
|
||||
5. Orchestrator enters **CrashLoopBackOff** with increasing delays between restarts
|
||||
|
||||
---
|
||||
|
||||
## 4. Remediation Actions Taken
|
||||
|
||||
| Action | Status |
|
||||
| :--- | :--- |
|
||||
| Increased container memory limit to 512MB | ✅ Done (hotfix) |
|
||||
| Restarted container with new limits | ✅ Done |
|
||||
| Confirmed service recovery and stability | ✅ Done |
|
||||
|
||||
---
|
||||
|
||||
## 5. Recommendations & Prevention
|
||||
|
||||
### 5.1 Application-Level Fixes
|
||||
|
||||
#### A. Limit Request Payload Size
|
||||
```javascript
|
||||
// In src/index.js — restrict incoming JSON body size
|
||||
app.use(express.json({ limit: '1mb' }));
|
||||
```
|
||||
> Prevents a single large request from consuming excessive memory.
|
||||
|
||||
#### B. Cap In-Memory Data Store
|
||||
```javascript
|
||||
// Limit the items array to prevent unbounded growth
|
||||
const MAX_ITEMS = 1000;
|
||||
|
||||
app.post('/api/v1/items', (req, res) => {
|
||||
if (items.length >= MAX_ITEMS) {
|
||||
return res.status(429).json({
|
||||
success: false,
|
||||
error: 'Maximum item limit reached'
|
||||
});
|
||||
}
|
||||
// ... rest of handler
|
||||
});
|
||||
```
|
||||
|
||||
#### C. Set Node.js Memory Ceiling
|
||||
```dockerfile
|
||||
# In Dockerfile — set explicit V8 heap limit
|
||||
CMD ["node", "--max-old-space-size=384", "src/index.js"]
|
||||
```
|
||||
> Ensures Node.js garbage collector runs more aggressively before hitting the container limit.
|
||||
|
||||
#### D. Use External Storage for Production
|
||||
Replace the in-memory `items` array with a proper database (Redis, PostgreSQL, MongoDB) so application memory remains constant regardless of data volume.
|
||||
|
||||
---
|
||||
|
||||
### 5.2 Container & Infrastructure Fixes
|
||||
|
||||
#### A. Set Proper Resource Requests & Limits
|
||||
|
||||
```yaml
|
||||
# Kubernetes Deployment example
|
||||
resources:
|
||||
requests:
|
||||
memory: "256Mi" # Guaranteed minimum
|
||||
cpu: "100m"
|
||||
limits:
|
||||
memory: "512Mi" # Hard ceiling
|
||||
cpu: "500m"
|
||||
```
|
||||
|
||||
> [!IMPORTANT]
|
||||
> **Rule of thumb:** Set `limits.memory` to at least **2x** the average working set. Set `requests.memory` to the **steady-state average**.
|
||||
|
||||
#### B. Add Graceful Shutdown Handling
|
||||
```javascript
|
||||
process.on('SIGTERM', () => {
|
||||
console.log('SIGTERM received. Shutting down gracefully...');
|
||||
server.close(() => {
|
||||
process.exit(0);
|
||||
});
|
||||
});
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Auto-Scaling Policy Design
|
||||
|
||||
### 6.1 Horizontal Pod Autoscaler (HPA) Policy
|
||||
|
||||
The following auto-scaling policy prevents a single container from being overwhelmed by distributing load across multiple replicas:
|
||||
|
||||
```yaml
|
||||
apiVersion: autoscaling/v2
|
||||
kind: HorizontalPodAutoscaler
|
||||
metadata:
|
||||
name: ghayma-api-hpa
|
||||
namespace: production
|
||||
spec:
|
||||
scaleTargetRef:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
name: ghayma-api
|
||||
|
||||
# Replica bounds
|
||||
minReplicas: 2 # Always run at least 2 for high availability
|
||||
maxReplicas: 10 # Cap to control costs
|
||||
|
||||
# Scaling metrics
|
||||
metrics:
|
||||
# Scale based on MEMORY usage (primary - prevents OOMKill)
|
||||
- type: Resource
|
||||
resource:
|
||||
name: memory
|
||||
target:
|
||||
type: Utilization
|
||||
averageUtilization: 70 # Scale up when memory > 70% of limit
|
||||
|
||||
# Scale based on CPU usage (secondary)
|
||||
- type: Resource
|
||||
resource:
|
||||
name: cpu
|
||||
target:
|
||||
type: Utilization
|
||||
averageUtilization: 75 # Scale up when CPU > 75%
|
||||
|
||||
# Scaling behavior (prevents flapping)
|
||||
behavior:
|
||||
scaleUp:
|
||||
stabilizationWindowSeconds: 30 # React quickly to spikes
|
||||
policies:
|
||||
- type: Pods
|
||||
value: 2 # Add up to 2 pods at a time
|
||||
periodSeconds: 60
|
||||
scaleDown:
|
||||
stabilizationWindowSeconds: 300 # Wait 5 min before scaling down
|
||||
policies:
|
||||
- type: Pods
|
||||
value: 1 # Remove 1 pod at a time
|
||||
periodSeconds: 120
|
||||
```
|
||||
|
||||
### 6.2 How This Policy Prevents Duplication of the Incident
|
||||
|
||||
```
|
||||
Traffic Spike Detected
|
||||
│
|
||||
▼
|
||||
┌───────────────────────┐
|
||||
│ Memory usage > 70% │──── YES ──▶ HPA adds new replicas (up to 10)
|
||||
│ on existing pods? │ Load balancer distributes traffic
|
||||
└───────────────────────┘ Memory per-pod stays under limit
|
||||
│ NO │
|
||||
▼ ▼
|
||||
Normal operation Spike absorbed across multiple pods
|
||||
No single pod reaches OOMKill threshold
|
||||
```
|
||||
|
||||
### 6.3 Cloud Platform Auto-Scaling (Non-Kubernetes)
|
||||
|
||||
For managed container platforms (e.g., Ghayma Cloud, AWS ECS, Google Cloud Run):
|
||||
|
||||
| Setting | Value | Rationale |
|
||||
| :--- | :--- | :--- |
|
||||
| **Min instances** | 2 | Avoid cold starts & single point of failure |
|
||||
| **Max instances** | 10 | Cost ceiling |
|
||||
| **Scale-up trigger** | Memory > 70% OR CPU > 75% OR Concurrent requests > 50 | Multi-signal scaling |
|
||||
| **Scale-up cooldown** | 30 seconds | React quickly to spikes |
|
||||
| **Scale-down cooldown** | 5 minutes | Prevent flapping during variable traffic |
|
||||
| **Memory per instance** | 512MB | 2x steady-state working set |
|
||||
|
||||
---
|
||||
|
||||
## 7. Early Detection — Monitoring & Alerting Strategy
|
||||
|
||||
### 7.1 Key Metrics to Monitor
|
||||
|
||||
| Metric | Source | Warning Threshold | Critical Threshold |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| **Container Memory Usage** | cAdvisor / Platform metrics | > 70% of limit | > 85% of limit |
|
||||
| **Container Restart Count** | Kubelet / Platform events | ≥ 1 restart in 5 min | ≥ 3 restarts in 10 min |
|
||||
| **OOMKilled Events** | Kernel / Container runtime | Any occurrence | N/A (always critical) |
|
||||
| **Response Time (p95)** | Application `/metrics` endpoint | > 500ms | > 2000ms |
|
||||
| **Error Rate (5xx)** | Load balancer / Application | > 1% | > 5% |
|
||||
| **Health Check Failures** | Platform health probe | 1 consecutive failure | 3 consecutive failures |
|
||||
| **Request Queue Depth** | Load balancer | > 50 pending | > 200 pending |
|
||||
|
||||
### 7.2 Alerting Rules (Prometheus / Grafana Example)
|
||||
|
||||
```yaml
|
||||
# Alert: Memory approaching container limit
|
||||
- alert: HighMemoryUsage
|
||||
expr: |
|
||||
(container_memory_usage_bytes{container="ghayma-api"}
|
||||
/ container_spec_memory_limit_bytes{container="ghayma-api"}) > 0.70
|
||||
for: 2m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Ghayma API memory usage above 70%"
|
||||
description: "Container {{ $labels.pod }} memory at {{ $value | humanizePercentage }}"
|
||||
|
||||
# Alert: OOMKill detected (CRITICAL — immediate page)
|
||||
- alert: OOMKillDetected
|
||||
expr: |
|
||||
increase(kube_pod_container_status_restarts_total{container="ghayma-api"}[5m]) > 0
|
||||
and kube_pod_container_status_last_terminated_reason{container="ghayma-api"} == "OOMKilled"
|
||||
for: 0m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "OOMKill detected on Ghayma API"
|
||||
description: "Pod {{ $labels.pod }} was OOMKilled. Immediate investigation required."
|
||||
|
||||
# Alert: High error rate
|
||||
- alert: HighErrorRate
|
||||
expr: |
|
||||
rate(http_requests_total{status=~"5.."}[5m])
|
||||
/ rate(http_requests_total[5m]) > 0.05
|
||||
for: 3m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Ghayma API error rate exceeds 5%"
|
||||
```
|
||||
|
||||
### 7.3 Monitoring Architecture
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ MONITORING STACK │
|
||||
│ │
|
||||
│ ┌──────────┐ ┌─────────────┐ ┌──────────────────────┐ │
|
||||
│ │ Ghayma │───▶│ Prometheus │───▶│ Grafana Dashboard │ │
|
||||
│ │ API │ │ (scrapes │ │ (visualization + │ │
|
||||
│ │ /metrics │ │ /metrics) │ │ alerting rules) │ │
|
||||
│ └──────────┘ └──────┬──────┘ └──────────────────────┘ │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ ┌─────────────────┐ │
|
||||
│ │ Alert Manager │ │
|
||||
│ │ (routing & │ │
|
||||
│ │ deduplication)│ │
|
||||
│ └────────┬────────┘ │
|
||||
│ │ │
|
||||
│ ┌────────────┼─────────────┐ │
|
||||
│ ▼ ▼ ▼ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌───────────┐ │
|
||||
│ │ Slack │ │ PagerDuty│ │ Email │ │
|
||||
│ │ #alerts │ │ on-call │ │ team │ │
|
||||
│ └──────────┘ └──────────┘ └───────────┘ │
|
||||
└─────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### 7.4 Cloud Platform Native Tools
|
||||
|
||||
| Cloud Platform | Monitoring Service | Key Feature for OOM Detection |
|
||||
| :--- | :--- | :--- |
|
||||
| **AWS ECS / Fargate** | CloudWatch Container Insights | `MemoryUtilization` metric + CloudWatch Alarms |
|
||||
| **Google Cloud Run** | Cloud Monitoring | `container/memory/utilization` + Alert Policies |
|
||||
| **Azure Container Apps** | Azure Monitor | `MemoryWorkingSet` + Metric Alerts |
|
||||
| **Kubernetes (any)** | Prometheus + Grafana | `container_memory_usage_bytes` + `kube_pod_container_status_last_terminated_reason` |
|
||||
|
||||
### 7.5 Recommended Grafana Dashboard Panels
|
||||
|
||||
1. **Memory Usage vs Limit** — Time series showing memory consumption relative to the container limit (with 70% and 85% threshold lines)
|
||||
2. **Container Restarts** — Counter panel showing restart events over time
|
||||
3. **Response Time Heatmap** — p50 / p95 / p99 latency distribution
|
||||
4. **Request Rate & Error Rate** — Dual-axis graph (total requests + 5xx rate)
|
||||
5. **Pod/Replica Count** — Shows HPA scaling activity over time
|
||||
|
||||
---
|
||||
|
||||
## 8. Lessons Learned
|
||||
|
||||
| # | Lesson | Action Item |
|
||||
| :--- | :--- | :--- |
|
||||
| 1 | Default memory limits were set too low without load testing | Conduct load tests before production deployment |
|
||||
| 2 | Single-replica deployment has no resilience | Always run **≥ 2 replicas** for production services |
|
||||
| 3 | Alert threshold (10 min) was too slow for a P1 outage | Reduce critical alert threshold to **2 minutes** |
|
||||
| 4 | In-memory data stores are dangerous without bounds | Use external databases or enforce collection size limits |
|
||||
| 5 | No runbook existed for OOMKilled incidents | Create and publish an OOMKilled response runbook |
|
||||
|
||||
---
|
||||
|
||||
## 9. Action Items Tracker
|
||||
|
||||
| # | Action Item | Owner | Priority | Deadline | Status |
|
||||
| :--- | :--- | :--- | :--- | :--- | :--- |
|
||||
| 1 | Increase container memory limit to 512MB | DevOps | P0 | Done | ✅ |
|
||||
| 2 | Add `express.json({ limit: '1mb' })` payload limit | Backend | P1 | +1 day | ⬜ |
|
||||
| 3 | Implement HPA with memory-based scaling | DevOps | P1 | +3 days | ⬜ |
|
||||
| 4 | Set up Prometheus memory alerts (70% / 85%) | DevOps | P1 | +3 days | ⬜ |
|
||||
| 5 | Add OOMKilled alert rule (immediate paging) | DevOps | P0 | +1 day | ⬜ |
|
||||
| 6 | Run minimum 2 replicas in production | DevOps | P1 | +1 day | ⬜ |
|
||||
| 7 | Replace in-memory store with Redis/DB | Backend | P2 | +1 week | ⬜ |
|
||||
| 8 | Conduct load testing with realistic traffic | QA | P2 | +2 weeks | ⬜ |
|
||||
| 9 | Create OOMKilled incident runbook | DevOps | P2 | +1 week | ⬜ |
|
||||
| 10 | Build Grafana dashboard for container metrics | DevOps | P2 | +1 week | ⬜ |
|
||||
|
||||
---
|
||||
|
||||
> **Report prepared by:** DevOps Team
|
||||
> **Review date:** July 26, 2026
|
||||
> **Next review:** August 2, 2026 (verify all P0/P1 items completed)
|
||||
المرجع في مشكلة جديدة
حظر مستخدم