3.6 KiB
Q2 - Postmortem Report
Incident Summary
Incident: Repeated OOMKilled causing service outage
Date: (Exam Scenario)
Duration: 45 minutes
Severity: High (SEV-1)
Impact:
- API became unavailable.
- Users received HTTP 5xx errors.
- Service uptime was affected.
- Customer requests could not be processed during the outage.
Timeline
| Time | Event |
|---|---|
| 10:00 | Increased memory usage observed. |
| 10:05 | First container terminated with OOMKilled. |
| 10:06 | Kubernetes restarted the container automatically. |
| 10:10 | Memory usage increased again causing another OOMKilled. |
| 10:15 | Continuous restart loop (CrashLoopBackOff). |
| 10:20 | Monitoring system generated alerts. |
| 10:30 | SRE team started investigation. |
| 10:40 | Memory limit increased and memory leak identified. |
| 10:45 | New container deployed successfully and service restored. |
Root Cause Analysis
The application suffered from excessive memory consumption caused by a memory leak under increased traffic.
The container memory limit was too low to handle the workload.
When memory usage exceeded the configured limit, Kubernetes terminated the container with an OOMKilled event.
Automatic restarts repeatedly created a CrashLoopBackOff situation, extending the outage.
Contributing Factors
- No Horizontal Pod Autoscaler (HPA).
- Memory limits configured too aggressively.
- No early memory usage alerts.
- Memory leak was not detected during testing.
- Lack of load testing before deployment.
Immediate Actions Taken
- Increased memory limits.
- Restarted affected containers.
- Fixed memory leak.
- Verified application health.
- Monitored memory utilization after recovery.
Preventive Recommendations
- Enable Horizontal Pod Autoscaler (HPA).
- Configure Vertical Pod Autoscaler (VPA) if appropriate.
- Define proper memory requests and limits.
- Add Prometheus memory monitoring.
- Configure Alertmanager notifications.
- Perform stress testing before production deployment.
- Review application memory usage regularly.
Auto-Scaling Policy
To prevent similar incidents on Ghaymah:
Horizontal Pod Autoscaler
Minimum Replicas: 2
Maximum Replicas: 10
Scale Out Conditions
- Memory usage > 70%
- CPU usage > 70%
- Request rate exceeds expected capacity
Scale In Conditions
- Memory usage < 40%
- CPU usage < 40%
- Stable traffic for at least 10 minutes
Cooldown Period
-
Scale Out: 60 seconds
-
Scale In: 300 seconds
Benefits
- Prevents memory exhaustion.
- Distributes incoming traffic.
- Improves availability.
- Reduces restart frequency.
Early Detection Using Ghaymah Monitoring
The incident could have been detected earlier using monitoring tools.
Metrics
- Container Memory Usage
- Memory Limit Percentage
- Container Restart Count
- OOMKilled Events
- CPU Utilization
- Request Rate
- Response Time
Alerts
Critical Alert
- Memory Usage > 85%
- More than 3 restarts within 5 minutes
- OOMKilled event detected
Warning Alert
- Memory Usage > 70%
- Rapid increase in memory consumption
Dashboards
Recommended dashboard should display:
- Memory Usage
- CPU Usage
- Pod Status
- Restart Count
- Response Time
- Error Rate
- Availability
Lessons Learned
- Memory monitoring must be proactive.
- Autoscaling should be enabled in production.
- Load testing should validate memory consumption.
- Alerting should notify engineers before service failure.
- Capacity planning should be reviewed regularly.
Status
✅ Incident resolved
No recurring issues observed after implementing corrective actions.