Q2
هذا الالتزام موجود في:
173
q2-postmortem/postmortem-report.md
Normal file
173
q2-postmortem/postmortem-report.md
Normal file
@@ -0,0 +1,173 @@
|
|||||||
|
# Q2 - Postmortem Report
|
||||||
|
|
||||||
|
## Incident Summary
|
||||||
|
|
||||||
|
**Incident:** Repeated OOMKilled causing service outage
|
||||||
|
|
||||||
|
**Date:** (Exam Scenario)
|
||||||
|
|
||||||
|
**Duration:** 45 minutes
|
||||||
|
|
||||||
|
**Severity:** High (SEV-1)
|
||||||
|
|
||||||
|
**Impact:**
|
||||||
|
- API became unavailable.
|
||||||
|
- Users received HTTP 5xx errors.
|
||||||
|
- Service uptime was affected.
|
||||||
|
- Customer requests could not be processed during the outage.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Timeline
|
||||||
|
|
||||||
|
| Time | Event |
|
||||||
|
|------|-------|
|
||||||
|
| 10:00 | Increased memory usage observed. |
|
||||||
|
| 10:05 | First container terminated with OOMKilled. |
|
||||||
|
| 10:06 | Kubernetes restarted the container automatically. |
|
||||||
|
| 10:10 | Memory usage increased again causing another OOMKilled. |
|
||||||
|
| 10:15 | Continuous restart loop (CrashLoopBackOff). |
|
||||||
|
| 10:20 | Monitoring system generated alerts. |
|
||||||
|
| 10:30 | SRE team started investigation. |
|
||||||
|
| 10:40 | Memory limit increased and memory leak identified. |
|
||||||
|
| 10:45 | New container deployed successfully and service restored. |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Root Cause Analysis
|
||||||
|
|
||||||
|
The application suffered from excessive memory consumption caused by a memory leak under increased traffic.
|
||||||
|
|
||||||
|
The container memory limit was too low to handle the workload.
|
||||||
|
|
||||||
|
When memory usage exceeded the configured limit, Kubernetes terminated the container with an **OOMKilled** event.
|
||||||
|
|
||||||
|
Automatic restarts repeatedly created a CrashLoopBackOff situation, extending the outage.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Contributing Factors
|
||||||
|
|
||||||
|
- No Horizontal Pod Autoscaler (HPA).
|
||||||
|
- Memory limits configured too aggressively.
|
||||||
|
- No early memory usage alerts.
|
||||||
|
- Memory leak was not detected during testing.
|
||||||
|
- Lack of load testing before deployment.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Immediate Actions Taken
|
||||||
|
|
||||||
|
- Increased memory limits.
|
||||||
|
- Restarted affected containers.
|
||||||
|
- Fixed memory leak.
|
||||||
|
- Verified application health.
|
||||||
|
- Monitored memory utilization after recovery.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Preventive Recommendations
|
||||||
|
|
||||||
|
- Enable Horizontal Pod Autoscaler (HPA).
|
||||||
|
- Configure Vertical Pod Autoscaler (VPA) if appropriate.
|
||||||
|
- Define proper memory requests and limits.
|
||||||
|
- Add Prometheus memory monitoring.
|
||||||
|
- Configure Alertmanager notifications.
|
||||||
|
- Perform stress testing before production deployment.
|
||||||
|
- Review application memory usage regularly.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Auto-Scaling Policy
|
||||||
|
|
||||||
|
To prevent similar incidents on Ghaymah:
|
||||||
|
|
||||||
|
### Horizontal Pod Autoscaler
|
||||||
|
|
||||||
|
Minimum Replicas: 2
|
||||||
|
|
||||||
|
Maximum Replicas: 10
|
||||||
|
|
||||||
|
Scale Out Conditions
|
||||||
|
|
||||||
|
- Memory usage > 70%
|
||||||
|
- CPU usage > 70%
|
||||||
|
- Request rate exceeds expected capacity
|
||||||
|
|
||||||
|
Scale In Conditions
|
||||||
|
|
||||||
|
- Memory usage < 40%
|
||||||
|
- CPU usage < 40%
|
||||||
|
- Stable traffic for at least 10 minutes
|
||||||
|
|
||||||
|
Cooldown Period
|
||||||
|
|
||||||
|
- Scale Out: 60 seconds
|
||||||
|
|
||||||
|
- Scale In: 300 seconds
|
||||||
|
|
||||||
|
Benefits
|
||||||
|
|
||||||
|
- Prevents memory exhaustion.
|
||||||
|
- Distributes incoming traffic.
|
||||||
|
- Improves availability.
|
||||||
|
- Reduces restart frequency.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Early Detection Using Ghaymah Monitoring
|
||||||
|
|
||||||
|
The incident could have been detected earlier using monitoring tools.
|
||||||
|
|
||||||
|
## Metrics
|
||||||
|
|
||||||
|
- Container Memory Usage
|
||||||
|
- Memory Limit Percentage
|
||||||
|
- Container Restart Count
|
||||||
|
- OOMKilled Events
|
||||||
|
- CPU Utilization
|
||||||
|
- Request Rate
|
||||||
|
- Response Time
|
||||||
|
|
||||||
|
## Alerts
|
||||||
|
|
||||||
|
Critical Alert
|
||||||
|
|
||||||
|
- Memory Usage > 85%
|
||||||
|
- More than 3 restarts within 5 minutes
|
||||||
|
- OOMKilled event detected
|
||||||
|
|
||||||
|
Warning Alert
|
||||||
|
|
||||||
|
- Memory Usage > 70%
|
||||||
|
- Rapid increase in memory consumption
|
||||||
|
|
||||||
|
## Dashboards
|
||||||
|
|
||||||
|
Recommended dashboard should display:
|
||||||
|
|
||||||
|
- Memory Usage
|
||||||
|
- CPU Usage
|
||||||
|
- Pod Status
|
||||||
|
- Restart Count
|
||||||
|
- Response Time
|
||||||
|
- Error Rate
|
||||||
|
- Availability
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Lessons Learned
|
||||||
|
|
||||||
|
- Memory monitoring must be proactive.
|
||||||
|
- Autoscaling should be enabled in production.
|
||||||
|
- Load testing should validate memory consumption.
|
||||||
|
- Alerting should notify engineers before service failure.
|
||||||
|
- Capacity planning should be reviewed regularly.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Status
|
||||||
|
|
||||||
|
✅ Incident resolved
|
||||||
|
|
||||||
|
No recurring issues observed after implementing corrective actions.
|
||||||
المرجع في مشكلة جديدة
حظر مستخدم