Q2
هذا الالتزام موجود في:
173
q2-postmortem/postmortem-report.md
Normal file
173
q2-postmortem/postmortem-report.md
Normal file
@@ -0,0 +1,173 @@
|
||||
# Q2 - Postmortem Report
|
||||
|
||||
## Incident Summary
|
||||
|
||||
**Incident:** Repeated OOMKilled causing service outage
|
||||
|
||||
**Date:** (Exam Scenario)
|
||||
|
||||
**Duration:** 45 minutes
|
||||
|
||||
**Severity:** High (SEV-1)
|
||||
|
||||
**Impact:**
|
||||
- API became unavailable.
|
||||
- Users received HTTP 5xx errors.
|
||||
- Service uptime was affected.
|
||||
- Customer requests could not be processed during the outage.
|
||||
|
||||
---
|
||||
|
||||
# Timeline
|
||||
|
||||
| Time | Event |
|
||||
|------|-------|
|
||||
| 10:00 | Increased memory usage observed. |
|
||||
| 10:05 | First container terminated with OOMKilled. |
|
||||
| 10:06 | Kubernetes restarted the container automatically. |
|
||||
| 10:10 | Memory usage increased again causing another OOMKilled. |
|
||||
| 10:15 | Continuous restart loop (CrashLoopBackOff). |
|
||||
| 10:20 | Monitoring system generated alerts. |
|
||||
| 10:30 | SRE team started investigation. |
|
||||
| 10:40 | Memory limit increased and memory leak identified. |
|
||||
| 10:45 | New container deployed successfully and service restored. |
|
||||
|
||||
---
|
||||
|
||||
# Root Cause Analysis
|
||||
|
||||
The application suffered from excessive memory consumption caused by a memory leak under increased traffic.
|
||||
|
||||
The container memory limit was too low to handle the workload.
|
||||
|
||||
When memory usage exceeded the configured limit, Kubernetes terminated the container with an **OOMKilled** event.
|
||||
|
||||
Automatic restarts repeatedly created a CrashLoopBackOff situation, extending the outage.
|
||||
|
||||
---
|
||||
|
||||
# Contributing Factors
|
||||
|
||||
- No Horizontal Pod Autoscaler (HPA).
|
||||
- Memory limits configured too aggressively.
|
||||
- No early memory usage alerts.
|
||||
- Memory leak was not detected during testing.
|
||||
- Lack of load testing before deployment.
|
||||
|
||||
---
|
||||
|
||||
# Immediate Actions Taken
|
||||
|
||||
- Increased memory limits.
|
||||
- Restarted affected containers.
|
||||
- Fixed memory leak.
|
||||
- Verified application health.
|
||||
- Monitored memory utilization after recovery.
|
||||
|
||||
---
|
||||
|
||||
# Preventive Recommendations
|
||||
|
||||
- Enable Horizontal Pod Autoscaler (HPA).
|
||||
- Configure Vertical Pod Autoscaler (VPA) if appropriate.
|
||||
- Define proper memory requests and limits.
|
||||
- Add Prometheus memory monitoring.
|
||||
- Configure Alertmanager notifications.
|
||||
- Perform stress testing before production deployment.
|
||||
- Review application memory usage regularly.
|
||||
|
||||
---
|
||||
|
||||
# Auto-Scaling Policy
|
||||
|
||||
To prevent similar incidents on Ghaymah:
|
||||
|
||||
### Horizontal Pod Autoscaler
|
||||
|
||||
Minimum Replicas: 2
|
||||
|
||||
Maximum Replicas: 10
|
||||
|
||||
Scale Out Conditions
|
||||
|
||||
- Memory usage > 70%
|
||||
- CPU usage > 70%
|
||||
- Request rate exceeds expected capacity
|
||||
|
||||
Scale In Conditions
|
||||
|
||||
- Memory usage < 40%
|
||||
- CPU usage < 40%
|
||||
- Stable traffic for at least 10 minutes
|
||||
|
||||
Cooldown Period
|
||||
|
||||
- Scale Out: 60 seconds
|
||||
|
||||
- Scale In: 300 seconds
|
||||
|
||||
Benefits
|
||||
|
||||
- Prevents memory exhaustion.
|
||||
- Distributes incoming traffic.
|
||||
- Improves availability.
|
||||
- Reduces restart frequency.
|
||||
|
||||
---
|
||||
|
||||
# Early Detection Using Ghaymah Monitoring
|
||||
|
||||
The incident could have been detected earlier using monitoring tools.
|
||||
|
||||
## Metrics
|
||||
|
||||
- Container Memory Usage
|
||||
- Memory Limit Percentage
|
||||
- Container Restart Count
|
||||
- OOMKilled Events
|
||||
- CPU Utilization
|
||||
- Request Rate
|
||||
- Response Time
|
||||
|
||||
## Alerts
|
||||
|
||||
Critical Alert
|
||||
|
||||
- Memory Usage > 85%
|
||||
- More than 3 restarts within 5 minutes
|
||||
- OOMKilled event detected
|
||||
|
||||
Warning Alert
|
||||
|
||||
- Memory Usage > 70%
|
||||
- Rapid increase in memory consumption
|
||||
|
||||
## Dashboards
|
||||
|
||||
Recommended dashboard should display:
|
||||
|
||||
- Memory Usage
|
||||
- CPU Usage
|
||||
- Pod Status
|
||||
- Restart Count
|
||||
- Response Time
|
||||
- Error Rate
|
||||
- Availability
|
||||
|
||||
---
|
||||
|
||||
# Lessons Learned
|
||||
|
||||
- Memory monitoring must be proactive.
|
||||
- Autoscaling should be enabled in production.
|
||||
- Load testing should validate memory consumption.
|
||||
- Alerting should notify engineers before service failure.
|
||||
- Capacity planning should be reviewed regularly.
|
||||
|
||||
---
|
||||
|
||||
# Status
|
||||
|
||||
✅ Incident resolved
|
||||
|
||||
No recurring issues observed after implementing corrective actions.
|
||||
المرجع في مشكلة جديدة
حظر مستخدم