Add OOMKilled incident postmortem and auto-scaling policy
هذا الالتزام موجود في:
195
docs/postmortem.md
Normal file
195
docs/postmortem.md
Normal file
@@ -0,0 +1,195 @@
|
||||
# Incident Postmortem: Repeated OOMKilled Causing Application Downtime
|
||||
|
||||
## Incident Summary
|
||||
|
||||
**Incident Title:** Repeated OOMKilled Causing Application Downtime
|
||||
|
||||
**Date:** July 28, 2026
|
||||
|
||||
**Duration:** 45 Minutes
|
||||
|
||||
**Severity:** High
|
||||
|
||||
**Affected Service:** Ghaymah SRE API
|
||||
|
||||
### Impact
|
||||
|
||||
- The application was unavailable for approximately 45 minutes.
|
||||
- Users experienced failed requests and service interruptions.
|
||||
- Multiple container restarts occurred due to repeated OOMKilled events.
|
||||
|
||||
---
|
||||
|
||||
# Timeline
|
||||
|
||||
| Time | Event |
|
||||
|------|-------|
|
||||
| 10:00 | Application deployed successfully |
|
||||
| 10:08 | Memory usage started increasing rapidly |
|
||||
| 10:12 | First OOMKilled event occurred |
|
||||
| 10:13 - 10:40 | Container repeatedly restarted due to memory exhaustion |
|
||||
| 10:20 | Monitoring system generated alerts |
|
||||
| 10:25 | SRE team started investigation |
|
||||
| 10:35 | Memory leak identified |
|
||||
| 10:42 | Memory limit increased and deployment restarted |
|
||||
| 10:45 | Service fully recovered |
|
||||
|
||||
---
|
||||
|
||||
# Root Cause Analysis
|
||||
|
||||
The application experienced a memory leak that continuously increased memory consumption.
|
||||
|
||||
Once the container exceeded its configured memory limit, Kubernetes terminated it with an **OOMKilled** event.
|
||||
|
||||
Because the underlying issue was not resolved immediately, the container repeatedly restarted, causing approximately 45 minutes of service downtime.
|
||||
|
||||
---
|
||||
|
||||
# Recommendations
|
||||
|
||||
## Immediate Actions
|
||||
|
||||
- Increase the container memory limit.
|
||||
- Restart the deployment after applying the fix.
|
||||
- Monitor memory usage during recovery.
|
||||
|
||||
## Long-Term Improvements
|
||||
|
||||
- Configure Horizontal Pod Autoscaler (HPA).
|
||||
- Set appropriate CPU and memory requests and limits.
|
||||
- Enable memory utilization alerts.
|
||||
- Monitor container restart counts.
|
||||
- Perform load testing before production deployments.
|
||||
- Regularly profile the application to detect memory leaks.
|
||||
|
||||
---
|
||||
|
||||
# Auto-Scaling Policy
|
||||
|
||||
## Objective
|
||||
|
||||
Automatically scale the application based on workload to maintain availability and reduce the risk of resource exhaustion.
|
||||
|
||||
| Configuration | Value |
|
||||
|--------------|-------|
|
||||
| Minimum Replicas | 2 |
|
||||
| Maximum Replicas | 10 |
|
||||
| CPU Target | 70% |
|
||||
| Memory Target | 75% |
|
||||
| Scale Up | Add 1–2 replicas after 2 minutes above threshold |
|
||||
| Scale Down | Remove 1 replica after 10 minutes of low utilization |
|
||||
|
||||
## Scale-Up Rules
|
||||
|
||||
Trigger scaling when:
|
||||
|
||||
- CPU usage exceeds 70%.
|
||||
- Memory usage exceeds 75%.
|
||||
- Traffic increases significantly.
|
||||
|
||||
Action:
|
||||
|
||||
- Create additional application replicas.
|
||||
- Distribute requests using the load balancer.
|
||||
|
||||
## Scale-Down Rules
|
||||
|
||||
Trigger scaling when:
|
||||
|
||||
- CPU remains below 30%.
|
||||
- Memory remains below 40%.
|
||||
- Low traffic continues for at least 10 minutes.
|
||||
|
||||
Action:
|
||||
|
||||
- Gradually remove unused replicas while keeping at least two running.
|
||||
|
||||
---
|
||||
|
||||
# Early Detection Using Monitoring
|
||||
|
||||
## Monitoring Metrics
|
||||
|
||||
### Memory Usage
|
||||
|
||||
Monitor:
|
||||
|
||||
- Container memory usage
|
||||
- Memory utilization percentage
|
||||
- Available memory
|
||||
|
||||
**Alert Rule**
|
||||
|
||||
```
|
||||
Memory usage > 80% for 5 minutes
|
||||
```
|
||||
|
||||
### OOMKilled Events
|
||||
|
||||
Monitor:
|
||||
|
||||
- Number of OOMKilled events
|
||||
- Container restart count
|
||||
|
||||
**Alert Rule**
|
||||
|
||||
```
|
||||
Restart count > 3 within 10 minutes
|
||||
```
|
||||
|
||||
### Container Health
|
||||
|
||||
Monitor:
|
||||
|
||||
- /health endpoint
|
||||
- Readiness probe
|
||||
- Liveness probe
|
||||
- Pod status
|
||||
|
||||
Alert if:
|
||||
|
||||
- Health endpoint returns a non-200 status.
|
||||
- Pods enter CrashLoopBackOff or Pending state.
|
||||
|
||||
### Application Performance
|
||||
|
||||
Monitor:
|
||||
|
||||
- HTTP response time
|
||||
- Request rate
|
||||
- Error rate (5xx)
|
||||
- Active requests
|
||||
|
||||
---
|
||||
|
||||
# Monitoring Dashboard
|
||||
|
||||
The dashboard should display:
|
||||
|
||||
- Application Status
|
||||
- CPU Usage
|
||||
- Memory Usage
|
||||
- Response Time
|
||||
- Total Requests
|
||||
- Container Restarts
|
||||
- OOMKilled Events
|
||||
- Running Replicas
|
||||
|
||||
---
|
||||
|
||||
# Alert Workflow
|
||||
|
||||
1. Prometheus collects application and container metrics.
|
||||
2. Alertmanager evaluates alert rules.
|
||||
3. Notifications are sent to the SRE team (Email, Slack, or Microsoft Teams).
|
||||
4. Engineers investigate the issue using Grafana dashboards and application logs.
|
||||
5. If auto-scaling is enabled, additional replicas are created automatically while the issue is being investigated.
|
||||
|
||||
---
|
||||
|
||||
# Conclusion
|
||||
|
||||
The outage was caused by excessive memory consumption that resulted in repeated OOMKilled events and continuous container restarts.
|
||||
|
||||
Implementing proactive monitoring, alerting, resource limits, and auto-scaling policies will significantly reduce the likelihood and impact of similar incidents in the future.
|
||||
المرجع في مشكلة جديدة
حظر مستخدم