الملفات
ghaymah-exam-hassan-yehia-sre/docs/postmortem.md

196 أسطر
4.3 KiB
Markdown
خام اللوم التاريخ

هذا الملف يحتوي على أحرف Unicode غامضة

هذا الملف يحتوي على أحرف Unicode قد تُخلط مع أحرف أخرى. إذا كنت تعتقد أن هذا مقصود، يمكنك تجاهل هذا التحذير بأمان. استخدم زر الهروب للكشف عنها.

# Incident Postmortem: Repeated OOMKilled Causing Application Downtime
## Incident Summary
**Incident Title:** Repeated OOMKilled Causing Application Downtime
**Date:** July 28, 2026
**Duration:** 45 Minutes
**Severity:** High
**Affected Service:** Ghaymah SRE API
### Impact
- The application was unavailable for approximately 45 minutes.
- Users experienced failed requests and service interruptions.
- Multiple container restarts occurred due to repeated OOMKilled events.
---
# Timeline
| Time | Event |
|------|-------|
| 10:00 | Application deployed successfully |
| 10:08 | Memory usage started increasing rapidly |
| 10:12 | First OOMKilled event occurred |
| 10:13 - 10:40 | Container repeatedly restarted due to memory exhaustion |
| 10:20 | Monitoring system generated alerts |
| 10:25 | SRE team started investigation |
| 10:35 | Memory leak identified |
| 10:42 | Memory limit increased and deployment restarted |
| 10:45 | Service fully recovered |
---
# Root Cause Analysis
The application experienced a memory leak that continuously increased memory consumption.
Once the container exceeded its configured memory limit, Kubernetes terminated it with an **OOMKilled** event.
Because the underlying issue was not resolved immediately, the container repeatedly restarted, causing approximately 45 minutes of service downtime.
---
# Recommendations
## Immediate Actions
- Increase the container memory limit.
- Restart the deployment after applying the fix.
- Monitor memory usage during recovery.
## Long-Term Improvements
- Configure Horizontal Pod Autoscaler (HPA).
- Set appropriate CPU and memory requests and limits.
- Enable memory utilization alerts.
- Monitor container restart counts.
- Perform load testing before production deployments.
- Regularly profile the application to detect memory leaks.
---
# Auto-Scaling Policy
## Objective
Automatically scale the application based on workload to maintain availability and reduce the risk of resource exhaustion.
| Configuration | Value |
|--------------|-------|
| Minimum Replicas | 2 |
| Maximum Replicas | 10 |
| CPU Target | 70% |
| Memory Target | 75% |
| Scale Up | Add 12 replicas after 2 minutes above threshold |
| Scale Down | Remove 1 replica after 10 minutes of low utilization |
## Scale-Up Rules
Trigger scaling when:
- CPU usage exceeds 70%.
- Memory usage exceeds 75%.
- Traffic increases significantly.
Action:
- Create additional application replicas.
- Distribute requests using the load balancer.
## Scale-Down Rules
Trigger scaling when:
- CPU remains below 30%.
- Memory remains below 40%.
- Low traffic continues for at least 10 minutes.
Action:
- Gradually remove unused replicas while keeping at least two running.
---
# Early Detection Using Monitoring
## Monitoring Metrics
### Memory Usage
Monitor:
- Container memory usage
- Memory utilization percentage
- Available memory
**Alert Rule**
```
Memory usage > 80% for 5 minutes
```
### OOMKilled Events
Monitor:
- Number of OOMKilled events
- Container restart count
**Alert Rule**
```
Restart count > 3 within 10 minutes
```
### Container Health
Monitor:
- /health endpoint
- Readiness probe
- Liveness probe
- Pod status
Alert if:
- Health endpoint returns a non-200 status.
- Pods enter CrashLoopBackOff or Pending state.
### Application Performance
Monitor:
- HTTP response time
- Request rate
- Error rate (5xx)
- Active requests
---
# Monitoring Dashboard
The dashboard should display:
- Application Status
- CPU Usage
- Memory Usage
- Response Time
- Total Requests
- Container Restarts
- OOMKilled Events
- Running Replicas
---
# Alert Workflow
1. Prometheus collects application and container metrics.
2. Alertmanager evaluates alert rules.
3. Notifications are sent to the SRE team (Email, Slack, or Microsoft Teams).
4. Engineers investigate the issue using Grafana dashboards and application logs.
5. If auto-scaling is enabled, additional replicas are created automatically while the issue is being investigated.
---
# Conclusion
The outage was caused by excessive memory consumption that resulted in repeated OOMKilled events and continuous container restarts.
Implementing proactive monitoring, alerting, resource limits, and auto-scaling policies will significantly reduce the likelihood and impact of similar incidents in the future.