- Created a GitHub Actions workflow for CI/CD to deploy to Ghyamah, including testing, building, and pushing Docker images. - Added architecture design document for handling 15,000 requests per second, detailing system components, capacity planning, and cold start strategies. - Introduced a Python-based uptime/latency/SSL monitor with a static dashboard, utilizing standard libraries only. - Included Dockerfile and entrypoint script for the monitoring application, ensuring it runs as a non-root user and handles process management. - Added a .dockerignore file to exclude unnecessary files from the Docker build context. - Created an HTML dashboard for visualizing monitoring metrics, including uptime, latency, and SSL certificate status.
271 أسطر
7.2 KiB
Plaintext
271 أسطر
7.2 KiB
Plaintext
# Incident Postmortem: Core Backend API Service Interruption
|
||
|
||
> **Incident Date:** July 27, 2026
|
||
> **Affected Service:** Core .NET Backend API (Ghaymah Platform)
|
||
> **Severity:** SEV-1 (Critical)
|
||
> **Status:** Resolved
|
||
|
||
---
|
||
|
||
# Table of Contents
|
||
|
||
- [Incident Overview](#incident-overview)
|
||
- [Incident Metadata](#incident-metadata)
|
||
- [Executive Summary](#executive-summary)
|
||
- [Incident Timeline](#incident-timeline)
|
||
- [Root Cause Analysis](#root-cause-analysis)
|
||
- [Proposed Auto-Scaling Architecture](#proposed-auto-scaling-architecture)
|
||
- [Horizontal Pod Autoscaler Configuration](#horizontal-pod-autoscaler-configuration)
|
||
- [Observability & Early Detection](#observability--early-detection)
|
||
- [Action Items](#action-items)
|
||
- [Lessons Learned](#lessons-learned)
|
||
|
||
---
|
||
|
||
# Incident Overview
|
||
|
||
On **July 27, 2026**, the **Core .NET Backend API** experienced a complete service outage lasting **45 minutes**.
|
||
|
||
The outage was caused by repeated **OOMKilled (Exit Code 137)** events after the application exhausted its available memory during a sudden traffic spike.
|
||
|
||
The platform entered a continuous **CrashLoop** state, resulting in **100% API failure** until memory resources were increased and the containers were restarted.
|
||
|
||
---
|
||
|
||
# Incident Metadata
|
||
|
||
| Category | Details |
|
||
|-----------|---------|
|
||
| **Affected Service** | Core .NET Backend API |
|
||
| **Platform** | Ghaymah |
|
||
| **Date** | 2026-07-27 |
|
||
| **Downtime** | 45 Minutes |
|
||
| **Time** | 14:00 – 14:45 EEST |
|
||
| **Severity** | SEV-1 (Critical) |
|
||
| **Customer Impact** | 100% API transaction failures |
|
||
| **Observed Errors** | 502 Bad Gateway |
|
||
|
||
---
|
||
|
||
# Executive Summary
|
||
|
||
A sudden **400% increase in traffic** caused the backend API to consume memory rapidly.
|
||
|
||
The application maintained an **unbounded in-memory cache**, storing increasingly large datasets without expiration or eviction.
|
||
|
||
Once the container reached its configured memory limit, the Linux kernel terminated the process (**OOMKilled - Exit Code 137**) to protect node stability.
|
||
|
||
The orchestration platform continuously restarted the containers, causing a **CrashLoopBackOff** cycle and preventing the service from recovering automatically.
|
||
|
||
Service was restored after:
|
||
|
||
- Increasing the container memory limit
|
||
- Restarting the backend pods
|
||
- Verifying successful application startup
|
||
|
||
Future mitigation requires:
|
||
|
||
- Memory-based auto-scaling
|
||
- Cache eviction policies
|
||
- API pagination
|
||
- Improved monitoring and alerting
|
||
|
||
---
|
||
|
||
# Incident Timeline
|
||
|
||
| Time | Event |
|
||
|------|-------|
|
||
| **13:50** | Traffic increased by approximately **400%** due to an unexpected external campaign. |
|
||
| **13:58** | Container memory utilization exceeded **95%** of allocated memory. |
|
||
| **14:00** | First container terminated with **OOMKilled (Exit Code 137)**. Service degradation begins. |
|
||
| **14:05 – 14:30** | Containers repeatedly restarted by the orchestrator, resulting in a CrashLoop and complete outage. |
|
||
| **14:30** | On-call infrastructure engineer identified repeated OOMKilled events from platform logs. |
|
||
| **14:35** | Memory limit increased from **512Mi** to **2Gi** and backend pods restarted manually. |
|
||
| **14:45** | Service stabilized and API responses returned **HTTP 200 OK**. Incident closed. |
|
||
|
||
---
|
||
|
||
# Root Cause Analysis
|
||
|
||
## Direct Cause
|
||
|
||
The Linux kernel terminated the backend process because the application attempted to allocate more memory than the container's configured memory limit.
|
||
|
||
```
|
||
Exit Code: 137
|
||
Reason: OOMKilled
|
||
```
|
||
|
||
---
|
||
|
||
## Underlying Causes
|
||
|
||
### Unbounded In-Memory Cache
|
||
|
||
The application stored data inside a local in-memory dictionary that had:
|
||
|
||
- No Time-To-Live (TTL)
|
||
- No maximum cache size
|
||
- No eviction policy
|
||
|
||
Large requests continuously expanded the cache until the process exhausted available memory.
|
||
|
||
---
|
||
|
||
### Missing API Pagination
|
||
|
||
Several endpoints returned very large datasets.
|
||
|
||
Without pagination:
|
||
|
||
- Large payloads were cached
|
||
- Memory usage increased rapidly
|
||
- Garbage collection became inefficient
|
||
|
||
---
|
||
|
||
### Lack of Horizontal Auto-Scaling
|
||
|
||
The deployment relied on static memory limits and a fixed number of replicas.
|
||
|
||
As traffic increased:
|
||
|
||
- No new replicas were created
|
||
- Existing containers absorbed all incoming traffic
|
||
- Memory utilization reached critical levels
|
||
|
||
---
|
||
|
||
# Proposed Auto-Scaling Architecture
|
||
|
||
To prevent similar incidents, backend workloads should scale automatically based on memory utilization.
|
||
|
||
| Parameter | Recommended Value |
|
||
|-----------|-------------------|
|
||
| **Scaling Metric** | Average Memory Utilization |
|
||
| **Target Utilization** | 75% |
|
||
| **Minimum Replicas** | 3 |
|
||
| **Maximum Replicas** | 12 |
|
||
| **Scale-Up Policy** | Add up to 4 replicas immediately |
|
||
| **Scale-Down Policy** | Wait 5 minutes after memory falls below 40% |
|
||
|
||
---
|
||
|
||
# Horizontal Pod Autoscaler Configuration
|
||
|
||
```yaml
|
||
apiVersion: autoscaling/v2
|
||
kind: HorizontalPodAutoscaler
|
||
metadata:
|
||
name: backend-api-scaler
|
||
|
||
spec:
|
||
minReplicas: 3
|
||
maxReplicas: 12
|
||
|
||
metrics:
|
||
- type: Resource
|
||
resource:
|
||
name: memory
|
||
target:
|
||
type: Utilization
|
||
averageUtilization: 75
|
||
```
|
||
|
||
---
|
||
|
||
# Observability & Early Detection
|
||
|
||
To move from reactive troubleshooting to proactive monitoring, the following observability improvements should be implemented.
|
||
|
||
## Alerting Rules
|
||
|
||
### Memory Utilization Alert
|
||
|
||
Trigger notification when:
|
||
|
||
- Memory utilization exceeds **70%**
|
||
- Sustained for **2 consecutive minutes**
|
||
|
||
Notification targets:
|
||
|
||
- Email
|
||
- Slack
|
||
- Webhook
|
||
|
||
---
|
||
|
||
### CrashLoop Detection
|
||
|
||
Create a high-priority incident whenever:
|
||
|
||
- Container restart count exceeds **3**
|
||
- Within a **10-minute** window
|
||
|
||
---
|
||
|
||
## Dashboard Metrics
|
||
|
||
Create a dedicated monitoring dashboard showing:
|
||
|
||
- Container memory usage
|
||
- Container memory limits
|
||
- Application heap size
|
||
- Garbage Collection duration
|
||
- API request rate
|
||
- Current replica count
|
||
- Container restart count
|
||
- CPU utilization
|
||
- Response latency
|
||
- Error rate (4xx / 5xx)
|
||
|
||
---
|
||
|
||
# Action Items
|
||
|
||
| Task | Owner | Priority | Status |
|
||
|------|-------|----------|--------|
|
||
| Implement cache eviction (LRU + TTL) | Development Team | Critical | In Progress |
|
||
| Enforce API pagination | Development Team | Critical | In Progress |
|
||
| Deploy memory-based Horizontal Pod Autoscaler | DevOps Team | High | Not Started |
|
||
| Create synthetic load tests (5× traffic) | QA Team | Medium | Not Started |
|
||
| Update on-call operational runbooks | Operations Team | Low | Completed |
|
||
|
||
---
|
||
|
||
# Lessons Learned
|
||
|
||
The incident highlighted several architectural improvements required to increase platform resilience.
|
||
|
||
## Infrastructure
|
||
|
||
- Configure Horizontal Pod Autoscaler (HPA)
|
||
- Define resource requests and limits carefully
|
||
- Monitor memory consumption continuously
|
||
|
||
## Application
|
||
|
||
- Implement cache eviction (LRU)
|
||
- Apply cache expiration (TTL)
|
||
- Enforce pagination on all large API endpoints
|
||
- Optimize memory allocation patterns
|
||
|
||
## Operations
|
||
|
||
- Improve proactive alerting
|
||
- Expand observability dashboards
|
||
- Regularly execute load and stress testing
|
||
- Update incident response runbooks
|
||
|
||
---
|
||
|
||
# Resolution Summary
|
||
|
||
| Item | Result |
|
||
|------|--------|
|
||
| **Root Cause** | Unbounded in-memory cache caused container OOM |
|
||
| **Immediate Fix** | Increased memory limit (512Mi → 2Gi) and restarted pods |
|
||
| **Long-Term Fixes** | HPA, cache eviction, pagination, monitoring improvements |
|
||
| **Incident Status** | ✅ Resolved |
|
||
``` |