الملفات
Gamal0909 439a31ba93 Add CI/CD pipeline, architecture documentation, and monitoring application
- Created a GitHub Actions workflow for CI/CD to deploy to Ghyamah, including testing, building, and pushing Docker images.
- Added architecture design document for handling 15,000 requests per second, detailing system components, capacity planning, and cold start strategies.
- Introduced a Python-based uptime/latency/SSL monitor with a static dashboard, utilizing standard libraries only.
- Included Dockerfile and entrypoint script for the monitoring application, ensuring it runs as a non-root user and handles process management.
- Added a .dockerignore file to exclude unnecessary files from the Docker build context.
- Created an HTML dashboard for visualizing monitoring metrics, including uptime, latency, and SSL certificate status.
2026-07-27 23:08:39 +03:00

271 أسطر
7.2 KiB
Plaintext
خام الرابط الدائم اللوم التاريخ

هذا الملف يحتوي على أحرف Unicode غامضة

هذا الملف يحتوي على أحرف Unicode قد تُخلط مع أحرف أخرى. إذا كنت تعتقد أن هذا مقصود، يمكنك تجاهل هذا التحذير بأمان. استخدم زر الهروب للكشف عنها.

# Incident Postmortem: Core Backend API Service Interruption
> **Incident Date:** July 27, 2026
> **Affected Service:** Core .NET Backend API (Ghaymah Platform)
> **Severity:** SEV-1 (Critical)
> **Status:** Resolved
---
# Table of Contents
- [Incident Overview](#incident-overview)
- [Incident Metadata](#incident-metadata)
- [Executive Summary](#executive-summary)
- [Incident Timeline](#incident-timeline)
- [Root Cause Analysis](#root-cause-analysis)
- [Proposed Auto-Scaling Architecture](#proposed-auto-scaling-architecture)
- [Horizontal Pod Autoscaler Configuration](#horizontal-pod-autoscaler-configuration)
- [Observability & Early Detection](#observability--early-detection)
- [Action Items](#action-items)
- [Lessons Learned](#lessons-learned)
---
# Incident Overview
On **July 27, 2026**, the **Core .NET Backend API** experienced a complete service outage lasting **45 minutes**.
The outage was caused by repeated **OOMKilled (Exit Code 137)** events after the application exhausted its available memory during a sudden traffic spike.
The platform entered a continuous **CrashLoop** state, resulting in **100% API failure** until memory resources were increased and the containers were restarted.
---
# Incident Metadata
| Category | Details |
|-----------|---------|
| **Affected Service** | Core .NET Backend API |
| **Platform** | Ghaymah |
| **Date** | 2026-07-27 |
| **Downtime** | 45 Minutes |
| **Time** | 14:00 14:45 EEST |
| **Severity** | SEV-1 (Critical) |
| **Customer Impact** | 100% API transaction failures |
| **Observed Errors** | 502 Bad Gateway |
---
# Executive Summary
A sudden **400% increase in traffic** caused the backend API to consume memory rapidly.
The application maintained an **unbounded in-memory cache**, storing increasingly large datasets without expiration or eviction.
Once the container reached its configured memory limit, the Linux kernel terminated the process (**OOMKilled - Exit Code 137**) to protect node stability.
The orchestration platform continuously restarted the containers, causing a **CrashLoopBackOff** cycle and preventing the service from recovering automatically.
Service was restored after:
- Increasing the container memory limit
- Restarting the backend pods
- Verifying successful application startup
Future mitigation requires:
- Memory-based auto-scaling
- Cache eviction policies
- API pagination
- Improved monitoring and alerting
---
# Incident Timeline
| Time | Event |
|------|-------|
| **13:50** | Traffic increased by approximately **400%** due to an unexpected external campaign. |
| **13:58** | Container memory utilization exceeded **95%** of allocated memory. |
| **14:00** | First container terminated with **OOMKilled (Exit Code 137)**. Service degradation begins. |
| **14:05 14:30** | Containers repeatedly restarted by the orchestrator, resulting in a CrashLoop and complete outage. |
| **14:30** | On-call infrastructure engineer identified repeated OOMKilled events from platform logs. |
| **14:35** | Memory limit increased from **512Mi** to **2Gi** and backend pods restarted manually. |
| **14:45** | Service stabilized and API responses returned **HTTP 200 OK**. Incident closed. |
---
# Root Cause Analysis
## Direct Cause
The Linux kernel terminated the backend process because the application attempted to allocate more memory than the container's configured memory limit.
```
Exit Code: 137
Reason: OOMKilled
```
---
## Underlying Causes
### Unbounded In-Memory Cache
The application stored data inside a local in-memory dictionary that had:
- No Time-To-Live (TTL)
- No maximum cache size
- No eviction policy
Large requests continuously expanded the cache until the process exhausted available memory.
---
### Missing API Pagination
Several endpoints returned very large datasets.
Without pagination:
- Large payloads were cached
- Memory usage increased rapidly
- Garbage collection became inefficient
---
### Lack of Horizontal Auto-Scaling
The deployment relied on static memory limits and a fixed number of replicas.
As traffic increased:
- No new replicas were created
- Existing containers absorbed all incoming traffic
- Memory utilization reached critical levels
---
# Proposed Auto-Scaling Architecture
To prevent similar incidents, backend workloads should scale automatically based on memory utilization.
| Parameter | Recommended Value |
|-----------|-------------------|
| **Scaling Metric** | Average Memory Utilization |
| **Target Utilization** | 75% |
| **Minimum Replicas** | 3 |
| **Maximum Replicas** | 12 |
| **Scale-Up Policy** | Add up to 4 replicas immediately |
| **Scale-Down Policy** | Wait 5 minutes after memory falls below 40% |
---
# Horizontal Pod Autoscaler Configuration
```yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: backend-api-scaler
spec:
minReplicas: 3
maxReplicas: 12
metrics:
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 75
```
---
# Observability & Early Detection
To move from reactive troubleshooting to proactive monitoring, the following observability improvements should be implemented.
## Alerting Rules
### Memory Utilization Alert
Trigger notification when:
- Memory utilization exceeds **70%**
- Sustained for **2 consecutive minutes**
Notification targets:
- Email
- Slack
- Webhook
---
### CrashLoop Detection
Create a high-priority incident whenever:
- Container restart count exceeds **3**
- Within a **10-minute** window
---
## Dashboard Metrics
Create a dedicated monitoring dashboard showing:
- Container memory usage
- Container memory limits
- Application heap size
- Garbage Collection duration
- API request rate
- Current replica count
- Container restart count
- CPU utilization
- Response latency
- Error rate (4xx / 5xx)
---
# Action Items
| Task | Owner | Priority | Status |
|------|-------|----------|--------|
| Implement cache eviction (LRU + TTL) | Development Team | Critical | In Progress |
| Enforce API pagination | Development Team | Critical | In Progress |
| Deploy memory-based Horizontal Pod Autoscaler | DevOps Team | High | Not Started |
| Create synthetic load tests (5× traffic) | QA Team | Medium | Not Started |
| Update on-call operational runbooks | Operations Team | Low | Completed |
---
# Lessons Learned
The incident highlighted several architectural improvements required to increase platform resilience.
## Infrastructure
- Configure Horizontal Pod Autoscaler (HPA)
- Define resource requests and limits carefully
- Monitor memory consumption continuously
## Application
- Implement cache eviction (LRU)
- Apply cache expiration (TTL)
- Enforce pagination on all large API endpoints
- Optimize memory allocation patterns
## Operations
- Improve proactive alerting
- Expand observability dashboards
- Regularly execute load and stress testing
- Update incident response runbooks
---
# Resolution Summary
| Item | Result |
|------|--------|
| **Root Cause** | Unbounded in-memory cache caused container OOM |
| **Immediate Fix** | Increased memory limit (512Mi → 2Gi) and restarted pods |
| **Long-Term Fixes** | HPA, cache eviction, pagination, monitoring improvements |
| **Incident Status** | ✅ Resolved |
```