Add CI/CD pipeline, architecture documentation, and monitoring application
- Created a GitHub Actions workflow for CI/CD to deploy to Ghyamah, including testing, building, and pushing Docker images. - Added architecture design document for handling 15,000 requests per second, detailing system components, capacity planning, and cold start strategies. - Introduced a Python-based uptime/latency/SSL monitor with a static dashboard, utilizing standard libraries only. - Included Dockerfile and entrypoint script for the monitoring application, ensuring it runs as a non-root user and handles process management. - Added a .dockerignore file to exclude unnecessary files from the Docker build context. - Created an HTML dashboard for visualizing monitoring metrics, including uptime, latency, and SSL certificate status.
هذا الالتزام موجود في:
271
Q2/readme,md
Normal file
271
Q2/readme,md
Normal file
@@ -0,0 +1,271 @@
|
||||
# Incident Postmortem: Core Backend API Service Interruption
|
||||
|
||||
> **Incident Date:** July 27, 2026
|
||||
> **Affected Service:** Core .NET Backend API (Ghaymah Platform)
|
||||
> **Severity:** SEV-1 (Critical)
|
||||
> **Status:** Resolved
|
||||
|
||||
---
|
||||
|
||||
# Table of Contents
|
||||
|
||||
- [Incident Overview](#incident-overview)
|
||||
- [Incident Metadata](#incident-metadata)
|
||||
- [Executive Summary](#executive-summary)
|
||||
- [Incident Timeline](#incident-timeline)
|
||||
- [Root Cause Analysis](#root-cause-analysis)
|
||||
- [Proposed Auto-Scaling Architecture](#proposed-auto-scaling-architecture)
|
||||
- [Horizontal Pod Autoscaler Configuration](#horizontal-pod-autoscaler-configuration)
|
||||
- [Observability & Early Detection](#observability--early-detection)
|
||||
- [Action Items](#action-items)
|
||||
- [Lessons Learned](#lessons-learned)
|
||||
|
||||
---
|
||||
|
||||
# Incident Overview
|
||||
|
||||
On **July 27, 2026**, the **Core .NET Backend API** experienced a complete service outage lasting **45 minutes**.
|
||||
|
||||
The outage was caused by repeated **OOMKilled (Exit Code 137)** events after the application exhausted its available memory during a sudden traffic spike.
|
||||
|
||||
The platform entered a continuous **CrashLoop** state, resulting in **100% API failure** until memory resources were increased and the containers were restarted.
|
||||
|
||||
---
|
||||
|
||||
# Incident Metadata
|
||||
|
||||
| Category | Details |
|
||||
|-----------|---------|
|
||||
| **Affected Service** | Core .NET Backend API |
|
||||
| **Platform** | Ghaymah |
|
||||
| **Date** | 2026-07-27 |
|
||||
| **Downtime** | 45 Minutes |
|
||||
| **Time** | 14:00 – 14:45 EEST |
|
||||
| **Severity** | SEV-1 (Critical) |
|
||||
| **Customer Impact** | 100% API transaction failures |
|
||||
| **Observed Errors** | 502 Bad Gateway |
|
||||
|
||||
---
|
||||
|
||||
# Executive Summary
|
||||
|
||||
A sudden **400% increase in traffic** caused the backend API to consume memory rapidly.
|
||||
|
||||
The application maintained an **unbounded in-memory cache**, storing increasingly large datasets without expiration or eviction.
|
||||
|
||||
Once the container reached its configured memory limit, the Linux kernel terminated the process (**OOMKilled - Exit Code 137**) to protect node stability.
|
||||
|
||||
The orchestration platform continuously restarted the containers, causing a **CrashLoopBackOff** cycle and preventing the service from recovering automatically.
|
||||
|
||||
Service was restored after:
|
||||
|
||||
- Increasing the container memory limit
|
||||
- Restarting the backend pods
|
||||
- Verifying successful application startup
|
||||
|
||||
Future mitigation requires:
|
||||
|
||||
- Memory-based auto-scaling
|
||||
- Cache eviction policies
|
||||
- API pagination
|
||||
- Improved monitoring and alerting
|
||||
|
||||
---
|
||||
|
||||
# Incident Timeline
|
||||
|
||||
| Time | Event |
|
||||
|------|-------|
|
||||
| **13:50** | Traffic increased by approximately **400%** due to an unexpected external campaign. |
|
||||
| **13:58** | Container memory utilization exceeded **95%** of allocated memory. |
|
||||
| **14:00** | First container terminated with **OOMKilled (Exit Code 137)**. Service degradation begins. |
|
||||
| **14:05 – 14:30** | Containers repeatedly restarted by the orchestrator, resulting in a CrashLoop and complete outage. |
|
||||
| **14:30** | On-call infrastructure engineer identified repeated OOMKilled events from platform logs. |
|
||||
| **14:35** | Memory limit increased from **512Mi** to **2Gi** and backend pods restarted manually. |
|
||||
| **14:45** | Service stabilized and API responses returned **HTTP 200 OK**. Incident closed. |
|
||||
|
||||
---
|
||||
|
||||
# Root Cause Analysis
|
||||
|
||||
## Direct Cause
|
||||
|
||||
The Linux kernel terminated the backend process because the application attempted to allocate more memory than the container's configured memory limit.
|
||||
|
||||
```
|
||||
Exit Code: 137
|
||||
Reason: OOMKilled
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Underlying Causes
|
||||
|
||||
### Unbounded In-Memory Cache
|
||||
|
||||
The application stored data inside a local in-memory dictionary that had:
|
||||
|
||||
- No Time-To-Live (TTL)
|
||||
- No maximum cache size
|
||||
- No eviction policy
|
||||
|
||||
Large requests continuously expanded the cache until the process exhausted available memory.
|
||||
|
||||
---
|
||||
|
||||
### Missing API Pagination
|
||||
|
||||
Several endpoints returned very large datasets.
|
||||
|
||||
Without pagination:
|
||||
|
||||
- Large payloads were cached
|
||||
- Memory usage increased rapidly
|
||||
- Garbage collection became inefficient
|
||||
|
||||
---
|
||||
|
||||
### Lack of Horizontal Auto-Scaling
|
||||
|
||||
The deployment relied on static memory limits and a fixed number of replicas.
|
||||
|
||||
As traffic increased:
|
||||
|
||||
- No new replicas were created
|
||||
- Existing containers absorbed all incoming traffic
|
||||
- Memory utilization reached critical levels
|
||||
|
||||
---
|
||||
|
||||
# Proposed Auto-Scaling Architecture
|
||||
|
||||
To prevent similar incidents, backend workloads should scale automatically based on memory utilization.
|
||||
|
||||
| Parameter | Recommended Value |
|
||||
|-----------|-------------------|
|
||||
| **Scaling Metric** | Average Memory Utilization |
|
||||
| **Target Utilization** | 75% |
|
||||
| **Minimum Replicas** | 3 |
|
||||
| **Maximum Replicas** | 12 |
|
||||
| **Scale-Up Policy** | Add up to 4 replicas immediately |
|
||||
| **Scale-Down Policy** | Wait 5 minutes after memory falls below 40% |
|
||||
|
||||
---
|
||||
|
||||
# Horizontal Pod Autoscaler Configuration
|
||||
|
||||
```yaml
|
||||
apiVersion: autoscaling/v2
|
||||
kind: HorizontalPodAutoscaler
|
||||
metadata:
|
||||
name: backend-api-scaler
|
||||
|
||||
spec:
|
||||
minReplicas: 3
|
||||
maxReplicas: 12
|
||||
|
||||
metrics:
|
||||
- type: Resource
|
||||
resource:
|
||||
name: memory
|
||||
target:
|
||||
type: Utilization
|
||||
averageUtilization: 75
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Observability & Early Detection
|
||||
|
||||
To move from reactive troubleshooting to proactive monitoring, the following observability improvements should be implemented.
|
||||
|
||||
## Alerting Rules
|
||||
|
||||
### Memory Utilization Alert
|
||||
|
||||
Trigger notification when:
|
||||
|
||||
- Memory utilization exceeds **70%**
|
||||
- Sustained for **2 consecutive minutes**
|
||||
|
||||
Notification targets:
|
||||
|
||||
- Email
|
||||
- Slack
|
||||
- Webhook
|
||||
|
||||
---
|
||||
|
||||
### CrashLoop Detection
|
||||
|
||||
Create a high-priority incident whenever:
|
||||
|
||||
- Container restart count exceeds **3**
|
||||
- Within a **10-minute** window
|
||||
|
||||
---
|
||||
|
||||
## Dashboard Metrics
|
||||
|
||||
Create a dedicated monitoring dashboard showing:
|
||||
|
||||
- Container memory usage
|
||||
- Container memory limits
|
||||
- Application heap size
|
||||
- Garbage Collection duration
|
||||
- API request rate
|
||||
- Current replica count
|
||||
- Container restart count
|
||||
- CPU utilization
|
||||
- Response latency
|
||||
- Error rate (4xx / 5xx)
|
||||
|
||||
---
|
||||
|
||||
# Action Items
|
||||
|
||||
| Task | Owner | Priority | Status |
|
||||
|------|-------|----------|--------|
|
||||
| Implement cache eviction (LRU + TTL) | Development Team | Critical | In Progress |
|
||||
| Enforce API pagination | Development Team | Critical | In Progress |
|
||||
| Deploy memory-based Horizontal Pod Autoscaler | DevOps Team | High | Not Started |
|
||||
| Create synthetic load tests (5× traffic) | QA Team | Medium | Not Started |
|
||||
| Update on-call operational runbooks | Operations Team | Low | Completed |
|
||||
|
||||
---
|
||||
|
||||
# Lessons Learned
|
||||
|
||||
The incident highlighted several architectural improvements required to increase platform resilience.
|
||||
|
||||
## Infrastructure
|
||||
|
||||
- Configure Horizontal Pod Autoscaler (HPA)
|
||||
- Define resource requests and limits carefully
|
||||
- Monitor memory consumption continuously
|
||||
|
||||
## Application
|
||||
|
||||
- Implement cache eviction (LRU)
|
||||
- Apply cache expiration (TTL)
|
||||
- Enforce pagination on all large API endpoints
|
||||
- Optimize memory allocation patterns
|
||||
|
||||
## Operations
|
||||
|
||||
- Improve proactive alerting
|
||||
- Expand observability dashboards
|
||||
- Regularly execute load and stress testing
|
||||
- Update incident response runbooks
|
||||
|
||||
---
|
||||
|
||||
# Resolution Summary
|
||||
|
||||
| Item | Result |
|
||||
|------|--------|
|
||||
| **Root Cause** | Unbounded in-memory cache caused container OOM |
|
||||
| **Immediate Fix** | Increased memory limit (512Mi → 2Gi) and restarted pods |
|
||||
| **Long-Term Fixes** | HPA, cache eviction, pagination, monitoring improvements |
|
||||
| **Incident Status** | ✅ Resolved |
|
||||
```
|
||||
المرجع في مشكلة جديدة
حظر مستخدم