الملفات
root 6acc75d6ae Q2
2026-07-26 22:54:42 +00:00

3.6 KiB

Q2 - Postmortem Report

Incident Summary

Incident: Repeated OOMKilled causing service outage

Date: (Exam Scenario)

Duration: 45 minutes

Severity: High (SEV-1)

Impact:

  • API became unavailable.
  • Users received HTTP 5xx errors.
  • Service uptime was affected.
  • Customer requests could not be processed during the outage.

Timeline

Time Event
10:00 Increased memory usage observed.
10:05 First container terminated with OOMKilled.
10:06 Kubernetes restarted the container automatically.
10:10 Memory usage increased again causing another OOMKilled.
10:15 Continuous restart loop (CrashLoopBackOff).
10:20 Monitoring system generated alerts.
10:30 SRE team started investigation.
10:40 Memory limit increased and memory leak identified.
10:45 New container deployed successfully and service restored.

Root Cause Analysis

The application suffered from excessive memory consumption caused by a memory leak under increased traffic.

The container memory limit was too low to handle the workload.

When memory usage exceeded the configured limit, Kubernetes terminated the container with an OOMKilled event.

Automatic restarts repeatedly created a CrashLoopBackOff situation, extending the outage.


Contributing Factors

  • No Horizontal Pod Autoscaler (HPA).
  • Memory limits configured too aggressively.
  • No early memory usage alerts.
  • Memory leak was not detected during testing.
  • Lack of load testing before deployment.

Immediate Actions Taken

  • Increased memory limits.
  • Restarted affected containers.
  • Fixed memory leak.
  • Verified application health.
  • Monitored memory utilization after recovery.

Preventive Recommendations

  • Enable Horizontal Pod Autoscaler (HPA).
  • Configure Vertical Pod Autoscaler (VPA) if appropriate.
  • Define proper memory requests and limits.
  • Add Prometheus memory monitoring.
  • Configure Alertmanager notifications.
  • Perform stress testing before production deployment.
  • Review application memory usage regularly.

Auto-Scaling Policy

To prevent similar incidents on Ghaymah:

Horizontal Pod Autoscaler

Minimum Replicas: 2

Maximum Replicas: 10

Scale Out Conditions

  • Memory usage > 70%
  • CPU usage > 70%
  • Request rate exceeds expected capacity

Scale In Conditions

  • Memory usage < 40%
  • CPU usage < 40%
  • Stable traffic for at least 10 minutes

Cooldown Period

  • Scale Out: 60 seconds

  • Scale In: 300 seconds

Benefits

  • Prevents memory exhaustion.
  • Distributes incoming traffic.
  • Improves availability.
  • Reduces restart frequency.

Early Detection Using Ghaymah Monitoring

The incident could have been detected earlier using monitoring tools.

Metrics

  • Container Memory Usage
  • Memory Limit Percentage
  • Container Restart Count
  • OOMKilled Events
  • CPU Utilization
  • Request Rate
  • Response Time

Alerts

Critical Alert

  • Memory Usage > 85%
  • More than 3 restarts within 5 minutes
  • OOMKilled event detected

Warning Alert

  • Memory Usage > 70%
  • Rapid increase in memory consumption

Dashboards

Recommended dashboard should display:

  • Memory Usage
  • CPU Usage
  • Pod Status
  • Restart Count
  • Response Time
  • Error Rate
  • Availability

Lessons Learned

  • Memory monitoring must be proactive.
  • Autoscaling should be enabled in production.
  • Load testing should validate memory consumption.
  • Alerting should notify engineers before service failure.
  • Capacity planning should be reviewed regularly.

Status

Incident resolved

No recurring issues observed after implementing corrective actions.