الملفات
ghaymah-exam-abdelrahman-sre/q2-postmortem/Auto-Scaling.md
2026-07-27 21:32:00 +03:00

2.2 KiB

Ghaymah Auto-Scaling Policy

Goal

Design an auto-scaling policy for the Ghaymah platform that reduces the impact of memory-related incidents (like the OOMKilled outage in the postmortem), while making clear that auto-scaling alone does not fix a memory leak — it only limits the damage.


Important Note

The original incident was caused by a memory leak, not by a real traffic increase. Auto-scaling can't fix a leak — a new pod created by auto-scaling will eventually leak too. So this policy is a safety net, not a replacement for fixing the leak itself.


1. Right-Size Requests and Limits First

Before any auto-scaling works well, the memory requests and limits need to be based on real usage instead of a guess. Otherwise the autoscaler is scaling based on wrong numbers.


2. Horizontal Pod Autoscaler (HPA) — based on memory

Scale the number of pods when average memory usage crosses a threshold, so no single pod gets close to its limit as fast.

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: ghaymah-app-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: ghaymah-app
  minReplicas: 2
  maxReplicas: 6
  metrics:
  - type: Resource
    resource:
      name: memory
      target:
        type: Utilization
        averageUtilization: 70

3. Cluster Autoscaler

If HPA creates more pods than the current nodes can fit, the Cluster Autoscaler adds more nodes automatically so pods aren't stuck pending.


4. Alerting Before Scaling Limits Are Hit

Add an alert when memory usage is trending up (not just when it hits the limit), so the team can catch a leak early instead of waiting for auto-scaling or a crash.


5. Restart / Readiness Behavior

Make sure the app has proper readiness probes, so traffic isn't sent to a pod that just restarted and isn't ready yet — this reduces user-facing errors during restart loops.


Summary

This policy helps absorb real load spikes and buys extra time before hitting memory limits, which reduces how often OOMKilled events happen. But it does not replace fixing the actual memory leak — that's still the real fix.