4.6 KiB
Q4 — Scalability & Load Balancing (Ghaymah Cloud)
1. Architecture
See architecture.png in this folder.
Flow: Clients (15,000 req/s) → Load Balancer (round-robin + health checks) → Container Auto-Scaling Pool (39 containers) → In-Memory Cache / Database + Ghaymah Block Storage.
- The front tier (load balancer + containers) is stateless and scales horizontally.
- The data tier (database) is stateful and relies on Ghaymah Block Storage for durability, decoupled from the container lifecycle.
2. Container Capacity Calculation
Given:
- Total incoming load:
15,000 req/s - Capacity per container:
500 req/s - Safety margin:
30%(to absorb traffic spikes and node/container failures)
Formula:
Required capacity = Total load × (1 + margin)
= 15,000 × 1.30
= 19,500 req/s
Containers needed = Required capacity / Capacity per container
= 19,500 / 500
= 39 containers
Result: 39 containers
| Tier | Count | Purpose |
|---|---|---|
| Min replicas (baseline) | 30 | Covers the raw 15,000 req/s |
| Target / desired | 39 | Includes the 30% margin |
| Max replicas (HPA ceiling) | 50–55 | Extra headroom for burst traffic |
Recommendation: configure the Horizontal Pod Autoscaler (HPA) on a
composite metric — requests-per-second (via a Prometheus adapter) as the
primary signal and CPU utilization as a secondary safeguard — rather than
CPU alone, since CPU can lag behind real request pressure.
3. Cold Start Strategy for New Containers
- Pre-warmed pool (warm standby) — keep 2–3 idle, already-running containers outside the active traffic path so bursts are absorbed instantly instead of waiting for a container to be built from scratch.
- Slim, optimized images — use minimal base images (Alpine/distroless), reduce layers, and pre-pull images on nodes to avoid registry pull latency during scale-out events.
- Strict readiness probes — a container only joins the load balancer
pool after passing a
readinessProbe(DB/cache connectivity check), preventing 503s from traffic sent to a not-yet-ready instance. - Predictive/scheduled scaling alongside reactive HPA — if peak hours are known (campaigns, daily traffic peaks), trigger scheduled scale-up 5–10 minutes ahead of the peak instead of relying solely on reactive HPA, which lags behind the metrics collection window.
- Connection pooling & pre-warmed init — open DB/cache connections in an init container or internal warm-up endpoint before the first real request arrives, instead of lazy-initializing on first use.
- Graceful scale-down — use a
preStophook and a properterminationGracePeriodSecondsto drain in-flight requests before terminating containers during scale-in, avoiding dropped requests on the way down.
4. Ghaymah Block Storage for Stateful Workloads
The containers above are intentionally stateless — any of them can be killed and replaced without losing data. Anything that needs durable state (database, message queues, uploaded files) must be decoupled from the container's lifecycle — this is where Block Storage comes in:
- Decouples data from container lifecycle — a block volume is attached to a container/VM as an independent raw block device. If the container is rescheduled or fails, the same volume is re-attached to the new instance/node without data loss — unlike ephemeral container storage, which is wiped on restart.
- Best fit — relational databases (PostgreSQL/MySQL), messaging systems (Kafka/RabbitMQ), and any workload needing low-latency random I/O, since block storage offers near-local-disk performance versus object storage.
- High availability — one block volume per database replica (never share a single volume across concurrently-active instances); pair it with application-level replication (e.g., Postgres streaming replication) rather than relying on storage replication alone.
- Snapshots & backup — schedule periodic snapshots (daily minimum, plus before any structural change/upgrade) to keep RPO/RTO low.
- Vertical scalability — volumes can typically be resized as data grows without downtime, complementing the horizontal scaling of the stateless container tier.
Summary: the front tier (load balancer + container pool) scales horizontally and fast because it's stateless, while the data tier relies on Ghaymah Block Storage for durability and consistency, with an in-memory cache absorbing read pressure and reducing latency.