33 KiB
Scalability & Load Distribution — Architecture Report
Project: Ghayma REST API
Platform: ghaymah.systems (Ghayma Cloud)
Target Load: 15,000 requests/second
Date: July 26, 2026
Table of Contents
- Architecture Diagram — 15,000 req/s
- Container Capacity Calculation
- Cold Start Strategy for New Containers
- Ghayma Block Storage for Stateful Workloads
1. Architecture Diagram — 15,000 req/s
High-Level System Architecture
15,000 req/s
│
▼
┌───────────────────────┐
│ DNS / CDN │
│ (dns.ghaymah.systems)│
└───────────┬───────────┘
│
▼
┌───────────────────────┐
│ LOAD BALANCER │
│ (Layer 7 - HTTPS) │
│ │
│ • SSL Termination │
│ • Health Checks │
│ • Round Robin / │
│ Least Connections │
│ • Rate Limiting │
└───────────┬───────────┘
│
┌────────────────────┼────────────────────┐
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Zone A │ │ Zone B │ │ Zone C │
│ (Rack/AZ 1) │ │ (Rack/AZ 2) │ │ (Rack/AZ 3) │
└──────┬──────┘ └──────┬──────┘ └──────┬──────┘
│ │ │
┌──────┴──────┐ ┌──────┴──────┐ ┌──────┴──────┐
│ ┌─────────┐ │ │ ┌─────────┐ │ │ ┌─────────┐ │
│ │Container│ │ │ │Container│ │ │ │Container│ │
│ │ 1-15 │ │ │ │ 16-30 │ │ │ │ 31-44 │ │
│ │ (15x) │ │ │ │ (15x) │ │ │ │ (14x) │ │
│ └─────────┘ │ │ └─────────┘ │ │ └─────────┘ │
└──────┬──────┘ └──────┬──────┘ └──────┬──────┘
│ │ │
└───────────────────┼───────────────────┘
│
┌─────────────────┼─────────────────┐
│ │ │
▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Database │ │ Redis │ │ Block Storage│
│ (Primary + │ │ (Session & │ │ (Persistent │
│ Replicas) │ │ Cache) │ │ Volumes) │
└──────────────┘ └──────────────┘ └──────────────┘
Detailed Component Architecture
┌──────────────────────────────────────────────────────────────────────────────────────┐
│ │
│ GHAYMA CLOUD PLATFORM │
│ │
│ ┌────────────────────────────────────────────────────────────────────────────────┐ │
│ │ NETWORKING LAYER │ │
│ │ │ │
│ │ Internet ──▶ [DNS] ──▶ [CDN/WAF] ──▶ [Load Balancer (L7)] │ │
│ │ │ │ │
│ │ ┌────────────────────┤ │ │
│ │ │ Health Checks │ │ │
│ │ │ GET /health (10s) │ │ │
│ │ │ Threshold: 3 fails │ │ │
│ │ └────────────────────┘ │ │
│ └────────────────────────────────┬───────────────────────────────────────────────┘ │
│ │ │
│ ┌────────────────────────────────┼───────────────────────────────────────────────┐ │
│ │ COMPUTE LAYER (44 Containers) │ │
│ │ │ │
│ │ ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ │
│ │ │ C-01 │ │ C-02 │ │ C-03 │ │ C-04 │ ... │ C-44 │ │ │
│ │ │ 500r/s │ │ 500r/s │ │ 500r/s │ │ 500r/s │ │ 500r/s │ │ │
│ │ │ 512MB │ │ 512MB │ │ 512MB │ │ 512MB │ │ 512MB │ │ │
│ │ │ 0.5 CPU │ │ 0.5 CPU │ │ 0.5 CPU │ │ 0.5 CPU │ │ 0.5 CPU │ │ │
│ │ └────┬────┘ └────┬────┘ └────┬────┘ └────┬────┘ └────┬────┘ │ │
│ │ └────────────┴────────────┴────────────┴──── ... ───────┘ │ │
│ │ │ │ │
│ │ ┌──────────────────────────────────────────────────────────────────┐ │ │
│ │ │ AUTO-SCALER (HPA) │ │ │
│ │ │ Min: 44 | Max: 66 | Scale at: CPU>75% or Mem>70% │ │ │
│ │ └──────────────────────────────────────────────────────────────────┘ │ │
│ └────────────────────────────────┬───────────────────────────────────────────────┘ │
│ │ │
│ ┌────────────────────────────────┼───────────────────────────────────────────────┐ │
│ │ DATA LAYER │ │
│ │ │ │
│ │ ┌──────────────────┐ ┌─────────────┐ ┌──────────────────────────────┐ │ │
│ │ │ PostgreSQL DB │ │ Redis │ │ Ghayma Block Storage │ │ │
│ │ │ Primary + 2 │ │ Cluster │ │ (Persistent Volumes) │ │ │
│ │ │ Read Replicas │ │ (Cache + │ │ • Database data (100GB) │ │ │
│ │ │ │ │ Sessions) │ │ • Logs & uploads (50GB) │ │ │
│ │ └──────────────────┘ └─────────────┘ └──────────────────────────────┘ │ │
│ └────────────────────────────────────────────────────────────────────────────────┘ │
│ │
│ ┌────────────────────────────────────────────────────────────────────────────────┐ │
│ │ OBSERVABILITY LAYER │ │
│ │ │ │
│ │ [Prometheus] ──▶ [Grafana Dashboards] ──▶ [AlertManager] ──▶ [PagerDuty] │ │
│ │ [Container Logs] ──▶ [Centralized Logging] ──▶ [Log Analytics] │ │
│ └────────────────────────────────────────────────────────────────────────────────┘ │
│ │
└──────────────────────────────────────────────────────────────────────────────────────┘
Load Balancer Distribution Strategy
15,000 req/s incoming
│
┌────────┴────────┐
│ LOAD BALANCER │
│ │
│ Algorithm: │
│ Least │
│ Connections │
└────────┬────────┘
│
┌────────────────┼────────────────┐
│ │ │
Zone A (33%) Zone B (33%) Zone C (34%)
~5,000 r/s ~5,000 r/s ~5,000 r/s
15 containers 15 containers 14 containers
~333 r/s each ~333 r/s each ~357 r/s each
Each container handles ~341 req/s on average — well within the 500 req/s capacity, leaving the 30% safety margin intact.
2. Container Capacity Calculation
Given Parameters
| Parameter | Value |
|---|---|
| Total traffic | 15,000 req/s |
| Capacity per container | 500 req/s |
| Safety margin | 30% |
Step-by-Step Calculation
Step 1: Calculate Effective Capacity per Container
The 30% margin means each container should only be utilized at 70% of its max capacity:
Effective capacity = Max capacity × (1 - Safety margin)
Effective capacity = 500 req/s × (1 - 0.30)
Effective capacity = 500 × 0.70
Effective capacity = 350 req/s per container
Step 2: Calculate Minimum Containers Required
Minimum containers = Total traffic ÷ Effective capacity
Minimum containers = 15,000 ÷ 350
Minimum containers = 42.86
Minimum containers = 43 containers (rounded up)
Step 3: Add Availability Buffer (+1)
To handle single-container failures without breaching capacity:
Recommended containers = 43 + 1 (failure tolerance)
Recommended containers = 44 containers
Summary Table
| Metric | Value | Formula |
|---|---|---|
| Total load | 15,000 req/s | Given |
| Max capacity/container | 500 req/s | Given |
| Safety margin | 30% | Given |
| Effective capacity/container | 350 req/s | 500 × 0.70 |
| Minimum containers | 43 | ⌈15,000 ÷ 350⌉ |
| Recommended (with buffer) | 44 | 43 + 1 |
| Actual utilization per container | ~341 req/s | 15,000 ÷ 44 |
| Utilization percentage | 68.2% | 341 ÷ 500 |
| Remaining headroom | 31.8% | ✅ Exceeds 30% margin |
Capacity Planning Visualization
Container Capacity: 500 req/s
├─────────────────────────────────────────────────┤
0 500
Actual load per container: ~341 req/s (68.2%)
├──────────────────────────────────┤
0 341
Safety margin zone (30%):
├──────────────────────────────────┤░░░░░░░░░░░░░░┤
0 350 500
▲
Threshold
(don't exceed)
Auto-Scaling Tiers
| Traffic Level | Containers Needed | Status |
|---|---|---|
| 5,000 req/s (low) | 15 | Minimum viable |
| 10,000 req/s (moderate) | 30 | Normal scaling |
| 15,000 req/s (target) | 44 | Design target |
| 20,000 req/s (spike) | 58 | Burst capacity |
| 25,000 req/s (peak) | 72 | Maximum scaling |
Resource Allocation per Container
| Resource | Value |
|---|---|
| Memory Request | 256 Mi |
| Memory Limit | 512 Mi |
| CPU Request | 250m (0.25 vCPU) |
| CPU Limit | 500m (0.5 vCPU) |
Total Cluster Resources (44 containers)
| Resource | Per Container | × 44 Containers | Total |
|---|---|---|---|
| Memory | 512 Mi | × 44 | 22 GB |
| CPU | 0.5 vCPU | × 44 | 22 vCPUs |
| Throughput | 350 req/s (effective) | × 44 | 15,400 req/s |
3. Cold Start Strategy for New Containers
What is a Cold Start?
A cold start occurs when a new container instance is created and must initialize before serving traffic. During this period:
- The container image is pulled from the registry
- The Node.js runtime starts
- Dependencies are loaded
- The application connects to databases/caches
- The first requests experience higher latency
Cold Start Timeline (without optimization)
──────────────────────────────────────────────────────
│ Image Pull │ Runtime Init │ App Boot │ DB Connect │ Ready │
0s 3s 4s 5s 7s 8s
▲
First request served
Cold Start Strategy — 5 Layers
Layer 1: Pre-Warmed Container Pool
Maintain a pool of idle but ready containers that can immediately accept traffic:
Active Containers (44) Pre-Warmed Pool (4-6)
┌────┬────┬────┬─ ... ─┬────┐ ┌────┬────┬────┬────┐
│ C1 │ C2 │ C3 │ │C44 │ │ W1 │ W2 │ W3 │ W4 │
│BUSY│BUSY│BUSY│ │BUSY│ │IDLE│IDLE│IDLE│IDLE│
└────┴────┴────┴─ ... ─┴────┘ └────┴────┴────┴────┘
│
Traffic spike detected
│
▼
W1, W2 immediately start
serving (0s cold start)
Implementation:
# HPA with pre-warmed minimum
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
minReplicas: 48 # 44 active + 4 pre-warmed
maxReplicas: 72
Layer 2: Optimized Docker Image (Fast Pulls)
Reduce image pull time by using lightweight images:
# ✅ GOOD: Alpine-based image (~150MB, pulls in ~2s)
FROM node:20-alpine
# ❌ BAD: Full Debian image (~1GB, pulls in ~12s)
FROM node:20
| Image Base | Size | Pull Time | Cold Start Impact |
|---|---|---|---|
node:20 (Debian) |
~1 GB | ~12s | Very slow |
node:20-slim |
~250 MB | ~5s | Moderate |
node:20-alpine |
~150 MB | ~2s | ✅ Fast |
Layer 3: Application-Level Readiness
Implement a readiness probe that only marks the container as "ready" after it has fully initialized:
// In src/index.js
let isReady = false;
// Readiness check - only returns 200 when app is fully initialized
app.get('/ready', (req, res) => {
if (isReady) {
res.status(200).json({ ready: true });
} else {
res.status(503).json({ ready: false, message: 'Warming up...' });
}
});
// App startup sequence
async function initialize() {
// 1. Connect to database
await connectDatabase();
// 2. Connect to Redis cache
await connectRedis();
// 3. Pre-load frequently accessed data into memory
await preloadCache();
// 4. Mark as ready ONLY after everything is initialized
isReady = true;
console.log('✅ Application fully initialized and ready');
}
app.listen(PORT, async () => {
console.log(`Server starting on port ${PORT}...`);
await initialize();
});
# Kubernetes readiness probe
readinessProbe:
httpGet:
path: /ready # Not /health — separate from liveness
port: 3000
initialDelaySeconds: 3
periodSeconds: 5
failureThreshold: 3 # Must pass 3 checks before receiving traffic
The load balancer will not route traffic to the container until
/readyreturns HTTP 200.
Layer 4: Proactive Scaling (Anticipate Demand)
Scale before the spike hits, not after:
# CronJob-based pre-scaling for known traffic patterns
apiVersion: batch/v1
kind: CronJob
metadata:
name: pre-scale-morning
spec:
# Scale up at 8:00 AM before morning traffic rush
schedule: "0 8 * * 1-5"
jobTemplate:
spec:
template:
spec:
containers:
- name: scaler
image: bitnami/kubectl:latest
command:
- kubectl
- scale
- deployment/ghayma-api
- --replicas=50
Layer 5: Connection Pooling & Warm-Up Requests
After a container starts, send internal warm-up requests to pre-establish connections:
async function warmUp() {
console.log('🔥 Running warm-up sequence...');
// Pre-establish database connection pool
await db.query('SELECT 1');
// Prime V8 JIT compiler with typical request patterns
const warmUpPaths = ['/health', '/api/v1/items'];
for (const path of warmUpPaths) {
await fetch(`http://localhost:${PORT}${path}`);
}
console.log('✅ Warm-up complete — ready for production traffic');
}
Cold Start Strategy Summary
┌─────────────────────────────────────────────────────────────────┐
│ COLD START MITIGATION LAYERS │
│ │
│ Layer 1: Pre-Warmed Pool → 0s cold start (instant) │
│ Layer 2: Alpine Docker Image → 2s image pull (vs 12s) │
│ Layer 3: Readiness Probes → No traffic until fully ready │
│ Layer 4: Proactive Scaling → Scale before the spike │
│ Layer 5: Warm-Up Requests → Pre-establish connections │
│ │
│ Combined Effect: New containers ready in ~3-5s │
│ Without strategy: New containers ready in ~15-20s │
└─────────────────────────────────────────────────────────────────┘
4. Ghayma Block Storage for Stateful Workloads
What is Ghayma Block Storage?
Ghayma Block Storage is a persistent, network-attached block storage service provided by the Ghayma Cloud platform. It provides raw storage volumes that can be attached to containers, similar to a virtual hard drive.
┌─────────────────────────────────────────────────────────────────┐
│ BLOCK STORAGE vs CONTAINER STORAGE │
│ │
│ Container (Ephemeral) Block Storage (Persistent) │
│ ┌──────────────────┐ ┌──────────────────────┐ │
│ │ Container FS │ │ Ghayma Block Volume │ │
│ │ │ │ │ │
│ │ ❌ Data lost │ │ ✅ Data survives │ │
│ │ when container│ │ container restarts│ │
│ │ stops/crashes │ │ and redeployments │ │
│ │ │ │ │ │
│ │ ❌ Not shared │ │ ✅ Can be reattached │ │
│ │ between │ │ to new containers │ │
│ │ containers │ │ │ │
│ └──────────────────┘ └──────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────┘
Why is Block Storage Needed?
Containers are stateless and ephemeral by design — when a container restarts, crashes, or is replaced during a deployment, all data written inside it is lost. This is a problem for:
| Workload Type | Data at Risk | Block Storage Solution |
|---|---|---|
| Databases (PostgreSQL, MySQL) | Table data, indexes, WAL logs | Persistent volume for /var/lib/postgresql/data |
| Cache (Redis with persistence) | RDB/AOF snapshots | Volume for /data |
| File Uploads | User-uploaded images, documents | Volume for /app/uploads |
| Application Logs | Log files for audit/compliance | Volume for /var/log/app |
| Session Storage | User session files | Volume for /tmp/sessions |
Architecture: Block Storage with Containers
┌─────────────────────────────────────────────────────────────────────┐
│ STATEFUL WORKLOAD ARCHITECTURE │
│ │
│ ┌───────────────────┐ ┌───────────────────────────────┐ │
│ │ API Container │ │ Database Container │ │
│ │ (Stateless) │ │ (Stateful — PostgreSQL) │ │
│ │ │──────▶│ │ │
│ │ No volume │ │ ┌─────────────────────┐ │ │
│ │ needed │ │ │ /var/lib/postgresql │ │ │
│ │ │ │ │ (mount point) │ │ │
│ └───────────────────┘ │ └──────────┬──────────┘ │ │
│ │ │ │ │
│ └──────────────┼───────────────┘ │
│ │ │
│ ┌─────────▼──────────┐ │
│ │ GHAYMA BLOCK │ │
│ │ STORAGE VOLUME │ │
│ │ │ │
│ │ Name: db-data │ │
│ │ Size: 100 GB │ │
│ │ Type: SSD (gp3) │ │
│ │ IOPS: 3,000 │ │
│ │ Encrypted: Yes │ │
│ │ │ │
│ │ ✅ Persists across │ │
│ │ restarts │ │
│ │ ✅ Automatic │ │
│ │ snapshots │ │
│ │ ✅ Reattachable │ │
│ └────────────────────┘ │
│ │
└────────────────────────────────────────────────────────────────────┘
How to Use Ghayma Block Storage
Step 1: Create a Block Storage Volume
# Using Ghayma CLI
ghaymah storage create db-data \
--size 100GB \
--type ssd \
--region me-central-1 \
--encrypted
Step 2: Attach to a Container / Deployment
# Kubernetes PersistentVolumeClaim using Ghayma storage class
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: db-data-pvc
spec:
accessModes:
- ReadWriteOnce # Single container read/write
storageClassName: ghaymah-block-ssd
resources:
requests:
storage: 100Gi
---
# Deployment using the volume
apiVersion: apps/v1
kind: Deployment
metadata:
name: postgres-db
spec:
replicas: 1 # Stateful — typically 1 primary
template:
spec:
containers:
- name: postgres
image: postgres:16-alpine
ports:
- containerPort: 5432
volumeMounts:
- name: db-storage
mountPath: /var/lib/postgresql/data
env:
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: db-secret
key: password
volumes:
- name: db-storage
persistentVolumeClaim:
claimName: db-data-pvc # Ghayma Block Storage volume
Step 3: Verify the Volume
# Check volume status
ghaymah storage list
# Output:
# NAME SIZE TYPE STATUS ATTACHED TO REGION
# db-data 100GB SSD attached postgres-db-pod-0 me-central-1
Block Storage Features on Ghayma Cloud
| Feature | Description |
|---|---|
| Persistence | Data survives container restarts, crashes, and redeployments |
| SSD Performance | Low-latency SSD storage with up to 16,000 IOPS |
| Encryption | At-rest encryption (AES-256) for security compliance |
| Snapshots | Point-in-time snapshots for backup and disaster recovery |
| Resizing | Expand volume size without downtime |
| Reattachment | Detach from one container and attach to another (for migration) |
| Availability Zones | Volume and container must be in the same zone |
When to Use Block Storage vs Other Options
| Storage Type | Use Case | Persistence | Shared Access | Speed |
|---|---|---|---|---|
| Container FS | Temp files, build artifacts | ❌ Ephemeral | ❌ No | ⚡ Fastest |
| Ghayma Block Storage | Databases, stateful apps, logs | ✅ Persistent | ❌ Single container | ⚡ Fast (SSD) |
| Object Storage (S3-like) | Images, videos, backups | ✅ Persistent | ✅ Multi-container | 🐌 Moderate |
| Shared File Storage (NFS) | Shared config, media libraries | ✅ Persistent | ✅ Multi-container | 🔄 Moderate |
Data Lifecycle with Block Storage
Container crashes or redeploys
│
▼
┌───────────────────────────────────────────┐
│ Old container is terminated │
│ ❌ Container filesystem is destroyed │
│ ✅ Block Storage volume is PRESERVED │
└────────────────────┬──────────────────────┘
│
▼
┌───────────────────────────────────────────┐
│ New container starts │
│ Block Storage volume is reattached │
│ Data is immediately available at the │
│ same mount path (/var/lib/postgresql) │
│ │
│ ✅ Zero data loss │
│ ✅ Database resumes from exact state │
└───────────────────────────────────────────┘
Summary
| Requirement | Solution |
|---|---|
| Architecture for 15K req/s | 3-zone architecture with L7 load balancer, 44 containers, auto-scaler, and data layer |
| Container count | 44 containers (15,000 ÷ 350 effective req/s + 1 buffer) at 68.2% utilization |
| Cold start strategy | 5-layer approach: pre-warmed pool, Alpine images, readiness probes, proactive scaling, warm-up requests |
| Block Storage | Persistent SSD volumes for databases and stateful workloads that survive container restarts |
Document prepared for: Ghayma Assessment — Scalability & Load Distribution section
Last updated: July 26, 2026