الملفات

33 KiB
خام الرابط الدائم اللوم التاريخ

Scalability & Load Distribution — Architecture Report

Project: Ghayma REST API
Platform: ghaymah.systems (Ghayma Cloud)
Target Load: 15,000 requests/second
Date: July 26, 2026


Table of Contents

  1. Architecture Diagram — 15,000 req/s
  2. Container Capacity Calculation
  3. Cold Start Strategy for New Containers
  4. Ghayma Block Storage for Stateful Workloads

1. Architecture Diagram — 15,000 req/s

High-Level System Architecture

                              15,000 req/s
                                   │
                                   ▼
                       ┌───────────────────────┐
                       │      DNS / CDN        │
                       │  (dns.ghaymah.systems)│
                       └───────────┬───────────┘
                                   │
                                   ▼
                       ┌───────────────────────┐
                       │    LOAD BALANCER      │
                       │  (Layer 7 - HTTPS)    │
                       │                       │
                       │  • SSL Termination    │
                       │  • Health Checks      │
                       │  • Round Robin /      │
                       │    Least Connections  │
                       │  • Rate Limiting      │
                       └───────────┬───────────┘
                                   │
              ┌────────────────────┼────────────────────┐
              │                    │                    │
              ▼                    ▼                    ▼
     ┌─────────────┐     ┌─────────────┐     ┌─────────────┐
     │   Zone A    │     │   Zone B    │     │   Zone C    │
     │ (Rack/AZ 1) │     │ (Rack/AZ 2) │     │ (Rack/AZ 3) │
     └──────┬──────┘     └──────┬──────┘     └──────┬──────┘
            │                   │                   │
     ┌──────┴──────┐     ┌──────┴──────┐     ┌──────┴──────┐
     │ ┌─────────┐ │     │ ┌─────────┐ │     │ ┌─────────┐ │
     │ │Container│ │     │ │Container│ │     │ │Container│ │
     │ │   1-15  │ │     │ │  16-30  │ │     │ │  31-44  │ │
     │ │ (15x)   │ │     │ │ (15x)   │ │     │ │ (14x)   │ │
     │ └─────────┘ │     │ └─────────┘ │     │ └─────────┘ │
     └──────┬──────┘     └──────┬──────┘     └──────┬──────┘
            │                   │                   │
            └───────────────────┼───────────────────┘
                                │
              ┌─────────────────┼─────────────────┐
              │                 │                 │
              ▼                 ▼                 ▼
     ┌──────────────┐  ┌──────────────┐  ┌──────────────┐
     │   Database   │  │    Redis     │  │ Block Storage│
     │  (Primary +  │  │  (Session &  │  │  (Persistent │
     │   Replicas)  │  │    Cache)    │  │    Volumes)  │
     └──────────────┘  └──────────────┘  └──────────────┘

Detailed Component Architecture

┌──────────────────────────────────────────────────────────────────────────────────────┐
│                                                                                      │
│                           GHAYMA CLOUD PLATFORM                                      │
│                                                                                      │
│  ┌────────────────────────────────────────────────────────────────────────────────┐   │
│  │                          NETWORKING LAYER                                     │   │
│  │                                                                                │   │
│  │   Internet ──▶ [DNS] ──▶ [CDN/WAF] ──▶ [Load Balancer (L7)]                  │   │
│  │                                              │                                 │   │
│  │                          ┌────────────────────┤                                │   │
│  │                          │ Health Checks      │                                │   │
│  │                          │ GET /health (10s)  │                                │   │
│  │                          │ Threshold: 3 fails │                                │   │
│  │                          └────────────────────┘                                │   │
│  └────────────────────────────────┬───────────────────────────────────────────────┘   │
│                                   │                                                   │
│  ┌────────────────────────────────┼───────────────────────────────────────────────┐   │
│  │                       COMPUTE LAYER (44 Containers)                            │   │
│  │                                                                                │   │
│  │   ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐        ┌─────────┐        │   │
│  │   │  C-01   │ │  C-02   │ │  C-03   │ │  C-04   │  ...   │  C-44   │        │   │
│  │   │ 500r/s  │ │ 500r/s  │ │ 500r/s  │ │ 500r/s  │        │ 500r/s  │        │   │
│  │   │ 512MB   │ │ 512MB   │ │ 512MB   │ │ 512MB   │        │ 512MB   │        │   │
│  │   │ 0.5 CPU │ │ 0.5 CPU │ │ 0.5 CPU │ │ 0.5 CPU │        │ 0.5 CPU │        │   │
│  │   └────┬────┘ └────┬────┘ └────┬────┘ └────┬────┘        └────┬────┘        │   │
│  │        └────────────┴────────────┴────────────┴──── ... ───────┘              │   │
│  │                                   │                                            │   │
│  │   ┌──────────────────────────────────────────────────────────────────┐        │   │
│  │   │              AUTO-SCALER (HPA)                                  │        │   │
│  │   │  Min: 44   |   Max: 66   |   Scale at: CPU>75% or Mem>70%     │        │   │
│  │   └──────────────────────────────────────────────────────────────────┘        │   │
│  └────────────────────────────────┬───────────────────────────────────────────────┘   │
│                                   │                                                   │
│  ┌────────────────────────────────┼───────────────────────────────────────────────┐   │
│  │                        DATA LAYER                                              │   │
│  │                                                                                │   │
│  │   ┌──────────────────┐  ┌─────────────┐  ┌──────────────────────────────┐    │   │
│  │   │  PostgreSQL DB   │  │   Redis     │  │  Ghayma Block Storage       │    │   │
│  │   │  Primary + 2     │  │   Cluster   │  │  (Persistent Volumes)       │    │   │
│  │   │  Read Replicas   │  │  (Cache +   │  │  • Database data (100GB)    │    │   │
│  │   │                  │  │   Sessions) │  │  • Logs & uploads (50GB)    │    │   │
│  │   └──────────────────┘  └─────────────┘  └──────────────────────────────┘    │   │
│  └────────────────────────────────────────────────────────────────────────────────┘   │
│                                                                                      │
│  ┌────────────────────────────────────────────────────────────────────────────────┐   │
│  │                     OBSERVABILITY LAYER                                        │   │
│  │                                                                                │   │
│  │   [Prometheus] ──▶ [Grafana Dashboards] ──▶ [AlertManager] ──▶ [PagerDuty]   │   │
│  │   [Container Logs] ──▶ [Centralized Logging] ──▶ [Log Analytics]              │   │
│  └────────────────────────────────────────────────────────────────────────────────┘   │
│                                                                                      │
└──────────────────────────────────────────────────────────────────────────────────────┘

Load Balancer Distribution Strategy

                    15,000 req/s incoming
                          │
                 ┌────────┴────────┐
                 │  LOAD BALANCER  │
                 │                 │
                 │  Algorithm:     │
                 │  Least          │
                 │  Connections    │
                 └────────┬────────┘
                          │
         ┌────────────────┼────────────────┐
         │                │                │
    Zone A (33%)     Zone B (33%)     Zone C (34%)
    ~5,000 r/s       ~5,000 r/s       ~5,000 r/s
    15 containers    15 containers    14 containers
    ~333 r/s each    ~333 r/s each    ~357 r/s each

Each container handles ~341 req/s on average — well within the 500 req/s capacity, leaving the 30% safety margin intact.


2. Container Capacity Calculation

Given Parameters

Parameter Value
Total traffic 15,000 req/s
Capacity per container 500 req/s
Safety margin 30%

Step-by-Step Calculation

Step 1: Calculate Effective Capacity per Container

The 30% margin means each container should only be utilized at 70% of its max capacity:

Effective capacity = Max capacity × (1 - Safety margin)
Effective capacity = 500 req/s × (1 - 0.30)
Effective capacity = 500 × 0.70
Effective capacity = 350 req/s per container

Step 2: Calculate Minimum Containers Required

Minimum containers = Total traffic ÷ Effective capacity
Minimum containers = 15,000 ÷ 350
Minimum containers = 42.86
Minimum containers = 43 containers (rounded up)

Step 3: Add Availability Buffer (+1)

To handle single-container failures without breaching capacity:

Recommended containers = 43 + 1 (failure tolerance)
Recommended containers = 44 containers

Summary Table

Metric Value Formula
Total load 15,000 req/s Given
Max capacity/container 500 req/s Given
Safety margin 30% Given
Effective capacity/container 350 req/s 500 × 0.70
Minimum containers 43 ⌈15,000 ÷ 350⌉
Recommended (with buffer) 44 43 + 1
Actual utilization per container ~341 req/s 15,000 ÷ 44
Utilization percentage 68.2% 341 ÷ 500
Remaining headroom 31.8% Exceeds 30% margin

Capacity Planning Visualization

Container Capacity: 500 req/s
                    ├─────────────────────────────────────────────────┤
                    0                                               500

Actual load per container: ~341 req/s (68.2%)
                    ├──────────────────────────────────┤
                    0                                 341

Safety margin zone (30%):
                    ├──────────────────────────────────┤░░░░░░░░░░░░░░┤
                    0                                 350            500
                                                      ▲
                                                  Threshold
                                                (don't exceed)

Auto-Scaling Tiers

Traffic Level Containers Needed Status
5,000 req/s (low) 15 Minimum viable
10,000 req/s (moderate) 30 Normal scaling
15,000 req/s (target) 44 Design target
20,000 req/s (spike) 58 Burst capacity
25,000 req/s (peak) 72 Maximum scaling

Resource Allocation per Container

Resource Value
Memory Request 256 Mi
Memory Limit 512 Mi
CPU Request 250m (0.25 vCPU)
CPU Limit 500m (0.5 vCPU)

Total Cluster Resources (44 containers)

Resource Per Container × 44 Containers Total
Memory 512 Mi × 44 22 GB
CPU 0.5 vCPU × 44 22 vCPUs
Throughput 350 req/s (effective) × 44 15,400 req/s

3. Cold Start Strategy for New Containers

What is a Cold Start?

A cold start occurs when a new container instance is created and must initialize before serving traffic. During this period:

  • The container image is pulled from the registry
  • The Node.js runtime starts
  • Dependencies are loaded
  • The application connects to databases/caches
  • The first requests experience higher latency
  Cold Start Timeline (without optimization)
  ──────────────────────────────────────────────────────
  │ Image Pull │ Runtime Init │ App Boot │ DB Connect │ Ready │
  0s          3s            4s         5s           7s      8s
                                                            ▲
                                                    First request served

Cold Start Strategy — 5 Layers

Layer 1: Pre-Warmed Container Pool

Maintain a pool of idle but ready containers that can immediately accept traffic:

Active Containers (44)          Pre-Warmed Pool (4-6)
┌────┬────┬────┬─ ... ─┬────┐  ┌────┬────┬────┬────┐
│ C1 │ C2 │ C3 │       │C44 │  │ W1 │ W2 │ W3 │ W4 │
│BUSY│BUSY│BUSY│       │BUSY│  │IDLE│IDLE│IDLE│IDLE│
└────┴────┴────┴─ ... ─┴────┘  └────┴────┴────┴────┘
                                  │
                         Traffic spike detected
                                  │
                                  ▼
                         W1, W2 immediately start
                         serving (0s cold start)

Implementation:

# HPA with pre-warmed minimum
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
  minReplicas: 48    # 44 active + 4 pre-warmed
  maxReplicas: 72

Layer 2: Optimized Docker Image (Fast Pulls)

Reduce image pull time by using lightweight images:

# ✅ GOOD: Alpine-based image (~150MB, pulls in ~2s)
FROM node:20-alpine

# ❌ BAD: Full Debian image (~1GB, pulls in ~12s)
FROM node:20
Image Base Size Pull Time Cold Start Impact
node:20 (Debian) ~1 GB ~12s Very slow
node:20-slim ~250 MB ~5s Moderate
node:20-alpine ~150 MB ~2s Fast

Layer 3: Application-Level Readiness

Implement a readiness probe that only marks the container as "ready" after it has fully initialized:

// In src/index.js
let isReady = false;

// Readiness check - only returns 200 when app is fully initialized
app.get('/ready', (req, res) => {
  if (isReady) {
    res.status(200).json({ ready: true });
  } else {
    res.status(503).json({ ready: false, message: 'Warming up...' });
  }
});

// App startup sequence
async function initialize() {
  // 1. Connect to database
  await connectDatabase();
  
  // 2. Connect to Redis cache
  await connectRedis();
  
  // 3. Pre-load frequently accessed data into memory
  await preloadCache();
  
  // 4. Mark as ready ONLY after everything is initialized
  isReady = true;
  console.log('✅ Application fully initialized and ready');
}

app.listen(PORT, async () => {
  console.log(`Server starting on port ${PORT}...`);
  await initialize();
});
# Kubernetes readiness probe
readinessProbe:
  httpGet:
    path: /ready        # Not /health — separate from liveness
    port: 3000
  initialDelaySeconds: 3
  periodSeconds: 5
  failureThreshold: 3   # Must pass 3 checks before receiving traffic

The load balancer will not route traffic to the container until /ready returns HTTP 200.

Layer 4: Proactive Scaling (Anticipate Demand)

Scale before the spike hits, not after:

# CronJob-based pre-scaling for known traffic patterns
apiVersion: batch/v1
kind: CronJob
metadata:
  name: pre-scale-morning
spec:
  # Scale up at 8:00 AM before morning traffic rush
  schedule: "0 8 * * 1-5"
  jobTemplate:
    spec:
      template:
        spec:
          containers:
            - name: scaler
              image: bitnami/kubectl:latest
              command:
                - kubectl
                - scale
                - deployment/ghayma-api
                - --replicas=50

Layer 5: Connection Pooling & Warm-Up Requests

After a container starts, send internal warm-up requests to pre-establish connections:

async function warmUp() {
  console.log('🔥 Running warm-up sequence...');
  
  // Pre-establish database connection pool
  await db.query('SELECT 1');
  
  // Prime V8 JIT compiler with typical request patterns
  const warmUpPaths = ['/health', '/api/v1/items'];
  for (const path of warmUpPaths) {
    await fetch(`http://localhost:${PORT}${path}`);
  }
  
  console.log('✅ Warm-up complete — ready for production traffic');
}

Cold Start Strategy Summary

┌─────────────────────────────────────────────────────────────────┐
│              COLD START MITIGATION LAYERS                        │
│                                                                  │
│  Layer 1: Pre-Warmed Pool      →  0s cold start (instant)       │
│  Layer 2: Alpine Docker Image  →  2s image pull (vs 12s)        │
│  Layer 3: Readiness Probes     →  No traffic until fully ready  │
│  Layer 4: Proactive Scaling    →  Scale before the spike        │
│  Layer 5: Warm-Up Requests     →  Pre-establish connections     │
│                                                                  │
│  Combined Effect: New containers ready in ~3-5s                 │
│  Without strategy: New containers ready in ~15-20s              │
└─────────────────────────────────────────────────────────────────┘

4. Ghayma Block Storage for Stateful Workloads

What is Ghayma Block Storage?

Ghayma Block Storage is a persistent, network-attached block storage service provided by the Ghayma Cloud platform. It provides raw storage volumes that can be attached to containers, similar to a virtual hard drive.

┌─────────────────────────────────────────────────────────────────┐
│                    BLOCK STORAGE vs CONTAINER STORAGE            │
│                                                                  │
│  Container (Ephemeral)              Block Storage (Persistent)  │
│  ┌──────────────────┐              ┌──────────────────────┐     │
│  │ Container FS     │              │ Ghayma Block Volume  │     │
│  │                  │              │                      │     │
│  │ ❌ Data lost     │              │ ✅ Data survives     │     │
│  │    when container│              │    container restarts│     │
│  │    stops/crashes │              │    and redeployments │     │
│  │                  │              │                      │     │
│  │ ❌ Not shared    │              │ ✅ Can be reattached │     │
│  │    between       │              │    to new containers │     │
│  │    containers    │              │                      │     │
│  └──────────────────┘              └──────────────────────┘     │
│                                                                  │
└─────────────────────────────────────────────────────────────────┘

Why is Block Storage Needed?

Containers are stateless and ephemeral by design — when a container restarts, crashes, or is replaced during a deployment, all data written inside it is lost. This is a problem for:

Workload Type Data at Risk Block Storage Solution
Databases (PostgreSQL, MySQL) Table data, indexes, WAL logs Persistent volume for /var/lib/postgresql/data
Cache (Redis with persistence) RDB/AOF snapshots Volume for /data
File Uploads User-uploaded images, documents Volume for /app/uploads
Application Logs Log files for audit/compliance Volume for /var/log/app
Session Storage User session files Volume for /tmp/sessions

Architecture: Block Storage with Containers

┌─────────────────────────────────────────────────────────────────────┐
│                      STATEFUL WORKLOAD ARCHITECTURE                 │
│                                                                     │
│   ┌───────────────────┐       ┌───────────────────────────────┐    │
│   │   API Container   │       │   Database Container          │    │
│   │   (Stateless)     │       │   (Stateful — PostgreSQL)     │    │
│   │                   │──────▶│                               │    │
│   │   No volume       │       │   ┌─────────────────────┐    │    │
│   │   needed          │       │   │ /var/lib/postgresql  │    │    │
│   │                   │       │   │      (mount point)   │    │    │
│   └───────────────────┘       │   └──────────┬──────────┘    │    │
│                               │              │               │    │
│                               └──────────────┼───────────────┘    │
│                                              │                     │
│                                    ┌─────────▼──────────┐         │
│                                    │  GHAYMA BLOCK      │         │
│                                    │  STORAGE VOLUME    │         │
│                                    │                    │         │
│                                    │  Name: db-data     │         │
│                                    │  Size: 100 GB      │         │
│                                    │  Type: SSD (gp3)   │         │
│                                    │  IOPS: 3,000       │         │
│                                    │  Encrypted: Yes    │         │
│                                    │                    │         │
│                                    │  ✅ Persists across │         │
│                                    │     restarts       │         │
│                                    │  ✅ Automatic       │         │
│                                    │     snapshots      │         │
│                                    │  ✅ Reattachable    │         │
│                                    └────────────────────┘         │
│                                                                    │
└────────────────────────────────────────────────────────────────────┘

How to Use Ghayma Block Storage

Step 1: Create a Block Storage Volume

# Using Ghayma CLI
ghaymah storage create db-data \
  --size 100GB \
  --type ssd \
  --region me-central-1 \
  --encrypted

Step 2: Attach to a Container / Deployment

# Kubernetes PersistentVolumeClaim using Ghayma storage class
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: db-data-pvc
spec:
  accessModes:
    - ReadWriteOnce          # Single container read/write
  storageClassName: ghaymah-block-ssd
  resources:
    requests:
      storage: 100Gi
---
# Deployment using the volume
apiVersion: apps/v1
kind: Deployment
metadata:
  name: postgres-db
spec:
  replicas: 1               # Stateful — typically 1 primary
  template:
    spec:
      containers:
        - name: postgres
          image: postgres:16-alpine
          ports:
            - containerPort: 5432
          volumeMounts:
            - name: db-storage
              mountPath: /var/lib/postgresql/data
          env:
            - name: POSTGRES_PASSWORD
              valueFrom:
                secretKeyRef:
                  name: db-secret
                  key: password
      volumes:
        - name: db-storage
          persistentVolumeClaim:
            claimName: db-data-pvc   # Ghayma Block Storage volume

Step 3: Verify the Volume

# Check volume status
ghaymah storage list

# Output:
# NAME       SIZE    TYPE    STATUS     ATTACHED TO          REGION
# db-data    100GB   SSD     attached   postgres-db-pod-0    me-central-1

Block Storage Features on Ghayma Cloud

Feature Description
Persistence Data survives container restarts, crashes, and redeployments
SSD Performance Low-latency SSD storage with up to 16,000 IOPS
Encryption At-rest encryption (AES-256) for security compliance
Snapshots Point-in-time snapshots for backup and disaster recovery
Resizing Expand volume size without downtime
Reattachment Detach from one container and attach to another (for migration)
Availability Zones Volume and container must be in the same zone

When to Use Block Storage vs Other Options

Storage Type Use Case Persistence Shared Access Speed
Container FS Temp files, build artifacts Ephemeral No Fastest
Ghayma Block Storage Databases, stateful apps, logs Persistent Single container Fast (SSD)
Object Storage (S3-like) Images, videos, backups Persistent Multi-container 🐌 Moderate
Shared File Storage (NFS) Shared config, media libraries Persistent Multi-container 🔄 Moderate

Data Lifecycle with Block Storage

 Container crashes or redeploys
         │
         ▼
 ┌───────────────────────────────────────────┐
 │  Old container is terminated              │
 │  ❌ Container filesystem is destroyed     │
 │  ✅ Block Storage volume is PRESERVED     │
 └────────────────────┬──────────────────────┘
                      │
                      ▼
 ┌───────────────────────────────────────────┐
 │  New container starts                     │
 │  Block Storage volume is reattached       │
 │  Data is immediately available at the     │
 │  same mount path (/var/lib/postgresql)    │
 │                                           │
 │  ✅ Zero data loss                        │
 │  ✅ Database resumes from exact state     │
 └───────────────────────────────────────────┘

Summary

Requirement Solution
Architecture for 15K req/s 3-zone architecture with L7 load balancer, 44 containers, auto-scaler, and data layer
Container count 44 containers (15,000 ÷ 350 effective req/s + 1 buffer) at 68.2% utilization
Cold start strategy 5-layer approach: pre-warmed pool, Alpine images, readiness probes, proactive scaling, warm-up requests
Block Storage Persistent SSD volumes for databases and stateful workloads that survive container restarts

Document prepared for: Ghayma Assessment — Scalability & Load Distribution section
Last updated: July 26, 2026