الملفات
ghaymah-exam-noran-sre/q4-scalability/calculations.md

3.6 KiB
خام اللوم التاريخ

Scalability and Load Balancing Design

Scenario

Design an architecture capable of handling:

  • 15,000 requests per second (req/s)

Each application container can process:

  • 500 req/s

A 30% capacity buffer is required to handle traffic spikes and unexpected load.


1. High-Level Architecture

The proposed architecture consists of the following components:

  • DNS
  • Load Balancer
  • Multiple API containers
  • Auto Scaling service
  • Container Orchestrator
  • Database
  • Ghaymah Block Storage
  • Monitoring and Alerting

Architecture Diagram

The following diagram illustrates the proposed scalable architecture for handling 15,000 requests per second using load balancing, auto scaling, centralized monitoring, and persistent storage.

Architecture


Request Flow

Users
   │
   ▼
DNS
   │
   ▼
Load Balancer
   │
   ▼
API Containers
   │
   ▼
Database

Monitoring components (Prometheus, Grafana, and centralized logging) collect metrics and logs from the Load Balancer, API containers, and database.


2. Container Capacity Calculation

Step 1: Calculate Required Containers

Traffic:

15,000 req/s

Container capacity:

500 req/s

Required containers:

15000 / 500 = 30 containers

Step 2: Add 30% Safety Buffer

30 × 1.30 = 39 containers

Minimum recommended deployment:

39 containers

This provides additional capacity for traffic spikes and reduces the risk of resource saturation during unexpected load increases.


3. Cold Start Strategy

To minimize startup delays during scaling events, the following strategies are recommended.

Warm Pool

Maintain 23 idle containers ready to receive traffic immediately.

Container Image Optimization

  • Use lightweight base images.
  • Remove unnecessary packages and dependencies.
  • Keep image size as small as possible.

Image Pre-Pulling

Pre-pull container images on worker nodes to reduce deployment time and avoid downloading images during scale-out events.

Health Checks

New containers must pass readiness and liveness checks before receiving production traffic.

Gradual Traffic Shift

The Load Balancer should gradually route traffic to newly started containers after successful health verification.

Predictive Scaling

Scale proactively based on CPU utilization, memory usage, and request-rate trends instead of waiting until the system reaches full capacity.


4. Using Ghaymah Block Storage for Stateful Workloads

Ghaymah Block Storage provides persistent storage volumes that remain available even if containers are restarted, replaced, or rescheduled.

This makes it suitable for workloads that require durable data storage.

Examples include:

  • Databases
  • Message queues
  • File storage
  • Application uploads
  • Persistent logs

Benefits

  • Persistent data survives container restarts.
  • Volumes can be attached to replacement containers.
  • Improves reliability for stateful applications.
  • Simplifies backup and disaster recovery.
  • Enhances data durability during scaling events.

Best Practices

  • Deploy multiple replicas behind a Load Balancer.
  • Enable Horizontal Auto Scaling based on CPU, memory, and request rate.
  • Configure readiness and liveness probes.
  • Use lightweight container images to reduce cold-start time.
  • Store persistent data using Ghaymah Block Storage.
  • Continuously monitor latency, error rate, CPU, and memory utilization.
  • Configure alerts for abnormal response times and infrastructure failures.
  • Regularly perform load testing to validate scalability assumptions.