5.2 KiB
Task 4 — Scalability & Load Distribution (15,000 Requests/Second)
Overview
This document proposes a scalable architecture for deploying an application on Ghaymah Cloud capable of handling 15,000 HTTP requests per second while maintaining high availability, fault tolerance, and low response latency.
1. High-Level Architecture
flowchart TD
Users[Users]
Users --> DNS[Ghaymah DNS / Edge]
DNS --> LB[Load Balancer]
LB --> A1[Container 1]
LB --> A2[Container 2]
LB --> A3[Container 3]
LB --> A4[...]
LB --> A39[Container 39]
A1 --> Redis[(Redis Cache)]
A2 --> Redis
A3 --> Redis
A39 --> Redis
Redis --> DB[(Primary Database)]
DB --> Storage[(Ghaymah Block Storage)]
subgraph AutoScaling
Metrics[CPU • Memory • RPS]
Decision{Scale?}
ScaleOut[Add Containers]
ScaleIn[Remove Containers]
Metrics --> Decision
Decision -->|High Load| ScaleOut
Decision -->|Low Load| ScaleIn
end
LB -. Metrics .-> Metrics
ScaleOut -. Register New Instances .-> LB
2. Capacity Calculation
Given
- Expected traffic = 15,000 requests/second
- One container capacity = 500 requests/second
Required Containers
15,000 / 500 = 30 Containers
To avoid saturation during sudden traffic spikes or node failures, a 30% safety margin is added.
30 × 1.30 = 39 Containers
Final Capacity
| Item | Value |
|---|---|
| Target Load | 15,000 req/s |
| Capacity per Container | 500 req/s |
| Base Containers | 30 |
| Safety Margin | 30% |
| Recommended Maximum | 39 Containers |
3. Auto Scaling Policy
Recommended configuration:
min_replicas: 12
max_replicas: 39
target_cpu: 65%
target_memory: 70%
target_requests_per_container: 400
scale_up:
increase: 4 containers
cooldown: 60s
scale_down:
decrease: 2 containers
cooldown: 300s
Scaling Rules
Scale Out when:
- CPU > 65%
- Memory > 70%
- Average Requests > 400 req/s per container
Scale In when:
- CPU < 35%
- Memory < 40%
- Traffic remains low for 5 minutes
This policy minimizes unnecessary scaling operations while maintaining performance.
4. Cold Start Strategy
Launching new containers requires image download, application startup, and service initialization.
To reduce startup latency:
1. Readiness Probe
New containers receive traffic only after passing repeated /health checks.
Example:
- Interval: 5 seconds
- Success Threshold: 3
2. Pre-Warmed Containers
Keep a minimum of 12 running replicas to absorb sudden traffic spikes without waiting for new containers to start.
3. Lightweight Docker Images
Use slim base images such as:
python:3.11-slim
Smaller images reduce image pull time significantly.
4. Predictive Scaling
Scale based on traffic trends rather than waiting until CPU reaches its maximum threshold.
5. Connection Draining
Before terminating a container:
- Stop accepting new requests.
- Finish active requests.
- Remove the container gracefully.
This prevents user-facing errors during scale-down.
5. Load Balancing Strategy
The Load Balancer should distribute requests evenly across all healthy containers.
Recommended algorithm:
- Round Robin
- Least Connections (preferred for variable workloads)
Health checks should continuously monitor:
/health- Container availability
- Response latency
Unhealthy containers should automatically be removed from rotation.
6. High Availability
To improve reliability:
- Deploy multiple container replicas.
- Eliminate single points of failure.
- Automatically replace unhealthy containers.
- Use horizontal scaling instead of vertical scaling.
- Keep the application stateless whenever possible.
7. Using Ghaymah Block Storage
The application layer remains stateless, allowing containers to be created or removed without affecting user requests.
Persistent storage is required for:
- User uploads
- Database storage
- Persistent logs
- Queue data
- Shared application files
Ghaymah Block Storage provides durable volumes that remain available even if containers are recreated.
Benefits include:
- Persistent data
- Easy attachment to new containers
- Independent lifecycle from application containers
- Simplified disaster recovery
8. Design Summary
| Component | Purpose |
|---|---|
| Ghaymah DNS | Entry point |
| Load Balancer | Traffic distribution |
| 39 Containers | Handle application traffic |
| Redis Cache | Reduce database load |
| Primary Database | Persistent data |
| Ghaymah Block Storage | Durable storage |
| Auto Scaling | Automatic scaling |
| Health Checks | Availability monitoring |
Conclusion
The proposed architecture can reliably handle 15,000 requests per second using horizontal scaling on Ghaymah Cloud.
Key design principles include:
- Horizontal scalability
- Automatic load balancing
- Redis caching
- Health checks
- Automatic scaling
- Graceful container lifecycle
- Persistent storage using Ghaymah Block Storage
This architecture delivers high availability, fault tolerance, and consistent performance under heavy production workloads.