16 KiB
Ghaymah Internship Test — Implementation Plan
Scope: 5 questions · 72 hours · 100 points (20 each) Execution model: Two agents working in parallel — Sol (Codex) and Opus 5 — with you as the integrator who does the manual platform steps (accounts, deployments, secrets).
0. Ground Rules & Repo Setup (do this first, yourself — 30 min)
Purpose: Both agents need a shared structure so their outputs merge cleanly without conflicts. Doing this before dispatching work prevents rework.
How:
- Create one Git repo (e.g.
ghaymah-test) with this layout:
ghaymah-test/
├── q1-deploy-monitor/
│ ├── app/ # API + Dockerfile
│ ├── monitor/ # monitoring script
│ └── dashboard/ # HTML/CSS/JS dashboard
├── q2-postmortem/
│ └── POSTMORTEM.md
├── q3-cicd/
│ ├── .github/workflows/ # (copied to repo root at the end)
│ └── CICD.md
├── q4-scalability/
│ └── SCALABILITY.md # diagram + math + strategy
├── q5-mithal-dashboard/
│ ├── collector/ # metrics script
│ ├── data/ # CSV/JSON output
│ └── dashboard/
└── README.md # index of all answers
- Sign up / log in to ghaymah.systems, explore: Container platform, Container Registry, CLI docs, monitoring tools, Block Storage docs. Take notes/screenshots — Q2, Q3, Q4 answers must reference real Ghaymah features, not generic cloud talk. This research is the single highest-value manual task: graders will notice platform-specific accuracy.
- Decide the stack once: Python + FastAPI for both APIs (Q1 and Q5 collector share language, smaller surface area) — or Node/Express if you prefer. The plan below assumes Python.
Why manual: account creation, credentials, and platform exploration cannot be delegated to agents (and shouldn't be — never give agents your passwords/API keys; store tokens as env vars/GitHub secrets yourself).
Parallelization Map (who does what)
| Track | Agent | Work | Rationale |
|---|---|---|---|
| A | Opus 5 | Q1 (app + Dockerfile + monitor + dashboard), Q5 (collector + dashboard) | Heaviest coding volume, frontend polish, end-to-end coherence — Opus is strongest at multi-file builds and design-quality dashboards |
| B | Sol | Q3 (CI/CD workflow + docs), Q2 (postmortem), Q4 (scaling doc + diagram) | Mostly structured writing + one YAML workflow — well-bounded, spec-driven tasks that run independently of Track A |
| — | You | Ghaymah account, deployments, registry creds, GitHub secrets, real screenshots, final review | Anything requiring credentials or the live platform |
The two tracks share zero files, so they can run truly simultaneously. Merge point: Day 2 evening, when you deploy Track A's artifacts and paste real values (URLs, registry names) into Track B's docs.
Prompting tip: give each agent the full question text plus your Ghaymah platform notes from step 0, and tell them the repo layout above so paths match.
Q1 — Deploy & Monitor an App on Ghaymah (20 pts)
1.1 The API app + /health endpoint — Opus
Purpose: The deliverable everything else in Q1 hangs off. /health is the standard liveness contract that orchestrators and your monitor script both consume.
How: Small FastAPI app with:
GET /— hello/info routeGET /health— returns{"status":"ok","uptime_s":...,"timestamp":...}with HTTP 200. Keep it dependency-free (no DB check) so it reflects process liveness only.GET /metrics— returns an in-memory request counter + simple stats (this feeds the dashboard's "عدد الطلبات" requirement). Implement with a middleware that increments a counter per request.
Implementation: app/main.py, app/requirements.txt (fastapi, uvicorn). Run with uvicorn main:app --host 0.0.0.0 --port 8080.
1.2 Dockerfile — Opus
Purpose: Shows you understand image hygiene, not just "it runs".
How: Multi-stage or slim single-stage:
python:3.12-slimbase, non-root user,COPY requirements.txtbefore code (layer caching),EXPOSE 8080,HEALTHCHECK CMD curl -f http://localhost:8080/health || exit 1,CMD uvicorn....- Add
.dockerignore.
Implementation check: you build locally: docker build -t ghaymah-api . && docker run -p 8080:8080 ghaymah-api, hit /health.
1.3 Deploy to ghaymah.systems — You (manual)
Purpose: The 20 points require a live deployment; agents can't do this.
How: Per platform docs — typically: push image to ghaymah Container Registry (docker login <registry>, docker tag, docker push), then create a container service in the dashboard/CLI pointing at the image, set port 8080, note the public URL. Screenshot the running service and the /health response for your submission.
1.4 Monitoring script (every 30s) — Opus
Purpose: Demonstrates operational thinking — a poller that produces machine-readable history the dashboard can consume.
How: monitor/monitor.py:
- Loop: every 30s,
GET <APP_URL>/healthwith timeout (e.g. 5s), recordtimestamp, status ("up"/"down"), http_code, response_ms; also pull/metricsfor request count. - Append each check to
monitor/data/checks.json(or JSONL) — this file is the dashboard's data source, so agree the schema up front:{"ts":"2026-07-26T12:00:00Z","status":"up","code":200,"latency_ms":42,"requests":1337} - Config via env var
APP_URL. Handle exceptions → status "down", latency null. Optional: alert line to stdout when 3 consecutive failures.
1.5 Dashboard (HTML/CSS/JS) — Opus
Purpose: Shows the three required metrics: status, response time, request count.
How: Single static page dashboard/index.html (inline CSS/JS, zero build step):
- Status badge (green/red) from latest check
- Line chart of
latency_ms(Chart.js from CDN, or hand-rolled SVG if you want zero dependencies) - Request counter tile
- Fetches
checks.json(served next to it, or the monitor writes into the dashboard folder) and refreshes every 30s withsetInterval.
Deliverable checklist Q1: Dockerfile ✔ live URL ✔ /health ✔ monitor script ✔ dashboard ✔ screenshots ✔
Q2 — OOMKilled Postmortem (20 pts) — Sol
Purpose: Pure documentation question testing incident-response maturity. No code; graded on structure, realism, and platform-specific recommendations.
How: q2-postmortem/POSTMORTEM.md in the classic blameless format:
- Summary — 45-min outage, repeated OOMKilled restarts (exit code 137), impact (error rate / downtime %), severity level.
- Timeline — invent a realistic minute-by-minute table: memory creep begins → first OOMKill → crash-loop (each restart re-accumulates memory faster under retry traffic) → alert fires → engineer raises memory limit + rolls back the leaking release → recovery. Timestamps, actor, action.
- Root cause — pick something concrete and defensible, e.g. a memory leak introduced in release X (unbounded in-process cache) combined with a memory limit sized for the old baseline; 5-Whys chain down to "no memory regression check in CI, no memory alerting."
- Recommendations — table with owner + priority: right-size limits, fix leak, add memory alerts at 80%, add auto-scaling (links to part 2), load-test with soak tests, add runbook.
- Auto-scaling policy design — a concrete policy for Ghaymah containers:
- Horizontal: scale out at avg memory > 70% or CPU > 65% for 2 min; min 2 / max N replicas; scale-in cooldown 5–10 min to avoid flapping.
- Explicitly explain why HPA alone doesn't fix a leak (leaks eat any replica count — scaling buys time for the fix, plus restart policy / vertical headroom). This nuance is what separates a top answer.
- Early detection with Ghaymah monitoring — reference the actual tools you found in step 0: memory-usage graphs per container, alert rules (memory > 80% for 5 min, restart-count > 3/10 min, OOMKilled event alerts), dashboards + notification channels.
Implementation: one well-formatted Markdown file; optionally a small Mermaid timeline/graph. You later swap in accurate Ghaymah tool names from your notes.
Q3 — CI/CD Pipeline (20 pts) — Sol
Purpose: Tests GitHub Actions fluency + registry integration + release discipline (manual gate).
3.1 Workflow — build & push to ghaymah Container Registry
How: .github/workflows/deploy.yml:
- Trigger:
pushtomain(+workflow_dispatch). - Job
build-push: checkout →docker/login-actionagainst the ghaymah registry usingsecrets.GHAYMAH_REGISTRY_USER/TOKEN→docker/build-push-actiontaggingregistry.ghaymah.systems/<user>/app:${{ github.sha }}and:latest. - Job
deploy-staging(needs build): deploy via ghaymah CLI to the staging service. - Job
deploy-production(needs staging): bound to a GitHub Environment namedproductionwith required reviewers — this is the manual-approval mechanism (job pauses until a human approves in the Actions UI). Sol should document that the environment + reviewers are configured in repo Settings → Environments, since that part isn't in YAML.
3.2 Staging vs Production explanation
How: Section in CICD.md: purpose (validation vs. live users), differences (data, scale/replicas, secrets, access control, alerting thresholds, deploy cadence), promotion flow diagram commit → build → staging (auto) → approval → production.
3.3 ghaymah CLI integration docs
How: CICD.md section: install CLI, ghaymah auth login with an API token stored as GHAYMAH_API_TOKEN GitHub secret, deploy/update-image command, how the workflow calls it in a step. You verify exact CLI command names from real docs and correct Sol's draft — flag every placeholder Sol writes with <!-- VERIFY --> so nothing invented ships.
You (manual): create the GitHub repo, add the secrets, create the production environment with yourself as reviewer, run the pipeline once and screenshot the approval gate.
Q4 — Scalability & Load Balancing (20 pts) — Sol
Purpose: Architecture reasoning + arithmetic + platform storage knowledge. Fully deterministic — ideal Sol task.
4.1 Architecture diagram (15,000 req/s)
How: Mermaid diagram in SCALABILITY.md (renders on GitHub): DNS → CDN/edge cache → Load Balancer (L7) → stateless API container fleet (auto-scaled) → cache layer (Redis) → DB (primary + replicas) → Block Storage for stateful pieces; plus monitoring/queue components. Annotate where the 15k req/s flows and what absorbs bursts.
4.2 Container count math — show the work
15,000 req/s ÷ 500 req/s per container = 30 containers
+30% headroom: 30 × 1.30 = 39 containers
State the answer plainly (39), then add operational notes: round up, N+1 for rolling deploys, and that headroom covers spikes + AZ loss.
4.3 Cold-start strategy
How: bullet strategy: keep a warm pool / min-replicas floor, pre-pull & slim images (small base, fewer layers), lazy-load nothing critical at boot, readiness probe gating so LB never routes to a cold container, predictive/scheduled scaling ahead of known peaks, gradual (step) scale-out policies.
4.4 ghaymah Block Storage for stateful workloads
How: explain: containers are ephemeral → attach Block Storage volumes for databases/queues/uploads; persistence across restarts/reschedules; one-writer-per-volume semantics (so stateful services scale differently than stateless API tier); snapshots/backups; IOPS considerations. Tie back to the diagram (DB nodes mount Block Storage; API tier stays diskless). You correct against real Ghaymah Block Storage docs.
Q5 — mithal.space Monitoring Dashboard (20 pts) — Opus
Purpose: The most integrated build: multi-metric collector + persisted history + richer dashboard + deployment. Assign to Opus because it shares patterns (and dashboard code style) with Q1 — one agent keeps them consistent.
5.1 Collector script (every minute) — collector/collect.py
How (one Python script, stdlib + requests):
- Latency:
requests.get("https://mithal.space", timeout=10)— measure elapsed ms. - Uptime: same request's status code → up if
200 ≤ code < 400. - SSL:
ssl+socket— open TLS connection to port 443, read certnotAfter, compute days remaining. - DNS: time
socket.getaddrinfo("mithal.space", 443)in ms. - Search response: since mithal.space is a search engine — send a query (inspect the site first to find the search URL pattern, e.g.
/search?q=testor its API endpoint) and time the response. You should check the actual URL format in a browser and give it to Opus; have Opus make it configurable. - Scheduling:
while True: run(); sleep(60)plus support one-shot mode (--once) so it can also run under cron / a scheduler.
5.2 Storage
How: append one JSON object per check to data/metrics.json (array) or JSONL; keep a rolling window (e.g. last 24–48h, prune older) so the file stays small. Schema fixed up-front (same discipline as Q1):
{"ts": "...", "up": true, "code": 200, "latency_ms": 120,
"dns_ms": 18, "ssl_days_left": 143, "search_ms": 210}
5.3 Dashboard — dashboard/index.html
How: static page, fetches metrics.json, renders:
- Uptime % (24h):
checks_up / checks_total * 100over last 24h — big number tile. - Latency line chart (last hour): Chart.js line of
latency_ms(optionally overlaysearch_ms). - SSL card: days remaining, color-coded (green > 30, yellow 8–30, red ≤ 7).
- Last-10-checks table: time, status ✅/❌, code, latency, DNS, search.
- Auto-refresh every 60s. Same visual language as Q1's dashboard (Opus keeps them consistent).
5.4 Deploy dashboard to ghaymah — You + Opus
How: Opus writes a tiny Dockerfile that serves the dashboard and runs the collector in the same container (simplest: FastAPI/nginx serving static files + collector as background process writing into the served data/ dir — a small supervisord or a shell entrypoint launching both). You push and deploy exactly as in Q1, screenshot the live URL.
Timeline (72h budget, comfortably front-loaded)
| When | You | Opus (Track A) | Sol (Track B) |
|---|---|---|---|
| Day 1 AM | Repo setup, Ghaymah account + docs research, notes | — | — |
| Day 1 PM | Feed notes + prompts to both agents | Q1 app, Dockerfile, monitor, dashboard | Q2 postmortem full draft |
| Day 2 AM | Build & deploy Q1, registry setup, GitHub secrets | Q5 collector + dashboard | Q3 workflow + CICD.md |
| Day 2 PM | Deploy Q5, run pipeline, capture approval screenshot | Fix anything found in deployment | Q4 scalability doc |
| Day 3 AM | Verify every <!-- VERIFY --> against real Ghaymah docs, insert real URLs/screenshots |
Polish dashboards with real data | Revisions from your review |
| Day 3 PM | Final README index, cross-check rubric (all sub-items × 5 questions), submit | — | — |
Quality Gates (your final pass)
- Every question's numbered sub-requirements are individually answered — graders score per item.
- No invented Ghaymah feature names survive — everything platform-specific is verified against real docs.
- Both apps are live with URLs + screenshots in the README.
- Monitor/collector scripts actually ran long enough to produce real data in the dashboards (start them Day 2 so charts aren't empty at submission).
- Repo README links every deliverable — one-click grading experience.