7.7 KiB
7.7 KiB
Q2 — Attack Simulation & Incident Response
Scenario: A successful brute force attack on the API login endpoint led to a data leak.
Part 1: Attack Timeline
T+00:00 — Reconnaissance
- Attacker scans the public IP using tools like
nmapor online port scanners. - Discovers the exposed API endpoint:
POST /api/v1/auth/login - Identifies no rate-limiting or lockout protection via repeated test requests.
T+00:15 — Credential Acquisition
- Attacker uses previously leaked or weak credentials (e.g., from public breach databases or common password lists).
- Builds a wordlist targeting usernames like
admin,deploy, or known email patterns.
T+00:30 — Brute Force Attack Begins
- An automated script (e.g.,
hydra,ffuf, or custom Python) starts sending POST requests to the login endpoint. - Sends hundreds of attempts per minute with different username/password combinations.
- No lockout policy is enforced — the server continues responding normally.
T+01:00 — Successful Authentication
- A valid credential pair is found.
- The attacker receives a valid JWT access token.
- First point of compromise confirmed.
T+01:05 — Session Takeover & Enumeration
- Attacker uses the JWT token to call authenticated API endpoints.
- Discovers
/api/v1/users/exportand/api/v1/reports/downloadvia path probing.
T+01:15 — Data Exfiltration Begins
- Attacker downloads a full CSV dump of user records including: names, emails, hashed passwords, internal metadata.
- No anomaly alert fires on the large data download.
T+01:45 — Exfiltration Complete
- All accessible data is exfiltrated to the attacker's remote server.
- The token remains valid for the full expiry window — no revocation triggered.
T+03:00 — Detection (Delayed)
- An administrator notices unusual API traffic in access logs.
- Investigation reveals the brute-force pattern in
/api/v1/auth/loginlogs.
T+03:30 — Incident Confirmed
- Security team confirms unauthorized access and data exfiltration.
- Incident Response Plan is activated.
Root Cause Summary
| Root Cause | Detail |
|---|---|
| No rate limiting | API accepted unlimited login attempts |
| No account lockout | No lockout after failed attempt threshold |
| Weak or reused credentials | Password matched a known leaked value |
| No anomaly alerting | No alert fired for burst of failed logins |
| Overly broad API authorization | Export endpoint accessible with any valid token |
| No egress monitoring | Large data download went undetected |
Part 2: Incident Response Plan
Classification: P1 — Critical Security Incident
System Affected: API Login Endpoint
Phase 1: Identification (0–15 min)
- Pull API access logs and filter for
POST /api/v1/auth/loginin the last 6 hours. - Identify source IPs with >50 failed attempts followed by a successful login.
- Confirm large API calls to export endpoints post-login.
- Assign an incident commander and open a dedicated communication channel.
Phase 2: Containment (15–60 min)
| Action | Method |
|---|---|
| Block attacker IP(s) | Firewall / WAF deny rule |
| Revoke all active JWT tokens | Rotate JWT signing secret or blacklist tokens |
| Force logout all active sessions | Flush session store (Redis/DB) |
| Disable the export endpoint | Feature flag or reverse proxy block |
| Enable emergency rate limiting | Nginx limit_req_zone or API gateway rule |
| Enable account lockout | Application config: 5 attempts → 15 min lock |
Phase 3: Eradication (1–2 hours)
- Patch the login endpoint with rate limiting, account lockout, and MFA enforcement.
- Audit all API endpoints for over-permissive authorization.
- Check for any persistence (e.g., new admin accounts created during the attack).
- Rotate all application secrets and API keys.
- Require password reset for all users whose data was exported.
Phase 4: Recovery (2–8 hours)
- Re-enable the export endpoint with admin-only authorization.
- Deploy the patched container image.
- Verify monitoring and alerting is active.
- Notify affected users per data breach disclosure requirements.
Phase 5: Lessons Learned (within 72 hours)
- Conduct a post-mortem with all stakeholders.
- Update the threat model and security runbooks.
- Schedule a penetration test to verify fixes.
Part 3: Prevention on the Cloud
The following controls can be implemented in a Ghaymah cloud deployment using its networking, container, and monitoring capabilities.
Network-Level Controls
Rate Limiting via Nginx:
http {
limit_req_zone $binary_remote_addr zone=login_limit:10m rate=5r/m;
server {
location /api/v1/auth/login {
limit_req zone=login_limit burst=10 nodelay;
limit_req_status 429;
proxy_pass http://app_backend;
}
}
}
Web Application Firewall:
- Deploy a WAF in front of the application layer.
- Configure rules to block IPs generating >100 requests/min to
/authendpoints. - Block known malicious user agents (e.g., scanner fingerprints).
IP Restriction for Admin Endpoints:
# Allow only internal network to access sensitive endpoints
iptables -A INPUT -p tcp --dport 8080 -s 10.0.0.0/8 -j ACCEPT
iptables -A INPUT -p tcp --dport 8080 -j DROP
Container Network Policy:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: restrict-login-access
namespace: production
spec:
podSelector:
matchLabels:
app: api-server
ingress:
- from:
- podSelector:
matchLabels:
role: frontend
ports:
- protocol: TCP
port: 8080
policyTypes:
- Ingress
Container Security Controls
Run as non-root user:
FROM python:3.12-slim
RUN useradd --create-home appuser
USER appuser
Read-only filesystem:
services:
api:
read_only: true
tmpfs:
- /tmp
Secrets via environment injection (never baked into images):
services:
api:
env_file:
- .env.production
Image vulnerability scanning in CI/CD:
- name: Scan Docker image
uses: aquasecurity/trivy-action@master
with:
image-ref: 'registry/api:latest'
exit-code: '1'
severity: 'CRITICAL,HIGH'
Application-Level Controls
| Control | Implementation |
|---|---|
| Account lockout | Lock after 5 failed attempts, notify user by email |
| MFA | TOTP required for all admin accounts |
| Password validation | Reject known-weak passwords, enforce minimum 12 chars |
| JWT expiry | Short-lived access tokens (15 min), refresh token rotation |
| Structured auth logging | Log all login attempts to SIEM for real-time analysis |
Part 4: Brute Force Alert Rule
The full deployable Prometheus rule file is in alert-rule.yml. It defines three rules:
Alert Rules Explanation
| Rule | Trigger | Severity | Recommended Action |
|---|---|---|---|
BruteForceLoginAttempt |
>10 failed logins/min from same IP for 1+ min | Critical | Auto-block IP at firewall |
SuspiciousLoginAfterFailures |
20+ failures then 1 success from same IP | Critical | Revoke token, notify security team |
AbnormalDataExport |
>10 MB exported in 5 min by single user | Warning | Flag for human review |
AlertManager Notification Config
# alertmanager.yml
receivers:
- name: secops-slack
slack_configs:
- api_url: ${SLACK_WEBHOOK_URL}
channel: '#security-alerts'
title: '{{ .CommonAnnotations.summary }}'
text: '{{ .CommonAnnotations.description }}'
- name: secops-pagerduty
pagerduty_configs:
- routing_key: ${PAGERDUTY_KEY}
severity: critical
route:
receiver: secops-slack
routes:
- match:
severity: critical
receiver: secops-pagerduty