added
هذا الالتزام موجود في:
248
q2-attack-simulation/incident-response.md
Normal file
248
q2-attack-simulation/incident-response.md
Normal file
@@ -0,0 +1,248 @@
|
||||
# Q2 — Attack Simulation & Incident Response
|
||||
|
||||
**Scenario:** A successful brute force attack on the API login endpoint led to a data leak.
|
||||
|
||||
---
|
||||
|
||||
## Part 1: Attack Timeline
|
||||
|
||||
### T+00:00 — Reconnaissance
|
||||
- Attacker scans the public IP using tools like `nmap` or online port scanners.
|
||||
- Discovers the exposed API endpoint: `POST /api/v1/auth/login`
|
||||
- Identifies no rate-limiting or lockout protection via repeated test requests.
|
||||
|
||||
### T+00:15 — Credential Acquisition
|
||||
- Attacker uses previously leaked or weak credentials (e.g., from public breach databases or common password lists).
|
||||
- Builds a wordlist targeting usernames like `admin`, `deploy`, or known email patterns.
|
||||
|
||||
### T+00:30 — Brute Force Attack Begins
|
||||
- An automated script (e.g., `hydra`, `ffuf`, or custom Python) starts sending POST requests to the login endpoint.
|
||||
- Sends hundreds of attempts per minute with different username/password combinations.
|
||||
- No lockout policy is enforced — the server continues responding normally.
|
||||
|
||||
### T+01:00 — Successful Authentication
|
||||
- A valid credential pair is found.
|
||||
- The attacker receives a valid JWT access token.
|
||||
- **First point of compromise confirmed.**
|
||||
|
||||
### T+01:05 — Session Takeover & Enumeration
|
||||
- Attacker uses the JWT token to call authenticated API endpoints.
|
||||
- Discovers `/api/v1/users/export` and `/api/v1/reports/download` via path probing.
|
||||
|
||||
### T+01:15 — Data Exfiltration Begins
|
||||
- Attacker downloads a full CSV dump of user records including: names, emails, hashed passwords, internal metadata.
|
||||
- No anomaly alert fires on the large data download.
|
||||
|
||||
### T+01:45 — Exfiltration Complete
|
||||
- All accessible data is exfiltrated to the attacker's remote server.
|
||||
- The token remains valid for the full expiry window — no revocation triggered.
|
||||
|
||||
### T+03:00 — Detection (Delayed)
|
||||
- An administrator notices unusual API traffic in access logs.
|
||||
- Investigation reveals the brute-force pattern in `/api/v1/auth/login` logs.
|
||||
|
||||
### T+03:30 — Incident Confirmed
|
||||
- Security team confirms unauthorized access and data exfiltration.
|
||||
- **Incident Response Plan is activated.**
|
||||
|
||||
---
|
||||
|
||||
### Root Cause Summary
|
||||
|
||||
| Root Cause | Detail |
|
||||
|---|---|
|
||||
| No rate limiting | API accepted unlimited login attempts |
|
||||
| No account lockout | No lockout after failed attempt threshold |
|
||||
| Weak or reused credentials | Password matched a known leaked value |
|
||||
| No anomaly alerting | No alert fired for burst of failed logins |
|
||||
| Overly broad API authorization | Export endpoint accessible with any valid token |
|
||||
| No egress monitoring | Large data download went undetected |
|
||||
|
||||
---
|
||||
|
||||
## Part 2: Incident Response Plan
|
||||
|
||||
**Classification:** P1 — Critical Security Incident
|
||||
**System Affected:** API Login Endpoint
|
||||
|
||||
### Phase 1: Identification (0–15 min)
|
||||
|
||||
1. Pull API access logs and filter for `POST /api/v1/auth/login` in the last 6 hours.
|
||||
2. Identify source IPs with >50 failed attempts followed by a successful login.
|
||||
3. Confirm large API calls to export endpoints post-login.
|
||||
4. Assign an incident commander and open a dedicated communication channel.
|
||||
|
||||
### Phase 2: Containment (15–60 min)
|
||||
|
||||
| Action | Method |
|
||||
|---|---|
|
||||
| Block attacker IP(s) | Firewall / WAF deny rule |
|
||||
| Revoke all active JWT tokens | Rotate JWT signing secret or blacklist tokens |
|
||||
| Force logout all active sessions | Flush session store (Redis/DB) |
|
||||
| Disable the export endpoint | Feature flag or reverse proxy block |
|
||||
| Enable emergency rate limiting | Nginx `limit_req_zone` or API gateway rule |
|
||||
| Enable account lockout | Application config: 5 attempts → 15 min lock |
|
||||
|
||||
### Phase 3: Eradication (1–2 hours)
|
||||
|
||||
1. Patch the login endpoint with rate limiting, account lockout, and MFA enforcement.
|
||||
2. Audit all API endpoints for over-permissive authorization.
|
||||
3. Check for any persistence (e.g., new admin accounts created during the attack).
|
||||
4. Rotate all application secrets and API keys.
|
||||
5. Require password reset for all users whose data was exported.
|
||||
|
||||
### Phase 4: Recovery (2–8 hours)
|
||||
|
||||
1. Re-enable the export endpoint with admin-only authorization.
|
||||
2. Deploy the patched container image.
|
||||
3. Verify monitoring and alerting is active.
|
||||
4. Notify affected users per data breach disclosure requirements.
|
||||
|
||||
### Phase 5: Lessons Learned (within 72 hours)
|
||||
|
||||
1. Conduct a post-mortem with all stakeholders.
|
||||
2. Update the threat model and security runbooks.
|
||||
3. Schedule a penetration test to verify fixes.
|
||||
|
||||
---
|
||||
|
||||
## Part 3: Prevention on the Cloud
|
||||
|
||||
> The following controls can be implemented in a Ghaymah cloud deployment using its networking, container, and monitoring capabilities.
|
||||
|
||||
### Network-Level Controls
|
||||
|
||||
**Rate Limiting via Nginx:**
|
||||
```nginx
|
||||
http {
|
||||
limit_req_zone $binary_remote_addr zone=login_limit:10m rate=5r/m;
|
||||
|
||||
server {
|
||||
location /api/v1/auth/login {
|
||||
limit_req zone=login_limit burst=10 nodelay;
|
||||
limit_req_status 429;
|
||||
proxy_pass http://app_backend;
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Web Application Firewall:**
|
||||
- Deploy a WAF in front of the application layer.
|
||||
- Configure rules to block IPs generating >100 requests/min to `/auth` endpoints.
|
||||
- Block known malicious user agents (e.g., scanner fingerprints).
|
||||
|
||||
**IP Restriction for Admin Endpoints:**
|
||||
```bash
|
||||
# Allow only internal network to access sensitive endpoints
|
||||
iptables -A INPUT -p tcp --dport 8080 -s 10.0.0.0/8 -j ACCEPT
|
||||
iptables -A INPUT -p tcp --dport 8080 -j DROP
|
||||
```
|
||||
|
||||
**Container Network Policy:**
|
||||
```yaml
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: restrict-login-access
|
||||
namespace: production
|
||||
spec:
|
||||
podSelector:
|
||||
matchLabels:
|
||||
app: api-server
|
||||
ingress:
|
||||
- from:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
role: frontend
|
||||
ports:
|
||||
- protocol: TCP
|
||||
port: 8080
|
||||
policyTypes:
|
||||
- Ingress
|
||||
```
|
||||
|
||||
### Container Security Controls
|
||||
|
||||
**Run as non-root user:**
|
||||
```dockerfile
|
||||
FROM python:3.12-slim
|
||||
RUN useradd --create-home appuser
|
||||
USER appuser
|
||||
```
|
||||
|
||||
**Read-only filesystem:**
|
||||
```yaml
|
||||
services:
|
||||
api:
|
||||
read_only: true
|
||||
tmpfs:
|
||||
- /tmp
|
||||
```
|
||||
|
||||
**Secrets via environment injection (never baked into images):**
|
||||
```yaml
|
||||
services:
|
||||
api:
|
||||
env_file:
|
||||
- .env.production
|
||||
```
|
||||
|
||||
**Image vulnerability scanning in CI/CD:**
|
||||
```yaml
|
||||
- name: Scan Docker image
|
||||
uses: aquasecurity/trivy-action@master
|
||||
with:
|
||||
image-ref: 'registry/api:latest'
|
||||
exit-code: '1'
|
||||
severity: 'CRITICAL,HIGH'
|
||||
```
|
||||
|
||||
### Application-Level Controls
|
||||
|
||||
| Control | Implementation |
|
||||
|---|---|
|
||||
| Account lockout | Lock after 5 failed attempts, notify user by email |
|
||||
| MFA | TOTP required for all admin accounts |
|
||||
| Password validation | Reject known-weak passwords, enforce minimum 12 chars |
|
||||
| JWT expiry | Short-lived access tokens (15 min), refresh token rotation |
|
||||
| Structured auth logging | Log all login attempts to SIEM for real-time analysis |
|
||||
|
||||
---
|
||||
|
||||
## Part 4: Brute Force Alert Rule
|
||||
|
||||
The full deployable Prometheus rule file is in [`alert-rule.yml`](./alert-rule.yml). It defines three rules:
|
||||
|
||||
### Alert Rules Explanation
|
||||
|
||||
| Rule | Trigger | Severity | Recommended Action |
|
||||
|---|---|---|---|
|
||||
| `BruteForceLoginAttempt` | >10 failed logins/min from same IP for 1+ min | Critical | Auto-block IP at firewall |
|
||||
| `SuspiciousLoginAfterFailures` | 20+ failures then 1 success from same IP | Critical | Revoke token, notify security team |
|
||||
| `AbnormalDataExport` | >10 MB exported in 5 min by single user | Warning | Flag for human review |
|
||||
|
||||
### AlertManager Notification Config
|
||||
|
||||
```yaml
|
||||
# alertmanager.yml
|
||||
receivers:
|
||||
- name: secops-slack
|
||||
slack_configs:
|
||||
- api_url: ${SLACK_WEBHOOK_URL}
|
||||
channel: '#security-alerts'
|
||||
title: '{{ .CommonAnnotations.summary }}'
|
||||
text: '{{ .CommonAnnotations.description }}'
|
||||
|
||||
- name: secops-pagerduty
|
||||
pagerduty_configs:
|
||||
- routing_key: ${PAGERDUTY_KEY}
|
||||
severity: critical
|
||||
|
||||
route:
|
||||
receiver: secops-slack
|
||||
routes:
|
||||
- match:
|
||||
severity: critical
|
||||
receiver: secops-pagerduty
|
||||
```
|
||||
المرجع في مشكلة جديدة
حظر مستخدم