JB/RUNBOOK.md

135 lines
5.7 KiB
Markdown

# JobsBoard Platform — Operational Runbook & Incident Response
This document establishes operational incident response protocols, disaster recovery procedures, and infrastructure operational workflows for engineers maintaining the JobsBoard platform.
---
## 1. System Topology & Infrastructure States
| Component | Reality Classification | Operational Characteristics |
| :--- | :--- | :--- |
| **Relational Database** | **ACTIVE AND VERIFIED** | **PostgreSQL 16** (container/managed cloud). Verified 1,425 jobs, 400 companies, 9 users. Dual SQLite schema compatibility maintained for local dev. |
| **Authentication & IAM** | **ACTIVE AND VERIFIED** | NextAuth JWT with DB session rehydration on every request. Immediate revocation upon account lockout (`lockedUntil > now`). |
| **Rate Limiting Engine** | **ACTIVE AND VERIFIED** | Process-local LRU memory limiter (active) + Distributed Redis REST client (`DistributedRedisRateLimiter`) with risk-aware graceful degradation. |
| **Object Storage** | **IMPLEMENTED BUT NOT CONFIGURED** | `IObjectStorage` with secure local filesystem adapter (active) + S3/Cloudflare R2 adapter. Path traversal sanitized. |
| **Background Processing**| **ACTIVE AND VERIFIED** | Priority-aware queue (`CRITICAL`, `IMPORTANT`, `BEST_EFFORT`) with idempotency deduplication keys and dead-letter handling. |
| **Explainable Matching**| **ACTIVE AND VERIFIED** | Deterministic canonical skill taxonomy matching with SHA-256 caching and prompt injection sanitization. |
---
## 2. Emergency Incident Response Workflows
### Incident A: Primary Database Outage (PostgreSQL Connection Refusal / Crash)
**Symptoms**: `/api/health?type=readiness` returns HTTP 503; error logs report `P1001: Can't reach database server`.
**Immediate Mitigations**:
1. Check container/cluster health:
```bash
podman ps -a | grep jobsboard-postgres
podman logs --tail 50 jobsboard-postgres
```
2. Verify network connectivity:
```bash
pg_isready -h localhost -p 5432 -U postgres
```
3. Restart PostgreSQL instance:
```bash
podman restart jobsboard-postgres
```
4. Verify readiness recovery:
```bash
curl -s http://localhost:3000/api/health?type=readiness
# Expected response: {"status":"ready", "dbLatencyMs": <number>}
```
---
### Incident B: Distributed Redis Outage / Network Timeout
**Symptoms**: Health probe reports Redis as degraded; Upstash REST API returns timeouts or HTTP 5xx.
**Expected Automatic System Behavior**:
- The platform **does not crash**.
- The `DistributedRedisRateLimiter` automatically falls back to local memory rate-limiting with 10k key LRU cache.
- The `backgroundQueue` automatically buffers jobs in-memory with exponential retry backoff.
**Operator Action**:
1. Check Redis REST credentials in `.env`:
```bash
curl -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" "$UPSTASH_REDIS_REST_URL/ping"
```
2. If Upstash service has an external outage, no action is needed; local degraded mode will protect the cluster until external connectivity recovers.
---
### Incident C: Corrupted Database State / Disaster Recovery
**RPO**: Automated daily logical dump + WAL archive.
**RTO**: Measured < 2 seconds for complete dataset restore.
**Restoration Procedure**:
1. Locate latest verified backup:
```bash
ls -la /backups/pg-backup-*.sql
```
2. Restore logical dump into PostgreSQL target:
```bash
psql -U postgres -h localhost -d jobsboard < "/backups/latest-backup.sql"
```
3. Run verification test suite:
```bash
npm run test:postgres
```
---
### Incident D: Queue Worker Failure / Dead-Letter Flooding
**Symptoms**: `/api/metrics` displays increasing `queue.failed` counts or dead-letter alerts in structured logs.
**Mitigation Protocol**:
1. Query metrics endpoint:
```bash
curl -s http://localhost:3000/api/metrics | jq .queue
```
2. Inspect dead-letter logs for unhandled exceptions:
```bash
grep -i "QUEUE_DEAD_LETTER" /var/log/jobsboard/app.log
```
3. Fix root cause in worker handler (`src/lib/queue.ts`) and trigger reprocessing. Duplicate submissions are automatically ignored due to the built-in `idempotencyKey` cache.
---
### Incident E: Suspicious Activity / Security Anomaly Detection
**Symptoms**: High 429 rates, repeated failed logins, suspicious job postings.
**Mitigation Steps**:
1. Check audit logs in database:
```sql
SELECT "createdAt", "action", "actorId", "ipAddress", "details"
FROM "AuditLog"
ORDER BY "createdAt" DESC
LIMIT 50;
```
2. Review flagged job submissions:
```sql
SELECT "id", "title", "company", "moderationStatus", "moderationReason"
FROM "Job"
WHERE "moderationStatus" = 'PENDING_REVIEW';
```
3. Suspend abusive company or user:
```sql
UPDATE "User" SET "lockedUntil" = NOW() + INTERVAL '24 HOURS' WHERE "id" = '<USER_ID>';
UPDATE "Company" SET "trustStatus" = 'SUSPENDED' WHERE "id" = '<COMPANY_ID>';
```
*Note: Session revocation takes effect immediately on the user's next request.*
---
## 3. High Availability & Horizontal Scaling Guidelines
When deploying multiple application instances behind an Application Load Balancer (ALB / Cloudflare):
1. **Readiness Probe**: Configure ALB health checks to query `/api/health?type=readiness`. Only instances connected to a functional database receive user traffic.
2. **Stateless App Nodes**:
- NextAuth sessions are stateless JWT tokens re-verified against PostgreSQL.
- Configure shared Redis (`UPSTASH_REDIS_REST_URL`) so rate-limiting quotas and background job queues are shared across instances.
- Configure S3/R2 Cloud Object Storage (`STORAGE_ACCESS_KEY` & `STORAGE_SECRET_KEY`) so uploaded resumes are accessible from any node.
3. **Graceful Shutdown**: On `SIGTERM`, allow 10 seconds for running queue jobs to finish before terminating the Node process.