# JobsBoard Platform — Operational Runbook & Incident Response This document establishes operational incident response protocols, disaster recovery procedures, and infrastructure operational workflows for engineers maintaining the JobsBoard platform. --- ## 1. System Topology & Infrastructure States | Component | Reality Classification | Operational Characteristics | | :--- | :--- | :--- | | **Relational Database** | **ACTIVE AND VERIFIED** | **PostgreSQL 16** (container/managed cloud). Verified 1,425 jobs, 400 companies, 9 users. Dual SQLite schema compatibility maintained for local dev. | | **Authentication & IAM** | **ACTIVE AND VERIFIED** | NextAuth JWT with DB session rehydration on every request. Immediate revocation upon account lockout (`lockedUntil > now`). | | **Rate Limiting Engine** | **ACTIVE AND VERIFIED** | Process-local LRU memory limiter (active) + Distributed Redis REST client (`DistributedRedisRateLimiter`) with risk-aware graceful degradation. | | **Object Storage** | **IMPLEMENTED BUT NOT CONFIGURED** | `IObjectStorage` with secure local filesystem adapter (active) + S3/Cloudflare R2 adapter. Path traversal sanitized. | | **Background Processing**| **ACTIVE AND VERIFIED** | Priority-aware queue (`CRITICAL`, `IMPORTANT`, `BEST_EFFORT`) with idempotency deduplication keys and dead-letter handling. | | **Explainable Matching**| **ACTIVE AND VERIFIED** | Deterministic canonical skill taxonomy matching with SHA-256 caching and prompt injection sanitization. | --- ## 2. Emergency Incident Response Workflows ### Incident A: Primary Database Outage (PostgreSQL Connection Refusal / Crash) **Symptoms**: `/api/health?type=readiness` returns HTTP 503; error logs report `P1001: Can't reach database server`. **Immediate Mitigations**: 1. Check container/cluster health: ```bash podman ps -a | grep jobsboard-postgres podman logs --tail 50 jobsboard-postgres ``` 2. Verify network connectivity: ```bash pg_isready -h localhost -p 5432 -U postgres ``` 3. Restart PostgreSQL instance: ```bash podman restart jobsboard-postgres ``` 4. Verify readiness recovery: ```bash curl -s http://localhost:3000/api/health?type=readiness # Expected response: {"status":"ready", "dbLatencyMs": } ``` --- ### Incident B: Distributed Redis Outage / Network Timeout **Symptoms**: Health probe reports Redis as degraded; Upstash REST API returns timeouts or HTTP 5xx. **Expected Automatic System Behavior**: - The platform **does not crash**. - The `DistributedRedisRateLimiter` automatically falls back to local memory rate-limiting with 10k key LRU cache. - The `backgroundQueue` automatically buffers jobs in-memory with exponential retry backoff. **Operator Action**: 1. Check Redis REST credentials in `.env`: ```bash curl -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" "$UPSTASH_REDIS_REST_URL/ping" ``` 2. If Upstash service has an external outage, no action is needed; local degraded mode will protect the cluster until external connectivity recovers. --- ### Incident C: Corrupted Database State / Disaster Recovery **RPO**: Automated daily logical dump + WAL archive. **RTO**: Measured < 2 seconds for complete dataset restore. **Restoration Procedure**: 1. Locate latest verified backup: ```bash ls -la /backups/pg-backup-*.sql ``` 2. Restore logical dump into PostgreSQL target: ```bash psql -U postgres -h localhost -d jobsboard < "/backups/latest-backup.sql" ``` 3. Run verification test suite: ```bash npm run test:postgres ``` --- ### Incident D: Queue Worker Failure / Dead-Letter Flooding **Symptoms**: `/api/metrics` displays increasing `queue.failed` counts or dead-letter alerts in structured logs. **Mitigation Protocol**: 1. Query metrics endpoint: ```bash curl -s http://localhost:3000/api/metrics | jq .queue ``` 2. Inspect dead-letter logs for unhandled exceptions: ```bash grep -i "QUEUE_DEAD_LETTER" /var/log/jobsboard/app.log ``` 3. Fix root cause in worker handler (`src/lib/queue.ts`) and trigger reprocessing. Duplicate submissions are automatically ignored due to the built-in `idempotencyKey` cache. --- ### Incident E: Suspicious Activity / Security Anomaly Detection **Symptoms**: High 429 rates, repeated failed logins, suspicious job postings. **Mitigation Steps**: 1. Check audit logs in database: ```sql SELECT "createdAt", "action", "actorId", "ipAddress", "details" FROM "AuditLog" ORDER BY "createdAt" DESC LIMIT 50; ``` 2. Review flagged job submissions: ```sql SELECT "id", "title", "company", "moderationStatus", "moderationReason" FROM "Job" WHERE "moderationStatus" = 'PENDING_REVIEW'; ``` 3. Suspend abusive company or user: ```sql UPDATE "User" SET "lockedUntil" = NOW() + INTERVAL '24 HOURS' WHERE "id" = ''; UPDATE "Company" SET "trustStatus" = 'SUSPENDED' WHERE "id" = ''; ``` *Note: Session revocation takes effect immediately on the user's next request.* --- ## 3. High Availability & Horizontal Scaling Guidelines When deploying multiple application instances behind an Application Load Balancer (ALB / Cloudflare): 1. **Readiness Probe**: Configure ALB health checks to query `/api/health?type=readiness`. Only instances connected to a functional database receive user traffic. 2. **Stateless App Nodes**: - NextAuth sessions are stateless JWT tokens re-verified against PostgreSQL. - Configure shared Redis (`UPSTASH_REDIS_REST_URL`) so rate-limiting quotas and background job queues are shared across instances. - Configure S3/R2 Cloud Object Storage (`STORAGE_ACCESS_KEY` & `STORAGE_SECRET_KEY`) so uploaded resumes are accessible from any node. 3. **Graceful Shutdown**: On `SIGTERM`, allow 10 seconds for running queue jobs to finish before terminating the Node process.