5.7 KiB
JobsBoard Platform — Operational Runbook & Incident Response
This document establishes operational incident response protocols, disaster recovery procedures, and infrastructure operational workflows for engineers maintaining the JobsBoard platform.
1. System Topology & Infrastructure States
| Component | Reality Classification | Operational Characteristics |
|---|---|---|
| Relational Database | ACTIVE AND VERIFIED | PostgreSQL 16 (container/managed cloud). Verified 1,425 jobs, 400 companies, 9 users. Dual SQLite schema compatibility maintained for local dev. |
| Authentication & IAM | ACTIVE AND VERIFIED | NextAuth JWT with DB session rehydration on every request. Immediate revocation upon account lockout (lockedUntil > now). |
| Rate Limiting Engine | ACTIVE AND VERIFIED | Process-local LRU memory limiter (active) + Distributed Redis REST client (DistributedRedisRateLimiter) with risk-aware graceful degradation. |
| Object Storage | IMPLEMENTED BUT NOT CONFIGURED | IObjectStorage with secure local filesystem adapter (active) + S3/Cloudflare R2 adapter. Path traversal sanitized. |
| Background Processing | ACTIVE AND VERIFIED | Priority-aware queue (CRITICAL, IMPORTANT, BEST_EFFORT) with idempotency deduplication keys and dead-letter handling. |
| Explainable Matching | ACTIVE AND VERIFIED | Deterministic canonical skill taxonomy matching with SHA-256 caching and prompt injection sanitization. |
2. Emergency Incident Response Workflows
Incident A: Primary Database Outage (PostgreSQL Connection Refusal / Crash)
Symptoms: /api/health?type=readiness returns HTTP 503; error logs report P1001: Can't reach database server.
Immediate Mitigations:
- Check container/cluster health:
podman ps -a | grep jobsboard-postgres podman logs --tail 50 jobsboard-postgres - Verify network connectivity:
pg_isready -h localhost -p 5432 -U postgres - Restart PostgreSQL instance:
podman restart jobsboard-postgres - Verify readiness recovery:
curl -s http://localhost:3000/api/health?type=readiness # Expected response: {"status":"ready", "dbLatencyMs": <number>}
Incident B: Distributed Redis Outage / Network Timeout
Symptoms: Health probe reports Redis as degraded; Upstash REST API returns timeouts or HTTP 5xx.
Expected Automatic System Behavior:
- The platform does not crash.
- The
DistributedRedisRateLimiterautomatically falls back to local memory rate-limiting with 10k key LRU cache. - The
backgroundQueueautomatically buffers jobs in-memory with exponential retry backoff.
Operator Action:
- Check Redis REST credentials in
.env:curl -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" "$UPSTASH_REDIS_REST_URL/ping" - If Upstash service has an external outage, no action is needed; local degraded mode will protect the cluster until external connectivity recovers.
Incident C: Corrupted Database State / Disaster Recovery
RPO: Automated daily logical dump + WAL archive.
RTO: Measured < 2 seconds for complete dataset restore.
Restoration Procedure:
- Locate latest verified backup:
ls -la /backups/pg-backup-*.sql - Restore logical dump into PostgreSQL target:
psql -U postgres -h localhost -d jobsboard < "/backups/latest-backup.sql" - Run verification test suite:
npm run test:postgres
Incident D: Queue Worker Failure / Dead-Letter Flooding
Symptoms: /api/metrics displays increasing queue.failed counts or dead-letter alerts in structured logs.
Mitigation Protocol:
- Query metrics endpoint:
curl -s http://localhost:3000/api/metrics | jq .queue - Inspect dead-letter logs for unhandled exceptions:
grep -i "QUEUE_DEAD_LETTER" /var/log/jobsboard/app.log - Fix root cause in worker handler (
src/lib/queue.ts) and trigger reprocessing. Duplicate submissions are automatically ignored due to the built-inidempotencyKeycache.
Incident E: Suspicious Activity / Security Anomaly Detection
Symptoms: High 429 rates, repeated failed logins, suspicious job postings.
Mitigation Steps:
- Check audit logs in database:
SELECT "createdAt", "action", "actorId", "ipAddress", "details" FROM "AuditLog" ORDER BY "createdAt" DESC LIMIT 50; - Review flagged job submissions:
SELECT "id", "title", "company", "moderationStatus", "moderationReason" FROM "Job" WHERE "moderationStatus" = 'PENDING_REVIEW'; - Suspend abusive company or user:
Note: Session revocation takes effect immediately on the user's next request.UPDATE "User" SET "lockedUntil" = NOW() + INTERVAL '24 HOURS' WHERE "id" = '<USER_ID>'; UPDATE "Company" SET "trustStatus" = 'SUSPENDED' WHERE "id" = '<COMPANY_ID>';
3. High Availability & Horizontal Scaling Guidelines
When deploying multiple application instances behind an Application Load Balancer (ALB / Cloudflare):
- Readiness Probe: Configure ALB health checks to query
/api/health?type=readiness. Only instances connected to a functional database receive user traffic. - Stateless App Nodes:
- NextAuth sessions are stateless JWT tokens re-verified against PostgreSQL.
- Configure shared Redis (
UPSTASH_REDIS_REST_URL) so rate-limiting quotas and background job queues are shared across instances. - Configure S3/R2 Cloud Object Storage (
STORAGE_ACCESS_KEY&STORAGE_SECRET_KEY) so uploaded resumes are accessible from any node.
- Graceful Shutdown: On
SIGTERM, allow 10 seconds for running queue jobs to finish before terminating the Node process.