JB/RUNBOOK.md

5.7 KiB

JobsBoard Platform — Operational Runbook & Incident Response

This document establishes operational incident response protocols, disaster recovery procedures, and infrastructure operational workflows for engineers maintaining the JobsBoard platform.


1. System Topology & Infrastructure States

Component Reality Classification Operational Characteristics
Relational Database ACTIVE AND VERIFIED PostgreSQL 16 (container/managed cloud). Verified 1,425 jobs, 400 companies, 9 users. Dual SQLite schema compatibility maintained for local dev.
Authentication & IAM ACTIVE AND VERIFIED NextAuth JWT with DB session rehydration on every request. Immediate revocation upon account lockout (lockedUntil > now).
Rate Limiting Engine ACTIVE AND VERIFIED Process-local LRU memory limiter (active) + Distributed Redis REST client (DistributedRedisRateLimiter) with risk-aware graceful degradation.
Object Storage IMPLEMENTED BUT NOT CONFIGURED IObjectStorage with secure local filesystem adapter (active) + S3/Cloudflare R2 adapter. Path traversal sanitized.
Background Processing ACTIVE AND VERIFIED Priority-aware queue (CRITICAL, IMPORTANT, BEST_EFFORT) with idempotency deduplication keys and dead-letter handling.
Explainable Matching ACTIVE AND VERIFIED Deterministic canonical skill taxonomy matching with SHA-256 caching and prompt injection sanitization.

2. Emergency Incident Response Workflows

Incident A: Primary Database Outage (PostgreSQL Connection Refusal / Crash)

Symptoms: /api/health?type=readiness returns HTTP 503; error logs report P1001: Can't reach database server.

Immediate Mitigations:

  1. Check container/cluster health:
    podman ps -a | grep jobsboard-postgres
    podman logs --tail 50 jobsboard-postgres
    
  2. Verify network connectivity:
    pg_isready -h localhost -p 5432 -U postgres
    
  3. Restart PostgreSQL instance:
    podman restart jobsboard-postgres
    
  4. Verify readiness recovery:
    curl -s http://localhost:3000/api/health?type=readiness
    # Expected response: {"status":"ready", "dbLatencyMs": <number>}
    

Incident B: Distributed Redis Outage / Network Timeout

Symptoms: Health probe reports Redis as degraded; Upstash REST API returns timeouts or HTTP 5xx.

Expected Automatic System Behavior:

  • The platform does not crash.
  • The DistributedRedisRateLimiter automatically falls back to local memory rate-limiting with 10k key LRU cache.
  • The backgroundQueue automatically buffers jobs in-memory with exponential retry backoff.

Operator Action:

  1. Check Redis REST credentials in .env:
    curl -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" "$UPSTASH_REDIS_REST_URL/ping"
    
  2. If Upstash service has an external outage, no action is needed; local degraded mode will protect the cluster until external connectivity recovers.

Incident C: Corrupted Database State / Disaster Recovery

RPO: Automated daily logical dump + WAL archive.
RTO: Measured < 2 seconds for complete dataset restore.

Restoration Procedure:

  1. Locate latest verified backup:
    ls -la /backups/pg-backup-*.sql
    
  2. Restore logical dump into PostgreSQL target:
    psql -U postgres -h localhost -d jobsboard < "/backups/latest-backup.sql"
    
  3. Run verification test suite:
    npm run test:postgres
    

Incident D: Queue Worker Failure / Dead-Letter Flooding

Symptoms: /api/metrics displays increasing queue.failed counts or dead-letter alerts in structured logs.

Mitigation Protocol:

  1. Query metrics endpoint:
    curl -s http://localhost:3000/api/metrics | jq .queue
    
  2. Inspect dead-letter logs for unhandled exceptions:
    grep -i "QUEUE_DEAD_LETTER" /var/log/jobsboard/app.log
    
  3. Fix root cause in worker handler (src/lib/queue.ts) and trigger reprocessing. Duplicate submissions are automatically ignored due to the built-in idempotencyKey cache.

Incident E: Suspicious Activity / Security Anomaly Detection

Symptoms: High 429 rates, repeated failed logins, suspicious job postings.

Mitigation Steps:

  1. Check audit logs in database:
    SELECT "createdAt", "action", "actorId", "ipAddress", "details"
    FROM "AuditLog"
    ORDER BY "createdAt" DESC
    LIMIT 50;
    
  2. Review flagged job submissions:
    SELECT "id", "title", "company", "moderationStatus", "moderationReason"
    FROM "Job"
    WHERE "moderationStatus" = 'PENDING_REVIEW';
    
  3. Suspend abusive company or user:
    UPDATE "User" SET "lockedUntil" = NOW() + INTERVAL '24 HOURS' WHERE "id" = '<USER_ID>';
    UPDATE "Company" SET "trustStatus" = 'SUSPENDED' WHERE "id" = '<COMPANY_ID>';
    
    Note: Session revocation takes effect immediately on the user's next request.

3. High Availability & Horizontal Scaling Guidelines

When deploying multiple application instances behind an Application Load Balancer (ALB / Cloudflare):

  1. Readiness Probe: Configure ALB health checks to query /api/health?type=readiness. Only instances connected to a functional database receive user traffic.
  2. Stateless App Nodes:
    • NextAuth sessions are stateless JWT tokens re-verified against PostgreSQL.
    • Configure shared Redis (UPSTASH_REDIS_REST_URL) so rate-limiting quotas and background job queues are shared across instances.
    • Configure S3/R2 Cloud Object Storage (STORAGE_ACCESS_KEY & STORAGE_SECRET_KEY) so uploaded resumes are accessible from any node.
  3. Graceful Shutdown: On SIGTERM, allow 10 seconds for running queue jobs to finish before terminating the Node process.