135 lines
5.7 KiB
Markdown
135 lines
5.7 KiB
Markdown
# JobsBoard Platform — Operational Runbook & Incident Response
|
|
|
|
This document establishes operational incident response protocols, disaster recovery procedures, and infrastructure operational workflows for engineers maintaining the JobsBoard platform.
|
|
|
|
---
|
|
|
|
## 1. System Topology & Infrastructure States
|
|
|
|
| Component | Reality Classification | Operational Characteristics |
|
|
| :--- | :--- | :--- |
|
|
| **Relational Database** | **ACTIVE AND VERIFIED** | **PostgreSQL 16** (container/managed cloud). Verified 1,425 jobs, 400 companies, 9 users. Dual SQLite schema compatibility maintained for local dev. |
|
|
| **Authentication & IAM** | **ACTIVE AND VERIFIED** | NextAuth JWT with DB session rehydration on every request. Immediate revocation upon account lockout (`lockedUntil > now`). |
|
|
| **Rate Limiting Engine** | **ACTIVE AND VERIFIED** | Process-local LRU memory limiter (active) + Distributed Redis REST client (`DistributedRedisRateLimiter`) with risk-aware graceful degradation. |
|
|
| **Object Storage** | **IMPLEMENTED BUT NOT CONFIGURED** | `IObjectStorage` with secure local filesystem adapter (active) + S3/Cloudflare R2 adapter. Path traversal sanitized. |
|
|
| **Background Processing**| **ACTIVE AND VERIFIED** | Priority-aware queue (`CRITICAL`, `IMPORTANT`, `BEST_EFFORT`) with idempotency deduplication keys and dead-letter handling. |
|
|
| **Explainable Matching**| **ACTIVE AND VERIFIED** | Deterministic canonical skill taxonomy matching with SHA-256 caching and prompt injection sanitization. |
|
|
|
|
---
|
|
|
|
## 2. Emergency Incident Response Workflows
|
|
|
|
### Incident A: Primary Database Outage (PostgreSQL Connection Refusal / Crash)
|
|
**Symptoms**: `/api/health?type=readiness` returns HTTP 503; error logs report `P1001: Can't reach database server`.
|
|
|
|
**Immediate Mitigations**:
|
|
1. Check container/cluster health:
|
|
```bash
|
|
podman ps -a | grep jobsboard-postgres
|
|
podman logs --tail 50 jobsboard-postgres
|
|
```
|
|
2. Verify network connectivity:
|
|
```bash
|
|
pg_isready -h localhost -p 5432 -U postgres
|
|
```
|
|
3. Restart PostgreSQL instance:
|
|
```bash
|
|
podman restart jobsboard-postgres
|
|
```
|
|
4. Verify readiness recovery:
|
|
```bash
|
|
curl -s http://localhost:3000/api/health?type=readiness
|
|
# Expected response: {"status":"ready", "dbLatencyMs": <number>}
|
|
```
|
|
|
|
---
|
|
|
|
### Incident B: Distributed Redis Outage / Network Timeout
|
|
**Symptoms**: Health probe reports Redis as degraded; Upstash REST API returns timeouts or HTTP 5xx.
|
|
|
|
**Expected Automatic System Behavior**:
|
|
- The platform **does not crash**.
|
|
- The `DistributedRedisRateLimiter` automatically falls back to local memory rate-limiting with 10k key LRU cache.
|
|
- The `backgroundQueue` automatically buffers jobs in-memory with exponential retry backoff.
|
|
|
|
**Operator Action**:
|
|
1. Check Redis REST credentials in `.env`:
|
|
```bash
|
|
curl -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" "$UPSTASH_REDIS_REST_URL/ping"
|
|
```
|
|
2. If Upstash service has an external outage, no action is needed; local degraded mode will protect the cluster until external connectivity recovers.
|
|
|
|
---
|
|
|
|
### Incident C: Corrupted Database State / Disaster Recovery
|
|
**RPO**: Automated daily logical dump + WAL archive.
|
|
**RTO**: Measured < 2 seconds for complete dataset restore.
|
|
|
|
**Restoration Procedure**:
|
|
1. Locate latest verified backup:
|
|
```bash
|
|
ls -la /backups/pg-backup-*.sql
|
|
```
|
|
2. Restore logical dump into PostgreSQL target:
|
|
```bash
|
|
psql -U postgres -h localhost -d jobsboard < "/backups/latest-backup.sql"
|
|
```
|
|
3. Run verification test suite:
|
|
```bash
|
|
npm run test:postgres
|
|
```
|
|
|
|
---
|
|
|
|
### Incident D: Queue Worker Failure / Dead-Letter Flooding
|
|
**Symptoms**: `/api/metrics` displays increasing `queue.failed` counts or dead-letter alerts in structured logs.
|
|
|
|
**Mitigation Protocol**:
|
|
1. Query metrics endpoint:
|
|
```bash
|
|
curl -s http://localhost:3000/api/metrics | jq .queue
|
|
```
|
|
2. Inspect dead-letter logs for unhandled exceptions:
|
|
```bash
|
|
grep -i "QUEUE_DEAD_LETTER" /var/log/jobsboard/app.log
|
|
```
|
|
3. Fix root cause in worker handler (`src/lib/queue.ts`) and trigger reprocessing. Duplicate submissions are automatically ignored due to the built-in `idempotencyKey` cache.
|
|
|
|
---
|
|
|
|
### Incident E: Suspicious Activity / Security Anomaly Detection
|
|
**Symptoms**: High 429 rates, repeated failed logins, suspicious job postings.
|
|
|
|
**Mitigation Steps**:
|
|
1. Check audit logs in database:
|
|
```sql
|
|
SELECT "createdAt", "action", "actorId", "ipAddress", "details"
|
|
FROM "AuditLog"
|
|
ORDER BY "createdAt" DESC
|
|
LIMIT 50;
|
|
```
|
|
2. Review flagged job submissions:
|
|
```sql
|
|
SELECT "id", "title", "company", "moderationStatus", "moderationReason"
|
|
FROM "Job"
|
|
WHERE "moderationStatus" = 'PENDING_REVIEW';
|
|
```
|
|
3. Suspend abusive company or user:
|
|
```sql
|
|
UPDATE "User" SET "lockedUntil" = NOW() + INTERVAL '24 HOURS' WHERE "id" = '<USER_ID>';
|
|
UPDATE "Company" SET "trustStatus" = 'SUSPENDED' WHERE "id" = '<COMPANY_ID>';
|
|
```
|
|
*Note: Session revocation takes effect immediately on the user's next request.*
|
|
|
|
---
|
|
|
|
## 3. High Availability & Horizontal Scaling Guidelines
|
|
|
|
When deploying multiple application instances behind an Application Load Balancer (ALB / Cloudflare):
|
|
|
|
1. **Readiness Probe**: Configure ALB health checks to query `/api/health?type=readiness`. Only instances connected to a functional database receive user traffic.
|
|
2. **Stateless App Nodes**:
|
|
- NextAuth sessions are stateless JWT tokens re-verified against PostgreSQL.
|
|
- Configure shared Redis (`UPSTASH_REDIS_REST_URL`) so rate-limiting quotas and background job queues are shared across instances.
|
|
- Configure S3/R2 Cloud Object Storage (`STORAGE_ACCESS_KEY` & `STORAGE_SECRET_KEY`) so uploaded resumes are accessible from any node.
|
|
3. **Graceful Shutdown**: On `SIGTERM`, allow 10 seconds for running queue jobs to finish before terminating the Node process.
|