who_need_help/docs/operations.md

6.6 KiB

Operations runbook

This runbook describes the commands that are implemented and verified in this repository. It does not claim a production recovery point objective, recovery time objective, retention period, storage capacity, or high-availability model; those values require product policy and measurements from the eventual production environment.

Compose database backup

Create a PostgreSQL 18 custom-format archive, validate its table of contents, and write a SHA-256 manifest:

./scripts/backup-compose.sh

The default destination is the ignored output/backups/ directory. An explicit new destination may be supplied as the only argument. The command refuses to overwrite either an archive or its checksum manifest and writes through temporary files before publishing the final pair. Files and newly created directories are restricted by umask 077.

The archive covers the configured application database. PostgreSQL cluster globals such as roles and tablespaces are not part of pg_dump; deployment credentials and database roles must be provisioned separately from secrets. Local backup files on the same workstation are not an off-site backup.

Isolated restore drill

Run a real restore into a uniquely named temporary database:

./scripts/restore-drill-compose.sh output/backups/compose-YYYYMMDD-HHMMSS.dump

The drill:

  1. validates the SHA-256 manifest;
  2. validates the archive table of contents;
  3. creates a pristine database from template0;
  4. restores with pg_restore --exit-on-error;
  5. reads every restored public application table and checks the PostGIS library;
  6. runs the current immutable release's migrations and migration-readiness check against only the temporary database;
  7. drops only the temporary drill database and verifies that it is gone.

A trap also attempts to drop the exact temporary database if a check fails. The source application database is never passed to pg_restore, dropdb, a clean operation, or the drill migration runner. The drill intentionally does not compare an older backup's row counts to the live source, because concurrent legitimate writes or a historical archive would make that comparison invalid.

Service checks

docker compose ps
curl --fail http://localhost:4010/healthz/live
curl --fail http://localhost:4010/healthz/ready
./scripts/verify-realtime-cluster.sh compose

live verifies that the web process can serve HTTP. ready additionally runs SELECT 1 through the configured Ecto repository. The cluster probe subscribes on one connected BEAM node and broadcasts from another.

Health checks do not replace alerting, database backups, restore drills, or application-level synthetic checks.

Local failure and rolling-replacement drills

The isolated load project can exercise process crashes, sequential container replacement, and a real Oban retry without touching the normal Compose project:

./scripts/load-stack-up.sh
./scripts/load-resilience-run.sh local-resilience

The resilience script refuses LOAD_PROJECT=who_need_help and verifies the Compose project/service labels of every container before stopping it. It:

  1. continuously calls readiness through the isolated Traefik route;
  2. terminates the BEAM process in one web and one worker container and requires Docker's observed restart count to increase;
  3. removes and replaces each web and worker replica one at a time;
  4. waits for every configured BEAM node, then runs the cross-node PubSub probe;
  5. enqueues a side-effect-free local worker that fails its first Oban attempt and succeeds on its second;
  6. removes that exact Oban row and requires no fixture domain rows to remain.

Traefik's retry middleware is attached to the HTTP and local TLS routers. Its attempt count is an environment input. Traefik retries transport failures and, with the checked configuration, does not opt in to retrying non-idempotent requests. This reduces a stale-backend window; it is not a claim of production availability.

For the project-owned kind cluster, run:

./scripts/kind-rolling-verify.sh local-kind-rollout

That script requires both the kind ownership marker and the control-plane cluster label before invoking rollout restart. It changes only the web and worker Deployment pod templates. It snapshots application-table counts before and after, continuously probes the observed Docker mapping for the chart's NodePort, requires all four old pod UIDs to disappear, waits for the exact BEAM peer count, and verifies cross-node PubSub. PostGIS, its hostPath, the Kubernetes Secret, and the namespace are not recreated.

The rollout timeout, probe interval/timeout/retry count, and cluster-join timeout are experiment inputs. They are not production SLOs or resource requirements.

Protected Prometheus metrics

The web role exposes Prometheus text format at /metrics. It requires the independent METRICS_TOKEN deployment secret:

curl --fail \
  --header "Authorization: Bearer $METRICS_TOKEN" \
  http://localhost:4010/metrics

The endpoint returns 401 without the exact token, disables response caching, and does not put the credential in a URL. The reporter exports cumulative HTTP request and duration, router exception, database query and duration, WebSocket connection, VM memory, and scheduler run-queue metrics. Cumulative durations are integer microseconds because the selected reporter's sum accumulator is integer-based; divide by 1_000_000 in PromQL when seconds are required. Definitions intentionally have no request path, user, request, or event-name labels that could create unbounded cardinality.

Prometheus itself, durable metrics retention, alert rules, notification destinations, and measured alert thresholds are deployment responsibilities and are not claimed by this repository. In Kubernetes, put the token in existingSecret; configure the external scraper to send it as a Bearer token.

Metrics are local to each BEAM process. Discover and scrape every web pod or container as a distinct target and preserve Prometheus's instance label. A request through the load-balanced public route reaches only one replica and is therefore useful as an authorization/smoke check, not as a cluster-wide aggregate.

Rollback boundary

The release image is immutable and migrations run as a separate one-shot role. Before a schema rollout, create and restore-test a current backup. Application rollback and database migration rollback are separate decisions: do not run an Ecto down migration merely because an image is rolled back. Inspect the exact migration and compatibility boundary first.

The repository intentionally does not ship an automatic destructive production restore command.