who_need_help/docs/operations.md

194 lines
8.2 KiB
Markdown

# Operations runbook
This runbook describes the commands that are implemented and verified in this
repository. It does not claim a production recovery point objective, recovery
time objective, retention period, storage capacity, or high-availability model;
those values require product policy and measurements from the eventual
production environment.
## Compose database backup
Create a PostgreSQL 18 custom-format archive, validate its table of contents,
and write a SHA-256 manifest:
```bash
./scripts/backup-compose.sh
```
The default destination is the ignored `output/backups/` directory. An explicit
new destination may be supplied as the only argument. The command refuses to
overwrite either an archive or its checksum manifest and writes through
temporary files before publishing the final pair. Files and newly created
directories are restricted by `umask 077`.
The archive covers the configured application database. PostgreSQL cluster
globals such as roles and tablespaces are not part of `pg_dump`; deployment
credentials and database roles must be provisioned separately from secrets.
Local backup files on the same workstation are not an off-site backup.
## Isolated restore drill
Run a real restore into a uniquely named temporary database:
```bash
./scripts/restore-drill-compose.sh output/backups/compose-YYYYMMDD-HHMMSS.dump
```
The drill:
1. validates the SHA-256 manifest;
2. validates the archive table of contents;
3. creates a pristine database from `template0`;
4. restores with `pg_restore --exit-on-error`;
5. reads every restored public application table and checks the PostGIS
library;
6. runs the current immutable release's migrations and migration-readiness check
against only the temporary database;
7. drops only the temporary drill database and verifies that it is gone.
A trap also attempts to drop the exact temporary database if a check fails.
The source application database is never passed to `pg_restore`, `dropdb`, a
clean operation, or the drill migration runner. The drill intentionally does
not compare an older backup's row counts to the live source, because concurrent
legitimate writes or a historical archive would make that comparison invalid.
## Service checks
```bash
docker compose ps
curl --fail http://localhost:4010/healthz/live
curl --fail http://localhost:4010/healthz/ready
./scripts/verify-realtime-cluster.sh compose
```
`live` verifies that the web process can serve HTTP. `ready` additionally runs
`SELECT 1` through the configured Ecto repository. The cluster probe subscribes
on one connected BEAM node and broadcasts from another.
Health checks do not replace alerting, database backups, restore drills, or
application-level synthetic checks.
## Local failure and rolling-replacement drills
The isolated load project can exercise process crashes, sequential container
replacement, and a real Oban retry without touching the normal Compose project:
```bash
./scripts/load-stack-up.sh
./scripts/load-resilience-run.sh local-resilience
```
The resilience script refuses `LOAD_PROJECT=who_need_help` and verifies the
Compose project/service labels of every container before stopping it. It:
1. continuously calls readiness through the isolated Traefik route;
2. terminates the BEAM process in one web and one worker container and requires
Docker's observed restart count to increase;
3. removes and replaces each web and worker replica one at a time;
4. waits for every configured BEAM node, then runs the cross-node PubSub probe;
5. enqueues a side-effect-free local worker that fails its first Oban attempt
and succeeds on its second;
6. removes that exact Oban row and requires no fixture domain rows to remain.
Traefik's retry middleware is attached to the HTTP and local TLS routers. Its
attempt count is an environment input. Traefik retries transport failures and,
with the checked configuration, does not opt in to retrying non-idempotent
requests. This reduces a stale-backend window; it is not a claim of production
availability.
For the project-owned kind cluster, run:
```bash
./scripts/kind-rolling-verify.sh local-kind-rollout
```
That script requires both the kind ownership marker and the control-plane
cluster label before invoking `rollout restart`. It changes only the web and
worker Deployment pod templates. It snapshots application-table counts before
and after, continuously probes the observed Docker mapping for the chart's
NodePort, requires all four old pod UIDs to disappear, waits for the exact BEAM
peer count, and verifies cross-node PubSub. PostGIS, its hostPath, the
Kubernetes Secret, and the namespace are not recreated.
The rollout timeout, probe interval/timeout/retry count, and cluster-join
timeout are experiment inputs. They are not production SLOs or resource
requirements.
## Protected Prometheus metrics
The web role exposes Prometheus text format at `/metrics`. It requires the
independent `METRICS_TOKEN` deployment secret:
```bash
curl --fail \
--header "Authorization: Bearer $METRICS_TOKEN" \
http://localhost:4010/metrics
```
The endpoint returns `401` without the exact token, disables response caching,
and does not put the credential in a URL. The reporter exports cumulative HTTP
request and duration, router exception, database query and duration, WebSocket
connection, VM memory, and scheduler run-queue metrics. Cumulative durations are
integer microseconds because the selected reporter's sum accumulator is
integer-based; divide by `1_000_000` in PromQL when seconds are required.
Definitions intentionally have no request path, user, request, or event-name
labels that could create unbounded cardinality.
Metrics are local to each BEAM process. Discover and scrape every web pod or
container as a distinct target and preserve Prometheus's `instance` label. A
request through the load-balanced public route reaches only one replica and is
therefore useful as an authorization/smoke check, not as a cluster-wide
aggregate.
The isolated load project includes a local observability profile:
```bash
./scripts/load-stack-up.sh
./scripts/observability-run.sh local-observability
```
The run script refuses the staging project, validates every current web
container's Compose labels, and writes a `file_sd` target for each observed
internal IP. Prometheus reads the Bearer value from a mode-`0600` runtime file,
not a tracked config or URL. Its direct request includes the internal
`X-Forwarded-Proto: https` signal required by the application's production SSL
rewrite while preserving the target's own `instance` label.
Prometheus, Alertmanager, and Grafana are pinned by tag and digest. Their host
ports default to Docker-assigned values bound only to `127.0.0.1`; the run
prints the observed URLs. Grafana uses the random admin password generated in
ignored `.env.load`, disables anonymous signup, update checks, suggested plugin
installation, and its unused built-in alert engine. The Prometheus datasource
and four-panel dashboard are provisioned from tracked files.
The verification stops exactly one scoped load web container. The
`WhoNeedHelpWebReplicaUnavailable` rule is based only on the factual
`up == 0` result; it is a local failure drill, not an invented latency,
capacity, or production SLO threshold. The script requires both firing and
resolved webhook payloads from Alertmanager, starts the same container, waits
for every direct target, and compares read-only database counts before and
after. Evidence is retained in `output/observability/` without the metrics or
Grafana secrets.
Stop only the monitoring services with:
```bash
./scripts/observability-stop.sh
```
Prometheus/Grafana/Alertmanager retention, production notification
destinations, production availability, and measured alert policies remain
deployment decisions. In Kubernetes, put the metrics token in
`existingSecret`; configure the external scraper to send it as a Bearer token.
## Rollback boundary
The release image is immutable and migrations run as a separate one-shot role.
Before a schema rollout, create and restore-test a current backup. Application
rollback and database migration rollback are separate decisions: do not run an
Ecto down migration merely because an image is rolled back. Inspect the exact
migration and compatibility boundary first.
The repository intentionally does not ship an automatic destructive production
restore command.