158 lines
6.6 KiB
Markdown
158 lines
6.6 KiB
Markdown
# Operations runbook
|
|
|
|
This runbook describes the commands that are implemented and verified in this
|
|
repository. It does not claim a production recovery point objective, recovery
|
|
time objective, retention period, storage capacity, or high-availability model;
|
|
those values require product policy and measurements from the eventual
|
|
production environment.
|
|
|
|
## Compose database backup
|
|
|
|
Create a PostgreSQL 18 custom-format archive, validate its table of contents,
|
|
and write a SHA-256 manifest:
|
|
|
|
```bash
|
|
./scripts/backup-compose.sh
|
|
```
|
|
|
|
The default destination is the ignored `output/backups/` directory. An explicit
|
|
new destination may be supplied as the only argument. The command refuses to
|
|
overwrite either an archive or its checksum manifest and writes through
|
|
temporary files before publishing the final pair. Files and newly created
|
|
directories are restricted by `umask 077`.
|
|
|
|
The archive covers the configured application database. PostgreSQL cluster
|
|
globals such as roles and tablespaces are not part of `pg_dump`; deployment
|
|
credentials and database roles must be provisioned separately from secrets.
|
|
Local backup files on the same workstation are not an off-site backup.
|
|
|
|
## Isolated restore drill
|
|
|
|
Run a real restore into a uniquely named temporary database:
|
|
|
|
```bash
|
|
./scripts/restore-drill-compose.sh output/backups/compose-YYYYMMDD-HHMMSS.dump
|
|
```
|
|
|
|
The drill:
|
|
|
|
1. validates the SHA-256 manifest;
|
|
2. validates the archive table of contents;
|
|
3. creates a pristine database from `template0`;
|
|
4. restores with `pg_restore --exit-on-error`;
|
|
5. reads every restored public application table and checks the PostGIS
|
|
library;
|
|
6. runs the current immutable release's migrations and migration-readiness check
|
|
against only the temporary database;
|
|
7. drops only the temporary drill database and verifies that it is gone.
|
|
|
|
A trap also attempts to drop the exact temporary database if a check fails.
|
|
The source application database is never passed to `pg_restore`, `dropdb`, a
|
|
clean operation, or the drill migration runner. The drill intentionally does
|
|
not compare an older backup's row counts to the live source, because concurrent
|
|
legitimate writes or a historical archive would make that comparison invalid.
|
|
|
|
## Service checks
|
|
|
|
```bash
|
|
docker compose ps
|
|
curl --fail http://localhost:4010/healthz/live
|
|
curl --fail http://localhost:4010/healthz/ready
|
|
./scripts/verify-realtime-cluster.sh compose
|
|
```
|
|
|
|
`live` verifies that the web process can serve HTTP. `ready` additionally runs
|
|
`SELECT 1` through the configured Ecto repository. The cluster probe subscribes
|
|
on one connected BEAM node and broadcasts from another.
|
|
|
|
Health checks do not replace alerting, database backups, restore drills, or
|
|
application-level synthetic checks.
|
|
|
|
## Local failure and rolling-replacement drills
|
|
|
|
The isolated load project can exercise process crashes, sequential container
|
|
replacement, and a real Oban retry without touching the normal Compose project:
|
|
|
|
```bash
|
|
./scripts/load-stack-up.sh
|
|
./scripts/load-resilience-run.sh local-resilience
|
|
```
|
|
|
|
The resilience script refuses `LOAD_PROJECT=who_need_help` and verifies the
|
|
Compose project/service labels of every container before stopping it. It:
|
|
|
|
1. continuously calls readiness through the isolated Traefik route;
|
|
2. terminates the BEAM process in one web and one worker container and requires
|
|
Docker's observed restart count to increase;
|
|
3. removes and replaces each web and worker replica one at a time;
|
|
4. waits for every configured BEAM node, then runs the cross-node PubSub probe;
|
|
5. enqueues a side-effect-free local worker that fails its first Oban attempt
|
|
and succeeds on its second;
|
|
6. removes that exact Oban row and requires no fixture domain rows to remain.
|
|
|
|
Traefik's retry middleware is attached to the HTTP and local TLS routers. Its
|
|
attempt count is an environment input. Traefik retries transport failures and,
|
|
with the checked configuration, does not opt in to retrying non-idempotent
|
|
requests. This reduces a stale-backend window; it is not a claim of production
|
|
availability.
|
|
|
|
For the project-owned kind cluster, run:
|
|
|
|
```bash
|
|
./scripts/kind-rolling-verify.sh local-kind-rollout
|
|
```
|
|
|
|
That script requires both the kind ownership marker and the control-plane
|
|
cluster label before invoking `rollout restart`. It changes only the web and
|
|
worker Deployment pod templates. It snapshots application-table counts before
|
|
and after, continuously probes the observed Docker mapping for the chart's
|
|
NodePort, requires all four old pod UIDs to disappear, waits for the exact BEAM
|
|
peer count, and verifies cross-node PubSub. PostGIS, its hostPath, the
|
|
Kubernetes Secret, and the namespace are not recreated.
|
|
|
|
The rollout timeout, probe interval/timeout/retry count, and cluster-join
|
|
timeout are experiment inputs. They are not production SLOs or resource
|
|
requirements.
|
|
|
|
## Protected Prometheus metrics
|
|
|
|
The web role exposes Prometheus text format at `/metrics`. It requires the
|
|
independent `METRICS_TOKEN` deployment secret:
|
|
|
|
```bash
|
|
curl --fail \
|
|
--header "Authorization: Bearer $METRICS_TOKEN" \
|
|
http://localhost:4010/metrics
|
|
```
|
|
|
|
The endpoint returns `401` without the exact token, disables response caching,
|
|
and does not put the credential in a URL. The reporter exports cumulative HTTP
|
|
request and duration, router exception, database query and duration, WebSocket
|
|
connection, VM memory, and scheduler run-queue metrics. Cumulative durations are
|
|
integer microseconds because the selected reporter's sum accumulator is
|
|
integer-based; divide by `1_000_000` in PromQL when seconds are required.
|
|
Definitions intentionally have no request path, user, request, or event-name
|
|
labels that could create unbounded cardinality.
|
|
|
|
Prometheus itself, durable metrics retention, alert rules, notification
|
|
destinations, and measured alert thresholds are deployment responsibilities and
|
|
are not claimed by this repository. In Kubernetes, put the token in
|
|
`existingSecret`; configure the external scraper to send it as a Bearer token.
|
|
|
|
Metrics are local to each BEAM process. Discover and scrape every web pod or
|
|
container as a distinct target and preserve Prometheus's `instance` label. A
|
|
request through the load-balanced public route reaches only one replica and is
|
|
therefore useful as an authorization/smoke check, not as a cluster-wide
|
|
aggregate.
|
|
|
|
## Rollback boundary
|
|
|
|
The release image is immutable and migrations run as a separate one-shot role.
|
|
Before a schema rollout, create and restore-test a current backup. Application
|
|
rollback and database migration rollback are separate decisions: do not run an
|
|
Ecto down migration merely because an image is rolled back. Inspect the exact
|
|
migration and compatibility boundary first.
|
|
|
|
The repository intentionally does not ship an automatic destructive production
|
|
restore command.
|