19 KiB
Operations runbook
This runbook describes the commands that are implemented and verified in this repository. It does not claim a production recovery point objective, recovery time objective, retention period, storage capacity, or high-availability model; those values require product policy and measurements from the eventual production environment.
Compose database backup
Create a PostgreSQL 18 custom-format archive, validate its table of contents, and write a SHA-256 manifest:
./scripts/backup-compose.sh
The default destination is the ignored output/backups/ directory. An explicit
new destination may be supplied as the only argument. The command refuses to
overwrite either an archive or its checksum manifest and writes through
temporary files before publishing the final pair. Files and newly created
directories are restricted by umask 077.
The archive covers the configured application database. PostgreSQL cluster
globals such as roles and tablespaces are not part of pg_dump; deployment
credentials and database roles must be provisioned separately from secrets.
Local backup files on the same workstation are not an off-site backup.
Encrypted local S3-compatible backup drill
The isolated load project can run a complete encrypted Restic/MinIO drill:
./scripts/load-stack-up.sh
./scripts/backup-s3-drill.sh local-encrypted-backup
The script refuses the staging Compose project and validates the project and
service labels of every pre-existing container in its scope. On first use,
scripts/ensure-local-load-env.sh generates independent random MinIO and
Restic credentials in ignored .env.load and restricts that file to mode
0600. MinIO publishes Docker-assigned ports only on 127.0.0.1; the observed
API and console URLs are printed after a successful run.
The backup tool combines the matching PostgreSQL 18 client with pinned Restic
rebuilt on Go 1.26.5. MinIO server and client are also rebuilt as non-root
Alpine images from checksum-pinned upstream source commits with the exact
dependency updates recorded in Dockerfile.minio. The quality gate verifies
their reported release, commit, Go runtime, configured user, and current
HIGH/CRITICAL vulnerability scan. restic backup --stdin-from-command runs a
custom-format pg_dump, checks the producer exit status, encrypts the data,
and uploads it directly to MinIO. No plaintext database dump is written to the
host. The drill then:
- runs
restic check --read-data; - streams
restic dumpintopg_restore --list; - restores into a uniquely named database created from
template0; - checks tables, current Ecto migrations, PostGIS, and release migration readiness before removing that exact database;
- clones and corrupts an isolated repository and requires both check and dump to fail;
- stops an exact scoped in-progress backup container only after encrypted objects reach MinIO, requires zero published snapshots, prunes unreferenced packs, and rechecks the repository;
- removes and verifies removal of the corruption/interruption buckets and requires source table counts to remain unchanged.
The successful encrypted bucket is deliberately retained in the named local
MinIO volume. Non-secret evidence is written under ignored
output/backups-s3/<run-label>/; the runtime scratch directory is under
ignored tmp/backup-s3/. Use a unique lowercase run label of at most 32
characters. The command refuses to replace an existing retained bucket.
MinIO server and client are AGPLv3. Dockerfile.minio identifies the exact
upstream source commits and contains the dependency changes and complete build
commands used here. Before distributing or publicly operating modified images,
review the license and make the corresponding source available as required;
this runbook does not provide legal advice.
This verifies encryption, local S3 protocol use, restore mechanics, and two failure paths on the observed workstation. A MinIO volume on that same workstation is not an off-site backup and does not establish production RPO, RTO, retention, capacity, key custody, object locking, or database HA.
Local external-service boundary drill
Run the OAuth, SMTP, and provider-neutral push protocol checks without public credentials, a server, or host-published ports:
./scripts/external-boundaries-run.sh local-boundaries
The script creates a uniquely named Compose project on an internal-only Docker
network. It generates independent one-run OAuth and push credentials in an
ignored mode-0600 environment file, builds the production release plus a
non-root standard-library Python protocol mock, and then verifies:
- GitHub-compatible OAuth authorization, PKCE S256, token exchange, normalized user lookup, state mismatch, provider rejection, one-time-code replay, a fresh flow after a temporary token error, and a token timeout;
- SMTP acceptance, permanent recipient rejection without retry, one retry after a temporary greeting failure, a greeting timeout, and the result of submitting the same message twice;
- the disabled default push boundary plus HTTP success, permanent rejection, one temporary retry, replay deduplication, and deduplication after an ambiguous timeout using the same idempotency key;
- a fresh PostGIS database, current migrations, and two Oban worker replicas; request acceptance and new-chat domain transactions enqueue stable user-recipient events, replay is deduplicated before HTTP, an injected temporary chat delivery fails its first Oban attempt and completes on its second, and private message text is absent from the push payload.
The evidence JSON and mock state contain counters, booleans, normalized identity fields, payload digests, run-scoped user/event identifiers, and the privacy-safe notification metadata asserted by the drill. The script fails if any generated secret appears in retained evidence. Its trap validates exact Compose labels, removes only that project, volumes and one-run images, and deletes the temporary credential file. Failure logs are retained under the same ignored evidence directory.
This drill uses the application's real Assent/Req and Swoosh/gen_smtp clients,
but the providers are local. It therefore verifies client-side protocol and
product integration, not GitHub, SMTP-provider, FCM, or APNs availability.
SMTP permits duplicate delivery after ambiguous outcomes, so the result
explicitly makes no exactly-once claim. Push currently targets a stable
user:<uuid> recipient; selecting a provider, registering device tokens, and
resolving that user to devices remain deployment/provider work.
The implementation follows the configured adapter interfaces and protocol semantics documented by Assent 0.3.1, GitHub OAuth, and SMTP RFC 5321.
Isolated restore drill
Run a real restore into a uniquely named temporary database:
./scripts/restore-drill-compose.sh output/backups/compose-YYYYMMDD-HHMMSS.dump
The drill:
- validates the SHA-256 manifest;
- validates the archive table of contents;
- creates a pristine database from
template0; - restores with
pg_restore --exit-on-error; - reads every restored public application table and checks the PostGIS library;
- runs the current immutable release's migrations and migration-readiness check against only the temporary database;
- drops only the temporary drill database and verifies that it is gone.
A trap also attempts to drop the exact temporary database if a check fails.
The source application database is never passed to pg_restore, dropdb, a
clean operation, or the drill migration runner. The drill intentionally does
not compare an older backup's row counts to the live source, because concurrent
legitimate writes or a historical archive would make that comparison invalid.
Isolated Compose upgrade rehearsal
After creating a current custom-format backup, run the complete current release against an isolated restored copy:
./scripts/upgrade-rehearsal-compose.sh \
output/backups/compose-YYYYMMDD-HHMMSS.dump
The rehearsal validates the checksum and archive catalog, reads the public
origin configuration from ignored .env, and uses only the independently
generated credentials in ignored mode-0600 .env.e2e. It builds a uniquely
tagged production release, creates a uniquely named Compose project and
database from template0, restores the archive, records application-table
counts, and then:
- applies every current timestamped Ecto migration;
- requires the localized category column and all 11 valid cursor indexes;
- starts 2 web and 2 worker replicas behind the isolated Traefik instance;
- requires the four-node BEAM cluster and cross-node PubSub probe;
- checks the production HTTP-to-HTTPS redirect, trusted-proxy public pages, and both health endpoints;
- requires an empty before/after diff for every public application table except the expected migration and Oban-internal tables;
- removes and verifies removal of the exact containers, networks, database volume, and one-run image.
The input archive is read-only and is not copied into the evidence directory.
The ordinary Compose project, source database, public route, and running
containers are outside the generated project scope. Non-secret evidence is
retained under ignored mode-0700
output/upgrade-rehearsal/<run-id>/, with files mode 0600.
Clean tracked-revision deployment drill
Run:
./scripts/clean-deploy-verify.sh
The script archives the current tracked Git revision into a new temporary
directory. It verifies that the archive contains no local .env, Git
metadata, generated output, or existing E2E/load environment. It generates
independent one-run PostgreSQL, Phoenix, handover, BEAM-cookie, and metrics
secrets with mode 0600; optional GitHub OAuth and push delivery stay
disabled.
The drill uses a unique Compose project, application image, router/service
name, internal network, PostgreSQL volume, database, and dynamic host ports.
It requires exactly two healthy web and two running worker replicas, all
tracked migrations, seeded categories, an idempotent repeated migration,
working readiness/home/registration/Mailpit routes, four connected BEAM
nodes, and a cross-replica PubSub message. It then removes the exact project
including volumes, its one-run image, and the temporary archive, and verifies
that those scoped resources are absent. Evidence is retained under
output/portability/<run-id>/.
This is a local Docker portability observation. It does not establish production SMTP, OAuth, push-provider, TLS, database-HA, storage, capacity, or jurisdictional readiness.
Service checks
docker compose ps
curl --fail http://localhost:4010/healthz/live
curl --fail http://localhost:4010/healthz/ready
./scripts/verify-realtime-cluster.sh compose
live verifies that the web process can serve HTTP. ready additionally runs
SELECT 1 through the configured Ecto repository. The cluster probe subscribes
on one connected BEAM node and broadcasts from another.
Health checks do not replace alerting, database backups, restore drills, or application-level synthetic checks.
Local failure and rolling-replacement drills
The isolated load project can exercise process crashes, sequential container replacement, and a real Oban retry without touching the normal Compose project:
./scripts/load-stack-up.sh
./scripts/load-resilience-run.sh local-resilience
The resilience script refuses LOAD_PROJECT=who_need_help and verifies the
Compose project/service labels of every container before stopping it. It:
- continuously calls readiness through the isolated Traefik route;
- terminates the BEAM process in one web and one worker container and requires Docker's observed restart count to increase;
- removes and replaces each web and worker replica one at a time;
- waits for every configured BEAM node, then runs the cross-node PubSub probe;
- enqueues a side-effect-free local worker that fails its first Oban attempt and succeeds on its second;
- removes that exact Oban row and requires no fixture domain rows to remain.
Traefik's retry middleware is attached to the HTTP and local TLS routers. Its attempt count is an environment input. Traefik retries transport failures and, with the checked configuration, does not opt in to retrying non-idempotent requests. This reduces a stale-backend window; it is not a claim of production availability.
For the project-owned kind cluster, run:
./scripts/kind-rolling-verify.sh local-kind-rollout
That script requires both the kind ownership marker and the control-plane
cluster label before invoking rollout restart. It changes only the web and
worker Deployment pod templates. It snapshots application-table counts before
and after, continuously probes the observed Docker mapping for the chart's
NodePort, requires all four old pod UIDs to disappear, waits for the exact BEAM
peer count, and verifies cross-node PubSub. PostGIS, its hostPath, the
Kubernetes Secret, and the namespace are not recreated.
The rollout timeout, probe interval/timeout/retry count, and cluster-join timeout are experiment inputs. They are not production SLOs or resource requirements.
Isolated Oban burst measurement
The worker role consumes only maintenance and push. Configure their
per-worker limits with OBAN_MAINTENANCE_CONCURRENCY and
OBAN_PUSH_CONCURRENCY; multiplying either value by the number of worker
replicas gives the configured cluster-wide concurrency for that queue.
After starting the isolated load project, run an explicitly sized experiment:
./scripts/load-stack-up.sh
./scripts/oban-burst-run.sh local-oban-burst 1000 120
The three required arguments are an evidence label, job count, and recorded
experiment timeout in seconds. They are not capacity thresholds. The script
refuses the ordinary who_need_help project, verifies the expected worker
replicas, inserts only confirmed LocalBurstProbe jobs with a unique run id,
records each worker's effective queue configuration and container samples,
requires every job to complete, compares domain-table row counts, and deletes
exactly its own jobs. The probe worker has no product side effects and no
product flow enqueues it.
Kubernetes ingress isolation
networkPolicy.enabled=true renders the chart's ingress NetworkPolicy.
Enforcement is a property of the cluster CNI, not of the YAML object alone.
Before relying on it, verify that the target cluster uses a NetworkPolicy-
capable plugin and run positive HTTP/cluster checks plus negative blocked-port
checks there. Egress remains intentionally unrestricted until the actual
database and external-service destinations are known.
BEAM runtime memory guard
Compose sets ERL_ZFLAGS="+Q ${ERLANG_PORT_LIMIT}"; Helm renders the same flag
from app.erlangPortLimit. Both default to 65536, OTP's normal port-table
limit. This makes the runtime independent of an unusually large host or nested
container nofile limit. Validate all live replicas after deployment:
./scripts/verify-beam-runtime.sh compose
./scripts/verify-beam-runtime.sh kind
The command records the effective port count/limit, allocated port-table bytes, BEAM memory, process count, cgroup memory, and RSS for each web and worker replica. It fails when a live VM does not use the configured limit; kind mode also requires the desired number of Ready application pods.
65536 is a concurrency ceiling for simultaneously existing Erlang ports
(files, sockets, and drivers), not a container memory limit. Do not lower or
raise it from a RAM estimate alone. A changed value must be validated against
measured peak port usage and the target environment.
Protected Prometheus metrics
The web role exposes Prometheus text format at /metrics. It requires the
independent METRICS_TOKEN deployment secret:
curl --fail \
--header "Authorization: Bearer $METRICS_TOKEN" \
http://localhost:4010/metrics
The endpoint returns 401 without the exact token, disables response caching,
and does not put the credential in a URL. The reporter exports cumulative HTTP
request and duration, router exception, database query and duration, WebSocket
connection, VM memory, and scheduler run-queue metrics. Cumulative durations are
integer microseconds because the selected reporter's sum accumulator is
integer-based; divide by 1_000_000 in PromQL when seconds are required.
Definitions intentionally have no request path, user, request, or event-name
labels that could create unbounded cardinality.
Metrics are local to each BEAM process. Discover and scrape every web pod or
container as a distinct target and preserve Prometheus's instance label. A
request through the load-balanced public route reaches only one replica and is
therefore useful as an authorization/smoke check, not as a cluster-wide
aggregate.
The isolated load project includes a local observability profile:
./scripts/load-stack-up.sh
./scripts/observability-run.sh local-observability
The run script refuses the staging project, validates every current web
container's Compose labels, and writes a file_sd target for each observed
internal IP. Prometheus reads the Bearer value from a mode-0600 runtime file,
not a tracked config or URL. Its direct request includes the internal
X-Forwarded-Proto: https signal required by the application's production SSL
rewrite while preserving the target's own instance label.
Prometheus, Alertmanager, and Grafana are pinned by tag and digest. Their host
ports default to Docker-assigned values bound only to 127.0.0.1; the run
prints the observed URLs. Grafana uses the random admin password generated in
ignored .env.load, disables anonymous signup, update checks, suggested plugin
installation, and its unused built-in alert engine. The Prometheus datasource
and four-panel dashboard are provisioned from tracked files.
The verification stops exactly one scoped load web container. The
WhoNeedHelpWebReplicaUnavailable rule is based only on the factual
up == 0 result; it is a local failure drill, not an invented latency,
capacity, or production SLO threshold. The script requires both firing and
resolved webhook payloads from Alertmanager, starts the same container, waits
for every direct target, and compares read-only database counts before and
after. Evidence is retained in output/observability/ without the metrics or
Grafana secrets.
Stop only the monitoring services with:
./scripts/observability-stop.sh
Prometheus/Grafana/Alertmanager retention, production notification
destinations, production availability, and measured alert policies remain
deployment decisions. In Kubernetes, put the metrics token in
existingSecret; configure the external scraper to send it as a Bearer token.
Rollback boundary
The release image is immutable and migrations run as a separate one-shot role. Before a schema rollout, create and restore-test a current backup. Application rollback and database migration rollback are separate decisions: do not run an Ecto down migration merely because an image is rolled back. Inspect the exact migration and compatibility boundary first.
The repository intentionally does not ship an automatic destructive production restore command.