who_need_help/docs/operations.md

69 KiB

Operations runbook

This runbook describes the commands that are implemented and verified in this repository. It does not claim a production recovery point objective, recovery time objective, retention period, storage capacity, or high-availability model; those values require product policy and measurements from the eventual production environment.

Private Gitea continuous integration

The observed private Git remote is hosted by Gitea 1.22.0. Gitea 1.22 reads repository workflows from .gitea/workflows/, so the repository keeps a Gitea-specific quality workflow in addition to the GitHub-compatible workflow. Both files are checked by the pinned local actionlint gate.

The Gitea workflow deliberately targets only the custom who-need-help-ci runner label. Its jobs build and run many Docker images and must not be scheduled onto an unrelated shared runner or onto the production application host. Its gates run sequentially in one job so a single push cannot start four heavy Docker suites concurrently. Before enabling it:

  1. enable Actions in the private repository settings;
  2. register a trusted runner controlled by this project;
  3. assign who-need-help-ci:host only to a Linux host where the exact Docker Engine and Compose commands used by the repository have been verified;
  4. install the runtime required by the pinned checkout action on that host;
  5. keep the runner registration state and any Docker socket access outside the repository and away from untrusted projects.

The runner's availability, labels, software versions, free resources and repository Actions setting cannot be established from the unauthenticated repository clone. Do not treat the presence of the workflow file as evidence that remote CI has run. After provisioning the runner, trigger the workflow manually, require the full job and all four gate steps to pass, and record the run URL and commit before protecting main.

First production Compose environment

The deployment environment selects topology and database ownership:

Setting Result
APP_TOPOLOGY=compact One Phoenix+Oban container, directly published to the host; default for the first server.
APP_TOPOLOGY=split Traefik, WEB_REPLICAS web containers, and WORKER_REPLICAS worker containers.
DATABASE_MODE=container Start the project-owned PostgreSQL/PostGIS container and volume.
DATABASE_MODE=external Do not create/start a Compose database; use only DATABASE_URL.

There is no Redis dependency. Queues, rate-limit counters, Oban leadership, and durable application state use PostgreSQL.

Development public route after a workstation restart

The current development gateway sends whoneedhelp.imalto.site to the workstation's assigned OpenVPN address 10.8.0.14:4010. If HTTPS connects to the gateway but returns no response after a workstation restart, first verify the local app and tunnel:

curl --fail --show-error --max-time 10 \
  http://127.0.0.1:4010/healthz/ready
ip -brief address show tun0
nmcli --wait 30 connection up openvpn__toha_nobara_pc
curl --fail --show-error --max-time 15 \
  https://whoneedhelp.imalto.site/healthz/ready

The observed profile requires remote-cert-tls=server; do not remove that server-certificate check. NetworkManager documents that connection.autoconnect is not implemented for VPN profiles and recommends a base connection's connection.secondaries instead: NetworkManager connection settings. The observed wired base profile currently has no secondary connection, so a standalone connection.autoconnect=yes value on the VPN does not prove that the development route will return after reboot.

Adding the VPN UUID as a wired/Wi-Fi secondary is a workstation networking change: this profile currently installs the default route and DNS through the VPN. Review that whole-host effect and obtain explicit approval before configuring automatic activation. The manual command above changes no project, test, production, or gateway configuration.

Two independent checkouts and one .env in each

The server uses exactly these independent Git clones:

/srv/who_need_help-test/.env
/srv/who_need_help-production/.env

No .env.test, .env.staging, or .env.production is used inside either application checkout. Each checkout can remain on a different commit. Test and production have different Compose projects, application/infrastructure image tags, database connections, Docker volumes, public aliases, Google OAuth clients, email delivery paths, and generated secrets. Oban queues are isolated by those different PostgreSQL databases. There is no Redis dependency.

Generated harness state is not a deployment environment. E2E, load, and Android instrumentation scripts keep their random local inputs below the ignored mode-0700 output/runtime/ directory:

output/runtime/e2e.env
output/runtime/load.env
output/runtime/android-test.env

Those files are generated automatically, never copied to a server, and do not represent dev, test, or production. The one ignored .env at each checkout root remains the only application/deployment configuration.

The current submitted deployment still has a legacy shared Caddy container created from the production checkout. Keep it running and unchanged while the judging test URL is frozen. Application release and rollback scripts do not build, recreate, or select an image for that container.

The replacement is a separate server-level server_edge project with its own single mode-0600 .env, route directory, Caddy image, certificate volumes, and Compose lifecycle. Each application keeps only its route contract in its existing .env: PHX_HOST, PUBLIC_ROUTE_ID, PUBLIC_EDGE_NETWORK, PUBLIC_UPSTREAM_NAME, PUBLIC_UPSTREAM_PORT, PUBLIC_HEALTH_PATH, and PUBLIC_WWW_REDIRECT. No application deployment is allowed to restart the server edge. Both applications intentionally share only the external public edge Docker network.

Android build inputs belong in the same ignored mode-0600 .env as the web checkout they target. Do not create .env.android-release, .env.production, or another permanent environment file:

WNH_BASE_URL=https://dev.example.com
WNH_ANDROID_VERSION_CODE=1
WNH_ANDROID_VERSION_NAME=0.1.0
WNH_ANDROID_DEVELOPMENT_SIGNING_KEY_ALIAS=who-need-help-development
ANDROID_APP_LINKS_PACKAGE_NAME=org.whoneedhelp.mobile.development
ANDROID_APP_LINKS_SHA256_CERT_FINGERPRINTS=AA:BB:...

The development checkout uses package org.whoneedhelp.mobile.development and a dedicated stable key under ~/.config/who_need_help/android-development/. Generate it once and place only its public identity into the existing ignored .env:

./scripts/init-android-development-signing.sh
./scripts/configure-android-development-env.sh
./scripts/android-development-build.sh

The build exports the package and certificate reports, verifies that both match the checkout .env, and checks the HTTPS /.well-known/assetlinks.json response. Losing this key changes the development certificate and breaks previously installed App Links, so back it up.

The independent test checkout uses package org.whoneedhelp.mobile.staging, its own key under ~/.config/who_need_help/android-staging/, and scripts/android-staging-build.sh. Do not copy either ignored .env or signing material between development and test.

The production checkout instead uses package org.whoneedhelp.mobile, the separate upload material under ~/.config/who_need_help/android-release/, and its own .env:

WNH_ANDROID_SIGNING_KEY_ALIAS=who-need-help-upload
ANDROID_APP_LINKS_PACKAGE_NAME=org.whoneedhelp.mobile
ANDROID_APP_LINKS_SHA256_CERT_FINGERPRINTS=UPLOAD_SHA256,PLAY_APP_SIGNING_SHA256
ANDROID_PLAY_APP_SIGNING_SHA256_CERT_FINGERPRINTS=PLAY_APP_SIGNING_SHA256

Run ./scripts/android-release-build.sh from that production release checkout. It produces an APK, Play AAB, package report, signing report, and lint report. Before the Play application exists, the build may use only the upload certificate in ANDROID_APP_LINKS_SHA256_CERT_FINGERPRINTS and leave ANDROID_PLAY_APP_SIGNING_SHA256_CERT_FINGERPRINTS empty. This is sufficient to create the first signed AAB, but check-environment-readiness.sh --require-release intentionally continues to report the Play identity as missing. Application-only server releases use --require-server-release plus the production validator's --allow-pre-play mode: they accept this upload-certificate-only state while continuing to reject every other missing production capability. These modes do not make an Android build ready for Google Play. After Play App Signing is enabled, add the Play signing certificate fingerprint to the comma-separated App Links value; the upload certificate alone does not describe Play-delivered APKs. Record the same Play fingerprint separately in ANDROID_PLAY_APP_SIGNING_SHA256_CERT_FINGERPRINTS; release readiness verifies that every Play identity is present in the public App Links list. The production environment validator accepts multiple SHA-256 fingerprints and rejects partial, malformed, upload-only, or inconsistent configuration.

Release capability inputs

Each environment owns distinct external-provider credentials. Seed a new production .env with scripts/init-production-env.sh and the corresponding PRODUCTION_* process variables; use TEST_* only when creating the separate test checkout. The generated file uses the ordinary runtime names:

Capability Values kept in that checkout's .env
Google sign-in GOOGLE_OAUTH_CLIENT_ID, GOOGLE_OAUTH_CLIENT_SECRET, GOOGLE_OAUTH_AUTHORIZED_PARTY_IDS
Browser push three WEB_PUSH_VAPID_* values
Android Firebase client four public WNH_FIREBASE_* values
Android delivery FCM_PROJECT_ID and one private service-account source
Verified Android links ANDROID_APP_LINKS_PACKAGE_NAME, ANDROID_APP_LINKS_SHA256_CERT_FINGERPRINTS
Transactional email SMTP_*, sender, and SUPPORT_INBOX_ADDRESS

SUPPORT_INBOX_ADDRESS configures Reply-To and the optional operator-alert destination. The environment readiness script validates only that this value is present and syntactically usable; it does not prove that the domain has inbound MX routing or that the mailbox is monitored. Check DNS and perform an actual inbound delivery/reply test before presenting the address as a public support contact.

Do not reuse a Google client, VAPID private key, Firebase project/service account, SMTP credential, or Android signing key between dev and production. The public Firebase Android values are build configuration; the Base64 FCM service-account JSON is a server secret and must never be passed into the Android build.

For Android Google sign-in, each environment's GOOGLE_OAUTH_CLIENT_ID is also the public server client ID supplied at runtime to Credential Manager. GOOGLE_OAUTH_CLIENT_SECRET never leaves Phoenix. The Android package and signing-certificate identity must be registered in the matching Google project: org.whoneedhelp.mobile.development plus the stable development certificate for development, org.whoneedhelp.mobile.staging plus the stable staging certificate for test, and org.whoneedhelp.mobile plus the Play-distributed certificate for production. Do not add a second Google secret file or Gradle property. Set GOOGLE_OAUTH_AUTHORIZED_PARTY_IDS to the comma-separated Android OAuth client IDs from that environment. They are accepted only as the signed azp claim; the signed aud must still contain the environment's Web client ID.

Check an environment without printing its secret values:

./scripts/check-environment-readiness.sh .env
./scripts/check-environment-readiness.sh .env --require-release

The complete promotion sequence and the human/provider decisions that these scripts cannot prove are listed in docs/public-launch-checklist.md.

Provider downloads can be imported into that same file without placing secrets on a command line or printing them:

./scripts/generate-vapid-env.sh \
  .env mailto:contact@YOUR_DOMAIN

chmod 600 /secure/downloads/google-oauth-client.json
./scripts/import-google-oauth-client.sh \
  .env /secure/downloads/google-oauth-client.json

./scripts/import-firebase-android-config.sh \
  .env /secure/downloads/google-services.json

chmod 600 /secure/downloads/fcm-service-account.json
./scripts/import-fcm-service-account.sh \
  .env /secure/downloads/fcm-service-account.json

After the first AAB has made Google Play generate its delivery identities, download a fresh production google-services.json after registering every Play App Signing SHA-1 in the production Firebase Android application. Record the SHA-1/SHA-256 pairs shown by Play in a protected temporary JSON document:

{
  "package_name": "org.whoneedhelp.mobile",
  "identities": [
    {"sha1": "PLAY_SHA1", "sha256": "PLAY_SHA256"}
  ]
}

Keep both provider downloads at mode 0400 or 0600. First run the importer without --apply: this validates the package, every SHA-1-to-Android-OAuth client match, Firebase/FCM project consistency, the resulting App Links set, and the production Android environment while leaving .env byte-for-byte unchanged. Apply the same validated candidate only after reviewing the counts:

chmod 600 /secure/downloads/google-services.json \
  /secure/downloads/play-identities.json

./scripts/import-play-android-config.sh \
  .env \
  /secure/downloads/google-services.json \
  /secure/downloads/play-identities.json

./scripts/import-play-android-config.sh \
  .env \
  /secure/downloads/google-services.json \
  /secure/downloads/play-identities.json \
  --apply

The apply step is atomic. It preserves the existing upload certificate and Android OAuth client, adds every Play delivery identity, refreshes the four public Firebase Android values, and never prints OAuth client IDs or the Firebase API key. It refuses a test/development environment, another package, an unmatched or duplicate certificate, and a Firebase project inconsistent with the configured FCM service account. Provider files remain on disk after the import and must be stored or removed deliberately.

Once the production association has been deployed and the internal-track build has been installed from Google Play, prove that the physical device is not still running a sideloaded upload-key build:

./scripts/verify-play-installed-android.sh \
  /secure/downloads/play-identities.json \
  DEVICE_SERIAL \
  EXPECTED_VERSION_CODE \
  EXPECTED_VERSION_NAME

The command does not launch, install, uninstall, clear, or reconfigure the app. It checks the installer, installed version, Play signing identity, Android's domain-verification state, and implicit production App Link resolution. Passing it is a prerequisite for the later authenticated Play-delivered smoke test, not a substitute for that test.

The VAPID helper runs the exact locked web_push_elixir generator in an isolated, network-disabled container, imports the result atomically, removes its one-run image tag and temporary files, and never prints either key. It refuses to replace an existing VAPID identity because an unplanned rotation invalidates existing browser subscriptions.

The importers validate the exact OAuth callback, Android package, Firebase project relationship, and required service-account fields before atomically replacing existing keys. They preserve mode 0600 and never create another permanent environment file. Provider and initializer values that cannot be represented as one unquoted Compose .env line are rejected before mutation. Literal dollar signs are stored as $$, which Compose resolves back to one $ inside the container. OAuth and service-account downloads still contain private credentials after import; deliberately move them to protected backup storage or remove them after verification.

The first command reports incomplete or local-only capabilities. The second is a blocking release preflight and exits nonzero until Google sign-in, external SMTP, browser Web Push, Android Firebase/FCM, App Links, support routing, Android version/signing aliases, and the core application configuration are all complete.

Create the test configuration inside the test checkout:

cd /srv/who_need_help-test
TEST_CODEX_SESSION_ID=YOUR_MAIN_CODEX_SESSION_ID \
TEST_GOOGLE_OAUTH_CLIENT_ID=YOUR_TEST_CLIENT_ID \
TEST_GOOGLE_OAUTH_CLIENT_SECRET=YOUR_TEST_CLIENT_SECRET \
TEST_GOOGLE_OAUTH_AUTHORIZED_PARTY_IDS=YOUR_TEST_ANDROID_CLIENT_ID \
  ./scripts/init-test-env.sh test.whoneedhelp.com
./scripts/validate-test-env.sh .env test.whoneedhelp.com
./scripts/deploy-up.sh .env

For later test updates, check out the desired clean revision and update only the image tags. Existing deployment secrets remain unchanged:

./scripts/set-deployment-revision.sh .env
./scripts/deploy-up.sh .env

Test always uses its own who_need_help_test PostGIS container/volume. Local development can use Mailpit; a public test deployment must use its own SMTP password and a visibly test-specific sender identity.

Create the production configuration inside the production checkout after loading only the production database/provider credentials into the process:

cd /srv/who_need_help-production
PRODUCTION_DATABASE_MODE=external \
PRODUCTION_DATABASE_URL='ecto://PRODUCTION_ROLE:PASSWORD@localhost/who_need_help_production' \
PRODUCTION_DATABASE_SOCKET_DIR=/var/run/postgresql \
PRODUCTION_EMAIL_DELIVERY_PROVIDER=smtp \
PRODUCTION_SMTP_RELAY=smtp.example.net \
PRODUCTION_SMTP_PORT=587 \
PRODUCTION_SMTP_USERNAME=YOUR_PRODUCTION_SMTP_LOGIN \
PRODUCTION_SMTP_PASSWORD=YOUR_PRODUCTION_SMTP_PASSWORD \
PRODUCTION_GOOGLE_OAUTH_CLIENT_ID=YOUR_PRODUCTION_CLIENT_ID \
PRODUCTION_GOOGLE_OAUTH_CLIENT_SECRET=YOUR_PRODUCTION_CLIENT_SECRET \
PRODUCTION_GOOGLE_OAUTH_AUTHORIZED_PARTY_IDS=YOUR_PRODUCTION_ANDROID_CLIENT_ID \
PRODUCTION_CODEX_SESSION_ID=YOUR_MAIN_CODEX_SESSION_ID \
  ./scripts/init-production-env.sh whoneedhelp.com
./scripts/validate-production-env.sh .env whoneedhelp.com

Promote the exact revision already verified in test by checking out that SHA in the independent production clone, then update only the production image tags:

git checkout --detach <verified-test-sha>
./scripts/set-deployment-revision.sh .env
./scripts/deploy-up.sh .env

Before either public switch, prove cross-environment isolation without printing credentials:

/srv/who_need_help-test/scripts/validate-deployment-isolation.sh \
  /srv/who_need_help-test /srv/who_need_help-production

Start/update test first and run the complete browser/API/Android verification. Only then check out that exact tested SHA in production. Start the application with ./scripts/deploy-up.sh .env. The application command joins the external public network but does not own or restart Caddy.

The legacy ./scripts/edge-up.sh .env command exists only for the currently frozen submitted deployment. Do not use it from an ordinary application release. After judging, render and validate each route through the independent server_edge workflow, transfer certificate-volume ownership in a reviewed maintenance window, and only then retire the legacy edge container.

The authoritative A records for whoneedhelp.com, www.whoneedhelp.com, and test.whoneedhelp.com must point to the verified server address before Caddy can obtain their certificates. Do not remove the edge volumes during an ordinary application deploy. Do not add --volumes to a stop command unless the exact database has been inspected and deletion is intended.

The initializer reads .env.example, derives the numeric group of that host's Docker socket, generates independent random values for PostgreSQL, SECRET_KEY_BASE, handover codes, the BEAM release cookie, and metrics access, and writes an ignored mode-0600 file. It does not print those values and refuses to overwrite an existing destination.

By default, the generated public proxy port binds to 127.0.0.1, which is appropriate only when the verified reverse proxy reaches the application on the same host. Set PRODUCTION_HTTP_BIND_ADDRESS when generating the file, or edit HTTP_BIND_ADDRESS afterward, to match the observed target topology. Replace TRAEFIK_TRUSTED_IPS with the exact source IP/CIDR observed at Traefik; do not copy the temporary VPN value into an unrelated server.

Configure the transactional SMTP relay and a sender accepted by it using EMAIL_DELIVERY_PROVIDER=smtp plus the SMTP_* settings. The initializer accepts the corresponding PRODUCTION_EMAIL_DELIVERY_PROVIDER and PRODUCTION_SMTP_* inputs. Set optional SUPPORT_INBOX_ADDRESS to a monitored address for support/removal queue alerts and email Reply-To. Public support creates an alert only after email verification; authenticated support is already verified, while content-removal notification rules remain separate. The alert is queued in mail; conversation messages remain in the permission-scoped staff workspace rather than producing email per message. Leaving the value empty disables operator email alerts, not the protected queues. If Google registration/sign-in is enabled, also set both Google Web client credentials and register https://YOUR_PHX_HOST/auth/google/callback as the exact authorized redirect URI. Leave both credentials empty to keep the feature disabled. Then run:

SUPPORT_INBOX_ADDRESS does not configure inbound delivery. To opt into the Brevo inbound-parse boundary, set SUPPORT_INBOUND_RECIPIENT to the exact mailbox on a dedicated receiving subdomain. Set SUPPORT_INBOUND_WEBHOOK_TOKEN to an independent secret, or let the production initializer generate it when a recipient is supplied. Configure Brevo to send a secured webhook with Authorization: Bearer <that token> to https://YOUR_PHX_HOST/webhooks/brevo/inbound-support. Follow Brevo's current MX instructions for the receiving subdomain, then verify a real inbound message and the subsequent mailbox-confirmation link. Inbound messages stay outside the staff queue until confirmation; provider attachment tokens are deliberately not downloaded. See docs/support-and-content-removal.md for the trust boundary.

./scripts/validate-production-env.sh .env whoneedhelp.com
./scripts/deploy-up.sh .env

The production override keeps Mailpit behind its inactive local-mail profile, so public registration cannot appear to succeed while mail is only retained locally. The validator checks file ownership/mode, origin consistency, template markers, independent generated secrets, settings required by the selected email provider, a complete-or-empty Google credential pair, the generated database URL, and the final Compose render without printing secrets. It does not contact DNS, TLS, the email provider, the reverse proxy, or the application. After deployment, verify /healthz/ready, inspect all replica health and logs, register a unique address through the public browser, receive its message at the real mailbox, follow the HTTPS confirmation link, and remove only that run-scoped account.

External PostgreSQL/PostGIS

PostgreSQL on the same Linux host

The server-local mode uses the PostgreSQL Unix socket instead of exposing the database on a Docker-reachable TCP address. Ecto/Postgrex receive DATABASE_SOCKET_DIR; scripts/compose.sh then adds compose.external-db-socket.yaml and mounts that exact directory read-only into the active application and migration services. An empty setting preserves the normal remote-provider TCP behavior.

After read-only inspection confirms the intended PostgreSQL 18 cluster, installed PostGIS package, socket directory, and absence of the project-scoped roles/databases, run the root-only provisioner:

sudo ./scripts/provision-host-postgres.sh "$USER"

Its exact mutation scope is:

  • prepend two database-and-role-specific local ... scram-sha-256 rules to the active pg_hba.conf, retaining a mode-0600 copy under /var/backups/who_need_help/, and reload that cluster;
  • create login roles wnh_production and wnh_test without superuser, database-creation, role-creation, replication, or row-security bypass rights;
  • create empty owner databases who_need_help_production and who_need_help_test, revoke public connect, and preload citext and PostGIS so the non-superuser Ecto migrations can run;
  • verify both credentials through the observed Unix socket and write separate mode-0600 initializer fragments under ~/.config/who_need_help/ without printing passwords.

The command refuses to overwrite credential fragments, refuses any matching pre-existing role/database or managed HBA marker, validates the candidate HBA rules before reload, and restores its HBA backup plus removes only objects it created if provisioning fails. It does not inspect or migrate application tables; the release migration runner remains authoritative for schema.

Generate test or production after loading only the matching trusted fragment:

set -a
. "$HOME/.config/who_need_help/database-production.env"
set +a
./scripts/init-production-env.sh whoneedhelp.com
unset PRODUCTION_DATABASE_MODE PRODUCTION_DATABASE_URL \
  PRODUCTION_DATABASE_SOCKET_DIR

For test external-database drills, load database-test.env; the public test deployment normally uses its own Compose database. Do not load both fragments into one shell.

Provision the database and role first, then generate the environment without placing its credentials on a command line that is retained in shell history:

export PRODUCTION_DATABASE_MODE=external
read -rsp 'External DATABASE_URL: ' PRODUCTION_DATABASE_URL && echo
export PRODUCTION_DATABASE_URL
./scripts/init-production-env.sh whoneedhelp.com
unset PRODUCTION_DATABASE_URL

Set PRODUCTION_APP_TOPOLOGY=split during generation when independent web and worker scaling is required. In external mode scripts/compose-up.sh builds the migration image, runs scripts/check-database.sh, and fails before migration or application startup unless PostgreSQL and PostGIS_Version() are reachable. The check also reports the observed TLS flag. The connection URL must follow the database provider's verified TLS/CA requirements.

scripts/compose.sh .env config --services is the canonical read-only render check. In external mode its output must not contain db. Use scripts/compose.sh .env ps -a and scripts/compose.sh .env logs web worker for the selected mode.

The repository's Compose backup, restore, rotation, and rollback-based staging scripts deliberately refuse DATABASE_MODE=external; they are scoped to the project-owned database container. For an external database, use the provider's verified backup/restore procedure and test recovery before public launch.

Compose database backup

Create a PostgreSQL 18 custom-format archive, validate its table of contents, and write a SHA-256 manifest:

./scripts/backup-compose.sh

The default destination is the ignored output/backups/ directory. An explicit new destination may be supplied as the only argument. The command refuses to overwrite either an archive or its checksum manifest and writes through temporary files before publishing the final pair. Files and newly created directories are restricted by umask 077.

The archive covers the configured application database. PostgreSQL cluster globals such as roles and tablespaces are not part of pg_dump; deployment credentials and database roles must be provisioned separately from secrets. Local backup files on the same workstation are not an off-site backup.

Encrypted local S3-compatible backup drill

The isolated load project can run a complete encrypted Restic/MinIO drill:

./scripts/load-stack-up.sh
./scripts/backup-s3-drill.sh local-encrypted-backup

The script refuses the staging Compose project and validates the project and service labels of every pre-existing container in its scope. On first use, scripts/ensure-local-load-env.sh generates independent random MinIO and Restic credentials in ignored output/runtime/load.env and restricts that file to mode 0600. MinIO publishes Docker-assigned ports only on 127.0.0.1; the observed API and console URLs are printed after a successful run.

The backup tool combines the PostgreSQL 18.6 client with pinned Restic rebuilt on Go 1.26.7. MinIO server and client are also rebuilt as non-root Alpine images from checksum-pinned upstream source commits with the exact dependency updates recorded in Dockerfile.minio. The quality gate verifies their reported release, commit, Go runtime, configured user, and current HIGH/CRITICAL vulnerability scan. restic backup --stdin-from-command runs a custom-format pg_dump, checks the producer exit status, encrypts the data, and uploads it directly to MinIO. No plaintext database dump is written to the host. The drill then:

  1. runs restic check --read-data;
  2. streams restic dump into pg_restore --list;
  3. restores into a uniquely named database created from template0;
  4. checks tables, current Ecto migrations, PostGIS, and release migration readiness before removing that exact database;
  5. clones and corrupts an isolated repository and requires both check and dump to fail;
  6. stops an exact scoped in-progress backup container only after encrypted objects reach MinIO, requires zero published snapshots, prunes unreferenced packs, and rechecks the repository;
  7. removes and verifies removal of the corruption/interruption buckets and requires source table counts to remain unchanged.

The successful encrypted bucket is deliberately retained in the named local MinIO volume. Non-secret evidence is written under ignored output/backups-s3/<run-label>/; the runtime scratch directory is under ignored tmp/backup-s3/. Use a unique lowercase run label of at most 32 characters. The command refuses to replace an existing retained bucket.

MinIO server and client are AGPLv3. Dockerfile.minio identifies the exact upstream source commits and contains the dependency changes and complete build commands used here. Before distributing or publicly operating modified images, review the license and make the corresponding source available as required; this runbook does not provide legal advice.

This verifies encryption, local S3 protocol use, restore mechanics, and two failure paths on the observed workstation. A MinIO volume on that same workstation is not an off-site backup and does not establish production RPO, RTO, retention, capacity, key custody, object locking, or database HA.

Encrypted production backup to an independent host

The production workflow streams an external PostgreSQL 18 custom-format dump through an encrypted Restic repository on an SSH target that must resolve to a host other than production. It then runs restic check --read-data and restores the newest snapshot into a uniquely named, temporary PostGIS container created from template0. The source production database is never a restore target.

Bootstrap the checksum-pinned Restic binary and create the ignored local operations configuration:

./scripts/bootstrap-restic.sh
./scripts/init-production-operations.sh

The initializer creates only these ignored local files:

  • tmp/production-operations/backup.env, mode 0600;
  • tmp/production-operations/restic-password, mode 0600.

The generated configuration uses the reviewed whoneedhelp production SSH target and the independent buyvm-maya SFTP target. Inspect the read-only plan before repository initialization:

./scripts/production-offsite-backup.sh plan

Repository initialization is an external mutation and requires the exact repository value printed by the plan:

WNH_OFFSITE_BACKUP_INIT_CONFIRM='sftp:buyvm-maya:backups/who_need_help-production' \
  ./scripts/production-offsite-backup.sh init

The routine operations are:

./scripts/production-offsite-backup.sh run
./scripts/production-offsite-backup.sh check
./scripts/production-offsite-backup.sh restore-drill

run creates a fresh source dump, copies it into the Restic input, publishes one encrypted snapshot only after the producer succeeds, performs a full-data repository check, and runs the isolated restore drill. It removes its exact temporary source dump, local scratch directory, Docker container, and Docker volume on success, failure, or interrupt. Non-secret catalogs, checksums, and restore observations remain under ignored output/production-operations/<run-id>/.

After all of those checks pass, run atomically publishes a mode-0600 heartbeat on the independent repository host. The heartbeat contains only the Restic snapshot identifier, verification time, and restore_verified: true; it contains no database rows, account identifiers, recipient addresses, or credentials. A failed backup or failed restore never advances the heartbeat.

Install the daily user-systemd timer on the operator workstation only after one manual run succeeds. The production host does not run this timer. The workstation connects to production for the source dump and publishes the restore-verified heartbeat to the independent repository host:

./scripts/install-production-backup-timer.sh
systemctl --user list-timers --all who-need-help-production-backup.timer
systemctl --user status who-need-help-production-backup.service --no-pager
journalctl --user -u who-need-help-production-backup.service --no-pager

The timer uses Persistent=true, so systemd runs a missed calendar invocation after the workstation and user manager return. A failed service invokes the external monitor's backup-failure notification. It does not retry, delete snapshots, or infer a retention policy.

The Restic password is required for every restore. Before relying on this as the only recovery copy, place it in an operator-controlled password manager or offline recovery record. Never print it into a terminal log, commit it, or copy it into the backup repository. There is intentionally no automatic forget/prune policy yet: retention, RPO, RTO, capacity, and key custody are operator decisions, and the script does not invent them.

Independent production monitor

The external monitor runs on an SSH target that must resolve to a host other than production. It requests the public readiness endpoint and accepts only HTTP 200 with the exact JSON object {"status":"ready"}. It also scrapes the authenticated /metrics endpoint with the independent METRICS_TOKEN copied from the production environment into the external host's mode-0600 configuration. Readiness or metrics-scrape failure changes the monitor state to down. Email is sent once when that state changes to down and once when it recovers; repeated down checks do not send repeated alerts.

The monitor can also treat a missing, invalid, unverified, or stale backup heartbeat as down. This is intentionally not enabled with an invented default: the maximum acceptable age is the operator's RPO/alerting decision. After that decision is recorded, reinstall the monitor with the chosen positive number of seconds:

MONITOR_BACKUP_MAX_AGE_SECONDS=<operator-selected-seconds> \
  ./scripts/install-production-external-monitor.sh

The installed configuration then records the exact threshold and heartbeat path. A stale backup produces one state-transition alert; a later successful, restore-verified backup produces one recovery notification.

For each observed application node, the monitor keeps a baseline of aggregate HTTP exception, failed Oban job, and failed/exceptional email-delivery counters. It sends one aggregate notification when one of those counters increases. A new node establishes a baseline without alerting, and a lower value is treated as a counter reset rather than a failure. The metric labels are deliberately bounded to queue, delivery status, and the fixed application-owned email purposes listed by WhoNeedHelp.EmailDelivery. This lets an operator distinguish an authentication-delivery failure from a support-update failure without exposing recipient addresses, message content, request payloads, or user identifiers. Unknown email-purpose labels are ignored rather than creating unbounded monitor state. With multiple application replicas, this external check samples only the node that answers each request; use the full Prometheus deployment when every replica must be scraped continuously.

Install it from the production SMTP configuration without printing the SMTP credential:

./scripts/install-production-external-monitor.sh
ssh buyvm-maya \
  'systemctl --user status who-need-help-production-monitor.timer --no-pager'
ssh buyvm-maya \
  'journalctl --user -u who-need-help-production-monitor.service --no-pager'

The default schedule is once per minute with separate three-second readiness and metrics timeouts. Those defaults match the current application container health timeout and were installed only after public readiness requests from the selected monitor host were observed to finish below one second. They are observations of the current path, not universal capacity or availability guarantees. Re-running the installer refreshes the exact script, mode-0600 configuration, and user units, then performs one check before enabling the timer.

By default the installer mirrors the production application's SMTP credential. For the pilot, use an independently revocable monitor SMTP key so revoking or rotating the application key does not also disable operational alerts. Keep the override outside the repository in a mode-0600 (or read-only mode-0400) file containing exactly these keys:

SMTP_RELAY=smtp-relay.example
SMTP_PORT=587
SMTP_USERNAME=independent-monitor-login
SMTP_PASSWORD=secret-from-the-provider
SMTP_TLS=always
SMTP_SSL=false
EMAIL_FROM_ADDRESS=monitor@whoneedhelp.com
EMAIL_FROM_NAME=Who Need Help monitor
SUPPORT_INBOX_ADDRESS=monitored-operator@example.com

Create that file through the password manager or an editor that does not place the secret in shell history. When the override is present, the installer reads only METRICS_TOKEN from production and never copies the application's SMTP credential into its local staging configuration. Then reinstall the monitor with the scoped override and send one test notification before revoking any previous key:

MONITOR_SMTP_VALUES_FILE="$HOME/.config/who-need-help/monitor-smtp.env" \
  ./scripts/install-production-external-monitor.sh
ssh buyvm-maya \
  '/usr/bin/python3 ~/.local/lib/who-need-help/production-external-monitor.py \
    --config ~/.config/who-need-help/monitor.json \
    --state ~/.local/state/who-need-help/monitor.json \
    send-test-notification'

Verify receipt in the monitored mailbox. Only then revoke the former monitor key. The application SMTP key is a separate rotation: validate the new key, atomically update the production .env, recreate only the application services so they read the new runtime value, verify an application-generated authentication message and readiness, and revoke the former application key last. If any check fails before revocation, restore the prior mode-0600 environment and recreate the same services; do not leave application and monitor configurations half-updated.

Local external-service boundary drill

Run the OAuth, SMTP, and provider-neutral push protocol checks without public credentials, a server, or host-published ports:

./scripts/external-boundaries-run.sh local-boundaries

The script creates a uniquely named Compose project on an internal-only Docker network. It generates independent one-run OAuth and push credentials in an ignored mode-0600 environment file, including a separate bearer token for the protected metrics endpoint, builds the production release plus a non-root standard-library Python protocol mock, and then verifies:

compose.external-boundaries.yaml is the only deployment file that sets ALLOW_INSECURE_EXTERNAL_HTTP=true. Ordinary runtime configuration requires HTTPS for OAuth and push endpoints so provider credentials are not sent over plain HTTP.

  1. GitHub-compatible OAuth authorization, PKCE S256, token exchange, normalized user lookup, state mismatch, provider rejection, one-time-code replay, a fresh flow after a temporary token error, and a token timeout;
  2. SMTP acceptance, permanent recipient rejection without retry, one retry after a temporary greeting failure, a greeting timeout, and the result of submitting the same message twice;
  3. the disabled default push boundary plus HTTP success, permanent rejection, one temporary retry, replay deduplication, and deduplication after an ambiguous timeout using the same idempotency key;
  4. a fresh PostGIS database, current migrations, and two Oban worker replicas; request acceptance and new-chat domain transactions enqueue stable user-recipient events, replay is deduplicated before HTTP, an injected temporary chat delivery fails its first Oban attempt and completes on its second, and private message text is absent from the push payload.

The evidence JSON and mock state contain counters, booleans, normalized identity fields, payload digests, run-scoped user/event identifiers, and the privacy-safe notification metadata asserted by the drill. The script fails if any generated secret appears in retained evidence. Its trap validates exact Compose labels, removes only that project, volumes and one-run images, and deletes the temporary credential file. Failure logs are retained under the same ignored evidence directory.

This drill uses the application's real Assent/Req and Swoosh/gen_smtp clients, but the providers are local. It therefore verifies client-side protocol and product integration, not GitHub, SMTP-provider, FCM, or APNs availability. SMTP permits duplicate delivery after ambiguous outcomes, so the result explicitly makes no exactly-once claim. Push currently targets a stable user:<uuid> recipient; selecting a provider, registering device tokens, and resolving that user to devices remain deployment/provider work.

The implementation follows the configured adapter interfaces and protocol semantics documented by Assent 0.3.1, GitHub OAuth, and SMTP RFC 5321.

The application uses Swoosh's SMTP adapter and gen_smtp. Verify the selected relay's authenticated delivery, sender-domain authentication, and delivered message headers before launch. Test and production may share an account-level SMTP login only when they use independently revocable SMTP passwords.

Isolated restore drill

Run a real restore into a uniquely named temporary database:

./scripts/restore-drill-compose.sh output/backups/compose-YYYYMMDD-HHMMSS.dump

The drill:

  1. validates the SHA-256 manifest;
  2. validates the archive table of contents;
  3. creates a pristine database from template0;
  4. restores with pg_restore --exit-on-error;
  5. reads every restored public application table and checks the PostGIS library;
  6. runs the current immutable release's migrations and migration-readiness check against only the temporary database;
  7. drops only the temporary drill database and verifies that it is gone.

A trap also attempts to drop the exact temporary database if a check fails. The source application database is never passed to pg_restore, dropdb, a clean operation, or the drill migration runner. The drill intentionally does not compare an older backup's row counts to the live source, because concurrent legitimate writes or a historical archive would make that comparison invalid.

Isolated Compose upgrade rehearsal

After creating a current custom-format backup, run the complete current release against an isolated restored copy:

./scripts/upgrade-rehearsal-compose.sh \
  output/backups/compose-YYYYMMDD-HHMMSS.dump

The rehearsal validates the checksum and archive catalog, reads the public origin configuration from ignored .env, and uses only the independently generated credentials in ignored mode-0600 output/runtime/e2e.env. It builds a uniquely tagged production release, creates a uniquely named Compose project and database from template0, restores the archive, records application-table counts, and then:

  1. applies every current timestamped Ecto migration;
  2. requires the localized category column and all 11 valid cursor indexes;
  3. starts 2 web and 2 worker replicas behind the isolated Traefik instance;
  4. requires the four-node BEAM cluster and cross-node PubSub probe;
  5. checks the production HTTP-to-HTTPS redirect, trusted-proxy public pages, and both health endpoints;
  6. requires an empty before/after diff for every public application table except the expected migration and Oban-internal tables;
  7. removes and verifies removal of the exact containers, networks, database volume, and one-run image.

The input archive is read-only and is not copied into the evidence directory. The ordinary Compose project, source database, public route, and running containers are outside the generated project scope. Non-secret evidence is retained under ignored mode-0700 output/upgrade-rehearsal/<run-id>/, with files mode 0600.

Clean tracked-revision deployment drill

Run:

./scripts/clean-deploy-verify.sh

The script archives the current tracked Git revision into a new temporary directory. It verifies that the archive contains no local .env, Git metadata, generated output, or existing E2E/load environment. It generates independent one-run PostgreSQL, Phoenix, handover, BEAM-cookie, and metrics secrets with mode 0600; optional GitHub OAuth and push delivery stay disabled.

The drill uses a unique Compose project, application image, Docker socket proxy image, PostGIS image, Traefik image, router/service name, internal network, PostgreSQL volume, database, and dynamic host ports. It requires exactly two healthy web and two running worker replicas, all tracked migrations, seeded categories, an idempotent repeated migration, working readiness/home/registration/Mailpit routes, four connected BEAM nodes, and a cross-replica PubSub message. It then removes the exact project including volumes, all four one-run images, and the temporary archive, and verifies that every scoped resource is absent. If Compose validation fails before project creation, cleanup avoids invoking an unrenderable Compose configuration and still verifies that the generated scope is absent. Evidence is retained under output/portability/<run-id>/.

This is a local Docker portability observation. It does not establish production SMTP, OAuth, push-provider, TLS, database-HA, storage, capacity, or jurisdictional readiness.

Service checks

docker compose ps
curl --fail http://localhost:4010/healthz/live
curl --fail http://localhost:4010/healthz/ready
./scripts/verify-realtime-cluster.sh compose

live verifies that the web process can serve HTTP. ready additionally runs SELECT 1 through the configured Ecto repository. The cluster probe subscribes on one connected BEAM node and broadcasts from another.

Health checks do not replace alerting, database backups, restore drills, or application-level synthetic checks.

Local failure and rolling-replacement drills

The isolated load project can exercise process crashes, sequential container replacement, and a real Oban retry without touching the normal Compose project:

./scripts/load-cycle.sh local-resilience resilience

The lifecycle wrapper assigns unique Compose and image names and removes that exact run's containers, networks, named volumes, temporary environment, and image tags on success, failure, or interruption. Use load-stack-up.sh plus load-resilience-run.sh only when the isolated stack must remain available for manual inspection.

The resilience script refuses LOAD_PROJECT=who_need_help and verifies the Compose project/service labels of every container before stopping it. It:

  1. continuously calls readiness through the isolated Traefik route;
  2. terminates the BEAM process in one web and one worker container and requires Docker's observed restart count to increase;
  3. removes and replaces each web and worker replica one at a time;
  4. waits for every configured BEAM node, then runs the cross-node PubSub probe;
  5. enqueues a side-effect-free local worker that fails its first Oban attempt and succeeds on its second;
  6. removes that exact Oban row and requires no fixture domain rows to remain.

Traefik's retry middleware is attached to the HTTP and local TLS routers. Its attempt count is an environment input. Traefik retries transport failures and, with the checked configuration, does not opt in to retrying non-idempotent requests. This reduces a stale-backend window; it is not a claim of production availability.

For the project-owned kind cluster, run:

./scripts/kind-rolling-verify.sh local-kind-rollout

The kind control plane is an optional verification environment, not part of the ordinary public Compose route. Pause it without deleting its container or Kubernetes state when the cluster is not being tested:

./scripts/kind-stop.sh

The next ./scripts/kind-up.sh validates the ownership marker and kind labels, starts the stopped control plane, then continues with image loading and chart reconciliation.

That script requires both the kind ownership marker and the control-plane cluster label before invoking rollout restart. It changes only the web and worker Deployment pod templates. It snapshots application-table counts before and after, continuously probes the observed Docker mapping for the chart's NodePort, requires all four old pod UIDs to disappear, waits for the exact BEAM peer count, and verifies cross-node PubSub. PostGIS, its hostPath, the Kubernetes Secret, and the namespace are not recreated.

The rollout timeout, probe interval/timeout/retry count, and cluster-join timeout are experiment inputs. They are not production SLOs or resource requirements.

Isolated Oban burst measurement

The worker role consumes maintenance, push, and mail. Configure their per-worker limits with OBAN_MAINTENANCE_CONCURRENCY, OBAN_PUSH_CONCURRENCY, and OBAN_MAIL_CONCURRENCY; multiplying a value by the number of worker replicas gives the configured cluster-wide concurrency for that queue. Authentication links remain synchronous because their database token is committed only when delivery succeeds. Support-contact verification, support updates, content-removal confirmation/receipt/decision messages, and optional new-case operator alerts run in mail, so their SMTP latency does not occupy a web request or the push queue. Requester and staff conversation messages do not each generate email, and internal assignment-only legal updates do not notify the submitter.

After starting the isolated load project, run an explicitly sized experiment:

./scripts/load-stack-up.sh
./scripts/oban-burst-run.sh local-oban-burst 1000 120

The three required arguments are an evidence label, job count, and recorded experiment timeout in seconds. They are not capacity thresholds. The script refuses the ordinary who_need_help project, verifies the expected worker replicas, inserts only confirmed LocalBurstProbe jobs with a unique run id, records each worker's effective queue configuration and container samples, requires every job to complete, compares domain-table row counts, and deletes exactly its own jobs. The probe worker has no product side effects and no product flow enqueues it.

Kubernetes ingress isolation

networkPolicy.enabled=true renders the chart's ingress NetworkPolicy. Enforcement is a property of the cluster CNI, not of the YAML object alone. Before relying on it, verify that the target cluster uses a NetworkPolicy- capable plugin and run positive HTTP/cluster checks plus negative blocked-port checks there. Egress remains intentionally unrestricted until the actual database and external-service destinations are known.

BEAM runtime memory guard

Compose sets ERL_ZFLAGS="+Q ${ERLANG_PORT_LIMIT}"; Helm renders the same flag from app.erlangPortLimit. Both default to 65536, OTP's normal port-table limit. This makes the runtime independent of an unusually large host or nested container nofile limit. Validate all live replicas after deployment:

./scripts/verify-beam-runtime.sh compose
./scripts/verify-beam-runtime.sh kind

The command records the effective port count/limit, allocated port-table bytes, BEAM memory, process count, cgroup memory, and RSS for each web and worker replica. It fails when a live VM does not use the configured limit; kind mode also requires the desired number of Ready application pods.

65536 is a concurrency ceiling for simultaneously existing Erlang ports (files, sockets, and drivers), not a container memory limit. Do not lower or raise it from a RAM estimate alone. A changed value must be validated against measured peak port usage and the target environment.

Protected Prometheus metrics

The web role exposes Prometheus text format at /metrics. It requires the independent METRICS_TOKEN deployment secret:

curl --fail \
  --header "Authorization: Bearer $METRICS_TOKEN" \
  http://localhost:4010/metrics

The endpoint returns 401 without the exact token, disables response caching, and does not put the credential in a URL. The reporter exports cumulative HTTP request and duration, router exception, database query and duration, WebSocket connection, Oban attempt and stop outcome, aggregate single-email delivery outcome, VM memory, and scheduler run-queue metrics. The Oban outcome series distinguishes the fixed success, cancelled, discard, and snoozed stop states without job arguments or identifiers; exception attempts remain a separate counter. Email metrics retain the adapter-level ok/error counter and exception counter, and also expose the application's fixed allow-listed delivery purpose together with ok, error, or exception. Purpose labels are code-defined values such as auth_login, support_confirmation, or content_removal_update; recipients, subjects, references, and message bodies are never labels. Cumulative durations are integer microseconds because the selected reporter's sum accumulator is integer-based; divide by 1_000_000 in PromQL when seconds are required. Definitions intentionally have no request path, user, request, recipient, message, or event-name labels that could create unbounded cardinality or expose private data.

The locked telemetry_metrics_prometheus_core reporter aggregates distribution samples only when a scrape occurs. Each application VM therefore also runs a supervised internal aggregation every ten seconds. This keeps the reporter's raw distribution table independent of whether an external Prometheus server is currently configured, while preserving the cumulative histograms exposed by the authenticated endpoint. The generated exposition text from the internal aggregation is discarded.

After a deployment, the raw table can be inspected without exposing metric credentials:

/app/bin/who_need_help eval \
  'IO.inspect(:ets.info(:prometheus_metrics_dist, :size), label: "pending_distribution_samples")'

Under idle conditions the value returns to zero after the next aggregation interval. During traffic it represents only samples received since the most recent internal or external scrape; it is not the cumulative histogram count.

Metrics are local to each BEAM process. Discover and scrape every web pod or container as a distinct target and preserve Prometheus's instance label. A request through the load-balanced public route reaches only one replica and is therefore useful as an authorization/smoke check, not as a cluster-wide aggregate.

The isolated load project includes a local observability profile:

./scripts/load-stack-up.sh
./scripts/observability-run.sh local-observability

The run script refuses the staging project, validates every current web container's Compose labels, and writes a file_sd target for each observed internal IP. Prometheus reads the Bearer value from a mode-0600 runtime file, not a tracked config or URL. Its direct request includes the internal X-Forwarded-Proto: https signal required by the application's production SSL rewrite while preserving the target's own instance label.

Prometheus, Alertmanager, and Grafana are pinned by tag and digest. Their host ports default to Docker-assigned values bound only to 127.0.0.1; the run prints the observed URLs. Grafana uses the random admin password generated in ignored output/runtime/load.env, disables anonymous signup, update checks, suggested plugin installation, and its unused built-in alert engine. The Prometheus datasource and ten-panel dashboard are provisioned from tracked files. The dashboard separates all discovered web and worker replicas and includes HTTP traffic/latency, BEAM memory and scheduler run queues, Ecto execution and pool-wait latency, and Oban outcomes and queue wait by queue.

The verification stops exactly one scoped load web container. The WhoNeedHelpWebReplicaUnavailable rule is based only on the factual up == 0 result; it is a local failure drill, not an invented latency, capacity, or production SLO threshold. The script requires both firing and resolved webhook payloads from Alertmanager, starts the same container, waits for every direct target, and compares read-only database counts before and after. Evidence is retained in output/observability/ without the metrics or Grafana secrets.

Stop only the monitoring services with:

./scripts/observability-stop.sh

Prometheus/Grafana/Alertmanager retention, production notification destinations, production availability, and measured alert policies remain deployment decisions. In Kubernetes, put the metrics token in existingSecret; configure the external scraper to send it as a Bearer token.

The provisioned Grafana overview also shows one-hour transactional-email attempts grouped only by the fixed application-owned purpose and outcome. It does not expose recipient addresses, subjects, message bodies, support references, or user identifiers. The panel is intended to answer which bounded workflow is creating delivery volume or failures; it is not a user-activity log.

The metrics endpoint also exposes cumulative rate-limit bucket checks grouped only by the configured action name and the bounded allowed or limited outcome. It never labels a metric with the hashed scope, email address, client address, user identifier, bucket count, or reset timestamp. Use these counters to observe which pilot policies affect real traffic before changing their limits; the counters do not by themselves establish an abuse policy or a safe capacity threshold.

Support conversations deliberately do not send one email per staff message. The requester receives an email for the first staff response and for later public status changes; additional messages while the status is unchanged stay in the private support conversation and the in-product notification flow. Anonymous submissions receive the address-confirmation message before they enter the operator queue. Operator email alerts are disabled unless SUPPORT_OPERATOR_EMAIL_MODE=immediate is explicitly configured.

Run-scoped production browser verification

The production E2E workflow is opt-in and starts with a read-only identity and scope check:

./scripts/production-full-e2e.sh plan production-e2e-YYYYMMDD

The plan verifies the production checkout, compact Compose project, external database identity, healthy application container, and absence of an existing fixture with the requested run ID. It prints an exact confirmation value but does not create users or records.

After that exact scope has been reviewed and explicitly authorised, use the printed value without changing the run ID:

WNH_PRODUCTION_E2E_CONFIRM='VALUE_PRINTED_BY_PLAN' \
  ./scripts/production-full-e2e.sh run production-e2e-YYYYMMDD

The run creates six uniquely prefixed synthetic users and only their associated mutual-aid, activity, notification, audit, support, and legal fixture records. The support and legal records are inserted directly without email jobs and are only read through the staff UI; the browser does not submit or moderate them. Cleanup uses the mode-0600 manifest of exact IDs on success, failure, or interrupt, refuses cross-fixture relationships, deletes only matching jobs and records, and verifies that the run prefix is absent. It never resets the database. Evidence is stored below output/production-full-e2e/<run-id>/.

Production release without pushing the frozen repository

The post-submission workflow keeps the public Git repository and test.whoneedhelp.com untouched. A clean local commit is packaged as a verified Git bundle and transferred directly over SSH to only /srv/who_need_help-production:

./scripts/production-release.sh plan whoneedhelp

The default plan action is read-only. It verifies the exact local and remote commits, requires a fast-forward history, checks the production checkout, Compose scope and healthy containers, checks public readiness, opens a read-only PostgreSQL connection, and runs the server-release environment capability preflight. The candidate validator is streamed over SSH and checks the existing server environment and Compose render before any bundle is uploaded, so a validator change does not depend on the previously deployed source tree. Before the first Google Play release, this preflight allows only the absent Play App Signing certificate; the stricter check-environment-readiness.sh .env --require-release remains the gate for publishing Android through Google Play. The plan neither uploads a bundle nor creates a backup.

Prepare and verify the immutable Git bundle and linux/amd64 image archive on the development workstation before authorising any production mutation:

./scripts/production-release-clean.sh prepare whoneedhelp

prepare repeats the read-only production and environment checks, reads the production image-build inputs into a mode-0600 temporary file, and then creates or verifies the commit-bound artifacts under output/releases/. It does not upload an artifact, create a production backup, change the remote checkout, load an image, run a migration, or restart a service. A later apply for the same commit verifies and reuses those exact artifacts instead of compiling them again.

After reviewing the exact commit printed by the plan, execution additionally requires an explicit per-commit confirmation:

WNH_PRODUCTION_RELEASE_CONFIRM=whoneedhelp.com:FULL_COMMIT \
  ./scripts/production-release.sh apply whoneedhelp

The application must still be released from a clean checkout. If a verified bundle and image archive were prepared in another clean checkout, point the release process at their absolute parent directory instead of rebuilding them:

WNH_PRODUCTION_RELEASE_ARTIFACT_ROOT=/absolute/path/to/output/releases \
WNH_PRODUCTION_RELEASE_CONFIRM=whoneedhelp.com:FULL_COMMIT \
  ./scripts/production-release.sh apply whoneedhelp

The selected directory must contain a child named with the exact full commit. The workflow rechecks the Git bundle HEAD, bundle checksum and manifest, image archive checksum, target platform, topology, immutable image names and image IDs before transfer. It does not accept a relative artifact path or silently fall back to a different commit. This allows a temporary clean worktree to reuse the already scanned workstation artifact while keeping production image compilation off the 4-GiB server.

The apply path refuses tracked local or remote modifications. It then:

  1. creates and verifies a full Git bundle for exactly that clean commit;
  2. reads the production environment over SSH into a mode-0600 temporary local file, selects the candidate's immutable image tags, builds the required linux/amd64 images on the development workstation, and deletes that temporary environment file;
  3. records every image ID in a manifest, creates a timestamp-free gzip-compressed Docker archive, verifies its SHA-256, and uploads the bundle, archive, checksum, and manifest to the production checkout's ignored output/releases/;
  4. creates a custom-format PostgreSQL 18 backup without exposing the database password in process arguments;
  5. verifies its archive catalog and SHA-256, then copies and verifies the backup again under local ignored output/production-backups/;
  6. fast-forwards the production checkout without accessing or changing the test checkout or public remote;
  7. verifies the transferred archive and manifest, loads the ready images without compiling on the production host, applies migrations, starts only the application topology, and verifies public readiness through the already-running shared edge plus the Android App Links endpoints.

Every newly added migration must have one reviewed entry in priv/repo/migration_application_compatibility.tsv. application_safe means the previously deployed application can run against the resulting schema; forward_only means that it cannot be guaranteed. The plan prints the aggregate policy. A forward-only apply requires a second exact confirmation:

WNH_PRODUCTION_FORWARD_ONLY_CONFIRM=whoneedhelp.com:FULL_COMMIT:forward-only \
  WNH_PRODUCTION_RELEASE_CONFIRM=whoneedhelp.com:FULL_COMMIT \
  ./scripts/production-release.sh apply whoneedhelp

For a forward-only release, all candidate images are built and transferred first, the old application is stopped before migration begins, and the target application is started only after the migration runner succeeds. If anything fails after the migration begins, the release deliberately leaves the old application stopped and records that boundary in the release manifest. Restarting an older image against a potentially incompatible schema is never automatic.

For an application_safe release, an application startup failure restores the previous immutable application image tags and attempts to recover public readiness through the unchanged edge. A forward_only release never starts the old application after migration begins. Neither path reverses Git source or Ecto migrations automatically. The per-release rollback manifest and backup paths are recorded below the production checkout's ignored output/releases/. The copied backup is separate from the production host, but a long-term encrypted off-site backup destination remains an operational requirement.

Rollback boundary

The release image is immutable and migrations run as a separate one-shot role.

If the main workspace contains intentional tracked edits that must not enter the release (for example local agent instructions), run the same workflow from an exact clean detached worktree instead of stashing, discarding, or silently including those edits:

./scripts/production-release-clean.sh plan whoneedhelp

./scripts/production-release-clean.sh prepare whoneedhelp

WNH_PRODUCTION_RELEASE_CONFIRM='whoneedhelp.com:EXACT_COMMIT' \
  ./scripts/production-release-clean.sh apply whoneedhelp

The wrapper selects the current HEAD, creates a task-owned detached worktree, runs production-release.sh there with the original immutable artifact root, and removes the worktree on success, failure, or interrupt. The underlying release still requires the exact production confirmation and any applicable forward-only migration confirmation. It does not stash, reset, stage, commit, or copy changes from the main workspace. Before a schema rollout, create and restore-test a current backup. Application rollback and database migration rollback are separate decisions: do not run an Ecto down migration merely because an image is rolled back. Inspect the exact migration and compatibility boundary first.

Every successful SSH release prints a mode-0600 rollback-manifest.txt. Before changing production, inspect that exact manifest with the read-only rollback plan:

./scripts/production-rollback.sh plan \
  /srv/who_need_help-production/output/releases/RELEASE_ID/rollback-manifest.txt \
  whoneedhelp

The rollback plan refuses manifests marked forward_only. Such releases must be repaired forward after inspecting the exact migration state; an application image rollback is available only when the manifest records migration_policy=application_safe.

The plan requires the manifest target to be the currently checked-out production commit, verifies the previous immutable application image still exists, checks the pre-release backup catalog and checksum, and prints the exact confirmation token. It does not change the remote environment or containers.

After separately reviewing application/schema backward compatibility and approving that exact scope, run:

WNH_PRODUCTION_ROLLBACK_CONFIRM=whoneedhelp.com:TARGET_COMMIT:PREVIOUS_COMMIT \
  ./scripts/production-rollback.sh apply \
  /srv/who_need_help-production/output/releases/RELEASE_ID/rollback-manifest.txt \
  whoneedhelp

This application rollback atomically restores the three previous application image selectors and recreates only the selected application topology with --no-build. It never changes or recreates the shared edge. It verifies the resulting image identities, container health, public readiness through the unchanged edge, and App Links. A failed rollback attempts to restore the pre-rollback image selection. The Git checkout intentionally remains at the newer source commit so the reviewed release tooling and manifest remain available.

The command never restores PostgreSQL, reverses Ecto migrations, changes the test deployment, or touches the public Git/Devpost submission. The repository intentionally does not ship an automatic destructive production database restore command.