who_need_help/docs/operations.md

1548 lines
73 KiB
Markdown

# Operations runbook
This runbook describes the commands that are implemented and verified in this
repository. It does not claim a production recovery point objective, recovery
time objective, retention period, storage capacity, or high-availability model;
those values require product policy and measurements from the eventual
production environment.
## Private Gitea continuous integration
The observed private Git remote is hosted by Gitea 1.22.0. Gitea 1.22 reads
repository workflows from `.gitea/workflows/`, so the repository keeps a
Gitea-specific quality workflow in addition to the GitHub-compatible workflow.
Both files are checked by the pinned local actionlint gate.
The Gitea workflow deliberately targets only the custom
`who-need-help-ci` runner label. Its jobs build and run many Docker images and
must not be scheduled onto an unrelated shared runner or onto the production
application host. Its gates run sequentially in one job so a single push cannot
start four heavy Docker suites concurrently. Before enabling it:
1. enable Actions in the private repository settings;
2. register a trusted runner controlled by this project;
3. assign `who-need-help-ci:host` only to a Linux host where the exact Docker
Engine and Compose commands used by the repository have been verified;
4. install the runtime required by the pinned checkout action on that host;
5. keep the runner registration state and any Docker socket access outside the
repository and away from untrusted projects.
The runner's availability, labels, software versions, free resources and
repository Actions setting cannot be established from the unauthenticated
repository clone. Do not treat the presence of the workflow file as evidence
that remote CI has run. After provisioning the runner, trigger the workflow
manually, require the full job and all four gate steps to pass, and record the
run URL and commit before protecting `main`.
## First production Compose environment
The deployment environment selects topology and database ownership:
| Setting | Result |
| --- | --- |
| `APP_TOPOLOGY=compact` | One Phoenix+Oban container, directly published to the host; default for the first server. |
| `APP_TOPOLOGY=split` | Traefik, `WEB_REPLICAS` web containers, and `WORKER_REPLICAS` worker containers. |
| `DATABASE_MODE=container` | Start the project-owned PostgreSQL/PostGIS container and volume. |
| `DATABASE_MODE=external` | Do not create/start a Compose database; use only `DATABASE_URL`. |
There is no Redis dependency. Queues, rate-limit counters, Oban leadership,
and durable application state use PostgreSQL.
### Development public route after a workstation restart
The current development gateway sends `whoneedhelp.imalto.site` to the
workstation's assigned OpenVPN address `10.8.0.14:4010`. If HTTPS connects to
the gateway but returns no response after a workstation restart, first verify
the local app and tunnel:
```bash
curl --fail --show-error --max-time 10 \
http://127.0.0.1:4010/healthz/ready
ip -brief address show tun0
nmcli --wait 30 connection up openvpn__toha_nobara_pc
curl --fail --show-error --max-time 15 \
https://whoneedhelp.imalto.site/healthz/ready
```
The observed profile requires `remote-cert-tls=server`; do not remove that
server-certificate check. NetworkManager documents that
`connection.autoconnect` is not implemented for VPN profiles and recommends a
base connection's `connection.secondaries` instead:
[NetworkManager connection settings](https://www.networkmanager.dev/docs/api/latest/nm-settings-nmcli.html).
The observed wired base profile currently has no secondary connection, so a
standalone `connection.autoconnect=yes` value on the VPN does not prove that
the development route will return after reboot.
Adding the VPN UUID as a wired/Wi-Fi secondary is a workstation networking
change: this profile currently installs the default route and DNS through the
VPN. Review that whole-host effect and obtain explicit approval before
configuring automatic activation. The manual command above changes no project,
test, production, or gateway configuration.
### Two independent checkouts and one `.env` in each
The server uses exactly these independent Git clones:
```text
/srv/who_need_help-test/.env
/srv/who_need_help-production/.env
```
No `.env.test`, `.env.staging`, or `.env.production` is used inside either
application checkout.
Each checkout can remain on a different commit. Test and production have
different Compose projects, application/infrastructure image tags, database
connections, Docker volumes, public aliases, Google OAuth clients, email
delivery paths, and generated secrets. Oban queues are isolated by those
different PostgreSQL databases. There is no Redis dependency.
Generated harness state is not a deployment environment. E2E, load, and
Android instrumentation scripts keep their random local inputs below the
ignored mode-`0700` `output/runtime/` directory:
```text
output/runtime/e2e.env
output/runtime/load.env
output/runtime/android-test.env
```
Those files are generated automatically, never copied to a server, and do not
represent dev, test, or production. The one ignored `.env` at each checkout
root remains the only application/deployment configuration.
The current submitted deployment still has a legacy shared Caddy container
created from the production checkout. Keep it running and unchanged while the
judging test URL is frozen. Application release and rollback scripts do not
build, recreate, or select an image for that container.
The replacement is a separate server-level `server_edge` project with its own
single mode-`0600` `.env`, route directory, Caddy image, certificate volumes,
and Compose lifecycle. Each application keeps only its route contract in its
existing `.env`: `PHX_HOST`, `PUBLIC_ROUTE_ID`, `PUBLIC_EDGE_NETWORK`,
`PUBLIC_UPSTREAM_NAME`, `PUBLIC_UPSTREAM_PORT`, `PUBLIC_HEALTH_PATH`, and
`PUBLIC_WWW_REDIRECT`. No application deployment is allowed to restart the
server edge. Both applications intentionally share only the external public
edge Docker network.
### Environment-specific Android builds and App Links
Android build inputs belong in the same ignored mode-`0600` `.env` as the web
checkout they target. Do not create `.env.android-release`,
`.env.production`, or another permanent environment file:
```dotenv
WNH_BASE_URL=https://dev.example.com
WNH_ANDROID_VERSION_CODE=1
WNH_ANDROID_VERSION_NAME=0.1.0
WNH_ANDROID_DEVELOPMENT_SIGNING_KEY_ALIAS=who-need-help-development
ANDROID_APP_LINKS_PACKAGE_NAME=org.whoneedhelp.mobile.development
ANDROID_APP_LINKS_SHA256_CERT_FINGERPRINTS=AA:BB:...
```
The development checkout uses package
`org.whoneedhelp.mobile.development` and a dedicated stable key under
`~/.config/who_need_help/android-development/`. Generate it once and place only
its public identity into the existing ignored `.env`:
```bash
./scripts/init-android-development-signing.sh
./scripts/configure-android-development-env.sh
./scripts/android-development-build.sh
```
The build exports the package and certificate reports, verifies that both match
the checkout `.env`, and checks the HTTPS
`/.well-known/assetlinks.json` response. Losing this key changes the development
certificate and breaks previously installed App Links, so back it up.
The independent test checkout uses package
`org.whoneedhelp.mobile.staging`, its own key under
`~/.config/who_need_help/android-staging/`, and
`scripts/android-staging-build.sh`. Do not copy either ignored `.env` or signing
material between development and test.
The production checkout instead uses package `org.whoneedhelp.mobile`, the
separate upload material under
`~/.config/who_need_help/android-release/`, and its own `.env`:
```dotenv
WNH_ANDROID_SIGNING_KEY_ALIAS=who-need-help-upload
ANDROID_APP_LINKS_PACKAGE_NAME=org.whoneedhelp.mobile
ANDROID_APP_LINKS_SHA256_CERT_FINGERPRINTS=UPLOAD_SHA256,PLAY_APP_SIGNING_SHA256
ANDROID_PLAY_APP_SIGNING_SHA256_CERT_FINGERPRINTS=PLAY_APP_SIGNING_SHA256
```
Run `./scripts/android-release-build.sh` from that production release checkout.
It produces an APK, Play AAB, package report, signing report, and lint report.
Before the Play application exists, the build may use only the upload
certificate in `ANDROID_APP_LINKS_SHA256_CERT_FINGERPRINTS` and leave
`ANDROID_PLAY_APP_SIGNING_SHA256_CERT_FINGERPRINTS` empty. This is sufficient
to create the first signed AAB, but `check-environment-readiness.sh
--require-release` intentionally continues to report the Play identity as
missing. Application-only server releases use
`--require-server-release` plus the production validator's
`--allow-pre-play` mode: they accept this upload-certificate-only state while
continuing to reject every other missing production capability. These modes do
not make an Android build ready for Google Play.
After Play App Signing is enabled, add the Play signing certificate fingerprint
to the comma-separated App Links value; the upload certificate alone does not
describe Play-delivered APKs. Record the same Play fingerprint separately in
`ANDROID_PLAY_APP_SIGNING_SHA256_CERT_FINGERPRINTS`; release readiness verifies
that every Play identity is present in the public App Links list. The production
environment validator accepts multiple SHA-256 fingerprints and rejects partial,
malformed, upload-only, or inconsistent configuration.
### Release capability inputs
Each environment owns distinct external-provider credentials. Seed a new
production `.env` with `scripts/init-production-env.sh` and the corresponding
`PRODUCTION_*` process variables; use `TEST_*` only when creating the separate
test checkout. The generated file uses the ordinary runtime names:
| Capability | Values kept in that checkout's `.env` |
| --- | --- |
| Google sign-in | `GOOGLE_OAUTH_CLIENT_ID`, `GOOGLE_OAUTH_CLIENT_SECRET`, `GOOGLE_OAUTH_AUTHORIZED_PARTY_IDS` |
| Browser push | three `WEB_PUSH_VAPID_*` values |
| Android Firebase client | four public `WNH_FIREBASE_*` values |
| Android delivery | `FCM_PROJECT_ID` and one private service-account source |
| Verified Android links | `ANDROID_APP_LINKS_PACKAGE_NAME`, `ANDROID_APP_LINKS_SHA256_CERT_FINGERPRINTS` |
| Transactional email | `SMTP_*`, sender, and `SUPPORT_INBOX_ADDRESS` |
`SUPPORT_INBOX_ADDRESS` configures `Reply-To` and the optional operator-alert
destination. The environment readiness script validates only that this value is
present and syntactically usable; it does not prove that the domain has inbound
MX routing or that the mailbox is monitored. Check DNS and perform an actual
inbound delivery/reply test before presenting the address as a public support
contact.
Do not reuse a Google client, VAPID private key, Firebase project/service
account, SMTP credential, or Android signing key between dev and production.
The public Firebase Android values are build configuration; the Base64 FCM
service-account JSON is a server secret and must never be passed into the
Android build.
For Android Google sign-in, each environment's `GOOGLE_OAUTH_CLIENT_ID` is also
the public server client ID supplied at runtime to Credential Manager.
`GOOGLE_OAUTH_CLIENT_SECRET` never leaves Phoenix. The Android package and
signing-certificate identity must be registered in the matching Google project:
`org.whoneedhelp.mobile.development` plus the stable development certificate
for development,
`org.whoneedhelp.mobile.staging` plus the stable staging certificate for test,
and `org.whoneedhelp.mobile` plus the Play-distributed certificate for
production. Do not add a second Google secret file or Gradle property.
Set `GOOGLE_OAUTH_AUTHORIZED_PARTY_IDS` to the comma-separated Android OAuth
client IDs from that environment. They are accepted only as the signed `azp`
claim; the signed `aud` must still contain the environment's Web client ID.
Check an environment without printing its secret values:
```bash
./scripts/check-environment-readiness.sh .env
./scripts/check-environment-readiness.sh .env --require-release
```
The complete promotion sequence and the human/provider decisions that these
scripts cannot prove are listed in
[`docs/public-launch-checklist.md`](public-launch-checklist.md).
Provider downloads can be imported into that same file without placing
secrets on a command line or printing them:
```bash
./scripts/generate-vapid-env.sh \
.env mailto:contact@YOUR_DOMAIN
chmod 600 /secure/downloads/google-oauth-client.json
./scripts/import-google-oauth-client.sh \
.env /secure/downloads/google-oauth-client.json
./scripts/import-firebase-android-config.sh \
.env /secure/downloads/google-services.json
chmod 600 /secure/downloads/fcm-service-account.json
./scripts/import-fcm-service-account.sh \
.env /secure/downloads/fcm-service-account.json
```
After the first AAB has made Google Play generate its delivery identities,
download a fresh production `google-services.json` after registering every
Play App Signing SHA-1 in the production Firebase Android application. Record
the SHA-1/SHA-256 pairs shown by Play in a protected temporary JSON document:
```json
{
"package_name": "org.whoneedhelp.mobile",
"identities": [
{"sha1": "PLAY_SHA1", "sha256": "PLAY_SHA256"}
]
}
```
Keep both provider downloads at mode `0400` or `0600`. First run the importer
without `--apply`: this validates the package, every SHA-1-to-Android-OAuth
client match, Firebase/FCM project consistency, the resulting App Links set,
and the production Android environment while leaving `.env` byte-for-byte
unchanged. Apply the same validated candidate only after reviewing the counts:
```bash
chmod 600 /secure/downloads/google-services.json \
/secure/downloads/play-identities.json
./scripts/import-play-android-config.sh \
.env \
/secure/downloads/google-services.json \
/secure/downloads/play-identities.json
./scripts/import-play-android-config.sh \
.env \
/secure/downloads/google-services.json \
/secure/downloads/play-identities.json \
--apply
```
The apply step is atomic. It preserves the existing upload certificate and
Android OAuth client, adds every Play delivery identity, refreshes the four
public Firebase Android values, and never prints OAuth client IDs or the
Firebase API key. It refuses a test/development environment, another package,
an unmatched or duplicate certificate, and a Firebase project inconsistent
with the configured FCM service account. Provider files remain on disk after
the import and must be stored or removed deliberately.
Once the production association has been deployed and the internal-track build
has been installed from Google Play, prove that the physical device is not
still running a sideloaded upload-key build:
```bash
./scripts/verify-play-installed-android.sh \
/secure/downloads/play-identities.json \
DEVICE_SERIAL \
EXPECTED_VERSION_CODE \
EXPECTED_VERSION_NAME
```
The command does not launch, install, uninstall, clear, or reconfigure the app.
It checks the installer, installed version, Play signing identity, Android's
domain-verification state, and implicit production App Link resolution. Passing
it is a prerequisite for the later authenticated Play-delivered smoke test, not
a substitute for that test.
The VAPID helper runs the exact locked `web_push_elixir` generator in an
isolated, network-disabled container, imports the result atomically, removes
its one-run image tag and temporary files, and never prints either key. It
refuses to replace an existing VAPID identity because an unplanned rotation
invalidates existing browser subscriptions.
The importers validate the exact OAuth callback, Android package, Firebase
project relationship, and required service-account fields before atomically
replacing existing keys. They preserve mode `0600` and never create another
permanent environment file. Provider and initializer values that cannot be
represented as one unquoted Compose `.env` line are rejected before mutation.
Literal dollar signs are stored as `$$`, which Compose resolves back to one
`$` inside the container. OAuth and service-account downloads still contain
private credentials after import; deliberately move them to protected backup
storage or remove them after verification.
The first command reports incomplete or local-only capabilities. The second is
a blocking release preflight and exits nonzero until Google sign-in, external
SMTP, browser Web Push, Android Firebase/FCM, App Links, support routing,
Android version/signing aliases, and the core application configuration are
all complete.
Create the test configuration inside the test checkout:
```bash
cd /srv/who_need_help-test
TEST_CODEX_SESSION_ID=YOUR_MAIN_CODEX_SESSION_ID \
TEST_GOOGLE_OAUTH_CLIENT_ID=YOUR_TEST_CLIENT_ID \
TEST_GOOGLE_OAUTH_CLIENT_SECRET=YOUR_TEST_CLIENT_SECRET \
TEST_GOOGLE_OAUTH_AUTHORIZED_PARTY_IDS=YOUR_TEST_ANDROID_CLIENT_ID \
./scripts/init-test-env.sh test.whoneedhelp.com
./scripts/validate-test-env.sh .env test.whoneedhelp.com
./scripts/deploy-up.sh .env
```
For later public test updates, release the exact clean revision with immutable
images built on the operator workstation. The test release workflow performs
a read-only scope plan, builds and verifies `linux/amd64` images outside the
4-GiB server, creates and copies an independently verified database backup,
runs a restore-and-migrate drill, and starts Compose with `--no-build`:
```bash
./scripts/test-release.sh plan whoneedhelp
./scripts/test-release.sh prepare whoneedhelp
# Make the selected commit available as origin/main before apply. The apply
# gate verifies that origin/main equals the exact local SHA.
candidate=$(git rev-parse HEAD)
WNH_TEST_RELEASE_CONFIRM="test.whoneedhelp.com:$candidate" \
WNH_TEST_FORWARD_ONLY_CONFIRM="test.whoneedhelp.com:$candidate:forward-only" \
./scripts/test-release.sh apply whoneedhelp
```
Omit `WNH_TEST_FORWARD_ONLY_CONFIRM` when the read-only plan reports
`migration_policy=application_safe`. The workflow refuses shared Caddy changes,
requires a fast-forward from the deployed test commit, preserves the single
test `.env`, and never addresses the production Compose project or database.
Do not use `deploy-up.sh` for an ordinary public test update on the small
server: that command intentionally includes `--build` and therefore builds the
release on the target host.
Test always uses its own `who_need_help_test` PostGIS container/volume. Local
development can use Mailpit; a public test deployment must use its own SMTP
password and a visibly test-specific sender identity.
Create the production configuration inside the production checkout after
loading only the production database/provider credentials into the process:
```bash
cd /srv/who_need_help-production
PRODUCTION_DATABASE_MODE=external \
PRODUCTION_DATABASE_URL='ecto://PRODUCTION_ROLE:PASSWORD@localhost/who_need_help_production' \
PRODUCTION_DATABASE_SOCKET_DIR=/var/run/postgresql \
PRODUCTION_EMAIL_DELIVERY_PROVIDER=smtp \
PRODUCTION_SMTP_RELAY=smtp.example.net \
PRODUCTION_SMTP_PORT=587 \
PRODUCTION_SMTP_USERNAME=YOUR_PRODUCTION_SMTP_LOGIN \
PRODUCTION_SMTP_PASSWORD=YOUR_PRODUCTION_SMTP_PASSWORD \
PRODUCTION_GOOGLE_OAUTH_CLIENT_ID=YOUR_PRODUCTION_CLIENT_ID \
PRODUCTION_GOOGLE_OAUTH_CLIENT_SECRET=YOUR_PRODUCTION_CLIENT_SECRET \
PRODUCTION_GOOGLE_OAUTH_AUTHORIZED_PARTY_IDS=YOUR_PRODUCTION_ANDROID_CLIENT_ID \
PRODUCTION_CODEX_SESSION_ID=YOUR_MAIN_CODEX_SESSION_ID \
./scripts/init-production-env.sh whoneedhelp.com
./scripts/validate-production-env.sh .env whoneedhelp.com
```
Promote the exact revision already verified in test by checking out that SHA in
the independent production clone, then update only the production image tags:
```bash
git checkout --detach <verified-test-sha>
./scripts/set-deployment-revision.sh .env
./scripts/deploy-up.sh .env
```
Before either public switch, prove cross-environment isolation without printing
credentials:
```bash
/srv/who_need_help-test/scripts/validate-deployment-isolation.sh \
/srv/who_need_help-test /srv/who_need_help-production
```
Start/update test first and run the complete browser/API/Android verification.
Only then check out that exact tested SHA in production. Start the application
with `./scripts/deploy-up.sh .env`. The application command joins the external
public network but does not own or restart Caddy.
The legacy `./scripts/edge-up.sh .env` command exists only for the currently
frozen submitted deployment. Do not use it from an ordinary application
release. After judging, render and validate each route through the independent
`server_edge` workflow, transfer certificate-volume ownership in a reviewed
maintenance window, and only then retire the legacy edge container.
The authoritative A records for `whoneedhelp.com`, `www.whoneedhelp.com`, and
`test.whoneedhelp.com` must point to the verified server address before Caddy
can obtain their certificates. Do not remove the edge volumes during an
ordinary application deploy. Do not add `--volumes` to a stop command unless
the exact database has been inspected and deletion is intended.
The initializer reads `.env.example`, derives the numeric group of that host's
Docker socket, generates independent random values for PostgreSQL,
`SECRET_KEY_BASE`, handover codes, the BEAM release cookie, and metrics access,
and writes an ignored mode-`0600` file. It does not print those values and
refuses to overwrite an existing destination.
By default, the generated public proxy port binds to `127.0.0.1`, which is
appropriate only when the verified reverse proxy reaches the application on
the same host. Set `PRODUCTION_HTTP_BIND_ADDRESS` when generating the file, or
edit `HTTP_BIND_ADDRESS` afterward, to match the observed target topology.
Replace `TRAEFIK_TRUSTED_IPS` with the exact source IP/CIDR observed at Traefik;
do not copy the temporary VPN value into an unrelated server.
Configure the transactional SMTP relay and a sender accepted by it using
`EMAIL_DELIVERY_PROVIDER=smtp` plus the `SMTP_*` settings. The initializer
accepts the corresponding `PRODUCTION_EMAIL_DELIVERY_PROVIDER` and
`PRODUCTION_SMTP_*` inputs.
Set optional `SUPPORT_INBOX_ADDRESS` to
a monitored address for support/removal queue alerts and email `Reply-To`.
Public support creates an alert only after email verification; authenticated
support is already verified, while content-removal notification rules remain
separate. The alert is queued in `mail`; conversation messages remain in the
permission-scoped staff workspace rather than producing email per message.
Leaving the value empty disables operator email alerts, not the protected queues. If
Google registration/sign-in is
enabled, also set both Google Web client credentials and register
`https://YOUR_PHX_HOST/auth/google/callback` as the exact authorized redirect
URI. Leave both credentials empty to keep the feature disabled. Then run:
`SUPPORT_INBOX_ADDRESS` does not configure inbound delivery. To opt into the
Brevo inbound-parse boundary, set `SUPPORT_INBOUND_RECIPIENT` to the exact
mailbox on a dedicated receiving subdomain. Set
`SUPPORT_INBOUND_WEBHOOK_TOKEN` to an independent secret, or let the production
initializer generate it when a recipient is supplied. Configure Brevo to send
a secured webhook with `Authorization: Bearer <that token>` to
`https://YOUR_PHX_HOST/webhooks/brevo/inbound-support`. Follow Brevo's current
MX instructions for the receiving subdomain, then verify a real inbound message
and the subsequent mailbox-confirmation link. Inbound messages stay outside the
staff queue until confirmation; provider attachment tokens are deliberately not
downloaded. See `docs/support-and-content-removal.md` for the trust boundary.
```bash
./scripts/validate-production-env.sh .env whoneedhelp.com
./scripts/deploy-up.sh .env
```
The production override keeps Mailpit behind its inactive `local-mail` profile,
so public registration cannot appear to succeed while mail is only retained
locally. The validator checks file ownership/mode, origin consistency,
template markers, independent generated secrets, settings required by the
selected email provider, a
complete-or-empty Google credential pair, the generated database URL, and the
final Compose render without printing secrets.
It does not contact DNS, TLS, the email provider, the reverse proxy, or the application.
After deployment, verify `/healthz/ready`, inspect all replica health and logs,
register a unique address through the public browser, receive its message at
the real mailbox, follow the HTTPS confirmation link, and remove only that
run-scoped account.
### External PostgreSQL/PostGIS
#### PostgreSQL on the same Linux host
The server-local mode uses the PostgreSQL Unix socket instead of exposing the
database on a Docker-reachable TCP address. Ecto/Postgrex receive
`DATABASE_SOCKET_DIR`; `scripts/compose.sh` then adds
`compose.external-db-socket.yaml` and mounts that exact directory read-only
into the active application and migration services. An empty setting preserves
the normal remote-provider TCP behavior.
After read-only inspection confirms the intended PostgreSQL 18 cluster,
installed PostGIS package, socket directory, and absence of the project-scoped
roles/databases, run the root-only provisioner:
```bash
sudo ./scripts/provision-host-postgres.sh "$USER"
```
Its exact mutation scope is:
- prepend two database-and-role-specific `local ... scram-sha-256` rules to the
active `pg_hba.conf`, retaining a mode-`0600` copy under
`/var/backups/who_need_help/`, and reload that cluster;
- create login roles `wnh_production` and `wnh_test` without superuser,
database-creation, role-creation, replication, or row-security bypass rights;
- create empty owner databases `who_need_help_production` and
`who_need_help_test`, revoke public connect, and preload `citext` and
PostGIS so the non-superuser Ecto migrations can run;
- verify both credentials through the observed Unix socket and write separate
mode-`0600` initializer fragments under
`~/.config/who_need_help/` without printing passwords.
The command refuses to overwrite credential fragments, refuses any matching
pre-existing role/database or managed HBA marker, validates the candidate HBA
rules before reload, and restores its HBA backup plus removes only objects it
created if provisioning fails. It does not inspect or migrate application
tables; the release migration runner remains authoritative for schema.
Generate test or production after loading only the matching trusted
fragment:
```bash
set -a
. "$HOME/.config/who_need_help/database-production.env"
set +a
./scripts/init-production-env.sh whoneedhelp.com
unset PRODUCTION_DATABASE_MODE PRODUCTION_DATABASE_URL \
PRODUCTION_DATABASE_SOCKET_DIR
```
For test external-database drills, load `database-test.env`; the public test
deployment normally uses its own Compose database. Do not load both fragments into
one shell.
Provision the database and role first, then generate the environment without
placing its credentials on a command line that is retained in shell history:
```bash
export PRODUCTION_DATABASE_MODE=external
read -rsp 'External DATABASE_URL: ' PRODUCTION_DATABASE_URL && echo
export PRODUCTION_DATABASE_URL
./scripts/init-production-env.sh whoneedhelp.com
unset PRODUCTION_DATABASE_URL
```
Set `PRODUCTION_APP_TOPOLOGY=split` during generation when independent web and
worker scaling is required. In external mode `scripts/compose-up.sh` builds the
migration image, runs `scripts/check-database.sh`, and fails before migration
or application startup unless PostgreSQL and `PostGIS_Version()` are reachable.
The check also reports the observed TLS flag. The connection URL must follow
the database provider's verified TLS/CA requirements.
`scripts/compose.sh .env config --services` is the canonical
read-only render check. In external mode its output must not contain `db`.
Use `scripts/compose.sh .env ps -a` and
`scripts/compose.sh .env logs web worker` for the selected mode.
The repository's Compose backup, restore, rotation, and rollback-based staging
scripts deliberately refuse `DATABASE_MODE=external`; they are scoped to the
project-owned database container. For an external database, use the provider's
verified backup/restore procedure and test recovery before public launch.
## Compose database backup
Create a PostgreSQL 18 custom-format archive, validate its table of contents,
and write a SHA-256 manifest:
```bash
./scripts/backup-compose.sh
```
The default destination is the ignored `output/backups/` directory. An explicit
new destination may be supplied as the only argument. The command refuses to
overwrite either an archive or its checksum manifest and writes through
temporary files before publishing the final pair. Files and newly created
directories are restricted by `umask 077`.
The archive covers the configured application database. PostgreSQL cluster
globals such as roles and tablespaces are not part of `pg_dump`; deployment
credentials and database roles must be provisioned separately from secrets.
Local backup files on the same workstation are not an off-site backup.
## Encrypted local S3-compatible backup drill
The isolated load project can run a complete encrypted Restic/MinIO drill:
```bash
./scripts/load-stack-up.sh
./scripts/backup-s3-drill.sh local-encrypted-backup
```
The script refuses the staging Compose project and validates the project and
service labels of every pre-existing container in its scope. On first use,
`scripts/ensure-local-load-env.sh` generates independent random MinIO and
Restic credentials in ignored `output/runtime/load.env` and restricts that file to mode
`0600`. MinIO publishes Docker-assigned ports only on `127.0.0.1`; the observed
API and console URLs are printed after a successful run.
The backup tool combines the PostgreSQL 18.6 client with pinned Restic
rebuilt on Go 1.26.7. MinIO server and client are also rebuilt as non-root
Alpine images from checksum-pinned upstream source commits with the exact
dependency updates recorded in `Dockerfile.minio`. The quality gate verifies
their reported release, commit, Go runtime, configured user, and current
HIGH/CRITICAL vulnerability scan. `restic backup --stdin-from-command` runs a
custom-format `pg_dump`, checks the producer exit status, encrypts the data,
and uploads it directly to MinIO. No plaintext database dump is written to the
host. The drill then:
1. runs `restic check --read-data`;
2. streams `restic dump` into `pg_restore --list`;
3. restores into a uniquely named database created from `template0`;
4. checks tables, current Ecto migrations, PostGIS, and release migration
readiness before removing that exact database;
5. clones and corrupts an isolated repository and requires both check and dump
to fail;
6. stops an exact scoped in-progress backup container only after encrypted
objects reach MinIO, requires zero published snapshots, prunes unreferenced
packs, and rechecks the repository;
7. removes and verifies removal of the corruption/interruption buckets and
requires source table counts to remain unchanged.
The successful encrypted bucket is deliberately retained in the named local
MinIO volume. Non-secret evidence is written under ignored
`output/backups-s3/<run-label>/`; the runtime scratch directory is under
ignored `tmp/backup-s3/`. Use a unique lowercase run label of at most 32
characters. The command refuses to replace an existing retained bucket.
MinIO server and client are AGPLv3. `Dockerfile.minio` identifies the exact
upstream source commits and contains the dependency changes and complete build
commands used here. Before distributing or publicly operating modified images,
review the license and make the corresponding source available as required;
this runbook does not provide legal advice.
This verifies encryption, local S3 protocol use, restore mechanics, and two
failure paths on the observed workstation. A MinIO volume on that same
workstation is not an off-site backup and does not establish production RPO,
RTO, retention, capacity, key custody, object locking, or database HA.
## Encrypted production backup to an independent host
The production workflow streams an external PostgreSQL 18 custom-format dump
through an encrypted Restic repository on an SSH target that must resolve to a
host other than production. It then runs `restic check --read-data` and restores
the newest snapshot into a uniquely named, temporary PostGIS container created
from `template0`. The source production database is never a restore target.
Bootstrap the checksum-pinned Restic binary and create the ignored local
operations configuration:
```bash
./scripts/bootstrap-restic.sh
./scripts/init-production-operations.sh
```
The initializer creates only these ignored local files:
- `tmp/production-operations/backup.env`, mode `0600`;
- `tmp/production-operations/restic-password`, mode `0600`.
The generated configuration uses the reviewed `whoneedhelp` production SSH
target and the independent `buyvm-maya` SFTP target. Inspect the read-only plan
before repository initialization:
```bash
./scripts/production-offsite-backup.sh plan
```
Repository initialization is an external mutation and requires the exact
repository value printed by the plan:
```bash
WNH_OFFSITE_BACKUP_INIT_CONFIRM='sftp:buyvm-maya:backups/who_need_help-production' \
./scripts/production-offsite-backup.sh init
```
The routine operations are:
```bash
./scripts/production-offsite-backup.sh run
./scripts/production-offsite-backup.sh check
./scripts/production-offsite-backup.sh restore-drill
```
`run` creates a fresh source dump, copies it into the Restic input, publishes
one encrypted snapshot only after the producer succeeds, performs a full-data
repository check, and runs the isolated restore drill. It removes its exact
temporary source dump, local scratch directory, Docker container, and Docker
volume on success, failure, or interrupt. Non-secret catalogs, checksums, and
restore observations remain under ignored
`output/production-operations/<run-id>/`.
After all of those checks pass, `run` atomically publishes a mode-`0600`
heartbeat on the independent repository host. The heartbeat contains only the
Restic snapshot identifier, verification time, and `restore_verified: true`;
it contains no database rows, account identifiers, recipient addresses, or
credentials. A failed backup or failed restore never advances the heartbeat.
Install the daily user-systemd timer on the operator workstation only after one
manual `run` succeeds. The production host does not run this timer. The
workstation connects to production for the source dump and publishes the
restore-verified heartbeat to the independent repository host:
```bash
./scripts/install-production-backup-timer.sh
systemctl --user list-timers --all who-need-help-production-backup.timer
systemctl --user status who-need-help-production-backup.service --no-pager
journalctl --user -u who-need-help-production-backup.service --no-pager
```
The timer uses `Persistent=true`, so systemd runs a missed calendar invocation
after the workstation and user manager return. A failed service invokes the
external monitor's backup-failure notification. It does not retry, delete
snapshots, or infer a retention policy.
The Restic password is required for every restore. Before relying on this as
the only recovery copy, place it in an operator-controlled password manager or
offline recovery record. Never print it into a terminal log, commit it, or
copy it into the backup repository. There is intentionally no automatic
`forget`/`prune` policy yet: retention, RPO, RTO, capacity, and key custody are
operator decisions, and the script does not invent them.
## Independent production monitor
The external monitor runs on an SSH target that must resolve to a host other
than production. It requests the public readiness endpoint and accepts only
HTTP 200 with the exact JSON object `{"status":"ready"}`. It also scrapes the
authenticated `/metrics` endpoint with the independent `METRICS_TOKEN` copied
from the production environment into the external host's mode-`0600`
configuration. Readiness or metrics-scrape failure changes the monitor state to
down. Email is sent once when that state changes to down and once when it
recovers; repeated down checks do not send repeated alerts.
The monitor can also treat a missing, invalid, unverified, or stale backup
heartbeat as down. This is intentionally not enabled with an invented default:
the maximum acceptable age is the operator's RPO/alerting decision. After that
decision is recorded, reinstall the monitor with the chosen positive number of
seconds:
```bash
MONITOR_BACKUP_MAX_AGE_SECONDS=<operator-selected-seconds> \
./scripts/install-production-external-monitor.sh
```
The installed configuration then records the exact threshold and heartbeat
path. A stale backup produces one state-transition alert; a later successful,
restore-verified backup produces one recovery notification.
For each observed application node, the monitor keeps a baseline of aggregate
HTTP exception, failed Oban job, and failed/exceptional email-delivery counters.
It sends one aggregate notification when one of those counters increases. A
new node establishes a baseline without alerting, and a lower value is treated
as a counter reset rather than a failure. The metric labels are deliberately
bounded to queue, delivery status, and the fixed application-owned email
purposes listed by `WhoNeedHelp.EmailDelivery`. This lets an operator
distinguish an authentication-delivery failure from a support-update failure
without exposing recipient addresses, message content, request payloads, or
user identifiers. Unknown email-purpose labels are ignored rather than
creating unbounded monitor state. With multiple application replicas, this
external check samples only the node that answers each request; use the full
Prometheus deployment when every replica must be scraped continuously.
Install it from the production SMTP configuration without printing the SMTP
credential:
```bash
ssh buyvm-maya 'loginctl show-user simple --property=Linger --value'
./scripts/install-production-external-monitor.sh
ssh buyvm-maya \
'systemctl --user status who-need-help-production-monitor.timer --no-pager'
ssh buyvm-maya \
'journalctl --user -u who-need-help-production-monitor.service --no-pager'
```
The first command must print `yes`. A user timer is not independent of SSH
sessions merely because it is enabled: the external host must keep that user's
systemd manager running after logout and start it at boot. The installer checks
this before copying configuration or unit files and refuses installation when
lingering is not enabled. An administrator on the external monitor host enables
it explicitly:
```bash
sudo loginctl enable-linger simple
loginctl show-user simple --property=Linger --value
```
This affects every enabled user unit for `simple`, not only the Who Need Help
monitor. To reverse it, first disable and remove the monitor user units, verify
that no other enabled user unit requires persistence, and only then run
`sudo loginctl disable-linger simple`.
The default schedule is once per minute with separate three-second readiness
and metrics timeouts. Those defaults match the current application container
health timeout and were installed only after public readiness requests from
the selected monitor host were observed to finish below one second. They are
observations of the current path, not universal capacity or availability
guarantees. Re-running the installer refreshes the exact script, mode-`0600`
configuration, and user units, then performs one check before enabling the
timer.
By default the installer mirrors the production application's SMTP credential.
For the pilot, use an independently revocable monitor SMTP key so revoking or
rotating the application key does not also disable operational alerts. Keep the
override outside the repository in a mode-`0600` (or read-only mode-`0400`)
file containing exactly these keys:
```dotenv
SMTP_RELAY=smtp-relay.example
SMTP_PORT=587
SMTP_USERNAME=independent-monitor-login
SMTP_PASSWORD=secret-from-the-provider
SMTP_TLS=always
SMTP_SSL=false
EMAIL_FROM_ADDRESS=monitor@whoneedhelp.com
EMAIL_FROM_NAME=Who Need Help monitor
SUPPORT_INBOX_ADDRESS=monitored-operator@example.com
```
Create that file through the password manager or an editor that does not place
the secret in shell history. When the override is present, the installer reads
only `METRICS_TOKEN` from production and never copies the application's SMTP
credential into its local staging configuration. Then reinstall the monitor
with the scoped override and send one test notification before revoking any
previous key:
```bash
MONITOR_SMTP_VALUES_FILE="$HOME/.config/who-need-help/monitor-smtp.env" \
./scripts/install-production-external-monitor.sh
ssh buyvm-maya \
'/usr/bin/python3 ~/.local/lib/who-need-help/production-external-monitor.py \
--config ~/.config/who-need-help/monitor.json \
--state ~/.local/state/who-need-help/monitor.json \
send-test-notification'
```
Verify receipt in the monitored mailbox. Only then revoke the former monitor
key. The application SMTP key is a separate rotation: validate the new key,
atomically update the production `.env`, recreate only the application
services so they read the new runtime value, verify an application-generated
authentication message and readiness, and revoke the former application key
last. If any check fails before revocation, restore the prior mode-`0600`
environment and recreate the same services; do not leave application and
monitor configurations half-updated.
## Local external-service boundary drill
Run the OAuth, SMTP, and provider-neutral push protocol checks without public
credentials, a server, or host-published ports:
```bash
./scripts/external-boundaries-run.sh local-boundaries
```
The script creates a uniquely named Compose project on an internal-only Docker
network. It generates independent one-run OAuth and push credentials in an
ignored mode-`0600` environment file, including a separate bearer token for the
protected metrics endpoint, builds the production release plus a non-root
standard-library Python protocol mock, and then verifies:
`compose.external-boundaries.yaml` is the only deployment file that sets
`ALLOW_INSECURE_EXTERNAL_HTTP=true`. Ordinary runtime configuration requires
HTTPS for OAuth and push endpoints so provider credentials are not sent over
plain HTTP.
1. GitHub-compatible OAuth authorization, PKCE S256, token exchange, normalized
user lookup, state mismatch, provider rejection, one-time-code replay,
a fresh flow after a temporary token error, and a token timeout;
2. SMTP acceptance, permanent recipient rejection without retry, one retry
after a temporary greeting failure, a greeting timeout, and the result of
submitting the same message twice;
3. the disabled default push boundary plus HTTP success, permanent rejection,
one temporary retry, replay deduplication, and deduplication after an
ambiguous timeout using the same idempotency key;
4. a fresh PostGIS database, current migrations, and two Oban worker replicas;
request acceptance and new-chat domain transactions enqueue stable
user-recipient events, replay is deduplicated before HTTP, an injected
temporary chat delivery fails its first Oban attempt and completes on its
second, and private message text is absent from the push payload.
The evidence JSON and mock state contain counters, booleans, normalized
identity fields, payload digests, run-scoped user/event identifiers, and the
privacy-safe notification metadata asserted by the drill. The script fails if
any generated secret appears in retained evidence. Its trap validates exact
Compose labels, removes only that project, volumes and one-run images, and
deletes the temporary credential file. Failure logs are retained under the
same ignored evidence directory.
This drill uses the application's real Assent/Req and Swoosh/gen_smtp clients,
but the providers are local. It therefore verifies client-side protocol and
product integration, not GitHub, SMTP-provider, FCM, or APNs availability.
SMTP permits duplicate delivery after ambiguous outcomes, so the result
explicitly makes no exactly-once claim. Push currently targets a stable
`user:<uuid>` recipient; selecting a provider, registering device tokens, and
resolving that user to devices remain deployment/provider work.
The implementation follows the configured adapter interfaces and protocol
semantics documented by
[Assent 0.3.1](https://hexdocs.pm/assent/0.3.1/Assent.HTTPAdapter.html),
[GitHub OAuth](https://docs.github.com/en/apps/oauth-apps/building-oauth-apps/authorizing-oauth-apps),
and [SMTP RFC 5321](https://datatracker.ietf.org/doc/html/rfc5321).
The application uses Swoosh's SMTP adapter and `gen_smtp`. Verify the selected
relay's authenticated delivery, sender-domain authentication, and delivered
message headers before launch. Test and production may share an account-level
SMTP login only when they use independently revocable SMTP passwords.
## Isolated restore drill
Run a real restore into a uniquely named temporary database:
```bash
./scripts/restore-drill-compose.sh output/backups/compose-YYYYMMDD-HHMMSS.dump
```
The drill:
1. validates the SHA-256 manifest;
2. validates the archive table of contents;
3. creates a pristine database from `template0`;
4. restores with `pg_restore --exit-on-error`;
5. reads every restored public application table and checks the PostGIS
library;
6. runs the current immutable release's migrations and migration-readiness check
against only the temporary database;
7. drops only the temporary drill database and verifies that it is gone.
A trap also attempts to drop the exact temporary database if a check fails.
The source application database is never passed to `pg_restore`, `dropdb`, a
clean operation, or the drill migration runner. The drill intentionally does
not compare an older backup's row counts to the live source, because concurrent
legitimate writes or a historical archive would make that comparison invalid.
## Isolated Compose upgrade rehearsal
After creating a current custom-format backup, run the complete current release
against an isolated restored copy:
```bash
./scripts/upgrade-rehearsal-compose.sh \
output/backups/compose-YYYYMMDD-HHMMSS.dump
```
The rehearsal validates the checksum and archive catalog, reads the public
origin configuration from ignored `.env`, and uses only the independently
generated credentials in ignored mode-`0600` `output/runtime/e2e.env`. It builds a uniquely
tagged production release, creates a uniquely named Compose project and
database from `template0`, restores the archive, records application-table
counts, and then:
1. applies every current timestamped Ecto migration;
2. requires the localized category column and all 11 valid cursor indexes;
3. starts 2 web and 2 worker replicas behind the isolated Traefik instance;
4. requires the four-node BEAM cluster and cross-node PubSub probe;
5. checks the production HTTP-to-HTTPS redirect, trusted-proxy public pages,
and both health endpoints;
6. requires an empty before/after diff for every public application table
except the expected migration and Oban-internal tables;
7. removes and verifies removal of the exact containers, networks, database
volume, and one-run image.
The input archive is read-only and is not copied into the evidence directory.
The ordinary Compose project, source database, public route, and running
containers are outside the generated project scope. Non-secret evidence is
retained under ignored mode-`0700`
`output/upgrade-rehearsal/<run-id>/`, with files mode `0600`.
## Clean tracked-revision deployment drill
Run:
```bash
./scripts/clean-deploy-verify.sh
```
The script archives the current tracked Git revision into a new temporary
directory. It verifies that the archive contains no local `.env`, Git
metadata, generated output, or existing E2E/load environment. It generates
independent one-run PostgreSQL, Phoenix, handover, BEAM-cookie, and metrics
secrets with mode `0600`; optional GitHub OAuth and push delivery stay
disabled.
The drill uses a unique Compose project, application image, Docker socket
proxy image, PostGIS image, Traefik image, router/service name, internal
network, PostgreSQL volume, database, and dynamic host ports.
It requires exactly two healthy web and two running worker replicas, all
tracked migrations, seeded categories, an idempotent repeated migration,
working readiness/home/registration/Mailpit routes, four connected BEAM
nodes, and a cross-replica PubSub message. It then removes the exact project
including volumes, all four one-run images, and the temporary archive, and
verifies that every scoped resource is absent. If Compose validation fails
before project creation, cleanup avoids invoking an unrenderable Compose
configuration and still verifies that the generated scope is absent. Evidence
is retained under `output/portability/<run-id>/`.
This is a local Docker portability observation. It does not establish
production SMTP, OAuth, push-provider, TLS, database-HA, storage, capacity, or
jurisdictional readiness.
## Service checks
```bash
docker compose ps
curl --fail http://localhost:4010/healthz/live
curl --fail http://localhost:4010/healthz/ready
./scripts/verify-realtime-cluster.sh compose
```
`live` verifies that the web process can serve HTTP. `ready` additionally runs
`SELECT 1` through the configured Ecto repository. The cluster probe subscribes
on one connected BEAM node and broadcasts from another.
Health checks do not replace alerting, database backups, restore drills, or
application-level synthetic checks.
## Local failure and rolling-replacement drills
The isolated load project can exercise process crashes, sequential container
replacement, and a real Oban retry without touching the normal Compose project:
```bash
./scripts/load-cycle.sh local-resilience resilience
```
The lifecycle wrapper assigns unique Compose and image names and removes that
exact run's containers, networks, named volumes, temporary environment, and
image tags on success, failure, or interruption. Use `load-stack-up.sh` plus
`load-resilience-run.sh` only when the isolated stack must remain available for
manual inspection.
The resilience script refuses `LOAD_PROJECT=who_need_help` and verifies the
Compose project/service labels of every container before stopping it. It:
1. continuously calls readiness through the isolated Traefik route;
2. terminates the BEAM process in one web and one worker container and requires
Docker's observed restart count to increase;
3. removes and replaces each web and worker replica one at a time;
4. waits for every configured BEAM node, then runs the cross-node PubSub probe;
5. enqueues a side-effect-free local worker that fails its first Oban attempt
and succeeds on its second;
6. removes that exact Oban row and requires no fixture domain rows to remain.
Traefik's retry middleware is attached to the HTTP and local TLS routers. Its
attempt count is an environment input. Traefik retries transport failures and,
with the checked configuration, does not opt in to retrying non-idempotent
requests. This reduces a stale-backend window; it is not a claim of production
availability.
For the project-owned kind cluster, run:
```bash
./scripts/kind-rolling-verify.sh local-kind-rollout
```
The kind control plane is an optional verification environment, not part of
the ordinary public Compose route. Pause it without deleting its container or
Kubernetes state when the cluster is not being tested:
```bash
./scripts/kind-stop.sh
```
The next `./scripts/kind-up.sh` validates the ownership marker and kind labels,
starts the stopped control plane, then continues with image loading and chart
reconciliation.
That script requires both the kind ownership marker and the control-plane
cluster label before invoking `rollout restart`. It changes only the web and
worker Deployment pod templates. It snapshots application-table counts before
and after, continuously probes the observed Docker mapping for the chart's
NodePort, requires all four old pod UIDs to disappear, waits for the exact BEAM
peer count, and verifies cross-node PubSub. PostGIS, its hostPath, the
Kubernetes Secret, and the namespace are not recreated.
The rollout timeout, probe interval/timeout/retry count, and cluster-join
timeout are experiment inputs. They are not production SLOs or resource
requirements.
## Isolated Oban burst measurement
The worker role consumes `maintenance`, `push`, and `mail`. Configure their
per-worker limits with `OBAN_MAINTENANCE_CONCURRENCY`,
`OBAN_PUSH_CONCURRENCY`, and `OBAN_MAIL_CONCURRENCY`; multiplying a value by
the number of worker replicas gives the configured cluster-wide concurrency
for that queue. Authentication links remain synchronous because their database
token is committed only when delivery succeeds. Support-contact verification,
support updates, content-removal confirmation/receipt/decision messages, and
optional new-case operator alerts run in `mail`, so their SMTP latency does not
occupy a web request or the `push` queue. Requester and staff conversation
messages do not each generate email, and internal assignment-only legal updates
do not notify the submitter.
After starting the isolated load project, run an explicitly sized experiment:
```bash
./scripts/load-stack-up.sh
./scripts/oban-burst-run.sh local-oban-burst 1000 120
```
The three required arguments are an evidence label, job count, and recorded
experiment timeout in seconds. They are not capacity thresholds. The script
refuses the ordinary `who_need_help` project, verifies the expected worker
replicas, inserts only confirmed `LocalBurstProbe` jobs with a unique run id,
records each worker's effective queue configuration and container samples,
requires every job to complete, compares domain-table row counts, and deletes
exactly its own jobs. The probe worker has no product side effects and no
product flow enqueues it.
## Kubernetes ingress isolation
`networkPolicy.enabled=true` renders the chart's ingress NetworkPolicy.
Enforcement is a property of the cluster CNI, not of the YAML object alone.
Before relying on it, verify that the target cluster uses a NetworkPolicy-
capable plugin and run positive HTTP/cluster checks plus negative blocked-port
checks there. Egress remains intentionally unrestricted until the actual
database and external-service destinations are known.
## BEAM runtime memory guard
Compose sets `ERL_ZFLAGS="+Q ${ERLANG_PORT_LIMIT}"`; Helm renders the same flag
from `app.erlangPortLimit`. Both default to `65536`, OTP's normal port-table
limit. This makes the runtime independent of an unusually large host or nested
container `nofile` limit. Validate all live replicas after deployment:
```bash
./scripts/verify-beam-runtime.sh compose
./scripts/verify-beam-runtime.sh kind
```
The command records the effective port count/limit, allocated port-table bytes,
BEAM memory, process count, cgroup memory, and RSS for each web and worker
replica. It fails when a live VM does not use the configured limit; kind mode
also requires the desired number of Ready application pods.
`65536` is a concurrency ceiling for simultaneously existing Erlang ports
(files, sockets, and drivers), not a container memory limit. Do not lower or
raise it from a RAM estimate alone. A changed value must be validated against
measured peak port usage and the target environment.
## Protected Prometheus metrics
The web role exposes Prometheus text format at `/metrics`. It requires the
independent `METRICS_TOKEN` deployment secret:
```bash
curl --fail \
--header "Authorization: Bearer $METRICS_TOKEN" \
http://localhost:4010/metrics
```
The endpoint returns `401` without the exact token, disables response caching,
and does not put the credential in a URL. The reporter exports cumulative HTTP
request and duration, router exception, database query and duration, WebSocket
connection, Oban attempt and stop outcome, aggregate single-email delivery
outcome, VM memory, and scheduler run-queue metrics. The Oban outcome series
distinguishes the fixed `success`, `cancelled`, `discard`, and `snoozed` stop
states without job arguments or identifiers; exception attempts remain a
separate counter. Email metrics retain the adapter-level `ok`/`error`
counter and exception counter, and also expose the application's fixed
allow-listed delivery purpose together with `ok`, `error`, or `exception`.
Purpose labels are code-defined values such as `auth_login`,
`support_confirmation`, or `content_removal_update`; recipients, subjects,
references, and message bodies are never labels. Cumulative durations are
integer microseconds because the selected reporter's sum accumulator is
integer-based; divide by `1_000_000` in PromQL when seconds are required.
Definitions intentionally have no request path, user, request, recipient,
message, or event-name labels that could create unbounded cardinality or expose
private data.
The locked `telemetry_metrics_prometheus_core` reporter aggregates
distribution samples only when a scrape occurs. Each application VM therefore
also runs a supervised internal aggregation every ten seconds. This keeps the
reporter's raw distribution table independent of whether an external
Prometheus server is currently configured, while preserving the cumulative
histograms exposed by the authenticated endpoint. The generated exposition
text from the internal aggregation is discarded.
After a deployment, the raw table can be inspected without exposing metric
credentials:
```bash
/app/bin/who_need_help eval \
'IO.inspect(:ets.info(:prometheus_metrics_dist, :size), label: "pending_distribution_samples")'
```
Under idle conditions the value returns to zero after the next aggregation
interval. During traffic it represents only samples received since the most
recent internal or external scrape; it is not the cumulative histogram count.
Metrics are local to each BEAM process. Discover and scrape every web pod or
container as a distinct target and preserve Prometheus's `instance` label. A
request through the load-balanced public route reaches only one replica and is
therefore useful as an authorization/smoke check, not as a cluster-wide
aggregate.
The isolated load project includes a local observability profile:
```bash
./scripts/load-stack-up.sh
./scripts/observability-run.sh local-observability
```
The run script refuses the staging project, validates every current web
container's Compose labels, and writes a `file_sd` target for each observed
internal IP. Prometheus reads the Bearer value from a mode-`0600` runtime file,
not a tracked config or URL. Its direct request includes the internal
`X-Forwarded-Proto: https` signal required by the application's production SSL
rewrite while preserving the target's own `instance` label.
Prometheus, Alertmanager, and Grafana are pinned by tag and digest. Their host
ports default to Docker-assigned values bound only to `127.0.0.1`; the run
prints the observed URLs. Grafana uses the random admin password generated in
ignored `output/runtime/load.env`, disables anonymous signup, update checks, suggested plugin
installation, and its unused built-in alert engine. The Prometheus datasource
and ten-panel dashboard are provisioned from tracked files. The dashboard
separates all discovered web and worker replicas and includes HTTP
traffic/latency, BEAM memory and scheduler run queues, Ecto execution and
pool-wait latency, and Oban outcomes and queue wait by queue.
The verification stops exactly one scoped load web container. The
`WhoNeedHelpWebReplicaUnavailable` rule is based only on the factual
`up == 0` result; it is a local failure drill, not an invented latency,
capacity, or production SLO threshold. The script requires both firing and
resolved webhook payloads from Alertmanager, starts the same container, waits
for every direct target, and compares read-only database counts before and
after. Evidence is retained in `output/observability/` without the metrics or
Grafana secrets.
Stop only the monitoring services with:
```bash
./scripts/observability-stop.sh
```
Prometheus/Grafana/Alertmanager retention, production notification
destinations, production availability, and measured alert policies remain
deployment decisions. In Kubernetes, put the metrics token in
`existingSecret`; configure the external scraper to send it as a Bearer token.
The provisioned Grafana overview also shows one-hour transactional-email
attempts grouped only by the fixed application-owned purpose and outcome. It
does not expose recipient addresses, subjects, message bodies, support
references, or user identifiers. The panel is intended to answer which bounded
workflow is creating delivery volume or failures; it is not a user-activity
log.
The metrics endpoint also exposes cumulative rate-limit bucket checks grouped
only by the configured action name and the bounded `allowed` or `limited`
outcome. It never labels a metric with the hashed scope, email address, client
address, user identifier, bucket count, or reset timestamp. Use these counters
to observe which pilot policies affect real traffic before changing their
limits; the counters do not by themselves establish an abuse policy or a safe
capacity threshold.
Support conversations deliberately do not send one email per staff message.
The requester receives an email for the first staff response and for later
public status changes; additional messages while the status is unchanged stay
in the private support conversation and the in-product notification flow.
Anonymous submissions receive the address-confirmation message before they
enter the operator queue. Operator email alerts are disabled unless
`SUPPORT_OPERATOR_EMAIL_MODE=immediate` is explicitly configured.
## Exact production Android FCM smoke
The Android FCM smoke is split into explicit phases so the operator can inspect
the notification on the exact physical device before cleanup. The plan is
read-only and verifies the production Compose identity, running image, healthy
application container, database name, and one exact active Android FCM device:
```bash
./scripts/production-android-fcm-smoke.sh \
plan <push-device-uuid> --check-only whoneedhelp.com
```
After the printed scope and exact device have been reviewed and separately
authorised, prepare one notification, inspect it on the phone, verify provider
delivery, and remove the run-owned records:
```bash
./scripts/production-android-fcm-smoke.sh \
prepare <push-device-uuid> --confirm whoneedhelp.com
./scripts/production-android-fcm-smoke.sh \
verify --from-state --confirm whoneedhelp.com
./scripts/production-android-fcm-smoke.sh \
cleanup --from-state --confirm whoneedhelp.com
```
`prepare` creates one notification and one exact FCM delivery job in a single
database transaction. Its mode-`0600` manifest is exclusive-created inside the
same transaction and contains a unique ownership token; a manifest failure
rolls the database transaction back, and failure cleanup never removes a file
that lacks that exact token. The smoke does not enqueue email, target Web Push,
change the frozen test deployment, update Caddy, or push Git.
The local state records the exact application container. If that container is
recreated between phases, the wrapper refuses to continue instead of silently
targeting a replacement. Inspect the saved state and production records before
performing an explicit recovery; do not delete notifications or Oban jobs by a
broad timestamp or worker query.
## Run-scoped production browser verification
The production E2E workflow is opt-in and starts with a read-only identity and
scope check:
```bash
./scripts/production-full-e2e.sh plan production-e2e-YYYYMMDD
```
The plan verifies the production checkout, compact Compose project, external
database identity, healthy application container, and absence of an existing
fixture with the requested run ID. It prints an exact confirmation value but
does not create users or records.
After that exact scope has been reviewed and explicitly authorised, use the
printed value without changing the run ID:
```bash
WNH_PRODUCTION_E2E_CONFIRM='VALUE_PRINTED_BY_PLAN' \
./scripts/production-full-e2e.sh run production-e2e-YYYYMMDD
```
The run creates six uniquely prefixed synthetic users and only their associated
mutual-aid, activity, notification, audit, support, and legal fixture records.
The support record, general-removal notice, and separate TAKE IT DOWN notice are
inserted directly without email jobs and are only read through the staff UI;
the browser does not submit or moderate them.
Cleanup uses the mode-`0600` manifest of exact IDs on success, failure, or
interrupt, refuses cross-fixture relationships, deletes only matching jobs and
records, and verifies that the run prefix is absent. It never resets the
database. Evidence is stored below `output/production-full-e2e/<run-id>/`.
## Production release without pushing the frozen repository
The post-submission workflow keeps the public Git repository and
`test.whoneedhelp.com` untouched. A clean local commit is packaged as a
verified Git bundle and transferred directly over SSH to only
`/srv/who_need_help-production`:
```bash
./scripts/production-release.sh plan whoneedhelp
```
The default `plan` action is read-only. It verifies the exact local and remote
commits, requires a fast-forward history, checks the production checkout,
Compose scope and healthy containers, checks public readiness, opens a
read-only PostgreSQL connection, and runs the server-release environment
capability preflight. The candidate validator is streamed over SSH and checks
the existing server environment and Compose render before any bundle is
uploaded, so a validator change does not depend on the previously deployed
source tree. Before the first Google Play release, this preflight
allows only the absent Play App Signing certificate; the stricter
`check-environment-readiness.sh .env --require-release` remains the gate for
publishing Android through Google Play. The plan neither uploads a bundle nor
creates a backup.
Prepare and verify the immutable Git bundle and `linux/amd64` image archive on
the development workstation before authorising any production mutation:
```bash
./scripts/production-release-clean.sh prepare whoneedhelp
```
`prepare` repeats the read-only production and environment checks, reads the
production image-build inputs into a mode-`0600` temporary file, and then
creates or verifies the commit-bound artifacts under `output/releases/`. It
does not upload an artifact, create a production backup, change the remote
checkout, load an image, run a migration, or restart a service. A later
`apply` for the same commit verifies and reuses those exact artifacts instead
of compiling them again.
After reviewing the exact commit printed by the plan, execution additionally
requires an explicit per-commit confirmation:
```bash
WNH_PRODUCTION_RELEASE_CONFIRM=whoneedhelp.com:FULL_COMMIT \
./scripts/production-release.sh apply whoneedhelp
```
The application must still be released from a clean checkout. If a verified
bundle and image archive were prepared in another clean checkout, point the
release process at their absolute parent directory instead of rebuilding them:
```bash
WNH_PRODUCTION_RELEASE_ARTIFACT_ROOT=/absolute/path/to/output/releases \
WNH_PRODUCTION_RELEASE_CONFIRM=whoneedhelp.com:FULL_COMMIT \
./scripts/production-release.sh apply whoneedhelp
```
The selected directory must contain a child named with the exact full commit.
The workflow rechecks the Git bundle HEAD, bundle checksum and manifest, image
archive checksum, target platform, topology, immutable image names and image
IDs before transfer. It does not accept a relative artifact path or silently
fall back to a different commit. This allows a temporary clean worktree to
reuse the already scanned workstation artifact while keeping production image
compilation off the 4-GiB server.
The apply path refuses tracked local or remote modifications. It then:
1. creates and verifies a full Git bundle for exactly that clean commit;
2. reads the production environment over SSH into a mode-0600 temporary local
file, selects the candidate's immutable image tags, builds the required
`linux/amd64` images on the development workstation, and deletes that
temporary environment file;
3. records every image ID in a manifest, creates a timestamp-free
gzip-compressed Docker archive, verifies its SHA-256, and uploads the
bundle, archive, checksum, and manifest to the production checkout's ignored
`output/releases/`;
4. creates a custom-format PostgreSQL 18 backup without exposing the database
password in process arguments;
5. verifies its archive catalog and SHA-256, then copies and verifies the
backup again under local ignored `output/production-backups/`;
6. fast-forwards the production checkout without accessing or changing the
test checkout or public remote;
7. verifies the transferred archive and manifest, loads the ready images
without compiling on the production host, applies migrations, starts only
the application topology, and verifies public readiness through the
already-running shared edge plus the Android App Links endpoints.
Every newly added migration must have one reviewed entry in
`priv/repo/migration_application_compatibility.tsv`. `application_safe` means
the previously deployed application can run against the resulting schema;
`forward_only` means that it cannot be guaranteed. The plan prints the
aggregate policy. A forward-only apply requires a second exact confirmation:
```bash
WNH_PRODUCTION_FORWARD_ONLY_CONFIRM=whoneedhelp.com:FULL_COMMIT:forward-only \
WNH_PRODUCTION_RELEASE_CONFIRM=whoneedhelp.com:FULL_COMMIT \
./scripts/production-release.sh apply whoneedhelp
```
For a forward-only release, all candidate images are built and transferred
first, the old application is stopped before migration begins, and the target
application is started only after the migration runner succeeds. If anything
fails after the migration begins, the release deliberately leaves the old
application stopped and records that boundary in the release manifest.
Restarting an older image against a potentially incompatible schema is never
automatic.
For an `application_safe` release, an application startup failure restores the
previous immutable application image tags and attempts to recover public
readiness through the unchanged edge. A `forward_only` release never starts the old application
after migration begins. Neither path reverses Git source or Ecto migrations
automatically. The per-release rollback manifest and backup paths are recorded
below the production checkout's ignored `output/releases/`. The copied backup
is separate from the production host, but a long-term encrypted off-site
backup destination remains an operational requirement.
## Rollback boundary
The release image is immutable and migrations run as a separate one-shot role.
If the main workspace contains intentional tracked edits that must not enter
the release (for example local agent instructions), run the same workflow from
an exact clean detached worktree instead of stashing, discarding, or silently
including those edits:
```bash
./scripts/production-release-clean.sh plan whoneedhelp
./scripts/production-release-clean.sh prepare whoneedhelp
WNH_PRODUCTION_RELEASE_CONFIRM='whoneedhelp.com:EXACT_COMMIT' \
./scripts/production-release-clean.sh apply whoneedhelp
```
The wrapper selects the current `HEAD`, creates a task-owned detached worktree,
runs `production-release.sh` there with the original immutable artifact root,
and removes the worktree on success, failure, or interrupt. The underlying
release still requires the exact production confirmation and any applicable
forward-only migration confirmation. It does not stash, reset, stage, commit,
or copy changes from the main workspace.
Before a schema rollout, create and restore-test a current backup. Application
rollback and database migration rollback are separate decisions: do not run an
Ecto down migration merely because an image is rolled back. Inspect the exact
migration and compatibility boundary first.
Every successful SSH release prints a mode-`0600` `rollback-manifest.txt`.
Before changing production, inspect that exact manifest with the read-only
rollback plan:
```bash
./scripts/production-rollback.sh plan \
/srv/who_need_help-production/output/releases/RELEASE_ID/rollback-manifest.txt \
whoneedhelp
```
The rollback plan refuses manifests marked `forward_only`. Such releases must
be repaired forward after inspecting the exact migration state; an application
image rollback is available only when the manifest records
`migration_policy=application_safe`.
The plan requires the manifest target to be the currently checked-out
production commit, verifies the previous immutable application image still
exists, checks the pre-release backup catalog and checksum, and prints the
exact confirmation token. It does not change the remote environment or
containers.
After separately reviewing application/schema backward compatibility and
approving that exact scope, run:
```bash
WNH_PRODUCTION_ROLLBACK_CONFIRM=whoneedhelp.com:TARGET_COMMIT:PREVIOUS_COMMIT \
./scripts/production-rollback.sh apply \
/srv/who_need_help-production/output/releases/RELEASE_ID/rollback-manifest.txt \
whoneedhelp
```
This application rollback atomically restores the three previous application
image selectors and recreates only the selected application topology with
`--no-build`. It never changes or recreates the shared edge. It verifies the
resulting image identities, container health, public readiness through the
unchanged edge, and App Links. A failed rollback attempts to restore the
pre-rollback image selection. The Git checkout intentionally remains at
the newer source commit so the reviewed release tooling and manifest remain
available.
The command never restores PostgreSQL, reverses Ecto migrations, changes the
test deployment, or touches the public Git/Devpost submission. The repository
intentionally does not ship an automatic destructive production database
restore command.