who_need_help/docs/public-launch-checklist.md

13 KiB

Public launch checklist

This checklist separates properties the repository can verify from decisions that require the operator, a provider, or jurisdiction-specific legal review. Checking a box is an operator record, not proof created by the application. Do not replace unknown values with estimates.

1. Freeze the candidate

  • Record the candidate Git revision and immutable image references for the services that actually exist in the target topology, plus the externally hosted database version. Production has no project-local PostGIS or socket-proxy container.
  • Confirm that the target is the independent production checkout, Compose project, database, volumes, secrets, Google/Firebase project, SMTP credentials, and Android identity.
  • Keep development, hackathon test, and production credentials separate.
  • Run the repository quality suite and retain its non-secret evidence.
git rev-parse HEAD
./scripts/quality.sh

The 2026-08-09 isolated run covering the E2E-cleanup implementation now committed as 3c731b5 exited successfully: 451 ExUnit tests passed, all configured compiler/format/xref/Credo/Sobelow/Dialyzer/dependency/container/ Compose/Helm/migration/rollback/observability gates passed, and the final runtime image scan reported zero detected vulnerabilities. The final manifest-transport-only shell adjustment then passed bash -n, ShellCheck, and the complete production browser E2E replay. Exact identities and evidence paths are recorded in docs/verification.md.

A later 2026-08-09 local browser replay passed all 17 Chromium scenarios, including the new end-to-end support and legal-removal workflow. Its focused and full structured results are recorded in docs/verification.md. This is local isolated evidence only and does not mark the production support/legal or SMTP checklist items complete.

The latest complete local rerun on 2026-08-10 passed 459 ExUnit tests, fourteen browser-asset tests, every configured quality/security gate, and the final image scans after updating Phoenix LiveView from advisory-affected 1.2.8 to patched 1.2.9. Exact unit, timing, memory, cleanup, and dependency evidence are recorded in docs/verification.md.

The final 2026-08-12 email-boundary rerun passed 465 ExUnit tests, fourteen browser map-asset tests, the migration-registry completeness check, all configured quality/security and production-release drills, and the final image scans with zero detected vulnerabilities. The isolated unit completed in 4 minutes 49.760 seconds with a measured 213.5 MiB memory peak. A subsequent support/legal browser replay passed in Chromium, Firefox, and WebKit and left no run-scoped container, network, or temporary image. Exact email-flow, migration, and cleanup evidence is recorded in docs/verification.md.

The later production-E2E harness candidate 23776e2 passed the complete isolated pipeline on 2026-08-14 with 480 ExUnit tests, every configured quality and security gate, and zero detected runtime-image vulnerabilities. Production remains on clean application revision f1947f2264a6; the local commits after that revision change only the throwaway production-E2E verifier and its evidence. The frozen hackathon test deployment remained unchanged. Exact current-candidate and 2026-08-14 operations evidence is recorded in docs/verification.md.

2. Verify production configuration without exposing secrets

Run both checks against the single ignored production .env. The first reports capabilities; the second additionally enforces the repository's release requirements.

./scripts/validate-production-env.sh .env whoneedhelp.com
./scripts/check-environment-readiness.sh .env --require-release

The release check covers the application origin and secrets, database selection, SMTP and support routing, Google OAuth, browser VAPID, Firebase/FCM, Android package/signing configuration, and production App Links. It does not prove that external providers will deliver successfully after deployment. In particular, a configured SUPPORT_INBOX_ADDRESS is not evidence of inbound mail delivery: verify its MX routing and complete a real receive-and-reply test before using it in Google Play or public support pages.

If the optional Brevo inbound path is selected, also verify the dedicated receiving subdomain, exact SUPPORT_INBOUND_RECIPIENT, authenticated webhook, provider MessageId deduplication, confirmation-gated staff visibility, and a real reply. Do not advertise that address while either inbound setting is missing or before this external test passes.

The following decisions are intentionally not generated by code:

  • Launch jurisdictions and the exact local emergency contacts shown to users have been approved.
  • Privacy notice, Terms, medicine/prohibited-items guidance, and voluntary payment/reimbursement wording have been reviewed for those jurisdictions.
  • A retention and anonymisation schedule defines each account, safety, dispute, moderation, support, legal, tracking-evidence, and backup record.
  • The account-deletion executor and operator procedure have been approved and tested against that schedule before destructive execution is enabled.
  • Moderation, support, privacy, account-deletion, content-removal, and incident queues have named owners and monitored contact paths.
  • Provider terms and capacity are approved for email, map tiles, Google identity, Web Push, Firebase/FCM, database hosting, backups, and monitoring.
  • Google Analytics for Firebase remains disabled until the operator has approved a purpose and event allow-list, updated the privacy notice, implemented an explicit user analytics preference, and verified that the Android client does not initialise Analytics before the user opts in. Analytics events must not contain email addresses, account or device identifiers owned by Who Need Help, request/chat/support text, medicine details, exact or approximate coordinates, handover codes, social-account data, or moderation and safety evidence. The initial useful event set is limited to coarse product milestones such as registration_completed, request_created, helper_joined, handover_completed, and push_opened; every event and parameter requires a privacy review before release. Enabling Firebase Analytics later is a separate change from the server-side ProductAnalytics context, which stores only allow-listed daily aggregate counters without user identifiers.
  • Representative load measurements from the intended host and traffic shape justify database pools, action-limit policies, replica counts, memory/CPU allocation, alerts, and any scaling thresholds.
  • Backup ownership, encryption-key custody, off-site destination, restore procedure, and measured recovery objectives have been approved.
  • The independent stale-backup alert uses the operator-selected maximum age of 36 hours (129600 seconds). This is an alert threshold for the latest restore-verified heartbeat, not approval of retention, RPO, RTO, capacity, or encryption-key custody.

Until the retention decision and tested executor exist, account-deletion cases remain an audited operator workflow and automatic erasure stays disabled. The software must not claim that an unresolved legal or operational item is complete.

4. Rehearse recovery before promotion

Create a current custom-format backup and restore it only into the isolated drill database. Then rehearse the current release against that backup.

backup_output=$(./scripts/backup-compose.sh)
backup_path=$(printf '%s\n' "$backup_output" | sed -n 's/^Backup: //p')
./scripts/restore-drill-compose.sh "$backup_path"
./scripts/upgrade-rehearsal-compose.sh "$backup_path"
./scripts/production-rollback-drill.sh

Local backups and local MinIO drills verify mechanics but are not proof of off-site durability or a production recovery policy.

The implemented external workflow can be rehearsed without restoring over the source database:

./scripts/bootstrap-restic.sh
./scripts/production-offsite-backup.sh plan
./scripts/production-offsite-backup.sh run
systemctl --user status who-need-help-production-backup.service --no-pager
ssh buyvm-maya \
  'loginctl show-user simple --property=Linger --value'
ssh buyvm-maya \
  'systemctl --user status who-need-help-production-monitor.timer --no-pager'

The linger check must print yes; otherwise the enabled user timer can depend on an active login session and is not a persistent external monitor.

The 2026-08-25 pre-change check returned linger=no for simple on the selected external host. The reviewed change then enabled lingering for that exact user. A separate SSH session observed linger=yes, the monitor timer enabled and active, and a manually triggered service result of success with readiness, aggregate metrics, and restore-verified backup freshness all up. Per the installed loginctl(1) documentation, the enabled lingering state starts the user manager at boot and keeps it after logout.

Do not mark backup ownership complete merely because this command passes. The operator must still store the Restic key independently and approve retention, RPO, RTO, capacity, and responsible owners.

5. Promote with an explicit rollback point

Use the release workflow in docs/operations.md; do not copy mutable source over a running production checkout. A successful release must retain the generated rollback manifest, the pre-release backup checksum/catalog, and the previous immutable image references. A forward-only migration must be treated as forward-only rather than receiving an invented database rollback.

6. Verify the deployed product

  • Public HTTPS home, liveness, readiness, WebSocket upgrade, manifest, and assetlinks.json return the expected production identity.
  • Registration, returning-user login, and settings linking complete against the production Google OAuth client and exact callback origin.
  • A production-generated authentication email reaches an external mailbox and is DKIM-signed by the production domain.
  • Authentication-email action URLs use the production domain directly. The app-only production release at dcac2fa removed the HTML action anchor and leaves one visible copy-and-paste URL. A newly generated production message was inspected in Gmail on 2026-08-25: Gmail's detected link targeted https://whoneedhelp.com directly, no Brevo or UniSender tracking host appeared, and the one-time link completed the explicit production reauthentication flow. The follow-up app-only release at 305bdeb preserved the authenticated account context after an invalid or already-consumed link. A headed Chromium replay confirmed the link once, reopened the same link, and observed the invalid/expired notice together with the signed-in account email in the replacement-link form.
  • The Web Push provider accepts a production notification for a real active browser subscription without disabling the device.
  • A person has observed the resulting operating-system browser notification and its navigation target on the subscribed workstation.
  • The production Android build signs in, opens verified App Links, receives FCM, and performs user-started foreground location sharing on a physical device.
  • The full two-person help flow passes: create, discover, accept, chat, optional tracking, start, handover code, both confirmations, blind review, report/block, and notification delivery.
  • The Activity flow passes: create, request to join, approve/decline, withdraw, rejoin, group chat, exact-location disclosure only after approval, and reporting.
  • Support, privacy/data, account-deletion, general content-removal, and TAKE IT DOWN submissions reach the correct production operator queues. A 2026-08-21 run-scoped check proved authenticated support, privacy, data-export and account-deletion submissions plus separate general content-removal and TAKE IT DOWN records. Staff found every record in its permission-scoped production queue, and exact cleanup left zero records with the run prefix. A separate 2026-08-25 production mail smoke proved that an anonymous request remains outside staff search until its contact address is confirmed, then enters the staff queue exactly once. Reusing the confirmation link was idempotent, operator email remained disabled, and exact cleanup left no matching request, history, audit, or mail-job record.
  • Database, application, worker, email, push, backup, and edge monitoring are visible through public readiness plus bounded aggregate metrics and the restore-verified backup heartbeat. The independent monitor timer is enabled and active, its latest forced check reported every configured boundary up, and the external user has linger=yes so its user manager starts at boot and remains after logout.

Record observed timestamps, revision/image identities, and non-secret evidence paths in docs/verification.md. Record failed checks as failed; do not convert them into documentation-only success.