diff --git a/Dockerfile b/Dockerfile index faf25f2..1e71455 100644 --- a/Dockerfile +++ b/Dockerfile @@ -163,6 +163,7 @@ COPY priv priv COPY lib lib COPY test test COPY scripts/production-play-physical-fixture.exs scripts/production-play-physical-fixture.exs +COPY scripts/production-android-fcm-smoke.exs scripts/production-android-fcm-smoke.exs RUN mix compile diff --git a/docs/operations.md b/docs/operations.md index 3a63090..9c27d8f 100644 --- a/docs/operations.md +++ b/docs/operations.md @@ -1268,6 +1268,44 @@ Anonymous submissions receive the address-confirmation message before they enter the operator queue. Operator email alerts are disabled unless `SUPPORT_OPERATOR_EMAIL_MODE=immediate` is explicitly configured. +## Exact production Android FCM smoke + +The Android FCM smoke is split into explicit phases so the operator can inspect +the notification on the exact physical device before cleanup. The plan is +read-only and verifies the production Compose identity, running image, healthy +application container, database name, and one exact active Android FCM device: + +```bash +./scripts/production-android-fcm-smoke.sh \ + plan --check-only whoneedhelp.com +``` + +After the printed scope and exact device have been reviewed and separately +authorised, prepare one notification, inspect it on the phone, verify provider +delivery, and remove the run-owned records: + +```bash +./scripts/production-android-fcm-smoke.sh \ + prepare --confirm whoneedhelp.com +./scripts/production-android-fcm-smoke.sh \ + verify --from-state --confirm whoneedhelp.com +./scripts/production-android-fcm-smoke.sh \ + cleanup --from-state --confirm whoneedhelp.com +``` + +`prepare` creates one notification and one exact FCM delivery job in a single +database transaction. Its mode-`0600` manifest is exclusive-created inside the +same transaction and contains a unique ownership token; a manifest failure +rolls the database transaction back, and failure cleanup never removes a file +that lacks that exact token. The smoke does not enqueue email, target Web Push, +change the frozen test deployment, update Caddy, or push Git. + +The local state records the exact application container. If that container is +recreated between phases, the wrapper refuses to continue instead of silently +targeting a replacement. Inspect the saved state and production records before +performing an explicit recovery; do not delete notifications or Oban jobs by a +broad timestamp or worker query. + ## Run-scoped production browser verification The production E2E workflow is opt-in and starts with a read-only identity and diff --git a/docs/public-launch-checklist.md b/docs/public-launch-checklist.md index b71ff52..87d1c94 100644 --- a/docs/public-launch-checklist.md +++ b/docs/public-launch-checklist.md @@ -164,6 +164,14 @@ ssh buyvm-maya \ The linger check must print `yes`; otherwise the enabled user timer can depend on an active login session and is not a persistent external monitor. +A read-only check on 2026-08-25 returned `linger=no` for `simple` on the +selected external host. The timer was loaded, enabled, and active while the +user manager was running, and its latest completed check reported readiness, +metrics, and backup freshness as `up`; those facts do not prove that the user +manager survives logout or starts at boot. Persistent off-site monitoring +therefore remains open until an administrator enables lingering and the timer +is rechecked after a login-independent restart. + Do not mark backup ownership complete merely because this command passes. The operator must still store the Restic key independently and approve retention, RPO, RTO, capacity, and responsible owners. @@ -212,8 +220,10 @@ as forward-only rather than receiving an invented database rollback. with the run prefix. Anonymous contact verification still requires a controlled production mail smoke before this combined item can be checked. -- [x] Database, application, worker, email, push, backup, and edge monitoring - are visible to the responsible operator. +- [ ] Database, application, worker, email, push, backup, and edge monitoring + are visible to the responsible operator and persist independently of an + interactive login. Current checks prove visibility while the external + user manager is running, but that host still reports `linger=no`. Record observed timestamps, revision/image identities, and non-secret evidence paths in `docs/verification.md`. Record failed checks as failed; do not convert diff --git a/docs/verification.md b/docs/verification.md index 7f33b16..de3acbc 100644 --- a/docs/verification.md +++ b/docs/verification.md @@ -3427,3 +3427,36 @@ promoted. image carrying run identifier `20260820232333-2842586`. The frozen hackathon deployment, production deployment, Caddy, and public Git remote were not changed by this local verification. + +# 2026-08-25 manifest ownership and external-monitor final recheck + +- The production Android FCM smoke and production browser E2E preparation now + exclusive-create their mode-`0600` manifests inside the same database + transaction as the run-owned records. Each manifest contains a UUID ownership + token. Failure cleanup removes a manifest only when that exact token matches, + so a concurrent process cannot cause one run to delete another run's state. +- The external-monitor installer now reads and validates the selected remote + account's actual home directory. Remote systemd, state, configuration, and + backup-heartbeat paths are derived from that observed home instead of assuming + `/home/simple`. The installer regression suite passed all six cases, including + a non-default `/home/monitor-user` account. +- A read-only external-host check observed account `simple` with home + `/home/simple` and `linger=no`. Its user timer was loaded, enabled, and active + while the user manager was running; the latest completed monitor check + reported readiness, authenticated metrics, and backup freshness as `up`. + This proves the current check completed, but it does not prove persistence + after logout or boot. No lingering setting was changed. +- The final isolated quality unit + `codex-heavy-wnh-final-quality-manifest-monitor-20260825-20260825-175845-502195.service` + exited successfully after 4 minutes 55.976 seconds. It reported 487 passing + ExUnit tests, 18 passing browser-asset tests, all configured shell/Python + regression suites, and all isolated quality and security gates passing. The + final Debian 13.6 production image and the pinned infrastructure images + reported zero detected vulnerabilities at the configured scan severity. +- The unit's measured memory peak was 438.9 MiB and swap use was zero. Exact + post-run inspection found no container, network, volume, or temporary image + from Compose project `wnh_quality_20260825145846502200` or run identifier + `20260825145846-502200`. +- This was local verification only. Production, the frozen hackathon test + deployment, shared Caddy, and the public Git remote were not changed or + pushed. diff --git a/lib/mix/tasks/wnh.staging_full_e2e.ex b/lib/mix/tasks/wnh.staging_full_e2e.ex index 182a3a4..ec981ab 100644 --- a/lib/mix/tasks/wnh.staging_full_e2e.ex +++ b/lib/mix/tasks/wnh.staging_full_e2e.ex @@ -139,65 +139,72 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do assert_prefix_unused!(context) - password_hash = Bcrypt.hash_pwd_salt(password) - now = DateTime.utc_now(:second) - - {:ok, - %{ - users: users, - support_request: support_request, - removal_notice: removal_notice, - take_it_down_notice: take_it_down_notice - }} = - Repo.transaction(fn -> - users = - Map.new(@precreated_roles, fn role -> - user = - insert_user!( - context.emails[role], - display_name(role), - password_hash, - now - ) - - if role == "admin" do - %StaffRoleAssignment{} - |> StaffRoleAssignment.changeset(%{user_id: user.id, role: :admin}) - |> Repo.insert!() - end - - {role, %{"id" => user.id, "email" => user.email}} - end) - - requester = users["requester"] - - support_request = insert_support_fixture!(context, requester, now) - removal_notice = insert_removal_fixture!(context, requester, now) - take_it_down_notice = insert_take_it_down_fixture!(context, requester, now) - - %{ - users: users, - support_request: support_request, - removal_notice: removal_notice, - take_it_down_notice: take_it_down_notice - } - end) - - manifest = %{ - "schema_version" => 3, - "run_id" => context.run_id, - "database" => context.database, - "precreated_users" => users, - "precreated_records" => %{ - "support_request" => support_request.id, - "content_removal_notice" => removal_notice.id, - "take_it_down_notice" => take_it_down_notice.id - }, - "allowed_emails" => context.emails |> Map.values() |> Enum.sort() - } + if File.exists?(context.manifest_path) do + Mix.raise("full staging E2E manifest already exists") + end context.manifest_path |> Path.dirname() |> File.mkdir_p!() - write_manifest!(context.manifest_path, manifest) + + password_hash = Bcrypt.hash_pwd_salt(password) + now = DateTime.utc_now(:second) + manifest_ownership_token = Ecto.UUID.generate() + + try do + {:ok, _fixture} = + Repo.transaction(fn -> + users = + Map.new(@precreated_roles, fn role -> + user = + insert_user!( + context.emails[role], + display_name(role), + password_hash, + now + ) + + if role == "admin" do + %StaffRoleAssignment{} + |> StaffRoleAssignment.changeset(%{user_id: user.id, role: :admin}) + |> Repo.insert!() + end + + {role, %{"id" => user.id, "email" => user.email}} + end) + + requester = users["requester"] + + support_request = insert_support_fixture!(context, requester, now) + removal_notice = insert_removal_fixture!(context, requester, now) + take_it_down_notice = insert_take_it_down_fixture!(context, requester, now) + + manifest = %{ + "schema_version" => 3, + "run_id" => context.run_id, + "database" => context.database, + "ownership_token" => manifest_ownership_token, + "precreated_users" => users, + "precreated_records" => %{ + "support_request" => support_request.id, + "content_removal_notice" => removal_notice.id, + "take_it_down_notice" => take_it_down_notice.id + }, + "allowed_emails" => context.emails |> Map.values() |> Enum.sort() + } + + write_new_manifest!(context.manifest_path, manifest) + + %{ + users: users, + support_request: support_request, + removal_notice: removal_notice, + take_it_down_notice: take_it_down_notice + } + end) + rescue + exception -> + remove_owned_prepare_manifest(context, manifest_ownership_token) + reraise exception, __STACKTRACE__ + end Mix.shell().info("prepared exact full staging E2E users for #{context.run_id}") end @@ -489,6 +496,7 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do unless manifest["schema_version"] == 3 and manifest["run_id"] == context.run_id and manifest["database"] == context.database and + uuid?(manifest["ownership_token"]) and manifest["allowed_emails"] == expected_emails and valid_precreated? and valid_precreated_records? do Mix.raise("full staging E2E manifest does not match the requested run and database") @@ -770,6 +778,37 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do File.rename!(temporary_path, path) end + defp write_new_manifest!(path, manifest) do + encoded = Jason.encode_to_iodata!(manifest, pretty: true) + + File.open!(path, [:write, :binary, :exclusive], fn file -> + IO.binwrite(file, encoded) + end) + + try do + File.chmod!(path, 0o600) + rescue + exception -> + File.rm(path) + reraise exception, __STACKTRACE__ + end + end + + defp remove_owned_prepare_manifest(context, ownership_token) do + expected_emails = context.emails |> Map.values() |> Enum.sort() + + with {:ok, encoded} <- File.read(context.manifest_path), + {:ok, manifest} <- Jason.decode(encoded), + true <- manifest["run_id"] == context.run_id, + true <- manifest["database"] == context.database, + true <- manifest["allowed_emails"] == expected_emails, + true <- manifest["ownership_token"] == ownership_token do + File.rm(context.manifest_path) + else + _missing_or_different_manifest -> :ok + end + end + defp insert_user!(email, display_name, password_hash, now) do %User{} |> User.registration_changeset(%{ diff --git a/scripts/install-production-external-monitor.sh b/scripts/install-production-external-monitor.sh index 84f938d..7ffecb4 100755 --- a/scripts/install-production-external-monitor.sh +++ b/scripts/install-production-external-monitor.sh @@ -6,9 +6,9 @@ ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd) source_target=${PRODUCTION_SSH_TARGET:-whoneedhelp} monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya} production_env=${PRODUCTION_ENV_PATH:-/srv/who_need_help-production/.env} -remote_root=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help} -remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json} -remote_state=${MONITOR_REMOTE_STATE:-/home/simple/.local/state/who-need-help/monitor.json} +remote_root=${MONITOR_REMOTE_ROOT:-} +remote_config=${MONITOR_REMOTE_CONFIG:-} +remote_state=${MONITOR_REMOTE_STATE:-} monitor_smtp_values_file=${MONITOR_SMTP_VALUES_FILE:-} monitor_install_work_root=${MONITOR_INSTALL_WORK_ROOT:-} monitor_url=${MONITOR_URL:-https://whoneedhelp.com/healthz/ready} @@ -16,7 +16,7 @@ metrics_url=${MONITOR_METRICS_URL:-https://whoneedhelp.com/metrics} monitor_calendar=${MONITOR_ON_CALENDAR:-'*:0/1'} health_timeout=${MONITOR_HEALTH_TIMEOUT_SECONDS:-3} metrics_timeout=${MONITOR_METRICS_TIMEOUT_SECONDS:-3} -backup_heartbeat_path=${MONITOR_BACKUP_HEARTBEAT_PATH:-/home/simple/.local/state/who-need-help/production-backup.json} +backup_heartbeat_path=${MONITOR_BACKUP_HEARTBEAT_PATH:-} backup_max_age=${MONITOR_BACKUP_MAX_AGE_SECONDS:-} for command in awk install mktemp python3 realpath scp ssh stat; do @@ -77,31 +77,38 @@ case "$metrics_timeout" in 0) echo "MONITOR_METRICS_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;; esac +if ! monitor_identity=$(ssh -o BatchMode=yes "$monitor_target" \ + 'printf "%s\n%s\n%s\n" "$(id -un)" "$HOME" "$(loginctl show-user "$(id -un)" --property=Linger --value)"'); then + echo "Unable to verify systemd user lingering on the external monitor host." >&2 + exit 2 +fi +monitor_user=$(printf '%s\n' "$monitor_identity" | awk 'NR == 1 {print; exit}') +monitor_home=$(printf '%s\n' "$monitor_identity" | awk 'NR == 2 {print; exit}') +monitor_linger=$(printf '%s\n' "$monitor_identity" | awk 'NR == 3 {print; exit}') +if [ -z "$monitor_user" ] || [ -z "$monitor_home" ] || [ "$monitor_linger" != yes ]; then + echo "External monitor installation requires systemd user lingering for '$monitor_user' on '$monitor_target'." >&2 + echo "An administrator must run: sudo loginctl enable-linger '$monitor_user'" >&2 + echo "Verify with: loginctl show-user '$monitor_user' --property=Linger --value" >&2 + exit 2 +fi + +remote_root=${remote_root:-$monitor_home/.local/lib/who-need-help} +remote_config=${remote_config:-$monitor_home/.config/who-need-help/monitor.json} +remote_state=${remote_state:-$monitor_home/.local/state/who-need-help/monitor.json} +backup_heartbeat_path=${backup_heartbeat_path:-$monitor_home/.local/state/who-need-help/production-backup.json} +remote_systemd_dir=$monitor_home/.config/systemd/user + if [ -n "$backup_max_age" ]; then case "$backup_max_age" in *[!0-9]*) echo "MONITOR_BACKUP_MAX_AGE_SECONDS must be a positive integer." >&2; exit 2 ;; 0) echo "MONITOR_BACKUP_MAX_AGE_SECONDS must be a positive integer." >&2; exit 2 ;; esac - if [ "$backup_heartbeat_path" != /home/simple/.local/state/who-need-help/production-backup.json ]; then + if [ "$backup_heartbeat_path" != "$monitor_home/.local/state/who-need-help/production-backup.json" ]; then echo "MONITOR_BACKUP_HEARTBEAT_PATH is outside the reviewed monitor state path." >&2 exit 2 fi fi -if ! monitor_identity=$(ssh -o BatchMode=yes "$monitor_target" \ - 'printf "%s:%s\n" "$(id -un)" "$(loginctl show-user "$(id -un)" --property=Linger --value)"'); then - echo "Unable to verify systemd user lingering on the external monitor host." >&2 - exit 2 -fi -monitor_user=${monitor_identity%%:*} -monitor_linger=${monitor_identity#*:} -if [ -z "$monitor_user" ] || [ "$monitor_linger" != yes ]; then - echo "External monitor installation requires systemd user lingering for '$monitor_user' on '$monitor_target'." >&2 - echo "An administrator must run: sudo loginctl enable-linger '$monitor_user'" >&2 - echo "Verify with: loginctl show-user '$monitor_user' --property=Linger --value" >&2 - exit 2 -fi - work_dir=$(mktemp -d "$monitor_install_work_root/monitor-install.XXXXXX") config="$work_dir/monitor.json" service="$work_dir/who-need-help-production-monitor.service" @@ -224,15 +231,15 @@ WantedBy=timers.target EOF ssh -o BatchMode=yes "$monitor_target" \ - "install -d -m 700 '$remote_root' /home/simple/.config/who-need-help /home/simple/.local/state/who-need-help /home/simple/.config/systemd/user" + "install -d -m 700 '$remote_root' '$monitor_home/.config/who-need-help' '$monitor_home/.local/state/who-need-help' '$remote_systemd_dir'" scp -q "$ROOT/scripts/production-external-monitor.py" \ "$monitor_target:$remote_root/production-external-monitor.py" scp -q "$config" "$monitor_target:$remote_config" -scp -q "$service" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.service" -scp -q "$timer" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.timer" +scp -q "$service" "$monitor_target:$remote_systemd_dir/who-need-help-production-monitor.service" +scp -q "$timer" "$monitor_target:$remote_systemd_dir/who-need-help-production-monitor.timer" ssh -o BatchMode=yes "$monitor_target" \ - "chmod 700 '$remote_root/production-external-monitor.py'; chmod 600 '$remote_config' /home/simple/.config/systemd/user/who-need-help-production-monitor.service /home/simple/.config/systemd/user/who-need-help-production-monitor.timer; systemctl --user daemon-reload; systemctl --user start who-need-help-production-monitor.service; systemctl --user enable --now who-need-help-production-monitor.timer" + "chmod 700 '$remote_root/production-external-monitor.py'; chmod 600 '$remote_config' '$remote_systemd_dir/who-need-help-production-monitor.service' '$remote_systemd_dir/who-need-help-production-monitor.timer'; systemctl --user daemon-reload; systemctl --user start who-need-help-production-monitor.service; systemctl --user enable --now who-need-help-production-monitor.timer" printf 'External monitor installed on %s (%s).\n' "$monitor_target" "$monitor_host" printf 'Persistent systemd user manager verified for %s (linger=yes).\n' "$monitor_user" diff --git a/scripts/lib/postgres-readiness.sh b/scripts/lib/postgres-readiness.sh index 90d64cf..db3c94f 100644 --- a/scripts/lib/postgres-readiness.sh +++ b/scripts/lib/postgres-readiness.sh @@ -10,30 +10,21 @@ capture_postgres_readiness_evidence() { wait_for_postgres_final_ready() { local container=$1 local evidence_file=$2 - local max_attempts=$3 - local interval_seconds=$4 - local attempt=0 + local interval_seconds=$3 local running local health while ! docker logs "$container" 2>&1 | grep -Fq 'PostgreSQL init process complete; ready for start up.'; do - attempt=$((attempt + 1)) running=$(docker inspect --format '{{.State.Running}}' "$container" 2>/dev/null || printf missing) if [[ "$running" != true ]]; then capture_postgres_readiness_evidence "$container" "$evidence_file" echo "Isolated restore PostgreSQL exited during initialization." >&2 return 1 fi - if ((attempt >= max_attempts)); then - capture_postgres_readiness_evidence "$container" "$evidence_file" - echo "Isolated restore PostgreSQL did not finish initialization." >&2 - return 1 - fi sleep "$interval_seconds" done - attempt=0 while true; do # Docker can retain a cached healthy state from the temporary bootstrap # postmaster while that server shuts down. Query the final server directly @@ -47,7 +38,6 @@ wait_for_postgres_final_ready() { return 0 fi - attempt=$((attempt + 1)) running=$(docker inspect --format '{{.State.Running}}' "$container" 2>/dev/null || printf missing) health=$(docker inspect \ --format '{{if .State.Health}}{{.State.Health.Status}}{{else}}missing{{end}}' \ @@ -58,7 +48,7 @@ wait_for_postgres_final_ready() { echo "Isolated restore PostgreSQL exited before accepting SQL." >&2 return 1 fi - if [[ "$health" == unhealthy ]] || ((attempt >= max_attempts)); then + if [[ "$health" == unhealthy ]]; then capture_postgres_readiness_evidence "$container" "$evidence_file" echo "Isolated restore PostgreSQL did not become SQL-ready." >&2 return 1 diff --git a/scripts/production-android-fcm-smoke.exs b/scripts/production-android-fcm-smoke.exs index 7a2a0a6..b16c939 100644 --- a/scripts/production-android-fcm-smoke.exs +++ b/scripts/production-android-fcm-smoke.exs @@ -50,6 +50,8 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do defp prepare(context) do if File.exists?(context.manifest_path), do: raise("manifest already exists") + manifest_ownership_token = Ecto.UUID.generate() + device = Repo.get(PushDevice, context.device_id) unless match?( @@ -67,41 +69,49 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do raise "run-scoped notification already exists" end - {:ok, fixture} = - Repo.transaction(fn -> - notification = - %Notification{} - |> Notification.changeset(%{ - user_id: device.user_id, - kind: :support_update, - title: "Who Need Help notification check", - body: "Android notifications are working.", - path: "/notifications", - data: %{"run_id" => context.run_id, "synthetic" => true}, - idempotency_key: context.idempotency_key - }) - |> Repo.insert!() + fixture = + try do + {:ok, committed_fixture} = + Repo.transaction(fn -> + notification = + %Notification{} + |> Notification.changeset(%{ + user_id: device.user_id, + kind: :support_update, + title: "Who Need Help notification check", + body: "Android notifications are working.", + path: "/notifications", + data: %{"run_id" => context.run_id, "synthetic" => true}, + idempotency_key: context.idempotency_key + }) + |> Repo.insert!() - job = - %{"notification_id" => notification.id, "device_id" => device.id} - |> DeviceDeliveryWorker.new() - |> Oban.insert!() + job = + %{"notification_id" => notification.id, "device_id" => device.id} + |> DeviceDeliveryWorker.new() + |> Oban.insert!() - %{notification: notification, job: job} - end) + manifest = %{ + "schema_version" => 1, + "run_id" => context.run_id, + "database" => context.database, + "ownership_token" => manifest_ownership_token, + "idempotency_key" => context.idempotency_key, + "notification_id" => notification.id, + "device_id" => device.id, + "job_id" => job.id + } - manifest = %{ - "schema_version" => 1, - "run_id" => context.run_id, - "database" => context.database, - "idempotency_key" => context.idempotency_key, - "notification_id" => fixture.notification.id, - "device_id" => device.id, - "job_id" => fixture.job.id - } + persist_manifest!(context.manifest_path, manifest) + %{notification: notification, job: job} + end) - File.write!(context.manifest_path, Jason.encode_to_iodata!(manifest, pretty: true)) - File.chmod!(context.manifest_path, 0o600) + committed_fixture + rescue + exception -> + remove_owned_manifest(context, manifest_ownership_token) + reraise exception, __STACKTRACE__ + end IO.puts("android_fcm_smoke_prepared=true") IO.puts("job_id=#{fixture.job.id}") @@ -184,6 +194,7 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do valid? = manifest["schema_version"] == 1 and manifest["run_id"] == context.run_id and manifest["database"] == context.database and + uuid?(manifest["ownership_token"]) and manifest["device_id"] == context.device_id and manifest["idempotency_key"] == context.idempotency_key and uuid?(manifest["notification_id"]) and is_integer(manifest["job_id"]) @@ -192,6 +203,36 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do manifest end + defp persist_manifest!(path, manifest) do + encoded = Jason.encode_to_iodata!(manifest, pretty: true) + + File.open!(path, [:write, :binary, :exclusive], fn file -> + IO.binwrite(file, encoded) + end) + + try do + File.chmod!(path, 0o600) + rescue + exception -> + File.rm(path) + reraise exception, __STACKTRACE__ + end + end + + defp remove_owned_manifest(context, ownership_token) do + with {:ok, encoded} <- File.read(context.manifest_path), + {:ok, manifest} <- Jason.decode(encoded), + true <- manifest["run_id"] == context.run_id, + true <- manifest["database"] == context.database, + true <- manifest["device_id"] == context.device_id, + true <- manifest["idempotency_key"] == context.idempotency_key, + true <- manifest["ownership_token"] == ownership_token do + File.rm(context.manifest_path) + else + _missing_or_different_manifest -> :ok + end + end + defp required_option!(options, name) do case Map.get(options, name) do value when is_binary(value) and value != "" -> value diff --git a/scripts/production-offsite-backup.sh b/scripts/production-offsite-backup.sh index 8806f42..63ecdf7 100755 --- a/scripts/production-offsite-backup.sh +++ b/scripts/production-offsite-backup.sh @@ -255,7 +255,6 @@ restore_snapshot() { wait_for_postgres_final_ready \ "$restore_container" \ "$evidence_dir/restore-postgres.log" \ - 120 \ 1 # The PostGIS image initializes its requested database with PostGIS already diff --git a/test/scripts/install_production_external_monitor_test.py b/test/scripts/install_production_external_monitor_test.py index 0c7bf42..6e1434f 100644 --- a/test/scripts/install_production_external_monitor_test.py +++ b/test/scripts/install_production_external_monitor_test.py @@ -18,6 +18,7 @@ class InstallProductionExternalMonitorTest(unittest.TestCase): self.fake_bin = self.root / "bin" self.fake_bin.mkdir() self.capture = self.root / "captured-monitor.json" + self.scp_log = self.root / "scp.log" self.work_root = self.root / "work" self.work_root.mkdir() self.production_env = self.root / "production.env" @@ -85,7 +86,9 @@ class InstallProductionExternalMonitorTest(unittest.TestCase): command = sys.argv[-1] if "loginctl show-user" in command: - print(f"monitor-user:{os.environ.get('FAKE_MONITOR_LINGER', 'yes')}") + print("monitor-user") + print(os.environ.get("FAKE_MONITOR_HOME", "/home/monitor-user")) + print(os.environ.get("FAKE_MONITOR_LINGER", "yes")) raise SystemExit(0) if command.startswith("python3 - "): result = subprocess.run( @@ -112,6 +115,9 @@ class InstallProductionExternalMonitorTest(unittest.TestCase): import sys from pathlib import Path + with open(os.environ["FAKE_SCP_LOG"], "a", encoding="utf-8") as handle: + handle.write(" ".join(sys.argv[1:]) + "\\n") + sources = [Path(value) for value in sys.argv[1:-1] if not value.startswith("-")] for source in sources: if source.name == "monitor.json": @@ -122,6 +128,7 @@ class InstallProductionExternalMonitorTest(unittest.TestCase): def run_installer(self, *, override, backup_max_age=None, linger="yes"): self.capture.unlink(missing_ok=True) + self.scp_log.unlink(missing_ok=True) env = os.environ.copy() env.update( { @@ -130,8 +137,10 @@ class InstallProductionExternalMonitorTest(unittest.TestCase): "MONITOR_SSH_TARGET": "monitor", "PRODUCTION_ENV_PATH": str(self.production_env), "FAKE_MONITOR_CAPTURE": str(self.capture), + "FAKE_SCP_LOG": str(self.scp_log), "MONITOR_INSTALL_WORK_ROOT": str(self.work_root), "FAKE_MONITOR_LINGER": linger, + "FAKE_MONITOR_HOME": "/home/monitor-user", } ) if override: @@ -182,7 +191,7 @@ class InstallProductionExternalMonitorTest(unittest.TestCase): self.assertEqual( captured["backup"], { - "heartbeat_path": "/home/simple/.local/state/who-need-help/production-backup.json", + "heartbeat_path": "/home/monitor-user/.local/state/who-need-help/production-backup.json", "max_age_seconds": 172800, }, ) @@ -211,6 +220,13 @@ class InstallProductionExternalMonitorTest(unittest.TestCase): "Persistent systemd user manager verified for monitor-user (linger=yes).", result.stdout, ) + scp_log = self.scp_log.read_text(encoding="utf-8") + self.assertIn( + "monitor:/home/monitor-user/.config/systemd/user/" + "who-need-help-production-monitor.service", + scp_log, + ) + self.assertNotIn("/home/simple", scp_log) if __name__ == "__main__": diff --git a/test/scripts/production_android_fcm_smoke_test.exs b/test/scripts/production_android_fcm_smoke_test.exs new file mode 100644 index 0000000..9c7eb1e --- /dev/null +++ b/test/scripts/production_android_fcm_smoke_test.exs @@ -0,0 +1,107 @@ +Code.require_file(Path.expand("../../scripts/production-android-fcm-smoke.exs", __DIR__)) + +defmodule WhoNeedHelp.ProductionAndroidFCMSmokeTest do + use WhoNeedHelp.DataCase, async: false + use Oban.Testing, repo: WhoNeedHelp.Repo + + import Ecto.Query + import ExUnit.CaptureIO + import WhoNeedHelp.AccountsFixtures + + alias WhoNeedHelp.Notifications + alias WhoNeedHelp.Notifications.Notification + alias WhoNeedHelp.Push.DeviceDeliveryWorker + alias WhoNeedHelp.Repo + + setup do + user = user_fixture(display_name: "Production FCM smoke test user") + scope = user_scope_fixture(user) + + {:ok, device} = + Notifications.register_device(scope, %{ + "platform" => "android", + "provider" => "fcm", + "token" => "test-fcm-token-#{Ecto.UUID.generate()}", + "installation_id" => Ecto.UUID.generate(), + "device_label" => "Production FCM smoke test" + }) + + %Postgrex.Result{rows: [[database]]} = + Repo.query!("SELECT current_database()", [], log: false) + + %{device: device, database: database} + end + + test "manifest write failure rolls back the notification and exact delivery job", context do + run_id = run_id() + missing_parent = "/tmp/wnh-production-android-fcm-missing-#{run_id}" + manifest_path = Path.join(missing_parent, "manifest.json") + File.rm_rf!(missing_parent) + + options = options(context, run_id, manifest_path) + worker = to_string(DeviceDeliveryWorker) + jobs_before = Repo.aggregate(from(job in Oban.Job, where: job.worker == ^worker), :count) + + assert_raise File.Error, fn -> + WhoNeedHelp.ProductionAndroidFCMSmoke.run("prepare", options) + end + + refute Repo.get_by(Notification, + idempotency_key: "production-android-fcm-smoke:#{run_id}" + ) + + assert Repo.aggregate(from(job in Oban.Job, where: job.worker == ^worker), :count) == + jobs_before + + refute File.exists?(manifest_path) + end + + test "prepare and cleanup retain an exact manifest-bound scope", context do + run_id = run_id() + manifest_path = "/tmp/wnh-production-android-fcm-#{run_id}.json" + File.rm(manifest_path) + on_exit(fn -> File.rm(manifest_path) end) + options = options(context, run_id, manifest_path) + + assert capture_io(fn -> + WhoNeedHelp.ProductionAndroidFCMSmoke.run("prepare", options) + end) =~ "android_fcm_smoke_prepared=true" + + assert {:ok, stat} = File.stat(manifest_path) + assert stat.mode |> Bitwise.band(0o777) == 0o600 + + manifest = manifest_path |> File.read!() |> Jason.decode!() + assert {:ok, _ownership_token} = Ecto.UUID.cast(manifest["ownership_token"]) + + assert %Notification{id: notification_id} = + Repo.get_by(Notification, + idempotency_key: "production-android-fcm-smoke:#{run_id}" + ) + + assert notification_id == manifest["notification_id"] + assert context.device.id == manifest["device_id"] + assert %Oban.Job{id: job_id} = Repo.get(Oban.Job, manifest["job_id"]) + assert job_id == manifest["job_id"] + + assert capture_io(fn -> + WhoNeedHelp.ProductionAndroidFCMSmoke.run("cleanup", options) + end) =~ "android_fcm_smoke_cleanup_verified=true" + + refute Repo.get(Notification, notification_id) + refute Repo.get(Oban.Job, job_id) + refute File.exists?(manifest_path) + end + + defp options(context, run_id, manifest_path) do + %{ + run_id: run_id, + expected_database: context.database, + device_id: context.device.id, + manifest_path: manifest_path + } + end + + defp run_id do + "test-#{System.unique_integer([:positive, :monotonic])}-#{Ecto.UUID.generate()}" + end +end diff --git a/test/scripts/production_offsite_backup_readiness_test.sh b/test/scripts/production_offsite_backup_readiness_test.sh index da64211..25288b6 100644 --- a/test/scripts/production_offsite_backup_readiness_test.sh +++ b/test/scripts/production_offsite_backup_readiness_test.sh @@ -48,7 +48,7 @@ sleep() { evidence=$(mktemp) trap 'rm -f "$evidence"' EXIT -wait_for_postgres_final_ready restore-probe "$evidence" 3 0 +wait_for_postgres_final_ready restore-probe "$evidence" 0 if ((sql_attempts != 2)); then printf 'expected two direct SQL readiness attempts, observed %s\n' "$sql_attempts" >&2 @@ -64,7 +64,7 @@ mode=unhealthy sql_attempts=0 : >"$evidence" -if wait_for_postgres_final_ready restore-probe "$evidence" 3 0; then +if wait_for_postgres_final_ready restore-probe "$evidence" 0; then echo 'unhealthy final PostgreSQL unexpectedly passed readiness' >&2 exit 1 fi diff --git a/test/who_need_help/staging_full_e2e_cleanup_test.exs b/test/who_need_help/staging_full_e2e_cleanup_test.exs index 0eff203..ccd39d6 100644 --- a/test/who_need_help/staging_full_e2e_cleanup_test.exs +++ b/test/who_need_help/staging_full_e2e_cleanup_test.exs @@ -2,6 +2,7 @@ defmodule WhoNeedHelp.StagingFullE2ECleanupTest do use WhoNeedHelp.DataCase, async: false use Oban.Testing, repo: WhoNeedHelp.Repo + import Ecto.Query import ExUnit.CaptureIO alias Oban.Job @@ -9,6 +10,50 @@ defmodule WhoNeedHelp.StagingFullE2ECleanupTest do alias WhoNeedHelp.Mail.{ContentRemovalEmailWorker, SupportConfirmationWorker} alias WhoNeedHelp.Push.NearbyMatchWorker + test "prepare refuses a pre-existing manifest before creating fixture users" do + run_id = "existing-manifest-#{System.unique_integer([:positive])}" + manifest_path = "/output/wnh-staging-full-e2e-#{run_id}.json" + password = "staging-existing-#{System.unique_integer([:positive])}" + original_manifest = "operator-owned manifest\n" + + %Postgrex.Result{rows: [[database]]} = + Repo.query!("SELECT current_database()", [], log: false) + + env = %{ + "WNH_STAGING_E2E_CONFIRM" => "public-staging-full-e2e", + "WNH_STAGING_E2E_EXPECTED_DATABASE" => database, + "WNH_STAGING_E2E_MANIFEST_PATH" => manifest_path, + "WNH_STAGING_E2E_PASSWORD" => password, + "WNH_STAGING_E2E_RUN_ID" => run_id + } + + previous_env = Map.new(env, fn {name, _value} -> {name, System.get_env(name)} end) + Enum.each(env, fn {name, value} -> System.put_env(name, value) end) + File.mkdir_p!(Path.dirname(manifest_path)) + File.write!(manifest_path, original_manifest) + + on_exit(fn -> + File.rm(manifest_path) + + Enum.each(previous_env, fn + {name, nil} -> System.delete_env(name) + {name, value} -> System.put_env(name, value) + end) + end) + + assert_raise Mix.Error, "full staging E2E manifest already exists", fn -> + run_task("prepare") + end + + assert File.read!(manifest_path) == original_manifest + + refute Repo.exists?( + from(user in WhoNeedHelp.Accounts.User, + where: like(user.email, ^"wnh-staging-e2e-#{run_id}-%") + ) + ) + end + test "cleanup removes only run-owned support and legal mail jobs" do run_id = "cleanup-#{System.unique_integer([:positive])}" manifest_path = "/output/wnh-staging-full-e2e-#{run_id}.json" @@ -76,6 +121,7 @@ defmodule WhoNeedHelp.StagingFullE2ECleanupTest do |> Repo.insert!() assert manifest["schema_version"] == 3 + assert {:ok, _ownership_token} = Ecto.UUID.cast(manifest["ownership_token"]) assert prepared_support_request.subject == "Production E2E support #{run_id}" assert prepared_support_request.contact_verified_at assert prepared_removal_notice.regime == :general