Harden production verification ownership

This commit is contained in:
SimpleTest 2026-08-25 18:07:09 +03:00
parent 4f81c38717
commit 7094f21394
13 changed files with 457 additions and 130 deletions

View File

@ -163,6 +163,7 @@ COPY priv priv
COPY lib lib COPY lib lib
COPY test test COPY test test
COPY scripts/production-play-physical-fixture.exs scripts/production-play-physical-fixture.exs COPY scripts/production-play-physical-fixture.exs scripts/production-play-physical-fixture.exs
COPY scripts/production-android-fcm-smoke.exs scripts/production-android-fcm-smoke.exs
RUN mix compile RUN mix compile

View File

@ -1268,6 +1268,44 @@ Anonymous submissions receive the address-confirmation message before they
enter the operator queue. Operator email alerts are disabled unless enter the operator queue. Operator email alerts are disabled unless
`SUPPORT_OPERATOR_EMAIL_MODE=immediate` is explicitly configured. `SUPPORT_OPERATOR_EMAIL_MODE=immediate` is explicitly configured.
## Exact production Android FCM smoke
The Android FCM smoke is split into explicit phases so the operator can inspect
the notification on the exact physical device before cleanup. The plan is
read-only and verifies the production Compose identity, running image, healthy
application container, database name, and one exact active Android FCM device:
```bash
./scripts/production-android-fcm-smoke.sh \
plan <push-device-uuid> --check-only whoneedhelp.com
```
After the printed scope and exact device have been reviewed and separately
authorised, prepare one notification, inspect it on the phone, verify provider
delivery, and remove the run-owned records:
```bash
./scripts/production-android-fcm-smoke.sh \
prepare <push-device-uuid> --confirm whoneedhelp.com
./scripts/production-android-fcm-smoke.sh \
verify --from-state --confirm whoneedhelp.com
./scripts/production-android-fcm-smoke.sh \
cleanup --from-state --confirm whoneedhelp.com
```
`prepare` creates one notification and one exact FCM delivery job in a single
database transaction. Its mode-`0600` manifest is exclusive-created inside the
same transaction and contains a unique ownership token; a manifest failure
rolls the database transaction back, and failure cleanup never removes a file
that lacks that exact token. The smoke does not enqueue email, target Web Push,
change the frozen test deployment, update Caddy, or push Git.
The local state records the exact application container. If that container is
recreated between phases, the wrapper refuses to continue instead of silently
targeting a replacement. Inspect the saved state and production records before
performing an explicit recovery; do not delete notifications or Oban jobs by a
broad timestamp or worker query.
## Run-scoped production browser verification ## Run-scoped production browser verification
The production E2E workflow is opt-in and starts with a read-only identity and The production E2E workflow is opt-in and starts with a read-only identity and

View File

@ -164,6 +164,14 @@ ssh buyvm-maya \
The linger check must print `yes`; otherwise the enabled user timer can depend The linger check must print `yes`; otherwise the enabled user timer can depend
on an active login session and is not a persistent external monitor. on an active login session and is not a persistent external monitor.
A read-only check on 2026-08-25 returned `linger=no` for `simple` on the
selected external host. The timer was loaded, enabled, and active while the
user manager was running, and its latest completed check reported readiness,
metrics, and backup freshness as `up`; those facts do not prove that the user
manager survives logout or starts at boot. Persistent off-site monitoring
therefore remains open until an administrator enables lingering and the timer
is rechecked after a login-independent restart.
Do not mark backup ownership complete merely because this command passes. The Do not mark backup ownership complete merely because this command passes. The
operator must still store the Restic key independently and approve retention, operator must still store the Restic key independently and approve retention,
RPO, RTO, capacity, and responsible owners. RPO, RTO, capacity, and responsible owners.
@ -212,8 +220,10 @@ as forward-only rather than receiving an invented database rollback.
with the run prefix. Anonymous contact verification still requires a with the run prefix. Anonymous contact verification still requires a
controlled production mail smoke before this combined item can be controlled production mail smoke before this combined item can be
checked. checked.
- [x] Database, application, worker, email, push, backup, and edge monitoring - [ ] Database, application, worker, email, push, backup, and edge monitoring
are visible to the responsible operator. are visible to the responsible operator and persist independently of an
interactive login. Current checks prove visibility while the external
user manager is running, but that host still reports `linger=no`.
Record observed timestamps, revision/image identities, and non-secret evidence Record observed timestamps, revision/image identities, and non-secret evidence
paths in `docs/verification.md`. Record failed checks as failed; do not convert paths in `docs/verification.md`. Record failed checks as failed; do not convert

View File

@ -3427,3 +3427,36 @@ promoted.
image carrying run identifier `20260820232333-2842586`. The frozen hackathon image carrying run identifier `20260820232333-2842586`. The frozen hackathon
deployment, production deployment, Caddy, and public Git remote were not deployment, production deployment, Caddy, and public Git remote were not
changed by this local verification. changed by this local verification.
# 2026-08-25 manifest ownership and external-monitor final recheck
- The production Android FCM smoke and production browser E2E preparation now
exclusive-create their mode-`0600` manifests inside the same database
transaction as the run-owned records. Each manifest contains a UUID ownership
token. Failure cleanup removes a manifest only when that exact token matches,
so a concurrent process cannot cause one run to delete another run's state.
- The external-monitor installer now reads and validates the selected remote
account's actual home directory. Remote systemd, state, configuration, and
backup-heartbeat paths are derived from that observed home instead of assuming
`/home/simple`. The installer regression suite passed all six cases, including
a non-default `/home/monitor-user` account.
- A read-only external-host check observed account `simple` with home
`/home/simple` and `linger=no`. Its user timer was loaded, enabled, and active
while the user manager was running; the latest completed monitor check
reported readiness, authenticated metrics, and backup freshness as `up`.
This proves the current check completed, but it does not prove persistence
after logout or boot. No lingering setting was changed.
- The final isolated quality unit
`codex-heavy-wnh-final-quality-manifest-monitor-20260825-20260825-175845-502195.service`
exited successfully after 4 minutes 55.976 seconds. It reported 487 passing
ExUnit tests, 18 passing browser-asset tests, all configured shell/Python
regression suites, and all isolated quality and security gates passing. The
final Debian 13.6 production image and the pinned infrastructure images
reported zero detected vulnerabilities at the configured scan severity.
- The unit's measured memory peak was 438.9 MiB and swap use was zero. Exact
post-run inspection found no container, network, volume, or temporary image
from Compose project `wnh_quality_20260825145846502200` or run identifier
`20260825145846-502200`.
- This was local verification only. Production, the frozen hackathon test
deployment, shared Caddy, and the public Git remote were not changed or
pushed.

View File

@ -139,65 +139,72 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do
assert_prefix_unused!(context) assert_prefix_unused!(context)
password_hash = Bcrypt.hash_pwd_salt(password) if File.exists?(context.manifest_path) do
now = DateTime.utc_now(:second) Mix.raise("full staging E2E manifest already exists")
end
{:ok,
%{
users: users,
support_request: support_request,
removal_notice: removal_notice,
take_it_down_notice: take_it_down_notice
}} =
Repo.transaction(fn ->
users =
Map.new(@precreated_roles, fn role ->
user =
insert_user!(
context.emails[role],
display_name(role),
password_hash,
now
)
if role == "admin" do
%StaffRoleAssignment{}
|> StaffRoleAssignment.changeset(%{user_id: user.id, role: :admin})
|> Repo.insert!()
end
{role, %{"id" => user.id, "email" => user.email}}
end)
requester = users["requester"]
support_request = insert_support_fixture!(context, requester, now)
removal_notice = insert_removal_fixture!(context, requester, now)
take_it_down_notice = insert_take_it_down_fixture!(context, requester, now)
%{
users: users,
support_request: support_request,
removal_notice: removal_notice,
take_it_down_notice: take_it_down_notice
}
end)
manifest = %{
"schema_version" => 3,
"run_id" => context.run_id,
"database" => context.database,
"precreated_users" => users,
"precreated_records" => %{
"support_request" => support_request.id,
"content_removal_notice" => removal_notice.id,
"take_it_down_notice" => take_it_down_notice.id
},
"allowed_emails" => context.emails |> Map.values() |> Enum.sort()
}
context.manifest_path |> Path.dirname() |> File.mkdir_p!() context.manifest_path |> Path.dirname() |> File.mkdir_p!()
write_manifest!(context.manifest_path, manifest)
password_hash = Bcrypt.hash_pwd_salt(password)
now = DateTime.utc_now(:second)
manifest_ownership_token = Ecto.UUID.generate()
try do
{:ok, _fixture} =
Repo.transaction(fn ->
users =
Map.new(@precreated_roles, fn role ->
user =
insert_user!(
context.emails[role],
display_name(role),
password_hash,
now
)
if role == "admin" do
%StaffRoleAssignment{}
|> StaffRoleAssignment.changeset(%{user_id: user.id, role: :admin})
|> Repo.insert!()
end
{role, %{"id" => user.id, "email" => user.email}}
end)
requester = users["requester"]
support_request = insert_support_fixture!(context, requester, now)
removal_notice = insert_removal_fixture!(context, requester, now)
take_it_down_notice = insert_take_it_down_fixture!(context, requester, now)
manifest = %{
"schema_version" => 3,
"run_id" => context.run_id,
"database" => context.database,
"ownership_token" => manifest_ownership_token,
"precreated_users" => users,
"precreated_records" => %{
"support_request" => support_request.id,
"content_removal_notice" => removal_notice.id,
"take_it_down_notice" => take_it_down_notice.id
},
"allowed_emails" => context.emails |> Map.values() |> Enum.sort()
}
write_new_manifest!(context.manifest_path, manifest)
%{
users: users,
support_request: support_request,
removal_notice: removal_notice,
take_it_down_notice: take_it_down_notice
}
end)
rescue
exception ->
remove_owned_prepare_manifest(context, manifest_ownership_token)
reraise exception, __STACKTRACE__
end
Mix.shell().info("prepared exact full staging E2E users for #{context.run_id}") Mix.shell().info("prepared exact full staging E2E users for #{context.run_id}")
end end
@ -489,6 +496,7 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do
unless manifest["schema_version"] == 3 and manifest["run_id"] == context.run_id and unless manifest["schema_version"] == 3 and manifest["run_id"] == context.run_id and
manifest["database"] == context.database and manifest["database"] == context.database and
uuid?(manifest["ownership_token"]) and
manifest["allowed_emails"] == expected_emails and valid_precreated? and manifest["allowed_emails"] == expected_emails and valid_precreated? and
valid_precreated_records? do valid_precreated_records? do
Mix.raise("full staging E2E manifest does not match the requested run and database") Mix.raise("full staging E2E manifest does not match the requested run and database")
@ -770,6 +778,37 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do
File.rename!(temporary_path, path) File.rename!(temporary_path, path)
end end
defp write_new_manifest!(path, manifest) do
encoded = Jason.encode_to_iodata!(manifest, pretty: true)
File.open!(path, [:write, :binary, :exclusive], fn file ->
IO.binwrite(file, encoded)
end)
try do
File.chmod!(path, 0o600)
rescue
exception ->
File.rm(path)
reraise exception, __STACKTRACE__
end
end
defp remove_owned_prepare_manifest(context, ownership_token) do
expected_emails = context.emails |> Map.values() |> Enum.sort()
with {:ok, encoded} <- File.read(context.manifest_path),
{:ok, manifest} <- Jason.decode(encoded),
true <- manifest["run_id"] == context.run_id,
true <- manifest["database"] == context.database,
true <- manifest["allowed_emails"] == expected_emails,
true <- manifest["ownership_token"] == ownership_token do
File.rm(context.manifest_path)
else
_missing_or_different_manifest -> :ok
end
end
defp insert_user!(email, display_name, password_hash, now) do defp insert_user!(email, display_name, password_hash, now) do
%User{} %User{}
|> User.registration_changeset(%{ |> User.registration_changeset(%{

View File

@ -6,9 +6,9 @@ ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
source_target=${PRODUCTION_SSH_TARGET:-whoneedhelp} source_target=${PRODUCTION_SSH_TARGET:-whoneedhelp}
monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya} monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya}
production_env=${PRODUCTION_ENV_PATH:-/srv/who_need_help-production/.env} production_env=${PRODUCTION_ENV_PATH:-/srv/who_need_help-production/.env}
remote_root=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help} remote_root=${MONITOR_REMOTE_ROOT:-}
remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json} remote_config=${MONITOR_REMOTE_CONFIG:-}
remote_state=${MONITOR_REMOTE_STATE:-/home/simple/.local/state/who-need-help/monitor.json} remote_state=${MONITOR_REMOTE_STATE:-}
monitor_smtp_values_file=${MONITOR_SMTP_VALUES_FILE:-} monitor_smtp_values_file=${MONITOR_SMTP_VALUES_FILE:-}
monitor_install_work_root=${MONITOR_INSTALL_WORK_ROOT:-} monitor_install_work_root=${MONITOR_INSTALL_WORK_ROOT:-}
monitor_url=${MONITOR_URL:-https://whoneedhelp.com/healthz/ready} monitor_url=${MONITOR_URL:-https://whoneedhelp.com/healthz/ready}
@ -16,7 +16,7 @@ metrics_url=${MONITOR_METRICS_URL:-https://whoneedhelp.com/metrics}
monitor_calendar=${MONITOR_ON_CALENDAR:-'*:0/1'} monitor_calendar=${MONITOR_ON_CALENDAR:-'*:0/1'}
health_timeout=${MONITOR_HEALTH_TIMEOUT_SECONDS:-3} health_timeout=${MONITOR_HEALTH_TIMEOUT_SECONDS:-3}
metrics_timeout=${MONITOR_METRICS_TIMEOUT_SECONDS:-3} metrics_timeout=${MONITOR_METRICS_TIMEOUT_SECONDS:-3}
backup_heartbeat_path=${MONITOR_BACKUP_HEARTBEAT_PATH:-/home/simple/.local/state/who-need-help/production-backup.json} backup_heartbeat_path=${MONITOR_BACKUP_HEARTBEAT_PATH:-}
backup_max_age=${MONITOR_BACKUP_MAX_AGE_SECONDS:-} backup_max_age=${MONITOR_BACKUP_MAX_AGE_SECONDS:-}
for command in awk install mktemp python3 realpath scp ssh stat; do for command in awk install mktemp python3 realpath scp ssh stat; do
@ -77,31 +77,38 @@ case "$metrics_timeout" in
0) echo "MONITOR_METRICS_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;; 0) echo "MONITOR_METRICS_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;;
esac esac
if ! monitor_identity=$(ssh -o BatchMode=yes "$monitor_target" \
'printf "%s\n%s\n%s\n" "$(id -un)" "$HOME" "$(loginctl show-user "$(id -un)" --property=Linger --value)"'); then
echo "Unable to verify systemd user lingering on the external monitor host." >&2
exit 2
fi
monitor_user=$(printf '%s\n' "$monitor_identity" | awk 'NR == 1 {print; exit}')
monitor_home=$(printf '%s\n' "$monitor_identity" | awk 'NR == 2 {print; exit}')
monitor_linger=$(printf '%s\n' "$monitor_identity" | awk 'NR == 3 {print; exit}')
if [ -z "$monitor_user" ] || [ -z "$monitor_home" ] || [ "$monitor_linger" != yes ]; then
echo "External monitor installation requires systemd user lingering for '$monitor_user' on '$monitor_target'." >&2
echo "An administrator must run: sudo loginctl enable-linger '$monitor_user'" >&2
echo "Verify with: loginctl show-user '$monitor_user' --property=Linger --value" >&2
exit 2
fi
remote_root=${remote_root:-$monitor_home/.local/lib/who-need-help}
remote_config=${remote_config:-$monitor_home/.config/who-need-help/monitor.json}
remote_state=${remote_state:-$monitor_home/.local/state/who-need-help/monitor.json}
backup_heartbeat_path=${backup_heartbeat_path:-$monitor_home/.local/state/who-need-help/production-backup.json}
remote_systemd_dir=$monitor_home/.config/systemd/user
if [ -n "$backup_max_age" ]; then if [ -n "$backup_max_age" ]; then
case "$backup_max_age" in case "$backup_max_age" in
*[!0-9]*) echo "MONITOR_BACKUP_MAX_AGE_SECONDS must be a positive integer." >&2; exit 2 ;; *[!0-9]*) echo "MONITOR_BACKUP_MAX_AGE_SECONDS must be a positive integer." >&2; exit 2 ;;
0) echo "MONITOR_BACKUP_MAX_AGE_SECONDS must be a positive integer." >&2; exit 2 ;; 0) echo "MONITOR_BACKUP_MAX_AGE_SECONDS must be a positive integer." >&2; exit 2 ;;
esac esac
if [ "$backup_heartbeat_path" != /home/simple/.local/state/who-need-help/production-backup.json ]; then if [ "$backup_heartbeat_path" != "$monitor_home/.local/state/who-need-help/production-backup.json" ]; then
echo "MONITOR_BACKUP_HEARTBEAT_PATH is outside the reviewed monitor state path." >&2 echo "MONITOR_BACKUP_HEARTBEAT_PATH is outside the reviewed monitor state path." >&2
exit 2 exit 2
fi fi
fi fi
if ! monitor_identity=$(ssh -o BatchMode=yes "$monitor_target" \
'printf "%s:%s\n" "$(id -un)" "$(loginctl show-user "$(id -un)" --property=Linger --value)"'); then
echo "Unable to verify systemd user lingering on the external monitor host." >&2
exit 2
fi
monitor_user=${monitor_identity%%:*}
monitor_linger=${monitor_identity#*:}
if [ -z "$monitor_user" ] || [ "$monitor_linger" != yes ]; then
echo "External monitor installation requires systemd user lingering for '$monitor_user' on '$monitor_target'." >&2
echo "An administrator must run: sudo loginctl enable-linger '$monitor_user'" >&2
echo "Verify with: loginctl show-user '$monitor_user' --property=Linger --value" >&2
exit 2
fi
work_dir=$(mktemp -d "$monitor_install_work_root/monitor-install.XXXXXX") work_dir=$(mktemp -d "$monitor_install_work_root/monitor-install.XXXXXX")
config="$work_dir/monitor.json" config="$work_dir/monitor.json"
service="$work_dir/who-need-help-production-monitor.service" service="$work_dir/who-need-help-production-monitor.service"
@ -224,15 +231,15 @@ WantedBy=timers.target
EOF EOF
ssh -o BatchMode=yes "$monitor_target" \ ssh -o BatchMode=yes "$monitor_target" \
"install -d -m 700 '$remote_root' /home/simple/.config/who-need-help /home/simple/.local/state/who-need-help /home/simple/.config/systemd/user" "install -d -m 700 '$remote_root' '$monitor_home/.config/who-need-help' '$monitor_home/.local/state/who-need-help' '$remote_systemd_dir'"
scp -q "$ROOT/scripts/production-external-monitor.py" \ scp -q "$ROOT/scripts/production-external-monitor.py" \
"$monitor_target:$remote_root/production-external-monitor.py" "$monitor_target:$remote_root/production-external-monitor.py"
scp -q "$config" "$monitor_target:$remote_config" scp -q "$config" "$monitor_target:$remote_config"
scp -q "$service" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.service" scp -q "$service" "$monitor_target:$remote_systemd_dir/who-need-help-production-monitor.service"
scp -q "$timer" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.timer" scp -q "$timer" "$monitor_target:$remote_systemd_dir/who-need-help-production-monitor.timer"
ssh -o BatchMode=yes "$monitor_target" \ ssh -o BatchMode=yes "$monitor_target" \
"chmod 700 '$remote_root/production-external-monitor.py'; chmod 600 '$remote_config' /home/simple/.config/systemd/user/who-need-help-production-monitor.service /home/simple/.config/systemd/user/who-need-help-production-monitor.timer; systemctl --user daemon-reload; systemctl --user start who-need-help-production-monitor.service; systemctl --user enable --now who-need-help-production-monitor.timer" "chmod 700 '$remote_root/production-external-monitor.py'; chmod 600 '$remote_config' '$remote_systemd_dir/who-need-help-production-monitor.service' '$remote_systemd_dir/who-need-help-production-monitor.timer'; systemctl --user daemon-reload; systemctl --user start who-need-help-production-monitor.service; systemctl --user enable --now who-need-help-production-monitor.timer"
printf 'External monitor installed on %s (%s).\n' "$monitor_target" "$monitor_host" printf 'External monitor installed on %s (%s).\n' "$monitor_target" "$monitor_host"
printf 'Persistent systemd user manager verified for %s (linger=yes).\n' "$monitor_user" printf 'Persistent systemd user manager verified for %s (linger=yes).\n' "$monitor_user"

View File

@ -10,30 +10,21 @@ capture_postgres_readiness_evidence() {
wait_for_postgres_final_ready() { wait_for_postgres_final_ready() {
local container=$1 local container=$1
local evidence_file=$2 local evidence_file=$2
local max_attempts=$3 local interval_seconds=$3
local interval_seconds=$4
local attempt=0
local running local running
local health local health
while ! docker logs "$container" 2>&1 | while ! docker logs "$container" 2>&1 |
grep -Fq 'PostgreSQL init process complete; ready for start up.'; do grep -Fq 'PostgreSQL init process complete; ready for start up.'; do
attempt=$((attempt + 1))
running=$(docker inspect --format '{{.State.Running}}' "$container" 2>/dev/null || printf missing) running=$(docker inspect --format '{{.State.Running}}' "$container" 2>/dev/null || printf missing)
if [[ "$running" != true ]]; then if [[ "$running" != true ]]; then
capture_postgres_readiness_evidence "$container" "$evidence_file" capture_postgres_readiness_evidence "$container" "$evidence_file"
echo "Isolated restore PostgreSQL exited during initialization." >&2 echo "Isolated restore PostgreSQL exited during initialization." >&2
return 1 return 1
fi fi
if ((attempt >= max_attempts)); then
capture_postgres_readiness_evidence "$container" "$evidence_file"
echo "Isolated restore PostgreSQL did not finish initialization." >&2
return 1
fi
sleep "$interval_seconds" sleep "$interval_seconds"
done done
attempt=0
while true; do while true; do
# Docker can retain a cached healthy state from the temporary bootstrap # Docker can retain a cached healthy state from the temporary bootstrap
# postmaster while that server shuts down. Query the final server directly # postmaster while that server shuts down. Query the final server directly
@ -47,7 +38,6 @@ wait_for_postgres_final_ready() {
return 0 return 0
fi fi
attempt=$((attempt + 1))
running=$(docker inspect --format '{{.State.Running}}' "$container" 2>/dev/null || printf missing) running=$(docker inspect --format '{{.State.Running}}' "$container" 2>/dev/null || printf missing)
health=$(docker inspect \ health=$(docker inspect \
--format '{{if .State.Health}}{{.State.Health.Status}}{{else}}missing{{end}}' \ --format '{{if .State.Health}}{{.State.Health.Status}}{{else}}missing{{end}}' \
@ -58,7 +48,7 @@ wait_for_postgres_final_ready() {
echo "Isolated restore PostgreSQL exited before accepting SQL." >&2 echo "Isolated restore PostgreSQL exited before accepting SQL." >&2
return 1 return 1
fi fi
if [[ "$health" == unhealthy ]] || ((attempt >= max_attempts)); then if [[ "$health" == unhealthy ]]; then
capture_postgres_readiness_evidence "$container" "$evidence_file" capture_postgres_readiness_evidence "$container" "$evidence_file"
echo "Isolated restore PostgreSQL did not become SQL-ready." >&2 echo "Isolated restore PostgreSQL did not become SQL-ready." >&2
return 1 return 1

View File

@ -50,6 +50,8 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do
defp prepare(context) do defp prepare(context) do
if File.exists?(context.manifest_path), do: raise("manifest already exists") if File.exists?(context.manifest_path), do: raise("manifest already exists")
manifest_ownership_token = Ecto.UUID.generate()
device = Repo.get(PushDevice, context.device_id) device = Repo.get(PushDevice, context.device_id)
unless match?( unless match?(
@ -67,41 +69,49 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do
raise "run-scoped notification already exists" raise "run-scoped notification already exists"
end end
{:ok, fixture} = fixture =
Repo.transaction(fn -> try do
notification = {:ok, committed_fixture} =
%Notification{} Repo.transaction(fn ->
|> Notification.changeset(%{ notification =
user_id: device.user_id, %Notification{}
kind: :support_update, |> Notification.changeset(%{
title: "Who Need Help notification check", user_id: device.user_id,
body: "Android notifications are working.", kind: :support_update,
path: "/notifications", title: "Who Need Help notification check",
data: %{"run_id" => context.run_id, "synthetic" => true}, body: "Android notifications are working.",
idempotency_key: context.idempotency_key path: "/notifications",
}) data: %{"run_id" => context.run_id, "synthetic" => true},
|> Repo.insert!() idempotency_key: context.idempotency_key
})
|> Repo.insert!()
job = job =
%{"notification_id" => notification.id, "device_id" => device.id} %{"notification_id" => notification.id, "device_id" => device.id}
|> DeviceDeliveryWorker.new() |> DeviceDeliveryWorker.new()
|> Oban.insert!() |> Oban.insert!()
%{notification: notification, job: job} manifest = %{
end) "schema_version" => 1,
"run_id" => context.run_id,
"database" => context.database,
"ownership_token" => manifest_ownership_token,
"idempotency_key" => context.idempotency_key,
"notification_id" => notification.id,
"device_id" => device.id,
"job_id" => job.id
}
manifest = %{ persist_manifest!(context.manifest_path, manifest)
"schema_version" => 1, %{notification: notification, job: job}
"run_id" => context.run_id, end)
"database" => context.database,
"idempotency_key" => context.idempotency_key,
"notification_id" => fixture.notification.id,
"device_id" => device.id,
"job_id" => fixture.job.id
}
File.write!(context.manifest_path, Jason.encode_to_iodata!(manifest, pretty: true)) committed_fixture
File.chmod!(context.manifest_path, 0o600) rescue
exception ->
remove_owned_manifest(context, manifest_ownership_token)
reraise exception, __STACKTRACE__
end
IO.puts("android_fcm_smoke_prepared=true") IO.puts("android_fcm_smoke_prepared=true")
IO.puts("job_id=#{fixture.job.id}") IO.puts("job_id=#{fixture.job.id}")
@ -184,6 +194,7 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do
valid? = valid? =
manifest["schema_version"] == 1 and manifest["run_id"] == context.run_id and manifest["schema_version"] == 1 and manifest["run_id"] == context.run_id and
manifest["database"] == context.database and manifest["database"] == context.database and
uuid?(manifest["ownership_token"]) and
manifest["device_id"] == context.device_id and manifest["device_id"] == context.device_id and
manifest["idempotency_key"] == context.idempotency_key and manifest["idempotency_key"] == context.idempotency_key and
uuid?(manifest["notification_id"]) and is_integer(manifest["job_id"]) uuid?(manifest["notification_id"]) and is_integer(manifest["job_id"])
@ -192,6 +203,36 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do
manifest manifest
end end
defp persist_manifest!(path, manifest) do
encoded = Jason.encode_to_iodata!(manifest, pretty: true)
File.open!(path, [:write, :binary, :exclusive], fn file ->
IO.binwrite(file, encoded)
end)
try do
File.chmod!(path, 0o600)
rescue
exception ->
File.rm(path)
reraise exception, __STACKTRACE__
end
end
defp remove_owned_manifest(context, ownership_token) do
with {:ok, encoded} <- File.read(context.manifest_path),
{:ok, manifest} <- Jason.decode(encoded),
true <- manifest["run_id"] == context.run_id,
true <- manifest["database"] == context.database,
true <- manifest["device_id"] == context.device_id,
true <- manifest["idempotency_key"] == context.idempotency_key,
true <- manifest["ownership_token"] == ownership_token do
File.rm(context.manifest_path)
else
_missing_or_different_manifest -> :ok
end
end
defp required_option!(options, name) do defp required_option!(options, name) do
case Map.get(options, name) do case Map.get(options, name) do
value when is_binary(value) and value != "" -> value value when is_binary(value) and value != "" -> value

View File

@ -255,7 +255,6 @@ restore_snapshot() {
wait_for_postgres_final_ready \ wait_for_postgres_final_ready \
"$restore_container" \ "$restore_container" \
"$evidence_dir/restore-postgres.log" \ "$evidence_dir/restore-postgres.log" \
120 \
1 1
# The PostGIS image initializes its requested database with PostGIS already # The PostGIS image initializes its requested database with PostGIS already

View File

@ -18,6 +18,7 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
self.fake_bin = self.root / "bin" self.fake_bin = self.root / "bin"
self.fake_bin.mkdir() self.fake_bin.mkdir()
self.capture = self.root / "captured-monitor.json" self.capture = self.root / "captured-monitor.json"
self.scp_log = self.root / "scp.log"
self.work_root = self.root / "work" self.work_root = self.root / "work"
self.work_root.mkdir() self.work_root.mkdir()
self.production_env = self.root / "production.env" self.production_env = self.root / "production.env"
@ -85,7 +86,9 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
command = sys.argv[-1] command = sys.argv[-1]
if "loginctl show-user" in command: if "loginctl show-user" in command:
print(f"monitor-user:{os.environ.get('FAKE_MONITOR_LINGER', 'yes')}") print("monitor-user")
print(os.environ.get("FAKE_MONITOR_HOME", "/home/monitor-user"))
print(os.environ.get("FAKE_MONITOR_LINGER", "yes"))
raise SystemExit(0) raise SystemExit(0)
if command.startswith("python3 - "): if command.startswith("python3 - "):
result = subprocess.run( result = subprocess.run(
@ -112,6 +115,9 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
import sys import sys
from pathlib import Path from pathlib import Path
with open(os.environ["FAKE_SCP_LOG"], "a", encoding="utf-8") as handle:
handle.write(" ".join(sys.argv[1:]) + "\\n")
sources = [Path(value) for value in sys.argv[1:-1] if not value.startswith("-")] sources = [Path(value) for value in sys.argv[1:-1] if not value.startswith("-")]
for source in sources: for source in sources:
if source.name == "monitor.json": if source.name == "monitor.json":
@ -122,6 +128,7 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
def run_installer(self, *, override, backup_max_age=None, linger="yes"): def run_installer(self, *, override, backup_max_age=None, linger="yes"):
self.capture.unlink(missing_ok=True) self.capture.unlink(missing_ok=True)
self.scp_log.unlink(missing_ok=True)
env = os.environ.copy() env = os.environ.copy()
env.update( env.update(
{ {
@ -130,8 +137,10 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
"MONITOR_SSH_TARGET": "monitor", "MONITOR_SSH_TARGET": "monitor",
"PRODUCTION_ENV_PATH": str(self.production_env), "PRODUCTION_ENV_PATH": str(self.production_env),
"FAKE_MONITOR_CAPTURE": str(self.capture), "FAKE_MONITOR_CAPTURE": str(self.capture),
"FAKE_SCP_LOG": str(self.scp_log),
"MONITOR_INSTALL_WORK_ROOT": str(self.work_root), "MONITOR_INSTALL_WORK_ROOT": str(self.work_root),
"FAKE_MONITOR_LINGER": linger, "FAKE_MONITOR_LINGER": linger,
"FAKE_MONITOR_HOME": "/home/monitor-user",
} }
) )
if override: if override:
@ -182,7 +191,7 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
self.assertEqual( self.assertEqual(
captured["backup"], captured["backup"],
{ {
"heartbeat_path": "/home/simple/.local/state/who-need-help/production-backup.json", "heartbeat_path": "/home/monitor-user/.local/state/who-need-help/production-backup.json",
"max_age_seconds": 172800, "max_age_seconds": 172800,
}, },
) )
@ -211,6 +220,13 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
"Persistent systemd user manager verified for monitor-user (linger=yes).", "Persistent systemd user manager verified for monitor-user (linger=yes).",
result.stdout, result.stdout,
) )
scp_log = self.scp_log.read_text(encoding="utf-8")
self.assertIn(
"monitor:/home/monitor-user/.config/systemd/user/"
"who-need-help-production-monitor.service",
scp_log,
)
self.assertNotIn("/home/simple", scp_log)
if __name__ == "__main__": if __name__ == "__main__":

View File

@ -0,0 +1,107 @@
Code.require_file(Path.expand("../../scripts/production-android-fcm-smoke.exs", __DIR__))
defmodule WhoNeedHelp.ProductionAndroidFCMSmokeTest do
use WhoNeedHelp.DataCase, async: false
use Oban.Testing, repo: WhoNeedHelp.Repo
import Ecto.Query
import ExUnit.CaptureIO
import WhoNeedHelp.AccountsFixtures
alias WhoNeedHelp.Notifications
alias WhoNeedHelp.Notifications.Notification
alias WhoNeedHelp.Push.DeviceDeliveryWorker
alias WhoNeedHelp.Repo
setup do
user = user_fixture(display_name: "Production FCM smoke test user")
scope = user_scope_fixture(user)
{:ok, device} =
Notifications.register_device(scope, %{
"platform" => "android",
"provider" => "fcm",
"token" => "test-fcm-token-#{Ecto.UUID.generate()}",
"installation_id" => Ecto.UUID.generate(),
"device_label" => "Production FCM smoke test"
})
%Postgrex.Result{rows: [[database]]} =
Repo.query!("SELECT current_database()", [], log: false)
%{device: device, database: database}
end
test "manifest write failure rolls back the notification and exact delivery job", context do
run_id = run_id()
missing_parent = "/tmp/wnh-production-android-fcm-missing-#{run_id}"
manifest_path = Path.join(missing_parent, "manifest.json")
File.rm_rf!(missing_parent)
options = options(context, run_id, manifest_path)
worker = to_string(DeviceDeliveryWorker)
jobs_before = Repo.aggregate(from(job in Oban.Job, where: job.worker == ^worker), :count)
assert_raise File.Error, fn ->
WhoNeedHelp.ProductionAndroidFCMSmoke.run("prepare", options)
end
refute Repo.get_by(Notification,
idempotency_key: "production-android-fcm-smoke:#{run_id}"
)
assert Repo.aggregate(from(job in Oban.Job, where: job.worker == ^worker), :count) ==
jobs_before
refute File.exists?(manifest_path)
end
test "prepare and cleanup retain an exact manifest-bound scope", context do
run_id = run_id()
manifest_path = "/tmp/wnh-production-android-fcm-#{run_id}.json"
File.rm(manifest_path)
on_exit(fn -> File.rm(manifest_path) end)
options = options(context, run_id, manifest_path)
assert capture_io(fn ->
WhoNeedHelp.ProductionAndroidFCMSmoke.run("prepare", options)
end) =~ "android_fcm_smoke_prepared=true"
assert {:ok, stat} = File.stat(manifest_path)
assert stat.mode |> Bitwise.band(0o777) == 0o600
manifest = manifest_path |> File.read!() |> Jason.decode!()
assert {:ok, _ownership_token} = Ecto.UUID.cast(manifest["ownership_token"])
assert %Notification{id: notification_id} =
Repo.get_by(Notification,
idempotency_key: "production-android-fcm-smoke:#{run_id}"
)
assert notification_id == manifest["notification_id"]
assert context.device.id == manifest["device_id"]
assert %Oban.Job{id: job_id} = Repo.get(Oban.Job, manifest["job_id"])
assert job_id == manifest["job_id"]
assert capture_io(fn ->
WhoNeedHelp.ProductionAndroidFCMSmoke.run("cleanup", options)
end) =~ "android_fcm_smoke_cleanup_verified=true"
refute Repo.get(Notification, notification_id)
refute Repo.get(Oban.Job, job_id)
refute File.exists?(manifest_path)
end
defp options(context, run_id, manifest_path) do
%{
run_id: run_id,
expected_database: context.database,
device_id: context.device.id,
manifest_path: manifest_path
}
end
defp run_id do
"test-#{System.unique_integer([:positive, :monotonic])}-#{Ecto.UUID.generate()}"
end
end

View File

@ -48,7 +48,7 @@ sleep() {
evidence=$(mktemp) evidence=$(mktemp)
trap 'rm -f "$evidence"' EXIT trap 'rm -f "$evidence"' EXIT
wait_for_postgres_final_ready restore-probe "$evidence" 3 0 wait_for_postgres_final_ready restore-probe "$evidence" 0
if ((sql_attempts != 2)); then if ((sql_attempts != 2)); then
printf 'expected two direct SQL readiness attempts, observed %s\n' "$sql_attempts" >&2 printf 'expected two direct SQL readiness attempts, observed %s\n' "$sql_attempts" >&2
@ -64,7 +64,7 @@ mode=unhealthy
sql_attempts=0 sql_attempts=0
: >"$evidence" : >"$evidence"
if wait_for_postgres_final_ready restore-probe "$evidence" 3 0; then if wait_for_postgres_final_ready restore-probe "$evidence" 0; then
echo 'unhealthy final PostgreSQL unexpectedly passed readiness' >&2 echo 'unhealthy final PostgreSQL unexpectedly passed readiness' >&2
exit 1 exit 1
fi fi

View File

@ -2,6 +2,7 @@ defmodule WhoNeedHelp.StagingFullE2ECleanupTest do
use WhoNeedHelp.DataCase, async: false use WhoNeedHelp.DataCase, async: false
use Oban.Testing, repo: WhoNeedHelp.Repo use Oban.Testing, repo: WhoNeedHelp.Repo
import Ecto.Query
import ExUnit.CaptureIO import ExUnit.CaptureIO
alias Oban.Job alias Oban.Job
@ -9,6 +10,50 @@ defmodule WhoNeedHelp.StagingFullE2ECleanupTest do
alias WhoNeedHelp.Mail.{ContentRemovalEmailWorker, SupportConfirmationWorker} alias WhoNeedHelp.Mail.{ContentRemovalEmailWorker, SupportConfirmationWorker}
alias WhoNeedHelp.Push.NearbyMatchWorker alias WhoNeedHelp.Push.NearbyMatchWorker
test "prepare refuses a pre-existing manifest before creating fixture users" do
run_id = "existing-manifest-#{System.unique_integer([:positive])}"
manifest_path = "/output/wnh-staging-full-e2e-#{run_id}.json"
password = "staging-existing-#{System.unique_integer([:positive])}"
original_manifest = "operator-owned manifest\n"
%Postgrex.Result{rows: [[database]]} =
Repo.query!("SELECT current_database()", [], log: false)
env = %{
"WNH_STAGING_E2E_CONFIRM" => "public-staging-full-e2e",
"WNH_STAGING_E2E_EXPECTED_DATABASE" => database,
"WNH_STAGING_E2E_MANIFEST_PATH" => manifest_path,
"WNH_STAGING_E2E_PASSWORD" => password,
"WNH_STAGING_E2E_RUN_ID" => run_id
}
previous_env = Map.new(env, fn {name, _value} -> {name, System.get_env(name)} end)
Enum.each(env, fn {name, value} -> System.put_env(name, value) end)
File.mkdir_p!(Path.dirname(manifest_path))
File.write!(manifest_path, original_manifest)
on_exit(fn ->
File.rm(manifest_path)
Enum.each(previous_env, fn
{name, nil} -> System.delete_env(name)
{name, value} -> System.put_env(name, value)
end)
end)
assert_raise Mix.Error, "full staging E2E manifest already exists", fn ->
run_task("prepare")
end
assert File.read!(manifest_path) == original_manifest
refute Repo.exists?(
from(user in WhoNeedHelp.Accounts.User,
where: like(user.email, ^"wnh-staging-e2e-#{run_id}-%")
)
)
end
test "cleanup removes only run-owned support and legal mail jobs" do test "cleanup removes only run-owned support and legal mail jobs" do
run_id = "cleanup-#{System.unique_integer([:positive])}" run_id = "cleanup-#{System.unique_integer([:positive])}"
manifest_path = "/output/wnh-staging-full-e2e-#{run_id}.json" manifest_path = "/output/wnh-staging-full-e2e-#{run_id}.json"
@ -76,6 +121,7 @@ defmodule WhoNeedHelp.StagingFullE2ECleanupTest do
|> Repo.insert!() |> Repo.insert!()
assert manifest["schema_version"] == 3 assert manifest["schema_version"] == 3
assert {:ok, _ownership_token} = Ecto.UUID.cast(manifest["ownership_token"])
assert prepared_support_request.subject == "Production E2E support #{run_id}" assert prepared_support_request.subject == "Production E2E support #{run_id}"
assert prepared_support_request.contact_verified_at assert prepared_support_request.contact_verified_at
assert prepared_removal_notice.regime == :general assert prepared_removal_notice.regime == :general