Harden production verification ownership

This commit is contained in:
SimpleTest 2026-08-25 18:07:09 +03:00
parent 4f81c38717
commit 7094f21394
13 changed files with 457 additions and 130 deletions

View File

@ -163,6 +163,7 @@ COPY priv priv
COPY lib lib
COPY test test
COPY scripts/production-play-physical-fixture.exs scripts/production-play-physical-fixture.exs
COPY scripts/production-android-fcm-smoke.exs scripts/production-android-fcm-smoke.exs
RUN mix compile

View File

@ -1268,6 +1268,44 @@ Anonymous submissions receive the address-confirmation message before they
enter the operator queue. Operator email alerts are disabled unless
`SUPPORT_OPERATOR_EMAIL_MODE=immediate` is explicitly configured.
## Exact production Android FCM smoke
The Android FCM smoke is split into explicit phases so the operator can inspect
the notification on the exact physical device before cleanup. The plan is
read-only and verifies the production Compose identity, running image, healthy
application container, database name, and one exact active Android FCM device:
```bash
./scripts/production-android-fcm-smoke.sh \
plan <push-device-uuid> --check-only whoneedhelp.com
```
After the printed scope and exact device have been reviewed and separately
authorised, prepare one notification, inspect it on the phone, verify provider
delivery, and remove the run-owned records:
```bash
./scripts/production-android-fcm-smoke.sh \
prepare <push-device-uuid> --confirm whoneedhelp.com
./scripts/production-android-fcm-smoke.sh \
verify --from-state --confirm whoneedhelp.com
./scripts/production-android-fcm-smoke.sh \
cleanup --from-state --confirm whoneedhelp.com
```
`prepare` creates one notification and one exact FCM delivery job in a single
database transaction. Its mode-`0600` manifest is exclusive-created inside the
same transaction and contains a unique ownership token; a manifest failure
rolls the database transaction back, and failure cleanup never removes a file
that lacks that exact token. The smoke does not enqueue email, target Web Push,
change the frozen test deployment, update Caddy, or push Git.
The local state records the exact application container. If that container is
recreated between phases, the wrapper refuses to continue instead of silently
targeting a replacement. Inspect the saved state and production records before
performing an explicit recovery; do not delete notifications or Oban jobs by a
broad timestamp or worker query.
## Run-scoped production browser verification
The production E2E workflow is opt-in and starts with a read-only identity and

View File

@ -164,6 +164,14 @@ ssh buyvm-maya \
The linger check must print `yes`; otherwise the enabled user timer can depend
on an active login session and is not a persistent external monitor.
A read-only check on 2026-08-25 returned `linger=no` for `simple` on the
selected external host. The timer was loaded, enabled, and active while the
user manager was running, and its latest completed check reported readiness,
metrics, and backup freshness as `up`; those facts do not prove that the user
manager survives logout or starts at boot. Persistent off-site monitoring
therefore remains open until an administrator enables lingering and the timer
is rechecked after a login-independent restart.
Do not mark backup ownership complete merely because this command passes. The
operator must still store the Restic key independently and approve retention,
RPO, RTO, capacity, and responsible owners.
@ -212,8 +220,10 @@ as forward-only rather than receiving an invented database rollback.
with the run prefix. Anonymous contact verification still requires a
controlled production mail smoke before this combined item can be
checked.
- [x] Database, application, worker, email, push, backup, and edge monitoring
are visible to the responsible operator.
- [ ] Database, application, worker, email, push, backup, and edge monitoring
are visible to the responsible operator and persist independently of an
interactive login. Current checks prove visibility while the external
user manager is running, but that host still reports `linger=no`.
Record observed timestamps, revision/image identities, and non-secret evidence
paths in `docs/verification.md`. Record failed checks as failed; do not convert

View File

@ -3427,3 +3427,36 @@ promoted.
image carrying run identifier `20260820232333-2842586`. The frozen hackathon
deployment, production deployment, Caddy, and public Git remote were not
changed by this local verification.
# 2026-08-25 manifest ownership and external-monitor final recheck
- The production Android FCM smoke and production browser E2E preparation now
exclusive-create their mode-`0600` manifests inside the same database
transaction as the run-owned records. Each manifest contains a UUID ownership
token. Failure cleanup removes a manifest only when that exact token matches,
so a concurrent process cannot cause one run to delete another run's state.
- The external-monitor installer now reads and validates the selected remote
account's actual home directory. Remote systemd, state, configuration, and
backup-heartbeat paths are derived from that observed home instead of assuming
`/home/simple`. The installer regression suite passed all six cases, including
a non-default `/home/monitor-user` account.
- A read-only external-host check observed account `simple` with home
`/home/simple` and `linger=no`. Its user timer was loaded, enabled, and active
while the user manager was running; the latest completed monitor check
reported readiness, authenticated metrics, and backup freshness as `up`.
This proves the current check completed, but it does not prove persistence
after logout or boot. No lingering setting was changed.
- The final isolated quality unit
`codex-heavy-wnh-final-quality-manifest-monitor-20260825-20260825-175845-502195.service`
exited successfully after 4 minutes 55.976 seconds. It reported 487 passing
ExUnit tests, 18 passing browser-asset tests, all configured shell/Python
regression suites, and all isolated quality and security gates passing. The
final Debian 13.6 production image and the pinned infrastructure images
reported zero detected vulnerabilities at the configured scan severity.
- The unit's measured memory peak was 438.9 MiB and swap use was zero. Exact
post-run inspection found no container, network, volume, or temporary image
from Compose project `wnh_quality_20260825145846502200` or run identifier
`20260825145846-502200`.
- This was local verification only. Production, the frozen hackathon test
deployment, shared Caddy, and the public Git remote were not changed or
pushed.

View File

@ -139,16 +139,18 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do
assert_prefix_unused!(context)
if File.exists?(context.manifest_path) do
Mix.raise("full staging E2E manifest already exists")
end
context.manifest_path |> Path.dirname() |> File.mkdir_p!()
password_hash = Bcrypt.hash_pwd_salt(password)
now = DateTime.utc_now(:second)
manifest_ownership_token = Ecto.UUID.generate()
{:ok,
%{
users: users,
support_request: support_request,
removal_notice: removal_notice,
take_it_down_notice: take_it_down_notice
}} =
try do
{:ok, _fixture} =
Repo.transaction(fn ->
users =
Map.new(@precreated_roles, fn role ->
@ -175,18 +177,11 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do
removal_notice = insert_removal_fixture!(context, requester, now)
take_it_down_notice = insert_take_it_down_fixture!(context, requester, now)
%{
users: users,
support_request: support_request,
removal_notice: removal_notice,
take_it_down_notice: take_it_down_notice
}
end)
manifest = %{
"schema_version" => 3,
"run_id" => context.run_id,
"database" => context.database,
"ownership_token" => manifest_ownership_token,
"precreated_users" => users,
"precreated_records" => %{
"support_request" => support_request.id,
@ -196,8 +191,20 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do
"allowed_emails" => context.emails |> Map.values() |> Enum.sort()
}
context.manifest_path |> Path.dirname() |> File.mkdir_p!()
write_manifest!(context.manifest_path, manifest)
write_new_manifest!(context.manifest_path, manifest)
%{
users: users,
support_request: support_request,
removal_notice: removal_notice,
take_it_down_notice: take_it_down_notice
}
end)
rescue
exception ->
remove_owned_prepare_manifest(context, manifest_ownership_token)
reraise exception, __STACKTRACE__
end
Mix.shell().info("prepared exact full staging E2E users for #{context.run_id}")
end
@ -489,6 +496,7 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do
unless manifest["schema_version"] == 3 and manifest["run_id"] == context.run_id and
manifest["database"] == context.database and
uuid?(manifest["ownership_token"]) and
manifest["allowed_emails"] == expected_emails and valid_precreated? and
valid_precreated_records? do
Mix.raise("full staging E2E manifest does not match the requested run and database")
@ -770,6 +778,37 @@ defmodule Mix.Tasks.Wnh.StagingFullE2e do
File.rename!(temporary_path, path)
end
defp write_new_manifest!(path, manifest) do
encoded = Jason.encode_to_iodata!(manifest, pretty: true)
File.open!(path, [:write, :binary, :exclusive], fn file ->
IO.binwrite(file, encoded)
end)
try do
File.chmod!(path, 0o600)
rescue
exception ->
File.rm(path)
reraise exception, __STACKTRACE__
end
end
defp remove_owned_prepare_manifest(context, ownership_token) do
expected_emails = context.emails |> Map.values() |> Enum.sort()
with {:ok, encoded} <- File.read(context.manifest_path),
{:ok, manifest} <- Jason.decode(encoded),
true <- manifest["run_id"] == context.run_id,
true <- manifest["database"] == context.database,
true <- manifest["allowed_emails"] == expected_emails,
true <- manifest["ownership_token"] == ownership_token do
File.rm(context.manifest_path)
else
_missing_or_different_manifest -> :ok
end
end
defp insert_user!(email, display_name, password_hash, now) do
%User{}
|> User.registration_changeset(%{

View File

@ -6,9 +6,9 @@ ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
source_target=${PRODUCTION_SSH_TARGET:-whoneedhelp}
monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya}
production_env=${PRODUCTION_ENV_PATH:-/srv/who_need_help-production/.env}
remote_root=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help}
remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json}
remote_state=${MONITOR_REMOTE_STATE:-/home/simple/.local/state/who-need-help/monitor.json}
remote_root=${MONITOR_REMOTE_ROOT:-}
remote_config=${MONITOR_REMOTE_CONFIG:-}
remote_state=${MONITOR_REMOTE_STATE:-}
monitor_smtp_values_file=${MONITOR_SMTP_VALUES_FILE:-}
monitor_install_work_root=${MONITOR_INSTALL_WORK_ROOT:-}
monitor_url=${MONITOR_URL:-https://whoneedhelp.com/healthz/ready}
@ -16,7 +16,7 @@ metrics_url=${MONITOR_METRICS_URL:-https://whoneedhelp.com/metrics}
monitor_calendar=${MONITOR_ON_CALENDAR:-'*:0/1'}
health_timeout=${MONITOR_HEALTH_TIMEOUT_SECONDS:-3}
metrics_timeout=${MONITOR_METRICS_TIMEOUT_SECONDS:-3}
backup_heartbeat_path=${MONITOR_BACKUP_HEARTBEAT_PATH:-/home/simple/.local/state/who-need-help/production-backup.json}
backup_heartbeat_path=${MONITOR_BACKUP_HEARTBEAT_PATH:-}
backup_max_age=${MONITOR_BACKUP_MAX_AGE_SECONDS:-}
for command in awk install mktemp python3 realpath scp ssh stat; do
@ -77,31 +77,38 @@ case "$metrics_timeout" in
0) echo "MONITOR_METRICS_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;;
esac
if ! monitor_identity=$(ssh -o BatchMode=yes "$monitor_target" \
'printf "%s\n%s\n%s\n" "$(id -un)" "$HOME" "$(loginctl show-user "$(id -un)" --property=Linger --value)"'); then
echo "Unable to verify systemd user lingering on the external monitor host." >&2
exit 2
fi
monitor_user=$(printf '%s\n' "$monitor_identity" | awk 'NR == 1 {print; exit}')
monitor_home=$(printf '%s\n' "$monitor_identity" | awk 'NR == 2 {print; exit}')
monitor_linger=$(printf '%s\n' "$monitor_identity" | awk 'NR == 3 {print; exit}')
if [ -z "$monitor_user" ] || [ -z "$monitor_home" ] || [ "$monitor_linger" != yes ]; then
echo "External monitor installation requires systemd user lingering for '$monitor_user' on '$monitor_target'." >&2
echo "An administrator must run: sudo loginctl enable-linger '$monitor_user'" >&2
echo "Verify with: loginctl show-user '$monitor_user' --property=Linger --value" >&2
exit 2
fi
remote_root=${remote_root:-$monitor_home/.local/lib/who-need-help}
remote_config=${remote_config:-$monitor_home/.config/who-need-help/monitor.json}
remote_state=${remote_state:-$monitor_home/.local/state/who-need-help/monitor.json}
backup_heartbeat_path=${backup_heartbeat_path:-$monitor_home/.local/state/who-need-help/production-backup.json}
remote_systemd_dir=$monitor_home/.config/systemd/user
if [ -n "$backup_max_age" ]; then
case "$backup_max_age" in
*[!0-9]*) echo "MONITOR_BACKUP_MAX_AGE_SECONDS must be a positive integer." >&2; exit 2 ;;
0) echo "MONITOR_BACKUP_MAX_AGE_SECONDS must be a positive integer." >&2; exit 2 ;;
esac
if [ "$backup_heartbeat_path" != /home/simple/.local/state/who-need-help/production-backup.json ]; then
if [ "$backup_heartbeat_path" != "$monitor_home/.local/state/who-need-help/production-backup.json" ]; then
echo "MONITOR_BACKUP_HEARTBEAT_PATH is outside the reviewed monitor state path." >&2
exit 2
fi
fi
if ! monitor_identity=$(ssh -o BatchMode=yes "$monitor_target" \
'printf "%s:%s\n" "$(id -un)" "$(loginctl show-user "$(id -un)" --property=Linger --value)"'); then
echo "Unable to verify systemd user lingering on the external monitor host." >&2
exit 2
fi
monitor_user=${monitor_identity%%:*}
monitor_linger=${monitor_identity#*:}
if [ -z "$monitor_user" ] || [ "$monitor_linger" != yes ]; then
echo "External monitor installation requires systemd user lingering for '$monitor_user' on '$monitor_target'." >&2
echo "An administrator must run: sudo loginctl enable-linger '$monitor_user'" >&2
echo "Verify with: loginctl show-user '$monitor_user' --property=Linger --value" >&2
exit 2
fi
work_dir=$(mktemp -d "$monitor_install_work_root/monitor-install.XXXXXX")
config="$work_dir/monitor.json"
service="$work_dir/who-need-help-production-monitor.service"
@ -224,15 +231,15 @@ WantedBy=timers.target
EOF
ssh -o BatchMode=yes "$monitor_target" \
"install -d -m 700 '$remote_root' /home/simple/.config/who-need-help /home/simple/.local/state/who-need-help /home/simple/.config/systemd/user"
"install -d -m 700 '$remote_root' '$monitor_home/.config/who-need-help' '$monitor_home/.local/state/who-need-help' '$remote_systemd_dir'"
scp -q "$ROOT/scripts/production-external-monitor.py" \
"$monitor_target:$remote_root/production-external-monitor.py"
scp -q "$config" "$monitor_target:$remote_config"
scp -q "$service" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.service"
scp -q "$timer" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.timer"
scp -q "$service" "$monitor_target:$remote_systemd_dir/who-need-help-production-monitor.service"
scp -q "$timer" "$monitor_target:$remote_systemd_dir/who-need-help-production-monitor.timer"
ssh -o BatchMode=yes "$monitor_target" \
"chmod 700 '$remote_root/production-external-monitor.py'; chmod 600 '$remote_config' /home/simple/.config/systemd/user/who-need-help-production-monitor.service /home/simple/.config/systemd/user/who-need-help-production-monitor.timer; systemctl --user daemon-reload; systemctl --user start who-need-help-production-monitor.service; systemctl --user enable --now who-need-help-production-monitor.timer"
"chmod 700 '$remote_root/production-external-monitor.py'; chmod 600 '$remote_config' '$remote_systemd_dir/who-need-help-production-monitor.service' '$remote_systemd_dir/who-need-help-production-monitor.timer'; systemctl --user daemon-reload; systemctl --user start who-need-help-production-monitor.service; systemctl --user enable --now who-need-help-production-monitor.timer"
printf 'External monitor installed on %s (%s).\n' "$monitor_target" "$monitor_host"
printf 'Persistent systemd user manager verified for %s (linger=yes).\n' "$monitor_user"

View File

@ -10,30 +10,21 @@ capture_postgres_readiness_evidence() {
wait_for_postgres_final_ready() {
local container=$1
local evidence_file=$2
local max_attempts=$3
local interval_seconds=$4
local attempt=0
local interval_seconds=$3
local running
local health
while ! docker logs "$container" 2>&1 |
grep -Fq 'PostgreSQL init process complete; ready for start up.'; do
attempt=$((attempt + 1))
running=$(docker inspect --format '{{.State.Running}}' "$container" 2>/dev/null || printf missing)
if [[ "$running" != true ]]; then
capture_postgres_readiness_evidence "$container" "$evidence_file"
echo "Isolated restore PostgreSQL exited during initialization." >&2
return 1
fi
if ((attempt >= max_attempts)); then
capture_postgres_readiness_evidence "$container" "$evidence_file"
echo "Isolated restore PostgreSQL did not finish initialization." >&2
return 1
fi
sleep "$interval_seconds"
done
attempt=0
while true; do
# Docker can retain a cached healthy state from the temporary bootstrap
# postmaster while that server shuts down. Query the final server directly
@ -47,7 +38,6 @@ wait_for_postgres_final_ready() {
return 0
fi
attempt=$((attempt + 1))
running=$(docker inspect --format '{{.State.Running}}' "$container" 2>/dev/null || printf missing)
health=$(docker inspect \
--format '{{if .State.Health}}{{.State.Health.Status}}{{else}}missing{{end}}' \
@ -58,7 +48,7 @@ wait_for_postgres_final_ready() {
echo "Isolated restore PostgreSQL exited before accepting SQL." >&2
return 1
fi
if [[ "$health" == unhealthy ]] || ((attempt >= max_attempts)); then
if [[ "$health" == unhealthy ]]; then
capture_postgres_readiness_evidence "$container" "$evidence_file"
echo "Isolated restore PostgreSQL did not become SQL-ready." >&2
return 1

View File

@ -50,6 +50,8 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do
defp prepare(context) do
if File.exists?(context.manifest_path), do: raise("manifest already exists")
manifest_ownership_token = Ecto.UUID.generate()
device = Repo.get(PushDevice, context.device_id)
unless match?(
@ -67,7 +69,9 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do
raise "run-scoped notification already exists"
end
{:ok, fixture} =
fixture =
try do
{:ok, committed_fixture} =
Repo.transaction(fn ->
notification =
%Notification{}
@ -87,21 +91,27 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do
|> DeviceDeliveryWorker.new()
|> Oban.insert!()
%{notification: notification, job: job}
end)
manifest = %{
"schema_version" => 1,
"run_id" => context.run_id,
"database" => context.database,
"ownership_token" => manifest_ownership_token,
"idempotency_key" => context.idempotency_key,
"notification_id" => fixture.notification.id,
"notification_id" => notification.id,
"device_id" => device.id,
"job_id" => fixture.job.id
"job_id" => job.id
}
File.write!(context.manifest_path, Jason.encode_to_iodata!(manifest, pretty: true))
File.chmod!(context.manifest_path, 0o600)
persist_manifest!(context.manifest_path, manifest)
%{notification: notification, job: job}
end)
committed_fixture
rescue
exception ->
remove_owned_manifest(context, manifest_ownership_token)
reraise exception, __STACKTRACE__
end
IO.puts("android_fcm_smoke_prepared=true")
IO.puts("job_id=#{fixture.job.id}")
@ -184,6 +194,7 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do
valid? =
manifest["schema_version"] == 1 and manifest["run_id"] == context.run_id and
manifest["database"] == context.database and
uuid?(manifest["ownership_token"]) and
manifest["device_id"] == context.device_id and
manifest["idempotency_key"] == context.idempotency_key and
uuid?(manifest["notification_id"]) and is_integer(manifest["job_id"])
@ -192,6 +203,36 @@ defmodule WhoNeedHelp.ProductionAndroidFCMSmoke do
manifest
end
defp persist_manifest!(path, manifest) do
encoded = Jason.encode_to_iodata!(manifest, pretty: true)
File.open!(path, [:write, :binary, :exclusive], fn file ->
IO.binwrite(file, encoded)
end)
try do
File.chmod!(path, 0o600)
rescue
exception ->
File.rm(path)
reraise exception, __STACKTRACE__
end
end
defp remove_owned_manifest(context, ownership_token) do
with {:ok, encoded} <- File.read(context.manifest_path),
{:ok, manifest} <- Jason.decode(encoded),
true <- manifest["run_id"] == context.run_id,
true <- manifest["database"] == context.database,
true <- manifest["device_id"] == context.device_id,
true <- manifest["idempotency_key"] == context.idempotency_key,
true <- manifest["ownership_token"] == ownership_token do
File.rm(context.manifest_path)
else
_missing_or_different_manifest -> :ok
end
end
defp required_option!(options, name) do
case Map.get(options, name) do
value when is_binary(value) and value != "" -> value

View File

@ -255,7 +255,6 @@ restore_snapshot() {
wait_for_postgres_final_ready \
"$restore_container" \
"$evidence_dir/restore-postgres.log" \
120 \
1
# The PostGIS image initializes its requested database with PostGIS already

View File

@ -18,6 +18,7 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
self.fake_bin = self.root / "bin"
self.fake_bin.mkdir()
self.capture = self.root / "captured-monitor.json"
self.scp_log = self.root / "scp.log"
self.work_root = self.root / "work"
self.work_root.mkdir()
self.production_env = self.root / "production.env"
@ -85,7 +86,9 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
command = sys.argv[-1]
if "loginctl show-user" in command:
print(f"monitor-user:{os.environ.get('FAKE_MONITOR_LINGER', 'yes')}")
print("monitor-user")
print(os.environ.get("FAKE_MONITOR_HOME", "/home/monitor-user"))
print(os.environ.get("FAKE_MONITOR_LINGER", "yes"))
raise SystemExit(0)
if command.startswith("python3 - "):
result = subprocess.run(
@ -112,6 +115,9 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
import sys
from pathlib import Path
with open(os.environ["FAKE_SCP_LOG"], "a", encoding="utf-8") as handle:
handle.write(" ".join(sys.argv[1:]) + "\\n")
sources = [Path(value) for value in sys.argv[1:-1] if not value.startswith("-")]
for source in sources:
if source.name == "monitor.json":
@ -122,6 +128,7 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
def run_installer(self, *, override, backup_max_age=None, linger="yes"):
self.capture.unlink(missing_ok=True)
self.scp_log.unlink(missing_ok=True)
env = os.environ.copy()
env.update(
{
@ -130,8 +137,10 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
"MONITOR_SSH_TARGET": "monitor",
"PRODUCTION_ENV_PATH": str(self.production_env),
"FAKE_MONITOR_CAPTURE": str(self.capture),
"FAKE_SCP_LOG": str(self.scp_log),
"MONITOR_INSTALL_WORK_ROOT": str(self.work_root),
"FAKE_MONITOR_LINGER": linger,
"FAKE_MONITOR_HOME": "/home/monitor-user",
}
)
if override:
@ -182,7 +191,7 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
self.assertEqual(
captured["backup"],
{
"heartbeat_path": "/home/simple/.local/state/who-need-help/production-backup.json",
"heartbeat_path": "/home/monitor-user/.local/state/who-need-help/production-backup.json",
"max_age_seconds": 172800,
},
)
@ -211,6 +220,13 @@ class InstallProductionExternalMonitorTest(unittest.TestCase):
"Persistent systemd user manager verified for monitor-user (linger=yes).",
result.stdout,
)
scp_log = self.scp_log.read_text(encoding="utf-8")
self.assertIn(
"monitor:/home/monitor-user/.config/systemd/user/"
"who-need-help-production-monitor.service",
scp_log,
)
self.assertNotIn("/home/simple", scp_log)
if __name__ == "__main__":

View File

@ -0,0 +1,107 @@
Code.require_file(Path.expand("../../scripts/production-android-fcm-smoke.exs", __DIR__))
defmodule WhoNeedHelp.ProductionAndroidFCMSmokeTest do
use WhoNeedHelp.DataCase, async: false
use Oban.Testing, repo: WhoNeedHelp.Repo
import Ecto.Query
import ExUnit.CaptureIO
import WhoNeedHelp.AccountsFixtures
alias WhoNeedHelp.Notifications
alias WhoNeedHelp.Notifications.Notification
alias WhoNeedHelp.Push.DeviceDeliveryWorker
alias WhoNeedHelp.Repo
setup do
user = user_fixture(display_name: "Production FCM smoke test user")
scope = user_scope_fixture(user)
{:ok, device} =
Notifications.register_device(scope, %{
"platform" => "android",
"provider" => "fcm",
"token" => "test-fcm-token-#{Ecto.UUID.generate()}",
"installation_id" => Ecto.UUID.generate(),
"device_label" => "Production FCM smoke test"
})
%Postgrex.Result{rows: [[database]]} =
Repo.query!("SELECT current_database()", [], log: false)
%{device: device, database: database}
end
test "manifest write failure rolls back the notification and exact delivery job", context do
run_id = run_id()
missing_parent = "/tmp/wnh-production-android-fcm-missing-#{run_id}"
manifest_path = Path.join(missing_parent, "manifest.json")
File.rm_rf!(missing_parent)
options = options(context, run_id, manifest_path)
worker = to_string(DeviceDeliveryWorker)
jobs_before = Repo.aggregate(from(job in Oban.Job, where: job.worker == ^worker), :count)
assert_raise File.Error, fn ->
WhoNeedHelp.ProductionAndroidFCMSmoke.run("prepare", options)
end
refute Repo.get_by(Notification,
idempotency_key: "production-android-fcm-smoke:#{run_id}"
)
assert Repo.aggregate(from(job in Oban.Job, where: job.worker == ^worker), :count) ==
jobs_before
refute File.exists?(manifest_path)
end
test "prepare and cleanup retain an exact manifest-bound scope", context do
run_id = run_id()
manifest_path = "/tmp/wnh-production-android-fcm-#{run_id}.json"
File.rm(manifest_path)
on_exit(fn -> File.rm(manifest_path) end)
options = options(context, run_id, manifest_path)
assert capture_io(fn ->
WhoNeedHelp.ProductionAndroidFCMSmoke.run("prepare", options)
end) =~ "android_fcm_smoke_prepared=true"
assert {:ok, stat} = File.stat(manifest_path)
assert stat.mode |> Bitwise.band(0o777) == 0o600
manifest = manifest_path |> File.read!() |> Jason.decode!()
assert {:ok, _ownership_token} = Ecto.UUID.cast(manifest["ownership_token"])
assert %Notification{id: notification_id} =
Repo.get_by(Notification,
idempotency_key: "production-android-fcm-smoke:#{run_id}"
)
assert notification_id == manifest["notification_id"]
assert context.device.id == manifest["device_id"]
assert %Oban.Job{id: job_id} = Repo.get(Oban.Job, manifest["job_id"])
assert job_id == manifest["job_id"]
assert capture_io(fn ->
WhoNeedHelp.ProductionAndroidFCMSmoke.run("cleanup", options)
end) =~ "android_fcm_smoke_cleanup_verified=true"
refute Repo.get(Notification, notification_id)
refute Repo.get(Oban.Job, job_id)
refute File.exists?(manifest_path)
end
defp options(context, run_id, manifest_path) do
%{
run_id: run_id,
expected_database: context.database,
device_id: context.device.id,
manifest_path: manifest_path
}
end
defp run_id do
"test-#{System.unique_integer([:positive, :monotonic])}-#{Ecto.UUID.generate()}"
end
end

View File

@ -48,7 +48,7 @@ sleep() {
evidence=$(mktemp)
trap 'rm -f "$evidence"' EXIT
wait_for_postgres_final_ready restore-probe "$evidence" 3 0
wait_for_postgres_final_ready restore-probe "$evidence" 0
if ((sql_attempts != 2)); then
printf 'expected two direct SQL readiness attempts, observed %s\n' "$sql_attempts" >&2
@ -64,7 +64,7 @@ mode=unhealthy
sql_attempts=0
: >"$evidence"
if wait_for_postgres_final_ready restore-probe "$evidence" 3 0; then
if wait_for_postgres_final_ready restore-probe "$evidence" 0; then
echo 'unhealthy final PostgreSQL unexpectedly passed readiness' >&2
exit 1
fi

View File

@ -2,6 +2,7 @@ defmodule WhoNeedHelp.StagingFullE2ECleanupTest do
use WhoNeedHelp.DataCase, async: false
use Oban.Testing, repo: WhoNeedHelp.Repo
import Ecto.Query
import ExUnit.CaptureIO
alias Oban.Job
@ -9,6 +10,50 @@ defmodule WhoNeedHelp.StagingFullE2ECleanupTest do
alias WhoNeedHelp.Mail.{ContentRemovalEmailWorker, SupportConfirmationWorker}
alias WhoNeedHelp.Push.NearbyMatchWorker
test "prepare refuses a pre-existing manifest before creating fixture users" do
run_id = "existing-manifest-#{System.unique_integer([:positive])}"
manifest_path = "/output/wnh-staging-full-e2e-#{run_id}.json"
password = "staging-existing-#{System.unique_integer([:positive])}"
original_manifest = "operator-owned manifest\n"
%Postgrex.Result{rows: [[database]]} =
Repo.query!("SELECT current_database()", [], log: false)
env = %{
"WNH_STAGING_E2E_CONFIRM" => "public-staging-full-e2e",
"WNH_STAGING_E2E_EXPECTED_DATABASE" => database,
"WNH_STAGING_E2E_MANIFEST_PATH" => manifest_path,
"WNH_STAGING_E2E_PASSWORD" => password,
"WNH_STAGING_E2E_RUN_ID" => run_id
}
previous_env = Map.new(env, fn {name, _value} -> {name, System.get_env(name)} end)
Enum.each(env, fn {name, value} -> System.put_env(name, value) end)
File.mkdir_p!(Path.dirname(manifest_path))
File.write!(manifest_path, original_manifest)
on_exit(fn ->
File.rm(manifest_path)
Enum.each(previous_env, fn
{name, nil} -> System.delete_env(name)
{name, value} -> System.put_env(name, value)
end)
end)
assert_raise Mix.Error, "full staging E2E manifest already exists", fn ->
run_task("prepare")
end
assert File.read!(manifest_path) == original_manifest
refute Repo.exists?(
from(user in WhoNeedHelp.Accounts.User,
where: like(user.email, ^"wnh-staging-e2e-#{run_id}-%")
)
)
end
test "cleanup removes only run-owned support and legal mail jobs" do
run_id = "cleanup-#{System.unique_integer([:positive])}"
manifest_path = "/output/wnh-staging-full-e2e-#{run_id}.json"
@ -76,6 +121,7 @@ defmodule WhoNeedHelp.StagingFullE2ECleanupTest do
|> Repo.insert!()
assert manifest["schema_version"] == 3
assert {:ok, _ownership_token} = Ecto.UUID.cast(manifest["ownership_token"])
assert prepared_support_request.subject == "Production E2E support #{run_id}"
assert prepared_support_request.contact_verified_at
assert prepared_removal_notice.regime == :general