Add verified offsite production recovery

This commit is contained in:
SimpleTest 2026-08-03 04:37:41 +03:00
parent c085620762
commit ead8a67d0f
9 changed files with 1103 additions and 1 deletions

View File

@ -564,6 +564,107 @@ failure paths on the observed workstation. A MinIO volume on that same
workstation is not an off-site backup and does not establish production RPO,
RTO, retention, capacity, key custody, object locking, or database HA.
## Encrypted production backup to an independent host
The production workflow streams an external PostgreSQL 18 custom-format dump
through an encrypted Restic repository on an SSH target that must resolve to a
host other than production. It then runs `restic check --read-data` and restores
the newest snapshot into a uniquely named, temporary PostGIS container created
from `template0`. The source production database is never a restore target.
Bootstrap the checksum-pinned Restic binary and create the ignored local
operations configuration:
```bash
./scripts/bootstrap-restic.sh
./scripts/init-production-operations.sh
```
The initializer creates only these ignored local files:
- `tmp/production-operations/backup.env`, mode `0600`;
- `tmp/production-operations/restic-password`, mode `0600`.
The generated configuration uses the reviewed `whoneedhelp` production SSH
target and the independent `buyvm-maya` SFTP target. Inspect the read-only plan
before repository initialization:
```bash
./scripts/production-offsite-backup.sh plan
```
Repository initialization is an external mutation and requires the exact
repository value printed by the plan:
```bash
WNH_OFFSITE_BACKUP_INIT_CONFIRM='sftp:buyvm-maya:backups/who_need_help-production' \
./scripts/production-offsite-backup.sh init
```
The routine operations are:
```bash
./scripts/production-offsite-backup.sh run
./scripts/production-offsite-backup.sh check
./scripts/production-offsite-backup.sh restore-drill
```
`run` creates a fresh source dump, copies it into the Restic input, publishes
one encrypted snapshot only after the producer succeeds, performs a full-data
repository check, and runs the isolated restore drill. It removes its exact
temporary source dump, local scratch directory, Docker container, and Docker
volume on success, failure, or interrupt. Non-secret catalogs, checksums, and
restore observations remain under ignored
`output/production-operations/<run-id>/`.
Install the daily user-systemd timer only after one manual `run` succeeds:
```bash
./scripts/install-production-backup-timer.sh
systemctl --user list-timers --all who-need-help-production-backup.timer
systemctl --user status who-need-help-production-backup.service --no-pager
journalctl --user -u who-need-help-production-backup.service --no-pager
```
The timer uses `Persistent=true`, so systemd runs a missed calendar invocation
after the workstation and user manager return. A failed service invokes the
external monitor's backup-failure notification. It does not retry, delete
snapshots, or infer a retention policy.
The Restic password is required for every restore. Before relying on this as
the only recovery copy, place it in an operator-controlled password manager or
offline recovery record. Never print it into a terminal log, commit it, or
copy it into the backup repository. There is intentionally no automatic
`forget`/`prune` policy yet: retention, RPO, RTO, capacity, and key custody are
operator decisions, and the script does not invent them.
## Independent public readiness monitor
The external monitor runs on an SSH target that must resolve to a host other
than production. It requests only the public readiness endpoint and accepts
only HTTP 200 with the exact JSON object `{"status":"ready"}`. State is kept
in a mode-`0600` JSON file. Email is sent once when the state changes to down
and once when it recovers; repeated down checks do not send repeated alerts.
Install it from the production SMTP configuration without printing the SMTP
credential:
```bash
./scripts/install-production-external-monitor.sh
ssh buyvm-maya \
'systemctl --user status who-need-help-production-monitor.timer --no-pager'
ssh buyvm-maya \
'journalctl --user -u who-need-help-production-monitor.service --no-pager'
```
The default schedule is once per minute with a three-second HTTP timeout. Those
defaults match the current application container health timeout and were
installed only after public readiness requests from the selected monitor host
were observed to finish below one second. They are observations of the current
path, not universal capacity or availability guarantees. Re-running the
installer refreshes the exact script, mode-`0600` configuration, and user
units, then performs one check before enabling the timer.
## Local external-service boundary drill
Run the OAuth, SMTP, and provider-neutral push protocol checks without public

View File

@ -93,6 +93,22 @@ backup_path=$(printf '%s\n' "$backup_output" | sed -n 's/^Backup: //p')
Local backups and local MinIO drills verify mechanics but are not proof of
off-site durability or a production recovery policy.
The implemented external workflow can be rehearsed without restoring over the
source database:
```bash
./scripts/bootstrap-restic.sh
./scripts/production-offsite-backup.sh plan
./scripts/production-offsite-backup.sh run
systemctl --user status who-need-help-production-backup.service --no-pager
ssh buyvm-maya \
'systemctl --user status who-need-help-production-monitor.timer --no-pager'
```
Do not mark backup ownership complete merely because this command passes. The
operator must still store the Restic key independently and approve retention,
RPO, RTO, capacity, and responsible owners.
## 5. Promote with an explicit rollback point
Use the release workflow in `docs/operations.md`; do not copy mutable source

View File

@ -1,8 +1,42 @@
# Who Need Help — implementation verification
Observed through 2026-07-28 in the local workspace. This report separates observed
Observed through 2026-08-03 in the local workspace. This report separates observed
results from product limits and unknown production properties.
## Independent production backup and monitoring on 2026-08-03
This verification changed only the independent backup repository, operator
user-systemd units, and independent monitor host. It did not deploy application
source, change production data, modify the hackathon test deployment, push Git,
or edit the Devpost submission.
- An encrypted Restic 0.19.1 repository was initialized at
`sftp:buyvm-maya:backups/who_need_help-production`. SSH resolution verified
that this target and the production target are different hosts. The Restic
password and workflow configuration remain only in ignored mode-`0600`
local files; password-manager/offline key custody is still an operator action.
- Two production snapshots exist. The installed user-systemd execution produced
snapshot `98d7d2e8e3e7c8dbf7908bbaceeb2acedf202a7e1dfd77c5398ba668a7981236`,
passed `restic check --read-data`, and restored into a temporary database
created from `template0`. The isolated restore observed 34 public application
tables, 24 applied schema migrations, and PostGIS 3.6.4. Non-secret evidence
is retained at
`output/production-operations/20260803T012727Z-1694894/`.
- That exact systemd service exited `0/SUCCESS` after 4 minutes 31.643 seconds
wall time, 4.115 seconds local CPU time, and a 78.2 MiB local memory peak.
These are observations of one run over the selected SFTP path, not a minimum
resource requirement or recovery objective. The scheduled source dump and
every run-scoped temporary restore container and volume were absent afterward.
- The daily backup timer is enabled with `Persistent=true`. Its failure unit
routes one alert through the independent monitor host. No retention deletion,
automatic pruning, or database rollback is configured.
- The independent BuyVM monitor performs the public readiness check once per
minute. The observed check returned HTTP 200 with the exact ready payload and
the service exited successfully. Its SMTP configuration and state are mode
`0600`. A one-time monitoring test message was accepted by Brevo; mailbox
receipt has not yet been independently observed, so end-to-end alert delivery
is not claimed complete.
## Production Android bundle replay on 2026-07-28
The release tooling and application inputs below are exact local commit

53
scripts/bootstrap-restic.sh Executable file
View File

@ -0,0 +1,53 @@
#!/usr/bin/env bash
set -euo pipefail
umask 077
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
target_dir="$ROOT/.tools/restic"
target="$target_dir/restic"
image=${WNH_BACKUP_TOOLS_IMAGE:-who-need-help:backup-tools}
expected='restic 0.19.1 compiled with go1.26.5'
if [[ -x "$target" ]] && [[ "$($target version)" == "$expected"* ]]; then
printf 'Pinned Restic is already available: %s\n' "$target"
exit 0
fi
for command in docker install mktemp; do
command -v "$command" >/dev/null 2>&1 || {
echo "Required command is unavailable: $command" >&2
exit 2
}
done
mkdir -p "$target_dir"
chmod 700 "$ROOT/.tools" "$target_dir"
docker build --tag "$image" --file "$ROOT/Dockerfile.backup" "$ROOT"
container=$(docker create "$image" restic version)
partial=$(mktemp "$target_dir/restic.XXXXXX")
cleanup() {
docker rm --force "$container" >/dev/null 2>&1 || true
rm -f "$partial"
}
trap cleanup EXIT HUP INT TERM
docker cp "$container:/usr/local/bin/restic" "$partial"
chmod 700 "$partial"
actual=$($partial version)
if [[ "$actual" != "$expected"* ]]; then
printf 'Unexpected Restic build: %s\n' "$actual" >&2
exit 1
fi
mv "$partial" "$target"
partial=
trap - EXIT HUP INT TERM
docker rm --force "$container" >/dev/null
printf 'Installed pinned Restic: %s\n' "$target"
printf 'Version: %s\n' "$actual"

View File

@ -0,0 +1,67 @@
#!/usr/bin/env bash
set -euo pipefail
umask 077
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
runtime_dir="$ROOT/tmp/production-operations"
config=${1:-"$runtime_dir/backup.env"}
if [[ "$config" != /* ]]; then
config="$ROOT/$config"
fi
password_file="$runtime_dir/restic-password"
for command in openssl ssh; do
command -v "$command" >/dev/null 2>&1 || {
echo "Required command is unavailable: $command" >&2
exit 2
}
done
if [[ -e "$config" || -e "$password_file" ]]; then
echo "Refusing to replace existing production operations configuration." >&2
printf 'Config: %s\nPassword file: %s\n' "$config" "$password_file" >&2
exit 2
fi
mkdir -p "$runtime_dir"
chmod 700 "$ROOT/tmp" "$runtime_dir"
source_target=whoneedhelp
repository_target=buyvm-maya
source_host=$(ssh -G "$source_target" | awk '$1 == "hostname" {print $2; exit}')
repository_host=$(ssh -G "$repository_target" | awk '$1 == "hostname" {print $2; exit}')
if [[ -z "$source_host" || -z "$repository_host" ]]; then
echo "Could not resolve both SSH targets." >&2
exit 2
fi
if [[ "$source_host" == "$repository_host" ]]; then
echo "The backup repository must not resolve to the production host." >&2
exit 2
fi
openssl rand -base64 48 | tr -d '\n' >"$password_file"
printf '\n' >>"$password_file"
chmod 600 "$password_file"
cat >"$config" <<EOF
PRODUCTION_SSH_TARGET=$source_target
PRODUCTION_REMOTE_ROOT=/srv/who_need_help-production
PRODUCTION_EXPECTED_ENVIRONMENT=production
OFFSITE_RESTIC_REPOSITORY=sftp:$repository_target:backups/who_need_help-production
OFFSITE_RESTIC_PASSWORD_FILE=$password_file
OFFSITE_RESTIC_HOST=who-need-help-production
OFFSITE_RESTIC_TAG=who-need-help-production
BACKUP_ON_CALENDAR=daily
EOF
chmod 600 "$config"
printf 'Created mode-0600 operations configuration: %s\n' "$config"
printf 'Created mode-0600 Restic key file: %s\n' "$password_file"
printf 'Production resolves to: %s\n' "$source_host"
printf 'Encrypted repository resolves to a different host: %s\n' "$repository_host"
echo "The Restic key is required for every restore. Copy it to an operator-controlled password manager before relying on this backup as the only recovery copy."

View File

@ -0,0 +1,79 @@
#!/usr/bin/env bash
set -euo pipefail
umask 077
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
config=${1:-"$ROOT/tmp/production-operations/backup.env"}
unit_dir=${SYSTEMD_USER_UNIT_DIR:-"$HOME/.config/systemd/user"}
if [[ "$config" != /* ]]; then
config="$ROOT/$config"
fi
if [[ ! -f "$config" || "$(stat -c '%a' "$config")" != 600 ]]; then
echo "The backup configuration must exist with mode 0600: $config" >&2
exit 2
fi
# shellcheck source=/dev/null
source "$config"
: "${BACKUP_ON_CALENDAR:?Set BACKUP_ON_CALENDAR}"
case "$BACKUP_ON_CALENDAR" in
*$'\n'* | *$'\r'*) echo "BACKUP_ON_CALENDAR must be one line." >&2; exit 2 ;;
esac
monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya}
remote_monitor=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help}/production-external-monitor.py
remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json}
ssh_path=$(command -v ssh)
mkdir -p "$unit_dir"
chmod 700 "$HOME/.config" "$HOME/.config/systemd" "$unit_dir"
cat >"$unit_dir/who-need-help-production-backup.service" <<EOF
[Unit]
Description=Who Need Help encrypted off-server production backup and restore drill
Wants=network-online.target
After=network-online.target
OnFailure=who-need-help-production-backup-alert.service
[Service]
Type=oneshot
WorkingDirectory=$ROOT
ExecStart=$ROOT/scripts/production-offsite-backup.sh run $config
EOF
cat >"$unit_dir/who-need-help-production-backup-alert.service" <<EOF
[Unit]
Description=Notify operators that the Who Need Help production backup failed
Wants=network-online.target
After=network-online.target
[Service]
Type=oneshot
ExecStart=$ssh_path -o BatchMode=yes $monitor_target /usr/bin/python3 $remote_monitor --config $remote_config notify-backup-failure --unit who-need-help-production-backup.service
EOF
cat >"$unit_dir/who-need-help-production-backup.timer" <<EOF
[Unit]
Description=Schedule Who Need Help encrypted production backups
[Timer]
OnCalendar=$BACKUP_ON_CALENDAR
Persistent=true
Unit=who-need-help-production-backup.service
[Install]
WantedBy=timers.target
EOF
chmod 600 \
"$unit_dir/who-need-help-production-backup.service" \
"$unit_dir/who-need-help-production-backup-alert.service" \
"$unit_dir/who-need-help-production-backup.timer"
systemctl --user daemon-reload
systemctl --user enable --now who-need-help-production-backup.timer
echo "Production backup timer installed without starting or changing the application."
systemctl --user list-timers --all --no-pager who-need-help-production-backup.timer

View File

@ -0,0 +1,137 @@
#!/usr/bin/env bash
set -euo pipefail
umask 077
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
source_target=${PRODUCTION_SSH_TARGET:-whoneedhelp}
monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya}
production_env=${PRODUCTION_ENV_PATH:-/srv/who_need_help-production/.env}
remote_root=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help}
remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json}
remote_state=${MONITOR_REMOTE_STATE:-/home/simple/.local/state/who-need-help/monitor.json}
monitor_url=${MONITOR_URL:-https://whoneedhelp.com/healthz/ready}
monitor_calendar=${MONITOR_ON_CALENDAR:-'*:0/1'}
health_timeout=${MONITOR_HEALTH_TIMEOUT_SECONDS:-3}
for command in mktemp scp ssh; do
command -v "$command" >/dev/null 2>&1 || {
echo "Required command is unavailable: $command" >&2
exit 2
}
done
source_host=$(ssh -G "$source_target" | awk '$1 == "hostname" {print $2; exit}')
monitor_host=$(ssh -G "$monitor_target" | awk '$1 == "hostname" {print $2; exit}')
if [[ -z "$source_host" || -z "$monitor_host" || "$source_host" == "$monitor_host" ]]; then
echo "The external monitor must resolve and run on a host other than production." >&2
exit 2
fi
case "$health_timeout" in
'' | *[!0-9]*) echo "MONITOR_HEALTH_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;;
0) echo "MONITOR_HEALTH_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;;
esac
work_dir=$(mktemp -d "$ROOT/tmp/production-operations/monitor-install.XXXXXX")
config="$work_dir/monitor.json"
service="$work_dir/who-need-help-production-monitor.service"
timer="$work_dir/who-need-help-production-monitor.timer"
cleanup() {
rm -rf "$work_dir"
}
trap cleanup EXIT HUP INT TERM
ssh -o BatchMode=yes "$source_target" \
"python3 - '$production_env' '$monitor_url' '$health_timeout'" >"$config" <<'PY'
import json
import shlex
import sys
path, health_url, timeout = sys.argv[1:]
wanted = {
"SMTP_RELAY",
"SMTP_PORT",
"SMTP_USERNAME",
"SMTP_PASSWORD",
"SMTP_TLS",
"SMTP_SSL",
"EMAIL_FROM_ADDRESS",
"EMAIL_FROM_NAME",
"SUPPORT_INBOX_ADDRESS",
}
values = {}
with open(path, encoding="utf-8") as handle:
for raw_line in handle:
line = raw_line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, value = line.split("=", 1)
if key not in wanted:
continue
parsed = shlex.split(value, comments=False, posix=True)
values[key] = parsed[0] if parsed else ""
missing = sorted(key for key in wanted if not values.get(key))
if missing:
raise SystemExit("Missing production mail settings: " + ", ".join(missing))
payload = {
"health_timeout_seconds": int(timeout),
"health_url": health_url,
"smtp": {
"from_address": values["EMAIL_FROM_ADDRESS"],
"from_name": values["EMAIL_FROM_NAME"],
"implicit_ssl": values["SMTP_SSL"].lower() == "true",
"password": values["SMTP_PASSWORD"],
"port": int(values["SMTP_PORT"]),
"recipient": values["SUPPORT_INBOX_ADDRESS"],
"relay": values["SMTP_RELAY"],
"starttls": values["SMTP_TLS"].lower() == "always",
"username": values["SMTP_USERNAME"],
},
}
json.dump(payload, sys.stdout, ensure_ascii=False, indent=2, sort_keys=True)
sys.stdout.write("\n")
PY
chmod 600 "$config"
cat >"$service" <<EOF
[Unit]
Description=Who Need Help independent production readiness monitor
Wants=network-online.target
After=network-online.target
[Service]
Type=oneshot
ExecStart=/usr/bin/python3 $remote_root/production-external-monitor.py --config $remote_config --state $remote_state check
EOF
cat >"$timer" <<EOF
[Unit]
Description=Run the Who Need Help independent production readiness monitor
[Timer]
OnCalendar=$monitor_calendar
Persistent=true
Unit=who-need-help-production-monitor.service
[Install]
WantedBy=timers.target
EOF
ssh -o BatchMode=yes "$monitor_target" \
"install -d -m 700 '$remote_root' /home/simple/.config/who-need-help /home/simple/.local/state/who-need-help /home/simple/.config/systemd/user"
scp -q "$ROOT/scripts/production-external-monitor.py" \
"$monitor_target:$remote_root/production-external-monitor.py"
scp -q "$config" "$monitor_target:$remote_config"
scp -q "$service" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.service"
scp -q "$timer" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.timer"
ssh -o BatchMode=yes "$monitor_target" \
"chmod 700 '$remote_root/production-external-monitor.py'; chmod 600 '$remote_config' /home/simple/.config/systemd/user/who-need-help-production-monitor.service /home/simple/.config/systemd/user/who-need-help-production-monitor.timer; systemctl --user daemon-reload; systemctl --user start who-need-help-production-monitor.service; systemctl --user enable --now who-need-help-production-monitor.timer"
printf 'External monitor installed on %s (%s).\n' "$monitor_target" "$monitor_host"
printf 'Health URL: %s\n' "$monitor_url"
printf 'Schedule: %s; request timeout: %ss.\n' "$monitor_calendar" "$health_timeout"
echo "The SMTP credential is stored only in a mode-0600 configuration on the external monitor host."

View File

@ -0,0 +1,231 @@
#!/usr/bin/env python3
"""Stateful external production health and operations notifications."""
from __future__ import annotations
import argparse
import json
import os
import smtplib
import ssl
import sys
import tempfile
import urllib.error
import urllib.request
from datetime import datetime, timezone
from email.message import EmailMessage
from pathlib import Path
from typing import Any
def utc_now() -> str:
return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z")
def read_json(path: Path) -> dict[str, Any]:
with path.open("r", encoding="utf-8") as handle:
value = json.load(handle)
if not isinstance(value, dict):
raise ValueError(f"Expected a JSON object in {path}")
return value
def write_json_atomic(path: Path, value: dict[str, Any]) -> None:
path.parent.mkdir(mode=0o700, parents=True, exist_ok=True)
descriptor, temporary_name = tempfile.mkstemp(dir=path.parent, prefix=f".{path.name}.")
temporary = Path(temporary_name)
try:
with os.fdopen(descriptor, "w", encoding="utf-8") as handle:
json.dump(value, handle, ensure_ascii=False, indent=2, sort_keys=True)
handle.write("\n")
temporary.chmod(0o600)
temporary.replace(path)
finally:
temporary.unlink(missing_ok=True)
def required_text(mapping: dict[str, Any], name: str) -> str:
value = mapping.get(name)
if not isinstance(value, str) or not value.strip():
raise ValueError(f"Missing non-empty configuration value: {name}")
return value.strip()
def send_message(config: dict[str, Any], subject: str, body: str) -> None:
smtp = config.get("smtp")
if not isinstance(smtp, dict):
raise ValueError("Missing SMTP configuration")
relay = required_text(smtp, "relay")
port = int(smtp.get("port"))
username = required_text(smtp, "username")
password = required_text(smtp, "password")
from_address = required_text(smtp, "from_address")
from_name = required_text(smtp, "from_name")
recipient = required_text(smtp, "recipient")
implicit_ssl = bool(smtp.get("implicit_ssl", False))
starttls = bool(smtp.get("starttls", True))
message = EmailMessage()
message["From"] = f"{from_name} <{from_address}>"
message["To"] = recipient
message["Subject"] = subject
message.set_content(body)
context = ssl.create_default_context()
if implicit_ssl:
client_context = smtplib.SMTP_SSL(relay, port, timeout=30, context=context)
else:
client_context = smtplib.SMTP(relay, port, timeout=30)
with client_context as client:
if not implicit_ssl:
client.ehlo()
if starttls:
client.starttls(context=context)
client.ehlo()
client.login(username, password)
client.send_message(message)
def check_health(config: dict[str, Any]) -> tuple[str, str]:
url = required_text(config, "health_url")
timeout = float(config.get("health_timeout_seconds"))
request = urllib.request.Request(
url,
headers={"User-Agent": "WhoNeedHelp-External-Monitor/1.0"},
)
try:
with urllib.request.urlopen(request, timeout=timeout) as response:
status = response.status
payload = response.read(4096)
parsed = json.loads(payload.decode("utf-8"))
if status == 200 and parsed == {"status": "ready"}:
return "up", f"HTTP {status}; ready payload matched"
return "down", f"HTTP {status}; unexpected readiness payload"
except (OSError, ValueError, urllib.error.URLError) as error:
return "down", f"{type(error).__name__}: {str(error)[:500]}"
def monitor(config_path: Path, state_path: Path) -> int:
config = read_json(config_path)
previous: dict[str, Any] = {}
if state_path.exists():
previous = read_json(state_path)
status, detail = check_health(config)
previous_status = previous.get("status")
if status != previous_status:
if status == "down":
send_message(
config,
"[Who Need Help] Production readiness is DOWN",
"\n".join(
[
"The independent production readiness check failed.",
f"URL: {required_text(config, 'health_url')}",
f"Observed at: {utc_now()}",
f"Result: {detail}",
"This alert is sent once per state transition.",
]
),
)
elif previous_status == "down":
send_message(
config,
"[Who Need Help] Production readiness recovered",
"\n".join(
[
"The independent production readiness check recovered.",
f"URL: {required_text(config, 'health_url')}",
f"Observed at: {utc_now()}",
f"Result: {detail}",
]
),
)
write_json_atomic(
state_path,
{
"checked_at": utc_now(),
"detail": detail,
"status": status,
},
)
print(json.dumps({"status": status, "detail": detail}, sort_keys=True))
return 0 if status == "up" else 1
def notify_backup_failure(config_path: Path, unit: str) -> int:
config = read_json(config_path)
send_message(
config,
"[Who Need Help] Production backup or restore verification failed",
"\n".join(
[
"The scheduled encrypted production backup did not finish successfully.",
f"Unit: {unit}",
f"Observed at: {utc_now()}",
"Inspect the local user-systemd journal and do not treat the newest snapshot as verified until a restore drill passes.",
]
),
)
print("Backup failure notification sent.")
return 0
def send_test_notification(config_path: Path) -> int:
config = read_json(config_path)
send_message(
config,
"[Who Need Help] Operations monitoring test",
"\n".join(
[
"This is a one-time delivery verification for the independent production monitor.",
f"Health URL: {required_text(config, 'health_url')}",
f"Sent at: {utc_now()}",
"No production incident was detected and no application data was changed.",
]
),
)
print("Operations monitoring test notification sent.")
return 0
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument(
"--config",
type=Path,
default=Path.home() / ".config/who-need-help/monitor.json",
)
parser.add_argument(
"--state",
type=Path,
default=Path.home() / ".local/state/who-need-help/monitor.json",
)
subparsers = parser.add_subparsers(dest="action", required=True)
subparsers.add_parser("check")
subparsers.add_parser("send-test-notification")
backup = subparsers.add_parser("notify-backup-failure")
backup.add_argument("--unit", required=True)
return parser.parse_args()
def main() -> int:
args = parse_args()
try:
if args.action == "check":
return monitor(args.config, args.state)
if args.action == "send-test-notification":
return send_test_notification(args.config)
return notify_backup_failure(args.config, args.unit)
except (OSError, ValueError, smtplib.SMTPException, json.JSONDecodeError) as error:
print(f"Operations monitor failed: {type(error).__name__}: {error}", file=sys.stderr)
return 1
if __name__ == "__main__":
raise SystemExit(main())

View File

@ -0,0 +1,384 @@
#!/usr/bin/env bash
set -euo pipefail
umask 077
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
action=${1:-plan}
config=${2:-"$ROOT/tmp/production-operations/backup.env"}
restic="$ROOT/.tools/restic/restic"
restore_image='postgis/postgis:18-3.6-alpine@sha256:05d68c7f0f19b9aa0bf7c4a2049b2e8b38b44a63116392b95726a4c913766cf6'
case "$action" in
plan | init | run | check | restore-drill) ;;
*)
echo "Usage: $0 [plan|init|run|check|restore-drill] [CONFIG]" >&2
exit 2
;;
esac
if [[ "$config" != /* ]]; then
config="$ROOT/$config"
fi
if [[ ! -f "$config" ]]; then
echo "Operations configuration does not exist: $config" >&2
exit 2
fi
if [[ "$(stat -c '%a' "$config")" != 600 ]]; then
echo "Operations configuration must have mode 0600: $config" >&2
exit 2
fi
# shellcheck source=/dev/null
source "$config"
: "${PRODUCTION_SSH_TARGET:?Set PRODUCTION_SSH_TARGET}"
: "${PRODUCTION_REMOTE_ROOT:?Set PRODUCTION_REMOTE_ROOT}"
: "${PRODUCTION_EXPECTED_ENVIRONMENT:?Set PRODUCTION_EXPECTED_ENVIRONMENT}"
: "${OFFSITE_RESTIC_REPOSITORY:?Set OFFSITE_RESTIC_REPOSITORY}"
: "${OFFSITE_RESTIC_PASSWORD_FILE:?Set OFFSITE_RESTIC_PASSWORD_FILE}"
: "${OFFSITE_RESTIC_HOST:?Set OFFSITE_RESTIC_HOST}"
: "${OFFSITE_RESTIC_TAG:?Set OFFSITE_RESTIC_TAG}"
if [[ "$PRODUCTION_EXPECTED_ENVIRONMENT" != production ]]; then
echo "The source environment must be exactly production." >&2
exit 2
fi
case "$PRODUCTION_REMOTE_ROOT" in
/srv/who_need_help-production) ;;
*)
echo "The production root is outside the reviewed deployment path." >&2
exit 2
;;
esac
case "$OFFSITE_RESTIC_REPOSITORY" in
sftp:*) ;;
*)
echo "This workflow currently accepts only an encrypted Restic SFTP repository." >&2
exit 2
;;
esac
if [[ ! -f "$OFFSITE_RESTIC_PASSWORD_FILE" ]] ||
[[ "$(stat -c '%a' "$OFFSITE_RESTIC_PASSWORD_FILE")" != 600 ]]; then
echo "The Restic password file must exist with mode 0600." >&2
exit 2
fi
if [[ ! -x "$restic" ]]; then
echo "Pinned Restic is unavailable. Run ./scripts/bootstrap-restic.sh first." >&2
exit 2
fi
for command in docker flock grep jq openssl pg_restore scp sha256sum ssh; do
command -v "$command" >/dev/null 2>&1 || {
echo "Required command is unavailable: $command" >&2
exit 2
}
done
export RESTIC_REPOSITORY="$OFFSITE_RESTIC_REPOSITORY"
export RESTIC_PASSWORD_FILE="$OFFSITE_RESTIC_PASSWORD_FILE"
export RESTIC_CACHE_DIR="$ROOT/tmp/production-operations/restic-cache"
mkdir -p "$RESTIC_CACHE_DIR"
chmod 700 "$ROOT/tmp" "$ROOT/tmp/production-operations" "$RESTIC_CACHE_DIR"
source_host=$(ssh -G "$PRODUCTION_SSH_TARGET" | awk '$1 == "hostname" {print $2; exit}')
repository_alias=${OFFSITE_RESTIC_REPOSITORY#sftp:}
repository_alias=${repository_alias%%:*}
repository_host=$(ssh -G "$repository_alias" | awk '$1 == "hostname" {print $2; exit}')
if [[ -z "$source_host" || -z "$repository_host" ]] ||
[[ "$source_host" == "$repository_host" ]]; then
echo "The repository host must resolve and differ from production." >&2
exit 2
fi
check_source() {
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
"bash -s -- '$PRODUCTION_REMOTE_ROOT/.env' --check-only '$PRODUCTION_EXPECTED_ENVIRONMENT'" \
<"$ROOT/scripts/backup-external-postgres.sh"
}
repo_initialized() {
"$restic" snapshots --json >/dev/null 2>&1
}
printf 'Production SSH target: %s (%s)\n' "$PRODUCTION_SSH_TARGET" "$source_host"
printf 'Production root: %s\n' "$PRODUCTION_REMOTE_ROOT"
printf 'Encrypted repository: %s (%s)\n' "$OFFSITE_RESTIC_REPOSITORY" "$repository_host"
printf 'Restic: %s\n' "$($restic version)"
if [[ "$action" == plan ]]; then
check_source
if repo_initialized; then
echo "Encrypted repository opened successfully."
else
echo "Encrypted repository is not initialized or cannot be opened."
fi
echo "Plan completed without changing production or the repository."
exit 0
fi
if [[ "$action" == init ]]; then
confirmation="$OFFSITE_RESTIC_REPOSITORY"
if [[ "${WNH_OFFSITE_BACKUP_INIT_CONFIRM:-}" != "$confirmation" ]]; then
echo "Repository initialization requires exact confirmation:" >&2
echo "WNH_OFFSITE_BACKUP_INIT_CONFIRM=$confirmation $0 init '$config'" >&2
exit 2
fi
check_source
if repo_initialized; then
echo "Encrypted repository is already initialized; no change was made."
exit 0
fi
"$restic" init
"$restic" check
echo "Encrypted off-server Restic repository initialized and opened successfully."
exit 0
fi
repo_initialized || {
echo "Encrypted repository cannot be opened. Initialize it first." >&2
exit 2
}
if [[ "$action" == check ]]; then
"$restic" check --read-data
echo "Encrypted repository full-data check passed."
exit 0
fi
if [[ "$action" == restore-drill ]]; then
snapshot_id=$($restic snapshots --json --host "$OFFSITE_RESTIC_HOST" \
--tag "$OFFSITE_RESTIC_TAG" --latest 1 | jq -r '.[0].short_id // empty')
if [[ -z "$snapshot_id" ]]; then
echo "No production backup snapshot is available for a restore drill." >&2
exit 1
fi
else
snapshot_id=
fi
lock_file="$ROOT/tmp/production-operations/backup.lock"
exec 9>"$lock_file"
flock --nonblock 9 || {
echo "Another production backup or restore drill is already running." >&2
exit 1
}
run_id=$(date -u +%Y%m%dT%H%M%SZ)-$$
work_dir=$(mktemp -d "$ROOT/tmp/production-operations/run-$run_id.XXXXXX")
evidence_dir="$ROOT/output/production-operations/$run_id"
remote_dump="$PRODUCTION_REMOTE_ROOT/output/backups/production/scheduled-$run_id.dump"
remote_created=false
restore_container="wnh-production-restore-$run_id"
restore_volume="wnh_production_restore_${run_id//[^a-zA-Z0-9]/_}"
restore_started=false
restore_volume_created=false
mkdir -p "$evidence_dir"
chmod 700 "$ROOT/output" "$ROOT/output/production-operations" "$evidence_dir" "$work_dir"
cleanup() {
trap - EXIT HUP INT TERM
if [[ "$restore_started" == true ]]; then
docker rm --force "$restore_container" >/dev/null 2>&1 || true
fi
if [[ "$restore_volume_created" == true ]]; then
docker volume rm "$restore_volume" >/dev/null 2>&1 || true
fi
if [[ "$remote_created" == true ]]; then
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
"rm -f -- '$remote_dump' '$remote_dump.sha256' '$remote_dump.metadata'" \
>/dev/null 2>&1 || true
fi
rm -rf "$work_dir"
}
trap cleanup EXIT HUP INT TERM
restore_snapshot() {
local selected_snapshot=$1
local dump="$work_dir/production.dump"
local password
local health
local tables
local migrations
local postgis
"$restic" dump "$selected_snapshot" /production.dump >"$dump"
[[ -s "$dump" ]] || {
echo "Restic restored an empty database dump." >&2
return 1
}
pg_restore --list "$dump" >"$evidence_dir/pg-restore-catalog.txt"
if ! docker image inspect "$restore_image" >/dev/null 2>&1; then
docker pull "$restore_image" >"$evidence_dir/postgis-pull.txt"
fi
password=$(openssl rand -base64 36 | tr -d '\n')
docker volume create "$restore_volume" >/dev/null
restore_volume_created=true
docker run --detach \
--name "$restore_container" \
--env POSTGRES_PASSWORD="$password" \
--env POSTGRES_DB=postgres \
--volume "$restore_volume:/var/lib/postgresql" \
--health-cmd='pg_isready --username postgres --dbname postgres' \
--health-interval=1s \
--health-timeout=2s \
--health-retries=120 \
"$restore_image" >/dev/null
restore_started=true
# A freshly initialized postgres/postgis container briefly accepts
# connections through its temporary bootstrap server. Wait until the image
# has completed that bootstrap and started the final server before restoring.
while ! docker logs "$restore_container" 2>&1 |
grep -Fq 'PostgreSQL init process complete; ready for start up.'; do
if [[ "$(docker inspect --format '{{.State.Running}}' "$restore_container")" != true ]]; then
docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true
echo "Isolated restore PostgreSQL exited during initialization." >&2
return 1
fi
sleep 1
done
while true; do
health=$(docker inspect --format '{{if .State.Health}}{{.State.Health.Status}}{{else}}missing{{end}}' "$restore_container")
case "$health" in
healthy) break ;;
unhealthy)
docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true
echo "Isolated restore PostgreSQL became unhealthy." >&2
return 1
;;
esac
sleep 1
done
# The PostGIS image initializes its requested database with PostGIS already
# installed. Restore into a database created from template0 instead, so the
# archive remains responsible for recreating its own extensions and schema.
docker exec "$restore_container" createdb \
--username postgres \
--template template0 \
restore_check
if ! docker exec --interactive "$restore_container" pg_restore \
--username postgres \
--dbname restore_check \
--exit-on-error \
--no-owner \
--no-privileges <"$dump"; then
docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true
docker inspect "$restore_container" >"$evidence_dir/restore-container.json" 2>&1 || true
echo "Isolated restore failed; container evidence was retained." >&2
return 1
fi
tables=$(docker exec "$restore_container" psql \
--username postgres --dbname restore_check --tuples-only --no-align \
--command="select count(*) from pg_tables where schemaname='public' and tablename <> 'spatial_ref_sys';")
migrations=$(docker exec "$restore_container" psql \
--username postgres --dbname restore_check --tuples-only --no-align \
--command='select count(*) from schema_migrations;')
postgis=$(docker exec "$restore_container" psql \
--username postgres --dbname restore_check --tuples-only --no-align \
--command='select PostGIS_Lib_Version();')
if [[ ! "$tables" =~ ^[1-9][0-9]*$ ]] || [[ ! "$migrations" =~ ^[1-9][0-9]*$ ]]; then
echo "The isolated restore does not contain the expected application schema." >&2
return 1
fi
{
printf 'snapshot_id=%s\n' "$selected_snapshot"
printf 'application_tables=%s\n' "$tables"
printf 'schema_migrations=%s\n' "$migrations"
printf 'postgis_version=%s\n' "$postgis"
printf 'verified_at=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
} >"$evidence_dir/restore-drill.txt"
chmod 600 "$evidence_dir/restore-drill.txt" "$evidence_dir/pg-restore-catalog.txt"
docker rm --force "$restore_container" >/dev/null
restore_started=false
docker volume rm "$restore_volume" >/dev/null
restore_volume_created=false
rm -f "$dump"
}
if [[ "$action" == restore-drill ]]; then
restore_snapshot "$snapshot_id"
echo "Isolated restore drill passed for snapshot: $snapshot_id"
printf 'Non-secret evidence: %s\n' "$evidence_dir"
exit 0
fi
check_source
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
"bash -s -- '$PRODUCTION_REMOTE_ROOT/.env' '$remote_dump' '$PRODUCTION_EXPECTED_ENVIRONMENT'" \
<"$ROOT/scripts/backup-external-postgres.sh"
remote_created=true
scp -p \
"$PRODUCTION_SSH_TARGET:$remote_dump" \
"$PRODUCTION_SSH_TARGET:$remote_dump.sha256" \
"$PRODUCTION_SSH_TARGET:$remote_dump.metadata" \
"$work_dir/"
dump="$work_dir/$(basename -- "$remote_dump")"
checksum="$dump.sha256"
metadata="$dump.metadata"
(
cd "$work_dir"
sha256sum --check "$(basename -- "$checksum")"
) >"$evidence_dir/source-checksum.txt"
pg_restore --list "$dump" >"$evidence_dir/source-catalog.txt"
cp "$checksum" "$metadata" "$evidence_dir/"
chmod 600 "$evidence_dir"/*
run_tag="$OFFSITE_RESTIC_TAG-$run_id"
backup_json="$evidence_dir/restic-backup.jsonl"
"$restic" backup \
--json \
--host "$OFFSITE_RESTIC_HOST" \
--tag "$OFFSITE_RESTIC_TAG" \
--tag "$run_tag" \
--stdin-filename production.dump \
--stdin-from-command -- cat "$dump" >"$backup_json"
snapshot_id=$(jq -r 'select(.message_type == "summary") | .snapshot_id // empty' "$backup_json" | tail -n 1)
if [[ -z "$snapshot_id" ]]; then
echo "Restic did not report a completed snapshot." >&2
exit 1
fi
snapshot_matches=$($restic snapshots --json --tag "$run_tag" | jq 'length')
if [[ "$snapshot_matches" != 1 ]]; then
echo "The run tag did not resolve to exactly one snapshot." >&2
exit 1
fi
"$restic" check --read-data >"$evidence_dir/restic-check.txt"
restore_snapshot "$snapshot_id"
printf '%s\n' "$snapshot_id" >"$evidence_dir/snapshot-id.txt"
chmod 600 "$evidence_dir"/*
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
"rm -f -- '$remote_dump' '$remote_dump.sha256' '$remote_dump.metadata'"
remote_created=false
echo "Encrypted off-server backup and isolated restore drill passed."
printf 'Snapshot: %s\n' "$snapshot_id"
printf 'Non-secret evidence: %s\n' "$evidence_dir"