diff --git a/docs/operations.md b/docs/operations.md index e5a8fa5..bdd7dcf 100644 --- a/docs/operations.md +++ b/docs/operations.md @@ -564,6 +564,107 @@ failure paths on the observed workstation. A MinIO volume on that same workstation is not an off-site backup and does not establish production RPO, RTO, retention, capacity, key custody, object locking, or database HA. +## Encrypted production backup to an independent host + +The production workflow streams an external PostgreSQL 18 custom-format dump +through an encrypted Restic repository on an SSH target that must resolve to a +host other than production. It then runs `restic check --read-data` and restores +the newest snapshot into a uniquely named, temporary PostGIS container created +from `template0`. The source production database is never a restore target. + +Bootstrap the checksum-pinned Restic binary and create the ignored local +operations configuration: + +```bash +./scripts/bootstrap-restic.sh +./scripts/init-production-operations.sh +``` + +The initializer creates only these ignored local files: + +- `tmp/production-operations/backup.env`, mode `0600`; +- `tmp/production-operations/restic-password`, mode `0600`. + +The generated configuration uses the reviewed `whoneedhelp` production SSH +target and the independent `buyvm-maya` SFTP target. Inspect the read-only plan +before repository initialization: + +```bash +./scripts/production-offsite-backup.sh plan +``` + +Repository initialization is an external mutation and requires the exact +repository value printed by the plan: + +```bash +WNH_OFFSITE_BACKUP_INIT_CONFIRM='sftp:buyvm-maya:backups/who_need_help-production' \ + ./scripts/production-offsite-backup.sh init +``` + +The routine operations are: + +```bash +./scripts/production-offsite-backup.sh run +./scripts/production-offsite-backup.sh check +./scripts/production-offsite-backup.sh restore-drill +``` + +`run` creates a fresh source dump, copies it into the Restic input, publishes +one encrypted snapshot only after the producer succeeds, performs a full-data +repository check, and runs the isolated restore drill. It removes its exact +temporary source dump, local scratch directory, Docker container, and Docker +volume on success, failure, or interrupt. Non-secret catalogs, checksums, and +restore observations remain under ignored +`output/production-operations//`. + +Install the daily user-systemd timer only after one manual `run` succeeds: + +```bash +./scripts/install-production-backup-timer.sh +systemctl --user list-timers --all who-need-help-production-backup.timer +systemctl --user status who-need-help-production-backup.service --no-pager +journalctl --user -u who-need-help-production-backup.service --no-pager +``` + +The timer uses `Persistent=true`, so systemd runs a missed calendar invocation +after the workstation and user manager return. A failed service invokes the +external monitor's backup-failure notification. It does not retry, delete +snapshots, or infer a retention policy. + +The Restic password is required for every restore. Before relying on this as +the only recovery copy, place it in an operator-controlled password manager or +offline recovery record. Never print it into a terminal log, commit it, or +copy it into the backup repository. There is intentionally no automatic +`forget`/`prune` policy yet: retention, RPO, RTO, capacity, and key custody are +operator decisions, and the script does not invent them. + +## Independent public readiness monitor + +The external monitor runs on an SSH target that must resolve to a host other +than production. It requests only the public readiness endpoint and accepts +only HTTP 200 with the exact JSON object `{"status":"ready"}`. State is kept +in a mode-`0600` JSON file. Email is sent once when the state changes to down +and once when it recovers; repeated down checks do not send repeated alerts. + +Install it from the production SMTP configuration without printing the SMTP +credential: + +```bash +./scripts/install-production-external-monitor.sh +ssh buyvm-maya \ + 'systemctl --user status who-need-help-production-monitor.timer --no-pager' +ssh buyvm-maya \ + 'journalctl --user -u who-need-help-production-monitor.service --no-pager' +``` + +The default schedule is once per minute with a three-second HTTP timeout. Those +defaults match the current application container health timeout and were +installed only after public readiness requests from the selected monitor host +were observed to finish below one second. They are observations of the current +path, not universal capacity or availability guarantees. Re-running the +installer refreshes the exact script, mode-`0600` configuration, and user +units, then performs one check before enabling the timer. + ## Local external-service boundary drill Run the OAuth, SMTP, and provider-neutral push protocol checks without public diff --git a/docs/public-launch-checklist.md b/docs/public-launch-checklist.md index b0c082b..9f851d1 100644 --- a/docs/public-launch-checklist.md +++ b/docs/public-launch-checklist.md @@ -93,6 +93,22 @@ backup_path=$(printf '%s\n' "$backup_output" | sed -n 's/^Backup: //p') Local backups and local MinIO drills verify mechanics but are not proof of off-site durability or a production recovery policy. +The implemented external workflow can be rehearsed without restoring over the +source database: + +```bash +./scripts/bootstrap-restic.sh +./scripts/production-offsite-backup.sh plan +./scripts/production-offsite-backup.sh run +systemctl --user status who-need-help-production-backup.service --no-pager +ssh buyvm-maya \ + 'systemctl --user status who-need-help-production-monitor.timer --no-pager' +``` + +Do not mark backup ownership complete merely because this command passes. The +operator must still store the Restic key independently and approve retention, +RPO, RTO, capacity, and responsible owners. + ## 5. Promote with an explicit rollback point Use the release workflow in `docs/operations.md`; do not copy mutable source diff --git a/docs/verification.md b/docs/verification.md index 14d6bcf..7646c9c 100644 --- a/docs/verification.md +++ b/docs/verification.md @@ -1,8 +1,42 @@ # Who Need Help — implementation verification -Observed through 2026-07-28 in the local workspace. This report separates observed +Observed through 2026-08-03 in the local workspace. This report separates observed results from product limits and unknown production properties. +## Independent production backup and monitoring on 2026-08-03 + +This verification changed only the independent backup repository, operator +user-systemd units, and independent monitor host. It did not deploy application +source, change production data, modify the hackathon test deployment, push Git, +or edit the Devpost submission. + +- An encrypted Restic 0.19.1 repository was initialized at + `sftp:buyvm-maya:backups/who_need_help-production`. SSH resolution verified + that this target and the production target are different hosts. The Restic + password and workflow configuration remain only in ignored mode-`0600` + local files; password-manager/offline key custody is still an operator action. +- Two production snapshots exist. The installed user-systemd execution produced + snapshot `98d7d2e8e3e7c8dbf7908bbaceeb2acedf202a7e1dfd77c5398ba668a7981236`, + passed `restic check --read-data`, and restored into a temporary database + created from `template0`. The isolated restore observed 34 public application + tables, 24 applied schema migrations, and PostGIS 3.6.4. Non-secret evidence + is retained at + `output/production-operations/20260803T012727Z-1694894/`. +- That exact systemd service exited `0/SUCCESS` after 4 minutes 31.643 seconds + wall time, 4.115 seconds local CPU time, and a 78.2 MiB local memory peak. + These are observations of one run over the selected SFTP path, not a minimum + resource requirement or recovery objective. The scheduled source dump and + every run-scoped temporary restore container and volume were absent afterward. +- The daily backup timer is enabled with `Persistent=true`. Its failure unit + routes one alert through the independent monitor host. No retention deletion, + automatic pruning, or database rollback is configured. +- The independent BuyVM monitor performs the public readiness check once per + minute. The observed check returned HTTP 200 with the exact ready payload and + the service exited successfully. Its SMTP configuration and state are mode + `0600`. A one-time monitoring test message was accepted by Brevo; mailbox + receipt has not yet been independently observed, so end-to-end alert delivery + is not claimed complete. + ## Production Android bundle replay on 2026-07-28 The release tooling and application inputs below are exact local commit diff --git a/scripts/bootstrap-restic.sh b/scripts/bootstrap-restic.sh new file mode 100755 index 0000000..5edd868 --- /dev/null +++ b/scripts/bootstrap-restic.sh @@ -0,0 +1,53 @@ +#!/usr/bin/env bash +set -euo pipefail +umask 077 + +ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd) +target_dir="$ROOT/.tools/restic" +target="$target_dir/restic" +image=${WNH_BACKUP_TOOLS_IMAGE:-who-need-help:backup-tools} +expected='restic 0.19.1 compiled with go1.26.5' + +if [[ -x "$target" ]] && [[ "$($target version)" == "$expected"* ]]; then + printf 'Pinned Restic is already available: %s\n' "$target" + exit 0 +fi + +for command in docker install mktemp; do + command -v "$command" >/dev/null 2>&1 || { + echo "Required command is unavailable: $command" >&2 + exit 2 + } +done + +mkdir -p "$target_dir" +chmod 700 "$ROOT/.tools" "$target_dir" + +docker build --tag "$image" --file "$ROOT/Dockerfile.backup" "$ROOT" + +container=$(docker create "$image" restic version) +partial=$(mktemp "$target_dir/restic.XXXXXX") + +cleanup() { + docker rm --force "$container" >/dev/null 2>&1 || true + rm -f "$partial" +} +trap cleanup EXIT HUP INT TERM + +docker cp "$container:/usr/local/bin/restic" "$partial" +chmod 700 "$partial" + +actual=$($partial version) +if [[ "$actual" != "$expected"* ]]; then + printf 'Unexpected Restic build: %s\n' "$actual" >&2 + exit 1 +fi + +mv "$partial" "$target" +partial= + +trap - EXIT HUP INT TERM +docker rm --force "$container" >/dev/null + +printf 'Installed pinned Restic: %s\n' "$target" +printf 'Version: %s\n' "$actual" diff --git a/scripts/init-production-operations.sh b/scripts/init-production-operations.sh new file mode 100755 index 0000000..968f857 --- /dev/null +++ b/scripts/init-production-operations.sh @@ -0,0 +1,67 @@ +#!/usr/bin/env bash +set -euo pipefail +umask 077 + +ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd) +runtime_dir="$ROOT/tmp/production-operations" +config=${1:-"$runtime_dir/backup.env"} + +if [[ "$config" != /* ]]; then + config="$ROOT/$config" +fi + +password_file="$runtime_dir/restic-password" + +for command in openssl ssh; do + command -v "$command" >/dev/null 2>&1 || { + echo "Required command is unavailable: $command" >&2 + exit 2 + } +done + +if [[ -e "$config" || -e "$password_file" ]]; then + echo "Refusing to replace existing production operations configuration." >&2 + printf 'Config: %s\nPassword file: %s\n' "$config" "$password_file" >&2 + exit 2 +fi + +mkdir -p "$runtime_dir" +chmod 700 "$ROOT/tmp" "$runtime_dir" + +source_target=whoneedhelp +repository_target=buyvm-maya + +source_host=$(ssh -G "$source_target" | awk '$1 == "hostname" {print $2; exit}') +repository_host=$(ssh -G "$repository_target" | awk '$1 == "hostname" {print $2; exit}') + +if [[ -z "$source_host" || -z "$repository_host" ]]; then + echo "Could not resolve both SSH targets." >&2 + exit 2 +fi + +if [[ "$source_host" == "$repository_host" ]]; then + echo "The backup repository must not resolve to the production host." >&2 + exit 2 +fi + +openssl rand -base64 48 | tr -d '\n' >"$password_file" +printf '\n' >>"$password_file" +chmod 600 "$password_file" + +cat >"$config" <&2 + exit 2 +fi + +# shellcheck source=/dev/null +source "$config" +: "${BACKUP_ON_CALENDAR:?Set BACKUP_ON_CALENDAR}" + +case "$BACKUP_ON_CALENDAR" in + *$'\n'* | *$'\r'*) echo "BACKUP_ON_CALENDAR must be one line." >&2; exit 2 ;; +esac + +monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya} +remote_monitor=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help}/production-external-monitor.py +remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json} +ssh_path=$(command -v ssh) + +mkdir -p "$unit_dir" +chmod 700 "$HOME/.config" "$HOME/.config/systemd" "$unit_dir" + +cat >"$unit_dir/who-need-help-production-backup.service" <"$unit_dir/who-need-help-production-backup-alert.service" <"$unit_dir/who-need-help-production-backup.timer" </dev/null 2>&1 || { + echo "Required command is unavailable: $command" >&2 + exit 2 + } +done + +source_host=$(ssh -G "$source_target" | awk '$1 == "hostname" {print $2; exit}') +monitor_host=$(ssh -G "$monitor_target" | awk '$1 == "hostname" {print $2; exit}') +if [[ -z "$source_host" || -z "$monitor_host" || "$source_host" == "$monitor_host" ]]; then + echo "The external monitor must resolve and run on a host other than production." >&2 + exit 2 +fi + +case "$health_timeout" in + '' | *[!0-9]*) echo "MONITOR_HEALTH_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;; + 0) echo "MONITOR_HEALTH_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;; +esac + +work_dir=$(mktemp -d "$ROOT/tmp/production-operations/monitor-install.XXXXXX") +config="$work_dir/monitor.json" +service="$work_dir/who-need-help-production-monitor.service" +timer="$work_dir/who-need-help-production-monitor.timer" + +cleanup() { + rm -rf "$work_dir" +} +trap cleanup EXIT HUP INT TERM + +ssh -o BatchMode=yes "$source_target" \ + "python3 - '$production_env' '$monitor_url' '$health_timeout'" >"$config" <<'PY' +import json +import shlex +import sys + +path, health_url, timeout = sys.argv[1:] +wanted = { + "SMTP_RELAY", + "SMTP_PORT", + "SMTP_USERNAME", + "SMTP_PASSWORD", + "SMTP_TLS", + "SMTP_SSL", + "EMAIL_FROM_ADDRESS", + "EMAIL_FROM_NAME", + "SUPPORT_INBOX_ADDRESS", +} +values = {} +with open(path, encoding="utf-8") as handle: + for raw_line in handle: + line = raw_line.strip() + if not line or line.startswith("#") or "=" not in line: + continue + key, value = line.split("=", 1) + if key not in wanted: + continue + parsed = shlex.split(value, comments=False, posix=True) + values[key] = parsed[0] if parsed else "" + +missing = sorted(key for key in wanted if not values.get(key)) +if missing: + raise SystemExit("Missing production mail settings: " + ", ".join(missing)) + +payload = { + "health_timeout_seconds": int(timeout), + "health_url": health_url, + "smtp": { + "from_address": values["EMAIL_FROM_ADDRESS"], + "from_name": values["EMAIL_FROM_NAME"], + "implicit_ssl": values["SMTP_SSL"].lower() == "true", + "password": values["SMTP_PASSWORD"], + "port": int(values["SMTP_PORT"]), + "recipient": values["SUPPORT_INBOX_ADDRESS"], + "relay": values["SMTP_RELAY"], + "starttls": values["SMTP_TLS"].lower() == "always", + "username": values["SMTP_USERNAME"], + }, +} +json.dump(payload, sys.stdout, ensure_ascii=False, indent=2, sort_keys=True) +sys.stdout.write("\n") +PY +chmod 600 "$config" + +cat >"$service" <"$timer" < str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def read_json(path: Path) -> dict[str, Any]: + with path.open("r", encoding="utf-8") as handle: + value = json.load(handle) + if not isinstance(value, dict): + raise ValueError(f"Expected a JSON object in {path}") + return value + + +def write_json_atomic(path: Path, value: dict[str, Any]) -> None: + path.parent.mkdir(mode=0o700, parents=True, exist_ok=True) + descriptor, temporary_name = tempfile.mkstemp(dir=path.parent, prefix=f".{path.name}.") + temporary = Path(temporary_name) + try: + with os.fdopen(descriptor, "w", encoding="utf-8") as handle: + json.dump(value, handle, ensure_ascii=False, indent=2, sort_keys=True) + handle.write("\n") + temporary.chmod(0o600) + temporary.replace(path) + finally: + temporary.unlink(missing_ok=True) + + +def required_text(mapping: dict[str, Any], name: str) -> str: + value = mapping.get(name) + if not isinstance(value, str) or not value.strip(): + raise ValueError(f"Missing non-empty configuration value: {name}") + return value.strip() + + +def send_message(config: dict[str, Any], subject: str, body: str) -> None: + smtp = config.get("smtp") + if not isinstance(smtp, dict): + raise ValueError("Missing SMTP configuration") + + relay = required_text(smtp, "relay") + port = int(smtp.get("port")) + username = required_text(smtp, "username") + password = required_text(smtp, "password") + from_address = required_text(smtp, "from_address") + from_name = required_text(smtp, "from_name") + recipient = required_text(smtp, "recipient") + implicit_ssl = bool(smtp.get("implicit_ssl", False)) + starttls = bool(smtp.get("starttls", True)) + + message = EmailMessage() + message["From"] = f"{from_name} <{from_address}>" + message["To"] = recipient + message["Subject"] = subject + message.set_content(body) + + context = ssl.create_default_context() + if implicit_ssl: + client_context = smtplib.SMTP_SSL(relay, port, timeout=30, context=context) + else: + client_context = smtplib.SMTP(relay, port, timeout=30) + + with client_context as client: + if not implicit_ssl: + client.ehlo() + if starttls: + client.starttls(context=context) + client.ehlo() + client.login(username, password) + client.send_message(message) + + +def check_health(config: dict[str, Any]) -> tuple[str, str]: + url = required_text(config, "health_url") + timeout = float(config.get("health_timeout_seconds")) + request = urllib.request.Request( + url, + headers={"User-Agent": "WhoNeedHelp-External-Monitor/1.0"}, + ) + + try: + with urllib.request.urlopen(request, timeout=timeout) as response: + status = response.status + payload = response.read(4096) + parsed = json.loads(payload.decode("utf-8")) + if status == 200 and parsed == {"status": "ready"}: + return "up", f"HTTP {status}; ready payload matched" + return "down", f"HTTP {status}; unexpected readiness payload" + except (OSError, ValueError, urllib.error.URLError) as error: + return "down", f"{type(error).__name__}: {str(error)[:500]}" + + +def monitor(config_path: Path, state_path: Path) -> int: + config = read_json(config_path) + previous: dict[str, Any] = {} + if state_path.exists(): + previous = read_json(state_path) + + status, detail = check_health(config) + previous_status = previous.get("status") + + if status != previous_status: + if status == "down": + send_message( + config, + "[Who Need Help] Production readiness is DOWN", + "\n".join( + [ + "The independent production readiness check failed.", + f"URL: {required_text(config, 'health_url')}", + f"Observed at: {utc_now()}", + f"Result: {detail}", + "This alert is sent once per state transition.", + ] + ), + ) + elif previous_status == "down": + send_message( + config, + "[Who Need Help] Production readiness recovered", + "\n".join( + [ + "The independent production readiness check recovered.", + f"URL: {required_text(config, 'health_url')}", + f"Observed at: {utc_now()}", + f"Result: {detail}", + ] + ), + ) + + write_json_atomic( + state_path, + { + "checked_at": utc_now(), + "detail": detail, + "status": status, + }, + ) + print(json.dumps({"status": status, "detail": detail}, sort_keys=True)) + return 0 if status == "up" else 1 + + +def notify_backup_failure(config_path: Path, unit: str) -> int: + config = read_json(config_path) + send_message( + config, + "[Who Need Help] Production backup or restore verification failed", + "\n".join( + [ + "The scheduled encrypted production backup did not finish successfully.", + f"Unit: {unit}", + f"Observed at: {utc_now()}", + "Inspect the local user-systemd journal and do not treat the newest snapshot as verified until a restore drill passes.", + ] + ), + ) + print("Backup failure notification sent.") + return 0 + + +def send_test_notification(config_path: Path) -> int: + config = read_json(config_path) + send_message( + config, + "[Who Need Help] Operations monitoring test", + "\n".join( + [ + "This is a one-time delivery verification for the independent production monitor.", + f"Health URL: {required_text(config, 'health_url')}", + f"Sent at: {utc_now()}", + "No production incident was detected and no application data was changed.", + ] + ), + ) + print("Operations monitoring test notification sent.") + return 0 + + +def parse_args() -> argparse.Namespace: + parser = argparse.ArgumentParser() + parser.add_argument( + "--config", + type=Path, + default=Path.home() / ".config/who-need-help/monitor.json", + ) + parser.add_argument( + "--state", + type=Path, + default=Path.home() / ".local/state/who-need-help/monitor.json", + ) + subparsers = parser.add_subparsers(dest="action", required=True) + subparsers.add_parser("check") + subparsers.add_parser("send-test-notification") + backup = subparsers.add_parser("notify-backup-failure") + backup.add_argument("--unit", required=True) + return parser.parse_args() + + +def main() -> int: + args = parse_args() + try: + if args.action == "check": + return monitor(args.config, args.state) + if args.action == "send-test-notification": + return send_test_notification(args.config) + return notify_backup_failure(args.config, args.unit) + except (OSError, ValueError, smtplib.SMTPException, json.JSONDecodeError) as error: + print(f"Operations monitor failed: {type(error).__name__}: {error}", file=sys.stderr) + return 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/scripts/production-offsite-backup.sh b/scripts/production-offsite-backup.sh new file mode 100755 index 0000000..894a849 --- /dev/null +++ b/scripts/production-offsite-backup.sh @@ -0,0 +1,384 @@ +#!/usr/bin/env bash +set -euo pipefail +umask 077 + +ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd) +action=${1:-plan} +config=${2:-"$ROOT/tmp/production-operations/backup.env"} +restic="$ROOT/.tools/restic/restic" +restore_image='postgis/postgis:18-3.6-alpine@sha256:05d68c7f0f19b9aa0bf7c4a2049b2e8b38b44a63116392b95726a4c913766cf6' + +case "$action" in + plan | init | run | check | restore-drill) ;; + *) + echo "Usage: $0 [plan|init|run|check|restore-drill] [CONFIG]" >&2 + exit 2 + ;; +esac + +if [[ "$config" != /* ]]; then + config="$ROOT/$config" +fi + +if [[ ! -f "$config" ]]; then + echo "Operations configuration does not exist: $config" >&2 + exit 2 +fi + +if [[ "$(stat -c '%a' "$config")" != 600 ]]; then + echo "Operations configuration must have mode 0600: $config" >&2 + exit 2 +fi + +# shellcheck source=/dev/null +source "$config" + +: "${PRODUCTION_SSH_TARGET:?Set PRODUCTION_SSH_TARGET}" +: "${PRODUCTION_REMOTE_ROOT:?Set PRODUCTION_REMOTE_ROOT}" +: "${PRODUCTION_EXPECTED_ENVIRONMENT:?Set PRODUCTION_EXPECTED_ENVIRONMENT}" +: "${OFFSITE_RESTIC_REPOSITORY:?Set OFFSITE_RESTIC_REPOSITORY}" +: "${OFFSITE_RESTIC_PASSWORD_FILE:?Set OFFSITE_RESTIC_PASSWORD_FILE}" +: "${OFFSITE_RESTIC_HOST:?Set OFFSITE_RESTIC_HOST}" +: "${OFFSITE_RESTIC_TAG:?Set OFFSITE_RESTIC_TAG}" + +if [[ "$PRODUCTION_EXPECTED_ENVIRONMENT" != production ]]; then + echo "The source environment must be exactly production." >&2 + exit 2 +fi + +case "$PRODUCTION_REMOTE_ROOT" in + /srv/who_need_help-production) ;; + *) + echo "The production root is outside the reviewed deployment path." >&2 + exit 2 + ;; +esac + +case "$OFFSITE_RESTIC_REPOSITORY" in + sftp:*) ;; + *) + echo "This workflow currently accepts only an encrypted Restic SFTP repository." >&2 + exit 2 + ;; +esac + +if [[ ! -f "$OFFSITE_RESTIC_PASSWORD_FILE" ]] || + [[ "$(stat -c '%a' "$OFFSITE_RESTIC_PASSWORD_FILE")" != 600 ]]; then + echo "The Restic password file must exist with mode 0600." >&2 + exit 2 +fi + +if [[ ! -x "$restic" ]]; then + echo "Pinned Restic is unavailable. Run ./scripts/bootstrap-restic.sh first." >&2 + exit 2 +fi + +for command in docker flock grep jq openssl pg_restore scp sha256sum ssh; do + command -v "$command" >/dev/null 2>&1 || { + echo "Required command is unavailable: $command" >&2 + exit 2 + } +done + +export RESTIC_REPOSITORY="$OFFSITE_RESTIC_REPOSITORY" +export RESTIC_PASSWORD_FILE="$OFFSITE_RESTIC_PASSWORD_FILE" +export RESTIC_CACHE_DIR="$ROOT/tmp/production-operations/restic-cache" + +mkdir -p "$RESTIC_CACHE_DIR" +chmod 700 "$ROOT/tmp" "$ROOT/tmp/production-operations" "$RESTIC_CACHE_DIR" + +source_host=$(ssh -G "$PRODUCTION_SSH_TARGET" | awk '$1 == "hostname" {print $2; exit}') +repository_alias=${OFFSITE_RESTIC_REPOSITORY#sftp:} +repository_alias=${repository_alias%%:*} +repository_host=$(ssh -G "$repository_alias" | awk '$1 == "hostname" {print $2; exit}') + +if [[ -z "$source_host" || -z "$repository_host" ]] || + [[ "$source_host" == "$repository_host" ]]; then + echo "The repository host must resolve and differ from production." >&2 + exit 2 +fi + +check_source() { + ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \ + "bash -s -- '$PRODUCTION_REMOTE_ROOT/.env' --check-only '$PRODUCTION_EXPECTED_ENVIRONMENT'" \ + <"$ROOT/scripts/backup-external-postgres.sh" +} + +repo_initialized() { + "$restic" snapshots --json >/dev/null 2>&1 +} + +printf 'Production SSH target: %s (%s)\n' "$PRODUCTION_SSH_TARGET" "$source_host" +printf 'Production root: %s\n' "$PRODUCTION_REMOTE_ROOT" +printf 'Encrypted repository: %s (%s)\n' "$OFFSITE_RESTIC_REPOSITORY" "$repository_host" +printf 'Restic: %s\n' "$($restic version)" + +if [[ "$action" == plan ]]; then + check_source + if repo_initialized; then + echo "Encrypted repository opened successfully." + else + echo "Encrypted repository is not initialized or cannot be opened." + fi + echo "Plan completed without changing production or the repository." + exit 0 +fi + +if [[ "$action" == init ]]; then + confirmation="$OFFSITE_RESTIC_REPOSITORY" + if [[ "${WNH_OFFSITE_BACKUP_INIT_CONFIRM:-}" != "$confirmation" ]]; then + echo "Repository initialization requires exact confirmation:" >&2 + echo "WNH_OFFSITE_BACKUP_INIT_CONFIRM=$confirmation $0 init '$config'" >&2 + exit 2 + fi + + check_source + if repo_initialized; then + echo "Encrypted repository is already initialized; no change was made." + exit 0 + fi + + "$restic" init + "$restic" check + echo "Encrypted off-server Restic repository initialized and opened successfully." + exit 0 +fi + +repo_initialized || { + echo "Encrypted repository cannot be opened. Initialize it first." >&2 + exit 2 +} + +if [[ "$action" == check ]]; then + "$restic" check --read-data + echo "Encrypted repository full-data check passed." + exit 0 +fi + +if [[ "$action" == restore-drill ]]; then + snapshot_id=$($restic snapshots --json --host "$OFFSITE_RESTIC_HOST" \ + --tag "$OFFSITE_RESTIC_TAG" --latest 1 | jq -r '.[0].short_id // empty') + if [[ -z "$snapshot_id" ]]; then + echo "No production backup snapshot is available for a restore drill." >&2 + exit 1 + fi +else + snapshot_id= +fi + +lock_file="$ROOT/tmp/production-operations/backup.lock" +exec 9>"$lock_file" +flock --nonblock 9 || { + echo "Another production backup or restore drill is already running." >&2 + exit 1 +} + +run_id=$(date -u +%Y%m%dT%H%M%SZ)-$$ +work_dir=$(mktemp -d "$ROOT/tmp/production-operations/run-$run_id.XXXXXX") +evidence_dir="$ROOT/output/production-operations/$run_id" +remote_dump="$PRODUCTION_REMOTE_ROOT/output/backups/production/scheduled-$run_id.dump" +remote_created=false +restore_container="wnh-production-restore-$run_id" +restore_volume="wnh_production_restore_${run_id//[^a-zA-Z0-9]/_}" +restore_started=false +restore_volume_created=false + +mkdir -p "$evidence_dir" +chmod 700 "$ROOT/output" "$ROOT/output/production-operations" "$evidence_dir" "$work_dir" + +cleanup() { + trap - EXIT HUP INT TERM + if [[ "$restore_started" == true ]]; then + docker rm --force "$restore_container" >/dev/null 2>&1 || true + fi + if [[ "$restore_volume_created" == true ]]; then + docker volume rm "$restore_volume" >/dev/null 2>&1 || true + fi + if [[ "$remote_created" == true ]]; then + ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \ + "rm -f -- '$remote_dump' '$remote_dump.sha256' '$remote_dump.metadata'" \ + >/dev/null 2>&1 || true + fi + rm -rf "$work_dir" +} +trap cleanup EXIT HUP INT TERM + +restore_snapshot() { + local selected_snapshot=$1 + local dump="$work_dir/production.dump" + local password + local health + local tables + local migrations + local postgis + + "$restic" dump "$selected_snapshot" /production.dump >"$dump" + [[ -s "$dump" ]] || { + echo "Restic restored an empty database dump." >&2 + return 1 + } + pg_restore --list "$dump" >"$evidence_dir/pg-restore-catalog.txt" + + if ! docker image inspect "$restore_image" >/dev/null 2>&1; then + docker pull "$restore_image" >"$evidence_dir/postgis-pull.txt" + fi + + password=$(openssl rand -base64 36 | tr -d '\n') + docker volume create "$restore_volume" >/dev/null + restore_volume_created=true + docker run --detach \ + --name "$restore_container" \ + --env POSTGRES_PASSWORD="$password" \ + --env POSTGRES_DB=postgres \ + --volume "$restore_volume:/var/lib/postgresql" \ + --health-cmd='pg_isready --username postgres --dbname postgres' \ + --health-interval=1s \ + --health-timeout=2s \ + --health-retries=120 \ + "$restore_image" >/dev/null + restore_started=true + + # A freshly initialized postgres/postgis container briefly accepts + # connections through its temporary bootstrap server. Wait until the image + # has completed that bootstrap and started the final server before restoring. + while ! docker logs "$restore_container" 2>&1 | + grep -Fq 'PostgreSQL init process complete; ready for start up.'; do + if [[ "$(docker inspect --format '{{.State.Running}}' "$restore_container")" != true ]]; then + docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true + echo "Isolated restore PostgreSQL exited during initialization." >&2 + return 1 + fi + sleep 1 + done + + while true; do + health=$(docker inspect --format '{{if .State.Health}}{{.State.Health.Status}}{{else}}missing{{end}}' "$restore_container") + case "$health" in + healthy) break ;; + unhealthy) + docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true + echo "Isolated restore PostgreSQL became unhealthy." >&2 + return 1 + ;; + esac + sleep 1 + done + + # The PostGIS image initializes its requested database with PostGIS already + # installed. Restore into a database created from template0 instead, so the + # archive remains responsible for recreating its own extensions and schema. + docker exec "$restore_container" createdb \ + --username postgres \ + --template template0 \ + restore_check + + if ! docker exec --interactive "$restore_container" pg_restore \ + --username postgres \ + --dbname restore_check \ + --exit-on-error \ + --no-owner \ + --no-privileges <"$dump"; then + docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true + docker inspect "$restore_container" >"$evidence_dir/restore-container.json" 2>&1 || true + echo "Isolated restore failed; container evidence was retained." >&2 + return 1 + fi + + tables=$(docker exec "$restore_container" psql \ + --username postgres --dbname restore_check --tuples-only --no-align \ + --command="select count(*) from pg_tables where schemaname='public' and tablename <> 'spatial_ref_sys';") + migrations=$(docker exec "$restore_container" psql \ + --username postgres --dbname restore_check --tuples-only --no-align \ + --command='select count(*) from schema_migrations;') + postgis=$(docker exec "$restore_container" psql \ + --username postgres --dbname restore_check --tuples-only --no-align \ + --command='select PostGIS_Lib_Version();') + + if [[ ! "$tables" =~ ^[1-9][0-9]*$ ]] || [[ ! "$migrations" =~ ^[1-9][0-9]*$ ]]; then + echo "The isolated restore does not contain the expected application schema." >&2 + return 1 + fi + + { + printf 'snapshot_id=%s\n' "$selected_snapshot" + printf 'application_tables=%s\n' "$tables" + printf 'schema_migrations=%s\n' "$migrations" + printf 'postgis_version=%s\n' "$postgis" + printf 'verified_at=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" + } >"$evidence_dir/restore-drill.txt" + chmod 600 "$evidence_dir/restore-drill.txt" "$evidence_dir/pg-restore-catalog.txt" + + docker rm --force "$restore_container" >/dev/null + restore_started=false + docker volume rm "$restore_volume" >/dev/null + restore_volume_created=false + rm -f "$dump" +} + +if [[ "$action" == restore-drill ]]; then + restore_snapshot "$snapshot_id" + echo "Isolated restore drill passed for snapshot: $snapshot_id" + printf 'Non-secret evidence: %s\n' "$evidence_dir" + exit 0 +fi + +check_source + +ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \ + "bash -s -- '$PRODUCTION_REMOTE_ROOT/.env' '$remote_dump' '$PRODUCTION_EXPECTED_ENVIRONMENT'" \ + <"$ROOT/scripts/backup-external-postgres.sh" +remote_created=true + +scp -p \ + "$PRODUCTION_SSH_TARGET:$remote_dump" \ + "$PRODUCTION_SSH_TARGET:$remote_dump.sha256" \ + "$PRODUCTION_SSH_TARGET:$remote_dump.metadata" \ + "$work_dir/" + +dump="$work_dir/$(basename -- "$remote_dump")" +checksum="$dump.sha256" +metadata="$dump.metadata" + +( + cd "$work_dir" + sha256sum --check "$(basename -- "$checksum")" +) >"$evidence_dir/source-checksum.txt" +pg_restore --list "$dump" >"$evidence_dir/source-catalog.txt" +cp "$checksum" "$metadata" "$evidence_dir/" +chmod 600 "$evidence_dir"/* + +run_tag="$OFFSITE_RESTIC_TAG-$run_id" +backup_json="$evidence_dir/restic-backup.jsonl" +"$restic" backup \ + --json \ + --host "$OFFSITE_RESTIC_HOST" \ + --tag "$OFFSITE_RESTIC_TAG" \ + --tag "$run_tag" \ + --stdin-filename production.dump \ + --stdin-from-command -- cat "$dump" >"$backup_json" + +snapshot_id=$(jq -r 'select(.message_type == "summary") | .snapshot_id // empty' "$backup_json" | tail -n 1) +if [[ -z "$snapshot_id" ]]; then + echo "Restic did not report a completed snapshot." >&2 + exit 1 +fi + +snapshot_matches=$($restic snapshots --json --tag "$run_tag" | jq 'length') +if [[ "$snapshot_matches" != 1 ]]; then + echo "The run tag did not resolve to exactly one snapshot." >&2 + exit 1 +fi + +"$restic" check --read-data >"$evidence_dir/restic-check.txt" +restore_snapshot "$snapshot_id" + +printf '%s\n' "$snapshot_id" >"$evidence_dir/snapshot-id.txt" +chmod 600 "$evidence_dir"/* + +ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \ + "rm -f -- '$remote_dump' '$remote_dump.sha256' '$remote_dump.metadata'" +remote_created=false + +echo "Encrypted off-server backup and isolated restore drill passed." +printf 'Snapshot: %s\n' "$snapshot_id" +printf 'Non-secret evidence: %s\n' "$evidence_dir"