Add verified offsite production recovery
This commit is contained in:
parent
c085620762
commit
ead8a67d0f
|
|
@ -564,6 +564,107 @@ failure paths on the observed workstation. A MinIO volume on that same
|
|||
workstation is not an off-site backup and does not establish production RPO,
|
||||
RTO, retention, capacity, key custody, object locking, or database HA.
|
||||
|
||||
## Encrypted production backup to an independent host
|
||||
|
||||
The production workflow streams an external PostgreSQL 18 custom-format dump
|
||||
through an encrypted Restic repository on an SSH target that must resolve to a
|
||||
host other than production. It then runs `restic check --read-data` and restores
|
||||
the newest snapshot into a uniquely named, temporary PostGIS container created
|
||||
from `template0`. The source production database is never a restore target.
|
||||
|
||||
Bootstrap the checksum-pinned Restic binary and create the ignored local
|
||||
operations configuration:
|
||||
|
||||
```bash
|
||||
./scripts/bootstrap-restic.sh
|
||||
./scripts/init-production-operations.sh
|
||||
```
|
||||
|
||||
The initializer creates only these ignored local files:
|
||||
|
||||
- `tmp/production-operations/backup.env`, mode `0600`;
|
||||
- `tmp/production-operations/restic-password`, mode `0600`.
|
||||
|
||||
The generated configuration uses the reviewed `whoneedhelp` production SSH
|
||||
target and the independent `buyvm-maya` SFTP target. Inspect the read-only plan
|
||||
before repository initialization:
|
||||
|
||||
```bash
|
||||
./scripts/production-offsite-backup.sh plan
|
||||
```
|
||||
|
||||
Repository initialization is an external mutation and requires the exact
|
||||
repository value printed by the plan:
|
||||
|
||||
```bash
|
||||
WNH_OFFSITE_BACKUP_INIT_CONFIRM='sftp:buyvm-maya:backups/who_need_help-production' \
|
||||
./scripts/production-offsite-backup.sh init
|
||||
```
|
||||
|
||||
The routine operations are:
|
||||
|
||||
```bash
|
||||
./scripts/production-offsite-backup.sh run
|
||||
./scripts/production-offsite-backup.sh check
|
||||
./scripts/production-offsite-backup.sh restore-drill
|
||||
```
|
||||
|
||||
`run` creates a fresh source dump, copies it into the Restic input, publishes
|
||||
one encrypted snapshot only after the producer succeeds, performs a full-data
|
||||
repository check, and runs the isolated restore drill. It removes its exact
|
||||
temporary source dump, local scratch directory, Docker container, and Docker
|
||||
volume on success, failure, or interrupt. Non-secret catalogs, checksums, and
|
||||
restore observations remain under ignored
|
||||
`output/production-operations/<run-id>/`.
|
||||
|
||||
Install the daily user-systemd timer only after one manual `run` succeeds:
|
||||
|
||||
```bash
|
||||
./scripts/install-production-backup-timer.sh
|
||||
systemctl --user list-timers --all who-need-help-production-backup.timer
|
||||
systemctl --user status who-need-help-production-backup.service --no-pager
|
||||
journalctl --user -u who-need-help-production-backup.service --no-pager
|
||||
```
|
||||
|
||||
The timer uses `Persistent=true`, so systemd runs a missed calendar invocation
|
||||
after the workstation and user manager return. A failed service invokes the
|
||||
external monitor's backup-failure notification. It does not retry, delete
|
||||
snapshots, or infer a retention policy.
|
||||
|
||||
The Restic password is required for every restore. Before relying on this as
|
||||
the only recovery copy, place it in an operator-controlled password manager or
|
||||
offline recovery record. Never print it into a terminal log, commit it, or
|
||||
copy it into the backup repository. There is intentionally no automatic
|
||||
`forget`/`prune` policy yet: retention, RPO, RTO, capacity, and key custody are
|
||||
operator decisions, and the script does not invent them.
|
||||
|
||||
## Independent public readiness monitor
|
||||
|
||||
The external monitor runs on an SSH target that must resolve to a host other
|
||||
than production. It requests only the public readiness endpoint and accepts
|
||||
only HTTP 200 with the exact JSON object `{"status":"ready"}`. State is kept
|
||||
in a mode-`0600` JSON file. Email is sent once when the state changes to down
|
||||
and once when it recovers; repeated down checks do not send repeated alerts.
|
||||
|
||||
Install it from the production SMTP configuration without printing the SMTP
|
||||
credential:
|
||||
|
||||
```bash
|
||||
./scripts/install-production-external-monitor.sh
|
||||
ssh buyvm-maya \
|
||||
'systemctl --user status who-need-help-production-monitor.timer --no-pager'
|
||||
ssh buyvm-maya \
|
||||
'journalctl --user -u who-need-help-production-monitor.service --no-pager'
|
||||
```
|
||||
|
||||
The default schedule is once per minute with a three-second HTTP timeout. Those
|
||||
defaults match the current application container health timeout and were
|
||||
installed only after public readiness requests from the selected monitor host
|
||||
were observed to finish below one second. They are observations of the current
|
||||
path, not universal capacity or availability guarantees. Re-running the
|
||||
installer refreshes the exact script, mode-`0600` configuration, and user
|
||||
units, then performs one check before enabling the timer.
|
||||
|
||||
## Local external-service boundary drill
|
||||
|
||||
Run the OAuth, SMTP, and provider-neutral push protocol checks without public
|
||||
|
|
|
|||
|
|
@ -93,6 +93,22 @@ backup_path=$(printf '%s\n' "$backup_output" | sed -n 's/^Backup: //p')
|
|||
Local backups and local MinIO drills verify mechanics but are not proof of
|
||||
off-site durability or a production recovery policy.
|
||||
|
||||
The implemented external workflow can be rehearsed without restoring over the
|
||||
source database:
|
||||
|
||||
```bash
|
||||
./scripts/bootstrap-restic.sh
|
||||
./scripts/production-offsite-backup.sh plan
|
||||
./scripts/production-offsite-backup.sh run
|
||||
systemctl --user status who-need-help-production-backup.service --no-pager
|
||||
ssh buyvm-maya \
|
||||
'systemctl --user status who-need-help-production-monitor.timer --no-pager'
|
||||
```
|
||||
|
||||
Do not mark backup ownership complete merely because this command passes. The
|
||||
operator must still store the Restic key independently and approve retention,
|
||||
RPO, RTO, capacity, and responsible owners.
|
||||
|
||||
## 5. Promote with an explicit rollback point
|
||||
|
||||
Use the release workflow in `docs/operations.md`; do not copy mutable source
|
||||
|
|
|
|||
|
|
@ -1,8 +1,42 @@
|
|||
# Who Need Help — implementation verification
|
||||
|
||||
Observed through 2026-07-28 in the local workspace. This report separates observed
|
||||
Observed through 2026-08-03 in the local workspace. This report separates observed
|
||||
results from product limits and unknown production properties.
|
||||
|
||||
## Independent production backup and monitoring on 2026-08-03
|
||||
|
||||
This verification changed only the independent backup repository, operator
|
||||
user-systemd units, and independent monitor host. It did not deploy application
|
||||
source, change production data, modify the hackathon test deployment, push Git,
|
||||
or edit the Devpost submission.
|
||||
|
||||
- An encrypted Restic 0.19.1 repository was initialized at
|
||||
`sftp:buyvm-maya:backups/who_need_help-production`. SSH resolution verified
|
||||
that this target and the production target are different hosts. The Restic
|
||||
password and workflow configuration remain only in ignored mode-`0600`
|
||||
local files; password-manager/offline key custody is still an operator action.
|
||||
- Two production snapshots exist. The installed user-systemd execution produced
|
||||
snapshot `98d7d2e8e3e7c8dbf7908bbaceeb2acedf202a7e1dfd77c5398ba668a7981236`,
|
||||
passed `restic check --read-data`, and restored into a temporary database
|
||||
created from `template0`. The isolated restore observed 34 public application
|
||||
tables, 24 applied schema migrations, and PostGIS 3.6.4. Non-secret evidence
|
||||
is retained at
|
||||
`output/production-operations/20260803T012727Z-1694894/`.
|
||||
- That exact systemd service exited `0/SUCCESS` after 4 minutes 31.643 seconds
|
||||
wall time, 4.115 seconds local CPU time, and a 78.2 MiB local memory peak.
|
||||
These are observations of one run over the selected SFTP path, not a minimum
|
||||
resource requirement or recovery objective. The scheduled source dump and
|
||||
every run-scoped temporary restore container and volume were absent afterward.
|
||||
- The daily backup timer is enabled with `Persistent=true`. Its failure unit
|
||||
routes one alert through the independent monitor host. No retention deletion,
|
||||
automatic pruning, or database rollback is configured.
|
||||
- The independent BuyVM monitor performs the public readiness check once per
|
||||
minute. The observed check returned HTTP 200 with the exact ready payload and
|
||||
the service exited successfully. Its SMTP configuration and state are mode
|
||||
`0600`. A one-time monitoring test message was accepted by Brevo; mailbox
|
||||
receipt has not yet been independently observed, so end-to-end alert delivery
|
||||
is not claimed complete.
|
||||
|
||||
## Production Android bundle replay on 2026-07-28
|
||||
|
||||
The release tooling and application inputs below are exact local commit
|
||||
|
|
|
|||
53
scripts/bootstrap-restic.sh
Executable file
53
scripts/bootstrap-restic.sh
Executable file
|
|
@ -0,0 +1,53 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
|
||||
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||
target_dir="$ROOT/.tools/restic"
|
||||
target="$target_dir/restic"
|
||||
image=${WNH_BACKUP_TOOLS_IMAGE:-who-need-help:backup-tools}
|
||||
expected='restic 0.19.1 compiled with go1.26.5'
|
||||
|
||||
if [[ -x "$target" ]] && [[ "$($target version)" == "$expected"* ]]; then
|
||||
printf 'Pinned Restic is already available: %s\n' "$target"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
for command in docker install mktemp; do
|
||||
command -v "$command" >/dev/null 2>&1 || {
|
||||
echo "Required command is unavailable: $command" >&2
|
||||
exit 2
|
||||
}
|
||||
done
|
||||
|
||||
mkdir -p "$target_dir"
|
||||
chmod 700 "$ROOT/.tools" "$target_dir"
|
||||
|
||||
docker build --tag "$image" --file "$ROOT/Dockerfile.backup" "$ROOT"
|
||||
|
||||
container=$(docker create "$image" restic version)
|
||||
partial=$(mktemp "$target_dir/restic.XXXXXX")
|
||||
|
||||
cleanup() {
|
||||
docker rm --force "$container" >/dev/null 2>&1 || true
|
||||
rm -f "$partial"
|
||||
}
|
||||
trap cleanup EXIT HUP INT TERM
|
||||
|
||||
docker cp "$container:/usr/local/bin/restic" "$partial"
|
||||
chmod 700 "$partial"
|
||||
|
||||
actual=$($partial version)
|
||||
if [[ "$actual" != "$expected"* ]]; then
|
||||
printf 'Unexpected Restic build: %s\n' "$actual" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
mv "$partial" "$target"
|
||||
partial=
|
||||
|
||||
trap - EXIT HUP INT TERM
|
||||
docker rm --force "$container" >/dev/null
|
||||
|
||||
printf 'Installed pinned Restic: %s\n' "$target"
|
||||
printf 'Version: %s\n' "$actual"
|
||||
67
scripts/init-production-operations.sh
Executable file
67
scripts/init-production-operations.sh
Executable file
|
|
@ -0,0 +1,67 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
|
||||
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||
runtime_dir="$ROOT/tmp/production-operations"
|
||||
config=${1:-"$runtime_dir/backup.env"}
|
||||
|
||||
if [[ "$config" != /* ]]; then
|
||||
config="$ROOT/$config"
|
||||
fi
|
||||
|
||||
password_file="$runtime_dir/restic-password"
|
||||
|
||||
for command in openssl ssh; do
|
||||
command -v "$command" >/dev/null 2>&1 || {
|
||||
echo "Required command is unavailable: $command" >&2
|
||||
exit 2
|
||||
}
|
||||
done
|
||||
|
||||
if [[ -e "$config" || -e "$password_file" ]]; then
|
||||
echo "Refusing to replace existing production operations configuration." >&2
|
||||
printf 'Config: %s\nPassword file: %s\n' "$config" "$password_file" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
mkdir -p "$runtime_dir"
|
||||
chmod 700 "$ROOT/tmp" "$runtime_dir"
|
||||
|
||||
source_target=whoneedhelp
|
||||
repository_target=buyvm-maya
|
||||
|
||||
source_host=$(ssh -G "$source_target" | awk '$1 == "hostname" {print $2; exit}')
|
||||
repository_host=$(ssh -G "$repository_target" | awk '$1 == "hostname" {print $2; exit}')
|
||||
|
||||
if [[ -z "$source_host" || -z "$repository_host" ]]; then
|
||||
echo "Could not resolve both SSH targets." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
if [[ "$source_host" == "$repository_host" ]]; then
|
||||
echo "The backup repository must not resolve to the production host." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
openssl rand -base64 48 | tr -d '\n' >"$password_file"
|
||||
printf '\n' >>"$password_file"
|
||||
chmod 600 "$password_file"
|
||||
|
||||
cat >"$config" <<EOF
|
||||
PRODUCTION_SSH_TARGET=$source_target
|
||||
PRODUCTION_REMOTE_ROOT=/srv/who_need_help-production
|
||||
PRODUCTION_EXPECTED_ENVIRONMENT=production
|
||||
OFFSITE_RESTIC_REPOSITORY=sftp:$repository_target:backups/who_need_help-production
|
||||
OFFSITE_RESTIC_PASSWORD_FILE=$password_file
|
||||
OFFSITE_RESTIC_HOST=who-need-help-production
|
||||
OFFSITE_RESTIC_TAG=who-need-help-production
|
||||
BACKUP_ON_CALENDAR=daily
|
||||
EOF
|
||||
chmod 600 "$config"
|
||||
|
||||
printf 'Created mode-0600 operations configuration: %s\n' "$config"
|
||||
printf 'Created mode-0600 Restic key file: %s\n' "$password_file"
|
||||
printf 'Production resolves to: %s\n' "$source_host"
|
||||
printf 'Encrypted repository resolves to a different host: %s\n' "$repository_host"
|
||||
echo "The Restic key is required for every restore. Copy it to an operator-controlled password manager before relying on this backup as the only recovery copy."
|
||||
79
scripts/install-production-backup-timer.sh
Executable file
79
scripts/install-production-backup-timer.sh
Executable file
|
|
@ -0,0 +1,79 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
|
||||
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||
config=${1:-"$ROOT/tmp/production-operations/backup.env"}
|
||||
unit_dir=${SYSTEMD_USER_UNIT_DIR:-"$HOME/.config/systemd/user"}
|
||||
|
||||
if [[ "$config" != /* ]]; then
|
||||
config="$ROOT/$config"
|
||||
fi
|
||||
if [[ ! -f "$config" || "$(stat -c '%a' "$config")" != 600 ]]; then
|
||||
echo "The backup configuration must exist with mode 0600: $config" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# shellcheck source=/dev/null
|
||||
source "$config"
|
||||
: "${BACKUP_ON_CALENDAR:?Set BACKUP_ON_CALENDAR}"
|
||||
|
||||
case "$BACKUP_ON_CALENDAR" in
|
||||
*$'\n'* | *$'\r'*) echo "BACKUP_ON_CALENDAR must be one line." >&2; exit 2 ;;
|
||||
esac
|
||||
|
||||
monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya}
|
||||
remote_monitor=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help}/production-external-monitor.py
|
||||
remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json}
|
||||
ssh_path=$(command -v ssh)
|
||||
|
||||
mkdir -p "$unit_dir"
|
||||
chmod 700 "$HOME/.config" "$HOME/.config/systemd" "$unit_dir"
|
||||
|
||||
cat >"$unit_dir/who-need-help-production-backup.service" <<EOF
|
||||
[Unit]
|
||||
Description=Who Need Help encrypted off-server production backup and restore drill
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
OnFailure=who-need-help-production-backup-alert.service
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
WorkingDirectory=$ROOT
|
||||
ExecStart=$ROOT/scripts/production-offsite-backup.sh run $config
|
||||
EOF
|
||||
|
||||
cat >"$unit_dir/who-need-help-production-backup-alert.service" <<EOF
|
||||
[Unit]
|
||||
Description=Notify operators that the Who Need Help production backup failed
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=$ssh_path -o BatchMode=yes $monitor_target /usr/bin/python3 $remote_monitor --config $remote_config notify-backup-failure --unit who-need-help-production-backup.service
|
||||
EOF
|
||||
|
||||
cat >"$unit_dir/who-need-help-production-backup.timer" <<EOF
|
||||
[Unit]
|
||||
Description=Schedule Who Need Help encrypted production backups
|
||||
|
||||
[Timer]
|
||||
OnCalendar=$BACKUP_ON_CALENDAR
|
||||
Persistent=true
|
||||
Unit=who-need-help-production-backup.service
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
EOF
|
||||
|
||||
chmod 600 \
|
||||
"$unit_dir/who-need-help-production-backup.service" \
|
||||
"$unit_dir/who-need-help-production-backup-alert.service" \
|
||||
"$unit_dir/who-need-help-production-backup.timer"
|
||||
|
||||
systemctl --user daemon-reload
|
||||
systemctl --user enable --now who-need-help-production-backup.timer
|
||||
|
||||
echo "Production backup timer installed without starting or changing the application."
|
||||
systemctl --user list-timers --all --no-pager who-need-help-production-backup.timer
|
||||
137
scripts/install-production-external-monitor.sh
Executable file
137
scripts/install-production-external-monitor.sh
Executable file
|
|
@ -0,0 +1,137 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
|
||||
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||
source_target=${PRODUCTION_SSH_TARGET:-whoneedhelp}
|
||||
monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya}
|
||||
production_env=${PRODUCTION_ENV_PATH:-/srv/who_need_help-production/.env}
|
||||
remote_root=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help}
|
||||
remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json}
|
||||
remote_state=${MONITOR_REMOTE_STATE:-/home/simple/.local/state/who-need-help/monitor.json}
|
||||
monitor_url=${MONITOR_URL:-https://whoneedhelp.com/healthz/ready}
|
||||
monitor_calendar=${MONITOR_ON_CALENDAR:-'*:0/1'}
|
||||
health_timeout=${MONITOR_HEALTH_TIMEOUT_SECONDS:-3}
|
||||
|
||||
for command in mktemp scp ssh; do
|
||||
command -v "$command" >/dev/null 2>&1 || {
|
||||
echo "Required command is unavailable: $command" >&2
|
||||
exit 2
|
||||
}
|
||||
done
|
||||
|
||||
source_host=$(ssh -G "$source_target" | awk '$1 == "hostname" {print $2; exit}')
|
||||
monitor_host=$(ssh -G "$monitor_target" | awk '$1 == "hostname" {print $2; exit}')
|
||||
if [[ -z "$source_host" || -z "$monitor_host" || "$source_host" == "$monitor_host" ]]; then
|
||||
echo "The external monitor must resolve and run on a host other than production." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
case "$health_timeout" in
|
||||
'' | *[!0-9]*) echo "MONITOR_HEALTH_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;;
|
||||
0) echo "MONITOR_HEALTH_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;;
|
||||
esac
|
||||
|
||||
work_dir=$(mktemp -d "$ROOT/tmp/production-operations/monitor-install.XXXXXX")
|
||||
config="$work_dir/monitor.json"
|
||||
service="$work_dir/who-need-help-production-monitor.service"
|
||||
timer="$work_dir/who-need-help-production-monitor.timer"
|
||||
|
||||
cleanup() {
|
||||
rm -rf "$work_dir"
|
||||
}
|
||||
trap cleanup EXIT HUP INT TERM
|
||||
|
||||
ssh -o BatchMode=yes "$source_target" \
|
||||
"python3 - '$production_env' '$monitor_url' '$health_timeout'" >"$config" <<'PY'
|
||||
import json
|
||||
import shlex
|
||||
import sys
|
||||
|
||||
path, health_url, timeout = sys.argv[1:]
|
||||
wanted = {
|
||||
"SMTP_RELAY",
|
||||
"SMTP_PORT",
|
||||
"SMTP_USERNAME",
|
||||
"SMTP_PASSWORD",
|
||||
"SMTP_TLS",
|
||||
"SMTP_SSL",
|
||||
"EMAIL_FROM_ADDRESS",
|
||||
"EMAIL_FROM_NAME",
|
||||
"SUPPORT_INBOX_ADDRESS",
|
||||
}
|
||||
values = {}
|
||||
with open(path, encoding="utf-8") as handle:
|
||||
for raw_line in handle:
|
||||
line = raw_line.strip()
|
||||
if not line or line.startswith("#") or "=" not in line:
|
||||
continue
|
||||
key, value = line.split("=", 1)
|
||||
if key not in wanted:
|
||||
continue
|
||||
parsed = shlex.split(value, comments=False, posix=True)
|
||||
values[key] = parsed[0] if parsed else ""
|
||||
|
||||
missing = sorted(key for key in wanted if not values.get(key))
|
||||
if missing:
|
||||
raise SystemExit("Missing production mail settings: " + ", ".join(missing))
|
||||
|
||||
payload = {
|
||||
"health_timeout_seconds": int(timeout),
|
||||
"health_url": health_url,
|
||||
"smtp": {
|
||||
"from_address": values["EMAIL_FROM_ADDRESS"],
|
||||
"from_name": values["EMAIL_FROM_NAME"],
|
||||
"implicit_ssl": values["SMTP_SSL"].lower() == "true",
|
||||
"password": values["SMTP_PASSWORD"],
|
||||
"port": int(values["SMTP_PORT"]),
|
||||
"recipient": values["SUPPORT_INBOX_ADDRESS"],
|
||||
"relay": values["SMTP_RELAY"],
|
||||
"starttls": values["SMTP_TLS"].lower() == "always",
|
||||
"username": values["SMTP_USERNAME"],
|
||||
},
|
||||
}
|
||||
json.dump(payload, sys.stdout, ensure_ascii=False, indent=2, sort_keys=True)
|
||||
sys.stdout.write("\n")
|
||||
PY
|
||||
chmod 600 "$config"
|
||||
|
||||
cat >"$service" <<EOF
|
||||
[Unit]
|
||||
Description=Who Need Help independent production readiness monitor
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=/usr/bin/python3 $remote_root/production-external-monitor.py --config $remote_config --state $remote_state check
|
||||
EOF
|
||||
|
||||
cat >"$timer" <<EOF
|
||||
[Unit]
|
||||
Description=Run the Who Need Help independent production readiness monitor
|
||||
|
||||
[Timer]
|
||||
OnCalendar=$monitor_calendar
|
||||
Persistent=true
|
||||
Unit=who-need-help-production-monitor.service
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
EOF
|
||||
|
||||
ssh -o BatchMode=yes "$monitor_target" \
|
||||
"install -d -m 700 '$remote_root' /home/simple/.config/who-need-help /home/simple/.local/state/who-need-help /home/simple/.config/systemd/user"
|
||||
scp -q "$ROOT/scripts/production-external-monitor.py" \
|
||||
"$monitor_target:$remote_root/production-external-monitor.py"
|
||||
scp -q "$config" "$monitor_target:$remote_config"
|
||||
scp -q "$service" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.service"
|
||||
scp -q "$timer" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.timer"
|
||||
|
||||
ssh -o BatchMode=yes "$monitor_target" \
|
||||
"chmod 700 '$remote_root/production-external-monitor.py'; chmod 600 '$remote_config' /home/simple/.config/systemd/user/who-need-help-production-monitor.service /home/simple/.config/systemd/user/who-need-help-production-monitor.timer; systemctl --user daemon-reload; systemctl --user start who-need-help-production-monitor.service; systemctl --user enable --now who-need-help-production-monitor.timer"
|
||||
|
||||
printf 'External monitor installed on %s (%s).\n' "$monitor_target" "$monitor_host"
|
||||
printf 'Health URL: %s\n' "$monitor_url"
|
||||
printf 'Schedule: %s; request timeout: %ss.\n' "$monitor_calendar" "$health_timeout"
|
||||
echo "The SMTP credential is stored only in a mode-0600 configuration on the external monitor host."
|
||||
231
scripts/production-external-monitor.py
Executable file
231
scripts/production-external-monitor.py
Executable file
|
|
@ -0,0 +1,231 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Stateful external production health and operations notifications."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import smtplib
|
||||
import ssl
|
||||
import sys
|
||||
import tempfile
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from datetime import datetime, timezone
|
||||
from email.message import EmailMessage
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
def utc_now() -> str:
|
||||
return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z")
|
||||
|
||||
|
||||
def read_json(path: Path) -> dict[str, Any]:
|
||||
with path.open("r", encoding="utf-8") as handle:
|
||||
value = json.load(handle)
|
||||
if not isinstance(value, dict):
|
||||
raise ValueError(f"Expected a JSON object in {path}")
|
||||
return value
|
||||
|
||||
|
||||
def write_json_atomic(path: Path, value: dict[str, Any]) -> None:
|
||||
path.parent.mkdir(mode=0o700, parents=True, exist_ok=True)
|
||||
descriptor, temporary_name = tempfile.mkstemp(dir=path.parent, prefix=f".{path.name}.")
|
||||
temporary = Path(temporary_name)
|
||||
try:
|
||||
with os.fdopen(descriptor, "w", encoding="utf-8") as handle:
|
||||
json.dump(value, handle, ensure_ascii=False, indent=2, sort_keys=True)
|
||||
handle.write("\n")
|
||||
temporary.chmod(0o600)
|
||||
temporary.replace(path)
|
||||
finally:
|
||||
temporary.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def required_text(mapping: dict[str, Any], name: str) -> str:
|
||||
value = mapping.get(name)
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise ValueError(f"Missing non-empty configuration value: {name}")
|
||||
return value.strip()
|
||||
|
||||
|
||||
def send_message(config: dict[str, Any], subject: str, body: str) -> None:
|
||||
smtp = config.get("smtp")
|
||||
if not isinstance(smtp, dict):
|
||||
raise ValueError("Missing SMTP configuration")
|
||||
|
||||
relay = required_text(smtp, "relay")
|
||||
port = int(smtp.get("port"))
|
||||
username = required_text(smtp, "username")
|
||||
password = required_text(smtp, "password")
|
||||
from_address = required_text(smtp, "from_address")
|
||||
from_name = required_text(smtp, "from_name")
|
||||
recipient = required_text(smtp, "recipient")
|
||||
implicit_ssl = bool(smtp.get("implicit_ssl", False))
|
||||
starttls = bool(smtp.get("starttls", True))
|
||||
|
||||
message = EmailMessage()
|
||||
message["From"] = f"{from_name} <{from_address}>"
|
||||
message["To"] = recipient
|
||||
message["Subject"] = subject
|
||||
message.set_content(body)
|
||||
|
||||
context = ssl.create_default_context()
|
||||
if implicit_ssl:
|
||||
client_context = smtplib.SMTP_SSL(relay, port, timeout=30, context=context)
|
||||
else:
|
||||
client_context = smtplib.SMTP(relay, port, timeout=30)
|
||||
|
||||
with client_context as client:
|
||||
if not implicit_ssl:
|
||||
client.ehlo()
|
||||
if starttls:
|
||||
client.starttls(context=context)
|
||||
client.ehlo()
|
||||
client.login(username, password)
|
||||
client.send_message(message)
|
||||
|
||||
|
||||
def check_health(config: dict[str, Any]) -> tuple[str, str]:
|
||||
url = required_text(config, "health_url")
|
||||
timeout = float(config.get("health_timeout_seconds"))
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
headers={"User-Agent": "WhoNeedHelp-External-Monitor/1.0"},
|
||||
)
|
||||
|
||||
try:
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
status = response.status
|
||||
payload = response.read(4096)
|
||||
parsed = json.loads(payload.decode("utf-8"))
|
||||
if status == 200 and parsed == {"status": "ready"}:
|
||||
return "up", f"HTTP {status}; ready payload matched"
|
||||
return "down", f"HTTP {status}; unexpected readiness payload"
|
||||
except (OSError, ValueError, urllib.error.URLError) as error:
|
||||
return "down", f"{type(error).__name__}: {str(error)[:500]}"
|
||||
|
||||
|
||||
def monitor(config_path: Path, state_path: Path) -> int:
|
||||
config = read_json(config_path)
|
||||
previous: dict[str, Any] = {}
|
||||
if state_path.exists():
|
||||
previous = read_json(state_path)
|
||||
|
||||
status, detail = check_health(config)
|
||||
previous_status = previous.get("status")
|
||||
|
||||
if status != previous_status:
|
||||
if status == "down":
|
||||
send_message(
|
||||
config,
|
||||
"[Who Need Help] Production readiness is DOWN",
|
||||
"\n".join(
|
||||
[
|
||||
"The independent production readiness check failed.",
|
||||
f"URL: {required_text(config, 'health_url')}",
|
||||
f"Observed at: {utc_now()}",
|
||||
f"Result: {detail}",
|
||||
"This alert is sent once per state transition.",
|
||||
]
|
||||
),
|
||||
)
|
||||
elif previous_status == "down":
|
||||
send_message(
|
||||
config,
|
||||
"[Who Need Help] Production readiness recovered",
|
||||
"\n".join(
|
||||
[
|
||||
"The independent production readiness check recovered.",
|
||||
f"URL: {required_text(config, 'health_url')}",
|
||||
f"Observed at: {utc_now()}",
|
||||
f"Result: {detail}",
|
||||
]
|
||||
),
|
||||
)
|
||||
|
||||
write_json_atomic(
|
||||
state_path,
|
||||
{
|
||||
"checked_at": utc_now(),
|
||||
"detail": detail,
|
||||
"status": status,
|
||||
},
|
||||
)
|
||||
print(json.dumps({"status": status, "detail": detail}, sort_keys=True))
|
||||
return 0 if status == "up" else 1
|
||||
|
||||
|
||||
def notify_backup_failure(config_path: Path, unit: str) -> int:
|
||||
config = read_json(config_path)
|
||||
send_message(
|
||||
config,
|
||||
"[Who Need Help] Production backup or restore verification failed",
|
||||
"\n".join(
|
||||
[
|
||||
"The scheduled encrypted production backup did not finish successfully.",
|
||||
f"Unit: {unit}",
|
||||
f"Observed at: {utc_now()}",
|
||||
"Inspect the local user-systemd journal and do not treat the newest snapshot as verified until a restore drill passes.",
|
||||
]
|
||||
),
|
||||
)
|
||||
print("Backup failure notification sent.")
|
||||
return 0
|
||||
|
||||
|
||||
def send_test_notification(config_path: Path) -> int:
|
||||
config = read_json(config_path)
|
||||
send_message(
|
||||
config,
|
||||
"[Who Need Help] Operations monitoring test",
|
||||
"\n".join(
|
||||
[
|
||||
"This is a one-time delivery verification for the independent production monitor.",
|
||||
f"Health URL: {required_text(config, 'health_url')}",
|
||||
f"Sent at: {utc_now()}",
|
||||
"No production incident was detected and no application data was changed.",
|
||||
]
|
||||
),
|
||||
)
|
||||
print("Operations monitoring test notification sent.")
|
||||
return 0
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument(
|
||||
"--config",
|
||||
type=Path,
|
||||
default=Path.home() / ".config/who-need-help/monitor.json",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--state",
|
||||
type=Path,
|
||||
default=Path.home() / ".local/state/who-need-help/monitor.json",
|
||||
)
|
||||
subparsers = parser.add_subparsers(dest="action", required=True)
|
||||
subparsers.add_parser("check")
|
||||
subparsers.add_parser("send-test-notification")
|
||||
backup = subparsers.add_parser("notify-backup-failure")
|
||||
backup.add_argument("--unit", required=True)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
try:
|
||||
if args.action == "check":
|
||||
return monitor(args.config, args.state)
|
||||
if args.action == "send-test-notification":
|
||||
return send_test_notification(args.config)
|
||||
return notify_backup_failure(args.config, args.unit)
|
||||
except (OSError, ValueError, smtplib.SMTPException, json.JSONDecodeError) as error:
|
||||
print(f"Operations monitor failed: {type(error).__name__}: {error}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
384
scripts/production-offsite-backup.sh
Executable file
384
scripts/production-offsite-backup.sh
Executable file
|
|
@ -0,0 +1,384 @@
|
|||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
|
||||
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||
action=${1:-plan}
|
||||
config=${2:-"$ROOT/tmp/production-operations/backup.env"}
|
||||
restic="$ROOT/.tools/restic/restic"
|
||||
restore_image='postgis/postgis:18-3.6-alpine@sha256:05d68c7f0f19b9aa0bf7c4a2049b2e8b38b44a63116392b95726a4c913766cf6'
|
||||
|
||||
case "$action" in
|
||||
plan | init | run | check | restore-drill) ;;
|
||||
*)
|
||||
echo "Usage: $0 [plan|init|run|check|restore-drill] [CONFIG]" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
|
||||
if [[ "$config" != /* ]]; then
|
||||
config="$ROOT/$config"
|
||||
fi
|
||||
|
||||
if [[ ! -f "$config" ]]; then
|
||||
echo "Operations configuration does not exist: $config" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
if [[ "$(stat -c '%a' "$config")" != 600 ]]; then
|
||||
echo "Operations configuration must have mode 0600: $config" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# shellcheck source=/dev/null
|
||||
source "$config"
|
||||
|
||||
: "${PRODUCTION_SSH_TARGET:?Set PRODUCTION_SSH_TARGET}"
|
||||
: "${PRODUCTION_REMOTE_ROOT:?Set PRODUCTION_REMOTE_ROOT}"
|
||||
: "${PRODUCTION_EXPECTED_ENVIRONMENT:?Set PRODUCTION_EXPECTED_ENVIRONMENT}"
|
||||
: "${OFFSITE_RESTIC_REPOSITORY:?Set OFFSITE_RESTIC_REPOSITORY}"
|
||||
: "${OFFSITE_RESTIC_PASSWORD_FILE:?Set OFFSITE_RESTIC_PASSWORD_FILE}"
|
||||
: "${OFFSITE_RESTIC_HOST:?Set OFFSITE_RESTIC_HOST}"
|
||||
: "${OFFSITE_RESTIC_TAG:?Set OFFSITE_RESTIC_TAG}"
|
||||
|
||||
if [[ "$PRODUCTION_EXPECTED_ENVIRONMENT" != production ]]; then
|
||||
echo "The source environment must be exactly production." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
case "$PRODUCTION_REMOTE_ROOT" in
|
||||
/srv/who_need_help-production) ;;
|
||||
*)
|
||||
echo "The production root is outside the reviewed deployment path." >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
|
||||
case "$OFFSITE_RESTIC_REPOSITORY" in
|
||||
sftp:*) ;;
|
||||
*)
|
||||
echo "This workflow currently accepts only an encrypted Restic SFTP repository." >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
|
||||
if [[ ! -f "$OFFSITE_RESTIC_PASSWORD_FILE" ]] ||
|
||||
[[ "$(stat -c '%a' "$OFFSITE_RESTIC_PASSWORD_FILE")" != 600 ]]; then
|
||||
echo "The Restic password file must exist with mode 0600." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
if [[ ! -x "$restic" ]]; then
|
||||
echo "Pinned Restic is unavailable. Run ./scripts/bootstrap-restic.sh first." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
for command in docker flock grep jq openssl pg_restore scp sha256sum ssh; do
|
||||
command -v "$command" >/dev/null 2>&1 || {
|
||||
echo "Required command is unavailable: $command" >&2
|
||||
exit 2
|
||||
}
|
||||
done
|
||||
|
||||
export RESTIC_REPOSITORY="$OFFSITE_RESTIC_REPOSITORY"
|
||||
export RESTIC_PASSWORD_FILE="$OFFSITE_RESTIC_PASSWORD_FILE"
|
||||
export RESTIC_CACHE_DIR="$ROOT/tmp/production-operations/restic-cache"
|
||||
|
||||
mkdir -p "$RESTIC_CACHE_DIR"
|
||||
chmod 700 "$ROOT/tmp" "$ROOT/tmp/production-operations" "$RESTIC_CACHE_DIR"
|
||||
|
||||
source_host=$(ssh -G "$PRODUCTION_SSH_TARGET" | awk '$1 == "hostname" {print $2; exit}')
|
||||
repository_alias=${OFFSITE_RESTIC_REPOSITORY#sftp:}
|
||||
repository_alias=${repository_alias%%:*}
|
||||
repository_host=$(ssh -G "$repository_alias" | awk '$1 == "hostname" {print $2; exit}')
|
||||
|
||||
if [[ -z "$source_host" || -z "$repository_host" ]] ||
|
||||
[[ "$source_host" == "$repository_host" ]]; then
|
||||
echo "The repository host must resolve and differ from production." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
check_source() {
|
||||
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
|
||||
"bash -s -- '$PRODUCTION_REMOTE_ROOT/.env' --check-only '$PRODUCTION_EXPECTED_ENVIRONMENT'" \
|
||||
<"$ROOT/scripts/backup-external-postgres.sh"
|
||||
}
|
||||
|
||||
repo_initialized() {
|
||||
"$restic" snapshots --json >/dev/null 2>&1
|
||||
}
|
||||
|
||||
printf 'Production SSH target: %s (%s)\n' "$PRODUCTION_SSH_TARGET" "$source_host"
|
||||
printf 'Production root: %s\n' "$PRODUCTION_REMOTE_ROOT"
|
||||
printf 'Encrypted repository: %s (%s)\n' "$OFFSITE_RESTIC_REPOSITORY" "$repository_host"
|
||||
printf 'Restic: %s\n' "$($restic version)"
|
||||
|
||||
if [[ "$action" == plan ]]; then
|
||||
check_source
|
||||
if repo_initialized; then
|
||||
echo "Encrypted repository opened successfully."
|
||||
else
|
||||
echo "Encrypted repository is not initialized or cannot be opened."
|
||||
fi
|
||||
echo "Plan completed without changing production or the repository."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [[ "$action" == init ]]; then
|
||||
confirmation="$OFFSITE_RESTIC_REPOSITORY"
|
||||
if [[ "${WNH_OFFSITE_BACKUP_INIT_CONFIRM:-}" != "$confirmation" ]]; then
|
||||
echo "Repository initialization requires exact confirmation:" >&2
|
||||
echo "WNH_OFFSITE_BACKUP_INIT_CONFIRM=$confirmation $0 init '$config'" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
check_source
|
||||
if repo_initialized; then
|
||||
echo "Encrypted repository is already initialized; no change was made."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
"$restic" init
|
||||
"$restic" check
|
||||
echo "Encrypted off-server Restic repository initialized and opened successfully."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
repo_initialized || {
|
||||
echo "Encrypted repository cannot be opened. Initialize it first." >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
if [[ "$action" == check ]]; then
|
||||
"$restic" check --read-data
|
||||
echo "Encrypted repository full-data check passed."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [[ "$action" == restore-drill ]]; then
|
||||
snapshot_id=$($restic snapshots --json --host "$OFFSITE_RESTIC_HOST" \
|
||||
--tag "$OFFSITE_RESTIC_TAG" --latest 1 | jq -r '.[0].short_id // empty')
|
||||
if [[ -z "$snapshot_id" ]]; then
|
||||
echo "No production backup snapshot is available for a restore drill." >&2
|
||||
exit 1
|
||||
fi
|
||||
else
|
||||
snapshot_id=
|
||||
fi
|
||||
|
||||
lock_file="$ROOT/tmp/production-operations/backup.lock"
|
||||
exec 9>"$lock_file"
|
||||
flock --nonblock 9 || {
|
||||
echo "Another production backup or restore drill is already running." >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
run_id=$(date -u +%Y%m%dT%H%M%SZ)-$$
|
||||
work_dir=$(mktemp -d "$ROOT/tmp/production-operations/run-$run_id.XXXXXX")
|
||||
evidence_dir="$ROOT/output/production-operations/$run_id"
|
||||
remote_dump="$PRODUCTION_REMOTE_ROOT/output/backups/production/scheduled-$run_id.dump"
|
||||
remote_created=false
|
||||
restore_container="wnh-production-restore-$run_id"
|
||||
restore_volume="wnh_production_restore_${run_id//[^a-zA-Z0-9]/_}"
|
||||
restore_started=false
|
||||
restore_volume_created=false
|
||||
|
||||
mkdir -p "$evidence_dir"
|
||||
chmod 700 "$ROOT/output" "$ROOT/output/production-operations" "$evidence_dir" "$work_dir"
|
||||
|
||||
cleanup() {
|
||||
trap - EXIT HUP INT TERM
|
||||
if [[ "$restore_started" == true ]]; then
|
||||
docker rm --force "$restore_container" >/dev/null 2>&1 || true
|
||||
fi
|
||||
if [[ "$restore_volume_created" == true ]]; then
|
||||
docker volume rm "$restore_volume" >/dev/null 2>&1 || true
|
||||
fi
|
||||
if [[ "$remote_created" == true ]]; then
|
||||
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
|
||||
"rm -f -- '$remote_dump' '$remote_dump.sha256' '$remote_dump.metadata'" \
|
||||
>/dev/null 2>&1 || true
|
||||
fi
|
||||
rm -rf "$work_dir"
|
||||
}
|
||||
trap cleanup EXIT HUP INT TERM
|
||||
|
||||
restore_snapshot() {
|
||||
local selected_snapshot=$1
|
||||
local dump="$work_dir/production.dump"
|
||||
local password
|
||||
local health
|
||||
local tables
|
||||
local migrations
|
||||
local postgis
|
||||
|
||||
"$restic" dump "$selected_snapshot" /production.dump >"$dump"
|
||||
[[ -s "$dump" ]] || {
|
||||
echo "Restic restored an empty database dump." >&2
|
||||
return 1
|
||||
}
|
||||
pg_restore --list "$dump" >"$evidence_dir/pg-restore-catalog.txt"
|
||||
|
||||
if ! docker image inspect "$restore_image" >/dev/null 2>&1; then
|
||||
docker pull "$restore_image" >"$evidence_dir/postgis-pull.txt"
|
||||
fi
|
||||
|
||||
password=$(openssl rand -base64 36 | tr -d '\n')
|
||||
docker volume create "$restore_volume" >/dev/null
|
||||
restore_volume_created=true
|
||||
docker run --detach \
|
||||
--name "$restore_container" \
|
||||
--env POSTGRES_PASSWORD="$password" \
|
||||
--env POSTGRES_DB=postgres \
|
||||
--volume "$restore_volume:/var/lib/postgresql" \
|
||||
--health-cmd='pg_isready --username postgres --dbname postgres' \
|
||||
--health-interval=1s \
|
||||
--health-timeout=2s \
|
||||
--health-retries=120 \
|
||||
"$restore_image" >/dev/null
|
||||
restore_started=true
|
||||
|
||||
# A freshly initialized postgres/postgis container briefly accepts
|
||||
# connections through its temporary bootstrap server. Wait until the image
|
||||
# has completed that bootstrap and started the final server before restoring.
|
||||
while ! docker logs "$restore_container" 2>&1 |
|
||||
grep -Fq 'PostgreSQL init process complete; ready for start up.'; do
|
||||
if [[ "$(docker inspect --format '{{.State.Running}}' "$restore_container")" != true ]]; then
|
||||
docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true
|
||||
echo "Isolated restore PostgreSQL exited during initialization." >&2
|
||||
return 1
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
|
||||
while true; do
|
||||
health=$(docker inspect --format '{{if .State.Health}}{{.State.Health.Status}}{{else}}missing{{end}}' "$restore_container")
|
||||
case "$health" in
|
||||
healthy) break ;;
|
||||
unhealthy)
|
||||
docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true
|
||||
echo "Isolated restore PostgreSQL became unhealthy." >&2
|
||||
return 1
|
||||
;;
|
||||
esac
|
||||
sleep 1
|
||||
done
|
||||
|
||||
# The PostGIS image initializes its requested database with PostGIS already
|
||||
# installed. Restore into a database created from template0 instead, so the
|
||||
# archive remains responsible for recreating its own extensions and schema.
|
||||
docker exec "$restore_container" createdb \
|
||||
--username postgres \
|
||||
--template template0 \
|
||||
restore_check
|
||||
|
||||
if ! docker exec --interactive "$restore_container" pg_restore \
|
||||
--username postgres \
|
||||
--dbname restore_check \
|
||||
--exit-on-error \
|
||||
--no-owner \
|
||||
--no-privileges <"$dump"; then
|
||||
docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true
|
||||
docker inspect "$restore_container" >"$evidence_dir/restore-container.json" 2>&1 || true
|
||||
echo "Isolated restore failed; container evidence was retained." >&2
|
||||
return 1
|
||||
fi
|
||||
|
||||
tables=$(docker exec "$restore_container" psql \
|
||||
--username postgres --dbname restore_check --tuples-only --no-align \
|
||||
--command="select count(*) from pg_tables where schemaname='public' and tablename <> 'spatial_ref_sys';")
|
||||
migrations=$(docker exec "$restore_container" psql \
|
||||
--username postgres --dbname restore_check --tuples-only --no-align \
|
||||
--command='select count(*) from schema_migrations;')
|
||||
postgis=$(docker exec "$restore_container" psql \
|
||||
--username postgres --dbname restore_check --tuples-only --no-align \
|
||||
--command='select PostGIS_Lib_Version();')
|
||||
|
||||
if [[ ! "$tables" =~ ^[1-9][0-9]*$ ]] || [[ ! "$migrations" =~ ^[1-9][0-9]*$ ]]; then
|
||||
echo "The isolated restore does not contain the expected application schema." >&2
|
||||
return 1
|
||||
fi
|
||||
|
||||
{
|
||||
printf 'snapshot_id=%s\n' "$selected_snapshot"
|
||||
printf 'application_tables=%s\n' "$tables"
|
||||
printf 'schema_migrations=%s\n' "$migrations"
|
||||
printf 'postgis_version=%s\n' "$postgis"
|
||||
printf 'verified_at=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
||||
} >"$evidence_dir/restore-drill.txt"
|
||||
chmod 600 "$evidence_dir/restore-drill.txt" "$evidence_dir/pg-restore-catalog.txt"
|
||||
|
||||
docker rm --force "$restore_container" >/dev/null
|
||||
restore_started=false
|
||||
docker volume rm "$restore_volume" >/dev/null
|
||||
restore_volume_created=false
|
||||
rm -f "$dump"
|
||||
}
|
||||
|
||||
if [[ "$action" == restore-drill ]]; then
|
||||
restore_snapshot "$snapshot_id"
|
||||
echo "Isolated restore drill passed for snapshot: $snapshot_id"
|
||||
printf 'Non-secret evidence: %s\n' "$evidence_dir"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
check_source
|
||||
|
||||
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
|
||||
"bash -s -- '$PRODUCTION_REMOTE_ROOT/.env' '$remote_dump' '$PRODUCTION_EXPECTED_ENVIRONMENT'" \
|
||||
<"$ROOT/scripts/backup-external-postgres.sh"
|
||||
remote_created=true
|
||||
|
||||
scp -p \
|
||||
"$PRODUCTION_SSH_TARGET:$remote_dump" \
|
||||
"$PRODUCTION_SSH_TARGET:$remote_dump.sha256" \
|
||||
"$PRODUCTION_SSH_TARGET:$remote_dump.metadata" \
|
||||
"$work_dir/"
|
||||
|
||||
dump="$work_dir/$(basename -- "$remote_dump")"
|
||||
checksum="$dump.sha256"
|
||||
metadata="$dump.metadata"
|
||||
|
||||
(
|
||||
cd "$work_dir"
|
||||
sha256sum --check "$(basename -- "$checksum")"
|
||||
) >"$evidence_dir/source-checksum.txt"
|
||||
pg_restore --list "$dump" >"$evidence_dir/source-catalog.txt"
|
||||
cp "$checksum" "$metadata" "$evidence_dir/"
|
||||
chmod 600 "$evidence_dir"/*
|
||||
|
||||
run_tag="$OFFSITE_RESTIC_TAG-$run_id"
|
||||
backup_json="$evidence_dir/restic-backup.jsonl"
|
||||
"$restic" backup \
|
||||
--json \
|
||||
--host "$OFFSITE_RESTIC_HOST" \
|
||||
--tag "$OFFSITE_RESTIC_TAG" \
|
||||
--tag "$run_tag" \
|
||||
--stdin-filename production.dump \
|
||||
--stdin-from-command -- cat "$dump" >"$backup_json"
|
||||
|
||||
snapshot_id=$(jq -r 'select(.message_type == "summary") | .snapshot_id // empty' "$backup_json" | tail -n 1)
|
||||
if [[ -z "$snapshot_id" ]]; then
|
||||
echo "Restic did not report a completed snapshot." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
snapshot_matches=$($restic snapshots --json --tag "$run_tag" | jq 'length')
|
||||
if [[ "$snapshot_matches" != 1 ]]; then
|
||||
echo "The run tag did not resolve to exactly one snapshot." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
"$restic" check --read-data >"$evidence_dir/restic-check.txt"
|
||||
restore_snapshot "$snapshot_id"
|
||||
|
||||
printf '%s\n' "$snapshot_id" >"$evidence_dir/snapshot-id.txt"
|
||||
chmod 600 "$evidence_dir"/*
|
||||
|
||||
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
|
||||
"rm -f -- '$remote_dump' '$remote_dump.sha256' '$remote_dump.metadata'"
|
||||
remote_created=false
|
||||
|
||||
echo "Encrypted off-server backup and isolated restore drill passed."
|
||||
printf 'Snapshot: %s\n' "$snapshot_id"
|
||||
printf 'Non-secret evidence: %s\n' "$evidence_dir"
|
||||
Loading…
Reference in New Issue
Block a user