Add verified offsite production recovery
This commit is contained in:
parent
c085620762
commit
ead8a67d0f
|
|
@ -564,6 +564,107 @@ failure paths on the observed workstation. A MinIO volume on that same
|
||||||
workstation is not an off-site backup and does not establish production RPO,
|
workstation is not an off-site backup and does not establish production RPO,
|
||||||
RTO, retention, capacity, key custody, object locking, or database HA.
|
RTO, retention, capacity, key custody, object locking, or database HA.
|
||||||
|
|
||||||
|
## Encrypted production backup to an independent host
|
||||||
|
|
||||||
|
The production workflow streams an external PostgreSQL 18 custom-format dump
|
||||||
|
through an encrypted Restic repository on an SSH target that must resolve to a
|
||||||
|
host other than production. It then runs `restic check --read-data` and restores
|
||||||
|
the newest snapshot into a uniquely named, temporary PostGIS container created
|
||||||
|
from `template0`. The source production database is never a restore target.
|
||||||
|
|
||||||
|
Bootstrap the checksum-pinned Restic binary and create the ignored local
|
||||||
|
operations configuration:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./scripts/bootstrap-restic.sh
|
||||||
|
./scripts/init-production-operations.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
The initializer creates only these ignored local files:
|
||||||
|
|
||||||
|
- `tmp/production-operations/backup.env`, mode `0600`;
|
||||||
|
- `tmp/production-operations/restic-password`, mode `0600`.
|
||||||
|
|
||||||
|
The generated configuration uses the reviewed `whoneedhelp` production SSH
|
||||||
|
target and the independent `buyvm-maya` SFTP target. Inspect the read-only plan
|
||||||
|
before repository initialization:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./scripts/production-offsite-backup.sh plan
|
||||||
|
```
|
||||||
|
|
||||||
|
Repository initialization is an external mutation and requires the exact
|
||||||
|
repository value printed by the plan:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
WNH_OFFSITE_BACKUP_INIT_CONFIRM='sftp:buyvm-maya:backups/who_need_help-production' \
|
||||||
|
./scripts/production-offsite-backup.sh init
|
||||||
|
```
|
||||||
|
|
||||||
|
The routine operations are:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./scripts/production-offsite-backup.sh run
|
||||||
|
./scripts/production-offsite-backup.sh check
|
||||||
|
./scripts/production-offsite-backup.sh restore-drill
|
||||||
|
```
|
||||||
|
|
||||||
|
`run` creates a fresh source dump, copies it into the Restic input, publishes
|
||||||
|
one encrypted snapshot only after the producer succeeds, performs a full-data
|
||||||
|
repository check, and runs the isolated restore drill. It removes its exact
|
||||||
|
temporary source dump, local scratch directory, Docker container, and Docker
|
||||||
|
volume on success, failure, or interrupt. Non-secret catalogs, checksums, and
|
||||||
|
restore observations remain under ignored
|
||||||
|
`output/production-operations/<run-id>/`.
|
||||||
|
|
||||||
|
Install the daily user-systemd timer only after one manual `run` succeeds:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./scripts/install-production-backup-timer.sh
|
||||||
|
systemctl --user list-timers --all who-need-help-production-backup.timer
|
||||||
|
systemctl --user status who-need-help-production-backup.service --no-pager
|
||||||
|
journalctl --user -u who-need-help-production-backup.service --no-pager
|
||||||
|
```
|
||||||
|
|
||||||
|
The timer uses `Persistent=true`, so systemd runs a missed calendar invocation
|
||||||
|
after the workstation and user manager return. A failed service invokes the
|
||||||
|
external monitor's backup-failure notification. It does not retry, delete
|
||||||
|
snapshots, or infer a retention policy.
|
||||||
|
|
||||||
|
The Restic password is required for every restore. Before relying on this as
|
||||||
|
the only recovery copy, place it in an operator-controlled password manager or
|
||||||
|
offline recovery record. Never print it into a terminal log, commit it, or
|
||||||
|
copy it into the backup repository. There is intentionally no automatic
|
||||||
|
`forget`/`prune` policy yet: retention, RPO, RTO, capacity, and key custody are
|
||||||
|
operator decisions, and the script does not invent them.
|
||||||
|
|
||||||
|
## Independent public readiness monitor
|
||||||
|
|
||||||
|
The external monitor runs on an SSH target that must resolve to a host other
|
||||||
|
than production. It requests only the public readiness endpoint and accepts
|
||||||
|
only HTTP 200 with the exact JSON object `{"status":"ready"}`. State is kept
|
||||||
|
in a mode-`0600` JSON file. Email is sent once when the state changes to down
|
||||||
|
and once when it recovers; repeated down checks do not send repeated alerts.
|
||||||
|
|
||||||
|
Install it from the production SMTP configuration without printing the SMTP
|
||||||
|
credential:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./scripts/install-production-external-monitor.sh
|
||||||
|
ssh buyvm-maya \
|
||||||
|
'systemctl --user status who-need-help-production-monitor.timer --no-pager'
|
||||||
|
ssh buyvm-maya \
|
||||||
|
'journalctl --user -u who-need-help-production-monitor.service --no-pager'
|
||||||
|
```
|
||||||
|
|
||||||
|
The default schedule is once per minute with a three-second HTTP timeout. Those
|
||||||
|
defaults match the current application container health timeout and were
|
||||||
|
installed only after public readiness requests from the selected monitor host
|
||||||
|
were observed to finish below one second. They are observations of the current
|
||||||
|
path, not universal capacity or availability guarantees. Re-running the
|
||||||
|
installer refreshes the exact script, mode-`0600` configuration, and user
|
||||||
|
units, then performs one check before enabling the timer.
|
||||||
|
|
||||||
## Local external-service boundary drill
|
## Local external-service boundary drill
|
||||||
|
|
||||||
Run the OAuth, SMTP, and provider-neutral push protocol checks without public
|
Run the OAuth, SMTP, and provider-neutral push protocol checks without public
|
||||||
|
|
|
||||||
|
|
@ -93,6 +93,22 @@ backup_path=$(printf '%s\n' "$backup_output" | sed -n 's/^Backup: //p')
|
||||||
Local backups and local MinIO drills verify mechanics but are not proof of
|
Local backups and local MinIO drills verify mechanics but are not proof of
|
||||||
off-site durability or a production recovery policy.
|
off-site durability or a production recovery policy.
|
||||||
|
|
||||||
|
The implemented external workflow can be rehearsed without restoring over the
|
||||||
|
source database:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./scripts/bootstrap-restic.sh
|
||||||
|
./scripts/production-offsite-backup.sh plan
|
||||||
|
./scripts/production-offsite-backup.sh run
|
||||||
|
systemctl --user status who-need-help-production-backup.service --no-pager
|
||||||
|
ssh buyvm-maya \
|
||||||
|
'systemctl --user status who-need-help-production-monitor.timer --no-pager'
|
||||||
|
```
|
||||||
|
|
||||||
|
Do not mark backup ownership complete merely because this command passes. The
|
||||||
|
operator must still store the Restic key independently and approve retention,
|
||||||
|
RPO, RTO, capacity, and responsible owners.
|
||||||
|
|
||||||
## 5. Promote with an explicit rollback point
|
## 5. Promote with an explicit rollback point
|
||||||
|
|
||||||
Use the release workflow in `docs/operations.md`; do not copy mutable source
|
Use the release workflow in `docs/operations.md`; do not copy mutable source
|
||||||
|
|
|
||||||
|
|
@ -1,8 +1,42 @@
|
||||||
# Who Need Help — implementation verification
|
# Who Need Help — implementation verification
|
||||||
|
|
||||||
Observed through 2026-07-28 in the local workspace. This report separates observed
|
Observed through 2026-08-03 in the local workspace. This report separates observed
|
||||||
results from product limits and unknown production properties.
|
results from product limits and unknown production properties.
|
||||||
|
|
||||||
|
## Independent production backup and monitoring on 2026-08-03
|
||||||
|
|
||||||
|
This verification changed only the independent backup repository, operator
|
||||||
|
user-systemd units, and independent monitor host. It did not deploy application
|
||||||
|
source, change production data, modify the hackathon test deployment, push Git,
|
||||||
|
or edit the Devpost submission.
|
||||||
|
|
||||||
|
- An encrypted Restic 0.19.1 repository was initialized at
|
||||||
|
`sftp:buyvm-maya:backups/who_need_help-production`. SSH resolution verified
|
||||||
|
that this target and the production target are different hosts. The Restic
|
||||||
|
password and workflow configuration remain only in ignored mode-`0600`
|
||||||
|
local files; password-manager/offline key custody is still an operator action.
|
||||||
|
- Two production snapshots exist. The installed user-systemd execution produced
|
||||||
|
snapshot `98d7d2e8e3e7c8dbf7908bbaceeb2acedf202a7e1dfd77c5398ba668a7981236`,
|
||||||
|
passed `restic check --read-data`, and restored into a temporary database
|
||||||
|
created from `template0`. The isolated restore observed 34 public application
|
||||||
|
tables, 24 applied schema migrations, and PostGIS 3.6.4. Non-secret evidence
|
||||||
|
is retained at
|
||||||
|
`output/production-operations/20260803T012727Z-1694894/`.
|
||||||
|
- That exact systemd service exited `0/SUCCESS` after 4 minutes 31.643 seconds
|
||||||
|
wall time, 4.115 seconds local CPU time, and a 78.2 MiB local memory peak.
|
||||||
|
These are observations of one run over the selected SFTP path, not a minimum
|
||||||
|
resource requirement or recovery objective. The scheduled source dump and
|
||||||
|
every run-scoped temporary restore container and volume were absent afterward.
|
||||||
|
- The daily backup timer is enabled with `Persistent=true`. Its failure unit
|
||||||
|
routes one alert through the independent monitor host. No retention deletion,
|
||||||
|
automatic pruning, or database rollback is configured.
|
||||||
|
- The independent BuyVM monitor performs the public readiness check once per
|
||||||
|
minute. The observed check returned HTTP 200 with the exact ready payload and
|
||||||
|
the service exited successfully. Its SMTP configuration and state are mode
|
||||||
|
`0600`. A one-time monitoring test message was accepted by Brevo; mailbox
|
||||||
|
receipt has not yet been independently observed, so end-to-end alert delivery
|
||||||
|
is not claimed complete.
|
||||||
|
|
||||||
## Production Android bundle replay on 2026-07-28
|
## Production Android bundle replay on 2026-07-28
|
||||||
|
|
||||||
The release tooling and application inputs below are exact local commit
|
The release tooling and application inputs below are exact local commit
|
||||||
|
|
|
||||||
53
scripts/bootstrap-restic.sh
Executable file
53
scripts/bootstrap-restic.sh
Executable file
|
|
@ -0,0 +1,53 @@
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
umask 077
|
||||||
|
|
||||||
|
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||||
|
target_dir="$ROOT/.tools/restic"
|
||||||
|
target="$target_dir/restic"
|
||||||
|
image=${WNH_BACKUP_TOOLS_IMAGE:-who-need-help:backup-tools}
|
||||||
|
expected='restic 0.19.1 compiled with go1.26.5'
|
||||||
|
|
||||||
|
if [[ -x "$target" ]] && [[ "$($target version)" == "$expected"* ]]; then
|
||||||
|
printf 'Pinned Restic is already available: %s\n' "$target"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
for command in docker install mktemp; do
|
||||||
|
command -v "$command" >/dev/null 2>&1 || {
|
||||||
|
echo "Required command is unavailable: $command" >&2
|
||||||
|
exit 2
|
||||||
|
}
|
||||||
|
done
|
||||||
|
|
||||||
|
mkdir -p "$target_dir"
|
||||||
|
chmod 700 "$ROOT/.tools" "$target_dir"
|
||||||
|
|
||||||
|
docker build --tag "$image" --file "$ROOT/Dockerfile.backup" "$ROOT"
|
||||||
|
|
||||||
|
container=$(docker create "$image" restic version)
|
||||||
|
partial=$(mktemp "$target_dir/restic.XXXXXX")
|
||||||
|
|
||||||
|
cleanup() {
|
||||||
|
docker rm --force "$container" >/dev/null 2>&1 || true
|
||||||
|
rm -f "$partial"
|
||||||
|
}
|
||||||
|
trap cleanup EXIT HUP INT TERM
|
||||||
|
|
||||||
|
docker cp "$container:/usr/local/bin/restic" "$partial"
|
||||||
|
chmod 700 "$partial"
|
||||||
|
|
||||||
|
actual=$($partial version)
|
||||||
|
if [[ "$actual" != "$expected"* ]]; then
|
||||||
|
printf 'Unexpected Restic build: %s\n' "$actual" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
mv "$partial" "$target"
|
||||||
|
partial=
|
||||||
|
|
||||||
|
trap - EXIT HUP INT TERM
|
||||||
|
docker rm --force "$container" >/dev/null
|
||||||
|
|
||||||
|
printf 'Installed pinned Restic: %s\n' "$target"
|
||||||
|
printf 'Version: %s\n' "$actual"
|
||||||
67
scripts/init-production-operations.sh
Executable file
67
scripts/init-production-operations.sh
Executable file
|
|
@ -0,0 +1,67 @@
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
umask 077
|
||||||
|
|
||||||
|
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||||
|
runtime_dir="$ROOT/tmp/production-operations"
|
||||||
|
config=${1:-"$runtime_dir/backup.env"}
|
||||||
|
|
||||||
|
if [[ "$config" != /* ]]; then
|
||||||
|
config="$ROOT/$config"
|
||||||
|
fi
|
||||||
|
|
||||||
|
password_file="$runtime_dir/restic-password"
|
||||||
|
|
||||||
|
for command in openssl ssh; do
|
||||||
|
command -v "$command" >/dev/null 2>&1 || {
|
||||||
|
echo "Required command is unavailable: $command" >&2
|
||||||
|
exit 2
|
||||||
|
}
|
||||||
|
done
|
||||||
|
|
||||||
|
if [[ -e "$config" || -e "$password_file" ]]; then
|
||||||
|
echo "Refusing to replace existing production operations configuration." >&2
|
||||||
|
printf 'Config: %s\nPassword file: %s\n' "$config" "$password_file" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
mkdir -p "$runtime_dir"
|
||||||
|
chmod 700 "$ROOT/tmp" "$runtime_dir"
|
||||||
|
|
||||||
|
source_target=whoneedhelp
|
||||||
|
repository_target=buyvm-maya
|
||||||
|
|
||||||
|
source_host=$(ssh -G "$source_target" | awk '$1 == "hostname" {print $2; exit}')
|
||||||
|
repository_host=$(ssh -G "$repository_target" | awk '$1 == "hostname" {print $2; exit}')
|
||||||
|
|
||||||
|
if [[ -z "$source_host" || -z "$repository_host" ]]; then
|
||||||
|
echo "Could not resolve both SSH targets." >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ "$source_host" == "$repository_host" ]]; then
|
||||||
|
echo "The backup repository must not resolve to the production host." >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
openssl rand -base64 48 | tr -d '\n' >"$password_file"
|
||||||
|
printf '\n' >>"$password_file"
|
||||||
|
chmod 600 "$password_file"
|
||||||
|
|
||||||
|
cat >"$config" <<EOF
|
||||||
|
PRODUCTION_SSH_TARGET=$source_target
|
||||||
|
PRODUCTION_REMOTE_ROOT=/srv/who_need_help-production
|
||||||
|
PRODUCTION_EXPECTED_ENVIRONMENT=production
|
||||||
|
OFFSITE_RESTIC_REPOSITORY=sftp:$repository_target:backups/who_need_help-production
|
||||||
|
OFFSITE_RESTIC_PASSWORD_FILE=$password_file
|
||||||
|
OFFSITE_RESTIC_HOST=who-need-help-production
|
||||||
|
OFFSITE_RESTIC_TAG=who-need-help-production
|
||||||
|
BACKUP_ON_CALENDAR=daily
|
||||||
|
EOF
|
||||||
|
chmod 600 "$config"
|
||||||
|
|
||||||
|
printf 'Created mode-0600 operations configuration: %s\n' "$config"
|
||||||
|
printf 'Created mode-0600 Restic key file: %s\n' "$password_file"
|
||||||
|
printf 'Production resolves to: %s\n' "$source_host"
|
||||||
|
printf 'Encrypted repository resolves to a different host: %s\n' "$repository_host"
|
||||||
|
echo "The Restic key is required for every restore. Copy it to an operator-controlled password manager before relying on this backup as the only recovery copy."
|
||||||
79
scripts/install-production-backup-timer.sh
Executable file
79
scripts/install-production-backup-timer.sh
Executable file
|
|
@ -0,0 +1,79 @@
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
umask 077
|
||||||
|
|
||||||
|
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||||
|
config=${1:-"$ROOT/tmp/production-operations/backup.env"}
|
||||||
|
unit_dir=${SYSTEMD_USER_UNIT_DIR:-"$HOME/.config/systemd/user"}
|
||||||
|
|
||||||
|
if [[ "$config" != /* ]]; then
|
||||||
|
config="$ROOT/$config"
|
||||||
|
fi
|
||||||
|
if [[ ! -f "$config" || "$(stat -c '%a' "$config")" != 600 ]]; then
|
||||||
|
echo "The backup configuration must exist with mode 0600: $config" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
# shellcheck source=/dev/null
|
||||||
|
source "$config"
|
||||||
|
: "${BACKUP_ON_CALENDAR:?Set BACKUP_ON_CALENDAR}"
|
||||||
|
|
||||||
|
case "$BACKUP_ON_CALENDAR" in
|
||||||
|
*$'\n'* | *$'\r'*) echo "BACKUP_ON_CALENDAR must be one line." >&2; exit 2 ;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya}
|
||||||
|
remote_monitor=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help}/production-external-monitor.py
|
||||||
|
remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json}
|
||||||
|
ssh_path=$(command -v ssh)
|
||||||
|
|
||||||
|
mkdir -p "$unit_dir"
|
||||||
|
chmod 700 "$HOME/.config" "$HOME/.config/systemd" "$unit_dir"
|
||||||
|
|
||||||
|
cat >"$unit_dir/who-need-help-production-backup.service" <<EOF
|
||||||
|
[Unit]
|
||||||
|
Description=Who Need Help encrypted off-server production backup and restore drill
|
||||||
|
Wants=network-online.target
|
||||||
|
After=network-online.target
|
||||||
|
OnFailure=who-need-help-production-backup-alert.service
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=oneshot
|
||||||
|
WorkingDirectory=$ROOT
|
||||||
|
ExecStart=$ROOT/scripts/production-offsite-backup.sh run $config
|
||||||
|
EOF
|
||||||
|
|
||||||
|
cat >"$unit_dir/who-need-help-production-backup-alert.service" <<EOF
|
||||||
|
[Unit]
|
||||||
|
Description=Notify operators that the Who Need Help production backup failed
|
||||||
|
Wants=network-online.target
|
||||||
|
After=network-online.target
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=oneshot
|
||||||
|
ExecStart=$ssh_path -o BatchMode=yes $monitor_target /usr/bin/python3 $remote_monitor --config $remote_config notify-backup-failure --unit who-need-help-production-backup.service
|
||||||
|
EOF
|
||||||
|
|
||||||
|
cat >"$unit_dir/who-need-help-production-backup.timer" <<EOF
|
||||||
|
[Unit]
|
||||||
|
Description=Schedule Who Need Help encrypted production backups
|
||||||
|
|
||||||
|
[Timer]
|
||||||
|
OnCalendar=$BACKUP_ON_CALENDAR
|
||||||
|
Persistent=true
|
||||||
|
Unit=who-need-help-production-backup.service
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=timers.target
|
||||||
|
EOF
|
||||||
|
|
||||||
|
chmod 600 \
|
||||||
|
"$unit_dir/who-need-help-production-backup.service" \
|
||||||
|
"$unit_dir/who-need-help-production-backup-alert.service" \
|
||||||
|
"$unit_dir/who-need-help-production-backup.timer"
|
||||||
|
|
||||||
|
systemctl --user daemon-reload
|
||||||
|
systemctl --user enable --now who-need-help-production-backup.timer
|
||||||
|
|
||||||
|
echo "Production backup timer installed without starting or changing the application."
|
||||||
|
systemctl --user list-timers --all --no-pager who-need-help-production-backup.timer
|
||||||
137
scripts/install-production-external-monitor.sh
Executable file
137
scripts/install-production-external-monitor.sh
Executable file
|
|
@ -0,0 +1,137 @@
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
umask 077
|
||||||
|
|
||||||
|
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||||
|
source_target=${PRODUCTION_SSH_TARGET:-whoneedhelp}
|
||||||
|
monitor_target=${MONITOR_SSH_TARGET:-buyvm-maya}
|
||||||
|
production_env=${PRODUCTION_ENV_PATH:-/srv/who_need_help-production/.env}
|
||||||
|
remote_root=${MONITOR_REMOTE_ROOT:-/home/simple/.local/lib/who-need-help}
|
||||||
|
remote_config=${MONITOR_REMOTE_CONFIG:-/home/simple/.config/who-need-help/monitor.json}
|
||||||
|
remote_state=${MONITOR_REMOTE_STATE:-/home/simple/.local/state/who-need-help/monitor.json}
|
||||||
|
monitor_url=${MONITOR_URL:-https://whoneedhelp.com/healthz/ready}
|
||||||
|
monitor_calendar=${MONITOR_ON_CALENDAR:-'*:0/1'}
|
||||||
|
health_timeout=${MONITOR_HEALTH_TIMEOUT_SECONDS:-3}
|
||||||
|
|
||||||
|
for command in mktemp scp ssh; do
|
||||||
|
command -v "$command" >/dev/null 2>&1 || {
|
||||||
|
echo "Required command is unavailable: $command" >&2
|
||||||
|
exit 2
|
||||||
|
}
|
||||||
|
done
|
||||||
|
|
||||||
|
source_host=$(ssh -G "$source_target" | awk '$1 == "hostname" {print $2; exit}')
|
||||||
|
monitor_host=$(ssh -G "$monitor_target" | awk '$1 == "hostname" {print $2; exit}')
|
||||||
|
if [[ -z "$source_host" || -z "$monitor_host" || "$source_host" == "$monitor_host" ]]; then
|
||||||
|
echo "The external monitor must resolve and run on a host other than production." >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
case "$health_timeout" in
|
||||||
|
'' | *[!0-9]*) echo "MONITOR_HEALTH_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;;
|
||||||
|
0) echo "MONITOR_HEALTH_TIMEOUT_SECONDS must be a positive integer." >&2; exit 2 ;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
work_dir=$(mktemp -d "$ROOT/tmp/production-operations/monitor-install.XXXXXX")
|
||||||
|
config="$work_dir/monitor.json"
|
||||||
|
service="$work_dir/who-need-help-production-monitor.service"
|
||||||
|
timer="$work_dir/who-need-help-production-monitor.timer"
|
||||||
|
|
||||||
|
cleanup() {
|
||||||
|
rm -rf "$work_dir"
|
||||||
|
}
|
||||||
|
trap cleanup EXIT HUP INT TERM
|
||||||
|
|
||||||
|
ssh -o BatchMode=yes "$source_target" \
|
||||||
|
"python3 - '$production_env' '$monitor_url' '$health_timeout'" >"$config" <<'PY'
|
||||||
|
import json
|
||||||
|
import shlex
|
||||||
|
import sys
|
||||||
|
|
||||||
|
path, health_url, timeout = sys.argv[1:]
|
||||||
|
wanted = {
|
||||||
|
"SMTP_RELAY",
|
||||||
|
"SMTP_PORT",
|
||||||
|
"SMTP_USERNAME",
|
||||||
|
"SMTP_PASSWORD",
|
||||||
|
"SMTP_TLS",
|
||||||
|
"SMTP_SSL",
|
||||||
|
"EMAIL_FROM_ADDRESS",
|
||||||
|
"EMAIL_FROM_NAME",
|
||||||
|
"SUPPORT_INBOX_ADDRESS",
|
||||||
|
}
|
||||||
|
values = {}
|
||||||
|
with open(path, encoding="utf-8") as handle:
|
||||||
|
for raw_line in handle:
|
||||||
|
line = raw_line.strip()
|
||||||
|
if not line or line.startswith("#") or "=" not in line:
|
||||||
|
continue
|
||||||
|
key, value = line.split("=", 1)
|
||||||
|
if key not in wanted:
|
||||||
|
continue
|
||||||
|
parsed = shlex.split(value, comments=False, posix=True)
|
||||||
|
values[key] = parsed[0] if parsed else ""
|
||||||
|
|
||||||
|
missing = sorted(key for key in wanted if not values.get(key))
|
||||||
|
if missing:
|
||||||
|
raise SystemExit("Missing production mail settings: " + ", ".join(missing))
|
||||||
|
|
||||||
|
payload = {
|
||||||
|
"health_timeout_seconds": int(timeout),
|
||||||
|
"health_url": health_url,
|
||||||
|
"smtp": {
|
||||||
|
"from_address": values["EMAIL_FROM_ADDRESS"],
|
||||||
|
"from_name": values["EMAIL_FROM_NAME"],
|
||||||
|
"implicit_ssl": values["SMTP_SSL"].lower() == "true",
|
||||||
|
"password": values["SMTP_PASSWORD"],
|
||||||
|
"port": int(values["SMTP_PORT"]),
|
||||||
|
"recipient": values["SUPPORT_INBOX_ADDRESS"],
|
||||||
|
"relay": values["SMTP_RELAY"],
|
||||||
|
"starttls": values["SMTP_TLS"].lower() == "always",
|
||||||
|
"username": values["SMTP_USERNAME"],
|
||||||
|
},
|
||||||
|
}
|
||||||
|
json.dump(payload, sys.stdout, ensure_ascii=False, indent=2, sort_keys=True)
|
||||||
|
sys.stdout.write("\n")
|
||||||
|
PY
|
||||||
|
chmod 600 "$config"
|
||||||
|
|
||||||
|
cat >"$service" <<EOF
|
||||||
|
[Unit]
|
||||||
|
Description=Who Need Help independent production readiness monitor
|
||||||
|
Wants=network-online.target
|
||||||
|
After=network-online.target
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=oneshot
|
||||||
|
ExecStart=/usr/bin/python3 $remote_root/production-external-monitor.py --config $remote_config --state $remote_state check
|
||||||
|
EOF
|
||||||
|
|
||||||
|
cat >"$timer" <<EOF
|
||||||
|
[Unit]
|
||||||
|
Description=Run the Who Need Help independent production readiness monitor
|
||||||
|
|
||||||
|
[Timer]
|
||||||
|
OnCalendar=$monitor_calendar
|
||||||
|
Persistent=true
|
||||||
|
Unit=who-need-help-production-monitor.service
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=timers.target
|
||||||
|
EOF
|
||||||
|
|
||||||
|
ssh -o BatchMode=yes "$monitor_target" \
|
||||||
|
"install -d -m 700 '$remote_root' /home/simple/.config/who-need-help /home/simple/.local/state/who-need-help /home/simple/.config/systemd/user"
|
||||||
|
scp -q "$ROOT/scripts/production-external-monitor.py" \
|
||||||
|
"$monitor_target:$remote_root/production-external-monitor.py"
|
||||||
|
scp -q "$config" "$monitor_target:$remote_config"
|
||||||
|
scp -q "$service" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.service"
|
||||||
|
scp -q "$timer" "$monitor_target:/home/simple/.config/systemd/user/who-need-help-production-monitor.timer"
|
||||||
|
|
||||||
|
ssh -o BatchMode=yes "$monitor_target" \
|
||||||
|
"chmod 700 '$remote_root/production-external-monitor.py'; chmod 600 '$remote_config' /home/simple/.config/systemd/user/who-need-help-production-monitor.service /home/simple/.config/systemd/user/who-need-help-production-monitor.timer; systemctl --user daemon-reload; systemctl --user start who-need-help-production-monitor.service; systemctl --user enable --now who-need-help-production-monitor.timer"
|
||||||
|
|
||||||
|
printf 'External monitor installed on %s (%s).\n' "$monitor_target" "$monitor_host"
|
||||||
|
printf 'Health URL: %s\n' "$monitor_url"
|
||||||
|
printf 'Schedule: %s; request timeout: %ss.\n' "$monitor_calendar" "$health_timeout"
|
||||||
|
echo "The SMTP credential is stored only in a mode-0600 configuration on the external monitor host."
|
||||||
231
scripts/production-external-monitor.py
Executable file
231
scripts/production-external-monitor.py
Executable file
|
|
@ -0,0 +1,231 @@
|
||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Stateful external production health and operations notifications."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import smtplib
|
||||||
|
import ssl
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import urllib.error
|
||||||
|
import urllib.request
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from email.message import EmailMessage
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
def utc_now() -> str:
|
||||||
|
return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z")
|
||||||
|
|
||||||
|
|
||||||
|
def read_json(path: Path) -> dict[str, Any]:
|
||||||
|
with path.open("r", encoding="utf-8") as handle:
|
||||||
|
value = json.load(handle)
|
||||||
|
if not isinstance(value, dict):
|
||||||
|
raise ValueError(f"Expected a JSON object in {path}")
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
def write_json_atomic(path: Path, value: dict[str, Any]) -> None:
|
||||||
|
path.parent.mkdir(mode=0o700, parents=True, exist_ok=True)
|
||||||
|
descriptor, temporary_name = tempfile.mkstemp(dir=path.parent, prefix=f".{path.name}.")
|
||||||
|
temporary = Path(temporary_name)
|
||||||
|
try:
|
||||||
|
with os.fdopen(descriptor, "w", encoding="utf-8") as handle:
|
||||||
|
json.dump(value, handle, ensure_ascii=False, indent=2, sort_keys=True)
|
||||||
|
handle.write("\n")
|
||||||
|
temporary.chmod(0o600)
|
||||||
|
temporary.replace(path)
|
||||||
|
finally:
|
||||||
|
temporary.unlink(missing_ok=True)
|
||||||
|
|
||||||
|
|
||||||
|
def required_text(mapping: dict[str, Any], name: str) -> str:
|
||||||
|
value = mapping.get(name)
|
||||||
|
if not isinstance(value, str) or not value.strip():
|
||||||
|
raise ValueError(f"Missing non-empty configuration value: {name}")
|
||||||
|
return value.strip()
|
||||||
|
|
||||||
|
|
||||||
|
def send_message(config: dict[str, Any], subject: str, body: str) -> None:
|
||||||
|
smtp = config.get("smtp")
|
||||||
|
if not isinstance(smtp, dict):
|
||||||
|
raise ValueError("Missing SMTP configuration")
|
||||||
|
|
||||||
|
relay = required_text(smtp, "relay")
|
||||||
|
port = int(smtp.get("port"))
|
||||||
|
username = required_text(smtp, "username")
|
||||||
|
password = required_text(smtp, "password")
|
||||||
|
from_address = required_text(smtp, "from_address")
|
||||||
|
from_name = required_text(smtp, "from_name")
|
||||||
|
recipient = required_text(smtp, "recipient")
|
||||||
|
implicit_ssl = bool(smtp.get("implicit_ssl", False))
|
||||||
|
starttls = bool(smtp.get("starttls", True))
|
||||||
|
|
||||||
|
message = EmailMessage()
|
||||||
|
message["From"] = f"{from_name} <{from_address}>"
|
||||||
|
message["To"] = recipient
|
||||||
|
message["Subject"] = subject
|
||||||
|
message.set_content(body)
|
||||||
|
|
||||||
|
context = ssl.create_default_context()
|
||||||
|
if implicit_ssl:
|
||||||
|
client_context = smtplib.SMTP_SSL(relay, port, timeout=30, context=context)
|
||||||
|
else:
|
||||||
|
client_context = smtplib.SMTP(relay, port, timeout=30)
|
||||||
|
|
||||||
|
with client_context as client:
|
||||||
|
if not implicit_ssl:
|
||||||
|
client.ehlo()
|
||||||
|
if starttls:
|
||||||
|
client.starttls(context=context)
|
||||||
|
client.ehlo()
|
||||||
|
client.login(username, password)
|
||||||
|
client.send_message(message)
|
||||||
|
|
||||||
|
|
||||||
|
def check_health(config: dict[str, Any]) -> tuple[str, str]:
|
||||||
|
url = required_text(config, "health_url")
|
||||||
|
timeout = float(config.get("health_timeout_seconds"))
|
||||||
|
request = urllib.request.Request(
|
||||||
|
url,
|
||||||
|
headers={"User-Agent": "WhoNeedHelp-External-Monitor/1.0"},
|
||||||
|
)
|
||||||
|
|
||||||
|
try:
|
||||||
|
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||||
|
status = response.status
|
||||||
|
payload = response.read(4096)
|
||||||
|
parsed = json.loads(payload.decode("utf-8"))
|
||||||
|
if status == 200 and parsed == {"status": "ready"}:
|
||||||
|
return "up", f"HTTP {status}; ready payload matched"
|
||||||
|
return "down", f"HTTP {status}; unexpected readiness payload"
|
||||||
|
except (OSError, ValueError, urllib.error.URLError) as error:
|
||||||
|
return "down", f"{type(error).__name__}: {str(error)[:500]}"
|
||||||
|
|
||||||
|
|
||||||
|
def monitor(config_path: Path, state_path: Path) -> int:
|
||||||
|
config = read_json(config_path)
|
||||||
|
previous: dict[str, Any] = {}
|
||||||
|
if state_path.exists():
|
||||||
|
previous = read_json(state_path)
|
||||||
|
|
||||||
|
status, detail = check_health(config)
|
||||||
|
previous_status = previous.get("status")
|
||||||
|
|
||||||
|
if status != previous_status:
|
||||||
|
if status == "down":
|
||||||
|
send_message(
|
||||||
|
config,
|
||||||
|
"[Who Need Help] Production readiness is DOWN",
|
||||||
|
"\n".join(
|
||||||
|
[
|
||||||
|
"The independent production readiness check failed.",
|
||||||
|
f"URL: {required_text(config, 'health_url')}",
|
||||||
|
f"Observed at: {utc_now()}",
|
||||||
|
f"Result: {detail}",
|
||||||
|
"This alert is sent once per state transition.",
|
||||||
|
]
|
||||||
|
),
|
||||||
|
)
|
||||||
|
elif previous_status == "down":
|
||||||
|
send_message(
|
||||||
|
config,
|
||||||
|
"[Who Need Help] Production readiness recovered",
|
||||||
|
"\n".join(
|
||||||
|
[
|
||||||
|
"The independent production readiness check recovered.",
|
||||||
|
f"URL: {required_text(config, 'health_url')}",
|
||||||
|
f"Observed at: {utc_now()}",
|
||||||
|
f"Result: {detail}",
|
||||||
|
]
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
write_json_atomic(
|
||||||
|
state_path,
|
||||||
|
{
|
||||||
|
"checked_at": utc_now(),
|
||||||
|
"detail": detail,
|
||||||
|
"status": status,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
print(json.dumps({"status": status, "detail": detail}, sort_keys=True))
|
||||||
|
return 0 if status == "up" else 1
|
||||||
|
|
||||||
|
|
||||||
|
def notify_backup_failure(config_path: Path, unit: str) -> int:
|
||||||
|
config = read_json(config_path)
|
||||||
|
send_message(
|
||||||
|
config,
|
||||||
|
"[Who Need Help] Production backup or restore verification failed",
|
||||||
|
"\n".join(
|
||||||
|
[
|
||||||
|
"The scheduled encrypted production backup did not finish successfully.",
|
||||||
|
f"Unit: {unit}",
|
||||||
|
f"Observed at: {utc_now()}",
|
||||||
|
"Inspect the local user-systemd journal and do not treat the newest snapshot as verified until a restore drill passes.",
|
||||||
|
]
|
||||||
|
),
|
||||||
|
)
|
||||||
|
print("Backup failure notification sent.")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
def send_test_notification(config_path: Path) -> int:
|
||||||
|
config = read_json(config_path)
|
||||||
|
send_message(
|
||||||
|
config,
|
||||||
|
"[Who Need Help] Operations monitoring test",
|
||||||
|
"\n".join(
|
||||||
|
[
|
||||||
|
"This is a one-time delivery verification for the independent production monitor.",
|
||||||
|
f"Health URL: {required_text(config, 'health_url')}",
|
||||||
|
f"Sent at: {utc_now()}",
|
||||||
|
"No production incident was detected and no application data was changed.",
|
||||||
|
]
|
||||||
|
),
|
||||||
|
)
|
||||||
|
print("Operations monitoring test notification sent.")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser()
|
||||||
|
parser.add_argument(
|
||||||
|
"--config",
|
||||||
|
type=Path,
|
||||||
|
default=Path.home() / ".config/who-need-help/monitor.json",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--state",
|
||||||
|
type=Path,
|
||||||
|
default=Path.home() / ".local/state/who-need-help/monitor.json",
|
||||||
|
)
|
||||||
|
subparsers = parser.add_subparsers(dest="action", required=True)
|
||||||
|
subparsers.add_parser("check")
|
||||||
|
subparsers.add_parser("send-test-notification")
|
||||||
|
backup = subparsers.add_parser("notify-backup-failure")
|
||||||
|
backup.add_argument("--unit", required=True)
|
||||||
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
args = parse_args()
|
||||||
|
try:
|
||||||
|
if args.action == "check":
|
||||||
|
return monitor(args.config, args.state)
|
||||||
|
if args.action == "send-test-notification":
|
||||||
|
return send_test_notification(args.config)
|
||||||
|
return notify_backup_failure(args.config, args.unit)
|
||||||
|
except (OSError, ValueError, smtplib.SMTPException, json.JSONDecodeError) as error:
|
||||||
|
print(f"Operations monitor failed: {type(error).__name__}: {error}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
384
scripts/production-offsite-backup.sh
Executable file
384
scripts/production-offsite-backup.sh
Executable file
|
|
@ -0,0 +1,384 @@
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
umask 077
|
||||||
|
|
||||||
|
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
|
||||||
|
action=${1:-plan}
|
||||||
|
config=${2:-"$ROOT/tmp/production-operations/backup.env"}
|
||||||
|
restic="$ROOT/.tools/restic/restic"
|
||||||
|
restore_image='postgis/postgis:18-3.6-alpine@sha256:05d68c7f0f19b9aa0bf7c4a2049b2e8b38b44a63116392b95726a4c913766cf6'
|
||||||
|
|
||||||
|
case "$action" in
|
||||||
|
plan | init | run | check | restore-drill) ;;
|
||||||
|
*)
|
||||||
|
echo "Usage: $0 [plan|init|run|check|restore-drill] [CONFIG]" >&2
|
||||||
|
exit 2
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
if [[ "$config" != /* ]]; then
|
||||||
|
config="$ROOT/$config"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ ! -f "$config" ]]; then
|
||||||
|
echo "Operations configuration does not exist: $config" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ "$(stat -c '%a' "$config")" != 600 ]]; then
|
||||||
|
echo "Operations configuration must have mode 0600: $config" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
# shellcheck source=/dev/null
|
||||||
|
source "$config"
|
||||||
|
|
||||||
|
: "${PRODUCTION_SSH_TARGET:?Set PRODUCTION_SSH_TARGET}"
|
||||||
|
: "${PRODUCTION_REMOTE_ROOT:?Set PRODUCTION_REMOTE_ROOT}"
|
||||||
|
: "${PRODUCTION_EXPECTED_ENVIRONMENT:?Set PRODUCTION_EXPECTED_ENVIRONMENT}"
|
||||||
|
: "${OFFSITE_RESTIC_REPOSITORY:?Set OFFSITE_RESTIC_REPOSITORY}"
|
||||||
|
: "${OFFSITE_RESTIC_PASSWORD_FILE:?Set OFFSITE_RESTIC_PASSWORD_FILE}"
|
||||||
|
: "${OFFSITE_RESTIC_HOST:?Set OFFSITE_RESTIC_HOST}"
|
||||||
|
: "${OFFSITE_RESTIC_TAG:?Set OFFSITE_RESTIC_TAG}"
|
||||||
|
|
||||||
|
if [[ "$PRODUCTION_EXPECTED_ENVIRONMENT" != production ]]; then
|
||||||
|
echo "The source environment must be exactly production." >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
case "$PRODUCTION_REMOTE_ROOT" in
|
||||||
|
/srv/who_need_help-production) ;;
|
||||||
|
*)
|
||||||
|
echo "The production root is outside the reviewed deployment path." >&2
|
||||||
|
exit 2
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
case "$OFFSITE_RESTIC_REPOSITORY" in
|
||||||
|
sftp:*) ;;
|
||||||
|
*)
|
||||||
|
echo "This workflow currently accepts only an encrypted Restic SFTP repository." >&2
|
||||||
|
exit 2
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
if [[ ! -f "$OFFSITE_RESTIC_PASSWORD_FILE" ]] ||
|
||||||
|
[[ "$(stat -c '%a' "$OFFSITE_RESTIC_PASSWORD_FILE")" != 600 ]]; then
|
||||||
|
echo "The Restic password file must exist with mode 0600." >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ ! -x "$restic" ]]; then
|
||||||
|
echo "Pinned Restic is unavailable. Run ./scripts/bootstrap-restic.sh first." >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
for command in docker flock grep jq openssl pg_restore scp sha256sum ssh; do
|
||||||
|
command -v "$command" >/dev/null 2>&1 || {
|
||||||
|
echo "Required command is unavailable: $command" >&2
|
||||||
|
exit 2
|
||||||
|
}
|
||||||
|
done
|
||||||
|
|
||||||
|
export RESTIC_REPOSITORY="$OFFSITE_RESTIC_REPOSITORY"
|
||||||
|
export RESTIC_PASSWORD_FILE="$OFFSITE_RESTIC_PASSWORD_FILE"
|
||||||
|
export RESTIC_CACHE_DIR="$ROOT/tmp/production-operations/restic-cache"
|
||||||
|
|
||||||
|
mkdir -p "$RESTIC_CACHE_DIR"
|
||||||
|
chmod 700 "$ROOT/tmp" "$ROOT/tmp/production-operations" "$RESTIC_CACHE_DIR"
|
||||||
|
|
||||||
|
source_host=$(ssh -G "$PRODUCTION_SSH_TARGET" | awk '$1 == "hostname" {print $2; exit}')
|
||||||
|
repository_alias=${OFFSITE_RESTIC_REPOSITORY#sftp:}
|
||||||
|
repository_alias=${repository_alias%%:*}
|
||||||
|
repository_host=$(ssh -G "$repository_alias" | awk '$1 == "hostname" {print $2; exit}')
|
||||||
|
|
||||||
|
if [[ -z "$source_host" || -z "$repository_host" ]] ||
|
||||||
|
[[ "$source_host" == "$repository_host" ]]; then
|
||||||
|
echo "The repository host must resolve and differ from production." >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
check_source() {
|
||||||
|
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
|
||||||
|
"bash -s -- '$PRODUCTION_REMOTE_ROOT/.env' --check-only '$PRODUCTION_EXPECTED_ENVIRONMENT'" \
|
||||||
|
<"$ROOT/scripts/backup-external-postgres.sh"
|
||||||
|
}
|
||||||
|
|
||||||
|
repo_initialized() {
|
||||||
|
"$restic" snapshots --json >/dev/null 2>&1
|
||||||
|
}
|
||||||
|
|
||||||
|
printf 'Production SSH target: %s (%s)\n' "$PRODUCTION_SSH_TARGET" "$source_host"
|
||||||
|
printf 'Production root: %s\n' "$PRODUCTION_REMOTE_ROOT"
|
||||||
|
printf 'Encrypted repository: %s (%s)\n' "$OFFSITE_RESTIC_REPOSITORY" "$repository_host"
|
||||||
|
printf 'Restic: %s\n' "$($restic version)"
|
||||||
|
|
||||||
|
if [[ "$action" == plan ]]; then
|
||||||
|
check_source
|
||||||
|
if repo_initialized; then
|
||||||
|
echo "Encrypted repository opened successfully."
|
||||||
|
else
|
||||||
|
echo "Encrypted repository is not initialized or cannot be opened."
|
||||||
|
fi
|
||||||
|
echo "Plan completed without changing production or the repository."
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ "$action" == init ]]; then
|
||||||
|
confirmation="$OFFSITE_RESTIC_REPOSITORY"
|
||||||
|
if [[ "${WNH_OFFSITE_BACKUP_INIT_CONFIRM:-}" != "$confirmation" ]]; then
|
||||||
|
echo "Repository initialization requires exact confirmation:" >&2
|
||||||
|
echo "WNH_OFFSITE_BACKUP_INIT_CONFIRM=$confirmation $0 init '$config'" >&2
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
|
|
||||||
|
check_source
|
||||||
|
if repo_initialized; then
|
||||||
|
echo "Encrypted repository is already initialized; no change was made."
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
"$restic" init
|
||||||
|
"$restic" check
|
||||||
|
echo "Encrypted off-server Restic repository initialized and opened successfully."
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
repo_initialized || {
|
||||||
|
echo "Encrypted repository cannot be opened. Initialize it first." >&2
|
||||||
|
exit 2
|
||||||
|
}
|
||||||
|
|
||||||
|
if [[ "$action" == check ]]; then
|
||||||
|
"$restic" check --read-data
|
||||||
|
echo "Encrypted repository full-data check passed."
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [[ "$action" == restore-drill ]]; then
|
||||||
|
snapshot_id=$($restic snapshots --json --host "$OFFSITE_RESTIC_HOST" \
|
||||||
|
--tag "$OFFSITE_RESTIC_TAG" --latest 1 | jq -r '.[0].short_id // empty')
|
||||||
|
if [[ -z "$snapshot_id" ]]; then
|
||||||
|
echo "No production backup snapshot is available for a restore drill." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
snapshot_id=
|
||||||
|
fi
|
||||||
|
|
||||||
|
lock_file="$ROOT/tmp/production-operations/backup.lock"
|
||||||
|
exec 9>"$lock_file"
|
||||||
|
flock --nonblock 9 || {
|
||||||
|
echo "Another production backup or restore drill is already running." >&2
|
||||||
|
exit 1
|
||||||
|
}
|
||||||
|
|
||||||
|
run_id=$(date -u +%Y%m%dT%H%M%SZ)-$$
|
||||||
|
work_dir=$(mktemp -d "$ROOT/tmp/production-operations/run-$run_id.XXXXXX")
|
||||||
|
evidence_dir="$ROOT/output/production-operations/$run_id"
|
||||||
|
remote_dump="$PRODUCTION_REMOTE_ROOT/output/backups/production/scheduled-$run_id.dump"
|
||||||
|
remote_created=false
|
||||||
|
restore_container="wnh-production-restore-$run_id"
|
||||||
|
restore_volume="wnh_production_restore_${run_id//[^a-zA-Z0-9]/_}"
|
||||||
|
restore_started=false
|
||||||
|
restore_volume_created=false
|
||||||
|
|
||||||
|
mkdir -p "$evidence_dir"
|
||||||
|
chmod 700 "$ROOT/output" "$ROOT/output/production-operations" "$evidence_dir" "$work_dir"
|
||||||
|
|
||||||
|
cleanup() {
|
||||||
|
trap - EXIT HUP INT TERM
|
||||||
|
if [[ "$restore_started" == true ]]; then
|
||||||
|
docker rm --force "$restore_container" >/dev/null 2>&1 || true
|
||||||
|
fi
|
||||||
|
if [[ "$restore_volume_created" == true ]]; then
|
||||||
|
docker volume rm "$restore_volume" >/dev/null 2>&1 || true
|
||||||
|
fi
|
||||||
|
if [[ "$remote_created" == true ]]; then
|
||||||
|
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
|
||||||
|
"rm -f -- '$remote_dump' '$remote_dump.sha256' '$remote_dump.metadata'" \
|
||||||
|
>/dev/null 2>&1 || true
|
||||||
|
fi
|
||||||
|
rm -rf "$work_dir"
|
||||||
|
}
|
||||||
|
trap cleanup EXIT HUP INT TERM
|
||||||
|
|
||||||
|
restore_snapshot() {
|
||||||
|
local selected_snapshot=$1
|
||||||
|
local dump="$work_dir/production.dump"
|
||||||
|
local password
|
||||||
|
local health
|
||||||
|
local tables
|
||||||
|
local migrations
|
||||||
|
local postgis
|
||||||
|
|
||||||
|
"$restic" dump "$selected_snapshot" /production.dump >"$dump"
|
||||||
|
[[ -s "$dump" ]] || {
|
||||||
|
echo "Restic restored an empty database dump." >&2
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
pg_restore --list "$dump" >"$evidence_dir/pg-restore-catalog.txt"
|
||||||
|
|
||||||
|
if ! docker image inspect "$restore_image" >/dev/null 2>&1; then
|
||||||
|
docker pull "$restore_image" >"$evidence_dir/postgis-pull.txt"
|
||||||
|
fi
|
||||||
|
|
||||||
|
password=$(openssl rand -base64 36 | tr -d '\n')
|
||||||
|
docker volume create "$restore_volume" >/dev/null
|
||||||
|
restore_volume_created=true
|
||||||
|
docker run --detach \
|
||||||
|
--name "$restore_container" \
|
||||||
|
--env POSTGRES_PASSWORD="$password" \
|
||||||
|
--env POSTGRES_DB=postgres \
|
||||||
|
--volume "$restore_volume:/var/lib/postgresql" \
|
||||||
|
--health-cmd='pg_isready --username postgres --dbname postgres' \
|
||||||
|
--health-interval=1s \
|
||||||
|
--health-timeout=2s \
|
||||||
|
--health-retries=120 \
|
||||||
|
"$restore_image" >/dev/null
|
||||||
|
restore_started=true
|
||||||
|
|
||||||
|
# A freshly initialized postgres/postgis container briefly accepts
|
||||||
|
# connections through its temporary bootstrap server. Wait until the image
|
||||||
|
# has completed that bootstrap and started the final server before restoring.
|
||||||
|
while ! docker logs "$restore_container" 2>&1 |
|
||||||
|
grep -Fq 'PostgreSQL init process complete; ready for start up.'; do
|
||||||
|
if [[ "$(docker inspect --format '{{.State.Running}}' "$restore_container")" != true ]]; then
|
||||||
|
docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true
|
||||||
|
echo "Isolated restore PostgreSQL exited during initialization." >&2
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
|
|
||||||
|
while true; do
|
||||||
|
health=$(docker inspect --format '{{if .State.Health}}{{.State.Health.Status}}{{else}}missing{{end}}' "$restore_container")
|
||||||
|
case "$health" in
|
||||||
|
healthy) break ;;
|
||||||
|
unhealthy)
|
||||||
|
docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true
|
||||||
|
echo "Isolated restore PostgreSQL became unhealthy." >&2
|
||||||
|
return 1
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
|
|
||||||
|
# The PostGIS image initializes its requested database with PostGIS already
|
||||||
|
# installed. Restore into a database created from template0 instead, so the
|
||||||
|
# archive remains responsible for recreating its own extensions and schema.
|
||||||
|
docker exec "$restore_container" createdb \
|
||||||
|
--username postgres \
|
||||||
|
--template template0 \
|
||||||
|
restore_check
|
||||||
|
|
||||||
|
if ! docker exec --interactive "$restore_container" pg_restore \
|
||||||
|
--username postgres \
|
||||||
|
--dbname restore_check \
|
||||||
|
--exit-on-error \
|
||||||
|
--no-owner \
|
||||||
|
--no-privileges <"$dump"; then
|
||||||
|
docker logs "$restore_container" >"$evidence_dir/restore-postgres.log" 2>&1 || true
|
||||||
|
docker inspect "$restore_container" >"$evidence_dir/restore-container.json" 2>&1 || true
|
||||||
|
echo "Isolated restore failed; container evidence was retained." >&2
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
tables=$(docker exec "$restore_container" psql \
|
||||||
|
--username postgres --dbname restore_check --tuples-only --no-align \
|
||||||
|
--command="select count(*) from pg_tables where schemaname='public' and tablename <> 'spatial_ref_sys';")
|
||||||
|
migrations=$(docker exec "$restore_container" psql \
|
||||||
|
--username postgres --dbname restore_check --tuples-only --no-align \
|
||||||
|
--command='select count(*) from schema_migrations;')
|
||||||
|
postgis=$(docker exec "$restore_container" psql \
|
||||||
|
--username postgres --dbname restore_check --tuples-only --no-align \
|
||||||
|
--command='select PostGIS_Lib_Version();')
|
||||||
|
|
||||||
|
if [[ ! "$tables" =~ ^[1-9][0-9]*$ ]] || [[ ! "$migrations" =~ ^[1-9][0-9]*$ ]]; then
|
||||||
|
echo "The isolated restore does not contain the expected application schema." >&2
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
{
|
||||||
|
printf 'snapshot_id=%s\n' "$selected_snapshot"
|
||||||
|
printf 'application_tables=%s\n' "$tables"
|
||||||
|
printf 'schema_migrations=%s\n' "$migrations"
|
||||||
|
printf 'postgis_version=%s\n' "$postgis"
|
||||||
|
printf 'verified_at=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
||||||
|
} >"$evidence_dir/restore-drill.txt"
|
||||||
|
chmod 600 "$evidence_dir/restore-drill.txt" "$evidence_dir/pg-restore-catalog.txt"
|
||||||
|
|
||||||
|
docker rm --force "$restore_container" >/dev/null
|
||||||
|
restore_started=false
|
||||||
|
docker volume rm "$restore_volume" >/dev/null
|
||||||
|
restore_volume_created=false
|
||||||
|
rm -f "$dump"
|
||||||
|
}
|
||||||
|
|
||||||
|
if [[ "$action" == restore-drill ]]; then
|
||||||
|
restore_snapshot "$snapshot_id"
|
||||||
|
echo "Isolated restore drill passed for snapshot: $snapshot_id"
|
||||||
|
printf 'Non-secret evidence: %s\n' "$evidence_dir"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
check_source
|
||||||
|
|
||||||
|
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
|
||||||
|
"bash -s -- '$PRODUCTION_REMOTE_ROOT/.env' '$remote_dump' '$PRODUCTION_EXPECTED_ENVIRONMENT'" \
|
||||||
|
<"$ROOT/scripts/backup-external-postgres.sh"
|
||||||
|
remote_created=true
|
||||||
|
|
||||||
|
scp -p \
|
||||||
|
"$PRODUCTION_SSH_TARGET:$remote_dump" \
|
||||||
|
"$PRODUCTION_SSH_TARGET:$remote_dump.sha256" \
|
||||||
|
"$PRODUCTION_SSH_TARGET:$remote_dump.metadata" \
|
||||||
|
"$work_dir/"
|
||||||
|
|
||||||
|
dump="$work_dir/$(basename -- "$remote_dump")"
|
||||||
|
checksum="$dump.sha256"
|
||||||
|
metadata="$dump.metadata"
|
||||||
|
|
||||||
|
(
|
||||||
|
cd "$work_dir"
|
||||||
|
sha256sum --check "$(basename -- "$checksum")"
|
||||||
|
) >"$evidence_dir/source-checksum.txt"
|
||||||
|
pg_restore --list "$dump" >"$evidence_dir/source-catalog.txt"
|
||||||
|
cp "$checksum" "$metadata" "$evidence_dir/"
|
||||||
|
chmod 600 "$evidence_dir"/*
|
||||||
|
|
||||||
|
run_tag="$OFFSITE_RESTIC_TAG-$run_id"
|
||||||
|
backup_json="$evidence_dir/restic-backup.jsonl"
|
||||||
|
"$restic" backup \
|
||||||
|
--json \
|
||||||
|
--host "$OFFSITE_RESTIC_HOST" \
|
||||||
|
--tag "$OFFSITE_RESTIC_TAG" \
|
||||||
|
--tag "$run_tag" \
|
||||||
|
--stdin-filename production.dump \
|
||||||
|
--stdin-from-command -- cat "$dump" >"$backup_json"
|
||||||
|
|
||||||
|
snapshot_id=$(jq -r 'select(.message_type == "summary") | .snapshot_id // empty' "$backup_json" | tail -n 1)
|
||||||
|
if [[ -z "$snapshot_id" ]]; then
|
||||||
|
echo "Restic did not report a completed snapshot." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
snapshot_matches=$($restic snapshots --json --tag "$run_tag" | jq 'length')
|
||||||
|
if [[ "$snapshot_matches" != 1 ]]; then
|
||||||
|
echo "The run tag did not resolve to exactly one snapshot." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
"$restic" check --read-data >"$evidence_dir/restic-check.txt"
|
||||||
|
restore_snapshot "$snapshot_id"
|
||||||
|
|
||||||
|
printf '%s\n' "$snapshot_id" >"$evidence_dir/snapshot-id.txt"
|
||||||
|
chmod 600 "$evidence_dir"/*
|
||||||
|
|
||||||
|
ssh -o BatchMode=yes "$PRODUCTION_SSH_TARGET" \
|
||||||
|
"rm -f -- '$remote_dump' '$remote_dump.sha256' '$remote_dump.metadata'"
|
||||||
|
remote_created=false
|
||||||
|
|
||||||
|
echo "Encrypted off-server backup and isolated restore drill passed."
|
||||||
|
printf 'Snapshot: %s\n' "$snapshot_id"
|
||||||
|
printf 'Non-secret evidence: %s\n' "$evidence_dir"
|
||||||
Loading…
Reference in New Issue
Block a user