Add guarded production load profiling

This commit is contained in:
SimpleTest 2026-07-28 13:18:45 +03:00
parent b519f902e8
commit cd7f6ea45c
7 changed files with 1049 additions and 5 deletions

View File

@ -444,12 +444,12 @@ The same external-service protocol drill used by CI can be run independently:
Run the isolated HTTP/WebSocket/authenticated chat/tracking load profile with
resource, PostgreSQL statement, database-connection, and Ecto pool-wait
measurements:
measurements. The lifecycle wrapper is preferred because it removes only its
unique containers, networks, volumes, environment, and image tags on success,
failure, or interruption:
```bash
./scripts/load-stack-up.sh
./scripts/load-run.sh local-load
./scripts/load-stack-stop.sh
./scripts/load-cycle.sh local-load load
```
The profile has its own generated mode-`0600` environment, Compose project,
@ -458,6 +458,24 @@ database. Exact inputs and threshold-free evidence are retained below
`output/performance/local-load/`; see
[Performance measurement](docs/performance.md) for scope and interpretation.
A separate production probe is limited to a compiled allowlist of public GET
pages and Phoenix heartbeat frames. First inspect a read-only plan with every
experiment input supplied explicitly:
```bash
WNH_PRODUCTION_LOAD_HTTP_VUS=<chosen-count> \
WNH_PRODUCTION_LOAD_WS_VUS=<chosen-count> \
WNH_PRODUCTION_LOAD_DURATION=<chosen-duration> \
WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS=<chosen-seconds> \
WNH_PRODUCTION_LOAD_WS_HOLD_MS=<chosen-milliseconds> \
WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS=<chosen-milliseconds> \
./scripts/production-readonly-load.sh plan
```
It does not start load in `plan` mode. The separately approved `run` mode
requires the exact confirmation printed by that plan and still cannot test
authenticated writes or establish a production capacity limit.
The cursor-pagination database benchmark also creates a one-run Compose
project, random database credentials, and a separate PostgreSQL volume:

View File

@ -28,7 +28,120 @@ queries.
This isolates the limiter on a fast local tmpfs database. It does not include
SMTP delivery, account lookup, production storage latency, internet latency,
or production contention, and therefore is not a production capacity claim.
It provides no evidence that Redis is currently required.
It verifies the one-statement implementation shape. It provides no evidence
that Redis is currently required or unnecessary on the production host.
## Two different load profiles
The repository deliberately separates two experiments:
1. `scripts/load-cycle.sh` is the full write-capable profile. It creates a
unique Compose project, database, credentials, images, and fixtures. It
exercises login, database writes, private chat, tracking, WebSockets,
multiple web/worker replicas, SQL statements, and connection pools. Use
this profile to find application and query bottlenecks without touching
real users.
2. `scripts/production-readonly-load.sh` is pinned to
`https://whoneedhelp.com`. Its compiled allowlist contains only public GET
pages, readiness, and Phoenix WebSocket heartbeats. It cannot exercise the
authenticated write path and therefore cannot determine write capacity or
whether Redis is warranted.
The production runner has separate `plan` and `run` modes. Every VU, duration,
think-time, and socket value is mandatory; the script supplies no hidden load
defaults. `plan` verifies the remote checkout identity, exact commit, Compose
project, running application containers, and current host resources without
starting k6. `run` additionally requires the exact confirmation string emitted
by that verified plan. If the production commit or any experiment input
changes, the confirmation no longer matches.
Example plan only:
```sh
WNH_PRODUCTION_LOAD_HTTP_VUS=<chosen-count> \
WNH_PRODUCTION_LOAD_WS_VUS=<chosen-count> \
WNH_PRODUCTION_LOAD_DURATION=<chosen-duration> \
WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS=<chosen-seconds> \
WNH_PRODUCTION_LOAD_WS_HOLD_MS=<chosen-milliseconds> \
WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS=<chosen-milliseconds> \
./scripts/production-readonly-load.sh plan
```
Do not copy a previous run's values as production limits. Choose and record a
specific experiment, inspect the printed target and scope, then use the printed
confirmation only when that production probe has been explicitly approved.
When run, evidence is written only on the operator workstation below
`output/performance/<label>/`. It includes the exact scope, machine-readable k6
summary, console report, readiness before/after, read-only database counter
snapshots, application metrics before/after, one-second host/container samples,
and a threshold-free resource summary. The sampler is stopped on success,
failure, or interruption and creates no remote file.
The host sampler reports the combined RSS of every PostgreSQL process on that
server as `host_postgres_rss_kib`. It is intentionally labelled host-wide
because multiple PostgreSQL clusters or databases may coexist; it must not be
presented as production-database-only memory.
The production profile applies pass/fail checks only to functional correctness:
HTTP 200 responses, successful WebSocket heartbeats, and zero observed request
or socket errors. It records latency, CPU, RAM, process counts, and PostgreSQL
counters but does not compare them to invented SLOs or capacity thresholds.
Read-only production inventory observed on 2026-07-28 before implementing this
runner:
- commit `3d6f9c5004f9218c55a040bc31e72b56b22d4470`;
- one healthy compact application container;
- 2 online CPUs;
- 4,004,224 KiB total memory and 1,282,436 KiB available in that sample;
- five idle application database connections;
- no Docker CPU or memory hard limit on the application container.
These values are a point-in-time observation, not a server minimum. A later
`plan` rechecks the mutable facts rather than trusting this paragraph.
## Current isolated harness verification
Observed on 2026-07-28 from the application tree at commit
`b519f902e8d30f49a3475ac07311f69722bba188`, with the uncommitted load-harness
additions documented on this page. The isolated profile used its recorded
inputs: 3 web replicas, 2 worker replicas, 40 public HTTP VUs, 40 Phoenix
WebSocket VUs, 8 authenticated mutual-aid VUs, and a 30-second duration.
| Observation | Result |
| --- | ---: |
| HTTP requests / failures | 38,222 / 0 |
| Public checks | 34,806 / 34,806 passed |
| Overall HTTP average / p95 / maximum | 2.94 / 7.03 / 182.37 ms |
| Authenticated chain average / p95 | 7.35 / 10.47 ms |
| Authenticated chat/tracking chains | 1,700 |
| WebSocket opened / heartbeat replies | 240 / 240 |
| PostgreSQL peak client / active backends | 21 / 3 |
| Configured Repo pool slots | 16 |
| Minimum observed connection headroom | 76 |
| Ecto average total / execution / queue | 257.46 / 234.79 / 21.75 µs |
| Rollbacks / deadlocks / temp files | 0 / 0 / 0 |
The SQL statement with the largest individual maximum took 8.793 ms and was a
primary-key request lookup. The highest mean was 0.761 ms across five session
token inserts. No statement used temporary blocks. This particular isolated
dataset and workload exposed no database connection exhaustion, deadlock,
temporary-file spill, or failed application operation. It does not prove that
the same result holds for production data cardinality, storage latency, or a
different request mix.
The load profile recorded exactly 1,700 messages, tracking sessions, and
tracking samples. Explicit tracking stop left zero current tracking positions.
Its cleanup diff was empty, and the lifecycle wrapper then removed all
containers, networks, named volumes, and unique image tags belonging to that
run. Evidence remains in the ignored local directory
`output/performance/load-harness-20260728/`.
The run used a 32-CPU, 100,289,273,856-byte workstation without container CPU
or RAM limits. Its latency, throughput, and CPU percentages must not be used as
production capacity evidence. The production server observed above has a
different CPU and memory topology.
## Latest local candidate measurement

View File

@ -0,0 +1,194 @@
import { check, sleep } from "k6";
import http from "k6/http";
import exec from "k6/execution";
import { Counter } from "k6/metrics";
import { WebSocket } from "k6/websockets";
function required(name) {
const value = __ENV[name];
if (value === undefined || value === "") {
throw new Error(`${name} is required`);
}
return value;
}
function nonNegativeInteger(name) {
const value = Number(required(name));
if (!Number.isInteger(value) || value < 0) {
throw new Error(`${name} must be a non-negative integer`);
}
return value;
}
function positiveInteger(name) {
const value = nonNegativeInteger(name);
if (value === 0) {
throw new Error(`${name} must be a positive integer`);
}
return value;
}
function nonNegativeNumber(name) {
const value = Number(required(name));
if (!Number.isFinite(value) || value < 0) {
throw new Error(`${name} must be a non-negative number`);
}
return value;
}
const baseUrl = required("BASE_URL").replace(/\/+$/, "");
if (baseUrl !== "https://whoneedhelp.com") {
throw new Error("The production read-only profile is pinned to https://whoneedhelp.com");
}
const httpVus = nonNegativeInteger("HTTP_VUS");
const websocketVus = nonNegativeInteger("WS_VUS");
if (httpVus + websocketVus === 0) {
throw new Error("At least one HTTP or WebSocket VU is required");
}
const duration = required("DURATION");
const httpThinkSeconds = nonNegativeNumber("HTTP_THINK_SECONDS");
const websocketHoldMs = positiveInteger("WS_HOLD_MS");
const websocketConnectTimeoutMs = positiveInteger("WS_CONNECT_TIMEOUT_MS");
const websocketUrl =
baseUrl.replace(/^http/, "ws") + "/live/websocket?vsn=2.0.0";
// This allowlist is intentionally compiled into the read-only profile. It
// contains GET-only public pages and the readiness endpoint. Authentication,
// registration, support submission, email delivery, and every mutating route
// are outside this file.
const publicRoutes = [
{ path: "/", endpoint: "home" },
{ path: "/safety", endpoint: "safety" },
{ path: "/privacy", endpoint: "privacy" },
{ path: "/terms", endpoint: "terms" },
{ path: "/users/log-in", endpoint: "login_page" },
{ path: "/users/register", endpoint: "registration_page" },
{ path: "/healthz/ready", endpoint: "readiness" },
];
const websocketOpened = new Counter("wnh_readonly_websocket_opened");
const websocketErrors = new Counter("wnh_readonly_websocket_errors");
const heartbeatReplies = new Counter("wnh_readonly_heartbeat_replies");
const scenarios = {};
const thresholds = {};
if (httpVus > 0) {
scenarios.public_http = {
executor: "constant-vus",
exec: "publicHttp",
vus: httpVus,
duration,
};
thresholds.checks = [{ threshold: "rate==1", abortOnFail: true }];
thresholds.http_req_failed = [{ threshold: "rate==0", abortOnFail: true }];
}
if (websocketVus > 0) {
scenarios.phoenix_websocket = {
executor: "constant-vus",
exec: "phoenixWebsocket",
vus: websocketVus,
duration,
};
thresholds.wnh_readonly_websocket_errors = [
{ threshold: "count==0", abortOnFail: true },
];
}
export const options = {
discardResponseBodies: true,
scenarios,
thresholds,
};
export function publicHttp() {
const route = publicRoutes[exec.scenario.iterationInTest % publicRoutes.length];
const response = http.get(`${baseUrl}${route.path}`, {
redirects: 0,
tags: { endpoint: route.endpoint, safety_scope: "production_readonly" },
headers: { "User-Agent": "WhoNeedHelp-production-readonly-load/1" },
});
check(response, {
[`GET ${route.path} returned 200`]: (candidate) => candidate.status === 200,
});
sleep(httpThinkSeconds);
}
export function phoenixWebsocket() {
const socket = new WebSocket(websocketUrl, [], {
tags: {
endpoint: "phoenix_live_socket",
safety_scope: "production_readonly",
},
headers: {
Origin: baseUrl,
"User-Agent": "WhoNeedHelp-production-readonly-load/1",
},
});
const ref = `${__VU}-${__ITER}`;
let opened = false;
let closeTimer;
const connectTimer = setTimeout(() => {
if (!opened) {
websocketErrors.add(1);
socket.close();
}
}, websocketConnectTimeoutMs);
socket.addEventListener("open", () => {
opened = true;
clearTimeout(connectTimer);
websocketOpened.add(1);
socket.send(JSON.stringify([null, ref, "phoenix", "heartbeat", {}]));
closeTimer = setTimeout(() => socket.close(), websocketHoldMs);
});
socket.addEventListener("message", (event) => {
let frame;
try {
frame = JSON.parse(event.data);
} catch (_error) {
return;
}
if (
Array.isArray(frame) &&
frame[1] === ref &&
frame[2] === "phoenix" &&
frame[3] === "phx_reply" &&
frame[4]?.status === "ok"
) {
heartbeatReplies.add(1);
}
});
socket.addEventListener("error", () => {
websocketErrors.add(1);
clearTimeout(connectTimer);
socket.close();
});
socket.addEventListener("close", () => {
clearTimeout(connectTimer);
if (closeTimer !== undefined) {
clearTimeout(closeTimer);
}
});
}

View File

@ -0,0 +1,116 @@
#!/usr/bin/env bash
set -euo pipefail
umask 077
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
K6_IMAGE="grafana/k6@sha256:65c920dc067d5e2e00befbf982af6ad6ad0117034e8b1c65817c7975c52d4669"
temporary_directory=$(mktemp -d "${TMPDIR:-/tmp}/wnh-production-load-drill.XXXXXX")
cleanup() {
trap - EXIT HUP INT TERM
rm -rf "$temporary_directory"
}
trap cleanup EXIT HUP INT TERM
mkdir -p "$temporary_directory/bin"
# The single-quoted text below is intentionally written into the SSH stub and
# expands only when that generated script runs.
# shellcheck disable=SC2016
printf '%s\n' \
'#!/bin/sh' \
'if [ "${WNH_LOAD_STUB_MODE:-good}" = bad ]; then' \
' printf "%s\n" "commit=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa" "project=unexpected" "deployment_env=test" "phx_host=invalid.example" "base_url=https://invalid.example" "containers=invalid-app-1" "cpu_count=1" "mem_total_kib=1" "mem_available_kib=1" "load_average=0 0 0"' \
'else' \
' printf "%s\n" "commit=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa" "project=who_need_help_production" "deployment_env=production" "phx_host=whoneedhelp.com" "base_url=https://whoneedhelp.com" "containers=who_need_help_production-app-1" "cpu_count=2" "mem_total_kib=4000000" "mem_available_kib=1000000" "load_average=0.1 0.1 0.1"' \
'fi' >"$temporary_directory/bin/ssh"
chmod +x "$temporary_directory/bin/ssh"
export PATH="$temporary_directory/bin:$PATH"
export WNH_PRODUCTION_LOAD_HTTP_VUS=1
export WNH_PRODUCTION_LOAD_WS_VUS=1
export WNH_PRODUCTION_LOAD_DURATION=1s
export WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS=1
export WNH_PRODUCTION_LOAD_WS_HOLD_MS=1000
export WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS=3000
"$ROOT/scripts/production-readonly-load.sh" plan drill \
>"$temporary_directory/plan.txt"
grep -F 'URL: https://whoneedhelp.com' "$temporary_directory/plan.txt" >/dev/null
grep -F 'No POST, login attempt, registration, email, support request' \
"$temporary_directory/plan.txt" >/dev/null
grep -F \
"WNH_PRODUCTION_LOAD_CONFIRM='whoneedhelp.com:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa:http=1:ws=1:duration=1s'" \
"$temporary_directory/plan.txt" >/dev/null
if WNH_LOAD_STUB_MODE=bad \
"$ROOT/scripts/production-readonly-load.sh" plan drill-bad-target \
>"$temporary_directory/bad-target.txt" 2>&1; then
echo "Production runner accepted an unexpected deployment identity." >&2
exit 1
fi
grep -F 'does not match the pinned production identity' \
"$temporary_directory/bad-target.txt" >/dev/null
if WNH_PRODUCTION_LOAD_CONFIRM=wrong \
"$ROOT/scripts/production-readonly-load.sh" run drill-refusal \
>"$temporary_directory/refusal.txt" 2>&1; then
echo "Production runner accepted an incorrect confirmation." >&2
exit 1
fi
grep -F 'does not match this verified plan' \
"$temporary_directory/refusal.txt" >/dev/null
test ! -e "$ROOT/output/performance/drill-refusal"
printf '%s\n' \
'{"observed_at":"2026-07-28T00:00:00Z","host":{"load_1":0.5,"mem_total_kib":4000000,"mem_available_kib":1000000,"host_postgres_rss_kib":200000},"containers":[{"Name":"app-1","CPUPerc":"10.0%","MemUsage":"200MiB / 4GiB","PIDs":"25"}]}' \
'{"observed_at":"2026-07-28T00:00:01Z","host":{"load_1":0.7,"mem_total_kib":4000000,"mem_available_kib":900000,"host_postgres_rss_kib":210000},"containers":[{"Name":"app-1","CPUPerc":"20.0%","MemUsage":"210MiB / 4GiB","PIDs":"26"}]}' \
>"$temporary_directory/server-stats.jsonl"
"$ROOT/scripts/summarize-production-load.py" \
"$temporary_directory/server-stats.jsonl" \
"$temporary_directory/server-stats-summary.json"
jq -e '
.sample_count == 2 and
.host.load_1.average == 0.6 and
.host.mem_available_kib.minimum == 900000 and
.host.host_postgres_rss_kib.maximum == 210000 and
.containers[0].memory_bytes.first_to_last_delta == 10485760
' "$temporary_directory/server-stats-summary.json" >/dev/null
docker run --rm \
--volume "$ROOT/load/k6:/scripts:ro" \
"$K6_IMAGE" inspect \
-e BASE_URL=https://whoneedhelp.com \
-e HTTP_VUS=1 \
-e WS_VUS=1 \
-e DURATION=1s \
-e HTTP_THINK_SECONDS=1 \
-e WS_HOLD_MS=1000 \
-e WS_CONNECT_TIMEOUT_MS=3000 \
/scripts/production-readonly.js \
>"$temporary_directory/k6-inspect.json"
jq -e '
.scenarios.public_http.vus == 1 and
.scenarios.phoenix_websocket.vus == 1 and
.thresholds.http_req_failed[0].threshold == "rate==0"
' "$temporary_directory/k6-inspect.json" >/dev/null
docker run --rm \
--volume "$ROOT/load/k6:/scripts:ro" \
"$K6_IMAGE" inspect \
-e BASE_URL=https://whoneedhelp.com \
-e HTTP_VUS=0 \
-e WS_VUS=1 \
-e DURATION=1s \
-e HTTP_THINK_SECONDS=1 \
-e WS_HOLD_MS=1000 \
-e WS_CONNECT_TIMEOUT_MS=3000 \
/scripts/production-readonly.js \
>"$temporary_directory/k6-websocket-only-inspect.json"
jq -e '
(.scenarios | has("public_http") | not) and
.scenarios.phoenix_websocket.vus == 1 and
(.thresholds | has("checks") | not) and
(.thresholds | has("http_req_failed") | not)
' "$temporary_directory/k6-websocket-only-inspect.json" >/dev/null
echo "Production read-only load safety drill passed."

View File

@ -0,0 +1,417 @@
#!/usr/bin/env bash
set -euo pipefail
umask 077
ROOT=$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)
K6_IMAGE="grafana/k6@sha256:65c920dc067d5e2e00befbf982af6ad6ad0117034e8b1c65817c7975c52d4669"
BASE_URL="https://whoneedhelp.com"
SSH_TARGET=${WNH_PRODUCTION_LOAD_SSH_TARGET:-whoneedhelp}
REMOTE_DIRECTORY=${WNH_PRODUCTION_LOAD_REMOTE_DIRECTORY:-/srv/who_need_help-production}
MODE=${1:-plan}
LABEL=${2:-"production-readonly-$(date -u +%Y%m%dT%H%M%SZ)"}
usage() {
cat <<'EOF'
Usage:
WNH_PRODUCTION_LOAD_HTTP_VUS=... \
WNH_PRODUCTION_LOAD_WS_VUS=... \
WNH_PRODUCTION_LOAD_DURATION=... \
WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS=... \
WNH_PRODUCTION_LOAD_WS_HOLD_MS=... \
WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS=... \
./scripts/production-readonly-load.sh plan [label]
The run mode additionally requires the exact confirmation string printed by
plan in WNH_PRODUCTION_LOAD_CONFIRM.
EOF
}
case "$MODE" in
plan | run) ;;
*)
usage >&2
exit 1
;;
esac
if [[ ! "$LABEL" =~ ^[A-Za-z0-9._-]+$ ]]; then
echo "Run label may contain only letters, numbers, dot, underscore, and dash." >&2
exit 1
fi
required=(
WNH_PRODUCTION_LOAD_HTTP_VUS
WNH_PRODUCTION_LOAD_WS_VUS
WNH_PRODUCTION_LOAD_DURATION
WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS
WNH_PRODUCTION_LOAD_WS_HOLD_MS
WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS
)
for name in "${required[@]}"; do
if [[ -z "${!name:-}" ]]; then
echo "$name must be supplied explicitly for this experiment." >&2
exit 1
fi
done
for name in WNH_PRODUCTION_LOAD_HTTP_VUS WNH_PRODUCTION_LOAD_WS_VUS; do
if [[ ! "${!name}" =~ ^[0-9]+$ ]]; then
echo "$name must be a non-negative integer." >&2
exit 1
fi
done
if ((WNH_PRODUCTION_LOAD_HTTP_VUS + WNH_PRODUCTION_LOAD_WS_VUS == 0)); then
echo "At least one HTTP or WebSocket VU is required." >&2
exit 1
fi
if [[ ! "$WNH_PRODUCTION_LOAD_DURATION" =~ ^[1-9][0-9]*(ms|s|m|h)$ ]]; then
echo "WNH_PRODUCTION_LOAD_DURATION must be a positive k6 duration." >&2
exit 1
fi
if [[ ! "$WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS" =~ ^[0-9]+([.][0-9]+)?$ ]]; then
echo "WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS must be non-negative." >&2
exit 1
fi
for name in \
WNH_PRODUCTION_LOAD_WS_HOLD_MS \
WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS; do
if [[ ! "${!name}" =~ ^[1-9][0-9]*$ ]]; then
echo "$name must be a positive integer." >&2
exit 1
fi
done
inventory=$(
ssh "$SSH_TARGET" bash -s -- "$REMOTE_DIRECTORY" <<'REMOTE'
set -euo pipefail
directory=$1
env_file="$directory/.env"
if [[ ! -f "$env_file" ]]; then
echo "Missing production environment: $env_file" >&2
exit 1
fi
read_env() {
local name=$1
sed -n "s/^${name}=//p" "$env_file" | tail -n 1
}
deployment_env=$(read_env DEPLOYMENT_ENV)
phx_host=$(read_env PHX_HOST)
base_url=$(read_env WNH_BASE_URL)
project=$(read_env COMPOSE_PROJECT_NAME)
if [[ "$deployment_env" != "production" ]] ||
[[ "$phx_host" != "whoneedhelp.com" ]] ||
[[ "$base_url" != "https://whoneedhelp.com" ]] ||
[[ "$project" != "who_need_help_production" ]]; then
echo "The remote checkout did not match the pinned production identity." >&2
exit 1
fi
commit=$(git -C "$directory" rev-parse HEAD)
mapfile -t containers < <(
docker ps \
--filter "label=com.docker.compose.project=$project" \
--filter "label=com.docker.compose.service=app" \
--format '{{.Names}}'
)
if [[ "${#containers[@]}" -eq 0 ]]; then
mapfile -t containers < <(
docker ps \
--filter "label=com.docker.compose.project=$project" \
--filter "label=com.docker.compose.service=web" \
--format '{{.Names}}'
)
fi
if [[ "${#containers[@]}" -eq 0 ]]; then
echo "No running production application containers were found." >&2
exit 1
fi
printf 'commit=%s\n' "$commit"
printf 'project=%s\n' "$project"
printf 'deployment_env=%s\n' "$deployment_env"
printf 'phx_host=%s\n' "$phx_host"
printf 'base_url=%s\n' "$base_url"
printf 'containers=%s\n' "$(IFS=,; echo "${containers[*]}")"
printf 'cpu_count=%s\n' "$(getconf _NPROCESSORS_ONLN)"
printf 'mem_total_kib=%s\n' "$(awk '/^MemTotal:/ {print $2}' /proc/meminfo)"
printf 'mem_available_kib=%s\n' "$(awk '/^MemAvailable:/ {print $2}' /proc/meminfo)"
printf 'load_average=%s\n' "$(cut -d' ' -f1-3 /proc/loadavg)"
REMOTE
)
value_from_inventory() {
local name=$1
printf '%s\n' "$inventory" | sed -n "s/^${name}=//p" | tail -n 1
}
commit=$(value_from_inventory commit)
project=$(value_from_inventory project)
containers_csv=$(value_from_inventory containers)
deployment_env=$(value_from_inventory deployment_env)
phx_host=$(value_from_inventory phx_host)
verified_base_url=$(value_from_inventory base_url)
if [[ ! "$commit" =~ ^[0-9a-f]{40}$ ]]; then
echo "Could not verify the production commit." >&2
exit 1
fi
if [[ "$project" != "who_need_help_production" ]] ||
[[ "$deployment_env" != "production" ]] ||
[[ "$phx_host" != "whoneedhelp.com" ]] ||
[[ "$verified_base_url" != "$BASE_URL" ]] ||
[[ -z "$containers_csv" ]]; then
echo "The verified inventory does not match the pinned production identity." >&2
exit 1
fi
confirmation="whoneedhelp.com:${commit}:http=${WNH_PRODUCTION_LOAD_HTTP_VUS}:ws=${WNH_PRODUCTION_LOAD_WS_VUS}:duration=${WNH_PRODUCTION_LOAD_DURATION}"
cat <<EOF
Verified target:
SSH target: $SSH_TARGET
directory: $REMOTE_DIRECTORY
URL: $BASE_URL
Compose project: $project
commit: $commit
application containers: $containers_csv
CPU count: $(value_from_inventory cpu_count)
MemTotal: $(value_from_inventory mem_total_kib) KiB
MemAvailable now: $(value_from_inventory mem_available_kib) KiB
load average now: $(value_from_inventory load_average)
Exact request scope:
GET /, /safety, /privacy, /terms, /users/log-in,
/users/register, /healthz/ready
WebSocket /live/websocket with Phoenix heartbeat only
No POST, login attempt, registration, email, support request, chat,
location update, application-data fixture, migration, or database reset.
Experiment inputs:
HTTP VUs: $WNH_PRODUCTION_LOAD_HTTP_VUS
WebSocket VUs: $WNH_PRODUCTION_LOAD_WS_VUS
duration: $WNH_PRODUCTION_LOAD_DURATION
HTTP think seconds: $WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS
WebSocket hold ms: $WNH_PRODUCTION_LOAD_WS_HOLD_MS
WebSocket connect timeout ms: $WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS
To execute exactly this plan, set:
WNH_PRODUCTION_LOAD_CONFIRM='$confirmation'
EOF
if [[ "$MODE" == "plan" ]]; then
exit 0
fi
if [[ "${WNH_PRODUCTION_LOAD_CONFIRM:-}" != "$confirmation" ]]; then
echo "Run refused: WNH_PRODUCTION_LOAD_CONFIRM does not match this verified plan." >&2
exit 1
fi
output_dir="$ROOT/output/performance/$LABEL"
if [[ -e "$output_dir" ]]; then
echo "Output path already exists: $output_dir" >&2
exit 1
fi
mkdir -p "$output_dir"
chmod 700 "$ROOT/output" "$ROOT/output/performance" "$output_dir"
printf '%s\n' "$inventory" >"$output_dir/environment.txt"
cat >"$output_dir/scope.txt" <<EOF
target=$BASE_URL
commit=$commit
confirmation=$confirmation
http_vus=$WNH_PRODUCTION_LOAD_HTTP_VUS
websocket_vus=$WNH_PRODUCTION_LOAD_WS_VUS
duration=$WNH_PRODUCTION_LOAD_DURATION
http_think_seconds=$WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS
websocket_hold_ms=$WNH_PRODUCTION_LOAD_WS_HOLD_MS
websocket_connect_timeout_ms=$WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS
scope=public GET allowlist and Phoenix heartbeat only
latency_or_capacity_thresholds_applied=false
EOF
IFS=, read -r -a containers <<<"$containers_csv"
primary_container=${containers[0]}
snapshot_database() {
local destination=$1
ssh "$SSH_TARGET" bash -s -- "$primary_container" <<'REMOTE' | tail -n 1 >"$destination"
set -euo pipefail
container=$1
docker exec "$container" /app/bin/who_need_help rpc '
alias WhoNeedHelp.Repo
database = Repo.query!("""
SELECT numbackends, xact_commit, xact_rollback, blks_read, blks_hit,
temp_files, temp_bytes, deadlocks
FROM pg_stat_database
WHERE datname = current_database()
""").rows
counts = Repo.query!("""
SELECT
(SELECT count(*) FROM users),
(SELECT count(*) FROM users_tokens),
(SELECT count(*) FROM rate_limit_buckets)
""").rows
IO.puts(Jason.encode!(%{
captured_at: DateTime.utc_now(),
pg_stat_database: database,
counts: counts,
tables: ["users", "users_tokens", "rate_limit_buckets"]
}))
'
REMOTE
}
snapshot_metrics() {
local destination=$1
# Expansion of METRICS_TOKEN occurs only inside the application container.
# shellcheck disable=SC2016
ssh "$SSH_TARGET" bash -s -- "$primary_container" <<'REMOTE' >"$destination"
set -euo pipefail
container=$1
docker exec "$container" sh -c \
'curl --fail --silent --show-error \
--header "Authorization: Bearer $METRICS_TOKEN" \
http://127.0.0.1:4000/metrics'
REMOTE
}
resource_log="$output_dir/server-stats.jsonl"
sampler_pid=
cleanup() {
local status=$?
trap - EXIT HUP INT TERM
if [[ -n "$sampler_pid" ]] && kill -0 "$sampler_pid" 2>/dev/null; then
kill "$sampler_pid" 2>/dev/null || true
wait "$sampler_pid" 2>/dev/null || true
fi
exit "$status"
}
trap cleanup EXIT HUP INT TERM
curl --fail --silent --show-error "$BASE_URL/healthz/ready" \
>"$output_dir/readiness-before.json"
snapshot_database "$output_dir/database-before.json"
snapshot_metrics "$output_dir/metrics-before.prom"
ssh "$SSH_TARGET" bash -s -- "$project" <<'REMOTE' >"$resource_log" &
set -euo pipefail
project=$1
while :; do
observed_at=$(date -u +%Y-%m-%dT%H:%M:%SZ)
load_1=$(awk '{print $1}' /proc/loadavg)
mem_total=$(awk '/^MemTotal:/ {print $2}' /proc/meminfo)
mem_available=$(awk '/^MemAvailable:/ {print $2}' /proc/meminfo)
host_postgres_rss=$(
ps -C postgres -o rss= 2>/dev/null |
awk '{sum += $1} END {print sum + 0}'
)
mapfile -t ids < <(
docker ps -q --filter "label=com.docker.compose.project=$project"
)
if [[ "${#ids[@]}" -eq 0 ]]; then
exit 1
fi
containers=$(
docker stats --no-stream --format '{{json .}}' "${ids[@]}" |
jq -s .
)
jq -cn \
--arg observed_at "$observed_at" \
--arg load_1 "$load_1" \
--arg mem_total "$mem_total" \
--arg mem_available "$mem_available" \
--arg host_postgres_rss "$host_postgres_rss" \
--argjson containers "$containers" \
'{
observed_at: $observed_at,
host: {
load_1: ($load_1 | tonumber),
mem_total_kib: ($mem_total | tonumber),
mem_available_kib: ($mem_available | tonumber),
host_postgres_rss_kib: ($host_postgres_rss | tonumber)
},
containers: $containers
}'
sleep 1
done
REMOTE
sampler_pid=$!
set +e
docker run --rm \
--user "$(id -u):$(id -g)" \
--volume "$ROOT/load/k6:/scripts:ro" \
--volume "$output_dir:/results" \
--env BASE_URL="$BASE_URL" \
--env HTTP_VUS="$WNH_PRODUCTION_LOAD_HTTP_VUS" \
--env WS_VUS="$WNH_PRODUCTION_LOAD_WS_VUS" \
--env DURATION="$WNH_PRODUCTION_LOAD_DURATION" \
--env HTTP_THINK_SECONDS="$WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS" \
--env WS_HOLD_MS="$WNH_PRODUCTION_LOAD_WS_HOLD_MS" \
--env WS_CONNECT_TIMEOUT_MS="$WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS" \
--env K6_NO_USAGE_REPORT=true \
"$K6_IMAGE" run \
--no-usage-report \
--summary-mode=full \
--summary-export /results/k6-summary.json \
--new-machine-readable-summary \
/scripts/production-readonly.js |
tee "$output_dir/k6-console.txt"
k6_status=${PIPESTATUS[0]}
set -e
kill "$sampler_pid" 2>/dev/null || true
wait "$sampler_pid" 2>/dev/null || true
sampler_pid=
curl --fail --silent --show-error "$BASE_URL/healthz/ready" \
>"$output_dir/readiness-after.json"
snapshot_database "$output_dir/database-after.json"
snapshot_metrics "$output_dir/metrics-after.prom"
"$ROOT/scripts/summarize-production-load.py" \
"$resource_log" "$output_dir/server-stats-summary.json"
jq -e \
--argjson http_vus "$WNH_PRODUCTION_LOAD_HTTP_VUS" \
--argjson websocket_vus "$WNH_PRODUCTION_LOAD_WS_VUS" '
def metric($name):
([.results.metrics[] | select(.name == $name) | .values][0] // {});
(
$http_vus == 0 or
(
(.results.checks.metrics[] |
select(.name == "checks_failed") |
.values.matches) == 0 and
metric("http_req_failed").matches == 0
)
) and
(
$websocket_vus == 0 or
(
(metric("wnh_readonly_websocket_errors").count // 0) == 0 and
metric("wnh_readonly_websocket_opened").count > 0 and
metric("wnh_readonly_websocket_opened").count ==
metric("wnh_readonly_heartbeat_replies").count
)
)
' "$output_dir/k6-summary.json" >/dev/null
if [[ "$k6_status" -ne 0 ]]; then
echo "k6 exited with status $k6_status; evidence is in $output_dir." >&2
exit "$k6_status"
fi
echo "Production read-only probe completed. Evidence: $output_dir"

View File

@ -66,6 +66,9 @@ docker run --rm \
"$SHELLCHECK_IMAGE" \
$(find scripts -type f -name '*.sh' -print | sort)
echo "Checking the production read-only load safety boundary"
./scripts/production-readonly-load-drill.sh
echo "Checking isolated production application rollback plan/apply"
./scripts/production-rollback-drill.sh

View File

@ -0,0 +1,183 @@
#!/usr/bin/env python3
"""Summarize production read-only probe samples without inventing limits."""
from __future__ import annotations
import argparse
import json
import re
from collections import defaultdict
from decimal import Decimal
from pathlib import Path
from statistics import fmean
UNIT_BYTES = {
"B": Decimal(1),
"kB": Decimal(1_000),
"MB": Decimal(1_000_000),
"GB": Decimal(1_000_000_000),
"KiB": Decimal(1_024),
"MiB": Decimal(1_048_576),
"GiB": Decimal(1_073_741_824),
}
def parse_size(value: str) -> int:
match = re.fullmatch(r"([0-9]+(?:\.[0-9]+)?)\s*([A-Za-z]+)", value.strip())
if match is None:
raise ValueError(f"invalid Docker size value: {value}")
number, unit = match.groups()
try:
multiplier = UNIT_BYTES[unit]
except KeyError as error:
raise ValueError(f"unsupported Docker size unit: {unit}") from error
return int(Decimal(number) * multiplier)
def parse_memory_usage(value: str) -> int:
used, separator, _limit = value.partition("/")
if not separator:
raise ValueError(f"invalid Docker MemUsage value: {value}")
return parse_size(used)
def parse_cpu(value: str) -> float:
if not value.endswith("%"):
raise ValueError(f"invalid Docker CPUPerc value: {value}")
return float(value.removesuffix("%"))
def summarize(source: Path) -> dict[str, object]:
hosts: list[dict[str, object]] = []
containers: dict[str, list[dict[str, object]]] = defaultdict(list)
with source.open(encoding="utf-8") as stream:
for line_number, line in enumerate(stream, start=1):
if not line.strip():
continue
try:
row = json.loads(line)
observed_at = str(row["observed_at"])
host = row["host"]
hosts.append(
{
"observed_at": observed_at,
"load_1": float(host["load_1"]),
"mem_total_kib": int(host["mem_total_kib"]),
"mem_available_kib": int(host["mem_available_kib"]),
"host_postgres_rss_kib": int(
host["host_postgres_rss_kib"]
),
}
)
for container in row["containers"]:
name = str(container["Name"])
containers[name].append(
{
"observed_at": observed_at,
"cpu_percent": parse_cpu(str(container["CPUPerc"])),
"memory_bytes": parse_memory_usage(
str(container["MemUsage"])
),
"pids": int(container["PIDs"]),
}
)
except (
KeyError,
TypeError,
ValueError,
json.JSONDecodeError,
) as error:
raise ValueError(f"{source}:{line_number}: {error}") from error
if not hosts:
raise ValueError(f"{source}: no host samples")
if not containers:
raise ValueError(f"{source}: no container samples")
host_load = [float(row["load_1"]) for row in hosts]
available_memory = [int(row["mem_available_kib"]) for row in hosts]
postgres_memory = [int(row["host_postgres_rss_kib"]) for row in hosts]
container_summaries: list[dict[str, object]] = []
for name in sorted(containers):
rows = containers[name]
cpu = [float(row["cpu_percent"]) for row in rows]
memory = [int(row["memory_bytes"]) for row in rows]
pids = [int(row["pids"]) for row in rows]
container_summaries.append(
{
"name": name,
"sample_count": len(rows),
"cpu_percent": {
"average": fmean(cpu),
"minimum": min(cpu),
"maximum": max(cpu),
},
"memory_bytes": {
"average": fmean(memory),
"minimum": min(memory),
"maximum": max(memory),
"first": memory[0],
"last": memory[-1],
"first_to_last_delta": memory[-1] - memory[0],
},
"pids": {
"minimum": min(pids),
"maximum": max(pids),
"first": pids[0],
"last": pids[-1],
},
}
)
return {
"schema_version": 1,
"measurement": (
"one-second read-only host and docker stats samples collected over SSH"
),
"thresholds_applied": False,
"first_observed_at": str(hosts[0]["observed_at"]),
"last_observed_at": str(hosts[-1]["observed_at"]),
"sample_count": len(hosts),
"host": {
"load_1": {
"average": fmean(host_load),
"minimum": min(host_load),
"maximum": max(host_load),
},
"mem_total_kib": {
"minimum": min(int(row["mem_total_kib"]) for row in hosts),
"maximum": max(int(row["mem_total_kib"]) for row in hosts),
},
"mem_available_kib": {
"average": fmean(available_memory),
"minimum": min(available_memory),
"maximum": max(available_memory),
},
"host_postgres_rss_kib": {
"average": fmean(postgres_memory),
"minimum": min(postgres_memory),
"maximum": max(postgres_memory),
},
},
"containers": container_summaries,
}
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("source", type=Path)
parser.add_argument("destination", type=Path)
args = parser.parse_args()
result = summarize(args.source)
args.destination.write_text(
json.dumps(result, indent=2, sort_keys=True) + "\n",
encoding="utf-8",
)
if __name__ == "__main__":
main()