who_need_help/docs/performance.md

54 KiB
Raw Blame History

Performance measurement

No production capacity, minimum resource requirement, SLO, alert threshold, pool size, or autoscaling threshold is known yet. The repository therefore contains a reproducible measurement profile, not a capacity claim or blocking resource preflight.

Authentication limiter micro-measurement

Observed locally on 2026-07-28 with Elixir 1.20.2 / OTP 29.0.3, PostgreSQL 18.4, Docker 29.6.2, and the current uncommitted authentication-limiter candidate on an AMD Ryzen 9 7950X3D workstation. PostgreSQL used an isolated tmpfs data directory and the application connected over a Docker bridge.

Each operation incremented an email bucket and an IP bucket together. A telemetry-backed integration test confirmed that the operation executes one PostgreSQL INSERT ... ON CONFLICT statement, not two independent counter queries.

Observation Result
Sequential operations 1,000
Sequential p50 / p95 / p99 155 / 237 / 273 µs
Concurrent operations / concurrency 1,000 / 50
Concurrent elapsed time 107.4 ms
Concurrent observed throughput 9,308 operations/s

This isolates the limiter on a fast local tmpfs database. It does not include SMTP delivery, account lookup, production storage latency, internet latency, or production contention, and therefore is not a production capacity claim. It verifies the one-statement implementation shape. It provides no evidence that Redis is currently required or unnecessary on the production host.

Two different load profiles

The repository deliberately separates two experiments:

  1. scripts/load-cycle.sh is the full write-capable profile. It creates a unique Compose project, database, credentials, images, and fixtures. It exercises login, database writes, private chat, tracking, WebSockets, multiple web/worker replicas, SQL statements, and connection pools. Use this profile to find application and query bottlenecks without touching real users.
  2. scripts/production-readonly-load.sh is pinned to https://whoneedhelp.com. Its compiled allowlist contains only public GET pages, readiness, and Phoenix WebSocket heartbeats. It cannot exercise the authenticated write path and therefore cannot determine write capacity or whether Redis is warranted.

The production runner has separate plan and run modes. Every VU, duration, think-time, and socket value is mandatory; the script supplies no hidden load defaults. plan verifies the remote checkout identity, exact commit, Compose project, running application containers, and current host resources without starting k6. run additionally requires the exact confirmation string emitted by that verified plan. If the production commit or any experiment input changes, the confirmation no longer matches.

Example plan only:

WNH_PRODUCTION_LOAD_HTTP_VUS=<chosen-count> \
WNH_PRODUCTION_LOAD_WS_VUS=<chosen-count> \
WNH_PRODUCTION_LOAD_DURATION=<chosen-duration> \
WNH_PRODUCTION_LOAD_HTTP_THINK_SECONDS=<chosen-seconds> \
WNH_PRODUCTION_LOAD_WS_HOLD_MS=<chosen-milliseconds> \
WNH_PRODUCTION_LOAD_WS_CONNECT_TIMEOUT_MS=<chosen-milliseconds> \
  ./scripts/production-readonly-load.sh plan

Do not copy a previous run's values as production limits. Choose and record a specific experiment, inspect the printed target and scope, then use the printed confirmation only when that production probe has been explicitly approved.

When run, evidence is written only on the operator workstation below output/performance/<label>/. It includes the exact scope, machine-readable k6 summary, console report, readiness before/after, read-only database counter snapshots, application metrics before/after, one-second host/container samples, and a threshold-free resource summary. The sampler is stopped on success, failure, or interruption and creates no remote file.

The host sampler reports the combined RSS of every PostgreSQL process on that server as host_postgres_rss_kib. It is intentionally labelled host-wide because multiple PostgreSQL clusters or databases may coexist; it must not be presented as production-database-only memory.

The production profile applies pass/fail checks only to functional correctness: HTTP 200 responses, successful WebSocket heartbeats, and zero observed request or socket errors. It records latency, CPU, RAM, process counts, and PostgreSQL counters but does not compare them to invented SLOs or capacity thresholds.

Read-only production inventory observed on 2026-07-28 before implementing this runner:

  • commit 3d6f9c5004f9218c55a040bc31e72b56b22d4470;
  • one healthy compact application container;
  • 2 online CPUs;
  • 4,004,224 KiB total memory and 1,282,436 KiB available in that sample;
  • five idle application database connections;
  • no Docker CPU or memory hard limit on the application container.

These values are a point-in-time observation, not a server minimum. A later plan rechecks the mutable facts rather than trusting this paragraph.

Current isolated harness verification

Observed on 2026-07-28 from the application tree at commit b519f902e8d30f49a3475ac07311f69722bba188, with the uncommitted load-harness additions documented on this page. The isolated profile used its recorded inputs: 3 web replicas, 2 worker replicas, 40 public HTTP VUs, 40 Phoenix WebSocket VUs, 8 authenticated mutual-aid VUs, and a 30-second duration.

Observation Result
HTTP requests / failures 38,222 / 0
Public checks 34,806 / 34,806 passed
Overall HTTP average / p95 / maximum 2.94 / 7.03 / 182.37 ms
Authenticated chain average / p95 7.35 / 10.47 ms
Authenticated chat/tracking chains 1,700
WebSocket opened / heartbeat replies 240 / 240
PostgreSQL peak client / active backends 21 / 3
Configured Repo pool slots 16
Minimum observed connection headroom 76
Ecto average total / execution / queue 257.46 / 234.79 / 21.75 µs
Rollbacks / deadlocks / temp files 0 / 0 / 0

The SQL statement with the largest individual maximum took 8.793 ms and was a primary-key request lookup. The highest mean was 0.761 ms across five session token inserts. No statement used temporary blocks. This particular isolated dataset and workload exposed no database connection exhaustion, deadlock, temporary-file spill, or failed application operation. It does not prove that the same result holds for production data cardinality, storage latency, or a different request mix.

The load profile recorded exactly 1,700 messages, tracking sessions, and tracking samples. Explicit tracking stop left zero current tracking positions. Its cleanup diff was empty, and the lifecycle wrapper then removed all containers, networks, named volumes, and unique image tags belonging to that run. Evidence remains in the ignored local directory output/performance/load-harness-20260728/.

The run used a 32-CPU, 100,289,273,856-byte workstation without container CPU or RAM limits. Its latency, throughput, and CPU percentages must not be used as production capacity evidence. The production server observed above has a different CPU and memory topology.

Latest local candidate measurement

Observed on 2026-07-25 at application commit 4f2ab7d205ad1e137b07a8365e30b7cf007e2a19 with 3 web and 2 worker replicas, 40 public HTTP VUs, 40 Phoenix WebSocket VUs, and 8 authenticated mutual-aid VUs for 30 seconds:

Observation Result
HTTP requests / failures 38,240 / 0
Public checks 34,866 / 34,866 passed
WebSocket opened / heartbeat replies 240 / 240
Authenticated chat/tracking iterations 1,679
Overall HTTP duration average / p95 / maximum 2.90 / 7.28 / 184.51 ms
Authenticated HTTP duration average / p95 7.63 / 10.81 ms
WebSocket connect p95 4.76 ms
PostgreSQL commits / rollbacks 141,740 / 0
Peak client / active backends 21 / 5
Minimum observed connection headroom 76
Ecto queries / average total / average queue 181,796 / 0.275 / 0.022 ms

Ten sequential resource samples observed maximum memory of 312.4, 319.8, and 331.6 MiB across web replicas; 204.2 and 204.6 MiB across workers; 140.7 MiB for PostgreSQL; and 323.1 MiB for the proxy. The sum of each container's individual maximum was 1,836.4 MiB. PostgreSQL recorded no deadlock, conflict, temporary file, or lock waiter, and exact fixture cleanup restored every tracked application-table count.

These measurements describe this specific concurrent workstation replay. They do not establish a minimum server size, production capacity, SLO, alert threshold, database pool size, or autoscaling policy. Machine-readable evidence is retained under output/performance/final-4f2ab7d.

Kind port-table memory investigation

On 2026-07-20, the local kind control-plane container used 9.748 GiB while its four otherwise idle application beam.smp processes each had roughly 2.3 GiB RSS. Read-only runtime probes found an effective Erlang port limit of 134,217,727 and an allocated port table of 1,610,612,736 bytes in every kind BEAM VM. The nested runtime exposed nofile=2,147,483,584; OTP had derived the maximum port limit from that value.

The same release image in ordinary Compose used the normal 65,536 port limit and a 786,432-byte port table. A separate VM probe with ERL_ZFLAGS="+Q 65536" reproduced those smaller values inside the kind pod before any rollout. Compose and Helm now set that value explicitly, and the startup and rolling-verification scripts check the effective value in every running web and worker VM.

On a host with more CPUs than the intended server, the optional compose.cpu-replay.yaml override can pin the application, worker, database, proxy, and supporting services to the same explicit host CPU set. The external load generator remains outside that set:

CPU_REPLAY_CPUSET=0 \
CPU_REPLAY_WEB_CPUS=1 CPU_REPLAY_WORKER_CPUS=1 \
CPU_REPLAY_WEB_SCHEDULERS=1 CPU_REPLAY_WORKER_SCHEDULERS=1 \
  docker compose -f compose.yaml -f compose.cpu-replay.yaml up -d --wait

All five values are explicit experiment inputs. Pinning every stack service to the same one-element CPU set reproduces CPU contention, but does not reproduce an unknown server's memory, storage, network, kernel, or neighbouring workloads. The scheduler inputs make total and online BEAM scheduler counts match the intended replay instead of leaving offline scheduler threads sized from the development host. None of these inputs is a production sizing recommendation. Running ordinary docker compose up without the override removes the replay constraints.

The isolated lifecycle accepts the same override without touching the ordinary Compose project. The explicit VU and replica values below are experiment inputs:

WNH_LOAD_CPU_REPLAY=true \
CPU_REPLAY_CPUSET=0 \
CPU_REPLAY_WEB_CPUS=1 CPU_REPLAY_WORKER_CPUS=1 \
CPU_REPLAY_WEB_SCHEDULERS=1 CPU_REPLAY_WORKER_SCHEDULERS=1 \
WNH_LOAD_WEB_REPLICAS_OVERRIDE=1 \
WNH_LOAD_WORKER_REPLICAS_OVERRIDE=1 \
WNH_LOAD_HTTP_VUS_OVERRIDE=5 \
WNH_LOAD_WS_VUS_OVERRIDE=5 \
WNH_LOAD_AUTH_VUS_OVERRIDE=1 \
LOAD_DURATION_OVERRIDE=30s \
  ./scripts/load-cycle.sh one-cpu-baseline load

After the Helm rollout, the unchanged kind control-plane container measured 1.796 GiB, a reduction of 7.952 GiB (81.6%). Its four application pod cgroups measured 188.5–215.6 MiB in the first post-change sample. A subsequent four-pod rolling replacement retained the 65,536 limit and 786,432-byte table; the replacement cgroups measured 194.3–211.4 MiB, all four replacement pods were Ready, cross-node PubSub passed, 184 readiness samples had no failed sample after retry handling, and the application-table count diff was empty. Evidence is retained under output/memory/20260720-port-table-fix-005420 and output/resilience/beam-port-limit-final-20260720.

In the ordinary 2-web/2-worker Compose profile, the immediate post-recreation sample measured 190.1–205.5 MiB per application container and 168.3 MiB for PostgreSQL. These local observations do not include a target server's OS and do not establish that a 1 GiB or 2 GiB machine is sufficient for production.

On 2026-07-21, an isolated production-compiled compact Compose run measured 198.3 MiB for the combined Phoenix+Oban container and 70.9 MiB for its fresh PostgreSQL/PostGIS container in one no-load snapshot. An isolated split run against a separately started external-database fixture measured 207.2–219.6 MiB for each of two web and two worker containers, 20.05 MiB for the Docker API proxy, 41.66 MiB for Traefik, and 117.6 MiB for that database fixture. These are short workstation observations, not minimums, limits, or capacity claims.

The profile uses a separate Compose project, generated independent secrets, and a separate PostgreSQL volume. Its Traefik instance is constrained to that exact Compose project; its router, service, middleware, and Host rule are unique, so the load profile cannot join the staging router. It exercises:

  • dynamic public home and Safety pages;
  • database-backed readiness queries;
  • Phoenix WebSocket upgrades and protocol heartbeats;
  • real password login and authenticated LiveView joins;
  • consent-driven tracking start, authenticated location writes, private chat writes, and explicit tracking stop/position deletion;
  • the configured number of web and worker replicas;
  • cross-node Phoenix PubSub after each run.

It does not render a real browser, retain a browser tracking session for an unknown real-world trip duration, or reproduce an unknown future production traffic mix. It also does not sample the k6 load-generator container itself. Its results only describe the recorded host, container versions, replica counts, database state, and experiment inputs.

Create the isolated profile

For an ordinary one-run measurement, prefer the lifecycle wrapper. It creates unique Compose and image names, then removes only that run's containers, networks, named volumes, temporary environment, and image tags on success, failure, or interruption:

./scripts/load-cycle.sh local-load load
./scripts/load-cycle.sh local-resilience resilience
./scripts/load-cycle.sh local-complete both

Evidence remains under output/performance/ and output/resilience/. The commands below are the manual diagnostic workflow; unlike the lifecycle wrapper, it deliberately keeps the isolated database volume between commands.

./scripts/ensure-local-load-env.sh
./scripts/load-stack-up.sh

The generated output/runtime/load.env is ignored, restricted to mode 600, and contains independent PostgreSQL and application secrets. Edit its LOAD_* inputs to define a specific experiment. Values in .env.load.example are reproducible measurement points, not recommendations.

For a one-run duration that must not rewrite output/runtime/load.env, pass an explicit experiment override:

LOAD_DURATION_OVERRIDE=10m ./scripts/load-run.sh local-soak-YYYYMMDD

The effective value and whether it came from output/runtime/load.env or the override are recorded in that run's environment.txt.

Run and compare replica counts

./scripts/load-run.sh two-web
./scripts/load-stack-up.sh 3
# Set LOAD_WEB_REPLICAS=3 in output/runtime/load.env so the recorded expected topology
# matches the running topology, then:
./scripts/load-run.sh three-web

Each run writes ignored evidence below output/performance/<label>/:

  • the exact inputs and Docker host observation;
  • a machine-readable k6 summary and complete console report;
  • periodic Docker CPU/memory/network/block-I/O/PID samples for every load stack container, plus a threshold-free CPU/RAM/PID JSON summary;
  • read-only PostgreSQL database size and pg_stat_database snapshots before and after k6, plus their machine-readable counter deltas;
  • pg_stat_statements rankings by total, mean, and maximum execution time, including calls, rows, buffer activity, temporary blocks, and WAL activity;
  • periodic pg_stat_activity connection samples with configured pool slots, observed connection headroom, transaction state, lock waits, and named active wait events;
  • per-web Ecto telemetry deltas separating database execution, pool queue, and result decode time;
  • before/after application table counts, exact authenticated-write evidence, fixture cleanup output, and a cleanup diff;
  • final Compose state, readiness response, application logs, and cross-node PubSub result.

The load-only PostgreSQL command preloads pg_stat_statements and creates the extension only in the isolated load database. The application exposes cumulative Ecto measurements through its authenticated metrics endpoint; the runner snapshots every web replica before and after k6 and records the delta. Neither mechanism changes the ordinary public Compose database.

Before k6 starts, the production-compiled fixture task verifies the exact configured load database name and a deliberate confirmation marker. It creates isolated requester/helper/request/assignment pairs through the application contexts. The pair count equals the maximum simultaneous VU count across all scenarios, because k6 may reuse its global VU pool between parallel scenarios; mapping each global VU identifier directly to its own pair prevents concurrent users from sharing a tracking assignment. The random fixture password exists only in ignored mode-600 output/runtime/load.env and is never written to the manifest or console.

On success and on trapped failure, cleanup deletes only the UUIDs recorded in the validated run manifest. A successful run additionally requires every tracked application-table count after cleanup to equal its pre-fixture count. In a separate negative check, replacing one helper UUID with another valid UUID made cleanup exit nonzero before mutation; all 2 users, 1 request, and 1 assignment remained. Restoring the original manifest then removed those exact rows and returned all three counts to zero.

The k6 image is pinned to version 2.1.0 by digest and runs locally with anonymous usage reporting disabled. No Grafana Cloud account, API token, OpenAI API, or usage-based service is involved.

Observed production public read-only profile

Observed on 2026-07-28 against https://whoneedhelp.com, after the runner verified /srv/who_need_help-production, Compose project who_need_help_production, and deployed commit 3d6f9c5004f9218c55a040bc31e72b56b22d4470. The host exposed 2 CPUs and 4,004,224 KiB total memory. The runner sent only the documented public GET allowlist and Phoenix transport heartbeats; it did not submit forms, log in, register, send email, open product LiveViews, or write application data.

Experiment input / observation Result
Baseline input 1 HTTP VU + 1 WebSocket VU, 30 seconds
Baseline HTTP checks / failures 28 / 0
Baseline HTTP average / p95 91.64 / 107.75 ms
Baseline WebSocket opens / heartbeat replies / errors 6 / 6 / 0
Baseline app average / max CPU 1.26% / 4.60%
Five-client input 5 HTTP VUs + 5 WebSocket VUs, 60 seconds
Five-client HTTP checks / failures / observed rate 1,055 / 0 / 17.27 requests/s
Five-client HTTP average / p95 83.85 / 129.63 ms
Five-client WebSocket opens / heartbeat replies / errors 20 / 20 / 0
Five-client app average / max CPU 7.37% / 10.77%
HTTP-only input 20 HTTP VUs, 30 seconds
HTTP-only checks / failures / observed rate 5,557 / 0 / 184.47 requests/s
HTTP-only average / p95 / maximum 57.18 / 79.25 / 673.45 ms
HTTP-only app average / max CPU 54.70% / 61.53%
HTTP-only app first-to-last memory delta +1 MiB
WebSocket-only input 20 sockets held for 60 seconds
WebSocket opens / heartbeat replies / errors 20 / 40 / 0
WebSocket connect average / p95 184.38 / 191.53 ms
WebSocket-only app average / max CPU 0.79% / 3.82%
WebSocket-only app first-to-last memory delta -0.6 MiB

Readiness returned {"status":"ready"} before and after every valid run. The application container remained healthy with zero restarts. The lowest observed host MemAvailable across the recorded valid runs was 1,241,432 KiB. The reported PostgreSQL RSS is explicitly the sum of every PostgreSQL process on the shared host, not an attribution to this application.

The HTTP-only run recorded 5,562 application HTTP completions and 994 Ecto queries. Their cumulative database execution time increased by 388,598 microseconds and pool queue time by 558,053 microseconds. The two sequential scans of the empty rate_limit_buckets table were the runner's own before/after count snapshots; the HTTP traffic added no rate-limit-table scan. Application table counts were unchanged.

This is not a product-query capacity result. Production contained no help requests or activities during the observation, and /requests and /activities require an authenticated session. PostgreSQL also did not have pg_stat_statements installed. Viewport clustering, authenticated LiveView joins, chat, tracking, handover, and write throughput therefore remain covered by the isolated profiles rather than this public production probe.

The environment measured in this historical production observation had RATE_LIMIT_POLICIES_JSON={}. In that configuration RateLimiter.check/2 returns :not_configured without a database query, so this observation neither exercises nor validates authentication throttling. It also provides no evidence that Redis is needed. The current candidate keeps cross-replica counters in PostgreSQL and ships a pilot email/IP policy; its limiter overhead still requires a separate measured run.

Ignored evidence:

  • output/performance/production-baseline-20260728/
  • output/performance/production-tier2-20260728/
  • output/performance/production-http-db-stats-20260728/
  • output/performance/production-ws-heartbeat-final-20260728/

Isolated rate-limiter observation

Observed locally on 2026-08-03 with the checked-in scripts/rate-limit-benchmark.sh. The application container was limited to one CPU; PostgreSQL reported version 18.4. Both profiles used 5,000 calls with concurrency 20. The first profile deliberately contended on one bucket; the second used 5,000 independent scopes.

Profile Successful / failed / limited Elapsed Observed rate Buckets
One hot scope 5,000 / 0 / 0 12,769.306 ms 391.56 operations/s 1
Distributed scopes 5,000 / 0 / 0 1,312.608 ms 3,809.21 operations/s 5,000

Each run persisted exactly 5,000 increments. The task then deleted its unique benchmark action: cleanup left zero benchmark buckets. The generated JSON is kept in the ignored local directory output/rate-limit/20260802233523-2242041/.

This observation measures the atomic counter implementation on this local host. It is not a production capacity claim and does not justify a hard startup guard. The much slower hot-scope result is expected contention on one database row; ordinary email/IP policy traffic distributes attempts across many scopes.

Observed local comparison

Observed on 2026-07-18 with Docker Engine 29.6.2 on the recorded 32-CPU, 100289265664-byte Docker host. Neither application containers nor the load generator had CPU or memory limits. Both runs used 40 HTTP VUs, 40 WebSocket VUs, a 30-second duration, a 0.1-second HTTP think time, and five-second socket sessions. Each HTTP iteration requested the home page, Safety page, and database-backed readiness endpoint.

Observation 2 web replicas 3 web replicas
HTTP checks 35,160 / 35,160 passed 35,163 / 35,163 passed
HTTP failed rate 0% 0%
Observed HTTP rate 1,169.12 requests/s 1,169.42 requests/s
HTTP duration p95 2.87 ms 3.36 ms
WebSocket opened / heartbeat replies 240 / 240 240 / 240
WebSocket errors 0 0
WebSocket connect p95 46.65 ms 24.12 ms
Average CPU per web container 66.57%, 66.63% 50.25%, 49.77%, 49.18%
Maximum CPU per web container 79.80%, 79.21% 58.16%, 59.23%, 59.08%
Maximum observed memory per web container 268.3, 238.3 MiB 221.0, 223.2, 219.4 MiB

Each resource series contained 11 sequential docker stats --no-stream samples. Both before/after database diffs were empty, both cross-node PubSub probes passed, readiness remained healthy, and the captured proxy/web/worker logs contained no application error or warning. The public staging root and readiness endpoint also returned HTTP 200 while the isolated project was running.

An earlier harness smoke exposed that fixed Traefik router/service names are global to the Docker provider: starting a second Compose project temporarily made staging readiness return HTTP 404. The load project was stopped immediately and staging returned HTTP 200. The final Compose configuration parameterizes the project constraint, router/service name, Docker network, and router rule. Merged-config inspection found no shared router/service label key, the two running proxies showed different exact project constraints, and both origins returned HTTP 200 concurrently before the recorded final runs.

The nearly identical request rate is expected from the selected paced workload: 40 HTTP VUs each issue three requests and then wait 0.1 seconds, giving an approximate workload ceiling of 1,200 requests/s before response time is added. The result therefore does not show the saturation point of two replicas and does not justify an HPA threshold. The lower per-container CPU and lower WebSocket connect p95 with three replicas are observations for this run only; the slightly higher HTTP p95 also shows why one short comparison is not a production sizing study.

Ignored evidence:

  • output/performance/two-web-final-20260718/
  • output/performance/three-web-final-20260718/

Observed authenticated write run

Observed on 2026-07-19 on the same recorded Docker host with 3 web replicas, 2 worker replicas, 40 paced public HTTP VUs, 40 five-second Phoenix heartbeat VUs, and 8 authenticated mutual-aid VUs for 30 seconds. The fixture tool created 88 isolated pairs so every possible global VU had a unique account and assignment.

Observation Result
Public HTTP checks 35,316 / 35,316 passed
All HTTP failures 0 / 39,108
Phoenix heartbeat opened / replies 240 / 240
Authenticated logins 8
Authenticated pages / LiveView joins 1,888 / 1,888
Tracking starts / position writes / stops 1,888 / 1,888 / 1,888
Private messages written 1,888
Authenticated or WebSocket errors 0
HTTP request duration p95 5.29 ms
Authenticated HTTP duration p95 8.10 ms
WebSocket connect p95 3.16 ms

The database evidence counted 1,888 chat rows, 1,888 tracking sessions with 1,888 total samples, 8 login tokens, and zero current positions after explicit stop. Cleanup removed the 176 fixture users, 88 requests, 88 assignments, 1,888 messages, 1,888 tracking sessions, and their audit rows. The post-cleanup application-table diff was empty. Readiness and the five-node web/worker PubSub probe passed, and the run-scoped proxy/web/worker logs contained no application error, warning, HTTP 4xx/5xx, or database deadlock.

Eleven sequential resource samples observed average CPU of 135.63%, 140.01%, and 136.88% for the three web containers, with maximum memory of 273.4, 271.2, and 272.9 MiB respectively. PostgreSQL averaged 37.55% CPU and reached 249.2 MiB; Traefik averaged 47.11% CPU and reached 79.25 MiB. These are short local observations without container resource limits, not minimums, production capacity, or autoscaling inputs.

Ignored evidence:

  • output/performance/final-load-20260719a/

Observed SQL and connection-pool run

Observed on 2026-07-20 with 3 web replicas, 2 worker replicas, PostgreSQL 18.4, 40 paced public HTTP VUs, 40 Phoenix heartbeat VUs, and 8 authenticated mutual-aid VUs for 30 seconds.

Observation Result
HTTP requests / failures 38,429 / 0
Public checks 35,163 / 35,163 passed
HTTP duration average / p95 / maximum 2.42 / 7.74 / 187.11 ms
Authenticated chat/tracking chains 1,664
Phoenix heartbeat opened / replies 240 / 240
Configured Ecto Repo pool slots 50
Peak PostgreSQL client backends 55
Minimum observed non-reserved connection headroom 42
Idle-in-transaction / lock-waiting backends 0 / 0
Ecto average execution / queue / decode time 308.58 / 31.15 / 0.27 µs
PostgreSQL rollback / deadlock / conflict / temp files 0 / 0 / 0 / 0

The connection sampler observed active PostgreSQL waits named Client:ClientRead and IO:WalSync; it observed no lock wait or idle-in-transaction backend. The Ecto queue measurement is the application-side evidence for time waiting to check out a Repo connection. It averaged 31.15 microseconds over 166,459 queries across the three web replicas. This short run did not produce a connection timeout or establish a maximum connection capacity.

The SQL ranking exposed an unbounded preload of every message attached to a request even though the LiveView immediately loaded a separate 50-message cursor page. In the before run that redundant statement executed 3,318 times, returned 342,378 rows in aggregate, and consumed 392.378 ms of PostgreSQL execution time. The request preload was removed and a regression test now requires the chat association to remain unloaded. In the repeat run the unbounded statement was absent; the bounded cursor statement retained its LIMIT, executed 3,328 times, returned 148,512 rows, and averaged 0.122 ms.

No normalized statement in the repeat run used temporary blocks. The largest mean statement time was 0.481 ms for six session-token inserts; the highest-call request lookup averaged 0.031 ms over 39,936 calls. These are observations against the small generated load dataset, not proof that every query remains fast at arbitrary data size. The separate 50,000-row database profile below checks the indexed list and cursor queries at a larger cardinality.

Across the seven measured application/database/proxy containers, simultaneous memory samples totalled 1,501.60–1,509.86 MiB during this run. This excludes the host OS, the k6 generator, Mailpit, and optional observability services. Therefore it is not a claim that this 3-web/2-worker topology fits a 1 GiB or 2 GiB server. It does show that the measured workload used about 1.47 GiB for those seven containers rather than the former kind port-table failure's roughly 10 GiB.

The run created exactly 1,664 messages, tracking sessions, and tracking samples, plus 8 session tokens; explicit tracking stop left zero current positions. Cleanup restored every tracked application-table count, readiness and the five-node PubSub probe passed, and the captured application log had no database timeout, Postgrex error, deadlock, ERROR, FATAL, or PANIC match.

Ignored before/after evidence:

  • output/performance/pool-sql-final-20260720/
  • output/performance/pool-sql-chat-fixed-canonical-20260720/

Observed role-specific connection pools

On 2026-07-20 the same 3-web/2-worker, 40 public HTTP VU, 40 Phoenix heartbeat VU, and 8 authenticated mutual-aid VU profile was repeated with four Repo connections per web replica and two per worker replica. The migrate role also started and completed with a two-connection pool. These are measured development defaults, not production minimums.

The 16 configured long-running Repo slots produced 38,678 HTTP requests with zero failures, 1,760 complete authenticated page/chat/tracking chains, and 240 Phoenix heartbeat sessions. The application recorded 131,409 web-side queries with 34.82 microseconds average pool queue time, compared with 31.15 microseconds in the earlier 10-connections-per-replica run. PostgreSQL observed 21 peak client backends, 6 peak active backends, 76 minimum non-reserved connection headroom, no idle transaction, no lock wait, no rollback, no deadlock, no conflict, and no temporary file.

The run-scoped application log had no match for DBConnection/Postgrex connection errors, pool timeouts, exceptions, deadlocks, FATAL, or PANIC. All 1,760 messages, tracking sessions, and tracking samples were recorded; explicit stop left zero current positions, and fixture cleanup restored the tracked table counts exactly. The result supports using web=4, worker=2, and migrate=2 as the repository's current measured starting profile. It does not establish saturation, an SLO, or a production capacity limit.

Ignored evidence:

  • output/performance/role-pools-4-2-20260720/

Observed 10-minute authenticated soak

Observed on 2026-07-19 with the same 3-web/2-worker topology and the same 40 public HTTP, 40 five-second Phoenix heartbeat, and 8 authenticated mutual-aid VUs. The duration came from the recorded LOAD_DURATION_OVERRIDE=10m; it is an experiment input, not a recommended soak duration or capacity target.

Observation Result
Completed / interrupted iterations 271,072 / 0
Public checks 703,995 / 703,995 passed
HTTP requests / failures 767,225 / 0
HTTP duration average / p95 / maximum 2.34 / 5.71 / 175.96 ms
Phoenix heartbeat opened / replies 4,800 / 4,800
Authenticated logins 8
Authenticated pages / LiveView joins 31,607 / 31,607
Tracking starts / position writes / stops 31,607 / 31,607 / 31,607
Private messages written 31,607
Authenticated or WebSocket errors 0

The database evidence counted exactly those 31,607 messages, tracking sessions, and samples, plus 8 login tokens and zero current positions after explicit stop. Its read-only pg_stat_database delta recorded 2,921,283 commits, zero rollbacks, 12 block reads, 33,967,340 block hits, zero temporary files/bytes, zero deadlocks, and zero conflicts. Database size increased by 16,678,912 bytes during the write workload. The physical database size was recorded but was not treated as a cleanup assertion; the post-cleanup count diff across all tracked application tables was empty. The five-node PubSub probe passed. All seven measured app/database/proxy containers remained running, had zero observed restarts, and were not OOM-killed. The 75,058,636 byte run-scoped log had zero matches for the recorded error, warning, exception, crash, HTTP 4xx/5xx, and deadlock search.

Each container has 199 sequential resource samples. No container resource limit or pass/fail resource threshold was applied.

Container Average / max CPU First 10 / last 10 memory average Max memory
PostgreSQL 61.44% / 85.30% 236.15 / 261.81 MiB 263.60 MiB
Traefik 57.25% / 73.14% 93.44 / 103.98 MiB 108.50 MiB
web 16 198.67% / 281.58% 264.24 / 376.33 MiB 397.10 MiB
web 17 195.66% / 247.46% 259.30 / 373.79 MiB 406.10 MiB
web 18 195.36% / 266.74% 265.06 / 388.06 MiB 410.80 MiB
worker 10 1.47% / 5.09% 201.63 / 202.64 MiB 203.80 MiB
worker 11 2.21% / 5.56% 199.30 / 201.05 MiB 201.80 MiB

The web RSS increased across most of the loaded interval, so this run does not establish a steady-state memory plateau. A separate read-only observation at 2026-07-19T16:06:21Z, 5 minutes 23 seconds after the k6 summary timestamp, measured web RSS at 278.2, 269.8, and 285.6 MiB. At that point the BEAM runtime reported 142,250,344, 147,808,440, and 147,403,872 total bytes with 613, 612, and 612 processes. This observed post-load reduction shows that a substantial part of the loaded RSS was released after the connections closed; it does not prove that the application is free of a longer-term memory leak. A production-like duration and traffic distribution remain unknown.

Ignored evidence:

  • output/performance/authenticated-soak-10m-20260719/

Observed local resilience drills

The canonical Compose drill on 2026-07-19 used the isolated 3-web/2-worker profile. One web and one worker BEAM process exited with status 1 and each container's observed restart count increased to 1. All original replicas were then replaced sequentially. The route returned 758 successful readiness responses with zero final failure and responses from all three web nodes. After replacement, the observed cluster contained all five BEAM nodes and the cross-node PubSub probe passed.

The local Oban probe completed with state completed, attempt 2, max_attempts=2, and exactly one recorded first-attempt error. Its exact row was removed afterward. The load database then contained zero probe jobs, users, requests, messages, and tracking sessions. Run-scoped available logs contained no unexpected application error, warning, HTTP 4xx/5xx, or database deadlock.

The canonical kind drill rolled both 2-replica Deployments from revision 19 to 20 with the chart's observed maxUnavailable=0 and maxSurge=1. All four pod UIDs changed, all replacements were Ready with zero container restart, the four-node BEAM cluster and PubSub probe passed, and the application-table count diff was empty. Of 362 readiness samples, all ultimately returned HTTP 200. Two samples needed one transport retry each while kind's single-node NodePort reset connections during endpoint changes. Those retries are retained in evidence rather than reported as uninterrupted raw TCP connections. This is a local kind observation, not a production availability guarantee.

Ignored evidence:

  • output/resilience/final-resilience-20260719a/
  • output/resilience/final-kind-20260719a/

Observed local monitoring drill

The canonical observability run on 2026-07-19 generated one direct Prometheus target for each of the three current load web containers. The exact expected and observed instance lists matched, and all three targets were up before the induced failure. Grafana reported a healthy Prometheus datasource and served the tracked provisioned dashboard with four panels.

The script then stopped the exact scoped container who_need_help_load-web-16. Prometheus fired WhoNeedHelpWebReplicaUnavailable, and Alertmanager delivered a webhook whose labels identified that instance and the run. After the same container became healthy again, all three direct targets returned to up, the alert cleared, and a resolved webhook with the same instance/run labels arrived. The application-table count diff was empty. These are observed local protocol results, not a production monitoring or notification guarantee.

Ignored evidence:

  • output/observability/final-observability-20260719a/

Stop the isolated containers without deleting their database volume:

./scripts/load-stack-stop.sh

The stop script reads the actual LOAD_PROJECT, refuses the staging project name, and does not delete the volume. Volume deletion is intentionally not automated.

Persistent million-row discovery profile

The write-capable load lifecycle above intentionally cleans its fixtures. For repeatable map, clustering, pagination, and responsive-UI work at high cardinality, a second profile keeps its synthetic database:

./scripts/scale-stack-up.sh
./scripts/scale-seed.sh
./scripts/scale-verify.sh

The profile is pinned to Compose project who_need_help_scale, database who_need_help_scale, and named volume who_need_help_scale_postgres_data. Its generated credentials live only in ignored output/runtime/scale.env. Fixture mutation additionally requires the exact database name and WNH_SCALE_FIXTURE_CONFIRM=persistent-local-scale. The seed task refuses a partially populated database instead of mixing data sets. scale-stack-stop.sh stops only this Compose project and deliberately preserves the named volume. scale-stack-down.sh removes only the profile's containers and networks while also preserving that volume for an explicit operator decision.

The canonical local fixture contains 1,000,000 requests, 1,000,000 activities, 1,000 synthetic request owners, and one scale viewer. Records are distributed across 20 cities and include varied categories, states, urgency, hidden locations, exact public points, and privacy radii. The recorded database size was 1,237,898,943 bytes: 709,197,824 bytes for requests and 505,962,496 bytes for activities including their indexes and TOAST data.

Every scale service is pinned to CPU 0 in this experiment. Three warm application-level samples recorded the following medians:

Viewport Request list / map Activity list / map Map rows / represented records
Kyiv 171.139 / 297.860 ms 265.442 / 700.397 ms 1 / 49,485
Europe 23.072 / 2,302.602 ms 1,544.453 / 4,130.415 ms 8 / 395,878
World 56.182 / 3,975.605 ms 3,910.191 / 7,416.714 ms 10 / 791,753

Server clustering kept each measured map payload between 135 and 1,463 bytes instead of returning hundreds of thousands of points. The broad low-zoom map queries remain the measured optimization target: clustering bounds the network/browser payload, but PostgreSQL still reads and groups the matching records.

A headed Chrome check used the same persistent profile. It verified 24-card initial pages, duplicate-free 48-card pagination, List/Map/Together layouts, 390-by-844 mobile rendering, cluster zoom, and request/activity switching. That check exposed a functional defect: grid-cell coordinates could pan to an empty area after selecting a cluster. Cluster coordinates now use the average of the public points in that group while the snapped cell remains only the grouping and stable-ID key. The full suite then exposed a second privacy regression: an activity close to a narrow viewport could pass the expanded index candidate envelope even when its rounded public point was outside that viewport. The query now retains the indexed candidate filter and also requires an exact intersection against the rounded public point. The final headed check confirmed cluster zoom, an explicit Search this area refresh, 390-by-844 mobile rendering without horizontal overflow, and no application or map console errors. The only console warning was the expected inability to register a service worker under the local self-signed TLS certificate.

The first authenticated HTML response now renders an honest loading shell and starts discovery only after LiveView connects. The discovery runs asynchronously, so a disconnected client no longer leaves its LiveView blocked inside a broad spatial query. During the headed verification, reconnects initially caused three identical logical world queries to overlap (nine active PostgreSQL worker processes). The application now coalesces an identical in-flight discovery key within each BEAM node and reuses a completed result for 10 seconds, which covers the observed 4.5- and 6.5-second local reconnect intervals. Each key includes the viewer, filters, view mode, and viewport; PubSub mutations invalidate the corresponding key. This is a node-local single-flight optimization, not a cluster-wide cache or a capacity guarantee.

The final headed request and activity map checks each represented 791,753 records as 10 clusters. No fresh application 500/503, database pool timeout, discovery error, or container restart was observed. Screenshots are retained in ignored local evidence:

  • output/playwright/scale-million-request-map-desktop-stable.png
  • output/playwright/scale-million-activity-map-desktop-stable.png

Authenticated HTTP load can reuse a restricted browser state without a login or fixture write:

WNH_SCALE_READONLY_VUS=<chosen-count> \
WNH_SCALE_READONLY_DURATION=<chosen-duration> \
WNH_SCALE_READONLY_THINK_SECONDS=<chosen-seconds> \
WNH_SCALE_READONLY_LABEL=<unique-label> \
  ./scripts/scale-readonly-load.sh

The k6 file has a compiled four-route GET allowlist for the request/activity static LiveView shells in Europe and the world. It intentionally does not claim to measure connected LiveView discovery: scripts/scale-verify.sh measures the actual list and clustering queries separately. The runner verifies the exact Compose project and database, uses the stack's internal network, and compares counts for users, tokens, requests, activities, assignments, participants, messages, and rate-limit buckets before and after.

The final 2026-07-28 one-VU, 30-second, one-CPU static-shell observation completed 60 authenticated GETs with 0 HTTP failures and 120/120 checks.

Keeping the fixtures in the ordinary development site

The isolated profile is useful for destructive load rehearsals. For routine development of the actual whoneedhelp.imalto.site UI, the same deterministic fixtures can instead be appended to the already running who_need_help development database:

./scripts/dev-scale-seed.sh

This guarded path requires the running Compose project and services to be named who_need_help, who_need_help-web-1, and who_need_help-db-1; verifies that the target database is exactly who_need_help; and uses the separate confirmation value persistent-dev-scale. It preserves ordinary development users and records. Synthetic rows carry structured_data.scale_fixture=true, use deterministic UUIDs, and are inserted with conflict handling, so rerunning the command fills missing fixtures without duplicating them. The resulting summary is ignored at output/dev-scale/summary.json. Overall HTTP duration was 2.566 ms median, 3.196 ms average, and 8.709 ms maximum. Every tracked table count was identical before and after. The discovery-query medians are the independently measured values in the table above. These are measurements of one local CPU-contention replay, not a production SLO, capacity claim, or minimum resource requirement. Ignored evidence is retained under output/scale/verification.json and output/performance/scale-million-onecpu-20260728-async-shell/.

Isolated database scale and cursor plans

Run the database-only before/after measurement with:

./scripts/db-scale-benchmark.sh

The script builds the test image and creates a uniquely named Compose project, random one-run PostgreSQL credentials, and a separate PostGIS volume. It migrates only through 20260719004249, seeds the isolated baseline, records machine-readable PostgreSQL EXPLAIN (ANALYZE, BUFFERS, FORMAT JSON) plans, applies the generated 20260719013320_add_keyset_pagination_indexes migration, and records the same queries again. It also verifies that two 25-row keyset pages exactly equal the first 50 ordered rows without duplicates or gaps. The exact project, network, volume, and image are removed by its trap.

DB_SCALE_ROWS controls the configured sample size and defaults to 50,000. That default is a reproducible experiment input, not a minimum database size, capacity claim, latency target, or production traffic model. The structural gate checks the canonical sample's chosen cursor indexes and query correctness; it does not fail on an arbitrary millisecond threshold.

Observed locally on 2026-07-19 with PostgreSQL 18.4, the canonical run created 50,000 rows in each of users, requests, assignments, messages, reviews, reports, abuse signals, proposals, activities, and activity messages; it created 50,500 activity participants and 49,998 block rows. The isolated database occupied 231,241,407 bytes after the cursor migration.

Query Baseline observed Cursor migration observed Selected cursor index
Urgent-help discovery 8.010 ms 0.110 ms help_requests_discovery_cursor_index
Requester history 7.969 ms 0.036 ms help_requests_requester_cursor_index
Activity discovery 8.137 ms 0.051 ms activities_discovery_cursor_index
Visible reviews 8.264 ms 0.052 ms reviews_visible_cursor_index
All reports 8.021 ms 0.025 ms reports_cursor_index
All category proposals 6.501 ms 0.025 ms category_proposals_cursor_index
Moderation users 5.367 ms 0.030 ms users_moderation_cursor_index
Blocks 5.165 ms 0.036 ms blocks_blocker_cursor_index

These are measurements of one warm local run and must not be interpreted as an SLO or portable speedup. The measurement also exposed redundant candidate indexes and an OR-based Activity membership query. The final migration keeps the indexes PostgreSQL selected, uses partial discovery indexes for open, non-hidden records, and the application joins the existing participant membership invariant directly. Personal reputation and the leaderboard now aggregate in PostgreSQL instead of loading all completed assignments into the BEAM; the leaderboard itself uses composite keyset pagination.

Ignored evidence for the recorded run:

  • output/db-scale/20260719054559-2339191/

Current final replay

The same threshold-free profiles were replayed on 2026-07-20 against application commit b96d443.

  • The 30-second 3-web/2-worker load run used 40 public HTTP, 40 heartbeat WebSocket, and 8 authenticated VUs. It completed 13,765 iterations and 38,826 HTTP requests; all 35,280 checks passed and HTTP failures were 0. The authenticated scenario completed 1,765 page/LiveView/chat/tracking start-update-stop chains, and the heartbeat scenario opened and answered 240 sockets. PostgreSQL recorded no rollback, deadlock, conflict, or temporary file. The run-specific fixtures were removed and the cleanup count diff was empty. Evidence is output/performance/final-load-20260720.
  • The database benchmark again used 50,000 configured rows per large table on PostgreSQL 18.4. After migration, the asserted list queries selected their cursor indexes; for example, observed execution time changed from 8.139 ms to 0.127 ms for urgent-help discovery and from 8.035 ms to 0.050 ms for Activity discovery. These are single local observations, not thresholds or portable performance claims. Evidence is output/db-scale/20260719233526-97476.
  • The corrected kind rolling run replaced all four application pod UIDs, retained four Ready zero-restart replacements, passed the four-node BEAM and PubSub probes, and produced an empty database diff. All 395 readiness samples ultimately returned HTTP 200; two required one retry during local NodePort endpoint replacement. Evidence is output/resilience/final-kind-rollout-fixed-20260720.

Runtime hardening and queue replay on 2026-07-20

The ordinary and isolated load releases were rebuilt after separating internal cluster, ingress, and outbound networks. All web and worker node names used internal 172.16.52.x or 172.16.55.x addresses respectively, every node observed the expected peers, and cross-node PubSub passed. The worker runtime no longer starts the unused default queue. Both measured load workers reported exactly maintenance: 2 and push: 1; their BEAM process counts were 574 after the change, compared with the earlier observation of 577/578. This three-to-four-process difference is an observation, not a memory or capacity guarantee.

The isolated Oban burst experiment inserted 1,000 side-effect-free jobs in one batch. With two workers and configured maintenance concurrency two per worker, all 1,000 reached completed in an observed 4,516 ms. Three resource samples per measured container recorded maxima of 80.80 MiB for PostgreSQL and 218.80/186.90 MiB for the two worker cgroups. No error matched the run-scoped logs; the database reported zero deadlocks, temporary files, or conflicts. The script removed exactly 1,000 probe rows, verified zero remaining probe jobs, and found byte-identical before/after row-count documents for all domain tables. These short local samples do not establish a production throughput, steady-state memory plateau, or minimum server size. Evidence is output/performance/oban-burst-1000-final-20260720.

The database-scale harness was repaired to build and pass its own hardened PostGIS image rather than relying on an image variable owned by another script. Its fresh PostgreSQL 18.4 replay again used 50,000 configured rows per large table and identical before/after table counts. Every asserted list plan selected its intended cursor index. Observed examples were 8.806 to 0.123 ms for urgent-help discovery, 8.348 to 0.061 ms for Activity discovery, and 9.084 to 0.048 ms for visible reviews. The aggregate leaderboard and single hot helper reputation plans still scanned their 50,000-row completed-assignment working sets and took 53.949 and 41.247 ms in this deliberately concentrated fixture. That is a measured future optimization target, not evidence of a current failure or a portable latency. Evidence is output/db-scale/20260720035817-2416501.

Final queue/runtime load replay

After the queue and network changes, the same isolated 3-web/2-worker profile was replayed on 2026-07-20 with 40 paced public HTTP VUs, 40 Phoenix heartbeat VUs, and 8 authenticated mutual-aid VUs for 30 seconds.

Observation Result
HTTP requests / failures 38,962 / 0
Public checks 35,280 / 35,280 passed
Authenticated page/chat/tracking chains 1,833
Phoenix heartbeat opened / replies 240 / 240
HTTP duration average / p95 / maximum 1.91 / 4.90 / 179.46 ms
Authenticated HTTP duration p95 7.77 ms
Configured long-running Repo pool slots 16
Peak PostgreSQL client / active backends 21 / 4
Minimum observed non-reserved connection headroom 76
Ecto average execution / queue / decode time 313.90 / 29.75 / 0.14 µs
PostgreSQL rollback / deadlock / conflict / temp files 0 / 0 / 0 / 0

The load generated exactly 1,833 chat rows, tracking sessions, and tracking samples; explicit stop left zero current positions. Cleanup removed the 176 fixture users, 88 requests, 88 assignments, 1,833 messages, 1,833 tracking sessions, and their audit rows. The before/after application-table count diff was empty, all five BEAM nodes participated in the PubSub probe, and every load service remained running.

Ten sequential resource samples recorded maxima of 268.5, 260.4, and 284.3 MiB for the three web cgroups; 207.2 and 198.7 MiB for the workers; 119.9 MiB for PostgreSQL; and 94.85 MiB for Traefik. These short samples had no memory limit and do not establish a steady-state plateau or server minimum. No captured query used temporary blocks. The highest mean SQL statement was a single 5.249 ms tracking-session maintenance update; the highest-total statement was the indexed session-token lookup at 0.081 ms mean across 7,332 calls. No measured query result justifies another schema or pool change at this data size.

Ignored evidence:

  • output/performance/final-after-queue-20260720/