Files
felhom.eu/documentation/audits/update-night-2026-09-21/18-postgres-refusal-reproduced.txt
T
admin 9c69b3ff07
gates / gates (push) Successful in 27s
Update night: Phases 2-4 evidence — both engines, the unattended HOLD, and five new findings
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows.

PHASE 2 — the two database engines, through the REAL Update button:
- MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time.
  All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine
  itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says
  "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the
  `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own
  pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back.
- PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured.
  5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s.
  The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it
  was REPRODUCED INDEPENDENTLY with a control on every step (R-320).

PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the
caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed
nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed
there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in
`backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found
R-458's risk narrower than the row states.

PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions.

FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 —
two templates name a health probe the app does not answer, and because the guarded update waits
on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served
HTTP 200 on the new version at four samples across five minutes and was then stopped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:13:57 +02:00

33 lines
1.7 KiB
Plaintext

=== 0. find the live docmost 16 datadir, and PROVE it is 16 before touching anything
mounts of docmost-postgres:
volume docmost_docmost_postgres_data /var/lib/docker/volumes/docmost_docmost_postgres_data/_data /var/lib/postgresql/data
volume: docmost_docmost_postgres_data
PG_VERSION of the SOURCE: '16'
=== 1. copy it, and PROVE the copy is 16 too
copied
PG_VERSION of the COPY: '16'
size of the copy: 49.0M
=== 2. THE MEASUREMENT: postgres:17-alpine on that 16 datadir — the household's exact case
container: running=false exit=1 restarts=0
--- POSTGRESQL 17's OWN WORDS, VERBATIM:
PostgreSQL Database directory appears to contain a database; Skipping initialization
2026-09-21 22:11:32.601 CEST [1] FATAL: database files are incompatible with server
2026-09-21 22:11:32.601 CEST [1] DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.
--- the copy's PG_VERSION AFTER 17 refused it (must still be 16 — nothing was migrated):
16
=== 3. POSITIVE CONTROL: the same copy under postgres:16-alpine must start AND hold the data
container: running=true exit=0
2026-09-21 22:11:43.497 CEST [33] LOG: checkpoint starting: end-of-recovery immediate wait
2026-09-21 22:11:43.583 CEST [33] LOG: checkpoint complete: wrote 600 buffers (3.7%); 0 WAL file(s) added, 0 removed, 0 recycled; write=0.014 s, sync=0.061 s, total=0.089 s; sync files=409, longest=0.025 s, average=0.001 s; distance=2838 kB, estimate=2838 kB; lsn=0/1BEB798, redo lsn=0/1BEB798
2026-09-21 22:11:43.592 CEST [1] LOG: database system is ready to accept connections
--- the DATA, asked of the engine itself:
tables in the public schema: 48
=== 4. teardown, BY NAME
removed: DRILL-pg17-refuse, DRILL-pg16-ok, volume DRILL-pg16-copy