Files
felhom.eu/documentation/audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md
T
admin 8d786f7940
gates / gates (push) Successful in 28s
Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:`
line is proven identical to before.

WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own
guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at
any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN
front door, with a negative control on every readback. Ten of the fourteen printed a verbatim
migration line. Up from the three apps this project had ever measured.

THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not
answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by
STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples
across five minutes, docker's own healthcheck green, and was then stopped and the household sent
to a restore they did not need. zipline and wger are the same defect, both confirmed live. The
gate that catches all three is static and cheap: both health checks already sit in the same file.

WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) —
which needed a purpose-built image store, because the rule that makes automatic updates safe is
the same rule that refuses the obvious way to break one. MariaDB across a major through the real
button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted,
with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at
once and all ended honest. And the two EARLY power-cut phases nobody had cut in.

TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was
measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English
household in English and they do not, and R-446/R-458 are both narrower than their rows state.

Two instrument fixes were needed before anything could be trusted: the unattended caller turned
every success into a timeout (R-623), and one of my own reproductions was wrong and is kept
labelled with what it actually measured.

Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond
the floor the operator asked for.

Gates: repo_gates.py --fast, all 15 OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:34:24 +02:00

4.8 KiB

09 §3b Q5 — PostgreSQL 16 → 17: what a household sees today, and what a conversion costs

Both halves measured 2026-09-21, guest 9202, on a REAL app with REAL seeded data.


(a) What a household would see TODAY — measured, and it is what R-463 predicted

docmost, seeded through its own API and read back first (control C1). Drill-catalog bump of the postgres: sidecar alone, 16-alpine → 17-alpine; the app image did not move. The guarded Update was pressed.

time to the verdict 5.1 s — the engine does not try, it refuses at once
final phase failed, app stopped and held
the pin postgres:17-alpine — while nothing is running on 17
installed_images still postgres:16-alpine — the observation and the decision disagree, correctly (§5.2)
the data intact
the way out the restore the hold sentence names: 29.1 s, hold cleared, app back, health probe 200

The engine's own refusal line had to be REPRODUCED, because failAndHold removed the container before any probe could read it (R-621) and the controller log does not carry it either. Reproduced independently, with a control on every step — source proven 16, copy proven 16, 49 MB:

FATAL:  database files are incompatible with server
DETAIL: The data directory was initialized by PostgreSQL version 16,
        which is not compatible with this version 17.11.

And the datadir was still 16 afterwards — nothing was migrated, nothing was damaged. Positive control: the same copy under postgres:16-alpine starts and holds 48 tables.

So R-463's reading is confirmed on the box: the engine refuses to start, the update ends HELD, the data is intact. Nothing else happened, and the household's route back works.


(b) The conversion rehearsal, COSTED

Route: logical dump and restore. Plain docker beside the product — there is no product path for this, and pricing one is the point. On a fresh, seeded docmost: 49.0 MB datadir, 48 tables.

step time what it produced
dump with 16 (pg_dumpall) 2.6 s 132 201 bytes, 48 CREATE TABLE statements
fresh 17 datadir + restore 6.5 s PG_VERSION 17, 48 tables restored, 2 benign ERROR lines
point the app at 17 and start it 124.8 s „Database connection successful" — the app's own words
the seed read back on 17 — TRUE, through the app's own login
the 17 datadir afterwards — 49.1 MB (from 49.0 MB)
total 155.9 s of which ~9 s is the engine work; the rest is the app restarting

The two ERROR lines are benign and are named so nobody reads them as data loss: role "docmost" already exists and database "docmost" already exists — pg_dumpall recreates both, and the 17 container's entrypoint had already made them. The 48-table count after the restore is the positive control that the replay worked.

What Q5 now has that it did not

  • A price. Nine seconds of engine work for a 49 MB database, and under three minutes end to end including the app restart. For eleven apps that is a maintenance window, not a project.
  • A shape that works, walked once on real data: stop the app, keep the engine, pg_dumpall, fresh 17 volume, replay, re-point, start, read the data back through the app's own door.
  • What could lose data, named: the dump is the single point of failure. Nothing in the rehearsal verified the dump before the old datadir was left behind — because nothing had to, since the old volume was untouched throughout. Any real procedure must keep the 16 datadir until the app has been read back on 17, which is exactly what this rehearsal did by accident of being a rehearsal.
  • The pg_upgrade route was NOT run. It needs both majors' binaries in one image and no such image exists in this project. Naming it costs nothing; building it is the work Q5's first option is really asking for, and the logical route above may make it unnecessary at this size.

This is a rehearsal and a costing, not a procedure for the catalog. The engine-major gate stays.

One fact about the harness's own instrument, since Q5's option 1 rests on it

upgrade-test.py's PostgreSQL probe is cat /var/lib/postgresql/data/PG_VERSION inside the container. Against the converted datadir it answered 17, exit 0 — it works. But it is blind in exactly the case that matters: when PostgreSQL refuses, the container is not running, so docker exec cannot ask it anything. The probe's own honesty rule covers this ("a probe that cannot run records why"), and tonight it recorded No such container. A probe that can only speak when the engine is happy is worth having — but it must never be read as the engine is content.