The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:` line is proven identical to before. WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN front door, with a negative control on every readback. Ten of the fourteen printed a verbatim migration line. Up from the three apps this project had ever measured. THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples across five minutes, docker's own healthcheck green, and was then stopped and the household sent to a restore they did not need. zipline and wger are the same defect, both confirmed live. The gate that catches all three is static and cheap: both health checks already sit in the same file. WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) — which needed a purpose-built image store, because the rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. MariaDB across a major through the real button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted, with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at once and all ended honest. And the two EARLY power-cut phases nobody had cut in. TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English household in English and they do not, and R-446/R-458 are both narrower than their rows state. Two instrument fixes were needed before anything could be trusted: the unattended caller turned every success into a timeout (R-623), and one of my own reproductions was wrong and is kept labelled with what it actually measured. Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond the floor the operator asked for. Gates: repo_gates.py --fast, all 15 OK. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3.5 KiB
REPORT — UPDATE NIGHT, 2026-09-21
The full record is documentation/audits/DRILL-update-night-2026-09-21.md. This file is the
session report: what ran, what shipped, what is owed.
Not done, or changed from the brief
Nothing in the brief was skipped. Five things were changed, re-run or measured on a different
venue, each named with its reason in the audit's own first section. In short: the PostgreSQL
rehearsal ran on guest 9202 rather than a separate harness LXC; the pg_upgrade route was not run
(it needs an image that does not exist here); B5's safety-dump cut MISSED first and was recorded
as a miss before being retried and hit; B8 and the rehearsal were re-run after B1's own precondition
swept the app they needed; and the harness RUNS of the new catalog edges are owed although the code
is shipped.
One thing the brief asked for that this venue cannot produce at all: every event and every
customer mail. Guest 9202 runs hub.enabled: false and the notifier returns before it logs
(R-620). Stated on every row of the alarm truth table rather than left blank.
What ran
- Phase 0 — the fleet floor to 0.261.0 (both demo boxes in 13 s, hub
SERVED … from declared); a private drill catalog with a positive and two negative controls; a throwaway image store on the scratch guest; capacity measured; the upstream drift re-run. - Phase 1 — 21 edges across 19 apps walked on guest 9202 through the product's own guarded Update, each seeded and read back through the app's own front door.
- Phase 2 — the two database engines across a major, through the real Update button.
- Phase 3 — the bad days, B1–B9.
- Phase 4 — the morning after.
- Phase 5 — teardown, three layers, plus Gitea.
What shipped
app-catalog-felhom.eu@4463243f2e09 — TEST CODE ONLY: four new harness fixtures and seven new edges (U1–U7). No template changed; noimage:line moved. Gates green; CI job 830 = success.felhom.eu— this report, the audit, the evidence, the register rows, the architecture updates, and one correction toupdate-arc-gaps-2026-09-21/00-api-recipe.md(the app page is/apps/<n>, not/app/<n>) and one FIX tounattended-caller.py(R-623).- No controller, agent or hub code was written. The brief forbade it and none was needed.
What is owed
- The harness RUNS of edges U1–U7. The code is in the catalog repo and the gates are green; the runs, and with them the per-app ABORT answers, have not been performed.
- A cut inside
startingitself. Both EARLY phases were cut tonight;startinglasts well under a second and still needs an in-process fault injector rather than a faster shell. - The mail half of Q4, and every event: structurally unmeasurable on this venue (R-620).
- Fixtures for the four inconclusive apps — and for two of them (vaultwarden, zipline) the honest
maximum is
inconclusivewhile the catalog rightly closes their sign-up (R-624). - What re-created the removed
navidromecontainer (R-626): observed, not diagnosed, because the controller had restarted and its log no longer reached that moment. wger's own edge — it was deployed only to measure its probe and was then removed.
The live catalog
Never touched with a broken, dummy, cross-repo or engine-major reference — not once, not for
thirteen minutes. Its main moved only for the harness commit above, which changes scripts/
and zero image: lines; the teardown diff proves every image: line identical to the drill copy.