Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows. PHASE 2 — the two database engines, through the REAL Update button: - MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time. All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back. - PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured. 5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s. The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it was REPRODUCED INDEPENDENTLY with a control on every step (R-320). PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in `backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found R-458's risk narrower than the row states. PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions. FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 — two templates name a health probe the app does not answer, and because the guarded update waits on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served HTTP 200 on the new version at four samples across five minutes and was then stopped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2.7 KiB
Draft — the line to append to the capability map's UPDATE row (00-capability-map.md line 105)
(Appended to the existing "Update is GUARDED …" row's evidence column. <…> filled at the end.)
WIDENED 2026-09-21 (the update night) from 3 apps to <N>, and NARROWED in one place by the same
run. audits/DRILL-update-night-2026-09-21.md. On scratch guest 9202 (controller v0.261.0),
against a private drill catalog so the live catalog carried no test reference at any point,
<N> real within-a-major upstream edges were walked through the product's own guarded Update, each
app seeded and read back through its own front door (R-156) with a negative control on every
readback: <P> proven, <F> failed, <I> inconclusive.
What the PROVEN edges prove, precisely: the app moved, the four version observables agreed, and
the data the app itself was given came back through the app's own interface afterwards. <MIG> of
them printed a verbatim migration line.
What the FAILED edges prove, and they are the more valuable half. <FAILED_LIST>.
adventurelog v0.12.1 → v0.13.0 applied nine database migrations successfully and then never
bound its port; the update held after the full health wait, the hold sentence named the tier, the
date and what the copy holds, and the restore the sentence names brought the app back. That is
this row's own promise, exercised on a real upstream edge rather than a staged one.
AND THE NARROWING, which this row must carry because it is the same mechanism: the verifying
phase trusts the .felhom.yml probe absolutely, and two of the 53 templates name a probe the app
does not answer — tandoor (port 8080; it listens on 80) and zipline (/api/health; it answers
404 there, while the compose healthcheck in the same file uses /api/healthcheck and is green).
For those apps a successful update is stopped by its own health wait: tandoor was measured
serving HTTP 200 on the new version at four samples across five minutes, with docker's own
healthcheck green, and was then stopped by failAndHold and the household sent to a restore they
did not need. R-618, P1. No data was lost and the restore works — but "the update is guarded"
must not be read as "the guard is right about whether the app came up".
Still true and unchanged: no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly.
Not measured on this venue, and named rather than assumed: every event and every customer mail.
Guest 9202 runs hub.enabled: false and the notifier returns before it logs (R-620), so the
whole "who was told" half of 08 was structurally unobservable tonight.