INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320), not at the end of the session. Phases 2-5 follow in a later commit. Phase 0, all three mechanisms proven with their controls: - the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub logging `managed floor SERVED ... from declared (golden 0.258.0)`. - a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and engine-major edges can be measured without the live catalog ever carrying one. Positive control quoted, and two negative controls: the live catalog's main and both real boxes' caches unchanged. - a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD measurable at all: an edge that PASSES the within-a-major test and still fails. CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases + 1 negative control), not by reading it. Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door (R-156), with a per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`. TWO INSTRUMENT FIXES, both in this repo's own evidence code: - 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described 404s. Corrected, with the session-expiry note that cost the same time. - unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were always None and EVERY followed update ran to its 900 s timeout and was then recorded `timeout` and never-press-again. Fixed before B1 relied on it. R-623. No controller, agent or hub code was written. The live catalog carries no broken reference. Gates: repo_gates.py --fast — all 15 OK, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
4.3 KiB
Draft — facts to add under 09 §3b, Q2–Q6
The questions stay the operator's. No recommendation changes below; where one is strengthened or
weakened by tonight's measurement, that is said in those words and the recommendation itself is left
exactly as written. Numbers marked <…> are filled from the verdict records at the end of the run.
Under Q2 — May an automatic update run on a bind-data app when no copy holds its FILES?
MEASURED 2026-09-21 (update night). The hold sentence Q2 turns on was read verbatim off a real
failure, not from source. adventurelog v0.12.1 → v0.13.0 held and said:
„A(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."
So the machinery Q2's first option would key on exists and works: the sentence names the tier, the date, and what the copy holds, and it does so on a real edge with no prompting. The question of whether the AUTOMATIC rule should differ from the button's is untouched by this — it is still a choice, and it is still the operator's.
And one thing Q2 did not ask, which tonight makes urgent: after the hold, nobody can find out WHY.
failAndHold removes the containers, so the failing version's own output is gone within seconds
(R-621). With a person pressing, they at least watched it happen. With nobody pressing, the only
account of the night is a sentence that says the app did not come up.
Under Q3 — What counts as "within a major" when the tag is not a version number?
MEASURED 2026-09-21, by running the comparator rather than reading it. CompareImageRefs orders
a reference carrying a host:port/ prefix correctly — splitImageRef takes the LAST colon and
rejects it only when a / follows, so a registry port is never mistaken for a tag. Four positive
cases and one negative control (different repositories → not orderable). This matters because it is
what made the whole unattended-hold leg possible: the drill edge localhost:5000/drill/glance:1.0.0 → :1.0.1 passes the within-a-major test and still fails, which no real catalog move does.
The recommendation is unchanged. The extension it already names — expose the parsed major from the same normaliser — is still owed.
Under Q4 — A held app: who is told, when, and does the box try again?
MEASURED 2026-09-21 (update night), and this is the half that was missing.
Under Q5 — PostgreSQL: what has to exist before the catalog may move postgres:16 to 17?
MEASURED 2026-09-21, both halves, on a real seeded datadir.
(a) What a household would see today.
(b) The conversion rehearsal, costed.
(c) A fact about the instrument, not the engine. The harness's own PostgreSQL probe is
cat /var/lib/postgresql/data/PG_VERSION inside the container (upgrade-test.py ENGINE_PROBES).
The recommendation is unchanged — a scripted conversion edge proven on all eleven before the catalog may move, and the gate stays until it lands. Tonight gives it a price rather than a new opinion.
Under Q6 — Should the catalog record each pin's DIGEST at push time?
MEASURED ON A BOX 2026-09-21 (update night), leg B8. §8.1's numbers came from a registry sweep run on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders.
The recommendation is unchanged — yes, and it is still the cheapest real improvement on the list.
Not a question, but it belongs beside them
The night could not measure a single event or a single customer mail, because the scratch guest
runs with hub.enabled: false and every notifier entry point returns before it logs anything
(R-620). Every Q4-shaped question about who is told therefore rests tonight on what the
household READS — the app page, the dashboard, the backups page — and not on what is SENT. Recorded
as the limit it is: the venue that is safe enough to break apps on is the one that cannot mail
anybody, and that is not a coincidence to design around silently.