INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320), not at the end of the session. Phases 2-5 follow in a later commit. Phase 0, all three mechanisms proven with their controls: - the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub logging `managed floor SERVED ... from declared (golden 0.258.0)`. - a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and engine-major edges can be measured without the live catalog ever carrying one. Positive control quoted, and two negative controls: the live catalog's main and both real boxes' caches unchanged. - a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD measurable at all: an edge that PASSES the within-a-major test and still fails. CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases + 1 negative control), not by reading it. Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door (R-156), with a per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`. TWO INSTRUMENT FIXES, both in this repo's own evidence code: - 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described 404s. Corrected, with the session-expiry note that cost the same time. - unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were always None and EVERY followed update ran to its 900 s timeout and was then recorded `timeout` and never-press-again. Fixed before B1 relied on it. R-623. No controller, agent or hub code was written. The live catalog carries no broken reference. Gates: repo_gates.py --fast — all 15 OK, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3.9 KiB
Draft section for 09-update-architecture.md — the standing method for update drills
(To be inserted after §6.4. Written here first so the audit and the architecture stay in step.)
6.5 The drill catalog and the image store — the standing method for update drills
Why this section exists. On 2026-09-21 an afternoon session put a deliberately broken image into the LIVE catalog for thirteen minutes to produce a failing edge. It was reverted and nothing reached a customer, but the method was wrong and the brief that asked for it said so. This is the method that replaces it, proven the same night.
The rule, and it has no exception: nothing broken, dummy, cross-repo or engine-major ever enters the live catalog — not as a fallback, not for thirteen minutes. If a leg cannot be done without that, the leg is skipped and named.
The two mechanisms
| what it is | what it makes possible | |
|---|---|---|
| the drill catalog | admin/app-catalog-drill on Gitea — private, a copy of the live catalog's main |
a scratch box can be pointed at a catalog where a failing edge is committable, because it carries none of the live repo's gates |
| the image store | a registry:2 container on the scratch guest at 127.0.0.1:5000 |
an edge that passes the within-a-major test and still fails — the one shape a real catalog move cannot produce |
The image store is not a convenience. 09 §3b Q4 could not be measured for a year of drills
because the only failing edges available were across-a-major, and the within-a-major rule — correctly
— refuses those before the guarded update is ever reached. The rule that makes automatic updates
safe is the same rule that refuses the obvious way to break one. Measuring an unattended HOLD needs
drill/<app>:X.Y.Z (the real image, retagged) against drill/<app>:X.Y.(Z+1) (a built image that
starts, stays up and never serves) — same repository, same major, plain version tags. A third
flavour, a tag simply absent from the store, gives the pull-failure leg.
stacks.CompareImageRefs orders a host:port/ reference correctly: splitImageRef takes the last
colon and rejects it only when a / follows, so a registry port is never read as a tag. Proven by
running it, four positive cases and a negative control, 2026-09-21.
Pointing a box at the drill catalog — the step that is NOT obvious
git.repo_url alone is inert. Syncer.gitCloneOrPull clones only when the cache has no .git;
otherwise it fetches from the remote the clone already stores. The cache directory must be removed as
well, or the box goes on following the live catalog and reports success. Filed as R-615; until it
is fixed, the drill procedure is:
- save
controller.yamlascontroller.yaml.pre-update-night; - set
git.repo_url(andusername/token— the drill repo is private); - remove
<data>/catalog-cache; - restart the controller, sync, rescan (R-607: a sync can answer „nincs változás" while the catalog has moved, and the badge answers from the stale value until the rescan);
- three controls, all quoted in the report — the drill bump appears on the scratch box; the
other boxes' caches are unchanged; the live catalog's
mainhash is unchanged.
What the drill must leave behind
controller.yamlrestored from the saved copy, the controller restarted, andgit.repo_urlread back and quoted as the live catalog.- The registry container and its volume removed; drill images removed by name. Never
prune. - The drill repo kept, private, reset to the live catalog's
main, so the next drill starts clean. - A diff of every
image:line against the live catalog'smain— expected: identical.
The fence
Only a scratch guest is ever pointed at the drill catalog. The drill repo's README says so, and no customer box has credentials for it. The store listens on the guest's loopback only.