Files
felhom.eu/documentation/audits/update-night-2026-09-21/09-SECTION-drill-method.md
T
admin da20722e76
gates / gates (push) Successful in 27s
Update night 2026-09-21: Phase 0 and Phase 1 evidence, the drill method, and two instrument fixes
INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320),
not at the end of the session. Phases 2-5 follow in a later commit.

Phase 0, all three mechanisms proven with their controls:
- the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub
  logging `managed floor SERVED ... from declared (golden 0.258.0)`.
- a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and
  engine-major edges can be measured without the live catalog ever carrying one. Positive
  control quoted, and two negative controls: the live catalog's main and both real boxes'
  caches unchanged.
- a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD
  measurable at all: an edge that PASSES the within-a-major test and still fails.
  CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases
  + 1 negative control), not by reading it.

Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own
guarded Update, each app seeded and read back through its OWN front door (R-156), with a
per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`.

TWO INSTRUMENT FIXES, both in this repo's own evidence code:
- 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described
  404s. Corrected, with the session-expiry note that cost the same time.
- unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were
  always None and EVERY followed update ran to its 900 s timeout and was then recorded
  `timeout` and never-press-again. Fixed before B1 relied on it. R-623.

No controller, agent or hub code was written. The live catalog carries no broken reference.

Gates: repo_gates.py --fast — all 15 OK, exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 21:17:46 +02:00

3.9 KiB

Draft section for 09-update-architecture.md — the standing method for update drills

(To be inserted after §6.4. Written here first so the audit and the architecture stay in step.)


6.5 The drill catalog and the image store — the standing method for update drills

Why this section exists. On 2026-09-21 an afternoon session put a deliberately broken image into the LIVE catalog for thirteen minutes to produce a failing edge. It was reverted and nothing reached a customer, but the method was wrong and the brief that asked for it said so. This is the method that replaces it, proven the same night.

The rule, and it has no exception: nothing broken, dummy, cross-repo or engine-major ever enters the live catalog — not as a fallback, not for thirteen minutes. If a leg cannot be done without that, the leg is skipped and named.

The two mechanisms

what it is what it makes possible
the drill catalog admin/app-catalog-drill on Gitea — private, a copy of the live catalog's main a scratch box can be pointed at a catalog where a failing edge is committable, because it carries none of the live repo's gates
the image store a registry:2 container on the scratch guest at 127.0.0.1:5000 an edge that passes the within-a-major test and still fails — the one shape a real catalog move cannot produce

The image store is not a convenience. 09 §3b Q4 could not be measured for a year of drills because the only failing edges available were across-a-major, and the within-a-major rule — correctly — refuses those before the guarded update is ever reached. The rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. Measuring an unattended HOLD needs drill/<app>:X.Y.Z (the real image, retagged) against drill/<app>:X.Y.(Z+1) (a built image that starts, stays up and never serves) — same repository, same major, plain version tags. A third flavour, a tag simply absent from the store, gives the pull-failure leg.

stacks.CompareImageRefs orders a host:port/ reference correctly: splitImageRef takes the last colon and rejects it only when a / follows, so a registry port is never read as a tag. Proven by running it, four positive cases and a negative control, 2026-09-21.

Pointing a box at the drill catalog — the step that is NOT obvious

git.repo_url alone is inert. Syncer.gitCloneOrPull clones only when the cache has no .git; otherwise it fetches from the remote the clone already stores. The cache directory must be removed as well, or the box goes on following the live catalog and reports success. Filed as R-615; until it is fixed, the drill procedure is:

  1. save controller.yaml as controller.yaml.pre-update-night;
  2. set git.repo_url (and username/token — the drill repo is private);
  3. remove <data>/catalog-cache;
  4. restart the controller, sync, rescan (R-607: a sync can answer „nincs változás" while the catalog has moved, and the badge answers from the stale value until the rescan);
  5. three controls, all quoted in the report — the drill bump appears on the scratch box; the other boxes' caches are unchanged; the live catalog's main hash is unchanged.

What the drill must leave behind

  • controller.yaml restored from the saved copy, the controller restarted, and git.repo_url read back and quoted as the live catalog.
  • The registry container and its volume removed; drill images removed by name. Never prune.
  • The drill repo kept, private, reset to the live catalog's main, so the next drill starts clean.
  • A diff of every image: line against the live catalog's main — expected: identical.

The fence

Only a scratch guest is ever pointed at the drill catalog. The drill repo's README says so, and no customer box has credentials for it. The store listens on the guest's loopback only.