Files
felhom.eu/documentation/audits/update-night-2026-09-21/09-SECTION-drill-method.md
T
admin da20722e76
gates / gates (push) Successful in 27s
Update night 2026-09-21: Phase 0 and Phase 1 evidence, the drill method, and two instrument fixes
INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320),
not at the end of the session. Phases 2-5 follow in a later commit.

Phase 0, all three mechanisms proven with their controls:
- the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub
  logging `managed floor SERVED ... from declared (golden 0.258.0)`.
- a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and
  engine-major edges can be measured without the live catalog ever carrying one. Positive
  control quoted, and two negative controls: the live catalog's main and both real boxes'
  caches unchanged.
- a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD
  measurable at all: an edge that PASSES the within-a-major test and still fails.
  CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases
  + 1 negative control), not by reading it.

Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own
guarded Update, each app seeded and read back through its OWN front door (R-156), with a
per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`.

TWO INSTRUMENT FIXES, both in this repo's own evidence code:
- 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described
  404s. Corrected, with the session-expiry note that cost the same time.
- unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were
  always None and EVERY followed update ran to its 900 s timeout and was then recorded
  `timeout` and never-press-again. Fixed before B1 relied on it. R-623.

No controller, agent or hub code was written. The live catalog carries no broken reference.

Gates: repo_gates.py --fast — all 15 OK, exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 21:17:46 +02:00

64 lines
3.9 KiB
Markdown

# Draft section for `09-update-architecture.md` — the standing method for update drills
*(To be inserted after §6.4. Written here first so the audit and the architecture stay in step.)*
---
## 6.5 The drill catalog and the image store — the standing method for update drills
**Why this section exists.** On 2026-09-21 an afternoon session put a deliberately broken image into
the LIVE catalog for thirteen minutes to produce a failing edge. It was reverted and nothing reached
a customer, but the method was wrong and the brief that asked for it said so. This is the method that
replaces it, proven the same night.
**The rule, and it has no exception:** *nothing broken, dummy, cross-repo or engine-major ever enters
the live catalog — not as a fallback, not for thirteen minutes. If a leg cannot be done without that,
the leg is skipped and named.*
### The two mechanisms
| | what it is | what it makes possible |
|---|---|---|
| **the drill catalog** | `admin/app-catalog-drill` on Gitea — private, a copy of the live catalog's `main` | a scratch box can be pointed at a catalog where a failing edge is *committable*, because it carries none of the live repo's gates |
| **the image store** | a `registry:2` container on the scratch guest at `127.0.0.1:5000` | an edge that **passes the within-a-major test and still fails** — the one shape a real catalog move cannot produce |
**The image store is not a convenience.** `09` §3b Q4 could not be measured for a year of drills
because the only failing edges available were across-a-major, and the within-a-major rule — correctly
— refuses those before the guarded update is ever reached. *The rule that makes automatic updates
safe is the same rule that refuses the obvious way to break one.* Measuring an unattended HOLD needs
`drill/<app>:X.Y.Z` (the real image, retagged) against `drill/<app>:X.Y.(Z+1)` (a built image that
starts, stays up and never serves) — same repository, same major, plain version tags. A third
flavour, a tag simply **absent** from the store, gives the pull-failure leg.
`stacks.CompareImageRefs` orders a `host:port/` reference correctly: `splitImageRef` takes the last
colon and rejects it only when a `/` follows, so a registry port is never read as a tag. **Proven by
running it**, four positive cases and a negative control, 2026-09-21.
### Pointing a box at the drill catalog — the step that is NOT obvious
**`git.repo_url` alone is inert.** `Syncer.gitCloneOrPull` clones only when the cache has no `.git`;
otherwise it fetches from the remote the clone already stores. The cache directory must be removed as
well, or the box goes on following the live catalog and reports success. Filed as **R-615**; until it
is fixed, the drill procedure is:
1. save `controller.yaml` as `controller.yaml.pre-update-night`;
2. set `git.repo_url` (and `username`/`token` — the drill repo is private);
3. **remove `<data>/catalog-cache`**;
4. restart the controller, sync, **rescan** (R-607: a sync can answer „nincs változás" while the
catalog has moved, and the badge answers from the stale value until the rescan);
5. **three controls, all quoted in the report** — the drill bump appears on the scratch box; the
other boxes' caches are unchanged; the live catalog's `main` hash is unchanged.
### What the drill must leave behind
- `controller.yaml` restored from the saved copy, the controller restarted, and `git.repo_url` **read
back and quoted** as the live catalog.
- The registry container and its volume removed; drill images removed **by name**. Never `prune`.
- The drill repo **kept**, private, reset to the live catalog's `main`, so the next drill starts clean.
- A diff of every `image:` line against the live catalog's `main` — expected: identical.
### The fence
Only a scratch guest is ever pointed at the drill catalog. The drill repo's README says so, and no
customer box has credentials for it. The store listens on the guest's loopback only.