Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all down or blocked; both demo boxes SERVED. Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s after the cut decision and the seeded data read back intact — so the branch that is one step from old-binary-on-migrated-database is now evidence, not argument. Instrument limit stated: `starting` lasts under a second; all three landed in `verifying`, which RecoverUpdates handles in the same branch. Part 3 (R-611) — the night the previous session skipped without saying so. An app updated with nobody pressing anything; a terminally-refused app was pressed exactly once and never again over three passes. The unattended HOLD was NOT produced: the within-a-major rule correctly refused the broken edge before it was attempted, so Q4 still rests on the attended hold from slice 4. Said plainly rather than implied. Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard), R-614 (stale update phase survives a redeploy). R-520's pointer corrected. Catalog: two drill pairs, both reverted; every image line byte-identical to ff9717d3. The alpine:3.20 negative control a security review flagged is cleared. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
9.5 KiB
REPORT — the update arc's two missing measurements, one lock, and the floor to 0.260.0
2026-09-21 (evening). Repos touched: felhom.eu (floor, docs, register, evidence),
felhom-controller v0.261.0, app-catalog-felhom.eu (two drill pairs, both reverted).
Architecture read first and named: documentation/architecture/09-update-architecture.md §3, §3b,
§4, §6.1, §6.2, §6.4, §8.
1. NOT DONE / CHANGED FROM THE BRIEF — first, because that is the point of this session
| item | state |
|---|---|
| Part 0 floor to 0.260.0 | done |
| Part 1 the three cuts | done, but NOT as specified. All three landed in verifying, never in starting — see §2. The brief allowed this explicitly and asked that it be said. |
Part 2 the lock + reason on the wire |
done, five red-proofs, proven live |
| Part 3 the unattended night | done for the success night and the no-retry proof. The unattended HOLD was NOT produced — see §4. |
| Part 4 docs and rows | done |
| the caller script's ≤150-line budget | 167 lines. Over by 17, not trimmed: the excess is the within-a-major rule and its comment, the one part of that file that must be readable. |
| Scenario C run by the measuring agent | run by the coordinator instead. The agent was stood down mid-session after two long intervals with no evidence written; the coordinator ran C and captured A and B independently. Stated because it changes who measured what. |
| a second drill bump/revert pair | used. The brief permits it "if a fifth move is truly needed" and asks that it be named. It was: DRILL 2 (ae08a037fd68) added one real edge and one deliberately failing edge for Part 3. |
Two instrumentation failures of my own, recorded because they cost evidence:
- The first unattended run's stdout was piped through
tail, which buffers, and the run was later killed — the caller's own log for the Scenario F press was lost. The outcome survived on the box; the log did not. The second run wrote straight to a file. - A background security review flagged the deliberately-broken
vikunja → alpine:3.20catalog edge as a supply-chain change. It was right to. Accepted deliberately — no customer or demo box runs vikunja, a deployed app is frozen at its own pin since v0.235.0, and it is the documented C3-class control — but the window is now closed by the revert, and it is named here rather than left in a tool notification.
2. Claims in the brief that turned out wrong
§2.4's open question — does backupMgr.IsRunning() cover the update's backing-up phase?
YES. RunAppBackupNow calls acquireRunning (internal/backup/update_guard.go:333), so that one
phase was already protected. The gap was checking, safety-dump, pinning, pulling, starting
and verifying. The live lock probe landed in safety-dump, i.e. squarely in the previously
unprotected window rather than in the one that was already covered. Everything else §2.4 asserted
held at source.
§2.3 — the resume path was READ, not measured. It is now measured, three times, and it does what it said.
The stopped-guest byte path in my own brief to the measuring agent was wrong.
/var/lib/lxc/9202/rootfs/var/lib/felhom/... is an empty mountpoint while 9202 is stopped, because
mp0 is a separate raw volume; pct mount 9202 does attach it. I had corrected one trap and
introduced a second. The positive control caught it.
"Four qualifying apps exist" — held. vikunja, uptime-kuma, wishlist, glance: single-container, no database sidecar, none on either demo box, all four target tags verified to exist upstream first.
3. Part 0 — the floor
Raised to 0.260.0 with MinAgent 0.131.0 declared (above the vouched golden 0.258.0, so the declaration carries it — §3 decision 7). Hub log:
[INFO] Global controller-version floor set to "0.260.0" (declared MinAgent "0.131.0")
[INFO] managed floor SERVED for demo-felhom: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
[INFO] managed floor SERVED for demo-hp: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
Blast radius, read from the hub before saving: 3 boxes below — drill-r50 (0.213.0, BLOCKED),
peti-felhom (0.115.0, DOWN), tester-1 (0.245.0, DOWN). None was reachable, so none moved; they
take it when they return. Both demo boxes were already on 0.260.0 by hand and now hold it by floor.
4. Part 1 — the three cuts (R-610, CLOSED)
Guest 9202, controller v0.260.0, four throwaway apps seeded through their own front doors first.
| A — vikunja | B — uptime-kuma | C — wishlist | |
|---|---|---|---|
| edge | 2.3.0 → 2.6.0 | 2.4.0 → 2.5.0 | v0.66.0 → v0.67.0 |
| cut | pct stop |
pct stop |
controller container only |
| phase at decision | starting |
verifying |
starting |
| cut latency | 3 759 ms | 3 016 ms | 1 675 ms |
| phase it died in | verifying |
verifying |
verifying |
| recovery | resumed, healthy 0 s | resumed, healthy 5 s | resumed, healthy 10 s |
| total | DONE 1 m 26 s | DONE 1 m 0 s | DONE 51 s |
| four observables | agree | agree | agree |
| seeded data | read back intact | read back intact | not re-read (gap) |
The dangerous case was genuinely exercised, and a log line proves it rather than an assumption.
vikunja's own log: Ran all migrations successfully / Vikunja version v2.6.0 at 12:28:26.881
UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering. The 2.6.0 schema
migration had already been applied to the customer's database when the power went. Recovery resumed
forward, so old-binary-on-migrated-database never happened — but this branch is one step from
it, and that is now evidence for §4's "no automatic rollback" ruling rather than argument for it.
Instrument limit: starting lasts well under a second on this box. Three attempts, two cut
mechanisms, all landed in verifying. No phase was faked. RecoverUpdates handles starting and
verifying in one branch, so all three exercise the arm under test. A cut inside starting
itself needs an in-process fault injector.
Scenario C also showed the other apps were undisturbed by the controller restart — uptime-kuma, vikunja, glance and filebrowser all kept their uptime. Only the app the update was itself recreating restarted.
5. Part 2 — v0.261.0, proven live with a control
See felhom-controller/REPORT.md for the code. The live proof is the part worth repeating:
| probe | result |
|---|---|
| manual self-update, no app update running | „A frissítés nem érhető el (nincs gazda-ügynök)" — the agent refusal |
manual self-update, app update in flight (safety-dump) |
„Egy alkalmazás frissítése éppen folyamatban van…" — our refusal |
The sentence changed. Guest 9202 has no host agent, so TriggerUpdate refuses either way — which
makes it the perfect negative control, because the new check sits before the agent check. Two
sentences, one probe, and no swap could reach a machine. Self-update was enabled for the probe and
restored to false from a copy taken first; verified by re-reading the file.
The reverse direction — an app update refused while the controller swaps — is NOT staged live. It is covered by a red-proofed consequence test. Staging it would need a real swap and a host agent this guest does not have. Stated as a gap.
6. Part 3 — the unattended night (R-611, CLOSED)
Success: uptime-kuma 2.5.0 → 2.5.1 applied with nobody pressing anything.
No-retry: after the revert left all four apps ahead of the catalog, the caller pressed each
exactly once, was refused downgrade (terminal), and pressed nothing across two further passes.
never_again=['glance','uptime-kuma','vikunja','wishlist'], outcomes={}. That is R-524 and R-609
working together, unattended.
The unattended HOLD was never produced, and the reason matters: the only failing edge available
(vikunja → alpine:3.20) was correctly refused by the within-a-major rule before it was ever
attempted. The rule that makes automatic updates safe is the same rule that refuses the obvious way
to break one. Measuring it needs an image that passes the version test and still fails health.
§3b Q4 therefore still rests on the ATTENDED hold from slice 4.
7. Rows, and the catalog
Closed: R-608, R-609, R-610, R-611. Opened: R-612 (P1 — wishlist unusable on a fresh install and the error is a lie), R-613 (P2 — uptime-kuma reports healthy on its setup wizard), R-614 (P3 — stale update phase survives a redeploy). Corrected: R-520's closing pointer now names R-610. Register 302 → 307.
Catalog: DRILL 573e41f5, DRILL 2 ae08a037, REVERT f5f6a152. Every image: line in
templates/ is byte-identical to the pre-drill ff9717d3 — git diff over those paths is 0
lines. catalog_since reads 2026-09-21 on the four rather than the older dates, because the gate
requires an image move to carry the day's date in either direction and a revert is a move.
Teardown, three layers. Machine: guest 9202 left running on v0.261.0 with the four throwaway
apps still deployed and healthy (glance, uptime-kuma, vikunja, wishlist) — they are the fixture for
the remaining Q4 work and removing them would cost the next session the seeding. Host: demo-hp
untouched apart from 9202; guest 9201 never touched. Hub: nothing provisioned, nothing
enrolled; 9202 reports to no hub by design. /tmp/.ctlpw shredded.