Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
16 KiB
DRILL — night 2026-09-25: the fixed agent delivered, the HP box backs up again, automatic updates built, proven, and watched through the demo boxes' first real night
Brief: "NIGHT SHIFT 2026-09-24/25". Architecture read before any claim: 09-update-architecture.md §3 decisions
11–30 and §6.4.2 (the spec for Part B), 03-host-agent.md, 07-backup-architecture.md, 08-alarm-ladder.md,
04-control-plane-authorization.md §2–§3.1. Evidence: audits/night-2026-09-25/ (A, B, C, D, E, F, G, tools).
Not done, or changed
- An unplanned whole-box backup of demo-hp 9201 ran for 5 min 16 s (22:40:21–22:45:37 CEST) and was aborted by
me. It came from my own live test of agent v0.134.0 (R-685), which was meant to prove a REFUSAL: the first build
read free space from
GET /storage, which carries no usage, so the check failed open and a real vzdump started. I stopped the task; vzdump removed its temporary files; no archive, no lock; root disk back to 60 %; the guest was suspended 2 s (snapshot mode), its apps kept running. This breached the brief's fence "no whole-box backup on a demo box except Part A4". Fixed before release (free space fromGET /nodes/<node>/storage), red-proofed twice, and re-proven with safe builds that stop before any vzdump (F/F1-*,F/F2-*). - The floor was raised at 22:59 CEST, not "around 02:00". Part C had passed; raising early gave 3.5 hours to see both boxes arrive and read their switch and window before the night. Both arrived in ~15 s.
- Part F ran BEFORE Part D, and its controller half is not done. Agent v0.134.0 (the space check) was released, signed and delivered to both demo boxes (Part F's live proof passed). The controller's backup-page line for a space skip needs a controller release, and v0.271.0 was already the night's one — R-685 stays open for that half. R-671, R-670, R-677 were not started — each needs a controller release (one per repo per night).
- Both agent_update jobs of 0.134.0 were queued within a minute, not one box after the other's commit (each is per-box by signature; the 0.132.0 delivery did the same). 0.133.0 went strictly one after the other.
- Part C's accidents: two ran live, two could not on a scratch box without an off-site tier and an agent — the off-site leg FAILING and W+5h reached with steps left are unit-proven only (R-687). "The controller killed BETWEEN two apps" landed INSIDE the next app's step: the leg presses the next app in the same second (see Part C).
- Part D: no whole-box backup followed the leg tonight, so the backup gate never had to wait. demo-hp's last whole-box backup was Part A4's (21:59), so it is not due until this evening; demo-felhom's is due ~07:27, after the session. Both legs ended by 04:19, before the gate opened at 04:30. The gate's wait is unit-proven and red-proofed only (R-687 updated).
- One step per app per night (decision 33) follows the brief's own words; §6.4.2 (a)'s "rescan between steps" read as climbing several — recorded as a decision the operator may widen.
Interventions: 0 on product behaviour during the real night (nothing was touched on the demo boxes between the floor raise and the end of the night except to read). The demo boxes' real night: demo-felhom 1 step done (opengist 1.13 → 1.15, 20 s), demo-hp nothing to do (every app current); 0 undone, 0 held, 0 skipped. The one result that matters most: the automatic leg ran unattended on both demo boxes exactly where the chain puts it — after the off-site copy — pressed only the one tested step waiting, and reported itself to the hub.
Part B — controller v0.271.0 (9cf13a3, CI job 974 success)
Built exactly §6.4.2 point 6 (a)–(d) plus R-680, R-678, decision 13's marks, the older-than-ladder rule, the switch
and the events rule. As built: stacks/unattended.go (RunUpdateLeg), chainUpdateLeg in cmd/controller/main.go
(the offbox-backup job), quiesce.Options.UpdateLegFn, backupwindow.UpdateLegStopOffsetMin, settings.json
app_update.unattended (absent = ON) with its card on /settings, app.yaml failed_update_step and
last_auto_update, the report's update_leg, backup.FreshWholeCopy; stacks.update_window removed. Decisions
31–33 taken unattended (09 §3). Red-proofs — 11, each seen failing, tree restored after each
(night-2026-09-25/B/redproofs/): the leg on every off-site path (a deferred call: error and panic cases fail without
it); the gate deferring while the leg runs and not at W+5h; the leg starting nothing at W+5h; one step per app a
night; the failed step skipped; needs_person skipped; files_may_change without a whole copy skipped; no crash-loop
verdict during a step (the first version of this test passed under the mutation — its container was set before the
leg's rescan rebuilt the stack; fixed and re-seen red); the steps-left count fresh at done; the switch ON by
default; the leg wired at startup. Gates green; one parity snapshot regenerated for the new card (diff: additions only).
Part A — the agent, the restore test, the HP backup
A1 — agent v0.133.0 signed and delivered, one box after the other (ruling 1, 2026-09-16). Package sha256
3aa30345…e69b6 verified against the published package before signing (A/A1-*.txt).
| box | signed (UTC) | box took it | AUTHORIZED → COMPLETED | committed | hub reads |
|---|---|---|---|---|---|
| demo-felhom-8363b5 | 19:25:30 | 21:34:05 CEST | job c4b592ff7c9ff307 |
21:35:09 | 0.133.0 |
| demo-hp-bb76ea | 19:35:55 | 21:48:58 CEST | job 36aab1ed8ad00979 |
21:50:02 | 0.133.0 |
Peti's box: never signed for.
A2 — the restore test back on, both boxes. agent.json.pre-r672 copied back; the diff was only
restore_test_eval_interval_seconds: -1 (and the trailing newline). The -1 config is kept as
agent.json.night-0925-off. Start-up line on both: restore-test scheduler starting (per-archive due-check) eval_interval=6h0m0s settle=24h0m0s.
A3 — one on-demand restore test. demo-felhom: PASS — scratch 990000 restored, booted, verified, torn down in 1 min 25 s; pool 3.13 % → peak 5.68 % → 3.13 %; preflight "required 12.8 GB, avail 362.8 GB". demo-hp: REFUSED, correctly — "restoring 21.1 GiB (vzdump log: total bytes written) needs 30.3 GiB free, has 22.1 GiB", exit 4, pool 59.01 % before and after, nothing created.
A4 — HP whole-box backups, option A. Before: root disk (local) 89 % used, 4.3 GB free, three 9201 archives
(6.2, 6.3, 7.4 GB), retention 3. Changed: local_backup_retention 3 → 1 (saved copy
agent.json.pre-a4-retention), agent restarted, the two oldest archives removed with PVE's own
pvesm prune-backups local --vmid 9201 --keep-last 1 (dry-run shown first) → 58 % used, 16.8 GB free. Why the
prune by hand: PVE prunes AFTER a successful backup, so lowering the retention alone would not have let the next
backup fit. One whole-box backup by the product's own button (POST /api/guest-backup/trigger, the Mentések
page) — 9201's apps were stopped for it (snapshot mode, resumed at "snapshotted"): local archive 8.18 GB in
7 min, then the PBS tier (the button covers every tier) 23.2 GB in 4.5 min; the new prune kept 1 archive; root
60 %; all 20 app containers healthy. The product row is R-685.
Part C — the proof on 9202: six simulated nights
Controller v0.271.0 on 9202 only (C/C0-deploy-9202.txt), drill catalog (C/C0-repoint-drill.txt, health timeout
90 s). Four apps installed through the product at the ladder's first from, then the drill put back at the head:
wishlist 1 step behind, navidrome 2, romm 3, vikunja 1 with a failing step (its head probe on port 8999). Marks
set in the drill: navidrome step 1 needs_person, wishlist's step files_may_change. Data seeded through each
app's own front door and read back after EVERY night (C/data-readback.txt: 8 of 8 rounds, 4 of 4 apps, each
fixture with its negative control). The window was moved through the product's own form (POST /backups/window)
so the off-site job — and the leg chained to it — fired two minutes later.
| night | what it proved | leg summary (the box's own line) | pages (hu + en) |
|---|---|---|---|
| 1 | the chain on a box with NO off-site target ("the off-site leg does nothing; the update leg runs now"); one step per app; needs_person skipped; files_may_change taken WITH a whole copy (wishlist); a failing step undone |
done=2 undone=1 held=0 failed=0 skipped=1 in 3m6s [skipped: navidrome=needs_person] — romm 60.1 s, vikunja undone 105.2 s, wishlist 20.1 s |
„Automatikus frissítés 2026-09-24 22:31-kor — sikeres." / "Automatic update at … — done."; vikunja's undone sentence in both |
| 2 | R-680 live: the undone step is NOT pressed again; navidrome's files_may_change step taken (its music folder is excluded, so its own unit is whole) |
done=2 undone=0 … skipped=1 in 1m10s [skipped: vikunja=failed_before] |
badges current / behind as expected |
| 3 — accident | controller killed (kill -9) the moment navidrome's step ended | the leg had pressed romm IN THE SAME SECOND; the kill landed before romm's up → pin and definition put back, romm runs its previous version, page: „A frissítés megszakadt, mert a vezérlő újraindult…" / "The update was interrupted…"; the leg was not resumed (R-686) |
all data read back |
| 4 — accident | power cut (pct stop/pct start of 9202) while romm's step was verifying |
box up in 13 s; "resumed 1 interrupted update(s)"; romm healthy after 30 s, DONE in 1 m 32 s; no automatic-update line on the page (the leg died — R-686) | all data read back |
| 5 | the switch OFF (POST /settings/app-update, page reads unchecked) |
done=0 … stopped=switched_off in 0s — nothing pressed |
— |
| 6 | the switch ON again; the catalog RE-TESTED vikunja's step (probe fixed, new tested_at) → the ladder print changed → R-680's record no longer binds |
done=1 … in 10s — vikunja done |
„Automatikus frissítés … 22:58-kor — sikeres." |
Not run live, and why (R-687): W+5h reached with steps left (the leg starts at W+105m; needs a 3-hour leg — unit
test TestLeg_NoStepAtOrAfterW5h); the off-site leg FAILING (9202 has no off-site tier — unit test
TestChainUpdateLeg_EveryPath covers error and panic); a files_may_change step WITHOUT a whole copy (every app
given the mark was whole on 9202); the full-system gate's deferral (9202 has no agent, so no quiesce loop).
Verdict: Part C passed — every accident that could run on 9202 ended with the app healthy, the data read back, and a true sentence on its page; the four not-runnable cases are named, unit-proven and in the register.
Part D — the demo boxes' first real automatic night
Floor 0.271.0 (declared MinAgent 0.131.0) saved 20:59:04Z, read back; both boxes on 0.271.0 ~15 s later
(D/D1, D/D2). Before the night, per box (D/D3, D/D0): switch ON (key absent = the default; demo-hp's
settings page reads it checked in both languages), window 02:30 (off-site + update leg at 04:15, gate [04:30, 08:30)).
Tested steps waiting: demo-felhom — opengist 1.13 → 1.15; demo-hp — none (all 10 apps at the catalog head).
Agents: 0.134.0 on both (Part F). Nothing was touched on either box during the night except to read.
demo-felhom (N100), leg by leg (D/D6-demo-felhom-night-full.log): db-dump 02:30:00 (0.8 s) → Tier 2 03:30:00
→ off-site 04:15:00–04:15:46 (1 app, 11 snapshots, 42 s) → update leg 04:15:46–04:16:06: opengist pressed; safety
dump, undo copy (0.2 MiB), pin, pull, start, verify healthy after 10 s, DONE in 17 s; done=1 undone=0 held=0 failed=0 skipped=0 in 20s. After: container opengist:1.15 healthy; app.yaml last_auto_update: done 1.13 → 1.15 (the page
line). One WARN worth keeping: the volume "carries no compose label (recreated by a restore before v0.268.0 — R-658)
— copied by name" — the R-658 fallback working. No mail (a successful step sends none).
demo-hp (D/D6-demo-hp-night-full.log): db-dump 02:30:00–02:31:59 → Tier 2 03:30:00–03:30:24 → off-site
04:15:00–04:18:41 (9 apps, 90 snapshots, 3 min 35 s, OK) → update leg 04:18:41: nothing to press, done=0 … in 0s.
The hub holds both summaries in its stored reports (received 02:44Z; D/D7-hub-report-update-leg.txt, read from a
copy of the hub DB taken with its -wal, then deleted). Whole-box backups: none ran tonight on either box — not
due (see "Not done"); both agents alive (627 / 1,069 journal lines since 02:25, the 10-minute janitor on schedule).
No app ended held or stopped.
Part E — two apps moved
Venue: bench LXC 9401 on demo-hp (created and destroyed tonight — no permission check refused either; the
Debian template it needed was removed again), harness v3; box walk on 9202 through the real guarded Update.
Negative control: n8n=alpine:3.20 → failed. A bench setup miss is recorded: the first queue ran without
the traefik-public network and returned inconclusive — FROM deploy failed (kept as *-noNetwork); re-run.
| app | step | bench | box | memory peak (own) | catalog commit (CI) |
|---|---|---|---|---|---|
| n8n | 2.41.1 → 2.41.2 | proven | proven | 22.4 % | f14a608 (job 980) |
| mealie | v3.27.0 → v3.28.0 | proven | proven | 23.1 % | b996218 (job 981) |
Published at 02:31–02:32, after the demo boxes' night began and on neither box's app list. Not moved, with the
reason for each: catalog REPORT.md (drift E/E0-drift.log).
Part F — agent v0.134.0 (R-685 agent half)
Released by scripts/release-agent.sh (tag v0.134.0 at 0722b2c, sha256 7593bebe…81c72d, verified by download;
CI jobs 975–977). A vzdump to a local target needs free ≥ newest archive × 1.25 + 1 GiB; a shortfall is a named skip
before anything starts; fail-open on PBS / first backup / unknown usage. Live, safe builds (a hard stop before vzdump):
×10 → "local has 14.9 GiB free; … needs about 77.2 GiB" refused; ×1.25 → "space preflight passed" need 11.3 GB,
avail 16.0 GB. Signed to both demo boxes (F/F3-*): demo-felhom committed 22:52:45, demo-hp 22:58:53; hub reads 0.134.0.
The incident above belongs to this part.
Teardown — three layers
- Machine: 9202 back on the live catalog (
catalog-cacheatc802509at the time; the controller followsmain), window 02:30 again, the six drill apps removed through the product (0 containers, 0 volumes, 0 undo copies); left by name on the scratch drive:userdata/navidrome,userdata/romm(the remove kept drive data — R-442's fail-closed path on 9202). The switch key on 9202 is now explicittrue(was absent = ON). 9202 stays on controller 0.271.0 (the floor). Bench 9401 destroyed (hostname checked first), template removed. - Host: demo-hp
pct list9201 + 9202, local-lvm 60.05 %, root 60 %; demo-felhom 9201 only, local-lvm 3.16 %. Config changes kept on purpose: restore test ON (both),local_backup_retention: 1(demo-hp). No test binaries left in /tmp. - Hub: global floor 0.271.0 (MinAgent 0.131.0); both demo boxes 0.271.0 / agent 0.134.0. Drill repo reset to live
main(dd9c6ad), private, Actions off.
Claims in the brief that turned out wrong (or right)
- The off-site job function is the single place the leg can be chained from on every path — TRUE for every path
of the job, with one exception outside it: a box with backups disabled registers no off-site job at all, so no
leg (the update itself would refuse
no_backupthere anyway). The job's own early returns and the scheduler's "not configured" skip all return intochainUpdateLeg. A controller restart after the job fired loses the rest of the night (R-686). - A demo box has a tested step waiting tonight — TRUE for demo-felhom (opengist), FALSE for demo-hp.
- The whole-box backup lands on demo-hp's root disk — TRUE (
local=/var/lib/vz, root LV). - Decision 28's suppressions cover a deploy's first start — WRONG.
Deployingclears whencompose up -dreturns; a first start that restarts ≥ 6 times in 10 min is stopped (R-676 updated). An automatic step, its verify and its undo ARE covered (pinned by a test).
Register
Before: 341 rows / 685,148 B. After: 338 rows / 682,053 B. Opened R-685 (backup-fit warning; agent half shipped, controller page half open), R-686 (no resume of the leg after a restart), R-687 (live-proof gaps). Closed R-672, R-673 (delivered), R-684 (option A), R-680, R-678, R-643. Updated R-676 (deploy first start), R-450 (narrowed to part 10).