6b2176e480
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
194 lines
16 KiB
Markdown
194 lines
16 KiB
Markdown
# DRILL — night 2026-09-25: the fixed agent delivered, the HP box backs up again, automatic updates built, proven, and watched through the demo boxes' first real night
|
||
|
||
Brief: "NIGHT SHIFT 2026-09-24/25". Architecture read before any claim: `09-update-architecture.md` §3 decisions
|
||
11–30 and §6.4.2 (the spec for Part B), `03-host-agent.md`, `07-backup-architecture.md`, `08-alarm-ladder.md`,
|
||
`04-control-plane-authorization.md` §2–§3.1. Evidence: `audits/night-2026-09-25/` (A, B, C, D, E, F, G, tools).
|
||
|
||
## Not done, or changed
|
||
|
||
- **An unplanned whole-box backup of demo-hp 9201 ran for 5 min 16 s (22:40:21–22:45:37 CEST) and was aborted by
|
||
me.** It came from my own live test of agent v0.134.0 (R-685), which was meant to prove a REFUSAL: the first build
|
||
read free space from `GET /storage`, which carries no usage, so the check failed open and a real vzdump started.
|
||
I stopped the task; vzdump removed its temporary files; no archive, no lock; root disk back to 60 %; the guest was
|
||
suspended 2 s (snapshot mode), its apps kept running. This breached the brief's fence "no whole-box backup on a
|
||
demo box except Part A4". Fixed before release (free space from `GET /nodes/<node>/storage`), red-proofed twice,
|
||
and re-proven with safe builds that stop before any vzdump (`F/F1-*`, `F/F2-*`).
|
||
- **The floor was raised at 22:59 CEST, not "around 02:00".** Part C had passed; raising early gave 3.5 hours to
|
||
see both boxes arrive and read their switch and window before the night. Both arrived in ~15 s.
|
||
- **Part F ran BEFORE Part D, and its controller half is not done.** Agent v0.134.0 (the space check) was released,
|
||
signed and delivered to both demo boxes (Part F's live proof passed). The controller's backup-page line for a space
|
||
skip needs a controller release, and v0.271.0 was already the night's one — R-685 stays open for that half.
|
||
**R-671, R-670, R-677 were not started** — each needs a controller release (one per repo per night).
|
||
- **Both agent_update jobs of 0.134.0 were queued within a minute**, not one box after the other's commit (each is
|
||
per-box by signature; the 0.132.0 delivery did the same). 0.133.0 went strictly one after the other.
|
||
- **Part C's accidents: two ran live, two could not** on a scratch box without an off-site tier and an agent — the
|
||
off-site leg FAILING and W+5h reached with steps left are unit-proven only (R-687). "The controller killed BETWEEN
|
||
two apps" landed INSIDE the next app's step: the leg presses the next app in the same second (see Part C).
|
||
- **Part D: no whole-box backup followed the leg tonight, so the backup gate never had to wait.** demo-hp's last
|
||
whole-box backup was Part A4's (21:59), so it is not due until this evening; demo-felhom's is due ~07:27, after the
|
||
session. Both legs ended by 04:19, before the gate opened at 04:30. The gate's wait is unit-proven and red-proofed
|
||
only (R-687 updated).
|
||
- **One step per app per night** (decision 33) follows the brief's own words; §6.4.2 (a)'s "rescan between steps"
|
||
read as climbing several — recorded as a decision the operator may widen.
|
||
|
||
**Interventions: 0** on product behaviour during the real night (nothing was touched on the demo boxes between the
|
||
floor raise and the end of the night except to read). **The demo boxes' real night: demo-felhom 1 step done
|
||
(opengist 1.13 → 1.15, 20 s), demo-hp nothing to do (every app current); 0 undone, 0 held, 0 skipped.** **The one
|
||
result that matters most:** the automatic leg ran unattended on both demo boxes exactly where the chain puts it —
|
||
after the off-site copy — pressed only the one tested step waiting, and reported itself to the hub.
|
||
|
||
|
||
## Part B — controller v0.271.0 (`9cf13a3`, CI job 974 success)
|
||
|
||
Built exactly §6.4.2 point 6 (a)–(d) plus R-680, R-678, decision 13's marks, the older-than-ladder rule, the switch
|
||
and the events rule. As built: `stacks/unattended.go` (`RunUpdateLeg`), `chainUpdateLeg` in `cmd/controller/main.go`
|
||
(the `offbox-backup` job), `quiesce.Options.UpdateLegFn`, `backupwindow.UpdateLegStopOffsetMin`, `settings.json`
|
||
`app_update.unattended` (absent = ON) with its card on /settings, app.yaml `failed_update_step` and
|
||
`last_auto_update`, the report's `update_leg`, `backup.FreshWholeCopy`; `stacks.update_window` removed. Decisions
|
||
31–33 taken unattended (`09` §3). **Red-proofs — 11, each seen failing, tree restored after each**
|
||
(`night-2026-09-25/B/redproofs/`): the leg on every off-site path (a deferred call: error and panic cases fail without
|
||
it); the gate deferring while the leg runs and not at W+5h; the leg starting nothing at W+5h; one step per app a
|
||
night; the failed step skipped; `needs_person` skipped; `files_may_change` without a whole copy skipped; no crash-loop
|
||
verdict during a step (the first version of this test passed under the mutation — its container was set before the
|
||
leg's rescan rebuilt the stack; fixed and re-seen red); the steps-left count fresh at `done`; the switch ON by
|
||
default; the leg wired at startup. Gates green; one parity snapshot regenerated for the new card (diff: additions only).
|
||
|
||
## Part A — the agent, the restore test, the HP backup
|
||
|
||
**A1 — agent v0.133.0 signed and delivered, one box after the other** (ruling 1, 2026-09-16). Package sha256
|
||
`3aa30345…e69b6` verified against the published package before signing (`A/A1-*.txt`).
|
||
|
||
| box | signed (UTC) | box took it | AUTHORIZED → COMPLETED | committed | hub reads |
|
||
|---|---|---|---|---|---|
|
||
| demo-felhom-8363b5 | 19:25:30 | 21:34:05 CEST | job `c4b592ff7c9ff307` | 21:35:09 | 0.133.0 |
|
||
| demo-hp-bb76ea | 19:35:55 | 21:48:58 CEST | job `36aab1ed8ad00979` | 21:50:02 | 0.133.0 |
|
||
|
||
Peti's box: never signed for.
|
||
|
||
**A2 — the restore test back on, both boxes.** `agent.json.pre-r672` copied back; the diff was only
|
||
`restore_test_eval_interval_seconds: -1` (and the trailing newline). The `-1` config is kept as
|
||
`agent.json.night-0925-off`. Start-up line on both: `restore-test scheduler starting (per-archive due-check)
|
||
eval_interval=6h0m0s settle=24h0m0s`.
|
||
|
||
**A3 — one on-demand restore test.** demo-felhom: **PASS** — scratch 990000 restored, booted, verified,
|
||
torn down in 1 min 25 s; pool 3.13 % → peak 5.68 % → 3.13 %; preflight "required 12.8 GB, avail 362.8 GB".
|
||
demo-hp: **REFUSED, correctly** — "restoring 21.1 GiB (vzdump log: total bytes written) needs 30.3 GiB free, has
|
||
22.1 GiB", exit 4, pool 59.01 % before and after, nothing created.
|
||
|
||
**A4 — HP whole-box backups, option A.** Before: root disk (`local`) 89 % used, 4.3 GB free, three 9201 archives
|
||
(6.2, 6.3, 7.4 GB), retention 3. Changed: `local_backup_retention` 3 → 1 (saved copy
|
||
`agent.json.pre-a4-retention`), agent restarted, the two oldest archives removed with PVE's own
|
||
`pvesm prune-backups local --vmid 9201 --keep-last 1` (dry-run shown first) → 58 % used, 16.8 GB free. **Why the
|
||
prune by hand:** PVE prunes AFTER a successful backup, so lowering the retention alone would not have let the next
|
||
backup fit. One whole-box backup by the product's own button (`POST /api/guest-backup/trigger`, the Mentések
|
||
page) — 9201's apps were stopped for it (snapshot mode, resumed at "snapshotted"): local archive **8.18 GB** in
|
||
7 min, then the PBS tier (the button covers every tier) 23.2 GB in 4.5 min; the new prune kept 1 archive; root
|
||
60 %; all 20 app containers healthy. The product row is **R-685**.
|
||
## Part C — the proof on 9202: six simulated nights
|
||
|
||
Controller v0.271.0 on 9202 only (`C/C0-deploy-9202.txt`), drill catalog (`C/C0-repoint-drill.txt`, health timeout
|
||
90 s). Four apps installed through the product at the ladder's first `from`, then the drill put back at the head:
|
||
wishlist 1 step behind, navidrome 2, romm 3, vikunja 1 with a failing step (its head probe on port 8999). Marks
|
||
set in the drill: navidrome step 1 `needs_person`, wishlist's step `files_may_change`. Data seeded through each
|
||
app's own front door and read back after EVERY night (`C/data-readback.txt`: 8 of 8 rounds, 4 of 4 apps, each
|
||
fixture with its negative control). The window was moved through the product's own form (`POST /backups/window`)
|
||
so the off-site job — and the leg chained to it — fired two minutes later.
|
||
|
||
| night | what it proved | leg summary (the box's own line) | pages (hu + en) |
|
||
|---|---|---|---|
|
||
| 1 | the chain on a box with NO off-site target ("the off-site leg does nothing; the update leg runs now"); one step per app; `needs_person` skipped; `files_may_change` taken WITH a whole copy (wishlist); a failing step undone | `done=2 undone=1 held=0 failed=0 skipped=1 in 3m6s [skipped: navidrome=needs_person]` — romm 60.1 s, vikunja undone 105.2 s, wishlist 20.1 s | „Automatikus frissítés 2026-09-24 22:31-kor — sikeres." / "Automatic update at … — done."; vikunja's undone sentence in both |
|
||
| 2 | R-680 live: the undone step is NOT pressed again; navidrome's `files_may_change` step taken (its music folder is `excluded`, so its own unit is whole) | `done=2 undone=0 … skipped=1 in 1m10s [skipped: vikunja=failed_before]` | badges current / behind as expected |
|
||
| 3 — accident | controller killed (kill -9) the moment navidrome's step ended | the leg had pressed romm IN THE SAME SECOND; the kill landed before romm's `up` → pin and definition put back, romm runs its previous version, page: „A frissítés megszakadt, mert a vezérlő újraindult…" / "The update was interrupted…"; the leg was not resumed (R-686) | all data read back |
|
||
| 4 — accident | power cut (`pct stop`/`pct start` of 9202) while romm's step was `verifying` | box up in 13 s; "resumed 1 interrupted update(s)"; romm healthy after 30 s, DONE in 1 m 32 s; no automatic-update line on the page (the leg died — R-686) | all data read back |
|
||
| 5 | the switch OFF (`POST /settings/app-update`, page reads unchecked) | `done=0 … stopped=switched_off in 0s` — nothing pressed | — |
|
||
| 6 | the switch ON again; the catalog RE-TESTED vikunja's step (probe fixed, new `tested_at`) → the ladder print changed → R-680's record no longer binds | `done=1 … in 10s` — vikunja done | „Automatikus frissítés … 22:58-kor — sikeres." |
|
||
|
||
**Not run live, and why (R-687):** W+5h reached with steps left (the leg starts at W+105m; needs a 3-hour leg — unit
|
||
test `TestLeg_NoStepAtOrAfterW5h`); the off-site leg FAILING (9202 has no off-site tier — unit test
|
||
`TestChainUpdateLeg_EveryPath` covers error and panic); a `files_may_change` step WITHOUT a whole copy (every app
|
||
given the mark was whole on 9202); the full-system gate's deferral (9202 has no agent, so no quiesce loop).
|
||
|
||
**Verdict: Part C passed** — every accident that could run on 9202 ended with the app healthy, the data read back,
|
||
and a true sentence on its page; the four not-runnable cases are named, unit-proven and in the register.
|
||
|
||
## Part D — the demo boxes' first real automatic night
|
||
|
||
Floor 0.271.0 (declared MinAgent 0.131.0) saved 20:59:04Z, read back; both boxes on 0.271.0 ~15 s later
|
||
(`D/D1`, `D/D2`). Before the night, per box (`D/D3`, `D/D0`): switch ON (key absent = the default; demo-hp's
|
||
settings page reads it checked in both languages), window 02:30 (off-site + update leg at 04:15, gate [04:30, 08:30)).
|
||
Tested steps waiting: **demo-felhom — opengist 1.13 → 1.15; demo-hp — none** (all 10 apps at the catalog head).
|
||
Agents: 0.134.0 on both (Part F). Nothing was touched on either box during the night except to read.
|
||
|
||
**demo-felhom (N100), leg by leg** (`D/D6-demo-felhom-night-full.log`): db-dump 02:30:00 (0.8 s) → Tier 2 03:30:00
|
||
→ off-site 04:15:00–04:15:46 (1 app, 11 snapshots, 42 s) → **update leg 04:15:46–04:16:06: opengist pressed; safety
|
||
dump, undo copy (0.2 MiB), pin, pull, start, verify healthy after 10 s, DONE in 17 s; `done=1 undone=0 held=0 failed=0
|
||
skipped=0 in 20s`**. After: container `opengist:1.15` healthy; app.yaml `last_auto_update: done 1.13 → 1.15` (the page
|
||
line). One WARN worth keeping: the volume "carries no compose label (recreated by a restore before v0.268.0 — R-658)
|
||
— copied by name" — the R-658 fallback working. No mail (a successful step sends none).
|
||
|
||
**demo-hp** (`D/D6-demo-hp-night-full.log`): db-dump 02:30:00–02:31:59 → Tier 2 03:30:00–03:30:24 → off-site
|
||
04:15:00–04:18:41 (9 apps, 90 snapshots, 3 min 35 s, OK) → **update leg 04:18:41: nothing to press, `done=0 … in
|
||
0s`**.
|
||
|
||
**The hub** holds both summaries in its stored reports (received 02:44Z; `D/D7-hub-report-update-leg.txt`, read from a
|
||
copy of the hub DB taken with its -wal, then deleted). **Whole-box backups:** none ran tonight on either box — not
|
||
due (see "Not done"); both agents alive (627 / 1,069 journal lines since 02:25, the 10-minute janitor on schedule).
|
||
**No app ended held or stopped.**
|
||
|
||
## Part E — two apps moved
|
||
|
||
Venue: bench LXC **9401** on demo-hp (created and destroyed tonight — no permission check refused either; the
|
||
Debian template it needed was removed again), harness v3; box walk on 9202 through the real guarded Update.
|
||
**Negative control: `n8n=alpine:3.20` → `failed`.** A bench setup miss is recorded: the first queue ran without
|
||
the `traefik-public` network and returned `inconclusive — FROM deploy failed` (kept as `*-noNetwork`); re-run.
|
||
|
||
| app | step | bench | box | memory peak (own) | catalog commit (CI) |
|
||
|---|---|---|---|---|---|
|
||
| n8n | 2.41.1 → 2.41.2 | proven | proven | 22.4 % | `f14a608` (job 980) |
|
||
| mealie | v3.27.0 → v3.28.0 | proven | proven | 23.1 % | `b996218` (job 981) |
|
||
|
||
Published at 02:31–02:32, after the demo boxes' night began and on neither box's app list. Not moved, with the
|
||
reason for each: catalog `REPORT.md` (drift `E/E0-drift.log`).
|
||
|
||
## Part F — agent v0.134.0 (R-685 agent half)
|
||
|
||
Released by `scripts/release-agent.sh` (tag `v0.134.0` at `0722b2c`, sha256 `7593bebe…81c72d`, verified by download;
|
||
CI jobs 975–977). A vzdump to a local target needs free ≥ newest archive × 1.25 + 1 GiB; a shortfall is a named skip
|
||
before anything starts; fail-open on PBS / first backup / unknown usage. Live, safe builds (a hard stop before vzdump):
|
||
×10 → "local has 14.9 GiB free; … needs about 77.2 GiB" refused; ×1.25 → "space preflight passed" need 11.3 GB,
|
||
avail 16.0 GB. Signed to both demo boxes (`F/F3-*`): demo-felhom committed 22:52:45, demo-hp 22:58:53; hub reads 0.134.0.
|
||
The incident above belongs to this part.
|
||
|
||
## Teardown — three layers
|
||
|
||
- **Machine:** 9202 back on the live catalog (`catalog-cache` at `c802509` at the time; the controller follows `main`),
|
||
window 02:30 again, the six drill apps removed through the product (0 containers, 0 volumes, 0 undo copies); left by
|
||
name on the scratch drive: `userdata/navidrome`, `userdata/romm` (the remove kept drive data — R-442's fail-closed
|
||
path on 9202). The switch key on 9202 is now explicit `true` (was absent = ON). 9202 stays on controller 0.271.0 (the
|
||
floor). Bench 9401 destroyed (hostname checked first), template removed.
|
||
- **Host:** demo-hp `pct list` 9201 + 9202, local-lvm 60.05 %, root 60 %; demo-felhom 9201 only, local-lvm 3.16 %.
|
||
Config changes kept on purpose: restore test ON (both), `local_backup_retention: 1` (demo-hp). No test binaries left
|
||
in /tmp.
|
||
- **Hub:** global floor 0.271.0 (MinAgent 0.131.0); both demo boxes 0.271.0 / agent 0.134.0. Drill repo reset to live
|
||
`main` (`dd9c6ad`), private, Actions off.
|
||
|
||
## Claims in the brief that turned out wrong (or right)
|
||
|
||
1. *The off-site job function is the single place the leg can be chained from on every path* — **TRUE for every path
|
||
of the job, with one exception outside it:** a box with backups disabled registers no off-site job at all, so no
|
||
leg (the update itself would refuse `no_backup` there anyway). The job's own early returns and the scheduler's "not
|
||
configured" skip all return into `chainUpdateLeg`. A controller restart after the job fired loses the rest of the
|
||
night (R-686).
|
||
2. *A demo box has a tested step waiting tonight* — **TRUE for demo-felhom (opengist), FALSE for demo-hp.**
|
||
3. *The whole-box backup lands on demo-hp's root disk* — **TRUE** (`local` = `/var/lib/vz`, root LV).
|
||
4. *Decision 28's suppressions cover a deploy's first start* — **WRONG.** `Deploying` clears when `compose up -d`
|
||
returns; a first start that restarts ≥ 6 times in 10 min is stopped (R-676 updated). An automatic step, its verify
|
||
and its undo ARE covered (pinned by a test).
|
||
|
||
## Register
|
||
|
||
Before: **341 rows / 685,148 B**. After: **338 rows / 682,053 B**. Opened R-685 (backup-fit warning; agent half
|
||
shipped, controller page half open), R-686 (no resume of the leg after a restart), R-687 (live-proof gaps). Closed
|
||
R-672, R-673 (delivered), R-684 (option A), R-680, R-678, R-643. Updated R-676 (deploy first start), R-450 (narrowed
|
||
to part 10).
|