Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
15 KiB
DRILL — night 2026-09-24 (run 2026-09-24 from 11:07 CEST): the second drive brings a file app back whole; a crash loop is stopped; exact image fingerprints; automatic updates spiked; a chaos hour
Run 2026-09-24 from 11:07 CEST (daytime; the brief's clock times were applied relative to the start).
Evidence: night-2026-09-24/ (PROGRESS.md is the step log). Controller v0.269.0 → v0.269.1, hub
v0.123.0, catalog cf7cf84, floor 0.269.1 (MinAgent 0.131.0). Architecture read: 09 (§3
decisions 11–28, §6.4), 07 (the whole-copy table), 08 (the alarm ladder).
Not done, or changed
-
Part D (move more apps) was NOT run. Its bar needs a bench machine on demo-hp that is created and destroyed tonight; this session's permission check refused host-level destroys (it refused a
pct destroy), so a bench could not be torn down — a half-state the rules forbid. Part D is first in the brief's order of sacrifice. Apps moved: 0 of 0 tried. Part E's last step (Update on demo-hp 9201 the standing apps Part D moved) therefore had nothing to press. -
Two controller releases, not one — v0.269.1 fixes a defect v0.269.0's own live proof found (the sync moved an installed app's image digest, so a restart would have changed the image unguarded). Decided by CC unattended — operator may reverse. The floor went to 0.269.1.
-
The crash-loop threshold is 6 restarts in 10 minutes, not the brief's 10 — measured: Docker's back-off caps a steady loop at ~1 restart/min. Decided by CC unattended — operator may reverse.
-
Part E's "whole-box backup" accident was the box's own full backup run, not a vzdump: demo-hp's backup storage had ~4 GB free and its root was at 90 %.
-
Part C step 1 was measured on demo-hp 9201's real night, not on 9202: 9202 has no off-site target, so its off-site leg never runs. 9202 gave the db-dump leg and the window move.
-
demo-hp is damaged and needs the operator (R-672, P1): before tonight's run the agent's scheduled restore-test filled the host's thin pool; 9201's disks went READ-ONLY. The pool was freed (intervention 1); 9201 still needs stop + fsck + start, and did not take the 0.269.1 floor.
-
Round 2's install and round 9's remove were broken by their accidents (R-681, R-682) — recorded, not fixed.
-
A slip of mine: commit
4502af6carriedE/chaos-state.jsonwith the passwords of three drill test accounts (wishlist, navidrome, n8n on scratch guest 9202). All three apps and their data were removed at teardown, so the accounts no longer exist; the file was deleted and ignored inf2234c0. History keeps it — not rewritten (a shared repo). The runner now keeps its state outside the evidence tree.
Interventions: 1 (an agent restart on demo-hp so its own Recover tore down the leaked restore-test guest). Apps moved: 0 of 0 tried. The result that matters most: a file app now comes back WHOLE from the second drive — proven four times, twice under a power cut or a nearly full disk — and nothing the household had was deleted or overwritten.
A1 — the whole restore from the second drive (decision 26, R-661) — PROVEN LIVE
- Spike (
A1/00-spike*): the Tier-2 mirror keeps a file app's drive files as plain files underhdd/<rel>anduserdata/<rel>— readable file by file. R-668 found on the way: the Tier-2 target was a removed folder on the app's own disk (the same-disk check failed open) — fixed, fails closed. - File rules, tested (
tier2_whole_test.go): the whole tree fingerprinted before and after; never delete, never overwrite a newer live file, bring back the missing, keep an older replaced file beside. Red-proofs:redproofs/A1-file-rules.txt,A1-r668.txt. - Live, 9202 (
A1/21-readback.json): restored 1, replaced 0, kept-newer 1, unchanged 137; 3/3 volumes, 1/1 database; 38 s; the account, the deleted file (back), the edited file (the household's newer copy kept) and an untouched file read back. From a hold (A1/30-*): the page named „második meghajtó … a beállításokat, az adatbázist és a fájlokat" / "second drive … the settings, the database and the files"; its restore cleared the hold in 33.8 s. Chaos rounds 3 and 11 repeated it under a power cut and a disk 1 GB above the floor.
A2 — keep-data-only Remove while support is informed (decision 27, R-666) — PROVEN LIVE
One-drive hold (the second-drive mirror moved aside): page hu „…A Felhom ügyfélszolgálatát értesítettük —
addig ne indítsd újra, és ne töröld az adatait." / en "…Felhom support has been told — until then, do not
restart the app or delete its data."; keep_data_only: true; Remove with data or with backups → 409
hu „Amíg az ügyfélszolgálat foglalkozik az alkalmazással, az adatai nem törölhetők — az alkalmazást az
adatai megtartásával eltávolíthatod." / en "While support is working on this app, its data cannot be deleted
— you can remove the app and keep its data."; the app stayed. Positive control: a hold WITH a whole copy
→ keep_data_only absent (A2/10-*). Found (R-669, proven): that hold was reached because the undo
judged the right old version with the FAILED step's probe, left by the previous hold + restore.
A3 — the box stops a crash loop / OOM storm (decision 28, R-667) — PROVEN LIVE
gokapi (a real crash loop): stopped at „crash_loop — 6 in 10m0s", trip 1, page hu/en + Start, no
„restore needed" badge (A3/21); Start lifted it, the fresh container restarted 9 times in 32 s, stopped
again: „Az alkalmazást ismét leállítottuk … A Felhom ügyfélszolgálatát értesítettük." / "We stopped the app
again … Felhom support has been told." (A3/30). Detector red-proofed (redproofs/A3-r667-detector.txt).
Hub v0.123.0 routes app_stopped_unhealthy to the household and the operator (red-proofed, A3-hub.txt).
A4 — each step judged by its own file (R-664, R-665); R-662 removed — shipped, red-proofed
redproofs/A4-*. Live through Part C (three apps climbed with their own step files).
Part B — exact image fingerprints on the box — PROVEN LIVE (on v0.269.1)
- Compose accepts
tag@digestfor every template with a tested digest: 25 apps, 37 digests, 0 missing, 0 parse differences (B/01). - A floating tag (redis:7-alpine) re-tested at a new digest: badge „Frissítés elérhető" (hu/en), Update
pulled exactly that digest, badge „Naprakész", records digest-free (
B/10,B/21). - Found and fixed: v0.269.0's sync wrote the NEW digest into the RUNNING app's compose before Update —
a restart would have changed the image with no backup and no undo. v0.269.1
CarryDigestskeeps an installed app's digest; red-proofed (redproofs/B-sync-carry.txt); re-proven live (B/21: "the sync kept the running digest: True").
Part C — automatic updates, SPIKE — the build brief is 09 §6.4.2
Chain today: legs clock-scheduled, nothing waits. 9201's real night: off-site 04:15:05 → 04:18:33 (3 m 28 s),
gate opens 04:30 → ~11 min for an update leg (decision 20's wait is required). "Off-site finished" signal:
only the persisted offbox.LastStatus+LastRun (offbox.go:1003-1125), not written on five early returns
— build: chain, do not poll. Gate interlock: an Options func beside WindowStartFn in quiesce.runOnce.
Simulated night on 9202 (C/31): romm 3 steps 77/71/94 s; wishlist 56 s; navidrome 11/9 s; vikunja's failing
step undone in 104 s (90 s drill timeout; ≈ 6–11 min at the 5 min default) and set aside; total 7 m 18 s.
Found: stale steps-left after done (R-678), Update on a current app runs a full update (R-679), the box
does not remember a failed step (R-680). Default of the per-box switch: ON (decision 12).
Part D — not run (see "Not done")
Part E — the schedule (committed before round 1)
Drawn by night-2026-09-24/tools/chaos_schedule.py, seed 20260924 (constraints: every action at least
once; controller killed at most 3 times, ≥ 20 min apart; Start-after-unhealthy after a stop; Remove-on-held
after a hold). Guest 9202 (scratch), controller v0.269.1, drill catalog.
seed 20260924, 12 rounds
| round | action | accident |
|---|---|---|
| 1 | oom_app | power_cut |
| 2 | install | docker_restart |
| 3 | cutoff_undo_file_app | power_cut |
| 4 | failing_step | controller_killed |
| 5 | remove_held | none |
| 6 | failing_step | whole_box_backup |
| 7 | crash_loop | whole_box_backup |
| 8 | oom_app | disk_to_floor_plus_1g |
| 9 | remove_held | controller_killed |
| 10 | ladder_step | docker_restart |
| 11 | cutoff_undo_file_app | disk_to_floor_plus_1g |
| 12 | start_after_unhealthy | none |
The app each round acts on — fixed here, not drawn (an app's state after round N decides what N+1 can do):
| round | app | why |
|---|---|---|
| 1 oom_app | chaosoom (drill-only: alpine, 32M cap, a 200 MB eater in a loop) |
out-of-memory storm → the box must stop it |
| 2 install | n8n at its ladder's FIRST from (2 steps behind), seeded through its front door |
gives round 10 a step |
| 3 cutoff_undo_file_app | nextcloud (the file app), a failing step (redis 7-alpine → 7.4.0-alpine, probe 8999), the NEWEST undo copy's marker cut at verifying → the hold must name the second drive → its whole restore, then the seeds + files read back | decision 26 under a power cut |
| 4 failing_step | wishlist v0.67.1 → :latest with probe 8999 → undo |
the undo under a controller kill |
| 5 remove_held | gokapi (held since A3: the box stopped its crash loop) — Remove with data | Remove on a held app |
| 6 failing_step | navidrome 0.64.1 → :latest with probe 8999 → undo |
the undo under a backup run |
| 7 crash_loop | chaoscrash (drill-only: alpine that exits after 3 s) |
the box must stop it |
| 8 oom_app | chaosoomb (a second eater) |
the OOM stop under a full disk |
| 9 remove_held | chaoscrash (held since round 7) |
Remove on a held app under a controller kill |
| 10 ladder_step | n8n, one step (2 → 1) | a good step under a docker restart |
| 11 cutoff_undo_file_app | nextcloud again | decision 26 with the disk 1 GB above the floor (a disk refusal is a valid outcome) |
| 12 start_after_unhealthy | chaosoomb (held since round 8) — Start → one more try → the repeat stop says support is told |
decision 28's second half |
Changed before round 1: the accident "a whole-box backup started on 9202" is the box's own full backup run
(POST /api/backup/run, as on 2026-09-23), not a vzdump: demo-hp's backup storage had ~4 GB free and the host
root was at 90 %, so a vzdump of 9202 could have filled the HOST (demo-hp also had R-672's full pool today).
How accidents are fired: the runner arms the round and waits; the session fires each host-level accident
as its own visible command (power cut = pct stop 9202; pct start 9202, kill -9 of the controller PID,
systemctl restart docker in the guest, a fallocate to leave 3 GB free on /var/lib/docker).
The rounds
| # | action × accident | app | steady | end state | data read back | events (from the box's own log) |
|---|---|---|---|---|---|---|
| 1 | oom_app × power cut | chaosoom | 185 s | stopped, unhealthy_stop |
— | app_oom (w), app_oom_storm (e), app_stopped_unhealthy (w) |
| 2 | install × docker restart | n8n | — | not installed, silently (R-681); reinstalled after, 211 s | — | app_deploy_started only |
| 3 | cutoff file app × power cut at verifying | nextcloud | 204 s | the update RESUMED after boot → held (second drive) → whole restore → running | account ✓ files 2/2 ✓ | app_update_held (e) |
| 4 | failing step × controller kill | wishlist | 148 s | resumed → undone | ✓ | app_update_undone (w) |
| 5 | remove held × none | gokapi | 27 s | removed, 0 leftovers | — | app_removed (i) |
| 6 | failing step × backup run | navidrome | 110 s | undone (the backup skipped the updating app) | ✓ | app_update_undone (w) |
| 7 | crash loop × backup run | chaoscrash | 116 s | stopped, unhealthy_stop |
— | app_stopped_unhealthy (w) |
| 8 | oom_app × disk to floor + 1 GB | chaosoomb | 102 s | stopped, unhealthy_stop |
— | app_oom (w), app_stopped_unhealthy (w) |
| 9 | remove held × controller kill | chaoscrash | — | half-removed: 502 to the household, containers gone, still listed (R-682); a second Remove finished it | — | — |
| 10 | ladder step × docker restart | n8n | 155 s | resumed → done, 2.31.3 → 2.40.5, steps 2 → 1 | ✓ | — |
| 11 | cutoff file app × disk to floor + 1 GB | nextcloud | 194 s | held (fresh 14:34 second-drive copy) → whole restore → running | account ✓ files 2/2 ✓ | app_update_held (e) |
| 12 | Start after unhealthy × none | chaosoomb | 125 s | one more try → stopped again, „support is told" hu/en | — | app_oom (w), app_oom_storm (e), app_stopped_unhealthy (w) |
Controller kills: 2 of max 3, 20 min 32 s apart. Against 08: every stop sent app_stopped_unhealthy;
every undo app_update_undone, every hold app_update_held. What should have fired and did not: round 2
(an install lost — nothing) and round 9 (a remove half-done — nothing); round 3's hold named an hour-old copy
(R-683, watch). No user file was deleted or overwritten by any restore; no data was lost.
Teardown — three layers
- Machine (9202): tonight's 8 apps removed through the product (0 containers, 0 volumes, 0 undo copies);
the three drive folders the product kept (R-442 on a box without drive access) removed by name; backup
window back to 02:30 through the product's form;
controller.yamlrestored from.pre-night0924(live catalog, no drillupdate:settings); controller stays on 0.269.1 (the floor). Standing before tonight and kept: privatebin, paperless-ngx, filebrowser. gokapi removed in round 5 (R-644's crash loop).F/10-12. - Host (demo-hp):
pct list9201 + 9202 only (as at the start, minus R-672's leaked 990000); no bench was created; storage unchanged; the thin pool at 58.9 % after intervention 1.F/20. - Hub: floor 0.269.1 with MinAgent 0.131.0, read back; N100 9201 arrived in ~16 s (
managed floor SERVED … from declared); demo-hp 9201 did NOT arrive (read-only disks, R-672). Drill repo reset to livecf7cf8456f00,has_actions: false.F/13,F/30-31.
Claims in the brief that turned out wrong (or right), named
- The Tier-2 mirror holds a file app's files in a form a restore can read file by file — TRUE
(plain files under
hdd/anduserdata/; the merge walked 139). RestartCountsurvives the scans that resetrestarting_since— TRUE per container run; it resets when the container is recreated (a Start after the box's stop), so the detector treats a drop as a new window.- 10 restarts in 10 minutes separates a crash loop from a slow start — WRONG: a steady loop runs at ~1 restart/min (gokapi 7 in 7 min) and would never reach 10; 6 chosen. No healthy first start in the drill evidence reached 6 (immich's broken first start reached 12 — R-676).
- Docker Compose accepts
tag@digestin every template's image line — TRUE (25/25). What the brief did not foresee: rendering digests in the SYNC moved a running app's image (fixed v0.269.1). - The chain signal can be measured on 9202 — WRONG: 9202 has no off-site target.
- One release per repo could not hold (see "Not done").
Register
Before 335 rows / 679,393 B; after 344 rows / 685,662 B. Opened R-668…R-683 (16); closed R-661, R-662, R-664, R-665, R-666, R-667, R-668 (7, compressed into CLOSED-ITEMS).