e924a17632
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
202 lines
15 KiB
Markdown
202 lines
15 KiB
Markdown
# DRILL — night 2026-09-24 (run 2026-09-24 from 11:07 CEST): the second drive brings a file app back whole; a crash loop is stopped; exact image fingerprints; automatic updates spiked; a chaos hour
|
||
|
||
Run 2026-09-24 from 11:07 CEST (daytime; the brief's clock times were applied relative to the start).
|
||
Evidence: `night-2026-09-24/` (PROGRESS.md is the step log). Controller **v0.269.0 → v0.269.1**, hub
|
||
**v0.123.0**, catalog **cf7cf84**, floor **0.269.1** (MinAgent 0.131.0). Architecture read: `09` (§3
|
||
decisions 11–28, §6.4), `07` (the whole-copy table), `08` (the alarm ladder).
|
||
|
||
## Not done, or changed
|
||
|
||
- **Part D (move more apps) was NOT run.** Its bar needs a bench machine on demo-hp that is created and
|
||
destroyed tonight; this session's permission check refused host-level destroys (it refused a
|
||
`pct destroy`), so a bench could not be torn down — a half-state the rules forbid. Part D is first in
|
||
the brief's order of sacrifice. **Apps moved: 0 of 0 tried.** Part E's last step (Update on demo-hp
|
||
9201 the standing apps Part D moved) therefore had nothing to press.
|
||
- **Two controller releases, not one** — v0.269.1 fixes a defect v0.269.0's own live proof found (the
|
||
sync moved an installed app's image digest, so a restart would have changed the image unguarded).
|
||
*Decided by CC unattended — operator may reverse.* The floor went to 0.269.1.
|
||
- **The crash-loop threshold is 6 restarts in 10 minutes, not the brief's 10** — measured: Docker's
|
||
back-off caps a steady loop at ~1 restart/min. *Decided by CC unattended — operator may reverse.*
|
||
- **Part E's "whole-box backup" accident was the box's own full backup run**, not a vzdump: demo-hp's
|
||
backup storage had ~4 GB free and its root was at 90 %.
|
||
- **Part C step 1 was measured on demo-hp 9201's real night, not on 9202**: 9202 has no off-site target,
|
||
so its off-site leg never runs. 9202 gave the db-dump leg and the window move.
|
||
- **demo-hp is damaged and needs the operator (R-672, P1):** before tonight's run the agent's scheduled
|
||
restore-test filled the host's thin pool; 9201's disks went READ-ONLY. The pool was freed (intervention
|
||
1); 9201 still needs stop + fsck + start, and did not take the 0.269.1 floor.
|
||
- Round 2's install and round 9's remove were broken by their accidents (R-681, R-682) — recorded, not fixed.
|
||
|
||
- **A slip of mine:** commit `4502af6` carried `E/chaos-state.json` with the passwords of three drill test
|
||
accounts (wishlist, navidrome, n8n on scratch guest 9202). All three apps and their data were removed at
|
||
teardown, so the accounts no longer exist; the file was deleted and ignored in `f2234c0`. History keeps it —
|
||
not rewritten (a shared repo). The runner now keeps its state outside the evidence tree.
|
||
|
||
**Interventions: 1** (an agent restart on demo-hp so its own Recover tore down the leaked restore-test
|
||
guest). **Apps moved: 0 of 0 tried.** **The result that matters most:** a file app now comes back WHOLE
|
||
from the second drive — proven four times, twice under a power cut or a nearly full disk — and nothing the
|
||
household had was deleted or overwritten.
|
||
|
||
## A1 — the whole restore from the second drive (decision 26, R-661) — PROVEN LIVE
|
||
|
||
- **Spike** (`A1/00-spike*`): the Tier-2 mirror keeps a file app's drive files as plain files under
|
||
`hdd/<rel>` and `userdata/<rel>` — readable file by file. **R-668 found on the way:** the Tier-2 target
|
||
was a removed folder on the app's own disk (the same-disk check failed open) — fixed, fails closed.
|
||
- **File rules, tested** (`tier2_whole_test.go`): the whole tree fingerprinted before and after; never
|
||
delete, never overwrite a newer live file, bring back the missing, keep an older replaced file beside.
|
||
Red-proofs: `redproofs/A1-file-rules.txt`, `A1-r668.txt`.
|
||
- **Live, 9202** (`A1/21-readback.json`): restored 1, replaced 0, kept-newer 1, unchanged 137; 3/3 volumes,
|
||
1/1 database; 38 s; the account, the deleted file (back), the edited file (the household's newer copy kept)
|
||
and an untouched file read back. **From a hold** (`A1/30-*`): the page named „második meghajtó … a
|
||
beállításokat, az adatbázist és a fájlokat" / "second drive … the settings, the database and the files";
|
||
its restore cleared the hold in 33.8 s. **Chaos rounds 3 and 11** repeated it under a power cut and a disk
|
||
1 GB above the floor.
|
||
|
||
## A2 — keep-data-only Remove while support is informed (decision 27, R-666) — PROVEN LIVE
|
||
|
||
One-drive hold (the second-drive mirror moved aside): page hu „…A Felhom ügyfélszolgálatát értesítettük —
|
||
addig ne indítsd újra, és ne töröld az adatait." / en "…Felhom support has been told — until then, do not
|
||
restart the app or delete its data."; `keep_data_only: true`; Remove with data or with backups → **409**
|
||
hu „Amíg az ügyfélszolgálat foglalkozik az alkalmazással, az adatai nem törölhetők — az alkalmazást az
|
||
adatai megtartásával eltávolíthatod." / en "While support is working on this app, its data cannot be deleted
|
||
— you can remove the app and keep its data."; the app stayed. Positive control: a hold WITH a whole copy
|
||
→ `keep_data_only` absent (`A2/10-*`). **Found (R-669, proven):** that hold was reached because the undo
|
||
judged the right old version with the FAILED step's probe, left by the previous hold + restore.
|
||
|
||
## A3 — the box stops a crash loop / OOM storm (decision 28, R-667) — PROVEN LIVE
|
||
|
||
gokapi (a real crash loop): stopped at „crash_loop — 6 in 10m0s", trip 1, page hu/en + Start, no
|
||
„restore needed" badge (`A3/21`); **Start** lifted it, the fresh container restarted 9 times in 32 s, stopped
|
||
again: „Az alkalmazást ismét leállítottuk … A Felhom ügyfélszolgálatát értesítettük." / "We stopped the app
|
||
again … Felhom support has been told." (`A3/30`). Detector red-proofed (`redproofs/A3-r667-detector.txt`).
|
||
Hub v0.123.0 routes `app_stopped_unhealthy` to the household and the operator (red-proofed, `A3-hub.txt`).
|
||
|
||
## A4 — each step judged by its own file (R-664, R-665); R-662 removed — shipped, red-proofed
|
||
|
||
`redproofs/A4-*`. Live through Part C (three apps climbed with their own step files).
|
||
|
||
## Part B — exact image fingerprints on the box — PROVEN LIVE (on v0.269.1)
|
||
|
||
- Compose accepts `tag@digest` for **every** template with a tested digest: 25 apps, 37 digests, 0 missing,
|
||
0 parse differences (`B/01`).
|
||
- A floating tag (redis:7-alpine) re-tested at a new digest: badge „Frissítés elérhető" (hu/en), Update
|
||
pulled exactly that digest, badge „Naprakész", records digest-free (`B/10`, `B/21`).
|
||
- **Found and fixed:** v0.269.0's sync wrote the NEW digest into the RUNNING app's compose before Update —
|
||
a restart would have changed the image with no backup and no undo. v0.269.1 `CarryDigests` keeps an
|
||
installed app's digest; red-proofed (`redproofs/B-sync-carry.txt`); re-proven live (`B/21`: "the sync kept
|
||
the running digest: True").
|
||
|
||
## Part C — automatic updates, SPIKE — the build brief is `09` §6.4.2
|
||
|
||
Chain today: legs clock-scheduled, nothing waits. 9201's real night: off-site 04:15:05 → 04:18:33 (3 m 28 s),
|
||
gate opens 04:30 → ~11 min for an update leg (decision 20's wait is required). "Off-site finished" signal:
|
||
only the persisted `offbox.LastStatus`+`LastRun` (`offbox.go:1003-1125`), not written on five early returns
|
||
— **build: chain, do not poll.** Gate interlock: an `Options` func beside `WindowStartFn` in `quiesce.runOnce`.
|
||
Simulated night on 9202 (`C/31`): romm 3 steps 77/71/94 s; wishlist 56 s; navidrome 11/9 s; vikunja's failing
|
||
step undone in 104 s (90 s drill timeout; ≈ 6–11 min at the 5 min default) and set aside; total 7 m 18 s.
|
||
Found: stale steps-left after `done` (R-678), Update on a current app runs a full update (R-679), the box
|
||
does not remember a failed step (R-680). Default of the per-box switch: **ON** (decision 12).
|
||
|
||
## Part D — not run (see "Not done")
|
||
|
||
## Part E — the schedule (committed before round 1)
|
||
|
||
Drawn by `night-2026-09-24/tools/chaos_schedule.py`, seed **20260924** (constraints: every action at least
|
||
once; controller killed at most 3 times, ≥ 20 min apart; Start-after-unhealthy after a stop; Remove-on-held
|
||
after a hold). Guest **9202** (scratch), controller v0.269.1, drill catalog.
|
||
|
||
seed 20260924, 12 rounds
|
||
| round | action | accident |
|
||
|---|---|---|
|
||
| 1 | oom_app | power_cut |
|
||
| 2 | install | docker_restart |
|
||
| 3 | cutoff_undo_file_app | power_cut |
|
||
| 4 | failing_step | controller_killed |
|
||
| 5 | remove_held | none |
|
||
| 6 | failing_step | whole_box_backup |
|
||
| 7 | crash_loop | whole_box_backup |
|
||
| 8 | oom_app | disk_to_floor_plus_1g |
|
||
| 9 | remove_held | controller_killed |
|
||
| 10 | ladder_step | docker_restart |
|
||
| 11 | cutoff_undo_file_app | disk_to_floor_plus_1g |
|
||
| 12 | start_after_unhealthy | none |
|
||
|
||
**The app each round acts on — fixed here, not drawn** (an app's state after round N decides what N+1 can do):
|
||
|
||
| round | app | why |
|
||
|---|---|---|
|
||
| 1 oom_app | `chaosoom` (drill-only: alpine, 32M cap, a 200 MB eater in a loop) | out-of-memory storm → the box must stop it |
|
||
| 2 install | n8n at its ladder's FIRST `from` (2 steps behind), seeded through its front door | gives round 10 a step |
|
||
| 3 cutoff_undo_file_app | nextcloud (the file app), a failing step (redis 7-alpine → 7.4.0-alpine, probe 8999), the NEWEST undo copy's marker cut at verifying → the hold must name the second drive → its whole restore, then the seeds + files read back | decision 26 under a power cut |
|
||
| 4 failing_step | wishlist v0.67.1 → `:latest` with probe 8999 → undo | the undo under a controller kill |
|
||
| 5 remove_held | gokapi (held since A3: the box stopped its crash loop) — Remove with data | Remove on a held app |
|
||
| 6 failing_step | navidrome 0.64.1 → `:latest` with probe 8999 → undo | the undo under a backup run |
|
||
| 7 crash_loop | `chaoscrash` (drill-only: alpine that exits after 3 s) | the box must stop it |
|
||
| 8 oom_app | `chaosoomb` (a second eater) | the OOM stop under a full disk |
|
||
| 9 remove_held | `chaoscrash` (held since round 7) | Remove on a held app under a controller kill |
|
||
| 10 ladder_step | n8n, one step (2 → 1) | a good step under a docker restart |
|
||
| 11 cutoff_undo_file_app | nextcloud again | decision 26 with the disk 1 GB above the floor (a `disk` refusal is a valid outcome) |
|
||
| 12 start_after_unhealthy | `chaosoomb` (held since round 8) — Start → one more try → the repeat stop says support is told | decision 28's second half |
|
||
|
||
**Changed before round 1:** the accident "a whole-box backup started on 9202" is the box's own full backup run
|
||
(`POST /api/backup/run`, as on 2026-09-23), not a vzdump: demo-hp's backup storage had ~4 GB free and the host
|
||
root was at 90 %, so a vzdump of 9202 could have filled the HOST (demo-hp also had R-672's full pool today).
|
||
**How accidents are fired:** the runner arms the round and waits; the session fires each host-level accident
|
||
as its own visible command (power cut = `pct stop 9202; pct start 9202`, kill -9 of the controller PID,
|
||
`systemctl restart docker` in the guest, a `fallocate` to leave 3 GB free on `/var/lib/docker`).
|
||
|
||
|
||
### The rounds
|
||
|
||
| # | action × accident | app | steady | end state | data read back | events (from the box's own log) |
|
||
|---|---|---|---|---|---|---|
|
||
| 1 | oom_app × power cut | chaosoom | 185 s | stopped, `unhealthy_stop` | — | app_oom (w), app_oom_storm (e), app_stopped_unhealthy (w) |
|
||
| 2 | install × docker restart | n8n | — | **not installed, silently** (R-681); reinstalled after, 211 s | — | app_deploy_started only |
|
||
| 3 | cutoff file app × power cut at verifying | nextcloud | 204 s | the update RESUMED after boot → held (second drive) → whole restore → running | account ✓ files 2/2 ✓ | app_update_held (e) |
|
||
| 4 | failing step × controller kill | wishlist | 148 s | resumed → undone | ✓ | app_update_undone (w) |
|
||
| 5 | remove held × none | gokapi | 27 s | removed, 0 leftovers | — | app_removed (i) |
|
||
| 6 | failing step × backup run | navidrome | 110 s | undone (the backup skipped the updating app) | ✓ | app_update_undone (w) |
|
||
| 7 | crash loop × backup run | chaoscrash | 116 s | stopped, `unhealthy_stop` | — | app_stopped_unhealthy (w) |
|
||
| 8 | oom_app × disk to floor + 1 GB | chaosoomb | 102 s | stopped, `unhealthy_stop` | — | app_oom (w), app_stopped_unhealthy (w) |
|
||
| 9 | remove held × controller kill | chaoscrash | — | **half-removed**: 502 to the household, containers gone, still listed (R-682); a second Remove finished it | — | — |
|
||
| 10 | ladder step × docker restart | n8n | 155 s | resumed → done, 2.31.3 → 2.40.5, steps 2 → 1 | ✓ | — |
|
||
| 11 | cutoff file app × disk to floor + 1 GB | nextcloud | 194 s | held (fresh 14:34 second-drive copy) → whole restore → running | account ✓ files 2/2 ✓ | app_update_held (e) |
|
||
| 12 | Start after unhealthy × none | chaosoomb | 125 s | one more try → stopped again, „support is told" hu/en | — | app_oom (w), app_oom_storm (e), app_stopped_unhealthy (w) |
|
||
|
||
Controller kills: 2 of max 3, 20 min 32 s apart. **Against `08`:** every stop sent `app_stopped_unhealthy`;
|
||
every undo `app_update_undone`, every hold `app_update_held`. **What should have fired and did not:** round 2
|
||
(an install lost — nothing) and round 9 (a remove half-done — nothing); round 3's hold named an hour-old copy
|
||
(R-683, watch). No user file was deleted or overwritten by any restore; no data was lost.
|
||
|
||
## Teardown — three layers
|
||
|
||
- **Machine (9202):** tonight's 8 apps removed through the product (0 containers, 0 volumes, 0 undo copies);
|
||
the three drive folders the product kept (R-442 on a box without drive access) removed by name; backup
|
||
window back to 02:30 through the product's form; `controller.yaml` restored from `.pre-night0924` (live
|
||
catalog, no drill `update:` settings); controller stays on 0.269.1 (the floor). Standing before tonight and
|
||
kept: privatebin, paperless-ngx, filebrowser. gokapi removed in round 5 (R-644's crash loop). `F/10-12`.
|
||
- **Host (demo-hp):** `pct list` 9201 + 9202 only (as at the start, minus R-672's leaked 990000); no bench
|
||
was created; storage unchanged; the thin pool at 58.9 % after intervention 1. `F/20`.
|
||
- **Hub:** floor 0.269.1 with MinAgent 0.131.0, read back; N100 9201 arrived in ~16 s (`managed floor SERVED
|
||
… from declared`); **demo-hp 9201 did NOT arrive** (read-only disks, R-672). Drill repo reset to live
|
||
`cf7cf8456f00`, `has_actions: false`. `F/13`, `F/30-31`.
|
||
|
||
## Claims in the brief that turned out wrong (or right), named
|
||
|
||
1. *The Tier-2 mirror holds a file app's files in a form a restore can read file by file* — **TRUE**
|
||
(plain files under `hdd/` and `userdata/`; the merge walked 139).
|
||
2. *`RestartCount` survives the scans that reset `restarting_since`* — **TRUE per container run**; it resets
|
||
when the container is recreated (a Start after the box's stop), so the detector treats a drop as a new window.
|
||
3. *10 restarts in 10 minutes separates a crash loop from a slow start* — **WRONG**: a steady loop runs at
|
||
~1 restart/min (gokapi 7 in 7 min) and would never reach 10; 6 chosen. No healthy first start in the
|
||
drill evidence reached 6 (immich's broken first start reached 12 — R-676).
|
||
4. *Docker Compose accepts `tag@digest` in every template's image line* — **TRUE** (25/25). What the brief
|
||
did not foresee: rendering digests in the SYNC moved a running app's image (fixed v0.269.1).
|
||
5. *The chain signal can be measured on 9202* — **WRONG**: 9202 has no off-site target.
|
||
6. *One release per repo* could not hold (see "Not done").
|
||
|
||
## Register
|
||
|
||
Before **335 rows / 679,393 B**; after **344 rows / 685,662 B**. Opened R-668…R-683 (16); closed R-661,
|
||
R-662, R-664, R-665, R-666, R-667, R-668 (7, compressed into CLOSED-ITEMS).
|
||
|