DRILL-night-2026-09-23.md: Parts A-E. 09 §3 decisions 21 (operator word), 22 and 23 (CC unattended, operator may reverse); §6.4 parts 4 and 6 (catalog half) shipped; §6.1a residuals R-658/R-659. Register 330 -> 336: R-651..R-660 opened (R-658 and R-659 P1), R-650/R-640/R-499/R-626 closed. Capability map, nightly rotation (opengist), STATUS (one question: the floor), CONTEXT, REPORT. The floor stays 0.266.0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
29 KiB
DRILL — night shift 2026-09-23: DooPlex guarded from its own tests, the test record, the proven moves, the chaos hour
Evidence: audits/night-2026-09-23/ (PROGRESS.md is the step log). Architecture read first and named:
architecture/09-update-architecture.md (§3 decisions 11–20, §6.1a, §6.4 parts 4–6, the verdict record),
07-backup-architecture.md §6, 08-alarm-ladder.md. Method: endpoint-level — every product act is
the endpoint the UI invokes; pages are fetched as HTML in both languages; no browser.
Not done, or changed
- The floor stays at 0.266.0. The brief raises it only if every live proof of Parts A, B and D passed; Part D found two P1s (R-658, R-659). Both exist in 0.266.0 as well — v0.267.0 does not cause them and adds R-640's protection to the same restore path — so the recommendation is to raise it; that is the operator's one question this morning.
- Adventurelog v0.13.0 was not moved (R-655): our own template's healthcheck points at a node path the new image moved, and the new backend downloads world data from the internet at every start and crash-loops when the download is cut short.
- The stop rule fired in round 11 (R-659): all twelve rounds had already run; nothing further was injected.
- Rounds 2–5 did not land their accidents inside the action (the instrument, stated in Part D); round 9 is the clean mid-update kill, round 6 the mid-verify memory hog, rounds 8 and 11 the mid-verify backup and the concurrent second update.
- Decisions taken by CC unattended (operator may reverse,
09§3): 22 — the memory mark reads the app's own memory, not the file cache; 23 — the push-time gate asks the registry for moved refs only. - R-518's per-tier quiesce, R-612's silent-seed half and R-613's template sweep stay open, narrowed.
Three lines
- Interventions: none by the operator; 5 by me on my own instruments (a stale-template wait, a dropped event ring, a killed runner of my own, a broken local edit reverted before any commit, the stop rule on a held app) — each stated where it happened. No product state was set by hand.
- Apps moved tonight: 11 of 14 tried (12 steps of 15 edges), each with its test record; the 21 moves of 2026-09-22 backfilled.
- The result that matters most: after a restore, an app's volumes lose the label the undo finds them by — so the next failed update is "undone" with nothing put back (R-658, P1); and a held file-leg app can be pointed at a restore that refuses it (R-659, P1).
Part A — controller v0.267.0 (80e6ad8c4772, CI job 919 success)
| row | what shipped | red-proof (seen failing, then restored) | evidence |
|---|---|---|---|
| R-650 | internal/dockerexec: every docker exec (77 sites) goes through it; under go test a real docker is refused with an error naming the command, unless FELHOM_TEST_REAL_DOCKER=1 or the binary is a stub under the temp dir. TestR650_NoBareDockerExec sweeps the tree. |
guard disabled → the decoy's docker ps -a runs (got <nil>); one bare call put back → the sweep names dbdump.go:140 |
A1-r650-*.txt |
| R-650 sweep | 8 api tests FAILED on the refusal — they built a real stacks.Manager and ran docker ps on DooPlex; web ran 67× docker compose version, 19× docker ps, 19× docker info, 10× docker inspect felhom-samba, 2× docker exec felhom-samba; stacks ran 1× docker-compose down, 11× compose ps; backup 46× docker ps; appexport 7×; system 1×. api/stacks/web now run under RunWithStub (TestMain); the other three pass on the refusal. |
— | A1-r650-sweep.txt |
| R-640 | appbackup.CheckDumpComplete (the engine's end marker, last 4 KiB, gzip-aware). The unit restore and the off-site restore refuse a cut-off copy before the first mutation; every replay checks again before any load, whatever the import seam is. |
replay check removed → a cut-off copy was LOADED; unit gate removed → stopped=true calls=[stop recreate startsvc:immich-postgres start]; off-site gate removed → only the replay check caught it, after a stop and a rollback |
A2-r640-redproof.txt |
| R-626 | measured, not reproduced on v0.266.0: navidrome deployed and removed through the product; docker events 390 s — the remove's kill/stop/die/destroy seen (the positive observable), no create; controller restarted at +150 s; guest rebooted after → 0 containers, 0 volumes at every check. My script first said came_back: true — it had counted the catalog-mirror directory every catalog app has; corrected in the verdict file, both values kept. Leftover found: applied-compose.yml + applied-meta/ stay after a remove (R-651). |
— | A3-r626-* |
| R-499 | the Tier-2 page's system-disk sentence has four branches (own drive / same disk / drive gone / cannot ask); „(PBS)" and „nincs külön teendő" only where true. Live on 9202 (no agent): the unknown branch in hu + en, matching /api/storage/backup-target known:false. |
handler not passing the fact → 3 branches missing; the old sentence in the same-disk branch → promises the backup protects this app |
A4-*, A7-* |
| R-518 | the backup button states the measured stop (≈ 8 min on a 12-app box), both languages, page + confirm. The brief's „csak néhány másodpercre" was already gone (v0.243.0); the vague „általában néhány perc" is replaced. | bundles back to v0.243.0 → measured downtime appears 0 times | A5-* |
Gates: go build/vet rc 0; go test -count=1 ./... rc 0 (31 packages); controller_gates.py all OK.
Deployed to 9202 only (A6-*); the floor waits for Part E.
Part B — the test record and its gate (catalog 6db08a5, CI job 920 success)
Spike first (15 min): a drill navidrome carrying a top-level update_ladder: (one JSON entry per
line) on 9202 at v0.266.0 and v0.267.0: synced, deploy-fields read, deployed running, probe
API GET :4533/ping → 200, badge „Naprakész" / "Up to date", no YAML error. Source agrees (no
KnownFields anywhere). The brief's claim holds. B1-spike-*.
The record (scripts/ladder.py): from/to per service, digest per to ref (sha256 from
scripts/image_digest.py — stdlib, equals Docker's RepoDigests for privatebin/pdo:2.0.6 on 9202),
verdict (proven | unrecorded), tested_at, harness_version, evidence, box_evidence,
memory_peak_pct (+ memory_basis, memory_cgroup_peak_pct), marks {files_may_change, needs_person,
memory_tight}. Written ONLY by upgrade-test.py --write-ladder (bench AND box proven, template at
FROM, digests resolved) — test_ladder_writer.py 5/5.
The gate — two rows of catalog_gates.py, both --fast:
check-test-record.py (static; runs in CI too: a ladder well-formed, continuous, and its newest step IS
the compose) and check-test-record-move.py (history + the registry for MOVED refs only: a move adds a
PROVEN, non-backfilled entry whose from/to are the compose before/after and whose digests the
registry still serves; memory_tight needs the limit raised in the same range). 16 decoy cases, both
directions. Red-proofs: the no-entry refusal removed → a bare image move and the entry only in
README both pass (rc 0, expected 1); failed allowed → both failed-verdict cases pass; the digest
comparison removed → the registry now serves another digest passes. B2-gate-redproofs.txt.
Live, on the first real move (romm): the gate judged it against the real registry (rc 0), and with a
wrong digest table it REFUSED all three services.
Backfill — the 21 moves of 2026-09-22, each from the record its commit cited:
| app | step (the service that moved) | verdict | evidence | note |
|---|---|---|---|---|
| actualbudget | actualbudget: actual-server:26.7.0 → actual-server:26.9.0 | proven | audits/update-night-2026-09-21/apps/actualbudget/verdict.json |
— |
| audiobookshelf | audiobookshelf: audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | proven | audits/update-night-2026-09-21/apps/audiobookshelf/verdict.json |
— |
| bookstack | bookstack: bookstack:26.05.2 → bookstack:26.05.5 | proven | audits/update-night-2026-09-21/apps/bookstack/verdict.json |
— |
| docmost | docmost: docmost:0.95.0 → docmost:0.96.0 | proven | audits/update-night-2026-09-21/apps/docmost/verdict.json |
— |
| emby | emby: embyserver:4.10.0.20 → embyserver:4.11.0.1 | proven | audits/the-28-2026-09-22/apps/emby/verdict.json |
— |
| ghost | ghost: ghost:6.53.0-alpine → ghost:6.64.0-alpine | proven | audits/the-28-2026-09-22/apps/ghost/verdict.json |
— |
| grafana | grafana: grafana:13.1.0 → grafana:13.2.2 | proven | audits/update-night-2026-09-21/apps/grafana/verdict.json |
— |
| home-assistant | home-assistant: home-assistant:2026.7.2 → home-assistant:2026.9.3 | proven | audits/update-night-2026-09-21/apps/home-assistant/verdict.json |
— |
| immich | immich-server: immich-server:v3.0.3 → immich-server:v3.2.2 | proven | audits/the-28-2026-09-22/apps/immich/verdict.json |
— |
| mealie | mealie: mealie:v3.20.1 → mealie:v3.27.0 | proven | audits/update-night-2026-09-21/apps/mealie/verdict.json |
— |
| n8n | n8n: n8n:2.31.3 → n8n:2.40.5 | proven | audits/update-night-2026-09-21/apps/n8n/verdict.json |
— |
| navidrome | navidrome: navidrome:0.63.2 → navidrome:0.64.0 | proven | audits/update-night-2026-09-21/apps/navidrome/verdict.json |
— |
| nextcloud | nextcloud-db: mariadb:11.6 → mariadb:12.3 | proven | audits/update-night-2026-09-21/apps/nextcloud-engine-mariadb/verdict.json |
commit cited no record; this one named |
| papra | papra: papra:26.6.1-rootless → papra:26.6.2-rootless | proven | audits/update-night-2026-09-21/apps/papra/verdict.json |
— |
| privatebin | privatebin: pdo:2.0.5 → pdo:2.0.6 | proven | audits/update-night-2026-09-21/apps/privatebin/verdict.json |
— |
| radarr | radarr: radarr:6.3.0 → radarr:6.4.4 | proven | audits/the-28-2026-09-22/apps/radarr/verdict.json |
— |
| romm | romm: romm:5.0.0 → romm:5.3.0 | proven | audits/update-night-2026-09-21/apps/romm/verdict.json |
memory watch 80.9 % (M1) |
| sonarr | sonarr: sonarr:4.0.19 → sonarr:4.0.20 | proven | audits/the-28-2026-09-22/apps/sonarr/verdict.json |
— |
| tandoor | tandoor: recipes:2.6.13 → recipes:2.6.15 | proven | audits/update-night-2026-09-21/apps/tandoor/verdict.json |
— |
| termix | termix: termix:2.5.0 → termix:2.8.0 | proven | audits/the-28-2026-09-22/apps/termix/verdict.json |
— |
| vikunja | vikunja: vikunja:2.3.0 → vikunja:2.6.0 | proven | audits/update-night-2026-09-21/apps/vikunja/verdict.json |
— |
21 of 21 proven, 0 unrecorded. Every entry is marked backfilled, carries harness v1 (box walk only, no memory watch — except romm's M1 figure), and a digest resolved TONIGHT, which each entry's note says. scripts/ladder_backfill.py.
Part C — move every app that proves itself
Venues: the bench = throwaway LXC 9401 on demo-hp (Debian 13, docker 26.1.5, Docker Hub login),
created tonight and destroyed at teardown; the box = guest 9202, v0.267.0, drill catalog. Negative
control first: C3 (alpine:3.20 as the TO image) → failed — the bench measures (bench-C3/).
Box walks by run_edge.py; bench edges by upgrade-test.py --move; each edge's bench evidence copied off
the bench when it ended (benchq.py). Evidence per app: night-2026-09-23/apps/<app>/.
| app | from → to | class | bench verdict | box verdict | memory: own (cgroup) peak % | migration line | moved? |
|---|---|---|---|---|---|---|---|
| nextcloud | nextcloud 34.0.1 → 34.0.4 | MariaDB + file-leg | proven | proven (done) | 24.2 (100) | nextcloud / Upgrading nextcloud from 34.0.1.2 ... | moved 5d49a11 (files_may_change) |
| romm | romm 5.3.0 → 5.3.1 | MariaDB | proven | proven (done) | n/a — measured before the own-memory sample existed (cgroup 79.5) | romm / INFO: [RomM][init][2026-09-23 21:12:26][0 | moved 7ff4e68 |
| romm-engine | romm-db mariadb 11.4 → 11.8 (own edge) | MariaDB engine | proven | proven (done) | 78.0 (84) | romm / INFO: [RomM][init][2026-09-23 21:47:51][0 | moved 431cdec |
| kimai | kimai-db mariadb 11.6 → 11.8 (own edge) | MariaDB engine | proven | proven (done) | 37.9 (100) | kimai / [OK] Already at the latest version ("DoctrineMigrations\Version2 | moved 376b2c3 |
| adventurelog | adventurelog v0.12.1 → v0.13.0 (both images) | PostgreSQL (PostGIS) | failed ×2 (frontend never healthy; then backend crash-loop on download-countries) |
failed (undone at +196 s, drill timeout 90 s) | — | adventurelog / Apply all migrations: account, admin, adventures, aut | NOT moved — our template's frontend healthcheck points at /nodejs/bin/node, gone in v0.13.0; with it dropped, the backend's start-time internet download crash-looped (R-655) |
| immich | immich-machine-learning v3.0.3 → v3.2.2 | PostgreSQL (server unchanged) | proven | proven (done) | 51.4 (100) | immich-server / [Nest] 7 - 09/23/2026, 9:11:07 PM LO | moved b82b7c6 |
| n8n | n8n 2.40.5 → 2.41.1 | SQLite | proven | proven (done) | 22.7 (47) | n8n / Migrations in progress, please do NOT stop the process. | moved 04b63db |
| ghost | ghost 6.64.0 → 6.65.0 | SQLite | proven | proven (done) | 24.9 (43) | ghost / [2026-09-23 21:21:22] WARN Database state requires migration | moved e6afab9 (first watch had no load — re-run) |
| komga | komga 1.25.0 → 1.27.1 | embedded | proven | proven (done) | 60.0 (77) | komga / 2026-09-23T21:35:41.616+02:00 INFO 1 --- [ main] o.f.core | moved fc1becc with 768M (memory_tight at 512M) |
| opengist | opengist 1.13 → 1.15 | SQLite | proven | proven (done) | 77.8 (100) | — | moved 657c83f (pages under /-/, R-654) |
| wishlist | wishlist v0.66.0 → v0.67.1 | SQLite | proven | proven (done) | 35.9 (53) | wishlist / Prisma schema loaded from prisma/schema.prisma. | moved 1545c5f (after R-612's 512M) |
| navidrome | navidrome 0.64.0 → 0.64.1 | SQLite | proven | proven (done) | 6.3 (9) | navidrome / time="2026-09-23T19:31:04Z" level=info msg="goose: no migrations | moved 76cc1f5 (second ladder step) |
| emby | emby 4.11.0.1 → 4.11.0.3 | embedded | proven | proven (done) | 4.5 (8) | emby / Info SqliteUserRepository: Sqlite compiler options: ATOMIC_INTRINSICS | moved 8d573ad |
| gitea | gitea 1.27.0 → 1.27.3 | SQLite | inconclusive | inconclusive () | — | — | NOT moved — inconclusive on both: its web installer blocks a headless first account (R-624) |
12 steps across 11 apps published, each written by upgrade-test.py --write-ladder, each its own
commit, every push through the hook (the new gate judged each against the registry), CI green (jobs
922–937, matched on head_sha; two commits rode inside a later push and were judged there). Tried: 14
apps, 15 edges. night-2026-09-23/C-moves.txt lists the commits.
Listed, not moved:
- across a major (11): sparkyfitness v0.17.3→v1.7.2 (both images), gokapi v1.9.6→v2.2.4, claper 2.5→3.0, homepage v1.13.2→v2.4.0, paperless-ngx 2.20.15→3.2.1, mariadb →13.0 (romm, kimai, bookstack, nextcloud), nextcloud 34→35, plex 1.41→1.43.
- no front-door seed route (R-624): code-server, outline, rallly; sign-up closed by design: vaultwarden, zipline; installer: gitea (tried — inconclusive on both venues).
- no fixture tonight: bentopdf, glance, crafty-controller, wger (2.6→2.7), wanderer's meilisearch (v1.36→v1.54), uptime-kuma (2.4.0→2.5.5 — its first account exists only over socket.io).
- tried and failed: adventurelog (R-655, two causes measured).
What the night added to the bench: the box walk's fixtures run on the bench (upgrade_boxport.py); a
wishlist fixture; opengist, komga and nextcloud fixtures fixed; the memory watch's load no longer errors on
box-fixture apps (R-653) and samples the app's own memory (decision 22, R-652); files_may_change from a
bind-mount tree hash (nextcloud's is the first entry carrying it).
R-612 / R-613 — both fixed (catalog a5a729a): wishlist 512M — 128M kills the first-boot seed every
run, sometimes after the Role/Group rows (so the sign-up lie is timing-dependent; it did NOT reproduce on
9202 tonight); 512M completes it (peak 312M). uptime-kuma UPTIME_KUMA_DB_TYPE=sqlite — before, the box
read running over setup-database; after, the real server (/metrics 401) and the database in the
volume. apps/wishlist/r612-*, apps/uptime-kuma/r613-*.
A consequence of publishing tonight, seen on demo-hp: romm now has three steps (the backfilled
5.0.0→5.3.0, 5.3.0→5.3.1, mariadb 11.4→11.8). The box cannot climb one step at a time yet (part 5), so the
press on 9201 took BOTH new steps at once — each tested alone, never together. It ended done, front door
200, 0 OOM kills after (press9201/).
Part D — the chaos hour
Written BEFORE round 1 (committed with this section). Guest 9202, controller v0.267.0, drill catalog,
update.health_timeout: 90s (the drill setting — the product default is 5 min). Schedule drawn once by
night-2026-09-23/chaos_schedule.py from seed 20260923 (constraints in its docstring):
seed 20260923, drawn in 3 attempt(s)
| round | action | accident | at +s |
|---|---|---|---|
| 1 | install | none | 11 |
| 2 | good_update | box_backup | 46 |
| 3 | restore | controller_kill | 28 |
| 4 | fail_health | memory_hog | 43 |
| 5 | install | disk_fill | 28 |
| 6 | good_update | memory_hog | 46 |
| 7 | restore | power_cut | 56 |
| 8 | cutoff | box_backup | 23 |
| 9 | fail_health | controller_kill | 58 |
| 10 | remove | docker_restart | 10 |
| 11 | cutoff | two_updates | 51 |
| 12 | remove | none | 47 |
The app each round acts on (fixed by rule in chaos.py, not drawn):
| round | app | why this app |
|---|---|---|
| 1 install / 10 remove | actualbudget | a throwaway beside the six |
| 5 install / 12 remove | papra | a second throwaway |
| 2 good_update | docmost 0.95.0 → 0.96.0 | PostgreSQL, the BIG database (300 extra spaces) |
| 6 good_update | romm 5.3.0 → 5.3.1 | MariaDB |
| 3 restore | vikunja | volume-only (SQLite + files) |
| 7 restore | docmost | the big database, under a power cut |
| 4 fail_health | adventurelog v0.12.1 → v0.13.0, probe on a port it does not answer | the second PostgreSQL (PostGIS) |
| 9 fail_health | vikunja 2.5.0 → 2.6.0, wrong probe | volume-only, under a controller kill |
| 8 cutoff | navidrome 0.64.0 → 0.64.1, wrong probe, the undo copy's finished-marker removed during verifying |
volume-only + a drive folder |
| 11 cutoff + two_updates | nextcloud 34.0.1 → 34.0.4 (the file-leg), and romm's database-engine Update (11.4 → 11.8, a tested good step) pressed at +51 s | two updates at once. Changed before round 1: the first draft named adventurelog, whose v0.13.0 cannot pass health (Part C) |
Setup: the six at their FROM pins, each seeded through its front door (A); one whole-box backup on 9202 (the restore rounds need a copy); seed B on docmost, romm and vikunja after it.
The twelve rounds — the five things, and the data read back
Evidence per round: night-2026-09-23/chaos/round-NN.{json,log}, round-NN-run.txt; rounds 11–12 also
round-NN-controller.log. When the accident landed is stated, because in rounds 2–5 it did not land
where the schedule meant it to (below).
| # | action · app × accident | what the household saw (hu / en) | what the box did by itself | steady | events that fired | should have fired, did not | data read back |
|---|---|---|---|---|---|---|---|
| 1 | install · actualbudget × none | installed; badge none | deploy 202 → running | 26.6 s | app_deploy_started, app_deployed; app_start_failed (warning) 6 s in — its app cannot be named (see "instrument") | — | A ✓ |
| 2 | good update · docmost 0.95→0.96 (412 spaces) × whole-box backup | „Naprakész" / "Up to date" | safety-dump → pulling → copying (+76.6 s) → starting → verifying → done 108 s | 558 s (waited ~7 min for the catalog — instrument) | none | — | A ✓, B ✓ (paged re-read; the first read asked page 1 of 21) |
| 3 | restore · vikunja × controller kill (+28 s) | „Naprakész" / "Up to date" | restore 6 s, running; the kill landed AFTER it; controller back in 0.2 s | 31.5 s | none | — | A ✓, B ✓ |
| 4 | fail health · adventurelog × 1 GB hog | „A(z) adventurelog frissítése … nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el." / "The update of adventurelog … did not succeed. The box put back the previous version and its data automatically — nothing was lost." | verifying → undoing +110.9 s → undone +155.1 s (the hog had ended — instrument) | 605 s | health_change, app_update_undone (warning) ✓ | — | A ✓ |
| 5 | install · papra × disk to 3 GB free | installed | deploy → running in 27 s; the fill landed just after | 26.6 s | deploy pair | a disk warning: not measurable (the fill lasted ~20 s, the health job runs every 5 min) | A ✓ |
| 6 | good update · romm 5.3.0→5.3.1 × 1 GB hog (during verifying) | „Naprakész" / "Up to date" | … verifying → done 55.8 s | 61.7 s | none | — | A ✓ |
| 7 | restore · docmost × power cut (+56 s, after the 32 s restore) | „Naprakész" / "Up to date" | restore 32 s; cut; guest back, controller back 12 s, apps up | 103.5 s | controller_started | — | A ✓, B ✓ |
| 8 | cut-off undo copy · navidrome × whole-box backup (+23 s, during verifying) | „Megállítva — visszaállítás szükséges" + the hold sentence naming a restore / "Stopped — restore needed" + the same in English | verifying → undoing +94.9 → failed 95.4 s: the copy's marker missing → nothing poured back → HOLD | 101 s | app_update_held (error) ✓, then app_start_failed 11 s later | — (the second is extra: R-660) | held (A unreadable by design) → the household's restore brought it back, A ✓ |
| 9 | fail health · vikunja 2.5→2.6 × controller kill in verifying | „A(z) vikunja frissítése … nem sikerült. A doboz automatikusan visszaállította …" / "… put back the previous version and its data automatically …" | killed at +58 s → "update recovery: vikunja was interrupted in verifying … RESUMING the health wait" → undoing → undone 157.7 s — with the undo copy will hold 0 named volume(s) (R-658) |
163 s | controller_started, app_update_undone ✓ | — | A ✓, B ✓ (the old version happens to run on migrated data) |
| 10 | remove · actualbudget × docker restart (+10 s) | the remove press answered 502 | stop done; remove lost to the restart; app stays installed and stopped | cap 1500 s | controller_started; app_start_failed + health_change every 5 min | — | app stopped (A not readable); the household must press again |
| 11 | cut-off undo copy · nextcloud × a second Update (romm's engine step) at the same moment | hold as round 8 (both languages); romm „Naprakész" | nextcloud → failed/HOLD 112.6 s; romm ran at the same time → done (no single-flight, §3b Q4) | 143 s | app_update_held ✓, app_start_failed 13 s later | — | romm A ✓; nextcloud held → the named restore REFUSED → R-659, the stop rule |
| 12 | remove · papra × none | not installed | stop → remove 200 → verified | 40.9 s | app_removed; health_change (the box-wide job: a held app + a stopped one — true) | — | removed clean (0 containers, 0 volumes) |
The instrument, stated rather than hidden. (1) Rounds 2 and 4 waited ~7 min for the drill commit on
the wrong file (an installed app's stack dir holds the APPLIED compose), so their accidents fired before the
update; round 3's and 5's actions ended before theirs. From round 5 the accident is armed at the press, and
from round 6 the wait reads the badge's own catalog_images (seconds). (2) The event column first came from
the controller's debug ring, which overflowed in the long rounds and MISSED both app_update_held events; it
was rebuilt from the controller's full log (E8-events-*) for rounds 7–12. Rounds 1–6 keep the ring's view,
and round 1's app_start_failed cannot be tied to an app: round 7's power cut took the container's earlier
log. From round 11 each round saves the whole controller log. (3) My stop rule first fired on round 8 because
a held app cannot answer; a hold is by design, so recovery — not the readback — is the test.
The truth table (08-alarm-ladder.md)
| situation | the ladder says | fired? |
|---|---|---|
| an update undone (rounds 4, 9) | app_update_undone, warning, household + operator, one per app per outcome |
✓ both |
| an update held (rounds 8, 11) | app_update_held, error |
✓ both — and each was followed by an app_start_failed for the same moment (R-660: a held app is stopped by the product and should not ALSO read as down) |
| a controller kill / power cut / docker restart (3, 7, 9, 10) | controller_started, info; the boot grace suppresses app alarms for 90 s |
✓ controller_started each time; no app alarm inside the grace |
| a restore that finished (3, 7) | nothing (a refused or interrupted one: restore_interrupted) |
✓ nothing |
| a remove (12, teardown) | app_removed, info |
✓ |
| a disk filled to 1 GB above the floor (5) | a disk warning from the health job | not measurable — the fill lasted ~20 s |
| the box-wide health job with a held + a stopped app | health_change, warning |
✓ every 5 min |
demo-hp 9201 — the standing apps Part C moved, pressed through the product
night-2026-09-23/press9201/ (live catalog, a hub attached, controller v0.266.0; no whole-box backup pressed).
| app | from → to | phases | end | front door after | memory after |
|---|---|---|---|---|---|
| opengist | 1.13 → 1.15 | backing-up → … → done | done in ~16 s | /healthcheck 200 |
— |
| kimai | mariadb 11.6 → 11.8 (engine alone) | … → done | done in ~69 s | /en/login 200 |
— |
| romm | 5.3.0 → 5.3.1 and mariadb 11.4 → 11.8 in ONE press (no ladder on the box yet) | … → done | done in ~90 s | /api/heartbeat 200 |
600 MiB / 768, oom_kill 0, 0 restarts (checked 10 min later) |
Every standing app afterwards: 24 containers before, 24 after, all healthy (E7-9201-after.txt). Nothing
else on 9201 was touched.
Teardown
Machine — 9202: every app of the night removed THROUGH THE PRODUCT (docmost, navidrome, vikunja, romm,
nextcloud, adventurelog, actualbudget; papra in round 12): 0 containers, 0 volumes, 0 undo copies left
(E2-teardown-removes.*). For the three RESTORED apps the remove answered volumes_removed: [] while their
volumes did go (compose's down --volumes removes them by name) — the report is wrong, not the act (R-658).
The drive folders the product kept (R-442 on 9202, as every night) removed by name: navidrome, nextcloud, romm
(E3-*); older folders from earlier nights left as found. 54 test images removed by name before Part D (no
prune): /var/lib/docker 86 % → 24 %. controller.yaml restored and read back identical to
.pre-night0923; the catalog cache on 9202 reads live 5d49a11 (E4-*, E5-*). 9202 stays on controller
v0.267.0 (the scratch guest; self-update off).
Machine — the bench: LXC 9401 destroyed (pct destroy --purge, its hostname checked first); the Debian
template downloaded for it removed (E6-*).
Machine — 9201: only the three guarded Updates above; 24 containers before and after, all healthy.
Host — demo-hp: pct list = 9201, 9202 (as before); local 84.49 % before and after; nvme-scratch
7.45 % → 7.59 % and local-lvm 56.3 % → 57.7 % (the two guests' pulled images) (00-*, E6-*).
Drill repo: reset to live main 5d49a11b31fe, has_actions: false, read back (E5-*).
Hub: not touched. The floor was NOT raised — see the three lines.
Claims in the brief that turned out wrong, or held
| the brief's claim | what the night found |
|---|---|
the controller ignores an unknown .felhom.yml key (Part B spike) |
held — v0.266.0 and v0.267.0 synced, deployed, probed and badged navidrome with update_ladder: exactly as without it |
| R-633's window closes R-626 | held, as far as one measurement goes — 390 s of events + a restart + a reboot: nothing came back. Closed by measurement, not by a named creator |
| 21 backfill records exist | held — 21 of 21; one commit (nextcloud's engine move) cited none and its record was found by name |
| wishlist's fault is memory (read from R-612) | held, and sharpened — 128M kills the seed every run; whether the sign-up then fails depends on WHEN the kill lands (it did not fail on 9202 tonight) |
| R-518: the page promises „csak néhány másodpercre" | wrong — v0.243.0 had already changed it to „általában néhány perc"; tonight states the measured ≈ 8 min |
check-image-resolvable.py already resolves digests |
half — it asks whether a ref EXISTS and throws the digest away; image_digest.py was written |
| one chaos round ≈ one accident inside one action | not as drawn — actions of 6–30 s ended before accidents at +28–56 s; four rounds' accidents landed outside the action (stated per round) |
| the chaos schedule's round-11 second update (first draft: adventurelog) | changed BEFORE round 1 to romm's engine step — adventurelog v0.13.0 cannot pass health |