# DRILL — night shift 2026-09-23: DooPlex guarded from its own tests, the test record, the proven moves, the chaos hour Evidence: `audits/night-2026-09-23/` (`PROGRESS.md` is the step log). Architecture read first and named: `architecture/09-update-architecture.md` (§3 decisions 11–20, §6.1a, §6.4 parts 4–6, the verdict record), `07-backup-architecture.md` §6, `08-alarm-ladder.md`. **Method: endpoint-level** — every product act is the endpoint the UI invokes; pages are fetched as HTML in both languages; no browser. ## Not done, or changed - **The floor stays at 0.266.0.** The brief raises it only if every live proof of Parts A, B and D passed; Part D found two P1s (R-658, R-659). Both exist in 0.266.0 as well — v0.267.0 does not cause them and adds R-640's protection to the same restore path — so the recommendation is to raise it; that is the operator's one question this morning. - **Adventurelog v0.13.0 was not moved** (R-655): our own template's healthcheck points at a node path the new image moved, and the new backend downloads world data from the internet at every start and crash-loops when the download is cut short. - **The stop rule fired in round 11** (R-659): all twelve rounds had already run; nothing further was injected. - **Rounds 2–5 did not land their accidents inside the action** (the instrument, stated in Part D); round 9 is the clean mid-update kill, round 6 the mid-verify memory hog, rounds 8 and 11 the mid-verify backup and the concurrent second update. - **Decisions taken by CC unattended** (operator may reverse, `09` §3): 22 — the memory mark reads the app's own memory, not the file cache; 23 — the push-time gate asks the registry for moved refs only. - R-518's per-tier quiesce, R-612's silent-seed half and R-613's template sweep stay open, narrowed. ## Three lines - **Interventions:** none by the operator; 5 by me on my own instruments (a stale-template wait, a dropped event ring, a killed runner of my own, a broken local edit reverted before any commit, the stop rule on a held app) — each stated where it happened. No product state was set by hand. - **Apps moved tonight: 11 of 14 tried (12 steps of 15 edges)**, each with its test record; the 21 moves of 2026-09-22 backfilled. - **The result that matters most:** after a restore, an app's volumes lose the label the undo finds them by — so the next failed update is "undone" with **nothing** put back (R-658, P1); and a held file-leg app can be pointed at a restore that refuses it (R-659, P1). --- ## Part A — controller v0.267.0 (`80e6ad8c4772`, CI job 919 success) | row | what shipped | red-proof (seen failing, then restored) | evidence | |---|---|---|---| | **R-650** | `internal/dockerexec`: every docker exec (77 sites) goes through it; under `go test` a real docker is refused with an error naming the command, unless `FELHOM_TEST_REAL_DOCKER=1` or the binary is a stub under the temp dir. `TestR650_NoBareDockerExec` sweeps the tree. | guard disabled → the decoy's `docker ps -a` runs (`got `); one bare call put back → the sweep names `dbdump.go:140` | `A1-r650-*.txt` | | R-650 sweep | **8 `api` tests FAILED on the refusal** — they built a real `stacks.Manager` and ran `docker ps` on DooPlex; `web` ran 67× `docker compose version`, 19× `docker ps`, 19× `docker info`, 10× `docker inspect felhom-samba`, 2× `docker exec felhom-samba`; `stacks` ran **1× `docker-compose down`**, 11× `compose ps`; `backup` 46× `docker ps`; `appexport` 7×; `system` 1×. `api`/`stacks`/`web` now run under `RunWithStub` (TestMain); the other three pass on the refusal. | — | `A1-r650-sweep.txt` | | **R-640** | `appbackup.CheckDumpComplete` (the engine's end marker, last 4 KiB, gzip-aware). The unit restore and the off-site restore refuse a cut-off copy **before the first mutation**; every replay checks again before any load, whatever the import seam is. | replay check removed → `a cut-off copy was LOADED`; unit gate removed → `stopped=true calls=[stop recreate startsvc:immich-postgres start]`; off-site gate removed → only the replay check caught it, after a stop and a rollback | `A2-r640-redproof.txt` | | **R-626** | **measured, not reproduced** on v0.266.0: navidrome deployed and removed through the product; `docker events` 390 s — the remove's kill/stop/die/destroy seen (the positive observable), **no create**; controller restarted at +150 s; guest rebooted after → 0 containers, 0 volumes at every check. My script first said `came_back: true` — it had counted the catalog-mirror directory every catalog app has; corrected in the verdict file, both values kept. Leftover found: `applied-compose.yml` + `applied-meta/` stay after a remove (R-651). | — | `A3-r626-*` | | **R-499** | the Tier-2 page's system-disk sentence has four branches (own drive / same disk / drive gone / cannot ask); „(PBS)" and „nincs külön teendő" only where true. Live on 9202 (no agent): the `unknown` branch in hu + en, matching `/api/storage/backup-target` `known:false`. | handler not passing the fact → 3 branches missing; the old sentence in the same-disk branch → `promises the backup protects this app` | `A4-*`, `A7-*` | | **R-518** | the backup button states the measured stop (≈ 8 min on a 12-app box), both languages, page + confirm. **The brief's „csak néhány másodpercre" was already gone** (v0.243.0); the vague „általában néhány perc" is replaced. | bundles back to v0.243.0 → measured downtime appears 0 times | `A5-*` | Gates: `go build/vet` rc 0; `go test -count=1 ./...` rc 0 (31 packages); `controller_gates.py` all OK. Deployed to **9202 only** (`A6-*`); the floor waits for Part E. ## Part B — the test record and its gate (catalog `6db08a5`, CI job 920 success) **Spike first (15 min):** a drill navidrome carrying a top-level `update_ladder:` (one JSON entry per line) on 9202 at **v0.266.0 and v0.267.0**: synced, deploy-fields read, deployed `running`, probe `API GET :4533/ping → 200`, badge „Naprakész" / "Up to date", no YAML error. Source agrees (no `KnownFields` anywhere). The brief's claim holds. `B1-spike-*`. **The record** (`scripts/ladder.py`): `from`/`to` per service, `digest` per `to` ref (sha256 from `scripts/image_digest.py` — stdlib, equals Docker's `RepoDigests` for `privatebin/pdo:2.0.6` on 9202), `verdict` (proven | unrecorded), `tested_at`, `harness_version`, `evidence`, `box_evidence`, `memory_peak_pct` (+ `memory_basis`, `memory_cgroup_peak_pct`), `marks` {files_may_change, needs_person, memory_tight}. Written ONLY by `upgrade-test.py --write-ladder` (bench AND box `proven`, template at FROM, digests resolved) — `test_ladder_writer.py` 5/5. **The gate** — two rows of `catalog_gates.py`, both `--fast`: `check-test-record.py` (static; runs in CI too: a ladder well-formed, continuous, and its newest step IS the compose) and `check-test-record-move.py` (history + the registry for MOVED refs only: a move adds a PROVEN, non-backfilled entry whose `from`/`to` are the compose before/after and whose digests the registry still serves; `memory_tight` needs the limit raised in the same range). 16 decoy cases, both directions. **Red-proofs:** the no-entry refusal removed → *a bare image move* and *the entry only in README* both pass (rc 0, expected 1); `failed` allowed → both failed-verdict cases pass; the digest comparison removed → *the registry now serves another digest* passes. `B2-gate-redproofs.txt`. Live, on the first real move (romm): the gate judged it against the real registry (rc 0), and with a wrong digest table it REFUSED all three services. **Backfill** — the 21 moves of 2026-09-22, each from the record its commit cited: | app | step (the service that moved) | verdict | evidence | note | |---|---|---|---|---| | actualbudget | actualbudget: actual-server:26.7.0 → actual-server:26.9.0 | proven | `audits/update-night-2026-09-21/apps/actualbudget/verdict.json` | — | | audiobookshelf | audiobookshelf: audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | proven | `audits/update-night-2026-09-21/apps/audiobookshelf/verdict.json` | — | | bookstack | bookstack: bookstack:26.05.2 → bookstack:26.05.5 | proven | `audits/update-night-2026-09-21/apps/bookstack/verdict.json` | — | | docmost | docmost: docmost:0.95.0 → docmost:0.96.0 | proven | `audits/update-night-2026-09-21/apps/docmost/verdict.json` | — | | emby | emby: embyserver:4.10.0.20 → embyserver:4.11.0.1 | proven | `audits/the-28-2026-09-22/apps/emby/verdict.json` | — | | ghost | ghost: ghost:6.53.0-alpine → ghost:6.64.0-alpine | proven | `audits/the-28-2026-09-22/apps/ghost/verdict.json` | — | | grafana | grafana: grafana:13.1.0 → grafana:13.2.2 | proven | `audits/update-night-2026-09-21/apps/grafana/verdict.json` | — | | home-assistant | home-assistant: home-assistant:2026.7.2 → home-assistant:2026.9.3 | proven | `audits/update-night-2026-09-21/apps/home-assistant/verdict.json` | — | | immich | immich-server: immich-server:v3.0.3 → immich-server:v3.2.2 | proven | `audits/the-28-2026-09-22/apps/immich/verdict.json` | — | | mealie | mealie: mealie:v3.20.1 → mealie:v3.27.0 | proven | `audits/update-night-2026-09-21/apps/mealie/verdict.json` | — | | n8n | n8n: n8n:2.31.3 → n8n:2.40.5 | proven | `audits/update-night-2026-09-21/apps/n8n/verdict.json` | — | | navidrome | navidrome: navidrome:0.63.2 → navidrome:0.64.0 | proven | `audits/update-night-2026-09-21/apps/navidrome/verdict.json` | — | | nextcloud | nextcloud-db: mariadb:11.6 → mariadb:12.3 | proven | `audits/update-night-2026-09-21/apps/nextcloud-engine-mariadb/verdict.json` | commit cited no record; this one named | | papra | papra: papra:26.6.1-rootless → papra:26.6.2-rootless | proven | `audits/update-night-2026-09-21/apps/papra/verdict.json` | — | | privatebin | privatebin: pdo:2.0.5 → pdo:2.0.6 | proven | `audits/update-night-2026-09-21/apps/privatebin/verdict.json` | — | | radarr | radarr: radarr:6.3.0 → radarr:6.4.4 | proven | `audits/the-28-2026-09-22/apps/radarr/verdict.json` | — | | romm | romm: romm:5.0.0 → romm:5.3.0 | proven | `audits/update-night-2026-09-21/apps/romm/verdict.json` | memory watch 80.9 % (M1) | | sonarr | sonarr: sonarr:4.0.19 → sonarr:4.0.20 | proven | `audits/the-28-2026-09-22/apps/sonarr/verdict.json` | — | | tandoor | tandoor: recipes:2.6.13 → recipes:2.6.15 | proven | `audits/update-night-2026-09-21/apps/tandoor/verdict.json` | — | | termix | termix: termix:2.5.0 → termix:2.8.0 | proven | `audits/the-28-2026-09-22/apps/termix/verdict.json` | — | | vikunja | vikunja: vikunja:2.3.0 → vikunja:2.6.0 | proven | `audits/update-night-2026-09-21/apps/vikunja/verdict.json` | — | **21 of 21 `proven`, 0 `unrecorded`.** Every entry is marked `backfilled`, carries harness v1 (box walk only, no memory watch — except romm's M1 figure), and a digest resolved TONIGHT, which each entry's `note` says. `scripts/ladder_backfill.py`. ## Part C — move every app that proves itself **Venues:** the bench = throwaway LXC **9401** on demo-hp (Debian 13, docker 26.1.5, Docker Hub login), created tonight and destroyed at teardown; the box = guest **9202**, v0.267.0, drill catalog. **Negative control first:** C3 (`alpine:3.20` as the TO image) → `failed` — the bench measures (`bench-C3/`). Box walks by `run_edge.py`; bench edges by `upgrade-test.py --move`; each edge's bench evidence copied off the bench when it ended (`benchq.py`). Evidence per app: `night-2026-09-23/apps//`. | app | from → to | class | bench verdict | box verdict | memory: own (cgroup) peak % | migration line | moved? | |---|---|---|---|---|---|---|---| | nextcloud | nextcloud 34.0.1 → 34.0.4 | MariaDB + file-leg | proven | proven (done) | 24.2 (100) | nextcloud / Upgrading nextcloud from 34.0.1.2 ... | moved `5d49a11` (files_may_change) | | romm | romm 5.3.0 → 5.3.1 | MariaDB | proven | proven (done) | n/a — measured before the own-memory sample existed (cgroup 79.5) | romm / INFO: [RomM][init][2026-09-23 21:12:26][0 | moved `7ff4e68` | | romm-engine | romm-db mariadb 11.4 → 11.8 (own edge) | MariaDB engine | proven | proven (done) | 78.0 (84) | romm / INFO: [RomM][init][2026-09-23 21:47:51][0 | moved `431cdec` | | kimai | kimai-db mariadb 11.6 → 11.8 (own edge) | MariaDB engine | proven | proven (done) | 37.9 (100) | kimai / [OK] Already at the latest version ("DoctrineMigrations\Version2 | moved `376b2c3` | | adventurelog | adventurelog v0.12.1 → v0.13.0 (both images) | PostgreSQL (PostGIS) | failed ×2 (frontend never healthy; then backend crash-loop on `download-countries`) | failed (undone at +196 s, drill timeout 90 s) | — | adventurelog / Apply all migrations: account, admin, adventures, aut | **NOT moved** — our template's frontend healthcheck points at `/nodejs/bin/node`, gone in v0.13.0; with it dropped, the backend's start-time internet download crash-looped (R-655) | | immich | immich-machine-learning v3.0.3 → v3.2.2 | PostgreSQL (server unchanged) | proven | proven (done) | 51.4 (100) | immich-server / [Nest] 7 - 09/23/2026, 9:11:07 PM LO | moved `b82b7c6` | | n8n | n8n 2.40.5 → 2.41.1 | SQLite | proven | proven (done) | 22.7 (47) | n8n / Migrations in progress, please do NOT stop the process. | moved `04b63db` | | ghost | ghost 6.64.0 → 6.65.0 | SQLite | proven | proven (done) | 24.9 (43) | ghost / [2026-09-23 21:21:22] WARN Database state requires migration | moved `e6afab9` (first watch had no load — re-run) | | komga | komga 1.25.0 → 1.27.1 | embedded | proven | proven (done) | 60.0 (77) | komga / 2026-09-23T21:35:41.616+02:00 INFO 1 --- [ main] o.f.core | moved `fc1becc` with 768M (memory_tight at 512M) | | opengist | opengist 1.13 → 1.15 | SQLite | proven | proven (done) | 77.8 (100) | — | moved `657c83f` (pages under `/-/`, R-654) | | wishlist | wishlist v0.66.0 → v0.67.1 | SQLite | proven | proven (done) | 35.9 (53) | wishlist / Prisma schema loaded from prisma/schema.prisma. | moved `1545c5f` (after R-612's 512M) | | navidrome | navidrome 0.64.0 → 0.64.1 | SQLite | proven | proven (done) | 6.3 (9) | navidrome / time="2026-09-23T19:31:04Z" level=info msg="goose: no migrations | moved `76cc1f5` (second ladder step) | | emby | emby 4.11.0.1 → 4.11.0.3 | embedded | proven | proven (done) | 4.5 (8) | emby / Info SqliteUserRepository: Sqlite compiler options: ATOMIC_INTRINSICS | moved `8d573ad` | | gitea | gitea 1.27.0 → 1.27.3 | SQLite | inconclusive | inconclusive () | — | — | NOT moved — inconclusive on both: its web installer blocks a headless first account (R-624) | **12 steps across 11 apps published**, each written by `upgrade-test.py --write-ladder`, each its own commit, every push through the hook (the new gate judged each against the registry), CI green (jobs 922–937, matched on head_sha; two commits rode inside a later push and were judged there). Tried: 14 apps, 15 edges. `night-2026-09-23/C-moves.txt` lists the commits. **Listed, not moved:** - *across a major* (11): sparkyfitness v0.17.3→v1.7.2 (both images), gokapi v1.9.6→v2.2.4, claper 2.5→3.0, homepage v1.13.2→v2.4.0, paperless-ngx 2.20.15→3.2.1, mariadb →13.0 (romm, kimai, bookstack, nextcloud), nextcloud 34→35, plex 1.41→1.43. - *no front-door seed route* (R-624): code-server, outline, rallly; *sign-up closed by design*: vaultwarden, zipline; *installer*: gitea (tried — inconclusive on both venues). - *no fixture tonight*: bentopdf, glance, crafty-controller, wger (2.6→2.7), wanderer's meilisearch (v1.36→v1.54), uptime-kuma (2.4.0→2.5.5 — its first account exists only over socket.io). - *tried and failed*: adventurelog (R-655, two causes measured). **What the night added to the bench:** the box walk's fixtures run on the bench (`upgrade_boxport.py`); a wishlist fixture; opengist, komga and nextcloud fixtures fixed; the memory watch's load no longer errors on box-fixture apps (R-653) and samples the app's own memory (decision 22, R-652); `files_may_change` from a bind-mount tree hash (nextcloud's is the first entry carrying it). **R-612 / R-613 — both fixed** (catalog `a5a729a`): wishlist 512M — 128M kills the first-boot seed every run, sometimes after the Role/Group rows (so the sign-up lie is timing-dependent; it did NOT reproduce on 9202 tonight); 512M completes it (peak 312M). uptime-kuma `UPTIME_KUMA_DB_TYPE=sqlite` — before, the box read `running` over `setup-database`; after, the real server (`/metrics` 401) and the database in the volume. `apps/wishlist/r612-*`, `apps/uptime-kuma/r613-*`. **A consequence of publishing tonight, seen on demo-hp:** romm now has three steps (the backfilled 5.0.0→5.3.0, 5.3.0→5.3.1, mariadb 11.4→11.8). The box cannot climb one step at a time yet (part 5), so the press on 9201 took BOTH new steps at once — each tested alone, never together. It ended `done`, front door 200, 0 OOM kills after (`press9201/`). ## Part D — the chaos hour **Written BEFORE round 1 (committed with this section).** Guest 9202, controller v0.267.0, drill catalog, `update.health_timeout: 90s` (the drill setting — the product default is 5 min). Schedule drawn once by `night-2026-09-23/chaos_schedule.py` from seed **20260923** (constraints in its docstring): seed 20260923, drawn in 3 attempt(s) | round | action | accident | at +s | |---|---|---|---| | 1 | install | none | 11 | | 2 | good_update | box_backup | 46 | | 3 | restore | controller_kill | 28 | | 4 | fail_health | memory_hog | 43 | | 5 | install | disk_fill | 28 | | 6 | good_update | memory_hog | 46 | | 7 | restore | power_cut | 56 | | 8 | cutoff | box_backup | 23 | | 9 | fail_health | controller_kill | 58 | | 10 | remove | docker_restart | 10 | | 11 | cutoff | two_updates | 51 | | 12 | remove | none | 47 | **The app each round acts on** (fixed by rule in `chaos.py`, not drawn): | round | app | why this app | |---|---|---| | 1 install / 10 remove | actualbudget | a throwaway beside the six | | 5 install / 12 remove | papra | a second throwaway | | 2 good_update | docmost 0.95.0 → 0.96.0 | PostgreSQL, the BIG database (300 extra spaces) | | 6 good_update | romm 5.3.0 → 5.3.1 | MariaDB | | 3 restore | vikunja | volume-only (SQLite + files) | | 7 restore | docmost | the big database, under a power cut | | 4 fail_health | adventurelog v0.12.1 → v0.13.0, probe on a port it does not answer | the second PostgreSQL (PostGIS) | | 9 fail_health | vikunja 2.5.0 → 2.6.0, wrong probe | volume-only, under a controller kill | | 8 cutoff | navidrome 0.64.0 → 0.64.1, wrong probe, the undo copy's finished-marker removed during `verifying` | volume-only + a drive folder | | 11 cutoff + two_updates | nextcloud 34.0.1 → 34.0.4 (the file-leg), and romm's database-engine Update (11.4 → 11.8, a tested good step) pressed at +51 s | two updates at once. *Changed before round 1:* the first draft named adventurelog, whose v0.13.0 cannot pass health (Part C) | Setup: the six at their FROM pins, each seeded through its front door (A); one whole-box backup on 9202 (the restore rounds need a copy); seed B on docmost, romm and vikunja after it. ### The twelve rounds — the five things, and the data read back Evidence per round: `night-2026-09-23/chaos/round-NN.{json,log}`, `round-NN-run.txt`; rounds 11–12 also `round-NN-controller.log`. **When the accident landed** is stated, because in rounds 2–5 it did not land where the schedule meant it to (below). | # | action · app × accident | what the household saw (hu / en) | what the box did by itself | steady | events that fired | should have fired, did not | data read back | |---|---|---|---|---|---|---|---| | 1 | install · actualbudget × none | installed; badge none | deploy 202 → running | 26.6 s | app_deploy_started, app_deployed; **app_start_failed (warning)** 6 s in — its app cannot be named (see "instrument") | — | A ✓ | | 2 | good update · docmost 0.95→0.96 (412 spaces) × whole-box backup | „Naprakész" / "Up to date" | safety-dump → pulling → copying (+76.6 s) → starting → verifying → **done 108 s** | 558 s (waited ~7 min for the catalog — instrument) | none | — | A ✓, B ✓ (paged re-read; the first read asked page 1 of 21) | | 3 | restore · vikunja × controller kill (+28 s) | „Naprakész" / "Up to date" | restore 6 s, running; the kill landed AFTER it; controller back in 0.2 s | 31.5 s | none | — | A ✓, B ✓ | | 4 | fail health · adventurelog × 1 GB hog | „A(z) adventurelog frissítése … nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el." / "The update of adventurelog … did not succeed. The box put back the previous version and its data automatically — nothing was lost." | verifying → **undoing +110.9 s → undone +155.1 s** (the hog had ended — instrument) | 605 s | health_change, **app_update_undone (warning)** ✓ | — | A ✓ | | 5 | install · papra × disk to 3 GB free | installed | deploy → running in 27 s; the fill landed just after | 26.6 s | deploy pair | a disk warning: not measurable (the fill lasted ~20 s, the health job runs every 5 min) | A ✓ | | 6 | good update · romm 5.3.0→5.3.1 × 1 GB hog (**during verifying**) | „Naprakész" / "Up to date" | … verifying → **done 55.8 s** | 61.7 s | none | — | A ✓ | | 7 | restore · docmost × **power cut** (+56 s, after the 32 s restore) | „Naprakész" / "Up to date" | restore 32 s; cut; guest back, controller back 12 s, apps up | 103.5 s | controller_started | — | A ✓, B ✓ | | 8 | cut-off undo copy · navidrome × whole-box backup (+23 s, during verifying) | „Megállítva — visszaállítás szükséges" + the hold sentence naming a restore / "Stopped — restore needed" + the same in English | verifying → undoing +94.9 → **failed 95.4 s**: the copy's marker missing → nothing poured back → HOLD | 101 s | **app_update_held (error)** ✓, then **app_start_failed** 11 s later | — (the second is extra: R-660) | held (A unreadable by design) → **the household's restore brought it back, A ✓** | | 9 | fail health · vikunja 2.5→2.6 × **controller kill in verifying** | „A(z) vikunja frissítése … nem sikerült. A doboz automatikusan visszaállította …" / "… put back the previous version and its data automatically …" | killed at +58 s → *"update recovery: vikunja was interrupted in verifying … RESUMING the health wait"* → undoing → **undone 157.7 s** — with **`the undo copy will hold 0 named volume(s)`** (R-658) | 163 s | controller_started, **app_update_undone** ✓ | — | A ✓, B ✓ (the old version happens to run on migrated data) | | 10 | remove · actualbudget × docker restart (+10 s) | the remove press answered **502** | stop done; remove lost to the restart; app stays installed and stopped | cap 1500 s | controller_started; app_start_failed + health_change every 5 min | — | app stopped (A not readable); the household must press again | | 11 | cut-off undo copy · nextcloud × **a second Update (romm's engine step) at the same moment** | hold as round 8 (both languages); romm „Naprakész" | nextcloud → **failed/HOLD 112.6 s**; romm ran at the same time → **done** (no single-flight, §3b Q4) | 143 s | **app_update_held** ✓, app_start_failed 13 s later | — | romm A ✓; nextcloud held → **the named restore REFUSED → R-659, the stop rule** | | 12 | remove · papra × none | not installed | stop → remove 200 → verified | 40.9 s | app_removed; health_change (the box-wide job: a held app + a stopped one — true) | — | removed clean (0 containers, 0 volumes) | **The instrument, stated rather than hidden.** (1) Rounds 2 and 4 waited ~7 min for the drill commit on the wrong file (an installed app's stack dir holds the APPLIED compose), so their accidents fired before the update; round 3's and 5's actions ended before theirs. From round 5 the accident is armed at the press, and from round 6 the wait reads the badge's own `catalog_images` (seconds). (2) The event column first came from the controller's debug ring, which overflowed in the long rounds and MISSED both `app_update_held` events; it was rebuilt from the controller's full log (`E8-events-*`) for rounds 7–12. Rounds 1–6 keep the ring's view, and round 1's `app_start_failed` cannot be tied to an app: round 7's power cut took the container's earlier log. From round 11 each round saves the whole controller log. (3) My stop rule first fired on round 8 because a held app cannot answer; a hold is by design, so recovery — not the readback — is the test. ### The truth table (`08-alarm-ladder.md`) | situation | the ladder says | fired? | |---|---|---| | an update undone (rounds 4, 9) | `app_update_undone`, warning, household + operator, one per app per outcome | ✓ both | | an update held (rounds 8, 11) | `app_update_held`, error | ✓ both — and each was followed by an `app_start_failed` for the same moment (R-660: a held app is stopped by the product and should not ALSO read as down) | | a controller kill / power cut / docker restart (3, 7, 9, 10) | `controller_started`, info; the boot grace suppresses app alarms for 90 s | ✓ controller_started each time; no app alarm inside the grace | | a restore that finished (3, 7) | nothing (a refused or interrupted one: `restore_interrupted`) | ✓ nothing | | a remove (12, teardown) | `app_removed`, info | ✓ | | a disk filled to 1 GB above the floor (5) | a disk warning from the health job | not measurable — the fill lasted ~20 s | | the box-wide health job with a held + a stopped app | `health_change`, warning | ✓ every 5 min | ### demo-hp 9201 — the standing apps Part C moved, pressed through the product `night-2026-09-23/press9201/` (live catalog, a hub attached, controller v0.266.0; no whole-box backup pressed). | app | from → to | phases | end | front door after | memory after | |---|---|---|---|---|---| | opengist | 1.13 → 1.15 | backing-up → … → done | **done in ~16 s** | `/healthcheck` 200 | — | | kimai | mariadb 11.6 → 11.8 (engine alone) | … → done | **done in ~69 s** | `/en/login` 200 | — | | romm | 5.3.0 → 5.3.1 **and** mariadb 11.4 → 11.8 in ONE press (no ladder on the box yet) | … → done | **done in ~90 s** | `/api/heartbeat` 200 | 600 MiB / 768, `oom_kill` 0, 0 restarts (checked 10 min later) | Every standing app afterwards: 24 containers before, 24 after, all healthy (`E7-9201-after.txt`). Nothing else on 9201 was touched. ## Teardown **Machine — 9202:** every app of the night removed THROUGH THE PRODUCT (docmost, navidrome, vikunja, romm, nextcloud, adventurelog, actualbudget; papra in round 12): 0 containers, 0 volumes, 0 undo copies left (`E2-teardown-removes.*`). For the three RESTORED apps the remove answered `volumes_removed: []` while their volumes did go (compose's `down --volumes` removes them by name) — the report is wrong, not the act (R-658). The drive folders the product kept (R-442 on 9202, as every night) removed by name: navidrome, nextcloud, romm (`E3-*`); older folders from earlier nights left as found. 54 test images removed by name before Part D (no prune): `/var/lib/docker` 86 % → 24 %. `controller.yaml` restored and read back **identical** to `.pre-night0923`; the catalog cache on 9202 reads live `5d49a11` (`E4-*`, `E5-*`). **9202 stays on controller v0.267.0** (the scratch guest; self-update off). **Machine — the bench:** LXC 9401 destroyed (`pct destroy --purge`, its hostname checked first); the Debian template downloaded for it removed (`E6-*`). **Machine — 9201:** only the three guarded Updates above; 24 containers before and after, all healthy. **Host — demo-hp:** `pct list` = 9201, 9202 (as before); `local` 84.49 % before and after; `nvme-scratch` 7.45 % → 7.59 % and `local-lvm` 56.3 % → 57.7 % (the two guests' pulled images) (`00-*`, `E6-*`). **Drill repo:** reset to live `main` `5d49a11b31fe`, `has_actions: false`, read back (`E5-*`). **Hub:** not touched. **The floor was NOT raised** — see the three lines. ## Claims in the brief that turned out wrong, or held | the brief's claim | what the night found | |---|---| | the controller ignores an unknown `.felhom.yml` key (Part B spike) | **held** — v0.266.0 and v0.267.0 synced, deployed, probed and badged navidrome with `update_ladder:` exactly as without it | | R-633's window closes R-626 | **held, as far as one measurement goes** — 390 s of events + a restart + a reboot: nothing came back. Closed by measurement, not by a named creator | | 21 backfill records exist | **held** — 21 of 21; one commit (nextcloud's engine move) cited none and its record was found by name | | wishlist's fault is memory (read from R-612) | **held, and sharpened** — 128M kills the seed every run; whether the sign-up then fails depends on WHEN the kill lands (it did not fail on 9202 tonight) | | R-518: the page promises „csak néhány másodpercre" | **wrong** — v0.243.0 had already changed it to „általában néhány perc"; tonight states the measured ≈ 8 min | | `check-image-resolvable.py` already resolves digests | **half** — it asks whether a ref EXISTS and throws the digest away; `image_digest.py` was written | | one chaos round ≈ one accident inside one action | **not as drawn** — actions of 6–30 s ended before accidents at +28–56 s; four rounds' accidents landed outside the action (stated per round) | | the chaos schedule's round-11 second update (first draft: adventurelog) | changed BEFORE round 1 to romm's engine step — adventurelog v0.13.0 cannot pass health |