Files
felhom.eu/documentation/audits/DRILL-night-2026-09-23.md
T
admin 3e58c184f6
gates / gates (push) Successful in 27s
night shift 2026-09-23: the record, the register, the morning note
DRILL-night-2026-09-23.md: Parts A-E. 09 §3 decisions 21 (operator word),
22 and 23 (CC unattended, operator may reverse); §6.4 parts 4 and 6
(catalog half) shipped; §6.1a residuals R-658/R-659. Register 330 -> 336:
R-651..R-660 opened (R-658 and R-659 P1), R-650/R-640/R-499/R-626 closed.
Capability map, nightly rotation (opengist), STATUS (one question: the
floor), CONTEXT, REPORT. The floor stays 0.266.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 00:24:04 +02:00

290 lines
29 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DRILL — night shift 2026-09-23: DooPlex guarded from its own tests, the test record, the proven moves, the chaos hour
Evidence: `audits/night-2026-09-23/` (`PROGRESS.md` is the step log). Architecture read first and named:
`architecture/09-update-architecture.md` (§3 decisions 11–20, §6.1a, §6.4 parts 4–6, the verdict record),
`07-backup-architecture.md` §6, `08-alarm-ladder.md`. **Method: endpoint-level** — every product act is
the endpoint the UI invokes; pages are fetched as HTML in both languages; no browser.
## Not done, or changed
- **The floor stays at 0.266.0.** The brief raises it only if every live proof of Parts A, B and D passed; Part D
found two P1s (R-658, R-659). Both exist in 0.266.0 as well — v0.267.0 does not cause them and adds R-640's
protection to the same restore path — so the recommendation is to raise it; that is the operator's one
question this morning.
- **Adventurelog v0.13.0 was not moved** (R-655): our own template's healthcheck points at a node path the new
image moved, and the new backend downloads world data from the internet at every start and crash-loops when
the download is cut short.
- **The stop rule fired in round 11** (R-659): all twelve rounds had already run; nothing further was injected.
- **Rounds 2–5 did not land their accidents inside the action** (the instrument, stated in Part D); round 9 is
the clean mid-update kill, round 6 the mid-verify memory hog, rounds 8 and 11 the mid-verify backup and the
concurrent second update.
- **Decisions taken by CC unattended** (operator may reverse, `09` §3): 22 — the memory mark reads the app's own
memory, not the file cache; 23 — the push-time gate asks the registry for moved refs only.
- R-518's per-tier quiesce, R-612's silent-seed half and R-613's template sweep stay open, narrowed.
## Three lines
- **Interventions:** none by the operator; 5 by me on my own instruments (a stale-template wait, a dropped
event ring, a killed runner of my own, a broken local edit reverted before any commit, the stop rule on a held
app) — each stated where it happened. No product state was set by hand.
- **Apps moved tonight: 11 of 14 tried (12 steps of 15 edges)**, each with its test record; the 21 moves of
2026-09-22 backfilled.
- **The result that matters most:** after a restore, an app's volumes lose the label the undo finds them by —
so the next failed update is "undone" with **nothing** put back (R-658, P1); and a held file-leg app can be
pointed at a restore that refuses it (R-659, P1).
---
## Part A — controller v0.267.0 (`80e6ad8c4772`, CI job 919 success)
| row | what shipped | red-proof (seen failing, then restored) | evidence |
|---|---|---|---|
| **R-650** | `internal/dockerexec`: every docker exec (77 sites) goes through it; under `go test` a real docker is refused with an error naming the command, unless `FELHOM_TEST_REAL_DOCKER=1` or the binary is a stub under the temp dir. `TestR650_NoBareDockerExec` sweeps the tree. | guard disabled → the decoy's `docker ps -a` runs (`got <nil>`); one bare call put back → the sweep names `dbdump.go:140` | `A1-r650-*.txt` |
| R-650 sweep | **8 `api` tests FAILED on the refusal** — they built a real `stacks.Manager` and ran `docker ps` on DooPlex; `web` ran 67× `docker compose version`, 19× `docker ps`, 19× `docker info`, 10× `docker inspect felhom-samba`, 2× `docker exec felhom-samba`; `stacks` ran **1× `docker-compose down`**, 11× `compose ps`; `backup` 46× `docker ps`; `appexport` 7×; `system` 1×. `api`/`stacks`/`web` now run under `RunWithStub` (TestMain); the other three pass on the refusal. | — | `A1-r650-sweep.txt` |
| **R-640** | `appbackup.CheckDumpComplete` (the engine's end marker, last 4 KiB, gzip-aware). The unit restore and the off-site restore refuse a cut-off copy **before the first mutation**; every replay checks again before any load, whatever the import seam is. | replay check removed → `a cut-off copy was LOADED`; unit gate removed → `stopped=true calls=[stop recreate startsvc:immich-postgres start]`; off-site gate removed → only the replay check caught it, after a stop and a rollback | `A2-r640-redproof.txt` |
| **R-626** | **measured, not reproduced** on v0.266.0: navidrome deployed and removed through the product; `docker events` 390 s — the remove's kill/stop/die/destroy seen (the positive observable), **no create**; controller restarted at +150 s; guest rebooted after → 0 containers, 0 volumes at every check. My script first said `came_back: true` — it had counted the catalog-mirror directory every catalog app has; corrected in the verdict file, both values kept. Leftover found: `applied-compose.yml` + `applied-meta/` stay after a remove (R-651). | — | `A3-r626-*` |
| **R-499** | the Tier-2 page's system-disk sentence has four branches (own drive / same disk / drive gone / cannot ask); „(PBS)" and „nincs külön teendő" only where true. Live on 9202 (no agent): the `unknown` branch in hu + en, matching `/api/storage/backup-target` `known:false`. | handler not passing the fact → 3 branches missing; the old sentence in the same-disk branch → `promises the backup protects this app` | `A4-*`, `A7-*` |
| **R-518** | the backup button states the measured stop (≈ 8 min on a 12-app box), both languages, page + confirm. **The brief's „csak néhány másodpercre" was already gone** (v0.243.0); the vague „általában néhány perc" is replaced. | bundles back to v0.243.0 → measured downtime appears 0 times | `A5-*` |
Gates: `go build/vet` rc 0; `go test -count=1 ./...` rc 0 (31 packages); `controller_gates.py` all OK.
Deployed to **9202 only** (`A6-*`); the floor waits for Part E.
## Part B — the test record and its gate (catalog `6db08a5`, CI job 920 success)
**Spike first (15 min):** a drill navidrome carrying a top-level `update_ladder:` (one JSON entry per
line) on 9202 at **v0.266.0 and v0.267.0**: synced, deploy-fields read, deployed `running`, probe
`API GET :4533/ping → 200`, badge „Naprakész" / "Up to date", no YAML error. Source agrees (no
`KnownFields` anywhere). The brief's claim holds. `B1-spike-*`.
**The record** (`scripts/ladder.py`): `from`/`to` per service, `digest` per `to` ref (sha256 from
`scripts/image_digest.py` — stdlib, equals Docker's `RepoDigests` for `privatebin/pdo:2.0.6` on 9202),
`verdict` (proven | unrecorded), `tested_at`, `harness_version`, `evidence`, `box_evidence`,
`memory_peak_pct` (+ `memory_basis`, `memory_cgroup_peak_pct`), `marks` {files_may_change, needs_person,
memory_tight}. Written ONLY by `upgrade-test.py --write-ladder` (bench AND box `proven`, template at
FROM, digests resolved) — `test_ladder_writer.py` 5/5.
**The gate** — two rows of `catalog_gates.py`, both `--fast`:
`check-test-record.py` (static; runs in CI too: a ladder well-formed, continuous, and its newest step IS
the compose) and `check-test-record-move.py` (history + the registry for MOVED refs only: a move adds a
PROVEN, non-backfilled entry whose `from`/`to` are the compose before/after and whose digests the
registry still serves; `memory_tight` needs the limit raised in the same range). 16 decoy cases, both
directions. **Red-proofs:** the no-entry refusal removed → *a bare image move* and *the entry only in
README* both pass (rc 0, expected 1); `failed` allowed → both failed-verdict cases pass; the digest
comparison removed → *the registry now serves another digest* passes. `B2-gate-redproofs.txt`.
Live, on the first real move (romm): the gate judged it against the real registry (rc 0), and with a
wrong digest table it REFUSED all three services.
**Backfill** — the 21 moves of 2026-09-22, each from the record its commit cited:
| app | step (the service that moved) | verdict | evidence | note |
|---|---|---|---|---|
| actualbudget | actualbudget: actual-server:26.7.0 → actual-server:26.9.0 | proven | `audits/update-night-2026-09-21/apps/actualbudget/verdict.json` | — |
| audiobookshelf | audiobookshelf: audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | proven | `audits/update-night-2026-09-21/apps/audiobookshelf/verdict.json` | — |
| bookstack | bookstack: bookstack:26.05.2 → bookstack:26.05.5 | proven | `audits/update-night-2026-09-21/apps/bookstack/verdict.json` | — |
| docmost | docmost: docmost:0.95.0 → docmost:0.96.0 | proven | `audits/update-night-2026-09-21/apps/docmost/verdict.json` | — |
| emby | emby: embyserver:4.10.0.20 → embyserver:4.11.0.1 | proven | `audits/the-28-2026-09-22/apps/emby/verdict.json` | — |
| ghost | ghost: ghost:6.53.0-alpine → ghost:6.64.0-alpine | proven | `audits/the-28-2026-09-22/apps/ghost/verdict.json` | — |
| grafana | grafana: grafana:13.1.0 → grafana:13.2.2 | proven | `audits/update-night-2026-09-21/apps/grafana/verdict.json` | — |
| home-assistant | home-assistant: home-assistant:2026.7.2 → home-assistant:2026.9.3 | proven | `audits/update-night-2026-09-21/apps/home-assistant/verdict.json` | — |
| immich | immich-server: immich-server:v3.0.3 → immich-server:v3.2.2 | proven | `audits/the-28-2026-09-22/apps/immich/verdict.json` | — |
| mealie | mealie: mealie:v3.20.1 → mealie:v3.27.0 | proven | `audits/update-night-2026-09-21/apps/mealie/verdict.json` | — |
| n8n | n8n: n8n:2.31.3 → n8n:2.40.5 | proven | `audits/update-night-2026-09-21/apps/n8n/verdict.json` | — |
| navidrome | navidrome: navidrome:0.63.2 → navidrome:0.64.0 | proven | `audits/update-night-2026-09-21/apps/navidrome/verdict.json` | — |
| nextcloud | nextcloud-db: mariadb:11.6 → mariadb:12.3 | proven | `audits/update-night-2026-09-21/apps/nextcloud-engine-mariadb/verdict.json` | commit cited no record; this one named |
| papra | papra: papra:26.6.1-rootless → papra:26.6.2-rootless | proven | `audits/update-night-2026-09-21/apps/papra/verdict.json` | — |
| privatebin | privatebin: pdo:2.0.5 → pdo:2.0.6 | proven | `audits/update-night-2026-09-21/apps/privatebin/verdict.json` | — |
| radarr | radarr: radarr:6.3.0 → radarr:6.4.4 | proven | `audits/the-28-2026-09-22/apps/radarr/verdict.json` | — |
| romm | romm: romm:5.0.0 → romm:5.3.0 | proven | `audits/update-night-2026-09-21/apps/romm/verdict.json` | memory watch 80.9 % (M1) |
| sonarr | sonarr: sonarr:4.0.19 → sonarr:4.0.20 | proven | `audits/the-28-2026-09-22/apps/sonarr/verdict.json` | — |
| tandoor | tandoor: recipes:2.6.13 → recipes:2.6.15 | proven | `audits/update-night-2026-09-21/apps/tandoor/verdict.json` | — |
| termix | termix: termix:2.5.0 → termix:2.8.0 | proven | `audits/the-28-2026-09-22/apps/termix/verdict.json` | — |
| vikunja | vikunja: vikunja:2.3.0 → vikunja:2.6.0 | proven | `audits/update-night-2026-09-21/apps/vikunja/verdict.json` | — |
**21 of 21 `proven`, 0 `unrecorded`.** Every entry is marked `backfilled`, carries harness v1 (box walk only, no memory watch — except romm's M1 figure), and a digest resolved TONIGHT, which each entry's `note` says. `scripts/ladder_backfill.py`.
## Part C — move every app that proves itself
**Venues:** the bench = throwaway LXC **9401** on demo-hp (Debian 13, docker 26.1.5, Docker Hub login),
created tonight and destroyed at teardown; the box = guest **9202**, v0.267.0, drill catalog. **Negative
control first:** C3 (`alpine:3.20` as the TO image) → `failed` — the bench measures (`bench-C3/`).
Box walks by `run_edge.py`; bench edges by `upgrade-test.py --move`; each edge's bench evidence copied off
the bench when it ended (`benchq.py`). Evidence per app: `night-2026-09-23/apps/<app>/`.
| app | from → to | class | bench verdict | box verdict | memory: own (cgroup) peak % | migration line | moved? |
|---|---|---|---|---|---|---|---|
| nextcloud | nextcloud 34.0.1 → 34.0.4 | MariaDB + file-leg | proven | proven (done) | 24.2 (100) | nextcloud / Upgrading nextcloud from 34.0.1.2 ... | moved `5d49a11` (files_may_change) |
| romm | romm 5.3.0 → 5.3.1 | MariaDB | proven | proven (done) | n/a — measured before the own-memory sample existed (cgroup 79.5) | romm / INFO: [RomM][init][2026-09-23 21:12:26][0 | moved `7ff4e68` |
| romm-engine | romm-db mariadb 11.4 → 11.8 (own edge) | MariaDB engine | proven | proven (done) | 78.0 (84) | romm / INFO: [RomM][init][2026-09-23 21:47:51][0 | moved `431cdec` |
| kimai | kimai-db mariadb 11.6 → 11.8 (own edge) | MariaDB engine | proven | proven (done) | 37.9 (100) | kimai / [OK] Already at the latest version ("DoctrineMigrations\Version2 | moved `376b2c3` |
| adventurelog | adventurelog v0.12.1 → v0.13.0 (both images) | PostgreSQL (PostGIS) | failed ×2 (frontend never healthy; then backend crash-loop on `download-countries`) | failed (undone at +196 s, drill timeout 90 s) | — | adventurelog / Apply all migrations: account, admin, adventures, aut | **NOT moved** — our template's frontend healthcheck points at `/nodejs/bin/node`, gone in v0.13.0; with it dropped, the backend's start-time internet download crash-looped (R-655) |
| immich | immich-machine-learning v3.0.3 → v3.2.2 | PostgreSQL (server unchanged) | proven | proven (done) | 51.4 (100) | immich-server / [Nest] 7 - 09/23/2026, 9:11:07 PM LO | moved `b82b7c6` |
| n8n | n8n 2.40.5 → 2.41.1 | SQLite | proven | proven (done) | 22.7 (47) | n8n / Migrations in progress, please do NOT stop the process. | moved `04b63db` |
| ghost | ghost 6.64.0 → 6.65.0 | SQLite | proven | proven (done) | 24.9 (43) | ghost / [2026-09-23 21:21:22] WARN Database state requires migration | moved `e6afab9` (first watch had no load — re-run) |
| komga | komga 1.25.0 → 1.27.1 | embedded | proven | proven (done) | 60.0 (77) | komga / 2026-09-23T21:35:41.616+02:00 INFO 1 --- [ main] o.f.core | moved `fc1becc` with 768M (memory_tight at 512M) |
| opengist | opengist 1.13 → 1.15 | SQLite | proven | proven (done) | 77.8 (100) | — | moved `657c83f` (pages under `/-/`, R-654) |
| wishlist | wishlist v0.66.0 → v0.67.1 | SQLite | proven | proven (done) | 35.9 (53) | wishlist / Prisma schema loaded from prisma/schema.prisma. | moved `1545c5f` (after R-612's 512M) |
| navidrome | navidrome 0.64.0 → 0.64.1 | SQLite | proven | proven (done) | 6.3 (9) | navidrome / time="2026-09-23T19:31:04Z" level=info msg="goose: no migrations | moved `76cc1f5` (second ladder step) |
| emby | emby 4.11.0.1 → 4.11.0.3 | embedded | proven | proven (done) | 4.5 (8) | emby / Info SqliteUserRepository: Sqlite compiler options: ATOMIC_INTRINSICS | moved `8d573ad` |
| gitea | gitea 1.27.0 → 1.27.3 | SQLite | inconclusive | inconclusive () | — | — | NOT moved — inconclusive on both: its web installer blocks a headless first account (R-624) |
**12 steps across 11 apps published**, each written by `upgrade-test.py --write-ladder`, each its own
commit, every push through the hook (the new gate judged each against the registry), CI green (jobs
922–937, matched on head_sha; two commits rode inside a later push and were judged there). Tried: 14
apps, 15 edges. `night-2026-09-23/C-moves.txt` lists the commits.
**Listed, not moved:**
- *across a major* (11): sparkyfitness v0.17.3→v1.7.2 (both images), gokapi v1.9.6→v2.2.4, claper 2.5→3.0,
homepage v1.13.2→v2.4.0, paperless-ngx 2.20.15→3.2.1, mariadb →13.0 (romm, kimai, bookstack,
nextcloud), nextcloud 34→35, plex 1.41→1.43.
- *no front-door seed route* (R-624): code-server, outline, rallly; *sign-up closed by design*: vaultwarden,
zipline; *installer*: gitea (tried — inconclusive on both venues).
- *no fixture tonight*: bentopdf, glance, crafty-controller, wger (2.6→2.7), wanderer's meilisearch
(v1.36→v1.54), uptime-kuma (2.4.0→2.5.5 — its first account exists only over socket.io).
- *tried and failed*: adventurelog (R-655, two causes measured).
**What the night added to the bench:** the box walk's fixtures run on the bench (`upgrade_boxport.py`); a
wishlist fixture; opengist, komga and nextcloud fixtures fixed; the memory watch's load no longer errors on
box-fixture apps (R-653) and samples the app's own memory (decision 22, R-652); `files_may_change` from a
bind-mount tree hash (nextcloud's is the first entry carrying it).
**R-612 / R-613 — both fixed** (catalog `a5a729a`): wishlist 512M — 128M kills the first-boot seed every
run, sometimes after the Role/Group rows (so the sign-up lie is timing-dependent; it did NOT reproduce on
9202 tonight); 512M completes it (peak 312M). uptime-kuma `UPTIME_KUMA_DB_TYPE=sqlite` — before, the box
read `running` over `setup-database`; after, the real server (`/metrics` 401) and the database in the
volume. `apps/wishlist/r612-*`, `apps/uptime-kuma/r613-*`.
**A consequence of publishing tonight, seen on demo-hp:** romm now has three steps (the backfilled
5.0.0→5.3.0, 5.3.0→5.3.1, mariadb 11.4→11.8). The box cannot climb one step at a time yet (part 5), so the
press on 9201 took BOTH new steps at once — each tested alone, never together. It ended `done`, front door
200, 0 OOM kills after (`press9201/`).
## Part D — the chaos hour
**Written BEFORE round 1 (committed with this section).** Guest 9202, controller v0.267.0, drill catalog,
`update.health_timeout: 90s` (the drill setting — the product default is 5 min). Schedule drawn once by
`night-2026-09-23/chaos_schedule.py` from seed **20260923** (constraints in its docstring):
seed 20260923, drawn in 3 attempt(s)
| round | action | accident | at +s |
|---|---|---|---|
| 1 | install | none | 11 |
| 2 | good_update | box_backup | 46 |
| 3 | restore | controller_kill | 28 |
| 4 | fail_health | memory_hog | 43 |
| 5 | install | disk_fill | 28 |
| 6 | good_update | memory_hog | 46 |
| 7 | restore | power_cut | 56 |
| 8 | cutoff | box_backup | 23 |
| 9 | fail_health | controller_kill | 58 |
| 10 | remove | docker_restart | 10 |
| 11 | cutoff | two_updates | 51 |
| 12 | remove | none | 47 |
**The app each round acts on** (fixed by rule in `chaos.py`, not drawn):
| round | app | why this app |
|---|---|---|
| 1 install / 10 remove | actualbudget | a throwaway beside the six |
| 5 install / 12 remove | papra | a second throwaway |
| 2 good_update | docmost 0.95.0 → 0.96.0 | PostgreSQL, the BIG database (300 extra spaces) |
| 6 good_update | romm 5.3.0 → 5.3.1 | MariaDB |
| 3 restore | vikunja | volume-only (SQLite + files) |
| 7 restore | docmost | the big database, under a power cut |
| 4 fail_health | adventurelog v0.12.1 → v0.13.0, probe on a port it does not answer | the second PostgreSQL (PostGIS) |
| 9 fail_health | vikunja 2.5.0 → 2.6.0, wrong probe | volume-only, under a controller kill |
| 8 cutoff | navidrome 0.64.0 → 0.64.1, wrong probe, the undo copy's finished-marker removed during `verifying` | volume-only + a drive folder |
| 11 cutoff + two_updates | nextcloud 34.0.1 → 34.0.4 (the file-leg), and romm's database-engine Update (11.4 → 11.8, a tested good step) pressed at +51 s | two updates at once. *Changed before round 1:* the first draft named adventurelog, whose v0.13.0 cannot pass health (Part C) |
Setup: the six at their FROM pins, each seeded through its front door (A); one whole-box backup on 9202
(the restore rounds need a copy); seed B on docmost, romm and vikunja after it.
### The twelve rounds — the five things, and the data read back
Evidence per round: `night-2026-09-23/chaos/round-NN.{json,log}`, `round-NN-run.txt`; rounds 11–12 also
`round-NN-controller.log`. **When the accident landed** is stated, because in rounds 2–5 it did not land
where the schedule meant it to (below).
| # | action · app × accident | what the household saw (hu / en) | what the box did by itself | steady | events that fired | should have fired, did not | data read back |
|---|---|---|---|---|---|---|---|
| 1 | install · actualbudget × none | installed; badge none | deploy 202 → running | 26.6 s | app_deploy_started, app_deployed; **app_start_failed (warning)** 6 s in — its app cannot be named (see "instrument") | — | A ✓ |
| 2 | good update · docmost 0.95→0.96 (412 spaces) × whole-box backup | „Naprakész" / "Up to date" | safety-dump → pulling → copying (+76.6 s) → starting → verifying → **done 108 s** | 558 s (waited ~7 min for the catalog — instrument) | none | — | A ✓, B ✓ (paged re-read; the first read asked page 1 of 21) |
| 3 | restore · vikunja × controller kill (+28 s) | „Naprakész" / "Up to date" | restore 6 s, running; the kill landed AFTER it; controller back in 0.2 s | 31.5 s | none | — | A ✓, B ✓ |
| 4 | fail health · adventurelog × 1 GB hog | „A(z) adventurelog frissítése … nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el." / "The update of adventurelog … did not succeed. The box put back the previous version and its data automatically — nothing was lost." | verifying → **undoing +110.9 s → undone +155.1 s** (the hog had ended — instrument) | 605 s | health_change, **app_update_undone (warning)** ✓ | — | A ✓ |
| 5 | install · papra × disk to 3 GB free | installed | deploy → running in 27 s; the fill landed just after | 26.6 s | deploy pair | a disk warning: not measurable (the fill lasted ~20 s, the health job runs every 5 min) | A ✓ |
| 6 | good update · romm 5.3.0→5.3.1 × 1 GB hog (**during verifying**) | „Naprakész" / "Up to date" | … verifying → **done 55.8 s** | 61.7 s | none | — | A ✓ |
| 7 | restore · docmost × **power cut** (+56 s, after the 32 s restore) | „Naprakész" / "Up to date" | restore 32 s; cut; guest back, controller back 12 s, apps up | 103.5 s | controller_started | — | A ✓, B ✓ |
| 8 | cut-off undo copy · navidrome × whole-box backup (+23 s, during verifying) | „Megállítva — visszaállítás szükséges" + the hold sentence naming a restore / "Stopped — restore needed" + the same in English | verifying → undoing +94.9 → **failed 95.4 s**: the copy's marker missing → nothing poured back → HOLD | 101 s | **app_update_held (error)** ✓, then **app_start_failed** 11 s later | — (the second is extra: R-660) | held (A unreadable by design) → **the household's restore brought it back, A ✓** |
| 9 | fail health · vikunja 2.5→2.6 × **controller kill in verifying** | „A(z) vikunja frissítése … nem sikerült. A doboz automatikusan visszaállította …" / "… put back the previous version and its data automatically …" | killed at +58 s → *"update recovery: vikunja was interrupted in verifying … RESUMING the health wait"* → undoing → **undone 157.7 s** — with **`the undo copy will hold 0 named volume(s)`** (R-658) | 163 s | controller_started, **app_update_undone** ✓ | — | A ✓, B ✓ (the old version happens to run on migrated data) |
| 10 | remove · actualbudget × docker restart (+10 s) | the remove press answered **502** | stop done; remove lost to the restart; app stays installed and stopped | cap 1500 s | controller_started; app_start_failed + health_change every 5 min | — | app stopped (A not readable); the household must press again |
| 11 | cut-off undo copy · nextcloud × **a second Update (romm's engine step) at the same moment** | hold as round 8 (both languages); romm „Naprakész" | nextcloud → **failed/HOLD 112.6 s**; romm ran at the same time → **done** (no single-flight, §3b Q4) | 143 s | **app_update_held** ✓, app_start_failed 13 s later | — | romm A ✓; nextcloud held → **the named restore REFUSED → R-659, the stop rule** |
| 12 | remove · papra × none | not installed | stop → remove 200 → verified | 40.9 s | app_removed; health_change (the box-wide job: a held app + a stopped one — true) | — | removed clean (0 containers, 0 volumes) |
**The instrument, stated rather than hidden.** (1) Rounds 2 and 4 waited ~7 min for the drill commit on
the wrong file (an installed app's stack dir holds the APPLIED compose), so their accidents fired before the
update; round 3's and 5's actions ended before theirs. From round 5 the accident is armed at the press, and
from round 6 the wait reads the badge's own `catalog_images` (seconds). (2) The event column first came from
the controller's debug ring, which overflowed in the long rounds and MISSED both `app_update_held` events; it
was rebuilt from the controller's full log (`E8-events-*`) for rounds 7–12. Rounds 1–6 keep the ring's view,
and round 1's `app_start_failed` cannot be tied to an app: round 7's power cut took the container's earlier
log. From round 11 each round saves the whole controller log. (3) My stop rule first fired on round 8 because
a held app cannot answer; a hold is by design, so recovery — not the readback — is the test.
### The truth table (`08-alarm-ladder.md`)
| situation | the ladder says | fired? |
|---|---|---|
| an update undone (rounds 4, 9) | `app_update_undone`, warning, household + operator, one per app per outcome | ✓ both |
| an update held (rounds 8, 11) | `app_update_held`, error | ✓ both — and each was followed by an `app_start_failed` for the same moment (R-660: a held app is stopped by the product and should not ALSO read as down) |
| a controller kill / power cut / docker restart (3, 7, 9, 10) | `controller_started`, info; the boot grace suppresses app alarms for 90 s | ✓ controller_started each time; no app alarm inside the grace |
| a restore that finished (3, 7) | nothing (a refused or interrupted one: `restore_interrupted`) | ✓ nothing |
| a remove (12, teardown) | `app_removed`, info | ✓ |
| a disk filled to 1 GB above the floor (5) | a disk warning from the health job | not measurable — the fill lasted ~20 s |
| the box-wide health job with a held + a stopped app | `health_change`, warning | ✓ every 5 min |
### demo-hp 9201 — the standing apps Part C moved, pressed through the product
`night-2026-09-23/press9201/` (live catalog, a hub attached, controller v0.266.0; no whole-box backup pressed).
| app | from → to | phases | end | front door after | memory after |
|---|---|---|---|---|---|
| opengist | 1.13 → 1.15 | backing-up → … → done | **done in ~16 s** | `/healthcheck` 200 | — |
| kimai | mariadb 11.6 → 11.8 (engine alone) | … → done | **done in ~69 s** | `/en/login` 200 | — |
| romm | 5.3.0 → 5.3.1 **and** mariadb 11.4 → 11.8 in ONE press (no ladder on the box yet) | … → done | **done in ~90 s** | `/api/heartbeat` 200 | 600 MiB / 768, `oom_kill` 0, 0 restarts (checked 10 min later) |
Every standing app afterwards: 24 containers before, 24 after, all healthy (`E7-9201-after.txt`). Nothing
else on 9201 was touched.
## Teardown
**Machine — 9202:** every app of the night removed THROUGH THE PRODUCT (docmost, navidrome, vikunja, romm,
nextcloud, adventurelog, actualbudget; papra in round 12): 0 containers, 0 volumes, 0 undo copies left
(`E2-teardown-removes.*`). For the three RESTORED apps the remove answered `volumes_removed: []` while their
volumes did go (compose's `down --volumes` removes them by name) — the report is wrong, not the act (R-658).
The drive folders the product kept (R-442 on 9202, as every night) removed by name: navidrome, nextcloud, romm
(`E3-*`); older folders from earlier nights left as found. 54 test images removed by name before Part D (no
prune): `/var/lib/docker` 86 % → 24 %. `controller.yaml` restored and read back **identical** to
`.pre-night0923`; the catalog cache on 9202 reads live `5d49a11` (`E4-*`, `E5-*`). **9202 stays on controller
v0.267.0** (the scratch guest; self-update off).
**Machine — the bench:** LXC 9401 destroyed (`pct destroy --purge`, its hostname checked first); the Debian
template downloaded for it removed (`E6-*`).
**Machine — 9201:** only the three guarded Updates above; 24 containers before and after, all healthy.
**Host — demo-hp:** `pct list` = 9201, 9202 (as before); `local` 84.49 % before and after; `nvme-scratch`
7.45 % → 7.59 % and `local-lvm` 56.3 % → 57.7 % (the two guests' pulled images) (`00-*`, `E6-*`).
**Drill repo:** reset to live `main` `5d49a11b31fe`, `has_actions: false`, read back (`E5-*`).
**Hub:** not touched. **The floor was NOT raised** — see the three lines.
## Claims in the brief that turned out wrong, or held
| the brief's claim | what the night found |
|---|---|
| the controller ignores an unknown `.felhom.yml` key (Part B spike) | **held** — v0.266.0 and v0.267.0 synced, deployed, probed and badged navidrome with `update_ladder:` exactly as without it |
| R-633's window closes R-626 | **held, as far as one measurement goes** — 390 s of events + a restart + a reboot: nothing came back. Closed by measurement, not by a named creator |
| 21 backfill records exist | **held** — 21 of 21; one commit (nextcloud's engine move) cited none and its record was found by name |
| wishlist's fault is memory (read from R-612) | **held, and sharpened** — 128M kills the seed every run; whether the sign-up then fails depends on WHEN the kill lands (it did not fail on 9202 tonight) |
| R-518: the page promises „csak néhány másodpercre" | **wrong** — v0.243.0 had already changed it to „általában néhány perc"; tonight states the measured ≈ 8 min |
| `check-image-resolvable.py` already resolves digests | **half** — it asks whether a ref EXISTS and throws the digest away; `image_digest.py` was written |
| one chaos round ≈ one accident inside one action | **not as drawn** — actions of 6–30 s ended before accidents at +28–56 s; four rounds' accidents landed outside the action (stated per round) |
| the chaos schedule's round-11 second update (first draft: adventurelog) | changed BEFORE round 1 to romm's engine step — adventurelog v0.13.0 cannot pass health |