3e58c184f6
gates / gates (push) Successful in 27s
DRILL-night-2026-09-23.md: Parts A-E. 09 §3 decisions 21 (operator word), 22 and 23 (CC unattended, operator may reverse); §6.4 parts 4 and 6 (catalog half) shipped; §6.1a residuals R-658/R-659. Register 330 -> 336: R-651..R-660 opened (R-658 and R-659 P1), R-650/R-640/R-499/R-626 closed. Capability map, nightly rotation (opengist), STATUS (one question: the floor), CONTEXT, REPORT. The floor stays 0.266.0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
290 lines
29 KiB
Markdown
290 lines
29 KiB
Markdown
# DRILL — night shift 2026-09-23: DooPlex guarded from its own tests, the test record, the proven moves, the chaos hour
|
||
|
||
Evidence: `audits/night-2026-09-23/` (`PROGRESS.md` is the step log). Architecture read first and named:
|
||
`architecture/09-update-architecture.md` (§3 decisions 11–20, §6.1a, §6.4 parts 4–6, the verdict record),
|
||
`07-backup-architecture.md` §6, `08-alarm-ladder.md`. **Method: endpoint-level** — every product act is
|
||
the endpoint the UI invokes; pages are fetched as HTML in both languages; no browser.
|
||
|
||
## Not done, or changed
|
||
|
||
- **The floor stays at 0.266.0.** The brief raises it only if every live proof of Parts A, B and D passed; Part D
|
||
found two P1s (R-658, R-659). Both exist in 0.266.0 as well — v0.267.0 does not cause them and adds R-640's
|
||
protection to the same restore path — so the recommendation is to raise it; that is the operator's one
|
||
question this morning.
|
||
- **Adventurelog v0.13.0 was not moved** (R-655): our own template's healthcheck points at a node path the new
|
||
image moved, and the new backend downloads world data from the internet at every start and crash-loops when
|
||
the download is cut short.
|
||
- **The stop rule fired in round 11** (R-659): all twelve rounds had already run; nothing further was injected.
|
||
- **Rounds 2–5 did not land their accidents inside the action** (the instrument, stated in Part D); round 9 is
|
||
the clean mid-update kill, round 6 the mid-verify memory hog, rounds 8 and 11 the mid-verify backup and the
|
||
concurrent second update.
|
||
- **Decisions taken by CC unattended** (operator may reverse, `09` §3): 22 — the memory mark reads the app's own
|
||
memory, not the file cache; 23 — the push-time gate asks the registry for moved refs only.
|
||
- R-518's per-tier quiesce, R-612's silent-seed half and R-613's template sweep stay open, narrowed.
|
||
|
||
## Three lines
|
||
|
||
|
||
- **Interventions:** none by the operator; 5 by me on my own instruments (a stale-template wait, a dropped
|
||
event ring, a killed runner of my own, a broken local edit reverted before any commit, the stop rule on a held
|
||
app) — each stated where it happened. No product state was set by hand.
|
||
- **Apps moved tonight: 11 of 14 tried (12 steps of 15 edges)**, each with its test record; the 21 moves of
|
||
2026-09-22 backfilled.
|
||
- **The result that matters most:** after a restore, an app's volumes lose the label the undo finds them by —
|
||
so the next failed update is "undone" with **nothing** put back (R-658, P1); and a held file-leg app can be
|
||
pointed at a restore that refuses it (R-659, P1).
|
||
|
||
---
|
||
|
||
## Part A — controller v0.267.0 (`80e6ad8c4772`, CI job 919 success)
|
||
|
||
| row | what shipped | red-proof (seen failing, then restored) | evidence |
|
||
|---|---|---|---|
|
||
| **R-650** | `internal/dockerexec`: every docker exec (77 sites) goes through it; under `go test` a real docker is refused with an error naming the command, unless `FELHOM_TEST_REAL_DOCKER=1` or the binary is a stub under the temp dir. `TestR650_NoBareDockerExec` sweeps the tree. | guard disabled → the decoy's `docker ps -a` runs (`got <nil>`); one bare call put back → the sweep names `dbdump.go:140` | `A1-r650-*.txt` |
|
||
| R-650 sweep | **8 `api` tests FAILED on the refusal** — they built a real `stacks.Manager` and ran `docker ps` on DooPlex; `web` ran 67× `docker compose version`, 19× `docker ps`, 19× `docker info`, 10× `docker inspect felhom-samba`, 2× `docker exec felhom-samba`; `stacks` ran **1× `docker-compose down`**, 11× `compose ps`; `backup` 46× `docker ps`; `appexport` 7×; `system` 1×. `api`/`stacks`/`web` now run under `RunWithStub` (TestMain); the other three pass on the refusal. | — | `A1-r650-sweep.txt` |
|
||
| **R-640** | `appbackup.CheckDumpComplete` (the engine's end marker, last 4 KiB, gzip-aware). The unit restore and the off-site restore refuse a cut-off copy **before the first mutation**; every replay checks again before any load, whatever the import seam is. | replay check removed → `a cut-off copy was LOADED`; unit gate removed → `stopped=true calls=[stop recreate startsvc:immich-postgres start]`; off-site gate removed → only the replay check caught it, after a stop and a rollback | `A2-r640-redproof.txt` |
|
||
| **R-626** | **measured, not reproduced** on v0.266.0: navidrome deployed and removed through the product; `docker events` 390 s — the remove's kill/stop/die/destroy seen (the positive observable), **no create**; controller restarted at +150 s; guest rebooted after → 0 containers, 0 volumes at every check. My script first said `came_back: true` — it had counted the catalog-mirror directory every catalog app has; corrected in the verdict file, both values kept. Leftover found: `applied-compose.yml` + `applied-meta/` stay after a remove (R-651). | — | `A3-r626-*` |
|
||
| **R-499** | the Tier-2 page's system-disk sentence has four branches (own drive / same disk / drive gone / cannot ask); „(PBS)" and „nincs külön teendő" only where true. Live on 9202 (no agent): the `unknown` branch in hu + en, matching `/api/storage/backup-target` `known:false`. | handler not passing the fact → 3 branches missing; the old sentence in the same-disk branch → `promises the backup protects this app` | `A4-*`, `A7-*` |
|
||
| **R-518** | the backup button states the measured stop (≈ 8 min on a 12-app box), both languages, page + confirm. **The brief's „csak néhány másodpercre" was already gone** (v0.243.0); the vague „általában néhány perc" is replaced. | bundles back to v0.243.0 → measured downtime appears 0 times | `A5-*` |
|
||
|
||
Gates: `go build/vet` rc 0; `go test -count=1 ./...` rc 0 (31 packages); `controller_gates.py` all OK.
|
||
Deployed to **9202 only** (`A6-*`); the floor waits for Part E.
|
||
|
||
## Part B — the test record and its gate (catalog `6db08a5`, CI job 920 success)
|
||
|
||
**Spike first (15 min):** a drill navidrome carrying a top-level `update_ladder:` (one JSON entry per
|
||
line) on 9202 at **v0.266.0 and v0.267.0**: synced, deploy-fields read, deployed `running`, probe
|
||
`API GET :4533/ping → 200`, badge „Naprakész" / "Up to date", no YAML error. Source agrees (no
|
||
`KnownFields` anywhere). The brief's claim holds. `B1-spike-*`.
|
||
|
||
**The record** (`scripts/ladder.py`): `from`/`to` per service, `digest` per `to` ref (sha256 from
|
||
`scripts/image_digest.py` — stdlib, equals Docker's `RepoDigests` for `privatebin/pdo:2.0.6` on 9202),
|
||
`verdict` (proven | unrecorded), `tested_at`, `harness_version`, `evidence`, `box_evidence`,
|
||
`memory_peak_pct` (+ `memory_basis`, `memory_cgroup_peak_pct`), `marks` {files_may_change, needs_person,
|
||
memory_tight}. Written ONLY by `upgrade-test.py --write-ladder` (bench AND box `proven`, template at
|
||
FROM, digests resolved) — `test_ladder_writer.py` 5/5.
|
||
|
||
**The gate** — two rows of `catalog_gates.py`, both `--fast`:
|
||
`check-test-record.py` (static; runs in CI too: a ladder well-formed, continuous, and its newest step IS
|
||
the compose) and `check-test-record-move.py` (history + the registry for MOVED refs only: a move adds a
|
||
PROVEN, non-backfilled entry whose `from`/`to` are the compose before/after and whose digests the
|
||
registry still serves; `memory_tight` needs the limit raised in the same range). 16 decoy cases, both
|
||
directions. **Red-proofs:** the no-entry refusal removed → *a bare image move* and *the entry only in
|
||
README* both pass (rc 0, expected 1); `failed` allowed → both failed-verdict cases pass; the digest
|
||
comparison removed → *the registry now serves another digest* passes. `B2-gate-redproofs.txt`.
|
||
Live, on the first real move (romm): the gate judged it against the real registry (rc 0), and with a
|
||
wrong digest table it REFUSED all three services.
|
||
|
||
**Backfill** — the 21 moves of 2026-09-22, each from the record its commit cited:
|
||
|
||
| app | step (the service that moved) | verdict | evidence | note |
|
||
|---|---|---|---|---|
|
||
| actualbudget | actualbudget: actual-server:26.7.0 → actual-server:26.9.0 | proven | `audits/update-night-2026-09-21/apps/actualbudget/verdict.json` | — |
|
||
| audiobookshelf | audiobookshelf: audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | proven | `audits/update-night-2026-09-21/apps/audiobookshelf/verdict.json` | — |
|
||
| bookstack | bookstack: bookstack:26.05.2 → bookstack:26.05.5 | proven | `audits/update-night-2026-09-21/apps/bookstack/verdict.json` | — |
|
||
| docmost | docmost: docmost:0.95.0 → docmost:0.96.0 | proven | `audits/update-night-2026-09-21/apps/docmost/verdict.json` | — |
|
||
| emby | emby: embyserver:4.10.0.20 → embyserver:4.11.0.1 | proven | `audits/the-28-2026-09-22/apps/emby/verdict.json` | — |
|
||
| ghost | ghost: ghost:6.53.0-alpine → ghost:6.64.0-alpine | proven | `audits/the-28-2026-09-22/apps/ghost/verdict.json` | — |
|
||
| grafana | grafana: grafana:13.1.0 → grafana:13.2.2 | proven | `audits/update-night-2026-09-21/apps/grafana/verdict.json` | — |
|
||
| home-assistant | home-assistant: home-assistant:2026.7.2 → home-assistant:2026.9.3 | proven | `audits/update-night-2026-09-21/apps/home-assistant/verdict.json` | — |
|
||
| immich | immich-server: immich-server:v3.0.3 → immich-server:v3.2.2 | proven | `audits/the-28-2026-09-22/apps/immich/verdict.json` | — |
|
||
| mealie | mealie: mealie:v3.20.1 → mealie:v3.27.0 | proven | `audits/update-night-2026-09-21/apps/mealie/verdict.json` | — |
|
||
| n8n | n8n: n8n:2.31.3 → n8n:2.40.5 | proven | `audits/update-night-2026-09-21/apps/n8n/verdict.json` | — |
|
||
| navidrome | navidrome: navidrome:0.63.2 → navidrome:0.64.0 | proven | `audits/update-night-2026-09-21/apps/navidrome/verdict.json` | — |
|
||
| nextcloud | nextcloud-db: mariadb:11.6 → mariadb:12.3 | proven | `audits/update-night-2026-09-21/apps/nextcloud-engine-mariadb/verdict.json` | commit cited no record; this one named |
|
||
| papra | papra: papra:26.6.1-rootless → papra:26.6.2-rootless | proven | `audits/update-night-2026-09-21/apps/papra/verdict.json` | — |
|
||
| privatebin | privatebin: pdo:2.0.5 → pdo:2.0.6 | proven | `audits/update-night-2026-09-21/apps/privatebin/verdict.json` | — |
|
||
| radarr | radarr: radarr:6.3.0 → radarr:6.4.4 | proven | `audits/the-28-2026-09-22/apps/radarr/verdict.json` | — |
|
||
| romm | romm: romm:5.0.0 → romm:5.3.0 | proven | `audits/update-night-2026-09-21/apps/romm/verdict.json` | memory watch 80.9 % (M1) |
|
||
| sonarr | sonarr: sonarr:4.0.19 → sonarr:4.0.20 | proven | `audits/the-28-2026-09-22/apps/sonarr/verdict.json` | — |
|
||
| tandoor | tandoor: recipes:2.6.13 → recipes:2.6.15 | proven | `audits/update-night-2026-09-21/apps/tandoor/verdict.json` | — |
|
||
| termix | termix: termix:2.5.0 → termix:2.8.0 | proven | `audits/the-28-2026-09-22/apps/termix/verdict.json` | — |
|
||
| vikunja | vikunja: vikunja:2.3.0 → vikunja:2.6.0 | proven | `audits/update-night-2026-09-21/apps/vikunja/verdict.json` | — |
|
||
|
||
**21 of 21 `proven`, 0 `unrecorded`.** Every entry is marked `backfilled`, carries harness v1 (box walk only, no memory watch — except romm's M1 figure), and a digest resolved TONIGHT, which each entry's `note` says. `scripts/ladder_backfill.py`.
|
||
|
||
## Part C — move every app that proves itself
|
||
|
||
**Venues:** the bench = throwaway LXC **9401** on demo-hp (Debian 13, docker 26.1.5, Docker Hub login),
|
||
created tonight and destroyed at teardown; the box = guest **9202**, v0.267.0, drill catalog. **Negative
|
||
control first:** C3 (`alpine:3.20` as the TO image) → `failed` — the bench measures (`bench-C3/`).
|
||
Box walks by `run_edge.py`; bench edges by `upgrade-test.py --move`; each edge's bench evidence copied off
|
||
the bench when it ended (`benchq.py`). Evidence per app: `night-2026-09-23/apps/<app>/`.
|
||
|
||
| app | from → to | class | bench verdict | box verdict | memory: own (cgroup) peak % | migration line | moved? |
|
||
|---|---|---|---|---|---|---|---|
|
||
| nextcloud | nextcloud 34.0.1 → 34.0.4 | MariaDB + file-leg | proven | proven (done) | 24.2 (100) | nextcloud / Upgrading nextcloud from 34.0.1.2 ... | moved `5d49a11` (files_may_change) |
|
||
| romm | romm 5.3.0 → 5.3.1 | MariaDB | proven | proven (done) | n/a — measured before the own-memory sample existed (cgroup 79.5) | romm / INFO: [RomM][init][2026-09-23 21:12:26][0 | moved `7ff4e68` |
|
||
| romm-engine | romm-db mariadb 11.4 → 11.8 (own edge) | MariaDB engine | proven | proven (done) | 78.0 (84) | romm / INFO: [RomM][init][2026-09-23 21:47:51][0 | moved `431cdec` |
|
||
| kimai | kimai-db mariadb 11.6 → 11.8 (own edge) | MariaDB engine | proven | proven (done) | 37.9 (100) | kimai / [OK] Already at the latest version ("DoctrineMigrations\Version2 | moved `376b2c3` |
|
||
| adventurelog | adventurelog v0.12.1 → v0.13.0 (both images) | PostgreSQL (PostGIS) | failed ×2 (frontend never healthy; then backend crash-loop on `download-countries`) | failed (undone at +196 s, drill timeout 90 s) | — | adventurelog / Apply all migrations: account, admin, adventures, aut | **NOT moved** — our template's frontend healthcheck points at `/nodejs/bin/node`, gone in v0.13.0; with it dropped, the backend's start-time internet download crash-looped (R-655) |
|
||
| immich | immich-machine-learning v3.0.3 → v3.2.2 | PostgreSQL (server unchanged) | proven | proven (done) | 51.4 (100) | immich-server / [Nest] 7 - 09/23/2026, 9:11:07 PM LO | moved `b82b7c6` |
|
||
| n8n | n8n 2.40.5 → 2.41.1 | SQLite | proven | proven (done) | 22.7 (47) | n8n / Migrations in progress, please do NOT stop the process. | moved `04b63db` |
|
||
| ghost | ghost 6.64.0 → 6.65.0 | SQLite | proven | proven (done) | 24.9 (43) | ghost / [2026-09-23 21:21:22] WARN Database state requires migration | moved `e6afab9` (first watch had no load — re-run) |
|
||
| komga | komga 1.25.0 → 1.27.1 | embedded | proven | proven (done) | 60.0 (77) | komga / 2026-09-23T21:35:41.616+02:00 INFO 1 --- [ main] o.f.core | moved `fc1becc` with 768M (memory_tight at 512M) |
|
||
| opengist | opengist 1.13 → 1.15 | SQLite | proven | proven (done) | 77.8 (100) | — | moved `657c83f` (pages under `/-/`, R-654) |
|
||
| wishlist | wishlist v0.66.0 → v0.67.1 | SQLite | proven | proven (done) | 35.9 (53) | wishlist / Prisma schema loaded from prisma/schema.prisma. | moved `1545c5f` (after R-612's 512M) |
|
||
| navidrome | navidrome 0.64.0 → 0.64.1 | SQLite | proven | proven (done) | 6.3 (9) | navidrome / time="2026-09-23T19:31:04Z" level=info msg="goose: no migrations | moved `76cc1f5` (second ladder step) |
|
||
| emby | emby 4.11.0.1 → 4.11.0.3 | embedded | proven | proven (done) | 4.5 (8) | emby / Info SqliteUserRepository: Sqlite compiler options: ATOMIC_INTRINSICS | moved `8d573ad` |
|
||
| gitea | gitea 1.27.0 → 1.27.3 | SQLite | inconclusive | inconclusive () | — | — | NOT moved — inconclusive on both: its web installer blocks a headless first account (R-624) |
|
||
|
||
**12 steps across 11 apps published**, each written by `upgrade-test.py --write-ladder`, each its own
|
||
commit, every push through the hook (the new gate judged each against the registry), CI green (jobs
|
||
922–937, matched on head_sha; two commits rode inside a later push and were judged there). Tried: 14
|
||
apps, 15 edges. `night-2026-09-23/C-moves.txt` lists the commits.
|
||
|
||
**Listed, not moved:**
|
||
- *across a major* (11): sparkyfitness v0.17.3→v1.7.2 (both images), gokapi v1.9.6→v2.2.4, claper 2.5→3.0,
|
||
homepage v1.13.2→v2.4.0, paperless-ngx 2.20.15→3.2.1, mariadb →13.0 (romm, kimai, bookstack,
|
||
nextcloud), nextcloud 34→35, plex 1.41→1.43.
|
||
- *no front-door seed route* (R-624): code-server, outline, rallly; *sign-up closed by design*: vaultwarden,
|
||
zipline; *installer*: gitea (tried — inconclusive on both venues).
|
||
- *no fixture tonight*: bentopdf, glance, crafty-controller, wger (2.6→2.7), wanderer's meilisearch
|
||
(v1.36→v1.54), uptime-kuma (2.4.0→2.5.5 — its first account exists only over socket.io).
|
||
- *tried and failed*: adventurelog (R-655, two causes measured).
|
||
|
||
**What the night added to the bench:** the box walk's fixtures run on the bench (`upgrade_boxport.py`); a
|
||
wishlist fixture; opengist, komga and nextcloud fixtures fixed; the memory watch's load no longer errors on
|
||
box-fixture apps (R-653) and samples the app's own memory (decision 22, R-652); `files_may_change` from a
|
||
bind-mount tree hash (nextcloud's is the first entry carrying it).
|
||
|
||
**R-612 / R-613 — both fixed** (catalog `a5a729a`): wishlist 512M — 128M kills the first-boot seed every
|
||
run, sometimes after the Role/Group rows (so the sign-up lie is timing-dependent; it did NOT reproduce on
|
||
9202 tonight); 512M completes it (peak 312M). uptime-kuma `UPTIME_KUMA_DB_TYPE=sqlite` — before, the box
|
||
read `running` over `setup-database`; after, the real server (`/metrics` 401) and the database in the
|
||
volume. `apps/wishlist/r612-*`, `apps/uptime-kuma/r613-*`.
|
||
|
||
**A consequence of publishing tonight, seen on demo-hp:** romm now has three steps (the backfilled
|
||
5.0.0→5.3.0, 5.3.0→5.3.1, mariadb 11.4→11.8). The box cannot climb one step at a time yet (part 5), so the
|
||
press on 9201 took BOTH new steps at once — each tested alone, never together. It ended `done`, front door
|
||
200, 0 OOM kills after (`press9201/`).
|
||
|
||
## Part D — the chaos hour
|
||
|
||
**Written BEFORE round 1 (committed with this section).** Guest 9202, controller v0.267.0, drill catalog,
|
||
`update.health_timeout: 90s` (the drill setting — the product default is 5 min). Schedule drawn once by
|
||
`night-2026-09-23/chaos_schedule.py` from seed **20260923** (constraints in its docstring):
|
||
|
||
seed 20260923, drawn in 3 attempt(s)
|
||
|
||
| round | action | accident | at +s |
|
||
|---|---|---|---|
|
||
| 1 | install | none | 11 |
|
||
| 2 | good_update | box_backup | 46 |
|
||
| 3 | restore | controller_kill | 28 |
|
||
| 4 | fail_health | memory_hog | 43 |
|
||
| 5 | install | disk_fill | 28 |
|
||
| 6 | good_update | memory_hog | 46 |
|
||
| 7 | restore | power_cut | 56 |
|
||
| 8 | cutoff | box_backup | 23 |
|
||
| 9 | fail_health | controller_kill | 58 |
|
||
| 10 | remove | docker_restart | 10 |
|
||
| 11 | cutoff | two_updates | 51 |
|
||
| 12 | remove | none | 47 |
|
||
|
||
**The app each round acts on** (fixed by rule in `chaos.py`, not drawn):
|
||
|
||
| round | app | why this app |
|
||
|---|---|---|
|
||
| 1 install / 10 remove | actualbudget | a throwaway beside the six |
|
||
| 5 install / 12 remove | papra | a second throwaway |
|
||
| 2 good_update | docmost 0.95.0 → 0.96.0 | PostgreSQL, the BIG database (300 extra spaces) |
|
||
| 6 good_update | romm 5.3.0 → 5.3.1 | MariaDB |
|
||
| 3 restore | vikunja | volume-only (SQLite + files) |
|
||
| 7 restore | docmost | the big database, under a power cut |
|
||
| 4 fail_health | adventurelog v0.12.1 → v0.13.0, probe on a port it does not answer | the second PostgreSQL (PostGIS) |
|
||
| 9 fail_health | vikunja 2.5.0 → 2.6.0, wrong probe | volume-only, under a controller kill |
|
||
| 8 cutoff | navidrome 0.64.0 → 0.64.1, wrong probe, the undo copy's finished-marker removed during `verifying` | volume-only + a drive folder |
|
||
| 11 cutoff + two_updates | nextcloud 34.0.1 → 34.0.4 (the file-leg), and romm's database-engine Update (11.4 → 11.8, a tested good step) pressed at +51 s | two updates at once. *Changed before round 1:* the first draft named adventurelog, whose v0.13.0 cannot pass health (Part C) |
|
||
|
||
Setup: the six at their FROM pins, each seeded through its front door (A); one whole-box backup on 9202
|
||
(the restore rounds need a copy); seed B on docmost, romm and vikunja after it.
|
||
|
||
### The twelve rounds — the five things, and the data read back
|
||
|
||
Evidence per round: `night-2026-09-23/chaos/round-NN.{json,log}`, `round-NN-run.txt`; rounds 11–12 also
|
||
`round-NN-controller.log`. **When the accident landed** is stated, because in rounds 2–5 it did not land
|
||
where the schedule meant it to (below).
|
||
|
||
| # | action · app × accident | what the household saw (hu / en) | what the box did by itself | steady | events that fired | should have fired, did not | data read back |
|
||
|---|---|---|---|---|---|---|---|
|
||
| 1 | install · actualbudget × none | installed; badge none | deploy 202 → running | 26.6 s | app_deploy_started, app_deployed; **app_start_failed (warning)** 6 s in — its app cannot be named (see "instrument") | — | A ✓ |
|
||
| 2 | good update · docmost 0.95→0.96 (412 spaces) × whole-box backup | „Naprakész" / "Up to date" | safety-dump → pulling → copying (+76.6 s) → starting → verifying → **done 108 s** | 558 s (waited ~7 min for the catalog — instrument) | none | — | A ✓, B ✓ (paged re-read; the first read asked page 1 of 21) |
|
||
| 3 | restore · vikunja × controller kill (+28 s) | „Naprakész" / "Up to date" | restore 6 s, running; the kill landed AFTER it; controller back in 0.2 s | 31.5 s | none | — | A ✓, B ✓ |
|
||
| 4 | fail health · adventurelog × 1 GB hog | „A(z) adventurelog frissítése … nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el." / "The update of adventurelog … did not succeed. The box put back the previous version and its data automatically — nothing was lost." | verifying → **undoing +110.9 s → undone +155.1 s** (the hog had ended — instrument) | 605 s | health_change, **app_update_undone (warning)** ✓ | — | A ✓ |
|
||
| 5 | install · papra × disk to 3 GB free | installed | deploy → running in 27 s; the fill landed just after | 26.6 s | deploy pair | a disk warning: not measurable (the fill lasted ~20 s, the health job runs every 5 min) | A ✓ |
|
||
| 6 | good update · romm 5.3.0→5.3.1 × 1 GB hog (**during verifying**) | „Naprakész" / "Up to date" | … verifying → **done 55.8 s** | 61.7 s | none | — | A ✓ |
|
||
| 7 | restore · docmost × **power cut** (+56 s, after the 32 s restore) | „Naprakész" / "Up to date" | restore 32 s; cut; guest back, controller back 12 s, apps up | 103.5 s | controller_started | — | A ✓, B ✓ |
|
||
| 8 | cut-off undo copy · navidrome × whole-box backup (+23 s, during verifying) | „Megállítva — visszaállítás szükséges" + the hold sentence naming a restore / "Stopped — restore needed" + the same in English | verifying → undoing +94.9 → **failed 95.4 s**: the copy's marker missing → nothing poured back → HOLD | 101 s | **app_update_held (error)** ✓, then **app_start_failed** 11 s later | — (the second is extra: R-660) | held (A unreadable by design) → **the household's restore brought it back, A ✓** |
|
||
| 9 | fail health · vikunja 2.5→2.6 × **controller kill in verifying** | „A(z) vikunja frissítése … nem sikerült. A doboz automatikusan visszaállította …" / "… put back the previous version and its data automatically …" | killed at +58 s → *"update recovery: vikunja was interrupted in verifying … RESUMING the health wait"* → undoing → **undone 157.7 s** — with **`the undo copy will hold 0 named volume(s)`** (R-658) | 163 s | controller_started, **app_update_undone** ✓ | — | A ✓, B ✓ (the old version happens to run on migrated data) |
|
||
| 10 | remove · actualbudget × docker restart (+10 s) | the remove press answered **502** | stop done; remove lost to the restart; app stays installed and stopped | cap 1500 s | controller_started; app_start_failed + health_change every 5 min | — | app stopped (A not readable); the household must press again |
|
||
| 11 | cut-off undo copy · nextcloud × **a second Update (romm's engine step) at the same moment** | hold as round 8 (both languages); romm „Naprakész" | nextcloud → **failed/HOLD 112.6 s**; romm ran at the same time → **done** (no single-flight, §3b Q4) | 143 s | **app_update_held** ✓, app_start_failed 13 s later | — | romm A ✓; nextcloud held → **the named restore REFUSED → R-659, the stop rule** |
|
||
| 12 | remove · papra × none | not installed | stop → remove 200 → verified | 40.9 s | app_removed; health_change (the box-wide job: a held app + a stopped one — true) | — | removed clean (0 containers, 0 volumes) |
|
||
|
||
**The instrument, stated rather than hidden.** (1) Rounds 2 and 4 waited ~7 min for the drill commit on
|
||
the wrong file (an installed app's stack dir holds the APPLIED compose), so their accidents fired before the
|
||
update; round 3's and 5's actions ended before theirs. From round 5 the accident is armed at the press, and
|
||
from round 6 the wait reads the badge's own `catalog_images` (seconds). (2) The event column first came from
|
||
the controller's debug ring, which overflowed in the long rounds and MISSED both `app_update_held` events; it
|
||
was rebuilt from the controller's full log (`E8-events-*`) for rounds 7–12. Rounds 1–6 keep the ring's view,
|
||
and round 1's `app_start_failed` cannot be tied to an app: round 7's power cut took the container's earlier
|
||
log. From round 11 each round saves the whole controller log. (3) My stop rule first fired on round 8 because
|
||
a held app cannot answer; a hold is by design, so recovery — not the readback — is the test.
|
||
|
||
### The truth table (`08-alarm-ladder.md`)
|
||
|
||
| situation | the ladder says | fired? |
|
||
|---|---|---|
|
||
| an update undone (rounds 4, 9) | `app_update_undone`, warning, household + operator, one per app per outcome | ✓ both |
|
||
| an update held (rounds 8, 11) | `app_update_held`, error | ✓ both — and each was followed by an `app_start_failed` for the same moment (R-660: a held app is stopped by the product and should not ALSO read as down) |
|
||
| a controller kill / power cut / docker restart (3, 7, 9, 10) | `controller_started`, info; the boot grace suppresses app alarms for 90 s | ✓ controller_started each time; no app alarm inside the grace |
|
||
| a restore that finished (3, 7) | nothing (a refused or interrupted one: `restore_interrupted`) | ✓ nothing |
|
||
| a remove (12, teardown) | `app_removed`, info | ✓ |
|
||
| a disk filled to 1 GB above the floor (5) | a disk warning from the health job | not measurable — the fill lasted ~20 s |
|
||
| the box-wide health job with a held + a stopped app | `health_change`, warning | ✓ every 5 min |
|
||
|
||
### demo-hp 9201 — the standing apps Part C moved, pressed through the product
|
||
|
||
`night-2026-09-23/press9201/` (live catalog, a hub attached, controller v0.266.0; no whole-box backup pressed).
|
||
|
||
| app | from → to | phases | end | front door after | memory after |
|
||
|---|---|---|---|---|---|
|
||
| opengist | 1.13 → 1.15 | backing-up → … → done | **done in ~16 s** | `/healthcheck` 200 | — |
|
||
| kimai | mariadb 11.6 → 11.8 (engine alone) | … → done | **done in ~69 s** | `/en/login` 200 | — |
|
||
| romm | 5.3.0 → 5.3.1 **and** mariadb 11.4 → 11.8 in ONE press (no ladder on the box yet) | … → done | **done in ~90 s** | `/api/heartbeat` 200 | 600 MiB / 768, `oom_kill` 0, 0 restarts (checked 10 min later) |
|
||
|
||
Every standing app afterwards: 24 containers before, 24 after, all healthy (`E7-9201-after.txt`). Nothing
|
||
else on 9201 was touched.
|
||
|
||
## Teardown
|
||
|
||
**Machine — 9202:** every app of the night removed THROUGH THE PRODUCT (docmost, navidrome, vikunja, romm,
|
||
nextcloud, adventurelog, actualbudget; papra in round 12): 0 containers, 0 volumes, 0 undo copies left
|
||
(`E2-teardown-removes.*`). For the three RESTORED apps the remove answered `volumes_removed: []` while their
|
||
volumes did go (compose's `down --volumes` removes them by name) — the report is wrong, not the act (R-658).
|
||
The drive folders the product kept (R-442 on 9202, as every night) removed by name: navidrome, nextcloud, romm
|
||
(`E3-*`); older folders from earlier nights left as found. 54 test images removed by name before Part D (no
|
||
prune): `/var/lib/docker` 86 % → 24 %. `controller.yaml` restored and read back **identical** to
|
||
`.pre-night0923`; the catalog cache on 9202 reads live `5d49a11` (`E4-*`, `E5-*`). **9202 stays on controller
|
||
v0.267.0** (the scratch guest; self-update off).
|
||
**Machine — the bench:** LXC 9401 destroyed (`pct destroy --purge`, its hostname checked first); the Debian
|
||
template downloaded for it removed (`E6-*`).
|
||
**Machine — 9201:** only the three guarded Updates above; 24 containers before and after, all healthy.
|
||
**Host — demo-hp:** `pct list` = 9201, 9202 (as before); `local` 84.49 % before and after; `nvme-scratch`
|
||
7.45 % → 7.59 % and `local-lvm` 56.3 % → 57.7 % (the two guests' pulled images) (`00-*`, `E6-*`).
|
||
**Drill repo:** reset to live `main` `5d49a11b31fe`, `has_actions: false`, read back (`E5-*`).
|
||
**Hub:** not touched. **The floor was NOT raised** — see the three lines.
|
||
|
||
## Claims in the brief that turned out wrong, or held
|
||
|
||
| the brief's claim | what the night found |
|
||
|---|---|
|
||
| the controller ignores an unknown `.felhom.yml` key (Part B spike) | **held** — v0.266.0 and v0.267.0 synced, deployed, probed and badged navidrome with `update_ladder:` exactly as without it |
|
||
| R-633's window closes R-626 | **held, as far as one measurement goes** — 390 s of events + a restart + a reboot: nothing came back. Closed by measurement, not by a named creator |
|
||
| 21 backfill records exist | **held** — 21 of 21; one commit (nextcloud's engine move) cited none and its record was found by name |
|
||
| wishlist's fault is memory (read from R-612) | **held, and sharpened** — 128M kills the seed every run; whether the sign-up then fails depends on WHEN the kill lands (it did not fail on 9202 tonight) |
|
||
| R-518: the page promises „csak néhány másodpercre" | **wrong** — v0.243.0 had already changed it to „általában néhány perc"; tonight states the measured ≈ 8 min |
|
||
| `check-image-resolvable.py` already resolves digests | **half** — it asks whether a ref EXISTS and throws the digest away; `image_digest.py` was written |
|
||
| one chaos round ≈ one accident inside one action | **not as drawn** — actions of 6–30 s ended before accidents at +28–56 s; four rounds' accidents landed outside the action (stated per round) |
|
||
| the chaos schedule's round-11 second update (first draft: adventurelog) | changed BEFORE round 1 to romm's engine step — adventurelog v0.13.0 cannot pass health |
|