R-672 session: audit README (not done first, wrong claims), STATUS, capability map, register 344->341, topic REPORT
gates / gates (push) Successful in 28s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-24 17:06:09 +02:00
parent 54bff69f88
commit d53bd08442
29 changed files with 395 additions and 24 deletions
+13
View File
@@ -0,0 +1,13 @@
# REPORT — R-672/R-673: restore test off, 9201 repaired, agent v0.133.0 + hub v0.124.0, controller v0.270.0
Full record: `documentation/audits/r672-2026-09-24/README.md` (opens with "not done, or changed" and the brief's
wrong claims).
- **Hub v0.124.0** (`2d24931`, manifest `068e065`, live): a thin pool is judged on the worse of data and metadata,
critical at 90 %; `storage_fill_*` cooldown per pool (host/storage), 6 h. Two red-proofs; suite green.
- **Agent v0.133.0** released (tag `9bdb4da`), not delivered; **controller v0.270.0**, floor 0.270.0, both demo boxes
arrived.
- **Docs:** `03` §8 (the preflight, the operator rulings, the -1 trap), `08` §6.2 (thin pool at 90 %), capability map
row, register (R-684 opened; R-669/R-674/R-679/R-681 closed; R-672/R-673 updated), CONTEXT, STATUS.
- **Register:** 344 → 341 rows (685,662 → 685,148 B).
- `unproven.py --summary`: see the session's final message (unchanged — 35 of 55 not walked).
+17 -20
View File
@@ -1,28 +1,25 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-24 (night shift, run in daytime from 11:07). The demo boxes are on controller 0.269.1 — except demo-hp's customer box, which is damaged and needs you (first item under "What needs you").**
**Updated 2026-09-24 (evening). The demo boxes are on controller 0.270.0. The HP customer box is repaired. A fixed host agent (0.133.0) is released and waits for your signature. Until then the automatic restore test is off on both demo boxes.**
**Decisions I took on my own (you may reverse either).**
- **A second controller release tonight (0.269.1).** The first one (0.269.0) had a flaw I found in its own live test: when the catalog re-tested an image, the box wrote the new image into a running app's file before anyone pressed Update, so any restart would have changed the image with no backup and no undo. I fixed it and released again rather than send the flaw to the fleet. Now an installed app keeps its image until an Update moves it.
- **A crash loop is 6 restarts in 10 minutes, not the 10 you wrote.** Measured: Docker slows a steady crash loop to about one restart a minute, so 10 would never catch it. No healthy app in any test record restarted more than once while starting.
**Decisions I took on my own (you may reverse them).**
- **The space check uses the real, uncompressed size, not the backup file size.** The rule in your brief ("file size × 1.2 + 5 GB") would NOT have stopped last night's accident. The HP box's backup file is 7 GB, but restoring it writes 22.6 GB.
- **To switch the test off I used the value −1, not 0.** In this setting, 0 means "every 6 hours". Only a negative number turns it off.
- **The hub got a small release too (0.124.0).** A thin disk pool now counts as urgent at 90 % (data or its bookkeeping part), with one alarm per pool per 6 hours. Before, the alarm came only at 95 %, up to 15 minutes late, and one full pool could silence another.
**What I tested, and it worked.**
- **The second drive now brings an app with files back whole** (your ruling). Tested four times on the scratch machine, twice under a power cut or a nearly full disk. Every time: the account and the files came back, a file the household had edited later was kept, and nothing was deleted.
- **A stranded app can only be removed keeping its data** (your ruling). The box refuses to delete the data, in both languages.
- **The box stops a crash loop or a memory storm** (your ruling), tells the household and you, and Start gives one more try. A second stop within a day says support is informed.
- **Exact image fingerprints.** The box runs exactly the image the catalog tested, and says "update available" when a newer tested image exists for the same version name.
- **Automatic updates: measured, not built.** A test caller updated three apps (up to three steps each) and set a failing one aside in about 7 minutes. The build plan is written with the numbers.
- **A chaos hour, 12 rounds** (power cuts, killed controller, Docker restarts, a full disk, backups). Updates resumed or undid themselves correctly every time.
**What I did, and it worked.**
- **The HP customer box is repaired.** I stopped it, checked both disks (small damage fixed, second check clean), and started it. All apps came back, it took the current version by itself, and the hub shows it online. Two apps' cache files (Redis) had been cut off by the full disk; with your yes I kept copies and repaired them. Both apps run.
- **The box's own crash-loop stop worked for real:** while the cache was broken, the box stopped those two apps and told the hub, by itself.
- **The new agent never starts a restore test that does not fit.** Tested on the HP box: it refused, said why, and created nothing.
- **Four controller fixes (0.270.0):** Update on an app that is already current is refused, with no restart. An install cut off by a restart is now cleaned up, reported and shown on the page. After a restore, the box keeps the right health check. One log line is corrected.
**What broke.**
- **demo-hp's customer box (not caused by tonight's work).** This morning the host's automatic restore test copied a whole guest onto the same full disk pool. The pool filled, and the customer box's disks became read-only. I freed the pool with an agent restart (the product's own clean-up). I could not repair the box itself.
- **An install interrupted by a restart is lost without a word**, and a remove interrupted the same way is left half-done. Both recorded, not fixed.
- After a failed update is restored, the box can keep the failed version's health check, and a later undo then fails for no reason. Recorded.
- Not done: moving more apps in the catalog. It needs a throwaway test machine, and this session could not delete one afterwards.
**What broke, or is not done.**
- **The HP box's own backups have been failing every day since yesterday.** Its host disk has 4 GB free, but one backup needs 7 GB, and three old ones are kept. Nothing fixes this by itself.
- **No full restore test fits on the HP box today** (it needs 30 GB free, 22 GB is free). The new agent will refuse it every time, correctly. So a passing test on the HP box is not proven yet.
**Rows.** 16 opened, 7 closed. The list went from 335 to 344.
**Rows.** 1 opened, 4 closed, 2 updated. The list went from 344 to 341.
**What needs you.**
1. **Repair demo-hp's customer box:** stop it, check its two disks, start it. If you do nothing, it keeps running but cannot save anything: no backups, no logs, no updates, and it stays on the old controller.
2. **Decide how the restore test may use disk space** (for example: it must check free space first and never use the pool of the box it tests). If you do nothing, the next scheduled restore test can fill the pool again, on any box with a small disk.
3. **Optional:** the two decisions above. If you do nothing, they stay as built.
1. **Sign the agent update (0.133.0) for demo-hp and demo-felhom.** Version 0.133.0, sha256 `3aa303452b8c6be58573d00af01a0ab4a0d97e4f885ffecd0144a18ac24e69b6`, one `felhom-opsign -op agent_update` job per box, the same way as 0.131.0. If you do nothing, the boxes keep the old agent and the automatic restore test stays off.
2. **After the new agent is on a box, turn the restore test back on.** On that host, remove the key `restore_test_eval_interval_seconds` from `backup` in `/etc/felhom-agent/agent.json` (the saved copy `agent.json.pre-r672` has it absent), then `systemctl restart felhom-agent`. If you do nothing, no backup is proven restorable on that box.
3. **Decide what to do with the HP box's backup storage:** keep fewer old backups, back up somewhere else, or add disk. If you do nothing, its whole-box backup keeps failing every night.
File diff suppressed because one or more lines are too long
@@ -0,0 +1,2 @@
gitea.dooplex.hu/admin/felhom-controller:0.270.0
gitea.dooplex.hu/admin/felhom-controller:0.270.0 Up 20 seconds (healthy)
@@ -0,0 +1,9 @@
privatebin: installed= {"privatebin": {"ref": "privatebin/pdo:2.0.6", "digest": "sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5", "at": "2026-09-22T19:44:56Z"}} catalog= {'privatebin': 'privatebin/pdo:2.0.6'} badge= [{'title': 'This app is running the newest version available.', 'text': 'Up to date'}]
hu Update -> 409 {"ok": false, "data": {"reason": "already_current"}, "error": "Ez az alkalmazás már a legfrissebb elérhető változatot futtatja — nincs mit frissíteni."}
en Update -> 409 {"ok": false, "data": {"reason": "already_current"}, "error": "This app is already running the newest version available — there is nothing to update."}
after: updating= False phase= None state= running
2026/09/24 14:51:18 router.go:591: [INFO] [api] update requested for stack: privatebin
2026/09/24 14:51:18 update.go:359: [ERROR] [stacks] update privatebin REFUSED (already_current): installed equals the catalog head on every service and no newer tested digest (installed=map[privatebin:{privatebin/pdo:2.0.6 sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5 2026-09-22T19:44:56Z}] catalog=map[privatebin:privatebin/pdo:2.0.6])
2026/09/24 14:51:18 router.go:591: [INFO] [api] update requested for stack: privatebin
2026/09/24 14:51:18 update.go:359: [ERROR] [stacks] update privatebin REFUSED (already_current): installed equals the catalog head on every service and no newer tested digest (installed=map[privatebin:{privatebin/pdo:2.0.6 sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5 2026-09-22T19:44:56Z}] catalog=map[privatebin:privatebin/pdo:2.0.6])
@@ -0,0 +1,7 @@
before: deployed= False
no actualbudget image cached
16:51:41 deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
-rw-r--r-- 1 root root 21 Sep 24 14:51 .felhom-install-pending
-rw------- 1 root root 251 Sep 24 14:51 app.yaml
@@ -0,0 +1 @@
16:51:52 killed pid 157525 rc=0
@@ -0,0 +1,19 @@
stack: deployed= True state= running install_interrupted= None deploy_error= None
2026/09/24 14:51:53 manager.go:1554: [INFO] [stacks] actualbudget actualbudget/actual-server:26.9.0@sha256:552beab3dec8c93d46b8b9245612d63c3f123b8a45063a474f53e229b17621d3 running Up 3 seconds (health: starting)
2026/09/24 14:51:54 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
2026/09/24 14:51:54 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for actualbudget → /mnt/sys_drive/felhom-data/backups/primary/actualbudget (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0)
--- files
-rw------- 1 root root 523 Sep 24 14:51 app.yaml
-rw-r--r-- 1 root root 1484 Sep 24 14:51 applied-compose.yml
drwxr-xr-x 2 root root 4096 Sep 24 14:51 applied-meta
--- containers
actualbudget Up 13 seconds (healthy)
(end)
hu page: None
en page: None
@@ -0,0 +1 @@
16:53:05 deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
@@ -0,0 +1,2 @@
16:53:05 /opt/docker/stacks/mealie/.felhom-install-pending
16:53:05 killed pid 158798 rc=0
@@ -0,0 +1,17 @@
stack: deployed= False state= not_deployed install_interrupted= True deploy_error= interrupted by a controller restart before it finished — install it again
2026/09/24 14:51:54 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
2026/09/24 14:53:05 router.go:431: [INFO] [api] Deploy requested for stack: mealie
2026/09/24 14:53:05 [INFO] [stacks] SaveAppConfig: saved config for mealie
2026/09/24 14:53:05 deploy.go:383: [INFO] [stacks] Deploying stack mealie with 2 env vars: [DOMAIN, SUBDOMAIN]
2026/09/24 14:53:05 manager.go:1590: [INFO] [stacks] Deploying stack mealie — checking 1 images...
2026/09/24 14:53:06 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
2026/09/24 14:53:07 install_interrupted.go:70: [WARN] [stacks] install mealie was INTERRUPTED by a controller restart — removing what it started (volumes kept) and reporting it (R-681)
2026/09/24 14:53:07 [WARN] [stacks] 1 install(s) interrupted by the restart were resolved and reported: [mealie]
--- files
-rw-r--r-- 1 root root 21 Sep 24 14:53 .felhom-install-interrupted
-rw------- 1 root root 255 Sep 24 14:53 app.yaml
--- containers
(end)
hu page: A telepítés egy újraindítás miatt félbemaradt, és a doboz eltávolította, amit elkezdett. Nyomd meg újra a Telepítés gombot.
en page: The installation was interrupted by a restart, and the box removed what it had started. Press Install again.
@@ -0,0 +1,12 @@
16:53:51 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
16:55:11 [1] deployed, controller state=running, pinned={'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'}
reinstall ok= True deployed= True install_interrupted= None
no install markers
16:55:16 [X] stop -> 200 {'ok': True, 'message': 'Stack mealie stop completed'}
16:55:48 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'mealie', 'volumes_removed': ['mealie_mealie_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalm
16:55:56 [X] after remove: deployed=False leftovers='/opt/docker/stacks/mealie'
remove -> 200
no mealie volumes
no leftovers
@@ -0,0 +1,8 @@
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
sync_interval: 15m
token: <redacted>
username: "admin"
hub:
update:
health_timeout: 90s
@@ -0,0 +1,84 @@
{
"1-installed": "stack: \napplied: \nunit: ()",
"2-before-update": "stack: \napplied: \nunit: ()",
"2-update": {
"accepted": true,
"http": "202",
"phases": [
{
"t": 0.0,
"phase": "backing-up",
"label": "Biztonsági mentés készül a frissítés előtt…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 2.1,
"phase": "safety-dump",
"label": "Adatbázis pillanatkép…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 3.1,
"phase": "pulling",
"label": "Új verzió letöltése…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 4.1,
"phase": "copying",
"label": "Az adatok másolása a frissítés előtt…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 16.4,
"phase": "verifying",
"label": "Működés ellenőrzése…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 106.6,
"phase": "undoing",
"label": "Visszaállítás az előző változatra…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 125.0,
"phase": "undone",
"label": "Visszaállítva az előző változatra",
"updating": false,
"error": null,
"hold": null
}
],
"duration_s": 125.1,
"final_phase": "undone",
"update_error": null,
"hold_reason": null,
"state": "running"
},
"2-after-update": "stack: \napplied: \nunit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)",
"CHECK_A_unit_keeps_pinned_probe": false,
"3-before-restore": "stack: \napplied: \nunit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)",
"3-restore": {
"op": "tier2-unit-restore",
"stack": "wishlist",
"ok": true,
"message": "A(z) wishlist: 2 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-24 16:57).",
"finished_at": "2026-09-24T15:00:33.459543237Z"
},
"3-after-restore": "stack: \napplied: \nunit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)",
"CHECK_B_restore_resets_applied": false,
"log": "2026/09/24 14:57:56 pin.go:373: [INFO] [stacks] update wishlist: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/wishlist/docker-compose.yml (wishlist=ghcr.io/cmintey/wishlist:latest)\n2026/09/24 14:58:08 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_data → wishlist_wishlist_data.pre-update-20260924T145808Z in 459ms\n2026/09/24 14:58:09 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_uploads → wishlist_wishlist_uploads.pre-update-20260924T145808Z in 440ms\n2026/09/24 14:59:40 update.go:1274: [INFO] [stacks] update wishlist: phase undoing\n2026/09/24 14:59:40 undo.go:513: [WARN] [stacks] update wishlist: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health check failing))\n2026/09/24 14:59:42 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1\n2026/09/24 14:59:57 undo.go:633: [INFO] [stacks] update wishlist: event app_update_undone\n2026/09/24 14:59:57 undo.go:561: [INFO] [stacks] update wishlist: UNDONE in 18s — the previous version is running on the data from before the update (the app's health check passed)\n2026/09/24 15:00:19 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1\n2026/09/24 15:00:33 restore_unit.go:464: [INFO] [backup] Restore-from-unit completed: wishlist — 2 volume(s) of 2 listed, 0 database(s) of 0 listed\n"
}
@@ -0,0 +1,20 @@
16:57:21 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
16:57:41 [1] deployed, controller state=running, pinned={'wishlist': 'ghcr.io/cmintey/wishlist:v0.67.1'}
16:57:45 [1] stack: | applied: | unit: ()
16:57:46 drill: wishlist: failing step ghcr.io/cmintey/wishlist:v0.67.1 -> ghcr.io/cmintey/wishlist:latest + probe 8999
16:57:53 [2] stack: | applied: | unit: ()
16:57:53 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
16:57:53 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None
16:57:55 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
16:57:56 + 3.1s phase=pulling label=Új verzió letöltése… err=None hold=None
16:57:57 + 4.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None
16:58:10 + 16.4s phase=verifying label=Működés ellenőrzése… err=None hold=None
16:59:40 + 106.6s phase=undoing label=Visszaállítás az előző változatra… err=None hold=None
16:59:58 + 125.0s phase=undone label=Visszaállítva az előző változatra err=None hold=None
17:00:02 [2] after: stack: | applied: | unit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)
17:00:03 drill: wishlist image back to v0.67.1 (probe left at 8999)
17:00:18 [3] stack: | applied: | unit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)
17:00:18 [3] POST /backup/tier2/unit-restore -> 302
17:00:34 [3] restore: {'op': 'tier2-unit-restore', 'stack': 'wishlist', 'ok': True, 'message': 'A(z) wishlist: 2 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-24 16:57).', 'finished_at'
17:00:37 [3] after: stack: | applied: | unit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)
17:00:40 CHECK A (unit keeps the pinned probe): False CHECK B (restore resets the applied record): False
@@ -0,0 +1,9 @@
== .felhom.yml (2026-09-24 15:00:19)
- type: http
port: 3000
== applied-meta/.felhom.yml (2026-09-24 15:00:19)
- type: http
port: 3000
== /mnt/sys_drive/felhom-data/backups/primary/wishlist/compose/.felhom.yml (2026-09-24 14:57:55)
- type: http
port: 3000
@@ -0,0 +1,16 @@
2026/09/24 14:57:22 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1
2026/09/24 14:57:46 sync.go:416: [INFO] [sync] Updated wishlist/.felhom.yml
2026/09/24 14:57:53 update.go:1274: [INFO] [stacks] update wishlist: phase backing-up
2026/09/24 14:57:53 update_guard.go:360: [INFO] [backup] update pre-backup for wishlist: starting (DB dump → volume dump → unit capture → Tier 2)
2026/09/24 14:57:55 update_guard.go:410: [INFO] [backup] update pre-backup for wishlist: volume dump OK
2026/09/24 14:57:55 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for wishlist → /mnt/sys_drive/felhom-data/backups/primary/wishlist (images=1, secrets-referenced=0, data_keys=0, port
2026/09/24 14:57:55 update_guard.go:416: [INFO] [backup] update pre-backup for wishlist: recovery unit captured (0 database dump(s))
2026/09/24 14:57:55 update_guard.go:425: [INFO] [backup] update pre-backup for wishlist: complete in 2.038s
2026/09/24 14:57:56 undo.go:285: [INFO] [stacks] update wishlist: the undo copy will hold 2 named volume(s), 0.3 MiB
2026/09/24 14:58:08 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_data → wishlist_wishlist_data.pre-update-20260924T145808Z in 459ms
2026/09/24 14:58:09 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_uploads → wishlist_wishlist_uploads.pre-update-20260924T145808Z in 440ms
2026/09/24 14:59:40 update.go:1274: [INFO] [stacks] update wishlist: phase undoing
2026/09/24 14:59:40 undo.go:513: [WARN] [stacks] update wishlist: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health che
2026/09/24 14:59:42 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1
2026/09/24 14:59:57 undo.go:633: [INFO] [stacks] update wishlist: event app_update_undone
2026/09/24 14:59:57 undo.go:561: [INFO] [stacks] update wishlist: UNDONE in 18s — the previous version is running on the data from before the update (the app's health check passed)
@@ -0,0 +1,8 @@
BEFORE restore:
.felhom.yml 15:01:22: port: 8999
applied-meta/.felhom.yml 15:00:19: port: 3000
POST /backup/tier2/unit-restore -> 302
restore: {'op': 'tier2-unit-restore', 'stack': 'wishlist', 'ok': True, 'message': 'A(z) wishlist: 2 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás f
AFTER restore:
.felhom.yml 15:01:38: port: 3000
applied-meta/.felhom.yml 15:01:38: port: 3000
@@ -0,0 +1,14 @@
remove wishlist -> 200
no wishlist volumes
no leftovers
sync_interval: 15m
token: <redacted>
username: ""
hub:
0
drill=c8025093d23c live=c8025093d23c
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
0
gitea.dooplex.hu/admin/felhom-controller:0.270.0
@@ -0,0 +1,2 @@
POST floor 0.270.0 -> 303 /configuration?flash=floor_set
read back: ['0.270.0']
@@ -0,0 +1,11 @@
Thu Sep 24 17:03:45 CEST 2026
== N100 9201
LANG = "en_US.UTF-8"
gitea.dooplex.hu/admin/felhom-controller:0.270.0 Up 5 seconds (healthy)
== demo-hp 9201
LANG = "en_US.UTF-8"
gitea.dooplex.hu/admin/felhom-controller:0.270.0 Up 5 seconds (healthy)
== hub
2026/09/24 17:03:33 [INFO] Global controller-version floor set to "0.270.0" (declared MinAgent "0.131.0")
2026/09/24 17:03:35 [INFO] managed floor SERVED for demo-felhom: floor 0.270.0, agent requirement "0.131.0" from declared (golden 0.258.0)
2026/09/24 17:03:36 [INFO] managed floor SERVED for demo-hp: floor 0.270.0, agent requirement "0.131.0" from declared (golden 0.258.0)
@@ -0,0 +1,89 @@
# The restore test off, demo-hp 9201 repaired, agent v0.133.0 + hub v0.124.0, controller v0.270.0 (2026-09-24 evening)
Brief: "the restore test switched off on the demo boxes, the HP customer box repaired, the host agent fixed so a
restore test can never fill a box's disk again; then four controller leftovers". Architecture read: `03-host-agent.md`
§8/§10, `07`, `08` §6.2, `09` §3/§6.4, `runbooks/target-selection.md`.
## Not done, or changed
- **Agent v0.133.0 is released, NOT delivered.** It reaches a box only with the operator's signed `agent_update` job
(R-530). Until then the scheduled restore test stays OFF on both demo hosts.
- **Part A used `-1`, not the brief's `0`.** `restore_test_eval_interval_seconds: 0` means "the 6-hour default"
(`config.go` `RestoreTestEvalInterval`); only a negative value disables. The start-up line says "restore-test cadence
disabled" on both hosts. With the cadence off no evaluation runs, so there is no "nothing due" line to quote.
- **Part C rule 1 changed:** "archive size × 1.2 + 5 GiB" would NOT have prevented the incident. The size is the
UNCOMPRESSED one (vzdump log / PBS snapshot size).
- **Part C rule 4 needed a hub release (v0.124.0).** The existing `storage_fill_critical` fired only at 95 %, on data
only, per household per hour, and only at the next 15-minute report. Now: a thin pool is critical at 90 % of data
or metadata, one alarm per pool per 6 hours, and the agent asks for a report the moment a pool crosses 90 %.
- **Live case (b) refused and case (c) did not run.** On demo-hp, 9201's restore needs 30.3 GiB and the pool has 22.1
GiB free. So no full restore test fits under the 80 % limit. The forced clean-up failure needs a scratch guest,
which cannot be created there. Both are covered by unit tests only.
- **Part B found damage fsck could not see:** the Redis append-only files of docmost and romm were cut off by the full
pool. With the operator's yes, each Redis folder was copied aside and cut at the last complete write with
`redis-check-aof --fix` (2,943 B / 6,631 B dropped). Meanwhile the box had stopped both apps by itself (decision 28).
- **R-674 has no live proof:** R-679 now refuses the only case that reached that log line.
- **R-679 removed a "repair path"** that a test comment named: a same-version Update. Restart is the repair path.
**Interventions: 0** on product behaviour. The Redis repair was an operator-approved data repair. **One agent
release, one hub release, one controller release.**
## Claims in the brief that turned out wrong (or right)
1. *An eval interval of 0 disables only the schedule and leaves the on-demand test working* — **half wrong.** 0 is
the 6-hour default; negative disables. The on-demand path (`--selftest=restore-test`) does not read the cadence —
TRUE, and used live.
2. *`pct fsck` can check the `/var/lib/felhom` mount by `--device`* — **TRUE** (`--device mp0`).
3. *demo-hp has a second eligible storage for the restore test* — **WRONG.** `nvme-scratch` takes `rootdir`, but the
agent holds only the inherited `Datastore.Audit` there, not `Datastore.AllocateSpace`.
4. *R-673's lock came from the pool-full event* — **WRONG.** The 06:59 and 07:42 backups failed with "No space left on
device" on `local`, the host ROOT disk (R-684). The lock is from the 07:42 clean-up. The pool filled at 10:35.
5. *"Archive size × 1.2 + 5 GiB" is enough* — **WRONG.** 9201: a 6.9 GB file, a 22.6 GB restore.
6. *A pool nearly full is only a log line* — **partly wrong.** The hub DID mail `storage_fill_critical` at 100 %, but
late and at the generic bands.
## Part A — the scheduled restore test off (`A-restore-test-off.txt`)
Both hosts: only `backup.restore_test_eval_interval_seconds` changed (saved copy `/etc/felhom-agent/agent.json.pre-r672`,
verified equal apart from that key). Start-up: "backup: restore-test cadence disabled". Peti's box untouched: it gets
the fix only through a signed agent.
## Part B — 9201 repaired (`B1…B9`)
Pool 58.9 % data / 2.65 % metadata. rootfs = `vm-9201-disk-0`, `/var/lib/felhom` = `mp0` (`vm-9201-disk-1`). Stop 4.8 s.
`pct fsck`: rootfs replayed its journal; mp0 fixed 17 "deleted inode has zero dtime" and one orphan block; second
pass clean on both (rc 0). Start; the controller took the 0.269.1 floor by itself; both disks writable; 21 containers
up; hub `/hosts`: demo-hp ONLINE. Then the Redis repair above; docmost and romm started through the product (Start
lifted the box's hold), 0 restarts.
## Part C — agent v0.133.0 (tag `9bdb4da`, sha256 `3aa30345…e69b6`, verified by an anonymous download) + hub v0.124.0
Red-proofs (`redproofs/C-*`): no preflight → the 2026-09-24 restore issued again; file size used → "the compressed
file size was used"; no timer → "the leaked scratch was not destroyed by the timer"; sweep without the gate → "the
sweep ran while a backup held the gate"; no 90 % edge → "0 report requests, want 1"; skips dropped → "a space refusal
never reached the host report"; hub generic bands → a 91 % pool only warns, a metadata-full pool raises nothing; hub
old key → "a second pool filling in the same hour was silenced".
Live on demo-hp (on-demand self-test, cadence off, pool 58.99 % before and after, nothing created): (a) factor 10 →
"needs 215.5 GiB free, has 22.1 GiB", exit 4; (b) normal → "restoring 21.1 GiB (vzdump log: total bytes written)
needs 30.3 GiB free, has 22.1 GiB", exit 4.
R-673: the stale-lock sweep now runs every 10 minutes, under the one-heavy-operation gate.
## Part D — controller v0.270.0, floor 0.270.0 (`D/`)
R-679 (409 `already_current`, hu + en, live), R-681 (interrupted install reported and cleaned, live), R-669 (the unit
keeps the pinned health check; a restore resets the applied record; live), R-674 (unit only). Both demo boxes
arrived on 0.270.0 in ~12 s.
## Register
Before this session **344 rows / 685,662 B**; after **341 rows / 685,148 B**. Opened R-684; closed R-669, R-674,
R-679, R-681; R-672 and R-673 updated to "fixed in v0.133.0, awaiting delivery".
## Teardown — three layers
- **Machine:** 9202 back on the live catalog with its saved config, on controller 0.270.0 (the floor); the throwaway
apps (actualbudget, mealie, wishlist) removed through the product — no containers, volumes or markers left.
9201 (demo-hp) repaired, running, on 0.270.0.
- **Host:** demo-hp and demo-felhom: only the agent config key changed (saved copies beside them); the live-test
binary and its test config deleted from demo-hp `/tmp`; `pct list` 9201 + 9202, nothing created.
- **Hub:** v0.124.0 deployed; floor 0.270.0 (MinAgent 0.131.0); drill repo reset to the live catalog.
## To turn the restore test back on (after the signed agent arrives on a box)
On that host: `cp /etc/felhom-agent/agent.json.pre-r672 /etc/felhom-agent/agent.json && systemctl restart felhom-agent`.
The saved copy had no `restore_test_eval_interval_seconds` key (= the 6-hour default). Check the journal for "restore-test
scheduler starting". On demo-hp the test will then REFUSE 9201's restore for space, correctly, until the pool has room
(R-684 is about the separate backup storage).
@@ -0,0 +1,9 @@
# Red-proofs D3 — (1) the restore does not reset applied-meta; (2) the unit captures the stack dir's .felhom.yml (v0.269.1)
--- FAIL: TestR669_RestoreResetsTheAppliedRecord (0.00s)
r669_restore_meta_test.go:31: the next undo would judge with the failed step's probe:
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.009s
--- FAIL: TestR669_UnitCapturesThePinnedVersionsMeta (0.00s)
r669_applied_meta_test.go:43: the unit carries the failing step's probe, not the pinned version's:
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks (cached)
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.008s
@@ -0,0 +1,5 @@
# Red-proof D2 — the AT-THE-HEAD branch removed (v0.269.1)
--- FAIL: TestR674_HeadIsNotCalledOlder (0.00s)
r674_head_test.go:23: the head was called older than the ladder: "the installed version web=nextcloud:34.0.1-apache matches no update_ladder entry (2 entries) — an app older than the ladder has no record to climb; the catalog's current definition"
FAIL
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
@@ -0,0 +1,5 @@
# Red-proof D1 (controller) — the already-current refusal removed (v0.269.1)
--- FAIL: TestR679_CurrentAppIsRefused (0.00s)
r679_already_current_test.go:22: an app at the head was allowed to update
FAIL
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.015s
@@ -0,0 +1,5 @@
# Red-proof D5 — the interrupted-install line removed from stacks.html
--- FAIL: TestR681_PageSaysTheInstallWasInterrupted (0.10s)
r681_install_interrupted_page_test.go:24: the apps page is silent about an interrupted install (hu)
FAIL
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.217s
@@ -0,0 +1,5 @@
# Red-proof D4 — RecoverInterruptedInstalls does nothing (v0.269.1)
--- FAIL: TestR681_InterruptedInstallIsFinishedAndReported (0.00s)
r681_install_interrupted_test.go:33: an interrupted install was not reported: resolved=[] hooks=[]
FAIL
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.009s
+4
View File
@@ -382,4 +382,8 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-666** | **While support is informed, Remove offered to delete the data (decision 27).** v0.269.0: the dialog reads `keep_data_only` and offers only „remove the app, keep my data"; the API refuses data or backup deletion with 409 (hu + en); the no-whole-copy sentence is informal. Live on 9202 (a one-drive hold). | v0.269.0, 2026-09-24 | `git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-24/A2/10-held-no-copy-keep-data.*` |
| **R-667** | **A crash loop never reached the alarm, and nothing stopped it (decision 28).** v0.269.0 + hub v0.123.0: ≥ 6 restarts in 10 min (`RestartCount`, not the resettable `restarting_since`) or an OOM storm → the box stops the app, holds it (`unhealthy_stop`), tells household + operator (`app_stopped_unhealthy`); Start = one more try; a repeat in 24 h says support is informed. Live: gokapi trip 1 and 2; chaos rounds 1, 7, 8, 12. **Rule:** Docker's back-off caps a steady loop at ~1 restart/min, so a threshold must be below 10 per 10 min. | v0.269.0 / hub v0.123.0, 2026-09-24 | `git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-24/A3/`, `08` §6.2 |
| **R-668** | **The Tier-2 copy chose a registered path that no longer existed, on the app's own disk (P2).** v0.269.0: the same-disk check fails CLOSED. Live: the next copy went to the SSD. Residual: an older same-disk record counts until the next Tier-2 run replaces it. | v0.269.0, 2026-09-24 | `git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-24/A1/01-find-mirror.txt` |
| **R-669** | **After a failed update ended in a restore, the box kept the FAILED step's health check as the pinned version's, and the next undo held the app for nothing (P2).** v0.270.0: the recovery unit captures the pinned version's `.felhom.yml` (`applied-meta`), not the stack dir's file the sync may already have replaced; a restore makes the restored file the applied record. Live on 9202: the sync wrote the bad probe at 14:57:46, the unit captured at 14:57:55 kept the good one; a second-drive restore turned stack 8999 / applied 3000 into 3000 / 3000, the applied record rewritten at the restore. **Rule:** a copy of an app's definition carries the PINNED version's health check, never the catalog's newest. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/30-33*` |
| **R-674** | **The ladder log called a pin at the head "older than the ladder" (P3).** v0.270.0: it says "AT THE HEAD". Unit test + red-proof only — no product path reaches it since R-679 refuses a current app first. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/redproofs/D-r674.txt` |
| **R-679** | **An Update on an app already at the head ran the whole guarded update — dump, pull, restart (P2).** v0.270.0: `409 already_current` before anything moves, hu + en; a re-tested digest of a floating tag still updates. Live on 9202 (privatebin): both languages refused, no backup, no pull. A test comment had called a same-version Update "the repair path" — Restart is. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/10-r679-live.txt` |
| **R-681** | **An install cut off by a controller restart was lost silently (P2).** v0.270.0: an install marker before the compose-up, removed when the install ends; a marker at start → `compose down` (volumes kept), stale pin records cleared, `app_deploy_failed` with the reason, and the apps page says the install was interrupted (hu + en) until the next install; a finished install only loses the marker. Live on 9202: mealie killed 0 s into its pull → reported, page sentence, nothing left running; reinstall cleared the sentence; actualbudget (killed after it finished) left alone. **Rule:** a long-running act the customer started is journaled so a restart finishes or reports it. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/20-26*` |
-4
View File
@@ -806,19 +806,15 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** |
| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. | **READY — P2; owner: CC (catalog)** |
| **R-657** | **[P2-MEDIUM] "Remove the app, keep my data", then install it again: nextcloud never installs, and the box only says „unhealthy".** MEASURED 2026-09-23 night on 9202 (v0.267.0): nextcloud was removed through the product keeping its drive folder (the remove with data was refused — R-442's fail-closed guard, as on every drill on 9202 — and the product's keep-data remove taken). An hour later a fresh install of nextcloud on the same box: the template binds `${HDD_PATH}/appdata/nextcloud` to `/var/www/html/data`, the kept folder still holds `admin/`, `appdata_*`, `.ncdata` and a 145 MB `nextcloud.log`, and the image's installer loops **„Login is invalid because files already exist for this user — Retrying install..."**; `occ status` reads `installed: false`. The controller records the deploy as done and the app as `unhealthy`; nothing tells the household that their kept files are what blocks the new install, or what to do. **Why it matters:** keep-data is the choice the product OFFERS a household at remove time — and for nextcloud the kept data makes the app uninstallable. **Needs:** decide the product's promise for a reinstall over kept data, per app class (adopt the data? refuse with a sentence? offer to move it aside?); at minimum a deploy-time refusal or warning when the app's drive folder is not empty. Evidence: `audits/night-2026-09-23/chaos/00-nextcloud-reinstall-over-kept-data.txt`. | **READY — P2; owner: operator (the promise) / CC (the build)** |
| **R-669** | **[P2-MEDIUM] After a failed update ends in a restore, the box keeps the FAILED step's health check as the pinned version's — and the NEXT failed update's undo judges the correct old version with it and holds the app.** PROVEN LIVE 2026-09-24 on 9202 (v0.269.0): a failing step (probe port 8999) → hold → the second drive's whole restore cleared the hold, but `applied-meta/.felhom.yml` still carried port 8999 (written by `advancePinTo` at 10:39:21; no restore path rewrites it), and the stack's `.felhom.yml` came back from the unit, which the sync had already filled with the failing step's file. The next failed update's undo put the right version and the right data back and then called it unhealthy on port 8999 → **a false hold** (`not_started`). Data safe; the app stopped for nothing. **Fix direction:** a restore that recreates the definition also sets the applied record to the `.felhom.yml` of the RESTORED pin (the ladder step's `steps/<key>.felhom.yml` for those refs, or the catalog's when the catalog head equals the pin), never the unit's copy. Pre-existing since v0.263.2; not a v0.269.0 regression. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P2; owner: CC (controller)** |
| **R-670** | **[P3-LOW] Every undo (and every step-file verify) logs `[ERROR] .felhom.yml backup block rejected … docker-compose.yml unreadable`.** `LoadMetadata` validates the backup block against a compose file that the pre-update-meta directory (undo.go:528, since v0.263.0) and the scratch dir of `loadMetadataFile` (v0.269.0) never hold. The health check it feeds is unaffected; an operator reading ERROR lines after an undo is misled. Seen 10:24:41Z and 10:46:24Z on 9202. **Fix:** a probe-only loader that skips the backup-block validation, or copy the compose beside it. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P3; owner: CC (controller)** |
| **R-671** | **[P3-LOW] The undo copies kept by a hold survive the hold's clearing by a restore.** MEASURED 2026-09-24 on 9202: three `nextcloud_*.pre-update-20260924T103924Z` volumes (~0.9 GiB) were still present after the whole restore cleared that hold at 10:41:52Z, and a second set joined them 6 minutes later. Nothing names them on a page; on a small disk they are the difference between the next update's copy fitting or not. **Fix direction:** the restore that clears an update hold removes that hold's undo copies (they describe the state the restore just replaced), logged. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` (volume list) | **READY — P3; owner: CC (controller)** |
| **R-672** | **[P1-HIGH] The agent's SCHEDULED restore-test filled the production thin pool and turned a customer guest's disks read-only.** FOUND 2026-09-24 on demo-hp (agent v0.132.0): at 10:29 CEST the restore-test restored 9201's 22 GiB archive as scratch guest 990000 into `local-lvm` — the SAME pool as 9201 — with no free-space check. The pool reached 100 % at 10:35:24 (`out_of_data_space`, `error_if_no_space`); the scratch failed to start; its teardown failed (`lvremove … contains a filesystem in use`) and was "left for Recover" — **which runs only at agent start**, so the pool stayed full for 2.5 h. The agent logged *„a full pool corrupts every guest on it"* every few seconds and took no action. **Consequence, measured:** 9201's controller got `no space left on device` from 10:38, its log stops at 10:39:57, and at 13:12 both its rootfs and `/var/lib/felhom` were READ-ONLY (errors=remount-ro). **Intervention:** an agent restart ran Recover ("destroyed leaked restore-test scratch guest", pool 100 % → 58.9 %). **Still open:** 9201 needs a stop + fsck + start (operator — the session's permission check refused host-level guest operations). **Fix direction:** the restore-test refuses to start unless the target storage has the archive's size plus a margin free and never targets the pool of the guest it tests; a failed teardown retries on a timer, not only at start; the pool-fill WARN pages the operator. `audits/night-2026-09-24/C-02…C-07` **— 2026-09-24 (evening): FIXED IN RELEASE, NOT YET DELIVERED.** Agent **v0.133.0** (tag `9bdb4da`, package verified by download): space preflight before anything is created (UNCOMPRESSED size × 1.2 + 5 GiB, thin metadata, off the tested guest's pool when another storage is eligible — demo-hp has none, `nvme-scratch` carries no grant — unknown refuses, reported as a non-pass `skipped`); failed teardown retried every 10 min; thin pool ≥ 90 % requests an immediate report. Hub **v0.124.0** (live): thin pool critical at 90 % of data or metadata, one alarm per pool per 6 h. **The brief's "archive × 1.2 + 5 GiB" would NOT have prevented this incident** (6.9 GB file → 22.6 GB restore). Six red-proofs; live on demo-hp: both margins REFUSE (21.1 GiB restored needs 30.3 GiB, 22.1 GiB free), pool unchanged. The hub HAD alarmed on the day (`storage_fill_critical` at 100 %, operator mail) — late and at the generic bands. **9201 repaired** (stop, `pct fsck` both disks: 17 half-deleted inodes + an orphan block fixed on mp0; second pass clean; start; floor 0.269.1 taken by itself; ONLINE); two Redis AOF tails the pool cut (2,943 B / 6,631 B) truncated with `redis-check-aof --fix` after copies were kept as volumes — the box had meanwhile stopped docmost and romm itself (decision 28, live). **Scheduled restore-test OFF on both demo hosts** (operator ruling) until v0.133.0 is delivered. `audits/r672-2026-09-24/` | **FIXED IN v0.133.0 — awaiting the operator's signed delivery; restore test OFF on demo hosts; owner: operator (sign) / CC** |
| **R-673** | **[P2-MEDIUM] 9201's whole-box backups failed all morning on a stale `snapshot-delete` lock.** demo-hp 2026-09-24: vzdump of 9201 failed at 06:59, 07:17, 07:42 and 08:22 CEST (`CT is locked (snapshot-delete)`), leaving `snap_vm-9201-disk-1_vzdump` (07:42) behind, before R-672's pool fill. The agent's stale-lock scanner ran every 30 s and did not clear it; the lock was gone after the 13:09 agent restart. Cause not read — the session's permission check refused reading the task logs. `audits/night-2026-09-24/C-01-demo-hp-9201-vzdump-errors.txt` **— 2026-09-24 (evening): CAUSE READ, FIX RELEASED.** Not the pool fill (that came at 10:35). The 06:59 and 07:42 backups failed writing the archive to `local` — the host ROOT disk: `zstd: error 70 : … No space left on device` (R-684); the 07:42 failure's cleanup left `snap_vm-9201-disk-1_vzdump` and the `snapshot-delete` lock, and the 08:22 run hit the lock. The agent's own stale-lock recovery cleared it at 13:10, one minute after the agent restart — because it ran ONLY at start. v0.133.0 runs it every 10 min under the one-heavy-operation gate (red-proofed). `audits/r672-2026-09-24/C5-r673-vzdump-logs.txt` | **FIXED IN v0.133.0 — awaiting delivery; root cause R-684; owner: CC** |
| **R-674** | **[P3-LOW] The ladder log says an app's pin „matches no update_ladder entry … older than the ladder" when the pin EQUALS the head.** Seen on 9202 2026-09-24 for nextcloud at the head. Misleading to an operator reading why nothing climbed. **Fix:** say "at the head" when the pin equals the newest `to`. | **READY — P3; owner: CC (controller)** |
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** |
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` | **OPEN — P3; owner: CC (watch)** |
| **R-677** | **[P3-LOW] For a floating tag re-tested at a new digest, the Behind badge's age reads the TAG's catalog date („1 napja"), not when the new digest was tested (minutes).** Seen 2026-09-24 on 9202 (Part B). Harmless but confusing. **Fix:** for a digest-only move, age from the ladder entry's `tested_at`. `audits/night-2026-09-24/B/10-floating-tag.json` | **READY — P3; owner: CC (controller)** |
| **R-678** | **[P3-LOW] After an update step ends `done`, the app's „steps left" and its badge stay STALE until the next scan.** MEASURED 2026-09-24 on 9202 (Part C, first caller run): wishlist read `ladder_steps_left` 1 after its only step, navidrome 2 after both of its steps — for ~50 s, through six presses. A person sees „Frissítés elérhető" for an app that just updated; an automatic caller re-presses. **Fix:** the update's finish refreshes the app's catalog/ladder fields (the same read the scan does). `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P3; owner: CC (controller)** |
| **R-679** | **[P2-MEDIUM] An Update pressed on an app that is already current runs the whole guarded update — backup, pull, restart — and reports `done`.** MEASURED 2026-09-24 on 9202: navidrome at the head was pressed four times by the stale-read caller (R-678); each press made a safety dump and restarted the app (8.5 s of downtime each) and changed nothing. There is no `current` refusal (`UpdatePreflight`, `update.go:366`; `UpdateOrderCurrent` falls through). Harmless by hand, costly for the automatic leg (`09` §6.4.2). **Fix:** the preflight refuses with reason `current` when the pin equals the catalog head and no newer tested digest exists. `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P2; owner: CC (controller)** |
| **R-680** | **[P2-MEDIUM] The box does not remember a failed update step — after an undo it offers the same step again.** MEASURED 2026-09-24 on 9202: vikunja's failing step was undone (104 s) and its badge went straight back to „Frissítés elérhető"; only the test caller's own memory stopped a re-press. Decision 15 („a failed step is never pressed again") therefore holds for a person only by their judgement and not at all for the automatic leg. **Fix (part 7 (b), `09` §6.4.2):** record the failed `to` per app; the leg skips it until the catalog's ladder for that app changes; the page says the step was tried and put back. `audits/night-2026-09-24/C/31-night.json` | **READY — P2; owner: CC (controller, with part 7)** |
| **R-681** | **[P2-MEDIUM] An install interrupted by a controller restart is lost SILENTLY — no event, no page sentence, and its half-written files stay.** MEASURED 2026-09-24 on 9202 (Part E round 2, v0.269.1): n8n's install pressed at 11:48:42Z ("Deploying stack n8n … checking 1 images"); `systemctl restart docker` 20 s later took the controller down with it. After the restart the box logged NOTHING about n8n, the app read `not_deployed`, and the household's page showed it as never installed — while `app.yaml` (with its generated `N8N_ENCRYPTION_KEY`), `applied-compose.yml` and `applied-meta/` stayed in the stack dir. A fresh install through the product afterwards worked (211 s) and was not blocked by the leftovers. Compare the update path, which journals and RESUMES after the same accident (round 3/4). **Fix direction:** journal the install like the update; on boot either resume it or say on the page (and in an event) that it was interrupted, and clean the half-written files. `audits/night-2026-09-24/E/round-02*` | **READY — P2; owner: CC (controller)** |
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
| **R-684** | **[P2-MEDIUM] demo-hp's whole-box backups cannot fit on its backup storage — 9201 has had no new whole-box backup since 2026-09-23.** `local` is the host ROOT disk (40 GB, 84 % used, ~4 GB free); it holds three 9201 archives of 5.8–6.9 GB (retention 3) and a new one needs ~7 GB: the 2026-09-24 runs failed `No space left on device` (06:59, 07:42), leaving the stale lock of R-673. Every night now fails the same way until space is made or the retention/target changes. A Tier-0 box, but the same shape exists on any appliance whose local backup target is its root disk. **Needs a decision** (fewer kept archives, another target, or a larger disk). `audits/r672-2026-09-24/C5-r673-vzdump-logs.txt` | **OPEN — P2; owner: operator (the decision) / CC** |