R-672 session: audit README (not done first, wrong claims), STATUS, capability map, register 344->341, topic REPORT
gates / gates (push) Successful in 28s
gates / gates (push) Successful in 28s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.270.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.270.0 Up 20 seconds (healthy)
|
||||
@@ -0,0 +1,9 @@
|
||||
privatebin: installed= {"privatebin": {"ref": "privatebin/pdo:2.0.6", "digest": "sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5", "at": "2026-09-22T19:44:56Z"}} catalog= {'privatebin': 'privatebin/pdo:2.0.6'} badge= [{'title': 'This app is running the newest version available.', 'text': 'Up to date'}]
|
||||
hu Update -> 409 {"ok": false, "data": {"reason": "already_current"}, "error": "Ez az alkalmazás már a legfrissebb elérhető változatot futtatja — nincs mit frissíteni."}
|
||||
en Update -> 409 {"ok": false, "data": {"reason": "already_current"}, "error": "This app is already running the newest version available — there is nothing to update."}
|
||||
after: updating= False phase= None state= running
|
||||
2026/09/24 14:51:18 router.go:591: [INFO] [api] update requested for stack: privatebin
|
||||
2026/09/24 14:51:18 update.go:359: [ERROR] [stacks] update privatebin REFUSED (already_current): installed equals the catalog head on every service and no newer tested digest (installed=map[privatebin:{privatebin/pdo:2.0.6 sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5 2026-09-22T19:44:56Z}] catalog=map[privatebin:privatebin/pdo:2.0.6])
|
||||
2026/09/24 14:51:18 router.go:591: [INFO] [api] update requested for stack: privatebin
|
||||
2026/09/24 14:51:18 update.go:359: [ERROR] [stacks] update privatebin REFUSED (already_current): installed equals the catalog head on every service and no newer tested digest (installed=map[privatebin:{privatebin/pdo:2.0.6 sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5 2026-09-22T19:44:56Z}] catalog=map[privatebin:privatebin/pdo:2.0.6])
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
before: deployed= False
|
||||
no actualbudget image cached
|
||||
|
||||
16:51:41 deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
-rw-r--r-- 1 root root 21 Sep 24 14:51 .felhom-install-pending
|
||||
-rw------- 1 root root 251 Sep 24 14:51 app.yaml
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
16:51:52 killed pid 157525 rc=0
|
||||
@@ -0,0 +1,19 @@
|
||||
stack: deployed= True state= running install_interrupted= None deploy_error= None
|
||||
2026/09/24 14:51:53 manager.go:1554: [INFO] [stacks] actualbudget actualbudget/actual-server:26.9.0@sha256:552beab3dec8c93d46b8b9245612d63c3f123b8a45063a474f53e229b17621d3 running Up 3 seconds (health: starting)
|
||||
2026/09/24 14:51:54 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
|
||||
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
|
||||
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
|
||||
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
|
||||
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
|
||||
2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml
|
||||
2026/09/24 14:51:54 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for actualbudget → /mnt/sys_drive/felhom-data/backups/primary/actualbudget (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0)
|
||||
--- files
|
||||
-rw------- 1 root root 523 Sep 24 14:51 app.yaml
|
||||
-rw-r--r-- 1 root root 1484 Sep 24 14:51 applied-compose.yml
|
||||
drwxr-xr-x 2 root root 4096 Sep 24 14:51 applied-meta
|
||||
--- containers
|
||||
actualbudget Up 13 seconds (healthy)
|
||||
(end)
|
||||
|
||||
hu page: None
|
||||
en page: None
|
||||
@@ -0,0 +1 @@
|
||||
16:53:05 deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
@@ -0,0 +1,2 @@
|
||||
16:53:05 /opt/docker/stacks/mealie/.felhom-install-pending
|
||||
16:53:05 killed pid 158798 rc=0
|
||||
@@ -0,0 +1,17 @@
|
||||
stack: deployed= False state= not_deployed install_interrupted= True deploy_error= interrupted by a controller restart before it finished — install it again
|
||||
2026/09/24 14:51:54 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
|
||||
2026/09/24 14:53:05 router.go:431: [INFO] [api] Deploy requested for stack: mealie
|
||||
2026/09/24 14:53:05 [INFO] [stacks] SaveAppConfig: saved config for mealie
|
||||
2026/09/24 14:53:05 deploy.go:383: [INFO] [stacks] Deploying stack mealie with 2 env vars: [DOMAIN, SUBDOMAIN]
|
||||
2026/09/24 14:53:05 manager.go:1590: [INFO] [stacks] Deploying stack mealie — checking 1 images...
|
||||
2026/09/24 14:53:06 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
|
||||
2026/09/24 14:53:07 install_interrupted.go:70: [WARN] [stacks] install mealie was INTERRUPTED by a controller restart — removing what it started (volumes kept) and reporting it (R-681)
|
||||
2026/09/24 14:53:07 [WARN] [stacks] 1 install(s) interrupted by the restart were resolved and reported: [mealie]
|
||||
--- files
|
||||
-rw-r--r-- 1 root root 21 Sep 24 14:53 .felhom-install-interrupted
|
||||
-rw------- 1 root root 255 Sep 24 14:53 app.yaml
|
||||
--- containers
|
||||
(end)
|
||||
|
||||
hu page: A telepítés egy újraindítás miatt félbemaradt, és a doboz eltávolította, amit elkezdett. Nyomd meg újra a Telepítés gombot.
|
||||
en page: The installation was interrupted by a restart, and the box removed what it had started. Press Install again.
|
||||
@@ -0,0 +1,12 @@
|
||||
16:53:51 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
16:55:11 [1] deployed, controller state=running, pinned={'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'}
|
||||
reinstall ok= True deployed= True install_interrupted= None
|
||||
no install markers
|
||||
|
||||
16:55:16 [X] stop -> 200 {'ok': True, 'message': 'Stack mealie stop completed'}
|
||||
16:55:48 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'mealie', 'volumes_removed': ['mealie_mealie_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalm
|
||||
16:55:56 [X] after remove: deployed=False leftovers='/opt/docker/stacks/mealie'
|
||||
remove -> 200
|
||||
no mealie volumes
|
||||
no leftovers
|
||||
|
||||
@@ -0,0 +1,8 @@
|
||||
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
|
||||
sync_interval: 15m
|
||||
token: <redacted>
|
||||
username: "admin"
|
||||
hub:
|
||||
update:
|
||||
health_timeout: 90s
|
||||
|
||||
@@ -0,0 +1,84 @@
|
||||
{
|
||||
"1-installed": "stack: \napplied: \nunit: ()",
|
||||
"2-before-update": "stack: \napplied: \nunit: ()",
|
||||
"2-update": {
|
||||
"accepted": true,
|
||||
"http": "202",
|
||||
"phases": [
|
||||
{
|
||||
"t": 0.0,
|
||||
"phase": "backing-up",
|
||||
"label": "Biztonsági mentés készül a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 2.1,
|
||||
"phase": "safety-dump",
|
||||
"label": "Adatbázis pillanatkép…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 3.1,
|
||||
"phase": "pulling",
|
||||
"label": "Új verzió letöltése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 4.1,
|
||||
"phase": "copying",
|
||||
"label": "Az adatok másolása a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 16.4,
|
||||
"phase": "verifying",
|
||||
"label": "Működés ellenőrzése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 106.6,
|
||||
"phase": "undoing",
|
||||
"label": "Visszaállítás az előző változatra…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 125.0,
|
||||
"phase": "undone",
|
||||
"label": "Visszaállítva az előző változatra",
|
||||
"updating": false,
|
||||
"error": null,
|
||||
"hold": null
|
||||
}
|
||||
],
|
||||
"duration_s": 125.1,
|
||||
"final_phase": "undone",
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"state": "running"
|
||||
},
|
||||
"2-after-update": "stack: \napplied: \nunit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)",
|
||||
"CHECK_A_unit_keeps_pinned_probe": false,
|
||||
"3-before-restore": "stack: \napplied: \nunit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)",
|
||||
"3-restore": {
|
||||
"op": "tier2-unit-restore",
|
||||
"stack": "wishlist",
|
||||
"ok": true,
|
||||
"message": "A(z) wishlist: 2 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-24 16:57).",
|
||||
"finished_at": "2026-09-24T15:00:33.459543237Z"
|
||||
},
|
||||
"3-after-restore": "stack: \napplied: \nunit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)",
|
||||
"CHECK_B_restore_resets_applied": false,
|
||||
"log": "2026/09/24 14:57:56 pin.go:373: [INFO] [stacks] update wishlist: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/wishlist/docker-compose.yml (wishlist=ghcr.io/cmintey/wishlist:latest)\n2026/09/24 14:58:08 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_data → wishlist_wishlist_data.pre-update-20260924T145808Z in 459ms\n2026/09/24 14:58:09 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_uploads → wishlist_wishlist_uploads.pre-update-20260924T145808Z in 440ms\n2026/09/24 14:59:40 update.go:1274: [INFO] [stacks] update wishlist: phase undoing\n2026/09/24 14:59:40 undo.go:513: [WARN] [stacks] update wishlist: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health check failing))\n2026/09/24 14:59:42 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1\n2026/09/24 14:59:57 undo.go:633: [INFO] [stacks] update wishlist: event app_update_undone\n2026/09/24 14:59:57 undo.go:561: [INFO] [stacks] update wishlist: UNDONE in 18s — the previous version is running on the data from before the update (the app's health check passed)\n2026/09/24 15:00:19 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1\n2026/09/24 15:00:33 restore_unit.go:464: [INFO] [backup] Restore-from-unit completed: wishlist — 2 volume(s) of 2 listed, 0 database(s) of 0 listed\n"
|
||||
}
|
||||
@@ -0,0 +1,20 @@
|
||||
16:57:21 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
16:57:41 [1] deployed, controller state=running, pinned={'wishlist': 'ghcr.io/cmintey/wishlist:v0.67.1'}
|
||||
16:57:45 [1] stack: | applied: | unit: ()
|
||||
16:57:46 drill: wishlist: failing step ghcr.io/cmintey/wishlist:v0.67.1 -> ghcr.io/cmintey/wishlist:latest + probe 8999
|
||||
16:57:53 [2] stack: | applied: | unit: ()
|
||||
16:57:53 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
16:57:53 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None
|
||||
16:57:55 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
|
||||
16:57:56 + 3.1s phase=pulling label=Új verzió letöltése… err=None hold=None
|
||||
16:57:57 + 4.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None
|
||||
16:58:10 + 16.4s phase=verifying label=Működés ellenőrzése… err=None hold=None
|
||||
16:59:40 + 106.6s phase=undoing label=Visszaállítás az előző változatra… err=None hold=None
|
||||
16:59:58 + 125.0s phase=undone label=Visszaállítva az előző változatra err=None hold=None
|
||||
17:00:02 [2] after: stack: | applied: | unit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)
|
||||
17:00:03 drill: wishlist image back to v0.67.1 (probe left at 8999)
|
||||
17:00:18 [3] stack: | applied: | unit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)
|
||||
17:00:18 [3] POST /backup/tier2/unit-restore -> 302
|
||||
17:00:34 [3] restore: {'op': 'tier2-unit-restore', 'stack': 'wishlist', 'ok': True, 'message': 'A(z) wishlist: 2 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-24 16:57).', 'finished_at'
|
||||
17:00:37 [3] after: stack: | applied: | unit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)
|
||||
17:00:40 CHECK A (unit keeps the pinned probe): False CHECK B (restore resets the applied record): False
|
||||
@@ -0,0 +1,9 @@
|
||||
== .felhom.yml (2026-09-24 15:00:19)
|
||||
- type: http
|
||||
port: 3000
|
||||
== applied-meta/.felhom.yml (2026-09-24 15:00:19)
|
||||
- type: http
|
||||
port: 3000
|
||||
== /mnt/sys_drive/felhom-data/backups/primary/wishlist/compose/.felhom.yml (2026-09-24 14:57:55)
|
||||
- type: http
|
||||
port: 3000
|
||||
@@ -0,0 +1,16 @@
|
||||
2026/09/24 14:57:22 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1
|
||||
2026/09/24 14:57:46 sync.go:416: [INFO] [sync] Updated wishlist/.felhom.yml
|
||||
2026/09/24 14:57:53 update.go:1274: [INFO] [stacks] update wishlist: phase backing-up
|
||||
2026/09/24 14:57:53 update_guard.go:360: [INFO] [backup] update pre-backup for wishlist: starting (DB dump → volume dump → unit capture → Tier 2)
|
||||
2026/09/24 14:57:55 update_guard.go:410: [INFO] [backup] update pre-backup for wishlist: volume dump OK
|
||||
2026/09/24 14:57:55 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for wishlist → /mnt/sys_drive/felhom-data/backups/primary/wishlist (images=1, secrets-referenced=0, data_keys=0, port
|
||||
2026/09/24 14:57:55 update_guard.go:416: [INFO] [backup] update pre-backup for wishlist: recovery unit captured (0 database dump(s))
|
||||
2026/09/24 14:57:55 update_guard.go:425: [INFO] [backup] update pre-backup for wishlist: complete in 2.038s
|
||||
2026/09/24 14:57:56 undo.go:285: [INFO] [stacks] update wishlist: the undo copy will hold 2 named volume(s), 0.3 MiB
|
||||
2026/09/24 14:58:08 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_data → wishlist_wishlist_data.pre-update-20260924T145808Z in 459ms
|
||||
2026/09/24 14:58:09 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_uploads → wishlist_wishlist_uploads.pre-update-20260924T145808Z in 440ms
|
||||
2026/09/24 14:59:40 update.go:1274: [INFO] [stacks] update wishlist: phase undoing
|
||||
2026/09/24 14:59:40 undo.go:513: [WARN] [stacks] update wishlist: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health che
|
||||
2026/09/24 14:59:42 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1
|
||||
2026/09/24 14:59:57 undo.go:633: [INFO] [stacks] update wishlist: event app_update_undone
|
||||
2026/09/24 14:59:57 undo.go:561: [INFO] [stacks] update wishlist: UNDONE in 18s — the previous version is running on the data from before the update (the app's health check passed)
|
||||
@@ -0,0 +1,8 @@
|
||||
BEFORE restore:
|
||||
.felhom.yml 15:01:22: port: 8999
|
||||
applied-meta/.felhom.yml 15:00:19: port: 3000
|
||||
POST /backup/tier2/unit-restore -> 302
|
||||
restore: {'op': 'tier2-unit-restore', 'stack': 'wishlist', 'ok': True, 'message': 'A(z) wishlist: 2 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás f
|
||||
AFTER restore:
|
||||
.felhom.yml 15:01:38: port: 3000
|
||||
applied-meta/.felhom.yml 15:01:38: port: 3000
|
||||
@@ -0,0 +1,14 @@
|
||||
remove wishlist -> 200
|
||||
no wishlist volumes
|
||||
no leftovers
|
||||
|
||||
sync_interval: 15m
|
||||
token: <redacted>
|
||||
username: ""
|
||||
hub:
|
||||
0
|
||||
|
||||
drill=c8025093d23c live=c8025093d23c
|
||||
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
|
||||
0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.270.0
|
||||
@@ -0,0 +1,2 @@
|
||||
POST floor 0.270.0 -> 303 /configuration?flash=floor_set
|
||||
read back: ['0.270.0']
|
||||
@@ -0,0 +1,11 @@
|
||||
Thu Sep 24 17:03:45 CEST 2026
|
||||
== N100 9201
|
||||
LANG = "en_US.UTF-8"
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.270.0 Up 5 seconds (healthy)
|
||||
== demo-hp 9201
|
||||
LANG = "en_US.UTF-8"
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.270.0 Up 5 seconds (healthy)
|
||||
== hub
|
||||
2026/09/24 17:03:33 [INFO] Global controller-version floor set to "0.270.0" (declared MinAgent "0.131.0")
|
||||
2026/09/24 17:03:35 [INFO] managed floor SERVED for demo-felhom: floor 0.270.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
||||
2026/09/24 17:03:36 [INFO] managed floor SERVED for demo-hp: floor 0.270.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
||||
@@ -0,0 +1,89 @@
|
||||
# The restore test off, demo-hp 9201 repaired, agent v0.133.0 + hub v0.124.0, controller v0.270.0 (2026-09-24 evening)
|
||||
|
||||
Brief: "the restore test switched off on the demo boxes, the HP customer box repaired, the host agent fixed so a
|
||||
restore test can never fill a box's disk again; then four controller leftovers". Architecture read: `03-host-agent.md`
|
||||
§8/§10, `07`, `08` §6.2, `09` §3/§6.4, `runbooks/target-selection.md`.
|
||||
|
||||
## Not done, or changed
|
||||
|
||||
- **Agent v0.133.0 is released, NOT delivered.** It reaches a box only with the operator's signed `agent_update` job
|
||||
(R-530). Until then the scheduled restore test stays OFF on both demo hosts.
|
||||
- **Part A used `-1`, not the brief's `0`.** `restore_test_eval_interval_seconds: 0` means "the 6-hour default"
|
||||
(`config.go` `RestoreTestEvalInterval`); only a negative value disables. The start-up line says "restore-test cadence
|
||||
disabled" on both hosts. With the cadence off no evaluation runs, so there is no "nothing due" line to quote.
|
||||
- **Part C rule 1 changed:** "archive size × 1.2 + 5 GiB" would NOT have prevented the incident. The size is the
|
||||
UNCOMPRESSED one (vzdump log / PBS snapshot size).
|
||||
- **Part C rule 4 needed a hub release (v0.124.0).** The existing `storage_fill_critical` fired only at 95 %, on data
|
||||
only, per household per hour, and only at the next 15-minute report. Now: a thin pool is critical at 90 % of data
|
||||
or metadata, one alarm per pool per 6 hours, and the agent asks for a report the moment a pool crosses 90 %.
|
||||
- **Live case (b) refused and case (c) did not run.** On demo-hp, 9201's restore needs 30.3 GiB and the pool has 22.1
|
||||
GiB free. So no full restore test fits under the 80 % limit. The forced clean-up failure needs a scratch guest,
|
||||
which cannot be created there. Both are covered by unit tests only.
|
||||
- **Part B found damage fsck could not see:** the Redis append-only files of docmost and romm were cut off by the full
|
||||
pool. With the operator's yes, each Redis folder was copied aside and cut at the last complete write with
|
||||
`redis-check-aof --fix` (2,943 B / 6,631 B dropped). Meanwhile the box had stopped both apps by itself (decision 28).
|
||||
- **R-674 has no live proof:** R-679 now refuses the only case that reached that log line.
|
||||
- **R-679 removed a "repair path"** that a test comment named: a same-version Update. Restart is the repair path.
|
||||
|
||||
**Interventions: 0** on product behaviour. The Redis repair was an operator-approved data repair. **One agent
|
||||
release, one hub release, one controller release.**
|
||||
|
||||
## Claims in the brief that turned out wrong (or right)
|
||||
|
||||
1. *An eval interval of 0 disables only the schedule and leaves the on-demand test working* — **half wrong.** 0 is
|
||||
the 6-hour default; negative disables. The on-demand path (`--selftest=restore-test`) does not read the cadence —
|
||||
TRUE, and used live.
|
||||
2. *`pct fsck` can check the `/var/lib/felhom` mount by `--device`* — **TRUE** (`--device mp0`).
|
||||
3. *demo-hp has a second eligible storage for the restore test* — **WRONG.** `nvme-scratch` takes `rootdir`, but the
|
||||
agent holds only the inherited `Datastore.Audit` there, not `Datastore.AllocateSpace`.
|
||||
4. *R-673's lock came from the pool-full event* — **WRONG.** The 06:59 and 07:42 backups failed with "No space left on
|
||||
device" on `local`, the host ROOT disk (R-684). The lock is from the 07:42 clean-up. The pool filled at 10:35.
|
||||
5. *"Archive size × 1.2 + 5 GiB" is enough* — **WRONG.** 9201: a 6.9 GB file, a 22.6 GB restore.
|
||||
6. *A pool nearly full is only a log line* — **partly wrong.** The hub DID mail `storage_fill_critical` at 100 %, but
|
||||
late and at the generic bands.
|
||||
|
||||
## Part A — the scheduled restore test off (`A-restore-test-off.txt`)
|
||||
Both hosts: only `backup.restore_test_eval_interval_seconds` changed (saved copy `/etc/felhom-agent/agent.json.pre-r672`,
|
||||
verified equal apart from that key). Start-up: "backup: restore-test cadence disabled". Peti's box untouched: it gets
|
||||
the fix only through a signed agent.
|
||||
|
||||
## Part B — 9201 repaired (`B1…B9`)
|
||||
Pool 58.9 % data / 2.65 % metadata. rootfs = `vm-9201-disk-0`, `/var/lib/felhom` = `mp0` (`vm-9201-disk-1`). Stop 4.8 s.
|
||||
`pct fsck`: rootfs replayed its journal; mp0 fixed 17 "deleted inode has zero dtime" and one orphan block; second
|
||||
pass clean on both (rc 0). Start; the controller took the 0.269.1 floor by itself; both disks writable; 21 containers
|
||||
up; hub `/hosts`: demo-hp ONLINE. Then the Redis repair above; docmost and romm started through the product (Start
|
||||
lifted the box's hold), 0 restarts.
|
||||
|
||||
## Part C — agent v0.133.0 (tag `9bdb4da`, sha256 `3aa30345…e69b6`, verified by an anonymous download) + hub v0.124.0
|
||||
Red-proofs (`redproofs/C-*`): no preflight → the 2026-09-24 restore issued again; file size used → "the compressed
|
||||
file size was used"; no timer → "the leaked scratch was not destroyed by the timer"; sweep without the gate → "the
|
||||
sweep ran while a backup held the gate"; no 90 % edge → "0 report requests, want 1"; skips dropped → "a space refusal
|
||||
never reached the host report"; hub generic bands → a 91 % pool only warns, a metadata-full pool raises nothing; hub
|
||||
old key → "a second pool filling in the same hour was silenced".
|
||||
Live on demo-hp (on-demand self-test, cadence off, pool 58.99 % before and after, nothing created): (a) factor 10 →
|
||||
"needs 215.5 GiB free, has 22.1 GiB", exit 4; (b) normal → "restoring 21.1 GiB (vzdump log: total bytes written)
|
||||
needs 30.3 GiB free, has 22.1 GiB", exit 4.
|
||||
R-673: the stale-lock sweep now runs every 10 minutes, under the one-heavy-operation gate.
|
||||
|
||||
## Part D — controller v0.270.0, floor 0.270.0 (`D/`)
|
||||
R-679 (409 `already_current`, hu + en, live), R-681 (interrupted install reported and cleaned, live), R-669 (the unit
|
||||
keeps the pinned health check; a restore resets the applied record; live), R-674 (unit only). Both demo boxes
|
||||
arrived on 0.270.0 in ~12 s.
|
||||
|
||||
## Register
|
||||
Before this session **344 rows / 685,662 B**; after **341 rows / 685,148 B**. Opened R-684; closed R-669, R-674,
|
||||
R-679, R-681; R-672 and R-673 updated to "fixed in v0.133.0, awaiting delivery".
|
||||
|
||||
## Teardown — three layers
|
||||
- **Machine:** 9202 back on the live catalog with its saved config, on controller 0.270.0 (the floor); the throwaway
|
||||
apps (actualbudget, mealie, wishlist) removed through the product — no containers, volumes or markers left.
|
||||
9201 (demo-hp) repaired, running, on 0.270.0.
|
||||
- **Host:** demo-hp and demo-felhom: only the agent config key changed (saved copies beside them); the live-test
|
||||
binary and its test config deleted from demo-hp `/tmp`; `pct list` 9201 + 9202, nothing created.
|
||||
- **Hub:** v0.124.0 deployed; floor 0.270.0 (MinAgent 0.131.0); drill repo reset to the live catalog.
|
||||
|
||||
## To turn the restore test back on (after the signed agent arrives on a box)
|
||||
On that host: `cp /etc/felhom-agent/agent.json.pre-r672 /etc/felhom-agent/agent.json && systemctl restart felhom-agent`.
|
||||
The saved copy had no `restore_test_eval_interval_seconds` key (= the 6-hour default). Check the journal for "restore-test
|
||||
scheduler starting". On demo-hp the test will then REFUSE 9201's restore for space, correctly, until the pool has room
|
||||
(R-684 is about the separate backup storage).
|
||||
@@ -0,0 +1,9 @@
|
||||
# Red-proofs D3 — (1) the restore does not reset applied-meta; (2) the unit captures the stack dir's .felhom.yml (v0.269.1)
|
||||
--- FAIL: TestR669_RestoreResetsTheAppliedRecord (0.00s)
|
||||
r669_restore_meta_test.go:31: the next undo would judge with the failed step's probe:
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.009s
|
||||
--- FAIL: TestR669_UnitCapturesThePinnedVersionsMeta (0.00s)
|
||||
r669_applied_meta_test.go:43: the unit carries the failing step's probe, not the pinned version's:
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks (cached)
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.008s
|
||||
@@ -0,0 +1,5 @@
|
||||
# Red-proof D2 — the AT-THE-HEAD branch removed (v0.269.1)
|
||||
--- FAIL: TestR674_HeadIsNotCalledOlder (0.00s)
|
||||
r674_head_test.go:23: the head was called older than the ladder: "the installed version web=nextcloud:34.0.1-apache matches no update_ladder entry (2 entries) — an app older than the ladder has no record to climb; the catalog's current definition"
|
||||
FAIL
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
@@ -0,0 +1,5 @@
|
||||
# Red-proof D1 (controller) — the already-current refusal removed (v0.269.1)
|
||||
--- FAIL: TestR679_CurrentAppIsRefused (0.00s)
|
||||
r679_already_current_test.go:22: an app at the head was allowed to update
|
||||
FAIL
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.015s
|
||||
@@ -0,0 +1,5 @@
|
||||
# Red-proof D5 — the interrupted-install line removed from stacks.html
|
||||
--- FAIL: TestR681_PageSaysTheInstallWasInterrupted (0.10s)
|
||||
r681_install_interrupted_page_test.go:24: the apps page is silent about an interrupted install (hu)
|
||||
FAIL
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.217s
|
||||
@@ -0,0 +1,5 @@
|
||||
# Red-proof D4 — RecoverInterruptedInstalls does nothing (v0.269.1)
|
||||
--- FAIL: TestR681_InterruptedInstallIsFinishedAndReported (0.00s)
|
||||
r681_install_interrupted_test.go:33: an interrupted install was not reported: resolved=[] hooks=[]
|
||||
FAIL
|
||||
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.009s
|
||||
Reference in New Issue
Block a user