Peti retired (record complete); controller v0.272.0 live proofs + floor 0.272.0; register 339 -> 334; STATUS
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -79,3 +79,23 @@ R-530, R-244, R-600 annotated); CONTEXT; STATUS; memory notes. Historic audits,
|
||||
2. *The hub's host delete removes the WireGuard peer* — **not testable here:** the host was deleted in July and
|
||||
ep0 carries no peer for it now; whether that delete removed one is not recorded (R-600 annotated).
|
||||
3. *The backup holds no user data* — **TRUE, and stronger:** it held no backup at all (one 81-byte key file).
|
||||
|
||||
## Part B — controller v0.272.0 (`44ae4de`), floor 0.272.0
|
||||
|
||||
Red-proofs, each seen failing (`retire-peti-2026-09-25/B/redproof-*.txt`): R-670 (undo read with the validating
|
||||
loader → "the undo logged a false backup-block ERROR"), R-671 (remover call dropped → "copies were left behind"; wiring
|
||||
dropped → "main.go never calls SetUndoCopyRemover"), R-677 (always catalog_since → "got 3 days, want 0"), R-685 page
|
||||
(line dropped → "a space skip must be said on the page"). Full suite + controller gates green; parity unchanged.
|
||||
|
||||
Live on 9202, drill catalog (`B/live/`): navidrome installed at 0.64.0 with a wrong probe; the step to 0.64.1 failed and
|
||||
its undo failed → HOLD, 1 undo copy kept; **R-670:** 0 "backup block rejected" lines while the undo ran (my first count
|
||||
of 3 "unreadable" was disk-watch lines — corrected in `r670-verdict.txt`); **R-671:** the backup page's restore →
|
||||
"update hold … CLEARED", "removed 1 undo cop(y/ies)", 0 copies left, the seeded user read back; **R-677:** the head
|
||||
re-tested at a new digest → „Frissítés elérhető — ma" / "Update available — today" (catalog_since two days old).
|
||||
**R-685 page line: not live** — 9202 has no agent, and no demo box is short of space. Teardown: navidrome removed
|
||||
through the product (0 volumes, 0 copies), 9202 back on the live catalog, drill reset to live `main`.
|
||||
Floor saved 09:31:31Z; both demo boxes on 0.272.0 at 09:31:55Z; demo-hp's backup page renders in both languages.
|
||||
|
||||
## Part C — the night watch: NOT REACHED
|
||||
|
||||
The session ended in the daytime (≈11:40 CEST).
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
before: gitea.dooplex.hu/admin/felhom-controller:0.271.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.272.0 Up 26 seconds (healthy)
|
||||
@@ -0,0 +1,2 @@
|
||||
hu 47546 bytes; no-space line present: False | , amelyből az egész készülék visszaállítható. Ezt a host-ügynök készíti és kezeli. ✓ Utolsó teljes mentés 2026-09-24 21:59 (13 órája) – Helyi tároló (local) Naprakész Következő mentés 13 órája — a mentési ablakon belül ✗ Visszaállítás ellenőrizve 2026-09-25 10:57 Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-24 21:59 (13 órája) Naprakész Bi
|
||||
?lang=en 47178 bytes; no-space line present: False | her — that can bring back the entire device. The host agent makes and manages it. ✓ Last full backup 2026-09-24 21:59 (13 hours ago) – Local storage (local) Up to date Next backup 13 hours ago — within the backup window ✗ Restore checked 2026-09-25 10:57 Local storage (local) ✓ Last successful backup: 2026-09-24 21:59 (13 hours ago) Up to date Back
|
||||
@@ -0,0 +1,7 @@
|
||||
# floor 0.272.0 (declared MinAgent 0.131.0) — 2026-09-25T09:31:31+00:00
|
||||
impact: {"below":4,"valid":true,"version":"0.272.0"}
|
||||
POST -> 303 Location: /configuration?flash=floor_set
|
||||
read back: Effective floor: v0.272.0 — source: DB (
|
||||
arrival 2026-09-25T09:31:55+00:00
|
||||
felhom-pve: gitea.dooplex.hu/admin/felhom-controller:0.272.0 Up 16 seconds (healthy)
|
||||
hp: gitea.dooplex.hu/admin/felhom-controller:0.272.0 Up 16 seconds (healthy)
|
||||
@@ -0,0 +1,174 @@
|
||||
{
|
||||
"update": {
|
||||
"accepted": true,
|
||||
"http": "202",
|
||||
"phases": [
|
||||
{
|
||||
"t": 0.0,
|
||||
"phase": "backing-up",
|
||||
"label": "Biztonsági mentés készül a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 2.1,
|
||||
"phase": "pulling",
|
||||
"label": "Új verzió letöltése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 3.1,
|
||||
"phase": "copying",
|
||||
"label": "Az adatok másolása a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 4.1,
|
||||
"phase": "starting",
|
||||
"label": "Indítás az új verzióval…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 5.1,
|
||||
"phase": "verifying",
|
||||
"label": "Működés ellenőrzése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 95.3,
|
||||
"phase": "undoing",
|
||||
"label": "Visszaállítás az előző változatra…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 187.6,
|
||||
"phase": "failed",
|
||||
"label": "A frissítés nem sikerült",
|
||||
"updating": false,
|
||||
"error": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza.",
|
||||
"hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza."
|
||||
}
|
||||
],
|
||||
"duration_s": 187.7,
|
||||
"final_phase": "failed",
|
||||
"update_error": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza.",
|
||||
"hold_reason": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza.",
|
||||
"state": "stopped"
|
||||
},
|
||||
"restore": {
|
||||
"ok": true,
|
||||
"snapshot_id": "helyi",
|
||||
"snapshots": [
|
||||
{
|
||||
"time": "2026-09-25T09:23:50Z",
|
||||
"short_id": "helyi",
|
||||
"tier": 1,
|
||||
"drive_label": "Tárhely (navidrome)"
|
||||
}
|
||||
],
|
||||
"http": "HTTP/2 302",
|
||||
"location": [
|
||||
"location: /backups/restore?flash=flash.restore.started"
|
||||
],
|
||||
"seconds": 95.2,
|
||||
"state_after": "unhealthy",
|
||||
"hold_after": null,
|
||||
"observables_after": {
|
||||
"pinned_images": {
|
||||
"navidrome": "deluan/navidrome:0.64.0"
|
||||
},
|
||||
"installed_images": {
|
||||
"navidrome": "deluan/navidrome:0.64.0"
|
||||
},
|
||||
"catalog_images": {
|
||||
"navidrome": "deluan/navidrome:0.64.1"
|
||||
},
|
||||
"live_compose_image_lines": [
|
||||
"image: deluan/navidrome:0.64.0@sha256:a384948b81bd1529986c5960169e7fc4fa00f46bde6bd517971a4c36671db2af"
|
||||
],
|
||||
"docker_inspect": [
|
||||
"navidrome deluan/navidrome:0.64.0@sha256:a384948b81bd1529986c5960169e7fc4fa00f46bde6bd517971a4c36671db2af running=true restarts=0"
|
||||
]
|
||||
}
|
||||
},
|
||||
"copies_kept": [
|
||||
"navidrome_navidrome_data.pre-update-20260925T092352Z"
|
||||
],
|
||||
"copies_after": [],
|
||||
"update2": {
|
||||
"accepted": true,
|
||||
"http": "202",
|
||||
"phases": [
|
||||
{
|
||||
"t": 0.0,
|
||||
"phase": "backing-up",
|
||||
"label": "Biztonsági mentés készül a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 2.1,
|
||||
"phase": "pulling",
|
||||
"label": "Új verzió letöltése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 3.1,
|
||||
"phase": "copying",
|
||||
"label": "Az adatok másolása a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 4.1,
|
||||
"phase": "verifying",
|
||||
"label": "Működés ellenőrzése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 10.3,
|
||||
"phase": "done",
|
||||
"label": "Frissítve",
|
||||
"updating": false,
|
||||
"error": null,
|
||||
"hold": null
|
||||
}
|
||||
],
|
||||
"duration_s": 10.3,
|
||||
"final_phase": "done",
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"state": "running"
|
||||
},
|
||||
"badges": {
|
||||
"hu": [
|
||||
{
|
||||
"title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.",
|
||||
"text": "Frissítés elérhető — ma"
|
||||
}
|
||||
],
|
||||
"en": [
|
||||
{
|
||||
"title": "A newer version of this app is available. Select the Update button to start it.",
|
||||
"text": "Update available — today"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
2026/09/25 09:23:40 fillwatch.go:211: [DEBUG] [fillwatch] /mnt/felhom-drives/scratch_hdd/userdata/nextcloud: usage unreadable — skipped (an unreadable filesystem is not a full one)
|
||||
2026/09/25 09:23:40 fillwatch.go:256: [INFO] [fillwatch] checked 6 filesystem(s), 1 unreadable/skipped, 0 notification(s); bands: all ok
|
||||
2026/09/25 09:25:24 undo.go:513: [WARN] [stacks] update navidrome: UNDO — putting back the previous version and its 1 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health check failing))
|
||||
2026/09/25 09:26:56 update.go:1003: [ERROR] [stacks] update navidrome: the UNDO failed too (not_started) — HOLDING the app; the undo copies are kept: [{navidrome_navidrome_data navidrome_navidrome_data.pre-update-20260925T092352Z}]
|
||||
@@ -0,0 +1,9 @@
|
||||
# R-670 verdict (corrected reading) — 2026-09-25T09:30:06+00:00
|
||||
The run log's count of 3 'unreadable' matched fillwatch lines ('usage unreadable'), not the R-670 line.
|
||||
Lines containing 'backup block rejected' in the update+undo window: 0
|
||||
Positive observable — the undo ran: 1 line(s)
|
||||
navidrome's .felhom.yml carries a backup block (so v0.271.0 would have logged the line):
|
||||
23:backup:
|
||||
24- userdata:
|
||||
25- - path: media/music
|
||||
26- class: excluded # re-rippable bulk, :ro, shared tree
|
||||
@@ -0,0 +1,4 @@
|
||||
2026/09/25 09:23:40 fillwatch.go:211: [DEBUG] [fillwatch] /mnt/felhom-drives/scratch_hdd/userdata/nextcloud: usage unreadable — skipped (an unreadable filesystem is not a full one)
|
||||
2026/09/25 09:23:40 fillwatch.go:256: [INFO] [fillwatch] checked 6 filesystem(s), 1 unreadable/skipped, 0 notification(s); bands: all ok
|
||||
2026/09/25 09:28:38 update_guard.go:609: [INFO] [backup] navidrome: restore completed — the update hold (set 2026-09-25T09:26:56Z) is CLEARED
|
||||
2026/09/25 09:28:38 update_guard.go:617: [INFO] [backup] navidrome: removed 1 undo cop(y/ies) the lifted update hold had kept (R-671)
|
||||
@@ -0,0 +1,24 @@
|
||||
11:23:02 drill: navidrome at 0.64.0, probe port 9999 (wrong on purpose) (box at e70453c)
|
||||
11:23:20 deploy navidrome -> True; state=running
|
||||
11:23:20 navidrome: createAdmin http=200
|
||||
11:23:20 seed ok=True
|
||||
11:23:21 navidrome: login as the seeded user http=200 ok=True
|
||||
11:23:21 C1 read before: True
|
||||
11:23:45 drill: navidrome head 0.64.1 (probe still wrong) (box at 06fb564)
|
||||
11:26:57 update -> final failed hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza.'
|
||||
11:27:00 undo copies kept by the hold: ['navidrome_navidrome_data.pre-update-20260925T092352Z']
|
||||
11:27:03 R-670: 'docker-compose.yml unreadable' lines during the update+undo: 3; UNDO lines: 2
|
||||
11:28:42 restore -> {"ok": true, "http": "HTTP/2 302", "seconds": 95.2, "state_after": "unhealthy", "hold_after": null}
|
||||
11:28:44 undo copies AFTER the restore: []
|
||||
11:28:47 after-restore lines:
|
||||
2026/09/25 09:23:40 fillwatch.go:211: [DEBUG] [fillwatch] /mnt/felhom-drives/scratch_hdd/userdata/nextcloud: usage unreadable — skipped (an unreadable filesystem is not a full one)
|
||||
2026/09/25 09:23:40 fillwatch.go:256: [INFO] [fillwatch] checked 6 filesystem(s), 1 unreadable/skipped, 0 notification(s); bands: all ok
|
||||
2026/09/25 09:28:38 update_guard.go:609: [INFO] [backup] navidrome: restore completed — the update hold (set 2026-09-25T09:26:56Z) is CLEARED
|
||||
2026/09/25 09:28:38 update_guard.go:617: [INFO] [backup] navidrome: removed 1 undo cop(y/ies) the lifted update hold had kept (R-671)
|
||||
|
||||
11:28:48 navidrome: login as the seeded user http=200 ok=True
|
||||
11:28:48 read after restore: True
|
||||
11:28:56 drill: navidrome probe put right (box at 2f1863f)
|
||||
11:29:11 update to the head -> done
|
||||
11:29:40 drill: navidrome head re-tested at a new digest (tested now) (box at 2aec5c2)
|
||||
11:29:44 R-677 badges: {"hu": [{"title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.", "text": "Frissítés elérhető — ma"}], "en": [{"title": "A newer version of this app is available. Select the Update button to start it.", "text": "Update available — today"}]}
|
||||
@@ -0,0 +1,10 @@
|
||||
felhom-controller filebrowser paperless-postgres paperless-redis paperless-webserver privatebin traefik
|
||||
0
|
||||
0
|
||||
|
||||
hub:
|
||||
0
|
||||
|
||||
drill=f389756b4dc7 live=f389756b4dc7
|
||||
https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
|
||||
f389756 rules: Peti's box retired 2026-09-25 — the fence is DooPlex and ep0
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestR670_ProbeLoaderLogsNoFalseError
|
||||
r670_probe_meta_test.go:44: the undo logged a false backup-block ERROR:
|
||||
--- FAIL: TestR670_ProbeLoaderLogsNoFalseError (0.01s)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestUndoCopyRemoverIsWiredAtStartup
|
||||
update_leg_wiring_test.go:80: main.go never calls SetUndoCopyRemover — a lifted hold's undo copies would stay for ever
|
||||
--- FAIL: TestUndoCopyRemoverIsWiredAtStartup (0.01s)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestR671_RestoreThatClearsAnUpdateHoldRemovesItsUndoCopies
|
||||
r671_undo_copies_test.go:27: the lifted update hold's undo copies were left behind (remover calls: map[])
|
||||
--- FAIL: TestR671_RestoreThatClearsAnUpdateHoldRemovesItsUndoCopies (0.01s)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestR677_DigestOnlyAgeFromTestedAt
|
||||
r677_badge_age_test.go:32: the badge dates the TAG, not the tested digest: got 3 days (ok=true), want 0
|
||||
--- FAIL: TestR677_DigestOnlyAgeFromTestedAt (0.00s)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestR685_BackupPageSaysTheBackupDoesNotFit
|
||||
r685_no_space_test.go:67: space skip [hu]: a space skip must be said on the page — got "", want "A teljes rendszermentés nem fér el: 11.3 GiB kell, 14.9 GiB szabad."
|
||||
--- FAIL: TestR685_BackupPageSaysTheBackupDoesNotFit (0.06s)
|
||||
@@ -394,3 +394,7 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-643** | **The ruled chain left the automatic update leg at most 15 minutes a night (P2).** Decision 20, built in controller v0.271.0: the full-system backup's gate defers while the leg runs, until W+5h (then only for a step in flight, cap W+5h30m — decision 31); the leg starts no step at or after W+5h; one shared constant. Unit + red-proof (`TestD20_GateWaitsForTheLeg`); live on the demo boxes: see the night record Part D. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/B/redproofs/D20-*` |
|
||||
| **PETI** | **`peti-felhom` deliberately not migrated; parked until the tester reinstalls.** **RETIRED 2026-09-25 (operator ruling):** the tester wiped his server and the box will not return. Removed through the hub's customer delete (journal #20): the customer record and 1,519 residue rows, Storage Box sub-account `u629488-sub2` (id 269130) — which held only one 81-byte `authorized_keys`, never a repository (all 482 reports: 0 off-site snapshots, 0 bytes); ep0 held nothing of it (no PBS namespace, no WireGuard peer). The audit trail stays by design. | retired 2026-09-25 | `git show 6b2176e:documentation/backlog/OPEN-ITEMS.md`; `audits/RETIRE-peti-2026-09-25.md` |
|
||||
| **R-686** | **The automatic update leg is not resumed after a controller restart during the night.** **RULED 2026-09-25 (operator, option B):** the apps the leg had not reached wait for the next night; nothing is built — `09` §3 decision 34. The page-line side effect (a resumed step's `last_auto_update` not written) stays as measured. | ruled 2026-09-25 | `git show 6b2176e:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/C/night3-kill/`, `night4-power/` |
|
||||
| **R-685** | **A whole-box backup that cannot fit must say so before it fails (P2).** Agent v0.134.0: a named skip before any vzdump (free space from `GET /nodes/<n>/storage`), proven live with safe builds on demo-hp. Controller v0.272.0: the backup page's tier row says „A teljes rendszermentés nem fér el: %s kell, %s szabad…” / "The full system backup does not fit…" (render test through the real template, both languages). **The page line has no live proof:** 9202 has no agent and no demo box is short of space now; the operator event rides the existing `whole_guest_backup_failed`. | agent v0.134.0 + controller v0.272.0, 2026-09-25 | `git show eb1c56a:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/F/`; `audits/retire-peti-2026-09-25/B/` |
|
||||
| **R-671** | **The undo copies kept by a hold survived the hold's clearing by a restore (P3).** v0.272.0: the restore that lifts an update hold removes that hold's undo copies (never a restore hold, never mid-update). Live on 9202: navidrome held with 1 copy kept → the backup page's restore → hold CLEARED, "removed 1 undo cop(y/ies)", 0 copies left, data read back. | v0.272.0, 2026-09-25 | `git show eb1c56a:documentation/backlog/OPEN-ITEMS.md`; `audits/retire-peti-2026-09-25/B/live/` |
|
||||
| **R-670** | **Every undo logged a false `backup block rejected … docker-compose.yml unreadable` (P3).** v0.272.0: probe-only copies load with `stacks.LoadProbeMetadata`. Live on 9202: an undo of navidrome (a backup block in its `.felhom.yml`) logged 0 such lines. | v0.272.0, 2026-09-25 | `git show eb1c56a:documentation/backlog/OPEN-ITEMS.md`; `audits/retire-peti-2026-09-25/B/live/r670-verdict.txt` |
|
||||
| **R-677** | **A re-tested floating tag's badge age read the tag's date (P3).** v0.272.0: `stacks.BehindSinceAge` — a digest-only move counts from `tested_at`; both producers. Live on 9202: „Frissítés elérhető — ma” / "Update available — today" with `catalog_since` two days old. | v0.272.0, 2026-09-25 | `git show eb1c56a:documentation/backlog/OPEN-ITEMS.md`; `audits/retire-peti-2026-09-25/B/live/live272.json` |
|
||||
|
||||
@@ -804,14 +804,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** |
|
||||
| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. | **READY — P2; owner: CC (catalog)** |
|
||||
| **R-657** | **[P2-MEDIUM] "Remove the app, keep my data", then install it again: nextcloud never installs, and the box only says „unhealthy".** MEASURED 2026-09-23 night on 9202 (v0.267.0): nextcloud was removed through the product keeping its drive folder (the remove with data was refused — R-442's fail-closed guard, as on every drill on 9202 — and the product's keep-data remove taken). An hour later a fresh install of nextcloud on the same box: the template binds `${HDD_PATH}/appdata/nextcloud` to `/var/www/html/data`, the kept folder still holds `admin/`, `appdata_*`, `.ncdata` and a 145 MB `nextcloud.log`, and the image's installer loops **„Login is invalid because files already exist for this user — Retrying install..."**; `occ status` reads `installed: false`. The controller records the deploy as done and the app as `unhealthy`; nothing tells the household that their kept files are what blocks the new install, or what to do. **Why it matters:** keep-data is the choice the product OFFERS a household at remove time — and for nextcloud the kept data makes the app uninstallable. **Needs:** decide the product's promise for a reinstall over kept data, per app class (adopt the data? refuse with a sentence? offer to move it aside?); at minimum a deploy-time refusal or warning when the app's drive folder is not empty. Evidence: `audits/night-2026-09-23/chaos/00-nextcloud-reinstall-over-kept-data.txt`. | **READY — P2; owner: operator (the promise) / CC (the build)** |
|
||||
| **R-670** | **[P3-LOW] Every undo (and every step-file verify) logs `[ERROR] .felhom.yml backup block rejected … docker-compose.yml unreadable`.** `LoadMetadata` validates the backup block against a compose file that the pre-update-meta directory (undo.go:528, since v0.263.0) and the scratch dir of `loadMetadataFile` (v0.269.0) never hold. The health check it feeds is unaffected; an operator reading ERROR lines after an undo is misled. Seen 10:24:41Z and 10:46:24Z on 9202. **Fix:** a probe-only loader that skips the backup-block validation, or copy the compose beside it. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-671** | **[P3-LOW] The undo copies kept by a hold survive the hold's clearing by a restore.** MEASURED 2026-09-24 on 9202: three `nextcloud_*.pre-update-20260924T103924Z` volumes (~0.9 GiB) were still present after the whole restore cleared that hold at 10:41:52Z, and a second set joined them 6 minutes later. Nothing names them on a page; on a small disk they are the difference between the next update's copy fitting or not. **Fix direction:** the restore that clears an update hold removes that hold's undo copies (they describe the state the restore just replaced), logged. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` (volume list) | **READY — P3; owner: CC (controller)** |
|
||||
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-677** | **[P3-LOW] For a floating tag re-tested at a new digest, the Behind badge's age reads the TAG's catalog date („1 napja"), not when the new digest was tested (minutes).** Seen 2026-09-24 on 9202 (Part B). Harmless but confusing. **Fix:** for a digest-only move, age from the ladder entry's `tested_at`. `audits/night-2026-09-24/B/10-floating-tag.json` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-685** | **[P2-MEDIUM] A box whose whole-box backup cannot fit must say so BEFORE the night — an operator event and a line on the backup page — never only a nightly failure a log shows.** The product lesson of R-684 (operator ruling 2026-09-24 evening, option A for demo-hp): demo-hp's root-disk target held three 6–7 GB archives with ~4 GB free and failed `No space left on device` every night from 2026-09-23 while nothing but the vzdump log said why. Measured 2026-09-24 night: a 9201 archive is 8.18 GB (22.6 GB uncompressed); PVE prunes AFTER a successful backup, so a target must hold keep-last + 1 archives during the run — lowering the retention alone does not un-stick a full target. **Fix direction:** the agent predicts the archive size (last archive × margin) against the target's free space before a whole-box backup, skips with a reason reported to the hub (operator event), and the controller's backup page shows the sentence. `audits/night-2026-09-25/A/A4-hp-backup-space.txt` | **READY — P2; owner: CC (agent + controller)** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` | **OPEN — P3; owner: CC** |
|
||||
| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` | **READY — P3; owner: CC (hub) / operator (Peti's Cloudflare leftovers, if any)** |
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
|
||||
Reference in New Issue
Block a user