Peti retired (record complete); controller v0.272.0 live proofs + floor 0.272.0; register 339 -> 334; STATUS
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -14,6 +14,14 @@
|
||||
> language, one screen, no identifiers in the prose. Same subjects, different readers; merging them
|
||||
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
## 2026-09-25 (midday) — Peti retired in fact; controller v0.272.0
|
||||
|
||||
- Peti's hub customer deleted through the cascade (journal #20); Storage Box sub-account 269130 (`u629488-sub2`)
|
||||
deleted — it held one 81-byte `authorized_keys`, never a repo; ep0 held nothing of it. Record:
|
||||
`documentation/audits/RETIRE-peti-2026-09-25.md`. Cloudflare side not covered by the cascade: R-688.
|
||||
- Controller **v0.272.0**, floor 0.272.0 (MinAgent 0.131.0): backup-page „nem fér el" sentence (R-685 page half),
|
||||
R-671, R-670, R-677 — three proven live on 9202; the page sentence render-tested only (no box is short of space).
|
||||
|
||||
## Rulings 2026-09-25 (operator)
|
||||
|
||||
- **Peti's box (`peti-felhom`, host `peti-felhom-86d37d`) is RETIRED.** It will not return — the tester wiped his
|
||||
|
||||
@@ -1,8 +1,7 @@
|
||||
# REPORT — night shift 2026-09-24/25 (felhom.eu side)
|
||||
# REPORT — Peti's box retired; controller v0.272.0 (2026-09-25, felhom.eu side)
|
||||
|
||||
The full record is `documentation/audits/DRILL-night-2026-09-25.md` (opens with "Not done, or changed"). This repo
|
||||
carried: the evidence tree `documentation/audits/night-2026-09-25/`; `09` (part 7 shipped, decisions 31–33), `03`
|
||||
(delivery + R-685), `07` (the fourth leg; demo-hp retention), `08` (what decision 28's suppression covers), the
|
||||
capability map; the register (341 → 338 rows); `STATUS.md`; the hub's global floor raised to 0.271.0 (MinAgent
|
||||
0.131.0) through the operator UI endpoint; `scripts/wire_contract_gate.py` allowlists the report's new `update_leg`
|
||||
with its reason. No hub release.
|
||||
Record: `documentation/audits/RETIRE-peti-2026-09-25.md` (opens with "Not done, or changed"; inventory, listing,
|
||||
removal, before/after, documents, claims). Controller release evidence: `documentation/audits/retire-peti-2026-09-25/B/`.
|
||||
This repo carried: the rulings (`CONTEXT.md`, `09` §3 decision 34, `target-selection.md`), the document updates,
|
||||
the register (339 → 334 rows; opened R-688; closed PETI, R-686, R-685, R-671, R-670, R-677), `STATUS.md`, the hub
|
||||
customer delete of `peti-felhom` through the operator UI endpoint, and the global floor raised to 0.272.0. No hub release.
|
||||
|
||||
@@ -1,32 +1,24 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-25 (morning). Both demo boxes run controller 0.271.0 and host agent 0.134.0. Automatic app updates are built and ran on their own last night. The HP box can back up again.**
|
||||
**Updated 2026-09-25 (midday). Peti's box is retired. Both demo boxes run controller 0.272.0 and host agent 0.134.0.**
|
||||
|
||||
**Decisions I took on my own (you may reverse them).**
|
||||
- **The box takes one tested step per app per night.** Your brief said "one tested step per app". An app three steps behind now needs three nights.
|
||||
- **After the 5-hour mark, the full-system backup waits only for an update step that is already running, and never more than 30 minutes.** Starting the backup in the middle of a step would stop the app while it is being checked and undo a good update.
|
||||
- **The controller does not update itself while the night's app updates run.** A controller restart in the middle would stop the rest of that night's updates.
|
||||
**Decisions I took on my own.** None this time. I recorded your two rulings: Peti's box is retired, and a controller restart during the night's updates does not continue them (the apps not reached wait a night).
|
||||
|
||||
**What I did, and it worked.**
|
||||
- **The fixed host agent reached both demo boxes.** I signed it for each box. The restore test is on again. On the N100 box it passed (85 seconds). On the HP box it said "not enough space" and created nothing, which is correct.
|
||||
- **The HP box backs up again (your option A).** It now keeps one old whole-box backup instead of three. I removed the two oldest. A backup started with the backup page's own button fitted: 8 GB, and the disk is now 60 % full.
|
||||
- **Automatic updates are built.** Each night, after the off-site copy, the box updates its apps by itself: one app at a time, one tested step each, the previous version ready to put back. There is a switch on the settings page, in both languages, on by default. A successful update sends no mail; the app page shows a line.
|
||||
- **Proven on the scratch box first, over six simulated nights.** A failing update was put back and not tried again until the catalog re-tested it. A step marked "needs a person" was never taken. A power cut during an update: the box finished it after the restart. The controller killed during an update: the update was cleanly put back. Switch off: nothing happened. The app data read back after every night.
|
||||
- **The demo boxes' first real automatic night.** The N100 box updated opengist (1.13 → 1.15) by itself at 04:15, in 20 seconds. The HP box had nothing to update, and said so. Both reported to the hub.
|
||||
- **A new host agent (0.134.0) skips a whole-box backup that cannot fit**, and says why, before it starts. It is on both demo boxes.
|
||||
- **Two more apps moved in the catalog** (n8n, mealie), each tested twice before it moved.
|
||||
- **Peti's box is gone from everything we run.** Its hub customer and all its records went through the hub's own delete, and its off-site folder was removed. The hub keeps only its history (events and the deletion note).
|
||||
- **His off-site "backup" was never a backup.** His folder held one small key file and no backup at all. None of his 482 reports ever showed an off-site copy. That matches what you said: no user data.
|
||||
- **Nothing of his was on ep0**, the off-site server. I checked: its backups and its network peers are the same as before, and so are the demo boxes' off-site folders.
|
||||
- **The protected list is now DooPlex and ep0.** I updated the rules, the runbooks and the architecture pages. Old records keep their text.
|
||||
- **Controller 0.272.0 is on both demo boxes.** The backup page now says in plain words when a whole-box backup cannot fit. Three small fixes: a restore now also removes the extra copies an update hold kept; a false error line after every undo is gone; the "update available" age is right for a re-tested image. I proved the three fixes on the scratch box.
|
||||
|
||||
**What broke, or is not done.**
|
||||
- **My own test started a real whole-box backup on the HP box for 5 minutes.** I stopped it. Nothing was left behind, and the apps kept running. The cause was a bug in the new agent, and I fixed it before the release.
|
||||
- **If the controller restarts in the night, the rest of that night's updates wait for the next night.** Your decision is below.
|
||||
- **The backup-page sentence for "backup does not fit" is not built yet.** The agent half is done. The page half needs the next controller release.
|
||||
- **Last night no whole-box backup was due on either box**, so the "backup waits for the updates" rule did not happen for real yet. The tests prove it.
|
||||
- **The "does not fit" sentence has not appeared on a real box yet.** No box is short of space now. The tests prove it in both languages.
|
||||
- **The hub's customer delete promises to remove the Cloudflare tunnel and name, but it does not.** Peti's tokens were deleted with his record. Anything left on Cloudflare's side is not checked. I filed it as a row.
|
||||
- **Last night's watch did not run.** This session ended in the daytime.
|
||||
|
||||
**Rows.** 3 opened, 6 closed, 2 updated. The list went from 341 to 338.
|
||||
**Rows.** 1 opened, 6 closed. The list went from 339 to 334.
|
||||
|
||||
**What needs you.**
|
||||
1. **Peti's box: what should its automatic updates be before it comes back online?** It has been silent since 15 July, on a very old version. The switch is on by default. If it comes back and takes the new version, it updates its apps by itself from its first night.
|
||||
- **A (my choice): hold Peti's box on its current version until someone looks at it after it returns.** Cost: it stays behind until then, and the hold must be removed later by hand.
|
||||
- **B: leave it.** Cost: on its first night online it jumps many versions and updates its apps by itself, with nobody watching.
|
||||
- **If you do nothing, B happens.**
|
||||
2. **If the controller restarts in the night, should the box continue that night's updates?** (A) Yes: the box remembers the night and continues until the 5-hour mark. (B) No: it waits a day, as now. I would choose A. If you do nothing, B stays.
|
||||
1. **Remove Peti's box from the Claude project instructions** (your own text in the Claude project). I cannot edit those. If you leave it, new sessions still treat his box as protected.
|
||||
2. **Check Cloudflare for anything left of Peti's domain** (`sajatfelhom.hu`: a tunnel or DNS records). If you do nothing, it stays there unused.
|
||||
3. **The old Storage Box `PBS-storage-1`** (u629193) is still on the list for you to delete. Neither of our tokens can see it now; it may already be gone. If it still exists, it keeps old test leftovers, including a folder named for Peti.
|
||||
|
||||
@@ -79,3 +79,23 @@ R-530, R-244, R-600 annotated); CONTEXT; STATUS; memory notes. Historic audits,
|
||||
2. *The hub's host delete removes the WireGuard peer* — **not testable here:** the host was deleted in July and
|
||||
ep0 carries no peer for it now; whether that delete removed one is not recorded (R-600 annotated).
|
||||
3. *The backup holds no user data* — **TRUE, and stronger:** it held no backup at all (one 81-byte key file).
|
||||
|
||||
## Part B — controller v0.272.0 (`44ae4de`), floor 0.272.0
|
||||
|
||||
Red-proofs, each seen failing (`retire-peti-2026-09-25/B/redproof-*.txt`): R-670 (undo read with the validating
|
||||
loader → "the undo logged a false backup-block ERROR"), R-671 (remover call dropped → "copies were left behind"; wiring
|
||||
dropped → "main.go never calls SetUndoCopyRemover"), R-677 (always catalog_since → "got 3 days, want 0"), R-685 page
|
||||
(line dropped → "a space skip must be said on the page"). Full suite + controller gates green; parity unchanged.
|
||||
|
||||
Live on 9202, drill catalog (`B/live/`): navidrome installed at 0.64.0 with a wrong probe; the step to 0.64.1 failed and
|
||||
its undo failed → HOLD, 1 undo copy kept; **R-670:** 0 "backup block rejected" lines while the undo ran (my first count
|
||||
of 3 "unreadable" was disk-watch lines — corrected in `r670-verdict.txt`); **R-671:** the backup page's restore →
|
||||
"update hold … CLEARED", "removed 1 undo cop(y/ies)", 0 copies left, the seeded user read back; **R-677:** the head
|
||||
re-tested at a new digest → „Frissítés elérhető — ma" / "Update available — today" (catalog_since two days old).
|
||||
**R-685 page line: not live** — 9202 has no agent, and no demo box is short of space. Teardown: navidrome removed
|
||||
through the product (0 volumes, 0 copies), 9202 back on the live catalog, drill reset to live `main`.
|
||||
Floor saved 09:31:31Z; both demo boxes on 0.272.0 at 09:31:55Z; demo-hp's backup page renders in both languages.
|
||||
|
||||
## Part C — the night watch: NOT REACHED
|
||||
|
||||
The session ended in the daytime (≈11:40 CEST).
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
before: gitea.dooplex.hu/admin/felhom-controller:0.271.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.272.0 Up 26 seconds (healthy)
|
||||
@@ -0,0 +1,2 @@
|
||||
hu 47546 bytes; no-space line present: False | , amelyből az egész készülék visszaállítható. Ezt a host-ügynök készíti és kezeli. ✓ Utolsó teljes mentés 2026-09-24 21:59 (13 órája) – Helyi tároló (local) Naprakész Következő mentés 13 órája — a mentési ablakon belül ✗ Visszaállítás ellenőrizve 2026-09-25 10:57 Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-24 21:59 (13 órája) Naprakész Bi
|
||||
?lang=en 47178 bytes; no-space line present: False | her — that can bring back the entire device. The host agent makes and manages it. ✓ Last full backup 2026-09-24 21:59 (13 hours ago) – Local storage (local) Up to date Next backup 13 hours ago — within the backup window ✗ Restore checked 2026-09-25 10:57 Local storage (local) ✓ Last successful backup: 2026-09-24 21:59 (13 hours ago) Up to date Back
|
||||
@@ -0,0 +1,7 @@
|
||||
# floor 0.272.0 (declared MinAgent 0.131.0) — 2026-09-25T09:31:31+00:00
|
||||
impact: {"below":4,"valid":true,"version":"0.272.0"}
|
||||
POST -> 303 Location: /configuration?flash=floor_set
|
||||
read back: Effective floor: v0.272.0 — source: DB (
|
||||
arrival 2026-09-25T09:31:55+00:00
|
||||
felhom-pve: gitea.dooplex.hu/admin/felhom-controller:0.272.0 Up 16 seconds (healthy)
|
||||
hp: gitea.dooplex.hu/admin/felhom-controller:0.272.0 Up 16 seconds (healthy)
|
||||
@@ -0,0 +1,174 @@
|
||||
{
|
||||
"update": {
|
||||
"accepted": true,
|
||||
"http": "202",
|
||||
"phases": [
|
||||
{
|
||||
"t": 0.0,
|
||||
"phase": "backing-up",
|
||||
"label": "Biztonsági mentés készül a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 2.1,
|
||||
"phase": "pulling",
|
||||
"label": "Új verzió letöltése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 3.1,
|
||||
"phase": "copying",
|
||||
"label": "Az adatok másolása a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 4.1,
|
||||
"phase": "starting",
|
||||
"label": "Indítás az új verzióval…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 5.1,
|
||||
"phase": "verifying",
|
||||
"label": "Működés ellenőrzése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 95.3,
|
||||
"phase": "undoing",
|
||||
"label": "Visszaállítás az előző változatra…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 187.6,
|
||||
"phase": "failed",
|
||||
"label": "A frissítés nem sikerült",
|
||||
"updating": false,
|
||||
"error": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza.",
|
||||
"hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza."
|
||||
}
|
||||
],
|
||||
"duration_s": 187.7,
|
||||
"final_phase": "failed",
|
||||
"update_error": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza.",
|
||||
"hold_reason": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza.",
|
||||
"state": "stopped"
|
||||
},
|
||||
"restore": {
|
||||
"ok": true,
|
||||
"snapshot_id": "helyi",
|
||||
"snapshots": [
|
||||
{
|
||||
"time": "2026-09-25T09:23:50Z",
|
||||
"short_id": "helyi",
|
||||
"tier": 1,
|
||||
"drive_label": "Tárhely (navidrome)"
|
||||
}
|
||||
],
|
||||
"http": "HTTP/2 302",
|
||||
"location": [
|
||||
"location: /backups/restore?flash=flash.restore.started"
|
||||
],
|
||||
"seconds": 95.2,
|
||||
"state_after": "unhealthy",
|
||||
"hold_after": null,
|
||||
"observables_after": {
|
||||
"pinned_images": {
|
||||
"navidrome": "deluan/navidrome:0.64.0"
|
||||
},
|
||||
"installed_images": {
|
||||
"navidrome": "deluan/navidrome:0.64.0"
|
||||
},
|
||||
"catalog_images": {
|
||||
"navidrome": "deluan/navidrome:0.64.1"
|
||||
},
|
||||
"live_compose_image_lines": [
|
||||
"image: deluan/navidrome:0.64.0@sha256:a384948b81bd1529986c5960169e7fc4fa00f46bde6bd517971a4c36671db2af"
|
||||
],
|
||||
"docker_inspect": [
|
||||
"navidrome deluan/navidrome:0.64.0@sha256:a384948b81bd1529986c5960169e7fc4fa00f46bde6bd517971a4c36671db2af running=true restarts=0"
|
||||
]
|
||||
}
|
||||
},
|
||||
"copies_kept": [
|
||||
"navidrome_navidrome_data.pre-update-20260925T092352Z"
|
||||
],
|
||||
"copies_after": [],
|
||||
"update2": {
|
||||
"accepted": true,
|
||||
"http": "202",
|
||||
"phases": [
|
||||
{
|
||||
"t": 0.0,
|
||||
"phase": "backing-up",
|
||||
"label": "Biztonsági mentés készül a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 2.1,
|
||||
"phase": "pulling",
|
||||
"label": "Új verzió letöltése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 3.1,
|
||||
"phase": "copying",
|
||||
"label": "Az adatok másolása a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 4.1,
|
||||
"phase": "verifying",
|
||||
"label": "Működés ellenőrzése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 10.3,
|
||||
"phase": "done",
|
||||
"label": "Frissítve",
|
||||
"updating": false,
|
||||
"error": null,
|
||||
"hold": null
|
||||
}
|
||||
],
|
||||
"duration_s": 10.3,
|
||||
"final_phase": "done",
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"state": "running"
|
||||
},
|
||||
"badges": {
|
||||
"hu": [
|
||||
{
|
||||
"title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.",
|
||||
"text": "Frissítés elérhető — ma"
|
||||
}
|
||||
],
|
||||
"en": [
|
||||
{
|
||||
"title": "A newer version of this app is available. Select the Update button to start it.",
|
||||
"text": "Update available — today"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
2026/09/25 09:23:40 fillwatch.go:211: [DEBUG] [fillwatch] /mnt/felhom-drives/scratch_hdd/userdata/nextcloud: usage unreadable — skipped (an unreadable filesystem is not a full one)
|
||||
2026/09/25 09:23:40 fillwatch.go:256: [INFO] [fillwatch] checked 6 filesystem(s), 1 unreadable/skipped, 0 notification(s); bands: all ok
|
||||
2026/09/25 09:25:24 undo.go:513: [WARN] [stacks] update navidrome: UNDO — putting back the previous version and its 1 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health check failing))
|
||||
2026/09/25 09:26:56 update.go:1003: [ERROR] [stacks] update navidrome: the UNDO failed too (not_started) — HOLDING the app; the undo copies are kept: [{navidrome_navidrome_data navidrome_navidrome_data.pre-update-20260925T092352Z}]
|
||||
@@ -0,0 +1,9 @@
|
||||
# R-670 verdict (corrected reading) — 2026-09-25T09:30:06+00:00
|
||||
The run log's count of 3 'unreadable' matched fillwatch lines ('usage unreadable'), not the R-670 line.
|
||||
Lines containing 'backup block rejected' in the update+undo window: 0
|
||||
Positive observable — the undo ran: 1 line(s)
|
||||
navidrome's .felhom.yml carries a backup block (so v0.271.0 would have logged the line):
|
||||
23:backup:
|
||||
24- userdata:
|
||||
25- - path: media/music
|
||||
26- class: excluded # re-rippable bulk, :ro, shared tree
|
||||
@@ -0,0 +1,4 @@
|
||||
2026/09/25 09:23:40 fillwatch.go:211: [DEBUG] [fillwatch] /mnt/felhom-drives/scratch_hdd/userdata/nextcloud: usage unreadable — skipped (an unreadable filesystem is not a full one)
|
||||
2026/09/25 09:23:40 fillwatch.go:256: [INFO] [fillwatch] checked 6 filesystem(s), 1 unreadable/skipped, 0 notification(s); bands: all ok
|
||||
2026/09/25 09:28:38 update_guard.go:609: [INFO] [backup] navidrome: restore completed — the update hold (set 2026-09-25T09:26:56Z) is CLEARED
|
||||
2026/09/25 09:28:38 update_guard.go:617: [INFO] [backup] navidrome: removed 1 undo cop(y/ies) the lifted update hold had kept (R-671)
|
||||
@@ -0,0 +1,24 @@
|
||||
11:23:02 drill: navidrome at 0.64.0, probe port 9999 (wrong on purpose) (box at e70453c)
|
||||
11:23:20 deploy navidrome -> True; state=running
|
||||
11:23:20 navidrome: createAdmin http=200
|
||||
11:23:20 seed ok=True
|
||||
11:23:21 navidrome: login as the seeded user http=200 ok=True
|
||||
11:23:21 C1 read before: True
|
||||
11:23:45 drill: navidrome head 0.64.1 (probe still wrong) (box at 06fb564)
|
||||
11:26:57 update -> final failed hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) navidrome frissítése 2026-09-25 11:26-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-25 11:23 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza.'
|
||||
11:27:00 undo copies kept by the hold: ['navidrome_navidrome_data.pre-update-20260925T092352Z']
|
||||
11:27:03 R-670: 'docker-compose.yml unreadable' lines during the update+undo: 3; UNDO lines: 2
|
||||
11:28:42 restore -> {"ok": true, "http": "HTTP/2 302", "seconds": 95.2, "state_after": "unhealthy", "hold_after": null}
|
||||
11:28:44 undo copies AFTER the restore: []
|
||||
11:28:47 after-restore lines:
|
||||
2026/09/25 09:23:40 fillwatch.go:211: [DEBUG] [fillwatch] /mnt/felhom-drives/scratch_hdd/userdata/nextcloud: usage unreadable — skipped (an unreadable filesystem is not a full one)
|
||||
2026/09/25 09:23:40 fillwatch.go:256: [INFO] [fillwatch] checked 6 filesystem(s), 1 unreadable/skipped, 0 notification(s); bands: all ok
|
||||
2026/09/25 09:28:38 update_guard.go:609: [INFO] [backup] navidrome: restore completed — the update hold (set 2026-09-25T09:26:56Z) is CLEARED
|
||||
2026/09/25 09:28:38 update_guard.go:617: [INFO] [backup] navidrome: removed 1 undo cop(y/ies) the lifted update hold had kept (R-671)
|
||||
|
||||
11:28:48 navidrome: login as the seeded user http=200 ok=True
|
||||
11:28:48 read after restore: True
|
||||
11:28:56 drill: navidrome probe put right (box at 2f1863f)
|
||||
11:29:11 update to the head -> done
|
||||
11:29:40 drill: navidrome head re-tested at a new digest (tested now) (box at 2aec5c2)
|
||||
11:29:44 R-677 badges: {"hu": [{"title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.", "text": "Frissítés elérhető — ma"}], "en": [{"title": "A newer version of this app is available. Select the Update button to start it.", "text": "Update available — today"}]}
|
||||
@@ -0,0 +1,10 @@
|
||||
felhom-controller filebrowser paperless-postgres paperless-redis paperless-webserver privatebin traefik
|
||||
0
|
||||
0
|
||||
|
||||
hub:
|
||||
0
|
||||
|
||||
drill=f389756b4dc7 live=f389756b4dc7
|
||||
https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
|
||||
f389756 rules: Peti's box retired 2026-09-25 — the fence is DooPlex and ep0
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestR670_ProbeLoaderLogsNoFalseError
|
||||
r670_probe_meta_test.go:44: the undo logged a false backup-block ERROR:
|
||||
--- FAIL: TestR670_ProbeLoaderLogsNoFalseError (0.01s)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestUndoCopyRemoverIsWiredAtStartup
|
||||
update_leg_wiring_test.go:80: main.go never calls SetUndoCopyRemover — a lifted hold's undo copies would stay for ever
|
||||
--- FAIL: TestUndoCopyRemoverIsWiredAtStartup (0.01s)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestR671_RestoreThatClearsAnUpdateHoldRemovesItsUndoCopies
|
||||
r671_undo_copies_test.go:27: the lifted update hold's undo copies were left behind (remover calls: map[])
|
||||
--- FAIL: TestR671_RestoreThatClearsAnUpdateHoldRemovesItsUndoCopies (0.01s)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestR677_DigestOnlyAgeFromTestedAt
|
||||
r677_badge_age_test.go:32: the badge dates the TAG, not the tested digest: got 3 days (ok=true), want 0
|
||||
--- FAIL: TestR677_DigestOnlyAgeFromTestedAt (0.00s)
|
||||
@@ -0,0 +1,3 @@
|
||||
=== RUN TestR685_BackupPageSaysTheBackupDoesNotFit
|
||||
r685_no_space_test.go:67: space skip [hu]: a space skip must be said on the page — got "", want "A teljes rendszermentés nem fér el: 11.3 GiB kell, 14.9 GiB szabad."
|
||||
--- FAIL: TestR685_BackupPageSaysTheBackupDoesNotFit (0.06s)
|
||||
@@ -394,3 +394,7 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-643** | **The ruled chain left the automatic update leg at most 15 minutes a night (P2).** Decision 20, built in controller v0.271.0: the full-system backup's gate defers while the leg runs, until W+5h (then only for a step in flight, cap W+5h30m — decision 31); the leg starts no step at or after W+5h; one shared constant. Unit + red-proof (`TestD20_GateWaitsForTheLeg`); live on the demo boxes: see the night record Part D. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/B/redproofs/D20-*` |
|
||||
| **PETI** | **`peti-felhom` deliberately not migrated; parked until the tester reinstalls.** **RETIRED 2026-09-25 (operator ruling):** the tester wiped his server and the box will not return. Removed through the hub's customer delete (journal #20): the customer record and 1,519 residue rows, Storage Box sub-account `u629488-sub2` (id 269130) — which held only one 81-byte `authorized_keys`, never a repository (all 482 reports: 0 off-site snapshots, 0 bytes); ep0 held nothing of it (no PBS namespace, no WireGuard peer). The audit trail stays by design. | retired 2026-09-25 | `git show 6b2176e:documentation/backlog/OPEN-ITEMS.md`; `audits/RETIRE-peti-2026-09-25.md` |
|
||||
| **R-686** | **The automatic update leg is not resumed after a controller restart during the night.** **RULED 2026-09-25 (operator, option B):** the apps the leg had not reached wait for the next night; nothing is built — `09` §3 decision 34. The page-line side effect (a resumed step's `last_auto_update` not written) stays as measured. | ruled 2026-09-25 | `git show 6b2176e:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/C/night3-kill/`, `night4-power/` |
|
||||
| **R-685** | **A whole-box backup that cannot fit must say so before it fails (P2).** Agent v0.134.0: a named skip before any vzdump (free space from `GET /nodes/<n>/storage`), proven live with safe builds on demo-hp. Controller v0.272.0: the backup page's tier row says „A teljes rendszermentés nem fér el: %s kell, %s szabad…” / "The full system backup does not fit…" (render test through the real template, both languages). **The page line has no live proof:** 9202 has no agent and no demo box is short of space now; the operator event rides the existing `whole_guest_backup_failed`. | agent v0.134.0 + controller v0.272.0, 2026-09-25 | `git show eb1c56a:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/F/`; `audits/retire-peti-2026-09-25/B/` |
|
||||
| **R-671** | **The undo copies kept by a hold survived the hold's clearing by a restore (P3).** v0.272.0: the restore that lifts an update hold removes that hold's undo copies (never a restore hold, never mid-update). Live on 9202: navidrome held with 1 copy kept → the backup page's restore → hold CLEARED, "removed 1 undo cop(y/ies)", 0 copies left, data read back. | v0.272.0, 2026-09-25 | `git show eb1c56a:documentation/backlog/OPEN-ITEMS.md`; `audits/retire-peti-2026-09-25/B/live/` |
|
||||
| **R-670** | **Every undo logged a false `backup block rejected … docker-compose.yml unreadable` (P3).** v0.272.0: probe-only copies load with `stacks.LoadProbeMetadata`. Live on 9202: an undo of navidrome (a backup block in its `.felhom.yml`) logged 0 such lines. | v0.272.0, 2026-09-25 | `git show eb1c56a:documentation/backlog/OPEN-ITEMS.md`; `audits/retire-peti-2026-09-25/B/live/r670-verdict.txt` |
|
||||
| **R-677** | **A re-tested floating tag's badge age read the tag's date (P3).** v0.272.0: `stacks.BehindSinceAge` — a digest-only move counts from `tested_at`; both producers. Live on 9202: „Frissítés elérhető — ma” / "Update available — today" with `catalog_since` two days old. | v0.272.0, 2026-09-25 | `git show eb1c56a:documentation/backlog/OPEN-ITEMS.md`; `audits/retire-peti-2026-09-25/B/live/live272.json` |
|
||||
|
||||
@@ -804,14 +804,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** |
|
||||
| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. | **READY — P2; owner: CC (catalog)** |
|
||||
| **R-657** | **[P2-MEDIUM] "Remove the app, keep my data", then install it again: nextcloud never installs, and the box only says „unhealthy".** MEASURED 2026-09-23 night on 9202 (v0.267.0): nextcloud was removed through the product keeping its drive folder (the remove with data was refused — R-442's fail-closed guard, as on every drill on 9202 — and the product's keep-data remove taken). An hour later a fresh install of nextcloud on the same box: the template binds `${HDD_PATH}/appdata/nextcloud` to `/var/www/html/data`, the kept folder still holds `admin/`, `appdata_*`, `.ncdata` and a 145 MB `nextcloud.log`, and the image's installer loops **„Login is invalid because files already exist for this user — Retrying install..."**; `occ status` reads `installed: false`. The controller records the deploy as done and the app as `unhealthy`; nothing tells the household that their kept files are what blocks the new install, or what to do. **Why it matters:** keep-data is the choice the product OFFERS a household at remove time — and for nextcloud the kept data makes the app uninstallable. **Needs:** decide the product's promise for a reinstall over kept data, per app class (adopt the data? refuse with a sentence? offer to move it aside?); at minimum a deploy-time refusal or warning when the app's drive folder is not empty. Evidence: `audits/night-2026-09-23/chaos/00-nextcloud-reinstall-over-kept-data.txt`. | **READY — P2; owner: operator (the promise) / CC (the build)** |
|
||||
| **R-670** | **[P3-LOW] Every undo (and every step-file verify) logs `[ERROR] .felhom.yml backup block rejected … docker-compose.yml unreadable`.** `LoadMetadata` validates the backup block against a compose file that the pre-update-meta directory (undo.go:528, since v0.263.0) and the scratch dir of `loadMetadataFile` (v0.269.0) never hold. The health check it feeds is unaffected; an operator reading ERROR lines after an undo is misled. Seen 10:24:41Z and 10:46:24Z on 9202. **Fix:** a probe-only loader that skips the backup-block validation, or copy the compose beside it. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-671** | **[P3-LOW] The undo copies kept by a hold survive the hold's clearing by a restore.** MEASURED 2026-09-24 on 9202: three `nextcloud_*.pre-update-20260924T103924Z` volumes (~0.9 GiB) were still present after the whole restore cleared that hold at 10:41:52Z, and a second set joined them 6 minutes later. Nothing names them on a page; on a small disk they are the difference between the next update's copy fitting or not. **Fix direction:** the restore that clears an update hold removes that hold's undo copies (they describe the state the restore just replaced), logged. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` (volume list) | **READY — P3; owner: CC (controller)** |
|
||||
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-677** | **[P3-LOW] For a floating tag re-tested at a new digest, the Behind badge's age reads the TAG's catalog date („1 napja"), not when the new digest was tested (minutes).** Seen 2026-09-24 on 9202 (Part B). Harmless but confusing. **Fix:** for a digest-only move, age from the ladder entry's `tested_at`. `audits/night-2026-09-24/B/10-floating-tag.json` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-685** | **[P2-MEDIUM] A box whose whole-box backup cannot fit must say so BEFORE the night — an operator event and a line on the backup page — never only a nightly failure a log shows.** The product lesson of R-684 (operator ruling 2026-09-24 evening, option A for demo-hp): demo-hp's root-disk target held three 6–7 GB archives with ~4 GB free and failed `No space left on device` every night from 2026-09-23 while nothing but the vzdump log said why. Measured 2026-09-24 night: a 9201 archive is 8.18 GB (22.6 GB uncompressed); PVE prunes AFTER a successful backup, so a target must hold keep-last + 1 archives during the run — lowering the retention alone does not un-stick a full target. **Fix direction:** the agent predicts the archive size (last archive × margin) against the target's free space before a whole-box backup, skips with a reason reported to the hub (operator event), and the controller's backup page shows the sentence. `audits/night-2026-09-25/A/A4-hp-backup-space.txt` | **READY — P2; owner: CC (agent + controller)** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` | **OPEN — P3; owner: CC** |
|
||||
| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` | **READY — P3; owner: CC (hub) / operator (Peti's Cloudflare leftovers, if any)** |
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
|
||||
Reference in New Issue
Block a user