From d53bd08442db25ab5c828a1518ef7aeb841e9a1a Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 24 Sep 2026 17:06:09 +0200 Subject: [PATCH] R-672 session: audit README (not done first, wrong claims), STATUS, capability map, register 344->341, topic REPORT Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT-r672-2026-09-24.md | 13 +++ STATUS.md | 37 ++++---- .../architecture/00-capability-map.md | 1 + .../D/00-9202-deploy-0.270.0.txt | 2 + .../audits/r672-2026-09-24/D/10-r679-live.txt | 9 ++ .../r672-2026-09-24/D/20-r681-press.txt | 7 ++ .../audits/r672-2026-09-24/D/21-r681-kill.txt | 1 + .../r672-2026-09-24/D/22-r681-after.txt | 19 ++++ .../D/23-r681-press-mealie.txt | 1 + .../r672-2026-09-24/D/24-r681-kill-mealie.txt | 2 + .../D/25-r681-after-mealie.txt | 17 ++++ .../D/26-r681-reinstall-remove.txt | 12 +++ .../r672-2026-09-24/D/29-repoint-drill.txt | 8 ++ .../r672-2026-09-24/D/30-r669-live.json | 84 +++++++++++++++++ .../audits/r672-2026-09-24/D/30-r669-live.log | 20 +++++ .../r672-2026-09-24/D/31-r669-files.txt | 9 ++ .../r672-2026-09-24/D/32-r669-order.txt | 16 ++++ .../r672-2026-09-24/D/33-r669-checkB.txt | 8 ++ .../audits/r672-2026-09-24/D/40-teardown.txt | 14 +++ .../r672-2026-09-24/D/50-floor-0.270.0.txt | 2 + .../r672-2026-09-24/D/51-boxes-arriving.txt | 11 +++ .../audits/r672-2026-09-24/README.md | 89 +++++++++++++++++++ .../r672-2026-09-24/redproofs/D-r669.txt | 9 ++ .../r672-2026-09-24/redproofs/D-r674.txt | 5 ++ .../r672-2026-09-24/redproofs/D-r679.txt | 5 ++ .../r672-2026-09-24/redproofs/D-r681-page.txt | 5 ++ .../r672-2026-09-24/redproofs/D-r681.txt | 5 ++ documentation/backlog/CLOSED-ITEMS.md | 4 + documentation/backlog/OPEN-ITEMS.md | 4 - 29 files changed, 395 insertions(+), 24 deletions(-) create mode 100644 REPORT-r672-2026-09-24.md create mode 100644 documentation/audits/r672-2026-09-24/D/00-9202-deploy-0.270.0.txt create mode 100644 documentation/audits/r672-2026-09-24/D/10-r679-live.txt create mode 100644 documentation/audits/r672-2026-09-24/D/20-r681-press.txt create mode 100644 documentation/audits/r672-2026-09-24/D/21-r681-kill.txt create mode 100644 documentation/audits/r672-2026-09-24/D/22-r681-after.txt create mode 100644 documentation/audits/r672-2026-09-24/D/23-r681-press-mealie.txt create mode 100644 documentation/audits/r672-2026-09-24/D/24-r681-kill-mealie.txt create mode 100644 documentation/audits/r672-2026-09-24/D/25-r681-after-mealie.txt create mode 100644 documentation/audits/r672-2026-09-24/D/26-r681-reinstall-remove.txt create mode 100644 documentation/audits/r672-2026-09-24/D/29-repoint-drill.txt create mode 100644 documentation/audits/r672-2026-09-24/D/30-r669-live.json create mode 100644 documentation/audits/r672-2026-09-24/D/30-r669-live.log create mode 100644 documentation/audits/r672-2026-09-24/D/31-r669-files.txt create mode 100644 documentation/audits/r672-2026-09-24/D/32-r669-order.txt create mode 100644 documentation/audits/r672-2026-09-24/D/33-r669-checkB.txt create mode 100644 documentation/audits/r672-2026-09-24/D/40-teardown.txt create mode 100644 documentation/audits/r672-2026-09-24/D/50-floor-0.270.0.txt create mode 100644 documentation/audits/r672-2026-09-24/D/51-boxes-arriving.txt create mode 100644 documentation/audits/r672-2026-09-24/README.md create mode 100644 documentation/audits/r672-2026-09-24/redproofs/D-r669.txt create mode 100644 documentation/audits/r672-2026-09-24/redproofs/D-r674.txt create mode 100644 documentation/audits/r672-2026-09-24/redproofs/D-r679.txt create mode 100644 documentation/audits/r672-2026-09-24/redproofs/D-r681-page.txt create mode 100644 documentation/audits/r672-2026-09-24/redproofs/D-r681.txt diff --git a/REPORT-r672-2026-09-24.md b/REPORT-r672-2026-09-24.md new file mode 100644 index 00000000..976fb984 --- /dev/null +++ b/REPORT-r672-2026-09-24.md @@ -0,0 +1,13 @@ +# REPORT — R-672/R-673: restore test off, 9201 repaired, agent v0.133.0 + hub v0.124.0, controller v0.270.0 + +Full record: `documentation/audits/r672-2026-09-24/README.md` (opens with "not done, or changed" and the brief's +wrong claims). + +- **Hub v0.124.0** (`2d24931`, manifest `068e065`, live): a thin pool is judged on the worse of data and metadata, + critical at 90 %; `storage_fill_*` cooldown per pool (host/storage), 6 h. Two red-proofs; suite green. +- **Agent v0.133.0** released (tag `9bdb4da`), not delivered; **controller v0.270.0**, floor 0.270.0, both demo boxes + arrived. +- **Docs:** `03` §8 (the preflight, the operator rulings, the -1 trap), `08` §6.2 (thin pool at 90 %), capability map + row, register (R-684 opened; R-669/R-674/R-679/R-681 closed; R-672/R-673 updated), CONTEXT, STATUS. +- **Register:** 344 → 341 rows (685,662 → 685,148 B). +- `unproven.py --summary`: see the session's final message (unchanged — 35 of 55 not walked). diff --git a/STATUS.md b/STATUS.md index 4be97686..949056aa 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,28 +1,25 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-24 (night shift, run in daytime from 11:07). The demo boxes are on controller 0.269.1 — except demo-hp's customer box, which is damaged and needs you (first item under "What needs you").** +**Updated 2026-09-24 (evening). The demo boxes are on controller 0.270.0. The HP customer box is repaired. A fixed host agent (0.133.0) is released and waits for your signature. Until then the automatic restore test is off on both demo boxes.** -**Decisions I took on my own (you may reverse either).** -- **A second controller release tonight (0.269.1).** The first one (0.269.0) had a flaw I found in its own live test: when the catalog re-tested an image, the box wrote the new image into a running app's file before anyone pressed Update, so any restart would have changed the image with no backup and no undo. I fixed it and released again rather than send the flaw to the fleet. Now an installed app keeps its image until an Update moves it. -- **A crash loop is 6 restarts in 10 minutes, not the 10 you wrote.** Measured: Docker slows a steady crash loop to about one restart a minute, so 10 would never catch it. No healthy app in any test record restarted more than once while starting. +**Decisions I took on my own (you may reverse them).** +- **The space check uses the real, uncompressed size, not the backup file size.** The rule in your brief ("file size × 1.2 + 5 GB") would NOT have stopped last night's accident. The HP box's backup file is 7 GB, but restoring it writes 22.6 GB. +- **To switch the test off I used the value −1, not 0.** In this setting, 0 means "every 6 hours". Only a negative number turns it off. +- **The hub got a small release too (0.124.0).** A thin disk pool now counts as urgent at 90 % (data or its bookkeeping part), with one alarm per pool per 6 hours. Before, the alarm came only at 95 %, up to 15 minutes late, and one full pool could silence another. -**What I tested, and it worked.** -- **The second drive now brings an app with files back whole** (your ruling). Tested four times on the scratch machine, twice under a power cut or a nearly full disk. Every time: the account and the files came back, a file the household had edited later was kept, and nothing was deleted. -- **A stranded app can only be removed keeping its data** (your ruling). The box refuses to delete the data, in both languages. -- **The box stops a crash loop or a memory storm** (your ruling), tells the household and you, and Start gives one more try. A second stop within a day says support is informed. -- **Exact image fingerprints.** The box runs exactly the image the catalog tested, and says "update available" when a newer tested image exists for the same version name. -- **Automatic updates: measured, not built.** A test caller updated three apps (up to three steps each) and set a failing one aside in about 7 minutes. The build plan is written with the numbers. -- **A chaos hour, 12 rounds** (power cuts, killed controller, Docker restarts, a full disk, backups). Updates resumed or undid themselves correctly every time. +**What I did, and it worked.** +- **The HP customer box is repaired.** I stopped it, checked both disks (small damage fixed, second check clean), and started it. All apps came back, it took the current version by itself, and the hub shows it online. Two apps' cache files (Redis) had been cut off by the full disk; with your yes I kept copies and repaired them. Both apps run. +- **The box's own crash-loop stop worked for real:** while the cache was broken, the box stopped those two apps and told the hub, by itself. +- **The new agent never starts a restore test that does not fit.** Tested on the HP box: it refused, said why, and created nothing. +- **Four controller fixes (0.270.0):** Update on an app that is already current is refused, with no restart. An install cut off by a restart is now cleaned up, reported and shown on the page. After a restore, the box keeps the right health check. One log line is corrected. -**What broke.** -- **demo-hp's customer box (not caused by tonight's work).** This morning the host's automatic restore test copied a whole guest onto the same full disk pool. The pool filled, and the customer box's disks became read-only. I freed the pool with an agent restart (the product's own clean-up). I could not repair the box itself. -- **An install interrupted by a restart is lost without a word**, and a remove interrupted the same way is left half-done. Both recorded, not fixed. -- After a failed update is restored, the box can keep the failed version's health check, and a later undo then fails for no reason. Recorded. -- Not done: moving more apps in the catalog. It needs a throwaway test machine, and this session could not delete one afterwards. +**What broke, or is not done.** +- **The HP box's own backups have been failing every day since yesterday.** Its host disk has 4 GB free, but one backup needs 7 GB, and three old ones are kept. Nothing fixes this by itself. +- **No full restore test fits on the HP box today** (it needs 30 GB free, 22 GB is free). The new agent will refuse it every time, correctly. So a passing test on the HP box is not proven yet. -**Rows.** 16 opened, 7 closed. The list went from 335 to 344. +**Rows.** 1 opened, 4 closed, 2 updated. The list went from 344 to 341. **What needs you.** -1. **Repair demo-hp's customer box:** stop it, check its two disks, start it. If you do nothing, it keeps running but cannot save anything: no backups, no logs, no updates, and it stays on the old controller. -2. **Decide how the restore test may use disk space** (for example: it must check free space first and never use the pool of the box it tests). If you do nothing, the next scheduled restore test can fill the pool again, on any box with a small disk. -3. **Optional:** the two decisions above. If you do nothing, they stay as built. +1. **Sign the agent update (0.133.0) for demo-hp and demo-felhom.** Version 0.133.0, sha256 `3aa303452b8c6be58573d00af01a0ab4a0d97e4f885ffecd0144a18ac24e69b6`, one `felhom-opsign -op agent_update` job per box, the same way as 0.131.0. If you do nothing, the boxes keep the old agent and the automatic restore test stays off. +2. **After the new agent is on a box, turn the restore test back on.** On that host, remove the key `restore_test_eval_interval_seconds` from `backup` in `/etc/felhom-agent/agent.json` (the saved copy `agent.json.pre-r672` has it absent), then `systemctl restart felhom-agent`. If you do nothing, no backup is proven restorable on that box. +3. **Decide what to do with the HP box's backup storage:** keep fewer old backups, back up somewhere else, or add disk. If you do nothing, its whole-box backup keeps failing every night. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index bf547296..8e45f4a0 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -92,6 +92,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | **A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code** | controller **v0.242.0**, agent **v0.130.0**, hub **v0.112.0**, golden **0.242.0**, installer ISO **1.26.1** | **PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app** | `audits/DRILL-fresh-install-0242-2026-09-14.md`, evidence `audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md`. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back **byte-identical with its attachment** 32 s after restore; a hard power cut returned every app on the **same version** with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. **ONE intervention a volunteer could not make:** the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — **R-494**), so the dashboard was reached by LAN address. **And no instruction exists to begin with (R-493).** **WHAT IT DOES NOT CLAIM:** the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. **2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL.** `audits/DOORSTEP-walk-1270-2026-09-14.md`. Held again on a new customer (`tester-1`, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, `pvebanner` masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. **BIGNIGHT 2026-09-14/15 (ISO 1.27.1):** the first hour re-walked with the real mailbox — self-bind link and mailed setup code both worked; 2 interventions (no auto bind mail R-509; tunnel 502 R-510); then twelve apps, a month of routines and nine faults: `audits/BIGNIGHT-household-month-2026-09-14.md`. **DRILL 2026-09-16 on controller 0.243.0 / agent 0.131.0 / hub 0.115.0 / golden 0.243.0 / ISO 1.27.1 — the WALK half is now PROVEN-LIVE with ZERO interventions, and the BACKUP half is narrowed, not proven.** `audits/DRILL-prove-fixes-0243-2026-09-16.md`. A fresh nested box was installed from the published ISO, landed on this drill's own golden **by checksum** (`e2d1843c…c10a`, no self-update), and was walked as a volunteer: the connect e-mail and self-bind page, the claim, **the tunnel from outside (302→200, R-510 CLOSED)**, the data drive, the file manager with its own generated password (`admin`/`admin` refused 401 before any dashboard action), four apps deployed and used, and „Mentés most” with **26 s** of app downtime. **Interventions: 0** — the two operator presses (self-bind send, PBS re-issue) were pre-declared and counted apart, and the four moments that look like help were my own API-driving errors or my own damage, each named in the drill's interventions table. **The automatic connect e-mail is PROVEN with a real mailbox:** a host delete at 12:22:59Z produced the „Kösd össze a Felhom dobozodat” mail at **12:23:00Z**, with `selfbind_link_sent (host delete)` on the timeline (R-509). **WHAT THIS DOES NOT CLAIM, and it is the backup half of this row's own sentence:** on a one-drive box with no off-site tier — what every fresh install is — the household's files are in **no backup at all** (the whole-guest tiers exclude `mp8` by design; the file leg lives at tier 2/3, both unset), the app-backup page nevertheless reads „DB + Konfig + Adatok” (**R-537**), and a restore reports success while leaving Nextcloud listing five photos that return `Sabre\DAV\Exception\NotFound` — after making the app's own trash, which still held every byte, unreachable (**R-538**). **Also still open from this walk:** the off-site tier cannot be provisioned at all (**R-534**, P1, an ep0 grant), the console keeps its pairing banner after bind and claim (**R-535**), and „app installed” is emitted at accept time (**R-536**). Faults re-measured: controller killed during a deploy → back in **37 s**; three more kills 20 min apart → **61/41/61 s**, none accumulating (**R-531**); two reboots 60 s apart → everything back in **124 s** with the boots NOT counted against the brake; the claim page locks after the **second** wrong code and mails the operator truthfully. **2026-09-16 (evening) — THE BACKUP HALF IS NOW WALKED TOO, on controller 0.244.0 / hub 0.116.0 / golden 0.244.0 / ISO 1.28.0 (built, unpublished).** `audits/evidence-backup-promise-2026-09-16/`. A second fresh box (VM 335) was installed from the BUILT image and walked to the end of the sentence this row could not previously finish: **five photos in → deleted the way a child would → the old route REFUSED and touched nothing → the off-site restore returned them → they OPEN, byte-identical (sha256 5/5, negative control)**. The refusal reads „Ez a mentés nem tartalmazza az alkalmazás fájljait… a fájlok így a helyükön maradnak" and names the route that can help; the app was `running` before and after; the wastebasket was untouched. The off-site restore ran in two steps — a verification copy that states „A meglévő adatok változatlanok", then a reconstitution counting „5 fájl és 3 adatkötet és az adatbázis". **Delivery is part of it:** the box landed on agent 0.131.0 + controller 0.244.0 (the golden baked and vouched the same day) with no hand upgrade, and `app_deploy_started` (19:15:34) / `app_deployed` (19:16:23) finally mean different things. **The self-bind half needed NO operator press** — the box registered itself and the bind used the mail the hub sent itself after the morning's host delete (R-509, proven twice today: 12:23:00Z and 18:17:46Z, each one second after a host delete). **WHAT IT STILL DOES NOT CLAIM — one operator press and one day-one gap.** The PBS-DR cascade stopped at the refusal R-511 documents („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action") and needed the operator to press it; that press then SUCCEEDED because of the ep0 grant given this morning (R-534 CLOSED, R-511 CLOSED). And off-site ON by default is not off-site WORKING: a fresh box sits at „Kulcsletétre vár" until the household performs the escrow ceremony, which nothing asks them to do, while the tier-1 row already promises that copy (**R-543**, P1). **2026-09-16 (late evening) — that last gap is CLOSED, controller v0.245.0 (R-543).** The pause is the zero-knowledge escrow design and was not touched; what was missing was the ASK. Every authenticated page now carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking the ceremony (the R-241 bar, second instance, hung on the single render choke point), the tier-1 sentence renders by tier-3 STATE („védené … szünetel" while paused, „védi" when running), and the first-hour guide asks for the code right after the dashboard password and before the first app. **Measured on two boxes running 0.245.0:** paused box — bar on four pages, `POST /backup/offbox/run` refused by the fork-4 gate with no snapshot written, app row „védené" and „védi"=0; escrowed box — no bar anywhere, row „védi". `audits/evidence-recovery-code-2026-09-16/`. **WHAT IT STILL DOES NOT CLAIM:** the ask has not been walked by an actual volunteer from the written guide — the sentence is proven, the human following it is not. So the journey now reads: **files protected from day one, once the household writes down the recovery code the box asks them for on every page.** | Rows **R-493 … R-500**, R-534 … R-545 | | **Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts** | controller v0.197.0 | **PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it** | On demo-hp, an app on the **system-data fallback** bound `/mnt/sys_drive/userdata/media/books` while the capture set looked in `/mnt/sys_drive/felhom-data/userdata/media/books`: the declared-mandatory directory was in **no** snapshot and the run reported `ok` (R-203). After v0.197.0 the capture log reads `1 mandatory path(s)` and the file is **listed inside the snapshot** — `restic ls -l latest --tag calibre-web` → `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt`. Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 and the v0.197.0 CHANGELOG | **WHAT IT DOES NOT CLAIM.** It covers CAPTURE, not RESTORE: **no file has ever been restored from an off-site backup after a wipe** (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports **`incomplete`** rather than `ok`, so this row's guarantee is one the status can express. **Widened 2026-08-06 (controller v0.205.0, R-234):** the same verdict now also covers an app skipped ENTIRELY — until then a missing declared FOLDER made the run incomplete while an app with no recovery unit at all still reported `ok`, so the smaller gap moved the verdict and the bigger one did not. **This row still claims CAPTURE, not that a newly-selected app is protected by the next run** — for a DEPLOYED app it is (the run's own pre-dump phase writes the unit, measured 2026-08-06), and for an undeployed one it is not and the card now says so || **Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE** — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | **PROVEN-LIVE (2026-07-26)** | `audits/SPIKE-r82-phase0-2026-07-26.md`; per-repo CHANGELOGs/REPORTs. **Restore round-trip on demo-hp:** `--selftest=restore-test` against `felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z` → `pass:true`, `verified:"boot+running"`, **`mount_parity:"ok"`** (`mp0=/var/lib/docker 50G`, `mp1=/mnt/sys_drive 20G`, mp8/mp9 throwaway stand-ins for the archived binds), `source_tier:"pbs"`, 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | **This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL.** That tier was PROVEN-LIVE as `applied` since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held **zero, ever** — "applied and empty", the R-39 shape one level quieter. **What earns PROVEN-LIVE here:** (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it **restores into a bootable, mount-complete guest** — `mount_parity` is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) **the multi-tier quiesce ran through the real UI endpoint** (`POST /api/guest-backup/trigger`, authed+CSRF) and produced **exactly ONE stop/start pair with BOTH backups inside it** — `quiescing 1 stack(s)` 17:01:39 → local done 17:02:56 *"next tier may start (app still quiesced)"* → felhom-pbs snapshotted 17:03:06 → `unquiescing` 17:03:06. **App downtime 1m27s for both tiers**, and the app came back healthy. **Known gaps, recorded not hidden:** the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → **R-82** | | **Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard** | agent v0.104.0 → **v0.121.0**, hub v0.77.0 → **v0.91.0** | **PROVEN-LIVE (2026-08-03)** | per-repo CHANGELOGs; `backlog/SPEC-r85-phase4-5-2026-07-26.md`. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | **The row above is earned by a MANUAL `--selftest=restore-test`; this one is about the SCHEDULED path, and the distinction is the whole point.** Before R-85 the scheduler could only ever see `cfg.Backup.BackupTarget()`, so the offsite tier was never a candidate — and a failed restore-test was a `[WARN]` line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — `restore_test_failed` (broken now) and `restore_test_stale` (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. **Why this was NOT PROVEN-LIVE until now:** rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → **R-85**. **R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon**, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. **UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer.** The scheduler's own log: `15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven"` → `proxmox-backup-client restore --crypt-mode=encrypt` under the agent's own token → `15:25:08 gate decision class=guest_destroy guest=990000 allowed=true` → `15:25:14 scratch guest torn down` → **`15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1`**. A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored, booted, verified and destroyed in 635 s, unattended. **The three things a timer could not show**, all verified after it: the state names THAT archive (`{"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}`); a second evaluation reports `due=false … is already proven` and runs nothing; and **an agent restart runs nothing**, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from `pct list`, **zero** `990000` volumes in `lvs`, and the hub-side `restore_tests[]` entry deliberately RETAINED (it IS the proof the staleness check reads). **R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path:** the proof above reached the hub only because no restart intervened — `restore_tests[]` came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (`0 restore-tests` on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. **THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly.** Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was **unprovable**, and on demo-hp equally: the agent's token had no ACL on `/storage/felhom-backup`, the storage both boxes configure as `local_backup_target`, so the content API answered `{"data":[]}` through the token while root listed three archives. The scheduler skipped it as *"no settled archive yet"* — which is exactly what a brand-new tier reports — so nothing ever said so (**R-185**). **CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0.** The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now **asks whether it may read each tier it depends on** instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on **while the box was still blind**. **THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed.** Four SCHEDULED runs, nothing triggered by hand: **demo-felhom** host tier `…2026_08_02-04_42_14.tar.zst` **passed in 83.8 s** at 00:55, offsite `…2026-07-28T04:49:43Z` **passed in 540.4 s** at 06:55; **demo-hp** host tier `…2026_08_02-04_49_29.tar.zst` **passed in 109.3 s** at 02:05, offsite `…2026-07-28T19:19:45Z` **passed in 300.1 s** at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; `pct list` and `lvs` show zero `990000` afterwards on both, and both boxes' `local-lvm` returned to their pre-run figures (1.95 % and 30.83 %). **Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule:** never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. **The proofs reached the hub**, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds **two** entries, one per tier, and the `local` one can only have come from the persisted state because the in-memory store held only that morning's offsite run. **SCOPE, stated because one box proving something does not make it a fleet property:** this covers `demo-felhom` and `demo-hp`. The tester's box is untested and untouched. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md` | +| **A restore-test can never fill the box's disk: it sizes the restore first (uncompressed), keeps off the tested guest's pool when it can, refuses what does not fit, and retries a failed clean-up on a timer** | agent **v0.133.0** (RELEASED, NOT DELIVERED), hub **v0.124.0** | **BUILT + red-proofed; the refusal PROVEN on demo-hp (2026-09-24); the full-test path NOT proven live** | `audits/r672-2026-09-24/C7a…C7b`, `audits/r672-2026-09-24/redproofs/C-*` | The 2026-09-24 incident (R-672) filled demo-hp's pool and turned 9201 read-only. Live: both margins refuse 9201's 21.1 GiB restore (22.1 GiB free, needs 30.3). No full restore-test fits demo-hp under 80 % pool use, so a pass after the change is unproven; the scheduled test is OFF on both demo hosts until the agent is delivered (operator ruling). | | Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive | hub v0.61.0, felhom-tenantsync v1.1.0 | **PROVEN-LIVE (external teardown, incl. two real firings)** | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — two live firings, both host-delete-first, on two different customers** (`demo-vm-felhom` 15:49:57, `demo-felhom` 16:08:51): every leg `ok` (`claim`, `db_purge`, `descriptor`, `hetzner`, `pbs`), escrow acked separately, each completing in 8–9 s (`hub-state.txt` `customer_resets`). The **Hetzner sub-account destruction is now verified against the live pool box** — and produced the run's sharpest lesson: **a sub-account is an access-control object, not a data object.** Deleting it left its `/home` intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (**a finding by S7's own criterion**) and why RESET now needs a base-dir purge → **R-32**. Prior: hub v0.61.0 REPORT; **ep0 live drill 2026-07-17** (throwaway `drill-reset-01` with a real backup: deprovision `deleted:true` destroyed the namespace + backup group + token, idempotent re-run `deleted:false`, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests | External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. **Live-clicked 2026-07-18** (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. **R-25b CLOSED (hub v0.69.0, 2026-07-21):** the Danger-zone DELETE is now the guided full-teardown cascade that runs this very sequence as its middle leg — see the row below | | **Customer DELETE cascade** (top lifecycle tier): one guided operator action → `hosts → RESET → residue → purge`; host rows deleted (custody DEMOTED), full external teardown (Hetzner repo destroyed, PBS namespace/token revoked, tunnel + zone removed), then the customer record and ALL escrow ciphertext purged | hub v0.69.0 | **UNIT-PROVEN; live leg PENDING** | `hub/internal/web/customer_delete_test.go` — leg ORDER observed from inside leg 2 (hosts already gone, customer row still present, custody still retained); 9 fail-closed gate cases each asserting zero mutations + zero external calls + no journal row; resume-after-external-failure converges; `purgeEscrow` custody semantics; preview leaks no secret. **5 red-proofs** (ack gate, stale-preview gate, ONLINE-host gate, leg order inverted, `purgeEscrow=true`) | Three acknowledgements + typed customer-id + stale-preview check + ONLINE-host refusal, ALL before any write. Ruling-3 preserved BY CONSTRUCTION (leg 2 never sees a host row); custody purged EXACTLY ONCE, in leg 3. **Coupling:** hub-only — no agent/controller/catalog change; the cascade calls the same service paths as manual host-delete and standalone RESET, so their rules move together. **v0.70.0 (2026-07-21):** added the **residue** leg — `GetCustomers()` is REPORT-derived, so before it a fully deleted customer stayed on the Customers list and its report stream kept the staleness/offsite checkers alerting (live: `demo-vm-felhom` deleted 07-18, still emailing `offsite_stale` on 07-21). The leg also purges the credential-bearing `appliance_registrations` + `selfbind_tokens`. **Ghost customers (config row already gone) are now deletable** — 404 means "nothing here", not "no config row"; the Hetzner/descriptor legs record `skipped_no_config`. **Gap:** the end-to-end live leg on a scratch customer (external Hetzner teardown observed from outside) is not yet run | | Uninstall: KEPT-vs-WIPED statement, secret purge, enrolled-drive handling | installer | **PARTIAL** | `DRILL-GL6-2026-07-08` Phase 1/5 (KEPT-vs-WIPED printed verbatim; drive data intact ×3); GL-4 code | Secret purge (GL6-F1 `.bak` residue) fixed v1.12.0; enrolled-drive `mnt-*.mount` units survive (GL6-F2, open); cluster-aware `felhom_guests` guard + saferemove cost warning missing → R-9 | diff --git a/documentation/audits/r672-2026-09-24/D/00-9202-deploy-0.270.0.txt b/documentation/audits/r672-2026-09-24/D/00-9202-deploy-0.270.0.txt new file mode 100644 index 00000000..84094594 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/00-9202-deploy-0.270.0.txt @@ -0,0 +1,2 @@ +gitea.dooplex.hu/admin/felhom-controller:0.270.0 +gitea.dooplex.hu/admin/felhom-controller:0.270.0 Up 20 seconds (healthy) diff --git a/documentation/audits/r672-2026-09-24/D/10-r679-live.txt b/documentation/audits/r672-2026-09-24/D/10-r679-live.txt new file mode 100644 index 00000000..e23c1868 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/10-r679-live.txt @@ -0,0 +1,9 @@ +privatebin: installed= {"privatebin": {"ref": "privatebin/pdo:2.0.6", "digest": "sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5", "at": "2026-09-22T19:44:56Z"}} catalog= {'privatebin': 'privatebin/pdo:2.0.6'} badge= [{'title': 'This app is running the newest version available.', 'text': 'Up to date'}] +hu Update -> 409 {"ok": false, "data": {"reason": "already_current"}, "error": "Ez az alkalmazás már a legfrissebb elérhető változatot futtatja — nincs mit frissíteni."} +en Update -> 409 {"ok": false, "data": {"reason": "already_current"}, "error": "This app is already running the newest version available — there is nothing to update."} +after: updating= False phase= None state= running +2026/09/24 14:51:18 router.go:591: [INFO] [api] update requested for stack: privatebin +2026/09/24 14:51:18 update.go:359: [ERROR] [stacks] update privatebin REFUSED (already_current): installed equals the catalog head on every service and no newer tested digest (installed=map[privatebin:{privatebin/pdo:2.0.6 sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5 2026-09-22T19:44:56Z}] catalog=map[privatebin:privatebin/pdo:2.0.6]) +2026/09/24 14:51:18 router.go:591: [INFO] [api] update requested for stack: privatebin +2026/09/24 14:51:18 update.go:359: [ERROR] [stacks] update privatebin REFUSED (already_current): installed equals the catalog head on every service and no newer tested digest (installed=map[privatebin:{privatebin/pdo:2.0.6 sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5 2026-09-22T19:44:56Z}] catalog=map[privatebin:privatebin/pdo:2.0.6]) + diff --git a/documentation/audits/r672-2026-09-24/D/20-r681-press.txt b/documentation/audits/r672-2026-09-24/D/20-r681-press.txt new file mode 100644 index 00000000..1c43b7e5 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/20-r681-press.txt @@ -0,0 +1,7 @@ +before: deployed= False +no actualbudget image cached + +16:51:41 deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +-rw-r--r-- 1 root root 21 Sep 24 14:51 .felhom-install-pending +-rw------- 1 root root 251 Sep 24 14:51 app.yaml + diff --git a/documentation/audits/r672-2026-09-24/D/21-r681-kill.txt b/documentation/audits/r672-2026-09-24/D/21-r681-kill.txt new file mode 100644 index 00000000..e16bcd36 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/21-r681-kill.txt @@ -0,0 +1 @@ +16:51:52 killed pid 157525 rc=0 diff --git a/documentation/audits/r672-2026-09-24/D/22-r681-after.txt b/documentation/audits/r672-2026-09-24/D/22-r681-after.txt new file mode 100644 index 00000000..fda3efba --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/22-r681-after.txt @@ -0,0 +1,19 @@ +stack: deployed= True state= running install_interrupted= None deploy_error= None +2026/09/24 14:51:53 manager.go:1554: [INFO] [stacks] actualbudget actualbudget/actual-server:26.9.0@sha256:552beab3dec8c93d46b8b9245612d63c3f123b8a45063a474f53e229b17621d3 running Up 3 seconds (health: starting) +2026/09/24 14:51:54 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) +2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/24 14:51:54 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/24 14:51:54 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for actualbudget → /mnt/sys_drive/felhom-data/backups/primary/actualbudget (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0) +--- files +-rw------- 1 root root 523 Sep 24 14:51 app.yaml +-rw-r--r-- 1 root root 1484 Sep 24 14:51 applied-compose.yml +drwxr-xr-x 2 root root 4096 Sep 24 14:51 applied-meta +--- containers +actualbudget Up 13 seconds (healthy) +(end) + +hu page: None +en page: None diff --git a/documentation/audits/r672-2026-09-24/D/23-r681-press-mealie.txt b/documentation/audits/r672-2026-09-24/D/23-r681-press-mealie.txt new file mode 100644 index 00000000..a32c4739 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/23-r681-press-mealie.txt @@ -0,0 +1 @@ +16:53:05 deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} diff --git a/documentation/audits/r672-2026-09-24/D/24-r681-kill-mealie.txt b/documentation/audits/r672-2026-09-24/D/24-r681-kill-mealie.txt new file mode 100644 index 00000000..bfe37fa1 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/24-r681-kill-mealie.txt @@ -0,0 +1,2 @@ +16:53:05 /opt/docker/stacks/mealie/.felhom-install-pending +16:53:05 killed pid 158798 rc=0 diff --git a/documentation/audits/r672-2026-09-24/D/25-r681-after-mealie.txt b/documentation/audits/r672-2026-09-24/D/25-r681-after-mealie.txt new file mode 100644 index 00000000..c6b08be8 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/25-r681-after-mealie.txt @@ -0,0 +1,17 @@ +stack: deployed= False state= not_deployed install_interrupted= True deploy_error= interrupted by a controller restart before it finished — install it again +2026/09/24 14:51:54 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) +2026/09/24 14:53:05 router.go:431: [INFO] [api] Deploy requested for stack: mealie +2026/09/24 14:53:05 [INFO] [stacks] SaveAppConfig: saved config for mealie +2026/09/24 14:53:05 deploy.go:383: [INFO] [stacks] Deploying stack mealie with 2 env vars: [DOMAIN, SUBDOMAIN] +2026/09/24 14:53:05 manager.go:1590: [INFO] [stacks] Deploying stack mealie — checking 1 images... +2026/09/24 14:53:06 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) +2026/09/24 14:53:07 install_interrupted.go:70: [WARN] [stacks] install mealie was INTERRUPTED by a controller restart — removing what it started (volumes kept) and reporting it (R-681) +2026/09/24 14:53:07 [WARN] [stacks] 1 install(s) interrupted by the restart were resolved and reported: [mealie] +--- files +-rw-r--r-- 1 root root 21 Sep 24 14:53 .felhom-install-interrupted +-rw------- 1 root root 255 Sep 24 14:53 app.yaml +--- containers +(end) + +hu page: A telepítés egy újraindítás miatt félbemaradt, és a doboz eltávolította, amit elkezdett. Nyomd meg újra a Telepítés gombot. +en page: The installation was interrupted by a restart, and the box removed what it had started. Press Install again. diff --git a/documentation/audits/r672-2026-09-24/D/26-r681-reinstall-remove.txt b/documentation/audits/r672-2026-09-24/D/26-r681-reinstall-remove.txt new file mode 100644 index 00000000..b8b121fe --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/26-r681-reinstall-remove.txt @@ -0,0 +1,12 @@ +16:53:51 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +16:55:11 [1] deployed, controller state=running, pinned={'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'} +reinstall ok= True deployed= True install_interrupted= None +no install markers + +16:55:16 [X] stop -> 200 {'ok': True, 'message': 'Stack mealie stop completed'} +16:55:48 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'mealie', 'volumes_removed': ['mealie_mealie_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalm +16:55:56 [X] after remove: deployed=False leftovers='/opt/docker/stacks/mealie' +remove -> 200 +no mealie volumes +no leftovers + diff --git a/documentation/audits/r672-2026-09-24/D/29-repoint-drill.txt b/documentation/audits/r672-2026-09-24/D/29-repoint-drill.txt new file mode 100644 index 00000000..2c6fccaf --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/29-repoint-drill.txt @@ -0,0 +1,8 @@ + repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + sync_interval: 15m + token: + username: "admin" +hub: +update: + health_timeout: 90s + diff --git a/documentation/audits/r672-2026-09-24/D/30-r669-live.json b/documentation/audits/r672-2026-09-24/D/30-r669-live.json new file mode 100644 index 00000000..a019d531 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/30-r669-live.json @@ -0,0 +1,84 @@ +{ + "1-installed": "stack: \napplied: \nunit: ()", + "2-before-update": "stack: \napplied: \nunit: ()", + "2-update": { + "accepted": true, + "http": "202", + "phases": [ + { + "t": 0.0, + "phase": "backing-up", + "label": "Biztonsági mentés készül a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 2.1, + "phase": "safety-dump", + "label": "Adatbázis pillanatkép…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 3.1, + "phase": "pulling", + "label": "Új verzió letöltése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 4.1, + "phase": "copying", + "label": "Az adatok másolása a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 16.4, + "phase": "verifying", + "label": "Működés ellenőrzése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 106.6, + "phase": "undoing", + "label": "Visszaállítás az előző változatra…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 125.0, + "phase": "undone", + "label": "Visszaállítva az előző változatra", + "updating": false, + "error": null, + "hold": null + } + ], + "duration_s": 125.1, + "final_phase": "undone", + "update_error": null, + "hold_reason": null, + "state": "running" + }, + "2-after-update": "stack: \napplied: \nunit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)", + "CHECK_A_unit_keeps_pinned_probe": false, + "3-before-restore": "stack: \napplied: \nunit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)", + "3-restore": { + "op": "tier2-unit-restore", + "stack": "wishlist", + "ok": true, + "message": "A(z) wishlist: 2 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-24 16:57).", + "finished_at": "2026-09-24T15:00:33.459543237Z" + }, + "3-after-restore": "stack: \napplied: \nunit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose)", + "CHECK_B_restore_resets_applied": false, + "log": "2026/09/24 14:57:56 pin.go:373: [INFO] [stacks] update wishlist: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/wishlist/docker-compose.yml (wishlist=ghcr.io/cmintey/wishlist:latest)\n2026/09/24 14:58:08 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_data → wishlist_wishlist_data.pre-update-20260924T145808Z in 459ms\n2026/09/24 14:58:09 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_uploads → wishlist_wishlist_uploads.pre-update-20260924T145808Z in 440ms\n2026/09/24 14:59:40 update.go:1274: [INFO] [stacks] update wishlist: phase undoing\n2026/09/24 14:59:40 undo.go:513: [WARN] [stacks] update wishlist: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health check failing))\n2026/09/24 14:59:42 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1\n2026/09/24 14:59:57 undo.go:633: [INFO] [stacks] update wishlist: event app_update_undone\n2026/09/24 14:59:57 undo.go:561: [INFO] [stacks] update wishlist: UNDONE in 18s — the previous version is running on the data from before the update (the app's health check passed)\n2026/09/24 15:00:19 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1\n2026/09/24 15:00:33 restore_unit.go:464: [INFO] [backup] Restore-from-unit completed: wishlist — 2 volume(s) of 2 listed, 0 database(s) of 0 listed\n" +} \ No newline at end of file diff --git a/documentation/audits/r672-2026-09-24/D/30-r669-live.log b/documentation/audits/r672-2026-09-24/D/30-r669-live.log new file mode 100644 index 00000000..ef8f8825 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/30-r669-live.log @@ -0,0 +1,20 @@ +16:57:21 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +16:57:41 [1] deployed, controller state=running, pinned={'wishlist': 'ghcr.io/cmintey/wishlist:v0.67.1'} +16:57:45 [1] stack: | applied: | unit: () +16:57:46 drill: wishlist: failing step ghcr.io/cmintey/wishlist:v0.67.1 -> ghcr.io/cmintey/wishlist:latest + probe 8999 +16:57:53 [2] stack: | applied: | unit: () +16:57:53 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +16:57:53 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None +16:57:55 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +16:57:56 + 3.1s phase=pulling label=Új verzió letöltése… err=None hold=None +16:57:57 + 4.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +16:58:10 + 16.4s phase=verifying label=Működés ellenőrzése… err=None hold=None +16:59:40 + 106.6s phase=undoing label=Visszaállítás az előző változatra… err=None hold=None +16:59:58 + 125.0s phase=undone label=Visszaállítva az előző változatra err=None hold=None +17:00:02 [2] after: stack: | applied: | unit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose) +17:00:03 drill: wishlist image back to v0.67.1 (probe left at 8999) +17:00:18 [3] stack: | applied: | unit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose) +17:00:18 [3] POST /backup/tier2/unit-restore -> 302 +17:00:34 [3] restore: {'op': 'tier2-unit-restore', 'stack': 'wishlist', 'ok': True, 'message': 'A(z) wishlist: 2 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-24 16:57).', 'finished_at' +17:00:37 [3] after: stack: | applied: | unit: (/mnt/sys_drive/felhom-data/backups/primary/wishlist/compose) +17:00:40 CHECK A (unit keeps the pinned probe): False CHECK B (restore resets the applied record): False diff --git a/documentation/audits/r672-2026-09-24/D/31-r669-files.txt b/documentation/audits/r672-2026-09-24/D/31-r669-files.txt new file mode 100644 index 00000000..ab61126c --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/31-r669-files.txt @@ -0,0 +1,9 @@ +== .felhom.yml (2026-09-24 15:00:19) + - type: http + port: 3000 +== applied-meta/.felhom.yml (2026-09-24 15:00:19) + - type: http + port: 3000 +== /mnt/sys_drive/felhom-data/backups/primary/wishlist/compose/.felhom.yml (2026-09-24 14:57:55) + - type: http + port: 3000 diff --git a/documentation/audits/r672-2026-09-24/D/32-r669-order.txt b/documentation/audits/r672-2026-09-24/D/32-r669-order.txt new file mode 100644 index 00000000..5f2432e9 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/32-r669-order.txt @@ -0,0 +1,16 @@ +2026/09/24 14:57:22 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1 +2026/09/24 14:57:46 sync.go:416: [INFO] [sync] Updated wishlist/.felhom.yml +2026/09/24 14:57:53 update.go:1274: [INFO] [stacks] update wishlist: phase backing-up +2026/09/24 14:57:53 update_guard.go:360: [INFO] [backup] update pre-backup for wishlist: starting (DB dump → volume dump → unit capture → Tier 2) +2026/09/24 14:57:55 update_guard.go:410: [INFO] [backup] update pre-backup for wishlist: volume dump OK +2026/09/24 14:57:55 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for wishlist → /mnt/sys_drive/felhom-data/backups/primary/wishlist (images=1, secrets-referenced=0, data_keys=0, port +2026/09/24 14:57:55 update_guard.go:416: [INFO] [backup] update pre-backup for wishlist: recovery unit captured (0 database dump(s)) +2026/09/24 14:57:55 update_guard.go:425: [INFO] [backup] update pre-backup for wishlist: complete in 2.038s +2026/09/24 14:57:56 undo.go:285: [INFO] [stacks] update wishlist: the undo copy will hold 2 named volume(s), 0.3 MiB +2026/09/24 14:58:08 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_data → wishlist_wishlist_data.pre-update-20260924T145808Z in 459ms +2026/09/24 14:58:09 undo.go:435: [INFO] [stacks] update wishlist: copied wishlist_wishlist_uploads → wishlist_wishlist_uploads.pre-update-20260924T145808Z in 440ms +2026/09/24 14:59:40 update.go:1274: [INFO] [stacks] update wishlist: phase undoing +2026/09/24 14:59:40 undo.go:513: [WARN] [stacks] update wishlist: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health che +2026/09/24 14:59:42 pin.go:93: [INFO] [stacks] pin wishlist: wishlist=ghcr.io/cmintey/wishlist:v0.67.1 +2026/09/24 14:59:57 undo.go:633: [INFO] [stacks] update wishlist: event app_update_undone +2026/09/24 14:59:57 undo.go:561: [INFO] [stacks] update wishlist: UNDONE in 18s — the previous version is running on the data from before the update (the app's health check passed) diff --git a/documentation/audits/r672-2026-09-24/D/33-r669-checkB.txt b/documentation/audits/r672-2026-09-24/D/33-r669-checkB.txt new file mode 100644 index 00000000..412977c8 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/33-r669-checkB.txt @@ -0,0 +1,8 @@ +BEFORE restore: +.felhom.yml 15:01:22: port: 8999 +applied-meta/.felhom.yml 15:00:19: port: 3000 +POST /backup/tier2/unit-restore -> 302 +restore: {'op': 'tier2-unit-restore', 'stack': 'wishlist', 'ok': True, 'message': 'A(z) wishlist: 2 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás f +AFTER restore: +.felhom.yml 15:01:38: port: 3000 +applied-meta/.felhom.yml 15:01:38: port: 3000 diff --git a/documentation/audits/r672-2026-09-24/D/40-teardown.txt b/documentation/audits/r672-2026-09-24/D/40-teardown.txt new file mode 100644 index 00000000..de6da8a7 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/40-teardown.txt @@ -0,0 +1,14 @@ +remove wishlist -> 200 +no wishlist volumes +no leftovers + + sync_interval: 15m + token: + username: "" +hub: +0 + +drill=c8025093d23c live=c8025093d23c + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git +0 +gitea.dooplex.hu/admin/felhom-controller:0.270.0 diff --git a/documentation/audits/r672-2026-09-24/D/50-floor-0.270.0.txt b/documentation/audits/r672-2026-09-24/D/50-floor-0.270.0.txt new file mode 100644 index 00000000..1caa3ee4 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/50-floor-0.270.0.txt @@ -0,0 +1,2 @@ +POST floor 0.270.0 -> 303 /configuration?flash=floor_set +read back: ['0.270.0'] diff --git a/documentation/audits/r672-2026-09-24/D/51-boxes-arriving.txt b/documentation/audits/r672-2026-09-24/D/51-boxes-arriving.txt new file mode 100644 index 00000000..5f11b158 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/D/51-boxes-arriving.txt @@ -0,0 +1,11 @@ +Thu Sep 24 17:03:45 CEST 2026 +== N100 9201 + LANG = "en_US.UTF-8" +gitea.dooplex.hu/admin/felhom-controller:0.270.0 Up 5 seconds (healthy) +== demo-hp 9201 + LANG = "en_US.UTF-8" +gitea.dooplex.hu/admin/felhom-controller:0.270.0 Up 5 seconds (healthy) +== hub +2026/09/24 17:03:33 [INFO] Global controller-version floor set to "0.270.0" (declared MinAgent "0.131.0") +2026/09/24 17:03:35 [INFO] managed floor SERVED for demo-felhom: floor 0.270.0, agent requirement "0.131.0" from declared (golden 0.258.0) +2026/09/24 17:03:36 [INFO] managed floor SERVED for demo-hp: floor 0.270.0, agent requirement "0.131.0" from declared (golden 0.258.0) diff --git a/documentation/audits/r672-2026-09-24/README.md b/documentation/audits/r672-2026-09-24/README.md new file mode 100644 index 00000000..f3404453 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/README.md @@ -0,0 +1,89 @@ +# The restore test off, demo-hp 9201 repaired, agent v0.133.0 + hub v0.124.0, controller v0.270.0 (2026-09-24 evening) + +Brief: "the restore test switched off on the demo boxes, the HP customer box repaired, the host agent fixed so a +restore test can never fill a box's disk again; then four controller leftovers". Architecture read: `03-host-agent.md` +§8/§10, `07`, `08` §6.2, `09` §3/§6.4, `runbooks/target-selection.md`. + +## Not done, or changed + +- **Agent v0.133.0 is released, NOT delivered.** It reaches a box only with the operator's signed `agent_update` job + (R-530). Until then the scheduled restore test stays OFF on both demo hosts. +- **Part A used `-1`, not the brief's `0`.** `restore_test_eval_interval_seconds: 0` means "the 6-hour default" + (`config.go` `RestoreTestEvalInterval`); only a negative value disables. The start-up line says "restore-test cadence + disabled" on both hosts. With the cadence off no evaluation runs, so there is no "nothing due" line to quote. +- **Part C rule 1 changed:** "archive size × 1.2 + 5 GiB" would NOT have prevented the incident. The size is the + UNCOMPRESSED one (vzdump log / PBS snapshot size). +- **Part C rule 4 needed a hub release (v0.124.0).** The existing `storage_fill_critical` fired only at 95 %, on data + only, per household per hour, and only at the next 15-minute report. Now: a thin pool is critical at 90 % of data + or metadata, one alarm per pool per 6 hours, and the agent asks for a report the moment a pool crosses 90 %. +- **Live case (b) refused and case (c) did not run.** On demo-hp, 9201's restore needs 30.3 GiB and the pool has 22.1 + GiB free. So no full restore test fits under the 80 % limit. The forced clean-up failure needs a scratch guest, + which cannot be created there. Both are covered by unit tests only. +- **Part B found damage fsck could not see:** the Redis append-only files of docmost and romm were cut off by the full + pool. With the operator's yes, each Redis folder was copied aside and cut at the last complete write with + `redis-check-aof --fix` (2,943 B / 6,631 B dropped). Meanwhile the box had stopped both apps by itself (decision 28). +- **R-674 has no live proof:** R-679 now refuses the only case that reached that log line. +- **R-679 removed a "repair path"** that a test comment named: a same-version Update. Restart is the repair path. + +**Interventions: 0** on product behaviour. The Redis repair was an operator-approved data repair. **One agent +release, one hub release, one controller release.** + +## Claims in the brief that turned out wrong (or right) + +1. *An eval interval of 0 disables only the schedule and leaves the on-demand test working* — **half wrong.** 0 is + the 6-hour default; negative disables. The on-demand path (`--selftest=restore-test`) does not read the cadence — + TRUE, and used live. +2. *`pct fsck` can check the `/var/lib/felhom` mount by `--device`* — **TRUE** (`--device mp0`). +3. *demo-hp has a second eligible storage for the restore test* — **WRONG.** `nvme-scratch` takes `rootdir`, but the + agent holds only the inherited `Datastore.Audit` there, not `Datastore.AllocateSpace`. +4. *R-673's lock came from the pool-full event* — **WRONG.** The 06:59 and 07:42 backups failed with "No space left on + device" on `local`, the host ROOT disk (R-684). The lock is from the 07:42 clean-up. The pool filled at 10:35. +5. *"Archive size × 1.2 + 5 GiB" is enough* — **WRONG.** 9201: a 6.9 GB file, a 22.6 GB restore. +6. *A pool nearly full is only a log line* — **partly wrong.** The hub DID mail `storage_fill_critical` at 100 %, but + late and at the generic bands. + +## Part A — the scheduled restore test off (`A-restore-test-off.txt`) +Both hosts: only `backup.restore_test_eval_interval_seconds` changed (saved copy `/etc/felhom-agent/agent.json.pre-r672`, +verified equal apart from that key). Start-up: "backup: restore-test cadence disabled". Peti's box untouched: it gets +the fix only through a signed agent. + +## Part B — 9201 repaired (`B1…B9`) +Pool 58.9 % data / 2.65 % metadata. rootfs = `vm-9201-disk-0`, `/var/lib/felhom` = `mp0` (`vm-9201-disk-1`). Stop 4.8 s. +`pct fsck`: rootfs replayed its journal; mp0 fixed 17 "deleted inode has zero dtime" and one orphan block; second +pass clean on both (rc 0). Start; the controller took the 0.269.1 floor by itself; both disks writable; 21 containers +up; hub `/hosts`: demo-hp ONLINE. Then the Redis repair above; docmost and romm started through the product (Start +lifted the box's hold), 0 restarts. + +## Part C — agent v0.133.0 (tag `9bdb4da`, sha256 `3aa30345…e69b6`, verified by an anonymous download) + hub v0.124.0 +Red-proofs (`redproofs/C-*`): no preflight → the 2026-09-24 restore issued again; file size used → "the compressed +file size was used"; no timer → "the leaked scratch was not destroyed by the timer"; sweep without the gate → "the +sweep ran while a backup held the gate"; no 90 % edge → "0 report requests, want 1"; skips dropped → "a space refusal +never reached the host report"; hub generic bands → a 91 % pool only warns, a metadata-full pool raises nothing; hub +old key → "a second pool filling in the same hour was silenced". +Live on demo-hp (on-demand self-test, cadence off, pool 58.99 % before and after, nothing created): (a) factor 10 → +"needs 215.5 GiB free, has 22.1 GiB", exit 4; (b) normal → "restoring 21.1 GiB (vzdump log: total bytes written) +needs 30.3 GiB free, has 22.1 GiB", exit 4. +R-673: the stale-lock sweep now runs every 10 minutes, under the one-heavy-operation gate. + +## Part D — controller v0.270.0, floor 0.270.0 (`D/`) +R-679 (409 `already_current`, hu + en, live), R-681 (interrupted install reported and cleaned, live), R-669 (the unit +keeps the pinned health check; a restore resets the applied record; live), R-674 (unit only). Both demo boxes +arrived on 0.270.0 in ~12 s. + +## Register +Before this session **344 rows / 685,662 B**; after **341 rows / 685,148 B**. Opened R-684; closed R-669, R-674, +R-679, R-681; R-672 and R-673 updated to "fixed in v0.133.0, awaiting delivery". + +## Teardown — three layers +- **Machine:** 9202 back on the live catalog with its saved config, on controller 0.270.0 (the floor); the throwaway + apps (actualbudget, mealie, wishlist) removed through the product — no containers, volumes or markers left. + 9201 (demo-hp) repaired, running, on 0.270.0. +- **Host:** demo-hp and demo-felhom: only the agent config key changed (saved copies beside them); the live-test + binary and its test config deleted from demo-hp `/tmp`; `pct list` 9201 + 9202, nothing created. +- **Hub:** v0.124.0 deployed; floor 0.270.0 (MinAgent 0.131.0); drill repo reset to the live catalog. + +## To turn the restore test back on (after the signed agent arrives on a box) +On that host: `cp /etc/felhom-agent/agent.json.pre-r672 /etc/felhom-agent/agent.json && systemctl restart felhom-agent`. +The saved copy had no `restore_test_eval_interval_seconds` key (= the 6-hour default). Check the journal for "restore-test +scheduler starting". On demo-hp the test will then REFUSE 9201's restore for space, correctly, until the pool has room +(R-684 is about the separate backup storage). diff --git a/documentation/audits/r672-2026-09-24/redproofs/D-r669.txt b/documentation/audits/r672-2026-09-24/redproofs/D-r669.txt new file mode 100644 index 00000000..2f2c7399 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/redproofs/D-r669.txt @@ -0,0 +1,9 @@ +# Red-proofs D3 — (1) the restore does not reset applied-meta; (2) the unit captures the stack dir's .felhom.yml (v0.269.1) +--- FAIL: TestR669_RestoreResetsTheAppliedRecord (0.00s) + r669_restore_meta_test.go:31: the next undo would judge with the failed step's probe: +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.009s +--- FAIL: TestR669_UnitCapturesThePinnedVersionsMeta (0.00s) + r669_applied_meta_test.go:43: the unit carries the failing step's probe, not the pinned version's: +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks (cached) +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.008s diff --git a/documentation/audits/r672-2026-09-24/redproofs/D-r674.txt b/documentation/audits/r672-2026-09-24/redproofs/D-r674.txt new file mode 100644 index 00000000..fcba8ff9 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/redproofs/D-r674.txt @@ -0,0 +1,5 @@ +# Red-proof D2 — the AT-THE-HEAD branch removed (v0.269.1) +--- FAIL: TestR674_HeadIsNotCalledOlder (0.00s) + r674_head_test.go:23: the head was called older than the ladder: "the installed version web=nextcloud:34.0.1-apache matches no update_ladder entry (2 entries) — an app older than the ladder has no record to climb; the catalog's current definition" +FAIL +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s diff --git a/documentation/audits/r672-2026-09-24/redproofs/D-r679.txt b/documentation/audits/r672-2026-09-24/redproofs/D-r679.txt new file mode 100644 index 00000000..d1c2ff23 --- /dev/null +++ b/documentation/audits/r672-2026-09-24/redproofs/D-r679.txt @@ -0,0 +1,5 @@ +# Red-proof D1 (controller) — the already-current refusal removed (v0.269.1) +--- FAIL: TestR679_CurrentAppIsRefused (0.00s) + r679_already_current_test.go:22: an app at the head was allowed to update +FAIL +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.015s diff --git a/documentation/audits/r672-2026-09-24/redproofs/D-r681-page.txt b/documentation/audits/r672-2026-09-24/redproofs/D-r681-page.txt new file mode 100644 index 00000000..5a32e27b --- /dev/null +++ b/documentation/audits/r672-2026-09-24/redproofs/D-r681-page.txt @@ -0,0 +1,5 @@ +# Red-proof D5 — the interrupted-install line removed from stacks.html +--- FAIL: TestR681_PageSaysTheInstallWasInterrupted (0.10s) + r681_install_interrupted_page_test.go:24: the apps page is silent about an interrupted install (hu) +FAIL +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.217s diff --git a/documentation/audits/r672-2026-09-24/redproofs/D-r681.txt b/documentation/audits/r672-2026-09-24/redproofs/D-r681.txt new file mode 100644 index 00000000..a1579a8d --- /dev/null +++ b/documentation/audits/r672-2026-09-24/redproofs/D-r681.txt @@ -0,0 +1,5 @@ +# Red-proof D4 — RecoverInterruptedInstalls does nothing (v0.269.1) +--- FAIL: TestR681_InterruptedInstallIsFinishedAndReported (0.00s) + r681_install_interrupted_test.go:33: an interrupted install was not reported: resolved=[] hooks=[] +FAIL +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.009s diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 54eaccc6..0c2bf3b9 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -382,4 +382,8 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-666** | **While support is informed, Remove offered to delete the data (decision 27).** v0.269.0: the dialog reads `keep_data_only` and offers only „remove the app, keep my data"; the API refuses data or backup deletion with 409 (hu + en); the no-whole-copy sentence is informal. Live on 9202 (a one-drive hold). | v0.269.0, 2026-09-24 | `git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-24/A2/10-held-no-copy-keep-data.*` | | **R-667** | **A crash loop never reached the alarm, and nothing stopped it (decision 28).** v0.269.0 + hub v0.123.0: ≥ 6 restarts in 10 min (`RestartCount`, not the resettable `restarting_since`) or an OOM storm → the box stops the app, holds it (`unhealthy_stop`), tells household + operator (`app_stopped_unhealthy`); Start = one more try; a repeat in 24 h says support is informed. Live: gokapi trip 1 and 2; chaos rounds 1, 7, 8, 12. **Rule:** Docker's back-off caps a steady loop at ~1 restart/min, so a threshold must be below 10 per 10 min. | v0.269.0 / hub v0.123.0, 2026-09-24 | `git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-24/A3/`, `08` §6.2 | | **R-668** | **The Tier-2 copy chose a registered path that no longer existed, on the app's own disk (P2).** v0.269.0: the same-disk check fails CLOSED. Live: the next copy went to the SSD. Residual: an older same-disk record counts until the next Tier-2 run replaces it. | v0.269.0, 2026-09-24 | `git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-24/A1/01-find-mirror.txt` | +| **R-669** | **After a failed update ended in a restore, the box kept the FAILED step's health check as the pinned version's, and the next undo held the app for nothing (P2).** v0.270.0: the recovery unit captures the pinned version's `.felhom.yml` (`applied-meta`), not the stack dir's file the sync may already have replaced; a restore makes the restored file the applied record. Live on 9202: the sync wrote the bad probe at 14:57:46, the unit captured at 14:57:55 kept the good one; a second-drive restore turned stack 8999 / applied 3000 into 3000 / 3000, the applied record rewritten at the restore. **Rule:** a copy of an app's definition carries the PINNED version's health check, never the catalog's newest. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/30-33*` | +| **R-674** | **The ladder log called a pin at the head "older than the ladder" (P3).** v0.270.0: it says "AT THE HEAD". Unit test + red-proof only — no product path reaches it since R-679 refuses a current app first. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/redproofs/D-r674.txt` | +| **R-679** | **An Update on an app already at the head ran the whole guarded update — dump, pull, restart (P2).** v0.270.0: `409 already_current` before anything moves, hu + en; a re-tested digest of a floating tag still updates. Live on 9202 (privatebin): both languages refused, no backup, no pull. A test comment had called a same-version Update "the repair path" — Restart is. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/10-r679-live.txt` | +| **R-681** | **An install cut off by a controller restart was lost silently (P2).** v0.270.0: an install marker before the compose-up, removed when the install ends; a marker at start → `compose down` (volumes kept), stale pin records cleared, `app_deploy_failed` with the reason, and the apps page says the install was interrupted (hu + en) until the next install; a finished install only loses the marker. Live on 9202: mealie killed 0 s into its pull → reported, page sentence, nothing left running; reinstall cleared the sentence; actualbudget (killed after it finished) left alone. **Rule:** a long-running act the customer started is journaled so a restart finishes or reports it. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/20-26*` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index e67c864c..9febb67f 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -806,19 +806,15 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** | | **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. | **READY — P2; owner: CC (catalog)** | | **R-657** | **[P2-MEDIUM] "Remove the app, keep my data", then install it again: nextcloud never installs, and the box only says „unhealthy".** MEASURED 2026-09-23 night on 9202 (v0.267.0): nextcloud was removed through the product keeping its drive folder (the remove with data was refused — R-442's fail-closed guard, as on every drill on 9202 — and the product's keep-data remove taken). An hour later a fresh install of nextcloud on the same box: the template binds `${HDD_PATH}/appdata/nextcloud` to `/var/www/html/data`, the kept folder still holds `admin/`, `appdata_*`, `.ncdata` and a 145 MB `nextcloud.log`, and the image's installer loops **„Login is invalid because files already exist for this user — Retrying install..."**; `occ status` reads `installed: false`. The controller records the deploy as done and the app as `unhealthy`; nothing tells the household that their kept files are what blocks the new install, or what to do. **Why it matters:** keep-data is the choice the product OFFERS a household at remove time — and for nextcloud the kept data makes the app uninstallable. **Needs:** decide the product's promise for a reinstall over kept data, per app class (adopt the data? refuse with a sentence? offer to move it aside?); at minimum a deploy-time refusal or warning when the app's drive folder is not empty. Evidence: `audits/night-2026-09-23/chaos/00-nextcloud-reinstall-over-kept-data.txt`. | **READY — P2; owner: operator (the promise) / CC (the build)** | -| **R-669** | **[P2-MEDIUM] After a failed update ends in a restore, the box keeps the FAILED step's health check as the pinned version's — and the NEXT failed update's undo judges the correct old version with it and holds the app.** PROVEN LIVE 2026-09-24 on 9202 (v0.269.0): a failing step (probe port 8999) → hold → the second drive's whole restore cleared the hold, but `applied-meta/.felhom.yml` still carried port 8999 (written by `advancePinTo` at 10:39:21; no restore path rewrites it), and the stack's `.felhom.yml` came back from the unit, which the sync had already filled with the failing step's file. The next failed update's undo put the right version and the right data back and then called it unhealthy on port 8999 → **a false hold** (`not_started`). Data safe; the app stopped for nothing. **Fix direction:** a restore that recreates the definition also sets the applied record to the `.felhom.yml` of the RESTORED pin (the ladder step's `steps/.felhom.yml` for those refs, or the catalog's when the catalog head equals the pin), never the unit's copy. Pre-existing since v0.263.2; not a v0.269.0 regression. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P2; owner: CC (controller)** | | **R-670** | **[P3-LOW] Every undo (and every step-file verify) logs `[ERROR] .felhom.yml backup block rejected … docker-compose.yml unreadable`.** `LoadMetadata` validates the backup block against a compose file that the pre-update-meta directory (undo.go:528, since v0.263.0) and the scratch dir of `loadMetadataFile` (v0.269.0) never hold. The health check it feeds is unaffected; an operator reading ERROR lines after an undo is misled. Seen 10:24:41Z and 10:46:24Z on 9202. **Fix:** a probe-only loader that skips the backup-block validation, or copy the compose beside it. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P3; owner: CC (controller)** | | **R-671** | **[P3-LOW] The undo copies kept by a hold survive the hold's clearing by a restore.** MEASURED 2026-09-24 on 9202: three `nextcloud_*.pre-update-20260924T103924Z` volumes (~0.9 GiB) were still present after the whole restore cleared that hold at 10:41:52Z, and a second set joined them 6 minutes later. Nothing names them on a page; on a small disk they are the difference between the next update's copy fitting or not. **Fix direction:** the restore that clears an update hold removes that hold's undo copies (they describe the state the restore just replaced), logged. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` (volume list) | **READY — P3; owner: CC (controller)** | | **R-672** | **[P1-HIGH] The agent's SCHEDULED restore-test filled the production thin pool and turned a customer guest's disks read-only.** FOUND 2026-09-24 on demo-hp (agent v0.132.0): at 10:29 CEST the restore-test restored 9201's 22 GiB archive as scratch guest 990000 into `local-lvm` — the SAME pool as 9201 — with no free-space check. The pool reached 100 % at 10:35:24 (`out_of_data_space`, `error_if_no_space`); the scratch failed to start; its teardown failed (`lvremove … contains a filesystem in use`) and was "left for Recover" — **which runs only at agent start**, so the pool stayed full for 2.5 h. The agent logged *„a full pool corrupts every guest on it"* every few seconds and took no action. **Consequence, measured:** 9201's controller got `no space left on device` from 10:38, its log stops at 10:39:57, and at 13:12 both its rootfs and `/var/lib/felhom` were READ-ONLY (errors=remount-ro). **Intervention:** an agent restart ran Recover ("destroyed leaked restore-test scratch guest", pool 100 % → 58.9 %). **Still open:** 9201 needs a stop + fsck + start (operator — the session's permission check refused host-level guest operations). **Fix direction:** the restore-test refuses to start unless the target storage has the archive's size plus a margin free and never targets the pool of the guest it tests; a failed teardown retries on a timer, not only at start; the pool-fill WARN pages the operator. `audits/night-2026-09-24/C-02…C-07` **— 2026-09-24 (evening): FIXED IN RELEASE, NOT YET DELIVERED.** Agent **v0.133.0** (tag `9bdb4da`, package verified by download): space preflight before anything is created (UNCOMPRESSED size × 1.2 + 5 GiB, thin metadata, off the tested guest's pool when another storage is eligible — demo-hp has none, `nvme-scratch` carries no grant — unknown refuses, reported as a non-pass `skipped`); failed teardown retried every 10 min; thin pool ≥ 90 % requests an immediate report. Hub **v0.124.0** (live): thin pool critical at 90 % of data or metadata, one alarm per pool per 6 h. **The brief's "archive × 1.2 + 5 GiB" would NOT have prevented this incident** (6.9 GB file → 22.6 GB restore). Six red-proofs; live on demo-hp: both margins REFUSE (21.1 GiB restored needs 30.3 GiB, 22.1 GiB free), pool unchanged. The hub HAD alarmed on the day (`storage_fill_critical` at 100 %, operator mail) — late and at the generic bands. **9201 repaired** (stop, `pct fsck` both disks: 17 half-deleted inodes + an orphan block fixed on mp0; second pass clean; start; floor 0.269.1 taken by itself; ONLINE); two Redis AOF tails the pool cut (2,943 B / 6,631 B) truncated with `redis-check-aof --fix` after copies were kept as volumes — the box had meanwhile stopped docmost and romm itself (decision 28, live). **Scheduled restore-test OFF on both demo hosts** (operator ruling) until v0.133.0 is delivered. `audits/r672-2026-09-24/` | **FIXED IN v0.133.0 — awaiting the operator's signed delivery; restore test OFF on demo hosts; owner: operator (sign) / CC** | | **R-673** | **[P2-MEDIUM] 9201's whole-box backups failed all morning on a stale `snapshot-delete` lock.** demo-hp 2026-09-24: vzdump of 9201 failed at 06:59, 07:17, 07:42 and 08:22 CEST (`CT is locked (snapshot-delete)`), leaving `snap_vm-9201-disk-1_vzdump` (07:42) behind, before R-672's pool fill. The agent's stale-lock scanner ran every 30 s and did not clear it; the lock was gone after the 13:09 agent restart. Cause not read — the session's permission check refused reading the task logs. `audits/night-2026-09-24/C-01-demo-hp-9201-vzdump-errors.txt` **— 2026-09-24 (evening): CAUSE READ, FIX RELEASED.** Not the pool fill (that came at 10:35). The 06:59 and 07:42 backups failed writing the archive to `local` — the host ROOT disk: `zstd: error 70 : … No space left on device` (R-684); the 07:42 failure's cleanup left `snap_vm-9201-disk-1_vzdump` and the `snapshot-delete` lock, and the 08:22 run hit the lock. The agent's own stale-lock recovery cleared it at 13:10, one minute after the agent restart — because it ran ONLY at start. v0.133.0 runs it every 10 min under the one-heavy-operation gate (red-proofed). `audits/r672-2026-09-24/C5-r673-vzdump-logs.txt` | **FIXED IN v0.133.0 — awaiting delivery; root cause R-684; owner: CC** | -| **R-674** | **[P3-LOW] The ladder log says an app's pin „matches no update_ladder entry … older than the ladder" when the pin EQUALS the head.** Seen on 9202 2026-09-24 for nextcloud at the head. Misleading to an operator reading why nothing climbed. **Fix:** say "at the head" when the pin equals the newest `to`. | **READY — P3; owner: CC (controller)** | | **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** | | **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` | **OPEN — P3; owner: CC (watch)** | | **R-677** | **[P3-LOW] For a floating tag re-tested at a new digest, the Behind badge's age reads the TAG's catalog date („1 napja"), not when the new digest was tested (minutes).** Seen 2026-09-24 on 9202 (Part B). Harmless but confusing. **Fix:** for a digest-only move, age from the ladder entry's `tested_at`. `audits/night-2026-09-24/B/10-floating-tag.json` | **READY — P3; owner: CC (controller)** | | **R-678** | **[P3-LOW] After an update step ends `done`, the app's „steps left" and its badge stay STALE until the next scan.** MEASURED 2026-09-24 on 9202 (Part C, first caller run): wishlist read `ladder_steps_left` 1 after its only step, navidrome 2 after both of its steps — for ~50 s, through six presses. A person sees „Frissítés elérhető" for an app that just updated; an automatic caller re-presses. **Fix:** the update's finish refreshes the app's catalog/ladder fields (the same read the scan does). `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P3; owner: CC (controller)** | -| **R-679** | **[P2-MEDIUM] An Update pressed on an app that is already current runs the whole guarded update — backup, pull, restart — and reports `done`.** MEASURED 2026-09-24 on 9202: navidrome at the head was pressed four times by the stale-read caller (R-678); each press made a safety dump and restarted the app (8.5 s of downtime each) and changed nothing. There is no `current` refusal (`UpdatePreflight`, `update.go:366`; `UpdateOrderCurrent` falls through). Harmless by hand, costly for the automatic leg (`09` §6.4.2). **Fix:** the preflight refuses with reason `current` when the pin equals the catalog head and no newer tested digest exists. `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P2; owner: CC (controller)** | | **R-680** | **[P2-MEDIUM] The box does not remember a failed update step — after an undo it offers the same step again.** MEASURED 2026-09-24 on 9202: vikunja's failing step was undone (104 s) and its badge went straight back to „Frissítés elérhető"; only the test caller's own memory stopped a re-press. Decision 15 („a failed step is never pressed again") therefore holds for a person only by their judgement and not at all for the automatic leg. **Fix (part 7 (b), `09` §6.4.2):** record the failed `to` per app; the leg skips it until the catalog's ladder for that app changes; the page says the step was tried and put back. `audits/night-2026-09-24/C/31-night.json` | **READY — P2; owner: CC (controller, with part 7)** | -| **R-681** | **[P2-MEDIUM] An install interrupted by a controller restart is lost SILENTLY — no event, no page sentence, and its half-written files stay.** MEASURED 2026-09-24 on 9202 (Part E round 2, v0.269.1): n8n's install pressed at 11:48:42Z ("Deploying stack n8n … checking 1 images"); `systemctl restart docker` 20 s later took the controller down with it. After the restart the box logged NOTHING about n8n, the app read `not_deployed`, and the household's page showed it as never installed — while `app.yaml` (with its generated `N8N_ENCRYPTION_KEY`), `applied-compose.yml` and `applied-meta/` stayed in the stack dir. A fresh install through the product afterwards worked (211 s) and was not blocked by the leftovers. Compare the update path, which journals and RESUMES after the same accident (round 3/4). **Fix direction:** journal the install like the update; on boot either resume it or say on the page (and in an event) that it was interrupted, and clean the half-written files. `audits/night-2026-09-24/E/round-02*` | **READY — P2; owner: CC (controller)** | | **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** | | **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | | **R-684** | **[P2-MEDIUM] demo-hp's whole-box backups cannot fit on its backup storage — 9201 has had no new whole-box backup since 2026-09-23.** `local` is the host ROOT disk (40 GB, 84 % used, ~4 GB free); it holds three 9201 archives of 5.8–6.9 GB (retention 3) and a new one needs ~7 GB: the 2026-09-24 runs failed `No space left on device` (06:59, 07:42), leaving the stale lock of R-673. Every night now fails the same way until space is made or the retention/target changes. A Tier-0 box, but the same shape exists on any appliance whose local backup target is its root disk. **Needs a decision** (fewer kept archives, another target, or a larger disk). `audits/r672-2026-09-24/C5-r673-vzdump-logs.txt` | **OPEN — P2; owner: operator (the decision) / CC** |