chaos-night fixes: rulings recorded, guide reordered, three rows closed, two filed
gates / gates (push) Successful in 21s

Architecture: 08 records ruling A (45 m, the round-9 arithmetic, the cost) and
the two new event types with their audiences; 03 records the slow counter as
built; 07 records the restore-record persistence as a REVERSED design for the
restore record only; CONTEXT.md carries the day's rulings.

Guide: the recovery code moves after the first apps and waits for the yellow bar.

Register: R-549 and R-550 closed PROVEN-LIVE; R-546 closed on red-proofed tests
with its live walk owed by R-551 (no Tier-0 box is paused AND agent-connected).
R-552 filed: an interrupted-restore notice for a removed app never clears -
found in my own v0.246.0 after the release was built.

Evidence: Part A (hub prints 45m/1h30m), Part C delivery on HP and N100, B.4(a)
live proof and its teardown.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 11:00:29 +02:00
parent 06334e164c
commit 3c1882a0a4
11 changed files with 163 additions and 13 deletions
+3
View File
@@ -318,3 +318,6 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-487** | **A removed app whose backups were kept was listed on neither backup page (P2).** Closed in controller **v0.242.0** (`d698ce3`): the local lists are keyed on the DRIVES the way R-237 keyed the off-site list on the store — `ListRemovedAppUnits` walks `backups/primary/` on the system path and every connected registered drive; the Mentések page lists the unit after the deployed rows („Eltávolítva — visszaállítható", one action), the Visszaállítás picker lists it in its own group, `GET /api/backup/snapshots` answers for it, and the restore opens the unit where it sits (`primaryUnitDirFor` — a unit kept on a data drive was unreachable before, the fallback named the system path). Proven live on the scratch guest 9202 with an opengist throwaway: removed with data, backups kept → row + picker + API answered; „Visszaállítás a mentésből" reinstalled it running. Red-proofs: lister inert, picker 404, wrong unit dir, row not built, row not rendered, picker not rendered — all fail. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
| **R-490** | **The monitoring page's memory-distribution card never rendered — `/api/system/info` was 404 (P3).** Closed in controller **v0.242.0** (`d698ce3`): an exact-path mount ahead of the web layer's `/api/system/` prefix, and `systemInfo` reads the default storage path like every other reader of the empty global. Live on 9202: 200 with the drive figures. Red-proofs: mount removed, fallback removed — both fail. **The global's deletion stays deferred → R-492.** `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
| **R-491** | **Removing an app left its update hold in the store, so a reinstall started held (P2).** Closed in controller **v0.242.0** (`d698ce3`): `removeStack` clears an UPDATE hold (`Settings.ClearUpdateHold`, never an R-379 restore hold), logged. Proven live on 9202: a held opengist removed → the store no longer carries the hold, the app redeployed without refusal. Red-proof: the removal without the clear fails the wiring test. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
| **R-546** | **The first-hour guide and the reminder bar sent the household to create their recovery code before the box could (P2).** Closed in controller **v0.246.0** (`0fe315b`): the bar consults the agent's OWN preflight `ok` (all five blocking items, not a copy of `pbs_storage_id`), cached 60 s, probed only while paused, held back while not ready; `/backup/escrow` shows a waiting card that polls and reloads; `POST /api/escrow/start` refuses 409 before staging (the direct path that produced the raw `-storage` stderr); unknown readiness keeps the bar. The guide moves the step after the first apps: „amikor a sárga sáv megjelenik”. Red-proofs: bar held back, waiting card, start refusal; controls ready and unknown. **Proven by tests through ServeHTTP, NOT live** — no Tier-0 box is paused and agent-connected (**R-551**); chaos night measured live the ~17-minute red window and its self-heal. | **CLOSED 2026-09-17 — PROVEN (tests); live walk owed by R-551** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
| **R-549** | **The staleness alarm's budget was two report cycles, so one failed push spent all of it (P2).** Closed by operator ruling A (2026-09-17): `alerting.stale_threshold` 30 m → **45 m**, `node_down`/`host_down` at 90 m (`manifests/hub.yaml`, commit `06334e1`). The dashboard's customer status hardcoded 30 m / 1 h and would have disagreed with the alarms, so hub **v0.117.0** (`37ae31f`) makes `controllerStatus` read the same value — red-proof `TestControllerStatus_FollowsConfiguredThreshold` (report 40m old: status warn, want ok). **Proven live:** the running hub printed `node_stale after 45m0s, node_down after 1h30m0s` and `host_stale after 45m0s, host_down after 1h30m0s` at 08:22Z. **Reasoning kept:** the threshold is configuration and every reader — both checkers, host status, customer status — reads the one value; a dead box now pages 15 minutes later, a cost the ruling accepts. `audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
| **R-550** | **The restore record was in-memory only: after the machine stopped, nothing told the household their restore did not finish (P2).** Closed in controller **v0.246.0** (`0fe315b`) by operator ruling „fix” — a reversal of the in-memory design for the restore record only (cooldowns stay in memory): `restore-status.json` in DataDir, atomic at both ends of an op; a record still running at startup becomes a failed, interrupted result per app, shown on `/backups/restore` until that app's next restore, raised once as `restore_interrupted` (hub v0.117.0, household). Red-proofs: record across restart (`StartedAt:0001-01-01`), `main()` wiring (AST), startup helper, page card. **Proven live on demo-hp 9201:** a throwaway homebox restore killed 2 s in; after the supervisor's restart the status read `ok:false … megszakadt … interrupted:true`, the card showed, the event reached the hub (HTTP 200, stored under demo-hp); a second restore cleared the card. **Known gap filed: R-552** (a removed app keeps its notice). `audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
+2 -3
View File
@@ -728,11 +728,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. **CLOSED 2026-09-16 — controller v0.245.0, both halves proven live.** The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism. (a) **The household is asked:** while the off-site tier is configured and its escrow is not complete, every authenticated page carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking `/backup/escrow`. It is the R-241 bar, second instance — same session-cookie dismissal, back at the next visit, gone for good when escrowed; **no second banner system**. It hangs off `executeTemplate`, the single render choke point, so it cannot reach only the pages someone remembered. (b) **The tier-1 sentence renders by state:** `driveFilesNoteFor` takes `tier3State`'s own vocabulary — `active` → „védi", `escrow_pending` → „védené … a helyreállítási kód létrehozásáig szünetel" + the route, no off-site and no second drive → „nincs másolat" + both ways out. **Measured live on 0.245.0:** on a paused box (9202, off-site configured through the product's own endpoint, `escrow_state=pending`) the bar renders on /dashboard, /launcher, /backups/apps and /settings; a manual `POST /backup/offbox/run` is refused by the fork-4 gate („A távoli mentés a kulcs letétbe helyezésére vár.") with `last_run=None, snapshot_count=None`; and a throwaway class-A app's row reads „…védené … szünetel" with „védi"=0. On an escrowed box (9201) the bar is absent on all three pages and the row reads „védi". Both red-proofed (the bar test fails on BOTH pages with the one hook line removed; the sentence test quotes the exact v0.244.0 promise when the state is ignored). The first-hour guide now asks for the code right after the dashboard password and before the first app. Evidence: `audits/evidence-recovery-code-2026-09-16/`. | **CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live** |
| **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** |
| **R-545** | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** |
| **R-546** | **[P2-MEDIUM] The first-hour guide sends the household to create their recovery code at a moment when the box cannot yet do it — and the new reminder bar urges them there on every page.** MEASURED 2026-09-16/17 on a fresh box (`tester-1-022354`, guest 9201, controller 0.245.0, agent 0.131.0, installed from the published ISO 1.28.0). `VOLUNTEER-first-hour.md` §6 — added hours earlier in controller v0.245.0 — places „A helyreállítási kód" immediately after the dashboard password and **before the first app**, because until it is done the off-site copy does not run. At exactly that point the ceremony FAILS: `POST /api/escrow/start` → 200, then `GET /api/escrow/status` → `detail: "exit 2: … selftest=escrow-create requires -storage <pbs-storage-id> (or escrow.pbs_storage…"`, and `POST /api/escrow/claim` → **409** „A folyamat jelenlegi állapotában a kód nem kérhető le." **Cause, measured on both sides:** the hub had auto-provisioned the DR descriptor at 20:19 (no press — see R-534/R-511's acknowledged-delete path) and its Backup & DR panel itself read „descriptor provisioned … **waiting** · ceremony possible once the descriptor is applied on the box"; the box had no PBS storage (`pvesm status` = local + local-lvm only) and `/etc/felhom-agent/agent.json` had **no `escrow` section at all**. **It is a TIMING gap and it self-heals:** a watcher left the box alone and polled — `pbs_storage` and `escrow.pbs_storage_id` both became `felhom-pbs` at **20:35:16Z, ~17 minutes after the bind**; the retried ceremony then passed every preflight item and the claim returned 200 (83-character code, entropy 129.2 bits), and `escrow_state` flipped to `escrowed`. **Why it still matters:** for those ~17 minutes the R-543 reminder bar (also v0.245.0) is on *every* page telling the household to do the one thing that refuses, and nothing on the page says „wait a few minutes" — the volunteer meets a stderr fragment about a `-storage` flag. **Fix shape (one of):** have the escrow page/bar consult `preflight` and say „a doboz még készül — pár perc múlva próbáld újra" while `pbs_storage_id` is unset; or move the guide's step to after the first app; or make the bar appear only once preflight is green. **No product code was changed tonight** (validation run). Evidence: `audits/evidence-chaos-night-2026-09-17/phase0-escrow-failure.txt` and `phase0-escrow-retry.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller copy + guide timing)** |
| **R-547** | **[P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: `disk_critical` is defined at ≥95 % used, but the fill-watch runs once a day.** MEASURED 2026-09-17 (chaos night) on a fresh box (`tester-1-022354`, controller 0.245.0): the customer guest’s root filesystem was held at **96 % for ten minutes** (29 G used, 1.5 G free) and **no alarm of any kind fired** — checked twice, once by the round’s own runner and once independently after the fill was released. Cause, established from the ladder BEFORE the round rather than after: `fillwatch` runs **daily at 03:30** plus once ~90 s after a controller start, so a ten-minute window contains no check unless a restart lands inside it. The timing here was almost comic — the controller restarted at 21:28 after the previous round’s power cut, so its one opportunistic check ran about **twenty seconds before** the disk filled. Meanwhile all twelve apps kept serving and the background household loop logged 12 operations with 0 failures, so nothing else would have hinted at it either. **This is the ladder working as designed, not a missed alarm** — the fill-watch is a daily sweep, not a monitor. It is filed because the honest answer to „would the household be told their disk is full?” is **no, unless the controller happens to restart while it is full**, and that answer is not written down anywhere. **Fix shape (one of):** sample the fill more often than daily (a cheap `statfs` on the 5-minute health pass would do it); or say plainly in `08-alarm-ladder.md` that a transient full disk is out of scope. Evidence: `audits/evidence-chaos-night-2026-09-17/round-3.txt`. | **READY — rank P3-LOW; owner: CC** |
| **R-548** | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. | **READY — rank P3-LOW; owner: CC** |
| **R-549** | **[P2-MEDIUM] The staleness alarm's budget is exactly two report cycles, so ONE failed push spends all of it.** MEASURED 2026-09-17 (chaos night, round 9) on `tester-1-022354` (controller 0.245.0): hub reachability was removed for ten minutes from the VM's side. The controller built its 23:08:42Z report, retried the push **three times over 1 m 40.8 s**, and gave up at 23:10:23Z (`hub push failed after 3 attempts`) - **31 seconds before the link returned**. Nothing is queued, and that is correct: a report is a snapshot, so the next one carries the same truth, and the controller says so itself ("backing off (the 15-min cycle still reconciles)"). **The arithmetic is the finding.** Last good report 22:53:43Z; next scheduled 23:23:42Z; measured cadence **15m0s**; `node_stale` trips at **30 minutes**. The gap is **29 m 59 s** - one second inside the threshold. So a single missed push spends the whole staleness budget, and any ordinary jitter pages the operator about a box that is healthy, serving every app, and has already repaired itself unaided. **The product behaved correctly throughout:** no false alarm fired, no app stopped, and both the hub link and the host-agent link recovered by themselves the moment the block lifted. What is filed is the margin, not a misbehaviour. **Fix shape:** either set the staleness threshold to a clear multiple of the cadence (three cycles, not two), or let a push that has failed all three attempts retry once off-cycle instead of waiting for the next scheduled report. **Honest caveat:** the cut was injected by the drill and also severed the controller from its host agent, which a real ISP outage would not do - but the report arithmetic above depends only on the hub being unreachable. Evidence: `audits/evidence-chaos-night-2026-09-17/round-9.txt`. | **READY - rank P2-MEDIUM; owner: CC** |
| **R-550** | **[P2-MEDIUM] The restore record is in-memory only: after the machine stops, the status surface is blank and nothing tells the household their restore did not finish.** MEASURED 2026-09-17 (chaos night, round 10) on `tester-1-022354` (controller 0.245.0): an app restore was accepted at 23:26:08Z (`302`, flash "Visszaallitas elindult") and the guest's host was hard-reset **four seconds later**, mid-write. Afterwards `/api/backup/restore-status` - the endpoint the restore page itself polls - answers `{"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}}`: the Go **zero value**, with **no `last` field at all**, while the page's own script renders "<operation> sikertelen." from `st.last.message` and "<operation> folyamatban" from `st.op`. So the UI has a last-operation branch with nothing to populate it after a restart. On disk, in the real data directory, no restore, lock or state file exists and **no file at all was modified in the reset window**. **The box recovered perfectly** - 26/26 containers back in 150 s, boot reconciliation naming the app it recovered, every front door serving, one true `controller_started` alarm and no false one. **CORRECTED 2026-09-17T00:24Z:** this row first claimed no restore surface existed at all, citing four endpoints that 404'd - all four were paths I GUESSED, and all four were wrong. The real routes (`/backups/restore`, `/stacks/<name>/backup`, `/api/backup/restore-status`) came from the controller's own rendered links. The corrected claim is narrower and stands on the endpoint's own answer. **Honest limit:** only four seconds elapsed, so `started_at` may be zero because the restore never truly began rather than because the reboot erased it; the pre-reset log is unrecoverable (the stream holds zero lines before the reboot). Either way nothing tells the customer. **Fix shape:** persist the last restore outcome the way the backup tiers already persist theirs, and have boot reconciliation mark an in-flight restore as abandoned so the page can say so. Evidence: `audits/evidence-chaos-night-2026-09-17/round-10.txt`. | **READY - rank P2-MEDIUM; owner: CC** |
| **R-551** | **[P3-LOW] No Tier-0 box can put the escrow ceremony in the state R-546 fixes — paused AND connected to its agent — so the readiness branches are proven only by tests.** FOUND 2026-09-17 while live-validating controller v0.246.0 (R-546). The branches (reminder bar held back while the agent's preflight is not `ok`; the waiting card on `/backup/escrow`; `POST /api/escrow/start` refused 409 before staging) need a box whose off-site tier is configured, whose escrow is NOT done, and whose controller reaches the agent. Measured on demo-hp: **9201** reaches the agent but is escrowed (the bar is off by design there; making it paused would be hand-set state on the standing demo box, and a real ceremony would supersede its live escrow — the hub keeps ONE `host_escrow` row per host, `host_id PRIMARY KEY`); **9202** is paused-capable but has **no local-API token at all** (its `bootstrap.json` holds only `schema`, `customer.id`, `disposition`), so its readiness is always UNKNOWN and the bar always shows. The fresh-bind window where this state occurs naturally (~17 min, chaos night Phase 0) needs a fresh install. **What IS proven:** five tests driving the real pages and handler through `ServeHTTP` with a fake agent, each red-proofed; and chaos night measured live that the agent's preflight is red for ~17 minutes after a bind and turns green by itself (`evidence-chaos-night-2026-09-17/phase0-escrow-*`). **Fix shape:** give the scratch guest a local-API token through the agent's own provisioning path, or walk R-546 on the next fresh install. | **READY - rank P3-LOW; owner: CC** |
| **R-552** | **[P3-LOW] An interrupted-restore notice for an app that is then REMOVED stays on the restore page for ever.** FOUND 2026-09-17 by CC reviewing its own controller v0.246.0 (R-550) during live validation. The per-app notice (`Manager.opInterrupted`, persisted in `restore-status.json`) is cleared in exactly one place — `BeginRestoreOp` for that app (`internal/backup/opstatus.go`) — and `removeStack` (`internal/api/router.go`) never touches the restore record. So a household that answers „A visszaállítás megszakadt … indítsd el újra" by REMOVING the app instead of restoring it keeps a „Megszakadt visszaállítás" card about an app that no longer exists. Measured shape, not hypothetical: on 9201 the notice cleared only when homebox was restored again (08:54:50Z, card count 0) — the teardown deliberately took that path before removing it. **Fix shape:** `removeStack` clears the app's notice (a `ClearInterruptedRestore(stack)` beside the existing update-hold clear, R-491's precedent), with a wiring test. Not fixed in v0.246.0: found after the release was built; one release per repo per session. | **READY - rank P3-LOW; owner: CC** |
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |