From de3480fcd2bb1d7a3ab9987cbf4d0cedc5795f60 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 6 Oct 2026 20:59:35 +0200 Subject: [PATCH] Night 2026-10-06: R-542 fixed on controller main; R-554, R-562, R-724 need the operator or a design Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../audits/night-burndown-2026-10-06/NIGHT-LOG.md | 4 ++++ .../night-burndown-2026-10-06/ctrl/R-542-red.txt | 15 +++++++++++++++ documentation/backlog/OPEN-ITEMS.md | 8 ++++---- 3 files changed, 23 insertions(+), 4 deletions(-) create mode 100644 documentation/audits/night-burndown-2026-10-06/ctrl/R-542-red.txt diff --git a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md index d5ba74ec..b70225cc 100644 --- a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md +++ b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md @@ -27,3 +27,7 @@ no reboot. | R-895 | **OPENED** — the hub's clean-up-window check trusts the box's own snapshot counts (needs a design + a read-only measurement) | 5 | (this batch) | | (07 §6.5 area) | stale sentence fixed: „until that ships nothing prunes" — the window runs live since 2026-10-05 | 2 | (this batch) | | R-366 | **hub fix on main** (escrow retained when the backup key changes; red-proved) + **design** (`design-R-366.md`, pick B); the old wrong alarm was already gone (R-727) | 45 | (this batch) | +| R-542 | **fixed on controller main** `f65ace0` (registered drives dropped from the candidates; 2 red-proofs `ctrl/`) | 25 | controller `f65ace0` | +| R-554 | needs the operator — the household's recovery-info.txt names the wizard (a customer promise) | 10 | — | +| R-562 | needs the operator's word on the Hungarian number/date format (the row says so) | 2 | — | +| R-724 | needs a design — a new agent route for the guest network + a controller reader (two repos) | 10 | — | diff --git a/documentation/audits/night-burndown-2026-10-06/ctrl/R-542-red.txt b/documentation/audits/night-burndown-2026-10-06/ctrl/R-542-red.txt new file mode 100644 index 00000000..26ba46ad --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/ctrl/R-542-red.txt @@ -0,0 +1,15 @@ +## RED 1: dropRegisteredDrives returns resp unchanged (pre-R-542 pass-through) +--- FAIL: TestR542_RegisteredDriveNotOfferedForFormat (0.00s) + r542_registered_drive_test.go:34: a registered, in-use drive is offered for FORMATTING: initialize=[{Device:/dev/sdb SizeBytes:0 Model: FSType:ext4 DataBearing:true Mountable:false MountSource:/dev/sdb1 DurableID: AlreadyMounted:false} {Device:/dev/nvme0n1 SizeBytes:0 Model: FSType:ext4 DataBearing:true Mountable:false MountSource: DurableID: AlreadyMounted:false} {Device:/dev/sdc SizeBytes:0 Model: FSType: DataBearing:false Mountable:false MountSource: DurableID: AlreadyMounted:false}] (want only /dev/sdc) +--- FAIL: TestR542_UnreadableMountTableEmptiesInitialize (0.00s) + r542_registered_drive_test.go:62: unreadable mount table must empty initialize (non-nil, so JSON is []); got [{Device:/dev/sdb SizeBytes:0 Model: FSType: DataBearing:false Mountable:false MountSource: DurableID: AlreadyMounted:false}] +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.009s +FAIL + +## RED 2: the empty-mount-table branch removed +--- FAIL: TestR542_UnreadableMountTableEmptiesInitialize (0.00s) + r542_registered_drive_test.go:62: unreadable mount table must empty initialize (non-nil, so JSON is []); got [{Device:/dev/sdb SizeBytes:0 Model: FSType: DataBearing:false Mountable:false MountSource: DurableID: AlreadyMounted:false}] +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.008s +FAIL diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 18445fdb..9c1f4fac 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -118,7 +118,7 @@ stopping line that lies. |---|---|---|---|---|---|---|---| | **R-250** | Install & onboarding | P3 | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | — | — | operator | | **R-516** | Install & onboarding | P3 | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** **2026-10-05 (burn-down night): items 1 and 4 FIXED on controller `main`** (`d7efa7b`): the menu says „Hibakeresés” (English „Debug”); the storage pages use the te-form (formal-form ceiling 17 → 14; 116 parity fixtures changed by exactly those bytes). Item 2 was already fixed. **Left:** items 5–6 (apps' own screens) and the rest. **2026-10-06: items 1 and 4 DELIVERED** (controller v0.298.0). **2026-10-06 (burn-down night, later): items 7–10 and 12 FIXED on controller `main`** (`e4774e6`, `2e5ef7e`): the disconnect time in local time; one banner per unplugged drive; „nézd meg”; the disk/memory/CPU/temperature banners in the household's language (`health.*` keys; the wire text to the hub unchanged and pinned); the last eleven counted formal forms are te-form (ceiling 14 → 0; 119 parity fixtures changed by exactly those bytes). **Left:** items 5–6 (apps' own screens); about 20 formal forms the gate does not count yet („írja be”, „adja meg”, „hozzon létre”, …) on the debug, storage, network-storage and security pages. Item 11 moved to its own row **R-889**. **2026-10-06 02:40: the bundle has NO formal form left** (controller `4c3c203`, unreleased): 49 sentences in the te-form; the gate's stem list widened (it counted 61 on the previous bundle, 0 now) with decoys and one listed third-person exception. **Left (narrowed):** Felhom-owned formal forms in Go and template literals the gate cannot read yet — the network-storage attach errors, escrow and share handlers, the single-copy notice, a settings refusal, the SMART mail, a fill-watch „Kérjük”, and the setup wizard pages; each moves into the bundle in the te-form. Items 5–6 are the apps' own screens (not ours). **2026-10-06: items 7–10, 12 and the bundle's 49 formal forms DELIVERED** in controller v0.299.0. | — | — | CC | -| **R-554** | Install & onboarding | P3 | **[P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable.** OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). `02-controller-module-map.md` L56 calls `internal/setup/` obsolete; `cmd/controller/main.go` L322 still enters it when `setup.NeedsSetup(cfg)` — `customer.id` empty after bootstrap ingestion, or a `.needs-setup` marker (`internal/setup/setup.go` L17-25). Ingestion leaves `customer.id` empty on a missing/invalid `bootstrap.json`, a failed hub pull, or a failed merge/write/reload (`internal/bootstrap/bootstrap.go` L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). **Fix shape:** decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete `internal/setup/` and `runSetupMode`; red-proof that a failed ingestion renders the waiting page, not a 404. **Check first** whether any drill/golden path still relies on `.needs-setup`. | **READY - rank P3-LOW; owner: CC** | — | — | CC | +| **R-554** | Install & onboarding | P3 | **[P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable.** OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). `02-controller-module-map.md` L56 calls `internal/setup/` obsolete; `cmd/controller/main.go` L322 still enters it when `setup.NeedsSetup(cfg)` — `customer.id` empty after bootstrap ingestion, or a `.needs-setup` marker (`internal/setup/setup.go` L17-25). Ingestion leaves `customer.id` empty on a missing/invalid `bootstrap.json`, a failed hub pull, or a failed merge/write/reload (`internal/bootstrap/bootstrap.go` L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). **Fix shape:** decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete `internal/setup/` and `runSetupMode`; red-proof that a failed ingestion renders the waiting page, not a 404. **Check first** whether any drill/golden path still relies on `.needs-setup`. | **READY - rank P3-LOW; owner: CC** **2026-10-06 night: stopped — needs the operator:** deleting the wizard makes the household's `recovery-info.txt` (`internal/recovery/info.go:25-47`, „docker-setup.sh … :8081 … Visszaállítás mentésből") false; what that file should say instead is a customer promise. The debug route `debugTriggerSetupWizard` (`handler_debug.go:690`) is the only `.needs-setup` writer; no drill/golden uses it. | — | — | CC | | **R-494** | Install & onboarding | P4 | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).** *Original finding, kept:* **[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.** **What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: operator-side automation; the operator creates the tunnel by hand per the day-0 runbook.** | — | — | CC | | **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator | @@ -126,7 +126,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-562** | Apps & catalog | P3 | **[P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves.** FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (`2006. 01. 02. 15:04` Hungarian vs `2006-01-02 15:04` ISO); 10 layout literals in `internal/web` Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (`%.1f GB`, 4 helpers) where Hungarian uses a comma; `timeAgo`/`nextRunLabel`/`pruneLabel` produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). **Fix shape:** one date and one size formatter per language in `internal/i18n`, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | **READY - rank P3-LOW; owner: CC** | — | — | CC | +| **R-562** | Apps & catalog | P3 | **[P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves.** FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (`2006. 01. 02. 15:04` Hungarian vs `2006-01-02 15:04` ISO); 10 layout literals in `internal/web` Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (`%.1f GB`, 4 helpers) where Hungarian uses a comma; `timeAgo`/`nextRunLabel`/`pruneLabel` produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). **Fix shape:** one date and one size formatter per language in `internal/i18n`, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | **READY - rank P3-LOW; owner: CC** **2026-10-06 night: not started — the row needs the operator's word on the Hungarian format (decimal comma, date style) before any code.** | — | — | CC | | **R-612** | Apps & catalog | P3 | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** **— FIXED 2026-09-23 night (catalog `a5a729a`): 512M.** Measured on the bench: the seed peaks at 312–345M and was OOM-killed at 128M on every run (kernel `oom_kill` 3); running, the app sits at ~115M (90 % of the old limit alone). On 9202 at 128M the kill came AFTER the Role/Group rows this time, so the sign-up lie did NOT reproduce — the kill is timing-dependent, the fault is memory (the brief's claim held). At 512M the seed completes (`The seed command has been executed`), and wishlist then moved v0.66.0 → v0.67.1 with its test record. **Still open:** a failed first-boot seed is invisible to the box — nothing reads the seed's exit. | **READY — P3, narrowed to making a failed seed visible; owner: CC (catalog)** **Re-ranked 2026-10-03: P1->P3: the memory fix shipped (catalog a5a729a); only the silent-seed detection remains.** | — | — | CC | | **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. **-- 2026-10-06 (afternoon), measured again on 9202 (live template, wger 2.7):** `/static/css/workout-manager.css` 404 straight at the app, `/home/wger/static` 4 KB, settings `DEBUG False`; the image has gunicorn but no whitenoise and runs Django's `runserver` (no `WGER_USE_GUNICORN`, R-755). So `DJANGO_DEBUG=False` alone would collect the files and still serve none: the fix needs a server for `/static` + `/media` (a second container) — medium, not taken. `audits/r890-instructions-2026-10-06/C/wger.txt`. | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** **2026-10-06: NARROWED — the files are served.** Catalog `cf1ed43`: `DJANGO_DEBUG=False` + a `wger-files` nginx serving `/static` and `/media`, a definition step proven on the bench (harness v5) and on 9202 (photo 404 → 200, CSS 404 → 200). **What remains is R-755's merged half:** wger still runs Django's `runserver`; upstream's gunicorn runs 3 workers, which do not fit wger's 384 MB — a memory decision and a new proof. Known cost of the fix: the collected static files are 283 MB in a named volume, in every backup of wger. Ready to show? Its files and sign-up are fixed and proven; the server question is open — `audits/design-build-2026-10-06/`F/. | — | Decide the gunicorn worker count against the memory limit, prove it on both venues; then the operator decides whether wger is shown. If nothing: wger stays hidden | CC | @@ -192,7 +192,7 @@ stopping line that lies. | **R-298** | Storage & devices | P3 | **The `/storage` page's unregistered list is filtered by `role==='user-data'`, so a drive that is also the backup target can never be registered from it.** `storage.html:363` routes anything not `user-data` into the read-only protected group with NO actions. On the rebuilt `demo-hp` the NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`/etc/pve/storage.cfg`: `dir: felhom-backup` → `/mnt/nvme-1tb`), so it renders locked. **This is the SECOND reason that page was empty** during the reinstall rehearsal, independent of R-280's candidate-source defect, and R-280's fix does not touch it — attaching is non-destructive, so the format-wizard protection is the wrong gate for a REGISTER action | **READY (S) — NEW 2026-08-10** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Registering a drive the agent classes as the backup target conflicts with the agent's eject/decommission rule (403 for that role). First: is a drive holding both app data and backups a supported layout? | R-280 | Split the role gate: `user-data` keeps destructive actions; any mounted role may be REGISTERED | CC | | **R-330** | Storage & devices | P3 | **Disk health Phase 2 — the three SMART attributes the wire does not carry.** The failing drive's most telling counter was **187 `Reported_Uncorrect`**, sitting at normalized **1** against threshold **0** with a raw count of **1001** — one point from failing and structurally unable to get there. Also wanted: **199 `UDMA_CRC_Error_Count`** (cabling) and **188 `Command_Timeout`**. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART **samples** (the right home is `metrics.MetricsStore`, NOT the Phase-1 state file, which is one record per disk and must stay that way). **This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session** — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change | **READY (M) — NEW 2026-08-14** | R-328 (closed) | Add 187/199/188 to the agent's `SmartSummary` + hub model in one session; then persist samples | CC | | **R-332** | Storage & devices | P3 | **The new Hiba-from-counters path has never fired on real hardware.** v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the **negative** — three healthy disks still read Rendben across the deploy, no false alert — and the **severity wire** end to end, but no live disk has actually reached Hiba. **This is the honest gap and it must not be closed by pointing at the fixture tests**: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy | **WATCHING — NEW 2026-08-14, NARROWED same day.** One item originally in this gap is now PROVEN LIVE: the **persisted state surviving a controller restart**. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a `changed_at` written by the PREVIOUS version (`2026-08-14T07:23:14.640216851Z`, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. **What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting** | a real degrading disk, or an injection harness | **Closing condition:** a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent `/disks` → controller check → hub event), not a hand-set verdict | CC | -| **R-542** | Storage & devices | P3 | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** | — | — | CC | +| **R-542** | Storage & devices | P3 | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** **2026-10-06 night: fixed on controller main `f65ace0`** — `/api/disks/candidates` drops every drive backing a registered storage path from initialize and attach (joined through the guest mount table; an unreadable table empties initialize); `TestR542_*`, red-proofs `audits/night-burndown-2026-10-06/ctrl/R-542-red.txt`. Ships with the next controller release; closes when delivered. | — | — | CC | | **R-756** | Storage & devices | P3 | **[P3-LOW] On 9202, "remove with drive data" refuses calibre-web with 409 „…/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető", while the controller container lists that folder.** MEASURED twice on 2026-10-01 (`audits/lockouts-2026-10-01/B/B1…`, `audits/calibre-name-and-prune-2026-10-01/A/A1…`): `POST /api/stacks/calibre-web/remove` with `remove_hdd_data` → 409; `docker exec felhom-controller ls -ld /mnt/felhom-drives/scratch_hdd/userdata/calibre-web` → the directory (dated 2026-09-22). The walk then removed the app keeping the data (R-442's fail-closed answer). Either the drive is not a registered drive on this scratch box (a test-venue artefact) or the resolver reads another path than the one it names. Not measured which. **-- 2026-10-01 (night, new apps):** the same 409 for Grimmory (`remove_hdd_data`), and a drive app's per-app backup is NOT offered for restore after a remove — the controller logged `drive … not mounted — skipping ensure (held by drive gate)`; the unit lives on that drive. So on 9202 checklist row 2.5 cannot be shown for any drive app until its scratch drive is a registered drive (`audits/new-apps-2026-10-01/box/grimmory/restore-why.txt`). **Checked from source 2026-10-05 (burn-down round 2):** Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names .../scratch_hdd/use | **OPEN — rank P3-LOW; owner: CC** **2026-10-05 (burn-down night): NEEDS A LIVE READING.** The refusal shows that app's recorded HDD_PATH is a sub-folder of the drive (the R-839 shape); loosening the check would weaken the boot start gate. Next: that app's HDD_PATH and `findmnt` inside 9202. | — | — | CC | | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | @@ -242,7 +242,7 @@ stopping line that lies. | **R-388** | Monitoring & notifications | P3 | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor | | **R-435** | Monitoring & notifications | P3 | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Covering one-app deletions needs per-app snapshot counts from the controller and a per-tag threshold measured on real counts — two repos, a mechanism nobody has measured. | — | — | CC | | **R-521** | Monitoring & notifications | P3 | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** **2026-10-05 (burn-down night): FIXED on controller `main`** (`800b32c` — controller half: apps a lost drive stopped no longer mail app_start_failed one by one; test + red-proof). Ships with the next controller release; close after delivery. Left: the hub's cooldown outliving a recovery (F6/F7) and whether a household gets a mail for a lost drive (operator). **2026-10-06: controller half DELIVERED** (v0.298.0). | — | — | CC + operator | -| **R-724** | Monitoring & notifications | P3 | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). **FIXED 2026-09-30 (controller v0.283.0):** Beállítások shows the real schedule and a local time; the dashboard's last backup is local time (R-500); „az előző N órája készült" replaces „0 órája". **NARROWED — remaining:** „Helyi cím (LAN)" / „Átjáró" read through the file-sharing container, so a box without it says „nem állapítható meg" — not a text fix (the read needs another path). | **NARROWED — the LAN/gateway read only; owner: CC** | — | — | CC | +| **R-724** | Monitoring & notifications | P3 | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). **FIXED 2026-09-30 (controller v0.283.0):** Beállítások shows the real schedule and a local time; the dashboard's last backup is local time (R-500); „az előző N órája készült" replaces „0 órája". **NARROWED — remaining:** „Helyi cím (LAN)" / „Átjáró" read through the file-sharing container, so a box without it says „nem állapítható meg" — not a text fix (the read needs another path). | **NARROWED — the LAN/gateway read only; owner: CC** **2026-10-06 night: confirmed not a controller-only fix** — the only host-network container is samba (`infra/samba.go:123`) and the agent's local API has no guest-network route; the fix is a new agent route (e.g. the PVE lxc interfaces read) + a controller reader — a two-repo design. | — | — | CC | | **R-177** | Monitoring & notifications | P4 | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the same missing operator door as R-314 and R-279. | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | | **R-266** | Monitoring & notifications | P4 | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | — | — | operator | | **R-285** | Monitoring & notifications | P4 | **A planned, supervised reinstall pages the operator as if the machine had died — there is no notion of expected downtime anywhere.** During the 2026-08-09 rehearsal the hub sent, all `status: sent` to the operator channel: `host_stale` 08:58 UTC, `node_stale` 09:00, **`host_down` 09:28 (error)**, **`node_down` 09:30 (error)**, `host_leaf_changed` 09:31, `host_recovered` 09:31, `node_recovered` 09:34, `offsite_delivery_stuck` 09:34 — eight operator mails for work that was deliberate, attended and announced. **This is the OPPOSITE gap from the one R-281 filed:** the alarms are not missing, they are indiscriminate. `host_stale` at 30 min and `host_down` at 60 min (`monitor/host_staleness.go:22-23`, `downAfter = 2 * threshold`) cannot distinguish a wiped-on-purpose box from a dead one, and `host_leaf_changed` firing on a reinstall is correct-but-expected. **Note the interaction with the mute used on 2026-08-09 evening:** blocking a customer silences everything, so today the only two settings are *page me for planned work* and *tell me nothing at all*. **What is owed is a middle:** a maintenance window, or an operator-set expected-downtime flag, that suppresses staleness and leaf-change while leaving genuine faults audible | **READY (M) — NEW 2026-08-09** | — | The evidence is the operator's mailbox plus `events`/`notification_log` for 2026-08-09 | CC |