|
|
|
@@ -190,7 +190,7 @@ stopping line that lies.
|
|
|
|
|
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
|
|
|
|
|---|---|---|---|---|---|---|---|
|
|
|
|
|
| **R-298** | Storage & devices | P3 | **The `/storage` page's unregistered list is filtered by `role==='user-data'`, so a drive that is also the backup target can never be registered from it.** `storage.html:363` routes anything not `user-data` into the read-only protected group with NO actions. On the rebuilt `demo-hp` the NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`/etc/pve/storage.cfg`: `dir: felhom-backup` → `/mnt/nvme-1tb`), so it renders locked. **This is the SECOND reason that page was empty** during the reinstall rehearsal, independent of R-280's candidate-source defect, and R-280's fix does not touch it — attaching is non-destructive, so the format-wizard protection is the wrong gate for a REGISTER action | **READY (S) — NEW 2026-08-10** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Registering a drive the agent classes as the backup target conflicts with the agent's eject/decommission rule (403 for that role). First: is a drive holding both app data and backups a supported layout? | R-280 | Split the role gate: `user-data` keeps destructive actions; any mounted role may be REGISTERED | CC |
|
|
|
|
|
| **R-330** | Storage & devices | P3 | **Disk health Phase 2 — the three SMART attributes the wire does not carry.** The failing drive's most telling counter was **187 `Reported_Uncorrect`**, sitting at normalized **1** against threshold **0** with a raw count of **1001** — one point from failing and structurally unable to get there. Also wanted: **199 `UDMA_CRC_Error_Count`** (cabling) and **188 `Command_Timeout`**. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART **samples** (the right home is `metrics.MetricsStore`, NOT the Phase-1 state file, which is one record per disk and must stay that way). **This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session** — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change | **READY (M) — NEW 2026-08-14** | R-328 (closed) | Add 187/199/188 to the agent's `SmartSummary` + hub model in one session; then persist samples | CC |
|
|
|
|
|
| **R-330** | Storage & devices | P3 | **Disk health Phase 2 — the three SMART attributes the wire does not carry.** The failing drive's most telling counter was **187 `Reported_Uncorrect`**, sitting at normalized **1** against threshold **0** with a raw count of **1001** — one point from failing and structurally unable to get there. Also wanted: **199 `UDMA_CRC_Error_Count`** (cabling) and **188 `Command_Timeout`**. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART **samples** (the right home is `metrics.MetricsStore`, NOT the Phase-1 state file, which is one record per disk and must stay that way). **This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session** — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change | **READY (M) — NEW 2026-08-14** **2026-10-06 night: the wire half BUILT on main (unreleased):** the agent sends SMART 187 `reported_uncorrect`, 188 `command_timeout` (raw, vendor-packed on some drives — carried as reported), 199 `udma_crc_errors` (pointer + omitempty: unknown is absent, never 0); the controller decodes them and the hub's host-report mirror models them (G-1 wire gate green; it convicted the agent half alone). **Carried only** — no verdict, banner, mail or alarm reads them (`TestR330_CountersChangeNoVerdictYet`); using them changes what a household is told and is the next slice, with persisting samples. Red-proofs `audits/night-burndown-2026-10-06/r330/`. Ships with agent v0.150.0 and the next controller and hub releases. | R-328 (closed) | Add 187/199/188 to the agent's `SmartSummary` + hub model in one session; then persist samples | CC |
|
|
|
|
|
| **R-332** | Storage & devices | P3 | **The new Hiba-from-counters path has never fired on real hardware.** v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the **negative** — three healthy disks still read Rendben across the deploy, no false alert — and the **severity wire** end to end, but no live disk has actually reached Hiba. **This is the honest gap and it must not be closed by pointing at the fixture tests**: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy | **WATCHING — NEW 2026-08-14, NARROWED same day.** One item originally in this gap is now PROVEN LIVE: the **persisted state surviving a controller restart**. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a `changed_at` written by the PREVIOUS version (`2026-08-14T07:23:14.640216851Z`, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. **What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting** | a real degrading disk, or an injection harness | **Closing condition:** a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent `/disks` → controller check → hub event), not a hand-set verdict | CC |
|
|
|
|
|
| **R-542** | Storage & devices | P3 | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** **2026-10-06 night: fixed on controller main `f65ace0`** — `/api/disks/candidates` drops every drive backing a registered storage path from initialize and attach (joined through the guest mount table; an unreadable table empties initialize); `TestR542_*`, red-proofs `audits/night-burndown-2026-10-06/ctrl/R-542-red.txt`. Ships with the next controller release; closes when delivered. | — | — | CC |
|
|
|
|
|
| **R-756** | Storage & devices | P3 | **[P3-LOW] On 9202, "remove with drive data" refuses calibre-web with 409 „…/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető", while the controller container lists that folder.** MEASURED twice on 2026-10-01 (`audits/lockouts-2026-10-01/B/B1…`, `audits/calibre-name-and-prune-2026-10-01/A/A1…`): `POST /api/stacks/calibre-web/remove` with `remove_hdd_data` → 409; `docker exec felhom-controller ls -ld /mnt/felhom-drives/scratch_hdd/userdata/calibre-web` → the directory (dated 2026-09-22). The walk then removed the app keeping the data (R-442's fail-closed answer). Either the drive is not a registered drive on this scratch box (a test-venue artefact) or the resolver reads another path than the one it names. Not measured which. **-- 2026-10-01 (night, new apps):** the same 409 for Grimmory (`remove_hdd_data`), and a drive app's per-app backup is NOT offered for restore after a remove — the controller logged `drive … not mounted — skipping ensure (held by drive gate)`; the unit lives on that drive. So on 9202 checklist row 2.5 cannot be shown for any drive app until its scratch drive is a registered drive (`audits/new-apps-2026-10-01/box/grimmory/restore-why.txt`). **Checked from source 2026-10-05 (burn-down round 2):** Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names .../scratch_hdd/use | **OPEN — rank P3-LOW; owner: CC** **2026-10-05 (burn-down night): NEEDS A LIVE READING.** The refusal shows that app's recorded HDD_PATH is a sub-folder of the drive (the R-839 shape); loosening the check would weaken the boot start gate. Next: that app's HDD_PATH and `findmnt` inside 9202. | — | — | CC |
|
|
|
|
|