fix(disk-health): one physical disk must be evaluated once per run (R-335)
gates / gates (push) Successful in 9s

Found on live hardware two hours after the v0.215.0 deploy, by noticing the
release's own positive observable disagreed with its own persisted artefact:
the check logged '3 disk(s) evaluated' while disk-health-state.json held two
records. demo-hp's c11-scratch and felhom-backup are the same NVMe and share
a durable id, so one disk was walked twice per run.

Not cosmetic. The loop writes a disk's record before the next entry reads it,
so the second copy of an aliased disk consumed the FIRST copy's write as its
prior: the disk sustained against ITSELF and reached Hiba on a first sighting,
defeating truth-table row 6 — the rule that separates a one-hour benign
excursion from a false critical. It would also have emitted two identical
events for one drive. Latent on demo-hp only because all counters are zero.

Each diskKey is now evaluated once per run. Both entries stay marked seen so
neither looks like a disappeared disk, and the card still renders both rows —
the dedup is about state and alerts, not display.

Red-proof run and reverted: deleting the guard makes the first sighting emit
Kind:2 (Hiba-from-sectors) at 8 sectors.
This commit is contained in:
2026-08-14 10:30:22 +02:00
parent 8144a70a72
commit 90f2545679
3 changed files with 102 additions and 0 deletions
+31
View File
@@ -1,3 +1,34 @@
## v0.216.0 — one physical disk, one verdict (2026-08-14, R-335)
**MinAgent: 0.129.0**
**Found on live hardware within two hours of the v0.215.0 deploy, by reading the artefact the release
itself introduced.** The first two hourly checks on demo-hp logged *"3 disk(s) evaluated"* while the
persisted `disk-health-state.json` held only **two** records. The discrepancy was the bug: demo-hp's
`c11-scratch` and `felhom-backup` are the same physical NVMe (`/dev/nvme0n1`) and resolve to the same
`diskKey`, so one disk was walked twice in a single run.
**Why that was not cosmetic.** The loop writes a disk's new record before the next entry reads it, so
the SECOND copy of an aliased disk consumed the FIRST copy's write as its prior. The disk therefore
**sustained against itself** and reached **Hiba on a FIRST sighting** — defeating truth-table row 6,
the rule the whole v0.215.0 ladder is built on and the one thing standing between a one-hour benign
excursion and a false critical alert. It would also have emitted **two identical events** for one
drive. Nothing fired on demo-hp because all three entries are healthy with zero counters, so the fault
was latent, not active — but any aliased disk developing a single pending sector would have gone
straight to Hiba.
**Fix:** `RunDiskHealthCheck` evaluates each `diskKey` **once per run** (`internal/web/disk_health.go`).
Both entries are still marked `seen`, so neither is mistaken for a disappeared disk, and the card still
renders **both** storage rows — the dedup is about state and alerts, not display.
**Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`**, which asserts the first sighting stays
Figyelmeztetés and silent, that the stored prior is not this run's own write, that the second run emits
exactly ONE event, and that the card still shows two rows. Companion red-proof run and reverted:
deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors — the exact
false critical described above.
This is the shape §3 of the workspace rules warns about: the release's own positive observable
("N disks evaluated") disagreed with its own persisted artefact, and only reading BOTH exposed it.
## v0.215.0 — the disk alert that never sent (2026-08-14, R-328..R-333)
**MinAgent: 0.129.0** (unchanged — every field this reads has been on the wire since agent v0.94.0/0.95.0)