fix(disk-health): one physical disk must be evaluated once per run (R-335)
gates / gates (push) Successful in 9s
gates / gates (push) Successful in 9s
Found on live hardware two hours after the v0.215.0 deploy, by noticing the release's own positive observable disagreed with its own persisted artefact: the check logged '3 disk(s) evaluated' while disk-health-state.json held two records. demo-hp's c11-scratch and felhom-backup are the same NVMe and share a durable id, so one disk was walked twice per run. Not cosmetic. The loop writes a disk's record before the next entry reads it, so the second copy of an aliased disk consumed the FIRST copy's write as its prior: the disk sustained against ITSELF and reached Hiba on a first sighting, defeating truth-table row 6 — the rule that separates a one-hour benign excursion from a false critical. It would also have emitted two identical events for one drive. Latent on demo-hp only because all counters are zero. Each diskKey is now evaluated once per run. Both entries stay marked seen so neither looks like a disappeared disk, and the card still renders both rows — the dedup is about state and alerts, not display. Red-proof run and reverted: deleting the guard makes the first sighting emit Kind:2 (Hiba-from-sectors) at 8 sectors.
This commit is contained in:
@@ -1,3 +1,34 @@
|
||||
## v0.216.0 — one physical disk, one verdict (2026-08-14, R-335)
|
||||
**MinAgent: 0.129.0**
|
||||
|
||||
**Found on live hardware within two hours of the v0.215.0 deploy, by reading the artefact the release
|
||||
itself introduced.** The first two hourly checks on demo-hp logged *"3 disk(s) evaluated"* while the
|
||||
persisted `disk-health-state.json` held only **two** records. The discrepancy was the bug: demo-hp's
|
||||
`c11-scratch` and `felhom-backup` are the same physical NVMe (`/dev/nvme0n1`) and resolve to the same
|
||||
`diskKey`, so one disk was walked twice in a single run.
|
||||
|
||||
**Why that was not cosmetic.** The loop writes a disk's new record before the next entry reads it, so
|
||||
the SECOND copy of an aliased disk consumed the FIRST copy's write as its prior. The disk therefore
|
||||
**sustained against itself** and reached **Hiba on a FIRST sighting** — defeating truth-table row 6,
|
||||
the rule the whole v0.215.0 ladder is built on and the one thing standing between a one-hour benign
|
||||
excursion and a false critical alert. It would also have emitted **two identical events** for one
|
||||
drive. Nothing fired on demo-hp because all three entries are healthy with zero counters, so the fault
|
||||
was latent, not active — but any aliased disk developing a single pending sector would have gone
|
||||
straight to Hiba.
|
||||
|
||||
**Fix:** `RunDiskHealthCheck` evaluates each `diskKey` **once per run** (`internal/web/disk_health.go`).
|
||||
Both entries are still marked `seen`, so neither is mistaken for a disappeared disk, and the card still
|
||||
renders **both** storage rows — the dedup is about state and alerts, not display.
|
||||
|
||||
**Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`**, which asserts the first sighting stays
|
||||
Figyelmeztetés and silent, that the stored prior is not this run's own write, that the second run emits
|
||||
exactly ONE event, and that the card still shows two rows. Companion red-proof run and reverted:
|
||||
deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors — the exact
|
||||
false critical described above.
|
||||
|
||||
This is the shape §3 of the workspace rules warns about: the release's own positive observable
|
||||
("N disks evaluated") disagreed with its own persisted artefact, and only reading BOTH exposed it.
|
||||
|
||||
## v0.215.0 — the disk alert that never sent (2026-08-14, R-328..R-333)
|
||||
**MinAgent: 0.129.0** (unchanged — every field this reads has been on the wire since agent v0.94.0/0.95.0)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user