Commit the raw evidence from ST3000VX010 S/N Z6A07P2G (/dev/sdg on DooPlex), which went 8 -> 352 unreadable sectors 11-13 Aug while smart_status.passed stayed true throughout. - fixtures/smart-ST3000VX010-failing-2026-08-14.json: raw smartctl -a -j, verbatim - fixtures/smartd-history-sdg-2026-08-14.txt: 406 smartd journal lines, 11-14 Aug - DIAG-smart-passed-trap-2026-08-14.md: the mechanism (attrs 187/197/198 all carry thresh 0, so a normalized value that floors at 1 can never fail the overall verdict on unreadable sectors), the non-monotonic timeline, the three controller defects with locators, and the counterfactual: zero emails would have been sent.
9.5 KiB
DIAG — the SMART PASSED trap, and the disk alert that reached nobody
Date: 2026-08-14
Drive: Seagate ST3000VX010-2E3166, S/N Z6A07P2G, 3.0 TB, /dev/sdg on DooPlex
Fixtures: fixtures/smart-ST3000VX010-failing-2026-08-14.json (raw smartctl -a -j /dev/sdg),
fixtures/smartd-history-sdg-2026-08-14.txt (406 smartd journal lines for this device, 11–14 Aug)
Status: the three controller defects named here are FIXED in controller v0.215.0; the
counterfactual in §5 is derived from source, not reproduced live.
This is the project's first genuinely failing disk. Before it, the capability map recorded the disk-failure scenario as "Healthy path only — a genuinely failing disk has never been seen."
1. What happened
On 11 August 12:28 a 3 TB drive in DooPlex reported its first unreadable sectors. By 13 August
21:58 it was at 360, and it was taking down a running service. Throughout the entire episode the
drive's own overall self-assessment read PASSED, and it still does.
smartd was running the whole time and mailed local root — a mailbox nobody reads. The failure
was actually found by a crashlooping pod, not by any alert.
2. The mechanism — smart_status.passed cannot fail on unreadable sectors
Measured, from the committed fixture:
| ID | Attribute | value | worst | thresh | raw |
|---|---|---|---|---|---|
| 5 | Reallocated_Sector_Ct |
100 | 100 | 10 | 0 |
| 187 | Reported_Uncorrect |
1 | 1 | 0 | 1001 |
| 188 | Command_Timeout |
100 | 100 | 0 | 0 |
| 197 | Current_Pending_Sector |
98 | 98 | 0 | 352 |
| 198 | Offline_Uncorrectable |
98 | 98 | 0 | 352 |
| 199 | UDMA_CRC_Error_Count |
200 | 200 | 0 | 0 |
smart_status.passed is false only when some attribute's normalized value falls at or below
its threshold. Attributes 187, 197 and 198 — the three that record unreadable sectors — all carry
thresh: 0. A normalized SMART value floors at 1 and cannot reach 0.
Therefore
smart_status.passedis structurally incapable of failing on unreadable sectors. This is not a quirk of this drive or this vendor. It is how a zero threshold behaves.
Attribute 187 Reported_Uncorrect sits at normalized 1 against threshold 0 with a raw count
of 1001. It is one point from failing and has no remaining point to give.
Corroborating, from the same fixture: ata_smart_error_log.summary.count = 1001, power-on hours
60505 (~6.9 years), smartctl exit status 64 (bit 6 — the device error log contains
records of errors) while smart_status.passed is still true.
Any monitor built on the drive's overall verdict is blind to this entire class of failure.
Felhom already knows better — agentapi.DiskVerdictFor reads the raw counters — which is why it
would have noticed on 11 August, two days early.
3. Unreadable sectors are not monotonic
From smartd-history-sdg-2026-08-14.txt, Current_Pending_Sector over the episode (30-minute
sampling, host clock = CEST):
Aug 11 12:28 8 first sighting
Aug 11 12:58 16 (+8)
Aug 11 13:28 0 FULL CLEAR — "No more Currently unreadable (pending) sectors,
warning condition reset after 1 email"
Aug 11 20:28 8 returns
Aug 12 01:28 16 → 01:58 8 (-8)
Aug 12 02:28 32 → 02:58 24 (-8)
Aug 12 03:28 24 187 Reported_Uncorrect 100→97; ATA error count 0→3
Aug 12 03:58 16 (-8) → 04:28 24 (+8)
Aug 12 10:58 32 … steady 32 for ~11h …
Aug 12 21:58 24 (-8)
Aug 13 03:28 40 (+16) → 04:28 24 (-16)
Aug 13 11:28 64 terminal run begins — never returns below 64
Aug 13 11:58 72 12:58 80 13:28 112 15:58 120
Aug 13 21:58 360 (+240)
Aug 13 22:28 352 (-8) … steady 352 through 14 Aug …
Two measured facts carry design weight:
- The 11 August excursion cleared completely within one hour (12:28 → 13:28). A bare
> 0alarm would have fired on a drive that then looked fine for seven hours. This is why the ladder uses a sustain rule rather than a bare non-zero test. - The benign excursion peaked at 16; the terminal run crossed 64 at 13 Aug 11:28 and never came back. That is the entire empirical basis for the static count threshold of 64 — see §6.
4. The three controller defects
All three are in felhom-controller at 3e3ee94 (v0.214.0), the tree audited here.
D1 — the alert carries a severity the hub does not recognise (highest value)
internal/notify/notifier.go:565 — NotifyDiskHealthDegraded sets:
severity := "warn"
The hub accepts an exact-match lowercase vocabulary and coerces anything else to info:
felhom.eu/hub/internal/api/handler.go:2121-2126—case "info", "warning", "error", "critical":…default: payload.Severity = "info"felhom.eu/hub/internal/notify/dispatcher.go:89-96—severityNotifiesreturns true only forwarning/error/critical.
"warn" is not in the accepted set. So the Figyelmeztetés-level disk alert is stored as an
informational notice and emailed to nobody, on the customer leg and the operator leg alike.
The function's own doc comment reads "The hub applies its own per-event-type cooldown" — which presumes it routes. An invariant asserted in a comment with no test pinning it; this project's recurring shape.
D2 — no level above "worth keeping an eye on it"
internal/agentapi/diskverdict.go:28-41 returns Warn for any non-zero counter, and can only
reach Fail when Health == "FAILING" — which, by §2, this fault class cannot produce. A drive with
one aging sector and a drive at 352 unreadable sectors rendered the identical chip and the identical
mild sentence.
D3 — it speaks once, and forgets on restart
internal/web/disk_health.go:182 emits only on v > prev, against a in-memory baseline
(disk_health.go:26-29, "Lost on restart → the next check re-baselines silently"). Consequences:
- Between 8 and 352 pending sectors the verdict never changes level, so nothing further is emitted.
- A controller restart while a disk is already bad re-baselines it silently — that disk never alerts again.
5. Counterfactual — what a customer would have received
Derived from the source above plus the §3 timeline. NOT reproduced live.
| Date/time | Drive state | Felhom verdict at v0.214.0 | Emitted | Delivered |
|---|---|---|---|---|
| 11 Aug 12:28 | pending 8 | OK → Figyelmeztetés | disk_health_degraded, severity warn |
nothing — coerced to info, dropped by severityNotifies |
| 11 Aug 13:28 | pending 0 | Figyelmeztetés → Rendben | none (recovery is silent) | nothing |
| 11 Aug 20:28 → 13 Aug | 8 → 352 | Figyelmeztetés throughout | none (no level change) | nothing |
| 13 Aug 21:58 | pending 360 | Figyelmeztetés | none | nothing |
Felhom would have emitted zero emails about this drive. The one event it did produce was filed at
infoand delivered to no one.
Note the two defects compound: even had D1 been fixed alone, the customer would have received a single mild "Javasolt figyelemmel kísérni" at 8 sectors on 11 August and then silence through 352.
6. Provenance of the thresholds chosen in v0.215.0
- 64 unreadable sectors → Hiba. The observed benign excursion peaked at 16 and cleared inside an hour; the terminal run passed 64 at 13 Aug 11:28 and never returned below it. 64 sits above the one observed transient and below the observed terminal run. This is a judgement from ONE drive. It is a static backstop and is expected to be replaced in Phase 3 by growth-rate detection over real history.
- Sustain before count. The primary rule is "unreadable sectors still present at the next check"; the count is the backstop. On this drive sustain fires 12 Aug, the count not until 13 Aug — a full day earlier. The backstop exists for a box that was powered off or restarted across the sustain window.
- 55 / 60 °C. Adopted unchanged from the operator's existing Prometheus bands on DooPlex, so the two systems cannot disagree about the same drive.
7. Measured vs inferred
Measured (reproducible from the committed fixtures):
- The attribute table, thresholds and raw values in §2;
passed: trueat 352 pending sectors. - The full non-monotonic timeline in §3, including the one-hour full clear.
- The three defect locators in §4 — read from live source on both the controller and the hub side.
Inferred (source-derived, not executed):
- The §5 counterfactual. It follows from the §4 locators and the §3 timeline; no email path was
exercised against this drive, and no
disk_health_degradedevent for it exists in the hub.
Not covered here:
- Attributes 187, 199 and 188 are not on the agent→controller wire today. Adding them is Phase 2 (a declared wire change, so the hub models them in the same session under G-1). Everything the v0.215.0 fix needs was already on the wire.
- The new Fail-from-counters path has not fired on real hardware — only against this fixture's values in unit tests.
8. One layer out
smartd on DooPlex did its job and mailed local root, where nothing reads. That is the same shape
as D1 — a correct detection with a delivery path to nowhere — one layer outside the product. Tracked
separately as DooPlex hygiene.