# DIAG — the SMART `PASSED` trap, and the disk alert that reached nobody **Date:** 2026-08-14 **Drive:** Seagate `ST3000VX010-2E3166`, S/N `Z6A07P2G`, 3.0 TB, `/dev/sdg` on **DooPlex** **Fixtures:** `fixtures/smart-ST3000VX010-failing-2026-08-14.json` (raw `smartctl -a -j /dev/sdg`), `fixtures/smartd-history-sdg-2026-08-14.txt` (406 `smartd` journal lines for this device, 11–14 Aug) **Status:** the three controller defects named here are FIXED in controller **v0.215.0**; the counterfactual in §5 is derived from source, **not** reproduced live. This is the project's first genuinely failing disk. Before it, the capability map recorded the disk-failure scenario as *"Healthy path only — a genuinely failing disk has never been seen."* --- ## 1. What happened On **11 August 12:28** a 3 TB drive in DooPlex reported its first unreadable sectors. By **13 August 21:58** it was at 360, and it was taking down a running service. Throughout the entire episode the drive's own overall self-assessment read **`PASSED`**, and it still does. `smartd` was running the whole time and mailed **local root** — a mailbox nobody reads. The failure was actually found by a crashlooping pod, not by any alert. --- ## 2. The mechanism — `smart_status.passed` cannot fail on unreadable sectors Measured, from the committed fixture: | ID | Attribute | value | worst | thresh | raw | |-----|--------------------------|-------|-------|--------|--------| | 5 | `Reallocated_Sector_Ct` | 100 | 100 | **10** | 0 | | 187 | `Reported_Uncorrect` | **1** | **1** | **0** | **1001** | | 188 | `Command_Timeout` | 100 | 100 | **0** | 0 | | 197 | `Current_Pending_Sector` | 98 | 98 | **0** | **352** | | 198 | `Offline_Uncorrectable` | 98 | 98 | **0** | **352** | | 199 | `UDMA_CRC_Error_Count` | 200 | 200 | **0** | 0 | `smart_status.passed` is false only when some attribute's **normalized value** falls **at or below** its **threshold**. Attributes 187, 197 and 198 — the three that record unreadable sectors — all carry `thresh: 0`. A normalized SMART value floors at 1 and cannot reach 0. > **Therefore `smart_status.passed` is structurally incapable of failing on unreadable sectors.** > This is not a quirk of this drive or this vendor. It is how a zero threshold behaves. Attribute **187 `Reported_Uncorrect` sits at normalized `1` against threshold `0`** with a raw count of **1001**. It is one point from failing and has no remaining point to give. Corroborating, from the same fixture: `ata_smart_error_log.summary.count = 1001`, power-on hours **60505** (~6.9 years), `smartctl` exit status **64** (bit 6 — *the device error log contains records of errors*) while `smart_status.passed` is still `true`. **Any monitor built on the drive's overall verdict is blind to this entire class of failure.** Felhom already knows better — `agentapi.DiskVerdictFor` reads the raw counters — which is why it would have noticed on 11 August, two days early. --- ## 3. Unreadable sectors are not monotonic From `smartd-history-sdg-2026-08-14.txt`, `Current_Pending_Sector` over the episode (30-minute sampling, host clock = CEST): ``` Aug 11 12:28 8 first sighting Aug 11 12:58 16 (+8) Aug 11 13:28 0 FULL CLEAR — "No more Currently unreadable (pending) sectors, warning condition reset after 1 email" Aug 11 20:28 8 returns Aug 12 01:28 16 → 01:58 8 (-8) Aug 12 02:28 32 → 02:58 24 (-8) Aug 12 03:28 24 187 Reported_Uncorrect 100→97; ATA error count 0→3 Aug 12 03:58 16 (-8) → 04:28 24 (+8) Aug 12 10:58 32 … steady 32 for ~11h … Aug 12 21:58 24 (-8) Aug 13 03:28 40 (+16) → 04:28 24 (-16) Aug 13 11:28 64 terminal run begins — never returns below 64 Aug 13 11:58 72 12:58 80 13:28 112 15:58 120 Aug 13 21:58 360 (+240) Aug 13 22:28 352 (-8) … steady 352 through 14 Aug … ``` Two measured facts carry design weight: 1. **The 11 August excursion cleared completely within one hour** (12:28 → 13:28). A bare `> 0` alarm would have fired on a drive that then looked fine for seven hours. This is why the ladder uses a *sustain* rule rather than a bare non-zero test. 2. **The benign excursion peaked at 16; the terminal run crossed 64 at 13 Aug 11:28 and never came back.** That is the entire empirical basis for the static count threshold of 64 — see §6. --- ## 4. The three controller defects All three are in `felhom-controller` at `3e3ee94` (v0.214.0), the tree audited here. ### D1 — the alert carries a severity the hub does not recognise *(highest value)* `internal/notify/notifier.go:565` — `NotifyDiskHealthDegraded` sets: ```go severity := "warn" ``` The hub accepts an exact-match lowercase vocabulary and **coerces anything else to `info`**: - `felhom.eu/hub/internal/api/handler.go:2121-2126` — `case "info", "warning", "error", "critical":` … `default: payload.Severity = "info"` - `felhom.eu/hub/internal/notify/dispatcher.go:89-96` — `severityNotifies` returns true only for `warning` / `error` / `critical`. `"warn"` is not in the accepted set. So the Figyelmeztetés-level disk alert is **stored as an informational notice and emailed to nobody**, on the customer leg and the operator leg alike. The function's own doc comment reads *"The hub applies its own per-event-type cooldown"* — which presumes it routes. An invariant asserted in a comment with no test pinning it; this project's recurring shape. ### D2 — no level above "worth keeping an eye on it" `internal/agentapi/diskverdict.go:28-41` returns `Warn` for *any* non-zero counter, and can only reach `Fail` when `Health == "FAILING"` — which, by §2, this fault class cannot produce. A drive with one aging sector and a drive at 352 unreadable sectors rendered the identical chip and the identical mild sentence. ### D3 — it speaks once, and forgets on restart `internal/web/disk_health.go:182` emits only on `v > prev`, against a **in-memory** baseline (`disk_health.go:26-29`, *"Lost on restart → the next check re-baselines silently"*). Consequences: - Between 8 and 352 pending sectors the verdict never changes level, so **nothing further is emitted**. - A controller restart while a disk is already bad re-baselines it silently — that disk never alerts again. --- ## 5. Counterfactual — what a customer would have received **Derived from the source above plus the §3 timeline. NOT reproduced live.** | Date/time | Drive state | Felhom verdict at v0.214.0 | Emitted | Delivered | |-----------|-------------|-----------------------------|---------|-----------| | 11 Aug 12:28 | pending 8 | OK → Figyelmeztetés | `disk_health_degraded`, severity `warn` | **nothing** — coerced to `info`, dropped by `severityNotifies` | | 11 Aug 13:28 | pending 0 | Figyelmeztetés → Rendben | none (recovery is silent) | nothing | | 11 Aug 20:28 → 13 Aug | 8 → 352 | Figyelmeztetés throughout | none (no level change) | nothing | | 13 Aug 21:58 | pending 360 | Figyelmeztetés | none | nothing | > **Felhom would have emitted zero emails about this drive.** The one event it did produce was filed > at `info` and delivered to no one. Note the two defects compound: even had D1 been fixed alone, the customer would have received a single mild "Javasolt figyelemmel kísérni" at 8 sectors on 11 August and then silence through 352. --- ## 6. Provenance of the thresholds chosen in v0.215.0 - **64 unreadable sectors → Hiba.** The observed benign excursion peaked at **16** and cleared inside an hour; the terminal run passed **64** at 13 Aug 11:28 and never returned below it. 64 sits above the one observed transient and below the observed terminal run. **This is a judgement from ONE drive.** It is a static backstop and is expected to be replaced in Phase 3 by growth-rate detection over real history. - **Sustain before count.** The primary rule is "unreadable sectors still present at the next check"; the count is the backstop. On this drive sustain fires **12 Aug**, the count not until **13 Aug** — a full day earlier. The backstop exists for a box that was powered off or restarted across the sustain window. - **55 / 60 °C.** Adopted unchanged from the operator's existing Prometheus bands on DooPlex, so the two systems cannot disagree about the same drive. --- ## 7. Measured vs inferred **Measured** (reproducible from the committed fixtures): - The attribute table, thresholds and raw values in §2; `passed: true` at 352 pending sectors. - The full non-monotonic timeline in §3, including the one-hour full clear. - The three defect locators in §4 — read from live source on both the controller and the hub side. **Inferred** (source-derived, not executed): - The §5 counterfactual. It follows from the §4 locators and the §3 timeline; **no email path was exercised against this drive**, and no `disk_health_degraded` event for it exists in the hub. **Not covered here:** - Attributes **187**, **199** and **188** are not on the agent→controller wire today. Adding them is Phase 2 (a declared wire change, so the hub models them in the same session under G-1). Everything the v0.215.0 fix needs was already on the wire. - The new Fail-from-counters path has **not** fired on real hardware — only against this fixture's values in unit tests. --- ## 8. One layer out `smartd` on DooPlex did its job and mailed local root, where nothing reads. That is the same shape as D1 — a correct detection with a delivery path to nowhere — one layer outside the product. Tracked separately as DooPlex hygiene.