# REPORT — SPIKE R-117: a dead bind that reports healthy (2026-07-30) Written as `REPORT-.md`, not `REPORT.md`, per this repo's parallel-session rule and the established local pattern (`REPORT-r116-diag.md`, `REPORT-session-c.md`, …). The shared `REPORT.md` was not touched. **Class: Spike.** Deliverable is a findings document. **No production code was written; no `.go` file in either repo was modified; nothing was built for deployment, published, or version-bumped.** ## Deliverables | File | Action | |---|---| | `documentation/audits/SPIKE-r117-bind-liveness-2026-07-30.md` | **new** — the findings doc (Q1–Q7, evidence, recipe, probe comparison, recommendation) | | `documentation/backlog/OPEN-ITEMS.md` | **R-117 row rewritten** with the mechanism, the reproduction recipe and the fix constraint — the one register edit, per the brief | **No CHANGELOG entry.** This repo's changelogs are per-area (`hub/`, `scripts/`, `website/`); a documentation-only change belongs to none of them. Stated rather than silently skipped, per standing rule 4. ## Baselines agent `main` **v0.116.0** @ `d4eb259` · controller `main` **v0.186.0** @ `b331f18` · `felhom.eu` HEAD `29bcfeb` · hub **live 0.82.0** · golden **0.186.0** · demo-hp PVE 9.2.2 / kernel 7.0.2-6-pve, **live agent 0.113.0** (= manifest `MinAgent`; never used as the source of a behavioural claim — every predicate result came from a probe built from `main`). ## Results **All seven questions answered empirically.** - **Q1 — reproduced 3/3**, two device classes, on a purpose-built scratch LXC (9301) on demo-hp. **The device-node change is a consequence of the defect, not a precondition** — control test: with the stale bind held the drive returns as `sdc` (8:32); released, the letter is reused (`sdb`, 8:16). - **Q2 — two death states**: device removed ⇒ superopts gain `shutdown`, `EIO`(5) on read and write, host and guest; device errors in place ⇒ `emergency_ro`, write `EROFS`(30), reads served from cache. The raw host mount is genuinely healthy in both. **No cross-device mis-identification is possible** on this path — the unit is fs-UUID-keyed. - **Q3 — confirmed at source and live.** Both halves of the R-113 conjunction compare **field 5** of a mountinfo line and **never read field 3 (`major:minor`)**, so neither can see that the bind and the raw mount name different devices. Measured `BoundUnderParent = TRUE` over an `EIO` namespace. - **Q4 — a pure-`/proc` check costs 0.16–0.45 ms**, cannot hang, spins up no disk, writes nothing. **`statfs` and `getdents` both SUCCEED on a dead namespace** — probes built on either are hollow. The hang case is below; it is the sharpest result in the run. - **Q5 — the agent**, and not on balance: the controller runs inside the guest and cannot see the host mount tables the check needs. - **Q6 — recovery works in place, guest never restarted** (init PID identical). **The repair code already exists and three call sites already invoke it**, including the controller's `Return` branch *before* it restarts apps — all defeated by one early return. - **Q7 — a bind can die in steady state, no cycle at all.** The gate produces no action and **nothing is emitted on any channel.** A `Return`-branch fix cannot reach this half. ## Q4's hang case — measured, and it is the sharpest result Against a `dmsetup suspend`ed device (I/O queues instead of returning `EIO`): - **P1 and P2 completed in 364 µs / 206 µs.** They read `/proc`, so no block device is involved. - **`statfs` and `getdents` completed and reported HEALTHY** — on a wedged device they do not even hang. - **Every probe that touches the device blocked, including a buffered write with no `fsync`** — the `O_CREAT` metadata path needs journal access (`wchan=do_get_write_access`). There is no cheap-and-safe write probe. - **The blocked process survived `SIGTERM` and `SIGKILL`** (`stat=D`, still alive 3m50s after `kill -9`) and died only when the device was resumed. So **`systemctl restart felhom-agent` would hang**, leaving the agent unrecoverable until the device returns or the host reboots. The thread count does not reveal the leak (5→5, 5→6). **A timeout protects the caller's control flow and nothing else.** This turns "prefer a cheap probe" into a fence: **the fix must issue no block I/O.** ## Teardown — done, all three layers Guest 9301 destroyed; `r117scratch` removed; both dm devices and both loop devices gone; `scsi_debug` unloaded (`/dev/sd*` back to `sda1..3`); no `r117` mounts, `/mnt` and `/root` exactly as found; `local` **37.02 %** against a session-start **37.00 %**. Fences re-verified *after* teardown: 9201 `running`, `drill-r50` `stopped`, `local-lvm` **38.84 % byte-identical**, `felhom-backup` `content backup` unchanged, live `/mnt/felhom-drives` intact with both submounts, agent service `active`. **Layer 3 is genuinely empty** — 9301 had no network interface and ran no controller, so no hub-side record was ever created. **Ordering trap worth keeping:** a suspended dm device must be `dmsetup resume`d *before* any `umount`, or the teardown itself blocks on the same uninterruptible sleep. ## Not measured No load or duration testing of the recommended check — P1/P2 were single calls, not a sustained reconcile loop on a many-drive box. Nothing suggests a problem (they are two `/proc` reads the code already performs), but it was not measured. ## Findings filed, none fixed R-117 (mechanism + recipe), **R-117a** steady-state death with no event (HIGH, larger than R-117 as filed), **R-117b** `statfs`/`getdents` are hollow liveness probes, **R-117c** three untested comments promising "live + usable in the guest", **R-117d** the self-heal that already exists is short-circuited (HIGH), **R-117e** both demo boxes share one failure domain — no route survives the site losing internet, including the WireGuard OOB path, **R-117f** an I/O liveness probe turns a wedged drive into an unkillable agent (HIGH — disqualifies a whole probe class).