Completes the spike once the venue came back. Q4's hang case and teardown are now measurements, not plans. Against a dmsetup-suspended device (I/O queues instead of returning EIO): - P1 (devno compare) and P2 (ext4 abort flags) completed in 364us / 206us. They read /proc, so no block device is involved. - statfs and getdents completed and reported HEALTHY — on a wedged device they do not even hang. R-117b confirmed in a second failure mode. - EVERY probe that touches the device blocked, including a buffered write with no fsync: the O_CREAT metadata path needs journal access (wchan=do_get_write_access). There is no cheap-and-safe write probe. - The blocked process survived SIGTERM AND SIGKILL (stat=D, wchan=folio_wait_bit_common, still alive 3m50s after kill -9) and died only when the device was resumed. So `systemctl restart felhom-agent` would hang, leaving the agent unrecoverable until the device returns or the host reboots. The thread count does not reveal the leak (5->5, 5->6). Filed as R-117f. A timeout protects the caller's control flow and nothing else, so "the fix must issue no block I/O" is now a fence rather than a preference — the thread-leak hypothesis the probes were built to test turned out to be the weaker half of the result. Teardown done, all three layers: guest 9301 destroyed, r117scratch removed, both dm and both loop devices gone, scsi_debug unloaded, local back to 37.02% against a 37.00% session start. Fences re-verified AFTER teardown: 9201 running, drill-r50 stopped, local-lvm 38.84% byte-identical, felhom-backup content unchanged, live /mnt/felhom-drives intact with both submounts, agent active. Layer 3 genuinely empty — 9301 had no NIC and ran no controller. Trap recorded: a suspended dm device must be resumed BEFORE any umount, or the teardown blocks on the same uninterruptible sleep.
5.9 KiB
REPORT — SPIKE R-117: a dead bind that reports healthy (2026-07-30)
Written as REPORT-<topic>.md, not REPORT.md, per this repo's parallel-session rule and the
established local pattern (REPORT-r116-diag.md, REPORT-session-c.md, …). The shared REPORT.md
was not touched.
Class: Spike. Deliverable is a findings document. No production code was written; no .go file
in either repo was modified; nothing was built for deployment, published, or version-bumped.
Deliverables
| File | Action |
|---|---|
documentation/audits/SPIKE-r117-bind-liveness-2026-07-30.md |
new — the findings doc (Q1–Q7, evidence, recipe, probe comparison, recommendation) |
documentation/backlog/OPEN-ITEMS.md |
R-117 row rewritten with the mechanism, the reproduction recipe and the fix constraint — the one register edit, per the brief |
No CHANGELOG entry. This repo's changelogs are per-area (hub/, scripts/, website/); a
documentation-only change belongs to none of them. Stated rather than silently skipped, per standing
rule 4.
Baselines
agent main v0.116.0 @ d4eb259 · controller main v0.186.0 @ b331f18 ·
felhom.eu HEAD 29bcfeb · hub live 0.82.0 · golden 0.186.0 ·
demo-hp PVE 9.2.2 / kernel 7.0.2-6-pve, live agent 0.113.0 (= manifest MinAgent; never used as
the source of a behavioural claim — every predicate result came from a probe built from main).
Results
All seven questions answered empirically.
- Q1 — reproduced 3/3, two device classes, on a purpose-built scratch LXC (9301) on demo-hp.
The device-node change is a consequence of the defect, not a precondition — control test: with
the stale bind held the drive returns as
sdc(8:32); released, the letter is reused (sdb, 8:16). - Q2 — two death states: device removed ⇒ superopts gain
shutdown,EIO(5) on read and write, host and guest; device errors in place ⇒emergency_ro, writeEROFS(30), reads served from cache. The raw host mount is genuinely healthy in both. No cross-device mis-identification is possible on this path — the unit is fs-UUID-keyed. - Q3 — confirmed at source and live. Both halves of the R-113 conjunction compare field 5 of a
mountinfo line and never read field 3 (
major:minor), so neither can see that the bind and the raw mount name different devices. MeasuredBoundUnderParent = TRUEover anEIOnamespace. - Q4 — a pure-
/proccheck costs 0.16–0.45 ms, cannot hang, spins up no disk, writes nothing.statfsandgetdentsboth SUCCEED on a dead namespace — probes built on either are hollow. The hang case is below; it is the sharpest result in the run. - Q5 — the agent, and not on balance: the controller runs inside the guest and cannot see the host mount tables the check needs.
- Q6 — recovery works in place, guest never restarted (init PID identical). The repair code
already exists and three call sites already invoke it, including the controller's
Returnbranch before it restarts apps — all defeated by one early return. - Q7 — a bind can die in steady state, no cycle at all. The gate produces no action and nothing
is emitted on any channel. A
Return-branch fix cannot reach this half.
Q4's hang case — measured, and it is the sharpest result
Against a dmsetup suspended device (I/O queues instead of returning EIO):
- P1 and P2 completed in 364 µs / 206 µs. They read
/proc, so no block device is involved. statfsandgetdentscompleted and reported HEALTHY — on a wedged device they do not even hang.- Every probe that touches the device blocked, including a buffered write with no
fsync— theO_CREATmetadata path needs journal access (wchan=do_get_write_access). There is no cheap-and-safe write probe. - The blocked process survived
SIGTERMandSIGKILL(stat=D, still alive 3m50s afterkill -9) and died only when the device was resumed. Sosystemctl restart felhom-agentwould hang, leaving the agent unrecoverable until the device returns or the host reboots. The thread count does not reveal the leak (5→5, 5→6).
A timeout protects the caller's control flow and nothing else. This turns "prefer a cheap probe" into a fence: the fix must issue no block I/O.
Teardown — done, all three layers
Guest 9301 destroyed; r117scratch removed; both dm devices and both loop devices gone; scsi_debug
unloaded (/dev/sd* back to sda1..3); no r117 mounts, /mnt and /root exactly as found; local
37.02 % against a session-start 37.00 %. Fences re-verified after teardown: 9201 running,
drill-r50 stopped, local-lvm 38.84 % byte-identical, felhom-backup content backup
unchanged, live /mnt/felhom-drives intact with both submounts, agent service active. Layer 3 is
genuinely empty — 9301 had no network interface and ran no controller, so no hub-side record was ever
created.
Ordering trap worth keeping: a suspended dm device must be dmsetup resumed before any umount,
or the teardown itself blocks on the same uninterruptible sleep.
Not measured
No load or duration testing of the recommended check — P1/P2 were single calls, not a sustained
reconcile loop on a many-drive box. Nothing suggests a problem (they are two /proc reads the code
already performs), but it was not measured.
Findings filed, none fixed
R-117 (mechanism + recipe), R-117a steady-state death with no event (HIGH, larger than R-117 as
filed), R-117b statfs/getdents are hollow liveness probes, R-117c three untested comments
promising "live + usable in the guest", R-117d the self-heal that already exists is short-circuited
(HIGH), R-117e both demo boxes share one failure domain — no route survives the site losing internet,
including the WireGuard OOB path, R-117f an I/O liveness probe turns a wedged drive into an
unkillable agent (HIGH — disqualifies a whole probe class).