Completes the spike once the venue came back. Q4's hang case and teardown are
now measurements, not plans.
Against a dmsetup-suspended device (I/O queues instead of returning EIO):
- P1 (devno compare) and P2 (ext4 abort flags) completed in 364us / 206us.
They read /proc, so no block device is involved.
- statfs and getdents completed and reported HEALTHY — on a wedged device they
do not even hang. R-117b confirmed in a second failure mode.
- EVERY probe that touches the device blocked, including a buffered write with
no fsync: the O_CREAT metadata path needs journal access
(wchan=do_get_write_access). There is no cheap-and-safe write probe.
- The blocked process survived SIGTERM AND SIGKILL (stat=D,
wchan=folio_wait_bit_common, still alive 3m50s after kill -9) and died only
when the device was resumed. So `systemctl restart felhom-agent` would hang,
leaving the agent unrecoverable until the device returns or the host reboots.
The thread count does not reveal the leak (5->5, 5->6).
Filed as R-117f. A timeout protects the caller's control flow and nothing else,
so "the fix must issue no block I/O" is now a fence rather than a preference —
the thread-leak hypothesis the probes were built to test turned out to be the
weaker half of the result.
Teardown done, all three layers: guest 9301 destroyed, r117scratch removed, both
dm and both loop devices gone, scsi_debug unloaded, local back to 37.02% against
a 37.00% session start. Fences re-verified AFTER teardown: 9201 running,
drill-r50 stopped, local-lvm 38.84% byte-identical, felhom-backup content
unchanged, live /mnt/felhom-drives intact with both submounts, agent active.
Layer 3 genuinely empty — 9301 had no NIC and ran no controller.
Trap recorded: a suspended dm device must be resumed BEFORE any umount, or the
teardown blocks on the same uninterruptible sleep.
Both halves of the R-113 conjunction are path-presence tests: GuestSeesMount
(intermediary.go:276) and isHostMountpoint (:394) compare field 5 of a mountinfo
line and never read field 3, so neither can see that the bind and the raw mount
name different devices. Measured BoundUnderParent=TRUE over a namespace that
EIOs on every read and write.
Reproduced 3/3 on a purpose-built scratch LXC on demo-hp; predicates evaluated
by a throwaway probe calling the real localapi code from d4eb259.
Three results that change the shape of the fix:
- Q7: a bind can die in STEADY STATE with no detach/return cycle. The gate
produces no action and nothing is emitted on any channel. A Return-branch fix
cannot reach this half, and a devno comparison does not detect it.
- Q6/R-117d: AttachDrive's normalize leg already performs the repair, and three
call sites already invoke it - including the controller's Return branch before
it restarts apps. All defeated by one early return at :235. Unblock the
existing path; do not add a new one.
- Q1: the device-node change is a CONSEQUENCE, not a precondition. The stale
bind pins the dead superblock, forcing the returning device onto a new number.
Control test: released, the letter is reused.
Not established: the hang case. Venue and probes built, run lost to a site
internet outage; the thread-leak hypothesis is not claimed as a result.
Teardown of the spike venue is owed - commands in the findings doc; nothing
fenced was touched and no hub-side record was created.