e70b5feebe
Completes the spike once the venue came back. Q4's hang case and teardown are now measurements, not plans. Against a dmsetup-suspended device (I/O queues instead of returning EIO): - P1 (devno compare) and P2 (ext4 abort flags) completed in 364us / 206us. They read /proc, so no block device is involved. - statfs and getdents completed and reported HEALTHY — on a wedged device they do not even hang. R-117b confirmed in a second failure mode. - EVERY probe that touches the device blocked, including a buffered write with no fsync: the O_CREAT metadata path needs journal access (wchan=do_get_write_access). There is no cheap-and-safe write probe. - The blocked process survived SIGTERM AND SIGKILL (stat=D, wchan=folio_wait_bit_common, still alive 3m50s after kill -9) and died only when the device was resumed. So `systemctl restart felhom-agent` would hang, leaving the agent unrecoverable until the device returns or the host reboots. The thread count does not reveal the leak (5->5, 5->6). Filed as R-117f. A timeout protects the caller's control flow and nothing else, so "the fix must issue no block I/O" is now a fence rather than a preference — the thread-leak hypothesis the probes were built to test turned out to be the weaker half of the result. Teardown done, all three layers: guest 9301 destroyed, r117scratch removed, both dm and both loop devices gone, scsi_debug unloaded, local back to 37.02% against a 37.00% session start. Fences re-verified AFTER teardown: 9201 running, drill-r50 stopped, local-lvm 38.84% byte-identical, felhom-backup content unchanged, live /mnt/felhom-drives intact with both submounts, agent active. Layer 3 genuinely empty — 9301 had no NIC and ran no controller. Trap recorded: a suspended dm device must be resumed BEFORE any umount, or the teardown blocks on the same uninterruptible sleep.