Both halves of the disposition were run and neither survived as a finding.
The I1/I1-pair violations cluster at cycles 31-33 and nowhere else across 39
cycles; c34-c39 are clean, so it recovered with no intervention. Final tally I1
37 PASS / 2 VIOLATION, I1-pair 36 PASS / 3 VIOLATION. On the quiesced box one slow
detach with 4 minutes either side produced a perfect pair. And the alarming
false-healthy (mentes bound=False while degraded=false) does not survive
quiescence - I had been reading the two halves at different instants of a detach.
No R-n.
Separately cleared: the hub's SQLITE_BUSY event drops. 7 in 24h including one for
the real customer demo-felhom, and the hub does return 500 with notification
dispatch only after a successful save - so a lost event would be a lost alarm. But
the controller retries 3 times and ZERO events exhausted their attempts; the
07:04:39 drop landed at 07:04:42. Nothing lost. Only cosmetic note: the ERROR line
reads like data loss and is not.
The case R-117's spike called the worse half — a drive dying with no detach/return
cycle, which before agent v0.117.0 emitted nothing on any channel indefinitely.
Box runs 0.119.0. Aborted ext4 in place (abort,emergency_ro; device still present):
bound_under_parent went false, storage_disconnected fired, the storage page named
the stopped app, and calibre-web (whose library binds that drive) was STOPPED
rather than restarted onto the dead namespace. Recovery needed a full device close,
not a remount — exactly as the fix intends (BindAborted => quiet no-op).
Runner extended with the 7 atom families run 1 skipped: abort-fs-in-place,
kill-agent-mid-backup, hard-reset-VM-mid-write, reboot-VM, concurrent
backup+restore, concurrent backup+detach, fill-drive-near-full. Also fixes a run-1
flaw recorded in the audit: reboot was appended AFTER the shuffle so it never
interleaved with a detach; heavy atoms are now permuted in with the rest.
Run-1 evidence preserved as *-run1.* (cycle numbering restarts per run).
Ran the soak on the Phase A rig. Ended on its own deadline — no watchdog halt,
no atom exception, no I11 breach.
I1 28+28 pairs, I2 28+28 pairs, I3 56, I4 56, I5/I6 28 each, I7 28, I10 135,
I11 28. Zero violations. The row counts are themselves the no-silent-skip check:
I3/I4 twice per cycle (both drives), I10 = 5 secret-class fields x 27, REBOOT on
cycles 7/14/21 only.
I7 is the headline: 28 restores, 28 correct discriminators — never stale, never
empty. RTO (Tier 1, rallly, 66 MB): min 38.8s, median 42.0s, p90 42.5s, max
44.3s. That is the S band's lower end ONLY; the 5.5s spread over 28 runs says
fixed work dominates, so nothing extrapolates to M or L. RPO not measured.
Every atom and invariant was proven BY HAND before automation — the runner
asserts nothing that was not first observed live.
Caught a Phase A gap before starting: no app had HDD_PATH, so all data sat on the
system disk and I3 could never have fired. Deployed calibre-web onto adatok
first; otherwise the run would have produced 27 green cycles that tested nothing
cross-drive.
Investigated and DISPROVED a suspected defect (audit 5.2): /api/disks reports
state=attached for a physically absent drive, and intermediary.go:230 really does
compute presence from State=="attached". It is inert — planDriveGates only gates
paths under /mnt/felhom-drives/ and uses BoundUnderParent there, which was
correctly false. The gate fired; the storage page showed "Meghajtó leválasztva".
No R-n minted.
Honest gaps: 6 of ~12 atom families ran. Not run — Tier 3 (structurally
un-isolatable), abort-fs-in-place, kill-agent-mid-backup, hard-reset-mid-write,
reboot-VM, both concurrency atoms, fill-drive-near-full. I8 not checked, I9 not
automated (cited from the tester-gate run, not re-claimed). kill_controller is
NOT mid-backup and reboot_guest never interleaved with a detach. 27 cycles does
not answer the brief's question about drift at the thirty-eighth.
Teardown still OWED, including hub customer c10-soak (disposition: DELETE).