Both halves of the disposition were run and neither survived as a finding.
The I1/I1-pair violations cluster at cycles 31-33 and nowhere else across 39
cycles; c34-c39 are clean, so it recovered with no intervention. Final tally I1
37 PASS / 2 VIOLATION, I1-pair 36 PASS / 3 VIOLATION. On the quiesced box one slow
detach with 4 minutes either side produced a perfect pair. And the alarming
false-healthy (mentes bound=False while degraded=false) does not survive
quiescence - I had been reading the two halves at different instants of a detach.
No R-n.
Separately cleared: the hub's SQLITE_BUSY event drops. 7 in 24h including one for
the real customer demo-felhom, and the hub does return 500 with notification
dispatch only after a successful save - so a lost event would be a lost alarm. But
the controller retries 3 times and ZERO events exhausted their attempts; the
07:04:39 drop landed at 07:04:42. Nothing lost. Only cosmetic note: the ERROR line
reads like data loss and is not.
Run 2a hit its first two violations at cycle 10 and BOTH trace to my harness, not
the product. Recorded in full because a check that fails for the wrong reason is
as corrosive as one that passes for the wrong reason.
HARD-RESET VM returned=True canaries_intact=False
I7 want=C10-C010-A-194530 got=C10-C009-A-192929 restore_ok=True
Root cause, evidenced: the cc_proof table's highest row is C10-C009-A — there is
NO C010-A row at all, so the seed never landed. The hard-reset atom ran earlier in
the same cycle and left rallly Exited(255); atom_restore_verify called seed() and
never checked its return value, so an unwritten generation became a fake stale
The case R-117's spike called the worse half — a drive dying with no detach/return
cycle, which before agent v0.117.0 emitted nothing on any channel indefinitely.
Box runs 0.119.0. Aborted ext4 in place (abort,emergency_ro; device still present):
bound_under_parent went false, storage_disconnected fired, the storage page named
the stopped app, and calibre-web (whose library binds that drive) was STOPPED
rather than restarted onto the dead namespace. Recovery needed a full device close,
not a remount — exactly as the fix intends (BindAborted => quiet no-op).
Runner extended with the 7 atom families run 1 skipped: abort-fs-in-place,
kill-agent-mid-backup, hard-reset-VM-mid-write, reboot-VM, concurrent
backup+restore, concurrent backup+detach, fill-drive-near-full. Also fixes a run-1
flaw recorded in the audit: reboot was appended AFTER the shuffle so it never
interleaved with a detach; heavy atoms are now permuted in with the rest.
Run-1 evidence preserved as *-run1.* (cycle numbering restarts per run).
Ran the soak on the Phase A rig. Ended on its own deadline — no watchdog halt,
no atom exception, no I11 breach.
I1 28+28 pairs, I2 28+28 pairs, I3 56, I4 56, I5/I6 28 each, I7 28, I10 135,
I11 28. Zero violations. The row counts are themselves the no-silent-skip check:
I3/I4 twice per cycle (both drives), I10 = 5 secret-class fields x 27, REBOOT on
cycles 7/14/21 only.
I7 is the headline: 28 restores, 28 correct discriminators — never stale, never
empty. RTO (Tier 1, rallly, 66 MB): min 38.8s, median 42.0s, p90 42.5s, max
44.3s. That is the S band's lower end ONLY; the 5.5s spread over 28 runs says
fixed work dominates, so nothing extrapolates to M or L. RPO not measured.
Every atom and invariant was proven BY HAND before automation — the runner
asserts nothing that was not first observed live.
Caught a Phase A gap before starting: no app had HDD_PATH, so all data sat on the
system disk and I3 could never have fired. Deployed calibre-web onto adatok
first; otherwise the run would have produced 27 green cycles that tested nothing
cross-drive.
Investigated and DISPROVED a suspected defect (audit 5.2): /api/disks reports
state=attached for a physically absent drive, and intermediary.go:230 really does
compute presence from State=="attached". It is inert — planDriveGates only gates
paths under /mnt/felhom-drives/ and uses BoundUnderParent there, which was
correctly false. The gate fired; the storage page showed "Meghajtó leválasztva".
No R-n minted.
Honest gaps: 6 of ~12 atom families ran. Not run — Tier 3 (structurally
un-isolatable), abort-fs-in-place, kill-agent-mid-backup, hard-reset-mid-write,
reboot-VM, both concurrency atoms, fill-drive-near-full. I8 not checked, I9 not
automated (cited from the tester-gate run, not re-claimed). kill_controller is
NOT mid-backup and reboot_guest never interleaved with a detach. 27 cycles does
not answer the brief's question about drift at the thirty-eighth.
Teardown still OWED, including hub customer c10-soak (disposition: DELETE).