Files
admin 816c59c43a Campaign 10: R-117 Q7 (fs aborted in place, device present) proven PASS; extended atom set
The case R-117's spike called the worse half — a drive dying with no detach/return
cycle, which before agent v0.117.0 emitted nothing on any channel indefinitely.
Box runs 0.119.0. Aborted ext4 in place (abort,emergency_ro; device still present):
bound_under_parent went false, storage_disconnected fired, the storage page named
the stopped app, and calibre-web (whose library binds that drive) was STOPPED
rather than restarted onto the dead namespace. Recovery needed a full device close,
not a remount — exactly as the fix intends (BindAborted => quiet no-op).

Runner extended with the 7 atom families run 1 skipped: abort-fs-in-place,
kill-agent-mid-backup, hard-reset-VM-mid-write, reboot-VM, concurrent
backup+restore, concurrent backup+detach, fill-drive-near-full. Also fixes a run-1
flaw recorded in the audit: reboot was appended AFTER the shuffle so it never
interleaved with a detach; heavy atoms are now permuted in with the rest.

Run-1 evidence preserved as *-run1.* (cycle numbering restarts per run).
2026-08-01 17:11:58 +02:00

2.9 KiB

R-117 Q7 — filesystem aborted IN PLACE, device present: PASS (2026-08-01)

The case R-117 was filed for and the one its spike called "the worse half": a drive dies with no detach/return cycle at all, so the gate never runs and — before agent v0.117.0 — "nothing is emitted on any channel, indefinitely" (R-117a).

Box runs agent 0.119.0, so the v0.117.0 bindUsable fix is in. Proven by hand before automation.

Method

The spike used a dmsetup error table; these drives are raw SCSI passthrough, so the equivalent state was produced with ext4's own abort, which lands in exactly the vocabulary bindLiveness matches (shutdown / emergency_ro):

# mount -o remount,abort /mnt/adatok
/dev/sde /mnt/adatok ext4 rw,relatime,abort,emergency_ro 0 0     <-- fs dead
lsblk: sda sde sdf sr0                                            <-- DEVICE STILL PRESENT

Result — it surfaces, and the gate acts

Stable across three polls (~2 min), so not a transient:

Signal Value Meaning
bound_under_parent false the abort surfaced — the fix's stated mechanism
state attached correct: the drive is still physically there
hub event storage_disconnected (error) — Meghajtó váratlanul leválasztva: Adatok an event fired — R-117a's silence is gone
storage page Adatok … Leválasztva … **Leállított alkalmazások: calibre-web** the customer is told which app was stopped
calibre-web not running the gate STOPPED the app on the dead namespace rather than restarting it onto one

calibre-web's library is a bind onto this drive (/mnt/felhom-drives/adatok/userdata/media/books → /calibre-library), which is why it is the app that must stop. Before the fix this state left it running on a namespace that EIOs every byte.

Recovery — a remount is NOT enough, exactly as the fix intends

umount -l + mount UUID=… left abort,emergency_ro in /proc/mounts: the lazy umount leaves the aborted superblock held open by the stale bind (the spike hit the same thing as dmsetup remove --forceDevice or resource busy). The agent deliberately does not auto-repair here — BindAborted ⇒ quiet no-op, because a re-bind lands on the same aborted superblock and this runs every 20 s, so repairing would be an infinite silent retry. Clearing it is an operator decision.

A full device close — detach and re-attach the virtual disk, i.e. re-seating the drive — recovered it in one agent tick:

/dev/sdb /mnt/adatok ext4 rw,relatime      <-- abort flags gone, letter recycled after the close
bound_under_parent = True ; calibre-web = Up

The recycled device letter matches the spike's own observation that releasing the bind lets the letter be reused.

Verdict: I4's second clause ("an aborted filesystem surfaces and does not silently retry") and I3 ("no app is ever restarted onto a dead namespace") both hold on the hardest case in the register.