The audit said the union row "writes false" for the guest-path key. That is wrong, and the next reader would have inherited the error. Isolated during R-116's Phase 0: RoleForStorage returns RoleSystem whenever backingDevice == "" (felhom-agent internal/storage/role.go:180-181). When the device vanishes the target row's role flips to system and it loses its guest path, but KEEPS its MountPath -- and the union loop skips any drive whose MountPath is already seen, so the registry row is never emitted at all. /disks therefore carries NO row with that guest path: isTarget[guestPath] is a MISSING KEY, not a false value. The practical difference is decisive -- the obvious fix (set BackupTarget on the union row) could not have worked, because that row does not exist in the state where the alarm is needed. The section's own "not isolated" caveat is replaced by the isolated answer.
12 KiB
SESSION-C-2026-07-29 — R-113, R-114, R-112 proven on a real box; C5 fails on a new defect
Run: RUNBOOK-session-c-2026-07-29.md, executed by CC on DooPlex, 2026-07-29.
Session C of three. A → R-113 (agent v0.114.0) · B → R-114 + R-112 (controller v0.186.0) · C → this.
| Claim | Fix | Verdict |
|---|---|---|
| C4 — the offer appears and moves the target when accepted | R-112 | ✅ PASS |
| C3 — with the target absent the customer is told the drive is missing, and offered nothing | R-114 | ✅ PASS |
C5 — drive loss produces backup_target_absent (error) and backup_target_restored (info) |
R-113 | ❌ FAIL — the alarm fires, but as the generic event; the recovery is the specific one. Unmatched pair → R-116 |
All three fixes are proven to work. R-113's gate fires, R-114's message is correct, R-112's banner reaches the customer. C5 fails on a fourth, separate defect that only became reachable because R-113 made the gate fire at all — it could not have been seen before.
1. Baselines, and what the drill box actually ran
| Artifact | Vouched / used | How it got there |
|---|---|---|
| agent | 0.114.0 | from the Day-0 manifest — Artifact manifest served for sess-c (agent=0.114.0 golden=0.185.1). C5 therefore tested the shipped binary, closing R-115 observation 1 for R-113 |
| golden | 0.185.1 | vouched; bakes controller 0.185.1 |
| controller | 0.186.0 | hand-deployed after install (§3.1 ruling (a)) via the bootstrap path — 0.185.1 → 0.186.0, Up 21 seconds (healthy) |
| hub | 0.81.0 | unchanged |
| host-install | 1.22.0 | unchanged |
Repos at run time: controller b331f18, agent b58d7bc, felhom.eu af518ba, all clean.
§3.1 ruling — (a), deploy 0.186.0 by hand. A bake is its own session and would have inserted one before any verdict; the install path is already proven by E-2d. Caveat, as required: C3/C4 prove the CODE, not the SHIPPED GOLDEN. The vouched golden lagging a controller release is R-115's shape and is filed there, not as a new ID.
§3.2 cadence — driveGateLoop registered server.go:224, defined intermediary.go:328, ticker
30 s (:339) ⇒ C5 budget 60 s. §3.3 — SCSI hot-detach, volume survives as unused0.
Fence re-measured before and after: local-lvm 38.78 % both times.
2. Timeline (VM 9301 sessionc, nested PVE on demo-hp)
| UTC | Event |
|---|---|
| 18:29 | VM created — one disk, q35/OVMF SB-off, hotplug disk, outside the felhom pool |
| 18:33:47 | registered unclaimed, pairing code 2XT-3PV |
| 18:34:08 | bound to sess-c by CC — no operator STOP (E-2d retired it) |
| 18:34:57 | manifest served: agent 0.114.0, golden 0.185.1 |
| 18:38:03 | controller_started (0.185.1) — install complete, bind→controller in 3 m 55 s |
| ~18:40 | controller 0.186.0 deployed; dashboard claimed with the emailed code |
| 18:43:50 | target drive hot-detached |
| 18:43:54 | gate fired — 4 seconds |
| 18:47:01 | reattached |
| 18:47:24 | gate RETURNED (23 s) |
| 18:50:32 | non-target drive detached (mirror) |
| 18:50:54 | gate fired (22 s) |
3. C4 — PASS (R-112 + the assign flow)
The banner reached a customer's page for the first time. E-2d could only prove the JSON.
Never-configured box, before any second drive:
GET /backups http=200
banner element : 1
never-configured copy : 1 <- "ugyanazon a lemezen van" IS in the HTML
absent copy : 0
offer block : 0 <- correct: no eligible drive registered yet
After attaching a 100 GB disk and bringing it in through the dashboard wizard:
banner element : 1 offer block : 1 data-path="/mnt/felhom-drives/mentes"
Decline path — PROVEN. After registering and not accepting: target still local, no
felhom-backup storage, agent.json local_backup_target = local. Registration confers no role.
Accept: {"assigned":"/mnt/felhom-drives/mentes","restart_required":true}.
The agent did NOT self-restart — ActiveEnterTimestamp 20:35:15 CEST, unchanged at 20:42:20.
In-flight check performed and recorded before restarting: 0 vzdump processes, 0 backup lines in
the agent journal (5 min), 0 running PVE tasks. Wrapper created dir: felhom-backup, path /mnt/mentes, is_mountpoint 1 — the drive's own mountpoint.
After restart, healthy renders nothing — and it is proven POSITIVELY, not by an empty page:
banner element 0 · offer block 0 · degraded copy 0 · absent copy 0
/backup/tiers calls before 6 · after idle 8s 6 (delta 0) · after page load 7 (delta +1)
That call has exactly one caller (resolveBackupTargetState), so the seam ran and chose silence.
4. C3 — PASS (R-114)
With the target drive detached, the customer's page:
banner element : 1
ABSENT copy (correct) : 1 <- "vissza nem csatlakoztatod"
system-disk copy (BAD): 0 <- E-2d saw THIS
offer block (BAD) : 0 <- E-2d offered the VANISHED drive
API: {"degraded":true,"target":"felhom-backup","message":"A rendszermentés meghajtója nem érhető el — amíg vissza nem csatlakoztatod, a teljes rendszermentés nem készül el."}
Both of E-2d's falsehoods are gone.
5. C5 — FAIL, and R-113 is nonetheless PROVEN
R-113 works. This is the headline. E-2d measured zero gate activity over 4½ minutes. Here:
18:43:54 [INFO] [settings] Storage path /mnt/felhom-drives/mentes disconnected=true
18:43:54 [WARN] [gate] drive ABSENT /mnt/felhom-drives/mentes — stopped+blocked 0 app(s): []
4 seconds, and SetDisconnected reached — the flag R-114 and the offer suppression depend on.
And it fired on exactly the shape that defeated it before:
raw /mnt/mentes : NOT mounted <- device-bound, died with the device
bind /mnt/felhom-drives/mentes : /dev/sdb[/felhom-data] <- the stale bind SURVIVED
The bind outlived the device, as always; the gate now sees past it.
What fails is which event is sent.
absent : Event pushed: storage_disconnected (error) <- GENERIC
return : Event pushed: backup_target_restored (info) <- SPECIFIC
C5 required backup_target_absent. It never fired — backup_target_absent count 0 across the run.
Worse than a missing label: the alarm and its recovery are a mismatched pair, which is precisely
what notifyDriveReturned's own comment forbids — "a target that alarmed as backup_target_absent has
to recover as backup_target_restored, not as a generic reconnect the operator cannot match". Here it
is inverted: generic alarm, specific recovery.
Root cause, established at source
notifyDriveAbsent (intermediary.go:635-646) branches on isTarget[path], built by
driveTargetByPath (:602-616) as out[GuestPath] = d.BackupTarget / out[MountPath] = d.BackupTarget.
The gate's a.Path is the guest path.
The drive is represented by two /disks rows, and the flag and the guest path sit on different ones:
| Row | Source | BackupTarget |
GuestPath |
|---|---|---|---|
the felhom-backup storage |
Observe path, felhom-agent/internal/localapi/disks.go:211 (t.Name == primaryTargetID) |
true | set only when role is user-data |
| the registry drive | union path, disks.go:265-267 — never assigns BackupTarget at all |
false | /mnt/felhom-drives/mentes |
CORRECTION 2026-07-29 (R-116 Phase 0). This section originally said the union row "writes
false". That is wrong, and the real mechanism is worse — the union row is not written at all.The step left un-isolated above has now been isolated:
RoleForStoragereturnsRoleSystemwheneverbackingDevice == ""(felhom-agent/internal/storage/role.go:180-181). So when the device vanishes, Observe'sexactMountDevicefails,t.BackingDevicebecomes"", the row's role flips to system and it loses its guest path — but it keeps itsMountPath. The union loop skips any drive whoseMountPathis alreadyseen(disks.go:263), so the registry row is DEDUPED AWAY ENTIRELY.
/diskstherefore carries no row with that guest path at all:isTarget[guestPath]is a missing key, not afalsevalue. The practical difference is decisive — the obvious fix (setBackupTargeton the union row) could not have worked, because that row is not emitted in the state where the alarm is needed.
On return the rows rejoin and the specific event fires — which is exactly the asymmetry observed. v0.184.1 fixed the keying (both keys are mapped); it did not fix the flag being absent from the row that carries the guest path.
Filed as R-116. Not fixed: a defect found in a validation run is a finding.
6. Mirror scenario + over-correction guard — PASS, with a caveat
A non-target drive (adat, 70 GB, registered but never assigned) detached at 18:50:32:
18:50:54 [WARN] [gate] drive ABSENT /mnt/felhom-drives/adat — stopped+blocked 0 app(s)
18:50:54 [INFO] Event pushed: storage_disconnected (error) — Meghajtó váratlanul leválasztva: Adat meghajto
backup_target_absent count : 0 storage_disconnected count : 1
Over-correction guard — the real risk of R-113 — PASSES. With both drives present the gate produced
0 ABSENT lines, and the target stayed degraded:false with banner element 0 throughout the whole
mirror sequence. R-113's stricter presence did not make a healthy drive read absent.
Caveat, stated rather than glossed: because the target drive also produced the generic event, the mirror passes trivially. It confirms no over-correction; it cannot confirm specific-vs-generic discrimination, because nothing produced the specific event on absence at all.
7. Findings
- R-116 (new) — the absent alarm is generic while the recovery is specific; the pair cannot be matched by an operator. Root cause §5. This defect was unreachable before R-113: the gate never ran, so the branch was never taken. Fixing one bug exposed the next, which is what a validation run is for.
- The stale bind survives a return, still naming a dead device node. After the drive came back as
/dev/sdc,findmnt /mnt/felhom-drives/mentesstill read/dev/sdb[/felhom-data]while the state reporteddegraded:false. E-2d noticed this and did not chase it; it is now confirmed. Whether apps writing through that path reach the live device was not established — worth its own look. - The capability map still has no E-2 / backup-target rows at all, so no row could be moved to PROVEN-LIVE. Creating them is a design act, not a validation act; filed rather than invented (third session in a row this has been noted).
8. Teardown
VM 9301 destroyed --purge, sc-images storage removed, scratch dir removed.
pvesm status after == before (local-lvm 38.78 %, felhom-backup 931059224 KiB available).
Guest 9201 and drill-r50 untouched.
Customer ruling: DELETE (R-93's lesson — no half-real fixtures). Attempted and correctly refused:
delete sess-c REFUSED: host sess-c-8158e4 is ONLINE. The VM is destroyed so no further reports come;
the host ages OK → WARN (30 m) → DOWN (>1 h) and is then deletable. Command, unchanged from E-2d:
POST /configs/sess-c/delete ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=sess-c expect_hosts=1
9. What did not happen
backup_target_absentnever fired, so its severity, its Hungarian customer copy and its hub routing are all still unexercised end-to-end. R-116 blocks them.- The offer control was accepted via the endpoint it POSTs, not by a browser click — no browser
automation exists on DooPlex. The rendered control, its
data-pathand its non-auto-submission were verified in the HTML. - The stale-bind-after-return question (finding 2) was observed, not investigated.