Nothing in the repo said which machines are safe to break. The host table gave access and role and stopped there, so a session needing a victim had to guess -- and the guessing inverted: the two boxes that exist to be broken were treated as sacred, and DooPlex (the recovery chain) got used because it was the only box no spec had fenced. New documentation/runbooks/target-selection.md -- one page, three tiers, and per machine what is freely permitted / needs care / forbidden, each carrying its REASON so a rule can be correctly narrowed later instead of ossifying. States the selection rule positively (start at Tier 0; a Tier 2 box only when a task says so explicitly; an absent fence is not permission) and that fences name ACTS, not machines -- demo-hp's over-subscribed local-lvm is one dangerous storage, not a dangerous box. CLAUDE.md: host table gains a Blast radius column, gains the missing demo-hp row (it was where the drill VMs ran and it was not in the table at all), and a pointer line to the new runbook. CORRECTION to the spec's problem statement: the designation was not missing. The 2026-07-25 operator ruling naming the t740 as drill+build VM host -- explicitly "moved off DooPlex" -- already existed in operations/nodes.md. It sat where no session reads at start, while the prohibitions were repeated in every task spec. The defect is reachability of the ruling, not its absence, and the R-116 drill on DooPlex contradicted a written ruling rather than filling a vacuum. CORRECTION to the R-116 record, same commit: the baseline claimed controller 0.186.0 on both demo boxes. Only felhom-pve was sampled and generalised; demo-hp re-checked directly runs 0.185.1, so the fleet is split and R-114's TargetAbsent branch is absent from demo-hp. Fixed in the audit table and REPORT-r116-diag. Docs only -- no code, no build, no deploy, no host reconfigured, no host renamed.
33 KiB
DIAG-r116-disks-payload-2026-07-30 — the /disks payload, and the mechanism isolated from it
Run: read-only diagnosis by CC on DooPlex, 2026-07-30. No code written, nothing shipped.
Result: the payload was captured in both states and the mechanism is isolated. It is theory #1
— the registry-union row writes false — the theory that was raised, declared wrong, and retracted.
The retraction was the error.
1. Baselines
All confirmed this run, not carried from the brief.
| Artifact | Version | How confirmed |
|---|---|---|
| agent | 0.115.0 | felhom-agent --version on felhom-pve; felhom-agent/CHANGELOG.md:1 = ## v0.115.0; repo HEAD a58239f6ded367fe27b4eb57cf2fcb0c54414b80 |
| controller | 0.186.0 on demo-felhom, 0.185.1 on demo-hp ⚠ | pct exec 9201 -- docker ps: felhom-pve → felhom-controller:0.186.0, demo-hp → felhom-controller:0.185.1 (both re-verified 2026-07-30); CHANGELOG.md:3 = ### v0.186.0; HEAD b331f18424b7cf8d0883e7709f74595808264080 |
| hub | 0.81.0 | manifests/hub.yaml:128 pin and the live pod image (kubectl -n felhom-system get pod hub-… -o jsonpath=…) both felhom-hub:0.81.0; hub/CHANGELOG.md:1 = ## v0.81.0 |
| host-install | 1.22.0 | scripts/felhom-host-install.sh:187 SCRIPT_VERSION="1.22.0" |
felhom.eu HEAD |
c3ce4c7b20ce4c0d3673590d710d13e5d28b68bb |
git rev-parse HEAD |
⚠ The controller version is not one number. 0.185.1 is what the golden bakes (so it is what a
fresh drill box installs, which is why the brief and Part 5 both recorded it); 0.186.0 shipped
2026-07-29 as R-114 + R-112 and is what demo-felhom runs — but demo-hp is still on 0.185.1, so
the fleet is split. The distinction matters because R-114's TargetAbsent branch exists only in 0.186.0,
i.e. demo-hp does not have it.
Correction, 2026-07-30. As first written this row said 0.186.0 was what both demo boxes run. That was wrong: only felhom-pve's guest was sampled, and the claim was generalised to the fleet from one box. demo-hp re-checked directly →
felhom-controller:0.185.1. Corrected in the same commit that addedrunbooks/target-selection.md.
Hub host table, read this run: demo-felhom-8363b5 agent 0.115.0 ONLINE · demo-hp-bb76ea agent
0.113.0 ONLINE · drill-r50-0a4f9a DOWN · sess-c-8158e4 DOWN · sess-d-0452c4 STALE.
2. §3 reconfirmation — every cited fact, at file:line
| Cited fact | Line now | Holds? |
|---|---|---|
observe.go — mountPath comes from the mount table, assigned only on exact-mount success |
internal/storage/observe.go:184-190 |
HOLDS, verbatim as quoted |
role.go — if backingDevice == "" { return RoleSystem } for USB/LocalDir |
internal/storage/role.go:180-181 |
HOLDS, verbatim |
primaryTargetID := s.primaryTier().TargetID, in scope for both loops |
internal/localapi/disks.go:194 |
HOLDS |
BackupTarget: t.Name == primaryTargetID |
disks.go:211 |
HOLDS |
GuestPath set only inside if di.Role == RoleUserData |
disks.go:215-217 |
HOLDS |
first loop out = append(out, di) |
disks.go:284 (brief said :249) |
HOLDS — line moved, v0.115.0 inserted :226-260 |
seen[d.MountPath] = true built by iterating out |
disks.go:290-295 (brief said :258) |
HOLDS — line moved. Note the guard: seen is written only when d.MountPath != "" (:292) |
union dedup if d.MountPath == "" || seen[d.MountPath] { continue } |
disks.go:298 (brief said :263) |
HOLDS — line moved |
union row hardcodes Role: RoleUserData |
disks.go:304 |
HOLDS |
union row's GuestPath set from StablePathForRaw(d.MountPath) |
disks.go:309-310 |
HOLDS |
BackupTarget is absent from the union struct literal |
disks.go:301-306 |
HOLDS — the literal sets Name/Type/State/MountPath/DurableID/Role/GuestAttached and nothing else |
v0.115.0's fallback di.GuestPath = StablePathForRaw(t.MountPath) |
disks.go:258-260 |
HOLDS |
StablePathForRaw("") == "" (so the fallback sets nothing when MountPath is empty) |
internal/localapi/intermediary.go:69-75 → DriveNameFromRaw:79-88 (no /mnt/ prefix → "") |
HOLDS |
PVE reports the detached target dir inactive 0 0 0 |
reproduced live, §5 | HOLDS |
One fact the brief did not state, and it is load-bearing. /disks' union source is the
systemd-unit-file registry, not PVE and not the mount table:
storage.NewRegistryKnownTargets(driveUnitDir, …) (cmd/felhom-agent/main.go:605) is the value passed
as DriveTargets (main.go:764 → buildLocalAPIServer → disks.go:296 s.driveTargets.Known()).
Known() (internal/storage/registry_known.go:40-75) os.ReadDirs /etc/systemd/system, parses each
Felhom .mount unit, and returns MountPath: spec.Where (:70) — it never consults the mount table
or the device. So the union row's MountPath survives the device's disappearance by construction.
Verified on the live box (felhom-pve), unit present and device-bound:
/etc/systemd/system/mnt-hdd_1.mount
What=/dev/disk/by-uuid/47a3361a-91e0-4831-a69d-27f540ed3f48
Where=/mnt/hdd_1
/var/lib/felhom-agent/drive-intents.json → intent recorded (not `new`, so Known() includes it)
3. The working /disks invocation — verbatim
Two sessions were lost to this. The token's plaintext exists in exactly one place: the guest's
bootstrap mount, which the agent owns host-side. The agent's own token store holds only SHA-256
hashes (internal/localapi/tokenstore.go:26-32, :136), so it is not a source — that is what the
earlier extraction attempts were fighting.
Authoritative source, on the Proxmox host, as root:
/var/lib/felhom-agent/guests/<VMID>/bootstrap/bootstrap.json # 0600, owner 100000:100000
(written by internal/provision/backhalf.go:124-130; the in-guest view is
/etc/felhom-bootstrap/bootstrap.json, mp9, read-only — the host path avoids the guest hop entirely).
The invocation. Run it ON the Proxmox host so the token never crosses a hop or reaches a log:
ssh felhom-pve
export LC_ALL=C
B=/var/lib/felhom-agent/guests/9201/bootstrap/bootstrap.json
TOK=$(python3 -c "import json;print(json.load(open('$B'))['local_api']['token'])")
EP=$(python3 -c "import json;print(json.load(open('$B'))['local_api']['endpoint'])") # 169.254.253.1:8443
curl -sS -k -H "Authorization: Bearer $TOK" "https://$EP/disks" | python3 -m json.tool
-k is correct and not a shortcut: the leaf is self-signed and the controller pins it by SHA-256
(local_api.fingerprint in the same file, 2d325e89…c191 on felhom-pve) rather than by a CA. The
endpoint is the island bridge vmbr9 (169.254.253.1/30), reachable from the host itself.
3.1 The control run — it proves the query works
Present-drive control, live felhom-pve, before anything else:
endpoint=169.254.253.1:8443
HTTP 200 bytes=2483
rows: 4 (local-lvm, felhom-pbs, felhom-backup, local)
The backup-target row in that control run:
{ "name": "felhom-backup", "type": "local-dir", "state": "attached",
"backing_device": "/dev/sdb", "mount_path": "/mnt/hdd_1", "role": "user-data",
"backup_target": true, "guest_path": "/mnt/felhom-drives/hdd_1",
"bound_under_parent": true, "durable_id": "uuid:47a3361a-91e0-4831-a69d-27f540ed3f48" }
Non-empty, plausible, four rows. Part 5's failure mode (0 rows on a present drive) is excluded.
4. Where the absent state was staged, and why there
The demo boxes' targets are fenced, and drill-r50 is the drift fixture. No new box was built. The
absent state was staged on the existing nested-Proxmox drill VM on DooPlex
(/mnt/5_hdd/felhom.eu/drill/drill.qcow2, PVE 9.2, snapshot virgin, the golden-bake fixture) — my own
host, zero production exposure, and it can hot-unplug a disk for a genuine device loss rather than a
simulated one.
Fidelity — what was real:
- the byte-identical live agent binary:
scpfrom felhom-pve/usr/local/bin/felhom-agent,sha256 f48544ad5c4242a719c080ff83d7a6bc759e14272987416eb26c9683386803b7,--version0.115.0. - the real
configs/felhom-backup-target-apply(E-2a) andfelhom-mkfs-guardedwrappers. - a real LXC guest (9201), in the
felhompool, with the mp8 parent bind created exactly asprovision.BackHalfdoes it (backhalf.go:200-206:pct set 9201 -mp8 /mnt/felhom-drives,mp=/mnt/felhom-drives). - every state-producing step through the real endpoints, not by hand:
POST /disks/format→POST /disks/assign→POST /disks/guest-attach→POST /backup/target. That produced the.mountunit,drive-intents.json: enrolled, the guest bind, thedir: felhom-backup / path /mnt/cel / is_mountpoint 1storage and thelocal_backup_targetrewrite. - device loss by QMP
device_delon the virtio-blk frontend — the device node vanishes.
What was NOT production-shaped, stated plainly: (a) the agent ran as root with
privileged.mode: "direct" and a root@pam PVE token instead of non-root + sudoers + felhom-agent@pve
— this changes who runs the privileged ops, not what /disks computes; (b) hub.url pointed at a
dead local port, so nothing was reported to the hub; (c) the local-API bearer token record was
hand-written into local-tokens.log (one {"v":9201,"h":"<sha256>"} line) because minting is only
reachable through --selftest=provision, which needs a golden and a real customer. That is the key to
the door, not the state under measurement — every drive/storage/target fact was produced by the real
endpoints. (d) No controller ran, so §7 correlates against the two prior live runs rather than a fresh one.
The rig validated itself against production before being trusted: its present-state target row is
field-for-field identical in shape to felhom-pve's control run (role: user-data,
backup_target: true, guest_path set, bound_under_parent: true, durable_id: uuid:…) — only the
device node and mountpoint names differ.
5. Host state at each capture
| PRESENT (05:53:49Z) | ABSENT (05:54:32Z) | RETURNED (05:55:45Z) | |
|---|---|---|---|
| device node | /dev/vdb |
gone (ls /dev/vdb → No such file) |
/dev/vdc |
findmnt /mnt/cel |
/dev/vdb ext4 |
not mounted (rc=1) | /dev/vdc ext4 |
guest findmnt /mnt/felhom-drives/cel |
/dev/vdb[/felhom-data] |
/dev/vdb[/felhom-data] …,shutdown |
/dev/vdb[/felhom-data] …,shutdown |
pvesm status |
felhom-backup dir active |
felhom-backup dir inactive 0 0 0 |
felhom-backup dir active |
The detached state reproduced PVE's exact Part-5 string:
unable to activate storage 'felhom-backup' - directory is expected to be a mount point
but is not mounted: '/mnt/cel'
and the agent's exact Session-C / E-2d lines:
07:54:08 WARN storage: watchdog detected target state change from=attached to=disconnected
07:54:09 INFO storage: enrolled drive absent by UUID — not re-asserting (will retry when it enumerates)
07:54:09 WARN reconcile: enrolled drive not present (durable-id absent) — skipping
The R-113 asymmetry is present as measured before: raw /mnt/cel NOT mounted while the guest's bind
still names /dev/vdb[/felhom-data].
6. The raw payloads
6.1 PRESENT — GET /disks, HTTP 200, 1824 bytes, 3 rows
{"ok":true,"data":{"disks":[
{"name":"felhom-backup","type":"local-dir","state":"attached","backing_device":"/dev/vdb","mount_path":"/mnt/cel","class":"slow","role":"user-data","data_bearing":true,"data_reason":"device is mounted","total_bytes":4143677440,"used_bytes":1081344,"used_fraction":0.000260962397690878,"durable_id":"uuid:694034cc-eee6-4b17-901f-708f50df9b09","wipe_durable_id":"byid:virtio-R116CEL","guest_attached":false,"backup_target":true,"guest_path":"/mnt/felhom-drives/cel","bound_under_parent":true,"smart":{"health":"UNKNOWN","temperature_c":null,"power_on_hours":null,"reallocated_sectors":null,"pending_sectors":null,"offline_uncorrectable":null,"critical_warning":null,"media_errors":null,"percentage_used":null}},
{"name":"local-lvm","type":"lvmthin","state":"attached","backing_device":"","mount_path":"","class":"","role":"system","data_bearing":false,"total_bytes":81533075456,"used_bytes":807177447,"used_fraction":0.00989999999982339,"durable_id":"pve/data","guest_attached":false,"bound_under_parent":false,"smart":{"health":"UNKNOWN","temperature_c":null,"power_on_hours":null,"reallocated_sectors":null,"pending_sectors":null,"offline_uncorrectable":null,"critical_warning":null,"media_errors":null,"percentage_used":null}},
{"name":"local","type":"local","state":"attached","backing_device":"","mount_path":"","class":"","role":"system","data_bearing":false,"total_bytes":49675956224,"used_bytes":4584443904,"used_fraction":0.092286978499774,"durable_id":"path:/var/lib/vz","guest_attached":false,"bound_under_parent":false,"smart":{"health":"UNKNOWN","temperature_c":null,"power_on_hours":null,"reallocated_sectors":null,"pending_sectors":null,"offline_uncorrectable":null,"critical_warning":null,"media_errors":null,"percentage_used":null}}
],"guest_boot_id":"1785390256-46753","vmid":9201}}
6.2 ABSENT — GET /disks, HTTP 200, 2040 bytes, 4 rows (one MORE than present)
{"ok":true,"data":{"disks":[
{"name":"felhom-backup","type":"local-dir","state":"disconnected","backing_device":"","mount_path":"","class":"","role":"system","data_bearing":false,"total_bytes":0,"used_bytes":0,"used_fraction":0,"durable_id":"path:/mnt/cel","guest_attached":false,"backup_target":true,"bound_under_parent":false,"smart":{"health":"UNKNOWN","temperature_c":null,"power_on_hours":null,"reallocated_sectors":null,"pending_sectors":null,"offline_uncorrectable":null,"critical_warning":null,"media_errors":null,"percentage_used":null}},
{"name":"local-lvm","type":"lvmthin","state":"attached","backing_device":"","mount_path":"","class":"","role":"system","data_bearing":false,"total_bytes":81533075456,"used_bytes":807177447,"used_fraction":0.00989999999982339,"durable_id":"pve/data","guest_attached":false,"bound_under_parent":false,"smart":{"health":"UNKNOWN","temperature_c":null,"power_on_hours":null,"reallocated_sectors":null,"pending_sectors":null,"offline_uncorrectable":null,"critical_warning":null,"media_errors":null,"percentage_used":null}},
{"name":"local","type":"local","state":"attached","backing_device":"","mount_path":"","class":"","role":"system","data_bearing":false,"total_bytes":49675956224,"used_bytes":4584579072,"used_fraction":0.0922896994942001,"durable_id":"path:/var/lib/vz","guest_attached":false,"bound_under_parent":false,"smart":{"health":"UNKNOWN","temperature_c":null,"power_on_hours":null,"reallocated_sectors":null,"pending_sectors":null,"offline_uncorrectable":null,"critical_warning":null,"media_errors":null,"percentage_used":null}},
{"name":"694034cc-eee6-4b17-901f-708f50df9b09","type":"usb","state":"attached","backing_device":"","mount_path":"/mnt/cel","class":"","role":"user-data","data_bearing":false,"total_bytes":49675956224,"used_bytes":4584579072,"used_fraction":0.09228969949420011,"durable_id":"uuid:694034cc-eee6-4b17-901f-708f50df9b09","guest_attached":false,"guest_path":"/mnt/felhom-drives/cel","bound_under_parent":false}
],"guest_boot_id":"1785390256-46753","vmid":9201}}
6.3 RETURNED — GET /disks, HTTP 200, 1825 bytes, 3 rows (back to one row for the drive)
{"ok":true,"data":{"disks":[
{"name":"felhom-backup","type":"local-dir","state":"attached","backing_device":"/dev/vdc","mount_path":"/mnt/cel","class":"slow","role":"user-data","data_bearing":true,"data_reason":"device is mounted","total_bytes":4143677440,"used_bytes":1081344,"used_fraction":0.000260962397690878,"durable_id":"uuid:694034cc-eee6-4b17-901f-708f50df9b09","wipe_durable_id":"byid:virtio-R116CEL","guest_attached":false,"backup_target":true,"guest_path":"/mnt/felhom-drives/cel","bound_under_parent":true,"smart":{"health":"UNKNOWN", …}},
{"name":"local-lvm", … as above … },
{"name":"local", … as above … }
],"guest_boot_id":"1785390256-46753","vmid":9201}}
(the local-lvm / local rows in 6.3 are byte-identical to 6.1 apart from used_bytes; elided only
there, and nowhere else in this document.)
7. The per-row table for the backup-target drive
PAYLOAD. Every cell read from the JSON above.
PRESENT — 1 row mentions the drive
| Name | Type | State | Role | MountPath | GuestPath | BackingDevice | BackupTarget | BoundUnderParent |
|---|---|---|---|---|---|---|---|---|
felhom-backup |
local-dir |
attached |
user-data |
/mnt/cel |
/mnt/felhom-drives/cel |
/dev/vdb |
true |
true |
ABSENT — 2 rows mention the drive, and the two facts are on DIFFERENT rows
| Name | Type | State | Role | MountPath | GuestPath | BackingDevice | BackupTarget | BoundUnderParent | Source |
|---|---|---|---|---|---|---|---|---|---|
felhom-backup |
local-dir |
disconnected |
system |
"" |
"" (field absent) |
"" |
true |
false |
Observe loop, disks.go:196-284 |
694034cc-…9b09 |
usb |
attached (hardcoded :302) |
user-data (hardcoded :304) |
/mnt/cel |
/mnt/felhom-drives/cel |
"" |
field ABSENT ⇒ false |
false |
union loop, disks.go:297-339 |
durable_id also degrades on the Observe row: uuid:694034cc… → path:/mnt/cel.
RETURNED — 1 row again
| Name | Type | State | Role | MountPath | GuestPath | BackingDevice | BackupTarget | BoundUnderParent |
|---|---|---|---|---|---|---|---|---|
felhom-backup |
local-dir |
attached |
user-data |
/mnt/cel |
/mnt/felhom-drives/cel |
/dev/vdc |
true |
true |
8. The mechanism
Which row the lookup lands on, which key it uses, what that row says.
The controller's key is a.Path — the registered StoragePath, which for an external drive is the
GUEST path /mnt/felhom-drives/cel (intermediary.go:297, :318). driveTargetByPath
(intermediary.go:602-618) builds the map with plain assignment, one entry per non-empty
GuestPath (:612) and per non-empty MountPath (:615).
PAYLOAD — the map, computed by applying intermediary.go:602-618 verbatim to the captured JSON:
PRESENT rows=3 isTarget["/mnt/felhom-drives/cel"] = True → SPECIFIC
map: {'/mnt/felhom-drives/cel': True, '/mnt/cel': True}
ABSENT rows=4 isTarget["/mnt/felhom-drives/cel"] = False → GENERIC
map: {'/mnt/felhom-drives/cel': False, '/mnt/cel': False}
RETURNED rows=3 isTarget["/mnt/felhom-drives/cel"] = True → SPECIFIC
map: {'/mnt/felhom-drives/cel': True, '/mnt/cel': True}
That is the live asymmetry — generic on detach, specific on return — reproduced from the payload alone, matching Session C and Part 5 exactly.
Why it fails, step by step. Steps 1–4 are payload; step 5 is the map above.
- The device goes.
exactMountDevicefails, sobackingDeviceandmountPathare both""(observe.go:184-190). Payload: Observe row hasbacking_device:"",mount_path:"". RoleForStoragetherefore returnsRoleSystem(role.go:180-181), so theGuestPathblock atdisks.go:215-217is skipped. Payload:role:"system", noguest_path.- v0.115.0's fallback (
disks.go:258-260) fires — its three gates all pass — but it callsStablePathForRaw(t.MountPath)withMountPath == "", which returns""(intermediary.go:69-75). It assigns an empty string. Payload confirms: noguest_pathon that row. This is exactly why v0.115.0 is inert. seenis built only from rows whoseMountPath != ""(disks.go:290-295). The Observe row's is empty, so/mnt/celis not inseen, so the union row is not deduped at:298. It is emitted withMountPathfrom the.mountunit file andGuestPathfrom:309-310— andBackupTargetis not in its struct literal (:301-306), so it serialises asfalse. Payload: row 4 exists, carries both paths, andbackup_targetis absent.- So the row holding
backup_target:truecontributes no key at all (both its keys are empty strings and are skipped at:611and:614), while the row that owns the key/mnt/felhom-drives/celsaysfalse.isTarget[a.Path]→false→notifyDriveAbsenttakes the else branch (intermediary.go:645) → genericstorage_disconnected. - On return the two rows re-merge into one (the Observe row regains
MountPath, so the union row is deduped again) — that single row carriesbackup_target:trueand the guest path, sonotifyDriveReturnedtakes the specific branch (:656) →backup_target_restored. Hence the unmatchable pair.
INFERENCE (small, and labelled): step 5's map-construction and step 6's branch selection were computed by applying the controller's published source to the captured payload, not observed from a running controller this session. The observable they predict — generic out, specific back — is exactly what Session C and Part 5 did measure live on real boxes, twice. So the prediction is corroborated, but the controller was not itself re-run here.
There is no map-overwrite race. An earlier worry was that two rows might both write the same key
with the last one winning. They do not: the Observe row contributes nothing, so the union row is the
sole writer of /mnt/felhom-drives/cel. Row order is irrelevant.
9. Which theory this is
| Theory | Verdict from payload |
|---|---|
#1 — "the union row writes false" (raised, then declared wrong and retracted) |
✅ CORRECT. The union row is emitted (rows 3 → 4), owns the key, and its backup_target field is absent ⇒ false. The retraction was the error, not the theory. |
| #2 — "no row carries the guest path; the union row is deduped away" (the basis of the shipped v0.115.0 fix) | ❌ FALSE on both halves. A row does carry guest_path:"/mnt/felhom-drives/cel" (row 4), and the union row is not deduped — precisely because MountPath does not survive on the Observe row, which is the fact #2 was built on. #2 had the premise right and drew the opposite conclusion from it. |
#3 — "isTarget["/mnt/<n>"] is already true in the absent state, so the controller could fall back to it" |
❌ FALSE. isTarget["/mnt/cel"] = false as well — the same union row writes both keys. A fallback to the raw path changes nothing. |
So: #1, and it was available all along. The one genuinely new fact is why the dedup does not fire
— seen is keyed on the very field the absent state empties (disks.go:290-295 vs :298), which makes
the union row's emission a consequence of the vanished device rather than an independent condition.
10. What a fix would have to change — constraints for the next session's spec
Not implemented here. Stated as constraints, not as a design.
- The necessary and sufficient condition is that the row owning the key
/mnt/felhom-drives/<name>must reportbackup_target: truewhen that drive is the primary target. Everything else is a means. Whether that is achieved by carrying the flag onto the union row, by merging the two rows, or by makingdriveTargetByPathOR rather than assign, is the spec's choice — but it must be verified against a payload, not a fixture. disks.go:301-306omitsBackupTarget— and the omission is invisible. A struct literal that silently defaults a bool is the whole defect. Whatever the fix, this literal is the site. Note the union loop has no access tot(there is no Observe target for it) — it would have to compared.Name/d.MountPathagainstprimaryTargetID/the target's path, andd.Namefor a union row is the drive's UUID, not the storage id (payload:"694034cc-…"vs"felhom-backup"), so a naived.Name == primaryTargetIDcomparison will silently never match. That trap is worth stating in the spec explicitly.- Fixing only the agent may be enough, but the controller half is worth a look in the same pass.
driveTargetByPathuses assignment, so it is order-dependent the moment two rows ever carry the same key with different values — which is exactly the state a partial fix would create. An OR (or a documented single-writer invariant with a test) removes the hazard. - v0.115.0's fallback should not simply be extended. It is gated on
t.BackingDevice == ""and feedsStablePathForRaw(t.MountPath); in the failing stateMountPathis also"", so there is no value on that row to derive a guest path from. Any fix on the Observe row needs a path source that survives device loss — and the only one that does is the same.mountunit the union row already reads. That argues for fixing the union row, not the Observe row. - Do not let the fix make an absent drive read as present.
planDriveGatescomputespresent[gp] = present[gp] || d.BoundUnderParent(intermediary.go:226). The union row'sBoundUnderParentisfalsein the payload (R-113's conjunction holds), and it must stay false. The existingTestAbsentTargetRowDoesNotRegisterPresencepins this for the Observe row; the union row needs the equivalent. - The red-proof must be the payload shape, not the Session-C fixture. v0.115.0 was unit-tested and
red-proofed against a fixture that supplied a
MountPathand adisconnectedstate — a shape that does not occur. The fixture for the next fix is §6.2 of this document: 4 rows, the target's flag and its guest path on different rows,state:"attached"on the union row. - Whether to keep v0.115.0 is now answerable. It is provably inert (it assigns
""). It is not harmful. Reverting it is a separate change with its own risk; extending it is ruled out by (4). Recommendation, for the spec to accept or reject: leave it, and note in its comment that the case it targets cannot arise — or delete it in the same commit that fixes the union row, so the tree does not carry two explanations of one defect.
11. R-117 (new) — the returned drive's guest bind is a DEAD MOUNT, and everything reports healthy
The brief's §9 "stale bind, unchased after three runs". It is not cosmetic. Captured this run, and it is worse than a wrong device name.
PAYLOAD + live probe, RETURNED state (drive back, host healthy on /dev/vdc):
host findmnt -no SOURCE /mnt/cel → /dev/vdc
guest findmnt -no SOURCE,OPTIONS /mnt/felhom-drives/cel → /dev/vdb[/felhom-data] rw,relatime,shutdown
guest ls /mnt/felhom-drives/cel → ls: reading directory …: Input/output error
guest write into it → cannot create …: Input/output error (GUEST-WRITE-FAIL)
host ls /mnt/cel/felhom-data/ → fine, empty, live
/disks says: state "attached", role "user-data", bound_under_parent TRUE, backup_target true
The consequence, and this is why it outranks R-116. planDriveGates computes
present["/mnt/felhom-drives/cel"] = … || d.BoundUnderParent (intermediary.go:226), which is true.
The drive is currently Disconnected, so the gate takes the Return branch
(intermediary.go:281, :301-320): it clears the disconnected flag and restarts the gate-stopped
apps onto a mount that returns EIO for every operation — then reports the drive healthy and emails
backup_target_restored. A customer's apps come back up pointed at a dead namespace with no alarm on
any channel.
R-113's conjunction cannot catch it: half 1 (guest mountinfo has an entry at the guest path) is
satisfied by the stale entry, and half 2 (devicePresent("/mnt/cel")) is satisfied by the new
device. Neither half compares the two. disks.go:158-165's comment claims only that the conjunction
fixes the boot-ordering and outlived-bind cases — it does not claim this, so this is a gap rather than a
false invariant.
Reproduced deliberately and cleanly here; consistent with the same asymmetry seen (uninvestigated) in E-2d, Session C and Part 5. Filed, not chased, per the brief.
12. R-118 (new) — the absent drive's union row reports the ROOT filesystem's size as its own
PAYLOAD, §6.2 row 4: total_bytes: 49675956224, used_bytes: 4584579072 — byte-identical to the
local row (durable_id: path:/var/lib/vz, i.e. pve-root) in the same response. The real drive is
4 GB (4143677440 in §6.1).
Cause, from source: statfsCapacity(d.MountPath) at disks.go:335-338 statfs's /mnt/cel, which with
the device gone is a bare directory on the root filesystem. observe.go:176-183's comment warns
about exactly this trap and guards the Observe path against it ("an unmounted removable dir-storage's
mountpoint reverts to a bare directory on root … would be a catastrophic DR mis-id"); the union path
at :335 has no equivalent guard.
Not a DR mis-id — durable_id on that row is still the correct uuid:694034cc…, so the re-attach
identity is safe. It is a false capacity: a detached 4 GB drive advertises 46 GiB with 9.2 % used, to
any consumer of total_bytes/used_fraction (fill monitors, the storage cards). Same class as the
role.go:180-181 flip that drives R-116 — an absent drive's fields decaying to the root filesystem's.
13. Teardown
| Before | After | |
|---|---|---|
drill VM pvesm status |
felhom-backup active, local-lvm 0.99 %, local 9.23 % |
felhom-backup removed, local-lvm 0.00 %, local 9.23 % |
drill VM pct list |
9201 running |
empty |
- guest 9201
pct destroy --purge(logical volume removed, config purged);pvesm remove felhom-backup;mnt-cel.mountdisabled + deleted +daemon-reload;shred -uon the local-API token file and the PVE token file. poweroff, qemu exited (0procs),qemu-img snapshot -a virgin drill.qcow2→ the golden-bake fixture is exactly as found (virgin, ID 1, 2026-07-03 14:12:15).- scratch artifacts deleted:
r116-cel.qcow2,qmp.sock,qemu.pid, both console.ppmdumps. - DooPlex
df -h /mnt/5_hdd: 2.1T used / 6.6T avail, 24 % — same as at session start. - The disposable
root@pam!r116PVE token and the drill local-API token died with the snapshot revert; neither is recorded in any committed file.
Fences — re-verified after teardown, not assumed:
| Fence | State |
|---|---|
demo-hp local-lvm |
38.83 % — identical before, after, and to Part 5's reading |
demo-hp drill-r50 (VM 300) |
stopped — untouched |
| demo-hp guest 9201 | running |
felhom-pve felhom-backup |
active, /mnt/hdd_1 on /dev/sdb — target not re-pointed |
| felhom-pve guest 9201 | running |
| Both demo boxes | read-only access only — GET /disks, file reads, pvesm status, pct list, one scp of the agent binary |
| Code | no .go file edited, nothing built, nothing published, v0.115.0 untouched |
sess-d, and a leftover the prior runs did not record
sess-d-0452c4 now reads STALE, not ONLINE. The delete gate refuses only on
s.hostStatus(...) == "ok" (hub/internal/web/customer_delete.go:220-228) — ONLINE — so sess-d is
now deletable. Not executed here: deleting a customer runs external teardown plus a DB purge, and the
brief scoped this session to recording the command.
POST /configs/sess-d/delete ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=sess-d expect_hosts=1
Also still present, and not mentioned in the Session-C audit: sess-c / host sess-c-8158e4
(agent 0.114.0, DOWN). Same shape, same gate, deletable on the same terms —
confirm_id=sess-c expect_hosts=1. Two drill customers are now accumulating in the hub; that is the
operator's cleanup call, not a register row.
14. What did not happen, and why
- No controller was run, so no fresh gate-log / hub-event correlation was captured (brief §6.4). Deliberate: that observable was already captured live and identically twice (Session C, Part 5), and the payload was the only missing input. Staging a controller meant a hub customer, a pairing, a golden fetch and a claim — the work that consumed the three prior sessions. §8 correlates the payload to those existing measurements instead, and labels that as the inference it is.
- No new drill box was built — the brief's stated preference. The existing DooPlex nested-PVE fixture was reused and restored.
- R-117 and R-118 were filed, not investigated. R-117 in particular deserves its own run; on the evidence here it is more customer-affecting than R-116.
- v0.115.0 was neither reverted nor extended, per the fence.
- Two boot attempts of the drill VM failed before the run started and are recorded because the cause
is a reusable trap: adding
-deviceentries at launch shifts PCI slot assignment, which renames the guest NIC (ens3→ gone), so vmbr0 loses its port and SSH is unreachable while the console shows a healthy PVE login.bootindexfixes the boot-order half but not the NIC half. The fix is to boot the memory's known-good command line unchanged and hot-plug the extra disk over QMP afterwards — which is also the more faithful device-loss mechanism.