Files
felhom.eu/documentation/audits/R116-PART5-2026-07-30.md
T
admin c3ce4c7b20 R-116 Part 5 FAILED: the fix shipped, C5 still fails, mechanism NOT isolated
A fresh box running the fully shipped stack -- agent 0.115.0 from the Day-0
manifest plus controller 0.185.1 from the vouched golden, no hand-deploy -- still
fired the GENERIC storage_disconnected on detach and the SPECIFIC
backup_target_restored on return. backup_target_absent count 0. Identical to
Session C. The v0.115.0 fix changed nothing observable.

Part 4's three positive observables were all obtained before the run (registry
newest 0.115.0, hub vouches 0.115.0, felhom-pve running 0.115.0 clean), so the
publish step forgotten twice was not forgotten a third time, and the box
demonstrably installed the fix under test.

Discrimination FAILS: the target itself produced the generic event, so the two
cannot be told apart regardless of the non-target leg -- which was therefore not
staged. Reported as a fail, not as Session C's trivial pass.

Over-correction guard PASSES: 0 ABSENT lines with the drive present, target
degraded:false.

THE HONEST PART. The fix targets a shape that does not occur live, and which
shape does occur is NOT ISOLATED. With the drive detached PVE reports the
storage inactive with zeroed fields -- a shape the unit fixture did not model.
Three attempts to read the real /disks payload failed on token extraction across
the ssh -> guest -> container layers, and a present-drive CONTROL query also
returned 0 rows, proving the query was broken rather than the payload. Without
that control this run would have recorded a third false mechanism, after "the
union row writes false" (wrong, corrected yesterday) and "no row carries the
guest path" (unverified). The leading hypothesis -- an inactive storage reaching
Observe with an empty MountPath, so StablePathForRaw returns "" -- is consistent
with the pvesm output but is NOT evidence and is recorded as such.

Next session's first job is a working /disks read, with a present-drive control
run FIRST, before any further code.

agent v0.115.0 is published, vouched and INERT. Not reverted: reverting is
itself a change, the runbook forbids fixing mid-run, and the code is tested and
harmless.

Capability-map row F stays PARTIAL, now citing the re-test.
Teardown clean: pvesm status after == before (local-lvm 38.83%), guest 9201 and
drill-r50 untouched. Customer sess-d pending the usual ONLINE-ages-to-DOWN gate.
2026-07-30 07:15:31 +02:00

5.4 KiB
Raw Blame History

R116-PART5-2026-07-30 — the fix shipped, C5 still fails, and the mechanism is NOT isolated

Run: RUNBOOK R-116 Part 5, executed by CC on DooPlex, 2026-07-30. Result: FAILED. Agent 0.115.0 — the R-116 fix, installed from the Day-0 manifest — did not change the outcome. The absent-drive alarm is still the generic one.

Claim Verdict
C5backup_target_absent on detach, backup_target_restored on return, paired FAIL — detach still fires storage_disconnected; return fires the specific event. Same asymmetry as Session C
Discrimination — target ⇒ specific, non-target ⇒ generic FAIL — the target produced the generic event, so the two are still indistinguishable. Not achieved
No over-correction PASS — 0 ABSENT lines with the drive present; target degraded:false

1. What ran — the fully shipped stack

Artifact Version How
agent 0.115.0 from the Day-0 manifestArtifact manifest served for customer sess-d (agent=0.115.0 golden=0.185.1); host row sess-d-0452c4 … 0.115.0 ONLINE
controller 0.185.1 from the vouched golden — no hand-deploy this run, deliberately, so the whole stack is the shipped article

Part 4's three positive observables were all obtained before the run: registry newest 0.115.0, hub vouches 0.115.0, felhom-pve reports felhom-agent 0.115.0 active with 0 errors. The publish step that was forgotten twice was not forgotten.

Fence reconfirmed: local-lvm 38.83 % before and after. Guest 9201 and drill-r50 untouched.

2. The observation

05:07:12  [gate] drive ABSENT /mnt/felhom-drives/cel — stopped+blocked 0 app(s)
05:07:15  Event pushed: storage_disconnected (error) — Meghajtó váratlanul leválasztva: Cel meghajto
          backup_target_absent : 0     storage_disconnected : 1

05:11:43  [gate] drive RETURNED /mnt/felhom-drives/cel
05:11:43  Event pushed: backup_target_restored (info)

Identical to Session C. The gate fires correctly (R-113 holds, 0 s3 s), SetDisconnected is reached, and the wrong event goes out.

3. Why the fix did not fire — NOT ISOLATED, and that is the honest answer

The v0.115.0 change sets di.GuestPath on the Observe row when di.GuestPath == "" && di.BackupTarget && t.BackingDevice == "". It is unit-tested and red-proofed against a fixture built from the Session-C shape. On the live box it changed nothing.

What was established:

  • With the drive detached, pvesm status reports the storage as felhom-backup dir inactive 0 0 0, and PVE emits unable to activate storage 'felhom-backup' - directory is expected to be a mount point but is not mounted: '/mnt/cel'. So the storage is listed but inactive with zeroed fields — a shape the unit fixture did not model (it supplied a MountPath and a disconnected state).

What was NOT established, and must not be guessed at again: the actual /disks payload. Three attempts to read it failed on token extraction across the ssh → guest → container layers. A control query with the drive PRESENT also returned 0 rows, which proves the query itself was broken, not the payload. Without that control this run would have recorded a third false mechanism — after "the union row writes false" (wrong, corrected) and "no row carries the guest path" (unverified).

So the honest state is: the fix addresses a shape that does not occur on the live box, and which shape does occur is unknown. The leading hypothesis — that an inactive storage reaches Observe with an empty MountPath, so StablePathForRaw("") returns "" and the fallback sets nothing — is consistent with the pvesm output but is not evidence.

The next session's first job is a working /disks read, with a present-drive control run first, before any further code. Everything after that is guesswork until it exists.

4. Also observed

  • The stale bind survives the return and still names a dead device node. After the drive came back as /dev/sdc, findmnt /mnt/felhom-drives/cel still read /dev/sdb[/felhom-data] while the state reported degraded:false. Third consecutive run showing this; still not investigated.
  • Discrimination could not be reached. Because the target itself produced the generic event, there was no point staging the non-target leg: the two cannot be told apart regardless of what it emits. Reported as a FAIL, not as Session C's trivial pass.

5. Teardown

VM 9302 destroyed --purge, p5-images removed, scratch dir removed. pvesm status after == before (local-lvm 38.83 %). Guest 9201 and drill-r50 untouched.

Customer sess-d + host sess-d-0452c4 remain — the delete gate refuses while the host still reads ONLINE, as in both prior runs. Deletable once it ages to DOWN (>1 h):

POST /configs/sess-d/delete  ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=sess-d expect_hosts=1

6. Standing of agent v0.115.0

It is published and vouched, and it is inert — it adds a guest path in a case that apparently does not arise, so it neither fixes nor breaks anything observable. It was NOT reverted: reverting is itself a change, the runbook forbids fixing mid-run, and the code is well-tested and harmless. Whether to keep, extend or revert it is the next session's decision, taken with a working /disks read in hand.