docs(R-117): REPORT.md — v0.117.0 shipped and proven live on demo-hp

Predicate validated on hardware for both dead states (stale-device and
filesystem-aborted), 340-497us per call, no block I/O proven by strace (only
/proc/self/mountinfo, 0 statfs). No regression through the real pipeline: the
live backup-target drive reads bound_under_parent=True via GET /disks with the
controller's own credential.

Records what was NOT covered: the stale-bind repair on hardware (StablePathForRaw
hardcodes the live parent, so it would write into guest 9201's namespace -
R-117h), and sustained-load behaviour, still unmeasured.
This commit is contained in:
2026-07-30 12:42:09 +02:00
parent 966d8f41ff
commit 1913e12031
+79 -119
View File
@@ -1,143 +1,103 @@
# REPORT — v0.116.0: R-116, the flag and the key share a row (2026-07-30) # REPORT — v0.117.0: R-117, the liveness signal now tests liveness (2026-07-30)
**Overwritten** per the standing rule. Agent-only fix: the controller was **not** touched (one fix, one **Overwritten** per the standing rule. Full evidence:
place) and `v0.115.0` was **not** reverted. Mechanism evidence: `felhom.eu` `felhom.eu/documentation/audits/R117-v0117-2026-07-30.md`. Spike this builds on:
`audits/DIAG-r116-disks-payload-2026-07-30.md`. `audits/SPIKE-r117-bind-liveness-2026-07-30.md`.
## Baselines as actually running (not as on `main`) | | |
|---|---|
| agent | 0.116.0 → **0.117.0** @ `966d8f4` |
| published + vouched | **0.117.0**, sha `e6768a12eba1c9e0…` (anonymous GET verified, manifest read back) |
| deployed | **demo-hp 0.117.0** (from **0.113.0** — see R-121) · felhom-pve still **0.115.0**, out of scope |
| controller / golden / hub | 0.186.0 / 0.186.0 / 0.82.0 — **untouched** |
| tests | **849 → 863**, 29/29 packages, `build`/`vet`/`test` each `rc=0` read separately |
| | `main` | running | a FRESH box gets (hub Day-0 manifest) | ## The defect
|---|---|---|---|
| agent | 0.115.0 → **0.116.0** | felhom-pve **0.115.0**, demo-hp **0.113.0** | **0.115.0** |
| controller | 0.186.0 | felhom-pve **0.186.0**, demo-hp **0.185.1** | golden bakes **0.185.1** |
| hub | 0.81.0 | 0.81.0 | `min_agent` 0.113.0, `min_controller` 0.156.0 |
| `felhom.eu` | `1aa1bd1` | — | — |
**The fleet is not uniform on either component.** The golden baking controller **0.185.1** while `BoundUnderParent` reported a namespace returning `EIO` on every read and write as healthy, and the gate
demo-felhom runs **0.186.0** is load-bearing for the ruling below: a controller-side fix would not reach restarted the customer's apps onto it. Both existing terms — `GuestSeesMount` and `isHostMountpoint`
a fresh box without a re-bake, whereas the agent channel already serves the newest published version. parse a mountinfo line and then test only `fields[4]`, the mount POINT. **Field 3, `major:minor`, sat in
the same parsed slice and was discarded.**
## The ruling: neither (a) nor (b) — a hybrid, because both regress R-114 ## The fix — one predicate, both halves mandatory
`primaryTargetID` is a PVE storage name (`felhom-backup`); the registry row's `Name` is the drive UUID. `bindLiveness` (`internal/localapi/intermediary.go`) becomes the third term of the conjunction at both
With the device gone the two records of one drive share **no runtime field** — no mount, no backing `/disks` construction sites. `/proc` only, **no block I/O**:
device, and the Observe row's `DurableID` has degraded off the fs-UUID to `path:/mnt/cel`.
**Rejected (a) — back-fill `MountPath` on the Observe row.** The spec's lean, and it does re-activate - **the bind must name the same device as the raw mount** — sound, not heuristic: a stale bind pins the
v0.115.0. Rejected on source evidence: `backup_target_offer.go:79` reads dead superblock, which keeps the old device index allocated, which forces the returning device onto a
`d.BackupTarget && d.MountPath != ""` as *"a real drive with its own mountpoint — healthy"* and returns different number;
**before** its `TargetAbsent` branch. Option (a) manufactures exactly that row while the drive is - **the filesystem must not have aborted** — `shutdown` or `emergency_ro`, both measured.
missing, so it would have silently flipped **R-114** — shipped 2026-07-29, proven live — back to a false
*healthy* banner. Also `MountPath`'s own contract is "the RAW host PVE mount" (`disks.go:152-153`); a path
that is not mounted is not that, so the downstream reading is correct and the change would be the bug.
**Rejected (b) — carry `BackupTarget` on the registry row.** A join **does** exist (below), so this was **The second is not optional.** A device that fails *without disappearing* gives the identical
available — but the registry row's `MountPath` is non-empty, read from the by-then-stale `.mount` unit all-signals-healthy state with the **devnos equal** and the drive never `Disconnected`, so the gate emits
file. It trips `:79` identically. Same regression, same reason. nothing at all, indefinitely (R-117a). A devno-only fix passes every payload test; red-proof RP3 exists for
exactly that.
**Rejected (c) — resolve controller-side.** The mapping is not there to find, and the *one state, one **Three states, never a bool:** `{Unknown, Live, StaleDevice, Aborted}`, `Unknown` the zero value, and
owner* ruling stands: the agent owns drive identity. `Usable()` is the single place where unknown counts as **present** — absent stops a working customer's apps.
**Taken (d) — give the Observe row the guest path (and NO MountPath), then dedup the registry row on ## The ruling (task §2.2) — and it is flagged for overrule
guest path.** The gate keys on **GuestPath**; R-114 keys on **MountPath**. Separating them satisfies both:
the surviving row carries the flag and the guest path with `mount_path: ""`, so `isTarget[guestPath]` is
`true` **and** `:79` does not match, leaving R-114's `TargetAbsent` branch reachable.
**The join:** with the device gone the shared identity is **configuration**, not state — the storage's `StaleDevice`**re-bind** (the raw mount is a healthy new superblock; repairs live, guest never restarts).
`path` from `storage.cfg` (Observe side) and the `.mount` unit's `Where` (registry side) are the same `Aborted`**quiet no-op, and surface via `BoundUnderParent=false`** — a re-bind lands on the *same* dead
path, so both derive the same stable guest path. That is the key both sides can still compute. superblock and this runs every 20 s, so repairing would be an infinite silent retry that also masks the
state. **No operator decision taken quietly:** it does change customer-visible behaviour (apps stop and an
alarm fires where today nothing happens), but it routes an already-broken state into the **existing** gate,
event types and Hungarian copy — no new concept reaches the customer — and the alternative is applications
writing documents into a filesystem that rejects every write.
**What (d)'s dedup removes in the absent state**, enumerated: the registry row's `uuid:` `durable_id`, its **No new recovery path.** `AttachDrive`'s normalize leg already did the repair and three call sites already
hardcoded and **false** `state: "attached"`, its `role: user-data`, and its `total_bytes`/`used_bytes` invoked it (20 s ticker, agent startup, and the controller's `Return` branch **before** `restartStacks`).
copied from the **root** filesystem. No consumer found — wipe/decommission require the drive present; All three died on `if n == 1 && GuestSeesMount(...)` logging *"fully live, no-op"* about a dead namespace.
`firstOfferableDrive` (`:138`) already requires `MountPath != ""` **and** `Role == "user-data"`; `diskKey`
(`disk_health.go:130`) changes key transiently but the health check excludes UNKNOWN verdicts, so no false
degradation. **R-118's symptom disappears incidentally — R-118 is NOT fixed and stays open.**
**Deliberately unchanged:** `Role` (the absent target still reads `system`), the `BoundUnderParent` **Ordering trap, caught by a test:** reading the abort flag before comparing devices classifies the real
conjunction, and every wire field. return state as aborted (its stale bind carries `shutdown` too) and refuses the repair **while still
reporting correctly** — so the abort flag is read off the RAW mount in the stale case.
## Files ## Live validation, demo-hp
`internal/hub/report.go` (`ConfigPath`, **`json:"-"`** — that struct is a cross-repo contract pinned by | State | raw | bind | verdict | usable |
`testdata/host-report.golden.json` + `contract_test.go`'s key-set comparison, so a wire-visible field |---|---|---|---|---|
would have needed a matching `felhom.eu/hub` change; nothing off-box needs the value) · | healthy | `8:16` | `8:16` | `live` | true |
`internal/storage/observe.go` (set from `s.Path`) · `internal/localapi/disks.go` (ConfigPath fallback + | **R-117 return** | `8:32` `/dev/sdc` | `8:16 …,shutdown` | **`stale-device`** | **false** |
guest-path dedup) · `internal/localapi/disks_backup_target_row_test.go` (fixture corrected + 4 tests) · | **Q7 in-place abort** | `252:11 …,emergency_ro` | `252:11 …,emergency_ro` | **`filesystem-aborted`** | **false** |
`CHANGELOG.md`.
`observe.go`'s change is **not** the fallthrough its `:176-183` comment forbids: that prohibition is about 340497 µs per call. **No block I/O proven by `strace`**: only `/proc/self/mountinfo`, **0** `statfs`.
resolving a *device or UUID* from the **containing** filesystem, which would hand back root's identity and **No regression through the real pipeline:** `GET /disks` with the controller's own credential shows the
mis-target a DR re-attach. `s.Path` is the storage's own declaration, identifies nothing but itself, and is live backup-target drive `bound_under_parent=True`, with 32 gate lines in 3 min as the positive observable
used for no resolution. `MountPath` stays empty, which is the truth. and zero spurious transitions.
## §2 reconfirmation ## Tests + red-proofs
| Cited | Line now | Holds | 14 new tests; the payload ones drive the real `/disks` handler and the `AttachDrive` ones record the actual
|---|---|---| privileged calls. Fixtures are the spike's captured mountinfo, injected as **data** via
| `mountPath` assigned only on exactMount success | `observe.go:184-190` | **HOLDS** verbatim | `procSelfMountinfo`/`procGuestMountinfo`, so the real parser, predicate, `AttachDrive` and handler all run.
| empty `backingDevice``RoleSystem` | `role.go:180-181` | **HOLDS** verbatim |
| guest-path guard skipped for non-user-data | `disks.go:215-217` | **HOLDS** |
| `Known()` reads unit FILES, never the mount table | `registry_known.go:40-75`, `MountPath: spec.Where` at `:70` | **HOLDS** |
| `driveTargetByPath` keys only from non-empty fields | `intermediary.go:602-618` (`:611`, `:614`) | **HOLDS** |
| the dedup keys on `d.MountPath` at `:263` | `:263` is `if t.BackingDevice != ""`. Actual: `seen` built `:290-295`, dedup `:298` | **HOLDS; the line number in the spec is wrong** |
| "`seen[""]` is set, so the union row is not skipped" | **DOES NOT HOLD as stated**`seen[d.MountPath] = true` is guarded by `if d.MountPath != ""`, so `seen` never gains `""`. The union row survives because `/mnt/<name>` is *absent* from `seen`, not because `""` is present. Same outcome, different reason | corrected |
## Tests — 845 → 849, suite green (`go build ./... && go vet ./... && go test ./...` rc=0) **6 red-proofs, each verified to have landed.** RP1/RP2 (term 3 removed per site), RP3 (**the P1-only
fix**), RP4 (`emergency_ro` dropped), RP5 (unknown reported dead), RP6 (aborted arm falls through to a
re-bind).
All PASS: `TestAbsentBackupTargetIsResolvableByGuestPath` (rewritten — real shape **and** the union loop), **A hollow test was caught, by RP1 failing to fail.** The aborted fixture first used a `/dev/mapper`
`TestAbsentTargetAppearsExactlyOnce` (new, the join), `TestAbsentTargetKeepsR114DegradedSignal` (new, the device, for which `RoleForStorage` derives `role="system"` — a system row never enters the block that
regression guard), `TestPresentTargetPayloadUnchanged` (new, fleet-wide parity), computes `BoundUnderParent`, so the field stayed false **by default**, the assertion passed, and no mutation
`TestAbsentNonTargetKeepsItsRegistryRowAndNoFlag` (new, negative + over-suppression), could ever fail it. Fixtures now assert the production row shape before asserting the field.
`TestAbsentTargetRowDoesNotRegisterPresence`, `TestCaseBLocalTargetGetsNoGuestPath`,
`TestSystemBackedTargetUnderMntGetsNoGuestPath`, `TestNonTargetDriveNeverCarriesTheFlag`.
**Red-proofs — each mutation asserted to have landed before the run:** ## Not yet done / not covered
1. v0.115.0's MountPath-only fallback → `isTarget[…] = FALSE … rows=2`. **The empirical proof v0.115.0 - **The stale-bind repair on hardware.** `StablePathForRaw` hardcodes the production
was inert**, rather than an argument for it. `/mnt/felhom-drives`, so exercising `AttachDrive` for a scratch drive writes a submount into guest
2. Drop the guest-path dedup → `the absent drive is carried by 2 rows, want exactly 1`. 9201's namespace, which the task fences. It happened once inadvertently, was reverted within a minute
3. Give the absent row a `MountPath` (what option (a) produces) → the R-114 guard fires. and 9201 verified unaffected, and was deliberately not repeated. Filed **R-117h**.
4. Over-broad dedup (`gp != ""` without `seenGuest`) → the non-target drive loses its own row, and two - **Sustained-load behaviour of the predicate** — flagged unmeasured by the spike and **still unmeasured**.
pre-existing R-113 union tests fail as well. - **felhom-pve not upgraded** (0.115.0). The fleet is non-uniform.
- **R-117g** — an aborted filesystem is never cleared automatically, by design, so it alarms until a human
acts; there is no guided recovery.
- **R-121** — a box's installed agent can sit releases behind the vouched one and nothing notices; the
R-120 gate is golden-vs-fleet-controller and cannot see it.
**The corrected fixture is the headline test change.** `theAbsentTarget` supplied ## Observations (not acted on)
`MountPath: "/mnt/mentes"`, which production never supplies, and `targetRowServer` left `DriveTargets`
**nil**, so the union loop never ran and the two-row shape was invisible to the entire suite. Both fixed —
that combination is why an inert fix shipped green.
## Process note Guest 9201 has 2 stacked entries at `/mnt/felhom-drives/nvme-1tb` while the host has 1 (pre-existing;
`AttachDrive` normalizes the host count only). 14 stale `felhom-agent.bak-*` binaries on demo-hp (~200 MB),
One red-proof was restored with `git checkout <file>`, which discarded the whole implementation rather nothing prunes them. 11 files in this repo fail `gofmt -l`, all pre-existing and none in `internal/localapi/`.
than the one-line mutation. Re-applied and re-verified; later red-proofs restored from pristine copies
kept outside the repo. `git checkout` is not an undo for a mutation on a work-in-progress file.
## LIVE VALIDATION — all claims PASS (`felhom.eu` `audits/R116-v0116-2026-07-30.md`)
Fresh box, real day-0 from the v1.25.0 ISO on a nested PVE on **demo-hp** (per `runbooks/target-selection.md`),
agent **installed unaided from the vouched Day-0 manifest**, both drives enrolled through the real endpoints,
device loss a real hot-detach.
```
07:20:04 backup_target_absent (error) Cel meghajto ← TARGET, specific (was generic through v0.115.0)
07:22:34 backup_target_restored (info) Cel meghajto ← its matching pair
07:24:04 storage_disconnected (error) Adat meghajto ← NON-target, generic, same box 4 min later
07:25:34 storage_reconnected (info) Adat meghajto
```
All four at the hub; gate fired in 3 s. **Discrimination proven non-trivially for the first time.**
Over-correction passes on a positive observable (0 ABSENT lines / 0 drive events over 2m14s with both drives
present, target `degraded:false`, while two `RETURNED` lines prove the gate ticked). Publish observables:
registry GET of the published bytes matched `b47c5c4d…`, the manifest was read back showing 0.116.0, and the
box reported it running.
**One caveat, not a regression:** the drill's controller was **0.185.1** from the golden, which predates
R-114, so its absent-state banner showed the old false "the backup is on the system disk" copy. The R-114
guard is pinned by unit test and by the payload (no row combines the flag with a mount path) but could not be
confirmed on that box. Filed as **R-120** — the golden is a release behind, which is R-115's class one layer up.
## Teardown — all three §13 layers
VM 9401 destroyed `--purge`; `r116-images` removed and the space measured back (`felhom-backup` 928787076 KiB
available vs 928787080 before); hub records **gate-blocked** — the cascade was attempted and correctly refused
`HTTP 409 host … is ONLINE`, with the command recorded for once it ages to DOWN. Secrets shredded.