diff --git a/REPORT-freboot-fleak-fobs.md b/REPORT-freboot-fleak-fobs.md new file mode 100644 index 0000000..6e9918c --- /dev/null +++ b/REPORT-freboot-fleak-fobs.md @@ -0,0 +1,132 @@ +# REPORT — F-REBOOT + F-LEAK + F-OBS, and two investigations (2026-07-28) + +Scope in this repo: **`scripts/felhom-host-install.sh` v1.20.0 → v1.21.0** (which is where F-LEAK's +*actual* fix lives), plus the Campaign 8 audit doc and `OPEN-ITEMS.md`. Written as +`REPORT-freboot-fleak-fobs.md` so the shared `REPORT.md` is not clobbered. + +Code companions: `felhom-agent` v0.106.0 → **v0.109.0**, `felhom-controller` v0.179.0 → **v0.180.0**. + +## Baselines (reconfirmed, not copied) +`felhom.eu d0cec9d`, `felhom-agent af1c21a`, `felhom-controller fb91c8d`, all clean. Agent `0.106.0` +and controller `0.179.0` live on both demo boxes. + +--- + +## host-install v1.21.0 — F-LEAK, and why the fix landed *here* rather than in the agent + +**The finding.** A restore-test whose restore **fails** leaves a scratch guest the agent cannot destroy +(`403 missing privilege VM.Allocate`), so a half-restored guest holds its disks until a human removes +it and the 10-slot scratch band shrinks silently. + +**The cause is structural, not a missing privilege in the role.** `FelhomAgentGuest` is granted at +`/pool/felhom`, and **a guest joins that pool only when its restore completes**. A failed restore +therefore produces a guest that exists, is in no pool, and is out of the token's reach entirely. + +**The first fix was wrong, and its own live replay is what proved it.** Agent v0.107.0 shipped a +teardown fallback that adopted the stranded guest into the pool and retried — reasoning from +`Pool.Allocate` on `/pool/felhom`. It fired exactly as designed and PVE refused it: + +``` +ERROR restore-test: pool adoption failed; left for Recover vmid=990000 + err="proxmox: PUT /pools/felhom -> HTTP 500: permission denied at /vms/990000 (missing privilege ...)" +``` + +`PUT /pools/{pool}` **also** requires `VM.Allocate` on the VM being added. **Pool membership cannot +bootstrap its own authority.** Removed in agent v0.108.0 rather than left in place — a path that +provably cannot work is worse than none, because it reads as a fix. + +**What shipped instead.** `apply_scoped_acl` now grants `FelhomAgentGuest` at each `/vms/` in +`PVE_SCRATCH_VMID_MIN..PVE_SCRATCH_VMID_MAX` (990000–990009 — the band the restore-test already picks +from), to **both** the user and the token, because the privsep-intersection rule applies here as +everywhere. + +Two supporting changes, both load-bearing rather than tidy-up: +- **`remove_scoped_acl` deletes the band grants before the role delete.** PVE refuses to delete a role + still referenced by any ACL, so omitting this would have broken the uninstall — a failure that would + only surface on a decommission. +- **`step_verify` asserts the band grants.** A missing one is otherwise invisible until a restore-test + *fails*, which is precisely the case that leaked a guest in the first place. + +### Why this is not a widening — proven live, at the seam the defect lives in + +A real PBS restore to `990000` **without `--pool`** reproduced the exact stranded state +(`990000 stopped`; `felhom pool members: [9201]`; `990000 in pool: False`). Then, with the agent's own +token, same guest, minutes apart: + +| | `DELETE /nodes//lxc/990000` | +|---|---| +| **grant removed** (the original defect) | `403 Permission check failed (/vms/990000, VM.Allocate)` — guest still present | +| **grant restored** (the fix) | `200 UPID:...:vzdestroy:990000:felhom-agent@pve!agent` — guest gone | + +And it still cannot reach anything else: + +| target | result | +|---|---| +| `/vms/990010` (one past the band) | **403** `Permission check failed (/vms/990010, VM.Allocate)` | +| `/vms/100` | **403** same | + +`990010` does not exist and PVE **still** answered 403 rather than "does not exist" — so PVE evaluates +**permission before existence**, which makes these genuine authorization refusals rather than artifacts +of a missing guest. Granting at `/vms` was considered and rejected: it would authorise destroying every +guest on the box, including a co-tenant's. + +Applied on **both** demo boxes (demo-hp and demo-felhom) so the live fleet matches the installer. + +**A careless step of mine, recorded rather than buried.** The probe loop also issued a live `DELETE` +against running guest **9201**. It was refused with `500 container is running` — but the *permission +check passed* (9201 is a pool member by design), so had the guest been stopped I would have destroyed +the live demo guest. The scratch-band probes were the safe ones; 9201 had no business in that list. + +--- + +## Documentation changes + +- **`documentation/audits/CAMPAIGN-8-backup-restore-2026-07-27.md`** — F-REBOOT, F-LEAK and F-OBS + written up as FIXED with their live evidence, including F-LEAK's refuted first attempt (recorded + precisely *because* it looked right), plus a new **§6b** for the follow-up investigation. +- **`documentation/backlog/OPEN-ITEMS.md`** — three findings closed, **R-100** filed. + +--- + +## R-100 — the investigation's finding, deliberately NOT fixed + +**A restic offsite tier that fails every night never goes stale on the hub.** This is **F-CRIT-2's +defect class one layer up and on the other tier** — a *failed* run resetting the freshness clock — and +it was found by asking whether F-CRIT-2's shape existed anywhere else. + +Both halves verified in the source, not inferred: +- **Controller:** `o.LastRun = time.Now()` is set **unconditionally** at + `controller/internal/backup/offbox.go:716`, *outside* the `runErr` branch. The failure is recorded + faithfully, but into a different field — `o.LastStatus = "error"` at `:725`. +- **Hub:** `isStale()` reads **only** `off.LastRun` (`hub/internal/monitor/offsite.go:120`, `:127`, + `:131`) and never consults `LastStatus`. + +So a nightly restic run that fails every night keeps `LastRun` fresh, `isStale` is permanently false, +and the staleness alarm never fires — with no successful offsite backup having occurred at all. + +**Scope of the silence, stated precisely rather than dramatically.** `LastStatus` *does* reach the hub — +it is parsed into the report struct and **only logged** (`offsite.go:270`); it drives no checker and no +notification. The controller's own guest UI surfaces `LastStatus="error"`, so the failure is visible to +someone who looks. What is missing is the **push**: the operator's fleet-wide alarm plane is silent, +which is the plane that matters for an unattended appliance. + +Not fixed, per this task's investigation-only scope. Fix direction: gate staleness on the last +*successful* run rather than the last attempt — exactly what F-CRIT-2's `NewestArchiveTime` fix did for +the PBS tier. + +--- + +## Verification + +`bash -n scripts/felhom-host-install.sh` clean. The ACL behaviour was verified **live on real hardware** +rather than by dry-run, since the whole finding is about what PVE's authorizer actually does — and the +live run is what refuted my first design. + +## Fleet state +Agent **0.109.0** and controller **0.180.0** on both demo boxes, all healthy. Scratch-band ACLs applied +on both. No leftover scratch guests. demo-hp's `restore_test_cadence_seconds` reverted **600 → 302400** +(a bounded change made for the replay). + +`felhom.eu`: this repo has a **foreign uncommitted WIP file** (`documentation/PROMPT-TEMPLATE.md`) from +another session in the shared worktree. Left untouched; my commits staged explicit paths only, per the +never-`git add -A` rule.