8.0 KiB
REPORT — F-REBOOT + F-LEAK + F-OBS, and two investigations (2026-07-28)
Scope in this repo: scripts/felhom-host-install.sh v1.20.0 → v1.21.0 (which is where F-LEAK's
actual fix lives), plus the Campaign 8 audit doc and OPEN-ITEMS.md. Written as
REPORT-freboot-fleak-fobs.md so the shared REPORT.md is not clobbered.
Code companions: felhom-agent v0.106.0 → v0.110.0, felhom-controller v0.179.0 → v0.180.0.
Correction to this repo's part of the story: v1.21.0's band-scoped ACL is not the final F-LEAK
fix. It works, but only once per slot — PVE's destroy path calls
AccessControl::remove_vm_access($vmid) (API2/LXC.pm:906), which deletes every ACL at /vms/<vmid>
(AccessControl.pm:1898), so the grant is consumed by the operation it authorises. Found by counting
ACL rows after the first successful teardown (/vms/990000 → 0 grants), not by reasoning about it. The
durable fix is agent v0.110.0's band-scoped fenced destroy; v1.21.0 remains valuable because it makes
the common case need no privileged call, and it is now the first of two layers rather than the only one.
Baselines (reconfirmed, not copied)
felhom.eu d0cec9d, felhom-agent af1c21a, felhom-controller fb91c8d, all clean. Agent 0.106.0
and controller 0.179.0 live on both demo boxes.
host-install v1.21.0 — F-LEAK, and why the fix landed here rather than in the agent
The finding. A restore-test whose restore fails leaves a scratch guest the agent cannot destroy
(403 missing privilege VM.Allocate), so a half-restored guest holds its disks until a human removes
it and the 10-slot scratch band shrinks silently.
The cause is structural, not a missing privilege in the role. FelhomAgentGuest is granted at
/pool/felhom, and a guest joins that pool only when its restore completes. A failed restore
therefore produces a guest that exists, is in no pool, and is out of the token's reach entirely.
The first fix was wrong, and its own live replay is what proved it. Agent v0.107.0 shipped a
teardown fallback that adopted the stranded guest into the pool and retried — reasoning from
Pool.Allocate on /pool/felhom. It fired exactly as designed and PVE refused it:
ERROR restore-test: pool adoption failed; left for Recover vmid=990000
err="proxmox: PUT /pools/felhom -> HTTP 500: permission denied at /vms/990000 (missing privilege ...)"
PUT /pools/{pool} also requires VM.Allocate on the VM being added. Pool membership cannot
bootstrap its own authority. Removed in agent v0.108.0 rather than left in place — a path that
provably cannot work is worse than none, because it reads as a fix.
What shipped instead. apply_scoped_acl now grants FelhomAgentGuest at each /vms/<id> in
PVE_SCRATCH_VMID_MIN..PVE_SCRATCH_VMID_MAX (990000–990009 — the band the restore-test already picks
from), to both the user and the token, because the privsep-intersection rule applies here as
everywhere.
Two supporting changes, both load-bearing rather than tidy-up:
remove_scoped_acldeletes the band grants before the role delete. PVE refuses to delete a role still referenced by any ACL, so omitting this would have broken the uninstall — a failure that would only surface on a decommission.step_verifyasserts the band grants. A missing one is otherwise invisible until a restore-test fails, which is precisely the case that leaked a guest in the first place.
Why the grant is still not a widening — proven live, at the seam the defect lives in
A real PBS restore to 990000 without --pool reproduced the exact stranded state
(990000 stopped; felhom pool members: [9201]; 990000 in pool: False). Then, with the agent's own
token, same guest, minutes apart:
DELETE /nodes/<node>/lxc/990000 |
|
|---|---|
| grant removed (the original defect) | 403 Permission check failed (/vms/990000, VM.Allocate) — guest still present |
| grant restored (the fix) | 200 UPID:...:vzdestroy:990000:felhom-agent@pve!agent — guest gone |
And it still cannot reach anything else:
| target | result |
|---|---|
/vms/990010 (one past the band) |
403 Permission check failed (/vms/990010, VM.Allocate) |
/vms/100 |
403 same |
990010 does not exist and PVE still answered 403 rather than "does not exist" — so PVE evaluates
permission before existence, which makes these genuine authorization refusals rather than artifacts
of a missing guest. Granting at /vms was considered and rejected: it would authorise destroying every
guest on the box, including a co-tenant's.
Applied on both demo boxes (demo-hp and demo-felhom) so the live fleet matches the installer.
A careless step of mine, recorded rather than buried. The probe loop also issued a live DELETE
against running guest 9201. It was refused with 500 container is running — but the permission
check passed (9201 is a pool member by design), so had the guest been stopped I would have destroyed
the live demo guest. The scratch-band probes were the safe ones; 9201 had no business in that list.
Documentation changes
documentation/audits/CAMPAIGN-8-backup-restore-2026-07-27.md— F-REBOOT, F-LEAK and F-OBS written up as FIXED with their live evidence, including F-LEAK's refuted first attempt (recorded precisely because it looked right), plus a new §6b for the follow-up investigation.documentation/backlog/OPEN-ITEMS.md— three findings closed, R-100 filed.
R-100 — the investigation's finding, deliberately NOT fixed
A restic offsite tier that fails every night never goes stale on the hub. This is F-CRIT-2's defect class one layer up and on the other tier — a failed run resetting the freshness clock — and it was found by asking whether F-CRIT-2's shape existed anywhere else.
Both halves verified in the source, not inferred:
- Controller:
o.LastRun = time.Now()is set unconditionally atcontroller/internal/backup/offbox.go:716, outside therunErrbranch. The failure is recorded faithfully, but into a different field —o.LastStatus = "error"at:725. - Hub:
isStale()reads onlyoff.LastRun(hub/internal/monitor/offsite.go:120,:127,:131) and never consultsLastStatus.
So a nightly restic run that fails every night keeps LastRun fresh, isStale is permanently false,
and the staleness alarm never fires — with no successful offsite backup having occurred at all.
Scope of the silence, stated precisely rather than dramatically. LastStatus does reach the hub —
it is parsed into the report struct and only logged (offsite.go:270); it drives no checker and no
notification. The controller's own guest UI surfaces LastStatus="error", so the failure is visible to
someone who looks. What is missing is the push: the operator's fleet-wide alarm plane is silent,
which is the plane that matters for an unattended appliance.
Not fixed, per this task's investigation-only scope. Fix direction: gate staleness on the last
successful run rather than the last attempt — exactly what F-CRIT-2's NewestArchiveTime fix did for
the PBS tier.
Verification
bash -n scripts/felhom-host-install.sh clean. The ACL behaviour was verified live on real hardware
rather than by dry-run, since the whole finding is about what PVE's authorizer actually does — and the
live run is what refuted my first design.
Fleet state
Agent 0.110.0 (with the updated sudoers) and controller 0.180.0 on both demo boxes, all healthy.
Scratch-band ACLs at 20 rows on both — re-applied after the attempt-2 destroy consumed one. No leftover scratch guests. demo-hp's restore_test_cadence_seconds reverted 600 → 302400
(a bounded change made for the replay).
felhom.eu: this repo has a foreign uncommitted WIP file (documentation/PROMPT-TEMPLATE.md) from
another session in the shared worktree. Left untouched; my commits staged explicit paths only, per the
never-git add -A rule.