141 lines
8.0 KiB
Markdown
141 lines
8.0 KiB
Markdown
# REPORT — F-REBOOT + F-LEAK + F-OBS, and two investigations (2026-07-28)
|
||
|
||
Scope in this repo: **`scripts/felhom-host-install.sh` v1.20.0 → v1.21.0** (which is where F-LEAK's
|
||
*actual* fix lives), plus the Campaign 8 audit doc and `OPEN-ITEMS.md`. Written as
|
||
`REPORT-freboot-fleak-fobs.md` so the shared `REPORT.md` is not clobbered.
|
||
|
||
Code companions: `felhom-agent` v0.106.0 → **v0.110.0**, `felhom-controller` v0.179.0 → **v0.180.0**.
|
||
|
||
**Correction to this repo's part of the story:** v1.21.0's band-scoped ACL is *not* the final F-LEAK
|
||
fix. It works, but only **once per slot** — PVE's destroy path calls
|
||
`AccessControl::remove_vm_access($vmid)` (`API2/LXC.pm:906`), which deletes every ACL at `/vms/<vmid>`
|
||
(`AccessControl.pm:1898`), so **the grant is consumed by the operation it authorises**. Found by counting
|
||
ACL rows after the first successful teardown (`/vms/990000` → 0 grants), not by reasoning about it. The
|
||
durable fix is agent **v0.110.0**'s band-scoped fenced destroy; v1.21.0 remains valuable because it makes
|
||
the common case need no privileged call, and it is now the *first* of two layers rather than the only one.
|
||
|
||
## Baselines (reconfirmed, not copied)
|
||
`felhom.eu d0cec9d`, `felhom-agent af1c21a`, `felhom-controller fb91c8d`, all clean. Agent `0.106.0`
|
||
and controller `0.179.0` live on both demo boxes.
|
||
|
||
---
|
||
|
||
## host-install v1.21.0 — F-LEAK, and why the fix landed *here* rather than in the agent
|
||
|
||
**The finding.** A restore-test whose restore **fails** leaves a scratch guest the agent cannot destroy
|
||
(`403 missing privilege VM.Allocate`), so a half-restored guest holds its disks until a human removes
|
||
it and the 10-slot scratch band shrinks silently.
|
||
|
||
**The cause is structural, not a missing privilege in the role.** `FelhomAgentGuest` is granted at
|
||
`/pool/felhom`, and **a guest joins that pool only when its restore completes**. A failed restore
|
||
therefore produces a guest that exists, is in no pool, and is out of the token's reach entirely.
|
||
|
||
**The first fix was wrong, and its own live replay is what proved it.** Agent v0.107.0 shipped a
|
||
teardown fallback that adopted the stranded guest into the pool and retried — reasoning from
|
||
`Pool.Allocate` on `/pool/felhom`. It fired exactly as designed and PVE refused it:
|
||
|
||
```
|
||
ERROR restore-test: pool adoption failed; left for Recover vmid=990000
|
||
err="proxmox: PUT /pools/felhom -> HTTP 500: permission denied at /vms/990000 (missing privilege ...)"
|
||
```
|
||
|
||
`PUT /pools/{pool}` **also** requires `VM.Allocate` on the VM being added. **Pool membership cannot
|
||
bootstrap its own authority.** Removed in agent v0.108.0 rather than left in place — a path that
|
||
provably cannot work is worse than none, because it reads as a fix.
|
||
|
||
**What shipped instead.** `apply_scoped_acl` now grants `FelhomAgentGuest` at each `/vms/<id>` in
|
||
`PVE_SCRATCH_VMID_MIN..PVE_SCRATCH_VMID_MAX` (990000–990009 — the band the restore-test already picks
|
||
from), to **both** the user and the token, because the privsep-intersection rule applies here as
|
||
everywhere.
|
||
|
||
Two supporting changes, both load-bearing rather than tidy-up:
|
||
- **`remove_scoped_acl` deletes the band grants before the role delete.** PVE refuses to delete a role
|
||
still referenced by any ACL, so omitting this would have broken the uninstall — a failure that would
|
||
only surface on a decommission.
|
||
- **`step_verify` asserts the band grants.** A missing one is otherwise invisible until a restore-test
|
||
*fails*, which is precisely the case that leaked a guest in the first place.
|
||
|
||
### Why the grant is still not a widening — proven live, at the seam the defect lives in
|
||
|
||
A real PBS restore to `990000` **without `--pool`** reproduced the exact stranded state
|
||
(`990000 stopped`; `felhom pool members: [9201]`; `990000 in pool: False`). Then, with the agent's own
|
||
token, same guest, minutes apart:
|
||
|
||
| | `DELETE /nodes/<node>/lxc/990000` |
|
||
|---|---|
|
||
| **grant removed** (the original defect) | `403 Permission check failed (/vms/990000, VM.Allocate)` — guest still present |
|
||
| **grant restored** (the fix) | `200 UPID:...:vzdestroy:990000:felhom-agent@pve!agent` — guest gone |
|
||
|
||
And it still cannot reach anything else:
|
||
|
||
| target | result |
|
||
|---|---|
|
||
| `/vms/990010` (one past the band) | **403** `Permission check failed (/vms/990010, VM.Allocate)` |
|
||
| `/vms/100` | **403** same |
|
||
|
||
`990010` does not exist and PVE **still** answered 403 rather than "does not exist" — so PVE evaluates
|
||
**permission before existence**, which makes these genuine authorization refusals rather than artifacts
|
||
of a missing guest. Granting at `/vms` was considered and rejected: it would authorise destroying every
|
||
guest on the box, including a co-tenant's.
|
||
|
||
Applied on **both** demo boxes (demo-hp and demo-felhom) so the live fleet matches the installer.
|
||
|
||
**A careless step of mine, recorded rather than buried.** The probe loop also issued a live `DELETE`
|
||
against running guest **9201**. It was refused with `500 container is running` — but the *permission
|
||
check passed* (9201 is a pool member by design), so had the guest been stopped I would have destroyed
|
||
the live demo guest. The scratch-band probes were the safe ones; 9201 had no business in that list.
|
||
|
||
---
|
||
|
||
## Documentation changes
|
||
|
||
- **`documentation/audits/CAMPAIGN-8-backup-restore-2026-07-27.md`** — F-REBOOT, F-LEAK and F-OBS
|
||
written up as FIXED with their live evidence, including F-LEAK's refuted first attempt (recorded
|
||
precisely *because* it looked right), plus a new **§6b** for the follow-up investigation.
|
||
- **`documentation/backlog/OPEN-ITEMS.md`** — three findings closed, **R-100** filed.
|
||
|
||
---
|
||
|
||
## R-100 — the investigation's finding, deliberately NOT fixed
|
||
|
||
**A restic offsite tier that fails every night never goes stale on the hub.** This is **F-CRIT-2's
|
||
defect class one layer up and on the other tier** — a *failed* run resetting the freshness clock — and
|
||
it was found by asking whether F-CRIT-2's shape existed anywhere else.
|
||
|
||
Both halves verified in the source, not inferred:
|
||
- **Controller:** `o.LastRun = time.Now()` is set **unconditionally** at
|
||
`controller/internal/backup/offbox.go:716`, *outside* the `runErr` branch. The failure is recorded
|
||
faithfully, but into a different field — `o.LastStatus = "error"` at `:725`.
|
||
- **Hub:** `isStale()` reads **only** `off.LastRun` (`hub/internal/monitor/offsite.go:120`, `:127`,
|
||
`:131`) and never consults `LastStatus`.
|
||
|
||
So a nightly restic run that fails every night keeps `LastRun` fresh, `isStale` is permanently false,
|
||
and the staleness alarm never fires — with no successful offsite backup having occurred at all.
|
||
|
||
**Scope of the silence, stated precisely rather than dramatically.** `LastStatus` *does* reach the hub —
|
||
it is parsed into the report struct and **only logged** (`offsite.go:270`); it drives no checker and no
|
||
notification. The controller's own guest UI surfaces `LastStatus="error"`, so the failure is visible to
|
||
someone who looks. What is missing is the **push**: the operator's fleet-wide alarm plane is silent,
|
||
which is the plane that matters for an unattended appliance.
|
||
|
||
Not fixed, per this task's investigation-only scope. Fix direction: gate staleness on the last
|
||
*successful* run rather than the last attempt — exactly what F-CRIT-2's `NewestArchiveTime` fix did for
|
||
the PBS tier.
|
||
|
||
---
|
||
|
||
## Verification
|
||
|
||
`bash -n scripts/felhom-host-install.sh` clean. The ACL behaviour was verified **live on real hardware**
|
||
rather than by dry-run, since the whole finding is about what PVE's authorizer actually does — and the
|
||
live run is what refuted my first design.
|
||
|
||
## Fleet state
|
||
Agent **0.110.0** (with the updated sudoers) and controller **0.180.0** on both demo boxes, all healthy.
|
||
Scratch-band ACLs at 20 rows on both — re-applied after the attempt-2 destroy consumed one. No leftover scratch guests. demo-hp's `restore_test_cadence_seconds` reverted **600 → 302400**
|
||
(a bounded change made for the replay).
|
||
|
||
`felhom.eu`: this repo has a **foreign uncommitted WIP file** (`documentation/PROMPT-TEMPLATE.md`) from
|
||
another session in the shared worktree. Left untouched; my commits staged explicit paths only, per the
|
||
never-`git add -A` rule.
|