Files
felhom-agent/REPORT.md
T

10 KiB

REPORT — F-REBOOT + F-LEAK (v0.107.0 → v0.109.0) (2026-07-28)

Overwritten per the standing rule. Agent v0.106.0 → v0.109.0 across three commits, two of them corrections to the first. Companions: felhom-controller v0.180.0 (F-OBS) and felhom.eu host-install v1.21.0 (F-LEAK's actual fix).

Baselines (reconfirmed, not copied)

Agent 0.106.0 live on both boxes at start; controller 0.179.0; felhom-agent af1c21a, felhom.eu d0cec9d, felhom-controller 8f46495, all clean.


F-REBOOT — a guest rebooted mid-backup never came back (v0.107.0, SHIPPED + PROVEN LIVE)

Fault 11: the backup succeeded, and the guest was found stopped with 0 containers, no lock, and nothing retrying — 9m47s of total appliance outage with every alarm silent, because nothing was broken except that the box was off.

Why the existing recovery missed it. RecoverStaleLockedGuests already does unlock → delete dangling snapshot → start iff onboot, and it is correct. Two narrow gaps: its predicate acts only on a guest holding a stale vzdump lock (fault 11's guest was unlocked), and it runs once at agent startup on the load-bearing invariant that a lock present then is stale by definition — a guest that goes down while the agent is already up is never re-examined.

internal/localapi/guestpower.go closes exactly those two and nothing more: periodic (60 s), acting on "should be running, is not, is not locked".

The trap, which is F-CRIT-1's shape

A deliberately stopped guest must not be auto-started; fighting the operator is worse than the outage. The discriminator is onboot, and it is deliberately not invented here — it is already what stalelock.go uses for this same decision (so the two paths cannot disagree), it is 0 on scratch/golden guests, and it is what pve-guests itself consults at host boot, so the agent agrees with the platform instead of keeping a second private definition of "should be running". The hub's desired-state Run is stronger but hub-dependent; onboot still works on a box that has lost hub contact, which is when an unattended appliance most needs to come back.

Guards, in order: running → forget history · !onboot → return (Scenario B) · lock != "" → leave to the stale-lock path · BackupRunning true or error → do not start (a stop-mode backup stops the guest on purpose) · bounded 3 attempts at 1m/2m/4m → then ERROR once (Scenario C).

Live, on demo-hp

guest stopped out from under the agent recovered unattended in 120 s
the finding's original incident 587 s, needing a human pct start
lock-deferral guard fired live during the replay
Scenario B, same box an onboot:0 guest left stopped across the whole window

F-LEAK — and the first fix was WRONG, refuted by its own live replay

What I shipped in v0.107.0 and then removed in v0.108.0. A teardown fallback that, on the 403, adopted the stranded scratch into the felhom pool and retried the destroy — reasoning that the token holds Pool.Allocate on /pool/felhom. It fired exactly as designed and PVE refused it:

WARN  restore-test: scratch teardown failed — adopting the stranded scratch ... vmid=990000 pool=felhom
ERROR restore-test: pool adoption failed; left for Recover  vmid=990000
      err="proxmox: PUT /pools/felhom -> HTTP 500: permission denied at /vms/990000 (missing privilege ...)"

PUT /pools/{pool} also requires VM.Allocate on the VM being added. Pool membership cannot bootstrap its own authority. The code and its guard were removed rather than left in place — a path that provably cannot work is worse than no path, because it looks like a fix.

The real cause is structural, not a missing privilege in the role: FelhomAgentGuest is granted at /pool/felhom, and a guest joins that pool only when its restore completes. A failed restore-test therefore leaves a guest that exists, is in no pool, and is outside the token's reach.

The fix is felhom-host-install.sh v1.21.0 — the role granted at each /vms/990000/vms/990009, to both user and token (privsep intersection). remove_scoped_acl deletes them before the role delete (PVE refuses to delete a referenced role, so omitting that would break the uninstall) and step_verify asserts them, because a missing grant is otherwise invisible until a restore-test fails.

Live A/B on demo-hp — same guest, same token, minutes apart

A real PBS restore to 990000 without --pool reproduced the exact stranded state (990000 stopped, felhom pool members: [9201], 990000 in pool: False).

DELETE /nodes/<node>/lxc/990000 with the agent's own token
grant removed (the original defect) 403 Permission check failed (/vms/990000, VM.Allocate) — guest still present
grant restored (the fix) 200 UPID:...:vzdestroy:990000:felhom-agent@pve!agent — guest gone

Scenario E — it still cannot destroy a non-scratch guest

target result
/vms/990010 (one past the band) 403 Permission check failed (/vms/990010, VM.Allocate)
/vms/100 403 same

990010 does not exist, and PVE still answered 403 rather than "does not exist" — so PVE checks permission before existence and these are genuine authorization refusals, not artifacts. Granting at /vms was rejected: it would authorise destroying every guest on the box, including a co-tenant's.

A careless step, recorded rather than buried. That probe loop also issued a live DELETE against running guest 9201. It was refused with 500 container is running — but the permission check passed (9201 is a pool member by design), so had the guest been stopped I would have destroyed the live demo guest. The scratch-band probes were the safe ones; 9201 had no business in that list.


v0.109.0 — the guest-power watchdog gets the observable it shipped without

Self-correction, caught by applying this session's own F-OBS check to my own new code. The v0.107.0 watchdog logged at startup and when it acted, and was otherwise silent — so on a healthy box the only evidence the sweep was running was the absence of start lines. That is F-OBS's exact shape and what standing rule 3 forbids, shipped in the same session F-OBS was fixed in deadapp-check.

INFO summary every 10th sweep (10 min) carrying sweeps_since_boot, guests_evaluated, currently_stopped — what it saw, not merely that it ran, because "alive, all up" and "alive, one guest deliberately left down" are different operator facts. Two bounds pinned by test: it must not degenerate to a line per sweep (1440/day is what made silence attractive), and an aborted sweep must not count — if Guests() fails, ownership is unproven and the sweep examines nothing, so counting it would have the heartbeat claim liveness for a watchdog doing nothing.


Red-proofs — 8 total, all observed failing

# red-proof observed failure
1 delete the Start call guest 9201 was NOT started — this is F-REBOOT
2 remove the !onboot guard a deliberately stopped guest (onboot:0) was started
3 remove the max-attempts branch start attempted N times, want at most 3 — infinite retry loop
4 drop the backoff retried inside the backoff window — the bound is not being honoured
5 make BackupRunning non-fatal started despite being unable to confirm no backup is running
6 unfiltered guest list on error acted with unproven ownership
7 remove the liveness observable no liveness observable after 10 sweeps — silence is indistinguishable from a dead watchdog
8 count aborted sweeps as healthy an aborted sweep was counted as healthy (sweeps=20)

go build, go vet ./..., go test ./... — 29 packages, rc=0, run separately from every commit.

A process failure of mine, corrected mid-task. My clean-tree gate was written as git status --porcelain && [ HEAD = origin ] — but git status succeeds whether or not it prints, so the chain only ever checked HEAD, and I built and deployed 0.109.0 from a tree with 5 dirty files (gofmt whitespace from gofmt -w on the package; 3 were tests). Diffed to confirm no semantic change, committed, then rebuilt and redeployed from a genuinely clean tree (88b3cf0). The gate is now if [ -n "$(git status --porcelain)" ]; then abort; fi.


Deployed & fleet state

Agent 0.109.0 on demo-hp and demo-felhom, both active. Scratch-band ACLs applied on both boxes. No leftover scratch guests — pct list shows only 9201(running) on demo-hp.

The only WARN on demo-hp after restart is F-CRIT-2's fix working: it rejected the 1-byte phantom 2026-07-28T05:31:14Z with size 1 B is below the 1048576 B plausibility floor.

Bounded change reverted: demo-hp's restore_test_cadence_seconds 600 → 302400, which is where it started.


NOT yet live-validated (carried forward)

  • Fault 4 — restic transport interruption; four injection approaches defeated by guest-bridged networking. Still the most valuable follow-up: F-CRIT-2 answered the phantom question for PBS and it remains open for restic.
  • R-100 (new, filed today) — a restic offsite tier failing nightly never goes stale on the hub; isStale reads only LastRun, which the controller writes unconditionally on failure. Investigated, deliberately not fixed.
  • R-99 — prune never removes phantom snapshots (no retention bug; they accumulate).
  • Fault 12 (host reboot mid-backup); the three-way backup + restore-test + GC overlap; the agent's DR bring-up path (ModeDRGuestLoss), still read-only-verified.
  • contentionAlarmAfter (3h) — unit-tested with an injected clock, not live.
  • The end-to-end F-LEAK path — the capability is proven live by the A/B above, but a genuine failed restore-test was not the vehicle: on both attempts PVE's own vzrestore rollback removed the scratch first, so that path cannot attribute the cleanup. The A/B at the seam was chosen precisely because it can.