# CONTEXT — felhom-agent working state > Snapshot of the current state + open threads. Authoritative history lives in `CHANGELOG.md` (top > entry = current); the end-of-task detail lives in `REPORT.md`. ## Current - **2026-07-28 — v0.107.0: F-REBOOT fixed — a guest rebooted mid-backup now comes back by itself.** New `internal/localapi/guestpower.go`: a 60 s watchdog that starts a guest which is `onboot:1`, stopped, unlocked, and has no vzdump in flight. It closes the two narrow gaps that let `RecoverStaleLockedGuests` miss campaign fault 11 — that recovery acts only on a **stale vzdump lock** (fault 11's guest was unlocked) and runs **once at agent startup** (fault 11's guest went down while the agent was already up). `onboot` is the deliberate-stop discriminator and is *not* invented here: it is already what `stalelock.go` uses for this decision, it is 0 on scratch/golden guests, and it is what `pve-guests` consults at host boot — so the agent agrees with the platform instead of keeping a second private definition of "should be running". Retry bounded at 3 (1m/2m/4m) then escalates **once**; an unbounded silent retry loop is the over-correction here. Live on demo-hp: **120 s unattended** recovery vs the incident's **587 s** with a human; Scenario B proven (an `onboot:0` guest left stopped throughout). Detail: `REPORT.md`. - **2026-07-28 — F-LEAK took THREE attempts; v0.108.0 and v0.110.0 are the corrections.** The cause is structural: `FelhomAgentGuest` is granted at `/pool/felhom` and a guest joins that pool only when its restore **completes**, so a *failed* restore-test leaves a pool-less guest out of reach (403). **(1) v0.107.0 pool adoption — REFUTED LIVE:** `PUT /pools/{pool}` also requires `VM.Allocate` on the VM being added, so membership cannot bootstrap its own authority; removed in **v0.108.0**. **(2) host-install v1.21.0 per-path `/vms/990000..990009` ACLs — works, but exactly ONCE per slot:** PVE's destroy calls `AccessControl::remove_vm_access` (`API2/LXC.pm:906`) which deletes every ACL at `/vms/` (`AccessControl.pm:1898`) — **the grant is consumed by the op it authorises**. Caught by counting ACL rows after the fix, not by reasoning. **(3) v0.110.0 SHIPPED — `Privileged.DestroyScratchLXC`, the FOURTH root-fenced exception** (was exactly three: keyctl `pct create`, USB mount/fstab, SMART/sensors). Band enforced in **sudoers literally** (`pct destroy 99000[0-9] --purge`) + re-checked in code + journal provenance at the caller; none is consumed by use. API destroy still tried FIRST; band ACLs stay provisioned so the common case needs no privileged call. **Ships with a sudoers change — deploy `configs/felhom-agent.sudoers` WITH the binary.** Live: token 403 on a stranded scratch → fenced path removed the guest and all 3 LVs; sudo PERMITS the band and REFUSES `9201`/`9100`/`9999`/`990010`/`1`, and refuses `pct start 990000` too. - **2026-07-28 — v0.109.0: the guest-power watchdog got the observable it shipped without.** A self-correction: v0.107.0's watchdog logged only at startup and when it *acted*, so on a healthy box its health could be read only from **absence** — F-OBS's exact shape, shipped in the same session F-OBS was fixed in the controller. Now an INFO summary every 10th sweep carrying `sweeps_since_boot`/`guests_evaluated`/`currently_stopped`. An **aborted** sweep (unproven ownership) does not count, or the heartbeat would claim liveness for a watchdog examining nothing. - **2026-07-28 — v0.106.0: F-CRIT-2 fixed — a failed backup no longer looks like a fresh one.** `NewestArchiveTime` counted an aborted PBS upload (1 byte, manifest-less, NEWEST) as a successful backup, so the tier reported fresh, went **not due**, and was never retried — 7 days of silence on the real 168h cadence, invisible to both the R-88 breaker (defers only DUE tiers) and the hub deadline monitor (reads the same freshness). Now only *plausibly complete* entries count, via a measured floor `minPlausibleArchiveBytes` = 1 MiB; undecidable ⇒ not counted. **Size is the only tier-agnostic discriminator** — `verification` and `encrypted` are absent on every local (dir) archive and on a good PBS snapshot until verify-new catches up, so gating on either would reject 100% of local backups and cause fleet-wide backup THRASH. Floor measured: smallest real backup on the fleet is 612,397,450 B, so 1 MiB leaves 584x headroom (asserted by a test). Rejections logged at WARN once per volid. Re-tested live by replaying campaign fault 2 on demo-hp — both directions, incl. a no-thrash window with 91 scheduler ticks as the positive observable. Deployed on both boxes. Detail: `REPORT.md`. **Also established:** server-side prune does NOT count phantoms toward `keep-last` (dry-run kept 2 real + the phantom) ⇒ **no retention/data-loss bug** — but it never removes them either, so they accumulate. Filed as R-99 (LOW). - **2026-07-25 — v0.95.0 (additive): SMART coverage fixes (spike B+A) + device model.** Union-path drives (USB/registry) now get SMART via `storage.SmartReader.SMARTForBacking` wired into the localapi `/disks` union (localapi `Smart` seam); `smartDeviceFor` resolves dm/LVM to the whole disk via `/sys/block//slaves` (recursive, skips >1-disk); the builtin `local` dir on the LVM root gets a **SMART-only** device from its containing filesystem (never touches backing/durable_id — the removable-safety guard in build() stays intact); `SmartSummary.ModelName` captured from smartctl. The watchdog `Known` path stays enrich-free. Consumed by controller v0.171.0. Source of WHERE: `felhom.eu/documentation/audits/SPIKE-smart-coverage-2026-07-25.md`. - **2026-07-24 — v0.94.0 (additive): SMART serialized into /disks.** `localapi.DiskInfo` gains `Smart *hub.SmartSummary` (omitempty), copied from the target's already-computed Observe-time enrichment when `Health != ""` — no new smartctl load, no endpoint, no sudoers/MinAgent change. The controller v0.169.0 renders a "Lemezek állapota" card + 6h degradation alert from it; old controllers ignore it. **NOTE: at the remote-site vacation window the agent is DOWN (localapi binds .162 → fails), so live /disks-from-real-agent validation is deferred — the field is unit-proven; publish only.** - **2026-07-22 — v0.93.0 is the FLEET AGENT.** Built, published (sha `a68b2ff73200622e…`), Day-0-manifest-vouched (MinAgent also 0.93.0, operator-ruled) and deployed to BOTH boxes (`demo-felhom-8363b5` + `demo-hp-bb76ea`, the latter over G1 break-glass — still no key baked); clean-restart 5/5 on both, `.bak-0.92.1` retained. Discharges the onboarding runbook §A5 ceremony gate. Record: `felhom.eu/documentation/pilot/RUNBOOK-publish-agent-0.93-2026-07-22.md`. **The bullet below ("agent is DOWN … deployed 0.90.0") is SUPERSEDED history** — vmbr0 was made static .162 on 2026-07-20 (F1 mitigation) and the agent has been up since; kept for the record. - **2026-07-20 — REMOTE SITE until ~2026-08-02; the agent is DOWN there and cannot self-recover.** felhom-pve moved off the home LAN; `ssh felhom-pve` = tailnet `100.70.170.35` (direct, ~37 ms). The host is on DHCP and holds `192.168.0.147`, so `localapi`'s literal `192.168.0.162` bind fails with `bind: cannot assign requested address` — the daemon exits ~1.1 s after start, systemd gave up after 4 retries, and a manual restart reproduces it exactly. Deployed binary is **0.90.0**. Fix needs `listen_addr` in `/etc/felhom-agent/agent.json` **and** the guest bootstrap endpoint (plus the pinned leaf's SAN) → **Viktor GO**; re-pinning to another literal just re-breaks on the next lease. Also re-observed each start: `pbs: cannot read token secret … /etc/pve/priv/storage/felhom-pbs.pw: permission denied` (R-39-adjacent). Evidence + ranked findings: `felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md` - **v0.90.0** (2026-07-17) — **agent train: guest RAM resize (R-24) + fast-tick (R-28); LIVE on BOTH demo hosts (felhom-pve + nested demo-vm-felhom-4846bc).** MinAgent coupling: felhom-controller v0.143.0 gates its resize UI on this agent. (1) **R-24 guest RAM resize (controller-direct)** — self-scoped `GET`/`POST /guest/memory` (`internal/localapi/guestmemory.go`); the AGENT enforces every bound fresh per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)) and applies via PVE `SetConfig` — **live cgroup apply, no reboot** (Phase-0 PROVEN on the nested box; the break-glass access path + the proof are in `~/.claude/.../nested-vm-access-breakglass.md`). Verify- after-apply re-reads maxmem before claiming success. New narrow `MemoryOps` seam (GuestAPI untouched); memory only. (2) **R-28 fast-tick** (`internal/fasttick/`) — while any desired-state item is unapplied (esp. the pre-tunnel WG-registration window a hub poke can't reach) pulse the shared out-of-band trigger every 30 s, self-disarm on convergence; four cached sources (desired-gen==0, reconcile Planned−Pending>0, pbsdr waiting_secret ONLY, wgtunnel desired-not-operational). Seams: `reconcile.Engine.LastResult()` + `wgtunnel.Manager.TunnelConvergence()` (cached — no per-tick exec). (3) **Guests-0/0** REFUTED live: the 0/0 was the pre-provision window (guest not yet created), not a pool-membership bug; the fast-tick shortens that window. **OPEN (operator GO):** publish 0.90.0 + hub Day-0 manifest vouch + MinAgent-floor raise to 0.90.0 (password-gated UI; the safety gate — both agents on 0.90.0 — is satisfied and the coupling is proven live via the version header). See REPORT.md. - **v0.89.0** (2026-07-16) — **agent train: three bundled agent-plane items; built + published to Gitea (sha256 `3969fd91…`); paired with hub 0.59.0 (LIVE).** (1) **pbsdr self-grant (R-22)** — closes the F4 self-deadlock: a 403 on the token-auth `StorageEntry` pre-check now self-grants via the root wrapper + re-reads instead of aborting before the grant (the demo's `felhom-offsite` case). (2) **escrow config live-reload** — `/escrow/preflight`'s `pbs_storage_id` row now reads the live agent.json (late-bound `CurrentPBSStorageID`) so a pbsdr-seeded id flips green with no restart. (3) **agent-plane poke listener (Direction-2a)** — `internal/poke`: contentless UDP poke bound to the box WG /32 (port **51822**), leading-edge debounced, fires the hub-loop out-of-band trigger for an immediate desired-state cycle; enabled with `wg_tunnel.enabled`; first slice of R-13. Red-proofs for all three (run-fail-revert). **ALL THREE LIVE LEGS PROVEN on the demo (2026-07-17), demo now LIVE on 0.89.0:** Scenario 4 floor-driven A/B train 0.88→0.89 (operator signed+enqueued the `agent_update` op — the vouch+floor alone does NOT trigger it; committed, no rollback); Scenario 1 R-22 self-heal (marker aside + ACLs revoked → `pre-check 403 … self-granting (R-22)` → `converged state=adopted` in ~3 s, ACLs restored, offsite active); Scenario 3 poke→tick ~31 ms ep0→box + immediate report cycle (save→tick ≈ ~0.45 s). Details: REPORT.md. - **v0.88.0** (2026-07-13 eve) — **controller-driven escrow ceremony (agent half), LIVE on demo host + drill VM (63/63 capabilities both).** `--output=json` machine mode (text mode byte-identical; extraction into `escrowCeremony()`); the ONE fixed argv (`escrow.CeremonyArgs()` — shared by the localapi exec + the `escrow-ceremony` capability (Critical, pbs_dr-gated EXPLICIT) + the new `FELHOM_ESCROW` sudoers alias, three-way pin-tested); localapi job endpoints (`POST /escrow/ceremony` single-flight 60 s, status, ONE-SHOT claim → 410, 10-min TTL → `unclaimed_void`, `GET /escrow/preflight`). R in-memory ONLY (never the job struct — snapshot-hygiene-tested; restart loses it safely). Live-proven on drill endpoint-exact: stage → preflight all-green (live FELHOM_ESCROW list-probe) → job ~4 s → hub blob `restic_pw_sha256` covering (repaired the spike's hash-less blob) → claim 200 once → 410. Coupled: controller v0.127.0 (MinAgent 0.88.0 for the wizard). **OPEN: publish 0.88.0 + Day-0 manifest vouch (operator) at the next train; deployed hosts got direct deploys.** Details: REPORT.md + felhom.eu RUNBOOK-escrow-ceremony.md (F1 threat model). - **v0.87.0** (2026-07-13) — **SystemDisks device-mapper walk (IA finding 2, MEDIUM): legacy-boot hosts get a working drive wizard.** Operator ruling (approved 2026-07-13, verbatim): *resolve device-mapper/raid parents — for the root filesystem's backing block device, walk `/sys/block//slaves` recursively down to physical disks; those, plus any ESP holder when present, are system. Disks outside that set become wizard candidates (still subject to the existing data-bearing guards). The all-system fail-safe remains ONLY for walk failure — it returns to being the error case, not the legacy-boot common case.* Implemented as `physicalDisksOf`/`walkSlaves` + `HostReader.BlockSlaves` (one seam method); per-branch conservatism (any unresolvable slave → ok=false → unchanged all-system path); signature test `TestSystemDisks_WalkTopologies` (root-backing disk ALWAYS system — never weaken). §3 spike transcripts: drill (legacy) dm-1→sda3→sda; felhom-pve (EFI+LVM) ESP+walk agree on sda → byte-identical regression. §13.2 wizard leg COMPLETE (offered → enrolled → formatted → torn down, boxes as found) + Day-0 manifest vouched to 0.87.0 (operator). The leg also surfaced two CONTROLLER bugs (fixed same-day: v0.126.3 claimed-box wizard CSRF, v0.126.4 502-through-CF + native-alert ban). - **v0.83.0** (2026-07-11, LIVE on felhom-pve; NOT published — Peti stays 0.81.0) — **observability pass** (pairs with controller v0.116.1 + hub v0.46.0). `applog.New` → `(logger, *Ring)`: slog fan-out, journald at the configured level, ~1000-entry ring FIXED at DEBUG. `GET /debug/logs` (local API, token-authed; the controller Debug page's Ügynök tab) + request-level DEBUG middleware. Heartbeat log-pull: envelope `log_tail_requested` → next heartbeat ships `log_tail` (128 KB, consume-once; failed push re-armed by the next envelope; `operator log pull served` INFO on fulfillment). Gap-fill sweep: netverify phase/verdict lines (job start, trigger outcome, /proc/mounts verdict, journal bytes, classification code, rollback outcome, durations), netmount unit steps, signedjobs op-received (class/host/expiry — never signatures) + fetch duration, selfupdate invariants + download sha/duration, disks outcome INFOs, controller-swap pre-pull + health verdicts, desired/loop per-exchange DEBUG. Logging conventions: `felhom.eu/documentation/runbooks/logging-conventions.md`. OPEN: the hub-side live pull awaits the operator's button click (hub UI password-gated); pre-existing lanresolver permission-denied WARN on /var/lib/felhom-agent/guests noted in REPORT. - **v0.77.0** (2026-07-09) — **fork-4: escrow the offsite restic repo password under R.** `IdentityBundle` gains `ResticRepoPassword` (rides the existing age-under-R `WrapIdentityBundle` path — validated by the custody spike `febdc56`). New `POST /escrow/stage-secret` (`withGuest`) transiently stages the controller-pushed password (0600, never logged), which the `--selftest=escrow-create` ceremony auto-injects into the bundle and then wipes. `AttachResticPassword`/`StagedResticPasswordPath`/ `WipeStagedResticPassword` added. Pairs with controller v0.105.0 (push + atomicity gate + DR inject + `DRResticCoord`). **NOT yet live-validated** — the supervised escrow ceremony is operator-run. - **v0.76.0** (2026-07-08, LIVE on felhom-pve + **PUBLISHED sha `9828c5f7…f50b`** — THE Day-0 manifest bump target; **0.75.0 superseded unpublished**) — **GL-5b / G12: restore-test full-fidelity**. Params derive from the ARCHIVE's embedded config (`drRestoreOverrides`, same as DR — the old live-source-config path verified the wrong object AND dropped storage mpN per PVE's all-or-nothing rule; deleted with `bindMountOverrides`/`archiveVMID`). NEW mount-parity assert (restored mpN vs archive; miss/mispath/undersize/extra = FAIL naming the delta) + `MountParity`/ `MountInventory` on the wire record (additive). Live-proven: scratch 990000 ← 6.5GB 9201 archive, parity ok, inventory mp0 200G+mp1 50G+2 throwaways, **3m4s local tier** (cheaper than feared); rotated-out archive volid → clean up-front refusal (nice failure mode). bringup.go untouched. - **v0.75.0** (2026-07-08, LIVE on felhom-pve) — **GL-5 / go-live G8: guest-loss DR bring-up actually restores** (closes the v0.74.0 OPEN item + SPIKE-dr-bindmount-source §8). DR passes the COMPLETE explicit restore param set derived from the archive's embedded config (NEW `Client.ExtractArchiveConfig`, 200 under the scoped token) — **two live-discovered PVE rules: mpN params need an explicit rootfs, AND unlisted mountpoints are silently DROPPED** (first run booted without mp0/mp1!) — storage mpN passed through, structural mp8/mp9 → throwaways, then step 4d swaps the REAL binds in via the host runner (root pct; new `EngineOptions.HostRunner`+`StateDir` seam) and deletes the unusedN residue. Scratch-DR live-proven end-to-end (9310 from a real 9201 archive: mp0 200G + mp1 50G + real binds + no residue + clean teardown). Provision = nil overrides (regression-tested). NOTE: published/vouch-pending agent is 0.74.0 — publish 0.75.0 before/with the manifest bump. OBSERVATION: the DR selftest hardcodes KeepMAC=true — a scratch DR while the SOURCE guest is live briefly duplicates its MAC on the bridge (pre-existing; fine for supervised runs, worth a -keep-mac flag someday). Full customer-data DR drill = GL-6/S5 family. - **2026-07-07 — v0.74.0 Gitea-PUBLISHED (RUNBOOK GL-1)** — the LIVE felhom-pve binary's exact bytes, sha256 `1ec3f58842edce1e…76af05`, anon-fetch-verified. This supersedes/closes every standing "publish 0.6x + Day-0 vouch" OPEN item below (0.64→0.73 were never published; 0.74.0 is the vouch target). Golden 0.103.0 published in the same run (felhom.eu execution record `documentation/pilot/RUNBOOK-GL1-publish-2026-07-07.md`). **Day-0 manifest vouch = operator step** (agent 0.74.0 / golden 0.103.0). - **v0.74.0** (2026-07-07) — **campaign-2 R2 CLOSED; the mislabelled "R1" was a symptom** (LIVE on felhom-pve). Pool membership is what lets the pool-scoped token reach a guest; `pct restore --pool` sets it only at CREATE, so a restore-over-existing dropped 9201 from the `felhom` pool → no `VM.Audit` → restore-test's *existing* `bindMountOverrides` never ran → "mp8 … only possible for root". Fix: `Client.PoolAddVMID` + bring-up re-asserts membership post-restore (warn-not-fail). Role/ACL + `bindMountOverrides` untouched (both correct). **Live restore-test PASSED for the first time** once the pool was healed (Part A one-liner): read config → neutralize 2 binds → restore → boot+running → clean teardown, 4m35s. B3 (scratch-teardown 403) confirmed a cascade — no code. OPEN: DR `bring-up -mode dr` bind-override gap (spike `SPIKE-dr-bindmount-source-2026-07-07.md`: small known-constant override reusing `bindMountOverrides`; mp8/mp9 are structural constants). - **v0.73.0** (2026-07-06) — **F2 mount-role fallback CLOSED** (LIVE on felhom-pve). `roleForMountPath` gained a mount-table fallback (Impl-2b style): a bind-mounted RAW enrolled user-data drive is not a PVE storage, so it fail-safe'd to `system` and the eject/decommission gates 403'd EVERY user-data drive (campaign F2, `where=/mnt/teszt_enroll role=system`). Device-keyed classification + whole-disk containment (`storage.SameWholeDisk`); Observe-error keeps the fail-safe BEFORE the fallback. Only `roleForMountPath` touched. Live-proven full lifecycle on teszt_enroll (eject/decommission 200, no-rebind across restart, end==pre). OPEN follow-up: the `deviceRole`/`roleForMountPath` unification refactor (deferred). - **v0.72.0** (2026-07-05) — **OOB operator access (merged E1+H1)** — TASK H1, provenance both `SPIKE-{felhom-sshd,oob-wg-operator-peer}-2026-07-05`. Operator `/32` RENDERED into wg-felhom AllowedIPs (survives self-heal, [OF-1]); dedicated `internal/felhomsshd` (port claim + config render→sshd -t→reload + operator authorized_keys + heal + oob heartbeat stanza); static `inet felhom_oob` belt (agent mutates SET ELEMENTS ONLY); `configs/felhom-sshd.service` (NO RuntimeDirectory [SF-1]) + `felhom-oob.nft` + `felhom-op.sudoers`; `FELHOM_SSHD`+`FELHOM_OOB` grants; `oob.enabled` DEFAULT FALSE. Live on felhom-pve (8822, belt filled, operator SSH as felhom-op with scoped sudo); hub v0.35.0. Rollback `.bak-0.71.0`. 5 live-found bugs fixed (port path, self-listen flip-flop, nil-block lockout, reachable-via-dial, operator-configured source). - **v0.71.0** (2026-07-05) — **management-plane break-glass: privsep-dir watchdog + mgmt_plane health** — TASK G1 (prereq for felhom-sshd/H1), provenance `SPIKE-felhom-sshd-2026-07-05` §8. Host artifacts (`configs/felhom-privsep.tmpfiles` + `felhom-mgmt-watchdog.{sh,service,timer}`) make `/run/sshd` boot-persistent AND auto-heal it every ~60s **agent-independently** (heals with the agent stopped — proven live: `/run/sshd` removed → restored in 30.0s, `:22` back, no login). `internal/mgmtplane` reports the additive `mgmt_plane` heartbeat stanza; hub v0.34.1 raises `mgmt_plane_healed`. **NO unit declares `RuntimeDirectory=`** (the incident cause). H1 may now assume `/run/sshd` is guaranteed present. Live on felhom-pve; rollback `.bak-0.70.0`. - **v0.70.0** (2026-07-05) — **agent self-update (operator-signed A/B slots + crash-loop auto-rollback)** — TASK D1, provenance `SPIKE-agent-selfupdate-2026-07-05`. An operator-signed `agent_update` op (version+sha256, sha is the only integrity root) rides the signed-jobs gate; `internal/selfupdate.Executor` downloads+verifies+hands to `felhom-selfupdate-guarded apply` (root re-verify → A/B atomic flip → pending marker → detached restart); the new binary commits after a 60s dwell; a crash-looping binary is auto-reverted by `OnFailure=felhom-agent-rollback.service` (first-crash trigger [SF-1]) with the tuned `[Unit]` start-limit (120s/4) as backstop. Host artifacts + sudoers `FELHOM_SELFUPDATE` + `felhom-host-install.sh` day-0 install + report field `selfupdate_pending`. Green tests + companions. **LIVE-VALIDATED on felhom-pve (2026-07-05): all 4 drills PASS** — happy path (0.70.0→0.70.1 signed op → download+verify+flip+commit), crash-rollback (0.70.2-crash → OnFailure → **~2s crash-to-recovered**, byte-identical revert, no loop), no-pending guard, gate refusal (non-pinned key). Full agent-side pipeline ran real (envelope injected into the hub `signed_jobs` queue — CC lacks the hub global operator key; hub enqueue-auth is hub-unit-tested). Box restored to canonical **v0.70.0** (host artifacts KEPT installed; scratch operator key REMOVED — self-update dormant until an operator pins a real key, a Day-0-vouch-style follow-up). Rollback `felhom-agent.bak-0.69.0`. OPEN (v1 scope-outs): no hub-floor auto-update, no failed-update auto-retry, no pending-timeout auto-rollback; per-crash OnFailure can double-fire (idempotent — future: serialize the rollback oneshot). Detail: REPORT.md. - **v0.69.0** (2026-07-04, live on felhom-pve) — **S5: host-loss DR — safe halves shipped**. **Part 1** `wgtunnel.InstallRecoveredKey` — writes an escrow-recovered WG privkey (create-only, refuse-overwrite) so the tunnel re-establishes with the SAME identity/pubkey (same /32), no keygen; wired into `--selftest=identity-consume -install-wg-key` (opt-in; pre-S3 blob → logged fresh-keygen fallback). **Part 2** new `internal/dr` — consumes the host_loss `restore_directive` (was logged-ignored) into an inspectable RestorePlan via AddConsumer: per-guest {vmid,archive,target, sizing} + per-drive {durable_id→mount} + offsite PBS coord; DERIVE-AND-SURFACE only (Consumer has no restore/destroy dep — execute-nothing is structural). Tests + red-proofs (WG create-only; plan mode-gate). **Part 3** hub escrow-GET NOT needed (operator exports the blob via `sqlite3 writefile` on a cp'd hub.db). **Part 4-A** re-attach wrong-disk safety already unit-proven (`ResolveStorageDevice`: match resolves, absent/mismatch ERRORS, non-uuid scheme refused — never a near disk). **Part 4-B (destructive in-place 9201 restore) PREPARED + OPERATOR-GATED, NOT executed** — pre-flight green (offsite ct/9201 restorable per S4.1); the operator runs the R-consume steps + confirms the destroy (§9-4a: CC never runs a consume/R command — see [[operator-present-one-time-secrets]]). OPEN: the operator-run 4-B drill; guest_loss DR; hub-driven full-auto DR. Rollback felhom-agent.bak-0.68.0. Detail: REPORT.md + doc-06 §3.5/S5. - **v0.68.0** (2026-07-04, live on felhom-pve) — **S4.1: unattended offsite restore-test**. **Tier-aware restore-task deadline:** `RestoreTestSpec.RestoreTaskTimeout` (0→10m default) from `config.RestoreTestPBSRestoreTimeoutSeconds` (accessor default **120m**), set only when `SourceTier=="pbs"` (`main.restoreTaskTimeout`); local tier UNCHANGED. Fixes the WAN restore being killed at 10m → mid-restore teardown → leaked scratch. **Teardown "VM.Allocate" follow-up = PHANTOM (diagnosed, not blind-fixed):** ran the restore-test on the AGENT-TOKEN path sourcing the offsite (pbs) backup → `pass:true verified:boot+running`, teardown succeeded (`torn down vmid=990000`, no 403), scratch band clean. The earlier 403 was the 10m-timeout consequence (guest not yet pool-associated); the scratch is restored INTO `/pool/felhom` (ACL already grants VM.Allocate) so teardown is authorized once the restore completes. **No ACL/host-install change.** OPEN: publish 0.68.0 + Day-0 vouch; Tier-1/Tier-2 split for offsite-as-default; S5 DR consume. Rollback `felhom-agent.bak-0.67.0`. Detail: REPORT.md. - **v0.66.0 + v0.67.0** (2026-07-04, live on felhom-pve) — **S4: PBS over the tunnel**. **v0.66.0**: wgtunnel **v4-pin** (renderConf writes the resolved A LITERAL, never DNS/AAAA; `Resolver` seam, lowest addr; cached → steady-state zero-DNS/zero-exec) + **re-resolve watchdog** (`Manager.Watchdog`, loop-only; handshake stale > `stale_after_seconds`=180 → re-resolve → IP-changed re-render+restart) + FELHOM_WG **Critical** flips (conf-install/enable/restart/handshake-read). **v0.67.0**: **namespace-aware PBS client** (Config.Namespace → `Snapshots ?ns=`, `Verify ns=`; root-ns unchanged) — the operator-approved fix after Phase-1 showed the ns-unaware datastore-root 403s a per-tenant token. **Live Scenario-D (all green):** real vzdump of 9201 → **ciphertext** in ns `demo-felhom-01` over the tunnel; ns-scoped verify=ok under the box's own `felhom@pbs!demo-felhom-01` token; WARN gone; restore round-tripped (decrypt with box-born key → boot → teardown). **Confirmed tenant ACL (felhom-hetzner):** `DatastoreBackup` on `/datastore/felhom-offsite/` (NOT `/ns/`) to BOTH user `felhom@pbs` AND token (privsep=intersection; cross-ns 403); DatastoreBackup can't prune (safety). **FINDINGS:** retarget field is `local_backup_target` (not `backup_target`); retarget REVERTED to `local` (controller backs up ~every 30 min → single-target offsite = near-continuous 20-min uploads; needs Tier-1/Tier-2 split); restore-test scheduler needs a WAN restore deadline + scratch-band `VM.Allocate` before it runs offsite unattended. **OPEN:** escrow-create (OPERATOR-PRESENT, new R); publish 0.66/0.67 + Day-0 vouch; S5 DR consume. Rollback: `felhom-agent.bak-0.65.0`/`.bak-0.66.0`. Detail: REPORT.md + doc-06 §3.4/§4.2 + runbook §4a/§4b. - **v0.65.0** (2026-07-04, live on felhom-pve) — **S3.1 offsite-tunnel client MTU 1420 → 1280**: resolves `06 §4.3`'s OPEN DECISION left by the CGNAT smoke test. 1420 **silently black-holed bulk TCP** on sub-~1480 paths (mobile ~1400, DS-Lite ~1452) — handshake+ping healthy, PBS TLS page (and at S4 the backup itself) drops. New `const clientMTU = 1280` (RFC 8200 IPv6-minimum floor; outer 1340 v4 / 1360 v6 fits every realistic path), **permanent + fleet-wide + family-agnostic**. **Client-only by construction** — interface MTU caps box→PBS, advertised MSS caps PBS→box, so the endpoint's `wg0` is untouched (zero live-endpoint risk). Golden pins exact `MTU = 1280` (red-proofed vs a 1420 flip); no wire/JSON change. Live: agent re-rendered on restart (hash-gated apply), conf + live iface both 1280, PBS page loads at 1280 (no regression on wired). OPEN: true-CGNAT-SIM retest (low risk); publish 0.65.0 + Day-0 vouch (operator); S4 PBS-over-tunnel. Rollback: `felhom-agent.bak-0.64.0` on the box. The v4-pin (§4.2 determinism) is a separate, optional future note — NOT needed for MTU correctness. - **v0.64.0** (2026-07-04, live on felhom-pve) — **S3 offsite WG tunnel**: new `internal/wgtunnel` (keygen 0600/0700, marker-gated one-shot registration, agent-managed `wg-quick@wg-felhom` from the hub's desired-state `wireguard` block via the new `desired.Syncer.AddConsumer` seam, revoked-stays-revoked teardown, report stanza) + `FELHOM_WG` sudoers/capabilities + `IdentityBundle.WGPrivateKey` escrow auto-inject. **`wg_tunnel.enabled` DEFAULTS FALSE** (safety gate — rollout to Peti's box is a no-op until the production endpoint exists; enabled explicitly on felhom-pve only). Live: tunnel to ep0.felhom.eu:443 up 3 s after enable (PBS page through 10.77.0.1:8007), reboot-persistent, revocation drill clean, 30-min keepalive soak. GOTCHAS: hub envelope poll_interval_seconds (hub-side const 900 s) silently overrides agent poll_seconds on cycle 1; `wg show dump` leaks the PRIVATE key (forbidden everywhere — sudoers only grants `latest-handshakes`). OPEN: CGNAT/mobile-hotspot smoke (operator-assisted appendix); publish 0.64.0 to Gitea + Day-0 vouch (operator); S4 points PBS at the tunnel. - **configs: build-golden.sh v2.0.0** (2026-07-03, @ `ceca355`; no agent version change) — **drill findings B5 + B1 FIXED** (`DRILL-golden-098-2026-07-03.md`): the controller tag is a MANDATORY argument (the default rotted twice — a fresh install booted a pre-floor controller, forcing the guide's manual D.1b update) and the golden now bakes a `felhom-controller-bootstrap.path` unit (controller deploys the moment the back-half hot-plugs the bootstrap mount — no reboot; installer v1.9.1's reboot is a redundant belt, kept). **Golden 0.98.3** baked on the drill VM, clean-room validated (bake integrity → isolated hot-plug proof → local-golden Day-0 → published-artifact Day-0), published (sha256 b9a02ef1…fd01) + operator-vouched — Day-0 manifest now vouches **agent 0.63.0 + golden 0.98.3** (the v0.63.0 vouch follow-up below is DONE). Fresh installs land current and self-manage. NEW operator follow-up (SECURITY): the customer-config `git.token` has Gitea package-WRITE rights — scope down + rotate (evidence-doc observation O1). - **v0.63.0** (2026-07-03, live on felhom-pve + Gitea-published sha256 b4a89c81…) — **drill findings B3 + B2 FIXED** (`DRILL-day0-cleanroom-2026-07-03.md`): `TokenStore.Lookup` reloads the append-only store once on a miss (cross-process coherence with the one-shot provisioner — no more fresh-install `/controller/swap` 401 / manual restart; size short-circuit bounds the cost; behind the `TokenAuthority` seam) + `guesthook.InstallSnippet` issues a fenced `mkdir -p /var/lib/vz/snippets` first (fresh boxes lacked the dir → the self-heal hook silently never installed). Sudoers gained exactly that one grant — **ship sudoers WITH the binary** (done on felhom-pve). Red-proofed both; Scenario-E method: compiled test suite run ON felhom-pve + live channel-health hit-path. **OPERATOR FOLLOW-UP: bump the hub Day-0 manifest to agent 0.63.0** — until then fresh installs get 0.62.0 and the guide's D.1b restart-first step still applies (narrowed to "< v0.63.0" in the guide). - **v0.62.0** (2026-07-03) — **audit A1 RESOLVED**: the stale-lock reaper's scan is now pool-intersected (`staleLockController.Guests()` = `ListLXC` ∩ `Client.Pool("felhom")` members), fail-safe skip on pool-read failure; `pve:pool-read` capability (non-critical) + `--selftest` "pool read" line. Companion host-install **v1.9.0** adds `Pool.Audit` to `FelhomAgentGuest` — **deploy order on any box: rescope ACL first, then this agent.** Per `SPIKE-a1-pool-membership-read-2026-07-03.md`; red-proofed tests in stalelock_pool_test.go. - **2026-07-03 — CLAUDE.md refreshed**: version narrative removed (state lives HERE + CHANGELOG top), layout completed (all 17 internal packages + cmd/felhom-opsign); deploy runbook now in the `felhom-build-deploy` skill (`felhom.eu/skills/`). - **2026-07-03 — `REUSE.md` exists at the repo root** (canonical helpers / format-safety guards / traps / seams, code-verified); maintenance rule active: update it in the same commit that changes a shared helper. - **v0.61.0** (2026-07-03) — blast-radius audit fixes **B1 + D1 + D2 + D3** from `felhom.eu/documentation/audits/AUDIT-blast-radius-hostroot-localapi-2026-07-02.md`: random temp staging for root-installed scripts (+ sudoers/manifest glob updates), mkfs-wrapper member/RO re-checks (validated by `scripts/mkfs-guarded-harness.sh`), classifyClaim empty-lsblk fail-safe, and the blank-format anti-retarget (durable-id-bound, AGENT-001's benign-branch twin). - Deployed on demo host `felhom-pve` (node `demo-felhom`), non-root `felhom-agent` service user, pool-scoped token (`felhom` pool). ## Open threads - Deferred audit items (housekeeping/design, all INFO): C1 (controller-swap version floor), C2 (NAS server allowlist), A2 (gate journal cross-check), B2–B5, E1/E2. - Drive-enrollment leftovers: (a) `runStorageInit` slow-device detached-format polling; (b) Impl-3 shared-box operator format gate. - BUNDLE leftover: non-root agent can't read the PBS key; migration must preserve cert/key/tokens. - Not run (needs a supervised session): the destructive D1/D3 live proofs (real mkfs on a crafted member; a live /dev re-enumeration race during a real format).