Files
felhom-agent/CONTEXT.md
T

34 KiB
Raw Blame History

CONTEXT — felhom-agent working state

Snapshot of the current state + open threads. Authoritative history lives in CHANGELOG.md (top entry = current); the end-of-task detail lives in REPORT.md.

Current

  • 2026-07-28 — v0.107.0: F-REBOOT fixed — a guest rebooted mid-backup now comes back by itself. New internal/localapi/guestpower.go: a 60 s watchdog that starts a guest which is onboot:1, stopped, unlocked, and has no vzdump in flight. It closes the two narrow gaps that let RecoverStaleLockedGuests miss campaign fault 11 — that recovery acts only on a stale vzdump lock (fault 11's guest was unlocked) and runs once at agent startup (fault 11's guest went down while the agent was already up). onboot is the deliberate-stop discriminator and is not invented here: it is already what stalelock.go uses for this decision, it is 0 on scratch/golden guests, and it is what pve-guests consults at host boot — so the agent agrees with the platform instead of keeping a second private definition of "should be running". Retry bounded at 3 (1m/2m/4m) then escalates once; an unbounded silent retry loop is the over-correction here. Live on demo-hp: 120 s unattended recovery vs the incident's 587 s with a human; Scenario B proven (an onboot:0 guest left stopped throughout). Detail: REPORT.md.

  • 2026-07-28 — F-LEAK took THREE attempts; v0.108.0 and v0.110.0 are the corrections. The cause is structural: FelhomAgentGuest is granted at /pool/felhom and a guest joins that pool only when its restore completes, so a failed restore-test leaves a pool-less guest out of reach (403). (1) v0.107.0 pool adoption — REFUTED LIVE: PUT /pools/{pool} also requires VM.Allocate on the VM being added, so membership cannot bootstrap its own authority; removed in v0.108.0. (2) host-install v1.21.0 per-path /vms/990000..990009 ACLs — works, but exactly ONCE per slot: PVE's destroy calls AccessControl::remove_vm_access (API2/LXC.pm:906) which deletes every ACL at /vms/<vmid> (AccessControl.pm:1898) — the grant is consumed by the op it authorises. Caught by counting ACL rows after the fix, not by reasoning. (3) v0.110.0 SHIPPED — Privileged.DestroyScratchLXC, the FOURTH root-fenced exception (was exactly three: keyctl pct create, USB mount/fstab, SMART/sensors). Band enforced in sudoers literally (pct destroy 99000[0-9] --purge) + re-checked in code + journal provenance at the caller; none is consumed by use. API destroy still tried FIRST; band ACLs stay provisioned so the common case needs no privileged call. Ships with a sudoers change — deploy configs/felhom-agent.sudoers WITH the binary. Live: token 403 on a stranded scratch → fenced path removed the guest and all 3 LVs; sudo PERMITS the band and REFUSES 9201/9100/9999/990010/1, and refuses pct start 990000 too.

  • 2026-07-28 — v0.109.0: the guest-power watchdog got the observable it shipped without. A self-correction: v0.107.0's watchdog logged only at startup and when it acted, so on a healthy box its health could be read only from absence — F-OBS's exact shape, shipped in the same session F-OBS was fixed in the controller. Now an INFO summary every 10th sweep carrying sweeps_since_boot/guests_evaluated/currently_stopped. An aborted sweep (unproven ownership) does not count, or the heartbeat would claim liveness for a watchdog examining nothing.

  • 2026-07-28 — v0.106.0: F-CRIT-2 fixed — a failed backup no longer looks like a fresh one. NewestArchiveTime counted an aborted PBS upload (1 byte, manifest-less, NEWEST) as a successful backup, so the tier reported fresh, went not due, and was never retried — 7 days of silence on the real 168h cadence, invisible to both the R-88 breaker (defers only DUE tiers) and the hub deadline monitor (reads the same freshness). Now only plausibly complete entries count, via a measured floor minPlausibleArchiveBytes = 1 MiB; undecidable ⇒ not counted. Size is the only tier-agnostic discriminatorverification and encrypted are absent on every local (dir) archive and on a good PBS snapshot until verify-new catches up, so gating on either would reject 100% of local backups and cause fleet-wide backup THRASH. Floor measured: smallest real backup on the fleet is 612,397,450 B, so 1 MiB leaves 584x headroom (asserted by a test). Rejections logged at WARN once per volid. Re-tested live by replaying campaign fault 2 on demo-hp — both directions, incl. a no-thrash window with 91 scheduler ticks as the positive observable. Deployed on both boxes. Detail: REPORT.md. Also established: server-side prune does NOT count phantoms toward keep-last (dry-run kept 2 real + the phantom) ⇒ no retention/data-loss bug — but it never removes them either, so they accumulate. Filed as R-99 (LOW).

  • 2026-07-25 — v0.95.0 (additive): SMART coverage fixes (spike B+A) + device model. Union-path drives (USB/registry) now get SMART via storage.SmartReader.SMARTForBacking wired into the localapi /disks union (localapi Smart seam); smartDeviceFor resolves dm/LVM to the whole disk via /sys/block/<dm>/slaves (recursive, skips >1-disk); the builtin local dir on the LVM root gets a SMART-only device from its containing filesystem (never touches backing/durable_id — the removable-safety guard in build() stays intact); SmartSummary.ModelName captured from smartctl. The watchdog Known path stays enrich-free. Consumed by controller v0.171.0. Source of WHERE: felhom.eu/documentation/audits/SPIKE-smart-coverage-2026-07-25.md.

  • 2026-07-24 — v0.94.0 (additive): SMART serialized into /disks. localapi.DiskInfo gains Smart *hub.SmartSummary (omitempty), copied from the target's already-computed Observe-time enrichment when Health != "" — no new smartctl load, no endpoint, no sudoers/MinAgent change. The controller v0.169.0 renders a "Lemezek állapota" card + 6h degradation alert from it; old controllers ignore it. NOTE: at the remote-site vacation window the agent is DOWN (localapi binds .162 → fails), so live /disks-from-real-agent validation is deferred — the field is unit-proven; publish only.

  • 2026-07-22 — v0.93.0 is the FLEET AGENT. Built, published (sha a68b2ff73200622e…), Day-0-manifest-vouched (MinAgent also 0.93.0, operator-ruled) and deployed to BOTH boxes (demo-felhom-8363b5 + demo-hp-bb76ea, the latter over G1 break-glass — still no key baked); clean-restart 5/5 on both, .bak-0.92.1 retained. Discharges the onboarding runbook §A5 ceremony gate. Record: felhom.eu/documentation/pilot/RUNBOOK-publish-agent-0.93-2026-07-22.md. The bullet below ("agent is DOWN … deployed 0.90.0") is SUPERSEDED history — vmbr0 was made static .162 on 2026-07-20 (F1 mitigation) and the agent has been up since; kept for the record.

  • 2026-07-20 — REMOTE SITE until ~2026-08-02; the agent is DOWN there and cannot self-recover. felhom-pve moved off the home LAN; ssh felhom-pve = tailnet 100.70.170.35 (direct, ~37 ms). The host is on DHCP and holds 192.168.0.147, so localapi's literal 192.168.0.162 bind fails with bind: cannot assign requested address — the daemon exits ~1.1 s after start, systemd gave up after 4 retries, and a manual restart reproduces it exactly. Deployed binary is 0.90.0. Fix needs listen_addr in /etc/felhom-agent/agent.json and the guest bootstrap endpoint (plus the pinned leaf's SAN) → Viktor GO; re-pinning to another literal just re-breaks on the next lease. Also re-observed each start: pbs: cannot read token secret … /etc/pve/priv/storage/felhom-pbs.pw: permission denied (R-39-adjacent). Evidence + ranked findings: felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md

  • v0.90.0 (2026-07-17) — agent train: guest RAM resize (R-24) + fast-tick (R-28); LIVE on BOTH demo hosts (felhom-pve + nested demo-vm-felhom-4846bc). MinAgent coupling: felhom-controller v0.143.0 gates its resize UI on this agent. (1) R-24 guest RAM resize (controller-direct) — self-scoped GET/POST /guest/memory (internal/localapi/guestmemory.go); the AGENT enforces every bound fresh per request (min 2048 / max host_total2048 / shrink floor max(2048, usage+512)) and applies via PVE SetConfiglive cgroup apply, no reboot (Phase-0 PROVEN on the nested box; the break-glass access path + the proof are in ~/.claude/.../nested-vm-access-breakglass.md). Verify- after-apply re-reads maxmem before claiming success. New narrow MemoryOps seam (GuestAPI untouched); memory only. (2) R-28 fast-tick (internal/fasttick/) — while any desired-state item is unapplied (esp. the pre-tunnel WG-registration window a hub poke can't reach) pulse the shared out-of-band trigger every 30 s, self-disarm on convergence; four cached sources (desired-gen==0, reconcile PlannedPending>0, pbsdr waiting_secret ONLY, wgtunnel desired-not-operational). Seams: reconcile.Engine.LastResult() + wgtunnel.Manager.TunnelConvergence() (cached — no per-tick exec). (3) Guests-0/0 REFUTED live: the 0/0 was the pre-provision window (guest not yet created), not a pool-membership bug; the fast-tick shortens that window. OPEN (operator GO): publish 0.90.0 + hub Day-0 manifest vouch + MinAgent-floor raise to 0.90.0 (password-gated UI; the safety gate — both agents on 0.90.0 — is satisfied and the coupling is proven live via the version header). See REPORT.md.

  • v0.89.0 (2026-07-16) — agent train: three bundled agent-plane items; built + published to Gitea (sha256 3969fd91…); paired with hub 0.59.0 (LIVE). (1) pbsdr self-grant (R-22) — closes the F4 self-deadlock: a 403 on the token-auth StorageEntry pre-check now self-grants via the root wrapper + re-reads instead of aborting before the grant (the demo's felhom-offsite case). (2) escrow config live-reload/escrow/preflight's pbs_storage_id row now reads the live agent.json (late-bound CurrentPBSStorageID) so a pbsdr-seeded id flips green with no restart. (3) agent-plane poke listener (Direction-2a)internal/poke: contentless UDP poke bound to the box WG /32 (port 51822), leading-edge debounced, fires the hub-loop out-of-band trigger for an immediate desired-state cycle; enabled with wg_tunnel.enabled; first slice of R-13. Red-proofs for all three (run-fail-revert). ALL THREE LIVE LEGS PROVEN on the demo (2026-07-17), demo now LIVE on 0.89.0: Scenario 4 floor-driven A/B train 0.88→0.89 (operator signed+enqueued the agent_update op — the vouch+floor alone does NOT trigger it; committed, no rollback); Scenario 1 R-22 self-heal (marker aside + ACLs revoked → pre-check 403 … self-granting (R-22)converged state=adopted in ~3 s, ACLs restored, offsite active); Scenario 3 poke→tick ~31 ms ep0→box + immediate report cycle (save→tick ≈ ~0.45 s). Details: REPORT.md.

  • v0.88.0 (2026-07-13 eve) — controller-driven escrow ceremony (agent half), LIVE on demo host + drill VM (63/63 capabilities both). --output=json machine mode (text mode byte-identical; extraction into escrowCeremony()); the ONE fixed argv (escrow.CeremonyArgs() — shared by the localapi exec + the escrow-ceremony capability (Critical, pbs_dr-gated EXPLICIT) + the new FELHOM_ESCROW sudoers alias, three-way pin-tested); localapi job endpoints (POST /escrow/ceremony single-flight 60 s, status, ONE-SHOT claim → 410, 10-min TTL → unclaimed_void, GET /escrow/preflight). R in-memory ONLY (never the job struct — snapshot-hygiene-tested; restart loses it safely). Live-proven on drill endpoint-exact: stage → preflight all-green (live FELHOM_ESCROW list-probe) → job ~4 s → hub blob restic_pw_sha256 covering (repaired the spike's hash-less blob) → claim 200 once → 410. Coupled: controller v0.127.0 (MinAgent 0.88.0 for the wizard). OPEN: publish 0.88.0 + Day-0 manifest vouch (operator) at the next train; deployed hosts got direct deploys. Details: REPORT.md + felhom.eu RUNBOOK-escrow-ceremony.md (F1 threat model).

  • v0.87.0 (2026-07-13) — SystemDisks device-mapper walk (IA finding 2, MEDIUM): legacy-boot hosts get a working drive wizard. Operator ruling (approved 2026-07-13, verbatim): resolve device-mapper/raid parents — for the root filesystem's backing block device, walk /sys/block/<dev>/slaves recursively down to physical disks; those, plus any ESP holder when present, are system. Disks outside that set become wizard candidates (still subject to the existing data-bearing guards). The all-system fail-safe remains ONLY for walk failure — it returns to being the error case, not the legacy-boot common case. Implemented as physicalDisksOf/walkSlaves + HostReader.BlockSlaves (one seam method); per-branch conservatism (any unresolvable slave → ok=false → unchanged all-system path); signature test TestSystemDisks_WalkTopologies (root-backing disk ALWAYS system — never weaken). §3 spike transcripts: drill (legacy) dm-1→sda3→sda; felhom-pve (EFI+LVM) ESP+walk agree on sda → byte-identical regression. §13.2 wizard leg COMPLETE (offered → enrolled → formatted → torn down, boxes as found) + Day-0 manifest vouched to 0.87.0 (operator). The leg also surfaced two CONTROLLER bugs (fixed same-day: v0.126.3 claimed-box wizard CSRF, v0.126.4 502-through-CF + native-alert ban).

  • v0.83.0 (2026-07-11, LIVE on felhom-pve; NOT published — Peti stays 0.81.0) — observability pass (pairs with controller v0.116.1 + hub v0.46.0). applog.New(logger, *Ring): slog fan-out, journald at the configured level, ~1000-entry ring FIXED at DEBUG. GET /debug/logs (local API, token-authed; the controller Debug page's Ügynök tab) + request-level DEBUG middleware. Heartbeat log-pull: envelope log_tail_requested → next heartbeat ships log_tail (128 KB, consume-once; failed push re-armed by the next envelope; operator log pull served INFO on fulfillment). Gap-fill sweep: netverify phase/verdict lines (job start, trigger outcome, /proc/mounts verdict, journal bytes, classification code, rollback outcome, durations), netmount unit steps, signedjobs op-received (class/host/expiry — never signatures) + fetch duration, selfupdate invariants + download sha/duration, disks outcome INFOs, controller-swap pre-pull + health verdicts, desired/loop per-exchange DEBUG. Logging conventions: felhom.eu/documentation/runbooks/logging-conventions.md. OPEN: the hub-side live pull awaits the operator's button click (hub UI password-gated); pre-existing lanresolver permission-denied WARN on /var/lib/felhom-agent/guests noted in REPORT.

  • v0.77.0 (2026-07-09) — fork-4: escrow the offsite restic repo password under R. IdentityBundle gains ResticRepoPassword (rides the existing age-under-R WrapIdentityBundle path — validated by the custody spike febdc56). New POST /escrow/stage-secret (withGuest) transiently stages the controller-pushed password (0600, never logged), which the --selftest=escrow-create ceremony auto-injects into the bundle and then wipes. AttachResticPassword/StagedResticPasswordPath/ WipeStagedResticPassword added. Pairs with controller v0.105.0 (push + atomicity gate + DR inject + DRResticCoord). NOT yet live-validated — the supervised escrow ceremony is operator-run.

  • v0.76.0 (2026-07-08, LIVE on felhom-pve + PUBLISHED sha 9828c5f7…f50b — THE Day-0 manifest bump target; 0.75.0 superseded unpublished) — GL-5b / G12: restore-test full-fidelity. Params derive from the ARCHIVE's embedded config (drRestoreOverrides, same as DR — the old live-source-config path verified the wrong object AND dropped storage mpN per PVE's all-or-nothing rule; deleted with bindMountOverrides/archiveVMID). NEW mount-parity assert (restored mpN vs archive; miss/mispath/undersize/extra = FAIL naming the delta) + MountParity/ MountInventory on the wire record (additive). Live-proven: scratch 990000 ← 6.5GB 9201 archive, parity ok, inventory mp0 200G+mp1 50G+2 throwaways, 3m4s local tier (cheaper than feared); rotated-out archive volid → clean up-front refusal (nice failure mode). bringup.go untouched.

  • v0.75.0 (2026-07-08, LIVE on felhom-pve) — GL-5 / go-live G8: guest-loss DR bring-up actually restores (closes the v0.74.0 OPEN item + SPIKE-dr-bindmount-source §8). DR passes the COMPLETE explicit restore param set derived from the archive's embedded config (NEW Client.ExtractArchiveConfig, 200 under the scoped token) — two live-discovered PVE rules: mpN params need an explicit rootfs, AND unlisted mountpoints are silently DROPPED (first run booted without mp0/mp1!) — storage mpN passed through, structural mp8/mp9 → throwaways, then step 4d swaps the REAL binds in via the host runner (root pct; new EngineOptions.HostRunner+StateDir seam) and deletes the unusedN residue. Scratch-DR live-proven end-to-end (9310 from a real 9201 archive: mp0 200G + mp1 50G + real binds + no residue + clean teardown). Provision = nil overrides (regression-tested). NOTE: published/vouch-pending agent is 0.74.0 — publish 0.75.0 before/with the manifest bump. OBSERVATION: the DR selftest hardcodes KeepMAC=true — a scratch DR while the SOURCE guest is live briefly duplicates its MAC on the bridge (pre-existing; fine for supervised runs, worth a -keep-mac flag someday). Full customer-data DR drill = GL-6/S5 family.

  • 2026-07-07 — v0.74.0 Gitea-PUBLISHED (RUNBOOK GL-1) — the LIVE felhom-pve binary's exact bytes, sha256 1ec3f58842edce1e…76af05, anon-fetch-verified. This supersedes/closes every standing "publish 0.6x + Day-0 vouch" OPEN item below (0.64→0.73 were never published; 0.74.0 is the vouch target). Golden 0.103.0 published in the same run (felhom.eu execution record documentation/pilot/RUNBOOK-GL1-publish-2026-07-07.md). Day-0 manifest vouch = operator step (agent 0.74.0 / golden 0.103.0).

  • v0.74.0 (2026-07-07) — campaign-2 R2 CLOSED; the mislabelled "R1" was a symptom (LIVE on felhom-pve). Pool membership is what lets the pool-scoped token reach a guest; pct restore --pool sets it only at CREATE, so a restore-over-existing dropped 9201 from the felhom pool → no VM.Audit → restore-test's existing bindMountOverrides never ran → "mp8 … only possible for root". Fix: Client.PoolAddVMID + bring-up re-asserts membership post-restore (warn-not-fail). Role/ACL + bindMountOverrides untouched (both correct). Live restore-test PASSED for the first time once the pool was healed (Part A one-liner): read config → neutralize 2 binds → restore → boot+running → clean teardown, 4m35s. B3 (scratch-teardown 403) confirmed a cascade — no code. OPEN: DR bring-up -mode dr bind-override gap (spike SPIKE-dr-bindmount-source-2026-07-07.md: small known-constant override reusing bindMountOverrides; mp8/mp9 are structural constants).

  • v0.73.0 (2026-07-06) — F2 mount-role fallback CLOSED (LIVE on felhom-pve). roleForMountPath gained a mount-table fallback (Impl-2b style): a bind-mounted RAW enrolled user-data drive is not a PVE storage, so it fail-safe'd to system and the eject/decommission gates 403'd EVERY user-data drive (campaign F2, where=/mnt/teszt_enroll role=system). Device-keyed classification + whole-disk containment (storage.SameWholeDisk); Observe-error keeps the fail-safe BEFORE the fallback. Only roleForMountPath touched. Live-proven full lifecycle on teszt_enroll (eject/decommission 200, no-rebind across restart, end==pre). OPEN follow-up: the deviceRole/roleForMountPath unification refactor (deferred).

  • v0.72.0 (2026-07-05) — OOB operator access (merged E1+H1) — TASK H1, provenance both SPIKE-{felhom-sshd,oob-wg-operator-peer}-2026-07-05. Operator /32 RENDERED into wg-felhom AllowedIPs (survives self-heal, [OF-1]); dedicated internal/felhomsshd (port claim + config render→sshd -t→reload + operator authorized_keys + heal + oob heartbeat stanza); static inet felhom_oob belt (agent mutates SET ELEMENTS ONLY); configs/felhom-sshd.service (NO RuntimeDirectory [SF-1]) + felhom-oob.nft + felhom-op.sudoers; FELHOM_SSHD+FELHOM_OOB grants; oob.enabled DEFAULT FALSE. Live on felhom-pve (8822, belt filled, operator SSH as felhom-op with scoped sudo); hub v0.35.0. Rollback .bak-0.71.0. 5 live-found bugs fixed (port path, self-listen flip-flop, nil-block lockout, reachable-via-dial, operator-configured source).

  • v0.71.0 (2026-07-05) — management-plane break-glass: privsep-dir watchdog + mgmt_plane health — TASK G1 (prereq for felhom-sshd/H1), provenance SPIKE-felhom-sshd-2026-07-05 §8. Host artifacts (configs/felhom-privsep.tmpfiles + felhom-mgmt-watchdog.{sh,service,timer}) make /run/sshd boot-persistent AND auto-heal it every ~60s agent-independently (heals with the agent stopped — proven live: /run/sshd removed → restored in 30.0s, :22 back, no login). internal/mgmtplane reports the additive mgmt_plane heartbeat stanza; hub v0.34.1 raises mgmt_plane_healed. NO unit declares RuntimeDirectory= (the incident cause). H1 may now assume /run/sshd is guaranteed present. Live on felhom-pve; rollback .bak-0.70.0.

  • v0.70.0 (2026-07-05) — agent self-update (operator-signed A/B slots + crash-loop auto-rollback) — TASK D1, provenance SPIKE-agent-selfupdate-2026-07-05. An operator-signed agent_update op (version+sha256, sha is the only integrity root) rides the signed-jobs gate; internal/selfupdate.Executor downloads+verifies+hands to felhom-selfupdate-guarded apply (root re-verify → A/B atomic flip → pending marker → detached restart); the new binary commits after a 60s dwell; a crash-looping binary is auto-reverted by OnFailure=felhom-agent-rollback.service (first-crash trigger [SF-1]) with the tuned [Unit] start-limit (120s/4) as backstop. Host artifacts + sudoers FELHOM_SELFUPDATE + felhom-host-install.sh day-0 install + report field selfupdate_pending. Green tests + companions. LIVE-VALIDATED on felhom-pve (2026-07-05): all 4 drills PASS — happy path (0.70.0→0.70.1 signed op → download+verify+flip+commit), crash-rollback (0.70.2-crash → OnFailure → ~2s crash-to-recovered, byte-identical revert, no loop), no-pending guard, gate refusal (non-pinned key). Full agent-side pipeline ran real (envelope injected into the hub signed_jobs queue — CC lacks the hub global operator key; hub enqueue-auth is hub-unit-tested). Box restored to canonical v0.70.0 (host artifacts KEPT installed; scratch operator key REMOVED — self-update dormant until an operator pins a real key, a Day-0-vouch-style follow-up). Rollback felhom-agent.bak-0.69.0. OPEN (v1 scope-outs): no hub-floor auto-update, no failed-update auto-retry, no pending-timeout auto-rollback; per-crash OnFailure can double-fire (idempotent — future: serialize the rollback oneshot). Detail: REPORT.md.

  • v0.69.0 (2026-07-04, live on felhom-pve) — S5: host-loss DR — safe halves shipped. Part 1 wgtunnel.InstallRecoveredKey — writes an escrow-recovered WG privkey (create-only, refuse-overwrite) so the tunnel re-establishes with the SAME identity/pubkey (same /32), no keygen; wired into --selftest=identity-consume -install-wg-key (opt-in; pre-S3 blob → logged fresh-keygen fallback). Part 2 new internal/dr — consumes the host_loss restore_directive (was logged-ignored) into an inspectable RestorePlan via AddConsumer: per-guest {vmid,archive,target, sizing} + per-drive {durable_id→mount} + offsite PBS coord; DERIVE-AND-SURFACE only (Consumer has no restore/destroy dep — execute-nothing is structural). Tests + red-proofs (WG create-only; plan mode-gate). Part 3 hub escrow-GET NOT needed (operator exports the blob via sqlite3 writefile on a cp'd hub.db). Part 4-A re-attach wrong-disk safety already unit-proven (ResolveStorageDevice: match resolves, absent/mismatch ERRORS, non-uuid scheme refused — never a near disk). Part 4-B (destructive in-place 9201 restore) PREPARED + OPERATOR-GATED, NOT executed — pre-flight green (offsite ct/9201 restorable per S4.1); the operator runs the R-consume steps + confirms the destroy (§9-4a: CC never runs a consume/R command — see operator-present-one-time-secrets). OPEN: the operator-run 4-B drill; guest_loss DR; hub-driven full-auto DR. Rollback felhom-agent.bak-0.68.0. Detail: REPORT.md + doc-06 §3.5/S5.

  • v0.68.0 (2026-07-04, live on felhom-pve) — S4.1: unattended offsite restore-test. Tier-aware restore-task deadline: RestoreTestSpec.RestoreTaskTimeout (0→10m default) from config.RestoreTestPBSRestoreTimeoutSeconds (accessor default 120m), set only when SourceTier=="pbs" (main.restoreTaskTimeout); local tier UNCHANGED. Fixes the WAN restore being killed at 10m → mid-restore teardown → leaked scratch. Teardown "VM.Allocate" follow-up = PHANTOM (diagnosed, not blind-fixed): ran the restore-test on the AGENT-TOKEN path sourcing the offsite (pbs) backup → pass:true verified:boot+running, teardown succeeded (torn down vmid=990000, no 403), scratch band clean. The earlier 403 was the 10m-timeout consequence (guest not yet pool-associated); the scratch is restored INTO /pool/felhom (ACL already grants VM.Allocate) so teardown is authorized once the restore completes. No ACL/host-install change. OPEN: publish 0.68.0 + Day-0 vouch; Tier-1/Tier-2 split for offsite-as-default; S5 DR consume. Rollback felhom-agent.bak-0.67.0. Detail: REPORT.md.

  • v0.66.0 + v0.67.0 (2026-07-04, live on felhom-pve) — S4: PBS over the tunnel. v0.66.0: wgtunnel v4-pin (renderConf writes the resolved A LITERAL, never DNS/AAAA; Resolver seam, lowest addr; cached → steady-state zero-DNS/zero-exec) + re-resolve watchdog (Manager.Watchdog, loop-only; handshake stale > stale_after_seconds=180 → re-resolve → IP-changed re-render+restart)

    • FELHOM_WG Critical flips (conf-install/enable/restart/handshake-read). v0.67.0: namespace-aware PBS client (Config.Namespace → Snapshots ?ns=, Verify ns=; root-ns unchanged) — the operator-approved fix after Phase-1 showed the ns-unaware datastore-root 403s a per-tenant token. Live Scenario-D (all green): real vzdump of 9201 → ciphertext in ns demo-felhom-01 over the tunnel; ns-scoped verify=ok under the box's own felhom@pbs!demo-felhom-01 token; WARN gone; restore round-tripped (decrypt with box-born key → boot → teardown). Confirmed tenant ACL (felhom-hetzner): DatastoreBackup on /datastore/felhom-offsite/<ns> (NOT /ns/<ns>) to BOTH user felhom@pbs AND token (privsep=intersection; cross-ns 403); DatastoreBackup can't prune (safety). FINDINGS: retarget field is local_backup_target (not backup_target); retarget REVERTED to local (controller backs up ~every 30 min → single-target offsite = near-continuous 20-min uploads; needs Tier-1/Tier-2 split); restore-test scheduler needs a WAN restore deadline + scratch-band VM.Allocate before it runs offsite unattended. OPEN: escrow-create (OPERATOR-PRESENT, new R); publish 0.66/0.67 + Day-0 vouch; S5 DR consume. Rollback: felhom-agent.bak-0.65.0/.bak-0.66.0. Detail: REPORT.md + doc-06 §3.4/§4.2 + runbook §4a/§4b.
  • v0.65.0 (2026-07-04, live on felhom-pve) — S3.1 offsite-tunnel client MTU 1420 → 1280: resolves 06 §4.3's OPEN DECISION left by the CGNAT smoke test. 1420 silently black-holed bulk TCP on sub-~1480 paths (mobile ~1400, DS-Lite ~1452) — handshake+ping healthy, PBS TLS page (and at S4 the backup itself) drops. New const clientMTU = 1280 (RFC 8200 IPv6-minimum floor; outer 1340 v4 / 1360 v6 fits every realistic path), permanent + fleet-wide + family-agnostic. Client-only by construction — interface MTU caps box→PBS, advertised MSS caps PBS→box, so the endpoint's wg0 is untouched (zero live-endpoint risk). Golden pins exact MTU = 1280 (red-proofed vs a 1420 flip); no wire/JSON change. Live: agent re-rendered on restart (hash-gated apply), conf + live iface both 1280, PBS page loads at 1280 (no regression on wired). OPEN: true-CGNAT-SIM retest (low risk); publish 0.65.0 + Day-0 vouch (operator); S4 PBS-over-tunnel. Rollback: felhom-agent.bak-0.64.0 on the box. The v4-pin (§4.2 determinism) is a separate, optional future note — NOT needed for MTU correctness.

  • v0.64.0 (2026-07-04, live on felhom-pve) — S3 offsite WG tunnel: new internal/wgtunnel (keygen 0600/0700, marker-gated one-shot registration, agent-managed wg-quick@wg-felhom from the hub's desired-state wireguard block via the new desired.Syncer.AddConsumer seam, revoked-stays-revoked teardown, report stanza) + FELHOM_WG sudoers/capabilities + IdentityBundle.WGPrivateKey escrow auto-inject. wg_tunnel.enabled DEFAULTS FALSE (safety gate — rollout to Peti's box is a no-op until the production endpoint exists; enabled explicitly on felhom-pve only). Live: tunnel to ep0.felhom.eu:443 up 3 s after enable (PBS page through 10.77.0.1:8007), reboot-persistent, revocation drill clean, 30-min keepalive soak. GOTCHAS: hub envelope poll_interval_seconds (hub-side const 900 s) silently overrides agent poll_seconds on cycle 1; wg show <if> dump leaks the PRIVATE key (forbidden everywhere — sudoers only grants latest-handshakes). OPEN: CGNAT/mobile-hotspot smoke (operator-assisted appendix); publish 0.64.0 to Gitea + Day-0 vouch (operator); S4 points PBS at the tunnel.

  • configs: build-golden.sh v2.0.0 (2026-07-03, @ ceca355; no agent version change) — drill findings B5 + B1 FIXED (DRILL-golden-098-2026-07-03.md): the controller tag is a MANDATORY argument (the default rotted twice — a fresh install booted a pre-floor controller, forcing the guide's manual D.1b update) and the golden now bakes a felhom-controller-bootstrap.path unit (controller deploys the moment the back-half hot-plugs the bootstrap mount — no reboot; installer v1.9.1's reboot is a redundant belt, kept). Golden 0.98.3 baked on the drill VM, clean-room validated (bake integrity → isolated hot-plug proof → local-golden Day-0 → published-artifact Day-0), published (sha256 b9a02ef1…fd01) + operator-vouched — Day-0 manifest now vouches agent 0.63.0 + golden 0.98.3 (the v0.63.0 vouch follow-up below is DONE). Fresh installs land current and self-manage. NEW operator follow-up (SECURITY): the customer-config git.token has Gitea package-WRITE rights — scope down + rotate (evidence-doc observation O1).

  • v0.63.0 (2026-07-03, live on felhom-pve + Gitea-published sha256 b4a89c81…) — drill findings B3 + B2 FIXED (DRILL-day0-cleanroom-2026-07-03.md): TokenStore.Lookup reloads the append-only store once on a miss (cross-process coherence with the one-shot provisioner — no more fresh-install /controller/swap 401 / manual restart; size short-circuit bounds the cost; behind the TokenAuthority seam) + guesthook.InstallSnippet issues a fenced mkdir -p /var/lib/vz/snippets first (fresh boxes lacked the dir → the self-heal hook silently never installed). Sudoers gained exactly that one grant — ship sudoers WITH the binary (done on felhom-pve). Red-proofed both; Scenario-E method: compiled test suite run ON felhom-pve + live channel-health hit-path. OPERATOR FOLLOW-UP: bump the hub Day-0 manifest to agent 0.63.0 — until then fresh installs get 0.62.0 and the guide's D.1b restart-first step still applies (narrowed to "< v0.63.0" in the guide).

  • v0.62.0 (2026-07-03) — audit A1 RESOLVED: the stale-lock reaper's scan is now pool-intersected (staleLockController.Guests() = ListLXCClient.Pool("felhom") members), fail-safe skip on pool-read failure; pve:pool-read capability (non-critical) + --selftest "pool read" line. Companion host-install v1.9.0 adds Pool.Audit to FelhomAgentGuestdeploy order on any box: rescope ACL first, then this agent. Per SPIKE-a1-pool-membership-read-2026-07-03.md; red-proofed tests in stalelock_pool_test.go.

  • 2026-07-03 — CLAUDE.md refreshed: version narrative removed (state lives HERE + CHANGELOG top), layout completed (all 17 internal packages + cmd/felhom-opsign); deploy runbook now in the felhom-build-deploy skill (felhom.eu/skills/).

  • 2026-07-03 — REUSE.md exists at the repo root (canonical helpers / format-safety guards / traps / seams, code-verified); maintenance rule active: update it in the same commit that changes a shared helper.

  • v0.61.0 (2026-07-03) — blast-radius audit fixes B1 + D1 + D2 + D3 from felhom.eu/documentation/audits/AUDIT-blast-radius-hostroot-localapi-2026-07-02.md: random temp staging for root-installed scripts (+ sudoers/manifest glob updates), mkfs-wrapper member/RO re-checks (validated by scripts/mkfs-guarded-harness.sh), classifyClaim empty-lsblk fail-safe, and the blank-format anti-retarget (durable-id-bound, AGENT-001's benign-branch twin).

  • Deployed on demo host felhom-pve (node demo-felhom), non-root felhom-agent service user, pool-scoped token (felhom pool).

Open threads

  • Deferred audit items (housekeeping/design, all INFO): C1 (controller-swap version floor), C2 (NAS server allowlist), A2 (gate journal cross-check), B2B5, E1/E2.
  • Drive-enrollment leftovers: (a) runStorageInit slow-device detached-format polling; (b) Impl-3 shared-box operator format gate.
  • BUNDLE leftover: non-root agent can't read the PBS key; migration must preserve cert/key/tokens.
  • Not run (needs a supervised session): the destructive D1/D3 live proofs (real mkfs on a crafted member; a live /dev re-enumeration race during a real format).