Files
felhom-agent/CHANGELOG.md
T
admin aaa276a7b9 build-golden.sh: default controller image → current (0.85.1); golden rebuilt
The CONTROLLER_IMAGE default (arg 6) was a stale :0.43.0, so an argument-less
golden build baked an ancient controller (fresh Day-0 boxes started at 0.77).
Bumped the default to the current :0.85.1; always pass it explicitly per rebuild.
Golden rebuilt at 0.85.1 on felhom-pve (volid vzdump-lxc-9100-2026_06_27-11_42_51);
baked-image verify confirmed :0.85.1 in the build guest. No agent binary change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSZmmSFVzGwEzhYmxbkgBK
2026-06-27 11:56:40 +02:00

1718 lines
133 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Changelog
All notable changes to **felhom-agent** are recorded here. Update on every code
change that gets pushed.
## build-golden.sh — default controller image bumped to current; golden rebuilt at 0.85.1 (2026-06-27)
**Operational + a default fix (no agent binary change — version stays v0.42.0).**
- `configs/build-golden.sh`: the `CONTROLLER_IMAGE` default (positional arg 6) was a stale
`…/felhom-controller:0.43.0` — an argument-less golden build baked a wildly old controller, so fresh
Day-0 boxes started old (the demo started at 0.77). Bumped the default to the **current**
`…/felhom-controller:0.85.1` so the worst case (no explicit arg) is merely "current", not ancient.
- **Always pass the controller version explicitly at each rebuild** — this default only bounds the
worst case. A future `make golden` that resolves the latest pullable tag would remove the need for a
hand-bumped default (Observation, not this task).
- **Golden rebuilt at 0.85.1** on `felhom-pve` with the image passed **explicitly**
(`build-golden.sh 9100 … gitea.dooplex.hu/admin/felhom-controller:0.85.1`). New archive volid:
`local:backup/vzdump-lxc-9100-2026_06_27-11_42_51.tar.zst` (rootfs 32G + Docker-data 16G + user-data
8G, all in the archive; mp0+mp1 inclusion confirmed in the vzdump log).
- **Baked-image verify (cheap, mandatory):** in the build guest `/etc/felhom-controller-image` =
`…:0.85.1` and `docker images` showed it baked (379 MB). New Day-0 provisions now ship current.
- The host-bootstrap script auto-discovers the newest golden, so it picks up this rebuild
automatically. The real demo 9201 was **not** re-provisioned (it is the Phase-2 floor test box).
## v0.42.0 — agentic controller update: in-guest image swap + rollback (Phase 1) (2026-06-26)
The host agent now owns the in-guest controller image **swap** — the new-architecture replacement for
the controller's dead in-container `docker compose` self-update. The controller pre-pulls the target
image (shared docker socket, its own registry token) then asks the agent to swap; the agent — external
to the controller container, so it survives the controller being killed mid-swap — does the rest and
**rolls back** if the new controller doesn't come up healthy.
- **New local-API routes** (`internal/localapi/controllerswap.go`, token-scoped via `withGuest`):
- `POST /controller/swap {image}`**202** `{status:"swapping", previous_image, target_image}`, then
async: record previous (crash-safety state file `/var/lib/felhom-agent/controller-swap-<vmid>.json`)
→ confirm the target image is present in the guest (else abort, **no swap**) → write
`/etc/felhom-controller-image``systemctl restart felhom-controller-bootstrap.service` → poll the
new controller to **healthy** (`docker inspect`, ≤90s) → **roll back** to the previous image + restart
if it doesn't (the guest is never left without a controller). Single-flight per guest (409 if busy).
Image ref is strict-validated (`gitea.dooplex.hu/admin/felhom-controller:<semver>`) before any action.
- `GET /controller/swap/status``{state: swapping|done|failed, current, previous, target, error}`.
- **`GuestBinder.GuestExec`** (`internal/localapi/guestbind.go`): the one `pct exec` seam the swap
composes over (cat/inspect/write/restart), reusing the fenced root runner.
- **`--selftest=controller-swap -vmid -image <ref>`**: exercise the primitive directly (the target image
must already be pulled in the guest).
- Wired `ControllerSwap: guestBinder` into the local-API server (`cmd/felhom-agent/main.go`).
- Tests (`controllerswap_test.go`): happy swap, **rollback-on-unhealthy** (+ companion red-proof:
dropping the rollback leaves the guest on the bad image and fails the test), image-absent no-swap,
no-healthcheck-running, bad-image 400, single-flight 409.
## v0.41.0 — provisioned customer guests auto-start after a host reboot (`onboot:1`) (2026-06-24)
**F3 fix.** The provision back-half now sets **`onboot:1`** on the customer guest, so after a host
reboot/power-cut the customer's whole home-server (controller + apps) comes back **on its own**
previously every provisioned guest inherited the golden's `--onboot 0` and stayed **stopped** until a
manual `pct start` (confirmed live in the stable-path/sys-drive restart campaign, Phase 4.1). The new
step is a fatal `pct set <vmid> -onboot 1` placed right after the config-mount attach (`backhalf.go`),
mirroring the config-mount/parent-bind `pct set` ops. **No `startup`/boot-order/delay** — the v0.75
mountpoint-gate already covers the drive-bind race at boot (Phase 4.4), so the controller won't write
app data onto the rootfs while the agent re-binds drives.
The **golden stays `onboot:0`** (`build-golden.sh` unchanged): a template must not auto-start, and
`onboot` is a per-guest property the back-half is the right place to set. Unit-tested
(`TestProvision_SetsOnbootOne` asserts the exact `pct set … -onboot 1` invocation, with a red-proof
against removing the call). The pre-existing demo guest 9201 (provisioned pre-fix) was remediated
non-destructively with `pct set 9201 -onboot 1`. **Capstone live-validated (2026-06-24):** destroyed +
re-provisioned 9201 through the real provision chain with v0.41.0 → fresh `pct config` showed `onboot: 1`
with no manual set; a subsequent **felhom-pve host reboot** brought 9201 back **running with no manual
`pct start`** (the `onboot:0` scratch guests correctly stayed stopped), controller + base infra healthy,
drives re-bound at stable, sys_drive separate — the exact Phase-4.1 failure now passes.
## v0.40.0 — third CT volume: SSD user-data (`/mnt/sys_drive`, mp1) baked + `-sysdata-grow` (2026-06-23)
**The third golden volume.** Extends the OS/Docker-data split (v0.29.x) to a **three-volume layout**:
rootfs + Docker-data (`mp0`) + **SSD user-data (`mp1` @ `/mnt/sys_drive`, `backup=1`)** — the
controller's `system_data_path`. Until now `/mnt/sys_drive` was a plain directory on the 32 GB OS
rootfs, so the controller correctly warned that SSD app data (`<sys_drive>/felhom-data`) lands on the
OS drive. Baking it as its own thin volume clears that warning with **zero controller change** (the
controller already auto-discovers `<sys_drive>/felhom-data` and warns via `system.IsMountPoint`); the
`mp` under the guest's `/mnt` reaches the controller container through the existing
`-v /mnt:/mnt:rslave` bind.
- **`configs/build-golden.sh`** — `pct create` gains
`--mp1 ${ROOTFS_STORAGE}:${GOLDEN_SYSDATA_GB},mp=/mnt/sys_drive,backup=1` (new env
`GOLDEN_SYSDATA_GB=8`, near-empty; provision grows it). The resilience guards are mirrored for `mp1`:
a `findmnt /mnt/sys_drive` separate-mount assertion, and the vzdump-inclusion guard now aborts if
**either** `mp0` **or** `mp1` is EXCLUDED (the B3 trap — extra mountpoints default `backup=0`). The
golden does NOT pre-create `felhom-data`; the controller does once it's a real mountpoint.
- **`internal/reconcile/bringup.go`** — `const DefaultSysDataMount = "mp1"`; `BringUpSpec` gains
`SysDataGrowGB int` + `SysDataMount string`; a new **"4c"** grow block (online, grow-only `ResizeLXC`,
its own task) mirrors the "4b" Docker-data grow. `0 = skip` (separateness comes from the golden, not
the grow — the warning clears regardless of size).
- **`cmd/felhom-agent/main.go`** — `-sysdata-grow` / `-sysdata-mount` flags (mirror
`-datavol-grow`/`-datavol-mount`); `bringUpSizing` carries them into all three bring-up/provision call
sites; `--selftest=provision` help text updated.
- **Static volume, NOT an enrolled drive.** `/mnt/sys_drive` is part of the baked golden layout; it
never enrolls/ejects/decommissions and is deliberately kept off the drive-intent machinery.
`freeMountSlot` auto-skips the baked `mp0`/`mp1` so enrolled drives never collide.
- Tests: `TestRunBringUp_StorageSplit_SysDataGrow` (asserts `ResizeLXC(vmid,"mp1","+42G")`) +
`…_SysDataGrowZeroNoResize` (0 → no mp1 resize). RUNBOOK-provisioning-storage.md extended to the
three-volume layout (default ~512 GB SSD: 32 rootfs + 200 docker-data + 50 user-data).
## v0.39.0 — DR recipe completion: live PBS coord + drop the two unfillable drive fields (2026-06-16)
**DR-recipe agent-half completion.** A live eyeball of the demo recipe (v0.38.0) found three host-half
problems; all three are resolved here. No behavior change outside the recipe path.
- **PBS coord now resolved LIVE each collect.** New `internal/pbs/live_reporter.go`
`LiveSnapshotReporter` implements `hub.PBSReporter` by doing the cheap `Client.Snapshots()` list
itself, with **last-known-good fallback**, instead of reading only the verify-loop's `SnapshotStore`.
Previously the recipe's `pbs` block was omitted whenever the store was empty — which a one-shot
collect (`--selftest=hub`) and the first ~6 h window of every daemon after a restart always saw (the
verify loop populates the store on its own 6 h cadence). The restore SOURCE must not depend on a
maintenance cadence. Per-datastore: a live error/timeout → that datastore's last-known-good; a
successful (even empty) response is authoritative and updates the shared store. Targets-resolution
failure → the full LKG aggregate. Bounded by `DefaultLiveSnapshotTimeout` (8 s) so a hung PBS never
stalls the heartbeat. List only — it never triggers a `Verify`. The verify loop keeps Recording into
the SAME store (shared last-known-good); both use one hoisted `pbsTargets` closure.
- `SnapshotStore.Get(datastore)` added (per-datastore LKG copy) — the only `SnapshotStore` change.
- Wired into the collector in BOTH `runDaemon` and `runSelftestHub` (the selftest built its own
collector with a `nil` reporter — that is why the live `--selftest=hub` showed `pbs_snapshots:[]`).
- Intended side effect: `report.pbs_snapshots` is now live too (fresher hub PBS view).
- **`drives[].role` DROPPED from the v1 host-half shape.** A drive's purpose is a hub/operator-owned
manifest concept, not cleanly derivable host-side (both demo externals are `content=backup`, yet one
is the primary data drive and the other holds no apps). Deferred until the hub/operator stamps it.
- **`drives[].restic_repo_coord` DROPPED from the v1 host-half shape.** It named a backup tier that does
not exist — cross-drive backup is rsync to the SAME internal SSD; there is no offsite/second-failure-
domain bulk copy. RESERVED for a future tier (see the BACKLOG note in REPORT). v1 drive shape is now
`{durable_id, mount_path, intent, fs_type?, total_bytes}` — identifiers/intent/size only.
- The hub reads drives as `json.RawMessage`, so dropping fields needs NO hub struct change — only
golden + test sync. Cross-repo golden (`host-report.golden.json` here + the hub's copy) re-pinned and
verified **byte-identical** (sha256 `57f2a5e7…18b2f2b5` — manual checksum-diff discipline): the hub copy
previously lacked the `dr_recipe` section entirely; it is now a verbatim copy of the agent golden.
- Tests: new `internal/pbs/live_reporter_test.go` (T1 coord-present-without-prior-verify [load-bearing] +
inline bare-store companion, T2 error→LKG fallback, T3 success-warms-store, T4 targets-error→aggregate,
T5 bounded-by-timeout, T6 empty-success-authoritative); `TestDRRecipeHostHalf_V1DriveShape` (drive
object carries neither `role` nor `restic_repo_coord`); `TestBuildDRRecipeHostHalf` /
`TestHostReport_ContractMatchesGolden` updated to the v1 drive shape. Each companion was demonstrated
to FAIL on the pre-fix/mutated code, then reverted (see REPORT).
## v0.38.0 — DR recipe: emit the secret-free storage/guest/PBS half in the host-report (2026-06-16)
**DR recipe slice (agent half).** Additive `dr_recipe` section on the host-report — the agent half of the
secret-free reconstruction recipe (`SPIKE-dr-recipe-2026-06-16.md`) that complements escrow (keys) +
PBS/restic (bytes): the non-secret SCAFFOLDING an operator must rebuild before the PBS bytes can land.
The hub assembles it with the controller's app half into one customer recipe.
- `internal/hub/dr_recipe.go``DRRecipeHostHalf{recipe_version, guests[], pbs, drives[], pve_storage[]}`
built by the pure `BuildDRRecipeHostHalf(guests, targets, pbs)` from facts the report ALREADY collects
(no new privileged reads): `guests[]` = each guest's sizing (`GuestSpec`, skip status-unknown);
`drives[]` = the user-data external drives (usb/local-dir with a `uuid:` durable-id + mount path) with
`{durable_id, role, mount_path, intent, total_bytes}`; `pve_storage[]` = every storage target
`{name, type, content}` (the `storage.cfg` scaffolding); `pbs` = the latest snapshot's coordinates
`{repo_id (the pbs storage id), namespace, latest_snapshot_id}`. Wired into `Collect()` after the facts
are gathered; `HostReport.DRRecipe` (always set, never null).
- **BOUNDARY (the Phase-1 lesson):** every field is an identifier / intent / size / coordinate — NEVER a
key, password, token, hash, or `ENC:` value. The PBS encryption key stays in escrow; the access token in
identity-escrow; the restic password in escrow — the recipe names only the `repo_id`/`namespace`/
`durable_id`/`restic_repo_coord` the restore TARGETS. `recipe_version=1`; read is ignore-unknown
(forward-compat). The wire shape is pinned in the cross-repo golden (`host-report.golden.json` here +
the hub's copy — keep them byte-identical; manual checksum-diff on any change).
- Tests: `TestBuildDRRecipeHostHalf` (drives = only user-data; pve_storage = all; pbs = latest; guests
skip nil-spec), `..._NoPBS` (omitted, non-nil slices), `TestDRRecipeHostHalf_NoSecrets` (the lighter
boundary mirror — serialized half carries NO credential-shaped key; the load-bearing version is on the
controller emitter), and the `dr_recipe` key-set added to `TestHostReport_ContractMatchesGolden`.
## v0.37.0 — host-reboot remount re-resolves enrolled drives by filesystem UUID (2026-06-16)
**TASK A — close out the reboot story (agent half).** On a host reboot the kernel can re-enumerate block
devices and move a drive's node (felhom-usb `/dev/sdb``/dev/sdc`), and a `.mount` unit left `disabled`
by a prior detach never auto-mounts at boot — so an enrolled drive could stay unmounted (or, with any
node-trusting remount, mount the WRONG device). Root cause pinned LIVE: felhom-usb's systemd mount unit
was `disabled` (no `multi-user.target.wants` symlink) while felhom-flash's was `enabled`; `What=` was
already correct (by-UUID), but nothing re-asserted the unit at startup.
- `storage.ResolveStorageDevice(durableID)` — resolves the enrolled `uuid:<fs-uuid>` storage scheme to its
CURRENT backing `/dev` node by re-scanning `/dev/disk/by-uuid` (never a cached node); errors if the UUID
is genuinely absent so a caller skips a gone drive instead of fail-mounting a stale node.
- `storage.parseFelhomMountUnit` — pure inverse of `renderMountUnit` (Name/UUID/Where/Type/Options) keyed
on a `Managed by felhom-agent` marker; ignores any foreign `.mount` unit.
- `(*SudoHostOps).ReassertEnrolledMounts(ctx)` — at startup (BEFORE binding into the guest) and on the
periodic 20s tick: for each enrolled `.mount` unit, re-resolve by UUID and re-run `EnsureMount`
(idempotent `systemctl enable --now`) — re-enables a disabled unit AND mounts the CURRENT device by
UUID, so a `/dev/sdX` reshuffle is a no-op. Skips ONLY the durable steady state (mounted AND enabled),
via the pure `shouldReassertMount`; a **mounted-but-DISABLED** unit (the exact live felhom-usb bug — it
serves now but a reboot would not auto-mount it) is still re-asserted to re-create the wants-symlink.
Enabled-state is read with a privilege-free `os.Lstat` of the `multi-user.target.wants` symlink
(`unitEnabled`) — no `systemctl is-enabled` subprocess, no new sudoers entry. An absent UUID is skipped
(re-asserts on a later tick).
- Wired in `main.go` ahead of `ReassertGuestBinds` so mounts are live before the guest binds re-assert.
- Tests (Linux, seam the device-resolution): `TestResolveStorageDevice_ToleratesDeviceLetterMove`
(UUID symlink moved sdb→sdc → resolves sdc; companion asserts the cached enroll-time node differs from
the freshly-resolved one — a node-based remount would target the wrong device), `..._AbsentAndScheme`
(absent UUID errors; only the `uuid:` scheme resolvable), `TestParseFelhomMountUnit` (render→parse
round-trip + rejects a foreign unit), `TestShouldReassertMount` (the four mounted/enabled combos — pins
the mounted-but-disabled re-assert), `TestUnitEnabled` (wants-symlink detection).
**TASK A2 — verdict: enrolling a NEW drive does NOT need an LXC restart.** The enroll path lands on the
live intermediary-mount `AttachDrive` (`/disks/guest-attach``handleDiskGuestAttach``AttachDrive`,
"no pct, no reboot") under the single shared parent — unbounded named live slots — NOT the legacy
`RebootGuest` branch. The operator's pre-created-slot-pool idea is therefore unnecessary.
## v0.36.7 — isolate the shared parent only on CREATE (no peer-group churn) (2026-06-15)
Follow-up to v0.36.6: make-private+make-shared must run ONLY when the self-bind is first created, not on
every reconcile — re-doing it churns the peer-group id and ORPHANS the guest`s already-established slave
(propagation silently dies, guest sees empty). Guarded on the mountpoint check; on a fresh boot it runs
once before pve-guests so the guest slaves the right group.
## v0.36.6 — shared parent gets its OWN peer group (make-private first) — ROOT CAUSE of double-bind (2026-06-15)
The shared-parent self-bind INHERITED the root mount`s shared peer group (`/mnt/felhom-drives` was
`shared:1` same as `/`), so every drive bind under it propagated back via the root peer and DOUBLED
(2 stacked binds per drive — the real cause behind v0.36.3-.5). EnsureSharedParent + the boot script now
`make-private` (detach from the root group) BEFORE `make-shared` (own group whose only slave is the
guest), so a drive bind propagates to the guest exactly once.
## v0.36.5 — AttachDrive normalizes to exactly one bind (2026-06-15)
AttachDrive now COUNTS the binds at a stable path (countHostMounts) and normalizes to exactly one: it is
a no-op only when there is exactly ONE bind the guest sees; otherwise it strips ALL existing binds
(bounded loop) and lays down one fresh bind. This converges a stacked double-bind to one — the old
umount-one+mount-one force-rebind never did. Caught when a double-bind survived a guest reboot.
## v0.36.4 — serialize AttachDrive/DetachDrive (no double-bind race) (2026-06-15)
A mutex on GuestBinder serializes AttachDrive/DetachDrive so a controller-triggered reconnect and the
agent`s periodic reconcile can no longer both pass the isHostMountpoint check and double-bind the same
stable path (a TOCTOU race observed live as 2 stacked binds during rapid eject/reconnect).
## v0.36.3 — DetachDrive loop-umounts stacked binds (2026-06-15)
DetachDrive now umounts ALL stacked binds at a stable path (bounded loop), not just one layer — so an
eject/detach fully detaches even if more than one bind accumulated (operator bind on top, or a rare
attach race), keeping the fail-close intact. Caught in the E13 rapid eject/reconnect sweep.
## v0.36.2 — eject also keeps the raw mounted (reconnectable) (2026-06-15)
Extends v0.36.1 to EJECT: eject now DetachDrive`s the bind under the parent but LEAVES the raw
/mnt/<name> mounted (consistent with decommission), so the H1 disconnect→reconnect roundtrip re-binds
cleanly on a non-removable drive. Physical removal is the separate "remove from system" action. Test:
eject calls DetachDrive + does NOT unmount the raw.
## v0.36.1 — decommission keeps the raw mounted (re-enrollable) (2026-06-15)
Fix caught in the E10 acceptance test: the self-serve decommission unmounted the RAW /mnt/<name> host
mount, which orphaned a non-removable drive (no re-plug) so a one-click re-enroll bound an empty dir. On
the intermediary model decommission is now a LOGICAL retire — it DetachDrive`s the bind under the parent
(drive invisible to the guest) but LEAVES the raw mounted, so re-enroll re-binds cleanly. Physical
removal stays the separate "remove from system" action. Test updated.
## v0.36.0 — guest boot-id on /disks (deterministic guest-reboot recreate) (2026-06-15)
The agent now emits `guest_boot_id` on GET /disks: `<host-btime>-<guest-init-starttime>` — changes on
every guest boot (host reboot OR guest reboot) but is STABLE across a controller-only restart. The
controller persists the last-seen value and DETERMINISTICALLY recreates drive-backed apps when it
changes (replacing the fragile timed state-sample that could miss an app stopped at the sample instant).
`GuestBootID` reads `/proc/stat` btime + field 22 of `/proc/<init-pid>/stat` (parsed after the last
`)` so a comm with spaces/parens does not break it).
## v0.35.1 — shared-parent unit: run before pve-guests on host boot (2026-06-15)
Fix for the host-reboot ordering (the shared-parent oneshot never ran before pve-guests on the live
host, so the guest bound a not-yet-shared parent → private bind → propagation broken). The unit now uses
`WantedBy=pve-guests.service` (pve-guests PULLS IT IN + Before= orders it first) instead of the
unreliable `WantedBy=multi-user.target`, and drops `DefaultDependencies=no`. `EnsureSharedParent`
reinstalls the unit when its content differs (so the fix deploys on the next agent start/reconcile).
## v0.35.0 — intermediary mount: guest-reboot re-propagation (load-bearing) (2026-06-15)
Fix for the guest-reboot gap (caught in the live demo migration). A guest's parent bind is
NON-RECURSIVE, so on a guest reboot it does NOT carry the pre-existing drive submount, and mount
propagation only delivers events created AFTER the bind exists — so an enrolled drive is bound on the
HOST but INVISIBLE in the fresh guest namespace until re-bound. Without this, every guest reboot left
the apps on empty dirs.
- `AttachDrive` now takes `vmid` and checks GUEST visibility (`GuestSeesMount`, reading
`/proc/<guest-init-pid>/mountinfo`): if the host has the bind but the guest doesn't see it
(post-reboot), it FORCE re-binds (umount + mount) to re-fire propagation into the current guest ns.
- A periodic reconcile (20s ticker in main) re-runs `ReassertGuestBinds`, so a guest reboot self-heals
without an agent restart. `EnsureSharedParent` skips the unit re-install when already present (cheap
on repeat).
- `/disks` `BoundUnderParent` now reflects GUEST visibility (not the host mount) — the accurate signal
the controller's drive-absent gate keys on to stop/restart apps across a guest reboot.
## v0.34.0 — intermediary mount model: shared-parent + host-side attach/detach + reconcile (2026-06-15)
The drive hot-swap re-architecture (SPIKE-intermediary-mount). Replaces the per-drive `pct set -mpN`
bind (which needed a guest reboot to activate and bricked the guest when a drive was absent at boot)
with a SINGLE permanent parent bind `/mnt/felhom-drives` plus host-side swaps underneath it.
- `internal/localapi/intermediary.go` — `GuestBinder.EnsureSharedParent` (mkdir + self-bind +
`--make-shared` + installs/enables a `felhom-shared-parent.service` ordered **Before=pve-guests** so
the guest's parent bind inherits the shared peer group as `slave`); `AttachDrive` (`mount --bind
/mnt/<name>/felhom-data /mnt/felhom-drives/<name>` — propagates into the RUNNING guest live, no pct,
no reboot; confined to felhom-data; the stable dir stays host-root-owned = fail-closed); `DetachDrive`
(`umount`, leaving the bare fail-closed dir); `StablePathForRaw`/`DriveNameFromRaw`; `isHostMountpoint`.
- `ReassertGuestBinds` is now a pure HOST-SIDE reconcile: for each enrolled+present drive ensure its
felhom-data is bound under the parent (no guest-config read, no slot, no reboot) — fixes F9 and
drive-reconnect for free. Runs at startup (ensures the shared parent first).
- `handleDiskGuestAttach` uses `AttachDrive` (returns the stable `guest_path`); eject + decommission
call `DetachDrive`. Legacy `AttachBind`/`DetachBind` retained for the transition (decommission still
`--delete`s any lingering legacy mp).
- `/disks` reporting adds `GuestPath` (the stable `/mnt/felhom-drives/<name>` the controller repoints
HDD_PATH to) and `BoundUnderParent` (live-in-guest signal for the controller's drive-absent gate).
- Provision adds the one permanent parent bind (`-mp8 /mnt/felhom-drives,mp=/mnt/felhom-drives`).
Tests (non-hollow + companions): `TestGuestAttach_BindsUnderParent` (uses AttachDrive not legacy pct),
`TestReassertGuestBinds_RestoresMissingBind` (host-side reconcile, legacy AttachBind never called),
`TestStablePathForRaw_DriveName`, `TestDisks_GuestPathAndBoundUnderParent`. Sudoers: new
`FELHOM_INTERMEDIARY` alias (mount/umount under /mnt/felhom-drives, the unit install, the parent bind).
## v0.33.0 — C1 net: pre-start self-heal hook + decommission mp-delete (2026-06-15)
The transitional defense for the C1 brick (B3 critical bug) ahead of the intermediary-mount
re-architecture (which makes C1 structural). Two independent nets:
- **Pre-start self-heal hook** (`internal/guesthook`): a PVE `pre-start` hookscript runs
`felhom-agent guest-hook <vmid> <phase>` which, for every BIND mountpoint whose source path is
missing, creates an empty **host-root-owned** placeholder dir so the bind succeeds and the guest
always boots — fail-closed (host uid 0 is unmapped in the unprivileged-LXC userns, so the guest
can't write to the placeholder; a returning drive shadows it). It CREATES rather than DELETEs
because `pct set --delete` in pre-start would take the config lock the start task already holds
(dead-times-out → still bricks); the heal logic is in unit-tested Go, the wrapper just delegates.
Installed + registered per-guest by the provision back-half (`InstallSnippet`/`Register`).
- **Decommission mp-delete** (`GuestBinder.DetachBind` + `handleDiskDecommission`): decommission now
runs `pct set <vmid> --delete mpN` on the slot binding the drive (lock-safe on the running guest),
so its now-missing source can't brick the next reboot. The old handler unmounted but left the dead
`mpN` in config — the exact B3 C1 bug. Eject keeps its mp (temporary; the hook covers a
reboot-while-ejected).
Tests (non-hollow, each with a companion that fails the pre-fix/trivial impl):
`internal/guesthook/heal_test.go` (selector ignores storage volumes + present binds, heals only the
absent one; "return nothing"/"return all" both fail) and `TestDecommission_DeletesGuestMount`
(asserts the correct slot is `--delete`d; pre-fix never calls DetachBind → fails).
Sudoers: new `FELHOM_GUESTHOOK` alias (snippet install, `pct set --hookscript`, `pct set --delete mpN`).
## v0.32.0 — self-serve decommission + intent-aware re-assert (B2a) (2026-06-14)
Customer-self-serve storage decommission (no operator signature; non-destructive — never formats),
plus the load-bearing fix that keeps a decommissioned drive from auto-rebinding into the guest.
- **`POST /disks/decommission`** (`internal/localapi/disks.go` `handleDiskDecommission`, route in
`server.go`) — mirrors `handleDiskEject` exactly: `withGuest` self-scoping, `scopedFromBody`, and the
same **user-data role gate** (`roleForMountPath` must be `RoleUserData`, else 403; fail-safe-to-
protected on ambiguity) so a compromised controller can't decommission system/backup storage. It
records a PERMANENT `IntentDecommissioned`, prunes the `GuestBindStore` entry (hygiene), and unmounts
(so the drive is physically removable). It **NEVER** calls any format/mkfs path — the data stays on
the drive. The operator-signed `DecommissionExecutor` + `reconcile.Classify` classification are
untouched (the absent-drive/DR route).
- **`ReassertGuestBinds` is now intent-aware** (THE correctness fix): the startup re-assert skips any
durable-id whose intent is not `enrolled`, so a decommissioned- (or ejected-) but-still-present drive
is never auto-rebound into the guest on agent restart. A nil intent store falls back to legacy
bind-all (matching the watchdog's nil-intent rule). Covers both the self-serve and the operator-
signed decommission paths (both land on `IntentDecommissioned`).
- **`GuestBindStore.Remove(vmid, durableID)`** (`internal/localapi/guestbindstore.go`) — idempotent
(absent = no-op), atomic tmp+rename like `Record`; drops the vmid key when its set empties. Re-enroll
re-`Record`s via the existing `recordGuestBind` on guest-attach, so Remove doesn't break re-commission.
- `IntentRecorder` extended with `SetDecommissioned` + `Get` (both already on `*storage.IntentStore`).
- Non-hollow tests (`internal/localapi/decommission_test.go`): role-gate refuses system/backup (403,
no unmount); decommission sets intent + removes the bind + unmounts + never formats; intent-aware
re-assert does NOT rebind a decommissioned-but-present drive (companion: enrolled DOES rebind; the
intent-blind pre-fix code fails this); re-commission re-records; `Remove` idempotency + persistence.
## v0.31.0 — live-drive F9 + F20-BUG2 + F20-BUG3 (disk bind/wipe) (2026-06-14)
The last live-drive findings, all disk/`localapi`-side, implemented + deployed on `felhom-pve` and
validated live on guest 9201 (approach: attach-to-existing, no re-provision — see the audit fixspec).
- **F9 — guest data-drive bind survives a re-provision** (`4cd1d02`). The in-guest bind (`pct set -mpN`)
is config state a destroy+re-provision drops, and nothing restored it → a re-provisioned guest came up
with its enrolled HDD unattached. New `GuestBindStore` (durable-id-keyed, per guest, recorded at
guest-attach) + `ReassertGuestBinds` on agent startup re-adds any bind a guest is missing — only when
the durable-id still resolves to a present drive (a swapped/absent disk is never auto-bound), idempotent.
Plus `DiskInfo.GuestAttached` — the missing "bound into THIS guest" signal (vs mere host presence;
resolves the F2 `hdd_configured` disagreement). **Live-proven:** dropped the bind, restarted the agent
(real trigger) → re-attached with no manual call; reboot activated it; an HDD app then deployed onto
the drive with data on `/dev/sdb1`.
- **F20-BUG2 — one wipe durable-id scheme** (`a2a76e7`). `/disks` advertised only `durable_id` (`uuid:`,
used for assign), but the wipe gate resolves `byid:`/`byuuid:` → confirming a wipe with the advertised
id was a `binding_mismatch`. New `DiskInfo.WipeDurableID` via a shared `s.deviceDurableID` seam used by
BOTH the list and the gate, so the id the customer copies is the id the gate accepts. **Live-proven:** a
confirmed wipe using `/api/disks`'s `wipe_durable_id` is accepted (no mismatch).
- **F20-BUG3 — format runs detached; survives a request deadline AND an agent restart** (`4777f8a`). mkfs
ran under the HTTP request context, so a client deadline SIGKILLed it mid-write → corrupt disk. Now mkfs
runs off `s.baseCtx` via a persisted `formatJob` record; the handler still returns the synchronous
result (backward-compatible) but a dropped request no longer kills it. New `GET /disks/format/status`;
`RecoverFormatJob` on startup re-runs an interrupted durable-id-bound format (re-resolved; anti-retarget
— a blank/path-bound or unresolvable job is not auto-re-run). **Live-proven on the 916 GB felhom-usb:** a
2 s client timeout left a ~30 s mkfs running to a clean ext4 (the live-drive corruption is gone); an
agent restart mid-format was recovered + completed to a clean fs.
**Security fix (from the 2026-06-13 deep-sweep audit).** The inline customer-confirmed wipe in
`internal/localapi/disks.go` `handleDiskFormat` inspected and gate-bound the device by its durable id
but then ran `mkfs` on the caller-supplied mutable `/dev` path (`req.Device`). A USB re-enumeration
reassigning that `/dev` node to a different physical disk between inspection and `mkfs` (a
classify→mkfs TOCTOU) could wipe the wrong drive.
- New `internal/localapi/wipe_reresolve.go`: `antiRetargetResolve` (injected-deps, unit-tested) mirrors
`signedjobs.WipeExecutor.Execute` — resolve the confirmed durable id → current device, re-derive the
device's durable id and require an exact match, re-inspect (still data-bearing), and return the
re-resolved device. `(*Server).reresolveDurableForWipe` wires the real storage funcs.
- `handleDiskFormat` now formats the **re-resolved** device, never `req.Device`; any refusal →
`409 Conflict`, no `mkfs`. Injectable `reresolveWipe` seam on `Server` (defaults to the real path).
- Tests: `wipe_reresolve_test.go` covers happy-path, empty/gone/blank, re-inspect-error, and the core
`retarget-mismatch-refused` case. Round-trip safe for legitimate wipes (`DeviceDurableID` ↔
`ResolveDurableDevice` schemes match). Agent-only deploy; no golden rebake. See `AGENT-001-FIX-NOTES.md`.
## v0.29.1 — lanresolver: RESTART dnsmasq on change (not reload) — fixes stale split-horizon IP (2026-06-13)
**Bug:** after a guest's DHCP IP moved (e.g. the v0.29.0 9201 re-provision: .151 → .141), the LAN
split-horizon resolver kept answering the OLD IP, so LAN clients (via Pi-hole's conditional forward to
the host dnsmasq) resolved `*.demo-felhom.eu` to the dead IP. Root cause: `lanresolver.Manager` updated
the per-customer drop-in (`address=/<domain>/<ip>`) correctly but then ran `systemctl reload dnsmasq`
(SIGHUP) — and **dnsmasq's SIGHUP does NOT re-read its config files** (`/etc/dnsmasq.d/*.conf`); it only
clears the cache + re-reads `/etc/hosts`/addn-hosts. So the changed `address=` directive never took
effect until a restart. **Fix:** `reload()` → `restartDnsmasq()` (`systemctl restart dnsmasq`) for every
config-drop-in change (ReconcileGuest IP change, EnsureDnsmasq base change, Remove/decommission). Restart
is sub-second and the records carry local-ttl 0, so downstream forwarders don't cache a stale answer.
(Live: after the fix + a one-time host dnsmasq restart + a Pi-hole cache flush, `*.demo-felhom.eu`
resolves to the live guest IP again; future IP moves now self-heal on the loop's next tick.)
## v0.29.0 — OS / Docker-data storage split: golden + provision (2026-06-13)
Phase 1 of the storage-split slice (Phase 2 = felhom-controller v0.58.0 prevention layer). The
controller guest's OS rootfs and Docker data are carved onto separate `local-lvm` volumes for
RESILIENCE — an isolated OS rootfs stays bootable + agent-recoverable if the Docker volume fills.
- **`configs/build-golden.sh` — split baked in:** `--rootfs ${ROOTFS_STORAGE}:${OS_SIZE_GB}` (default
**32**, was hardcoded 8) **plus** `--mp0 ${ROOTFS_STORAGE}:${GOLDEN_DOCKER_GB},mp=/var/lib/docker,backup=1`
(default 16). The baked controller + infra images land on the data volume and travel inside the
golden archive (no empty-volume shadowing, no deploy-time pull). `backup=1` is MANDATORY — extra LXC
mountpoints default to `backup=0` = EXCLUDED from vzdump (spike B3), which would drop the images from
the archive entirely. The script now also bakes Docker **log rotation** into `daemon.json`
(`max-size 10m`, `max-file 3` — prevention layer 2D), asserts `/var/lib/docker` is a separate mount,
and **aborts if vzdump excludes mp0**.
- **`internal/reconcile/bringup.go` — sized provision:** `GuestMount` gains `Backup` (emits `,backup=1`
— closes the spike-B3/B5 silent-DB-loss trap at the mount builder). `BringUpSpec` gains
`DataVolGrowGB` + `DataVolMount` (default `mp0`): provision GROWS the golden-carried Docker-data
volume online to the per-customer target (grow-only, spike B4) rather than attaching a fresh empty
volume that would shadow the baked images. Plus `RootfsGrowGB` for the OS rootfs.
- **CLI seam:** `--selftest=bring-up|provision` gain `-rootfs-grow` / `-datavol-grow` / `-datavol-mount`
flags. Per-customer sizing source = flags now, the slice-10 hub storage manifest later.
- **`RUNBOOK-provisioning-storage.md`** (new): the split provisioning procedure + fresh-PVE-install
thin-pool carving knobs (`hdsize`/`maxroot`/`maxvz`, spike B4) + the per-customer sizing seam.
- Tests: `buildBringUpConfig` backup=1 emission; bring-up issues rootfs + data-volume resizes.
## (no version) — storage OS/data-split spike findings (2026-06-13)
Investigation only — **no code changed**. Findings report: `REPORT-storage-split-spike.md` (gates the
provisioning spec for splitting the controller guest's OS rootfs from its Docker/data onto separate
`local-lvm` volumes). Proven on a throwaway unprivileged LXC (9300, since destroyed): Docker `data-root`
on a second `local-lvm` mountpoint works (overlayfs/ext4, no idmap issue, reboot-survives); the
move-then-verify migration is safe (copy-not-move). **Key finding:** additional LXC mountpoints are
**excluded from vzdump by default** — they need `backup=1` set **and a CT restart** — so the docker-data
mount must be attached with `,backup=1` or named-volume DBs silently fall out of PBS. The exact seam is
`internal/reconcile/bringup.go:313` (`buildConfigParams`), which today builds `mpN` without a `backup=`
flag; `GuestMount` should carry the flag. Per-customer sizes belong in the slice-10 hub storage manifest
(marked at `bringup.go:49-50`); the golden rootfs is hardcoded `8` at `configs/build-golden.sh:40`.
## v0.28.0 — backup re-target → felhom-pbs (offsite DR) + operator-signed decommission (2026-06-12)
**Whole-guest backup now defaults to the offsite PBS tier (real DR).** `BackupConfig.BackupTarget()`
returns the configured `backup.local_backup_target` or, when empty, the new default `felhom-pbs` — a
PBS datastore on SEPARATE HARDWARE (the DooPlex box), so a host disk/hardware failure no longer takes
the backups with it. The target stays fully configurable (set `local_backup_target` to `local`/other
to override); no call site hardcodes it. All `NewBackupRunner` sites (restore-test scheduler, local-API,
`--selftest=backup`/`restore-test`) route through `BackupTarget()`.
Proven live on demo-felhom before the re-point (PHASE 0 gate):
- snapshot-mode `vzdump → felhom-pbs` still fires the `create storage snapshot 'vzdump'` marker, so the
8B.2 early-resume/quiesce signal survives a PBS target (the marker is mode-driven, not target-driven);
- the restore-test enumerates PBS backups through the SAME generic `StorageContent`
(`/nodes/<node>/storage/felhom-pbs/content` returns `content:"backup"` + ctime/vmid/volid), so
`PickRestoreCandidate`/`latestArchive` need NO PBS-client change;
- `pct restore` from a PBS volid round-trips cleanly (storage.cfg encryption key applied transparently);
- PBS gotchas (`ignore-verified`, node-from-UPID, privsep) touch only the verify-API path, not vzdump/restore.
**Operator-signed `decommission` now reachable (slice 10 P3 completion).** The previously-unreachable
`IntentDecommissioned` state (no production caller) is now reached ONLY via a gate-VERIFIED operator
signature — never customer-confirmable, distinct from a safe eject. New `internal/signedjobs`
`DecommissionExecutor` (op `decommission`, classified destructive in `reconcile.Classify`) calls
`IntentStore.SetDecommissioned`, keyed by the drive's STORAGE durable-id (the watchdog's key, e.g.
`uuid:<fs-uuid>` — NOT the device-level `byid:/byuuid:` scheme `storage_wipe` uses), so the recorded
intent actually gates future remounts. New `ExecutorChain` lets the signed-jobs runner serve both
`storage_wipe` and `decommission`; the runner wiring moved below the intent-store open in `main.go`.
`felhom-opsign` builds decommission params from `-durable-id`. No controller/customer UI — the operator
path is hub jobs-queue → signed-jobs runner.
**Restore-test now boot-verifies slice-10 enrolled guests (bind-mount mountpoints).** A guest whose
data drive is a host BIND mount (slice-10 P2 `mp0`) could not be vzrestore'd by the privsep token
("restoring 'mpN' to bind mount is only possible for root") — so the restore-test failed for every
enrolled guest, regardless of backup tier (surfaced during the felhom-pbs live validation). The
restore-test now reads the SOURCE guest config (vmid parsed from the archive volid — PBS `ct/<vmid>/`
and vzdump `vzdump-lxc-<vmid>-` forms) and passes `RestoreLXCOptions.MountOverrides` that neutralize
each bind-mount `mpN` to a throwaway 1G volume on the restore storage (needs no root; the boot-verify
doesn't need the drive's data, and the host paths would otherwise collide). Storage-backed mountpoints
are restored normally; best-effort (an unreadable source config restores as-is). `proxmox.RestoreLXC`
gained `MountOverrides`. Verified live: restore-test from felhom-pbs of bind-mounted guest 9201 →
boot+running PASS.
## v0.27.0 — slice 10 P3: self-heal watchdog reconcile + 4-state intent model (2026-06-12)
The storage watchdog goes from detect-only → detect-and-reconcile: the agent autonomously re-mounts an
enrolled external drive that dropped out-of-band (the colleague's Proxmox unmount), gated by a persisted
INTENT model so it never auto-adopts an unknown drive or fights an official eject.
- **`internal/storage/intent.go` — `IntentStore`** — durable, **durable-id-keyed** (UUID/WWN, never
sdX/path), atomic-write 4-state model: `new` (not recorded → never auto-mount), `enrolled` (desired
mounted → reconcile drift), `ejected` (intentional unmount → leave alone), `decommissioned`
(permanent). `OnAbsent` clears `ejected`→`enrolled` so a replug auto-mounts (the replug rule).
Records intent ONLY through the official enroll/eject paths — an out-of-band unmount records nothing
and is healed. Tests cover the states, persistence, the replug rule, and the reconcile gate.
- **`watchdog.go` — intent-gated reconcile + flapping guard (3C)** — the re-mount candidate (device
present, not mounted) now fires ONLY for an `enrolled` drive (via `IntentReader`); a present→absent
transition (device gone) calls `OnAbsent`. Exponential backoff (`debounce·2^fails`) + an alert after
4 failed cycles + a hard stop after 8 (no infinite loop). Failure = "still not present a full backoff
window after we dispatched" (a slow async re-mount isn't miscounted). Tests: colleague-unmount→
reconciled; ejected/new/decommissioned→left alone; ejected→absent→replug→auto-mount; flapping→caps.
- **`internal/localapi`** — `POST /disks/guest-attach` records `enrolled`; `POST /disks/eject` records
`ejected` (BEFORE unmount, while the durable-id still resolves) via the new `IntentRecorder`. `main.go`
opens one `IntentStore` (`<StateDir>/drive-intents.json`) shared by the watchdog + local API; open
failure degrades to ungated legacy remount (logged).
## v0.26.0 — slice 10 P2 activation: guest-reboot endpoint (user-triggered drive activation) (2026-06-12)
A drive enrolled into a RUNNING unprivileged guest can't be live-activated (proven: `pct set` won't
hot-apply; `/proc/<pid>/root` bind → mount-locking refusal; `nsenter -m` loses the host source). So the
bind activates at the next guest boot. This adds the user-triggered restart path.
- **`POST /guest/reboot` (`internal/localapi`)** — self-scoped (vmid from token). Runs `pct reboot
<vmid>` **detached** (it blocks ~30s until the guest is back) and returns **202** immediately, so the
calling controller gets a clean response before the reboot takes it down (the agent is host-side and
survives). `GuestBinder.RebootGuest` over the fenced runner. Tests: `TestGuestReboot_Accepted`
(202 + RebootGuest invoked for the token's vmid), `TestGuestReboot_CrossGuest403` (body vmid mismatch
refused, no reboot). Pairs with controller v0.49.0 (pending-activation detection + "Újraindítás most").
## v0.25.0 — slice 10 P2: bind enrolled user-data drives into the guest (passthrough) (2026-06-12)
External user-data drives are mounted on the HOST but were never passed INTO the guest (diagnosed
Branch A), so apps silently wrote to the rootfs and the controller couldn't see them. This adds the
guest passthrough. Spike-proven on 9201 first (see REPORT / the usb-passthrough-spike findings):
`pct set` **bind form** (host path, never `storage:size`), `chown` to the guest base (idmap not clean
for mixed-ownership data), `shared:49` propagation host↔guest automatic.
- **`POST /disks/guest-attach` (`internal/localapi`)** — self-scoped (vmid from token). Binds an
enrolled drive's **felhom-data namespace** into the guest at `/mnt/<name>` (**Model A**: the
felhom-data dir is the bind source mounted AT `/mnt/<name>`, so only Felhom's namespace crosses into
the guest — the customer's other data on the drive never does). Idempotent (returns the existing slot
if already bound); picks the lowest free `mpN`; validates `where` is `/mnt/<name>` (no traversal).
- **`GuestBinder` (`internal/localapi/guestbind.go`)** — the host-root steps over the fenced
`proxmox.Runner` (same pattern as the provision back-half's bind): `mkdir -p <drive>/felhom-data` →
`chown 100000:100000` the namespace ROOT (not -R; per-app subdirs are chowned at deploy) → `pct set
<vmid> -mpN <drive>/felhom-data,mp=/mnt/<name>` (RW bind). The namespace is created fresh + uniformly
owned, which sidesteps the drive's pre-existing mixed-ownership data entirely.
- **Tests** — `TestGuestAttach_*`: free-slot selection (mp0 when mp9 taken), idempotency (no re-bind +
`already:true`), bad-path rejection (traversal/non-/mnt/multi-component), not-configured 503.
Pairs with felhom-controller P2C (enroll triggers attach) + the golden's `/mnt:rslave` controller bind
(P2B). Self-heal reconcile (P3) and dual-role (P4) follow.
## v0.24.0 — role-gate the eject path (system/backup mounts are unmount-protected at the agent) (2026-06-12)
Closes the eject gap in the storage-authorization redesign: `POST /disks/eject` now **refuses to
unmount a system or backup storage**, enforced at the agent — not just hidden in the controller UI.
A direct API call (or a compromised controller) trying to `eject {where:"/var/lib/vz"}` or the PBS
mount is refused 403; only `user-data` mounts are ejectable.
- **`handleDiskEject` (`internal/localapi/disks.go`)** — before `Unmount`, resolves the AUTHORITATIVE
protection role of the storage mounted at `where` (the agent's own storage-view + host-topology
classification, never the caller's claim) via the new `roleForMountPath`. Refuses (403, no
`Unmount`) unless the role is `user-data`. **Fails SAFE**: an unresolvable mount (view error or no
storage target at that path) → treated as protected → refused (the same most-protected-on-ambiguity
default the wipe gate uses). Mirrors the wipe path's "protected — eject refused by role" logging.
- **`roleForMountPath` + `hostReader` seam** — `roleForMountPath` keys `RoleForStorage` on the mount
path (the eject input), mirroring `deviceRole`. `Options.HostReader` (optional; defaults to the
production `*storage.ProcHostReader`) injects the root-free topology reader so the role-gate is unit-
testable. `handleDisks`/`deviceRole` now share the same seam.
- **Tests** — `TestEject_RoleGated` asserts a `system` and a `backup` mount are refused with **no
`Unmount`**, a `user-data` mount ejects, and an unresolvable mount fails safe to refused (the same
non-hollowness the wipe tests use). `TestEject_UnmountAndDependents` updated to a user-data target.
## v0.23.0 — device-ROLE classification + tiered storage-wipe gate (system/backup operator-only, user-data customer-confirmable) (2026-06-11)
The storage-authorization redesign (agent half). The gate's destructive-wipe path is now **tiered by
the device's protection ROLE**, which the agent classifies from its OWN inspection — never the
caller's claim (the storage analog of classify.go's data-bearing verdict).
- **`internal/storage/role.go`** — `DeviceRole` (`system` | `backup` | `user-data`) + the
authoritative classifier. `RoleForStorage` (storage-view targets) and `RoleForRawDevice` (a raw
device, e.g. a fresh disk in the init flow) map a device to its tier via `SystemDisks` (the
whole-disks backing `/`, `/boot`, `/boot/efi`, root-free reads). Rules: `pbs` → backup; `lvmthin` /
builtin `local` / nfs / cifs / unknown → system; `usb` / `local-dir` on a **non-system external
device** → user-data. **Fail-safe**: any ambiguity (system disks unknown, or an unrecognizable
device topology) → **system** (most-protected) — never silently user-data.
- **`GET /disks`** — each `DiskInfo` now carries `role`. The controller drives the UI from it
(system/backup get a lock + no destructive controls; user-data is customer-manageable).
- **Gate tier (`reconcile`)** — new `CustomerConfirmable` disposition + `Gate.AuthorizeStorageWipe`:
- role=**user-data** → **customer-confirmable**: allowed iff the request carries an explicit
customer confirmation **bound to the device's durable id** (the agent re-resolves the durable id
and matches; a confirmation for one disk can't wipe another). **No operator signature.** A
user-data drive is already within the in-guest controller's blast radius (it bind-mounts `/mnt`),
so customer-confirmation adds no new reach. Recorded in the **audit log** with the durable id
(`AuditRecord.DurableID`).
- role=**system**/**backup** → unchanged **operator-signature** (`pending_signature`). The
`confirmed` flag is **IGNORED** — a compromised controller asserting `confirmed:true` on a
protected device is refused **by role**. Every other destructive class (`guest_destroy`,
`decommission`, `restore_overwrite`, `key_rotation`) keeps operator-signature exactly as before.
- **`POST /disks/format`** — accepts `confirmed` + `durable_id` (inert for system/backup). The
data-bearing path tiers by role: user-data customer-confirmed → `mkfs`; user-data unconfirmed →
403 `needs_confirmation` (+ the durable id to confirm against, NOT an opsign command); system/backup
→ 403 with the operator-signature pending op (as before). Blank devices stay benign `mkfs`.
- **Tests** — `role_test.go` (demo-storage mapping + fail-safe), `storage_wipe_test.go` (the gate
refuses a `confirmed` wipe on system/backup → no exec; durable-id mismatch / missing-durable
refused; unknown role fails safe), and the localapi format-handler branches (user-data confirmed →
mkfs; user-data unconfirmed → needs_confirmation, no opsign; confirmed-but-protected → still refused).
Pairs with the controller's lockout + type-to-confirm UX + drive-list restyle.
## v0.22.0 — expose durable_id in GET /disks (enable controller-side guided storage) (2026-06-11)
One-line, read-only addition: `localapi.DiskInfo` gains `durable_id` (mapped from
`StorageTarget.DurableID`, e.g. `"uuid:<fs-uuid>"` for usb/local-dir). The de-privileged controller
cannot read a device's fs UUID itself, yet `POST /disks/assign` mounts strictly by UUID — so without
this it could not complete the guided init/attach flows. The controller strips the `uuid:` prefix to
get the assign key. No new privilege, no behaviour change to format/assign/eject or the data-bearing
gate. Pairs with `felhom-controller` v0.43.0 (the storage-management UI rebuild).
## v0.21.0 — agent-managed split-horizon LAN resolver (internal/lanresolver) (2026-06-11)
LAN clients can now reach their guest **directly** at the same public hostname with the same real
wildcard cert (no Cloudflare hairpin), via a host-side dnsmasq the agent manages. The host is the
stable anchor (static LAN IP); the guest stays DHCP/ephemeral and the agent tracks its live IP.
- **`internal/lanresolver`** — renders a dnsmasq base drop-in (bind to the host LAN IP, no-resolv,
upstreams) + a per-customer drop-in `local=/<domain>/` + `address=/<domain>/<guest-ip>`. The proven
two-line shape: `local=` makes dnsmasq authoritative for the zone so **AAAA returns NODATA** (no
Cloudflare-AAAA split-brain — the guest has only link-local v6), `address=` is the wildcard A; all
other names (and their AAAA) forward upstream unchanged.
- **`Manager`** ensures dnsmasq present (apt) + the base config + enabled, discovers the guest's live
IPv4 (`pct exec <vmid> -- ip -4 -o addr show dev eth0`) and domain (read from the guest controller's
pulled `controller.yaml` — the v2 bootstrap omits it), writes drop-ins **write-if-changed**, and
**reloads** (not restarts) dnsmasq. Tolerates the early-boot pre-lease window (empty IP → skip+retry,
never a blank record). Logs IP transitions.
- **`Loop`** — a 7th daemon goroutine: every interval (default 300s) it enumerates provisioned guests
(`/var/lib/felhom-agent/guests/<vmid>/`) and reconciles each, so the resolver follows DHCP IP changes.
Config `lan_resolver.{enable,host_ip,upstreams,interval_seconds}` (host_ip defaults to the local-API
bridge IP). `--selftest=lanresolver -vmid N`.
- **`configs/felhom-agent.sudoers`** — new `FELHOM_DNSMASQ` alias (apt install dnsmasq; install
felhom-*.conf drop-ins; systemctl enable/reload dnsmasq; rm felhom-*.conf; the two FIXED `pct exec`
reads). The agent never touches `/etc/resolv.conf` (host's own resolution unaffected).
- **Box-down robustness** is a documented **router config** (DNS = [host-IP primary, upstream
secondary]) so a box reboot degrades to the Cloudflare path, not total DNS loss — see REPORT install step.
- Spiked live on felhom-pve first (`:53` free, host IP static `192.168.0.162`, host DNS intact, full
loop from a real LAN client returned the guest IP + AAAA NODATA + the real wildcard cert `200 0`).
## v0.20.0 — golden: stacks-dir bind + per-guest hostname/CT name + bake base-infra images (2026-06-11)
Lockstep with `felhom-controller` v0.41.0 + a golden rebake. Changes in `configs/build-golden.sh` and
the provision path; no change to the proxmox/authz/token fences.
- **Section-G mount fix (the load-bearing one):** the in-guest controller writes app/infra compose
stacks under `/opt/docker/stacks` *inside its container*, but the baked controller-bootstrap `docker run`
never bind-mounted that path. So `docker compose up` (run by the GUEST daemon over the shared socket)
resolved every relative bind source on the guest filesystem — silently creating empty dirs — which
broke **every** bind-mounted stack (base infra AND customer apps like immich/nextcloud). The bootstrap
unit now `mkdir -p /opt/docker/stacks` and adds a **same-path host bind**
`-v /opt/docker/stacks:/opt/docker/stacks` (a named volume would NOT fix this). Empirically confirmed on
guest 9201 before writing the fix.
- **Per-guest container hostname (3A):** the bootstrap unit derives `customer.id` from
`/etc/felhom-bootstrap/bootstrap.json` with a portable `sed` parse (NO jq in the golden) and passes
`--hostname <customer-id>` to `docker run`, so the controller's `os.Hostname()` (its hub-reported
hostname) is the customer id, not the Docker container ID. Fail-safe: no parse → no `--hostname`.
- **Per-guest CT/LXC name (3B):** `--selftest=provision` now defaults `-hostname` to the (DNS-safe
sanitized) `-customer-id` when not given, so the bring-up's existing `SetConfig hostname` step
(`bringup.go`) names the CT meaningfully (e.g. `demo-felhom`) instead of inheriting the golden's
`felhom-golden`. New `sanitizeHostname` (lowercase, collapse invalid → `-`, trim, ≤63).
- **Bake base-infra images:** the golden now also pulls the three PINNED, PUBLIC base-infra images
(`traefik:v3.6.7`, `cloudflare/cloudflared:2026.6.0`, `gtstef/filebrowser:1.3.3-stable`) into its Docker
storage so the controller's first-boot bring-up is OFFLINE-capable. A hard gate (`docker manifest
inspect`) fails the bake early on a bad pin. Tags MUST match the controller's `internal/infra` constants.
## v0.19.0 — bootstrap contract v2: agent relays the hub retrieval passphrase (no host key in the guest) (2026-06-11)
Lockstep with `felhom-controller` v0.40.0. Fixes the onboarding 401: a freshly provisioned guest's
controller used to come up with the agent's **host** hub key baked in, which the hub's `/api/v1/report`
(customer-scoped auth) rejects. The agent now bakes a **v2 bootstrap** carrying only what the controller
needs to **pull** its own config from the hub — the agent never touches the customer-scoped key or CF
tokens.
### Changed — bootstrap contract `v1 → v2` (`internal/provision`)
- `SchemaV1 → SchemaV2 = "felhom.bootstrap/v2"`. **`DocCustomer`** drops `name`/`domain`/`email` (keeps
`id`). **`DocHub`** drops `api_key`/`host_id`, adds **`retrieval_password`** (the customer's hub
retrieval passphrase — SECRET). `DocLocalAPI` unchanged. The contract is byte-compatible with the
controller's `internal/bootstrap.Bootstrap` (cross-repo round-trip verified).
- `backhalf.go`: renders the v2 Doc; validation now requires `customer.id` + `hub.url` +
`hub.retrieval_password` (was `customer.id` + `customer.domain`). Write/0600/chown/`pct set` unchanged.
- `cmd/felhom-agent/main.go` `--selftest=provision`: **new required `-hub-password`** flag (the customer's
hub retrieval passphrase; the customer must already exist in the hub). Stops baking `cfg.Hub.APIKey` /
`cfg.Hub.HostID`. `-customer-domain/-name/-email` still accepted (bring-up may use them) but NOT baked.
### Changed — `configs/build-golden.sh`
- Default `CONTROLLER_IMAGE` bumped off the stale `:v0.35.0` → `:0.40.0` (matches the registry's no-`v`
tag convention; latent footgun fixed).
### Tests
- `doc_test.go`/`backhalf_test.go` updated to the v2 shape (assert no `api_key`/`host_id`,
`retrieval_password` present, `customer` carries only `id`). `go build ./... && go test ./...` green.
## v0.18.0 — slice 10D: DR capstone — identity escrow + restore-mode consumption (agent side) (2026-06-10)
The agent half of the slice-10 DR capstone (closes slice 10). Grounded by both 10-series spikes
(escrow-consumption + identity-restore). The hub half (recovery-mode toggle, re-enroll + credential
rotation, directive serving) is hub v0.11.0. **Operator-side rotation model (locked):** the hub holds
no Cloudflare write-power; the destructive tunnel/PBS rotation is the operator's step from a trusted
environment (same spirit as 10B).
### Added (`internal/escrow`)
- **Identity escrow** (`identity.go`): `WrapIdentity`/`UnwrapIdentity` (+ `…Bundle`) wrap the
`{tunnel_token, pbs_token}` bundle under the SAME recovery code `R` via **`age`** (scrypt +
ChaCha20-Poly1305 — a vetted passphrase-AEAD, not hand-rolled), reusing the K-escrow pty mechanism
(passphrase via the tty, data via files; `R`/tokens never logged). Same two-factor, zero-knowledge
shape as the K-escrow. A **wrong R fails closed** (no bundle). `age` is a runtime dep for the
identity path (analogous to proxmox-backup-client for K).
- **`escrow.Create`** gains an optional `IdentityBundle` → also emits an `IdentityBlob` under the same
R (additive; the K-escrow + 10C `Consume` paths are byte-unchanged). Self-verifies the identity
round-trip before shipping.
- **`--selftest=escrow-create -identity-bundle <file> -directive <file>`** — also wrap + upload the
identity blob + the **non-secret** DR directive (pbs repo/ns, expected key fingerprint, tunnel id).
- **`--selftest=identity-consume -blob <file> -keydest <file>`** (R via `FELHOM_RECOVERY_CODE`) —
recover the identity bundle through the real code; tokens written 0600, never logged.
### Tests
- identity bundle round-trips (wrap→unwrap byte-identical; blob is opaque ciphertext); wrong R fails
closed + the blob stays retryable; input validation. K-escrow/10C tests byte-unchanged (additive).
(age integration tests gated to a host with the `age` CLI.)
## v0.17.0 — slice 10C: escrow consumption (productionize the spike) (2026-06-10)
Turns the throwaway 10C spike harness into a real, tested **`Consume`** path: recover the PBS key
`K` from an R-wrapped escrow blob, **gate it on the expected fingerprint**, and install it for the
restore. The spike already proved the crypto + real-data restore; this bakes its findings into
production code. **Agent-only** — 10C *reads* the four inputs as parameters (so it stays
standalone-testable); 10D sources blob/fingerprint/PBS-connection from the hub and prompts for R.
**Zero-knowledge holds**: the hub serves everything except **R** (by hand from the customer), so a
hub compromise alone still can't decrypt.
### Added
- **`escrow.Consume(ctx, blob, R, expectedFingerprint, keyDest)`** — the consumption contract:
1. **Unwrap** the blob (a copy — F-C6: the input blob is read-only → a failed Consume is
**retryable**) with `R`; a **wrong R fails closed** at the scrypt KDF (F-C3) → a clear,
R-free error, **nothing written**.
2. **Fingerprint gate (F-C4)** — `KeyFingerprint(recovered)` must equal the expected (the hub
knows it); a mismatch **fails fast + loud, no install, no restore attempted**.
3. **Atomic install (F-C2)** at `keyDest` (`0600`, write-temp-sibling→rename); any failure leaves
**no partial install**. The recovered key lives only in a `0700` tempdir that is always removed.
**Secret discipline:** `R` and key bytes are never logged/persisted (only fingerprint prefixes);
`K` is never mutated.
- **`--selftest=escrow-consume`** (`-blob -fingerprint -keydest`, R via env `FELHOM_RECOVERY_CODE`
to keep it off the command line) — invokes the real `Consume` live (the spike's S3 via the
production path, not a harness).
### Tests (non-hollow)
- valid → key installed + `KeyFingerprint(dest) == expected` + `0600` + blob byte-unchanged;
**wrong R** → error, **no file at dest**, blob unchanged; **fingerprint mismatch** → fail fast,
**no install** (the gate runs before any restore); input validation; format-tolerant fingerprint
compare (no empty-fingerprint gate-bypass); atomic-install permissions (integration tests gated to
a host with `proxmox-backup-client`).
## v0.16.0 — slice 10B: operator-signed destructive completion (offline key + signing CLI) (2026-06-10)
The security centerpiece: a destructive op runs ONLY on a verified, operator-signed authorization
— signature valid against a **pinned** operator pubkey (never the hub's or the blob's), nonce
unseen + durably burned, in-window, host-bound, and **resource-bound to a DURABLE device id** that
execution re-resolves + re-inspects. Decision (a): **offline operator key + signing CLI**,
hardware-key-ready (`sk-`/YubiKey via ssh-keygen). The key floor holds: the signing key is NOT in
the hub and NOT in the agent. Concrete consumer: this **closes the 8C data-bearing-wipe
`pending_signature` gap**. Pairs with hub v0.10.0.
### Added
- **`cmd/felhom-opsign`** — the operator's offline signing CLI. Builds the canonical `OpBlob` by
**reusing `authz.CanonicalBlob`** (the exact production path the verifier authenticates over — so
signer + verifier can never drift) and signs it with **`ssh-keygen -Y sign -n felhom-op-v1`**
(hardware-ready). Output: a `{op_blob_b64, sig_armored}` envelope to hand to the hub jobs queue
(optional `--upload`). Touches ONLY the operator's signing key.
- **`authz.CanonicalBlob`** — promoted to production (was test-only) so the CLI + verifier share one
canonical-bytes source; params canonicalized (sorted keys, compact).
- **`internal/storage` durable device identity** (`durable_device.go`): `DeviceDurableID` (derive a
stable `byid:`(wwn/serial)/`byuuid:` id from the world-readable udev symlinks — no privilege, no
subprocess) + `ResolveDurableDevice` (re-resolve to the current `/dev` path; a path-only/unknown
scheme is REFUSED). The resource-level anti-retarget.
- **`internal/signedjobs`** (new): the queue consumer. `Runner` fetches each opaque job → runs it
through the **gate** (the LOCKED authz pipeline) → on all-pass hands the verified op to an
`Executor`; the order is **verify → nonce-burn (durable, in Verify) → execute → clear job**. The
**`WipeExecutor`** is the 8C consumer: resolve the signed durable id → **re-derive + match**
(anti-retarget) → **re-inspect (8C classifier)** the device is still the data-bearing target →
`mkfs`. A vanished/changed/non-data-bearing device or a path-only binding is refused **even with a
valid signature**. Wired as a second `EnvelopeObserver` (runs on `HasSignedOps`).
- **`hub.Client.Jobs` / `CompleteJob`** + `hub.MultiObserver`; the 8C format refusal now **surfaces
the bound op** (op + durable id + host) in its 403 `pending_op` + a `felhom-opsign …` hint.
### Pinning / rotation
- Operator pubkeys are pinned via `authz.signers` (config, trusted path — provision/agent config,
NEVER hub-alone), **multiple** keys (KeyID selects; role-scoped), so a backup/rotation key exists
without a flag-day. Unchanged from the slice-4 verifier wiring; 10B activates the execute path.
### Tests (real crypto, non-hollow)
- `signedjobs` runner over the **real** gate+verifier (in-Go minted SSHSIGs): valid → executor runs
once + job cleared; **replay** (nonce burned) / **non-pinned signer** / **expired** / **retarget**
(other host) / **forged sig** / **no pinned signer** → all rejected, **executor never called**;
malformed envelope cleared.
- `WipeExecutor`: valid → `mkfs` runs; **path-only**, **durable-id mismatch**, **device gone**,
**re-inspect non-data-bearing**, **not-probed** → all refused, `Format` not called.
- `storage` durable: wwn-preference, uuid-fallback, path-only/traversal refusal, round-trip,
missing-device error (symlink tests gated to Linux — the agent's OS).
## v0.15.0 — slice 10A: hub desired-state serving — the "Down" channel (2026-06-10)
The agent half of slice 10A. The control envelope (`hub.ControlEnvelope`) stops being "reserved — ignored" and becomes the live **Down channel**: a cheap change-notification on every heartbeat. The agent caches the hub's desired-state + its generation; only when **`DesiredGeneration` advances** does it fetch the full state (the heartbeat stays light, the heavy state moves on change). The engine then reconciles **benign** deltas and the gate marks an explicit **destructive** delta `pending_signature` (no signer in 10A → never executed; signed execution is 10B). Pairs with hub v0.9.0.
### Added / changed
- **`internal/reconcile`**: `DesiredGuest.Decommission` — the canonical **destructive desired-state delta** (an EXPLICIT flag, not "absent from the list", so a partial hub list can never mass-destroy). The planner emits `ActionDecommission` → `ClassDecommission` → Destructive → the gate refuses it `pending_signature`. `Reconcile` now counts a `pending_signature` refusal as **`Result.Pending`** (expected, logged INFO) rather than a failure; any other refusal stays a real failure. `ActionDecommission` has **no executor** (slice 10B) — a defensive guard refuses to run it. New **`CachingProvider`** (thread-safe DesiredState + generation cache; `Desired`/`Update`/`Generation`) — the production `DesiredProvider`, replacing `EmptyProvider` in the daemon engine (empty until the hub serves intent → cold-start is a live no-op, unchanged).
- **`internal/hub`**: the **`ControlEnvelope`** fields are now active (DesiredGeneration drives the fetch, HasSignedOps noted). New wire types **`DesiredStateResponse`** + **`WireDesiredState`** (guests + forward-compat `restore_directive` (10D) / `pbs_namespace` / opaque `storage_manifest`+`backup_policy`) + **`WireDesiredGuest`** (vmid/run/spec/description/decommission). New **`Client.FetchDesiredState`** (GET `/api/v1/hosts/{host_id}/desired-state`, self-scoped to the client's own host). New **`EnvelopeObserver`** loop seam + `SetEnvelopeObserver` — the loop hands the envelope to the sync layer each cycle (hub does not import reconcile/desired).
- **`internal/desired`** (new): the **`Syncer`** — implements `hub.EnvelopeObserver`, fetches desired-state on a generation advance, maps the wire shape to the reconcile domain, and updates the `CachingProvider`. Caches the **fetched** generation (robust to a generation that advanced mid-fetch); a fetch failure keeps the last-known state. `restore_directive` is carried + logged, not acted on (10D). Wired in `cmd/felhom-agent` (daemon): provider → engine, syncer → loop.
### Tests
- reconcile: a desired-state with one benign + one decommission delta → **benign applied, destructive gated pending (not executed)**; `Plan` emits decommission-only for a decommissioned guest + classifies Destructive; `CachingProvider` update/isolation.
- desired: **fetch-once-on-advance** (no re-fetch on an unchanged generation), fetch-failure-keeps-cache, caches-the-fetched-generation.
- hub client: `FetchDesiredState` hits the self-scoped path with the bearer + decodes (incl. `restore_directive`); a 403 is a typed `HTTPError`.
- loop: the cycle notifies the observer + adopts `PollIntervalSeconds`; a report error skips the observer.
- cross-repo golden: `testdata/desired-state.golden.json` + `control-envelope.golden.json` decode + key-set guard, **byte-identical** with felhom.eu/hub.
## v0.14.0 — slice 9: host metrics to the controller (`GET /host/metrics` + CPU-temp collector) (2026-06-10)
The de-privileged controller (slice 8C) sees only its own cgroup, so it can't read host health itself. Slice 9 **re-serves** the slice-4 collector's host + per-storage view to the customer over the local API, plus the one missing collector — CPU/chassis temperature — so the customer sees their box's health in the controller. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot). Assumption: **one customer per host** (the home-server model); if a host ever serves multiple customers, host-wide CPU/mem would leak cross-customer load → revisit then.
### Added / changed
- **CPU/chassis-temp collector** (`internal/hub/cputemp.go`): `SysfsTempReader` reads the CPU package temperature straight from sysfs — hwmon (`coretemp`/`k10temp`/`zenpower`/`cpu_thermal`, preferring the `Package id 0` input) then the thermal zones (preferring `x86_pkg_temp`/`coretemp`/`cpu-thermal`, falling back to `acpitz`). **No external binary, no privilege** (sysfs nodes are world-readable), so the root-CLI fence is untouched. **Graceful-null**: a missing sensor, an unsupported board, an implausible reading (outside 5150 °C), or any read error all degrade to `null` ("n/a") — a missing sensor never fails the report. Wired into the collector via the new `TempReader` seam (nil-safe).
- **`HostMetrics.CPUTempC *int` (`cpu_temp_c`)** — new nullable wire field on the **shared** `HostMetrics` struct (same nullable contract as the disk `SmartSummary.TemperatureC`). It rides the **hub report too** (operator freebie) → cross-repo host-report golden updated.
- **`Collector.HostMetricsNow(ctx)`** — a fresh `NodeStatus` + CPU-temp read returning just the host block, the source for the local API (current cpu%/temp, not the 15-min snapshot). `Collect()` now also populates `cpu_temp_c` on the hub report. `Collector.SetTempReader` injects a fake in tests.
- **`GET /host/metrics`** (`internal/localapi/host_metrics.go`): host-wide health (cpu%/mem/load/uptime/`cpu_temp_c`) + per-storage capacity targets (total/used/fraction, thin-pool, SMART temp+wear). Token-authed via `withGuest` (host-wide data; cross-guest `?vmid=` still 403). Best-effort on storage (a view error still returns the host block). Served only when the `HostMetrics` provider (the shared collector) is wired — else 503 "not configured". Wired in `buildLocalAPIServer`.
### Tests
- `cputemp_test.go`: a fake `/sys` layout proves hwmon package-preference, hwmon first-input fallback, thermal-zone-by-type selection over a non-CPU hwmon, **graceful-null on a sensorless host** (no error), and rejection of implausible (0 m°C) readings.
- `hostmetrics_test.go`: `HostMetricsNow` populates the temp, gracefully nulls it, hard-errors on `NodeStatus` failure; `Collect()` carries the temp.
- `host_metrics_test.go` (localapi): populated host+storage with a valid token; `cpu_temp_c:null` serializes; **401 without a token** (collector never invoked); 403 on a cross-guest `?vmid=`; 503 when not configured.
## v0.13.0 — slice 8B.2: quiesce downtime optimization (`snapshotted` phase) (2026-06-10)
The agent half of slice 8B.2. In snapshot mode, vzdump only needs the app-stopped state captured at
the **storage-snapshot moment**; after that it reads from the snapshot and the app can resume. The
agent now emits a **`snapshotted`** phase on `GET /backup/status` when the snapshot is taken, so the
controller (v0.38.0) resumes its app early — app downtime drops from *whole-backup* to
*until-snapshot* with no loss of app-consistency. Validated Phase-0 first on PVE 9.2.2: the marker is
`INFO: create storage snapshot 'vzdump'`; downtime ~24s→~1s for a 934 MB guest.
### Added / changed (`internal/backup` + `internal/localapi`)
- **`BackupRunner.BackupWithSnapshotHook(ctx, vmid, onSnapshot)`** — while the vzdump runs, a watcher
tails the task log (`TaskLogTail`) for the **`create storage snapshot`** marker and fires
`onSnapshot` **once**. The marker only appears in snapshot mode (stop/downgraded takes no storage
snapshot), and the watcher also bails on `backup mode: stop` — so it never fires in stop mode.
(`Backup` keeps its signature for the scheduler/selftest; both share one body.)
- **`/backup/status` phase `snapshotted`** (between `running` and `done`): `handleBackup` passes the
hook → `markSnapshotted` flips the running job to `snapshotted`. `done`/`failed` semantics unchanged.
### Tests
- localapi: snapshot mode → phase reaches `snapshotted` before `done` (gated fake holds the backup
open); stop mode → `snapshotted` **never** emitted (stays running → done). runner: the watcher
fires `onSnapshot` on the marker; in stop-mode log it never fires. `snapshotWatchInterval` is a
package var so tests run fast.
## v0.12.0 — slice 8C Phase A: disk endpoints + data-bearing classifier gate + mkfs executor (2026-06-10)
The agent half of slice 8C, Phase A (additive). Adds the host disk-management endpoints the
controller's disk UI drives — with the **8C security invariant**: the agent decides
data-bearing-ness by **inspecting the actual device** (agent-internal evidence), NEVER from the
caller's claim. A compromised controller asserting "this drive is blank" cannot wipe a data-bearing
drive. (Controller rewire + disk-subsystem retirement + de-privilege are Phases B/C, `felhom-controller`.)
### Added
- **`internal/storage``mkfs` executor + data-bearing inspection.** `SudoHostOps.Format(device,
fstype)` (device-pinned, `ValidateBlockDevice`+`ValidateFSType`, narrow `FELHOM_FORMAT` sudoers —
`mkfs.ext4 -F` / `mkfs.xfs -f` on a `/dev/*` path the agent fine-validates first).
`SudoHostOps.InspectDevice(device)` → `DeviceProbe` (filesystem signature via `blkid -p`, partition
table / partitions / mount via `lsblk -J`). **`DeviceProbe.DataBearing()` is conservative**: any
signature / partition table / partition / mount — OR a probe that did not read cleanly — is
data-bearing (fail-safe; an unreadable device is never called blank).
- **`internal/localapi` — the §6 disk endpoints**, all self-scoped (token→guest; cross-guest 403):
- `GET /disks` — host drives + a **data-bearing flag** (UI hint). Read-only/benign.
- `POST /disks/assign` — attach a drive as a mount (benign, additive → `EnsureMount`). Self-serve.
- `POST /disks/eject` — safe-unmount (benign, data preserved) + the **dependent guests** that
mount it (so the controller can warn which apps lose that storage).
- `POST /disks/format` — **the security centerpiece**: the agent **inspects the device itself**;
blank → benign → `mkfs`; **data-bearing → ClassStorageWipe → the slice-4 gate → refused
`pending_signature`** (the operator-signed completion is slice 10). The caller's claim is
ignored — only a device the agent reads as blank is formatted.
- `storageGateAdapter` bridges the format path to the slice-4 reversibility gate (no new gate/crypto).
### Tests
- localapi (security matrix): blank device → **mkfs called, gate not consulted**; a **data-bearing
device → 403, mkfs NEVER called**, gate consulted (`pending_signature`); an **ambiguous/unprobed
device → treated destructive** (fail-safe); even a gate that *allows* does not format data-bearing
in 8C; assign → `EnsureMount`; eject → `Unmount` + dependent guests; cross-guest → 403; bad
device/fstype → 400; unconfigured → 503.
- storage: `ValidateBlockDevice`/`ValidateFSType` (whitelist + injection rejection); `InspectDevice`
blank/filesystem/partition-table/mounted/failed-probe-fail-safe; `Format` invokes the right `mkfs.*`.
## v0.11.0 — slice 8B: app-consistent backup — /backup/due policy + /backup/status phases (2026-06-10)
The agent half of slice 8B (doc 03 §8). Turns the 8A thin backup stubs into the real policy the
in-guest controller's quiesce loop drives (controller half: `felhom-controller` v0.36.0). No hub
change. The downtime optimization (`vzdump --mode snapshot` + a `snapshotted` phase) is the 8B.2
fast-follow; the hub-served per-guest policy is slice 10.
### Changed (`internal/localapi`)
- **`GET /backup/due`** — real **cadence** policy (replaces the 8A "never backed up" stub): a guest
is due when no **successful** backup is recorded OR the newest one is older than the agent-local
cadence (`backup.backup_cadence_seconds`, default 24h). A successful `POST /backup` flips due to
**false** for the window, so the controller won't re-quiesce in a loop. A failed backup does not
satisfy the cadence. Returns `age_seconds` for diagnosis.
- **`GET /backup/status`** — real **phases** `idle | running | done | failed` + the job id, so the
controller can poll a backup to completion (was: just the latest stored backup).
- **`POST /backup`** — returns a **job id** + `running` phase; tracks the in-flight job and is
**single-flight per guest** (a second POST while one runs returns the same job — no concurrent
vzdump). On completion the job transitions done/failed and the result is recorded to the store.
- Config: `backup.backup_cadence_seconds` + `BackupCadence()`; the local-API server takes the cadence.
### Tests
- `/backup/due`: due when stale / no backup, **not due within the window after a success**, due again
past the cadence, **a failed backup does not count**. `/backup/status`: running→done and
running→failed (gated fake to observe the running phase). `POST /backup` single-flight (one vzdump
for concurrent POSTs). All still self-scoped (token→guest).
## v0.10.0 — slice 8A: agent local-API server + provisioning back-half (2026-06-10)
The host-agent half of slice 8A (doc 03 §6). Adds the per-guest **local API** the in-guest
controller calls over the bridge, and the **provisioning back-half** that follows the slice-7
bring-up front half. Grounded by `felhom.eu/documentation/tests/slice8a-channel-deploy-spike-findings.md`
(commit `4a81a96` — channel + deploy plumbing proven; the 5 gotchas resolved here). Controller half
is `felhom-controller` v0.35.0. No hub change.
### Added
- **`internal/localapi`** — the HTTPS local-API server (doc 03 §6), the **per-guest authorization
gate**. Serves a **persisted self-signed leaf** with a **stable SHA-256 fingerprint** (generated
once; a fresh cert each boot would invalidate every baked bootstrap pin). The **7 §6 endpoints**,
all **self-scoped to the caller's own guest**: `GET /storage` (this guest's mpN mounts + fast/slow
class from the slice-5/7 storage view), `POST /snapshot`, `POST /rollback`, `POST /backup`
(enqueued, crash-consistent — the app-consistent quiesce loop is 8B), `GET /backup/due` (thin in
8A), `GET /backup/status`, `GET /restore-test/status`.
- **Token store** (`tokenstore.go`): durable, crash-safe per-guest token→guest map that persists
only a **SHA-256 hash** of each token (the plaintext exists transiently at mint→write-to-mount,
then is discarded), last-write-wins per guest, fsync'd append-only JSONL (mirrors the nonce store).
- **Self-scoping**: the VMID is resolved ONLY from the token; an explicit `vmid` (query/body) that
disagrees → **403 and the proxmox op is never issued for the other guest**; absent/unknown → 401.
- **`internal/provision`** — the back-half: mint the per-guest token → render the stable
**`bootstrap.json`** contract (schema `felhom.bootstrap/v1`; **no registry credential** — the
controller image is baked into the golden) → write it `0600` → **`chown 100000:100000`** (the
unprivileged-LXC mapped guest-root, spike gotcha 1) → attach a **read-only bind mount** via
`pct set`. Host-side only (F3 — the agent never enters the guest; **no `pct exec`**). The token
plaintext is never logged and never returned.
- **`--selftest=provision`** — the full chain on-demand: bring-up (provision) front half + the
back half; keeps the guest for the golden's baked controller-bootstrap unit to deploy.
- **`config.LocalAPIConfig`** (`local_api`) — enable + bridge `listen_addr` + cert/key paths + token
store path. The server is an optional 6th daemon goroutine, disabled cleanly when unconfigured or
on a token-store/cert failure (the daemon still reports/reconciles).
- **`configs/build-golden.sh`** now **bakes the controller image** (pulled once on the trusted build
host, then `docker logout` — no cred baked) + a **controller-bootstrap unit** that deploys the
**baked** image from the config mount on boot (no login/pull at deploy).
- **`configs/felhom-localapi-firewall.example`** — host firewall narrowing of the local-API port to
the guest bridge subnet (nft/iptables/PVE variants; defense-in-depth — the token stays the gate).
- **`configs/felhom-agent.sudoers`** — a narrow `FELHOM_PROVISION` alias (`chown 100000:100000` +
`pct set` bind-mount, both confined to the agent-owned `/var/lib/felhom-agent/guests/*` path) for
the non-root least-privilege deployment.
### Security / design notes
- The local-API leaf is pinned by **leaf-cert SHA-256** (decision: consistency with the agent's
PVE/PBS pinning); the fingerprint is baked into each guest's bootstrap.
- The back-half's host-root ops (chown + bind-mount attach) are **NOT** added to `proxmox.Privileged`
(which is fenced to its 3 exceptions) — they live in `internal/provision` and run through the shared
`Runner` (direct as root, or `sudo -n` with the new sudoers alias). This is the per-guest
provisioning host-root surface, host-side and F3-compliant.
### Tests
- localapi: self-scoping (cross-guest snapshot/rollback/backup → 403, op never issued for the other
guest; own-guest uses the token's VMID), 401 paths, `/storage` class mapping, `/backup` enqueue,
the thin `/backup/due`, status scoping; the token store persists only the hash (plaintext never on
disk), last-write-wins, survives reopen, uniqueness; the leaf fingerprint is stable across reload.
- provision: writes `0600` + chowns + attaches the bind mount with the right args; the **token never
appears in the Result**; the cross-repo `bootstrap.json` contract key-set is pinned.
## v0.9.0 — slice 7 close-out: PBS recovery-code escrow creation (2026-06-10)
The first code that touches the PBS client encryption key `K` and introduces the customer recovery
code `R`. Default posture is **zero-knowledge**: Felhom holds an opaque `R`-wrapped blob (cannot
open it), the customer holds `R`. Grounded by `felhom.eu/documentation/tests/slice7-escrow-spike-findings.md`
(round-trip proven on a throwaway: the `R`-recovered key restores a real encrypted snapshot). Hub
opaque storage is the `felhom.eu` half (hub v0.8.0); consumption/serving is slice 10.
### Secret discipline (overriding)
`R` is `crypto/rand`, ≥128 bits, surfaced **exactly once** and **never** logged/persisted/committed;
the wrap pty's echo is discarded so `R` can't leak. `K` is read by location, **never modified** (the
live key file is byte-unchanged — Wrap operates on a copy), never logged.
### Added
- **`internal/escrow`** — `Create` generates `R` (10 EFF-wordlist words ≈ 129 bits), wraps `K` under
`R` via the **PBS-native** `proxmox-backup-client key change-passphrase --kdf scrypt`, and
**self-verifies** the blob recovers `K` (fingerprint match) before shipping. The wrap is driven
over a **stdlib pty** (`x/sys/unix`; spike F-A1 — the command is TTY-only) with **output discarded**
(F-A2 — the pty echoes the passphrase). Opt-in outputs: **(b)** `R`-wrapped offline copy (two-factor,
no extra trust) and **(a)** raw paperkey (single-factor, unrevocable — loud caveat).
- **`--selftest=escrow-create`** (`-storage`, `-paperkey`, `-offline`, `-upload`): surfaces `R` once
to stdout (never the logger), prints the opaque blob's size/fingerprint/posture, and with
`-upload` PUTs the blob to the hub (`/api/v1/hosts/{host_id}/escrow`, per-host key).
- Config: `escrow` section (`posture` default `zero_knowledge`, `pbs_storage_id`); `PBSEncKeyPath`
helper (the `<id>.enc` key K).
- Runtime dependency on the `proxmox-backup-client` CLI (the PBS key+passphrase KDF).
### Tests
- `R` entropy ≥128 / 10-word format / uniqueness; integration round-trip (wrap→unwrap fingerprint
match, **wrong-`R` fails**, **live `K` byte-unchanged**, blob ≠ plaintext key) guarded to
linux+`proxmox-backup-client`; the agent→hub wire-contract key-set (mirrors the hub's).
- **Live-validated** (demo): `escrow-create` → `R` (10 words) surfaced once, blob 383 B opaque,
self-verify ok, **live `K` sha256 unchanged**, exact `R` absent from stderr/journal.
## v0.8.0 — slice 7 Phase 1: unified bring-up reconcile job (provision + guest-loss DR) (2026-06-09)
The shared FRONT HALF of provision and guest-loss DR, as a journaled reconcile job mirroring the
slice-6 restore-test's crash-safety — but it KEEPS the guest on success and applies a
scenario-specific identity policy. Agent-only; no hub/wire change (the new guest auto-appears in
the host-report via `ListLXC`). Grounded by the slice-7 bring-up spike findings (commit `3342993`):
F1 (restore preserves the archived MAC → provision reset is unconditional), F3 (SSH host keys do
not auto-regenerate → a baked golden first-boot unit, not an agent guest-internal op), F4 (the
transient PVE config-lock 500 → bounded retry).
### Added
- **`reconcile.RunBringUp`** (`bringup.go`) — `BringUpSpec` (Mode `provision`|`dr_guest_loss`,
Archive, VMID, RestoreStorage, Hostname, Cores/MemoryMB, RootfsGrowGB, Mounts, KeepMAC,
BootTimeout) → `BringUpResult` (VMID, AssignedMAC, Pass, Verified, StartWarnings/Recognized).
Sequence (each mutation preceded by journaling the owning entry): restore → identity reset →
size → attach mounts → start LINK-UP. **Verdict is liveness (`waitRunning`), never the start
exitstatus** (reuses the v0.7.0 WARNINGS surface). **Success KEEPS the guest** (no teardown).
- **Scenario-specific identity reset** (doc 03 §9): *provision* → fresh MAC unconditionally
(`PUT net0` with `hwaddr` omitted → PVE regenerates, F1) + hostname; machine-id + SSH host keys
regenerate guest-side on first boot (golden bake + the new unit) — the agent does NOT touch
guest internals. *dr_guest_loss* → preserve continuity (keep hostname; keep MAC unless
`KeepMAC=false`); never resets restic/tunnel/hub identity.
- **Compensating rollback** — any mid-flight failure destroys the just-created guest
(`ClassGuestDestroy`, benign via `Provenance{SameTxnCreated:true}`, gated); on teardown failure
the entry is left in-flight for `Recover`. New journal flag **`Rollback`** + `Recover`'s
`recoverBringUp` reap a half-built guest left by a mid-job crash (idempotent, via `ListLXC`).
- **F4 config-lock retry** — steps 3+5 coalesced into ONE `PUT config` (net0+hostname+cores+
memory+mpN); rootfs grow stays its own call. `setConfigWithLockRetry` retries ONLY the transient
PVE config-lock 500 (`pveConfigLock`: 500 + "can't lock file"/"got timeout"); any other error
fails immediately — never retried.
- **`--selftest=bring-up`** (`-mode provision|dr -archive -vmid -hostname [-keep]`) — runs the real
journaled job (after a `Recover`), then tears the guest down unless `-keep`.
- **`configs/build-golden.sh`** — the validated golden recipe as a script, incl. the F3
first-boot `felhom-regen-hostkeys.service` unit (Condition-gated: fires on provision, no-ops on
DR). The slice-7 spike archive (which lacks the unit) is superseded.
### Deferred (stated, not built)
- Provisioning BACK HALF (controller deploy, bootstrap, per-guest token mint) → **slice 8**.
- Host-loss DR + PBS escrow consumption → **slice 10**.
- The SOURCE of a `BringUpSpec` (hub desired-state: which archive/VMID/mounts) → **slice 10**;
this job takes the spec as input. `GuestMount` is defined minimally (no hub coupling).
### Tests
- provision happy path (fresh MAC = net0 without hwaddr, hostname, coalesced sizing+mount, rootfs
grow separate, started, **guest NOT destroyed**); compensating rollback at each step (restore /
config / start-task / waitRunning — asserts the guest WAS destroyed); DR continuity (MAC kept,
hostname not reset) + DR `KeepMAC=false` resets MAC; liveness verdict (warnings+running pass /
not-running fail); F4 (lock-500→retry→proceed; non-lock-500→fail without retry); owning entry
journaled BEFORE restore; reserved/existing VMID refused; `Recover` rolls back / clean.
### Live-validated (demo-felhom)
- provision: fresh MAC + hostname; **SSH host keys regenerated by the baked golden unit** (agent
issued no `ssh-keygen`), machine-id unique, Docker runs, clean DHCP lease → torn down.
- dr: continuity preserved (hostname + host keys kept). Recover: a killed mid-restore left an
orphan; the re-run's `Recover` rolled it back (idempotent).
- **Live caught a bug, then fixed:** the host-key unit's `ExecStart` was `/usr/sbin/ssh-keygen`
(203/EXEC); on Debian 13 it is `/usr/bin/ssh-keygen` — corrected in `build-golden.sh`, golden
rebuilt, re-validated. (Mocked unit tests couldn't surface this; the live run did.)
## v0.7.0 — restore-test: verdict is liveness, not start-task exitstatus (2026-06-09)
Fixes a correctness bug found by the live hub-enrollment runbook: the self-restore-test reported
`pass:false` on **every** modern-distro guest. PVE's guest-start task exits `"WARNINGS: 1"` for the
benign systemd-nesting advisory (`WARN: Systemd 257 detected. You may need to enable nesting.`), and
`WaitTask` treated any non-`"OK"` exitstatus as a hard failure — so the verdict was decided by an
advisory exit code instead of by observed liveness, *before* the real boot check ran. A crying-wolf
test got it disabled on the demo host; this re-enables it. **Single bump (0.6.0→0.7.0) covering the
agent's part of both task phases**; the wire fields below are consumed by hub from **v0.7.5**.
Design invariant (in code): **warning classification affects *visibility only*; pass/fail is
liveness-only.** A wrong/stale recognizer can at worst over-notice a benign warning — it can never
false-fail and never hide a real warning.
### Added
- **`proxmox.WaitOptions.AllowWarnings`** — opt-in per call. When set, a task that completes
`"WARNINGS: N"` is success with the `TaskStatus` (ExitStatus intact) returned so the caller can
read/surface it. Default (`false`) keeps **every existing caller strict** (vzdump/restore/destroy
warnings can be meaningful — relaxing them is a future per-call decision with evidence). Any
non-WARNINGS non-OK exit is still a `*TaskError`.
- **`reconcile.RestoreTestResult.StartWarnings` / `.WarningsRecognized`** + a version-free recognizer
(`benignWarningAnchor = "enable nesting"`, case-insensitive substring — contains no systemd version
number, so it can't rot back into the bug at systemd 258+). `extractWarningLines` pulls `WARN…`
lines from the start-task log.
- **`reconcile.GuestAPI.TaskLogTail`** — the engine fetches the start task's log to surface warnings.
- **`hub.RestoreTest.warnings` / `.warnings_recognized`** wire fields (`omitempty`), populated by
`ToHubRestoreTest`. Additive: the deployed v0.7.4 hub ignores them; hub v0.7.5 consumes them
(passed-with-warnings INFO, or WARN when not recognized). Cross-repo golden updated with the hub side.
### Changed
- **Restore-test start step** (`reconcile/restoretest.go`) now waits with `AllowWarnings:true`,
surfaces any start warnings, and **continues to `waitRunning` as the verdict** — boot+running is the
pass, exactly as before; a real (non-WARNINGS) start-task error still fails. The restore and
scratch-teardown WaitTasks stay strict.
- **Restore-test scheduler logging** distinguishes a clean pass, *passed-with-recognized-warnings*
(INFO), and *passed-with-unrecognized-warnings* (WARN) — nothing silent.
### Tests
- `WaitTask`: AllowWarnings accepts `WARNINGS` (status returned intact); AllowWarnings still fails a
real error; default still fails on `WARNINGS` (existing callers unaffected).
- Restore-test (engine, mock proxmox): start-with-warnings + running → **pass** with warnings
surfaced+recognized; unrecognized warning + running → pass, not-recognized; **not-running → fail
regardless of warnings** (verdict is liveness); teardown still runs.
- **Regression guard:** the `"enable nesting"` recognizer matches the advisory for systemd 256300,
proving it's version-independent and can't silently rot back into the false-fail.
## v0.6.0 — slice 6 Phase B: PBS offsite tier (verify + PBS-API client + reporting) (2026-06-09)
Completes slice 6. The PBS spike (felhom.eu phase5-pbs-spike-findings.md) proved backup-to-PBS
and restore-from-PBS reuse Phase A UNCHANGED (PBS is just a storage target + a volid), and the
operator token needs no widening. So the only new agent code is the **verify capability + a
small PBS-API client + PBSSnapshot reporting**. Escrow + host-loss DR stay slices 7/10.
### Added
- **`internal/pbs` — the PBS-API client** (the agent's SECOND privileged external surface,
slice-1 discipline): TLS **fingerprint-pinned** to the PBS leaf cert (a spoofed PBS →
rejected, mirroring the PVE pin), **token auth** (`PBSAPIToken=<id>:<secret>`; id from the
storage `username`, secret read at runtime from `/etc/pve/priv/storage/<id>.pw` — referenced
by location, never logged/committed), typed, no shell. Methods: `Verify` (POST
`/admin/datastore/<ds>/verify` → UPID), `Snapshots` (incl. the `verification` field),
`TaskStatus`/`WaitVerify` (node extracted from the UPID — `localhost` returns "unknown", the
spike B4 gotcha), `NodeFromUPID`.
- **The verify maintenance loop** (`pbs/verify.go`) — the cheap, key-free, ciphertext-level
integrity check (§8) on its OWN cadence (default 6h, the 5th daemon goroutine). It is a
reporting/maintenance task like the slice-5 watchdog: it does NOT go through the reconcile
gate/journal. Each cycle: trigger verify → poll task → re-list snapshots → record
per-snapshot `verify_state`. A failed verify is logged loudly.
- **`PBSSnapshot` reporting** — filled the stub (`namespace`/`backup_type`/`backup_id`/
`backup_time`(RFC3339)/`size_bytes`/`owner`/`protected`/`encrypted` (from `files[].crypt-mode`)
/`verify_state` (ok|failed|**none** until verified)/`verify_upid`). New `PBSReporter`
collector seam + an in-memory `SnapshotStore`. Cross-repo golden (both repos, byte-identical)
+ bidirectional key-set tests; hub `handler.go` parses `pbs_snapshots` and logs a **failed
verify `[WARN]`** (loudest offsite-DR signal).
- **Truthful backup mode** (`backup/runner.go`) — `Backup.mode` now reflects the ACTUAL vzdump
mode read from the task log (`backup mode: <x>`), since PVE may downgrade snapshot→stop for a
stopped guest (spike B1); falls back to the requested mode if unparseable.
- **proxmox**: `Storage.Username` (parsed from the pbs storage config — the token id).
- **config** `BackupConfig.{PBSVerifyCadenceSeconds, PBSSecretDir}` (cadence 0→6h, <0 disabled).
- **`--selftest=pbs-verify`** — discover pbs storages → verify each → print the PBSSnapshot
records (covers the runbook's verify + list). Standalone on the host.
### Notes
- Backup/restore-to-PBS reuse Phase A with no change (the restore-test runs with
`source_tier="pbs"` when fed a pbs volid). Zero-knowledge holds: verify is ciphertext-level,
the encryption key is never read here, and the PBS server has no client key (spike B6).
- Daemon runs cleanly with no pbs storage / verify disabled. `go test -race` covers the new
goroutine. Slice-3/4/5/6A surfaces, goldens, and adversarial tests intact.
## v0.6.0-rc1 — slice 6 Phase A: backup + the self-restore-test (local target) (2026-06-09)
Phase A of the backup/restore slice (doc 03 §8) — the agent's guest-level backup layer and
the **self-restore-test**, which closes "a backup you haven't restored isn't a backup".
Everything here is BENIGN (backup, restore-to-NEW, scratch teardown): reuses the slice-4
classifier/gate/journal — no new destructive class, no new crypto. Local target only; PBS is
Phase B. Restore is to a NEW guest only (no overwrite). Backups are crash-consistent only
(app-consistency needs the controller quiesce, slice 8) — marked so in the report.
### Added
- **proxmox** (`mutate.go`/`query.go`): `DestroyLXC` (DELETE …/lxc/{vmid}?purge=1&destroy-
unreferenced-disks=1 → UPID; the scratch-teardown primitive); `VzdumpOptions.Notes` →
`notes-template` (verified on PVE 9.2.2); `LatestBackupVolID` (resolve a produced archive
from the backup-storage listing — the task status carries no result volid).
- **reconcile self-restore-test** (`restoretest.go`) — `Engine.RunRestoreTest`: pick a free
scratch VMID (configured band, excludes 9999; full band → skip, never out-of-band) →
**journal a Scratch-owned entry BEFORE any mutation** → restore-to-new → benign net
**link-down** SetConfig (so the clone can't conflict with a running source's MAC/IP; this
is test-safety, NOT slice-7 identity reset) → boot → verify **reaches `running`** → ALWAYS
teardown (defer; benign `ClassGuestDestroy` + agent-tagged-scratch provenance, gated). Runs
on the scratch VMID's queue lane. Reuses the journal/gate; result feeds the report.
- **Crash-safe recovery** (`recover.go`): a Scratch journal entry is resolved by TEARDOWN,
not by re-checking the restore sub-task's UPID — special-cased BEFORE the generic path
(else the restore task's OK would mark it succeeded while the guest leaks). `Recover` now
destroys a leaked scratch guest (idempotent: already-gone → clean; list-unreadable → left
in-flight for a later pass). `JournalEntry.Scratch` flag; `RecoverResult.ScratchClean/
ScratchDestroyed`. GuestAPI gains `RestoreLXC`/`DestroyLXC`/`GuestStatus`.
- **`internal/backup` package**: `BackupRunner.Backup` (vzdump + archive/size resolve +
bulk-volume gap — a mountpoint is UNCOVERED unless it carries an explicit `backup=1`, so
an unset `backup=` is reported uncovered too, the safe DR direction); `PickRestoreCandidate`
(newest backup); an in-memory `Store` (latest-backup-per-target + latest-restore-test)
implementing the hub `BackupReporter`/`RestoreTestReporter` seams; a cadence `Scheduler`
(default 24h; the fourth daemon goroutine; disabled cleanly when off/misconfigured).
- **hub report** (`report.go`): filled the `Backup` + `RestoreTest` stubs (`PBSSnapshot`
stays a Phase-B stub); collector `BackupReporter`/`RestoreTestReporter` seams. Cross-repo
golden updated in BOTH repos (byte-identical) + bidirectional key-set tests for
`backups[0]`/`restore_tests[0]`. Hub `handler.go` parses + persists them (report_json; no
new columns) and logs a **FAILED restore-test prominently** (the loudest DR signal).
- **config** `BackupConfig` (local target, restore storage, restore-test cadence, scratch
VMID band 990000990009 default) + accessors + env overlay + cadence-gated validation.
- **`--selftest=backup -vmid N`** (one-shot backup → print the Backup record) and
**`--selftest=restore-test [-archive volid]`** (Recover-then restore→boot→verify→teardown,
print the RestoreTest record). Standalone on the Proxmox host.
### Notes
- The daemon runs cleanly with the cadence off or misconfigured (logs + disables, never
crashes); a leaked scratch guest from a mid-test crash is reaped by `engine.Recover` on
restart. `go test -race` covers the new scheduler goroutine.
- Slice-3/4/5 exported surfaces, goldens, and adversarial tests intact. Version bumps to
**v0.6.0** when Phase B (PBS) lands.
## v0.5.1 — slice 5 live-validation prep: durable_id mis-id fix + re-mount UUID memory (2026-06-09)
Two correctness fixes surfaced while preparing the live USB validation on `demo-felhom`
(a real 1TB USB HDD, sdb1, ext4). Both are DR-load-bearing — exactly the "false-id →
re-attach the wrong disk" failure mode the slice warned about.
### Fixed
- **Unmounted dir-storage no longer inherits the ROOT filesystem's UUID** (`observe.go`).
Previously, when a removable dir-storage was unmounted, the observer fell through to the
*containing* mount (root) for the backing device, so its `durable_id` became
`uuid:<root-uuid>` — a catastrophic DR mis-id (the hub would re-attach the wrong disk).
Now the backing device/UUID/`durable_id` are derived ONLY from the target's OWN
mountpoint; an unmounted target reports no device and a stable `store:<name>` durable_id,
never another filesystem's UUID. (Removed the `containingMountDevice` root-fallthrough.)
- **Watchdog remembers the fs-UUID observed while attached** (`watchdog.go`) so a re-mount
works even after the known-set cache refreshes mid-drop (an unmounted target can't resolve
its own UUID). The re-mount key is backfilled from this memory — aligning with doc 03 §7's
"sourced from the existing definition, no hub manifest needed": the agent learns the UUID
while the target is attached, then re-mounts by it on return.
### Tests
- Observer: an unmounted dir-storage asserts NO `uuid:` durable_id and no backing device.
- Watchdog: a drop where the cache lost the UUID still re-mounts using the remembered UUID.
## v0.5.0 — slice 5 Phase B: the host-root surface (mounts + SMART + grow + destructive gate) (2026-06-09)
The write surface — the agent's first step outside its Proxmox API token into OS-root.
Isolated behind a narrow, argument-validated, adversarially-tested seam, exactly like the
slice-4 gate. Completes slice 5 (Phase A = read-only observe/report/watchdog at v0.5.0-rc1).
### Added
- **`HostOps` seam + `SudoHostOps`** (`internal/storage/hostops.go`) — the one privileged
host surface: persistent mounts via **systemd `.mount` units keyed by fs-UUID** (enabled to
survive reboot), detach (stop+disable), SMART, and thin-pool metadata. Shells out via the
fenced Runner (`sudo -n`, **fixed arg vectors, no shell**); a fake backs the tests (no real
root in the suite). `NoopHostOps` is the safe fallback when the surface is unavailable.
- **The argument validator** (`internal/storage/validate.go`) — the security boundary:
`ValidateUUID` (strict hex), `ValidateMountPath` (absolute, no traversal, no metacharacters),
`ValidateSMARTDevice` (raw-disk whitelist), `ValidateLVMName`, and an in-process
`systemdEscapePath` (no `systemd-escape` shell-out). **Every argument is validated BEFORE a
command is constructed.** Headline test (`validate_test.go`): an adversarial matrix of
shell metacharacters / `../` traversal / malformed inputs is rejected with **zero exec**.
- **SMART** (`internal/storage/smart.go`) — parses `smartctl -a -j` into `StorageTarget.smart`:
**SATA** (reallocated/pending/offline-uncorrectable, temp, power-on-hours) **and NVMe**
(critical_warning, media_errors, percentage_used, temp), degrading to `UNKNOWN` for devices
with no SMART (USB-SATA bridges). **`lvs`** fills the lvmthin thin-pool **metadata** fill
(the value Phase A left null). Wired into the Observer's enrichment (Observe only, not the
watchdog's fast Known path).
- **Watchdog re-mount response** (`internal/storage/watchdog.go`) — on a known mount-backed
target's device returning **unmounted** (a new `DevicePresent` liveness probe), the watchdog
**dispatches a benign by-UUID re-mount off the poll path** (a goroutine, never under the
lock), rate-limited per target to the debounce window. The mount is routed through the gate
as benign (`gateRemounter` in `main.go`, so `storage` stays decoupled from `reconcile`).
- **Disk-grow executor** (`internal/reconcile`) — `ActionResize` (benign `ClassResize`), planned
**grow-only** (desired DiskBytes > actual → `pct resize rootfs +<n>M`; a shrink is refused,
never silently grown) + a defensive executor guard (size must start with `+`). New
`proxmox.Client.ResizeLXC` (API; `VM.Config.Disk`+`Datastore.AllocateSpace`; async→UPID).
Built + fixture-tested; **unfed** live (no hub spec until slice 10).
- **Destructive storage ops through the slice-4 gate** (`internal/reconcile/storage_ops.go`) —
`IntentForStorageMount` (benign) and `IntentForStorageDestructive` (`ClassStorageWipe`/
`ClassDecommission`). Host/target-scoped: the op binds on the storage **target identity**
(carried in `target.guest_id`). Reuses the existing verifier/role-scoping/binding/audit — no
new gate, no new crypto. Storage cases added to the adversarial matrix (`storage_test.go`):
unsigned wipe → `pending_signature`; "wipe A" signature vs "wipe B" → `binding_mismatch`;
valid → accepted. **Inert** live.
- **`--selftest=storage` [`-watch <dur>`]** — the live USB-runbook harness: an observe pass
(full table incl. SMART + thin-pool data+metadata), and a bounded watchdog window with the
re-mount response live. Runs standalone on the Proxmox host (no hub).
- **`configs/felhom-agent.sudoers`** — the documented narrow allowlist (install unit / systemctl
manage / smartctl / lvs), with the agent-side fine validation noted.
- **Config**: `privileged.{unit_dir,stage_dir,systemctl,install,smartctl,lvs}` (paths must match
the sudoers entries).
### Notes
- Daemon still runs cleanly with no removable storage / no signers / no hub manifest, and a
missing/declined sudoers entry degrades with a warning (SMART→UNKNOWN, mount→logged error),
not a crash. `go test -race` passes (the watchdog re-mount dispatches off the poll path).
- Slice-3/4 + Phase-A exported surfaces, goldens, and adversarial tests intact. `authz`
untouched. The destructive-storage executor + grow are built/tested but unfed live until
slice 10.
## v0.5.0-rc1 — slice 5 Phase A: storage observe + report + watchdog (read-only, live) (2026-06-09)
Phase A of the storage slice (doc 03 §7). Read-only and live: the agent now observes every
host storage target, reports it into the host-report's `storage_targets` (previously an empty
stub), and runs a fast-poll watchdog that pushes a disconnect to the hub in seconds. No
host-root writes this phase (mounts/SMART/grow/destructive-gate are Phase B). The hub-owned
desired manifest (class/role/policy/creds) is not served until slice 10, so reconcile against
it is built-but-unfed — this phase ships only the genuinely-useful read-only footprint.
### Added
- **`internal/storage` package** (new):
- **`StorageTarget` wire contract** (`internal/hub/report.go`) — filled the slice-3 stub:
`name`/`type`/`durable_id`/`state`/`reachable`, usage (`total`/`used`/`avail`/
`used_fraction`), `content`, `mount_path`/`backing_device`, `class_hint` (rotational HINT
— never authoritative; class is hub-owned), `role` (empty until slice 10), a `thin_pool`
sub-object (lvmthin data fill; metadata fill is Phase B/`lvs`), and a `smart` sub-object
(`UNKNOWN` until Phase B). Cross-repo golden kept byte-identical with `felhom.eu/hub` and
guarded by the bidirectional key-set test (`contract_test.go`).
- **`durable_id` derivation** (`durableid.go`) — deterministic per type (the DR-load-bearing
re-attach key): fs-UUID (usb/local-dir), `server:export` (nfs/cifs), `repo+fingerprint`
(pbs), `vg/pool` (lvmthin); never empty (falls back to a stable store id).
- **`HostReader` seam + `ProcHostReader`** (`hostread.go`) — non-privileged `/proc/mounts`,
`/dev/disk/by-uuid`, `/sys/.../rotational` + `removable` reads. Root-free by construction.
- **`Observer`** (`observe.go`) — builds `[]hub.StorageTarget` from `ListStorage`/`NodeStorage`
joined with host reads; surfaces the lvmthin thin-pool data fill prominently (warns ≥85%).
- **Storage watchdog** (`watchdog.go`) — a third daemon goroutine fast-polling the *known*
target set (a defined Proxmox storage and/or a previously-seen one) for
`attached↔disconnected` transitions; on a transition it triggers an immediate, **debounced**
out-of-band host-report. Only flags a *known* target's change (never a never-attached
device); coalesces flaps within the debounce window (leading + trailing edge).
`CachingKnownTargets` rate-limits the Proxmox-derived known set; `HostLiveness` probes
device/mount presence (local) + a reachability dial (network), all non-privileged.
- **Proxmox `Storage` type** (`internal/proxmox/types.go`) — additive parse-only config fields
(`server`/`export`/`share`/`datastore`/`fingerprint`/`vgname`/`thinpool`) feeding durable_id.
- **Collector `StorageObserver` seam** (`internal/hub/collect.go`) — populates `storage_targets`
via the observer; a nil observer or an observe error degrades to empty (never sinks the
heartbeat). Hub does not import storage (storage imports hub for the wire type).
- **Out-of-band report trigger** (`internal/hub/loop.go`) — `Loop.SetTrigger`: a watchdog
signal runs one extra collect→report immediately without disturbing the regular cadence.
- **`StorageConfig`** (`internal/config`) — watchdog interval / debounce / known-refresh knobs
(all optional; package defaults otherwise).
- **Hub ingest** (`felhom.eu/hub`) — `hostReportPayload` now parses `storage_targets`
(full mirror struct), persists them via `report_json`, counts + warns on disconnected
targets, and has its own half of the bidirectional golden key-set test.
### Notes
- The daemon still runs cleanly with no removable storage, no signers, and no hub manifest —
the watchdog finds nothing to flag; storage reporting is best-effort.
- `proxmox`/`hub`/`authz`/`reconcile` exported surfaces + their golden/adversarial tests are
intact. No host-root writes, no destructive paths, no SMART this phase (all Phase B).
- Version: **v0.5.0-rc1** at the Phase-A checkpoint; **v0.5.0** when Phase B lands.
## v0.4.0 — slice 4 Phase B: reversibility gate + signed-op consuming layer (2026-06-08)
The security core of slice 4: hub-supplied intent stops being trusted for destructive
change. Layered in front of the per-guest queue's executor — **every** mutation now
passes the gate. Reuses `internal/authz` for all crypto (untouched surface). Inert
this slice: no destructive deltas are served until slice 10, so the destructive path is
classified, gated, and adversarially tested but not wired to live execution.
### Added
- **Classifier (`classify.go`, doc 03 §4)** — benign vs destructive by **provenance +
data-bearing-ness, NOT by verb**. The `OpClass` vocabulary (seeded by the committed
slice-2 `op_blob.json`: `guest_destroy`) is the agent-side contract slice 10 matches.
Destroy/overwrite of customer data is destructive UNLESS **agent-internal**
provenance (same-journaled-transaction create → compensating rollback, or
agent-tagged scratch) makes it benign. `Provenance` is journal-recorded and **never
populated from the hub** (its zero value is the only thing an external intent may
carry). Unknown op class fails safe → destructive.
- **Reversibility gate (`gate.go`)** — `Gate.Authorize(intent, signed)`: benign →
allowed unsigned; destructive → requires a verified, role-authorized, action-bound
operator signature, else refused **`pending_signature`**, never executed. Every
decision is written to an `AuditSink` (audit is a signal, never the guard).
- **Signed-op consuming layer over `authz`** — verifies via `authz.Verifier.Verify`
(the locked pipeline, untouched), then enforces on the `VerifiedOp`:
- **Role-scoping (doc 04 §4)** — recovery key authorizes key-rotation re-pins ONLY;
operational key authorizes ordinary destructive ops + planned rotation.
- **Op-to-action binding** — verified `op` + host + guest + `params` must match the
gated action (a signature for guest X / op A can't authorize guest Y / op B);
params compared semantically (key-order/whitespace independent).
- **Signed-job orchestration (`job.go`)** — `RunSignedJob`: idempotency dedupe (the
op nonce as the journal key — a redelivered completed op is skipped, not re-run),
gate authorization, then journal-wrapped execution via an injected
`DestructiveExecutor` (nil this slice — authorized destructive ops are inert, no
executor wired until 6/7).
- **Crash-recovery consumer (`recover.go`, Note 1 / doc 03 §10)** — `Engine.Recover`
consumes the journal's `InFlight()` at startup: an op that crashed AFTER the Proxmox
POST and BEFORE its terminal record (`OpTaskRunning`, nonce already consumed) is NOT
covered by idempotency dedupe — only this resume-or-rollback resolves it (re-read the
task via the new `TaskStatusOnce`, record the real outcome; a no-task-id op is
abandoned fail-safe). Landed together with the signed-op executor, as Note 1 required.
- **Daemon wiring** — `runDaemon` builds the verifier from `config.Authz.Signers` (a
bad key / missing nonce-store path is a fatal misconfig; **no signers = nil verifier**,
the common slice-4 state), constructs the gate (+ `SlogAudit`), runs `Recover` before
issuing any mutation, and routes every reconcile action through the gate.
### Changed
- **Memory comparison canonicalized (Note 2)** — `desiredMemoryMiB` makes the
desired↔actual memory compare in the same MiB unit that is then written, so a
non-MiB-aligned `MemoryBytes` converges in one pass instead of re-issuing SetConfig
forever (the numeric cousin of the description-newline normalization). Test proves
convergence. Slice 10 should still serve MiB-aligned specs at the source.
### Tests (the security proof — each independently rejected)
- **Adversarial matrix** via the REAL `authz.Verifier` with in-test-minted SSHSIGs
(framing replicated in reconcile's test binary; production authz untouched, no signing
added to the verify-only package): unsigned destructive **job** → pending_signature;
unsigned destructive **desired-state delta** → pending_signature (distrusts hub
desired state, not just jobs); forged/unknown signer → `ErrUnknownSigner`; expired →
`ErrExpired`; **replayed nonce across an agent restart** (durable `FileNonceStore`) →
`ErrReplay`; wrong host → `ErrTarget`; wrong guest / wrong op / wrong params →
binding_mismatch; **recovery key on ordinary destructive** → role_denied;
**hub-supplied "scratch" tag ignored** → still destructive → refused; **valid + role +
target + fresh nonce → accepted**, and a second presentation → `ErrReplay` (nonce
consumed).
- Classifier (benign/destructive/provenance/key-rotation/fail-safe), role-scoping,
params binding, crash-recovery (resume OK / fail / still-running / no-task rollback /
unreadable / one-shot key applied on resume), signed-job idempotency (execute once,
dedupe redelivery, refused-not-executed, no-executor-inert, executor-error).
- Full module **race-clean** (`go test -race`) + vet clean on the Linux build server.
## v0.4.0-rc1 — slice 4 Phase A: reconcile engine (structural; runs live, unfed) (2026-06-08)
The agent-side control core's structural half. **Checkpoint marker** — `-rc1` is the
Phase-A push; awaiting validation before Phase B (the reversibility gate + signed-op
consuming layer) lands the final **v0.4.0**. Runs LIVE but UNFED: with no desired-state
provider until slice 10, the live engine computes an empty action set and performs
**zero mutations**.
### Added
- **`internal/reconcile`** package — the engine, the per-guest serializer, the
desired-state model, the normalization layer, and the durable op journal:
- **Per-guest serializer (`Queue`, doc 03 §10)** — the single choke point ALL
mutation sources funnel through. Same-vmid jobs run strictly one-at-a-time in
submit order; independent vmids run in parallel. Each vmid is a cond-var FIFO lane
(unbounded, non-blocking, order-preserving); graceful drain on `Close`.
- **Desired-state model + `DesiredProvider` seam** — `DesiredGuest` (per-field
optional: run-state / `*hub.GuestSpec` / `*description`), `DesiredState`. The only
live provider is **`EmptyProvider`** (slice 4 has no source); `StaticProvider`
feeds fixtures. The seam is where slice 10's hub-serving plugs in — no hub/local
source invented here.
- **Normalization layer (`FieldNormalizers`)** — reconcile compares *normalized*
desired-vs-actual so Proxmox round-trip quirks don't read as drift. `description`'s
trailing newline is the first registered case; the registry takes more (boolean
coercion, list ordering) as discovered. `normDesc` **promoted** out of
`cmd/felhom-agent/main.go` to **`reconcile.NormDescription`**; the `--selftest=task`
description round-trip now uses that shared helper (one source of truth for the quirk).
- **Plan engine (`Plan`, pure function)** — computes the minimal **benign** action set
(`Start`/`Stop`/`SetConfig`) for guests present in both desired and actual, with
normalized comparison, deterministic vmid ordering, config-before-run-state. Skips
provision (desired-absent-in-actual, slice 7) and destroy (actual-absent-in-desired,
gated, slice 10); never writes a config it couldn't first read (`SpecKnown`). Disk
(rootfs grow) intentionally not reconciled here.
- **Reconcile engine (`Engine`)** — reads desired+actual, plans, dispatches each action
onto the shared queue. Every Proxmox op handled per the mutate.go contract: non-empty
UPID → `WaitTask` + assert `exitstatus`; empty UPID → clean **synchronous** success
(slice-4 proven). Per-action failures are counted, not fatal (other guests still
converge).
- **Operation journal (`Journal`)** — durable fsync'd append-only JSONL mirroring
`authz.FileNonceStore`: records each op's lifecycle (started → task_running →
succeeded/failed) with its Proxmox task id (crash mid-op is detected and re-checkable
on restart via `InFlight()`), plus an **idempotency-key store** (`AlreadyApplied`) so
a one-shot op never re-runs across retries/restarts. Reconcile actions carry no
idempotency key (convergent — must re-run on real drift).
- **Daemon wiring (`runDaemon`)** — reconcile runs alongside the hub loop on the poll
cadence, **sharing the per-guest queue**. Journal path is a `journal.log` sibling of the
nonce store. The daemon runs cleanly with **no desired state and no signers** (reconcile
is a logged live no-op; a journal-open failure degrades to journal-less, never crashes).
### Tests
- Serializer: same-guest serialized (max-concurrency 1, submit order preserved) and
different-guests parallel (cross-waiting jobs both complete — would deadlock if not);
error propagation; drain-pending-on-close; submit-after-close.
- Normalization: description round-trip; unknown-field identity; extensibility seam
(synthetic boolean-coercion + list-ordering normalizers).
- Plan: run-state start/stop, spec drift (cores/memory), disk-not-reconciled,
description-newline-not-drift, unmanaged fields, spec-unknown skips config keeps
run-state, desired-absent skipped, combined ordering, empty-desired no-op, deterministic
vmid order.
- Engine: empty-provider zero mutations; async start (WaitTask); synchronous SetConfig
(no WaitTask); WaitTask failure + POST error counted failed; list error = pass failure.
- Journal: lifecycle latest-wins; in-flight survives restart; idempotency dedupe across
restart; failed key not applied; torn-trailing-line skipped.
- Full module **race-clean** (`go test -race`) on the Linux build server; vet clean.
### Not in this phase (Phase B)
- The benign/destructive classifier, the reversibility gate, and the signed-op consuming
layer over `internal/authz` (doc 03 §4 / doc 04) — added next, in front of the queue's
executor, landing **v0.4.0**.
## v0.3.2 — SetConfig selftest extension (slice-4 pre-check) (2026-06-08)
The gate before slice 4: prove `SetConfig` works live under the scoped token before
reconcile is built on it. **Self-gated live run PASSED** on `demo-felhom`/guest 9999.
### Added
- **Reversible `SetConfig` step appended to `--selftest=task`** (`cmd/felhom-agent/main.go`,
`selftestSetConfig`): read `GuestConfig` → write a `description` marker
(`felhom-selftest <RFC3339>`) → verify it landed → restore the original value (or
`delete` the key if it was absent) → verify the restore. Handles PVE's dual-mode
`SetConfig` return per the `mutate.go` contract: empty UPID = synchronous success
(printed `synchronous`); non-empty UPID = `WaitTask` + assert `exitstatus=OK`.
The existing snapshot → rollback → delete-snapshot steps are unchanged. First live
exercise of the **`VM.Config.*`** privilege cluster.
- **`normDesc` / `extraString` helpers** — `extraString` decodes a string-valued key
from `GuestConfig.Extra` (raw JSON); `normDesc` strips the trailing newline PVE
appends to `description` on read, so a written value round-trips equal.
### Finding (live)
- The LXC `description` write returned **synchronous (empty UPID)** — PVE applied it
inline, no task. The agent's dual-mode `SetConfig` modeling is correct: the
empty-string path is real and must not be treated as an error.
- PVE **appends a trailing `\n` to `description`** on read (stored URL-encoded as
`%0A`). A naive exact-match reconcile would see perpetual drift — slice-4 reconcile
must normalize `description` comparisons (hence `normDesc`).
### Ops
- Standing operator token (`felhom-agent@pve!agent`, privsep) **rotated** during this
run (the prior secret was not retrievable); role + both user/token ACL rows
re-confirmed at `/`. New secret stored out-of-band, **not persisted to the repo**.
Guest 9999 left pristine (stopped, no `description`, no leftover snapshot). Version → 0.3.2.
## Docs + live validation — no version bump (2026-06-08)
### Changed
- **Reflowed `CLAUDE.md`** — removed hard mid-paragraph line wraps (prose, list items, blockquotes now single-line, soft-wrapped); code blocks and tables untouched; rendered output unchanged.
- **Unified the REPORT/CHANGELOG convention** in `CLAUDE.md`: `CHANGELOG.md` is the cumulative log (newest on top); `REPORT.md` is overwritten with the most-recent implementation/validation only. Added an explicit **no-secrets** rule (never write tokens/passwords/keys into committed files; reference them as stored out-of-band).
### Added
- **`REPORT.md`** rewritten for the live `--selftest=task` validation on the demo host (`demo-felhom`): snapshot → rollback → delete-snapshot on guest 9999, each polled to `exitstatus=OK` under the `felhom-agent@pve!agent` privsep token (UPIDs name the token actor — privsep path genuinely exercised); 16-privilege `FelhomAgent` role + both user & token ACLs confirmed; `--selftest=read` clean. Closes the slice-1 "mutating ops unit-tested only" gap; `WaitTask` async foundation validated live → **slice 4 unblocked**. (Token secret stored out-of-band, not in the repo.)
## v0.3.1 — slice-3 validation follow-ups (2026-06-08)
### Changed
- **Collector keeps the known run-status on a `GuestConfig` failure** (`internal/hub/collect.go`):
previously a per-guest config-read error forced `status="unknown"`; now the run-status from
`ListLXC` is preserved (only the `spec` is dropped). An empty status is still normalized to
`unknown` (wire value is always `running|stopped|unknown`). Test renamed to
`TestCollect_GuestConfigFailureKeepsStatusOmitsSpec` and asserts the preserved `running` + nil spec.
- **`--selftest` usage** error string now reads `(want read|task|hub)`.
### Added
- **Cross-repo contract fixture** `internal/hub/testdata/host-report.golden.json` +
`TestHostReport_ContractMatchesGolden` — compares the marshaled `HostReport` field-name sets
(top level + `host` + `guests[0]`) against the golden, failing on any json-tag drift. The file is
**kept byte-identical** with felhom-hub's copy (duplicated contract until a shared types module;
revisit when slices 5/6 populate the empty collections). Version → 0.3.1.
## v0.3.0 — hub client + host-report + first daemon loop (slice 3) (2026-06-08)
The agent's first daemon: a periodic read-only host-report POSTed to the hub (the
heartbeat). No Proxmox mutations, no desired-state/signed-op consumption, no
storage/backup collection yet — those are slices 4/5/6.
### Added
- **`internal/hub`** package:
- **`HostReport`** wire contract (`report.go`) shared field-for-field with the hub
ingest: host metrics, guests (`vmid` + spec), `cloudflared` status, and the
`storage_targets`/`backups`/`restore_tests`/`pbs_snapshots`/`audit_tail`
collections **defined but emitted empty** (typed `[]`, slices 5/6 fill them).
- **`Collector`** (`collect.go`) builds the report from a read-only `proxmoxReader`
(adapted to the real `internal/proxmox` surface — node held by the client, value
returns, `proxmox.Guest`) + a `CloudflaredProber`. Partial-failure policy: a
failed `NodeStatus` is a hard error (skip the POST); a failed per-guest
`GuestConfig` degrades that guest to `status="unknown"` (spec omitted) but still
sends; a cloudflared probe failure → `"unknown"`, never fatal.
- **`CloudflaredProber`** + `SystemctlProber` (`systemctl is-active cloudflared`;
read-only — NOT a Privileged/root op; tunnel management is a later slice).
- **`Client`** (`client.go`): `POST /api/v1/host-report` with
`Authorization: Bearer <key>`, standard TLS (system roots or optional `ca_file`;
verification always on). Typed `*TransportError` / `*HTTPError`; the bearer token
never appears in any error.
- **`Loop`** (`loop.go`): the daemon — immediate first report then tick; adopts the
hub's `poll_interval_seconds` clamped to [60,3600]; resilient (a collect/report
error is logged and the loop continues); clean shutdown on context cancel.
- **`ControlEnvelope`**: only `poll_interval_seconds` is acted on; `blocked` /
`desired_generation` / `has_signed_ops` are parsed-but-ignored (logged at most)
pending reconcile (slice 4).
- **Config**: `HubConfig` (url/host_id/api_key/poll_seconds/timeout_seconds/ca_file),
`FELHOM_AGENT_HUB_*` env overlay, `HubConfig.Validate()` (mode-aware — proxmox-only
`--selftest=read|task` still runs without hub config), `WithDefaults()`, and
`Redacted()` now also blanks the hub key. `configs/agent.example.json` gains `hub`
(and `authz`) blocks.
- **`cmd/felhom-agent`**: the no-`--selftest` mode is now the **daemon** (poll loop);
added **`--selftest=hub`** (one collect+report, prints the report + envelope).
Version 0.2.0 → 0.3.0.
### Tests
- Report serialization (field names; empty collections are `[]` not `null`; spec
omitted when unknown); client (Bearer header, non-2xx→`*HTTPError`,
transport→`*TransportError`, **token never in error**); collector (host mapping,
guest spec, per-guest failure degrades-but-still-reports, NodeStatus hard error,
cloudflared error→unknown); loop (immediate first report, continuation after an
injected error, interval adoption + clamp); config (hub validate/redact/env).
### Notes
- `internal/proxmox` and `internal/authz` were **not touched** — no new proxmox
surface was needed (`ListLXC` already exposes status/maxmem/maxdisk; `GuestConfig`
exposes cores). The task's `proxmoxReader` sketch (node-arg/pointer/`LXC`) was
adapted to the real exports as instructed.
- **Defined-but-empty** this slice: `storage_targets`, `backups`, `restore_tests`,
`pbs_snapshots`, `audit_tail` (slices 5/6). **Parsed-but-ignored**: the envelope's
`blocked`/`desired_generation`/`has_signed_ops` (slice 4).
## v0.2.0 — `authz` signed-op verifier (slice 2) (2026-06-08)
Production form of the Phase-4 signing primitive: a key-type-agnostic SSHSIG
verifier for operator-signed destructive ops, with the full anti-replay/
authorization pipeline and a durable, crash-safe nonce store. What slice 4
(reconcile) will call to gate destructive desired-state deltas. No hub, no signing
CLI, no reconcile loop.
### Added
- **`internal/authz` — `Verifier`**: `New(signers, store, hostID)` + `Verify(blob,
sigArmored) (*VerifiedOp, error)`. Runs the LOCKED pipeline (order is
load-bearing): parse armor → namespace → parse pubkey → allow-list (by key
**material**, `pub.Marshal()` equality, not key_id) → crypto verify (over the
**raw received bytes**, never re-canonicalized) → parse blob → target → time
window → **nonce recorded LAST**. Each post-crypto stage rejects even with a
valid signature.
- **SSHSIG framing** (`sshsig.go`) via `golang.org/x/crypto/ssh` — `pem.Decode` →
strip 6-byte magic → `ssh.Unmarshal` → `ssh.ParsePublicKey` → recompute signed
data with the named hash → `pub.Verify` (dispatches on key algorithm). No
hand-rolled crypto. Key-type-agnostic: ed25519 / **sk-ssh-ed25519 (FIDO2)** /
rsa / ecdsa via the one path.
- **Fixed namespace** `felhom-op-v1` (package constant, never caller-supplied).
- **`OpBlob`** (corrected `host_id`/`guest_id` json tags) + **`VerifiedOp`** (op,
host/guest, params, key_id, matched signer). key_id is advisory/audit only —
never an authz input.
- **Typed errors**: `ErrMalformed, ErrNamespace, ErrUnknownSigner, ErrBadSignature,
ErrTarget, ErrExpired, ErrNotYetValid, ErrReplay` (errors.Is-friendly).
- **`NonceStore`** + two impls: `MemoryNonceStore` (tests) and **`FileNonceStore`**
— durable, crash-safe (fsync'd append log, replayed into an index on open,
periodic compaction, expiry-only pruning). A nonce is fsync'd to disk before
`SeenOrRecord` returns false; replay protection survives restart; I/O failure
fails safe (reports seen=true). Target generalization: host_id matched strictly,
guest_id surfaced for the caller to route.
- **Config**: `AuthzConfig` (nonce-store path + pinned operator `signers` tagged
`operational`/`recovery` with a key_id, as authorized_keys lines).
- **Version 0.2.0.**
### Tests
- Real OpenSSH interop via a committed `ssh-keygen -Y sign` vector (hermetic CI);
per-stage rejection (each with an otherwise-valid sig); the headline
**invalid-sig-does-not-burn-the-nonce** invariant; replay; **persistence across
restart**; synthetic **sk-ssh-ed25519** through the unchanged path; byte-exactness
(a re-serialized blob fails crypto — not re-canonicalized).
### Notes / corrections to the Phase-4 reference
- §7's `Target` lacked json tags (`host_id`/`guest_id`) — fixed.
- The doc paired "Go 1.24.4 / x/crypto v0.52.0", but v0.52.0 declares `go 1.25.0`
and does **not** build on Go 1.24. Resolved by upgrading the build server to
go1.26.0 (backward-compatible; felhom-controller/hub unaffected); the module is
`go 1.25.0` on x/crypto v0.52.0.
- Free function → constructed `Verifier`; returns the full `VerifiedOp`; typed
errors; clock-skew tolerance added; durable nonce store is the net-new work.
- **Shared-contract dependency flagged** (not built): the hub and the `felhom-sign`
CLI must emit byte-identical canonical JSON or signatures won't verify; a shared
canonicalizer both import would be the right home.
## v0.1.0 — Scaffold + `proxmox` interaction layer (slice 1) (2026-06-08)
First slice: stand up the host-agent project and its foundation — the typed
Proxmox interaction layer every other module will call. No reconcile loop, hub
client, signing, or storage/backup orchestration yet (later slices).
### Added
- **Project scaffold**: module `gitea.dooplex.hu/admin/felhom-agent`, binary
`felhom-agent` (`cmd/felhom-agent/`), Go 1.24, zero external dependencies
(pure stdlib). `--version` flag; `version` var overridable via
`-ldflags "-X main.version=<v>"`.
- **`internal/proxmox` — API backend (`Client`)**: hand-rolled REST client over
`https://<host>:8006/api2/json` with `PVEAPIToken` auth. Typed read ops
(`Version`, `Nodes`, `NodeStatus`, `ListLXC`, `GuestStatus`, `GuestConfig`,
`ListStorage`, `NodeStorage`, `StorageContent`) and async mutating ops
returning a UPID (`RestoreLXC` — the primary create path, `Vzdump`, `Snapshot`,
`Rollback`, `DeleteSnapshot`, `SetConfig`, `Start`, `Stop`).
- **`WaitTask`**: polls `GET /nodes/{node}/tasks/{upid}/status` until stopped, then
asserts `exitstatus == "OK"` (authorization can surface at task execution, not
the POST — phase1-2 §1.3). Exponential backoff (1s→5s cap), context
cancellation + timeout. `*APIError` parses the offending privilege from a 403;
`*TaskError` parses it from a failed task exitstatus + log tail.
- **`internal/proxmox` — fenced root-CLI backend (`Privileged`)**: limited to the
three proven OS-root exceptions only — `CreateGoldenLXC` (keyctl `pct create`),
`MountUSBByUUID`, `SMART`, `Sensors`; each cites why it can't be the API. Fence
is structural (Client never shells out, Privileged never makes an HTTP call) and
asserted in tests.
- **TLS trust**: SHA-256 leaf-cert pinning (the host serves a self-signed cert) or
a CA file; an explicitly-named `insecure_skip_verify` that is off by default. No
blanket verification disable.
- **`internal/config`**: JSON config file + `FELHOM_AGENT_*` env overrides; the
token secret is never logged (`Redacted()`).
- **`internal/log`**: slog setup (text, stderr, configurable level).
- **`cmd/felhom-agent --selftest`**: read-only health report against a live host
(version/nodes/status/guests/storage); `--selftest=task --vmid N` exercises
`WaitTask` on a reversible snapshot→rollback→delete op (gated; default selftest
mutates nothing).
- **Tests**: unit tests with a mock HTTP transport + mock runner (UPID parse,
`WaitTask` running→OK / failed-403 / timeout / ctx-cancel, 403→privilege error,
response decoding against shapes captured live from `demo-felhom`, config
redaction, and the API-vs-root routing fence).
### Notes
- Types are grounded in the spike findings
(`felhom.eu/documentation/proxmox-platform.md`, `tests/phase{0,1-2,3}-findings.md`)
and the exact JSON shapes captured live from `demo-felhom` (PVE 9.2.2).
- Verified: `go build/vet/test` green on Go 1.24.4 (build server) and a live
read-only `--selftest` against the demo host with TLS fingerprint pinning.
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
token) is provisioned out-of-band; the agent only consumes the token.