77fa5af592
After a guest DHCP IP move, the split-horizon resolver kept serving the old IP: the drop-in (address=/domain/ip) updated but 'systemctl reload dnsmasq' (SIGHUP) does NOT re-read /etc/dnsmasq.d config — only /etc/hosts + cache. Changed reload() -> restartDnsmasq() so address= changes actually take effect. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1311 lines
100 KiB
Markdown
1311 lines
100 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to **felhom-agent** are recorded here. Update on every code
|
||
change that gets pushed.
|
||
|
||
## v0.29.1 — lanresolver: RESTART dnsmasq on change (not reload) — fixes stale split-horizon IP (2026-06-13)
|
||
|
||
**Bug:** after a guest's DHCP IP moved (e.g. the v0.29.0 9201 re-provision: .151 → .141), the LAN
|
||
split-horizon resolver kept answering the OLD IP, so LAN clients (via Pi-hole's conditional forward to
|
||
the host dnsmasq) resolved `*.demo-felhom.eu` to the dead IP. Root cause: `lanresolver.Manager` updated
|
||
the per-customer drop-in (`address=/<domain>/<ip>`) correctly but then ran `systemctl reload dnsmasq`
|
||
(SIGHUP) — and **dnsmasq's SIGHUP does NOT re-read its config files** (`/etc/dnsmasq.d/*.conf`); it only
|
||
clears the cache + re-reads `/etc/hosts`/addn-hosts. So the changed `address=` directive never took
|
||
effect until a restart. **Fix:** `reload()` → `restartDnsmasq()` (`systemctl restart dnsmasq`) for every
|
||
config-drop-in change (ReconcileGuest IP change, EnsureDnsmasq base change, Remove/decommission). Restart
|
||
is sub-second and the records carry local-ttl 0, so downstream forwarders don't cache a stale answer.
|
||
(Live: after the fix + a one-time host dnsmasq restart + a Pi-hole cache flush, `*.demo-felhom.eu`
|
||
resolves to the live guest IP again; future IP moves now self-heal on the loop's next tick.)
|
||
|
||
## v0.29.0 — OS / Docker-data storage split: golden + provision (2026-06-13)
|
||
|
||
Phase 1 of the storage-split slice (Phase 2 = felhom-controller v0.58.0 prevention layer). The
|
||
controller guest's OS rootfs and Docker data are carved onto separate `local-lvm` volumes for
|
||
RESILIENCE — an isolated OS rootfs stays bootable + agent-recoverable if the Docker volume fills.
|
||
|
||
- **`configs/build-golden.sh` — split baked in:** `--rootfs ${ROOTFS_STORAGE}:${OS_SIZE_GB}` (default
|
||
**32**, was hardcoded 8) **plus** `--mp0 ${ROOTFS_STORAGE}:${GOLDEN_DOCKER_GB},mp=/var/lib/docker,backup=1`
|
||
(default 16). The baked controller + infra images land on the data volume and travel inside the
|
||
golden archive (no empty-volume shadowing, no deploy-time pull). `backup=1` is MANDATORY — extra LXC
|
||
mountpoints default to `backup=0` = EXCLUDED from vzdump (spike B3), which would drop the images from
|
||
the archive entirely. The script now also bakes Docker **log rotation** into `daemon.json`
|
||
(`max-size 10m`, `max-file 3` — prevention layer 2D), asserts `/var/lib/docker` is a separate mount,
|
||
and **aborts if vzdump excludes mp0**.
|
||
- **`internal/reconcile/bringup.go` — sized provision:** `GuestMount` gains `Backup` (emits `,backup=1`
|
||
— closes the spike-B3/B5 silent-DB-loss trap at the mount builder). `BringUpSpec` gains
|
||
`DataVolGrowGB` + `DataVolMount` (default `mp0`): provision GROWS the golden-carried Docker-data
|
||
volume online to the per-customer target (grow-only, spike B4) rather than attaching a fresh empty
|
||
volume that would shadow the baked images. Plus `RootfsGrowGB` for the OS rootfs.
|
||
- **CLI seam:** `--selftest=bring-up|provision` gain `-rootfs-grow` / `-datavol-grow` / `-datavol-mount`
|
||
flags. Per-customer sizing source = flags now, the slice-10 hub storage manifest later.
|
||
- **`RUNBOOK-provisioning-storage.md`** (new): the split provisioning procedure + fresh-PVE-install
|
||
thin-pool carving knobs (`hdsize`/`maxroot`/`maxvz`, spike B4) + the per-customer sizing seam.
|
||
- Tests: `buildBringUpConfig` backup=1 emission; bring-up issues rootfs + data-volume resizes.
|
||
|
||
## (no version) — storage OS/data-split spike findings (2026-06-13)
|
||
|
||
Investigation only — **no code changed**. Findings report: `REPORT-storage-split-spike.md` (gates the
|
||
provisioning spec for splitting the controller guest's OS rootfs from its Docker/data onto separate
|
||
`local-lvm` volumes). Proven on a throwaway unprivileged LXC (9300, since destroyed): Docker `data-root`
|
||
on a second `local-lvm` mountpoint works (overlayfs/ext4, no idmap issue, reboot-survives); the
|
||
move-then-verify migration is safe (copy-not-move). **Key finding:** additional LXC mountpoints are
|
||
**excluded from vzdump by default** — they need `backup=1` set **and a CT restart** — so the docker-data
|
||
mount must be attached with `,backup=1` or named-volume DBs silently fall out of PBS. The exact seam is
|
||
`internal/reconcile/bringup.go:313` (`buildConfigParams`), which today builds `mpN` without a `backup=`
|
||
flag; `GuestMount` should carry the flag. Per-customer sizes belong in the slice-10 hub storage manifest
|
||
(marked at `bringup.go:49-50`); the golden rootfs is hardcoded `8` at `configs/build-golden.sh:40`.
|
||
|
||
## v0.28.0 — backup re-target → felhom-pbs (offsite DR) + operator-signed decommission (2026-06-12)
|
||
|
||
**Whole-guest backup now defaults to the offsite PBS tier (real DR).** `BackupConfig.BackupTarget()`
|
||
returns the configured `backup.local_backup_target` or, when empty, the new default `felhom-pbs` — a
|
||
PBS datastore on SEPARATE HARDWARE (the DooPlex box), so a host disk/hardware failure no longer takes
|
||
the backups with it. The target stays fully configurable (set `local_backup_target` to `local`/other
|
||
to override); no call site hardcodes it. All `NewBackupRunner` sites (restore-test scheduler, local-API,
|
||
`--selftest=backup`/`restore-test`) route through `BackupTarget()`.
|
||
|
||
Proven live on demo-felhom before the re-point (PHASE 0 gate):
|
||
- snapshot-mode `vzdump → felhom-pbs` still fires the `create storage snapshot 'vzdump'` marker, so the
|
||
8B.2 early-resume/quiesce signal survives a PBS target (the marker is mode-driven, not target-driven);
|
||
- the restore-test enumerates PBS backups through the SAME generic `StorageContent`
|
||
(`/nodes/<node>/storage/felhom-pbs/content` returns `content:"backup"` + ctime/vmid/volid), so
|
||
`PickRestoreCandidate`/`latestArchive` need NO PBS-client change;
|
||
- `pct restore` from a PBS volid round-trips cleanly (storage.cfg encryption key applied transparently);
|
||
- PBS gotchas (`ignore-verified`, node-from-UPID, privsep) touch only the verify-API path, not vzdump/restore.
|
||
|
||
**Operator-signed `decommission` now reachable (slice 10 P3 completion).** The previously-unreachable
|
||
`IntentDecommissioned` state (no production caller) is now reached ONLY via a gate-VERIFIED operator
|
||
signature — never customer-confirmable, distinct from a safe eject. New `internal/signedjobs`
|
||
`DecommissionExecutor` (op `decommission`, classified destructive in `reconcile.Classify`) calls
|
||
`IntentStore.SetDecommissioned`, keyed by the drive's STORAGE durable-id (the watchdog's key, e.g.
|
||
`uuid:<fs-uuid>` — NOT the device-level `byid:/byuuid:` scheme `storage_wipe` uses), so the recorded
|
||
intent actually gates future remounts. New `ExecutorChain` lets the signed-jobs runner serve both
|
||
`storage_wipe` and `decommission`; the runner wiring moved below the intent-store open in `main.go`.
|
||
`felhom-opsign` builds decommission params from `-durable-id`. No controller/customer UI — the operator
|
||
path is hub jobs-queue → signed-jobs runner.
|
||
|
||
**Restore-test now boot-verifies slice-10 enrolled guests (bind-mount mountpoints).** A guest whose
|
||
data drive is a host BIND mount (slice-10 P2 `mp0`) could not be vzrestore'd by the privsep token
|
||
("restoring 'mpN' to bind mount is only possible for root") — so the restore-test failed for every
|
||
enrolled guest, regardless of backup tier (surfaced during the felhom-pbs live validation). The
|
||
restore-test now reads the SOURCE guest config (vmid parsed from the archive volid — PBS `ct/<vmid>/`
|
||
and vzdump `vzdump-lxc-<vmid>-` forms) and passes `RestoreLXCOptions.MountOverrides` that neutralize
|
||
each bind-mount `mpN` to a throwaway 1G volume on the restore storage (needs no root; the boot-verify
|
||
doesn't need the drive's data, and the host paths would otherwise collide). Storage-backed mountpoints
|
||
are restored normally; best-effort (an unreadable source config restores as-is). `proxmox.RestoreLXC`
|
||
gained `MountOverrides`. Verified live: restore-test from felhom-pbs of bind-mounted guest 9201 →
|
||
boot+running PASS.
|
||
|
||
## v0.27.0 — slice 10 P3: self-heal watchdog reconcile + 4-state intent model (2026-06-12)
|
||
|
||
The storage watchdog goes from detect-only → detect-and-reconcile: the agent autonomously re-mounts an
|
||
enrolled external drive that dropped out-of-band (the colleague's Proxmox unmount), gated by a persisted
|
||
INTENT model so it never auto-adopts an unknown drive or fights an official eject.
|
||
|
||
- **`internal/storage/intent.go` — `IntentStore`** — durable, **durable-id-keyed** (UUID/WWN, never
|
||
sdX/path), atomic-write 4-state model: `new` (not recorded → never auto-mount), `enrolled` (desired
|
||
mounted → reconcile drift), `ejected` (intentional unmount → leave alone), `decommissioned`
|
||
(permanent). `OnAbsent` clears `ejected`→`enrolled` so a replug auto-mounts (the replug rule).
|
||
Records intent ONLY through the official enroll/eject paths — an out-of-band unmount records nothing
|
||
and is healed. Tests cover the states, persistence, the replug rule, and the reconcile gate.
|
||
- **`watchdog.go` — intent-gated reconcile + flapping guard (3C)** — the re-mount candidate (device
|
||
present, not mounted) now fires ONLY for an `enrolled` drive (via `IntentReader`); a present→absent
|
||
transition (device gone) calls `OnAbsent`. Exponential backoff (`debounce·2^fails`) + an alert after
|
||
4 failed cycles + a hard stop after 8 (no infinite loop). Failure = "still not present a full backoff
|
||
window after we dispatched" (a slow async re-mount isn't miscounted). Tests: colleague-unmount→
|
||
reconciled; ejected/new/decommissioned→left alone; ejected→absent→replug→auto-mount; flapping→caps.
|
||
- **`internal/localapi`** — `POST /disks/guest-attach` records `enrolled`; `POST /disks/eject` records
|
||
`ejected` (BEFORE unmount, while the durable-id still resolves) via the new `IntentRecorder`. `main.go`
|
||
opens one `IntentStore` (`<StateDir>/drive-intents.json`) shared by the watchdog + local API; open
|
||
failure degrades to ungated legacy remount (logged).
|
||
|
||
## v0.26.0 — slice 10 P2 activation: guest-reboot endpoint (user-triggered drive activation) (2026-06-12)
|
||
|
||
A drive enrolled into a RUNNING unprivileged guest can't be live-activated (proven: `pct set` won't
|
||
hot-apply; `/proc/<pid>/root` bind → mount-locking refusal; `nsenter -m` loses the host source). So the
|
||
bind activates at the next guest boot. This adds the user-triggered restart path.
|
||
|
||
- **`POST /guest/reboot` (`internal/localapi`)** — self-scoped (vmid from token). Runs `pct reboot
|
||
<vmid>` **detached** (it blocks ~30s until the guest is back) and returns **202** immediately, so the
|
||
calling controller gets a clean response before the reboot takes it down (the agent is host-side and
|
||
survives). `GuestBinder.RebootGuest` over the fenced runner. Tests: `TestGuestReboot_Accepted`
|
||
(202 + RebootGuest invoked for the token's vmid), `TestGuestReboot_CrossGuest403` (body vmid mismatch
|
||
refused, no reboot). Pairs with controller v0.49.0 (pending-activation detection + "Újraindítás most").
|
||
|
||
## v0.25.0 — slice 10 P2: bind enrolled user-data drives into the guest (passthrough) (2026-06-12)
|
||
|
||
External user-data drives are mounted on the HOST but were never passed INTO the guest (diagnosed
|
||
Branch A), so apps silently wrote to the rootfs and the controller couldn't see them. This adds the
|
||
guest passthrough. Spike-proven on 9201 first (see REPORT / the usb-passthrough-spike findings):
|
||
`pct set` **bind form** (host path, never `storage:size`), `chown` to the guest base (idmap not clean
|
||
for mixed-ownership data), `shared:49` propagation host↔guest automatic.
|
||
|
||
- **`POST /disks/guest-attach` (`internal/localapi`)** — self-scoped (vmid from token). Binds an
|
||
enrolled drive's **felhom-data namespace** into the guest at `/mnt/<name>` (**Model A**: the
|
||
felhom-data dir is the bind source mounted AT `/mnt/<name>`, so only Felhom's namespace crosses into
|
||
the guest — the customer's other data on the drive never does). Idempotent (returns the existing slot
|
||
if already bound); picks the lowest free `mpN`; validates `where` is `/mnt/<name>` (no traversal).
|
||
- **`GuestBinder` (`internal/localapi/guestbind.go`)** — the host-root steps over the fenced
|
||
`proxmox.Runner` (same pattern as the provision back-half's bind): `mkdir -p <drive>/felhom-data` →
|
||
`chown 100000:100000` the namespace ROOT (not -R; per-app subdirs are chowned at deploy) → `pct set
|
||
<vmid> -mpN <drive>/felhom-data,mp=/mnt/<name>` (RW bind). The namespace is created fresh + uniformly
|
||
owned, which sidesteps the drive's pre-existing mixed-ownership data entirely.
|
||
- **Tests** — `TestGuestAttach_*`: free-slot selection (mp0 when mp9 taken), idempotency (no re-bind +
|
||
`already:true`), bad-path rejection (traversal/non-/mnt/multi-component), not-configured 503.
|
||
|
||
Pairs with felhom-controller P2C (enroll triggers attach) + the golden's `/mnt:rslave` controller bind
|
||
(P2B). Self-heal reconcile (P3) and dual-role (P4) follow.
|
||
|
||
## v0.24.0 — role-gate the eject path (system/backup mounts are unmount-protected at the agent) (2026-06-12)
|
||
|
||
Closes the eject gap in the storage-authorization redesign: `POST /disks/eject` now **refuses to
|
||
unmount a system or backup storage**, enforced at the agent — not just hidden in the controller UI.
|
||
A direct API call (or a compromised controller) trying to `eject {where:"/var/lib/vz"}` or the PBS
|
||
mount is refused 403; only `user-data` mounts are ejectable.
|
||
|
||
- **`handleDiskEject` (`internal/localapi/disks.go`)** — before `Unmount`, resolves the AUTHORITATIVE
|
||
protection role of the storage mounted at `where` (the agent's own storage-view + host-topology
|
||
classification, never the caller's claim) via the new `roleForMountPath`. Refuses (403, no
|
||
`Unmount`) unless the role is `user-data`. **Fails SAFE**: an unresolvable mount (view error or no
|
||
storage target at that path) → treated as protected → refused (the same most-protected-on-ambiguity
|
||
default the wipe gate uses). Mirrors the wipe path's "protected — eject refused by role" logging.
|
||
- **`roleForMountPath` + `hostReader` seam** — `roleForMountPath` keys `RoleForStorage` on the mount
|
||
path (the eject input), mirroring `deviceRole`. `Options.HostReader` (optional; defaults to the
|
||
production `*storage.ProcHostReader`) injects the root-free topology reader so the role-gate is unit-
|
||
testable. `handleDisks`/`deviceRole` now share the same seam.
|
||
- **Tests** — `TestEject_RoleGated` asserts a `system` and a `backup` mount are refused with **no
|
||
`Unmount`**, a `user-data` mount ejects, and an unresolvable mount fails safe to refused (the same
|
||
non-hollowness the wipe tests use). `TestEject_UnmountAndDependents` updated to a user-data target.
|
||
|
||
## v0.23.0 — device-ROLE classification + tiered storage-wipe gate (system/backup operator-only, user-data customer-confirmable) (2026-06-11)
|
||
|
||
The storage-authorization redesign (agent half). The gate's destructive-wipe path is now **tiered by
|
||
the device's protection ROLE**, which the agent classifies from its OWN inspection — never the
|
||
caller's claim (the storage analog of classify.go's data-bearing verdict).
|
||
|
||
- **`internal/storage/role.go`** — `DeviceRole` (`system` | `backup` | `user-data`) + the
|
||
authoritative classifier. `RoleForStorage` (storage-view targets) and `RoleForRawDevice` (a raw
|
||
device, e.g. a fresh disk in the init flow) map a device to its tier via `SystemDisks` (the
|
||
whole-disks backing `/`, `/boot`, `/boot/efi`, root-free reads). Rules: `pbs` → backup; `lvmthin` /
|
||
builtin `local` / nfs / cifs / unknown → system; `usb` / `local-dir` on a **non-system external
|
||
device** → user-data. **Fail-safe**: any ambiguity (system disks unknown, or an unrecognizable
|
||
device topology) → **system** (most-protected) — never silently user-data.
|
||
- **`GET /disks`** — each `DiskInfo` now carries `role`. The controller drives the UI from it
|
||
(system/backup get a lock + no destructive controls; user-data is customer-manageable).
|
||
- **Gate tier (`reconcile`)** — new `CustomerConfirmable` disposition + `Gate.AuthorizeStorageWipe`:
|
||
- role=**user-data** → **customer-confirmable**: allowed iff the request carries an explicit
|
||
customer confirmation **bound to the device's durable id** (the agent re-resolves the durable id
|
||
and matches; a confirmation for one disk can't wipe another). **No operator signature.** A
|
||
user-data drive is already within the in-guest controller's blast radius (it bind-mounts `/mnt`),
|
||
so customer-confirmation adds no new reach. Recorded in the **audit log** with the durable id
|
||
(`AuditRecord.DurableID`).
|
||
- role=**system**/**backup** → unchanged **operator-signature** (`pending_signature`). The
|
||
`confirmed` flag is **IGNORED** — a compromised controller asserting `confirmed:true` on a
|
||
protected device is refused **by role**. Every other destructive class (`guest_destroy`,
|
||
`decommission`, `restore_overwrite`, `key_rotation`) keeps operator-signature exactly as before.
|
||
- **`POST /disks/format`** — accepts `confirmed` + `durable_id` (inert for system/backup). The
|
||
data-bearing path tiers by role: user-data customer-confirmed → `mkfs`; user-data unconfirmed →
|
||
403 `needs_confirmation` (+ the durable id to confirm against, NOT an opsign command); system/backup
|
||
→ 403 with the operator-signature pending op (as before). Blank devices stay benign `mkfs`.
|
||
- **Tests** — `role_test.go` (demo-storage mapping + fail-safe), `storage_wipe_test.go` (the gate
|
||
refuses a `confirmed` wipe on system/backup → no exec; durable-id mismatch / missing-durable
|
||
refused; unknown role fails safe), and the localapi format-handler branches (user-data confirmed →
|
||
mkfs; user-data unconfirmed → needs_confirmation, no opsign; confirmed-but-protected → still refused).
|
||
|
||
Pairs with the controller's lockout + type-to-confirm UX + drive-list restyle.
|
||
|
||
## v0.22.0 — expose durable_id in GET /disks (enable controller-side guided storage) (2026-06-11)
|
||
|
||
One-line, read-only addition: `localapi.DiskInfo` gains `durable_id` (mapped from
|
||
`StorageTarget.DurableID`, e.g. `"uuid:<fs-uuid>"` for usb/local-dir). The de-privileged controller
|
||
cannot read a device's fs UUID itself, yet `POST /disks/assign` mounts strictly by UUID — so without
|
||
this it could not complete the guided init/attach flows. The controller strips the `uuid:` prefix to
|
||
get the assign key. No new privilege, no behaviour change to format/assign/eject or the data-bearing
|
||
gate. Pairs with `felhom-controller` v0.43.0 (the storage-management UI rebuild).
|
||
|
||
## v0.21.0 — agent-managed split-horizon LAN resolver (internal/lanresolver) (2026-06-11)
|
||
|
||
LAN clients can now reach their guest **directly** at the same public hostname with the same real
|
||
wildcard cert (no Cloudflare hairpin), via a host-side dnsmasq the agent manages. The host is the
|
||
stable anchor (static LAN IP); the guest stays DHCP/ephemeral and the agent tracks its live IP.
|
||
|
||
- **`internal/lanresolver`** — renders a dnsmasq base drop-in (bind to the host LAN IP, no-resolv,
|
||
upstreams) + a per-customer drop-in `local=/<domain>/` + `address=/<domain>/<guest-ip>`. The proven
|
||
two-line shape: `local=` makes dnsmasq authoritative for the zone so **AAAA returns NODATA** (no
|
||
Cloudflare-AAAA split-brain — the guest has only link-local v6), `address=` is the wildcard A; all
|
||
other names (and their AAAA) forward upstream unchanged.
|
||
- **`Manager`** ensures dnsmasq present (apt) + the base config + enabled, discovers the guest's live
|
||
IPv4 (`pct exec <vmid> -- ip -4 -o addr show dev eth0`) and domain (read from the guest controller's
|
||
pulled `controller.yaml` — the v2 bootstrap omits it), writes drop-ins **write-if-changed**, and
|
||
**reloads** (not restarts) dnsmasq. Tolerates the early-boot pre-lease window (empty IP → skip+retry,
|
||
never a blank record). Logs IP transitions.
|
||
- **`Loop`** — a 7th daemon goroutine: every interval (default 300s) it enumerates provisioned guests
|
||
(`/var/lib/felhom-agent/guests/<vmid>/`) and reconciles each, so the resolver follows DHCP IP changes.
|
||
Config `lan_resolver.{enable,host_ip,upstreams,interval_seconds}` (host_ip defaults to the local-API
|
||
bridge IP). `--selftest=lanresolver -vmid N`.
|
||
- **`configs/felhom-agent.sudoers`** — new `FELHOM_DNSMASQ` alias (apt install dnsmasq; install
|
||
felhom-*.conf drop-ins; systemctl enable/reload dnsmasq; rm felhom-*.conf; the two FIXED `pct exec`
|
||
reads). The agent never touches `/etc/resolv.conf` (host's own resolution unaffected).
|
||
- **Box-down robustness** is a documented **router config** (DNS = [host-IP primary, upstream
|
||
secondary]) so a box reboot degrades to the Cloudflare path, not total DNS loss — see REPORT install step.
|
||
- Spiked live on felhom-pve first (`:53` free, host IP static `192.168.0.162`, host DNS intact, full
|
||
loop from a real LAN client returned the guest IP + AAAA NODATA + the real wildcard cert `200 0`).
|
||
|
||
## v0.20.0 — golden: stacks-dir bind + per-guest hostname/CT name + bake base-infra images (2026-06-11)
|
||
|
||
Lockstep with `felhom-controller` v0.41.0 + a golden rebake. Changes in `configs/build-golden.sh` and
|
||
the provision path; no change to the proxmox/authz/token fences.
|
||
|
||
- **Section-G mount fix (the load-bearing one):** the in-guest controller writes app/infra compose
|
||
stacks under `/opt/docker/stacks` *inside its container*, but the baked controller-bootstrap `docker run`
|
||
never bind-mounted that path. So `docker compose up` (run by the GUEST daemon over the shared socket)
|
||
resolved every relative bind source on the guest filesystem — silently creating empty dirs — which
|
||
broke **every** bind-mounted stack (base infra AND customer apps like immich/nextcloud). The bootstrap
|
||
unit now `mkdir -p /opt/docker/stacks` and adds a **same-path host bind**
|
||
`-v /opt/docker/stacks:/opt/docker/stacks` (a named volume would NOT fix this). Empirically confirmed on
|
||
guest 9201 before writing the fix.
|
||
- **Per-guest container hostname (3A):** the bootstrap unit derives `customer.id` from
|
||
`/etc/felhom-bootstrap/bootstrap.json` with a portable `sed` parse (NO jq in the golden) and passes
|
||
`--hostname <customer-id>` to `docker run`, so the controller's `os.Hostname()` (its hub-reported
|
||
hostname) is the customer id, not the Docker container ID. Fail-safe: no parse → no `--hostname`.
|
||
- **Per-guest CT/LXC name (3B):** `--selftest=provision` now defaults `-hostname` to the (DNS-safe
|
||
sanitized) `-customer-id` when not given, so the bring-up's existing `SetConfig hostname` step
|
||
(`bringup.go`) names the CT meaningfully (e.g. `demo-felhom`) instead of inheriting the golden's
|
||
`felhom-golden`. New `sanitizeHostname` (lowercase, collapse invalid → `-`, trim, ≤63).
|
||
- **Bake base-infra images:** the golden now also pulls the three PINNED, PUBLIC base-infra images
|
||
(`traefik:v3.6.7`, `cloudflare/cloudflared:2026.6.0`, `gtstef/filebrowser:1.3.3-stable`) into its Docker
|
||
storage so the controller's first-boot bring-up is OFFLINE-capable. A hard gate (`docker manifest
|
||
inspect`) fails the bake early on a bad pin. Tags MUST match the controller's `internal/infra` constants.
|
||
|
||
## v0.19.0 — bootstrap contract v2: agent relays the hub retrieval passphrase (no host key in the guest) (2026-06-11)
|
||
|
||
Lockstep with `felhom-controller` v0.40.0. Fixes the onboarding 401: a freshly provisioned guest's
|
||
controller used to come up with the agent's **host** hub key baked in, which the hub's `/api/v1/report`
|
||
(customer-scoped auth) rejects. The agent now bakes a **v2 bootstrap** carrying only what the controller
|
||
needs to **pull** its own config from the hub — the agent never touches the customer-scoped key or CF
|
||
tokens.
|
||
|
||
### Changed — bootstrap contract `v1 → v2` (`internal/provision`)
|
||
- `SchemaV1 → SchemaV2 = "felhom.bootstrap/v2"`. **`DocCustomer`** drops `name`/`domain`/`email` (keeps
|
||
`id`). **`DocHub`** drops `api_key`/`host_id`, adds **`retrieval_password`** (the customer's hub
|
||
retrieval passphrase — SECRET). `DocLocalAPI` unchanged. The contract is byte-compatible with the
|
||
controller's `internal/bootstrap.Bootstrap` (cross-repo round-trip verified).
|
||
- `backhalf.go`: renders the v2 Doc; validation now requires `customer.id` + `hub.url` +
|
||
`hub.retrieval_password` (was `customer.id` + `customer.domain`). Write/0600/chown/`pct set` unchanged.
|
||
- `cmd/felhom-agent/main.go` `--selftest=provision`: **new required `-hub-password`** flag (the customer's
|
||
hub retrieval passphrase; the customer must already exist in the hub). Stops baking `cfg.Hub.APIKey` /
|
||
`cfg.Hub.HostID`. `-customer-domain/-name/-email` still accepted (bring-up may use them) but NOT baked.
|
||
|
||
### Changed — `configs/build-golden.sh`
|
||
- Default `CONTROLLER_IMAGE` bumped off the stale `:v0.35.0` → `:0.40.0` (matches the registry's no-`v`
|
||
tag convention; latent footgun fixed).
|
||
|
||
### Tests
|
||
- `doc_test.go`/`backhalf_test.go` updated to the v2 shape (assert no `api_key`/`host_id`,
|
||
`retrieval_password` present, `customer` carries only `id`). `go build ./... && go test ./...` green.
|
||
|
||
## v0.18.0 — slice 10D: DR capstone — identity escrow + restore-mode consumption (agent side) (2026-06-10)
|
||
|
||
The agent half of the slice-10 DR capstone (closes slice 10). Grounded by both 10-series spikes
|
||
(escrow-consumption + identity-restore). The hub half (recovery-mode toggle, re-enroll + credential
|
||
rotation, directive serving) is hub v0.11.0. **Operator-side rotation model (locked):** the hub holds
|
||
no Cloudflare write-power; the destructive tunnel/PBS rotation is the operator's step from a trusted
|
||
environment (same spirit as 10B).
|
||
|
||
### Added (`internal/escrow`)
|
||
- **Identity escrow** (`identity.go`): `WrapIdentity`/`UnwrapIdentity` (+ `…Bundle`) wrap the
|
||
`{tunnel_token, pbs_token}` bundle under the SAME recovery code `R` via **`age`** (scrypt +
|
||
ChaCha20-Poly1305 — a vetted passphrase-AEAD, not hand-rolled), reusing the K-escrow pty mechanism
|
||
(passphrase via the tty, data via files; `R`/tokens never logged). Same two-factor, zero-knowledge
|
||
shape as the K-escrow. A **wrong R fails closed** (no bundle). `age` is a runtime dep for the
|
||
identity path (analogous to proxmox-backup-client for K).
|
||
- **`escrow.Create`** gains an optional `IdentityBundle` → also emits an `IdentityBlob` under the same
|
||
R (additive; the K-escrow + 10C `Consume` paths are byte-unchanged). Self-verifies the identity
|
||
round-trip before shipping.
|
||
- **`--selftest=escrow-create -identity-bundle <file> -directive <file>`** — also wrap + upload the
|
||
identity blob + the **non-secret** DR directive (pbs repo/ns, expected key fingerprint, tunnel id).
|
||
- **`--selftest=identity-consume -blob <file> -keydest <file>`** (R via `FELHOM_RECOVERY_CODE`) —
|
||
recover the identity bundle through the real code; tokens written 0600, never logged.
|
||
|
||
### Tests
|
||
- identity bundle round-trips (wrap→unwrap byte-identical; blob is opaque ciphertext); wrong R fails
|
||
closed + the blob stays retryable; input validation. K-escrow/10C tests byte-unchanged (additive).
|
||
(age integration tests gated to a host with the `age` CLI.)
|
||
|
||
## v0.17.0 — slice 10C: escrow consumption (productionize the spike) (2026-06-10)
|
||
|
||
Turns the throwaway 10C spike harness into a real, tested **`Consume`** path: recover the PBS key
|
||
`K` from an R-wrapped escrow blob, **gate it on the expected fingerprint**, and install it for the
|
||
restore. The spike already proved the crypto + real-data restore; this bakes its findings into
|
||
production code. **Agent-only** — 10C *reads* the four inputs as parameters (so it stays
|
||
standalone-testable); 10D sources blob/fingerprint/PBS-connection from the hub and prompts for R.
|
||
**Zero-knowledge holds**: the hub serves everything except **R** (by hand from the customer), so a
|
||
hub compromise alone still can't decrypt.
|
||
|
||
### Added
|
||
- **`escrow.Consume(ctx, blob, R, expectedFingerprint, keyDest)`** — the consumption contract:
|
||
1. **Unwrap** the blob (a copy — F-C6: the input blob is read-only → a failed Consume is
|
||
**retryable**) with `R`; a **wrong R fails closed** at the scrypt KDF (F-C3) → a clear,
|
||
R-free error, **nothing written**.
|
||
2. **Fingerprint gate (F-C4)** — `KeyFingerprint(recovered)` must equal the expected (the hub
|
||
knows it); a mismatch **fails fast + loud, no install, no restore attempted**.
|
||
3. **Atomic install (F-C2)** at `keyDest` (`0600`, write-temp-sibling→rename); any failure leaves
|
||
**no partial install**. The recovered key lives only in a `0700` tempdir that is always removed.
|
||
**Secret discipline:** `R` and key bytes are never logged/persisted (only fingerprint prefixes);
|
||
`K` is never mutated.
|
||
- **`--selftest=escrow-consume`** (`-blob -fingerprint -keydest`, R via env `FELHOM_RECOVERY_CODE`
|
||
to keep it off the command line) — invokes the real `Consume` live (the spike's S3 via the
|
||
production path, not a harness).
|
||
|
||
### Tests (non-hollow)
|
||
- valid → key installed + `KeyFingerprint(dest) == expected` + `0600` + blob byte-unchanged;
|
||
**wrong R** → error, **no file at dest**, blob unchanged; **fingerprint mismatch** → fail fast,
|
||
**no install** (the gate runs before any restore); input validation; format-tolerant fingerprint
|
||
compare (no empty-fingerprint gate-bypass); atomic-install permissions (integration tests gated to
|
||
a host with `proxmox-backup-client`).
|
||
|
||
## v0.16.0 — slice 10B: operator-signed destructive completion (offline key + signing CLI) (2026-06-10)
|
||
|
||
The security centerpiece: a destructive op runs ONLY on a verified, operator-signed authorization
|
||
— signature valid against a **pinned** operator pubkey (never the hub's or the blob's), nonce
|
||
unseen + durably burned, in-window, host-bound, and **resource-bound to a DURABLE device id** that
|
||
execution re-resolves + re-inspects. Decision (a): **offline operator key + signing CLI**,
|
||
hardware-key-ready (`sk-`/YubiKey via ssh-keygen). The key floor holds: the signing key is NOT in
|
||
the hub and NOT in the agent. Concrete consumer: this **closes the 8C data-bearing-wipe
|
||
`pending_signature` gap**. Pairs with hub v0.10.0.
|
||
|
||
### Added
|
||
- **`cmd/felhom-opsign`** — the operator's offline signing CLI. Builds the canonical `OpBlob` by
|
||
**reusing `authz.CanonicalBlob`** (the exact production path the verifier authenticates over — so
|
||
signer + verifier can never drift) and signs it with **`ssh-keygen -Y sign -n felhom-op-v1`**
|
||
(hardware-ready). Output: a `{op_blob_b64, sig_armored}` envelope to hand to the hub jobs queue
|
||
(optional `--upload`). Touches ONLY the operator's signing key.
|
||
- **`authz.CanonicalBlob`** — promoted to production (was test-only) so the CLI + verifier share one
|
||
canonical-bytes source; params canonicalized (sorted keys, compact).
|
||
- **`internal/storage` durable device identity** (`durable_device.go`): `DeviceDurableID` (derive a
|
||
stable `byid:`(wwn/serial)/`byuuid:` id from the world-readable udev symlinks — no privilege, no
|
||
subprocess) + `ResolveDurableDevice` (re-resolve to the current `/dev` path; a path-only/unknown
|
||
scheme is REFUSED). The resource-level anti-retarget.
|
||
- **`internal/signedjobs`** (new): the queue consumer. `Runner` fetches each opaque job → runs it
|
||
through the **gate** (the LOCKED authz pipeline) → on all-pass hands the verified op to an
|
||
`Executor`; the order is **verify → nonce-burn (durable, in Verify) → execute → clear job**. The
|
||
**`WipeExecutor`** is the 8C consumer: resolve the signed durable id → **re-derive + match**
|
||
(anti-retarget) → **re-inspect (8C classifier)** the device is still the data-bearing target →
|
||
`mkfs`. A vanished/changed/non-data-bearing device or a path-only binding is refused **even with a
|
||
valid signature**. Wired as a second `EnvelopeObserver` (runs on `HasSignedOps`).
|
||
- **`hub.Client.Jobs` / `CompleteJob`** + `hub.MultiObserver`; the 8C format refusal now **surfaces
|
||
the bound op** (op + durable id + host) in its 403 `pending_op` + a `felhom-opsign …` hint.
|
||
|
||
### Pinning / rotation
|
||
- Operator pubkeys are pinned via `authz.signers` (config, trusted path — provision/agent config,
|
||
NEVER hub-alone), **multiple** keys (KeyID selects; role-scoped), so a backup/rotation key exists
|
||
without a flag-day. Unchanged from the slice-4 verifier wiring; 10B activates the execute path.
|
||
|
||
### Tests (real crypto, non-hollow)
|
||
- `signedjobs` runner over the **real** gate+verifier (in-Go minted SSHSIGs): valid → executor runs
|
||
once + job cleared; **replay** (nonce burned) / **non-pinned signer** / **expired** / **retarget**
|
||
(other host) / **forged sig** / **no pinned signer** → all rejected, **executor never called**;
|
||
malformed envelope cleared.
|
||
- `WipeExecutor`: valid → `mkfs` runs; **path-only**, **durable-id mismatch**, **device gone**,
|
||
**re-inspect non-data-bearing**, **not-probed** → all refused, `Format` not called.
|
||
- `storage` durable: wwn-preference, uuid-fallback, path-only/traversal refusal, round-trip,
|
||
missing-device error (symlink tests gated to Linux — the agent's OS).
|
||
|
||
## v0.15.0 — slice 10A: hub desired-state serving — the "Down" channel (2026-06-10)
|
||
|
||
The agent half of slice 10A. The control envelope (`hub.ControlEnvelope`) stops being "reserved — ignored" and becomes the live **Down channel**: a cheap change-notification on every heartbeat. The agent caches the hub's desired-state + its generation; only when **`DesiredGeneration` advances** does it fetch the full state (the heartbeat stays light, the heavy state moves on change). The engine then reconciles **benign** deltas and the gate marks an explicit **destructive** delta `pending_signature` (no signer in 10A → never executed; signed execution is 10B). Pairs with hub v0.9.0.
|
||
|
||
### Added / changed
|
||
- **`internal/reconcile`**: `DesiredGuest.Decommission` — the canonical **destructive desired-state delta** (an EXPLICIT flag, not "absent from the list", so a partial hub list can never mass-destroy). The planner emits `ActionDecommission` → `ClassDecommission` → Destructive → the gate refuses it `pending_signature`. `Reconcile` now counts a `pending_signature` refusal as **`Result.Pending`** (expected, logged INFO) rather than a failure; any other refusal stays a real failure. `ActionDecommission` has **no executor** (slice 10B) — a defensive guard refuses to run it. New **`CachingProvider`** (thread-safe DesiredState + generation cache; `Desired`/`Update`/`Generation`) — the production `DesiredProvider`, replacing `EmptyProvider` in the daemon engine (empty until the hub serves intent → cold-start is a live no-op, unchanged).
|
||
- **`internal/hub`**: the **`ControlEnvelope`** fields are now active (DesiredGeneration drives the fetch, HasSignedOps noted). New wire types **`DesiredStateResponse`** + **`WireDesiredState`** (guests + forward-compat `restore_directive` (10D) / `pbs_namespace` / opaque `storage_manifest`+`backup_policy`) + **`WireDesiredGuest`** (vmid/run/spec/description/decommission). New **`Client.FetchDesiredState`** (GET `/api/v1/hosts/{host_id}/desired-state`, self-scoped to the client's own host). New **`EnvelopeObserver`** loop seam + `SetEnvelopeObserver` — the loop hands the envelope to the sync layer each cycle (hub does not import reconcile/desired).
|
||
- **`internal/desired`** (new): the **`Syncer`** — implements `hub.EnvelopeObserver`, fetches desired-state on a generation advance, maps the wire shape to the reconcile domain, and updates the `CachingProvider`. Caches the **fetched** generation (robust to a generation that advanced mid-fetch); a fetch failure keeps the last-known state. `restore_directive` is carried + logged, not acted on (10D). Wired in `cmd/felhom-agent` (daemon): provider → engine, syncer → loop.
|
||
|
||
### Tests
|
||
- reconcile: a desired-state with one benign + one decommission delta → **benign applied, destructive gated pending (not executed)**; `Plan` emits decommission-only for a decommissioned guest + classifies Destructive; `CachingProvider` update/isolation.
|
||
- desired: **fetch-once-on-advance** (no re-fetch on an unchanged generation), fetch-failure-keeps-cache, caches-the-fetched-generation.
|
||
- hub client: `FetchDesiredState` hits the self-scoped path with the bearer + decodes (incl. `restore_directive`); a 403 is a typed `HTTPError`.
|
||
- loop: the cycle notifies the observer + adopts `PollIntervalSeconds`; a report error skips the observer.
|
||
- cross-repo golden: `testdata/desired-state.golden.json` + `control-envelope.golden.json` decode + key-set guard, **byte-identical** with felhom.eu/hub.
|
||
|
||
## v0.14.0 — slice 9: host metrics to the controller (`GET /host/metrics` + CPU-temp collector) (2026-06-10)
|
||
|
||
The de-privileged controller (slice 8C) sees only its own cgroup, so it can't read host health itself. Slice 9 **re-serves** the slice-4 collector's host + per-storage view to the customer over the local API, plus the one missing collector — CPU/chassis temperature — so the customer sees their box's health in the controller. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot). Assumption: **one customer per host** (the home-server model); if a host ever serves multiple customers, host-wide CPU/mem would leak cross-customer load → revisit then.
|
||
|
||
### Added / changed
|
||
- **CPU/chassis-temp collector** (`internal/hub/cputemp.go`): `SysfsTempReader` reads the CPU package temperature straight from sysfs — hwmon (`coretemp`/`k10temp`/`zenpower`/`cpu_thermal`, preferring the `Package id 0` input) then the thermal zones (preferring `x86_pkg_temp`/`coretemp`/`cpu-thermal`, falling back to `acpitz`). **No external binary, no privilege** (sysfs nodes are world-readable), so the root-CLI fence is untouched. **Graceful-null**: a missing sensor, an unsupported board, an implausible reading (outside 5–150 °C), or any read error all degrade to `null` ("n/a") — a missing sensor never fails the report. Wired into the collector via the new `TempReader` seam (nil-safe).
|
||
- **`HostMetrics.CPUTempC *int` (`cpu_temp_c`)** — new nullable wire field on the **shared** `HostMetrics` struct (same nullable contract as the disk `SmartSummary.TemperatureC`). It rides the **hub report too** (operator freebie) → cross-repo host-report golden updated.
|
||
- **`Collector.HostMetricsNow(ctx)`** — a fresh `NodeStatus` + CPU-temp read returning just the host block, the source for the local API (current cpu%/temp, not the 15-min snapshot). `Collect()` now also populates `cpu_temp_c` on the hub report. `Collector.SetTempReader` injects a fake in tests.
|
||
- **`GET /host/metrics`** (`internal/localapi/host_metrics.go`): host-wide health (cpu%/mem/load/uptime/`cpu_temp_c`) + per-storage capacity targets (total/used/fraction, thin-pool, SMART temp+wear). Token-authed via `withGuest` (host-wide data; cross-guest `?vmid=` still 403). Best-effort on storage (a view error still returns the host block). Served only when the `HostMetrics` provider (the shared collector) is wired — else 503 "not configured". Wired in `buildLocalAPIServer`.
|
||
|
||
### Tests
|
||
- `cputemp_test.go`: a fake `/sys` layout proves hwmon package-preference, hwmon first-input fallback, thermal-zone-by-type selection over a non-CPU hwmon, **graceful-null on a sensorless host** (no error), and rejection of implausible (0 m°C) readings.
|
||
- `hostmetrics_test.go`: `HostMetricsNow` populates the temp, gracefully nulls it, hard-errors on `NodeStatus` failure; `Collect()` carries the temp.
|
||
- `host_metrics_test.go` (localapi): populated host+storage with a valid token; `cpu_temp_c:null` serializes; **401 without a token** (collector never invoked); 403 on a cross-guest `?vmid=`; 503 when not configured.
|
||
|
||
## v0.13.0 — slice 8B.2: quiesce downtime optimization (`snapshotted` phase) (2026-06-10)
|
||
|
||
The agent half of slice 8B.2. In snapshot mode, vzdump only needs the app-stopped state captured at
|
||
the **storage-snapshot moment**; after that it reads from the snapshot and the app can resume. The
|
||
agent now emits a **`snapshotted`** phase on `GET /backup/status` when the snapshot is taken, so the
|
||
controller (v0.38.0) resumes its app early — app downtime drops from *whole-backup* to
|
||
*until-snapshot* with no loss of app-consistency. Validated Phase-0 first on PVE 9.2.2: the marker is
|
||
`INFO: create storage snapshot 'vzdump'`; downtime ~24s→~1s for a 934 MB guest.
|
||
|
||
### Added / changed (`internal/backup` + `internal/localapi`)
|
||
- **`BackupRunner.BackupWithSnapshotHook(ctx, vmid, onSnapshot)`** — while the vzdump runs, a watcher
|
||
tails the task log (`TaskLogTail`) for the **`create storage snapshot`** marker and fires
|
||
`onSnapshot` **once**. The marker only appears in snapshot mode (stop/downgraded takes no storage
|
||
snapshot), and the watcher also bails on `backup mode: stop` — so it never fires in stop mode.
|
||
(`Backup` keeps its signature for the scheduler/selftest; both share one body.)
|
||
- **`/backup/status` phase `snapshotted`** (between `running` and `done`): `handleBackup` passes the
|
||
hook → `markSnapshotted` flips the running job to `snapshotted`. `done`/`failed` semantics unchanged.
|
||
|
||
### Tests
|
||
- localapi: snapshot mode → phase reaches `snapshotted` before `done` (gated fake holds the backup
|
||
open); stop mode → `snapshotted` **never** emitted (stays running → done). runner: the watcher
|
||
fires `onSnapshot` on the marker; in stop-mode log it never fires. `snapshotWatchInterval` is a
|
||
package var so tests run fast.
|
||
|
||
## v0.12.0 — slice 8C Phase A: disk endpoints + data-bearing classifier gate + mkfs executor (2026-06-10)
|
||
|
||
The agent half of slice 8C, Phase A (additive). Adds the host disk-management endpoints the
|
||
controller's disk UI drives — with the **8C security invariant**: the agent decides
|
||
data-bearing-ness by **inspecting the actual device** (agent-internal evidence), NEVER from the
|
||
caller's claim. A compromised controller asserting "this drive is blank" cannot wipe a data-bearing
|
||
drive. (Controller rewire + disk-subsystem retirement + de-privilege are Phases B/C, `felhom-controller`.)
|
||
|
||
### Added
|
||
- **`internal/storage` — `mkfs` executor + data-bearing inspection.** `SudoHostOps.Format(device,
|
||
fstype)` (device-pinned, `ValidateBlockDevice`+`ValidateFSType`, narrow `FELHOM_FORMAT` sudoers —
|
||
`mkfs.ext4 -F` / `mkfs.xfs -f` on a `/dev/*` path the agent fine-validates first).
|
||
`SudoHostOps.InspectDevice(device)` → `DeviceProbe` (filesystem signature via `blkid -p`, partition
|
||
table / partitions / mount via `lsblk -J`). **`DeviceProbe.DataBearing()` is conservative**: any
|
||
signature / partition table / partition / mount — OR a probe that did not read cleanly — is
|
||
data-bearing (fail-safe; an unreadable device is never called blank).
|
||
- **`internal/localapi` — the §6 disk endpoints**, all self-scoped (token→guest; cross-guest 403):
|
||
- `GET /disks` — host drives + a **data-bearing flag** (UI hint). Read-only/benign.
|
||
- `POST /disks/assign` — attach a drive as a mount (benign, additive → `EnsureMount`). Self-serve.
|
||
- `POST /disks/eject` — safe-unmount (benign, data preserved) + the **dependent guests** that
|
||
mount it (so the controller can warn which apps lose that storage).
|
||
- `POST /disks/format` — **the security centerpiece**: the agent **inspects the device itself**;
|
||
blank → benign → `mkfs`; **data-bearing → ClassStorageWipe → the slice-4 gate → refused
|
||
`pending_signature`** (the operator-signed completion is slice 10). The caller's claim is
|
||
ignored — only a device the agent reads as blank is formatted.
|
||
- `storageGateAdapter` bridges the format path to the slice-4 reversibility gate (no new gate/crypto).
|
||
|
||
### Tests
|
||
- localapi (security matrix): blank device → **mkfs called, gate not consulted**; a **data-bearing
|
||
device → 403, mkfs NEVER called**, gate consulted (`pending_signature`); an **ambiguous/unprobed
|
||
device → treated destructive** (fail-safe); even a gate that *allows* does not format data-bearing
|
||
in 8C; assign → `EnsureMount`; eject → `Unmount` + dependent guests; cross-guest → 403; bad
|
||
device/fstype → 400; unconfigured → 503.
|
||
- storage: `ValidateBlockDevice`/`ValidateFSType` (whitelist + injection rejection); `InspectDevice`
|
||
blank/filesystem/partition-table/mounted/failed-probe-fail-safe; `Format` invokes the right `mkfs.*`.
|
||
|
||
## v0.11.0 — slice 8B: app-consistent backup — /backup/due policy + /backup/status phases (2026-06-10)
|
||
|
||
The agent half of slice 8B (doc 03 §8). Turns the 8A thin backup stubs into the real policy the
|
||
in-guest controller's quiesce loop drives (controller half: `felhom-controller` v0.36.0). No hub
|
||
change. The downtime optimization (`vzdump --mode snapshot` + a `snapshotted` phase) is the 8B.2
|
||
fast-follow; the hub-served per-guest policy is slice 10.
|
||
|
||
### Changed (`internal/localapi`)
|
||
- **`GET /backup/due`** — real **cadence** policy (replaces the 8A "never backed up" stub): a guest
|
||
is due when no **successful** backup is recorded OR the newest one is older than the agent-local
|
||
cadence (`backup.backup_cadence_seconds`, default 24h). A successful `POST /backup` flips due to
|
||
**false** for the window, so the controller won't re-quiesce in a loop. A failed backup does not
|
||
satisfy the cadence. Returns `age_seconds` for diagnosis.
|
||
- **`GET /backup/status`** — real **phases** `idle | running | done | failed` + the job id, so the
|
||
controller can poll a backup to completion (was: just the latest stored backup).
|
||
- **`POST /backup`** — returns a **job id** + `running` phase; tracks the in-flight job and is
|
||
**single-flight per guest** (a second POST while one runs returns the same job — no concurrent
|
||
vzdump). On completion the job transitions done/failed and the result is recorded to the store.
|
||
- Config: `backup.backup_cadence_seconds` + `BackupCadence()`; the local-API server takes the cadence.
|
||
|
||
### Tests
|
||
- `/backup/due`: due when stale / no backup, **not due within the window after a success**, due again
|
||
past the cadence, **a failed backup does not count**. `/backup/status`: running→done and
|
||
running→failed (gated fake to observe the running phase). `POST /backup` single-flight (one vzdump
|
||
for concurrent POSTs). All still self-scoped (token→guest).
|
||
|
||
## v0.10.0 — slice 8A: agent local-API server + provisioning back-half (2026-06-10)
|
||
|
||
The host-agent half of slice 8A (doc 03 §6). Adds the per-guest **local API** the in-guest
|
||
controller calls over the bridge, and the **provisioning back-half** that follows the slice-7
|
||
bring-up front half. Grounded by `felhom.eu/documentation/tests/slice8a-channel-deploy-spike-findings.md`
|
||
(commit `4a81a96` — channel + deploy plumbing proven; the 5 gotchas resolved here). Controller half
|
||
is `felhom-controller` v0.35.0. No hub change.
|
||
|
||
### Added
|
||
- **`internal/localapi`** — the HTTPS local-API server (doc 03 §6), the **per-guest authorization
|
||
gate**. Serves a **persisted self-signed leaf** with a **stable SHA-256 fingerprint** (generated
|
||
once; a fresh cert each boot would invalidate every baked bootstrap pin). The **7 §6 endpoints**,
|
||
all **self-scoped to the caller's own guest**: `GET /storage` (this guest's mpN mounts + fast/slow
|
||
class from the slice-5/7 storage view), `POST /snapshot`, `POST /rollback`, `POST /backup`
|
||
(enqueued, crash-consistent — the app-consistent quiesce loop is 8B), `GET /backup/due` (thin in
|
||
8A), `GET /backup/status`, `GET /restore-test/status`.
|
||
- **Token store** (`tokenstore.go`): durable, crash-safe per-guest token→guest map that persists
|
||
only a **SHA-256 hash** of each token (the plaintext exists transiently at mint→write-to-mount,
|
||
then is discarded), last-write-wins per guest, fsync'd append-only JSONL (mirrors the nonce store).
|
||
- **Self-scoping**: the VMID is resolved ONLY from the token; an explicit `vmid` (query/body) that
|
||
disagrees → **403 and the proxmox op is never issued for the other guest**; absent/unknown → 401.
|
||
- **`internal/provision`** — the back-half: mint the per-guest token → render the stable
|
||
**`bootstrap.json`** contract (schema `felhom.bootstrap/v1`; **no registry credential** — the
|
||
controller image is baked into the golden) → write it `0600` → **`chown 100000:100000`** (the
|
||
unprivileged-LXC mapped guest-root, spike gotcha 1) → attach a **read-only bind mount** via
|
||
`pct set`. Host-side only (F3 — the agent never enters the guest; **no `pct exec`**). The token
|
||
plaintext is never logged and never returned.
|
||
- **`--selftest=provision`** — the full chain on-demand: bring-up (provision) front half + the
|
||
back half; keeps the guest for the golden's baked controller-bootstrap unit to deploy.
|
||
- **`config.LocalAPIConfig`** (`local_api`) — enable + bridge `listen_addr` + cert/key paths + token
|
||
store path. The server is an optional 6th daemon goroutine, disabled cleanly when unconfigured or
|
||
on a token-store/cert failure (the daemon still reports/reconciles).
|
||
- **`configs/build-golden.sh`** now **bakes the controller image** (pulled once on the trusted build
|
||
host, then `docker logout` — no cred baked) + a **controller-bootstrap unit** that deploys the
|
||
**baked** image from the config mount on boot (no login/pull at deploy).
|
||
- **`configs/felhom-localapi-firewall.example`** — host firewall narrowing of the local-API port to
|
||
the guest bridge subnet (nft/iptables/PVE variants; defense-in-depth — the token stays the gate).
|
||
- **`configs/felhom-agent.sudoers`** — a narrow `FELHOM_PROVISION` alias (`chown 100000:100000` +
|
||
`pct set` bind-mount, both confined to the agent-owned `/var/lib/felhom-agent/guests/*` path) for
|
||
the non-root least-privilege deployment.
|
||
|
||
### Security / design notes
|
||
- The local-API leaf is pinned by **leaf-cert SHA-256** (decision: consistency with the agent's
|
||
PVE/PBS pinning); the fingerprint is baked into each guest's bootstrap.
|
||
- The back-half's host-root ops (chown + bind-mount attach) are **NOT** added to `proxmox.Privileged`
|
||
(which is fenced to its 3 exceptions) — they live in `internal/provision` and run through the shared
|
||
`Runner` (direct as root, or `sudo -n` with the new sudoers alias). This is the per-guest
|
||
provisioning host-root surface, host-side and F3-compliant.
|
||
|
||
### Tests
|
||
- localapi: self-scoping (cross-guest snapshot/rollback/backup → 403, op never issued for the other
|
||
guest; own-guest uses the token's VMID), 401 paths, `/storage` class mapping, `/backup` enqueue,
|
||
the thin `/backup/due`, status scoping; the token store persists only the hash (plaintext never on
|
||
disk), last-write-wins, survives reopen, uniqueness; the leaf fingerprint is stable across reload.
|
||
- provision: writes `0600` + chowns + attaches the bind mount with the right args; the **token never
|
||
appears in the Result**; the cross-repo `bootstrap.json` contract key-set is pinned.
|
||
|
||
## v0.9.0 — slice 7 close-out: PBS recovery-code escrow creation (2026-06-10)
|
||
|
||
The first code that touches the PBS client encryption key `K` and introduces the customer recovery
|
||
code `R`. Default posture is **zero-knowledge**: Felhom holds an opaque `R`-wrapped blob (cannot
|
||
open it), the customer holds `R`. Grounded by `felhom.eu/documentation/tests/slice7-escrow-spike-findings.md`
|
||
(round-trip proven on a throwaway: the `R`-recovered key restores a real encrypted snapshot). Hub
|
||
opaque storage is the `felhom.eu` half (hub v0.8.0); consumption/serving is slice 10.
|
||
|
||
### Secret discipline (overriding)
|
||
`R` is `crypto/rand`, ≥128 bits, surfaced **exactly once** and **never** logged/persisted/committed;
|
||
the wrap pty's echo is discarded so `R` can't leak. `K` is read by location, **never modified** (the
|
||
live key file is byte-unchanged — Wrap operates on a copy), never logged.
|
||
|
||
### Added
|
||
- **`internal/escrow`** — `Create` generates `R` (10 EFF-wordlist words ≈ 129 bits), wraps `K` under
|
||
`R` via the **PBS-native** `proxmox-backup-client key change-passphrase --kdf scrypt`, and
|
||
**self-verifies** the blob recovers `K` (fingerprint match) before shipping. The wrap is driven
|
||
over a **stdlib pty** (`x/sys/unix`; spike F-A1 — the command is TTY-only) with **output discarded**
|
||
(F-A2 — the pty echoes the passphrase). Opt-in outputs: **(b)** `R`-wrapped offline copy (two-factor,
|
||
no extra trust) and **(a)** raw paperkey (single-factor, unrevocable — loud caveat).
|
||
- **`--selftest=escrow-create`** (`-storage`, `-paperkey`, `-offline`, `-upload`): surfaces `R` once
|
||
to stdout (never the logger), prints the opaque blob's size/fingerprint/posture, and with
|
||
`-upload` PUTs the blob to the hub (`/api/v1/hosts/{host_id}/escrow`, per-host key).
|
||
- Config: `escrow` section (`posture` default `zero_knowledge`, `pbs_storage_id`); `PBSEncKeyPath`
|
||
helper (the `<id>.enc` key K).
|
||
- Runtime dependency on the `proxmox-backup-client` CLI (the PBS key+passphrase KDF).
|
||
|
||
### Tests
|
||
- `R` entropy ≥128 / 10-word format / uniqueness; integration round-trip (wrap→unwrap fingerprint
|
||
match, **wrong-`R` fails**, **live `K` byte-unchanged**, blob ≠ plaintext key) guarded to
|
||
linux+`proxmox-backup-client`; the agent→hub wire-contract key-set (mirrors the hub's).
|
||
- **Live-validated** (demo): `escrow-create` → `R` (10 words) surfaced once, blob 383 B opaque,
|
||
self-verify ok, **live `K` sha256 unchanged**, exact `R` absent from stderr/journal.
|
||
|
||
## v0.8.0 — slice 7 Phase 1: unified bring-up reconcile job (provision + guest-loss DR) (2026-06-09)
|
||
|
||
The shared FRONT HALF of provision and guest-loss DR, as a journaled reconcile job mirroring the
|
||
slice-6 restore-test's crash-safety — but it KEEPS the guest on success and applies a
|
||
scenario-specific identity policy. Agent-only; no hub/wire change (the new guest auto-appears in
|
||
the host-report via `ListLXC`). Grounded by the slice-7 bring-up spike findings (commit `3342993`):
|
||
F1 (restore preserves the archived MAC → provision reset is unconditional), F3 (SSH host keys do
|
||
not auto-regenerate → a baked golden first-boot unit, not an agent guest-internal op), F4 (the
|
||
transient PVE config-lock 500 → bounded retry).
|
||
|
||
### Added
|
||
- **`reconcile.RunBringUp`** (`bringup.go`) — `BringUpSpec` (Mode `provision`|`dr_guest_loss`,
|
||
Archive, VMID, RestoreStorage, Hostname, Cores/MemoryMB, RootfsGrowGB, Mounts, KeepMAC,
|
||
BootTimeout) → `BringUpResult` (VMID, AssignedMAC, Pass, Verified, StartWarnings/Recognized).
|
||
Sequence (each mutation preceded by journaling the owning entry): restore → identity reset →
|
||
size → attach mounts → start LINK-UP. **Verdict is liveness (`waitRunning`), never the start
|
||
exitstatus** (reuses the v0.7.0 WARNINGS surface). **Success KEEPS the guest** (no teardown).
|
||
- **Scenario-specific identity reset** (doc 03 §9): *provision* → fresh MAC unconditionally
|
||
(`PUT net0` with `hwaddr` omitted → PVE regenerates, F1) + hostname; machine-id + SSH host keys
|
||
regenerate guest-side on first boot (golden bake + the new unit) — the agent does NOT touch
|
||
guest internals. *dr_guest_loss* → preserve continuity (keep hostname; keep MAC unless
|
||
`KeepMAC=false`); never resets restic/tunnel/hub identity.
|
||
- **Compensating rollback** — any mid-flight failure destroys the just-created guest
|
||
(`ClassGuestDestroy`, benign via `Provenance{SameTxnCreated:true}`, gated); on teardown failure
|
||
the entry is left in-flight for `Recover`. New journal flag **`Rollback`** + `Recover`'s
|
||
`recoverBringUp` reap a half-built guest left by a mid-job crash (idempotent, via `ListLXC`).
|
||
- **F4 config-lock retry** — steps 3+5 coalesced into ONE `PUT config` (net0+hostname+cores+
|
||
memory+mpN); rootfs grow stays its own call. `setConfigWithLockRetry` retries ONLY the transient
|
||
PVE config-lock 500 (`pveConfigLock`: 500 + "can't lock file"/"got timeout"); any other error
|
||
fails immediately — never retried.
|
||
- **`--selftest=bring-up`** (`-mode provision|dr -archive -vmid -hostname [-keep]`) — runs the real
|
||
journaled job (after a `Recover`), then tears the guest down unless `-keep`.
|
||
- **`configs/build-golden.sh`** — the validated golden recipe as a script, incl. the F3
|
||
first-boot `felhom-regen-hostkeys.service` unit (Condition-gated: fires on provision, no-ops on
|
||
DR). The slice-7 spike archive (which lacks the unit) is superseded.
|
||
|
||
### Deferred (stated, not built)
|
||
- Provisioning BACK HALF (controller deploy, bootstrap, per-guest token mint) → **slice 8**.
|
||
- Host-loss DR + PBS escrow consumption → **slice 10**.
|
||
- The SOURCE of a `BringUpSpec` (hub desired-state: which archive/VMID/mounts) → **slice 10**;
|
||
this job takes the spec as input. `GuestMount` is defined minimally (no hub coupling).
|
||
|
||
### Tests
|
||
- provision happy path (fresh MAC = net0 without hwaddr, hostname, coalesced sizing+mount, rootfs
|
||
grow separate, started, **guest NOT destroyed**); compensating rollback at each step (restore /
|
||
config / start-task / waitRunning — asserts the guest WAS destroyed); DR continuity (MAC kept,
|
||
hostname not reset) + DR `KeepMAC=false` resets MAC; liveness verdict (warnings+running pass /
|
||
not-running fail); F4 (lock-500→retry→proceed; non-lock-500→fail without retry); owning entry
|
||
journaled BEFORE restore; reserved/existing VMID refused; `Recover` rolls back / clean.
|
||
|
||
### Live-validated (demo-felhom)
|
||
- provision: fresh MAC + hostname; **SSH host keys regenerated by the baked golden unit** (agent
|
||
issued no `ssh-keygen`), machine-id unique, Docker runs, clean DHCP lease → torn down.
|
||
- dr: continuity preserved (hostname + host keys kept). Recover: a killed mid-restore left an
|
||
orphan; the re-run's `Recover` rolled it back (idempotent).
|
||
- **Live caught a bug, then fixed:** the host-key unit's `ExecStart` was `/usr/sbin/ssh-keygen`
|
||
(203/EXEC); on Debian 13 it is `/usr/bin/ssh-keygen` — corrected in `build-golden.sh`, golden
|
||
rebuilt, re-validated. (Mocked unit tests couldn't surface this; the live run did.)
|
||
|
||
## v0.7.0 — restore-test: verdict is liveness, not start-task exitstatus (2026-06-09)
|
||
|
||
Fixes a correctness bug found by the live hub-enrollment runbook: the self-restore-test reported
|
||
`pass:false` on **every** modern-distro guest. PVE's guest-start task exits `"WARNINGS: 1"` for the
|
||
benign systemd-nesting advisory (`WARN: Systemd 257 detected. You may need to enable nesting.`), and
|
||
`WaitTask` treated any non-`"OK"` exitstatus as a hard failure — so the verdict was decided by an
|
||
advisory exit code instead of by observed liveness, *before* the real boot check ran. A crying-wolf
|
||
test got it disabled on the demo host; this re-enables it. **Single bump (0.6.0→0.7.0) covering the
|
||
agent's part of both task phases**; the wire fields below are consumed by hub from **v0.7.5**.
|
||
|
||
Design invariant (in code): **warning classification affects *visibility only*; pass/fail is
|
||
liveness-only.** A wrong/stale recognizer can at worst over-notice a benign warning — it can never
|
||
false-fail and never hide a real warning.
|
||
|
||
### Added
|
||
- **`proxmox.WaitOptions.AllowWarnings`** — opt-in per call. When set, a task that completes
|
||
`"WARNINGS: N"` is success with the `TaskStatus` (ExitStatus intact) returned so the caller can
|
||
read/surface it. Default (`false`) keeps **every existing caller strict** (vzdump/restore/destroy
|
||
warnings can be meaningful — relaxing them is a future per-call decision with evidence). Any
|
||
non-WARNINGS non-OK exit is still a `*TaskError`.
|
||
- **`reconcile.RestoreTestResult.StartWarnings` / `.WarningsRecognized`** + a version-free recognizer
|
||
(`benignWarningAnchor = "enable nesting"`, case-insensitive substring — contains no systemd version
|
||
number, so it can't rot back into the bug at systemd 258+). `extractWarningLines` pulls `WARN…`
|
||
lines from the start-task log.
|
||
- **`reconcile.GuestAPI.TaskLogTail`** — the engine fetches the start task's log to surface warnings.
|
||
- **`hub.RestoreTest.warnings` / `.warnings_recognized`** wire fields (`omitempty`), populated by
|
||
`ToHubRestoreTest`. Additive: the deployed v0.7.4 hub ignores them; hub v0.7.5 consumes them
|
||
(passed-with-warnings INFO, or WARN when not recognized). Cross-repo golden updated with the hub side.
|
||
|
||
### Changed
|
||
- **Restore-test start step** (`reconcile/restoretest.go`) now waits with `AllowWarnings:true`,
|
||
surfaces any start warnings, and **continues to `waitRunning` as the verdict** — boot+running is the
|
||
pass, exactly as before; a real (non-WARNINGS) start-task error still fails. The restore and
|
||
scratch-teardown WaitTasks stay strict.
|
||
- **Restore-test scheduler logging** distinguishes a clean pass, *passed-with-recognized-warnings*
|
||
(INFO), and *passed-with-unrecognized-warnings* (WARN) — nothing silent.
|
||
|
||
### Tests
|
||
- `WaitTask`: AllowWarnings accepts `WARNINGS` (status returned intact); AllowWarnings still fails a
|
||
real error; default still fails on `WARNINGS` (existing callers unaffected).
|
||
- Restore-test (engine, mock proxmox): start-with-warnings + running → **pass** with warnings
|
||
surfaced+recognized; unrecognized warning + running → pass, not-recognized; **not-running → fail
|
||
regardless of warnings** (verdict is liveness); teardown still runs.
|
||
- **Regression guard:** the `"enable nesting"` recognizer matches the advisory for systemd 256–300,
|
||
proving it's version-independent and can't silently rot back into the false-fail.
|
||
|
||
## v0.6.0 — slice 6 Phase B: PBS offsite tier (verify + PBS-API client + reporting) (2026-06-09)
|
||
|
||
Completes slice 6. The PBS spike (felhom.eu phase5-pbs-spike-findings.md) proved backup-to-PBS
|
||
and restore-from-PBS reuse Phase A UNCHANGED (PBS is just a storage target + a volid), and the
|
||
operator token needs no widening. So the only new agent code is the **verify capability + a
|
||
small PBS-API client + PBSSnapshot reporting**. Escrow + host-loss DR stay slices 7/10.
|
||
|
||
### Added
|
||
- **`internal/pbs` — the PBS-API client** (the agent's SECOND privileged external surface,
|
||
slice-1 discipline): TLS **fingerprint-pinned** to the PBS leaf cert (a spoofed PBS →
|
||
rejected, mirroring the PVE pin), **token auth** (`PBSAPIToken=<id>:<secret>`; id from the
|
||
storage `username`, secret read at runtime from `/etc/pve/priv/storage/<id>.pw` — referenced
|
||
by location, never logged/committed), typed, no shell. Methods: `Verify` (POST
|
||
`/admin/datastore/<ds>/verify` → UPID), `Snapshots` (incl. the `verification` field),
|
||
`TaskStatus`/`WaitVerify` (node extracted from the UPID — `localhost` returns "unknown", the
|
||
spike B4 gotcha), `NodeFromUPID`.
|
||
- **The verify maintenance loop** (`pbs/verify.go`) — the cheap, key-free, ciphertext-level
|
||
integrity check (§8) on its OWN cadence (default 6h, the 5th daemon goroutine). It is a
|
||
reporting/maintenance task like the slice-5 watchdog: it does NOT go through the reconcile
|
||
gate/journal. Each cycle: trigger verify → poll task → re-list snapshots → record
|
||
per-snapshot `verify_state`. A failed verify is logged loudly.
|
||
- **`PBSSnapshot` reporting** — filled the stub (`namespace`/`backup_type`/`backup_id`/
|
||
`backup_time`(RFC3339)/`size_bytes`/`owner`/`protected`/`encrypted` (from `files[].crypt-mode`)
|
||
/`verify_state` (ok|failed|**none** until verified)/`verify_upid`). New `PBSReporter`
|
||
collector seam + an in-memory `SnapshotStore`. Cross-repo golden (both repos, byte-identical)
|
||
+ bidirectional key-set tests; hub `handler.go` parses `pbs_snapshots` and logs a **failed
|
||
verify `[WARN]`** (loudest offsite-DR signal).
|
||
- **Truthful backup mode** (`backup/runner.go`) — `Backup.mode` now reflects the ACTUAL vzdump
|
||
mode read from the task log (`backup mode: <x>`), since PVE may downgrade snapshot→stop for a
|
||
stopped guest (spike B1); falls back to the requested mode if unparseable.
|
||
- **proxmox**: `Storage.Username` (parsed from the pbs storage config — the token id).
|
||
- **config** `BackupConfig.{PBSVerifyCadenceSeconds, PBSSecretDir}` (cadence 0→6h, <0 disabled).
|
||
- **`--selftest=pbs-verify`** — discover pbs storages → verify each → print the PBSSnapshot
|
||
records (covers the runbook's verify + list). Standalone on the host.
|
||
|
||
### Notes
|
||
- Backup/restore-to-PBS reuse Phase A with no change (the restore-test runs with
|
||
`source_tier="pbs"` when fed a pbs volid). Zero-knowledge holds: verify is ciphertext-level,
|
||
the encryption key is never read here, and the PBS server has no client key (spike B6).
|
||
- Daemon runs cleanly with no pbs storage / verify disabled. `go test -race` covers the new
|
||
goroutine. Slice-3/4/5/6A surfaces, goldens, and adversarial tests intact.
|
||
|
||
## v0.6.0-rc1 — slice 6 Phase A: backup + the self-restore-test (local target) (2026-06-09)
|
||
|
||
Phase A of the backup/restore slice (doc 03 §8) — the agent's guest-level backup layer and
|
||
the **self-restore-test**, which closes "a backup you haven't restored isn't a backup".
|
||
Everything here is BENIGN (backup, restore-to-NEW, scratch teardown): reuses the slice-4
|
||
classifier/gate/journal — no new destructive class, no new crypto. Local target only; PBS is
|
||
Phase B. Restore is to a NEW guest only (no overwrite). Backups are crash-consistent only
|
||
(app-consistency needs the controller quiesce, slice 8) — marked so in the report.
|
||
|
||
### Added
|
||
- **proxmox** (`mutate.go`/`query.go`): `DestroyLXC` (DELETE …/lxc/{vmid}?purge=1&destroy-
|
||
unreferenced-disks=1 → UPID; the scratch-teardown primitive); `VzdumpOptions.Notes` →
|
||
`notes-template` (verified on PVE 9.2.2); `LatestBackupVolID` (resolve a produced archive
|
||
from the backup-storage listing — the task status carries no result volid).
|
||
- **reconcile self-restore-test** (`restoretest.go`) — `Engine.RunRestoreTest`: pick a free
|
||
scratch VMID (configured band, excludes 9999; full band → skip, never out-of-band) →
|
||
**journal a Scratch-owned entry BEFORE any mutation** → restore-to-new → benign net
|
||
**link-down** SetConfig (so the clone can't conflict with a running source's MAC/IP; this
|
||
is test-safety, NOT slice-7 identity reset) → boot → verify **reaches `running`** → ALWAYS
|
||
teardown (defer; benign `ClassGuestDestroy` + agent-tagged-scratch provenance, gated). Runs
|
||
on the scratch VMID's queue lane. Reuses the journal/gate; result feeds the report.
|
||
- **Crash-safe recovery** (`recover.go`): a Scratch journal entry is resolved by TEARDOWN,
|
||
not by re-checking the restore sub-task's UPID — special-cased BEFORE the generic path
|
||
(else the restore task's OK would mark it succeeded while the guest leaks). `Recover` now
|
||
destroys a leaked scratch guest (idempotent: already-gone → clean; list-unreadable → left
|
||
in-flight for a later pass). `JournalEntry.Scratch` flag; `RecoverResult.ScratchClean/
|
||
ScratchDestroyed`. GuestAPI gains `RestoreLXC`/`DestroyLXC`/`GuestStatus`.
|
||
- **`internal/backup` package**: `BackupRunner.Backup` (vzdump + archive/size resolve +
|
||
bulk-volume gap — a mountpoint is UNCOVERED unless it carries an explicit `backup=1`, so
|
||
an unset `backup=` is reported uncovered too, the safe DR direction); `PickRestoreCandidate`
|
||
(newest backup); an in-memory `Store` (latest-backup-per-target + latest-restore-test)
|
||
implementing the hub `BackupReporter`/`RestoreTestReporter` seams; a cadence `Scheduler`
|
||
(default 24h; the fourth daemon goroutine; disabled cleanly when off/misconfigured).
|
||
- **hub report** (`report.go`): filled the `Backup` + `RestoreTest` stubs (`PBSSnapshot`
|
||
stays a Phase-B stub); collector `BackupReporter`/`RestoreTestReporter` seams. Cross-repo
|
||
golden updated in BOTH repos (byte-identical) + bidirectional key-set tests for
|
||
`backups[0]`/`restore_tests[0]`. Hub `handler.go` parses + persists them (report_json; no
|
||
new columns) and logs a **FAILED restore-test prominently** (the loudest DR signal).
|
||
- **config** `BackupConfig` (local target, restore storage, restore-test cadence, scratch
|
||
VMID band 990000–990009 default) + accessors + env overlay + cadence-gated validation.
|
||
- **`--selftest=backup -vmid N`** (one-shot backup → print the Backup record) and
|
||
**`--selftest=restore-test [-archive volid]`** (Recover-then restore→boot→verify→teardown,
|
||
print the RestoreTest record). Standalone on the Proxmox host.
|
||
|
||
### Notes
|
||
- The daemon runs cleanly with the cadence off or misconfigured (logs + disables, never
|
||
crashes); a leaked scratch guest from a mid-test crash is reaped by `engine.Recover` on
|
||
restart. `go test -race` covers the new scheduler goroutine.
|
||
- Slice-3/4/5 exported surfaces, goldens, and adversarial tests intact. Version bumps to
|
||
**v0.6.0** when Phase B (PBS) lands.
|
||
|
||
## v0.5.1 — slice 5 live-validation prep: durable_id mis-id fix + re-mount UUID memory (2026-06-09)
|
||
|
||
Two correctness fixes surfaced while preparing the live USB validation on `demo-felhom`
|
||
(a real 1TB USB HDD, sdb1, ext4). Both are DR-load-bearing — exactly the "false-id →
|
||
re-attach the wrong disk" failure mode the slice warned about.
|
||
|
||
### Fixed
|
||
- **Unmounted dir-storage no longer inherits the ROOT filesystem's UUID** (`observe.go`).
|
||
Previously, when a removable dir-storage was unmounted, the observer fell through to the
|
||
*containing* mount (root) for the backing device, so its `durable_id` became
|
||
`uuid:<root-uuid>` — a catastrophic DR mis-id (the hub would re-attach the wrong disk).
|
||
Now the backing device/UUID/`durable_id` are derived ONLY from the target's OWN
|
||
mountpoint; an unmounted target reports no device and a stable `store:<name>` durable_id,
|
||
never another filesystem's UUID. (Removed the `containingMountDevice` root-fallthrough.)
|
||
- **Watchdog remembers the fs-UUID observed while attached** (`watchdog.go`) so a re-mount
|
||
works even after the known-set cache refreshes mid-drop (an unmounted target can't resolve
|
||
its own UUID). The re-mount key is backfilled from this memory — aligning with doc 03 §7's
|
||
"sourced from the existing definition, no hub manifest needed": the agent learns the UUID
|
||
while the target is attached, then re-mounts by it on return.
|
||
|
||
### Tests
|
||
- Observer: an unmounted dir-storage asserts NO `uuid:` durable_id and no backing device.
|
||
- Watchdog: a drop where the cache lost the UUID still re-mounts using the remembered UUID.
|
||
|
||
## v0.5.0 — slice 5 Phase B: the host-root surface (mounts + SMART + grow + destructive gate) (2026-06-09)
|
||
|
||
The write surface — the agent's first step outside its Proxmox API token into OS-root.
|
||
Isolated behind a narrow, argument-validated, adversarially-tested seam, exactly like the
|
||
slice-4 gate. Completes slice 5 (Phase A = read-only observe/report/watchdog at v0.5.0-rc1).
|
||
|
||
### Added
|
||
- **`HostOps` seam + `SudoHostOps`** (`internal/storage/hostops.go`) — the one privileged
|
||
host surface: persistent mounts via **systemd `.mount` units keyed by fs-UUID** (enabled to
|
||
survive reboot), detach (stop+disable), SMART, and thin-pool metadata. Shells out via the
|
||
fenced Runner (`sudo -n`, **fixed arg vectors, no shell**); a fake backs the tests (no real
|
||
root in the suite). `NoopHostOps` is the safe fallback when the surface is unavailable.
|
||
- **The argument validator** (`internal/storage/validate.go`) — the security boundary:
|
||
`ValidateUUID` (strict hex), `ValidateMountPath` (absolute, no traversal, no metacharacters),
|
||
`ValidateSMARTDevice` (raw-disk whitelist), `ValidateLVMName`, and an in-process
|
||
`systemdEscapePath` (no `systemd-escape` shell-out). **Every argument is validated BEFORE a
|
||
command is constructed.** Headline test (`validate_test.go`): an adversarial matrix of
|
||
shell metacharacters / `../` traversal / malformed inputs is rejected with **zero exec**.
|
||
- **SMART** (`internal/storage/smart.go`) — parses `smartctl -a -j` into `StorageTarget.smart`:
|
||
**SATA** (reallocated/pending/offline-uncorrectable, temp, power-on-hours) **and NVMe**
|
||
(critical_warning, media_errors, percentage_used, temp), degrading to `UNKNOWN` for devices
|
||
with no SMART (USB-SATA bridges). **`lvs`** fills the lvmthin thin-pool **metadata** fill
|
||
(the value Phase A left null). Wired into the Observer's enrichment (Observe only, not the
|
||
watchdog's fast Known path).
|
||
- **Watchdog re-mount response** (`internal/storage/watchdog.go`) — on a known mount-backed
|
||
target's device returning **unmounted** (a new `DevicePresent` liveness probe), the watchdog
|
||
**dispatches a benign by-UUID re-mount off the poll path** (a goroutine, never under the
|
||
lock), rate-limited per target to the debounce window. The mount is routed through the gate
|
||
as benign (`gateRemounter` in `main.go`, so `storage` stays decoupled from `reconcile`).
|
||
- **Disk-grow executor** (`internal/reconcile`) — `ActionResize` (benign `ClassResize`), planned
|
||
**grow-only** (desired DiskBytes > actual → `pct resize rootfs +<n>M`; a shrink is refused,
|
||
never silently grown) + a defensive executor guard (size must start with `+`). New
|
||
`proxmox.Client.ResizeLXC` (API; `VM.Config.Disk`+`Datastore.AllocateSpace`; async→UPID).
|
||
Built + fixture-tested; **unfed** live (no hub spec until slice 10).
|
||
- **Destructive storage ops through the slice-4 gate** (`internal/reconcile/storage_ops.go`) —
|
||
`IntentForStorageMount` (benign) and `IntentForStorageDestructive` (`ClassStorageWipe`/
|
||
`ClassDecommission`). Host/target-scoped: the op binds on the storage **target identity**
|
||
(carried in `target.guest_id`). Reuses the existing verifier/role-scoping/binding/audit — no
|
||
new gate, no new crypto. Storage cases added to the adversarial matrix (`storage_test.go`):
|
||
unsigned wipe → `pending_signature`; "wipe A" signature vs "wipe B" → `binding_mismatch`;
|
||
valid → accepted. **Inert** live.
|
||
- **`--selftest=storage` [`-watch <dur>`]** — the live USB-runbook harness: an observe pass
|
||
(full table incl. SMART + thin-pool data+metadata), and a bounded watchdog window with the
|
||
re-mount response live. Runs standalone on the Proxmox host (no hub).
|
||
- **`configs/felhom-agent.sudoers`** — the documented narrow allowlist (install unit / systemctl
|
||
manage / smartctl / lvs), with the agent-side fine validation noted.
|
||
- **Config**: `privileged.{unit_dir,stage_dir,systemctl,install,smartctl,lvs}` (paths must match
|
||
the sudoers entries).
|
||
|
||
### Notes
|
||
- Daemon still runs cleanly with no removable storage / no signers / no hub manifest, and a
|
||
missing/declined sudoers entry degrades with a warning (SMART→UNKNOWN, mount→logged error),
|
||
not a crash. `go test -race` passes (the watchdog re-mount dispatches off the poll path).
|
||
- Slice-3/4 + Phase-A exported surfaces, goldens, and adversarial tests intact. `authz`
|
||
untouched. The destructive-storage executor + grow are built/tested but unfed live until
|
||
slice 10.
|
||
|
||
## v0.5.0-rc1 — slice 5 Phase A: storage observe + report + watchdog (read-only, live) (2026-06-09)
|
||
|
||
Phase A of the storage slice (doc 03 §7). Read-only and live: the agent now observes every
|
||
host storage target, reports it into the host-report's `storage_targets` (previously an empty
|
||
stub), and runs a fast-poll watchdog that pushes a disconnect to the hub in seconds. No
|
||
host-root writes this phase (mounts/SMART/grow/destructive-gate are Phase B). The hub-owned
|
||
desired manifest (class/role/policy/creds) is not served until slice 10, so reconcile against
|
||
it is built-but-unfed — this phase ships only the genuinely-useful read-only footprint.
|
||
|
||
### Added
|
||
- **`internal/storage` package** (new):
|
||
- **`StorageTarget` wire contract** (`internal/hub/report.go`) — filled the slice-3 stub:
|
||
`name`/`type`/`durable_id`/`state`/`reachable`, usage (`total`/`used`/`avail`/
|
||
`used_fraction`), `content`, `mount_path`/`backing_device`, `class_hint` (rotational HINT
|
||
— never authoritative; class is hub-owned), `role` (empty until slice 10), a `thin_pool`
|
||
sub-object (lvmthin data fill; metadata fill is Phase B/`lvs`), and a `smart` sub-object
|
||
(`UNKNOWN` until Phase B). Cross-repo golden kept byte-identical with `felhom.eu/hub` and
|
||
guarded by the bidirectional key-set test (`contract_test.go`).
|
||
- **`durable_id` derivation** (`durableid.go`) — deterministic per type (the DR-load-bearing
|
||
re-attach key): fs-UUID (usb/local-dir), `server:export` (nfs/cifs), `repo+fingerprint`
|
||
(pbs), `vg/pool` (lvmthin); never empty (falls back to a stable store id).
|
||
- **`HostReader` seam + `ProcHostReader`** (`hostread.go`) — non-privileged `/proc/mounts`,
|
||
`/dev/disk/by-uuid`, `/sys/.../rotational` + `removable` reads. Root-free by construction.
|
||
- **`Observer`** (`observe.go`) — builds `[]hub.StorageTarget` from `ListStorage`/`NodeStorage`
|
||
joined with host reads; surfaces the lvmthin thin-pool data fill prominently (warns ≥85%).
|
||
- **Storage watchdog** (`watchdog.go`) — a third daemon goroutine fast-polling the *known*
|
||
target set (a defined Proxmox storage and/or a previously-seen one) for
|
||
`attached↔disconnected` transitions; on a transition it triggers an immediate, **debounced**
|
||
out-of-band host-report. Only flags a *known* target's change (never a never-attached
|
||
device); coalesces flaps within the debounce window (leading + trailing edge).
|
||
`CachingKnownTargets` rate-limits the Proxmox-derived known set; `HostLiveness` probes
|
||
device/mount presence (local) + a reachability dial (network), all non-privileged.
|
||
- **Proxmox `Storage` type** (`internal/proxmox/types.go`) — additive parse-only config fields
|
||
(`server`/`export`/`share`/`datastore`/`fingerprint`/`vgname`/`thinpool`) feeding durable_id.
|
||
- **Collector `StorageObserver` seam** (`internal/hub/collect.go`) — populates `storage_targets`
|
||
via the observer; a nil observer or an observe error degrades to empty (never sinks the
|
||
heartbeat). Hub does not import storage (storage imports hub for the wire type).
|
||
- **Out-of-band report trigger** (`internal/hub/loop.go`) — `Loop.SetTrigger`: a watchdog
|
||
signal runs one extra collect→report immediately without disturbing the regular cadence.
|
||
- **`StorageConfig`** (`internal/config`) — watchdog interval / debounce / known-refresh knobs
|
||
(all optional; package defaults otherwise).
|
||
- **Hub ingest** (`felhom.eu/hub`) — `hostReportPayload` now parses `storage_targets`
|
||
(full mirror struct), persists them via `report_json`, counts + warns on disconnected
|
||
targets, and has its own half of the bidirectional golden key-set test.
|
||
|
||
### Notes
|
||
- The daemon still runs cleanly with no removable storage, no signers, and no hub manifest —
|
||
the watchdog finds nothing to flag; storage reporting is best-effort.
|
||
- `proxmox`/`hub`/`authz`/`reconcile` exported surfaces + their golden/adversarial tests are
|
||
intact. No host-root writes, no destructive paths, no SMART this phase (all Phase B).
|
||
- Version: **v0.5.0-rc1** at the Phase-A checkpoint; **v0.5.0** when Phase B lands.
|
||
|
||
## v0.4.0 — slice 4 Phase B: reversibility gate + signed-op consuming layer (2026-06-08)
|
||
|
||
The security core of slice 4: hub-supplied intent stops being trusted for destructive
|
||
change. Layered in front of the per-guest queue's executor — **every** mutation now
|
||
passes the gate. Reuses `internal/authz` for all crypto (untouched surface). Inert
|
||
this slice: no destructive deltas are served until slice 10, so the destructive path is
|
||
classified, gated, and adversarially tested but not wired to live execution.
|
||
|
||
### Added
|
||
- **Classifier (`classify.go`, doc 03 §4)** — benign vs destructive by **provenance +
|
||
data-bearing-ness, NOT by verb**. The `OpClass` vocabulary (seeded by the committed
|
||
slice-2 `op_blob.json`: `guest_destroy`) is the agent-side contract slice 10 matches.
|
||
Destroy/overwrite of customer data is destructive UNLESS **agent-internal**
|
||
provenance (same-journaled-transaction create → compensating rollback, or
|
||
agent-tagged scratch) makes it benign. `Provenance` is journal-recorded and **never
|
||
populated from the hub** (its zero value is the only thing an external intent may
|
||
carry). Unknown op class fails safe → destructive.
|
||
- **Reversibility gate (`gate.go`)** — `Gate.Authorize(intent, signed)`: benign →
|
||
allowed unsigned; destructive → requires a verified, role-authorized, action-bound
|
||
operator signature, else refused **`pending_signature`**, never executed. Every
|
||
decision is written to an `AuditSink` (audit is a signal, never the guard).
|
||
- **Signed-op consuming layer over `authz`** — verifies via `authz.Verifier.Verify`
|
||
(the locked pipeline, untouched), then enforces on the `VerifiedOp`:
|
||
- **Role-scoping (doc 04 §4)** — recovery key authorizes key-rotation re-pins ONLY;
|
||
operational key authorizes ordinary destructive ops + planned rotation.
|
||
- **Op-to-action binding** — verified `op` + host + guest + `params` must match the
|
||
gated action (a signature for guest X / op A can't authorize guest Y / op B);
|
||
params compared semantically (key-order/whitespace independent).
|
||
- **Signed-job orchestration (`job.go`)** — `RunSignedJob`: idempotency dedupe (the
|
||
op nonce as the journal key — a redelivered completed op is skipped, not re-run),
|
||
gate authorization, then journal-wrapped execution via an injected
|
||
`DestructiveExecutor` (nil this slice — authorized destructive ops are inert, no
|
||
executor wired until 6/7).
|
||
- **Crash-recovery consumer (`recover.go`, Note 1 / doc 03 §10)** — `Engine.Recover`
|
||
consumes the journal's `InFlight()` at startup: an op that crashed AFTER the Proxmox
|
||
POST and BEFORE its terminal record (`OpTaskRunning`, nonce already consumed) is NOT
|
||
covered by idempotency dedupe — only this resume-or-rollback resolves it (re-read the
|
||
task via the new `TaskStatusOnce`, record the real outcome; a no-task-id op is
|
||
abandoned fail-safe). Landed together with the signed-op executor, as Note 1 required.
|
||
- **Daemon wiring** — `runDaemon` builds the verifier from `config.Authz.Signers` (a
|
||
bad key / missing nonce-store path is a fatal misconfig; **no signers = nil verifier**,
|
||
the common slice-4 state), constructs the gate (+ `SlogAudit`), runs `Recover` before
|
||
issuing any mutation, and routes every reconcile action through the gate.
|
||
|
||
### Changed
|
||
- **Memory comparison canonicalized (Note 2)** — `desiredMemoryMiB` makes the
|
||
desired↔actual memory compare in the same MiB unit that is then written, so a
|
||
non-MiB-aligned `MemoryBytes` converges in one pass instead of re-issuing SetConfig
|
||
forever (the numeric cousin of the description-newline normalization). Test proves
|
||
convergence. Slice 10 should still serve MiB-aligned specs at the source.
|
||
|
||
### Tests (the security proof — each independently rejected)
|
||
- **Adversarial matrix** via the REAL `authz.Verifier` with in-test-minted SSHSIGs
|
||
(framing replicated in reconcile's test binary; production authz untouched, no signing
|
||
added to the verify-only package): unsigned destructive **job** → pending_signature;
|
||
unsigned destructive **desired-state delta** → pending_signature (distrusts hub
|
||
desired state, not just jobs); forged/unknown signer → `ErrUnknownSigner`; expired →
|
||
`ErrExpired`; **replayed nonce across an agent restart** (durable `FileNonceStore`) →
|
||
`ErrReplay`; wrong host → `ErrTarget`; wrong guest / wrong op / wrong params →
|
||
binding_mismatch; **recovery key on ordinary destructive** → role_denied;
|
||
**hub-supplied "scratch" tag ignored** → still destructive → refused; **valid + role +
|
||
target + fresh nonce → accepted**, and a second presentation → `ErrReplay` (nonce
|
||
consumed).
|
||
- Classifier (benign/destructive/provenance/key-rotation/fail-safe), role-scoping,
|
||
params binding, crash-recovery (resume OK / fail / still-running / no-task rollback /
|
||
unreadable / one-shot key applied on resume), signed-job idempotency (execute once,
|
||
dedupe redelivery, refused-not-executed, no-executor-inert, executor-error).
|
||
- Full module **race-clean** (`go test -race`) + vet clean on the Linux build server.
|
||
|
||
## v0.4.0-rc1 — slice 4 Phase A: reconcile engine (structural; runs live, unfed) (2026-06-08)
|
||
|
||
The agent-side control core's structural half. **Checkpoint marker** — `-rc1` is the
|
||
Phase-A push; awaiting validation before Phase B (the reversibility gate + signed-op
|
||
consuming layer) lands the final **v0.4.0**. Runs LIVE but UNFED: with no desired-state
|
||
provider until slice 10, the live engine computes an empty action set and performs
|
||
**zero mutations**.
|
||
|
||
### Added
|
||
- **`internal/reconcile`** package — the engine, the per-guest serializer, the
|
||
desired-state model, the normalization layer, and the durable op journal:
|
||
- **Per-guest serializer (`Queue`, doc 03 §10)** — the single choke point ALL
|
||
mutation sources funnel through. Same-vmid jobs run strictly one-at-a-time in
|
||
submit order; independent vmids run in parallel. Each vmid is a cond-var FIFO lane
|
||
(unbounded, non-blocking, order-preserving); graceful drain on `Close`.
|
||
- **Desired-state model + `DesiredProvider` seam** — `DesiredGuest` (per-field
|
||
optional: run-state / `*hub.GuestSpec` / `*description`), `DesiredState`. The only
|
||
live provider is **`EmptyProvider`** (slice 4 has no source); `StaticProvider`
|
||
feeds fixtures. The seam is where slice 10's hub-serving plugs in — no hub/local
|
||
source invented here.
|
||
- **Normalization layer (`FieldNormalizers`)** — reconcile compares *normalized*
|
||
desired-vs-actual so Proxmox round-trip quirks don't read as drift. `description`'s
|
||
trailing newline is the first registered case; the registry takes more (boolean
|
||
coercion, list ordering) as discovered. `normDesc` **promoted** out of
|
||
`cmd/felhom-agent/main.go` to **`reconcile.NormDescription`**; the `--selftest=task`
|
||
description round-trip now uses that shared helper (one source of truth for the quirk).
|
||
- **Plan engine (`Plan`, pure function)** — computes the minimal **benign** action set
|
||
(`Start`/`Stop`/`SetConfig`) for guests present in both desired and actual, with
|
||
normalized comparison, deterministic vmid ordering, config-before-run-state. Skips
|
||
provision (desired-absent-in-actual, slice 7) and destroy (actual-absent-in-desired,
|
||
gated, slice 10); never writes a config it couldn't first read (`SpecKnown`). Disk
|
||
(rootfs grow) intentionally not reconciled here.
|
||
- **Reconcile engine (`Engine`)** — reads desired+actual, plans, dispatches each action
|
||
onto the shared queue. Every Proxmox op handled per the mutate.go contract: non-empty
|
||
UPID → `WaitTask` + assert `exitstatus`; empty UPID → clean **synchronous** success
|
||
(slice-4 proven). Per-action failures are counted, not fatal (other guests still
|
||
converge).
|
||
- **Operation journal (`Journal`)** — durable fsync'd append-only JSONL mirroring
|
||
`authz.FileNonceStore`: records each op's lifecycle (started → task_running →
|
||
succeeded/failed) with its Proxmox task id (crash mid-op is detected and re-checkable
|
||
on restart via `InFlight()`), plus an **idempotency-key store** (`AlreadyApplied`) so
|
||
a one-shot op never re-runs across retries/restarts. Reconcile actions carry no
|
||
idempotency key (convergent — must re-run on real drift).
|
||
- **Daemon wiring (`runDaemon`)** — reconcile runs alongside the hub loop on the poll
|
||
cadence, **sharing the per-guest queue**. Journal path is a `journal.log` sibling of the
|
||
nonce store. The daemon runs cleanly with **no desired state and no signers** (reconcile
|
||
is a logged live no-op; a journal-open failure degrades to journal-less, never crashes).
|
||
|
||
### Tests
|
||
- Serializer: same-guest serialized (max-concurrency 1, submit order preserved) and
|
||
different-guests parallel (cross-waiting jobs both complete — would deadlock if not);
|
||
error propagation; drain-pending-on-close; submit-after-close.
|
||
- Normalization: description round-trip; unknown-field identity; extensibility seam
|
||
(synthetic boolean-coercion + list-ordering normalizers).
|
||
- Plan: run-state start/stop, spec drift (cores/memory), disk-not-reconciled,
|
||
description-newline-not-drift, unmanaged fields, spec-unknown skips config keeps
|
||
run-state, desired-absent skipped, combined ordering, empty-desired no-op, deterministic
|
||
vmid order.
|
||
- Engine: empty-provider zero mutations; async start (WaitTask); synchronous SetConfig
|
||
(no WaitTask); WaitTask failure + POST error counted failed; list error = pass failure.
|
||
- Journal: lifecycle latest-wins; in-flight survives restart; idempotency dedupe across
|
||
restart; failed key not applied; torn-trailing-line skipped.
|
||
- Full module **race-clean** (`go test -race`) on the Linux build server; vet clean.
|
||
|
||
### Not in this phase (Phase B)
|
||
- The benign/destructive classifier, the reversibility gate, and the signed-op consuming
|
||
layer over `internal/authz` (doc 03 §4 / doc 04) — added next, in front of the queue's
|
||
executor, landing **v0.4.0**.
|
||
|
||
## v0.3.2 — SetConfig selftest extension (slice-4 pre-check) (2026-06-08)
|
||
|
||
The gate before slice 4: prove `SetConfig` works live under the scoped token before
|
||
reconcile is built on it. **Self-gated live run PASSED** on `demo-felhom`/guest 9999.
|
||
|
||
### Added
|
||
- **Reversible `SetConfig` step appended to `--selftest=task`** (`cmd/felhom-agent/main.go`,
|
||
`selftestSetConfig`): read `GuestConfig` → write a `description` marker
|
||
(`felhom-selftest <RFC3339>`) → verify it landed → restore the original value (or
|
||
`delete` the key if it was absent) → verify the restore. Handles PVE's dual-mode
|
||
`SetConfig` return per the `mutate.go` contract: empty UPID = synchronous success
|
||
(printed `synchronous`); non-empty UPID = `WaitTask` + assert `exitstatus=OK`.
|
||
The existing snapshot → rollback → delete-snapshot steps are unchanged. First live
|
||
exercise of the **`VM.Config.*`** privilege cluster.
|
||
- **`normDesc` / `extraString` helpers** — `extraString` decodes a string-valued key
|
||
from `GuestConfig.Extra` (raw JSON); `normDesc` strips the trailing newline PVE
|
||
appends to `description` on read, so a written value round-trips equal.
|
||
|
||
### Finding (live)
|
||
- The LXC `description` write returned **synchronous (empty UPID)** — PVE applied it
|
||
inline, no task. The agent's dual-mode `SetConfig` modeling is correct: the
|
||
empty-string path is real and must not be treated as an error.
|
||
- PVE **appends a trailing `\n` to `description`** on read (stored URL-encoded as
|
||
`%0A`). A naive exact-match reconcile would see perpetual drift — slice-4 reconcile
|
||
must normalize `description` comparisons (hence `normDesc`).
|
||
|
||
### Ops
|
||
- Standing operator token (`felhom-agent@pve!agent`, privsep) **rotated** during this
|
||
run (the prior secret was not retrievable); role + both user/token ACL rows
|
||
re-confirmed at `/`. New secret stored out-of-band, **not persisted to the repo**.
|
||
Guest 9999 left pristine (stopped, no `description`, no leftover snapshot). Version → 0.3.2.
|
||
|
||
## Docs + live validation — no version bump (2026-06-08)
|
||
|
||
### Changed
|
||
- **Reflowed `CLAUDE.md`** — removed hard mid-paragraph line wraps (prose, list items, blockquotes now single-line, soft-wrapped); code blocks and tables untouched; rendered output unchanged.
|
||
- **Unified the REPORT/CHANGELOG convention** in `CLAUDE.md`: `CHANGELOG.md` is the cumulative log (newest on top); `REPORT.md` is overwritten with the most-recent implementation/validation only. Added an explicit **no-secrets** rule (never write tokens/passwords/keys into committed files; reference them as stored out-of-band).
|
||
|
||
### Added
|
||
- **`REPORT.md`** rewritten for the live `--selftest=task` validation on the demo host (`demo-felhom`): snapshot → rollback → delete-snapshot on guest 9999, each polled to `exitstatus=OK` under the `felhom-agent@pve!agent` privsep token (UPIDs name the token actor — privsep path genuinely exercised); 16-privilege `FelhomAgent` role + both user & token ACLs confirmed; `--selftest=read` clean. Closes the slice-1 "mutating ops unit-tested only" gap; `WaitTask` async foundation validated live → **slice 4 unblocked**. (Token secret stored out-of-band, not in the repo.)
|
||
|
||
## v0.3.1 — slice-3 validation follow-ups (2026-06-08)
|
||
|
||
### Changed
|
||
- **Collector keeps the known run-status on a `GuestConfig` failure** (`internal/hub/collect.go`):
|
||
previously a per-guest config-read error forced `status="unknown"`; now the run-status from
|
||
`ListLXC` is preserved (only the `spec` is dropped). An empty status is still normalized to
|
||
`unknown` (wire value is always `running|stopped|unknown`). Test renamed to
|
||
`TestCollect_GuestConfigFailureKeepsStatusOmitsSpec` and asserts the preserved `running` + nil spec.
|
||
- **`--selftest` usage** error string now reads `(want read|task|hub)`.
|
||
|
||
### Added
|
||
- **Cross-repo contract fixture** `internal/hub/testdata/host-report.golden.json` +
|
||
`TestHostReport_ContractMatchesGolden` — compares the marshaled `HostReport` field-name sets
|
||
(top level + `host` + `guests[0]`) against the golden, failing on any json-tag drift. The file is
|
||
**kept byte-identical** with felhom-hub's copy (duplicated contract until a shared types module;
|
||
revisit when slices 5/6 populate the empty collections). Version → 0.3.1.
|
||
|
||
## v0.3.0 — hub client + host-report + first daemon loop (slice 3) (2026-06-08)
|
||
|
||
The agent's first daemon: a periodic read-only host-report POSTed to the hub (the
|
||
heartbeat). No Proxmox mutations, no desired-state/signed-op consumption, no
|
||
storage/backup collection yet — those are slices 4/5/6.
|
||
|
||
### Added
|
||
- **`internal/hub`** package:
|
||
- **`HostReport`** wire contract (`report.go`) shared field-for-field with the hub
|
||
ingest: host metrics, guests (`vmid` + spec), `cloudflared` status, and the
|
||
`storage_targets`/`backups`/`restore_tests`/`pbs_snapshots`/`audit_tail`
|
||
collections **defined but emitted empty** (typed `[]`, slices 5/6 fill them).
|
||
- **`Collector`** (`collect.go`) builds the report from a read-only `proxmoxReader`
|
||
(adapted to the real `internal/proxmox` surface — node held by the client, value
|
||
returns, `proxmox.Guest`) + a `CloudflaredProber`. Partial-failure policy: a
|
||
failed `NodeStatus` is a hard error (skip the POST); a failed per-guest
|
||
`GuestConfig` degrades that guest to `status="unknown"` (spec omitted) but still
|
||
sends; a cloudflared probe failure → `"unknown"`, never fatal.
|
||
- **`CloudflaredProber`** + `SystemctlProber` (`systemctl is-active cloudflared`;
|
||
read-only — NOT a Privileged/root op; tunnel management is a later slice).
|
||
- **`Client`** (`client.go`): `POST /api/v1/host-report` with
|
||
`Authorization: Bearer <key>`, standard TLS (system roots or optional `ca_file`;
|
||
verification always on). Typed `*TransportError` / `*HTTPError`; the bearer token
|
||
never appears in any error.
|
||
- **`Loop`** (`loop.go`): the daemon — immediate first report then tick; adopts the
|
||
hub's `poll_interval_seconds` clamped to [60,3600]; resilient (a collect/report
|
||
error is logged and the loop continues); clean shutdown on context cancel.
|
||
- **`ControlEnvelope`**: only `poll_interval_seconds` is acted on; `blocked` /
|
||
`desired_generation` / `has_signed_ops` are parsed-but-ignored (logged at most)
|
||
pending reconcile (slice 4).
|
||
- **Config**: `HubConfig` (url/host_id/api_key/poll_seconds/timeout_seconds/ca_file),
|
||
`FELHOM_AGENT_HUB_*` env overlay, `HubConfig.Validate()` (mode-aware — proxmox-only
|
||
`--selftest=read|task` still runs without hub config), `WithDefaults()`, and
|
||
`Redacted()` now also blanks the hub key. `configs/agent.example.json` gains `hub`
|
||
(and `authz`) blocks.
|
||
- **`cmd/felhom-agent`**: the no-`--selftest` mode is now the **daemon** (poll loop);
|
||
added **`--selftest=hub`** (one collect+report, prints the report + envelope).
|
||
Version 0.2.0 → 0.3.0.
|
||
|
||
### Tests
|
||
- Report serialization (field names; empty collections are `[]` not `null`; spec
|
||
omitted when unknown); client (Bearer header, non-2xx→`*HTTPError`,
|
||
transport→`*TransportError`, **token never in error**); collector (host mapping,
|
||
guest spec, per-guest failure degrades-but-still-reports, NodeStatus hard error,
|
||
cloudflared error→unknown); loop (immediate first report, continuation after an
|
||
injected error, interval adoption + clamp); config (hub validate/redact/env).
|
||
|
||
### Notes
|
||
- `internal/proxmox` and `internal/authz` were **not touched** — no new proxmox
|
||
surface was needed (`ListLXC` already exposes status/maxmem/maxdisk; `GuestConfig`
|
||
exposes cores). The task's `proxmoxReader` sketch (node-arg/pointer/`LXC`) was
|
||
adapted to the real exports as instructed.
|
||
- **Defined-but-empty** this slice: `storage_targets`, `backups`, `restore_tests`,
|
||
`pbs_snapshots`, `audit_tail` (slices 5/6). **Parsed-but-ignored**: the envelope's
|
||
`blocked`/`desired_generation`/`has_signed_ops` (slice 4).
|
||
|
||
## v0.2.0 — `authz` signed-op verifier (slice 2) (2026-06-08)
|
||
|
||
Production form of the Phase-4 signing primitive: a key-type-agnostic SSHSIG
|
||
verifier for operator-signed destructive ops, with the full anti-replay/
|
||
authorization pipeline and a durable, crash-safe nonce store. What slice 4
|
||
(reconcile) will call to gate destructive desired-state deltas. No hub, no signing
|
||
CLI, no reconcile loop.
|
||
|
||
### Added
|
||
- **`internal/authz` — `Verifier`**: `New(signers, store, hostID)` + `Verify(blob,
|
||
sigArmored) (*VerifiedOp, error)`. Runs the LOCKED pipeline (order is
|
||
load-bearing): parse armor → namespace → parse pubkey → allow-list (by key
|
||
**material**, `pub.Marshal()` equality, not key_id) → crypto verify (over the
|
||
**raw received bytes**, never re-canonicalized) → parse blob → target → time
|
||
window → **nonce recorded LAST**. Each post-crypto stage rejects even with a
|
||
valid signature.
|
||
- **SSHSIG framing** (`sshsig.go`) via `golang.org/x/crypto/ssh` — `pem.Decode` →
|
||
strip 6-byte magic → `ssh.Unmarshal` → `ssh.ParsePublicKey` → recompute signed
|
||
data with the named hash → `pub.Verify` (dispatches on key algorithm). No
|
||
hand-rolled crypto. Key-type-agnostic: ed25519 / **sk-ssh-ed25519 (FIDO2)** /
|
||
rsa / ecdsa via the one path.
|
||
- **Fixed namespace** `felhom-op-v1` (package constant, never caller-supplied).
|
||
- **`OpBlob`** (corrected `host_id`/`guest_id` json tags) + **`VerifiedOp`** (op,
|
||
host/guest, params, key_id, matched signer). key_id is advisory/audit only —
|
||
never an authz input.
|
||
- **Typed errors**: `ErrMalformed, ErrNamespace, ErrUnknownSigner, ErrBadSignature,
|
||
ErrTarget, ErrExpired, ErrNotYetValid, ErrReplay` (errors.Is-friendly).
|
||
- **`NonceStore`** + two impls: `MemoryNonceStore` (tests) and **`FileNonceStore`**
|
||
— durable, crash-safe (fsync'd append log, replayed into an index on open,
|
||
periodic compaction, expiry-only pruning). A nonce is fsync'd to disk before
|
||
`SeenOrRecord` returns false; replay protection survives restart; I/O failure
|
||
fails safe (reports seen=true). Target generalization: host_id matched strictly,
|
||
guest_id surfaced for the caller to route.
|
||
- **Config**: `AuthzConfig` (nonce-store path + pinned operator `signers` tagged
|
||
`operational`/`recovery` with a key_id, as authorized_keys lines).
|
||
- **Version 0.2.0.**
|
||
|
||
### Tests
|
||
- Real OpenSSH interop via a committed `ssh-keygen -Y sign` vector (hermetic CI);
|
||
per-stage rejection (each with an otherwise-valid sig); the headline
|
||
**invalid-sig-does-not-burn-the-nonce** invariant; replay; **persistence across
|
||
restart**; synthetic **sk-ssh-ed25519** through the unchanged path; byte-exactness
|
||
(a re-serialized blob fails crypto — not re-canonicalized).
|
||
|
||
### Notes / corrections to the Phase-4 reference
|
||
- §7's `Target` lacked json tags (`host_id`/`guest_id`) — fixed.
|
||
- The doc paired "Go 1.24.4 / x/crypto v0.52.0", but v0.52.0 declares `go 1.25.0`
|
||
and does **not** build on Go 1.24. Resolved by upgrading the build server to
|
||
go1.26.0 (backward-compatible; felhom-controller/hub unaffected); the module is
|
||
`go 1.25.0` on x/crypto v0.52.0.
|
||
- Free function → constructed `Verifier`; returns the full `VerifiedOp`; typed
|
||
errors; clock-skew tolerance added; durable nonce store is the net-new work.
|
||
- **Shared-contract dependency flagged** (not built): the hub and the `felhom-sign`
|
||
CLI must emit byte-identical canonical JSON or signatures won't verify; a shared
|
||
canonicalizer both import would be the right home.
|
||
|
||
## v0.1.0 — Scaffold + `proxmox` interaction layer (slice 1) (2026-06-08)
|
||
|
||
First slice: stand up the host-agent project and its foundation — the typed
|
||
Proxmox interaction layer every other module will call. No reconcile loop, hub
|
||
client, signing, or storage/backup orchestration yet (later slices).
|
||
|
||
### Added
|
||
- **Project scaffold**: module `gitea.dooplex.hu/admin/felhom-agent`, binary
|
||
`felhom-agent` (`cmd/felhom-agent/`), Go 1.24, zero external dependencies
|
||
(pure stdlib). `--version` flag; `version` var overridable via
|
||
`-ldflags "-X main.version=<v>"`.
|
||
- **`internal/proxmox` — API backend (`Client`)**: hand-rolled REST client over
|
||
`https://<host>:8006/api2/json` with `PVEAPIToken` auth. Typed read ops
|
||
(`Version`, `Nodes`, `NodeStatus`, `ListLXC`, `GuestStatus`, `GuestConfig`,
|
||
`ListStorage`, `NodeStorage`, `StorageContent`) and async mutating ops
|
||
returning a UPID (`RestoreLXC` — the primary create path, `Vzdump`, `Snapshot`,
|
||
`Rollback`, `DeleteSnapshot`, `SetConfig`, `Start`, `Stop`).
|
||
- **`WaitTask`**: polls `GET /nodes/{node}/tasks/{upid}/status` until stopped, then
|
||
asserts `exitstatus == "OK"` (authorization can surface at task execution, not
|
||
the POST — phase1-2 §1.3). Exponential backoff (1s→5s cap), context
|
||
cancellation + timeout. `*APIError` parses the offending privilege from a 403;
|
||
`*TaskError` parses it from a failed task exitstatus + log tail.
|
||
- **`internal/proxmox` — fenced root-CLI backend (`Privileged`)**: limited to the
|
||
three proven OS-root exceptions only — `CreateGoldenLXC` (keyctl `pct create`),
|
||
`MountUSBByUUID`, `SMART`, `Sensors`; each cites why it can't be the API. Fence
|
||
is structural (Client never shells out, Privileged never makes an HTTP call) and
|
||
asserted in tests.
|
||
- **TLS trust**: SHA-256 leaf-cert pinning (the host serves a self-signed cert) or
|
||
a CA file; an explicitly-named `insecure_skip_verify` that is off by default. No
|
||
blanket verification disable.
|
||
- **`internal/config`**: JSON config file + `FELHOM_AGENT_*` env overrides; the
|
||
token secret is never logged (`Redacted()`).
|
||
- **`internal/log`**: slog setup (text, stderr, configurable level).
|
||
- **`cmd/felhom-agent --selftest`**: read-only health report against a live host
|
||
(version/nodes/status/guests/storage); `--selftest=task --vmid N` exercises
|
||
`WaitTask` on a reversible snapshot→rollback→delete op (gated; default selftest
|
||
mutates nothing).
|
||
- **Tests**: unit tests with a mock HTTP transport + mock runner (UPID parse,
|
||
`WaitTask` running→OK / failed-403 / timeout / ctx-cancel, 403→privilege error,
|
||
response decoding against shapes captured live from `demo-felhom`, config
|
||
redaction, and the API-vs-root routing fence).
|
||
|
||
### Notes
|
||
- Types are grounded in the spike findings
|
||
(`felhom.eu/documentation/proxmox-platform.md`, `tests/phase{0,1-2,3}-findings.md`)
|
||
and the exact JSON shapes captured live from `demo-felhom` (PVE 9.2.2).
|
||
- Verified: `go build/vet/test` green on Go 1.24.4 (build server) and a live
|
||
read-only `--selftest` against the demo host with TLS fingerprint pinning.
|
||
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
|
||
token) is provisioned out-of-band; the agent only consumes the token.
|