# Changelog All notable changes to **felhom-agent** are recorded here. Update on every code change that gets pushed. ## v0.18.0 — slice 10D: DR capstone — identity escrow + restore-mode consumption (agent side) (2026-06-10) The agent half of the slice-10 DR capstone (closes slice 10). Grounded by both 10-series spikes (escrow-consumption + identity-restore). The hub half (recovery-mode toggle, re-enroll + credential rotation, directive serving) is hub v0.11.0. **Operator-side rotation model (locked):** the hub holds no Cloudflare write-power; the destructive tunnel/PBS rotation is the operator's step from a trusted environment (same spirit as 10B). ### Added (`internal/escrow`) - **Identity escrow** (`identity.go`): `WrapIdentity`/`UnwrapIdentity` (+ `…Bundle`) wrap the `{tunnel_token, pbs_token}` bundle under the SAME recovery code `R` via **`age`** (scrypt + ChaCha20-Poly1305 — a vetted passphrase-AEAD, not hand-rolled), reusing the K-escrow pty mechanism (passphrase via the tty, data via files; `R`/tokens never logged). Same two-factor, zero-knowledge shape as the K-escrow. A **wrong R fails closed** (no bundle). `age` is a runtime dep for the identity path (analogous to proxmox-backup-client for K). - **`escrow.Create`** gains an optional `IdentityBundle` → also emits an `IdentityBlob` under the same R (additive; the K-escrow + 10C `Consume` paths are byte-unchanged). Self-verifies the identity round-trip before shipping. - **`--selftest=escrow-create -identity-bundle -directive `** — also wrap + upload the identity blob + the **non-secret** DR directive (pbs repo/ns, expected key fingerprint, tunnel id). - **`--selftest=identity-consume -blob -keydest `** (R via `FELHOM_RECOVERY_CODE`) — recover the identity bundle through the real code; tokens written 0600, never logged. ### Tests - identity bundle round-trips (wrap→unwrap byte-identical; blob is opaque ciphertext); wrong R fails closed + the blob stays retryable; input validation. K-escrow/10C tests byte-unchanged (additive). (age integration tests gated to a host with the `age` CLI.) ## v0.17.0 — slice 10C: escrow consumption (productionize the spike) (2026-06-10) Turns the throwaway 10C spike harness into a real, tested **`Consume`** path: recover the PBS key `K` from an R-wrapped escrow blob, **gate it on the expected fingerprint**, and install it for the restore. The spike already proved the crypto + real-data restore; this bakes its findings into production code. **Agent-only** — 10C *reads* the four inputs as parameters (so it stays standalone-testable); 10D sources blob/fingerprint/PBS-connection from the hub and prompts for R. **Zero-knowledge holds**: the hub serves everything except **R** (by hand from the customer), so a hub compromise alone still can't decrypt. ### Added - **`escrow.Consume(ctx, blob, R, expectedFingerprint, keyDest)`** — the consumption contract: 1. **Unwrap** the blob (a copy — F-C6: the input blob is read-only → a failed Consume is **retryable**) with `R`; a **wrong R fails closed** at the scrypt KDF (F-C3) → a clear, R-free error, **nothing written**. 2. **Fingerprint gate (F-C4)** — `KeyFingerprint(recovered)` must equal the expected (the hub knows it); a mismatch **fails fast + loud, no install, no restore attempted**. 3. **Atomic install (F-C2)** at `keyDest` (`0600`, write-temp-sibling→rename); any failure leaves **no partial install**. The recovered key lives only in a `0700` tempdir that is always removed. **Secret discipline:** `R` and key bytes are never logged/persisted (only fingerprint prefixes); `K` is never mutated. - **`--selftest=escrow-consume`** (`-blob -fingerprint -keydest`, R via env `FELHOM_RECOVERY_CODE` to keep it off the command line) — invokes the real `Consume` live (the spike's S3 via the production path, not a harness). ### Tests (non-hollow) - valid → key installed + `KeyFingerprint(dest) == expected` + `0600` + blob byte-unchanged; **wrong R** → error, **no file at dest**, blob unchanged; **fingerprint mismatch** → fail fast, **no install** (the gate runs before any restore); input validation; format-tolerant fingerprint compare (no empty-fingerprint gate-bypass); atomic-install permissions (integration tests gated to a host with `proxmox-backup-client`). ## v0.16.0 — slice 10B: operator-signed destructive completion (offline key + signing CLI) (2026-06-10) The security centerpiece: a destructive op runs ONLY on a verified, operator-signed authorization — signature valid against a **pinned** operator pubkey (never the hub's or the blob's), nonce unseen + durably burned, in-window, host-bound, and **resource-bound to a DURABLE device id** that execution re-resolves + re-inspects. Decision (a): **offline operator key + signing CLI**, hardware-key-ready (`sk-`/YubiKey via ssh-keygen). The key floor holds: the signing key is NOT in the hub and NOT in the agent. Concrete consumer: this **closes the 8C data-bearing-wipe `pending_signature` gap**. Pairs with hub v0.10.0. ### Added - **`cmd/felhom-opsign`** — the operator's offline signing CLI. Builds the canonical `OpBlob` by **reusing `authz.CanonicalBlob`** (the exact production path the verifier authenticates over — so signer + verifier can never drift) and signs it with **`ssh-keygen -Y sign -n felhom-op-v1`** (hardware-ready). Output: a `{op_blob_b64, sig_armored}` envelope to hand to the hub jobs queue (optional `--upload`). Touches ONLY the operator's signing key. - **`authz.CanonicalBlob`** — promoted to production (was test-only) so the CLI + verifier share one canonical-bytes source; params canonicalized (sorted keys, compact). - **`internal/storage` durable device identity** (`durable_device.go`): `DeviceDurableID` (derive a stable `byid:`(wwn/serial)/`byuuid:` id from the world-readable udev symlinks — no privilege, no subprocess) + `ResolveDurableDevice` (re-resolve to the current `/dev` path; a path-only/unknown scheme is REFUSED). The resource-level anti-retarget. - **`internal/signedjobs`** (new): the queue consumer. `Runner` fetches each opaque job → runs it through the **gate** (the LOCKED authz pipeline) → on all-pass hands the verified op to an `Executor`; the order is **verify → nonce-burn (durable, in Verify) → execute → clear job**. The **`WipeExecutor`** is the 8C consumer: resolve the signed durable id → **re-derive + match** (anti-retarget) → **re-inspect (8C classifier)** the device is still the data-bearing target → `mkfs`. A vanished/changed/non-data-bearing device or a path-only binding is refused **even with a valid signature**. Wired as a second `EnvelopeObserver` (runs on `HasSignedOps`). - **`hub.Client.Jobs` / `CompleteJob`** + `hub.MultiObserver`; the 8C format refusal now **surfaces the bound op** (op + durable id + host) in its 403 `pending_op` + a `felhom-opsign …` hint. ### Pinning / rotation - Operator pubkeys are pinned via `authz.signers` (config, trusted path — provision/agent config, NEVER hub-alone), **multiple** keys (KeyID selects; role-scoped), so a backup/rotation key exists without a flag-day. Unchanged from the slice-4 verifier wiring; 10B activates the execute path. ### Tests (real crypto, non-hollow) - `signedjobs` runner over the **real** gate+verifier (in-Go minted SSHSIGs): valid → executor runs once + job cleared; **replay** (nonce burned) / **non-pinned signer** / **expired** / **retarget** (other host) / **forged sig** / **no pinned signer** → all rejected, **executor never called**; malformed envelope cleared. - `WipeExecutor`: valid → `mkfs` runs; **path-only**, **durable-id mismatch**, **device gone**, **re-inspect non-data-bearing**, **not-probed** → all refused, `Format` not called. - `storage` durable: wwn-preference, uuid-fallback, path-only/traversal refusal, round-trip, missing-device error (symlink tests gated to Linux — the agent's OS). ## v0.15.0 — slice 10A: hub desired-state serving — the "Down" channel (2026-06-10) The agent half of slice 10A. The control envelope (`hub.ControlEnvelope`) stops being "reserved — ignored" and becomes the live **Down channel**: a cheap change-notification on every heartbeat. The agent caches the hub's desired-state + its generation; only when **`DesiredGeneration` advances** does it fetch the full state (the heartbeat stays light, the heavy state moves on change). The engine then reconciles **benign** deltas and the gate marks an explicit **destructive** delta `pending_signature` (no signer in 10A → never executed; signed execution is 10B). Pairs with hub v0.9.0. ### Added / changed - **`internal/reconcile`**: `DesiredGuest.Decommission` — the canonical **destructive desired-state delta** (an EXPLICIT flag, not "absent from the list", so a partial hub list can never mass-destroy). The planner emits `ActionDecommission` → `ClassDecommission` → Destructive → the gate refuses it `pending_signature`. `Reconcile` now counts a `pending_signature` refusal as **`Result.Pending`** (expected, logged INFO) rather than a failure; any other refusal stays a real failure. `ActionDecommission` has **no executor** (slice 10B) — a defensive guard refuses to run it. New **`CachingProvider`** (thread-safe DesiredState + generation cache; `Desired`/`Update`/`Generation`) — the production `DesiredProvider`, replacing `EmptyProvider` in the daemon engine (empty until the hub serves intent → cold-start is a live no-op, unchanged). - **`internal/hub`**: the **`ControlEnvelope`** fields are now active (DesiredGeneration drives the fetch, HasSignedOps noted). New wire types **`DesiredStateResponse`** + **`WireDesiredState`** (guests + forward-compat `restore_directive` (10D) / `pbs_namespace` / opaque `storage_manifest`+`backup_policy`) + **`WireDesiredGuest`** (vmid/run/spec/description/decommission). New **`Client.FetchDesiredState`** (GET `/api/v1/hosts/{host_id}/desired-state`, self-scoped to the client's own host). New **`EnvelopeObserver`** loop seam + `SetEnvelopeObserver` — the loop hands the envelope to the sync layer each cycle (hub does not import reconcile/desired). - **`internal/desired`** (new): the **`Syncer`** — implements `hub.EnvelopeObserver`, fetches desired-state on a generation advance, maps the wire shape to the reconcile domain, and updates the `CachingProvider`. Caches the **fetched** generation (robust to a generation that advanced mid-fetch); a fetch failure keeps the last-known state. `restore_directive` is carried + logged, not acted on (10D). Wired in `cmd/felhom-agent` (daemon): provider → engine, syncer → loop. ### Tests - reconcile: a desired-state with one benign + one decommission delta → **benign applied, destructive gated pending (not executed)**; `Plan` emits decommission-only for a decommissioned guest + classifies Destructive; `CachingProvider` update/isolation. - desired: **fetch-once-on-advance** (no re-fetch on an unchanged generation), fetch-failure-keeps-cache, caches-the-fetched-generation. - hub client: `FetchDesiredState` hits the self-scoped path with the bearer + decodes (incl. `restore_directive`); a 403 is a typed `HTTPError`. - loop: the cycle notifies the observer + adopts `PollIntervalSeconds`; a report error skips the observer. - cross-repo golden: `testdata/desired-state.golden.json` + `control-envelope.golden.json` decode + key-set guard, **byte-identical** with felhom.eu/hub. ## v0.14.0 — slice 9: host metrics to the controller (`GET /host/metrics` + CPU-temp collector) (2026-06-10) The de-privileged controller (slice 8C) sees only its own cgroup, so it can't read host health itself. Slice 9 **re-serves** the slice-4 collector's host + per-storage view to the customer over the local API, plus the one missing collector — CPU/chassis temperature — so the customer sees their box's health in the controller. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot). Assumption: **one customer per host** (the home-server model); if a host ever serves multiple customers, host-wide CPU/mem would leak cross-customer load → revisit then. ### Added / changed - **CPU/chassis-temp collector** (`internal/hub/cputemp.go`): `SysfsTempReader` reads the CPU package temperature straight from sysfs — hwmon (`coretemp`/`k10temp`/`zenpower`/`cpu_thermal`, preferring the `Package id 0` input) then the thermal zones (preferring `x86_pkg_temp`/`coretemp`/`cpu-thermal`, falling back to `acpitz`). **No external binary, no privilege** (sysfs nodes are world-readable), so the root-CLI fence is untouched. **Graceful-null**: a missing sensor, an unsupported board, an implausible reading (outside 5–150 °C), or any read error all degrade to `null` ("n/a") — a missing sensor never fails the report. Wired into the collector via the new `TempReader` seam (nil-safe). - **`HostMetrics.CPUTempC *int` (`cpu_temp_c`)** — new nullable wire field on the **shared** `HostMetrics` struct (same nullable contract as the disk `SmartSummary.TemperatureC`). It rides the **hub report too** (operator freebie) → cross-repo host-report golden updated. - **`Collector.HostMetricsNow(ctx)`** — a fresh `NodeStatus` + CPU-temp read returning just the host block, the source for the local API (current cpu%/temp, not the 15-min snapshot). `Collect()` now also populates `cpu_temp_c` on the hub report. `Collector.SetTempReader` injects a fake in tests. - **`GET /host/metrics`** (`internal/localapi/host_metrics.go`): host-wide health (cpu%/mem/load/uptime/`cpu_temp_c`) + per-storage capacity targets (total/used/fraction, thin-pool, SMART temp+wear). Token-authed via `withGuest` (host-wide data; cross-guest `?vmid=` still 403). Best-effort on storage (a view error still returns the host block). Served only when the `HostMetrics` provider (the shared collector) is wired — else 503 "not configured". Wired in `buildLocalAPIServer`. ### Tests - `cputemp_test.go`: a fake `/sys` layout proves hwmon package-preference, hwmon first-input fallback, thermal-zone-by-type selection over a non-CPU hwmon, **graceful-null on a sensorless host** (no error), and rejection of implausible (0 m°C) readings. - `hostmetrics_test.go`: `HostMetricsNow` populates the temp, gracefully nulls it, hard-errors on `NodeStatus` failure; `Collect()` carries the temp. - `host_metrics_test.go` (localapi): populated host+storage with a valid token; `cpu_temp_c:null` serializes; **401 without a token** (collector never invoked); 403 on a cross-guest `?vmid=`; 503 when not configured. ## v0.13.0 — slice 8B.2: quiesce downtime optimization (`snapshotted` phase) (2026-06-10) The agent half of slice 8B.2. In snapshot mode, vzdump only needs the app-stopped state captured at the **storage-snapshot moment**; after that it reads from the snapshot and the app can resume. The agent now emits a **`snapshotted`** phase on `GET /backup/status` when the snapshot is taken, so the controller (v0.38.0) resumes its app early — app downtime drops from *whole-backup* to *until-snapshot* with no loss of app-consistency. Validated Phase-0 first on PVE 9.2.2: the marker is `INFO: create storage snapshot 'vzdump'`; downtime ~24s→~1s for a 934 MB guest. ### Added / changed (`internal/backup` + `internal/localapi`) - **`BackupRunner.BackupWithSnapshotHook(ctx, vmid, onSnapshot)`** — while the vzdump runs, a watcher tails the task log (`TaskLogTail`) for the **`create storage snapshot`** marker and fires `onSnapshot` **once**. The marker only appears in snapshot mode (stop/downgraded takes no storage snapshot), and the watcher also bails on `backup mode: stop` — so it never fires in stop mode. (`Backup` keeps its signature for the scheduler/selftest; both share one body.) - **`/backup/status` phase `snapshotted`** (between `running` and `done`): `handleBackup` passes the hook → `markSnapshotted` flips the running job to `snapshotted`. `done`/`failed` semantics unchanged. ### Tests - localapi: snapshot mode → phase reaches `snapshotted` before `done` (gated fake holds the backup open); stop mode → `snapshotted` **never** emitted (stays running → done). runner: the watcher fires `onSnapshot` on the marker; in stop-mode log it never fires. `snapshotWatchInterval` is a package var so tests run fast. ## v0.12.0 — slice 8C Phase A: disk endpoints + data-bearing classifier gate + mkfs executor (2026-06-10) The agent half of slice 8C, Phase A (additive). Adds the host disk-management endpoints the controller's disk UI drives — with the **8C security invariant**: the agent decides data-bearing-ness by **inspecting the actual device** (agent-internal evidence), NEVER from the caller's claim. A compromised controller asserting "this drive is blank" cannot wipe a data-bearing drive. (Controller rewire + disk-subsystem retirement + de-privilege are Phases B/C, `felhom-controller`.) ### Added - **`internal/storage` — `mkfs` executor + data-bearing inspection.** `SudoHostOps.Format(device, fstype)` (device-pinned, `ValidateBlockDevice`+`ValidateFSType`, narrow `FELHOM_FORMAT` sudoers — `mkfs.ext4 -F` / `mkfs.xfs -f` on a `/dev/*` path the agent fine-validates first). `SudoHostOps.InspectDevice(device)` → `DeviceProbe` (filesystem signature via `blkid -p`, partition table / partitions / mount via `lsblk -J`). **`DeviceProbe.DataBearing()` is conservative**: any signature / partition table / partition / mount — OR a probe that did not read cleanly — is data-bearing (fail-safe; an unreadable device is never called blank). - **`internal/localapi` — the §6 disk endpoints**, all self-scoped (token→guest; cross-guest 403): - `GET /disks` — host drives + a **data-bearing flag** (UI hint). Read-only/benign. - `POST /disks/assign` — attach a drive as a mount (benign, additive → `EnsureMount`). Self-serve. - `POST /disks/eject` — safe-unmount (benign, data preserved) + the **dependent guests** that mount it (so the controller can warn which apps lose that storage). - `POST /disks/format` — **the security centerpiece**: the agent **inspects the device itself**; blank → benign → `mkfs`; **data-bearing → ClassStorageWipe → the slice-4 gate → refused `pending_signature`** (the operator-signed completion is slice 10). The caller's claim is ignored — only a device the agent reads as blank is formatted. - `storageGateAdapter` bridges the format path to the slice-4 reversibility gate (no new gate/crypto). ### Tests - localapi (security matrix): blank device → **mkfs called, gate not consulted**; a **data-bearing device → 403, mkfs NEVER called**, gate consulted (`pending_signature`); an **ambiguous/unprobed device → treated destructive** (fail-safe); even a gate that *allows* does not format data-bearing in 8C; assign → `EnsureMount`; eject → `Unmount` + dependent guests; cross-guest → 403; bad device/fstype → 400; unconfigured → 503. - storage: `ValidateBlockDevice`/`ValidateFSType` (whitelist + injection rejection); `InspectDevice` blank/filesystem/partition-table/mounted/failed-probe-fail-safe; `Format` invokes the right `mkfs.*`. ## v0.11.0 — slice 8B: app-consistent backup — /backup/due policy + /backup/status phases (2026-06-10) The agent half of slice 8B (doc 03 §8). Turns the 8A thin backup stubs into the real policy the in-guest controller's quiesce loop drives (controller half: `felhom-controller` v0.36.0). No hub change. The downtime optimization (`vzdump --mode snapshot` + a `snapshotted` phase) is the 8B.2 fast-follow; the hub-served per-guest policy is slice 10. ### Changed (`internal/localapi`) - **`GET /backup/due`** — real **cadence** policy (replaces the 8A "never backed up" stub): a guest is due when no **successful** backup is recorded OR the newest one is older than the agent-local cadence (`backup.backup_cadence_seconds`, default 24h). A successful `POST /backup` flips due to **false** for the window, so the controller won't re-quiesce in a loop. A failed backup does not satisfy the cadence. Returns `age_seconds` for diagnosis. - **`GET /backup/status`** — real **phases** `idle | running | done | failed` + the job id, so the controller can poll a backup to completion (was: just the latest stored backup). - **`POST /backup`** — returns a **job id** + `running` phase; tracks the in-flight job and is **single-flight per guest** (a second POST while one runs returns the same job — no concurrent vzdump). On completion the job transitions done/failed and the result is recorded to the store. - Config: `backup.backup_cadence_seconds` + `BackupCadence()`; the local-API server takes the cadence. ### Tests - `/backup/due`: due when stale / no backup, **not due within the window after a success**, due again past the cadence, **a failed backup does not count**. `/backup/status`: running→done and running→failed (gated fake to observe the running phase). `POST /backup` single-flight (one vzdump for concurrent POSTs). All still self-scoped (token→guest). ## v0.10.0 — slice 8A: agent local-API server + provisioning back-half (2026-06-10) The host-agent half of slice 8A (doc 03 §6). Adds the per-guest **local API** the in-guest controller calls over the bridge, and the **provisioning back-half** that follows the slice-7 bring-up front half. Grounded by `felhom.eu/documentation/tests/slice8a-channel-deploy-spike-findings.md` (commit `4a81a96` — channel + deploy plumbing proven; the 5 gotchas resolved here). Controller half is `felhom-controller` v0.35.0. No hub change. ### Added - **`internal/localapi`** — the HTTPS local-API server (doc 03 §6), the **per-guest authorization gate**. Serves a **persisted self-signed leaf** with a **stable SHA-256 fingerprint** (generated once; a fresh cert each boot would invalidate every baked bootstrap pin). The **7 §6 endpoints**, all **self-scoped to the caller's own guest**: `GET /storage` (this guest's mpN mounts + fast/slow class from the slice-5/7 storage view), `POST /snapshot`, `POST /rollback`, `POST /backup` (enqueued, crash-consistent — the app-consistent quiesce loop is 8B), `GET /backup/due` (thin in 8A), `GET /backup/status`, `GET /restore-test/status`. - **Token store** (`tokenstore.go`): durable, crash-safe per-guest token→guest map that persists only a **SHA-256 hash** of each token (the plaintext exists transiently at mint→write-to-mount, then is discarded), last-write-wins per guest, fsync'd append-only JSONL (mirrors the nonce store). - **Self-scoping**: the VMID is resolved ONLY from the token; an explicit `vmid` (query/body) that disagrees → **403 and the proxmox op is never issued for the other guest**; absent/unknown → 401. - **`internal/provision`** — the back-half: mint the per-guest token → render the stable **`bootstrap.json`** contract (schema `felhom.bootstrap/v1`; **no registry credential** — the controller image is baked into the golden) → write it `0600` → **`chown 100000:100000`** (the unprivileged-LXC mapped guest-root, spike gotcha 1) → attach a **read-only bind mount** via `pct set`. Host-side only (F3 — the agent never enters the guest; **no `pct exec`**). The token plaintext is never logged and never returned. - **`--selftest=provision`** — the full chain on-demand: bring-up (provision) front half + the back half; keeps the guest for the golden's baked controller-bootstrap unit to deploy. - **`config.LocalAPIConfig`** (`local_api`) — enable + bridge `listen_addr` + cert/key paths + token store path. The server is an optional 6th daemon goroutine, disabled cleanly when unconfigured or on a token-store/cert failure (the daemon still reports/reconciles). - **`configs/build-golden.sh`** now **bakes the controller image** (pulled once on the trusted build host, then `docker logout` — no cred baked) + a **controller-bootstrap unit** that deploys the **baked** image from the config mount on boot (no login/pull at deploy). - **`configs/felhom-localapi-firewall.example`** — host firewall narrowing of the local-API port to the guest bridge subnet (nft/iptables/PVE variants; defense-in-depth — the token stays the gate). - **`configs/felhom-agent.sudoers`** — a narrow `FELHOM_PROVISION` alias (`chown 100000:100000` + `pct set` bind-mount, both confined to the agent-owned `/var/lib/felhom-agent/guests/*` path) for the non-root least-privilege deployment. ### Security / design notes - The local-API leaf is pinned by **leaf-cert SHA-256** (decision: consistency with the agent's PVE/PBS pinning); the fingerprint is baked into each guest's bootstrap. - The back-half's host-root ops (chown + bind-mount attach) are **NOT** added to `proxmox.Privileged` (which is fenced to its 3 exceptions) — they live in `internal/provision` and run through the shared `Runner` (direct as root, or `sudo -n` with the new sudoers alias). This is the per-guest provisioning host-root surface, host-side and F3-compliant. ### Tests - localapi: self-scoping (cross-guest snapshot/rollback/backup → 403, op never issued for the other guest; own-guest uses the token's VMID), 401 paths, `/storage` class mapping, `/backup` enqueue, the thin `/backup/due`, status scoping; the token store persists only the hash (plaintext never on disk), last-write-wins, survives reopen, uniqueness; the leaf fingerprint is stable across reload. - provision: writes `0600` + chowns + attaches the bind mount with the right args; the **token never appears in the Result**; the cross-repo `bootstrap.json` contract key-set is pinned. ## v0.9.0 — slice 7 close-out: PBS recovery-code escrow creation (2026-06-10) The first code that touches the PBS client encryption key `K` and introduces the customer recovery code `R`. Default posture is **zero-knowledge**: Felhom holds an opaque `R`-wrapped blob (cannot open it), the customer holds `R`. Grounded by `felhom.eu/documentation/tests/slice7-escrow-spike-findings.md` (round-trip proven on a throwaway: the `R`-recovered key restores a real encrypted snapshot). Hub opaque storage is the `felhom.eu` half (hub v0.8.0); consumption/serving is slice 10. ### Secret discipline (overriding) `R` is `crypto/rand`, ≥128 bits, surfaced **exactly once** and **never** logged/persisted/committed; the wrap pty's echo is discarded so `R` can't leak. `K` is read by location, **never modified** (the live key file is byte-unchanged — Wrap operates on a copy), never logged. ### Added - **`internal/escrow`** — `Create` generates `R` (10 EFF-wordlist words ≈ 129 bits), wraps `K` under `R` via the **PBS-native** `proxmox-backup-client key change-passphrase --kdf scrypt`, and **self-verifies** the blob recovers `K` (fingerprint match) before shipping. The wrap is driven over a **stdlib pty** (`x/sys/unix`; spike F-A1 — the command is TTY-only) with **output discarded** (F-A2 — the pty echoes the passphrase). Opt-in outputs: **(b)** `R`-wrapped offline copy (two-factor, no extra trust) and **(a)** raw paperkey (single-factor, unrevocable — loud caveat). - **`--selftest=escrow-create`** (`-storage`, `-paperkey`, `-offline`, `-upload`): surfaces `R` once to stdout (never the logger), prints the opaque blob's size/fingerprint/posture, and with `-upload` PUTs the blob to the hub (`/api/v1/hosts/{host_id}/escrow`, per-host key). - Config: `escrow` section (`posture` default `zero_knowledge`, `pbs_storage_id`); `PBSEncKeyPath` helper (the `.enc` key K). - Runtime dependency on the `proxmox-backup-client` CLI (the PBS key+passphrase KDF). ### Tests - `R` entropy ≥128 / 10-word format / uniqueness; integration round-trip (wrap→unwrap fingerprint match, **wrong-`R` fails**, **live `K` byte-unchanged**, blob ≠ plaintext key) guarded to linux+`proxmox-backup-client`; the agent→hub wire-contract key-set (mirrors the hub's). - **Live-validated** (demo): `escrow-create` → `R` (10 words) surfaced once, blob 383 B opaque, self-verify ok, **live `K` sha256 unchanged**, exact `R` absent from stderr/journal. ## v0.8.0 — slice 7 Phase 1: unified bring-up reconcile job (provision + guest-loss DR) (2026-06-09) The shared FRONT HALF of provision and guest-loss DR, as a journaled reconcile job mirroring the slice-6 restore-test's crash-safety — but it KEEPS the guest on success and applies a scenario-specific identity policy. Agent-only; no hub/wire change (the new guest auto-appears in the host-report via `ListLXC`). Grounded by the slice-7 bring-up spike findings (commit `3342993`): F1 (restore preserves the archived MAC → provision reset is unconditional), F3 (SSH host keys do not auto-regenerate → a baked golden first-boot unit, not an agent guest-internal op), F4 (the transient PVE config-lock 500 → bounded retry). ### Added - **`reconcile.RunBringUp`** (`bringup.go`) — `BringUpSpec` (Mode `provision`|`dr_guest_loss`, Archive, VMID, RestoreStorage, Hostname, Cores/MemoryMB, RootfsGrowGB, Mounts, KeepMAC, BootTimeout) → `BringUpResult` (VMID, AssignedMAC, Pass, Verified, StartWarnings/Recognized). Sequence (each mutation preceded by journaling the owning entry): restore → identity reset → size → attach mounts → start LINK-UP. **Verdict is liveness (`waitRunning`), never the start exitstatus** (reuses the v0.7.0 WARNINGS surface). **Success KEEPS the guest** (no teardown). - **Scenario-specific identity reset** (doc 03 §9): *provision* → fresh MAC unconditionally (`PUT net0` with `hwaddr` omitted → PVE regenerates, F1) + hostname; machine-id + SSH host keys regenerate guest-side on first boot (golden bake + the new unit) — the agent does NOT touch guest internals. *dr_guest_loss* → preserve continuity (keep hostname; keep MAC unless `KeepMAC=false`); never resets restic/tunnel/hub identity. - **Compensating rollback** — any mid-flight failure destroys the just-created guest (`ClassGuestDestroy`, benign via `Provenance{SameTxnCreated:true}`, gated); on teardown failure the entry is left in-flight for `Recover`. New journal flag **`Rollback`** + `Recover`'s `recoverBringUp` reap a half-built guest left by a mid-job crash (idempotent, via `ListLXC`). - **F4 config-lock retry** — steps 3+5 coalesced into ONE `PUT config` (net0+hostname+cores+ memory+mpN); rootfs grow stays its own call. `setConfigWithLockRetry` retries ONLY the transient PVE config-lock 500 (`pveConfigLock`: 500 + "can't lock file"/"got timeout"); any other error fails immediately — never retried. - **`--selftest=bring-up`** (`-mode provision|dr -archive -vmid -hostname [-keep]`) — runs the real journaled job (after a `Recover`), then tears the guest down unless `-keep`. - **`configs/build-golden.sh`** — the validated golden recipe as a script, incl. the F3 first-boot `felhom-regen-hostkeys.service` unit (Condition-gated: fires on provision, no-ops on DR). The slice-7 spike archive (which lacks the unit) is superseded. ### Deferred (stated, not built) - Provisioning BACK HALF (controller deploy, bootstrap, per-guest token mint) → **slice 8**. - Host-loss DR + PBS escrow consumption → **slice 10**. - The SOURCE of a `BringUpSpec` (hub desired-state: which archive/VMID/mounts) → **slice 10**; this job takes the spec as input. `GuestMount` is defined minimally (no hub coupling). ### Tests - provision happy path (fresh MAC = net0 without hwaddr, hostname, coalesced sizing+mount, rootfs grow separate, started, **guest NOT destroyed**); compensating rollback at each step (restore / config / start-task / waitRunning — asserts the guest WAS destroyed); DR continuity (MAC kept, hostname not reset) + DR `KeepMAC=false` resets MAC; liveness verdict (warnings+running pass / not-running fail); F4 (lock-500→retry→proceed; non-lock-500→fail without retry); owning entry journaled BEFORE restore; reserved/existing VMID refused; `Recover` rolls back / clean. ### Live-validated (demo-felhom) - provision: fresh MAC + hostname; **SSH host keys regenerated by the baked golden unit** (agent issued no `ssh-keygen`), machine-id unique, Docker runs, clean DHCP lease → torn down. - dr: continuity preserved (hostname + host keys kept). Recover: a killed mid-restore left an orphan; the re-run's `Recover` rolled it back (idempotent). - **Live caught a bug, then fixed:** the host-key unit's `ExecStart` was `/usr/sbin/ssh-keygen` (203/EXEC); on Debian 13 it is `/usr/bin/ssh-keygen` — corrected in `build-golden.sh`, golden rebuilt, re-validated. (Mocked unit tests couldn't surface this; the live run did.) ## v0.7.0 — restore-test: verdict is liveness, not start-task exitstatus (2026-06-09) Fixes a correctness bug found by the live hub-enrollment runbook: the self-restore-test reported `pass:false` on **every** modern-distro guest. PVE's guest-start task exits `"WARNINGS: 1"` for the benign systemd-nesting advisory (`WARN: Systemd 257 detected. You may need to enable nesting.`), and `WaitTask` treated any non-`"OK"` exitstatus as a hard failure — so the verdict was decided by an advisory exit code instead of by observed liveness, *before* the real boot check ran. A crying-wolf test got it disabled on the demo host; this re-enables it. **Single bump (0.6.0→0.7.0) covering the agent's part of both task phases**; the wire fields below are consumed by hub from **v0.7.5**. Design invariant (in code): **warning classification affects *visibility only*; pass/fail is liveness-only.** A wrong/stale recognizer can at worst over-notice a benign warning — it can never false-fail and never hide a real warning. ### Added - **`proxmox.WaitOptions.AllowWarnings`** — opt-in per call. When set, a task that completes `"WARNINGS: N"` is success with the `TaskStatus` (ExitStatus intact) returned so the caller can read/surface it. Default (`false`) keeps **every existing caller strict** (vzdump/restore/destroy warnings can be meaningful — relaxing them is a future per-call decision with evidence). Any non-WARNINGS non-OK exit is still a `*TaskError`. - **`reconcile.RestoreTestResult.StartWarnings` / `.WarningsRecognized`** + a version-free recognizer (`benignWarningAnchor = "enable nesting"`, case-insensitive substring — contains no systemd version number, so it can't rot back into the bug at systemd 258+). `extractWarningLines` pulls `WARN…` lines from the start-task log. - **`reconcile.GuestAPI.TaskLogTail`** — the engine fetches the start task's log to surface warnings. - **`hub.RestoreTest.warnings` / `.warnings_recognized`** wire fields (`omitempty`), populated by `ToHubRestoreTest`. Additive: the deployed v0.7.4 hub ignores them; hub v0.7.5 consumes them (passed-with-warnings INFO, or WARN when not recognized). Cross-repo golden updated with the hub side. ### Changed - **Restore-test start step** (`reconcile/restoretest.go`) now waits with `AllowWarnings:true`, surfaces any start warnings, and **continues to `waitRunning` as the verdict** — boot+running is the pass, exactly as before; a real (non-WARNINGS) start-task error still fails. The restore and scratch-teardown WaitTasks stay strict. - **Restore-test scheduler logging** distinguishes a clean pass, *passed-with-recognized-warnings* (INFO), and *passed-with-unrecognized-warnings* (WARN) — nothing silent. ### Tests - `WaitTask`: AllowWarnings accepts `WARNINGS` (status returned intact); AllowWarnings still fails a real error; default still fails on `WARNINGS` (existing callers unaffected). - Restore-test (engine, mock proxmox): start-with-warnings + running → **pass** with warnings surfaced+recognized; unrecognized warning + running → pass, not-recognized; **not-running → fail regardless of warnings** (verdict is liveness); teardown still runs. - **Regression guard:** the `"enable nesting"` recognizer matches the advisory for systemd 256–300, proving it's version-independent and can't silently rot back into the false-fail. ## v0.6.0 — slice 6 Phase B: PBS offsite tier (verify + PBS-API client + reporting) (2026-06-09) Completes slice 6. The PBS spike (felhom.eu phase5-pbs-spike-findings.md) proved backup-to-PBS and restore-from-PBS reuse Phase A UNCHANGED (PBS is just a storage target + a volid), and the operator token needs no widening. So the only new agent code is the **verify capability + a small PBS-API client + PBSSnapshot reporting**. Escrow + host-loss DR stay slices 7/10. ### Added - **`internal/pbs` — the PBS-API client** (the agent's SECOND privileged external surface, slice-1 discipline): TLS **fingerprint-pinned** to the PBS leaf cert (a spoofed PBS → rejected, mirroring the PVE pin), **token auth** (`PBSAPIToken=:`; id from the storage `username`, secret read at runtime from `/etc/pve/priv/storage/.pw` — referenced by location, never logged/committed), typed, no shell. Methods: `Verify` (POST `/admin/datastore//verify` → UPID), `Snapshots` (incl. the `verification` field), `TaskStatus`/`WaitVerify` (node extracted from the UPID — `localhost` returns "unknown", the spike B4 gotcha), `NodeFromUPID`. - **The verify maintenance loop** (`pbs/verify.go`) — the cheap, key-free, ciphertext-level integrity check (§8) on its OWN cadence (default 6h, the 5th daemon goroutine). It is a reporting/maintenance task like the slice-5 watchdog: it does NOT go through the reconcile gate/journal. Each cycle: trigger verify → poll task → re-list snapshots → record per-snapshot `verify_state`. A failed verify is logged loudly. - **`PBSSnapshot` reporting** — filled the stub (`namespace`/`backup_type`/`backup_id`/ `backup_time`(RFC3339)/`size_bytes`/`owner`/`protected`/`encrypted` (from `files[].crypt-mode`) /`verify_state` (ok|failed|**none** until verified)/`verify_upid`). New `PBSReporter` collector seam + an in-memory `SnapshotStore`. Cross-repo golden (both repos, byte-identical) + bidirectional key-set tests; hub `handler.go` parses `pbs_snapshots` and logs a **failed verify `[WARN]`** (loudest offsite-DR signal). - **Truthful backup mode** (`backup/runner.go`) — `Backup.mode` now reflects the ACTUAL vzdump mode read from the task log (`backup mode: `), since PVE may downgrade snapshot→stop for a stopped guest (spike B1); falls back to the requested mode if unparseable. - **proxmox**: `Storage.Username` (parsed from the pbs storage config — the token id). - **config** `BackupConfig.{PBSVerifyCadenceSeconds, PBSSecretDir}` (cadence 0→6h, <0 disabled). - **`--selftest=pbs-verify`** — discover pbs storages → verify each → print the PBSSnapshot records (covers the runbook's verify + list). Standalone on the host. ### Notes - Backup/restore-to-PBS reuse Phase A with no change (the restore-test runs with `source_tier="pbs"` when fed a pbs volid). Zero-knowledge holds: verify is ciphertext-level, the encryption key is never read here, and the PBS server has no client key (spike B6). - Daemon runs cleanly with no pbs storage / verify disabled. `go test -race` covers the new goroutine. Slice-3/4/5/6A surfaces, goldens, and adversarial tests intact. ## v0.6.0-rc1 — slice 6 Phase A: backup + the self-restore-test (local target) (2026-06-09) Phase A of the backup/restore slice (doc 03 §8) — the agent's guest-level backup layer and the **self-restore-test**, which closes "a backup you haven't restored isn't a backup". Everything here is BENIGN (backup, restore-to-NEW, scratch teardown): reuses the slice-4 classifier/gate/journal — no new destructive class, no new crypto. Local target only; PBS is Phase B. Restore is to a NEW guest only (no overwrite). Backups are crash-consistent only (app-consistency needs the controller quiesce, slice 8) — marked so in the report. ### Added - **proxmox** (`mutate.go`/`query.go`): `DestroyLXC` (DELETE …/lxc/{vmid}?purge=1&destroy- unreferenced-disks=1 → UPID; the scratch-teardown primitive); `VzdumpOptions.Notes` → `notes-template` (verified on PVE 9.2.2); `LatestBackupVolID` (resolve a produced archive from the backup-storage listing — the task status carries no result volid). - **reconcile self-restore-test** (`restoretest.go`) — `Engine.RunRestoreTest`: pick a free scratch VMID (configured band, excludes 9999; full band → skip, never out-of-band) → **journal a Scratch-owned entry BEFORE any mutation** → restore-to-new → benign net **link-down** SetConfig (so the clone can't conflict with a running source's MAC/IP; this is test-safety, NOT slice-7 identity reset) → boot → verify **reaches `running`** → ALWAYS teardown (defer; benign `ClassGuestDestroy` + agent-tagged-scratch provenance, gated). Runs on the scratch VMID's queue lane. Reuses the journal/gate; result feeds the report. - **Crash-safe recovery** (`recover.go`): a Scratch journal entry is resolved by TEARDOWN, not by re-checking the restore sub-task's UPID — special-cased BEFORE the generic path (else the restore task's OK would mark it succeeded while the guest leaks). `Recover` now destroys a leaked scratch guest (idempotent: already-gone → clean; list-unreadable → left in-flight for a later pass). `JournalEntry.Scratch` flag; `RecoverResult.ScratchClean/ ScratchDestroyed`. GuestAPI gains `RestoreLXC`/`DestroyLXC`/`GuestStatus`. - **`internal/backup` package**: `BackupRunner.Backup` (vzdump + archive/size resolve + bulk-volume gap — a mountpoint is UNCOVERED unless it carries an explicit `backup=1`, so an unset `backup=` is reported uncovered too, the safe DR direction); `PickRestoreCandidate` (newest backup); an in-memory `Store` (latest-backup-per-target + latest-restore-test) implementing the hub `BackupReporter`/`RestoreTestReporter` seams; a cadence `Scheduler` (default 24h; the fourth daemon goroutine; disabled cleanly when off/misconfigured). - **hub report** (`report.go`): filled the `Backup` + `RestoreTest` stubs (`PBSSnapshot` stays a Phase-B stub); collector `BackupReporter`/`RestoreTestReporter` seams. Cross-repo golden updated in BOTH repos (byte-identical) + bidirectional key-set tests for `backups[0]`/`restore_tests[0]`. Hub `handler.go` parses + persists them (report_json; no new columns) and logs a **FAILED restore-test prominently** (the loudest DR signal). - **config** `BackupConfig` (local target, restore storage, restore-test cadence, scratch VMID band 990000–990009 default) + accessors + env overlay + cadence-gated validation. - **`--selftest=backup -vmid N`** (one-shot backup → print the Backup record) and **`--selftest=restore-test [-archive volid]`** (Recover-then restore→boot→verify→teardown, print the RestoreTest record). Standalone on the Proxmox host. ### Notes - The daemon runs cleanly with the cadence off or misconfigured (logs + disables, never crashes); a leaked scratch guest from a mid-test crash is reaped by `engine.Recover` on restart. `go test -race` covers the new scheduler goroutine. - Slice-3/4/5 exported surfaces, goldens, and adversarial tests intact. Version bumps to **v0.6.0** when Phase B (PBS) lands. ## v0.5.1 — slice 5 live-validation prep: durable_id mis-id fix + re-mount UUID memory (2026-06-09) Two correctness fixes surfaced while preparing the live USB validation on `demo-felhom` (a real 1TB USB HDD, sdb1, ext4). Both are DR-load-bearing — exactly the "false-id → re-attach the wrong disk" failure mode the slice warned about. ### Fixed - **Unmounted dir-storage no longer inherits the ROOT filesystem's UUID** (`observe.go`). Previously, when a removable dir-storage was unmounted, the observer fell through to the *containing* mount (root) for the backing device, so its `durable_id` became `uuid:` — a catastrophic DR mis-id (the hub would re-attach the wrong disk). Now the backing device/UUID/`durable_id` are derived ONLY from the target's OWN mountpoint; an unmounted target reports no device and a stable `store:` durable_id, never another filesystem's UUID. (Removed the `containingMountDevice` root-fallthrough.) - **Watchdog remembers the fs-UUID observed while attached** (`watchdog.go`) so a re-mount works even after the known-set cache refreshes mid-drop (an unmounted target can't resolve its own UUID). The re-mount key is backfilled from this memory — aligning with doc 03 §7's "sourced from the existing definition, no hub manifest needed": the agent learns the UUID while the target is attached, then re-mounts by it on return. ### Tests - Observer: an unmounted dir-storage asserts NO `uuid:` durable_id and no backing device. - Watchdog: a drop where the cache lost the UUID still re-mounts using the remembered UUID. ## v0.5.0 — slice 5 Phase B: the host-root surface (mounts + SMART + grow + destructive gate) (2026-06-09) The write surface — the agent's first step outside its Proxmox API token into OS-root. Isolated behind a narrow, argument-validated, adversarially-tested seam, exactly like the slice-4 gate. Completes slice 5 (Phase A = read-only observe/report/watchdog at v0.5.0-rc1). ### Added - **`HostOps` seam + `SudoHostOps`** (`internal/storage/hostops.go`) — the one privileged host surface: persistent mounts via **systemd `.mount` units keyed by fs-UUID** (enabled to survive reboot), detach (stop+disable), SMART, and thin-pool metadata. Shells out via the fenced Runner (`sudo -n`, **fixed arg vectors, no shell**); a fake backs the tests (no real root in the suite). `NoopHostOps` is the safe fallback when the surface is unavailable. - **The argument validator** (`internal/storage/validate.go`) — the security boundary: `ValidateUUID` (strict hex), `ValidateMountPath` (absolute, no traversal, no metacharacters), `ValidateSMARTDevice` (raw-disk whitelist), `ValidateLVMName`, and an in-process `systemdEscapePath` (no `systemd-escape` shell-out). **Every argument is validated BEFORE a command is constructed.** Headline test (`validate_test.go`): an adversarial matrix of shell metacharacters / `../` traversal / malformed inputs is rejected with **zero exec**. - **SMART** (`internal/storage/smart.go`) — parses `smartctl -a -j` into `StorageTarget.smart`: **SATA** (reallocated/pending/offline-uncorrectable, temp, power-on-hours) **and NVMe** (critical_warning, media_errors, percentage_used, temp), degrading to `UNKNOWN` for devices with no SMART (USB-SATA bridges). **`lvs`** fills the lvmthin thin-pool **metadata** fill (the value Phase A left null). Wired into the Observer's enrichment (Observe only, not the watchdog's fast Known path). - **Watchdog re-mount response** (`internal/storage/watchdog.go`) — on a known mount-backed target's device returning **unmounted** (a new `DevicePresent` liveness probe), the watchdog **dispatches a benign by-UUID re-mount off the poll path** (a goroutine, never under the lock), rate-limited per target to the debounce window. The mount is routed through the gate as benign (`gateRemounter` in `main.go`, so `storage` stays decoupled from `reconcile`). - **Disk-grow executor** (`internal/reconcile`) — `ActionResize` (benign `ClassResize`), planned **grow-only** (desired DiskBytes > actual → `pct resize rootfs +M`; a shrink is refused, never silently grown) + a defensive executor guard (size must start with `+`). New `proxmox.Client.ResizeLXC` (API; `VM.Config.Disk`+`Datastore.AllocateSpace`; async→UPID). Built + fixture-tested; **unfed** live (no hub spec until slice 10). - **Destructive storage ops through the slice-4 gate** (`internal/reconcile/storage_ops.go`) — `IntentForStorageMount` (benign) and `IntentForStorageDestructive` (`ClassStorageWipe`/ `ClassDecommission`). Host/target-scoped: the op binds on the storage **target identity** (carried in `target.guest_id`). Reuses the existing verifier/role-scoping/binding/audit — no new gate, no new crypto. Storage cases added to the adversarial matrix (`storage_test.go`): unsigned wipe → `pending_signature`; "wipe A" signature vs "wipe B" → `binding_mismatch`; valid → accepted. **Inert** live. - **`--selftest=storage` [`-watch `]** — the live USB-runbook harness: an observe pass (full table incl. SMART + thin-pool data+metadata), and a bounded watchdog window with the re-mount response live. Runs standalone on the Proxmox host (no hub). - **`configs/felhom-agent.sudoers`** — the documented narrow allowlist (install unit / systemctl manage / smartctl / lvs), with the agent-side fine validation noted. - **Config**: `privileged.{unit_dir,stage_dir,systemctl,install,smartctl,lvs}` (paths must match the sudoers entries). ### Notes - Daemon still runs cleanly with no removable storage / no signers / no hub manifest, and a missing/declined sudoers entry degrades with a warning (SMART→UNKNOWN, mount→logged error), not a crash. `go test -race` passes (the watchdog re-mount dispatches off the poll path). - Slice-3/4 + Phase-A exported surfaces, goldens, and adversarial tests intact. `authz` untouched. The destructive-storage executor + grow are built/tested but unfed live until slice 10. ## v0.5.0-rc1 — slice 5 Phase A: storage observe + report + watchdog (read-only, live) (2026-06-09) Phase A of the storage slice (doc 03 §7). Read-only and live: the agent now observes every host storage target, reports it into the host-report's `storage_targets` (previously an empty stub), and runs a fast-poll watchdog that pushes a disconnect to the hub in seconds. No host-root writes this phase (mounts/SMART/grow/destructive-gate are Phase B). The hub-owned desired manifest (class/role/policy/creds) is not served until slice 10, so reconcile against it is built-but-unfed — this phase ships only the genuinely-useful read-only footprint. ### Added - **`internal/storage` package** (new): - **`StorageTarget` wire contract** (`internal/hub/report.go`) — filled the slice-3 stub: `name`/`type`/`durable_id`/`state`/`reachable`, usage (`total`/`used`/`avail`/ `used_fraction`), `content`, `mount_path`/`backing_device`, `class_hint` (rotational HINT — never authoritative; class is hub-owned), `role` (empty until slice 10), a `thin_pool` sub-object (lvmthin data fill; metadata fill is Phase B/`lvs`), and a `smart` sub-object (`UNKNOWN` until Phase B). Cross-repo golden kept byte-identical with `felhom.eu/hub` and guarded by the bidirectional key-set test (`contract_test.go`). - **`durable_id` derivation** (`durableid.go`) — deterministic per type (the DR-load-bearing re-attach key): fs-UUID (usb/local-dir), `server:export` (nfs/cifs), `repo+fingerprint` (pbs), `vg/pool` (lvmthin); never empty (falls back to a stable store id). - **`HostReader` seam + `ProcHostReader`** (`hostread.go`) — non-privileged `/proc/mounts`, `/dev/disk/by-uuid`, `/sys/.../rotational` + `removable` reads. Root-free by construction. - **`Observer`** (`observe.go`) — builds `[]hub.StorageTarget` from `ListStorage`/`NodeStorage` joined with host reads; surfaces the lvmthin thin-pool data fill prominently (warns ≥85%). - **Storage watchdog** (`watchdog.go`) — a third daemon goroutine fast-polling the *known* target set (a defined Proxmox storage and/or a previously-seen one) for `attached↔disconnected` transitions; on a transition it triggers an immediate, **debounced** out-of-band host-report. Only flags a *known* target's change (never a never-attached device); coalesces flaps within the debounce window (leading + trailing edge). `CachingKnownTargets` rate-limits the Proxmox-derived known set; `HostLiveness` probes device/mount presence (local) + a reachability dial (network), all non-privileged. - **Proxmox `Storage` type** (`internal/proxmox/types.go`) — additive parse-only config fields (`server`/`export`/`share`/`datastore`/`fingerprint`/`vgname`/`thinpool`) feeding durable_id. - **Collector `StorageObserver` seam** (`internal/hub/collect.go`) — populates `storage_targets` via the observer; a nil observer or an observe error degrades to empty (never sinks the heartbeat). Hub does not import storage (storage imports hub for the wire type). - **Out-of-band report trigger** (`internal/hub/loop.go`) — `Loop.SetTrigger`: a watchdog signal runs one extra collect→report immediately without disturbing the regular cadence. - **`StorageConfig`** (`internal/config`) — watchdog interval / debounce / known-refresh knobs (all optional; package defaults otherwise). - **Hub ingest** (`felhom.eu/hub`) — `hostReportPayload` now parses `storage_targets` (full mirror struct), persists them via `report_json`, counts + warns on disconnected targets, and has its own half of the bidirectional golden key-set test. ### Notes - The daemon still runs cleanly with no removable storage, no signers, and no hub manifest — the watchdog finds nothing to flag; storage reporting is best-effort. - `proxmox`/`hub`/`authz`/`reconcile` exported surfaces + their golden/adversarial tests are intact. No host-root writes, no destructive paths, no SMART this phase (all Phase B). - Version: **v0.5.0-rc1** at the Phase-A checkpoint; **v0.5.0** when Phase B lands. ## v0.4.0 — slice 4 Phase B: reversibility gate + signed-op consuming layer (2026-06-08) The security core of slice 4: hub-supplied intent stops being trusted for destructive change. Layered in front of the per-guest queue's executor — **every** mutation now passes the gate. Reuses `internal/authz` for all crypto (untouched surface). Inert this slice: no destructive deltas are served until slice 10, so the destructive path is classified, gated, and adversarially tested but not wired to live execution. ### Added - **Classifier (`classify.go`, doc 03 §4)** — benign vs destructive by **provenance + data-bearing-ness, NOT by verb**. The `OpClass` vocabulary (seeded by the committed slice-2 `op_blob.json`: `guest_destroy`) is the agent-side contract slice 10 matches. Destroy/overwrite of customer data is destructive UNLESS **agent-internal** provenance (same-journaled-transaction create → compensating rollback, or agent-tagged scratch) makes it benign. `Provenance` is journal-recorded and **never populated from the hub** (its zero value is the only thing an external intent may carry). Unknown op class fails safe → destructive. - **Reversibility gate (`gate.go`)** — `Gate.Authorize(intent, signed)`: benign → allowed unsigned; destructive → requires a verified, role-authorized, action-bound operator signature, else refused **`pending_signature`**, never executed. Every decision is written to an `AuditSink` (audit is a signal, never the guard). - **Signed-op consuming layer over `authz`** — verifies via `authz.Verifier.Verify` (the locked pipeline, untouched), then enforces on the `VerifiedOp`: - **Role-scoping (doc 04 §4)** — recovery key authorizes key-rotation re-pins ONLY; operational key authorizes ordinary destructive ops + planned rotation. - **Op-to-action binding** — verified `op` + host + guest + `params` must match the gated action (a signature for guest X / op A can't authorize guest Y / op B); params compared semantically (key-order/whitespace independent). - **Signed-job orchestration (`job.go`)** — `RunSignedJob`: idempotency dedupe (the op nonce as the journal key — a redelivered completed op is skipped, not re-run), gate authorization, then journal-wrapped execution via an injected `DestructiveExecutor` (nil this slice — authorized destructive ops are inert, no executor wired until 6/7). - **Crash-recovery consumer (`recover.go`, Note 1 / doc 03 §10)** — `Engine.Recover` consumes the journal's `InFlight()` at startup: an op that crashed AFTER the Proxmox POST and BEFORE its terminal record (`OpTaskRunning`, nonce already consumed) is NOT covered by idempotency dedupe — only this resume-or-rollback resolves it (re-read the task via the new `TaskStatusOnce`, record the real outcome; a no-task-id op is abandoned fail-safe). Landed together with the signed-op executor, as Note 1 required. - **Daemon wiring** — `runDaemon` builds the verifier from `config.Authz.Signers` (a bad key / missing nonce-store path is a fatal misconfig; **no signers = nil verifier**, the common slice-4 state), constructs the gate (+ `SlogAudit`), runs `Recover` before issuing any mutation, and routes every reconcile action through the gate. ### Changed - **Memory comparison canonicalized (Note 2)** — `desiredMemoryMiB` makes the desired↔actual memory compare in the same MiB unit that is then written, so a non-MiB-aligned `MemoryBytes` converges in one pass instead of re-issuing SetConfig forever (the numeric cousin of the description-newline normalization). Test proves convergence. Slice 10 should still serve MiB-aligned specs at the source. ### Tests (the security proof — each independently rejected) - **Adversarial matrix** via the REAL `authz.Verifier` with in-test-minted SSHSIGs (framing replicated in reconcile's test binary; production authz untouched, no signing added to the verify-only package): unsigned destructive **job** → pending_signature; unsigned destructive **desired-state delta** → pending_signature (distrusts hub desired state, not just jobs); forged/unknown signer → `ErrUnknownSigner`; expired → `ErrExpired`; **replayed nonce across an agent restart** (durable `FileNonceStore`) → `ErrReplay`; wrong host → `ErrTarget`; wrong guest / wrong op / wrong params → binding_mismatch; **recovery key on ordinary destructive** → role_denied; **hub-supplied "scratch" tag ignored** → still destructive → refused; **valid + role + target + fresh nonce → accepted**, and a second presentation → `ErrReplay` (nonce consumed). - Classifier (benign/destructive/provenance/key-rotation/fail-safe), role-scoping, params binding, crash-recovery (resume OK / fail / still-running / no-task rollback / unreadable / one-shot key applied on resume), signed-job idempotency (execute once, dedupe redelivery, refused-not-executed, no-executor-inert, executor-error). - Full module **race-clean** (`go test -race`) + vet clean on the Linux build server. ## v0.4.0-rc1 — slice 4 Phase A: reconcile engine (structural; runs live, unfed) (2026-06-08) The agent-side control core's structural half. **Checkpoint marker** — `-rc1` is the Phase-A push; awaiting validation before Phase B (the reversibility gate + signed-op consuming layer) lands the final **v0.4.0**. Runs LIVE but UNFED: with no desired-state provider until slice 10, the live engine computes an empty action set and performs **zero mutations**. ### Added - **`internal/reconcile`** package — the engine, the per-guest serializer, the desired-state model, the normalization layer, and the durable op journal: - **Per-guest serializer (`Queue`, doc 03 §10)** — the single choke point ALL mutation sources funnel through. Same-vmid jobs run strictly one-at-a-time in submit order; independent vmids run in parallel. Each vmid is a cond-var FIFO lane (unbounded, non-blocking, order-preserving); graceful drain on `Close`. - **Desired-state model + `DesiredProvider` seam** — `DesiredGuest` (per-field optional: run-state / `*hub.GuestSpec` / `*description`), `DesiredState`. The only live provider is **`EmptyProvider`** (slice 4 has no source); `StaticProvider` feeds fixtures. The seam is where slice 10's hub-serving plugs in — no hub/local source invented here. - **Normalization layer (`FieldNormalizers`)** — reconcile compares *normalized* desired-vs-actual so Proxmox round-trip quirks don't read as drift. `description`'s trailing newline is the first registered case; the registry takes more (boolean coercion, list ordering) as discovered. `normDesc` **promoted** out of `cmd/felhom-agent/main.go` to **`reconcile.NormDescription`**; the `--selftest=task` description round-trip now uses that shared helper (one source of truth for the quirk). - **Plan engine (`Plan`, pure function)** — computes the minimal **benign** action set (`Start`/`Stop`/`SetConfig`) for guests present in both desired and actual, with normalized comparison, deterministic vmid ordering, config-before-run-state. Skips provision (desired-absent-in-actual, slice 7) and destroy (actual-absent-in-desired, gated, slice 10); never writes a config it couldn't first read (`SpecKnown`). Disk (rootfs grow) intentionally not reconciled here. - **Reconcile engine (`Engine`)** — reads desired+actual, plans, dispatches each action onto the shared queue. Every Proxmox op handled per the mutate.go contract: non-empty UPID → `WaitTask` + assert `exitstatus`; empty UPID → clean **synchronous** success (slice-4 proven). Per-action failures are counted, not fatal (other guests still converge). - **Operation journal (`Journal`)** — durable fsync'd append-only JSONL mirroring `authz.FileNonceStore`: records each op's lifecycle (started → task_running → succeeded/failed) with its Proxmox task id (crash mid-op is detected and re-checkable on restart via `InFlight()`), plus an **idempotency-key store** (`AlreadyApplied`) so a one-shot op never re-runs across retries/restarts. Reconcile actions carry no idempotency key (convergent — must re-run on real drift). - **Daemon wiring (`runDaemon`)** — reconcile runs alongside the hub loop on the poll cadence, **sharing the per-guest queue**. Journal path is a `journal.log` sibling of the nonce store. The daemon runs cleanly with **no desired state and no signers** (reconcile is a logged live no-op; a journal-open failure degrades to journal-less, never crashes). ### Tests - Serializer: same-guest serialized (max-concurrency 1, submit order preserved) and different-guests parallel (cross-waiting jobs both complete — would deadlock if not); error propagation; drain-pending-on-close; submit-after-close. - Normalization: description round-trip; unknown-field identity; extensibility seam (synthetic boolean-coercion + list-ordering normalizers). - Plan: run-state start/stop, spec drift (cores/memory), disk-not-reconciled, description-newline-not-drift, unmanaged fields, spec-unknown skips config keeps run-state, desired-absent skipped, combined ordering, empty-desired no-op, deterministic vmid order. - Engine: empty-provider zero mutations; async start (WaitTask); synchronous SetConfig (no WaitTask); WaitTask failure + POST error counted failed; list error = pass failure. - Journal: lifecycle latest-wins; in-flight survives restart; idempotency dedupe across restart; failed key not applied; torn-trailing-line skipped. - Full module **race-clean** (`go test -race`) on the Linux build server; vet clean. ### Not in this phase (Phase B) - The benign/destructive classifier, the reversibility gate, and the signed-op consuming layer over `internal/authz` (doc 03 §4 / doc 04) — added next, in front of the queue's executor, landing **v0.4.0**. ## v0.3.2 — SetConfig selftest extension (slice-4 pre-check) (2026-06-08) The gate before slice 4: prove `SetConfig` works live under the scoped token before reconcile is built on it. **Self-gated live run PASSED** on `demo-felhom`/guest 9999. ### Added - **Reversible `SetConfig` step appended to `--selftest=task`** (`cmd/felhom-agent/main.go`, `selftestSetConfig`): read `GuestConfig` → write a `description` marker (`felhom-selftest `) → verify it landed → restore the original value (or `delete` the key if it was absent) → verify the restore. Handles PVE's dual-mode `SetConfig` return per the `mutate.go` contract: empty UPID = synchronous success (printed `synchronous`); non-empty UPID = `WaitTask` + assert `exitstatus=OK`. The existing snapshot → rollback → delete-snapshot steps are unchanged. First live exercise of the **`VM.Config.*`** privilege cluster. - **`normDesc` / `extraString` helpers** — `extraString` decodes a string-valued key from `GuestConfig.Extra` (raw JSON); `normDesc` strips the trailing newline PVE appends to `description` on read, so a written value round-trips equal. ### Finding (live) - The LXC `description` write returned **synchronous (empty UPID)** — PVE applied it inline, no task. The agent's dual-mode `SetConfig` modeling is correct: the empty-string path is real and must not be treated as an error. - PVE **appends a trailing `\n` to `description`** on read (stored URL-encoded as `%0A`). A naive exact-match reconcile would see perpetual drift — slice-4 reconcile must normalize `description` comparisons (hence `normDesc`). ### Ops - Standing operator token (`felhom-agent@pve!agent`, privsep) **rotated** during this run (the prior secret was not retrievable); role + both user/token ACL rows re-confirmed at `/`. New secret stored out-of-band, **not persisted to the repo**. Guest 9999 left pristine (stopped, no `description`, no leftover snapshot). Version → 0.3.2. ## Docs + live validation — no version bump (2026-06-08) ### Changed - **Reflowed `CLAUDE.md`** — removed hard mid-paragraph line wraps (prose, list items, blockquotes now single-line, soft-wrapped); code blocks and tables untouched; rendered output unchanged. - **Unified the REPORT/CHANGELOG convention** in `CLAUDE.md`: `CHANGELOG.md` is the cumulative log (newest on top); `REPORT.md` is overwritten with the most-recent implementation/validation only. Added an explicit **no-secrets** rule (never write tokens/passwords/keys into committed files; reference them as stored out-of-band). ### Added - **`REPORT.md`** rewritten for the live `--selftest=task` validation on the demo host (`demo-felhom`): snapshot → rollback → delete-snapshot on guest 9999, each polled to `exitstatus=OK` under the `felhom-agent@pve!agent` privsep token (UPIDs name the token actor — privsep path genuinely exercised); 16-privilege `FelhomAgent` role + both user & token ACLs confirmed; `--selftest=read` clean. Closes the slice-1 "mutating ops unit-tested only" gap; `WaitTask` async foundation validated live → **slice 4 unblocked**. (Token secret stored out-of-band, not in the repo.) ## v0.3.1 — slice-3 validation follow-ups (2026-06-08) ### Changed - **Collector keeps the known run-status on a `GuestConfig` failure** (`internal/hub/collect.go`): previously a per-guest config-read error forced `status="unknown"`; now the run-status from `ListLXC` is preserved (only the `spec` is dropped). An empty status is still normalized to `unknown` (wire value is always `running|stopped|unknown`). Test renamed to `TestCollect_GuestConfigFailureKeepsStatusOmitsSpec` and asserts the preserved `running` + nil spec. - **`--selftest` usage** error string now reads `(want read|task|hub)`. ### Added - **Cross-repo contract fixture** `internal/hub/testdata/host-report.golden.json` + `TestHostReport_ContractMatchesGolden` — compares the marshaled `HostReport` field-name sets (top level + `host` + `guests[0]`) against the golden, failing on any json-tag drift. The file is **kept byte-identical** with felhom-hub's copy (duplicated contract until a shared types module; revisit when slices 5/6 populate the empty collections). Version → 0.3.1. ## v0.3.0 — hub client + host-report + first daemon loop (slice 3) (2026-06-08) The agent's first daemon: a periodic read-only host-report POSTed to the hub (the heartbeat). No Proxmox mutations, no desired-state/signed-op consumption, no storage/backup collection yet — those are slices 4/5/6. ### Added - **`internal/hub`** package: - **`HostReport`** wire contract (`report.go`) shared field-for-field with the hub ingest: host metrics, guests (`vmid` + spec), `cloudflared` status, and the `storage_targets`/`backups`/`restore_tests`/`pbs_snapshots`/`audit_tail` collections **defined but emitted empty** (typed `[]`, slices 5/6 fill them). - **`Collector`** (`collect.go`) builds the report from a read-only `proxmoxReader` (adapted to the real `internal/proxmox` surface — node held by the client, value returns, `proxmox.Guest`) + a `CloudflaredProber`. Partial-failure policy: a failed `NodeStatus` is a hard error (skip the POST); a failed per-guest `GuestConfig` degrades that guest to `status="unknown"` (spec omitted) but still sends; a cloudflared probe failure → `"unknown"`, never fatal. - **`CloudflaredProber`** + `SystemctlProber` (`systemctl is-active cloudflared`; read-only — NOT a Privileged/root op; tunnel management is a later slice). - **`Client`** (`client.go`): `POST /api/v1/host-report` with `Authorization: Bearer `, standard TLS (system roots or optional `ca_file`; verification always on). Typed `*TransportError` / `*HTTPError`; the bearer token never appears in any error. - **`Loop`** (`loop.go`): the daemon — immediate first report then tick; adopts the hub's `poll_interval_seconds` clamped to [60,3600]; resilient (a collect/report error is logged and the loop continues); clean shutdown on context cancel. - **`ControlEnvelope`**: only `poll_interval_seconds` is acted on; `blocked` / `desired_generation` / `has_signed_ops` are parsed-but-ignored (logged at most) pending reconcile (slice 4). - **Config**: `HubConfig` (url/host_id/api_key/poll_seconds/timeout_seconds/ca_file), `FELHOM_AGENT_HUB_*` env overlay, `HubConfig.Validate()` (mode-aware — proxmox-only `--selftest=read|task` still runs without hub config), `WithDefaults()`, and `Redacted()` now also blanks the hub key. `configs/agent.example.json` gains `hub` (and `authz`) blocks. - **`cmd/felhom-agent`**: the no-`--selftest` mode is now the **daemon** (poll loop); added **`--selftest=hub`** (one collect+report, prints the report + envelope). Version 0.2.0 → 0.3.0. ### Tests - Report serialization (field names; empty collections are `[]` not `null`; spec omitted when unknown); client (Bearer header, non-2xx→`*HTTPError`, transport→`*TransportError`, **token never in error**); collector (host mapping, guest spec, per-guest failure degrades-but-still-reports, NodeStatus hard error, cloudflared error→unknown); loop (immediate first report, continuation after an injected error, interval adoption + clamp); config (hub validate/redact/env). ### Notes - `internal/proxmox` and `internal/authz` were **not touched** — no new proxmox surface was needed (`ListLXC` already exposes status/maxmem/maxdisk; `GuestConfig` exposes cores). The task's `proxmoxReader` sketch (node-arg/pointer/`LXC`) was adapted to the real exports as instructed. - **Defined-but-empty** this slice: `storage_targets`, `backups`, `restore_tests`, `pbs_snapshots`, `audit_tail` (slices 5/6). **Parsed-but-ignored**: the envelope's `blocked`/`desired_generation`/`has_signed_ops` (slice 4). ## v0.2.0 — `authz` signed-op verifier (slice 2) (2026-06-08) Production form of the Phase-4 signing primitive: a key-type-agnostic SSHSIG verifier for operator-signed destructive ops, with the full anti-replay/ authorization pipeline and a durable, crash-safe nonce store. What slice 4 (reconcile) will call to gate destructive desired-state deltas. No hub, no signing CLI, no reconcile loop. ### Added - **`internal/authz` — `Verifier`**: `New(signers, store, hostID)` + `Verify(blob, sigArmored) (*VerifiedOp, error)`. Runs the LOCKED pipeline (order is load-bearing): parse armor → namespace → parse pubkey → allow-list (by key **material**, `pub.Marshal()` equality, not key_id) → crypto verify (over the **raw received bytes**, never re-canonicalized) → parse blob → target → time window → **nonce recorded LAST**. Each post-crypto stage rejects even with a valid signature. - **SSHSIG framing** (`sshsig.go`) via `golang.org/x/crypto/ssh` — `pem.Decode` → strip 6-byte magic → `ssh.Unmarshal` → `ssh.ParsePublicKey` → recompute signed data with the named hash → `pub.Verify` (dispatches on key algorithm). No hand-rolled crypto. Key-type-agnostic: ed25519 / **sk-ssh-ed25519 (FIDO2)** / rsa / ecdsa via the one path. - **Fixed namespace** `felhom-op-v1` (package constant, never caller-supplied). - **`OpBlob`** (corrected `host_id`/`guest_id` json tags) + **`VerifiedOp`** (op, host/guest, params, key_id, matched signer). key_id is advisory/audit only — never an authz input. - **Typed errors**: `ErrMalformed, ErrNamespace, ErrUnknownSigner, ErrBadSignature, ErrTarget, ErrExpired, ErrNotYetValid, ErrReplay` (errors.Is-friendly). - **`NonceStore`** + two impls: `MemoryNonceStore` (tests) and **`FileNonceStore`** — durable, crash-safe (fsync'd append log, replayed into an index on open, periodic compaction, expiry-only pruning). A nonce is fsync'd to disk before `SeenOrRecord` returns false; replay protection survives restart; I/O failure fails safe (reports seen=true). Target generalization: host_id matched strictly, guest_id surfaced for the caller to route. - **Config**: `AuthzConfig` (nonce-store path + pinned operator `signers` tagged `operational`/`recovery` with a key_id, as authorized_keys lines). - **Version 0.2.0.** ### Tests - Real OpenSSH interop via a committed `ssh-keygen -Y sign` vector (hermetic CI); per-stage rejection (each with an otherwise-valid sig); the headline **invalid-sig-does-not-burn-the-nonce** invariant; replay; **persistence across restart**; synthetic **sk-ssh-ed25519** through the unchanged path; byte-exactness (a re-serialized blob fails crypto — not re-canonicalized). ### Notes / corrections to the Phase-4 reference - §7's `Target` lacked json tags (`host_id`/`guest_id`) — fixed. - The doc paired "Go 1.24.4 / x/crypto v0.52.0", but v0.52.0 declares `go 1.25.0` and does **not** build on Go 1.24. Resolved by upgrading the build server to go1.26.0 (backward-compatible; felhom-controller/hub unaffected); the module is `go 1.25.0` on x/crypto v0.52.0. - Free function → constructed `Verifier`; returns the full `VerifiedOp`; typed errors; clock-skew tolerance added; durable nonce store is the net-new work. - **Shared-contract dependency flagged** (not built): the hub and the `felhom-sign` CLI must emit byte-identical canonical JSON or signatures won't verify; a shared canonicalizer both import would be the right home. ## v0.1.0 — Scaffold + `proxmox` interaction layer (slice 1) (2026-06-08) First slice: stand up the host-agent project and its foundation — the typed Proxmox interaction layer every other module will call. No reconcile loop, hub client, signing, or storage/backup orchestration yet (later slices). ### Added - **Project scaffold**: module `gitea.dooplex.hu/admin/felhom-agent`, binary `felhom-agent` (`cmd/felhom-agent/`), Go 1.24, zero external dependencies (pure stdlib). `--version` flag; `version` var overridable via `-ldflags "-X main.version="`. - **`internal/proxmox` — API backend (`Client`)**: hand-rolled REST client over `https://:8006/api2/json` with `PVEAPIToken` auth. Typed read ops (`Version`, `Nodes`, `NodeStatus`, `ListLXC`, `GuestStatus`, `GuestConfig`, `ListStorage`, `NodeStorage`, `StorageContent`) and async mutating ops returning a UPID (`RestoreLXC` — the primary create path, `Vzdump`, `Snapshot`, `Rollback`, `DeleteSnapshot`, `SetConfig`, `Start`, `Stop`). - **`WaitTask`**: polls `GET /nodes/{node}/tasks/{upid}/status` until stopped, then asserts `exitstatus == "OK"` (authorization can surface at task execution, not the POST — phase1-2 §1.3). Exponential backoff (1s→5s cap), context cancellation + timeout. `*APIError` parses the offending privilege from a 403; `*TaskError` parses it from a failed task exitstatus + log tail. - **`internal/proxmox` — fenced root-CLI backend (`Privileged`)**: limited to the three proven OS-root exceptions only — `CreateGoldenLXC` (keyctl `pct create`), `MountUSBByUUID`, `SMART`, `Sensors`; each cites why it can't be the API. Fence is structural (Client never shells out, Privileged never makes an HTTP call) and asserted in tests. - **TLS trust**: SHA-256 leaf-cert pinning (the host serves a self-signed cert) or a CA file; an explicitly-named `insecure_skip_verify` that is off by default. No blanket verification disable. - **`internal/config`**: JSON config file + `FELHOM_AGENT_*` env overrides; the token secret is never logged (`Redacted()`). - **`internal/log`**: slog setup (text, stderr, configurable level). - **`cmd/felhom-agent --selftest`**: read-only health report against a live host (version/nodes/status/guests/storage); `--selftest=task --vmid N` exercises `WaitTask` on a reversible snapshot→rollback→delete op (gated; default selftest mutates nothing). - **Tests**: unit tests with a mock HTTP transport + mock runner (UPID parse, `WaitTask` running→OK / failed-403 / timeout / ctx-cancel, 403→privilege error, response decoding against shapes captured live from `demo-felhom`, config redaction, and the API-vs-root routing fence). ### Notes - Types are grounded in the spike findings (`felhom.eu/documentation/proxmox-platform.md`, `tests/phase{0,1-2,3}-findings.md`) and the exact JSON shapes captured live from `demo-felhom` (PVE 9.2.2). - Verified: `go build/vet/test` green on Go 1.24.4 (build server) and a live read-only `--selftest` against the demo host with TLS fingerprint pinning. - The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and token) is provisioned out-of-band; the agent only consumes the token.