bd4bced771
wgtunnel.InstallRecoveredKey: write an escrow-recovered WG private key (create-
only, refuse-overwrite) so the tunnel re-establishes with the same identity/pubkey
(same /32), no keygen. Wired into identity-consume -install-wg-key (opt-in;
pre-S3 blob → logged fresh-keygen fallback). Value never logged.
internal/dr (new): consume the host_loss restore_directive (was logged-ignored)
into an inspectable RestorePlan via the AddConsumer raw seam — per guest
{vmid,archive,target,sizing} + per drive {durable_id→mount} + offsite PBS coord.
DERIVE-AND-SURFACE only; the Consumer has no restore/destroy dependency (execute-
nothing is structural). guest_loss/absent → no plan.
Tests + red-proofs (WG create-only overwrite; plan mode-gate). No secrets on
argv/stdout/logs. The destructive in-place restore is a separate operator-present
STOP-gated drill.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
2524 lines
197 KiB
Markdown
2524 lines
197 KiB
Markdown
## v0.69.0 — S5: host-loss DR — recovered WG-key install + directive→restore-PLAN (safe halves) (2026-07-04)
|
||
|
||
The two safe, non-destructive mechanical links for host-loss DR (the destructive in-place restore is
|
||
a separate operator-present, STOP-gated drill).
|
||
|
||
- **`internal/wgtunnel.InstallRecoveredKey`** — writes an escrow-recovered WG private key (32-byte
|
||
base64, re-encoded canonical) to the key file so the tunnel re-establishes with the SAME
|
||
identity/pubkey (→ the same hub `/32`), no fresh keygen. **CREATE-ONLY** — refuses if a key file
|
||
exists (a present key may be a live identity); value never logged. Wired into `--selftest=identity-
|
||
consume -install-wg-key` (opt-in; after `UnwrapIdentityBundle`, installs `bundle.WGPrivateKey`;
|
||
pre-S3 blob with no WG key → logged fallback to fresh keygen + re-register, which keeps the /32).
|
||
- **`internal/dr`** (new) — consumes the host_loss `restore_directive` (was logged-and-ignored) into
|
||
an inspectable **RestorePlan** via the `desired.Syncer.AddConsumer` raw seam: per guest →
|
||
{vmid, archive, target storage, sizing}; per drive → {durable_id → expected mount}; + the offsite
|
||
PBS coord. **Derive-and-surface only** — the `Consumer` has NO restore/destroy dependency, so
|
||
"execute nothing" is structural. `guest_loss`/absent → no plan. Recipe fetched on-demand (rare
|
||
directive) via a fresh `Collect`.
|
||
- Tests + red-proofs: WG install (same pubkey/no-keygen; present-key refuse — red-proofed against
|
||
allow-overwrite); plan (host_loss builds; guest_loss/absent/nil-recipe → none — red-proofed against
|
||
a relaxed mode gate). No secrets on argv/stdout/logs (field names only).
|
||
|
||
## v0.68.0 — S4.1: tier-aware restore-task deadline (unattended offsite restore-test) (2026-07-04)
|
||
|
||
The offsite restore-test couldn't complete on the scheduler path because a WAN restore of a large
|
||
guest exceeds the restore-task wait's 10-minute default — the wait expired mid-restore, teardown
|
||
then fired against a still-restoring (not-yet-pool-associated) scratch guest, and it leaked. Make
|
||
the restore-task wait **tier-aware**.
|
||
|
||
- **`internal/reconcile`**: `RestoreTestSpec.RestoreTaskTimeout` (0 → the 10m `WaitOptions` default);
|
||
the restore-task `WaitTask` now passes it. Local-tier restores are UNCHANGED (10m — a local
|
||
restore hanging that long is a genuine fault).
|
||
- **`internal/config`**: `BackupConfig.RestoreTestPBSRestoreTimeoutSeconds` + accessor
|
||
`RestoreTestPBSRestoreTimeout()` (positive as-is, else **120m** — for an unattended nightly test a
|
||
false timeout is worse than a slow pass; very large guests may need more).
|
||
- **`cmd/felhom-agent`**: `restoreTaskTimeout(cfg, tier)` sets the field to the configured PBS
|
||
timeout only when `SourceTier=="pbs"` (both the scheduler + selftest spec builds), else 0.
|
||
- Tests: tier-aware `WaitOptions.Timeout` (pbs→120m, local→0; red-proofed against `WaitOptions{}`) +
|
||
the accessor contract. The "grant scratch-band `VM.Allocate`" follow-up was diagnosed, not
|
||
blind-applied — the scratch guest is restored INTO `/pool/felhom` (whose ACL already grants
|
||
`VM.Allocate`), so the earlier teardown 403 was a *consequence* of the timeout (a not-yet-pooled,
|
||
still-restoring guest), not a missing grant. **No ACL/host-install change.** (Live confirmation of
|
||
the phantom in REPORT.)
|
||
|
||
## v0.67.0 — S4: namespace-aware PBS client (per-customer offsite tenancy) (2026-07-04)
|
||
|
||
Phase-1 live probe on felhom-hetzner proved the offsite tenancy path (backup/restore/list/isolation
|
||
all green over the tunnel with a per-customer DatastoreBackup token) but surfaced that the agent's
|
||
PBS client was **namespace-unaware**: `Snapshots` hit the datastore root (403 for a scoped token)
|
||
and `Verify` was whole-datastore (needs Datastore.Verify ~ admin). Operator-approved small change to
|
||
make the client namespace-scoped so a properly-isolated token services its own tenant.
|
||
|
||
- **`internal/pbs`**: `Config.Namespace` (+ `Client.namespace`). `Snapshots` appends `?ns=<ns>`
|
||
(lists ONLY the tenant's namespace); `Verify` sends `ns=<ns>` (verifies ONLY that namespace —
|
||
Phase-1-confirmed to work with a **DatastoreBackup** token on its own ns, no Datastore.Verify /
|
||
admin widening). Root-ns clients (Namespace="") are unchanged → whole-datastore (the DooPlex
|
||
`felhom-pbs` n100 path). Test `TestClient_NamespaceScoping` pins both (red-proofed).
|
||
- **`internal/proxmox`**: `Storage.Namespace` (parsed from the PVE `/storage` config key `namespace`).
|
||
- **`cmd/felhom-agent`**: `pbsTargetsFromPVE` threads `s.Namespace` into the PBS client, so a PBS
|
||
storage configured with a namespace is verified/reported scoped to it automatically.
|
||
|
||
**Confirmed minimal tenant ACL (recorded live 2026-07-04, felhom-hetzner):** `DatastoreBackup` on
|
||
`/datastore/felhom-offsite/<ns>` (the namespace path — NOT `/ns/<ns>`) granted to **BOTH** the user
|
||
`felhom@pbs` **and** the token `felhom@pbs!<ns>` — PBS privsep tokens = intersection(user, token),
|
||
so both are required; isolation holds because each token's ACL is only its own ns (cross-ns
|
||
list/backup → 403, proven). No token exceeds DatastoreBackup; no admin on the endpoint for the box.
|
||
|
||
## v0.66.0 — S4 agent half: endpoint v4-pin + re-resolve watchdog + FELHOM_WG Critical flips (2026-07-04)
|
||
|
||
The two agent items S4 needs before offsite backups ride the tunnel (the tenancy + storage weight is
|
||
runbook-side). No new sudoers grants; no wire/JSON change.
|
||
|
||
- **`internal/wgtunnel` — v4-pin (doc 06 §4.2).** `renderConf` now takes the **pre-resolved IPv4
|
||
literal** and writes `Endpoint = <ip>:<port>` — never the DNS name, never an AAAA. A new `Resolver`
|
||
seam (`net.DefaultResolver.LookupNetIP(ctx, "ip4", …)` — A records only) resolves in the Manager;
|
||
multiple A records → the numerically **lowest** (deterministic fleet-wide). renderConf stays pure
|
||
(no DNS/IO inside). The resolved IP is cached: **steady-state Apply hits the cache — zero DNS, zero
|
||
execs** (the load-bearing negative). DNS failure on (re)resolve → keep the last-applied conf + throttled
|
||
ERROR — **never a teardown** (teardown stays revocation-only).
|
||
- **`internal/wgtunnel` — re-resolve watchdog (doc 06 §4.2, the slice-3 promise).** New
|
||
`Manager.Watchdog` (loop-driven only, so Apply's zero-exec steady state is untouched): when the
|
||
handshake age exceeds `wg_tunnel.stale_after_seconds` (default **180**) it re-resolves; **IP changed
|
||
→ re-render + restart** (endpoint re-IP recovery); IP unchanged → no churn, one throttled warn
|
||
(endpoint merely down). The staleness read is the existing `wg show … latest-handshakes` (never
|
||
`dump`).
|
||
- **`internal/config`** — `WGTunnelConfig.StaleAfterSeconds` (default 180 via `WithDefaults`).
|
||
- **`internal/capability` — FELHOM_WG Critical flips (S4).** Backups ride the tunnel now, so
|
||
`wg-conf-install`, `wg-enable`, `wg-restart`, `wg-handshake-read` are **Critical=true**
|
||
(operator-alert-worthy on degradation); `wg-tools-install` (one-time) + `wg-disable` (deliberate
|
||
revocation) stay non-critical. `TestWGCapabilityCriticality` pins the exact set (red-proofed).
|
||
- Tests: v4-pin golden (A literal, AAAA/dns_name refused), watchdog (healthy=no-DNS negative,
|
||
stale+re-IP restarts, stale+same-IP no-churn+throttle, resolver-failure keeps conf, initial-resolve-
|
||
failure no-teardown+recovery). Red-proofs a/b/d all fire.
|
||
|
||
## v0.65.0 — S3.1 offsite-tunnel client MTU 1420 → 1280 (resolve the CGNAT-smoke MTU open decision) (2026-07-04)
|
||
|
||
One-constant fix closing `06 §4.3`'s OPEN DECISION. The 2026-07-04 CGNAT smoke test found the
|
||
shipped interface MTU 1420 **silently black-holes bulk TCP** on any path below ~1480 B (mobile
|
||
~1400, DS-Lite ~1452): the WG handshake and ping stay healthy (small packets) while the PBS TLS
|
||
page — and, at S4, the backup itself — drops. "Looks green, loses backups." Must be safe before
|
||
S4 flows bulk TCP over the tunnel.
|
||
|
||
- **`internal/wgtunnel/manager.go`**: new `const clientMTU = 1280` (the RFC 8200 IPv6-minimum link
|
||
MTU — every path carries ≥1280; outer = 1280+60 v4 / +80 v6, fits every realistic path);
|
||
`renderConf` emits `MTU = %d` from it. **Fleet-wide, family-agnostic, permanent** — decouples
|
||
the fix from the v4/v6 endpoint-resolution question (§4.2). **Client-only by construction:** the
|
||
interface MTU caps box→PBS and the advertised MSS (=MTU−40) caps PBS→box, so the endpoint's `wg0`
|
||
is deliberately untouched (zero live-endpoint risk). Rejected: auto-probe / per-connection-type
|
||
(fragile moving part optimizing throughput, a non-metric here) and MSS-clamp (no forwarded flows).
|
||
- **`internal/wgtunnel/manager_test.go`**: golden pins exact `MTU = 1280`; red-proofed (flip const
|
||
→ 1420 fails the golden on the MTU line — non-vacuous).
|
||
- **`internal/hub/report.go`**: stale "MTU 1420" comment → 1280 (still NOT a wire field).
|
||
- No wire/JSON-golden change (MTU is a client-derived constant, never on the wire); no endpoint,
|
||
hub, controller, key, or desired-state change.
|
||
|
||
## v0.64.0 — S3 offsite WG tunnel: keygen + registration + agent-managed wg-quick@wg-felhom + escrow join (2026-07-04)
|
||
|
||
The agent half of doc 06 §3.3 (felhom.eu S1/S2 built the endpoint + hub half). **DEFAULT OFF —
|
||
the safety gate:** `wg_tunnel.enabled` defaults to false; a v0.64.0 rollout without explicit
|
||
config is a no-op (no keygen, no registration, no report stanza). Enabled explicitly on
|
||
felhom-pve only; the default flips when the production endpoint exists.
|
||
|
||
- **`internal/wgtunnel`** (new; the lanresolver host-service shape): pure-Go keygen
|
||
(x/crypto curve25519; key 0600 in 0700 StateDir/wg; corrupt file = refuse, NEVER overwrite —
|
||
it may be the escrowed identity; stored form canonical-clamped — x/crypto clamps derivation
|
||
internally, so the STORED bytes are the property that matters, red-proof-anchored); one-shot
|
||
registration (`POST /hosts/{id}/wg`, marker-gated, exponential backoff cap 15 m);
|
||
desired-state consumption via the new `desired.Syncer.AddConsumer` raw seam (panic-contained);
|
||
conf render golden-tested (MTU 1420, AllowedIPs = pbs_tunnel_ip/32, keepalive 25 — doc 06 §4
|
||
client constants; all inputs strictly validated); hash-gated apply (steady state = ZERO
|
||
execs), restart-not-reload on conf change, self-heal enable, adopt-lost-marker,
|
||
re-key-on-mismatch. **Revocation semantics (doc 06 §3.5 completed): block absent from a
|
||
PRESENT desired-state → disable + marker KEPT + never re-register** (the operator re-adds the
|
||
peer using the reported pubkey); absent DATA (failed fetch) is never a teardown signal.
|
||
- **`internal/hub`**: `WireDesiredState.Wireguard` (field-exact with the S2 cross-repo golden,
|
||
copied byte-identical + decode test), `RegisterWG` client (typed errors, token-free),
|
||
`HostReport.Wireguard` status stanza `{pubkey, registered, active, last_handshake_age_s,
|
||
assigned_ip}` via the collector's `WireguardReporter` seam.
|
||
- **Sudoers/capabilities**: `Cmnd_Alias FELHOM_WG` (fixed-path conf install, enable/restart/
|
||
disable, `wg show wg-felhom latest-handshakes` — the ONLY wg read; `wg show … dump` is
|
||
FORBIDDEN, its interface line carries the PRIVATE KEY) + 6 manifest entries (Critical=false
|
||
until S4 makes the tunnel load-bearing).
|
||
- **Escrow**: `IdentityBundle.WGPrivateKey` (omitempty) + escrow-create auto-inject when the key
|
||
file exists (`escrow.AttachWGKey`; field NAME only in logs). Honest limit: pre-S3 blobs cannot
|
||
be retro-fitted (R is never retained) — S5 DR falls back to fresh-key re-registration, which
|
||
keeps the box's /32 (hub S2 re-key-in-place).
|
||
- `--selftest=wgtunnel` single-shot for supervised bring-up.
|
||
- Live-validated on felhom-pve (agent restart tolerance, host reboot with unit persistence,
|
||
revocation drill, 30-min keepalive soak, escrow inject) — see REPORT.md. Five red-proofs run
|
||
+ reverted. GOTCHA learned: the hub envelope's `poll_interval_seconds` (hub-side constant
|
||
900 s) silently overrides the agent's configured cadence on the FIRST cycle — the agent-side
|
||
`poll_seconds` is only the pre-first-heartbeat default.
|
||
|
||
## configs: build-golden.sh v2.0.0 — mandatory controller tag + bootstrap .path unit (B5 + B1) (2026-07-03)
|
||
|
||
Golden-bake script only — **no agent code, no binary, no agent version bump** (docs/config
|
||
precedent). Resolves drill findings B5 + B1
|
||
(`felhom.eu/documentation/audits/DRILL-day0-cleanroom-2026-07-03.md` §9) and closes the stale
|
||
backlog note `FOLLOWUP-golden-default-controller-tag.md`.
|
||
|
||
- **B5 — `CONTROLLER_IMAGE` (arg 6) is now MANDATORY** — the hand-bumped default rotted twice
|
||
(0.43.0 → 0.85.1 → stale again; 0.85.1 predates the v0.86.0 floor-honoring code, which is why
|
||
every fresh install needed the manual D.1b update). No 6th arg → die with usage (red-proofed:
|
||
exits 1 before any `pct` op). Auto-resolving "latest" was rejected — it could bake an unvouched
|
||
tag.
|
||
- **B1 — baked `felhom-controller-bootstrap.path` unit** (`PathExists=/etc/felhom-bootstrap/
|
||
bootstrap.json`, enabled next to the service): the service's `ConditionPathExists` is evaluated
|
||
only at boot, but the agent back-half hot-plugs the bootstrap mount into the already-running
|
||
guest — the path unit starts the service when the file APPEARS, so the controller deploys with
|
||
NO reboot (isolated systemd proof + full Day-0 proof on the clean-room drill VM; the installer's
|
||
v1.9.1 post-provision reboot is now a redundant belt, kept). `RemainAfterExit=yes` on the service
|
||
prevents re-trigger loops; the service itself stays unchanged.
|
||
- `GOLDEN_SCRIPT_VERSION` (2.0.0) + a `[golden]` provenance line (script version + baked controller
|
||
tag) now open every bake transcript.
|
||
- Baked + published with controller **0.98.3**: golden `0.98.3` /
|
||
sha256 `b9a02ef1b6f02b9b58babc4c6aad9cf6c053ebdfba116c78c8e7830de757fd01` (Gitea generic package,
|
||
201 + sha round-trip). Clean-room validation: bake integrity (mp0+mp1 included), isolated
|
||
hot-plug proof, full local-golden Day-0 install (first boot = 0.98.3, self-update reports
|
||
up-to-date, app deploy OK) — evidence:
|
||
`felhom.eu/documentation/audits/DRILL-golden-098-2026-07-03.md`.
|
||
|
||
## v0.63.0 — B3 + B2: fresh-install fixes — token reload-on-miss + guesthook snippets dir (2026-07-03)
|
||
|
||
The two agent-side gaps the Day-0 clean-room drill surfaced
|
||
(`felhom.eu/documentation/audits/DRILL-day0-cleanroom-2026-07-03.md` findings B3/B2). Both are
|
||
fresh-install path fixes; no behavior change on a warm box.
|
||
|
||
- **B3 — `TokenStore.Lookup` reload-on-miss** (localapi/tokenstore.go): provisioning is a SEPARATE
|
||
one-shot process (`--selftest=provision`) that Mints the new guest's token into the shared
|
||
append-only JSONL, while the long-lived daemon serves Lookup from an index built once at open —
|
||
so the daemon 401'd every token minted after it started (the drill's
|
||
`POST /controller/swap: HTTP 401`, until a manual `systemctl restart felhom-agent`). Lookup now
|
||
re-reads the file ONCE on a miss (`reloadLocked()`, factored from `load()`; full re-read is
|
||
idempotent under `apply`'s last-write-wins) and re-checks. An append-only size short-circuit
|
||
bounds the cost: an unknown token on an unchanged file is one `stat`, no re-read — never a reload
|
||
loop. Fix is entirely behind the `TokenAuthority` seam (no server change). Fail-closed on an
|
||
unreadable store; missing file loads as empty. Tests (tokenstore_test.go): cross-process-mint
|
||
coherence (red-proofed: pre-fix shape returns (0,false)), exactly-once reload bound +
|
||
size short-circuit, cross-process re-mint rotation coherence, deleted-file no-crash
|
||
(linux-only; windows can't unlink an open handle).
|
||
- **B2 — `guesthook.InstallSnippet` ensures the snippets dir** (guesthook/install.go): a fresh PVE
|
||
has no `/var/lib/vz/snippets` and `install` (without `-D`) won't create parents — the pre-start
|
||
self-heal hook silently failed to install on every freshly-bootstrapped box (warn-only in the
|
||
back-half). A fenced `mkdir -p /var/lib/vz/snippets` now precedes the install; **sudoers gains
|
||
exactly that one grant** (FELHOM_GUESTHOOK — configs/felhom-agent.sudoers must ship WITH this
|
||
binary, as always). Test: mkdir-precedes-install argv assertion (red-proofed: pre-fix has no
|
||
mkdir op).
|
||
- Guide follow-through (felhom.eu, separate commit): D.1b's "restart the agent first" step drops
|
||
once B3 is live-verified.
|
||
|
||
## v0.62.0 — A1: pool-membership ownership check for the stale-lock reaper (2026-07-03)
|
||
|
||
Implements audit finding A1 (`AUDIT-blast-radius-hostroot-localapi-2026-07-02` §A) per the spike
|
||
verdict (`SPIKE-a1-pool-membership-read-2026-07-03` — enumeration is pool-filtered under the scoped
|
||
token, so this is defense-in-depth: a future broad-token deployment can no longer re-arm the reaper
|
||
against co-tenant guests). Companion: host-install **v1.9.0** (`Pool.Audit` added to
|
||
`FelhomAgentGuest`) — **rescope BEFORE deploying this agent**, else the reaper fail-safes (skips)
|
||
until the ACL catches up.
|
||
|
||
- **`Client.Pool`** (proxmox/query.go): `GET /pools/{name}` → `PoolInfo{PoolID, Members[]{VMID,Type}}`.
|
||
Requires `Pool.Audit` at `/pool/{name}`; `Pool.Allocate` does NOT satisfy the read (spike T2).
|
||
- **`staleLockController.Guests()`** (localapi/stalelock.go): now returns `ListLXC ∩ pool members`
|
||
(nonzero-vmid, non-storage entries only). Ownership is PROVEN via the pool registry, never assumed
|
||
from enumeration scope. A pool-read failure returns a wrapped error ("pool membership read
|
||
(pool=felhom): …") that rides the existing "guest list unavailable — skipping recovery" guard —
|
||
fail-safe: NO unlock/snapshot-delete/start on ANY guest, never a fallback to the unfiltered list.
|
||
One new INFO line per scan: `stale-lock: scanning pool guests` (pool, listed, scanned) — emitted
|
||
by the controller (the unchanged `StaleLockController` seam can't carry the pre-intersect count).
|
||
- **`NewStaleLockController`** gains `(pool string, logger *slog.Logger)`; main.go threads
|
||
`reconcile.DefaultPool`.
|
||
- **Capability surfacing**: the hub-report prober is now a composed closure — the sudo manifest
|
||
probe + one `pve:pool-read` status (non-critical; degraded ⇒ reaper is fail-safed, visible on the
|
||
report, no operator page). Composed in main.go; `internal/capability/` untouched.
|
||
- **`--selftest`**: new "pool read" line (pool id + member count + guest members).
|
||
- **Tests** (stalelock_pool_test.go, driving the REAL controller over a broad-token-shaped fake):
|
||
`TestStaleLock_ForeignGuestNotReaped` (red-proved: intersect removed ⇒ FAILS with
|
||
`pct unlock 5000` recorded), `TestStaleLock_PoolGuestStillReaped` (anti-over-filter),
|
||
`TestStaleLock_PoolReadFails_SkipsAll` (red-proved: fallback-to-unfiltered ⇒ FAILS with mutations
|
||
recorded), `TestStaleLockController_GuestsIntersect` (storage-member + empty-pool edges). The 9
|
||
existing Server-level stalelock tests pass unmodified.
|
||
|
||
## docs — CLAUDE.md refresh: stable orientation, complete layout (2026-07-03)
|
||
|
||
No code change, no version bump. Deleted the version-pinned "Current: v0.31.0" narrative and the
|
||
per-slice history (stale by 30 versions — current state lives in CONTEXT.md/CHANGELOG top); Layout
|
||
completed with the 8 missing packages (capability, desired, escrow, guesthook, lanresolver, localapi,
|
||
provision, signedjobs + cmd/felhom-opsign — verified against the tree); build/deploy compressed to a
|
||
summary table pointing at the `felhom-build-deploy` skill (commands verified live on felhom-pve:
|
||
non-root `felhom-agent` user, `/usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json`).
|
||
Load-bearing Proxmox-model rules kept verbatim. Standing rule: no version-pinned state in CLAUDE.md.
|
||
|
||
## docs — REUSE.md introduced (2026-07-03)
|
||
|
||
Cross-repo reuse-map rollout (docs-only, no code change, no version bump). New `REUSE.md` at the
|
||
repo root: curated map of canonical helpers (48 rows — exec/sudoers surface, format-safety guards,
|
||
durable-id seams, stores, local-API plumbing), patterns, dangerous lookalikes (req.Device TOCTOU,
|
||
raw mkfs, uuid:-vs-byid: scheme confusion, MemoryNonceStore, pool-blind stale-lock scan…), test
|
||
seams, extension points, and observed duplication (7 clusters, NOT fixed). Every entry code-verified
|
||
at file+symbol; cited paths machine-checked by `felhom.eu/scripts/reuse_refs_check.py` (green).
|
||
CLAUDE.md gains the "See REUSE.md before writing new code" pointer + the same-commit maintenance rule.
|
||
|
||
## v0.61.0 — blast-radius audit fixes B1 + D1 + D2 + D3 (2026-07-03)
|
||
|
||
Four LOW/INFO fixes from `AUDIT-blast-radius-hostroot-localapi-2026-07-02.md` — the "second gate must
|
||
mirror the first" batch. Each shipped with a non-hollow test AND a companion red-proof (test shown
|
||
failing on the pre-fix implementation). A1 (stale-lock pool-membership) is deliberately NOT here — it
|
||
needs a pool-read spike (the role lacks `Pool.Audit`); C1/C2/A2/B2–B5/E1/E2 deferred.
|
||
|
||
- **B1 (LOW) — random temp staging for root-installed scripts.** `internal/guesthook/install.go`
|
||
(`InstallSnippet`) and `internal/localapi/intermediary.go` (`installSharedParentUnit`, new shared
|
||
`stageTemp`) staged root-executed scripts through FIXED, predictable `/tmp` names via `os.WriteFile`
|
||
(no O_EXCL, follows symlinks) — a local TOCTOU into a root-run PVE hookscript / boot script. Both now
|
||
use the `os.CreateTemp` random-name pattern `lanresolver` already used. `configs/felhom-agent.sudoers`
|
||
(`FELHOM_GUESTHOOK`/`FELHOM_INTERMEDIARY`) install-SOURCE grants became globs
|
||
(`/tmp/felhom-guest-hook-*.sh`, `/tmp/felhom-shared-parent-*.{sh,service}`; destinations stay pinned);
|
||
`internal/capability/manifest.go` representative vectors updated to match. Tests:
|
||
`TestInstallSnippet_RandomTempName`, `TestInstallSharedParent_RandomTempName` (fake runner records the
|
||
install source: random pattern, two calls differ, content + cleanup asserted).
|
||
- **D1 (LOW) — the guarded mkfs wrapper now mirrors the classifier's member/RO classes.**
|
||
`configs/felhom-mkfs-guarded.sh` re-checked only system-disk / LVM-PV (PATH-dependent `command -v
|
||
pvs`) / foreign-mount — a bypassed agent could mkfs a ZFS/mdraid/LUKS/swap member or a read-only
|
||
disk. Added (additive; nothing removed/reordered): `/sys/block/<disk>/ro` == 1 → die; an
|
||
lsblk-FSTYPE loop over the whole disk refusing exactly `claim.go`'s `memberFSTypes`
|
||
(LVM2_member/zfs_member/linux_raid_member/crypto_LUKS/swap — the FSTYPE catch works with pvs
|
||
absent); pvs resolved via absolute candidates (/usr/sbin/pvs, /sbin/pvs). Validated by the new
|
||
`scripts/mkfs-guarded-harness.sh` on felhom-pve: throwaway loop devices + PATH-shimmed lsblk + a
|
||
recorder bind-mounted over mkfs.ext4 in a private mount namespace (no real mkfs possible) — fixed
|
||
wrapper 8/8 (incl. plain-blank-disk still formats); pre-fix wrapper red-proof: 7/8 hostile fixtures
|
||
reached mkfs.
|
||
- **D2 (INFO) — `classifyClaim` empty-lsblk fail-safe.** A successful-but-empty `lsblk`
|
||
(`{"blockdevices":[]}`) skipped the member/mount loop and returned `unclaimed`.
|
||
`internal/storage/claim.go` now refuses when the node tree is empty OR the target whole-disk is
|
||
absent from it (undeterminable topology ⇒ claimed). Tests: `TestClassifyClaim_EmptyNodesRefused`,
|
||
`TestClassifyClaim_TargetAbsentFromTree`.
|
||
- **D3 (INFO) — blank-format anti-retarget (AGENT-001's benign-branch twin).** `handleDiskFormat`'s
|
||
blank branch formatted the mutable caller-supplied `req.Device` with no durable-id binding — a /dev
|
||
re-enumeration between inspect and mkfs could format a data-bearing disk that inherited the node.
|
||
The blank branch (`internal/localapi/disks.go`) now derives the device's durable id (no durable id ⇒
|
||
409 refuse — path-only formats are not permitted), re-resolves it via the new
|
||
`antiRetargetResolveBlank` (`wipe_reresolve.go`: shared `antiRetargetResolveExpect` core; the blank
|
||
variant asserts the device is STILL !DataBearing), and formats the RE-RESOLVED device. The format
|
||
job record (`formatjob.go`) carries `blank`; restart recovery re-checks blank jobs with the blank
|
||
variant (durable-id-bound, fail-safe refuse). Confirmed/data-bearing branch untouched. Tests:
|
||
`TestFormatBlankPath_AntiRetarget_{ReassignedDataBearingRefused,ReassignedDifferentDiskRefused,UnresolvableRefused,SameBlankProceeds}`,
|
||
`TestFormat_Blank_{FormatsReresolvedDeviceNotCallerPath,ReresolveRefusalNoMkfs,NoDurableIDRefused}`.
|
||
- Deploy note: the host's `/etc/sudoers.d/felhom-agent` MUST be updated together with the v0.61.0
|
||
binary (the old fixed-name grants deny the new random-name installs, and vice versa).
|
||
|
||
## v0.60.0 — proof-of-launch destroy gating + restore-test band-advance (campaign F1/F2) (2026-07-02)
|
||
|
||
Fixes the pool-effects campaign's HIGH finding (F1, `CAMPAIGN-pool-effects-2026-07-01.md`): the bring-up
|
||
compensating rollback and the restore-test teardown fired `DestroyLXC` on the target vmid even when
|
||
`RestoreLXC` failed synchronously WITHOUT creating anything (PVE refusing a pre-existing vmid the
|
||
pool-blind duplicate guard / band scan couldn't see) — destroying a guest the transaction never made.
|
||
Only the pool ACL's 403 saved the non-pool subset; an in-pool pre-existing guest would have been
|
||
destroyed, and any broad-token deployment re-arms the bug. Root cause: SameTxnCreated/scratch provenance
|
||
was ASSUMED, never verified against proof-of-launch. The fix makes a `RestoreLXC` UPID the sole destroy
|
||
authorization, in ALL THREE destroy paths — the pool ACL is defense-in-depth again, not the guard.
|
||
|
||
- **F1a `internal/reconcile/bringup.go` `runBringUp`:** the compensating-rollback defer is gated on
|
||
`launched` (set only after the restore POST is accepted). A synchronous restore failure (no UPID)
|
||
closes the owning entry terminal-failed WITHOUT any destroy. The pre-restore `OpStarted` journal
|
||
append is kept (crash-safety); `rollbackBringUp` is now only ever called launch-proven.
|
||
- **F1b `internal/reconcile/restoretest.go` `runScratchTest`:** same `launched` gate on
|
||
`teardownScratch` — a synchronous restore refusal never destroys the picked band vmid.
|
||
- **F1c `internal/reconcile/recover.go` `Recover`:** the no-UPID "POST never confirmed → abandon
|
||
fail-safe" check now runs BEFORE the Scratch/Rollback dispatch — a no-UPID Scratch/Rollback entry is
|
||
abandoned (marked failed, NO destroy) instead of destroy-by-vmid-existence. Recover is now safe by
|
||
DESIGN, not by the pool-blind "already gone" accident the campaign observed.
|
||
- **F2 `restoretest.go` `RunRestoreTest` band-advance:** a band vmid PVE refuses with "already exists"
|
||
(an invisible squatter — new `pveAlreadyExists`, mirrors `pveConfigLock`, never misclassifies a real
|
||
restore failure) is skipped and the next free band vmid tried (bounded by the band width;
|
||
`pickScratchVMID` gained an exclude set). A fully-occupied band → `Skipped` (scheduler raises no
|
||
"backup unrestorable" alert), never FAIL — one squatter no longer permanently breaks the restore-test.
|
||
- **Accepted residual (by design):** a crash in the one-statement window between obtaining the UPID and
|
||
journaling it leaks a half-built guest Recover won't destroy — cleanable, and vastly preferable to
|
||
destroying an innocent guest.
|
||
- Tests: red-proof companions verified (gates reverted → `TestRunBringUp_NoLaunchNoDestroy`,
|
||
`TestRunRestoreTest_RestoreNoLaunchNoTeardown`, `TestRecover_{BringUp,Scratch}NoUPIDAbandoned` all
|
||
fail with the innocent-guest destroy); no-regression `…LaunchedTaskFailureStillTearsDown` + rollback
|
||
table now includes an explicit restore-task-failure case; F2 advance + squatter-full-band-skips tests.
|
||
Live-validated on felhom-pve (provision onto existing 9001 → no destroy armed; restore-test advances
|
||
past a 990000 decoy). `go build/vet/test ./...` clean.
|
||
|
||
## v0.59.0 — report backing device + capacity for a registry-sourced drive in /disks (2026-07-01)
|
||
|
||
Completes the `/disks` representation for a registry-sourced (raw, no-PVE-storage) drive: the agent-view
|
||
showed "—" for the device and no size bar, because the union row never populated `backing_device` or
|
||
`total_bytes`/`used_bytes` (Observe drives get those from `pvesm status`, which a raw drive has none of).
|
||
|
||
- **`internal/localapi/disks.go` handleDisks registry union:** resolve `BackingDevice` from the fs-UUID
|
||
(`storage.ByUUIDDevicePath`) and read capacity via `statfsCapacity` (new build-tagged
|
||
`capacity_linux.go` = `syscall.Statfs` on the mount; `capacity_other.go` = no-op for dev builds).
|
||
- `go build/vet/test ./...` clean (Linux + Windows dev). Live: the registry drive now shows its device +
|
||
size in the agent-view, matching the Observe-sourced drives.
|
||
|
||
## v0.58.0 — report GuestPath/BoundUnderParent for a registry-sourced drive in /disks (2026-07-01)
|
||
|
||
Last piece of first-class raw-drive support: the `/disks` union row for a registry-sourced drive (Impl-2a
|
||
— a drive with no PVE storage) omitted `GuestPath` + `BoundUnderParent`, so the controller read it as
|
||
"Leválasztva" (disconnected) even though it was mounted + bound + live in the guest.
|
||
|
||
- **`internal/localapi/disks.go` handleDisks registry union:** populate `GuestPath` (`StablePathForRaw`)
|
||
+ `BoundUnderParent` (`boundUnderParent`) on the registry row, identical to the Observe path — so a
|
||
registry-only drive reports its true bound/active state. `go build/vet/test ./...` clean.
|
||
|
||
## v0.57.0 — re-assert a RAW drive's guest-bind (ReassertGuestBinds mount-table fallback) (2026-07-01)
|
||
|
||
Completes the raw-drive durability the v0.56.0 fix started. `ReassertGuestBinds` (the startup / drive-
|
||
returned reconcile that re-binds an enrolled drive's felhom-data under the shared parent so it's live in
|
||
the guest) built its durable-id→mount map from `Observe()` only — so a RAW enrolled drive was never found
|
||
("enrolled drive not present"), and its in-guest bind was not re-asserted after a reboot or a watchdog
|
||
re-mount (the drive would show "Leválasztva" in the controller).
|
||
|
||
- **`internal/localapi/disks.go` `ReassertGuestBinds`:** augment the durable-id→mount map from the mount
|
||
table — each raw `/mnt/<name>` mount → its device fs-UUID (`HostReader.Mounts`+`ResolveUUID`), skipping
|
||
the `/mnt/felhom-drives` bind (AttachDrive wants the raw path). Observe entries still win. Also: an
|
||
Observe failure is no longer fatal (fall through to the mount-table scan) so raw drives re-assert even
|
||
if the PVE view is momentarily unavailable.
|
||
- **Wiring fix (latent):** `buildLocalAPIServer` never passed `Options.HostReader`, so the local-API
|
||
server's `host` was nil in production — the v0.56.0 `durableIDForMount` raw fallback (and the role gate's
|
||
host classification) silently no-op'd. Now wired to `storage.NewProcHostReader()`. This is what makes
|
||
the v0.56.0 + v0.57.0 raw-mount resolutions actually fire live.
|
||
- `go build/vet/test ./...` clean. With v0.56.0 (guest-bind now RECORDED for raw drives) this closes the
|
||
reboot/reconnect guest-bind durability gap for raw drives end-to-end.
|
||
|
||
## v0.56.0 — record intent/guest-bind for a RAW enrolled drive (durableIDForMount fallback) (2026-07-01)
|
||
|
||
Surfaced by the first live raw enrollment (Impl-2b): a raw drive is not a PVE storage, so
|
||
`durableIDForMount` (Observe-based) returned "" for it → the enroll's intent + guest-bind recording
|
||
logged "durable-id unresolved" and silently skipped. Result: the drive mounted + bound + usable, but was
|
||
NOT intent-tracked (so `RegistryKnownTargets` — which gates on intent ≠ new — didn't health-track it) and
|
||
its guest-bind wasn't persisted.
|
||
|
||
- **`internal/localapi/disks.go` `durableIDForMount`:** after the Observe lookup, fall back to resolving
|
||
the fs-UUID directly from the mount table — the device mounted at `where` (via `HostReader.Mounts`) →
|
||
its by-uuid identity (`HostReader.ResolveUUID`) → `uuid:<fs-uuid>` (the SAME scheme Observe derives, so
|
||
intent keys stay consistent). Fixes intent recording (enroll/eject) AND guest-bind recording for raw
|
||
drives; the PVE-storage path is unchanged.
|
||
- Test `TestDurableIDForMount_RawFallback` (+ red-proof: Observe-only → ""). `go build/vet/test ./...` clean.
|
||
- **(Residual noted here fixed in v0.57.0:** `ReassertGuestBinds` raw-drive guest-bind re-assert.)
|
||
|
||
## v0.55.0 — raw-device discovery + registry-sourced drive tracking (Impl-2a) (2026-07-01)
|
||
|
||
Agent backend for drive enrollment (SPIKE-drive-enrollment §SQ1/SQ4/SQ5). Makes raw (non-PVE-storage)
|
||
drives (a) discoverable for enrollment and (b) health-tracked WITHOUT being a PVE storage — so a drive
|
||
enrolled the new way isn't enrolled-but-untracked (the 3b-fix false-detach class). No mkfs here (Impl-1
|
||
owns it); the controller wizard rewiring is Impl-2b.
|
||
|
||
- **`GET /disks/candidates`** (`internal/storage/candidates.go` + `internal/localapi/disks.go`):
|
||
enumerates host whole-disks from `/sys/block`, runs the Impl-1 unclaimed filter, and returns the free
|
||
ones with probe info (size/model/FS/data-bearing/durable-id preview), split into `initialize` (all
|
||
unclaimed) and `attach` (the subset carrying a mountable ext4/xfs FS). Fail-safe carries through (a
|
||
device not provably unclaimed is omitted).
|
||
- **`RegistryKnownTargets`** (`internal/storage/registry_known.go`): the watchdog's known-DRIVE set now
|
||
comes from the intent registry + Felhom `.mount` units, NOT `Observe()` (PVE storages). A unit is
|
||
tracked iff its intent ≠ `new` (enrolled/ejected/decommissioned; the watchdog's existing IntentReader
|
||
gate still decides re-mount). `main.go` swaps the watchdog `KnownTargets` source. **`Observe()` is
|
||
KEPT** for real PVE storages (local/local-lvm/pbs) — reports + the `/disks` view.
|
||
- **`handleDisks` union:** additive + deduped-by-mount-path — appends registry drives Observe doesn't
|
||
surface (a registry-only drive now appears in the agent-view) without dropping any Observe row (can't
|
||
regress the current view).
|
||
- **Existing-drive migration** (`ReconcileExistingDrives`, idempotent, at agent start): records each
|
||
currently-mounted Felhom-unit drive as `enrolled` so the registry-sourced `Known()` tracks it without
|
||
its legacy PVE dir-storage. Does NOT create/remove PVE storages.
|
||
- **Tests:** RegistryKnownTargets (enrolled tracked / new excluded / ejected tracked) + the **red-proof**
|
||
(Observe-based `Known()` misses a drive with no PVE storage; the registry provider tracks it);
|
||
migration idempotency; candidates init/attach split. `go build`/`vet`/`test ./...` clean. Watchdog/
|
||
HostLiveness/Remounter unchanged (only the source injected).
|
||
|
||
## v0.54.0 — format-safety foundation: unclaimed-disk guard + guarded-mkfs wrapper (2026-07-01)
|
||
|
||
Impl-1 (SPIKE-drive-enrollment-2026-07-01). Hardens the destructive `Format`/mkfs path BEFORE the
|
||
enrollment feature: today `Format` delegates authorization to its caller and only checks `DataBearing`
|
||
(has-data), which is insufficient — the OS disk is data-bearing yet catastrophic — and the sudoers
|
||
permits `mkfs /dev/*`. Two independent, layered guards:
|
||
|
||
- **Part A — mandatory unclaimed-disk guard inside `Format` (the primary safety).** New
|
||
`internal/storage/claim.go`: `classifyClaim` (pure) + `gatherClaimFacts` refuse to format any device
|
||
not provably UNCLAIMED — reusing `SystemDisks` (OS disk) + lsblk FSTYPE (LVM2_member/zfs_member/
|
||
linux_raid_member/crypto_LUKS/swap) + foreign mounts + read-only + authoritative `pvs`/`zpool`. A
|
||
Felhom-owned mount under `/mnt/felhom-drives` is NOT a foreign claim (re-init stays allowed; the
|
||
DataBearing wipe-confirm still gates data loss). **FAIL-SAFE: any read error / undeterminable topology
|
||
→ CLAIMED → refuse.** The guard is in `Format` (lowest layer), not the handler, so no caller can
|
||
bypass it. Read-only sudoers additions: `pvs`, `zpool status` (in `FELHOM_DISK`).
|
||
- **Part B — guarded-mkfs wrapper below the agent.** `configs/felhom-mkfs-guarded.sh` (root, 0755) is
|
||
now the ONLY mkfs path the sudoers allows (`FELHOM_FORMAT` no longer allowlists raw `mkfs.*`). It
|
||
re-checks the cheap catastrophic cases (system disk / LVM PV / foreign mount) and refuses — so even
|
||
an agent bug/compromise can't mkfs the OS disk. `Format` execs the wrapper (`<device> <fstype>`) via
|
||
`Binaries.MkfsGuarded`.
|
||
- **Tests:** `claim_test.go` — table-driven `classifyClaim` (every claim signal + fail-safe + the two
|
||
allow cases) incl. the **red-proof** (a claimed, non-data-bearing OS disk: removing the isSystem check
|
||
flips it to allowed → test fails, proving the guard adds safety beyond `DataBearing`); Format-guard
|
||
integration tests (refuses system disk / LVM member, allows unclaimed → wrapper invoked); capability
|
||
manifest updated (mkfs sample → the wrapper). `go build`/`vet`/`test ./...` clean.
|
||
- **Deferred to Impl-3:** a raw disk passed through to ANOTHER VM looks unused to the host — a host-level
|
||
filter can't detect it; the operator gate (shared-box mode) closes that. Impl-1 closes everything
|
||
host-visible (a strict improvement over today's no-guard state).
|
||
|
||
## v0.53.0 — restore guests INTO the felhom pool (pool-scoped-ACL enabler) (2026-07-01)
|
||
|
||
Colleague-safety batch #4 phase b (agent half). Enables the agent token to be scoped from `/` to
|
||
`/pool/felhom` + `/storage/<targets>` (real blast-radius containment on a shared host) by making every
|
||
restore allocate the guest INTO the pool — the only way a fresh vmid authorizes under a pool-scoped
|
||
token. Grounded by `felhom.eu/documentation/audits/SPIKE-pool-scoped-acl-2026-07-01.md` (PASS).
|
||
|
||
- **`internal/proxmox/mutate.go`:** `RestoreLXCOptions` gains `Pool string`; `RestoreLXC` sends
|
||
`pool=<p>` only when non-empty (`pct restore --pool`). Omit-when-empty (a broad-token restore needs
|
||
no pool) — unit-tested + red-proofed.
|
||
- **`internal/reconcile`:** new `const DefaultPool = "felhom"` (single source of truth); `BringUpSpec`
|
||
gains `Pool`, threaded to the bring-up restore. **Both restore sites** now pool the guest: the
|
||
provision/DR bring-up (`Pool: spec.Pool`, set to `DefaultPool` by the CLI) AND the **restore-test**
|
||
scratch guest (`Pool: DefaultPool`) — the latter closes SPIKE residual #2 (a pool-scoped token would
|
||
otherwise 403 on the out-of-pool scratch guest).
|
||
- **`cmd/felhom-agent/main.go`:** both `BringUpSpec` literals (bring-up/DR + provision) set
|
||
`Pool: reconcile.DefaultPool`.
|
||
- **No ACL/priv change in the agent** — that ships in the host-install script (v1.6.0). The pool param
|
||
is INERT until the token is granted `Pool.Allocate` at `/pool/felhom` and the pool exists; publishing
|
||
is therefore safe ahead of the coordinated ACL swap.
|
||
- New tests: `proxmox.TestRestoreLXC_PoolParam` (set → `pool=felhom`; empty → omitted),
|
||
`reconcile.TestRestoreSitesUsePool` (both restore sites carry `DefaultPool`). Both red-proofed
|
||
(unconditional `Set` → omit test fails; drop either site's `Pool` → both-sites test fails). `go
|
||
build`/`vet`/`test ./...` clean.
|
||
|
||
## v0.52.0 — operator-opt-in CPU/RAM cap for the provisioned guest (`-cores` / `-memory`) (2026-07-01)
|
||
|
||
Colleague-safety batch #3. So a trial appliance guest on a colleague's SHARED production Proxmox does
|
||
not pressure his existing guests, the operator can now cap the guest's CPU cores + RAM **at provision
|
||
time, before the guest's first boot** (the peak container-pull moment). Pure CLI→spec plumbing — the
|
||
reconcile engine already applied the cap; this only wires the flags to it.
|
||
|
||
- **`cmd/felhom-agent/main.go`:** new `-cores N` / `-memory M` (MiB) flags for
|
||
`--selftest=bring-up|provision` (0 = keep the golden's baked size). They flow through `bringUpSizing`
|
||
(now carries `Cores`/`MemoryMB`) into the `reconcile.BringUpSpec{Cores,MemoryMB}` built by BOTH
|
||
`runSelftestBringUp` and `runSelftestProvision`. The `-selftest` usage string documents them.
|
||
- **No engine change.** `internal/reconcile/bringup.go` already carries `BringUpSpec.Cores/MemoryMB`
|
||
(0 = leave as restored) and `buildBringUpConfig` already emits `cores`/`memory` into the SAME
|
||
coalesced config PUT as the identity reset — which runs BEFORE `e.api.Start`, so the cap lands
|
||
pre-boot. Rejected the `pct set`-post-provision alternative (runs after boot = an uncapped window;
|
||
bypasses the token/audit; second config source of truth).
|
||
- **Omit-when-zero guarantee:** an unset cap (0) emits NEITHER `cores` NOR `memory`, so an uncapped
|
||
provision keeps the golden defaults (no regression for the normal single-purpose box) and can never
|
||
shrink the guest to 0 cores. New pure-function test `TestBuildBringUpConfig_ResourceCaps` asserts
|
||
both the set (`cores=2`,`memory=4096`) and the absent-when-unset cases; a red-proof (unconditional
|
||
emit) was run and confirmed to fail the omit assertion, then reverted.
|
||
- **Deploy dependency:** a FRESH host-install `--cores`/`--memory` (felhom.eu script v1.4.0) requires
|
||
the hub artifact manifest to serve **agent ≥ v0.52.0**, else the old agent rejects the unknown flag.
|
||
The flags are opt-in, so nobody hits this until they intentionally cap.
|
||
- `go build` / `go vet` / `go test ./...` clean.
|
||
|
||
## v0.51.0 — local vzdump retention default (`--prune-backups keep-last=3`) (2026-06-30)
|
||
|
||
The PREVENTIVE counterpart to the hub's host_disk + storage_fill detectors: the agent's periodic local
|
||
whole-guest vzdump now prunes its own old archives, so a box can't refill its own root via its own backups
|
||
(the felhom-pve incident's root cause — that vzdump carried no retention, ~18 dumps piled under
|
||
`/var/lib/vz/dump`).
|
||
|
||
- **`internal/proxmox/mutate.go`:** `VzdumpOptions.PruneBackups` → passed as PVE's `--prune-backups` on
|
||
the vzdump POST (vmid+storage scoped, so PVE prunes only THIS guest's archives on THIS storage).
|
||
- **`internal/backup/runner.go`:** `NewBackupRunner` gains a `retention` arg; the backup applies it via
|
||
`localPruneSpec` ONLY when the target is a **non-PBS** storage (resolved via `ListStorage`) — PBS offsite
|
||
retention is a separate lifecycle and is never pruned by the per-run flag. **Fail-safe:** if the target
|
||
type can't be confirmed (lookup error / not found) the run SKIPS pruning rather than risk pruning PBS
|
||
(the detectors remain the safety net). Only the periodic local-API runner sets retention; the
|
||
restore-test / selftest runners pass "".
|
||
- **`internal/config/config.go`:** `backup.local_backup_retention` (keep-last N) with `KeepLast()` clamped
|
||
to **≥1** (0/unset/negative → default 3) — a mis-config can NEVER prune the just-made backup —
|
||
+ `PruneBackupsSpec()` → `keep-last=N`. Wired into the local-API backup runner (`main.go`).
|
||
- **Seeding:** `felhom.eu scripts/felhom-host-install.sh` seeds `local_backup_retention: 3` in the agent
|
||
config; the code default also protects any box where it is unset (KeepLast → 3) from day 0.
|
||
- F2-b stale-vzdump-lock recovery untouched.
|
||
- Tests: the local vzdump carries `--prune-backups keep-last=3` (+ companion: no-retention runner emits no
|
||
prune); **PBS is never pruned** (+ companion: same retention on a local target IS applied);
|
||
fail-safe-on-unknown-target; the **keep-last≥1 clamp** companion. `go build/vet/test ./...` green.
|
||
|
||
## v0.50.0 — NAS network storage Part A1: NFS/SMB automount foundation (2026-06-30)
|
||
|
||
Agent foundation of the validated `SPIKE-nas-storage-2026-06-29.md` (verdict READY): a customer NAS can
|
||
serve **bulk media** to a media app. The agent mounts a NAS share **host-side** under
|
||
`/mnt/felhom-drives/<name>` via a systemd `.automount` (+ `.mount`) pair; it propagates into guest 9201 for
|
||
free through the existing shared `mp8` bind (no new mountpoint, no restart). A NAS is a **distinct storage
|
||
class** — it carries **no durable-id** and never enters the drive enroll/eject/decommission/wipe/SMART/
|
||
watchdog machinery. **Bulk-media class only; STOP before A2 (controller registry/UI) + B (restic-SFTP).**
|
||
|
||
- **`internal/storage/netmount.go` (NEW).** `NetworkMountSpec` + the locked SPIKE recipe:
|
||
- **NFS (preferred):** `What=server:/export`, `Type=nfs4`, `Options=vers=4.1,soft,timeo=50,retrans=2,noatime,_netdev`.
|
||
`soft` is the failure-isolation knob (clean EIO, never a `df`/guest wedge); a default `hard` mount is
|
||
never emitted. The `+100000` uid mapping is the **export's** job (`anonuid=101000`), so the client mount
|
||
carries no uid.
|
||
- **SMB (fallback):** `What=//server/share`, `Type=cifs`,
|
||
`Options=vers=3.0,credentials=<0600 file>,uid=<+100000>,gid=<+100000>,forceuid,forcegid,file_mode=0664,dir_mode=0775,_netdev`
|
||
(plain octal modes, never setgid 2775). **The +100000 rule** (container uid/gid N = host N+100000): a
|
||
container uid 1000 renders `uid=101000` so the guest sees its native id and reads+writes; a naïve `+0`
|
||
lands as `nobody:nogroup` (not writable) — the documented trap, asserted by a companion test.
|
||
- **`.automount` with `TimeoutIdleSec`** (on-demand + idle-unmount): an idle NAS reboot is a non-event.
|
||
- **per-share liveness** (`ListNetworkMounts`): TCP-probes the NAS endpoint (2049/445) + reads
|
||
`/proc/mounts` — it never `stat`s the (possibly EIO/D-state) mountpoint, so a black-holed NAS cannot
|
||
wedge a list. Health `ok | idle | unreachable`, scoped to the affected share, never box-wide.
|
||
- **role gate** `NetworkMountRole`: network storage is **bulk-userdata only** — confined to the
|
||
`/mnt/felhom-drives` namespace; any other target is refused (most-protected).
|
||
- Full validation (`ValidateNetworkMountSpec`) before any unit is rendered: share name (safe segment),
|
||
server, NFS export (absolute, no traversal) / SMB share name, uid/gid range, creds path.
|
||
- **Drive-machinery bypass (Scenario D).** `parseFelhomMountUnit` (the host-reboot drive re-assert's
|
||
classifier) explicitly refuses any unit carrying the network marker, so a NAS mount is never given a
|
||
durable-id, SMART-probed, or re-asserted as a drive. Companion red-proof: the same by-uuid-shaped unit
|
||
with the drive marker DOES parse — the guard is the discriminator, not luck.
|
||
- **`internal/localapi/netstorage.go` (NEW).** Self-scoped endpoints `POST /netstorage/add`,
|
||
`GET /netstorage`, `POST /netstorage/remove`. SMB credentials are written **out-of-band** to a 0600 file
|
||
the agent owns (never in git, never in a plaintext registry, never logged). Role-gated to the user-data
|
||
namespace.
|
||
- **sudoers:** new narrow `FELHOM_NETMOUNT` alias (install/enable/disable/stop the `.automount` + remove the
|
||
felhom mount-unit files; the `.mount` half reuses `FELHOM_MOUNT`, the mountpoint mkdir reuses
|
||
`FELHOM_INTERMEDIARY`). `visudo -cf` clean.
|
||
- **config:** `privileged.smb_creds_dir` (default `/var/lib/felhom-agent/smb-creds`).
|
||
- **Runtime deps:** `mount.nfs` (nfs-common) + `mount.cifs` (cifs-utils) present on the host (confirmed live).
|
||
- Tests: exact NFS/SMB option-set string-asserts + the +100000 companion; validation matrix; role gate;
|
||
unit round-trip + health; the drive-machinery guard + companion; Ensure/Remove command sequences.
|
||
|
||
## v0.49.0 — reboot-during-backup stale-lock recovery (F2-b) + shared-parent script redeploy fix (F2-a) (2026-06-30)
|
||
|
||
Closes the two host-reboot findings from `TESTRUN-fullstack-2026-06-29.md`.
|
||
|
||
- **F2-b — startup stale-lock recovery (`internal/localapi/stalelock.go`, NEW).** A host reboot DURING a
|
||
vzdump backup leaves the guest with a `snapshot-delete`/`backup` lock + a dangling `vzdump` snapshot;
|
||
`onboot:1` then can't start the locked CT → the customer box stays DOWN until a human runs `pct unlock`.
|
||
The agent now self-heals at startup (`Server.RecoverStaleLockedGuests`, called alongside
|
||
`ReassertGuestBinds`/`RecoverFormatJob`): for each guest carrying a backup lock, **only when no vzdump is
|
||
genuinely in-flight** (the load-bearing invariant — at startup the agent's own backup loop hasn't run, so
|
||
the lock is stale; the guard fails SAFE if it can't confirm), it `pct unlock`s → deletes the dangling
|
||
`vzdump` snapshot (API + WaitTask) → starts the CT **iff** `onboot` and not already running. Scope is
|
||
strictly the two vzdump locks; `migrate`/`disk`/`create`/… are left untouched. Idempotent.
|
||
- **`internal/proxmox`:** new reads `GuestConfig.Lock()`/`OnBoot()`, `Client.ListSnapshots`,
|
||
`Client.ListRunningTasks`, and the `Snapshot` type. Reads + snapshot-delete + start go through the API
|
||
token; only `pct unlock` shells out (no API equivalent).
|
||
- **sudoers + capability manifest:** new narrow grant `FELHOM_STALELOCK = /usr/sbin/pct unlock [0-9]*`
|
||
and Critical capability `stalelock-unlock` (a stuck-locked guest = customer box down). `visudo -cf` clean;
|
||
covered by the manifest↔sudoers build gate.
|
||
- Tests: recovery sequence + companions — no-lock touches nothing; non-backup lock left alone; onboot=0
|
||
unlocked-but-not-started; delsnapshot only when a snapshot exists; **invariant guard** (live backup →
|
||
not cleared; unconfirmable → fail-safe); already-running → not restarted.
|
||
|
||
- **F2-a — shared-parent boot script never redeployed (`internal/localapi/intermediary.go`).** The host's
|
||
`/mnt/felhom-drives` was still in root's `shared:1` peer group (so every drive bind DOUBLED) because the
|
||
live boot script predated the v0.36.6 `make-private` fix. Root cause: `EnsureSharedParent` gated the
|
||
(re)install on the **unit** file only, so a script-only change never deployed. Fixed: the new
|
||
`sharedParentInstallStale` helper compares **both** the script and the unit (missing or differing →
|
||
reinstall). Boot-time-only — it rewrites the on-disk script; it does NOT churn the live mount (the live
|
||
bind/make-private/make-shared stays guarded on `!isHostMountpoint`). Verified empirically on the host:
|
||
the correct `bind → make-private → make-shared` sequence gives the parent its own group + no doubling.
|
||
- Tests: a stale-script/current-unit case triggers reinstall (the F2-a regression); both-current is a
|
||
no-op; missing files are stale; a content guard asserts the shipped script keeps `make-private`.
|
||
|
||
- **Live-caught fixes (same version, found during felhom-pve validation):** PVE 9.x rejects
|
||
`GET /nodes/{node}/tasks?running=1` (HTTP 400 "property not defined in schema") — the invariant guard
|
||
now uses `?source=active`. And the unprivileged-LXC start emits a benign `WARNINGS: 1` (systemd-nesting)
|
||
advisory that false-failed the recovery's start — `Start` now uses `AllowWarnings` (matching the
|
||
restore-test's start step).
|
||
- **§D supervised reboot — both findings live-validated.** F2-a: after reboot `/mnt/felhom-drives` came
|
||
up as its OWN peer group (`shared:94`, not `shared:1`) with no doubling; guest sees both drives, apps
|
||
healthy. F2-b: a reboot with the exact stale state (induced `snapshot-delete` lock + a real dangling
|
||
`vzdump` snapshot) reproduced the stuck symptom (pve-guests "CT is locked (snapshot-delete)" → start
|
||
failed) and the agent auto-recovered — unlock → removed the real dangling snapshot → started the CT.
|
||
Zero spurious operator pages on the reboots. See
|
||
`felhom.eu/documentation/audits/TESTRUN-fullstack-2026-06-29.md`.
|
||
- Version `0.48.0 → 0.49.0`.
|
||
|
||
## v0.48.0 — report the served local-API leaf fingerprint (hub-side re-key detection, Part A) (2026-06-29)
|
||
|
||
The agent now rides its **served leaf fingerprint** on every host report so the hub can detect an
|
||
agent re-key fleet-wide (the last self-health leg — `host_leaf_changed`, hub v0.22.0).
|
||
|
||
- **`internal/hub/report.go`:** new `HostReport.LeafFingerprint string` (`leaf_fingerprint`) — the
|
||
SHA-256 of the leaf the agent currently serves. Empty when the local API is disabled (no leaf) → the
|
||
hub treats "" as unknown, never an alert. Not a secret.
|
||
- **`internal/hub/collect.go` + `cmd/felhom-agent/main.go`:** `Collector.SetLeafFingerprint(fp)` threads
|
||
the `fp` from `EnsureLeaf` (the SAME value the loud LOADED/REGENERATED log reports) into every report,
|
||
next to `Capabilities`.
|
||
- Tests: the report includes the fp when set, `""` when unset (local API disabled); golden + contract +
|
||
field-names tests updated (cross-repo golden mirrors `leaf_fingerprint`). Version `0.47.0 → 0.48.0`.
|
||
|
||
## v0.47.0 — controller-swap verify hardening: reject a crash-looping no-healthcheck image (F1) (2026-06-29)
|
||
|
||
Closes F1 from the no-mercy testrun: a controller image with **no HEALTHCHECK that crash-loops** could
|
||
land a single "Running" inspect poll → the swap marked it healthy → **no rollback** (alpine tagged as
|
||
the controller passed in ~4 s, then `Restarting (0)`). The real controller image has a healthcheck so
|
||
the live severity is low, but the rollback safety net had a hole.
|
||
|
||
- **`internal/localapi/controllerswap.go`:** `controllerHealthy` now also reads `{{.RestartCount}}` (a
|
||
4th `docker inspect -f` field) — `running && RestartCount>0` → not-ok (a process that has already
|
||
crash-restarted isn't stably up, regardless of healthcheck). It also signals `needsDwell` for the
|
||
no-healthcheck (`none`) case. `verify` adds a **stability dwell**: a no-healthcheck image must report
|
||
ok on `verifyDwell` (=3) **consecutive** polls before it's accepted; a real `healthy` result is
|
||
trusted immediately (Docker already gated it). Any not-ok resets the dwell. Timeout → existing
|
||
rollback path runs. No change to writeImage, the sudoers grants (the `*` in `docker inspect -f *`
|
||
spans the extended template — confirmed live), or the state-file/rollback orchestration.
|
||
- Tests: F1 **red-proof** (`RestartCount>0` → verify false; companion: rc=0+dwell=1 verifies → the rc
|
||
check is what blocks it); the **dwell** (single ok then crash → verify false; companion dwell=1
|
||
accepts it); a real `healthy` image verifies promptly (no false rollback). Existing
|
||
`RollbackOnUnhealthy` / `HealthyWithNoHealthcheck` stay green. Version `0.46.0 → 0.47.0`.
|
||
|
||
## v0.46.0 — leaf lifecycle: signal + loud-log a regenerated leaf (prevention, Part B.1) (2026-06-29)
|
||
|
||
Makes an accidental local-API leaf **regeneration** (the 2026-06-28 root→non-root migration class —
|
||
moving `/var/lib/felhom-agent` aside silently minted a new leaf → every controller's pin invalidated
|
||
for days) **visible immediately** instead of silent.
|
||
|
||
- **`EnsureLeaf` now returns `generated bool`** (`internal/localapi/cert.go`): false = an existing
|
||
pair was LOADED (stable fingerprint), true = a fresh leaf was GENERATED.
|
||
- **Loud call-site (`cmd/felhom-agent/main.go`):** a load logs `INFO local-api leaf LOADED`; a
|
||
regeneration logs **`WARN local-api leaf REGENERATED — any previously issued bootstrap pins are now
|
||
INVALID; controllers will fail the pin check until re-bootstrapped`** (with the new fingerprint).
|
||
- No new sudo/capability surface — pure return + log change. The companion install-script preservation
|
||
(`--preserve-state-from` + the populated-host guard) lives in `felhom.eu/scripts/felhom-host-install.sh`.
|
||
- Tests: `EnsureLeaf` first call `generated==true`, second `generated==false` AND **same fingerprint**
|
||
(persistence keeps the pin stable). Version `0.45.0 → 0.46.0`.
|
||
|
||
# Changelog
|
||
|
||
All notable changes to **felhom-agent** are recorded here. Update on every code
|
||
change that gets pushed.
|
||
|
||
## v0.45.0 — controller-swap under non-root: stdin `tee` write + narrow sudoers grants (Option A) (2026-06-29)
|
||
|
||
Restores fleet controller-swap / managed auto-update under the **non-root** agent — the one capability
|
||
the 2026-06-29 sudoers audit deliberately left broken because the old write vector needed arbitrary
|
||
in-guest execution. Mechanics spike-proven
|
||
(`felhom.eu/documentation/audits/SPIKE-controllerswap-narrow-grants-2026-06-29.md`, GO). **No controller
|
||
change** — the swap endpoint contract is unchanged; only the agent's internal write mechanism + the
|
||
allowlist.
|
||
|
||
- **`writeImage` no longer shells out.** Was `GuestExec("bash","-c","printf '%s\n' '<img>' > <file>")`
|
||
(the swap's only interpolated/shell vector). Now
|
||
`GuestExecStdin(strings.NewReader(img+"\n"), "tee", "/etc/felhom-controller-image")` — the image ref
|
||
is piped on **stdin** into an in-guest `tee`; no shell, no interpolation. The trailing `\n` keeps the
|
||
on-disk bytes byte-identical to the golden's `printf '%s\n'`, and the bootstrap reads `IMAGE=$(cat …)`
|
||
(newline-stripping), so the write is consumed identically. `ValidControllerImage` still gates upstream.
|
||
- **New stdin seam (no fenced-runner bypass):** `proxmox.Runner.RunStdin` / `ExecRunner.RunStdin` (Run
|
||
with `cmd.Stdin`), `GuestBinder.GuestExecStdin`, and `GuestExecutor.GuestExecStdin` — the swap routes
|
||
stdin through the SAME `sudo -n` fenced runner as every other privileged op.
|
||
- **`FELHOM_CONTROLLERSWAP` sudoers alias (5 narrow, auditable grants):** `cat <fixed file>`,
|
||
`docker image inspect *`, `docker inspect -f *`, `systemctl restart <fixed unit>`,
|
||
`tee <FIXED image file>`. **No general `pct exec`, no `bash -c`** — the spike's negative controls
|
||
(arbitrary exec, `tee` to any other path, `docker rm`, `rm -rf`) stay denied. The 5 are added to the
|
||
v0.44.0 **capability manifest** (Critical — a silently-broken fleet auto-update is operator-alert-worthy),
|
||
so the self-probe watches them and the build-test asserts grant↔code coverage (companion red-proof:
|
||
dropping the `tee` grant fails the gate — demonstrated red→green on the real file).
|
||
- Existing swap tests (happy / rollback-on-unhealthy / image-absent / no-healthcheck / bad-image /
|
||
single-flight) pass over the new write path; a new test asserts the write is stdin-`tee` with exact
|
||
`image\n` and **no** shell vector. Version `0.44.0 → 0.45.0`.
|
||
|
||
## v0.44.0 — privileged-capability self-probe (build-time manifest test + runtime probe + hub snapshot) (2026-06-29)
|
||
|
||
The agent now self-checks the `sudo -n` grants it depends on, so a missing allowlist entry (the
|
||
2026-06-28 cutover class: lxc-info/make-private/…) is caught LOUD — in CI at build time and on the
|
||
host at runtime — instead of surfacing days later as user-visible breakage. **First slice of agent
|
||
self-health; the controller↔agent channel check is a separate later task.**
|
||
|
||
- **`internal/capability` (NEW):** a `Manifest()` of the required `(binary, representative-arg)`
|
||
vectors (seeded from the 2026-06-29 audit — the OK + CLOSED rows; the SURFACED/DEFERRED rows
|
||
`pct exec *`/`pct create`/`mount UUID`/`sensors` are deliberately excluded). `Prober.Probe` lists
|
||
each against the live policy with `sudo -n -l -- <binary> <args>` (a policy LIST — **never
|
||
executes**, safe for mkfs/pct entries) via a DIRECT runner, plus an `os.Stat` existence check,
|
||
mapping to `ok` / `degraded` ("sudo policy denied" | "binary not found"). A total sudo failure
|
||
(drop-in missing) collapses to ONE aggregate signal. Serve-degraded: the probe never blocks
|
||
startup, panics, or errors.
|
||
- **Build-time gate (`manifest_test.go`):** parses `configs/felhom-agent.sudoers`, translates each
|
||
glob to a regex, and asserts **every manifest vector is covered by a grant** — exactly what would
|
||
have caught the dropped `lxc-info`/`make-private` lines in CI. Includes a **red-proof**: with the
|
||
`lxc-info` line removed from an in-memory copy, the check FAILS for `guest-init-pid` (and passes
|
||
on the real file) — proving the gate is not hollow.
|
||
- **Runtime wiring:** `Probe` runs once at startup (INFO `capabilities self-check N/N ok`, plus an
|
||
ERROR per degraded capability naming the gated feature) and on every hub-report cycle; the snapshot
|
||
rides the report as the new non-nil `HostReport.Capabilities []capability.Status` (golden +
|
||
contract test updated; cross-repo hub copy mirrors it).
|
||
- **No allowlist change**; the live host is post-audit complete, so the probe reports N/N ok — itself
|
||
a live proof the probe agrees with the fixed sudoers. Version `0.43.0 → 0.44.0`.
|
||
|
||
## (sudoers completeness audit, folded into v0.44.0) — close non-root allowlist gaps (2026-06-29)
|
||
|
||
A full audit of every privileged command the agent shells via `sudo -n` against
|
||
`configs/felhom-agent.sudoers`, closing the read-only/fixed-vector gaps left by the 2026-06-28
|
||
root→non-root cutover. **Sudoers-only change — no Go change, no version bump** (the file is fetched
|
||
canonically by the host-install script). Root cause of the multi-drive "attach one, the other drops"
|
||
symptom (audit `felhom.eu/documentation/audits/SPIKE-multidrive-mutual-exclusion-2026-06-29.md`): the
|
||
allowlist was incomplete, so several `sudo -n` calls were denied under the non-root user.
|
||
|
||
- **`lxc-info -n [0-9]* -p -H` → FELHOM_INTERMEDIARY (THE root-cause fix).** `guestInitPID`
|
||
(`intermediary.go:256`) shells this to resolve the guest init PID for
|
||
`GuestSeesMount`→`bound_under_parent`. It was absent from the allowlist → `sudo -n` denied → empty
|
||
PID → **every external drive reported absent** → the controller drive-gate stopped each drive's apps
|
||
(flapping). With the grant, `bound_under_parent` reports truthfully and the gate quiesces.
|
||
- **`mount --make-private /mnt/felhom-drives` → FELHOM_INTERMEDIARY.** `EnsureSharedParent`
|
||
(`intermediary.go:110`) calls it to isolate the shared parent's peer group on first setup; the
|
||
allowlist had only `--make-shared`, so the parent stayed in root's peer group and host submounts
|
||
"doubled". Guarded by a mountpoint check (never re-churns a live parent).
|
||
- **`systemctl restart dnsmasq` → FELHOM_DNSMASQ.** The v0.29.x LAN-DNS fix switched `reload`→`restart`
|
||
(`lanresolver.go restartDnsmasq`) but the allowlist still only permitted `reload` → split-horizon
|
||
DNS self-heal was silently denied under non-root. Added alongside the retained `reload`.
|
||
- **`pct set [0-9]* -onboot 1` → FELHOM_PROVISION.** The provision back-half (`backhalf.go`, F3
|
||
auto-start) sets onboot; only `-mp[0-9]*` was allowed → denied under non-root.
|
||
- **`pct reboot [0-9]*` → FELHOM_GUESTHOOK.** `RebootGuest` (`disks.go:448`, the enroll "activate
|
||
pending binds" fallback) was unmatched.
|
||
|
||
**Surfaced for operator decision (NOT added — would require arbitrary root-in-guest):** `GuestExec`'s
|
||
general `pct exec [0-9]* -- <…>` (controller-swap self-update / Phase-2 managed updates) runs variable
|
||
vectors incl. `bash -c "<interpolated>"` — granting it = arbitrary execution. **Controller-swap is
|
||
currently broken under the non-root agent** until narrow per-vector grants are decided. **Deferred:**
|
||
`sensors -j` (defined-but-unwired AND `lm-sensors` not installed on the host — no live caller, path
|
||
unverifiable). **Not added (no daemon caller):** `pct create …` (CreateGoldenLXC, maintenance/broad),
|
||
`mount UUID=… …` (MountUSBByUUID, legacy/unreferenced). Full audit table in `REPORT.md`.
|
||
|
||
## v0.43.0 — canonical systemd unit + binary published to Gitea (BUNDLE slice) (2026-06-28)
|
||
|
||
Day-0 no longer needs a hand-installed agent. The agent binary is now PUBLISHED to Gitea as a generic
|
||
package and the host-bootstrap script fetches → verifies (sha256 vs the hub-vouched manifest) →
|
||
installs it. This commit adds the **canonical systemd unit** (was hand-made per host) and the publish
|
||
tooling; the binary itself is a version-only rebuild (no behavioural change).
|
||
|
||
- **`configs/felhom-agent.service` (NEW, canonical):** `User=felhom-agent`/`Group=felhom-agent` (the
|
||
documented non-root production model — README "Process model"; `privileged.mode: "sudo"` + the
|
||
narrow sudoers allowlist), `ExecStart=/usr/local/bin/felhom-agent --config
|
||
/etc/felhom-agent/agent.json`, `After=network-online.target pve-cluster.service pveproxy.service`,
|
||
`Restart=on-failure`, `StateDirectory=felhom-agent`. **Deliberately NO sandboxing**, with the reasons
|
||
documented inline:
|
||
- `NoNewPrivileges` is NOT set — it would block the setuid `sudo` the agent needs for every host-root
|
||
op (mount/format/pct/dnsmasq), silently killing all privileged capability.
|
||
- NO mount-namespacing hardening (`ProtectHome`/`ProtectSystem`/`PrivateTmp`/…) — any of those give the
|
||
unit a PRIVATE mount namespace, and the intermediary-mount drive model relies on `mount
|
||
--make-shared`/`--bind` propagating into the running guest; in a private namespace every drive
|
||
enrollment would silently break. The agent shares the host mount namespace; the sudoers allowlist is
|
||
the security boundary.
|
||
- **`scripts/publish-agent.sh` (NEW):** builds (optional) + PUTs the binary to
|
||
`/api/packages/admin/generic/felhom-agent/<ver>/felhom-agent` (Gitea generic), prints
|
||
`AGENT_VERSION` + `AGENT_SHA256`, and does a GET round-trip (re-fetch + sha256 re-check) to prove the
|
||
artifact is fetchable + intact. Pinned to a version (never `:latest`); idempotent (delete-then-PUT);
|
||
asserts the binary's `--version` matches the publish version. Creds via `GITEA_USER/GITEA_TOKEN`
|
||
(falls back to `REGISTRY_USER/REGISTRY_TOKEN`).
|
||
- **`configs/build-golden.sh`:** after the vzdump archive is produced, computes its sha256 and PUTs it
|
||
to `/api/packages/admin/generic/felhom-golden/<golden-version>/golden.tar.zst` (`<golden-version>` =
|
||
the baked controller version), printing `GOLDEN_VERSION` + `GOLDEN_SHA256`. Opt-in (only when the
|
||
Gitea creds are set); the local-golden auto-discovery stays as a fallback.
|
||
- **`configs/felhom-agent.sudoers` (latent bug fix):** escaped the commas in the `lvs -o
|
||
lv_name\,data_percent\,metadata_percent` and `lsblk -o NAME\,FSTYPE\,PTTYPE\,MOUNTPOINT` argument
|
||
lists. Sudoers treats a bare comma as a command separator, so `visudo -cf` REJECTED the file — it had
|
||
never been visudo-validated live because the demo host ran the agent as root+`direct` (sudoers
|
||
unused). The escaped commas still match the agent's real comma-bearing args. Surfaced by the BUNDLE
|
||
live install (the host-install script `visudo -cf`-validates before installing).
|
||
- **`cmd/felhom-agent/main.go`:** `version` 0.42.0 → 0.43.0.
|
||
- The operator records the printed agent + golden version+sha256 in the hub (Configs → "Day-0
|
||
artifacts"); the host-bootstrap script verifies fetched artifacts against those before installing.
|
||
- `go build/vet/test ./...` green.
|
||
|
||
## build-golden.sh — default controller image bumped to current; golden rebuilt at 0.85.1 (2026-06-27)
|
||
|
||
**Operational + a default fix (no agent binary change — version stays v0.42.0).**
|
||
|
||
- `configs/build-golden.sh`: the `CONTROLLER_IMAGE` default (positional arg 6) was a stale
|
||
`…/felhom-controller:0.43.0` — an argument-less golden build baked a wildly old controller, so fresh
|
||
Day-0 boxes started old (the demo started at 0.77). Bumped the default to the **current**
|
||
`…/felhom-controller:0.85.1` so the worst case (no explicit arg) is merely "current", not ancient.
|
||
- **Always pass the controller version explicitly at each rebuild** — this default only bounds the
|
||
worst case. A future `make golden` that resolves the latest pullable tag would remove the need for a
|
||
hand-bumped default (Observation, not this task).
|
||
- **Golden rebuilt at 0.85.1** on `felhom-pve` with the image passed **explicitly**
|
||
(`build-golden.sh 9100 … gitea.dooplex.hu/admin/felhom-controller:0.85.1`). New archive volid:
|
||
`local:backup/vzdump-lxc-9100-2026_06_27-11_42_51.tar.zst` (rootfs 32G + Docker-data 16G + user-data
|
||
8G, all in the archive; mp0+mp1 inclusion confirmed in the vzdump log).
|
||
- **Baked-image verify (cheap, mandatory):** in the build guest `/etc/felhom-controller-image` =
|
||
`…:0.85.1` and `docker images` showed it baked (379 MB). New Day-0 provisions now ship current.
|
||
- The host-bootstrap script auto-discovers the newest golden, so it picks up this rebuild
|
||
automatically. The real demo 9201 was **not** re-provisioned (it is the Phase-2 floor test box).
|
||
|
||
## v0.42.0 — agentic controller update: in-guest image swap + rollback (Phase 1) (2026-06-26)
|
||
|
||
The host agent now owns the in-guest controller image **swap** — the new-architecture replacement for
|
||
the controller's dead in-container `docker compose` self-update. The controller pre-pulls the target
|
||
image (shared docker socket, its own registry token) then asks the agent to swap; the agent — external
|
||
to the controller container, so it survives the controller being killed mid-swap — does the rest and
|
||
**rolls back** if the new controller doesn't come up healthy.
|
||
|
||
- **New local-API routes** (`internal/localapi/controllerswap.go`, token-scoped via `withGuest`):
|
||
- `POST /controller/swap {image}` → **202** `{status:"swapping", previous_image, target_image}`, then
|
||
async: record previous (crash-safety state file `/var/lib/felhom-agent/controller-swap-<vmid>.json`)
|
||
→ confirm the target image is present in the guest (else abort, **no swap**) → write
|
||
`/etc/felhom-controller-image` → `systemctl restart felhom-controller-bootstrap.service` → poll the
|
||
new controller to **healthy** (`docker inspect`, ≤90s) → **roll back** to the previous image + restart
|
||
if it doesn't (the guest is never left without a controller). Single-flight per guest (409 if busy).
|
||
Image ref is strict-validated (`gitea.dooplex.hu/admin/felhom-controller:<semver>`) before any action.
|
||
- `GET /controller/swap/status` → `{state: swapping|done|failed, current, previous, target, error}`.
|
||
- **`GuestBinder.GuestExec`** (`internal/localapi/guestbind.go`): the one `pct exec` seam the swap
|
||
composes over (cat/inspect/write/restart), reusing the fenced root runner.
|
||
- **`--selftest=controller-swap -vmid -image <ref>`**: exercise the primitive directly (the target image
|
||
must already be pulled in the guest).
|
||
- Wired `ControllerSwap: guestBinder` into the local-API server (`cmd/felhom-agent/main.go`).
|
||
- Tests (`controllerswap_test.go`): happy swap, **rollback-on-unhealthy** (+ companion red-proof:
|
||
dropping the rollback leaves the guest on the bad image and fails the test), image-absent no-swap,
|
||
no-healthcheck-running, bad-image 400, single-flight 409.
|
||
|
||
## v0.41.0 — provisioned customer guests auto-start after a host reboot (`onboot:1`) (2026-06-24)
|
||
|
||
**F3 fix.** The provision back-half now sets **`onboot:1`** on the customer guest, so after a host
|
||
reboot/power-cut the customer's whole home-server (controller + apps) comes back **on its own** —
|
||
previously every provisioned guest inherited the golden's `--onboot 0` and stayed **stopped** until a
|
||
manual `pct start` (confirmed live in the stable-path/sys-drive restart campaign, Phase 4.1). The new
|
||
step is a fatal `pct set <vmid> -onboot 1` placed right after the config-mount attach (`backhalf.go`),
|
||
mirroring the config-mount/parent-bind `pct set` ops. **No `startup`/boot-order/delay** — the v0.75
|
||
mountpoint-gate already covers the drive-bind race at boot (Phase 4.4), so the controller won't write
|
||
app data onto the rootfs while the agent re-binds drives.
|
||
|
||
The **golden stays `onboot:0`** (`build-golden.sh` unchanged): a template must not auto-start, and
|
||
`onboot` is a per-guest property the back-half is the right place to set. Unit-tested
|
||
(`TestProvision_SetsOnbootOne` asserts the exact `pct set … -onboot 1` invocation, with a red-proof
|
||
against removing the call). The pre-existing demo guest 9201 (provisioned pre-fix) was remediated
|
||
non-destructively with `pct set 9201 -onboot 1`. **Capstone live-validated (2026-06-24):** destroyed +
|
||
re-provisioned 9201 through the real provision chain with v0.41.0 → fresh `pct config` showed `onboot: 1`
|
||
with no manual set; a subsequent **felhom-pve host reboot** brought 9201 back **running with no manual
|
||
`pct start`** (the `onboot:0` scratch guests correctly stayed stopped), controller + base infra healthy,
|
||
drives re-bound at stable, sys_drive separate — the exact Phase-4.1 failure now passes.
|
||
|
||
## v0.40.0 — third CT volume: SSD user-data (`/mnt/sys_drive`, mp1) baked + `-sysdata-grow` (2026-06-23)
|
||
|
||
**The third golden volume.** Extends the OS/Docker-data split (v0.29.x) to a **three-volume layout**:
|
||
rootfs + Docker-data (`mp0`) + **SSD user-data (`mp1` @ `/mnt/sys_drive`, `backup=1`)** — the
|
||
controller's `system_data_path`. Until now `/mnt/sys_drive` was a plain directory on the 32 GB OS
|
||
rootfs, so the controller correctly warned that SSD app data (`<sys_drive>/felhom-data`) lands on the
|
||
OS drive. Baking it as its own thin volume clears that warning with **zero controller change** (the
|
||
controller already auto-discovers `<sys_drive>/felhom-data` and warns via `system.IsMountPoint`); the
|
||
`mp` under the guest's `/mnt` reaches the controller container through the existing
|
||
`-v /mnt:/mnt:rslave` bind.
|
||
|
||
- **`configs/build-golden.sh`** — `pct create` gains
|
||
`--mp1 ${ROOTFS_STORAGE}:${GOLDEN_SYSDATA_GB},mp=/mnt/sys_drive,backup=1` (new env
|
||
`GOLDEN_SYSDATA_GB=8`, near-empty; provision grows it). The resilience guards are mirrored for `mp1`:
|
||
a `findmnt /mnt/sys_drive` separate-mount assertion, and the vzdump-inclusion guard now aborts if
|
||
**either** `mp0` **or** `mp1` is EXCLUDED (the B3 trap — extra mountpoints default `backup=0`). The
|
||
golden does NOT pre-create `felhom-data`; the controller does once it's a real mountpoint.
|
||
- **`internal/reconcile/bringup.go`** — `const DefaultSysDataMount = "mp1"`; `BringUpSpec` gains
|
||
`SysDataGrowGB int` + `SysDataMount string`; a new **"4c"** grow block (online, grow-only `ResizeLXC`,
|
||
its own task) mirrors the "4b" Docker-data grow. `0 = skip` (separateness comes from the golden, not
|
||
the grow — the warning clears regardless of size).
|
||
- **`cmd/felhom-agent/main.go`** — `-sysdata-grow` / `-sysdata-mount` flags (mirror
|
||
`-datavol-grow`/`-datavol-mount`); `bringUpSizing` carries them into all three bring-up/provision call
|
||
sites; `--selftest=provision` help text updated.
|
||
- **Static volume, NOT an enrolled drive.** `/mnt/sys_drive` is part of the baked golden layout; it
|
||
never enrolls/ejects/decommissions and is deliberately kept off the drive-intent machinery.
|
||
`freeMountSlot` auto-skips the baked `mp0`/`mp1` so enrolled drives never collide.
|
||
- Tests: `TestRunBringUp_StorageSplit_SysDataGrow` (asserts `ResizeLXC(vmid,"mp1","+42G")`) +
|
||
`…_SysDataGrowZeroNoResize` (0 → no mp1 resize). RUNBOOK-provisioning-storage.md extended to the
|
||
three-volume layout (default ~512 GB SSD: 32 rootfs + 200 docker-data + 50 user-data).
|
||
|
||
## v0.39.0 — DR recipe completion: live PBS coord + drop the two unfillable drive fields (2026-06-16)
|
||
|
||
**DR-recipe agent-half completion.** A live eyeball of the demo recipe (v0.38.0) found three host-half
|
||
problems; all three are resolved here. No behavior change outside the recipe path.
|
||
|
||
- **PBS coord now resolved LIVE each collect.** New `internal/pbs/live_reporter.go` —
|
||
`LiveSnapshotReporter` implements `hub.PBSReporter` by doing the cheap `Client.Snapshots()` list
|
||
itself, with **last-known-good fallback**, instead of reading only the verify-loop's `SnapshotStore`.
|
||
Previously the recipe's `pbs` block was omitted whenever the store was empty — which a one-shot
|
||
collect (`--selftest=hub`) and the first ~6 h window of every daemon after a restart always saw (the
|
||
verify loop populates the store on its own 6 h cadence). The restore SOURCE must not depend on a
|
||
maintenance cadence. Per-datastore: a live error/timeout → that datastore's last-known-good; a
|
||
successful (even empty) response is authoritative and updates the shared store. Targets-resolution
|
||
failure → the full LKG aggregate. Bounded by `DefaultLiveSnapshotTimeout` (8 s) so a hung PBS never
|
||
stalls the heartbeat. List only — it never triggers a `Verify`. The verify loop keeps Recording into
|
||
the SAME store (shared last-known-good); both use one hoisted `pbsTargets` closure.
|
||
- `SnapshotStore.Get(datastore)` added (per-datastore LKG copy) — the only `SnapshotStore` change.
|
||
- Wired into the collector in BOTH `runDaemon` and `runSelftestHub` (the selftest built its own
|
||
collector with a `nil` reporter — that is why the live `--selftest=hub` showed `pbs_snapshots:[]`).
|
||
- Intended side effect: `report.pbs_snapshots` is now live too (fresher hub PBS view).
|
||
- **`drives[].role` DROPPED from the v1 host-half shape.** A drive's purpose is a hub/operator-owned
|
||
manifest concept, not cleanly derivable host-side (both demo externals are `content=backup`, yet one
|
||
is the primary data drive and the other holds no apps). Deferred until the hub/operator stamps it.
|
||
- **`drives[].restic_repo_coord` DROPPED from the v1 host-half shape.** It named a backup tier that does
|
||
not exist — cross-drive backup is rsync to the SAME internal SSD; there is no offsite/second-failure-
|
||
domain bulk copy. RESERVED for a future tier (see the BACKLOG note in REPORT). v1 drive shape is now
|
||
`{durable_id, mount_path, intent, fs_type?, total_bytes}` — identifiers/intent/size only.
|
||
- The hub reads drives as `json.RawMessage`, so dropping fields needs NO hub struct change — only
|
||
golden + test sync. Cross-repo golden (`host-report.golden.json` here + the hub's copy) re-pinned and
|
||
verified **byte-identical** (sha256 `57f2a5e7…18b2f2b5` — manual checksum-diff discipline): the hub copy
|
||
previously lacked the `dr_recipe` section entirely; it is now a verbatim copy of the agent golden.
|
||
- Tests: new `internal/pbs/live_reporter_test.go` (T1 coord-present-without-prior-verify [load-bearing] +
|
||
inline bare-store companion, T2 error→LKG fallback, T3 success-warms-store, T4 targets-error→aggregate,
|
||
T5 bounded-by-timeout, T6 empty-success-authoritative); `TestDRRecipeHostHalf_V1DriveShape` (drive
|
||
object carries neither `role` nor `restic_repo_coord`); `TestBuildDRRecipeHostHalf` /
|
||
`TestHostReport_ContractMatchesGolden` updated to the v1 drive shape. Each companion was demonstrated
|
||
to FAIL on the pre-fix/mutated code, then reverted (see REPORT).
|
||
|
||
## v0.38.0 — DR recipe: emit the secret-free storage/guest/PBS half in the host-report (2026-06-16)
|
||
|
||
**DR recipe slice (agent half).** Additive `dr_recipe` section on the host-report — the agent half of the
|
||
secret-free reconstruction recipe (`SPIKE-dr-recipe-2026-06-16.md`) that complements escrow (keys) +
|
||
PBS/restic (bytes): the non-secret SCAFFOLDING an operator must rebuild before the PBS bytes can land.
|
||
The hub assembles it with the controller's app half into one customer recipe.
|
||
|
||
- `internal/hub/dr_recipe.go` — `DRRecipeHostHalf{recipe_version, guests[], pbs, drives[], pve_storage[]}`
|
||
built by the pure `BuildDRRecipeHostHalf(guests, targets, pbs)` from facts the report ALREADY collects
|
||
(no new privileged reads): `guests[]` = each guest's sizing (`GuestSpec`, skip status-unknown);
|
||
`drives[]` = the user-data external drives (usb/local-dir with a `uuid:` durable-id + mount path) with
|
||
`{durable_id, role, mount_path, intent, total_bytes}`; `pve_storage[]` = every storage target
|
||
`{name, type, content}` (the `storage.cfg` scaffolding); `pbs` = the latest snapshot's coordinates
|
||
`{repo_id (the pbs storage id), namespace, latest_snapshot_id}`. Wired into `Collect()` after the facts
|
||
are gathered; `HostReport.DRRecipe` (always set, never null).
|
||
- **BOUNDARY (the Phase-1 lesson):** every field is an identifier / intent / size / coordinate — NEVER a
|
||
key, password, token, hash, or `ENC:` value. The PBS encryption key stays in escrow; the access token in
|
||
identity-escrow; the restic password in escrow — the recipe names only the `repo_id`/`namespace`/
|
||
`durable_id`/`restic_repo_coord` the restore TARGETS. `recipe_version=1`; read is ignore-unknown
|
||
(forward-compat). The wire shape is pinned in the cross-repo golden (`host-report.golden.json` here +
|
||
the hub's copy — keep them byte-identical; manual checksum-diff on any change).
|
||
- Tests: `TestBuildDRRecipeHostHalf` (drives = only user-data; pve_storage = all; pbs = latest; guests
|
||
skip nil-spec), `..._NoPBS` (omitted, non-nil slices), `TestDRRecipeHostHalf_NoSecrets` (the lighter
|
||
boundary mirror — serialized half carries NO credential-shaped key; the load-bearing version is on the
|
||
controller emitter), and the `dr_recipe` key-set added to `TestHostReport_ContractMatchesGolden`.
|
||
|
||
## v0.37.0 — host-reboot remount re-resolves enrolled drives by filesystem UUID (2026-06-16)
|
||
|
||
**TASK A — close out the reboot story (agent half).** On a host reboot the kernel can re-enumerate block
|
||
devices and move a drive's node (felhom-usb `/dev/sdb`→`/dev/sdc`), and a `.mount` unit left `disabled`
|
||
by a prior detach never auto-mounts at boot — so an enrolled drive could stay unmounted (or, with any
|
||
node-trusting remount, mount the WRONG device). Root cause pinned LIVE: felhom-usb's systemd mount unit
|
||
was `disabled` (no `multi-user.target.wants` symlink) while felhom-flash's was `enabled`; `What=` was
|
||
already correct (by-UUID), but nothing re-asserted the unit at startup.
|
||
|
||
- `storage.ResolveStorageDevice(durableID)` — resolves the enrolled `uuid:<fs-uuid>` storage scheme to its
|
||
CURRENT backing `/dev` node by re-scanning `/dev/disk/by-uuid` (never a cached node); errors if the UUID
|
||
is genuinely absent so a caller skips a gone drive instead of fail-mounting a stale node.
|
||
- `storage.parseFelhomMountUnit` — pure inverse of `renderMountUnit` (Name/UUID/Where/Type/Options) keyed
|
||
on a `Managed by felhom-agent` marker; ignores any foreign `.mount` unit.
|
||
- `(*SudoHostOps).ReassertEnrolledMounts(ctx)` — at startup (BEFORE binding into the guest) and on the
|
||
periodic 20s tick: for each enrolled `.mount` unit, re-resolve by UUID and re-run `EnsureMount`
|
||
(idempotent `systemctl enable --now`) — re-enables a disabled unit AND mounts the CURRENT device by
|
||
UUID, so a `/dev/sdX` reshuffle is a no-op. Skips ONLY the durable steady state (mounted AND enabled),
|
||
via the pure `shouldReassertMount`; a **mounted-but-DISABLED** unit (the exact live felhom-usb bug — it
|
||
serves now but a reboot would not auto-mount it) is still re-asserted to re-create the wants-symlink.
|
||
Enabled-state is read with a privilege-free `os.Lstat` of the `multi-user.target.wants` symlink
|
||
(`unitEnabled`) — no `systemctl is-enabled` subprocess, no new sudoers entry. An absent UUID is skipped
|
||
(re-asserts on a later tick).
|
||
- Wired in `main.go` ahead of `ReassertGuestBinds` so mounts are live before the guest binds re-assert.
|
||
- Tests (Linux, seam the device-resolution): `TestResolveStorageDevice_ToleratesDeviceLetterMove`
|
||
(UUID symlink moved sdb→sdc → resolves sdc; companion asserts the cached enroll-time node differs from
|
||
the freshly-resolved one — a node-based remount would target the wrong device), `..._AbsentAndScheme`
|
||
(absent UUID errors; only the `uuid:` scheme resolvable), `TestParseFelhomMountUnit` (render→parse
|
||
round-trip + rejects a foreign unit), `TestShouldReassertMount` (the four mounted/enabled combos — pins
|
||
the mounted-but-disabled re-assert), `TestUnitEnabled` (wants-symlink detection).
|
||
|
||
**TASK A2 — verdict: enrolling a NEW drive does NOT need an LXC restart.** The enroll path lands on the
|
||
live intermediary-mount `AttachDrive` (`/disks/guest-attach` → `handleDiskGuestAttach` → `AttachDrive`,
|
||
"no pct, no reboot") under the single shared parent — unbounded named live slots — NOT the legacy
|
||
`RebootGuest` branch. The operator's pre-created-slot-pool idea is therefore unnecessary.
|
||
|
||
## v0.36.7 — isolate the shared parent only on CREATE (no peer-group churn) (2026-06-15)
|
||
|
||
Follow-up to v0.36.6: make-private+make-shared must run ONLY when the self-bind is first created, not on
|
||
every reconcile — re-doing it churns the peer-group id and ORPHANS the guest`s already-established slave
|
||
(propagation silently dies, guest sees empty). Guarded on the mountpoint check; on a fresh boot it runs
|
||
once before pve-guests so the guest slaves the right group.
|
||
|
||
## v0.36.6 — shared parent gets its OWN peer group (make-private first) — ROOT CAUSE of double-bind (2026-06-15)
|
||
|
||
The shared-parent self-bind INHERITED the root mount`s shared peer group (`/mnt/felhom-drives` was
|
||
`shared:1` same as `/`), so every drive bind under it propagated back via the root peer and DOUBLED
|
||
(2 stacked binds per drive — the real cause behind v0.36.3-.5). EnsureSharedParent + the boot script now
|
||
`make-private` (detach from the root group) BEFORE `make-shared` (own group whose only slave is the
|
||
guest), so a drive bind propagates to the guest exactly once.
|
||
|
||
## v0.36.5 — AttachDrive normalizes to exactly one bind (2026-06-15)
|
||
|
||
AttachDrive now COUNTS the binds at a stable path (countHostMounts) and normalizes to exactly one: it is
|
||
a no-op only when there is exactly ONE bind the guest sees; otherwise it strips ALL existing binds
|
||
(bounded loop) and lays down one fresh bind. This converges a stacked double-bind to one — the old
|
||
umount-one+mount-one force-rebind never did. Caught when a double-bind survived a guest reboot.
|
||
|
||
## v0.36.4 — serialize AttachDrive/DetachDrive (no double-bind race) (2026-06-15)
|
||
|
||
A mutex on GuestBinder serializes AttachDrive/DetachDrive so a controller-triggered reconnect and the
|
||
agent`s periodic reconcile can no longer both pass the isHostMountpoint check and double-bind the same
|
||
stable path (a TOCTOU race observed live as 2 stacked binds during rapid eject/reconnect).
|
||
|
||
## v0.36.3 — DetachDrive loop-umounts stacked binds (2026-06-15)
|
||
|
||
DetachDrive now umounts ALL stacked binds at a stable path (bounded loop), not just one layer — so an
|
||
eject/detach fully detaches even if more than one bind accumulated (operator bind on top, or a rare
|
||
attach race), keeping the fail-close intact. Caught in the E13 rapid eject/reconnect sweep.
|
||
|
||
## v0.36.2 — eject also keeps the raw mounted (reconnectable) (2026-06-15)
|
||
|
||
Extends v0.36.1 to EJECT: eject now DetachDrive`s the bind under the parent but LEAVES the raw
|
||
/mnt/<name> mounted (consistent with decommission), so the H1 disconnect→reconnect roundtrip re-binds
|
||
cleanly on a non-removable drive. Physical removal is the separate "remove from system" action. Test:
|
||
eject calls DetachDrive + does NOT unmount the raw.
|
||
|
||
## v0.36.1 — decommission keeps the raw mounted (re-enrollable) (2026-06-15)
|
||
|
||
Fix caught in the E10 acceptance test: the self-serve decommission unmounted the RAW /mnt/<name> host
|
||
mount, which orphaned a non-removable drive (no re-plug) so a one-click re-enroll bound an empty dir. On
|
||
the intermediary model decommission is now a LOGICAL retire — it DetachDrive`s the bind under the parent
|
||
(drive invisible to the guest) but LEAVES the raw mounted, so re-enroll re-binds cleanly. Physical
|
||
removal stays the separate "remove from system" action. Test updated.
|
||
|
||
## v0.36.0 — guest boot-id on /disks (deterministic guest-reboot recreate) (2026-06-15)
|
||
|
||
The agent now emits `guest_boot_id` on GET /disks: `<host-btime>-<guest-init-starttime>` — changes on
|
||
every guest boot (host reboot OR guest reboot) but is STABLE across a controller-only restart. The
|
||
controller persists the last-seen value and DETERMINISTICALLY recreates drive-backed apps when it
|
||
changes (replacing the fragile timed state-sample that could miss an app stopped at the sample instant).
|
||
`GuestBootID` reads `/proc/stat` btime + field 22 of `/proc/<init-pid>/stat` (parsed after the last
|
||
`)` so a comm with spaces/parens does not break it).
|
||
|
||
## v0.35.1 — shared-parent unit: run before pve-guests on host boot (2026-06-15)
|
||
|
||
Fix for the host-reboot ordering (the shared-parent oneshot never ran before pve-guests on the live
|
||
host, so the guest bound a not-yet-shared parent → private bind → propagation broken). The unit now uses
|
||
`WantedBy=pve-guests.service` (pve-guests PULLS IT IN + Before= orders it first) instead of the
|
||
unreliable `WantedBy=multi-user.target`, and drops `DefaultDependencies=no`. `EnsureSharedParent`
|
||
reinstalls the unit when its content differs (so the fix deploys on the next agent start/reconcile).
|
||
|
||
## v0.35.0 — intermediary mount: guest-reboot re-propagation (load-bearing) (2026-06-15)
|
||
|
||
Fix for the guest-reboot gap (caught in the live demo migration). A guest's parent bind is
|
||
NON-RECURSIVE, so on a guest reboot it does NOT carry the pre-existing drive submount, and mount
|
||
propagation only delivers events created AFTER the bind exists — so an enrolled drive is bound on the
|
||
HOST but INVISIBLE in the fresh guest namespace until re-bound. Without this, every guest reboot left
|
||
the apps on empty dirs.
|
||
|
||
- `AttachDrive` now takes `vmid` and checks GUEST visibility (`GuestSeesMount`, reading
|
||
`/proc/<guest-init-pid>/mountinfo`): if the host has the bind but the guest doesn't see it
|
||
(post-reboot), it FORCE re-binds (umount + mount) to re-fire propagation into the current guest ns.
|
||
- A periodic reconcile (20s ticker in main) re-runs `ReassertGuestBinds`, so a guest reboot self-heals
|
||
without an agent restart. `EnsureSharedParent` skips the unit re-install when already present (cheap
|
||
on repeat).
|
||
- `/disks` `BoundUnderParent` now reflects GUEST visibility (not the host mount) — the accurate signal
|
||
the controller's drive-absent gate keys on to stop/restart apps across a guest reboot.
|
||
|
||
## v0.34.0 — intermediary mount model: shared-parent + host-side attach/detach + reconcile (2026-06-15)
|
||
|
||
The drive hot-swap re-architecture (SPIKE-intermediary-mount). Replaces the per-drive `pct set -mpN`
|
||
bind (which needed a guest reboot to activate and bricked the guest when a drive was absent at boot)
|
||
with a SINGLE permanent parent bind `/mnt/felhom-drives` plus host-side swaps underneath it.
|
||
|
||
- `internal/localapi/intermediary.go` — `GuestBinder.EnsureSharedParent` (mkdir + self-bind +
|
||
`--make-shared` + installs/enables a `felhom-shared-parent.service` ordered **Before=pve-guests** so
|
||
the guest's parent bind inherits the shared peer group as `slave`); `AttachDrive` (`mount --bind
|
||
/mnt/<name>/felhom-data /mnt/felhom-drives/<name>` — propagates into the RUNNING guest live, no pct,
|
||
no reboot; confined to felhom-data; the stable dir stays host-root-owned = fail-closed); `DetachDrive`
|
||
(`umount`, leaving the bare fail-closed dir); `StablePathForRaw`/`DriveNameFromRaw`; `isHostMountpoint`.
|
||
- `ReassertGuestBinds` is now a pure HOST-SIDE reconcile: for each enrolled+present drive ensure its
|
||
felhom-data is bound under the parent (no guest-config read, no slot, no reboot) — fixes F9 and
|
||
drive-reconnect for free. Runs at startup (ensures the shared parent first).
|
||
- `handleDiskGuestAttach` uses `AttachDrive` (returns the stable `guest_path`); eject + decommission
|
||
call `DetachDrive`. Legacy `AttachBind`/`DetachBind` retained for the transition (decommission still
|
||
`--delete`s any lingering legacy mp).
|
||
- `/disks` reporting adds `GuestPath` (the stable `/mnt/felhom-drives/<name>` the controller repoints
|
||
HDD_PATH to) and `BoundUnderParent` (live-in-guest signal for the controller's drive-absent gate).
|
||
- Provision adds the one permanent parent bind (`-mp8 /mnt/felhom-drives,mp=/mnt/felhom-drives`).
|
||
|
||
Tests (non-hollow + companions): `TestGuestAttach_BindsUnderParent` (uses AttachDrive not legacy pct),
|
||
`TestReassertGuestBinds_RestoresMissingBind` (host-side reconcile, legacy AttachBind never called),
|
||
`TestStablePathForRaw_DriveName`, `TestDisks_GuestPathAndBoundUnderParent`. Sudoers: new
|
||
`FELHOM_INTERMEDIARY` alias (mount/umount under /mnt/felhom-drives, the unit install, the parent bind).
|
||
|
||
## v0.33.0 — C1 net: pre-start self-heal hook + decommission mp-delete (2026-06-15)
|
||
|
||
The transitional defense for the C1 brick (B3 critical bug) ahead of the intermediary-mount
|
||
re-architecture (which makes C1 structural). Two independent nets:
|
||
|
||
- **Pre-start self-heal hook** (`internal/guesthook`): a PVE `pre-start` hookscript runs
|
||
`felhom-agent guest-hook <vmid> <phase>` which, for every BIND mountpoint whose source path is
|
||
missing, creates an empty **host-root-owned** placeholder dir so the bind succeeds and the guest
|
||
always boots — fail-closed (host uid 0 is unmapped in the unprivileged-LXC userns, so the guest
|
||
can't write to the placeholder; a returning drive shadows it). It CREATES rather than DELETEs
|
||
because `pct set --delete` in pre-start would take the config lock the start task already holds
|
||
(dead-times-out → still bricks); the heal logic is in unit-tested Go, the wrapper just delegates.
|
||
Installed + registered per-guest by the provision back-half (`InstallSnippet`/`Register`).
|
||
- **Decommission mp-delete** (`GuestBinder.DetachBind` + `handleDiskDecommission`): decommission now
|
||
runs `pct set <vmid> --delete mpN` on the slot binding the drive (lock-safe on the running guest),
|
||
so its now-missing source can't brick the next reboot. The old handler unmounted but left the dead
|
||
`mpN` in config — the exact B3 C1 bug. Eject keeps its mp (temporary; the hook covers a
|
||
reboot-while-ejected).
|
||
|
||
Tests (non-hollow, each with a companion that fails the pre-fix/trivial impl):
|
||
`internal/guesthook/heal_test.go` (selector ignores storage volumes + present binds, heals only the
|
||
absent one; "return nothing"/"return all" both fail) and `TestDecommission_DeletesGuestMount`
|
||
(asserts the correct slot is `--delete`d; pre-fix never calls DetachBind → fails).
|
||
|
||
Sudoers: new `FELHOM_GUESTHOOK` alias (snippet install, `pct set --hookscript`, `pct set --delete mpN`).
|
||
|
||
## v0.32.0 — self-serve decommission + intent-aware re-assert (B2a) (2026-06-14)
|
||
|
||
Customer-self-serve storage decommission (no operator signature; non-destructive — never formats),
|
||
plus the load-bearing fix that keeps a decommissioned drive from auto-rebinding into the guest.
|
||
|
||
- **`POST /disks/decommission`** (`internal/localapi/disks.go` `handleDiskDecommission`, route in
|
||
`server.go`) — mirrors `handleDiskEject` exactly: `withGuest` self-scoping, `scopedFromBody`, and the
|
||
same **user-data role gate** (`roleForMountPath` must be `RoleUserData`, else 403; fail-safe-to-
|
||
protected on ambiguity) so a compromised controller can't decommission system/backup storage. It
|
||
records a PERMANENT `IntentDecommissioned`, prunes the `GuestBindStore` entry (hygiene), and unmounts
|
||
(so the drive is physically removable). It **NEVER** calls any format/mkfs path — the data stays on
|
||
the drive. The operator-signed `DecommissionExecutor` + `reconcile.Classify` classification are
|
||
untouched (the absent-drive/DR route).
|
||
- **`ReassertGuestBinds` is now intent-aware** (THE correctness fix): the startup re-assert skips any
|
||
durable-id whose intent is not `enrolled`, so a decommissioned- (or ejected-) but-still-present drive
|
||
is never auto-rebound into the guest on agent restart. A nil intent store falls back to legacy
|
||
bind-all (matching the watchdog's nil-intent rule). Covers both the self-serve and the operator-
|
||
signed decommission paths (both land on `IntentDecommissioned`).
|
||
- **`GuestBindStore.Remove(vmid, durableID)`** (`internal/localapi/guestbindstore.go`) — idempotent
|
||
(absent = no-op), atomic tmp+rename like `Record`; drops the vmid key when its set empties. Re-enroll
|
||
re-`Record`s via the existing `recordGuestBind` on guest-attach, so Remove doesn't break re-commission.
|
||
- `IntentRecorder` extended with `SetDecommissioned` + `Get` (both already on `*storage.IntentStore`).
|
||
- Non-hollow tests (`internal/localapi/decommission_test.go`): role-gate refuses system/backup (403,
|
||
no unmount); decommission sets intent + removes the bind + unmounts + never formats; intent-aware
|
||
re-assert does NOT rebind a decommissioned-but-present drive (companion: enrolled DOES rebind; the
|
||
intent-blind pre-fix code fails this); re-commission re-records; `Remove` idempotency + persistence.
|
||
|
||
## v0.31.0 — live-drive F9 + F20-BUG2 + F20-BUG3 (disk bind/wipe) (2026-06-14)
|
||
|
||
The last live-drive findings, all disk/`localapi`-side, implemented + deployed on `felhom-pve` and
|
||
validated live on guest 9201 (approach: attach-to-existing, no re-provision — see the audit fixspec).
|
||
|
||
- **F9 — guest data-drive bind survives a re-provision** (`4cd1d02`). The in-guest bind (`pct set -mpN`)
|
||
is config state a destroy+re-provision drops, and nothing restored it → a re-provisioned guest came up
|
||
with its enrolled HDD unattached. New `GuestBindStore` (durable-id-keyed, per guest, recorded at
|
||
guest-attach) + `ReassertGuestBinds` on agent startup re-adds any bind a guest is missing — only when
|
||
the durable-id still resolves to a present drive (a swapped/absent disk is never auto-bound), idempotent.
|
||
Plus `DiskInfo.GuestAttached` — the missing "bound into THIS guest" signal (vs mere host presence;
|
||
resolves the F2 `hdd_configured` disagreement). **Live-proven:** dropped the bind, restarted the agent
|
||
(real trigger) → re-attached with no manual call; reboot activated it; an HDD app then deployed onto
|
||
the drive with data on `/dev/sdb1`.
|
||
- **F20-BUG2 — one wipe durable-id scheme** (`a2a76e7`). `/disks` advertised only `durable_id` (`uuid:`,
|
||
used for assign), but the wipe gate resolves `byid:`/`byuuid:` → confirming a wipe with the advertised
|
||
id was a `binding_mismatch`. New `DiskInfo.WipeDurableID` via a shared `s.deviceDurableID` seam used by
|
||
BOTH the list and the gate, so the id the customer copies is the id the gate accepts. **Live-proven:** a
|
||
confirmed wipe using `/api/disks`'s `wipe_durable_id` is accepted (no mismatch).
|
||
- **F20-BUG3 — format runs detached; survives a request deadline AND an agent restart** (`4777f8a`). mkfs
|
||
ran under the HTTP request context, so a client deadline SIGKILLed it mid-write → corrupt disk. Now mkfs
|
||
runs off `s.baseCtx` via a persisted `formatJob` record; the handler still returns the synchronous
|
||
result (backward-compatible) but a dropped request no longer kills it. New `GET /disks/format/status`;
|
||
`RecoverFormatJob` on startup re-runs an interrupted durable-id-bound format (re-resolved; anti-retarget
|
||
— a blank/path-bound or unresolvable job is not auto-re-run). **Live-proven on the 916 GB felhom-usb:** a
|
||
2 s client timeout left a ~30 s mkfs running to a clean ext4 (the live-drive corruption is gone); an
|
||
agent restart mid-format was recovered + completed to a clean fs.
|
||
|
||
|
||
|
||
**Security fix (from the 2026-06-13 deep-sweep audit).** The inline customer-confirmed wipe in
|
||
`internal/localapi/disks.go` `handleDiskFormat` inspected and gate-bound the device by its durable id
|
||
but then ran `mkfs` on the caller-supplied mutable `/dev` path (`req.Device`). A USB re-enumeration
|
||
reassigning that `/dev` node to a different physical disk between inspection and `mkfs` (a
|
||
classify→mkfs TOCTOU) could wipe the wrong drive.
|
||
|
||
- New `internal/localapi/wipe_reresolve.go`: `antiRetargetResolve` (injected-deps, unit-tested) mirrors
|
||
`signedjobs.WipeExecutor.Execute` — resolve the confirmed durable id → current device, re-derive the
|
||
device's durable id and require an exact match, re-inspect (still data-bearing), and return the
|
||
re-resolved device. `(*Server).reresolveDurableForWipe` wires the real storage funcs.
|
||
- `handleDiskFormat` now formats the **re-resolved** device, never `req.Device`; any refusal →
|
||
`409 Conflict`, no `mkfs`. Injectable `reresolveWipe` seam on `Server` (defaults to the real path).
|
||
- Tests: `wipe_reresolve_test.go` covers happy-path, empty/gone/blank, re-inspect-error, and the core
|
||
`retarget-mismatch-refused` case. Round-trip safe for legitimate wipes (`DeviceDurableID` ↔
|
||
`ResolveDurableDevice` schemes match). Agent-only deploy; no golden rebake. See `AGENT-001-FIX-NOTES.md`.
|
||
|
||
## v0.29.1 — lanresolver: RESTART dnsmasq on change (not reload) — fixes stale split-horizon IP (2026-06-13)
|
||
|
||
**Bug:** after a guest's DHCP IP moved (e.g. the v0.29.0 9201 re-provision: .151 → .141), the LAN
|
||
split-horizon resolver kept answering the OLD IP, so LAN clients (via Pi-hole's conditional forward to
|
||
the host dnsmasq) resolved `*.demo-felhom.eu` to the dead IP. Root cause: `lanresolver.Manager` updated
|
||
the per-customer drop-in (`address=/<domain>/<ip>`) correctly but then ran `systemctl reload dnsmasq`
|
||
(SIGHUP) — and **dnsmasq's SIGHUP does NOT re-read its config files** (`/etc/dnsmasq.d/*.conf`); it only
|
||
clears the cache + re-reads `/etc/hosts`/addn-hosts. So the changed `address=` directive never took
|
||
effect until a restart. **Fix:** `reload()` → `restartDnsmasq()` (`systemctl restart dnsmasq`) for every
|
||
config-drop-in change (ReconcileGuest IP change, EnsureDnsmasq base change, Remove/decommission). Restart
|
||
is sub-second and the records carry local-ttl 0, so downstream forwarders don't cache a stale answer.
|
||
(Live: after the fix + a one-time host dnsmasq restart + a Pi-hole cache flush, `*.demo-felhom.eu`
|
||
resolves to the live guest IP again; future IP moves now self-heal on the loop's next tick.)
|
||
|
||
## v0.29.0 — OS / Docker-data storage split: golden + provision (2026-06-13)
|
||
|
||
Phase 1 of the storage-split slice (Phase 2 = felhom-controller v0.58.0 prevention layer). The
|
||
controller guest's OS rootfs and Docker data are carved onto separate `local-lvm` volumes for
|
||
RESILIENCE — an isolated OS rootfs stays bootable + agent-recoverable if the Docker volume fills.
|
||
|
||
- **`configs/build-golden.sh` — split baked in:** `--rootfs ${ROOTFS_STORAGE}:${OS_SIZE_GB}` (default
|
||
**32**, was hardcoded 8) **plus** `--mp0 ${ROOTFS_STORAGE}:${GOLDEN_DOCKER_GB},mp=/var/lib/docker,backup=1`
|
||
(default 16). The baked controller + infra images land on the data volume and travel inside the
|
||
golden archive (no empty-volume shadowing, no deploy-time pull). `backup=1` is MANDATORY — extra LXC
|
||
mountpoints default to `backup=0` = EXCLUDED from vzdump (spike B3), which would drop the images from
|
||
the archive entirely. The script now also bakes Docker **log rotation** into `daemon.json`
|
||
(`max-size 10m`, `max-file 3` — prevention layer 2D), asserts `/var/lib/docker` is a separate mount,
|
||
and **aborts if vzdump excludes mp0**.
|
||
- **`internal/reconcile/bringup.go` — sized provision:** `GuestMount` gains `Backup` (emits `,backup=1`
|
||
— closes the spike-B3/B5 silent-DB-loss trap at the mount builder). `BringUpSpec` gains
|
||
`DataVolGrowGB` + `DataVolMount` (default `mp0`): provision GROWS the golden-carried Docker-data
|
||
volume online to the per-customer target (grow-only, spike B4) rather than attaching a fresh empty
|
||
volume that would shadow the baked images. Plus `RootfsGrowGB` for the OS rootfs.
|
||
- **CLI seam:** `--selftest=bring-up|provision` gain `-rootfs-grow` / `-datavol-grow` / `-datavol-mount`
|
||
flags. Per-customer sizing source = flags now, the slice-10 hub storage manifest later.
|
||
- **`RUNBOOK-provisioning-storage.md`** (new): the split provisioning procedure + fresh-PVE-install
|
||
thin-pool carving knobs (`hdsize`/`maxroot`/`maxvz`, spike B4) + the per-customer sizing seam.
|
||
- Tests: `buildBringUpConfig` backup=1 emission; bring-up issues rootfs + data-volume resizes.
|
||
|
||
## (no version) — storage OS/data-split spike findings (2026-06-13)
|
||
|
||
Investigation only — **no code changed**. Findings report: `REPORT-storage-split-spike.md` (gates the
|
||
provisioning spec for splitting the controller guest's OS rootfs from its Docker/data onto separate
|
||
`local-lvm` volumes). Proven on a throwaway unprivileged LXC (9300, since destroyed): Docker `data-root`
|
||
on a second `local-lvm` mountpoint works (overlayfs/ext4, no idmap issue, reboot-survives); the
|
||
move-then-verify migration is safe (copy-not-move). **Key finding:** additional LXC mountpoints are
|
||
**excluded from vzdump by default** — they need `backup=1` set **and a CT restart** — so the docker-data
|
||
mount must be attached with `,backup=1` or named-volume DBs silently fall out of PBS. The exact seam is
|
||
`internal/reconcile/bringup.go:313` (`buildConfigParams`), which today builds `mpN` without a `backup=`
|
||
flag; `GuestMount` should carry the flag. Per-customer sizes belong in the slice-10 hub storage manifest
|
||
(marked at `bringup.go:49-50`); the golden rootfs is hardcoded `8` at `configs/build-golden.sh:40`.
|
||
|
||
## v0.28.0 — backup re-target → felhom-pbs (offsite DR) + operator-signed decommission (2026-06-12)
|
||
|
||
**Whole-guest backup now defaults to the offsite PBS tier (real DR).** `BackupConfig.BackupTarget()`
|
||
returns the configured `backup.local_backup_target` or, when empty, the new default `felhom-pbs` — a
|
||
PBS datastore on SEPARATE HARDWARE (the DooPlex box), so a host disk/hardware failure no longer takes
|
||
the backups with it. The target stays fully configurable (set `local_backup_target` to `local`/other
|
||
to override); no call site hardcodes it. All `NewBackupRunner` sites (restore-test scheduler, local-API,
|
||
`--selftest=backup`/`restore-test`) route through `BackupTarget()`.
|
||
|
||
Proven live on demo-felhom before the re-point (PHASE 0 gate):
|
||
- snapshot-mode `vzdump → felhom-pbs` still fires the `create storage snapshot 'vzdump'` marker, so the
|
||
8B.2 early-resume/quiesce signal survives a PBS target (the marker is mode-driven, not target-driven);
|
||
- the restore-test enumerates PBS backups through the SAME generic `StorageContent`
|
||
(`/nodes/<node>/storage/felhom-pbs/content` returns `content:"backup"` + ctime/vmid/volid), so
|
||
`PickRestoreCandidate`/`latestArchive` need NO PBS-client change;
|
||
- `pct restore` from a PBS volid round-trips cleanly (storage.cfg encryption key applied transparently);
|
||
- PBS gotchas (`ignore-verified`, node-from-UPID, privsep) touch only the verify-API path, not vzdump/restore.
|
||
|
||
**Operator-signed `decommission` now reachable (slice 10 P3 completion).** The previously-unreachable
|
||
`IntentDecommissioned` state (no production caller) is now reached ONLY via a gate-VERIFIED operator
|
||
signature — never customer-confirmable, distinct from a safe eject. New `internal/signedjobs`
|
||
`DecommissionExecutor` (op `decommission`, classified destructive in `reconcile.Classify`) calls
|
||
`IntentStore.SetDecommissioned`, keyed by the drive's STORAGE durable-id (the watchdog's key, e.g.
|
||
`uuid:<fs-uuid>` — NOT the device-level `byid:/byuuid:` scheme `storage_wipe` uses), so the recorded
|
||
intent actually gates future remounts. New `ExecutorChain` lets the signed-jobs runner serve both
|
||
`storage_wipe` and `decommission`; the runner wiring moved below the intent-store open in `main.go`.
|
||
`felhom-opsign` builds decommission params from `-durable-id`. No controller/customer UI — the operator
|
||
path is hub jobs-queue → signed-jobs runner.
|
||
|
||
**Restore-test now boot-verifies slice-10 enrolled guests (bind-mount mountpoints).** A guest whose
|
||
data drive is a host BIND mount (slice-10 P2 `mp0`) could not be vzrestore'd by the privsep token
|
||
("restoring 'mpN' to bind mount is only possible for root") — so the restore-test failed for every
|
||
enrolled guest, regardless of backup tier (surfaced during the felhom-pbs live validation). The
|
||
restore-test now reads the SOURCE guest config (vmid parsed from the archive volid — PBS `ct/<vmid>/`
|
||
and vzdump `vzdump-lxc-<vmid>-` forms) and passes `RestoreLXCOptions.MountOverrides` that neutralize
|
||
each bind-mount `mpN` to a throwaway 1G volume on the restore storage (needs no root; the boot-verify
|
||
doesn't need the drive's data, and the host paths would otherwise collide). Storage-backed mountpoints
|
||
are restored normally; best-effort (an unreadable source config restores as-is). `proxmox.RestoreLXC`
|
||
gained `MountOverrides`. Verified live: restore-test from felhom-pbs of bind-mounted guest 9201 →
|
||
boot+running PASS.
|
||
|
||
## v0.27.0 — slice 10 P3: self-heal watchdog reconcile + 4-state intent model (2026-06-12)
|
||
|
||
The storage watchdog goes from detect-only → detect-and-reconcile: the agent autonomously re-mounts an
|
||
enrolled external drive that dropped out-of-band (the colleague's Proxmox unmount), gated by a persisted
|
||
INTENT model so it never auto-adopts an unknown drive or fights an official eject.
|
||
|
||
- **`internal/storage/intent.go` — `IntentStore`** — durable, **durable-id-keyed** (UUID/WWN, never
|
||
sdX/path), atomic-write 4-state model: `new` (not recorded → never auto-mount), `enrolled` (desired
|
||
mounted → reconcile drift), `ejected` (intentional unmount → leave alone), `decommissioned`
|
||
(permanent). `OnAbsent` clears `ejected`→`enrolled` so a replug auto-mounts (the replug rule).
|
||
Records intent ONLY through the official enroll/eject paths — an out-of-band unmount records nothing
|
||
and is healed. Tests cover the states, persistence, the replug rule, and the reconcile gate.
|
||
- **`watchdog.go` — intent-gated reconcile + flapping guard (3C)** — the re-mount candidate (device
|
||
present, not mounted) now fires ONLY for an `enrolled` drive (via `IntentReader`); a present→absent
|
||
transition (device gone) calls `OnAbsent`. Exponential backoff (`debounce·2^fails`) + an alert after
|
||
4 failed cycles + a hard stop after 8 (no infinite loop). Failure = "still not present a full backoff
|
||
window after we dispatched" (a slow async re-mount isn't miscounted). Tests: colleague-unmount→
|
||
reconciled; ejected/new/decommissioned→left alone; ejected→absent→replug→auto-mount; flapping→caps.
|
||
- **`internal/localapi`** — `POST /disks/guest-attach` records `enrolled`; `POST /disks/eject` records
|
||
`ejected` (BEFORE unmount, while the durable-id still resolves) via the new `IntentRecorder`. `main.go`
|
||
opens one `IntentStore` (`<StateDir>/drive-intents.json`) shared by the watchdog + local API; open
|
||
failure degrades to ungated legacy remount (logged).
|
||
|
||
## v0.26.0 — slice 10 P2 activation: guest-reboot endpoint (user-triggered drive activation) (2026-06-12)
|
||
|
||
A drive enrolled into a RUNNING unprivileged guest can't be live-activated (proven: `pct set` won't
|
||
hot-apply; `/proc/<pid>/root` bind → mount-locking refusal; `nsenter -m` loses the host source). So the
|
||
bind activates at the next guest boot. This adds the user-triggered restart path.
|
||
|
||
- **`POST /guest/reboot` (`internal/localapi`)** — self-scoped (vmid from token). Runs `pct reboot
|
||
<vmid>` **detached** (it blocks ~30s until the guest is back) and returns **202** immediately, so the
|
||
calling controller gets a clean response before the reboot takes it down (the agent is host-side and
|
||
survives). `GuestBinder.RebootGuest` over the fenced runner. Tests: `TestGuestReboot_Accepted`
|
||
(202 + RebootGuest invoked for the token's vmid), `TestGuestReboot_CrossGuest403` (body vmid mismatch
|
||
refused, no reboot). Pairs with controller v0.49.0 (pending-activation detection + "Újraindítás most").
|
||
|
||
## v0.25.0 — slice 10 P2: bind enrolled user-data drives into the guest (passthrough) (2026-06-12)
|
||
|
||
External user-data drives are mounted on the HOST but were never passed INTO the guest (diagnosed
|
||
Branch A), so apps silently wrote to the rootfs and the controller couldn't see them. This adds the
|
||
guest passthrough. Spike-proven on 9201 first (see REPORT / the usb-passthrough-spike findings):
|
||
`pct set` **bind form** (host path, never `storage:size`), `chown` to the guest base (idmap not clean
|
||
for mixed-ownership data), `shared:49` propagation host↔guest automatic.
|
||
|
||
- **`POST /disks/guest-attach` (`internal/localapi`)** — self-scoped (vmid from token). Binds an
|
||
enrolled drive's **felhom-data namespace** into the guest at `/mnt/<name>` (**Model A**: the
|
||
felhom-data dir is the bind source mounted AT `/mnt/<name>`, so only Felhom's namespace crosses into
|
||
the guest — the customer's other data on the drive never does). Idempotent (returns the existing slot
|
||
if already bound); picks the lowest free `mpN`; validates `where` is `/mnt/<name>` (no traversal).
|
||
- **`GuestBinder` (`internal/localapi/guestbind.go`)** — the host-root steps over the fenced
|
||
`proxmox.Runner` (same pattern as the provision back-half's bind): `mkdir -p <drive>/felhom-data` →
|
||
`chown 100000:100000` the namespace ROOT (not -R; per-app subdirs are chowned at deploy) → `pct set
|
||
<vmid> -mpN <drive>/felhom-data,mp=/mnt/<name>` (RW bind). The namespace is created fresh + uniformly
|
||
owned, which sidesteps the drive's pre-existing mixed-ownership data entirely.
|
||
- **Tests** — `TestGuestAttach_*`: free-slot selection (mp0 when mp9 taken), idempotency (no re-bind +
|
||
`already:true`), bad-path rejection (traversal/non-/mnt/multi-component), not-configured 503.
|
||
|
||
Pairs with felhom-controller P2C (enroll triggers attach) + the golden's `/mnt:rslave` controller bind
|
||
(P2B). Self-heal reconcile (P3) and dual-role (P4) follow.
|
||
|
||
## v0.24.0 — role-gate the eject path (system/backup mounts are unmount-protected at the agent) (2026-06-12)
|
||
|
||
Closes the eject gap in the storage-authorization redesign: `POST /disks/eject` now **refuses to
|
||
unmount a system or backup storage**, enforced at the agent — not just hidden in the controller UI.
|
||
A direct API call (or a compromised controller) trying to `eject {where:"/var/lib/vz"}` or the PBS
|
||
mount is refused 403; only `user-data` mounts are ejectable.
|
||
|
||
- **`handleDiskEject` (`internal/localapi/disks.go`)** — before `Unmount`, resolves the AUTHORITATIVE
|
||
protection role of the storage mounted at `where` (the agent's own storage-view + host-topology
|
||
classification, never the caller's claim) via the new `roleForMountPath`. Refuses (403, no
|
||
`Unmount`) unless the role is `user-data`. **Fails SAFE**: an unresolvable mount (view error or no
|
||
storage target at that path) → treated as protected → refused (the same most-protected-on-ambiguity
|
||
default the wipe gate uses). Mirrors the wipe path's "protected — eject refused by role" logging.
|
||
- **`roleForMountPath` + `hostReader` seam** — `roleForMountPath` keys `RoleForStorage` on the mount
|
||
path (the eject input), mirroring `deviceRole`. `Options.HostReader` (optional; defaults to the
|
||
production `*storage.ProcHostReader`) injects the root-free topology reader so the role-gate is unit-
|
||
testable. `handleDisks`/`deviceRole` now share the same seam.
|
||
- **Tests** — `TestEject_RoleGated` asserts a `system` and a `backup` mount are refused with **no
|
||
`Unmount`**, a `user-data` mount ejects, and an unresolvable mount fails safe to refused (the same
|
||
non-hollowness the wipe tests use). `TestEject_UnmountAndDependents` updated to a user-data target.
|
||
|
||
## v0.23.0 — device-ROLE classification + tiered storage-wipe gate (system/backup operator-only, user-data customer-confirmable) (2026-06-11)
|
||
|
||
The storage-authorization redesign (agent half). The gate's destructive-wipe path is now **tiered by
|
||
the device's protection ROLE**, which the agent classifies from its OWN inspection — never the
|
||
caller's claim (the storage analog of classify.go's data-bearing verdict).
|
||
|
||
- **`internal/storage/role.go`** — `DeviceRole` (`system` | `backup` | `user-data`) + the
|
||
authoritative classifier. `RoleForStorage` (storage-view targets) and `RoleForRawDevice` (a raw
|
||
device, e.g. a fresh disk in the init flow) map a device to its tier via `SystemDisks` (the
|
||
whole-disks backing `/`, `/boot`, `/boot/efi`, root-free reads). Rules: `pbs` → backup; `lvmthin` /
|
||
builtin `local` / nfs / cifs / unknown → system; `usb` / `local-dir` on a **non-system external
|
||
device** → user-data. **Fail-safe**: any ambiguity (system disks unknown, or an unrecognizable
|
||
device topology) → **system** (most-protected) — never silently user-data.
|
||
- **`GET /disks`** — each `DiskInfo` now carries `role`. The controller drives the UI from it
|
||
(system/backup get a lock + no destructive controls; user-data is customer-manageable).
|
||
- **Gate tier (`reconcile`)** — new `CustomerConfirmable` disposition + `Gate.AuthorizeStorageWipe`:
|
||
- role=**user-data** → **customer-confirmable**: allowed iff the request carries an explicit
|
||
customer confirmation **bound to the device's durable id** (the agent re-resolves the durable id
|
||
and matches; a confirmation for one disk can't wipe another). **No operator signature.** A
|
||
user-data drive is already within the in-guest controller's blast radius (it bind-mounts `/mnt`),
|
||
so customer-confirmation adds no new reach. Recorded in the **audit log** with the durable id
|
||
(`AuditRecord.DurableID`).
|
||
- role=**system**/**backup** → unchanged **operator-signature** (`pending_signature`). The
|
||
`confirmed` flag is **IGNORED** — a compromised controller asserting `confirmed:true` on a
|
||
protected device is refused **by role**. Every other destructive class (`guest_destroy`,
|
||
`decommission`, `restore_overwrite`, `key_rotation`) keeps operator-signature exactly as before.
|
||
- **`POST /disks/format`** — accepts `confirmed` + `durable_id` (inert for system/backup). The
|
||
data-bearing path tiers by role: user-data customer-confirmed → `mkfs`; user-data unconfirmed →
|
||
403 `needs_confirmation` (+ the durable id to confirm against, NOT an opsign command); system/backup
|
||
→ 403 with the operator-signature pending op (as before). Blank devices stay benign `mkfs`.
|
||
- **Tests** — `role_test.go` (demo-storage mapping + fail-safe), `storage_wipe_test.go` (the gate
|
||
refuses a `confirmed` wipe on system/backup → no exec; durable-id mismatch / missing-durable
|
||
refused; unknown role fails safe), and the localapi format-handler branches (user-data confirmed →
|
||
mkfs; user-data unconfirmed → needs_confirmation, no opsign; confirmed-but-protected → still refused).
|
||
|
||
Pairs with the controller's lockout + type-to-confirm UX + drive-list restyle.
|
||
|
||
## v0.22.0 — expose durable_id in GET /disks (enable controller-side guided storage) (2026-06-11)
|
||
|
||
One-line, read-only addition: `localapi.DiskInfo` gains `durable_id` (mapped from
|
||
`StorageTarget.DurableID`, e.g. `"uuid:<fs-uuid>"` for usb/local-dir). The de-privileged controller
|
||
cannot read a device's fs UUID itself, yet `POST /disks/assign` mounts strictly by UUID — so without
|
||
this it could not complete the guided init/attach flows. The controller strips the `uuid:` prefix to
|
||
get the assign key. No new privilege, no behaviour change to format/assign/eject or the data-bearing
|
||
gate. Pairs with `felhom-controller` v0.43.0 (the storage-management UI rebuild).
|
||
|
||
## v0.21.0 — agent-managed split-horizon LAN resolver (internal/lanresolver) (2026-06-11)
|
||
|
||
LAN clients can now reach their guest **directly** at the same public hostname with the same real
|
||
wildcard cert (no Cloudflare hairpin), via a host-side dnsmasq the agent manages. The host is the
|
||
stable anchor (static LAN IP); the guest stays DHCP/ephemeral and the agent tracks its live IP.
|
||
|
||
- **`internal/lanresolver`** — renders a dnsmasq base drop-in (bind to the host LAN IP, no-resolv,
|
||
upstreams) + a per-customer drop-in `local=/<domain>/` + `address=/<domain>/<guest-ip>`. The proven
|
||
two-line shape: `local=` makes dnsmasq authoritative for the zone so **AAAA returns NODATA** (no
|
||
Cloudflare-AAAA split-brain — the guest has only link-local v6), `address=` is the wildcard A; all
|
||
other names (and their AAAA) forward upstream unchanged.
|
||
- **`Manager`** ensures dnsmasq present (apt) + the base config + enabled, discovers the guest's live
|
||
IPv4 (`pct exec <vmid> -- ip -4 -o addr show dev eth0`) and domain (read from the guest controller's
|
||
pulled `controller.yaml` — the v2 bootstrap omits it), writes drop-ins **write-if-changed**, and
|
||
**reloads** (not restarts) dnsmasq. Tolerates the early-boot pre-lease window (empty IP → skip+retry,
|
||
never a blank record). Logs IP transitions.
|
||
- **`Loop`** — a 7th daemon goroutine: every interval (default 300s) it enumerates provisioned guests
|
||
(`/var/lib/felhom-agent/guests/<vmid>/`) and reconciles each, so the resolver follows DHCP IP changes.
|
||
Config `lan_resolver.{enable,host_ip,upstreams,interval_seconds}` (host_ip defaults to the local-API
|
||
bridge IP). `--selftest=lanresolver -vmid N`.
|
||
- **`configs/felhom-agent.sudoers`** — new `FELHOM_DNSMASQ` alias (apt install dnsmasq; install
|
||
felhom-*.conf drop-ins; systemctl enable/reload dnsmasq; rm felhom-*.conf; the two FIXED `pct exec`
|
||
reads). The agent never touches `/etc/resolv.conf` (host's own resolution unaffected).
|
||
- **Box-down robustness** is a documented **router config** (DNS = [host-IP primary, upstream
|
||
secondary]) so a box reboot degrades to the Cloudflare path, not total DNS loss — see REPORT install step.
|
||
- Spiked live on felhom-pve first (`:53` free, host IP static `192.168.0.162`, host DNS intact, full
|
||
loop from a real LAN client returned the guest IP + AAAA NODATA + the real wildcard cert `200 0`).
|
||
|
||
## v0.20.0 — golden: stacks-dir bind + per-guest hostname/CT name + bake base-infra images (2026-06-11)
|
||
|
||
Lockstep with `felhom-controller` v0.41.0 + a golden rebake. Changes in `configs/build-golden.sh` and
|
||
the provision path; no change to the proxmox/authz/token fences.
|
||
|
||
- **Section-G mount fix (the load-bearing one):** the in-guest controller writes app/infra compose
|
||
stacks under `/opt/docker/stacks` *inside its container*, but the baked controller-bootstrap `docker run`
|
||
never bind-mounted that path. So `docker compose up` (run by the GUEST daemon over the shared socket)
|
||
resolved every relative bind source on the guest filesystem — silently creating empty dirs — which
|
||
broke **every** bind-mounted stack (base infra AND customer apps like immich/nextcloud). The bootstrap
|
||
unit now `mkdir -p /opt/docker/stacks` and adds a **same-path host bind**
|
||
`-v /opt/docker/stacks:/opt/docker/stacks` (a named volume would NOT fix this). Empirically confirmed on
|
||
guest 9201 before writing the fix.
|
||
- **Per-guest container hostname (3A):** the bootstrap unit derives `customer.id` from
|
||
`/etc/felhom-bootstrap/bootstrap.json` with a portable `sed` parse (NO jq in the golden) and passes
|
||
`--hostname <customer-id>` to `docker run`, so the controller's `os.Hostname()` (its hub-reported
|
||
hostname) is the customer id, not the Docker container ID. Fail-safe: no parse → no `--hostname`.
|
||
- **Per-guest CT/LXC name (3B):** `--selftest=provision` now defaults `-hostname` to the (DNS-safe
|
||
sanitized) `-customer-id` when not given, so the bring-up's existing `SetConfig hostname` step
|
||
(`bringup.go`) names the CT meaningfully (e.g. `demo-felhom`) instead of inheriting the golden's
|
||
`felhom-golden`. New `sanitizeHostname` (lowercase, collapse invalid → `-`, trim, ≤63).
|
||
- **Bake base-infra images:** the golden now also pulls the three PINNED, PUBLIC base-infra images
|
||
(`traefik:v3.6.7`, `cloudflare/cloudflared:2026.6.0`, `gtstef/filebrowser:1.3.3-stable`) into its Docker
|
||
storage so the controller's first-boot bring-up is OFFLINE-capable. A hard gate (`docker manifest
|
||
inspect`) fails the bake early on a bad pin. Tags MUST match the controller's `internal/infra` constants.
|
||
|
||
## v0.19.0 — bootstrap contract v2: agent relays the hub retrieval passphrase (no host key in the guest) (2026-06-11)
|
||
|
||
Lockstep with `felhom-controller` v0.40.0. Fixes the onboarding 401: a freshly provisioned guest's
|
||
controller used to come up with the agent's **host** hub key baked in, which the hub's `/api/v1/report`
|
||
(customer-scoped auth) rejects. The agent now bakes a **v2 bootstrap** carrying only what the controller
|
||
needs to **pull** its own config from the hub — the agent never touches the customer-scoped key or CF
|
||
tokens.
|
||
|
||
### Changed — bootstrap contract `v1 → v2` (`internal/provision`)
|
||
- `SchemaV1 → SchemaV2 = "felhom.bootstrap/v2"`. **`DocCustomer`** drops `name`/`domain`/`email` (keeps
|
||
`id`). **`DocHub`** drops `api_key`/`host_id`, adds **`retrieval_password`** (the customer's hub
|
||
retrieval passphrase — SECRET). `DocLocalAPI` unchanged. The contract is byte-compatible with the
|
||
controller's `internal/bootstrap.Bootstrap` (cross-repo round-trip verified).
|
||
- `backhalf.go`: renders the v2 Doc; validation now requires `customer.id` + `hub.url` +
|
||
`hub.retrieval_password` (was `customer.id` + `customer.domain`). Write/0600/chown/`pct set` unchanged.
|
||
- `cmd/felhom-agent/main.go` `--selftest=provision`: **new required `-hub-password`** flag (the customer's
|
||
hub retrieval passphrase; the customer must already exist in the hub). Stops baking `cfg.Hub.APIKey` /
|
||
`cfg.Hub.HostID`. `-customer-domain/-name/-email` still accepted (bring-up may use them) but NOT baked.
|
||
|
||
### Changed — `configs/build-golden.sh`
|
||
- Default `CONTROLLER_IMAGE` bumped off the stale `:v0.35.0` → `:0.40.0` (matches the registry's no-`v`
|
||
tag convention; latent footgun fixed).
|
||
|
||
### Tests
|
||
- `doc_test.go`/`backhalf_test.go` updated to the v2 shape (assert no `api_key`/`host_id`,
|
||
`retrieval_password` present, `customer` carries only `id`). `go build ./... && go test ./...` green.
|
||
|
||
## v0.18.0 — slice 10D: DR capstone — identity escrow + restore-mode consumption (agent side) (2026-06-10)
|
||
|
||
The agent half of the slice-10 DR capstone (closes slice 10). Grounded by both 10-series spikes
|
||
(escrow-consumption + identity-restore). The hub half (recovery-mode toggle, re-enroll + credential
|
||
rotation, directive serving) is hub v0.11.0. **Operator-side rotation model (locked):** the hub holds
|
||
no Cloudflare write-power; the destructive tunnel/PBS rotation is the operator's step from a trusted
|
||
environment (same spirit as 10B).
|
||
|
||
### Added (`internal/escrow`)
|
||
- **Identity escrow** (`identity.go`): `WrapIdentity`/`UnwrapIdentity` (+ `…Bundle`) wrap the
|
||
`{tunnel_token, pbs_token}` bundle under the SAME recovery code `R` via **`age`** (scrypt +
|
||
ChaCha20-Poly1305 — a vetted passphrase-AEAD, not hand-rolled), reusing the K-escrow pty mechanism
|
||
(passphrase via the tty, data via files; `R`/tokens never logged). Same two-factor, zero-knowledge
|
||
shape as the K-escrow. A **wrong R fails closed** (no bundle). `age` is a runtime dep for the
|
||
identity path (analogous to proxmox-backup-client for K).
|
||
- **`escrow.Create`** gains an optional `IdentityBundle` → also emits an `IdentityBlob` under the same
|
||
R (additive; the K-escrow + 10C `Consume` paths are byte-unchanged). Self-verifies the identity
|
||
round-trip before shipping.
|
||
- **`--selftest=escrow-create -identity-bundle <file> -directive <file>`** — also wrap + upload the
|
||
identity blob + the **non-secret** DR directive (pbs repo/ns, expected key fingerprint, tunnel id).
|
||
- **`--selftest=identity-consume -blob <file> -keydest <file>`** (R via `FELHOM_RECOVERY_CODE`) —
|
||
recover the identity bundle through the real code; tokens written 0600, never logged.
|
||
|
||
### Tests
|
||
- identity bundle round-trips (wrap→unwrap byte-identical; blob is opaque ciphertext); wrong R fails
|
||
closed + the blob stays retryable; input validation. K-escrow/10C tests byte-unchanged (additive).
|
||
(age integration tests gated to a host with the `age` CLI.)
|
||
|
||
## v0.17.0 — slice 10C: escrow consumption (productionize the spike) (2026-06-10)
|
||
|
||
Turns the throwaway 10C spike harness into a real, tested **`Consume`** path: recover the PBS key
|
||
`K` from an R-wrapped escrow blob, **gate it on the expected fingerprint**, and install it for the
|
||
restore. The spike already proved the crypto + real-data restore; this bakes its findings into
|
||
production code. **Agent-only** — 10C *reads* the four inputs as parameters (so it stays
|
||
standalone-testable); 10D sources blob/fingerprint/PBS-connection from the hub and prompts for R.
|
||
**Zero-knowledge holds**: the hub serves everything except **R** (by hand from the customer), so a
|
||
hub compromise alone still can't decrypt.
|
||
|
||
### Added
|
||
- **`escrow.Consume(ctx, blob, R, expectedFingerprint, keyDest)`** — the consumption contract:
|
||
1. **Unwrap** the blob (a copy — F-C6: the input blob is read-only → a failed Consume is
|
||
**retryable**) with `R`; a **wrong R fails closed** at the scrypt KDF (F-C3) → a clear,
|
||
R-free error, **nothing written**.
|
||
2. **Fingerprint gate (F-C4)** — `KeyFingerprint(recovered)` must equal the expected (the hub
|
||
knows it); a mismatch **fails fast + loud, no install, no restore attempted**.
|
||
3. **Atomic install (F-C2)** at `keyDest` (`0600`, write-temp-sibling→rename); any failure leaves
|
||
**no partial install**. The recovered key lives only in a `0700` tempdir that is always removed.
|
||
**Secret discipline:** `R` and key bytes are never logged/persisted (only fingerprint prefixes);
|
||
`K` is never mutated.
|
||
- **`--selftest=escrow-consume`** (`-blob -fingerprint -keydest`, R via env `FELHOM_RECOVERY_CODE`
|
||
to keep it off the command line) — invokes the real `Consume` live (the spike's S3 via the
|
||
production path, not a harness).
|
||
|
||
### Tests (non-hollow)
|
||
- valid → key installed + `KeyFingerprint(dest) == expected` + `0600` + blob byte-unchanged;
|
||
**wrong R** → error, **no file at dest**, blob unchanged; **fingerprint mismatch** → fail fast,
|
||
**no install** (the gate runs before any restore); input validation; format-tolerant fingerprint
|
||
compare (no empty-fingerprint gate-bypass); atomic-install permissions (integration tests gated to
|
||
a host with `proxmox-backup-client`).
|
||
|
||
## v0.16.0 — slice 10B: operator-signed destructive completion (offline key + signing CLI) (2026-06-10)
|
||
|
||
The security centerpiece: a destructive op runs ONLY on a verified, operator-signed authorization
|
||
— signature valid against a **pinned** operator pubkey (never the hub's or the blob's), nonce
|
||
unseen + durably burned, in-window, host-bound, and **resource-bound to a DURABLE device id** that
|
||
execution re-resolves + re-inspects. Decision (a): **offline operator key + signing CLI**,
|
||
hardware-key-ready (`sk-`/YubiKey via ssh-keygen). The key floor holds: the signing key is NOT in
|
||
the hub and NOT in the agent. Concrete consumer: this **closes the 8C data-bearing-wipe
|
||
`pending_signature` gap**. Pairs with hub v0.10.0.
|
||
|
||
### Added
|
||
- **`cmd/felhom-opsign`** — the operator's offline signing CLI. Builds the canonical `OpBlob` by
|
||
**reusing `authz.CanonicalBlob`** (the exact production path the verifier authenticates over — so
|
||
signer + verifier can never drift) and signs it with **`ssh-keygen -Y sign -n felhom-op-v1`**
|
||
(hardware-ready). Output: a `{op_blob_b64, sig_armored}` envelope to hand to the hub jobs queue
|
||
(optional `--upload`). Touches ONLY the operator's signing key.
|
||
- **`authz.CanonicalBlob`** — promoted to production (was test-only) so the CLI + verifier share one
|
||
canonical-bytes source; params canonicalized (sorted keys, compact).
|
||
- **`internal/storage` durable device identity** (`durable_device.go`): `DeviceDurableID` (derive a
|
||
stable `byid:`(wwn/serial)/`byuuid:` id from the world-readable udev symlinks — no privilege, no
|
||
subprocess) + `ResolveDurableDevice` (re-resolve to the current `/dev` path; a path-only/unknown
|
||
scheme is REFUSED). The resource-level anti-retarget.
|
||
- **`internal/signedjobs`** (new): the queue consumer. `Runner` fetches each opaque job → runs it
|
||
through the **gate** (the LOCKED authz pipeline) → on all-pass hands the verified op to an
|
||
`Executor`; the order is **verify → nonce-burn (durable, in Verify) → execute → clear job**. The
|
||
**`WipeExecutor`** is the 8C consumer: resolve the signed durable id → **re-derive + match**
|
||
(anti-retarget) → **re-inspect (8C classifier)** the device is still the data-bearing target →
|
||
`mkfs`. A vanished/changed/non-data-bearing device or a path-only binding is refused **even with a
|
||
valid signature**. Wired as a second `EnvelopeObserver` (runs on `HasSignedOps`).
|
||
- **`hub.Client.Jobs` / `CompleteJob`** + `hub.MultiObserver`; the 8C format refusal now **surfaces
|
||
the bound op** (op + durable id + host) in its 403 `pending_op` + a `felhom-opsign …` hint.
|
||
|
||
### Pinning / rotation
|
||
- Operator pubkeys are pinned via `authz.signers` (config, trusted path — provision/agent config,
|
||
NEVER hub-alone), **multiple** keys (KeyID selects; role-scoped), so a backup/rotation key exists
|
||
without a flag-day. Unchanged from the slice-4 verifier wiring; 10B activates the execute path.
|
||
|
||
### Tests (real crypto, non-hollow)
|
||
- `signedjobs` runner over the **real** gate+verifier (in-Go minted SSHSIGs): valid → executor runs
|
||
once + job cleared; **replay** (nonce burned) / **non-pinned signer** / **expired** / **retarget**
|
||
(other host) / **forged sig** / **no pinned signer** → all rejected, **executor never called**;
|
||
malformed envelope cleared.
|
||
- `WipeExecutor`: valid → `mkfs` runs; **path-only**, **durable-id mismatch**, **device gone**,
|
||
**re-inspect non-data-bearing**, **not-probed** → all refused, `Format` not called.
|
||
- `storage` durable: wwn-preference, uuid-fallback, path-only/traversal refusal, round-trip,
|
||
missing-device error (symlink tests gated to Linux — the agent's OS).
|
||
|
||
## v0.15.0 — slice 10A: hub desired-state serving — the "Down" channel (2026-06-10)
|
||
|
||
The agent half of slice 10A. The control envelope (`hub.ControlEnvelope`) stops being "reserved — ignored" and becomes the live **Down channel**: a cheap change-notification on every heartbeat. The agent caches the hub's desired-state + its generation; only when **`DesiredGeneration` advances** does it fetch the full state (the heartbeat stays light, the heavy state moves on change). The engine then reconciles **benign** deltas and the gate marks an explicit **destructive** delta `pending_signature` (no signer in 10A → never executed; signed execution is 10B). Pairs with hub v0.9.0.
|
||
|
||
### Added / changed
|
||
- **`internal/reconcile`**: `DesiredGuest.Decommission` — the canonical **destructive desired-state delta** (an EXPLICIT flag, not "absent from the list", so a partial hub list can never mass-destroy). The planner emits `ActionDecommission` → `ClassDecommission` → Destructive → the gate refuses it `pending_signature`. `Reconcile` now counts a `pending_signature` refusal as **`Result.Pending`** (expected, logged INFO) rather than a failure; any other refusal stays a real failure. `ActionDecommission` has **no executor** (slice 10B) — a defensive guard refuses to run it. New **`CachingProvider`** (thread-safe DesiredState + generation cache; `Desired`/`Update`/`Generation`) — the production `DesiredProvider`, replacing `EmptyProvider` in the daemon engine (empty until the hub serves intent → cold-start is a live no-op, unchanged).
|
||
- **`internal/hub`**: the **`ControlEnvelope`** fields are now active (DesiredGeneration drives the fetch, HasSignedOps noted). New wire types **`DesiredStateResponse`** + **`WireDesiredState`** (guests + forward-compat `restore_directive` (10D) / `pbs_namespace` / opaque `storage_manifest`+`backup_policy`) + **`WireDesiredGuest`** (vmid/run/spec/description/decommission). New **`Client.FetchDesiredState`** (GET `/api/v1/hosts/{host_id}/desired-state`, self-scoped to the client's own host). New **`EnvelopeObserver`** loop seam + `SetEnvelopeObserver` — the loop hands the envelope to the sync layer each cycle (hub does not import reconcile/desired).
|
||
- **`internal/desired`** (new): the **`Syncer`** — implements `hub.EnvelopeObserver`, fetches desired-state on a generation advance, maps the wire shape to the reconcile domain, and updates the `CachingProvider`. Caches the **fetched** generation (robust to a generation that advanced mid-fetch); a fetch failure keeps the last-known state. `restore_directive` is carried + logged, not acted on (10D). Wired in `cmd/felhom-agent` (daemon): provider → engine, syncer → loop.
|
||
|
||
### Tests
|
||
- reconcile: a desired-state with one benign + one decommission delta → **benign applied, destructive gated pending (not executed)**; `Plan` emits decommission-only for a decommissioned guest + classifies Destructive; `CachingProvider` update/isolation.
|
||
- desired: **fetch-once-on-advance** (no re-fetch on an unchanged generation), fetch-failure-keeps-cache, caches-the-fetched-generation.
|
||
- hub client: `FetchDesiredState` hits the self-scoped path with the bearer + decodes (incl. `restore_directive`); a 403 is a typed `HTTPError`.
|
||
- loop: the cycle notifies the observer + adopts `PollIntervalSeconds`; a report error skips the observer.
|
||
- cross-repo golden: `testdata/desired-state.golden.json` + `control-envelope.golden.json` decode + key-set guard, **byte-identical** with felhom.eu/hub.
|
||
|
||
## v0.14.0 — slice 9: host metrics to the controller (`GET /host/metrics` + CPU-temp collector) (2026-06-10)
|
||
|
||
The de-privileged controller (slice 8C) sees only its own cgroup, so it can't read host health itself. Slice 9 **re-serves** the slice-4 collector's host + per-storage view to the customer over the local API, plus the one missing collector — CPU/chassis temperature — so the customer sees their box's health in the controller. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot). Assumption: **one customer per host** (the home-server model); if a host ever serves multiple customers, host-wide CPU/mem would leak cross-customer load → revisit then.
|
||
|
||
### Added / changed
|
||
- **CPU/chassis-temp collector** (`internal/hub/cputemp.go`): `SysfsTempReader` reads the CPU package temperature straight from sysfs — hwmon (`coretemp`/`k10temp`/`zenpower`/`cpu_thermal`, preferring the `Package id 0` input) then the thermal zones (preferring `x86_pkg_temp`/`coretemp`/`cpu-thermal`, falling back to `acpitz`). **No external binary, no privilege** (sysfs nodes are world-readable), so the root-CLI fence is untouched. **Graceful-null**: a missing sensor, an unsupported board, an implausible reading (outside 5–150 °C), or any read error all degrade to `null` ("n/a") — a missing sensor never fails the report. Wired into the collector via the new `TempReader` seam (nil-safe).
|
||
- **`HostMetrics.CPUTempC *int` (`cpu_temp_c`)** — new nullable wire field on the **shared** `HostMetrics` struct (same nullable contract as the disk `SmartSummary.TemperatureC`). It rides the **hub report too** (operator freebie) → cross-repo host-report golden updated.
|
||
- **`Collector.HostMetricsNow(ctx)`** — a fresh `NodeStatus` + CPU-temp read returning just the host block, the source for the local API (current cpu%/temp, not the 15-min snapshot). `Collect()` now also populates `cpu_temp_c` on the hub report. `Collector.SetTempReader` injects a fake in tests.
|
||
- **`GET /host/metrics`** (`internal/localapi/host_metrics.go`): host-wide health (cpu%/mem/load/uptime/`cpu_temp_c`) + per-storage capacity targets (total/used/fraction, thin-pool, SMART temp+wear). Token-authed via `withGuest` (host-wide data; cross-guest `?vmid=` still 403). Best-effort on storage (a view error still returns the host block). Served only when the `HostMetrics` provider (the shared collector) is wired — else 503 "not configured". Wired in `buildLocalAPIServer`.
|
||
|
||
### Tests
|
||
- `cputemp_test.go`: a fake `/sys` layout proves hwmon package-preference, hwmon first-input fallback, thermal-zone-by-type selection over a non-CPU hwmon, **graceful-null on a sensorless host** (no error), and rejection of implausible (0 m°C) readings.
|
||
- `hostmetrics_test.go`: `HostMetricsNow` populates the temp, gracefully nulls it, hard-errors on `NodeStatus` failure; `Collect()` carries the temp.
|
||
- `host_metrics_test.go` (localapi): populated host+storage with a valid token; `cpu_temp_c:null` serializes; **401 without a token** (collector never invoked); 403 on a cross-guest `?vmid=`; 503 when not configured.
|
||
|
||
## v0.13.0 — slice 8B.2: quiesce downtime optimization (`snapshotted` phase) (2026-06-10)
|
||
|
||
The agent half of slice 8B.2. In snapshot mode, vzdump only needs the app-stopped state captured at
|
||
the **storage-snapshot moment**; after that it reads from the snapshot and the app can resume. The
|
||
agent now emits a **`snapshotted`** phase on `GET /backup/status` when the snapshot is taken, so the
|
||
controller (v0.38.0) resumes its app early — app downtime drops from *whole-backup* to
|
||
*until-snapshot* with no loss of app-consistency. Validated Phase-0 first on PVE 9.2.2: the marker is
|
||
`INFO: create storage snapshot 'vzdump'`; downtime ~24s→~1s for a 934 MB guest.
|
||
|
||
### Added / changed (`internal/backup` + `internal/localapi`)
|
||
- **`BackupRunner.BackupWithSnapshotHook(ctx, vmid, onSnapshot)`** — while the vzdump runs, a watcher
|
||
tails the task log (`TaskLogTail`) for the **`create storage snapshot`** marker and fires
|
||
`onSnapshot` **once**. The marker only appears in snapshot mode (stop/downgraded takes no storage
|
||
snapshot), and the watcher also bails on `backup mode: stop` — so it never fires in stop mode.
|
||
(`Backup` keeps its signature for the scheduler/selftest; both share one body.)
|
||
- **`/backup/status` phase `snapshotted`** (between `running` and `done`): `handleBackup` passes the
|
||
hook → `markSnapshotted` flips the running job to `snapshotted`. `done`/`failed` semantics unchanged.
|
||
|
||
### Tests
|
||
- localapi: snapshot mode → phase reaches `snapshotted` before `done` (gated fake holds the backup
|
||
open); stop mode → `snapshotted` **never** emitted (stays running → done). runner: the watcher
|
||
fires `onSnapshot` on the marker; in stop-mode log it never fires. `snapshotWatchInterval` is a
|
||
package var so tests run fast.
|
||
|
||
## v0.12.0 — slice 8C Phase A: disk endpoints + data-bearing classifier gate + mkfs executor (2026-06-10)
|
||
|
||
The agent half of slice 8C, Phase A (additive). Adds the host disk-management endpoints the
|
||
controller's disk UI drives — with the **8C security invariant**: the agent decides
|
||
data-bearing-ness by **inspecting the actual device** (agent-internal evidence), NEVER from the
|
||
caller's claim. A compromised controller asserting "this drive is blank" cannot wipe a data-bearing
|
||
drive. (Controller rewire + disk-subsystem retirement + de-privilege are Phases B/C, `felhom-controller`.)
|
||
|
||
### Added
|
||
- **`internal/storage` — `mkfs` executor + data-bearing inspection.** `SudoHostOps.Format(device,
|
||
fstype)` (device-pinned, `ValidateBlockDevice`+`ValidateFSType`, narrow `FELHOM_FORMAT` sudoers —
|
||
`mkfs.ext4 -F` / `mkfs.xfs -f` on a `/dev/*` path the agent fine-validates first).
|
||
`SudoHostOps.InspectDevice(device)` → `DeviceProbe` (filesystem signature via `blkid -p`, partition
|
||
table / partitions / mount via `lsblk -J`). **`DeviceProbe.DataBearing()` is conservative**: any
|
||
signature / partition table / partition / mount — OR a probe that did not read cleanly — is
|
||
data-bearing (fail-safe; an unreadable device is never called blank).
|
||
- **`internal/localapi` — the §6 disk endpoints**, all self-scoped (token→guest; cross-guest 403):
|
||
- `GET /disks` — host drives + a **data-bearing flag** (UI hint). Read-only/benign.
|
||
- `POST /disks/assign` — attach a drive as a mount (benign, additive → `EnsureMount`). Self-serve.
|
||
- `POST /disks/eject` — safe-unmount (benign, data preserved) + the **dependent guests** that
|
||
mount it (so the controller can warn which apps lose that storage).
|
||
- `POST /disks/format` — **the security centerpiece**: the agent **inspects the device itself**;
|
||
blank → benign → `mkfs`; **data-bearing → ClassStorageWipe → the slice-4 gate → refused
|
||
`pending_signature`** (the operator-signed completion is slice 10). The caller's claim is
|
||
ignored — only a device the agent reads as blank is formatted.
|
||
- `storageGateAdapter` bridges the format path to the slice-4 reversibility gate (no new gate/crypto).
|
||
|
||
### Tests
|
||
- localapi (security matrix): blank device → **mkfs called, gate not consulted**; a **data-bearing
|
||
device → 403, mkfs NEVER called**, gate consulted (`pending_signature`); an **ambiguous/unprobed
|
||
device → treated destructive** (fail-safe); even a gate that *allows* does not format data-bearing
|
||
in 8C; assign → `EnsureMount`; eject → `Unmount` + dependent guests; cross-guest → 403; bad
|
||
device/fstype → 400; unconfigured → 503.
|
||
- storage: `ValidateBlockDevice`/`ValidateFSType` (whitelist + injection rejection); `InspectDevice`
|
||
blank/filesystem/partition-table/mounted/failed-probe-fail-safe; `Format` invokes the right `mkfs.*`.
|
||
|
||
## v0.11.0 — slice 8B: app-consistent backup — /backup/due policy + /backup/status phases (2026-06-10)
|
||
|
||
The agent half of slice 8B (doc 03 §8). Turns the 8A thin backup stubs into the real policy the
|
||
in-guest controller's quiesce loop drives (controller half: `felhom-controller` v0.36.0). No hub
|
||
change. The downtime optimization (`vzdump --mode snapshot` + a `snapshotted` phase) is the 8B.2
|
||
fast-follow; the hub-served per-guest policy is slice 10.
|
||
|
||
### Changed (`internal/localapi`)
|
||
- **`GET /backup/due`** — real **cadence** policy (replaces the 8A "never backed up" stub): a guest
|
||
is due when no **successful** backup is recorded OR the newest one is older than the agent-local
|
||
cadence (`backup.backup_cadence_seconds`, default 24h). A successful `POST /backup` flips due to
|
||
**false** for the window, so the controller won't re-quiesce in a loop. A failed backup does not
|
||
satisfy the cadence. Returns `age_seconds` for diagnosis.
|
||
- **`GET /backup/status`** — real **phases** `idle | running | done | failed` + the job id, so the
|
||
controller can poll a backup to completion (was: just the latest stored backup).
|
||
- **`POST /backup`** — returns a **job id** + `running` phase; tracks the in-flight job and is
|
||
**single-flight per guest** (a second POST while one runs returns the same job — no concurrent
|
||
vzdump). On completion the job transitions done/failed and the result is recorded to the store.
|
||
- Config: `backup.backup_cadence_seconds` + `BackupCadence()`; the local-API server takes the cadence.
|
||
|
||
### Tests
|
||
- `/backup/due`: due when stale / no backup, **not due within the window after a success**, due again
|
||
past the cadence, **a failed backup does not count**. `/backup/status`: running→done and
|
||
running→failed (gated fake to observe the running phase). `POST /backup` single-flight (one vzdump
|
||
for concurrent POSTs). All still self-scoped (token→guest).
|
||
|
||
## v0.10.0 — slice 8A: agent local-API server + provisioning back-half (2026-06-10)
|
||
|
||
The host-agent half of slice 8A (doc 03 §6). Adds the per-guest **local API** the in-guest
|
||
controller calls over the bridge, and the **provisioning back-half** that follows the slice-7
|
||
bring-up front half. Grounded by `felhom.eu/documentation/tests/slice8a-channel-deploy-spike-findings.md`
|
||
(commit `4a81a96` — channel + deploy plumbing proven; the 5 gotchas resolved here). Controller half
|
||
is `felhom-controller` v0.35.0. No hub change.
|
||
|
||
### Added
|
||
- **`internal/localapi`** — the HTTPS local-API server (doc 03 §6), the **per-guest authorization
|
||
gate**. Serves a **persisted self-signed leaf** with a **stable SHA-256 fingerprint** (generated
|
||
once; a fresh cert each boot would invalidate every baked bootstrap pin). The **7 §6 endpoints**,
|
||
all **self-scoped to the caller's own guest**: `GET /storage` (this guest's mpN mounts + fast/slow
|
||
class from the slice-5/7 storage view), `POST /snapshot`, `POST /rollback`, `POST /backup`
|
||
(enqueued, crash-consistent — the app-consistent quiesce loop is 8B), `GET /backup/due` (thin in
|
||
8A), `GET /backup/status`, `GET /restore-test/status`.
|
||
- **Token store** (`tokenstore.go`): durable, crash-safe per-guest token→guest map that persists
|
||
only a **SHA-256 hash** of each token (the plaintext exists transiently at mint→write-to-mount,
|
||
then is discarded), last-write-wins per guest, fsync'd append-only JSONL (mirrors the nonce store).
|
||
- **Self-scoping**: the VMID is resolved ONLY from the token; an explicit `vmid` (query/body) that
|
||
disagrees → **403 and the proxmox op is never issued for the other guest**; absent/unknown → 401.
|
||
- **`internal/provision`** — the back-half: mint the per-guest token → render the stable
|
||
**`bootstrap.json`** contract (schema `felhom.bootstrap/v1`; **no registry credential** — the
|
||
controller image is baked into the golden) → write it `0600` → **`chown 100000:100000`** (the
|
||
unprivileged-LXC mapped guest-root, spike gotcha 1) → attach a **read-only bind mount** via
|
||
`pct set`. Host-side only (F3 — the agent never enters the guest; **no `pct exec`**). The token
|
||
plaintext is never logged and never returned.
|
||
- **`--selftest=provision`** — the full chain on-demand: bring-up (provision) front half + the
|
||
back half; keeps the guest for the golden's baked controller-bootstrap unit to deploy.
|
||
- **`config.LocalAPIConfig`** (`local_api`) — enable + bridge `listen_addr` + cert/key paths + token
|
||
store path. The server is an optional 6th daemon goroutine, disabled cleanly when unconfigured or
|
||
on a token-store/cert failure (the daemon still reports/reconciles).
|
||
- **`configs/build-golden.sh`** now **bakes the controller image** (pulled once on the trusted build
|
||
host, then `docker logout` — no cred baked) + a **controller-bootstrap unit** that deploys the
|
||
**baked** image from the config mount on boot (no login/pull at deploy).
|
||
- **`configs/felhom-localapi-firewall.example`** — host firewall narrowing of the local-API port to
|
||
the guest bridge subnet (nft/iptables/PVE variants; defense-in-depth — the token stays the gate).
|
||
- **`configs/felhom-agent.sudoers`** — a narrow `FELHOM_PROVISION` alias (`chown 100000:100000` +
|
||
`pct set` bind-mount, both confined to the agent-owned `/var/lib/felhom-agent/guests/*` path) for
|
||
the non-root least-privilege deployment.
|
||
|
||
### Security / design notes
|
||
- The local-API leaf is pinned by **leaf-cert SHA-256** (decision: consistency with the agent's
|
||
PVE/PBS pinning); the fingerprint is baked into each guest's bootstrap.
|
||
- The back-half's host-root ops (chown + bind-mount attach) are **NOT** added to `proxmox.Privileged`
|
||
(which is fenced to its 3 exceptions) — they live in `internal/provision` and run through the shared
|
||
`Runner` (direct as root, or `sudo -n` with the new sudoers alias). This is the per-guest
|
||
provisioning host-root surface, host-side and F3-compliant.
|
||
|
||
### Tests
|
||
- localapi: self-scoping (cross-guest snapshot/rollback/backup → 403, op never issued for the other
|
||
guest; own-guest uses the token's VMID), 401 paths, `/storage` class mapping, `/backup` enqueue,
|
||
the thin `/backup/due`, status scoping; the token store persists only the hash (plaintext never on
|
||
disk), last-write-wins, survives reopen, uniqueness; the leaf fingerprint is stable across reload.
|
||
- provision: writes `0600` + chowns + attaches the bind mount with the right args; the **token never
|
||
appears in the Result**; the cross-repo `bootstrap.json` contract key-set is pinned.
|
||
|
||
## v0.9.0 — slice 7 close-out: PBS recovery-code escrow creation (2026-06-10)
|
||
|
||
The first code that touches the PBS client encryption key `K` and introduces the customer recovery
|
||
code `R`. Default posture is **zero-knowledge**: Felhom holds an opaque `R`-wrapped blob (cannot
|
||
open it), the customer holds `R`. Grounded by `felhom.eu/documentation/tests/slice7-escrow-spike-findings.md`
|
||
(round-trip proven on a throwaway: the `R`-recovered key restores a real encrypted snapshot). Hub
|
||
opaque storage is the `felhom.eu` half (hub v0.8.0); consumption/serving is slice 10.
|
||
|
||
### Secret discipline (overriding)
|
||
`R` is `crypto/rand`, ≥128 bits, surfaced **exactly once** and **never** logged/persisted/committed;
|
||
the wrap pty's echo is discarded so `R` can't leak. `K` is read by location, **never modified** (the
|
||
live key file is byte-unchanged — Wrap operates on a copy), never logged.
|
||
|
||
### Added
|
||
- **`internal/escrow`** — `Create` generates `R` (10 EFF-wordlist words ≈ 129 bits), wraps `K` under
|
||
`R` via the **PBS-native** `proxmox-backup-client key change-passphrase --kdf scrypt`, and
|
||
**self-verifies** the blob recovers `K` (fingerprint match) before shipping. The wrap is driven
|
||
over a **stdlib pty** (`x/sys/unix`; spike F-A1 — the command is TTY-only) with **output discarded**
|
||
(F-A2 — the pty echoes the passphrase). Opt-in outputs: **(b)** `R`-wrapped offline copy (two-factor,
|
||
no extra trust) and **(a)** raw paperkey (single-factor, unrevocable — loud caveat).
|
||
- **`--selftest=escrow-create`** (`-storage`, `-paperkey`, `-offline`, `-upload`): surfaces `R` once
|
||
to stdout (never the logger), prints the opaque blob's size/fingerprint/posture, and with
|
||
`-upload` PUTs the blob to the hub (`/api/v1/hosts/{host_id}/escrow`, per-host key).
|
||
- Config: `escrow` section (`posture` default `zero_knowledge`, `pbs_storage_id`); `PBSEncKeyPath`
|
||
helper (the `<id>.enc` key K).
|
||
- Runtime dependency on the `proxmox-backup-client` CLI (the PBS key+passphrase KDF).
|
||
|
||
### Tests
|
||
- `R` entropy ≥128 / 10-word format / uniqueness; integration round-trip (wrap→unwrap fingerprint
|
||
match, **wrong-`R` fails**, **live `K` byte-unchanged**, blob ≠ plaintext key) guarded to
|
||
linux+`proxmox-backup-client`; the agent→hub wire-contract key-set (mirrors the hub's).
|
||
- **Live-validated** (demo): `escrow-create` → `R` (10 words) surfaced once, blob 383 B opaque,
|
||
self-verify ok, **live `K` sha256 unchanged**, exact `R` absent from stderr/journal.
|
||
|
||
## v0.8.0 — slice 7 Phase 1: unified bring-up reconcile job (provision + guest-loss DR) (2026-06-09)
|
||
|
||
The shared FRONT HALF of provision and guest-loss DR, as a journaled reconcile job mirroring the
|
||
slice-6 restore-test's crash-safety — but it KEEPS the guest on success and applies a
|
||
scenario-specific identity policy. Agent-only; no hub/wire change (the new guest auto-appears in
|
||
the host-report via `ListLXC`). Grounded by the slice-7 bring-up spike findings (commit `3342993`):
|
||
F1 (restore preserves the archived MAC → provision reset is unconditional), F3 (SSH host keys do
|
||
not auto-regenerate → a baked golden first-boot unit, not an agent guest-internal op), F4 (the
|
||
transient PVE config-lock 500 → bounded retry).
|
||
|
||
### Added
|
||
- **`reconcile.RunBringUp`** (`bringup.go`) — `BringUpSpec` (Mode `provision`|`dr_guest_loss`,
|
||
Archive, VMID, RestoreStorage, Hostname, Cores/MemoryMB, RootfsGrowGB, Mounts, KeepMAC,
|
||
BootTimeout) → `BringUpResult` (VMID, AssignedMAC, Pass, Verified, StartWarnings/Recognized).
|
||
Sequence (each mutation preceded by journaling the owning entry): restore → identity reset →
|
||
size → attach mounts → start LINK-UP. **Verdict is liveness (`waitRunning`), never the start
|
||
exitstatus** (reuses the v0.7.0 WARNINGS surface). **Success KEEPS the guest** (no teardown).
|
||
- **Scenario-specific identity reset** (doc 03 §9): *provision* → fresh MAC unconditionally
|
||
(`PUT net0` with `hwaddr` omitted → PVE regenerates, F1) + hostname; machine-id + SSH host keys
|
||
regenerate guest-side on first boot (golden bake + the new unit) — the agent does NOT touch
|
||
guest internals. *dr_guest_loss* → preserve continuity (keep hostname; keep MAC unless
|
||
`KeepMAC=false`); never resets restic/tunnel/hub identity.
|
||
- **Compensating rollback** — any mid-flight failure destroys the just-created guest
|
||
(`ClassGuestDestroy`, benign via `Provenance{SameTxnCreated:true}`, gated); on teardown failure
|
||
the entry is left in-flight for `Recover`. New journal flag **`Rollback`** + `Recover`'s
|
||
`recoverBringUp` reap a half-built guest left by a mid-job crash (idempotent, via `ListLXC`).
|
||
- **F4 config-lock retry** — steps 3+5 coalesced into ONE `PUT config` (net0+hostname+cores+
|
||
memory+mpN); rootfs grow stays its own call. `setConfigWithLockRetry` retries ONLY the transient
|
||
PVE config-lock 500 (`pveConfigLock`: 500 + "can't lock file"/"got timeout"); any other error
|
||
fails immediately — never retried.
|
||
- **`--selftest=bring-up`** (`-mode provision|dr -archive -vmid -hostname [-keep]`) — runs the real
|
||
journaled job (after a `Recover`), then tears the guest down unless `-keep`.
|
||
- **`configs/build-golden.sh`** — the validated golden recipe as a script, incl. the F3
|
||
first-boot `felhom-regen-hostkeys.service` unit (Condition-gated: fires on provision, no-ops on
|
||
DR). The slice-7 spike archive (which lacks the unit) is superseded.
|
||
|
||
### Deferred (stated, not built)
|
||
- Provisioning BACK HALF (controller deploy, bootstrap, per-guest token mint) → **slice 8**.
|
||
- Host-loss DR + PBS escrow consumption → **slice 10**.
|
||
- The SOURCE of a `BringUpSpec` (hub desired-state: which archive/VMID/mounts) → **slice 10**;
|
||
this job takes the spec as input. `GuestMount` is defined minimally (no hub coupling).
|
||
|
||
### Tests
|
||
- provision happy path (fresh MAC = net0 without hwaddr, hostname, coalesced sizing+mount, rootfs
|
||
grow separate, started, **guest NOT destroyed**); compensating rollback at each step (restore /
|
||
config / start-task / waitRunning — asserts the guest WAS destroyed); DR continuity (MAC kept,
|
||
hostname not reset) + DR `KeepMAC=false` resets MAC; liveness verdict (warnings+running pass /
|
||
not-running fail); F4 (lock-500→retry→proceed; non-lock-500→fail without retry); owning entry
|
||
journaled BEFORE restore; reserved/existing VMID refused; `Recover` rolls back / clean.
|
||
|
||
### Live-validated (demo-felhom)
|
||
- provision: fresh MAC + hostname; **SSH host keys regenerated by the baked golden unit** (agent
|
||
issued no `ssh-keygen`), machine-id unique, Docker runs, clean DHCP lease → torn down.
|
||
- dr: continuity preserved (hostname + host keys kept). Recover: a killed mid-restore left an
|
||
orphan; the re-run's `Recover` rolled it back (idempotent).
|
||
- **Live caught a bug, then fixed:** the host-key unit's `ExecStart` was `/usr/sbin/ssh-keygen`
|
||
(203/EXEC); on Debian 13 it is `/usr/bin/ssh-keygen` — corrected in `build-golden.sh`, golden
|
||
rebuilt, re-validated. (Mocked unit tests couldn't surface this; the live run did.)
|
||
|
||
## v0.7.0 — restore-test: verdict is liveness, not start-task exitstatus (2026-06-09)
|
||
|
||
Fixes a correctness bug found by the live hub-enrollment runbook: the self-restore-test reported
|
||
`pass:false` on **every** modern-distro guest. PVE's guest-start task exits `"WARNINGS: 1"` for the
|
||
benign systemd-nesting advisory (`WARN: Systemd 257 detected. You may need to enable nesting.`), and
|
||
`WaitTask` treated any non-`"OK"` exitstatus as a hard failure — so the verdict was decided by an
|
||
advisory exit code instead of by observed liveness, *before* the real boot check ran. A crying-wolf
|
||
test got it disabled on the demo host; this re-enables it. **Single bump (0.6.0→0.7.0) covering the
|
||
agent's part of both task phases**; the wire fields below are consumed by hub from **v0.7.5**.
|
||
|
||
Design invariant (in code): **warning classification affects *visibility only*; pass/fail is
|
||
liveness-only.** A wrong/stale recognizer can at worst over-notice a benign warning — it can never
|
||
false-fail and never hide a real warning.
|
||
|
||
### Added
|
||
- **`proxmox.WaitOptions.AllowWarnings`** — opt-in per call. When set, a task that completes
|
||
`"WARNINGS: N"` is success with the `TaskStatus` (ExitStatus intact) returned so the caller can
|
||
read/surface it. Default (`false`) keeps **every existing caller strict** (vzdump/restore/destroy
|
||
warnings can be meaningful — relaxing them is a future per-call decision with evidence). Any
|
||
non-WARNINGS non-OK exit is still a `*TaskError`.
|
||
- **`reconcile.RestoreTestResult.StartWarnings` / `.WarningsRecognized`** + a version-free recognizer
|
||
(`benignWarningAnchor = "enable nesting"`, case-insensitive substring — contains no systemd version
|
||
number, so it can't rot back into the bug at systemd 258+). `extractWarningLines` pulls `WARN…`
|
||
lines from the start-task log.
|
||
- **`reconcile.GuestAPI.TaskLogTail`** — the engine fetches the start task's log to surface warnings.
|
||
- **`hub.RestoreTest.warnings` / `.warnings_recognized`** wire fields (`omitempty`), populated by
|
||
`ToHubRestoreTest`. Additive: the deployed v0.7.4 hub ignores them; hub v0.7.5 consumes them
|
||
(passed-with-warnings INFO, or WARN when not recognized). Cross-repo golden updated with the hub side.
|
||
|
||
### Changed
|
||
- **Restore-test start step** (`reconcile/restoretest.go`) now waits with `AllowWarnings:true`,
|
||
surfaces any start warnings, and **continues to `waitRunning` as the verdict** — boot+running is the
|
||
pass, exactly as before; a real (non-WARNINGS) start-task error still fails. The restore and
|
||
scratch-teardown WaitTasks stay strict.
|
||
- **Restore-test scheduler logging** distinguishes a clean pass, *passed-with-recognized-warnings*
|
||
(INFO), and *passed-with-unrecognized-warnings* (WARN) — nothing silent.
|
||
|
||
### Tests
|
||
- `WaitTask`: AllowWarnings accepts `WARNINGS` (status returned intact); AllowWarnings still fails a
|
||
real error; default still fails on `WARNINGS` (existing callers unaffected).
|
||
- Restore-test (engine, mock proxmox): start-with-warnings + running → **pass** with warnings
|
||
surfaced+recognized; unrecognized warning + running → pass, not-recognized; **not-running → fail
|
||
regardless of warnings** (verdict is liveness); teardown still runs.
|
||
- **Regression guard:** the `"enable nesting"` recognizer matches the advisory for systemd 256–300,
|
||
proving it's version-independent and can't silently rot back into the false-fail.
|
||
|
||
## v0.6.0 — slice 6 Phase B: PBS offsite tier (verify + PBS-API client + reporting) (2026-06-09)
|
||
|
||
Completes slice 6. The PBS spike (felhom.eu phase5-pbs-spike-findings.md) proved backup-to-PBS
|
||
and restore-from-PBS reuse Phase A UNCHANGED (PBS is just a storage target + a volid), and the
|
||
operator token needs no widening. So the only new agent code is the **verify capability + a
|
||
small PBS-API client + PBSSnapshot reporting**. Escrow + host-loss DR stay slices 7/10.
|
||
|
||
### Added
|
||
- **`internal/pbs` — the PBS-API client** (the agent's SECOND privileged external surface,
|
||
slice-1 discipline): TLS **fingerprint-pinned** to the PBS leaf cert (a spoofed PBS →
|
||
rejected, mirroring the PVE pin), **token auth** (`PBSAPIToken=<id>:<secret>`; id from the
|
||
storage `username`, secret read at runtime from `/etc/pve/priv/storage/<id>.pw` — referenced
|
||
by location, never logged/committed), typed, no shell. Methods: `Verify` (POST
|
||
`/admin/datastore/<ds>/verify` → UPID), `Snapshots` (incl. the `verification` field),
|
||
`TaskStatus`/`WaitVerify` (node extracted from the UPID — `localhost` returns "unknown", the
|
||
spike B4 gotcha), `NodeFromUPID`.
|
||
- **The verify maintenance loop** (`pbs/verify.go`) — the cheap, key-free, ciphertext-level
|
||
integrity check (§8) on its OWN cadence (default 6h, the 5th daemon goroutine). It is a
|
||
reporting/maintenance task like the slice-5 watchdog: it does NOT go through the reconcile
|
||
gate/journal. Each cycle: trigger verify → poll task → re-list snapshots → record
|
||
per-snapshot `verify_state`. A failed verify is logged loudly.
|
||
- **`PBSSnapshot` reporting** — filled the stub (`namespace`/`backup_type`/`backup_id`/
|
||
`backup_time`(RFC3339)/`size_bytes`/`owner`/`protected`/`encrypted` (from `files[].crypt-mode`)
|
||
/`verify_state` (ok|failed|**none** until verified)/`verify_upid`). New `PBSReporter`
|
||
collector seam + an in-memory `SnapshotStore`. Cross-repo golden (both repos, byte-identical)
|
||
+ bidirectional key-set tests; hub `handler.go` parses `pbs_snapshots` and logs a **failed
|
||
verify `[WARN]`** (loudest offsite-DR signal).
|
||
- **Truthful backup mode** (`backup/runner.go`) — `Backup.mode` now reflects the ACTUAL vzdump
|
||
mode read from the task log (`backup mode: <x>`), since PVE may downgrade snapshot→stop for a
|
||
stopped guest (spike B1); falls back to the requested mode if unparseable.
|
||
- **proxmox**: `Storage.Username` (parsed from the pbs storage config — the token id).
|
||
- **config** `BackupConfig.{PBSVerifyCadenceSeconds, PBSSecretDir}` (cadence 0→6h, <0 disabled).
|
||
- **`--selftest=pbs-verify`** — discover pbs storages → verify each → print the PBSSnapshot
|
||
records (covers the runbook's verify + list). Standalone on the host.
|
||
|
||
### Notes
|
||
- Backup/restore-to-PBS reuse Phase A with no change (the restore-test runs with
|
||
`source_tier="pbs"` when fed a pbs volid). Zero-knowledge holds: verify is ciphertext-level,
|
||
the encryption key is never read here, and the PBS server has no client key (spike B6).
|
||
- Daemon runs cleanly with no pbs storage / verify disabled. `go test -race` covers the new
|
||
goroutine. Slice-3/4/5/6A surfaces, goldens, and adversarial tests intact.
|
||
|
||
## v0.6.0-rc1 — slice 6 Phase A: backup + the self-restore-test (local target) (2026-06-09)
|
||
|
||
Phase A of the backup/restore slice (doc 03 §8) — the agent's guest-level backup layer and
|
||
the **self-restore-test**, which closes "a backup you haven't restored isn't a backup".
|
||
Everything here is BENIGN (backup, restore-to-NEW, scratch teardown): reuses the slice-4
|
||
classifier/gate/journal — no new destructive class, no new crypto. Local target only; PBS is
|
||
Phase B. Restore is to a NEW guest only (no overwrite). Backups are crash-consistent only
|
||
(app-consistency needs the controller quiesce, slice 8) — marked so in the report.
|
||
|
||
### Added
|
||
- **proxmox** (`mutate.go`/`query.go`): `DestroyLXC` (DELETE …/lxc/{vmid}?purge=1&destroy-
|
||
unreferenced-disks=1 → UPID; the scratch-teardown primitive); `VzdumpOptions.Notes` →
|
||
`notes-template` (verified on PVE 9.2.2); `LatestBackupVolID` (resolve a produced archive
|
||
from the backup-storage listing — the task status carries no result volid).
|
||
- **reconcile self-restore-test** (`restoretest.go`) — `Engine.RunRestoreTest`: pick a free
|
||
scratch VMID (configured band, excludes 9999; full band → skip, never out-of-band) →
|
||
**journal a Scratch-owned entry BEFORE any mutation** → restore-to-new → benign net
|
||
**link-down** SetConfig (so the clone can't conflict with a running source's MAC/IP; this
|
||
is test-safety, NOT slice-7 identity reset) → boot → verify **reaches `running`** → ALWAYS
|
||
teardown (defer; benign `ClassGuestDestroy` + agent-tagged-scratch provenance, gated). Runs
|
||
on the scratch VMID's queue lane. Reuses the journal/gate; result feeds the report.
|
||
- **Crash-safe recovery** (`recover.go`): a Scratch journal entry is resolved by TEARDOWN,
|
||
not by re-checking the restore sub-task's UPID — special-cased BEFORE the generic path
|
||
(else the restore task's OK would mark it succeeded while the guest leaks). `Recover` now
|
||
destroys a leaked scratch guest (idempotent: already-gone → clean; list-unreadable → left
|
||
in-flight for a later pass). `JournalEntry.Scratch` flag; `RecoverResult.ScratchClean/
|
||
ScratchDestroyed`. GuestAPI gains `RestoreLXC`/`DestroyLXC`/`GuestStatus`.
|
||
- **`internal/backup` package**: `BackupRunner.Backup` (vzdump + archive/size resolve +
|
||
bulk-volume gap — a mountpoint is UNCOVERED unless it carries an explicit `backup=1`, so
|
||
an unset `backup=` is reported uncovered too, the safe DR direction); `PickRestoreCandidate`
|
||
(newest backup); an in-memory `Store` (latest-backup-per-target + latest-restore-test)
|
||
implementing the hub `BackupReporter`/`RestoreTestReporter` seams; a cadence `Scheduler`
|
||
(default 24h; the fourth daemon goroutine; disabled cleanly when off/misconfigured).
|
||
- **hub report** (`report.go`): filled the `Backup` + `RestoreTest` stubs (`PBSSnapshot`
|
||
stays a Phase-B stub); collector `BackupReporter`/`RestoreTestReporter` seams. Cross-repo
|
||
golden updated in BOTH repos (byte-identical) + bidirectional key-set tests for
|
||
`backups[0]`/`restore_tests[0]`. Hub `handler.go` parses + persists them (report_json; no
|
||
new columns) and logs a **FAILED restore-test prominently** (the loudest DR signal).
|
||
- **config** `BackupConfig` (local target, restore storage, restore-test cadence, scratch
|
||
VMID band 990000–990009 default) + accessors + env overlay + cadence-gated validation.
|
||
- **`--selftest=backup -vmid N`** (one-shot backup → print the Backup record) and
|
||
**`--selftest=restore-test [-archive volid]`** (Recover-then restore→boot→verify→teardown,
|
||
print the RestoreTest record). Standalone on the Proxmox host.
|
||
|
||
### Notes
|
||
- The daemon runs cleanly with the cadence off or misconfigured (logs + disables, never
|
||
crashes); a leaked scratch guest from a mid-test crash is reaped by `engine.Recover` on
|
||
restart. `go test -race` covers the new scheduler goroutine.
|
||
- Slice-3/4/5 exported surfaces, goldens, and adversarial tests intact. Version bumps to
|
||
**v0.6.0** when Phase B (PBS) lands.
|
||
|
||
## v0.5.1 — slice 5 live-validation prep: durable_id mis-id fix + re-mount UUID memory (2026-06-09)
|
||
|
||
Two correctness fixes surfaced while preparing the live USB validation on `demo-felhom`
|
||
(a real 1TB USB HDD, sdb1, ext4). Both are DR-load-bearing — exactly the "false-id →
|
||
re-attach the wrong disk" failure mode the slice warned about.
|
||
|
||
### Fixed
|
||
- **Unmounted dir-storage no longer inherits the ROOT filesystem's UUID** (`observe.go`).
|
||
Previously, when a removable dir-storage was unmounted, the observer fell through to the
|
||
*containing* mount (root) for the backing device, so its `durable_id` became
|
||
`uuid:<root-uuid>` — a catastrophic DR mis-id (the hub would re-attach the wrong disk).
|
||
Now the backing device/UUID/`durable_id` are derived ONLY from the target's OWN
|
||
mountpoint; an unmounted target reports no device and a stable `store:<name>` durable_id,
|
||
never another filesystem's UUID. (Removed the `containingMountDevice` root-fallthrough.)
|
||
- **Watchdog remembers the fs-UUID observed while attached** (`watchdog.go`) so a re-mount
|
||
works even after the known-set cache refreshes mid-drop (an unmounted target can't resolve
|
||
its own UUID). The re-mount key is backfilled from this memory — aligning with doc 03 §7's
|
||
"sourced from the existing definition, no hub manifest needed": the agent learns the UUID
|
||
while the target is attached, then re-mounts by it on return.
|
||
|
||
### Tests
|
||
- Observer: an unmounted dir-storage asserts NO `uuid:` durable_id and no backing device.
|
||
- Watchdog: a drop where the cache lost the UUID still re-mounts using the remembered UUID.
|
||
|
||
## v0.5.0 — slice 5 Phase B: the host-root surface (mounts + SMART + grow + destructive gate) (2026-06-09)
|
||
|
||
The write surface — the agent's first step outside its Proxmox API token into OS-root.
|
||
Isolated behind a narrow, argument-validated, adversarially-tested seam, exactly like the
|
||
slice-4 gate. Completes slice 5 (Phase A = read-only observe/report/watchdog at v0.5.0-rc1).
|
||
|
||
### Added
|
||
- **`HostOps` seam + `SudoHostOps`** (`internal/storage/hostops.go`) — the one privileged
|
||
host surface: persistent mounts via **systemd `.mount` units keyed by fs-UUID** (enabled to
|
||
survive reboot), detach (stop+disable), SMART, and thin-pool metadata. Shells out via the
|
||
fenced Runner (`sudo -n`, **fixed arg vectors, no shell**); a fake backs the tests (no real
|
||
root in the suite). `NoopHostOps` is the safe fallback when the surface is unavailable.
|
||
- **The argument validator** (`internal/storage/validate.go`) — the security boundary:
|
||
`ValidateUUID` (strict hex), `ValidateMountPath` (absolute, no traversal, no metacharacters),
|
||
`ValidateSMARTDevice` (raw-disk whitelist), `ValidateLVMName`, and an in-process
|
||
`systemdEscapePath` (no `systemd-escape` shell-out). **Every argument is validated BEFORE a
|
||
command is constructed.** Headline test (`validate_test.go`): an adversarial matrix of
|
||
shell metacharacters / `../` traversal / malformed inputs is rejected with **zero exec**.
|
||
- **SMART** (`internal/storage/smart.go`) — parses `smartctl -a -j` into `StorageTarget.smart`:
|
||
**SATA** (reallocated/pending/offline-uncorrectable, temp, power-on-hours) **and NVMe**
|
||
(critical_warning, media_errors, percentage_used, temp), degrading to `UNKNOWN` for devices
|
||
with no SMART (USB-SATA bridges). **`lvs`** fills the lvmthin thin-pool **metadata** fill
|
||
(the value Phase A left null). Wired into the Observer's enrichment (Observe only, not the
|
||
watchdog's fast Known path).
|
||
- **Watchdog re-mount response** (`internal/storage/watchdog.go`) — on a known mount-backed
|
||
target's device returning **unmounted** (a new `DevicePresent` liveness probe), the watchdog
|
||
**dispatches a benign by-UUID re-mount off the poll path** (a goroutine, never under the
|
||
lock), rate-limited per target to the debounce window. The mount is routed through the gate
|
||
as benign (`gateRemounter` in `main.go`, so `storage` stays decoupled from `reconcile`).
|
||
- **Disk-grow executor** (`internal/reconcile`) — `ActionResize` (benign `ClassResize`), planned
|
||
**grow-only** (desired DiskBytes > actual → `pct resize rootfs +<n>M`; a shrink is refused,
|
||
never silently grown) + a defensive executor guard (size must start with `+`). New
|
||
`proxmox.Client.ResizeLXC` (API; `VM.Config.Disk`+`Datastore.AllocateSpace`; async→UPID).
|
||
Built + fixture-tested; **unfed** live (no hub spec until slice 10).
|
||
- **Destructive storage ops through the slice-4 gate** (`internal/reconcile/storage_ops.go`) —
|
||
`IntentForStorageMount` (benign) and `IntentForStorageDestructive` (`ClassStorageWipe`/
|
||
`ClassDecommission`). Host/target-scoped: the op binds on the storage **target identity**
|
||
(carried in `target.guest_id`). Reuses the existing verifier/role-scoping/binding/audit — no
|
||
new gate, no new crypto. Storage cases added to the adversarial matrix (`storage_test.go`):
|
||
unsigned wipe → `pending_signature`; "wipe A" signature vs "wipe B" → `binding_mismatch`;
|
||
valid → accepted. **Inert** live.
|
||
- **`--selftest=storage` [`-watch <dur>`]** — the live USB-runbook harness: an observe pass
|
||
(full table incl. SMART + thin-pool data+metadata), and a bounded watchdog window with the
|
||
re-mount response live. Runs standalone on the Proxmox host (no hub).
|
||
- **`configs/felhom-agent.sudoers`** — the documented narrow allowlist (install unit / systemctl
|
||
manage / smartctl / lvs), with the agent-side fine validation noted.
|
||
- **Config**: `privileged.{unit_dir,stage_dir,systemctl,install,smartctl,lvs}` (paths must match
|
||
the sudoers entries).
|
||
|
||
### Notes
|
||
- Daemon still runs cleanly with no removable storage / no signers / no hub manifest, and a
|
||
missing/declined sudoers entry degrades with a warning (SMART→UNKNOWN, mount→logged error),
|
||
not a crash. `go test -race` passes (the watchdog re-mount dispatches off the poll path).
|
||
- Slice-3/4 + Phase-A exported surfaces, goldens, and adversarial tests intact. `authz`
|
||
untouched. The destructive-storage executor + grow are built/tested but unfed live until
|
||
slice 10.
|
||
|
||
## v0.5.0-rc1 — slice 5 Phase A: storage observe + report + watchdog (read-only, live) (2026-06-09)
|
||
|
||
Phase A of the storage slice (doc 03 §7). Read-only and live: the agent now observes every
|
||
host storage target, reports it into the host-report's `storage_targets` (previously an empty
|
||
stub), and runs a fast-poll watchdog that pushes a disconnect to the hub in seconds. No
|
||
host-root writes this phase (mounts/SMART/grow/destructive-gate are Phase B). The hub-owned
|
||
desired manifest (class/role/policy/creds) is not served until slice 10, so reconcile against
|
||
it is built-but-unfed — this phase ships only the genuinely-useful read-only footprint.
|
||
|
||
### Added
|
||
- **`internal/storage` package** (new):
|
||
- **`StorageTarget` wire contract** (`internal/hub/report.go`) — filled the slice-3 stub:
|
||
`name`/`type`/`durable_id`/`state`/`reachable`, usage (`total`/`used`/`avail`/
|
||
`used_fraction`), `content`, `mount_path`/`backing_device`, `class_hint` (rotational HINT
|
||
— never authoritative; class is hub-owned), `role` (empty until slice 10), a `thin_pool`
|
||
sub-object (lvmthin data fill; metadata fill is Phase B/`lvs`), and a `smart` sub-object
|
||
(`UNKNOWN` until Phase B). Cross-repo golden kept byte-identical with `felhom.eu/hub` and
|
||
guarded by the bidirectional key-set test (`contract_test.go`).
|
||
- **`durable_id` derivation** (`durableid.go`) — deterministic per type (the DR-load-bearing
|
||
re-attach key): fs-UUID (usb/local-dir), `server:export` (nfs/cifs), `repo+fingerprint`
|
||
(pbs), `vg/pool` (lvmthin); never empty (falls back to a stable store id).
|
||
- **`HostReader` seam + `ProcHostReader`** (`hostread.go`) — non-privileged `/proc/mounts`,
|
||
`/dev/disk/by-uuid`, `/sys/.../rotational` + `removable` reads. Root-free by construction.
|
||
- **`Observer`** (`observe.go`) — builds `[]hub.StorageTarget` from `ListStorage`/`NodeStorage`
|
||
joined with host reads; surfaces the lvmthin thin-pool data fill prominently (warns ≥85%).
|
||
- **Storage watchdog** (`watchdog.go`) — a third daemon goroutine fast-polling the *known*
|
||
target set (a defined Proxmox storage and/or a previously-seen one) for
|
||
`attached↔disconnected` transitions; on a transition it triggers an immediate, **debounced**
|
||
out-of-band host-report. Only flags a *known* target's change (never a never-attached
|
||
device); coalesces flaps within the debounce window (leading + trailing edge).
|
||
`CachingKnownTargets` rate-limits the Proxmox-derived known set; `HostLiveness` probes
|
||
device/mount presence (local) + a reachability dial (network), all non-privileged.
|
||
- **Proxmox `Storage` type** (`internal/proxmox/types.go`) — additive parse-only config fields
|
||
(`server`/`export`/`share`/`datastore`/`fingerprint`/`vgname`/`thinpool`) feeding durable_id.
|
||
- **Collector `StorageObserver` seam** (`internal/hub/collect.go`) — populates `storage_targets`
|
||
via the observer; a nil observer or an observe error degrades to empty (never sinks the
|
||
heartbeat). Hub does not import storage (storage imports hub for the wire type).
|
||
- **Out-of-band report trigger** (`internal/hub/loop.go`) — `Loop.SetTrigger`: a watchdog
|
||
signal runs one extra collect→report immediately without disturbing the regular cadence.
|
||
- **`StorageConfig`** (`internal/config`) — watchdog interval / debounce / known-refresh knobs
|
||
(all optional; package defaults otherwise).
|
||
- **Hub ingest** (`felhom.eu/hub`) — `hostReportPayload` now parses `storage_targets`
|
||
(full mirror struct), persists them via `report_json`, counts + warns on disconnected
|
||
targets, and has its own half of the bidirectional golden key-set test.
|
||
|
||
### Notes
|
||
- The daemon still runs cleanly with no removable storage, no signers, and no hub manifest —
|
||
the watchdog finds nothing to flag; storage reporting is best-effort.
|
||
- `proxmox`/`hub`/`authz`/`reconcile` exported surfaces + their golden/adversarial tests are
|
||
intact. No host-root writes, no destructive paths, no SMART this phase (all Phase B).
|
||
- Version: **v0.5.0-rc1** at the Phase-A checkpoint; **v0.5.0** when Phase B lands.
|
||
|
||
## v0.4.0 — slice 4 Phase B: reversibility gate + signed-op consuming layer (2026-06-08)
|
||
|
||
The security core of slice 4: hub-supplied intent stops being trusted for destructive
|
||
change. Layered in front of the per-guest queue's executor — **every** mutation now
|
||
passes the gate. Reuses `internal/authz` for all crypto (untouched surface). Inert
|
||
this slice: no destructive deltas are served until slice 10, so the destructive path is
|
||
classified, gated, and adversarially tested but not wired to live execution.
|
||
|
||
### Added
|
||
- **Classifier (`classify.go`, doc 03 §4)** — benign vs destructive by **provenance +
|
||
data-bearing-ness, NOT by verb**. The `OpClass` vocabulary (seeded by the committed
|
||
slice-2 `op_blob.json`: `guest_destroy`) is the agent-side contract slice 10 matches.
|
||
Destroy/overwrite of customer data is destructive UNLESS **agent-internal**
|
||
provenance (same-journaled-transaction create → compensating rollback, or
|
||
agent-tagged scratch) makes it benign. `Provenance` is journal-recorded and **never
|
||
populated from the hub** (its zero value is the only thing an external intent may
|
||
carry). Unknown op class fails safe → destructive.
|
||
- **Reversibility gate (`gate.go`)** — `Gate.Authorize(intent, signed)`: benign →
|
||
allowed unsigned; destructive → requires a verified, role-authorized, action-bound
|
||
operator signature, else refused **`pending_signature`**, never executed. Every
|
||
decision is written to an `AuditSink` (audit is a signal, never the guard).
|
||
- **Signed-op consuming layer over `authz`** — verifies via `authz.Verifier.Verify`
|
||
(the locked pipeline, untouched), then enforces on the `VerifiedOp`:
|
||
- **Role-scoping (doc 04 §4)** — recovery key authorizes key-rotation re-pins ONLY;
|
||
operational key authorizes ordinary destructive ops + planned rotation.
|
||
- **Op-to-action binding** — verified `op` + host + guest + `params` must match the
|
||
gated action (a signature for guest X / op A can't authorize guest Y / op B);
|
||
params compared semantically (key-order/whitespace independent).
|
||
- **Signed-job orchestration (`job.go`)** — `RunSignedJob`: idempotency dedupe (the
|
||
op nonce as the journal key — a redelivered completed op is skipped, not re-run),
|
||
gate authorization, then journal-wrapped execution via an injected
|
||
`DestructiveExecutor` (nil this slice — authorized destructive ops are inert, no
|
||
executor wired until 6/7).
|
||
- **Crash-recovery consumer (`recover.go`, Note 1 / doc 03 §10)** — `Engine.Recover`
|
||
consumes the journal's `InFlight()` at startup: an op that crashed AFTER the Proxmox
|
||
POST and BEFORE its terminal record (`OpTaskRunning`, nonce already consumed) is NOT
|
||
covered by idempotency dedupe — only this resume-or-rollback resolves it (re-read the
|
||
task via the new `TaskStatusOnce`, record the real outcome; a no-task-id op is
|
||
abandoned fail-safe). Landed together with the signed-op executor, as Note 1 required.
|
||
- **Daemon wiring** — `runDaemon` builds the verifier from `config.Authz.Signers` (a
|
||
bad key / missing nonce-store path is a fatal misconfig; **no signers = nil verifier**,
|
||
the common slice-4 state), constructs the gate (+ `SlogAudit`), runs `Recover` before
|
||
issuing any mutation, and routes every reconcile action through the gate.
|
||
|
||
### Changed
|
||
- **Memory comparison canonicalized (Note 2)** — `desiredMemoryMiB` makes the
|
||
desired↔actual memory compare in the same MiB unit that is then written, so a
|
||
non-MiB-aligned `MemoryBytes` converges in one pass instead of re-issuing SetConfig
|
||
forever (the numeric cousin of the description-newline normalization). Test proves
|
||
convergence. Slice 10 should still serve MiB-aligned specs at the source.
|
||
|
||
### Tests (the security proof — each independently rejected)
|
||
- **Adversarial matrix** via the REAL `authz.Verifier` with in-test-minted SSHSIGs
|
||
(framing replicated in reconcile's test binary; production authz untouched, no signing
|
||
added to the verify-only package): unsigned destructive **job** → pending_signature;
|
||
unsigned destructive **desired-state delta** → pending_signature (distrusts hub
|
||
desired state, not just jobs); forged/unknown signer → `ErrUnknownSigner`; expired →
|
||
`ErrExpired`; **replayed nonce across an agent restart** (durable `FileNonceStore`) →
|
||
`ErrReplay`; wrong host → `ErrTarget`; wrong guest / wrong op / wrong params →
|
||
binding_mismatch; **recovery key on ordinary destructive** → role_denied;
|
||
**hub-supplied "scratch" tag ignored** → still destructive → refused; **valid + role +
|
||
target + fresh nonce → accepted**, and a second presentation → `ErrReplay` (nonce
|
||
consumed).
|
||
- Classifier (benign/destructive/provenance/key-rotation/fail-safe), role-scoping,
|
||
params binding, crash-recovery (resume OK / fail / still-running / no-task rollback /
|
||
unreadable / one-shot key applied on resume), signed-job idempotency (execute once,
|
||
dedupe redelivery, refused-not-executed, no-executor-inert, executor-error).
|
||
- Full module **race-clean** (`go test -race`) + vet clean on the Linux build server.
|
||
|
||
## v0.4.0-rc1 — slice 4 Phase A: reconcile engine (structural; runs live, unfed) (2026-06-08)
|
||
|
||
The agent-side control core's structural half. **Checkpoint marker** — `-rc1` is the
|
||
Phase-A push; awaiting validation before Phase B (the reversibility gate + signed-op
|
||
consuming layer) lands the final **v0.4.0**. Runs LIVE but UNFED: with no desired-state
|
||
provider until slice 10, the live engine computes an empty action set and performs
|
||
**zero mutations**.
|
||
|
||
### Added
|
||
- **`internal/reconcile`** package — the engine, the per-guest serializer, the
|
||
desired-state model, the normalization layer, and the durable op journal:
|
||
- **Per-guest serializer (`Queue`, doc 03 §10)** — the single choke point ALL
|
||
mutation sources funnel through. Same-vmid jobs run strictly one-at-a-time in
|
||
submit order; independent vmids run in parallel. Each vmid is a cond-var FIFO lane
|
||
(unbounded, non-blocking, order-preserving); graceful drain on `Close`.
|
||
- **Desired-state model + `DesiredProvider` seam** — `DesiredGuest` (per-field
|
||
optional: run-state / `*hub.GuestSpec` / `*description`), `DesiredState`. The only
|
||
live provider is **`EmptyProvider`** (slice 4 has no source); `StaticProvider`
|
||
feeds fixtures. The seam is where slice 10's hub-serving plugs in — no hub/local
|
||
source invented here.
|
||
- **Normalization layer (`FieldNormalizers`)** — reconcile compares *normalized*
|
||
desired-vs-actual so Proxmox round-trip quirks don't read as drift. `description`'s
|
||
trailing newline is the first registered case; the registry takes more (boolean
|
||
coercion, list ordering) as discovered. `normDesc` **promoted** out of
|
||
`cmd/felhom-agent/main.go` to **`reconcile.NormDescription`**; the `--selftest=task`
|
||
description round-trip now uses that shared helper (one source of truth for the quirk).
|
||
- **Plan engine (`Plan`, pure function)** — computes the minimal **benign** action set
|
||
(`Start`/`Stop`/`SetConfig`) for guests present in both desired and actual, with
|
||
normalized comparison, deterministic vmid ordering, config-before-run-state. Skips
|
||
provision (desired-absent-in-actual, slice 7) and destroy (actual-absent-in-desired,
|
||
gated, slice 10); never writes a config it couldn't first read (`SpecKnown`). Disk
|
||
(rootfs grow) intentionally not reconciled here.
|
||
- **Reconcile engine (`Engine`)** — reads desired+actual, plans, dispatches each action
|
||
onto the shared queue. Every Proxmox op handled per the mutate.go contract: non-empty
|
||
UPID → `WaitTask` + assert `exitstatus`; empty UPID → clean **synchronous** success
|
||
(slice-4 proven). Per-action failures are counted, not fatal (other guests still
|
||
converge).
|
||
- **Operation journal (`Journal`)** — durable fsync'd append-only JSONL mirroring
|
||
`authz.FileNonceStore`: records each op's lifecycle (started → task_running →
|
||
succeeded/failed) with its Proxmox task id (crash mid-op is detected and re-checkable
|
||
on restart via `InFlight()`), plus an **idempotency-key store** (`AlreadyApplied`) so
|
||
a one-shot op never re-runs across retries/restarts. Reconcile actions carry no
|
||
idempotency key (convergent — must re-run on real drift).
|
||
- **Daemon wiring (`runDaemon`)** — reconcile runs alongside the hub loop on the poll
|
||
cadence, **sharing the per-guest queue**. Journal path is a `journal.log` sibling of the
|
||
nonce store. The daemon runs cleanly with **no desired state and no signers** (reconcile
|
||
is a logged live no-op; a journal-open failure degrades to journal-less, never crashes).
|
||
|
||
### Tests
|
||
- Serializer: same-guest serialized (max-concurrency 1, submit order preserved) and
|
||
different-guests parallel (cross-waiting jobs both complete — would deadlock if not);
|
||
error propagation; drain-pending-on-close; submit-after-close.
|
||
- Normalization: description round-trip; unknown-field identity; extensibility seam
|
||
(synthetic boolean-coercion + list-ordering normalizers).
|
||
- Plan: run-state start/stop, spec drift (cores/memory), disk-not-reconciled,
|
||
description-newline-not-drift, unmanaged fields, spec-unknown skips config keeps
|
||
run-state, desired-absent skipped, combined ordering, empty-desired no-op, deterministic
|
||
vmid order.
|
||
- Engine: empty-provider zero mutations; async start (WaitTask); synchronous SetConfig
|
||
(no WaitTask); WaitTask failure + POST error counted failed; list error = pass failure.
|
||
- Journal: lifecycle latest-wins; in-flight survives restart; idempotency dedupe across
|
||
restart; failed key not applied; torn-trailing-line skipped.
|
||
- Full module **race-clean** (`go test -race`) on the Linux build server; vet clean.
|
||
|
||
### Not in this phase (Phase B)
|
||
- The benign/destructive classifier, the reversibility gate, and the signed-op consuming
|
||
layer over `internal/authz` (doc 03 §4 / doc 04) — added next, in front of the queue's
|
||
executor, landing **v0.4.0**.
|
||
|
||
## v0.3.2 — SetConfig selftest extension (slice-4 pre-check) (2026-06-08)
|
||
|
||
The gate before slice 4: prove `SetConfig` works live under the scoped token before
|
||
reconcile is built on it. **Self-gated live run PASSED** on `demo-felhom`/guest 9999.
|
||
|
||
### Added
|
||
- **Reversible `SetConfig` step appended to `--selftest=task`** (`cmd/felhom-agent/main.go`,
|
||
`selftestSetConfig`): read `GuestConfig` → write a `description` marker
|
||
(`felhom-selftest <RFC3339>`) → verify it landed → restore the original value (or
|
||
`delete` the key if it was absent) → verify the restore. Handles PVE's dual-mode
|
||
`SetConfig` return per the `mutate.go` contract: empty UPID = synchronous success
|
||
(printed `synchronous`); non-empty UPID = `WaitTask` + assert `exitstatus=OK`.
|
||
The existing snapshot → rollback → delete-snapshot steps are unchanged. First live
|
||
exercise of the **`VM.Config.*`** privilege cluster.
|
||
- **`normDesc` / `extraString` helpers** — `extraString` decodes a string-valued key
|
||
from `GuestConfig.Extra` (raw JSON); `normDesc` strips the trailing newline PVE
|
||
appends to `description` on read, so a written value round-trips equal.
|
||
|
||
### Finding (live)
|
||
- The LXC `description` write returned **synchronous (empty UPID)** — PVE applied it
|
||
inline, no task. The agent's dual-mode `SetConfig` modeling is correct: the
|
||
empty-string path is real and must not be treated as an error.
|
||
- PVE **appends a trailing `\n` to `description`** on read (stored URL-encoded as
|
||
`%0A`). A naive exact-match reconcile would see perpetual drift — slice-4 reconcile
|
||
must normalize `description` comparisons (hence `normDesc`).
|
||
|
||
### Ops
|
||
- Standing operator token (`felhom-agent@pve!agent`, privsep) **rotated** during this
|
||
run (the prior secret was not retrievable); role + both user/token ACL rows
|
||
re-confirmed at `/`. New secret stored out-of-band, **not persisted to the repo**.
|
||
Guest 9999 left pristine (stopped, no `description`, no leftover snapshot). Version → 0.3.2.
|
||
|
||
## Docs + live validation — no version bump (2026-06-08)
|
||
|
||
### Changed
|
||
- **Reflowed `CLAUDE.md`** — removed hard mid-paragraph line wraps (prose, list items, blockquotes now single-line, soft-wrapped); code blocks and tables untouched; rendered output unchanged.
|
||
- **Unified the REPORT/CHANGELOG convention** in `CLAUDE.md`: `CHANGELOG.md` is the cumulative log (newest on top); `REPORT.md` is overwritten with the most-recent implementation/validation only. Added an explicit **no-secrets** rule (never write tokens/passwords/keys into committed files; reference them as stored out-of-band).
|
||
|
||
### Added
|
||
- **`REPORT.md`** rewritten for the live `--selftest=task` validation on the demo host (`demo-felhom`): snapshot → rollback → delete-snapshot on guest 9999, each polled to `exitstatus=OK` under the `felhom-agent@pve!agent` privsep token (UPIDs name the token actor — privsep path genuinely exercised); 16-privilege `FelhomAgent` role + both user & token ACLs confirmed; `--selftest=read` clean. Closes the slice-1 "mutating ops unit-tested only" gap; `WaitTask` async foundation validated live → **slice 4 unblocked**. (Token secret stored out-of-band, not in the repo.)
|
||
|
||
## v0.3.1 — slice-3 validation follow-ups (2026-06-08)
|
||
|
||
### Changed
|
||
- **Collector keeps the known run-status on a `GuestConfig` failure** (`internal/hub/collect.go`):
|
||
previously a per-guest config-read error forced `status="unknown"`; now the run-status from
|
||
`ListLXC` is preserved (only the `spec` is dropped). An empty status is still normalized to
|
||
`unknown` (wire value is always `running|stopped|unknown`). Test renamed to
|
||
`TestCollect_GuestConfigFailureKeepsStatusOmitsSpec` and asserts the preserved `running` + nil spec.
|
||
- **`--selftest` usage** error string now reads `(want read|task|hub)`.
|
||
|
||
### Added
|
||
- **Cross-repo contract fixture** `internal/hub/testdata/host-report.golden.json` +
|
||
`TestHostReport_ContractMatchesGolden` — compares the marshaled `HostReport` field-name sets
|
||
(top level + `host` + `guests[0]`) against the golden, failing on any json-tag drift. The file is
|
||
**kept byte-identical** with felhom-hub's copy (duplicated contract until a shared types module;
|
||
revisit when slices 5/6 populate the empty collections). Version → 0.3.1.
|
||
|
||
## v0.3.0 — hub client + host-report + first daemon loop (slice 3) (2026-06-08)
|
||
|
||
The agent's first daemon: a periodic read-only host-report POSTed to the hub (the
|
||
heartbeat). No Proxmox mutations, no desired-state/signed-op consumption, no
|
||
storage/backup collection yet — those are slices 4/5/6.
|
||
|
||
### Added
|
||
- **`internal/hub`** package:
|
||
- **`HostReport`** wire contract (`report.go`) shared field-for-field with the hub
|
||
ingest: host metrics, guests (`vmid` + spec), `cloudflared` status, and the
|
||
`storage_targets`/`backups`/`restore_tests`/`pbs_snapshots`/`audit_tail`
|
||
collections **defined but emitted empty** (typed `[]`, slices 5/6 fill them).
|
||
- **`Collector`** (`collect.go`) builds the report from a read-only `proxmoxReader`
|
||
(adapted to the real `internal/proxmox` surface — node held by the client, value
|
||
returns, `proxmox.Guest`) + a `CloudflaredProber`. Partial-failure policy: a
|
||
failed `NodeStatus` is a hard error (skip the POST); a failed per-guest
|
||
`GuestConfig` degrades that guest to `status="unknown"` (spec omitted) but still
|
||
sends; a cloudflared probe failure → `"unknown"`, never fatal.
|
||
- **`CloudflaredProber`** + `SystemctlProber` (`systemctl is-active cloudflared`;
|
||
read-only — NOT a Privileged/root op; tunnel management is a later slice).
|
||
- **`Client`** (`client.go`): `POST /api/v1/host-report` with
|
||
`Authorization: Bearer <key>`, standard TLS (system roots or optional `ca_file`;
|
||
verification always on). Typed `*TransportError` / `*HTTPError`; the bearer token
|
||
never appears in any error.
|
||
- **`Loop`** (`loop.go`): the daemon — immediate first report then tick; adopts the
|
||
hub's `poll_interval_seconds` clamped to [60,3600]; resilient (a collect/report
|
||
error is logged and the loop continues); clean shutdown on context cancel.
|
||
- **`ControlEnvelope`**: only `poll_interval_seconds` is acted on; `blocked` /
|
||
`desired_generation` / `has_signed_ops` are parsed-but-ignored (logged at most)
|
||
pending reconcile (slice 4).
|
||
- **Config**: `HubConfig` (url/host_id/api_key/poll_seconds/timeout_seconds/ca_file),
|
||
`FELHOM_AGENT_HUB_*` env overlay, `HubConfig.Validate()` (mode-aware — proxmox-only
|
||
`--selftest=read|task` still runs without hub config), `WithDefaults()`, and
|
||
`Redacted()` now also blanks the hub key. `configs/agent.example.json` gains `hub`
|
||
(and `authz`) blocks.
|
||
- **`cmd/felhom-agent`**: the no-`--selftest` mode is now the **daemon** (poll loop);
|
||
added **`--selftest=hub`** (one collect+report, prints the report + envelope).
|
||
Version 0.2.0 → 0.3.0.
|
||
|
||
### Tests
|
||
- Report serialization (field names; empty collections are `[]` not `null`; spec
|
||
omitted when unknown); client (Bearer header, non-2xx→`*HTTPError`,
|
||
transport→`*TransportError`, **token never in error**); collector (host mapping,
|
||
guest spec, per-guest failure degrades-but-still-reports, NodeStatus hard error,
|
||
cloudflared error→unknown); loop (immediate first report, continuation after an
|
||
injected error, interval adoption + clamp); config (hub validate/redact/env).
|
||
|
||
### Notes
|
||
- `internal/proxmox` and `internal/authz` were **not touched** — no new proxmox
|
||
surface was needed (`ListLXC` already exposes status/maxmem/maxdisk; `GuestConfig`
|
||
exposes cores). The task's `proxmoxReader` sketch (node-arg/pointer/`LXC`) was
|
||
adapted to the real exports as instructed.
|
||
- **Defined-but-empty** this slice: `storage_targets`, `backups`, `restore_tests`,
|
||
`pbs_snapshots`, `audit_tail` (slices 5/6). **Parsed-but-ignored**: the envelope's
|
||
`blocked`/`desired_generation`/`has_signed_ops` (slice 4).
|
||
|
||
## v0.2.0 — `authz` signed-op verifier (slice 2) (2026-06-08)
|
||
|
||
Production form of the Phase-4 signing primitive: a key-type-agnostic SSHSIG
|
||
verifier for operator-signed destructive ops, with the full anti-replay/
|
||
authorization pipeline and a durable, crash-safe nonce store. What slice 4
|
||
(reconcile) will call to gate destructive desired-state deltas. No hub, no signing
|
||
CLI, no reconcile loop.
|
||
|
||
### Added
|
||
- **`internal/authz` — `Verifier`**: `New(signers, store, hostID)` + `Verify(blob,
|
||
sigArmored) (*VerifiedOp, error)`. Runs the LOCKED pipeline (order is
|
||
load-bearing): parse armor → namespace → parse pubkey → allow-list (by key
|
||
**material**, `pub.Marshal()` equality, not key_id) → crypto verify (over the
|
||
**raw received bytes**, never re-canonicalized) → parse blob → target → time
|
||
window → **nonce recorded LAST**. Each post-crypto stage rejects even with a
|
||
valid signature.
|
||
- **SSHSIG framing** (`sshsig.go`) via `golang.org/x/crypto/ssh` — `pem.Decode` →
|
||
strip 6-byte magic → `ssh.Unmarshal` → `ssh.ParsePublicKey` → recompute signed
|
||
data with the named hash → `pub.Verify` (dispatches on key algorithm). No
|
||
hand-rolled crypto. Key-type-agnostic: ed25519 / **sk-ssh-ed25519 (FIDO2)** /
|
||
rsa / ecdsa via the one path.
|
||
- **Fixed namespace** `felhom-op-v1` (package constant, never caller-supplied).
|
||
- **`OpBlob`** (corrected `host_id`/`guest_id` json tags) + **`VerifiedOp`** (op,
|
||
host/guest, params, key_id, matched signer). key_id is advisory/audit only —
|
||
never an authz input.
|
||
- **Typed errors**: `ErrMalformed, ErrNamespace, ErrUnknownSigner, ErrBadSignature,
|
||
ErrTarget, ErrExpired, ErrNotYetValid, ErrReplay` (errors.Is-friendly).
|
||
- **`NonceStore`** + two impls: `MemoryNonceStore` (tests) and **`FileNonceStore`**
|
||
— durable, crash-safe (fsync'd append log, replayed into an index on open,
|
||
periodic compaction, expiry-only pruning). A nonce is fsync'd to disk before
|
||
`SeenOrRecord` returns false; replay protection survives restart; I/O failure
|
||
fails safe (reports seen=true). Target generalization: host_id matched strictly,
|
||
guest_id surfaced for the caller to route.
|
||
- **Config**: `AuthzConfig` (nonce-store path + pinned operator `signers` tagged
|
||
`operational`/`recovery` with a key_id, as authorized_keys lines).
|
||
- **Version 0.2.0.**
|
||
|
||
### Tests
|
||
- Real OpenSSH interop via a committed `ssh-keygen -Y sign` vector (hermetic CI);
|
||
per-stage rejection (each with an otherwise-valid sig); the headline
|
||
**invalid-sig-does-not-burn-the-nonce** invariant; replay; **persistence across
|
||
restart**; synthetic **sk-ssh-ed25519** through the unchanged path; byte-exactness
|
||
(a re-serialized blob fails crypto — not re-canonicalized).
|
||
|
||
### Notes / corrections to the Phase-4 reference
|
||
- §7's `Target` lacked json tags (`host_id`/`guest_id`) — fixed.
|
||
- The doc paired "Go 1.24.4 / x/crypto v0.52.0", but v0.52.0 declares `go 1.25.0`
|
||
and does **not** build on Go 1.24. Resolved by upgrading the build server to
|
||
go1.26.0 (backward-compatible; felhom-controller/hub unaffected); the module is
|
||
`go 1.25.0` on x/crypto v0.52.0.
|
||
- Free function → constructed `Verifier`; returns the full `VerifiedOp`; typed
|
||
errors; clock-skew tolerance added; durable nonce store is the net-new work.
|
||
- **Shared-contract dependency flagged** (not built): the hub and the `felhom-sign`
|
||
CLI must emit byte-identical canonical JSON or signatures won't verify; a shared
|
||
canonicalizer both import would be the right home.
|
||
|
||
## v0.1.0 — Scaffold + `proxmox` interaction layer (slice 1) (2026-06-08)
|
||
|
||
First slice: stand up the host-agent project and its foundation — the typed
|
||
Proxmox interaction layer every other module will call. No reconcile loop, hub
|
||
client, signing, or storage/backup orchestration yet (later slices).
|
||
|
||
### Added
|
||
- **Project scaffold**: module `gitea.dooplex.hu/admin/felhom-agent`, binary
|
||
`felhom-agent` (`cmd/felhom-agent/`), Go 1.24, zero external dependencies
|
||
(pure stdlib). `--version` flag; `version` var overridable via
|
||
`-ldflags "-X main.version=<v>"`.
|
||
- **`internal/proxmox` — API backend (`Client`)**: hand-rolled REST client over
|
||
`https://<host>:8006/api2/json` with `PVEAPIToken` auth. Typed read ops
|
||
(`Version`, `Nodes`, `NodeStatus`, `ListLXC`, `GuestStatus`, `GuestConfig`,
|
||
`ListStorage`, `NodeStorage`, `StorageContent`) and async mutating ops
|
||
returning a UPID (`RestoreLXC` — the primary create path, `Vzdump`, `Snapshot`,
|
||
`Rollback`, `DeleteSnapshot`, `SetConfig`, `Start`, `Stop`).
|
||
- **`WaitTask`**: polls `GET /nodes/{node}/tasks/{upid}/status` until stopped, then
|
||
asserts `exitstatus == "OK"` (authorization can surface at task execution, not
|
||
the POST — phase1-2 §1.3). Exponential backoff (1s→5s cap), context
|
||
cancellation + timeout. `*APIError` parses the offending privilege from a 403;
|
||
`*TaskError` parses it from a failed task exitstatus + log tail.
|
||
- **`internal/proxmox` — fenced root-CLI backend (`Privileged`)**: limited to the
|
||
three proven OS-root exceptions only — `CreateGoldenLXC` (keyctl `pct create`),
|
||
`MountUSBByUUID`, `SMART`, `Sensors`; each cites why it can't be the API. Fence
|
||
is structural (Client never shells out, Privileged never makes an HTTP call) and
|
||
asserted in tests.
|
||
- **TLS trust**: SHA-256 leaf-cert pinning (the host serves a self-signed cert) or
|
||
a CA file; an explicitly-named `insecure_skip_verify` that is off by default. No
|
||
blanket verification disable.
|
||
- **`internal/config`**: JSON config file + `FELHOM_AGENT_*` env overrides; the
|
||
token secret is never logged (`Redacted()`).
|
||
- **`internal/log`**: slog setup (text, stderr, configurable level).
|
||
- **`cmd/felhom-agent --selftest`**: read-only health report against a live host
|
||
(version/nodes/status/guests/storage); `--selftest=task --vmid N` exercises
|
||
`WaitTask` on a reversible snapshot→rollback→delete op (gated; default selftest
|
||
mutates nothing).
|
||
- **Tests**: unit tests with a mock HTTP transport + mock runner (UPID parse,
|
||
`WaitTask` running→OK / failed-403 / timeout / ctx-cancel, 403→privilege error,
|
||
response decoding against shapes captured live from `demo-felhom`, config
|
||
redaction, and the API-vs-root routing fence).
|
||
|
||
### Notes
|
||
- Types are grounded in the spike findings
|
||
(`felhom.eu/documentation/proxmox-platform.md`, `tests/phase{0,1-2,3}-findings.md`)
|
||
and the exact JSON shapes captured live from `demo-felhom` (PVE 9.2.2).
|
||
- Verified: `go build/vet/test` green on Go 1.24.4 (build server) and a live
|
||
read-only `--selftest` against the demo host with TLS fingerprint pinning.
|
||
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
|
||
token) is provisioned out-of-band; the agent only consumes the token.
|