a667c269c7
Found by live validation on demo-felhom, not by review. The first real PBS-targeted backup ran past the runner's hard-coded 30-minute WaitTask bound. The agent stopped waiting and recorded success=false WHILE THE VZDUMP KEPT RUNNING (still running 72 min later, 2.4 GB uploaded). Consequences: the tier stays permanently due, the next attempt collides with the guest lock the live vzdump holds, and the hub sees a DR tier that never succeeds — R-82's 'applied and empty' fault re-created by a timeout. Measured: ~33 MB/min over wg to Hetzner, so a first FULL ~10 GB snapshot projects to ~5h. - BackupTargetConfig.WaitTimeoutSeconds: per-tier bound. Primary 30m UNCHANGED (a local vzdump hanging 30m IS a real fault); additional tier 6h, sized from the measurement. - backup.NewBackupRunnerWithWait: per-instance (per-tier) bound. NewBackupRunner keeps its signature, so restore-test/selftest are untouched. - localapi.BackupTier.WaitTimeout: the fire-and-forget context is sized from the tier, not a fixed 2h. BOTH bounds had to move — a 6h runner bound under a 2h outer context reproduces the same false failure four hours later. Same direction as restore_test_pbs_restore_timeout_seconds: when in doubt wait LONGER. A slow backup is a slow backup; a false timeout is a corrupt status plus lock contention. Red-proof observed and restored; full suite green.
3642 lines
283 KiB
Markdown
3642 lines
283 KiB
Markdown
## v0.98.0 — R-82 Slice A fix: per-tier vzdump wait bound (the 30-minute false failure) (2026-07-26)
|
||
|
||
**Found by live validation on demo-felhom, not by review.** The first real PBS-targeted backup ran
|
||
past the runner's hard-coded 30-minute `WaitTask` bound. The agent stopped waiting, recorded
|
||
`success:false` — **while the vzdump kept running** (still running 72 minutes later, 2.4 GB
|
||
uploaded). That is not "the backup didn't happen"; it is worse:
|
||
|
||
- the tier stays permanently **due** (a failed backup never satisfies a cadence),
|
||
- the next attempt collides with the **guest lock** the live vzdump still holds,
|
||
- the hub sees a DR tier that never succeeds — R-82's "applied and empty" fault, re-created by a
|
||
timeout,
|
||
- the recorded `Backup` says `success:false, size_bytes:0` for a backup that may yet complete.
|
||
|
||
30 minutes is right for a LOCAL vzdump (minutes) and simply wrong for an offsite upload. Measured on
|
||
demo-felhom: ~33 MB/min over the wg link to Hetzner ⇒ a first FULL ~10 GB snapshot projects to ≈5 h.
|
||
|
||
### Changed
|
||
- **`BackupTargetConfig.WaitTimeoutSeconds`** — per-tier vzdump wait bound. Defaults:
|
||
**primary 30 m (UNCHANGED)**, additional tier **6 h**. The asymmetry is the point: the primary is
|
||
the local tier where a 30-minute hang IS a genuine fault worth surfacing; an additional tier is by
|
||
construction the offsite one, where the binding constraint is uplink speed, not health. 6 h is
|
||
sized from the measurement above, not guessed.
|
||
- **`backup.NewBackupRunnerWithWait`** — the runner's wait bound is per-instance (i.e. per tier).
|
||
`NewBackupRunner` keeps its signature and delegates with 0 ⇒ 30 m, so every other caller
|
||
(restore-test, selftest) is untouched.
|
||
- **`localapi.BackupTier.WaitTimeout`** — the fire-and-forget backup context is now sized from the
|
||
tier instead of a fixed 2 h. **Both bounds had to move**: a 6 h runner bound under a 2 h outer
|
||
context would have reproduced the same false failure four hours later.
|
||
|
||
Same reasoning, and the same direction, as `restore_test_pbs_restore_timeout_seconds` on the restore
|
||
side: **when in doubt wait LONGER.** A slow backup is a slow backup; a false timeout is a corrupt
|
||
status plus lock contention.
|
||
|
||
### Tests
|
||
`TestBackupTiers_WaitTimeoutIsPerTier` pins the asymmetric defaults, the override, and that the two
|
||
tiers do NOT share one bound. Red-proof observed: setting the extra tier's default back to 30 m
|
||
fails with `an offsite tier must default to a GENEROUS wait — a false timeout is worse than a slow
|
||
pass; got 30m0s`. Restored; full suite green.
|
||
|
||
## v0.97.0 — R-82 Slice A: per-target backup tiers (local daily + PBS weekly) (2026-07-26)
|
||
|
||
Additive; **MinAgent floor rises** for the multi-tier contract (a controller that wants per-tier
|
||
backups needs this agent — but see the compatibility rule: an OLD controller is unaffected).
|
||
|
||
Slice A of R-82. `BackupTarget()` returned ONE string and `BackupCadence()` ONE 24h window, so
|
||
"local daily **and** PBS weekly" was not expressible at all — which is why the DR promise is
|
||
currently unbacked (demo-felhom holds one PBS snapshot from 2026-07-18, demo-hp zero, ever).
|
||
Phase-0 gate results: `felhom.eu/documentation/audits/SPIKE-r82-phase0-2026-07-26.md`.
|
||
|
||
**This slice ships the mechanism only. No box's behaviour changes until a `backup_targets` entry is
|
||
added to its config (Slice D).** An untouched config resolves to exactly one tier and behaves
|
||
byte-identically to v0.96.0.
|
||
|
||
### The compatibility rule (load-bearing)
|
||
|
||
The agent and controller deploy independently, so the untargeted local-API contract is **frozen**:
|
||
|
||
- `GET /backup/due` with **no** `?target=` → the PRIMARY tier, same cadence, same response **bytes**.
|
||
`BackupDueResponse.Target` is `omitempty` and left EMPTY for untargeted requests, so an old
|
||
controller cannot tell this agent from the old one. Pinned by a red-proofed test.
|
||
- Same for `POST /backup` and `GET /backup/status`.
|
||
- The primary tier's **job-id format is unchanged**; only additive tiers carry a target segment.
|
||
|
||
### Added
|
||
- **`config.BackupTargetConfig` + `BackupConfig.ExtraTargets`** (`backup_targets`) — each tier
|
||
carries its OWN cadence and its OWN retention. Those are semantically different per tier:
|
||
`keep-last=3` is three DAYS on a daily tier and three WEEKS on a weekly one, so sharing one knob
|
||
guarantees one of them is wrong.
|
||
- **`BackupConfig.BackupTiers() ([]BackupTier, []string)`** — resolves the tier list, primary first,
|
||
plus warnings the caller MUST log. A tier is **rejected, not defaulted**, when its cadence is
|
||
missing: silently defaulting a weekly DR tier to the 24h local default would fill the 37.2 GB
|
||
datastore. Empty/duplicate targets are rejected too. `main.go` logs every rejection at **ERROR** —
|
||
a silently dropped backup tier is the "applied and empty" fault this task exists to fix.
|
||
- **`GET /backup/tiers`** — advertises the tiers, primary first. This is the controller's capability
|
||
probe: a **404 means a pre-R-82 agent**, and Slice B falls back to single-tier on it.
|
||
- **`localapi.BackupTier` + `normalizeBackupTiers`** — nil tiers synthesize the legacy single tier,
|
||
so every existing caller and test hits the pre-R-82 path untouched. A tier with a nil runner is
|
||
DROPPED rather than advertised (advertising one would be an applied-and-empty tier).
|
||
|
||
### Changed
|
||
- **`/backup/due?target=` judges that tier against ITS OWN newest successful backup**
|
||
(`latestSuccessfulBackupForTarget`). Without this filter a fresh daily local backup would satisfy
|
||
the weekly PBS cadence and the DR tier would never run — today's bug, re-created in code. The
|
||
store was already keyed by target, so this is a lookup change, not a data-model change.
|
||
- **Backup jobs are keyed by (vmid, target), not vmid.** Single-flight is now PER TIER, which is
|
||
what lets the weekly night run both backups inside ONE quiesce window (Slice B). Keying by vmid
|
||
alone handed the second caller the first tier's job id — **caught by its own test**, and it would
|
||
have made the controller believe a PBS backup ran when only the local one had.
|
||
- **Job ids are unique per tier by construction**, not by clock luck (two tiers can start in the
|
||
same nanosecond). The primary keeps the old format; additive tiers carry the target segment.
|
||
- **One runner per tier** (`main.go`). The runner holds target/mode/notes/retention as immutable
|
||
construction state and `localPruneSpec` reads that retention — parameterising a single runner by
|
||
target would risk pairing tier A's target with tier B's retention.
|
||
- An unknown `?target=` is a **400, never a silent fallback to the primary**. A controller asking
|
||
about a tier this agent does not serve must find out, not act on another tier's freshness.
|
||
|
||
### NOT changed (deliberate)
|
||
- **The local tier.** Same target, same 24h cadence, same keep-last=3 clamp. It is the only honest
|
||
tier today and this slice does not touch it.
|
||
- **PBS is still never pruned by the per-run `--prune-backups` flag** (`localPruneSpec`). Per-tier
|
||
retention is plumbed and a tier's `keep_last` defaults to 0 = never prune, but turning on
|
||
automatic pruning of the DR datastore is irreversible and needs an operator ruling — R-82 Phase 0
|
||
already flagged retention as needing one. Recorded, not decided here.
|
||
- The fail-safe-toward-due rule on an unparseable timestamp. A spurious backup is cheap; a skipped
|
||
one is not.
|
||
|
||
### Tests
|
||
768 (was 748), +20 across `internal/localapi/backup_tiers_test.go` and
|
||
`internal/config/backup_tiers_test.go`. Red-proof #1 (old controller ↔ new agent) observed:
|
||
removing the untargeted compat branch makes the untargeted body carry `"target":"local"` and the
|
||
test fails on it. Restored.
|
||
|
||
## v0.96.0 — R-50 island NIC: provision attaches the guest's island net1 (2026-07-25)
|
||
|
||
Additive; **MinAgent unchanged** (no controller coupling — the controller dials whatever `bootstrap.json`
|
||
says, and the pin is address-independent). Implements the provisioning half of the R-50 island control
|
||
plane, spiked GO in `felhom.eu/documentation/audits/SPIKE-island-bridge-2026-07-25.md`. A fresh install
|
||
is now born immune to F1 (`AUDIT-vacation-remote-ops-2026-07-20`: a LAN/DHCP move made the agent fail to
|
||
bind → storage/PBS/quiesce/restore-test/DR down, silently).
|
||
|
||
- **`LocalAPIConfig`** gains `island_bridge` + `island_guest_addr` (config.go). `IslandEnabled()` = both
|
||
set. `Validate()` enforces all-or-nothing + a CIDR guest addr — a half-set/malformed island fails at
|
||
load (a botched install), never a silent LAN fallback that would leave an island bind with no island NIC.
|
||
- **`buildBringUpConfig`** (reconcile/bringup.go): when both island fields are set, attaches a static
|
||
`net1=name=eth1,bridge=<vmbr9>,ip=<169.254.253.2/30>` (no hwaddr → fresh per-guest MAC) on BOTH
|
||
provision and DR bring-up. Empty = pre-R-50, no net1 (byte-for-byte the old config on non-island hosts).
|
||
Plumbed from `cfg.LocalAPI` at both `RunBringUp` call sites (cmd/main.go). The bootstrap `endpoint`
|
||
already derives from `listen_addr` (main.go), so moving the agent bind to the island moves the guest
|
||
dial for free — **no template change** (A0 determination).
|
||
- **Design note (A0/A3):** endpoint is config-derived, so no code was needed there; the guestnet healer
|
||
is eth0-only (`parseMode` is dev-scoped) so the static island `eth1` is outside its scope — a red-proof
|
||
test locks that in (`TestParseMode_IslandStaticNICDoesNotConfuseEth0`) rather than changing the healer.
|
||
- Tests (all non-hollow, red-proofed): `TestBuildBringUpConfig_IslandNIC`,
|
||
`TestLocalAPIConfig_IslandValidation`, plus the healer scoping test above.
|
||
- **Coupling (deploy order):** a host-install that writes the island config REQUIRES agent ≥ 0.96.0 to
|
||
read `island_guest_addr` and attach net1 — vouch 0.96.0 before island installs go live. Migration of
|
||
existing boxes is the Phase-B runbook (`felhom.eu/documentation/runbooks/RUNBOOK-island-migration.md`).
|
||
|
||
## v0.95.0 — SMART coverage: union-path drives + LVM/dm root + device model (2026-07-25)
|
||
|
||
Additive; **MinAgent unchanged**; hub untouched (unknown JSON fields ignored). Implements the graded
|
||
fixes from `felhom.eu/documentation/audits/SPIKE-smart-coverage-2026-07-25.md`, which proved both demo
|
||
disks answer the allowlisted `smartctl -a -j` with PASSED but the agent never asked.
|
||
|
||
- **Fix B — union-path SMART:** registry/USB drives ride the `/disks` union path (`driveTargets.Known`),
|
||
which skips Observe's `enrich`, so they showed "Nincs adat" despite working SMART. New
|
||
`storage.SmartReader` (`SMARTForBacking`, reuses `smartDeviceFor`) is wired into the union path via a
|
||
localapi seam — the USB drive now reports its real verdict. The watchdog `Known` path stays
|
||
enrich-free (asserted: zero smartctl calls).
|
||
- **Fix A — LVM/device-mapper resolution:** `smartDeviceFor` gains a dm branch that resolves
|
||
`/dev/dm-N` / `/dev/mapper/X` to the single backing whole disk via `/sys/block/<dm>/slaves`
|
||
(recursive; **skips** rather than guesses when slaves span >1 physical disk). And the builtin `local`
|
||
dir on the LVM root — whose `backing_device` is empty by design (removable-safety) — now gets a
|
||
**SMART-only** device resolved from its containing filesystem (mount table), never touching
|
||
`backing_device`/`durable_id`. The system SSD stops reading "Nincs adat".
|
||
- **Device model:** `SmartSummary.ModelName` captured from smartctl's own `model_name` (already parsed),
|
||
so the controller card can label a disk "TOSHIBA MQ04ABF100" instead of a raw UUID.
|
||
- Fix C (`-d sat`) stays rejected (disproven live; absent from sudoers). No sudoers/manifest change.
|
||
Consumed by controller v0.171.0. Tests: dm-resolution table (+multi-disk-skip red-proof), containing-fs
|
||
resolution (+red-proof), union-path SMART (+red-proof), model capture, Known-path-never-SMARTs.
|
||
|
||
## v0.94.0 — serialize per-disk SMART into the /disks payload (2026-07-24)
|
||
|
||
Additive, backward-compatible; **MinAgent floor unchanged** (the controller feature-detects by payload
|
||
presence). No new smartctl load, no new endpoint, no sudoers change — the SMART is already computed on
|
||
the request path (`storage.Observe` → `enrich` runs `smartctl -a -j` for dir-backed targets); this
|
||
release simply copies the target's already-populated `Smart` into `localapi.DiskInfo`.
|
||
|
||
- `DiskInfo` gains `Smart *hub.SmartSummary \`json:"smart,omitempty"\``. In `handleDisks` the summary is
|
||
copied **only when `t.Smart.Health != ""`** — a zero-value summary (SMART never read: no smartctl
|
||
device, USB bridge, or a failed read) stays omitted, so the controller sees *absent* and renders
|
||
"Nincs adat" rather than a misleading UNKNOWN. The union-in drives (driveTargets.Known path, no
|
||
enrichment) carry no SMART and are omitted by the same guard.
|
||
- The controller (v0.169.0) consumes this to render a "Lemezek állapota" card + a 6-hourly
|
||
degradation notification. Old controllers ignore the extra field.
|
||
- Test `TestDisks_SmartSerialized` (payload includes SATA counters + temperature for a fixture target;
|
||
absent-SMART target omits the field); red-proof: drop the copy → the serialized-Smart assertion fails.
|
||
|
||
## v0.93.0 — a recovery code can no longer contain a hyphenated word (2026-07-21)
|
||
|
||
> **SHIPPED 2026-07-22:** built + published (sha256 `a68b2ff73200622e…`), Day-0-manifest-vouched,
|
||
> deployed to both fleet boxes (felhom-pve + demo-hp), clean-restart verified. Record:
|
||
> `felhom.eu/documentation/pilot/RUNBOOK-publish-agent-0.93-2026-07-22.md`.
|
||
|
||
**Generation-only change. Every recovery code already issued remains valid** — R is consumed as a
|
||
whole passphrase by the PBS scrypt KDF (`Wrap`/`Unwrap`) and is never re-split, so nothing about
|
||
verification moves. Nothing in the KDF/consume path, the word count, or the joiner changed.
|
||
|
||
The EFF large wordlist contains exactly four entries that themselves contain the hyphen we join
|
||
words with: `drop-down`, `felt-tip`, `t-shirt`, `yo-yo`. Drawing one produced a code that reads as
|
||
11 words rather than 10 — ambiguous to transcribe in exactly the situation R exists for, a customer
|
||
reading a code back during a disaster. The generator now draws from the list filtered of those four
|
||
(`joinSafe`), so a code always segments back into exactly `RecoveryCodeWords`.
|
||
|
||
- **Entropy floor holds, with the numbers asserted in the tests:** the draw space goes 7776 → 7772,
|
||
so a 10-word code goes 129.248 → 129.241 bits. The cost is 0.007 bits against a 128-bit floor.
|
||
- **The long-standing ~1/5 test flake was this defect, not a flaky test.**
|
||
`TestGenerateRecoveryCode_EntropyAndFormat` counted words by splitting the joined string, which
|
||
conflates "how many words were drawn" with "how many segments the code has". It now counts what
|
||
the generator drew, and asserts the segmentation property separately — the property `joinSafe`
|
||
actually buys. (`felhom.eu/REPORT.md` §6 item 3 recorded it at 3/8 in one session.)
|
||
- **Deterministic red-proof, in-tree:** `TestGeneratedCodeSegments_FilteredVsUnfiltered` drives the
|
||
generator against a fixture list where every word is hyphenated, so the pre-fix defect reproduces
|
||
with probability 1 instead of ~1/5, and shows the same list through `joinSafe` refuses to generate.
|
||
- New for audit: `WordlistFilteredOut()`, `RecoveryCodeSep`. `WordlistSize()` now reports the
|
||
effective (filtered) draw space, 7772.
|
||
|
||
## v0.92.1 — ship the guestnet sudoers grant with the binary (supersedes v0.92.0) (2026-07-21)
|
||
|
||
**Supersedes v0.92.0; that artifact is materially incomplete — do not vouch it.** It was published
|
||
before live verification revealed the watchdog had no sudoers grant for three of its four probes,
|
||
so it carries neither the `FELHOM_GUESTNET` alias nor the `guestnet-*` capability rows that make a
|
||
host missing that alias visible. Left in place rather than overwritten — a published version stays
|
||
immutable (the v0.91.0 → v0.91.1 precedent).
|
||
|
||
No behaviour change beyond the capability rows; the watchdog code is byte-identical to v0.92.0. The
|
||
functional fix is `configs/felhom-agent.sudoers`, which **must be deployed with the binary**.
|
||
|
||
## v0.92.0 — the guest network gets a watchdog (R-54) (2026-07-21)
|
||
|
||
**Host-tier only — no controller coupling, no wire change the hub must understand today** (the
|
||
`guest_net` stanza is additive and stored opaquely, exactly like `pbs_dr` and `wireguard`).
|
||
|
||
Closes the OPEN RISK left by `INCIDENT-guest-dhclient-killed-2026-07-20` §5: **the guest's DHCP
|
||
client is started once by ifupdown at boot and nothing supervises it.** When it died on 2026-07-20
|
||
the guest kept working for another ~80 minutes on its unexpired lease; only when the lease expired
|
||
did the address and the default route vanish, taking the Cloudflare tunnel, the hub reports, the
|
||
catalog sync and the controller→agent channel with them — a 1h15m outage in which every observable
|
||
signal said healthy for the first 80 minutes.
|
||
|
||
**The design consequence, and the point of the whole package: liveness of the DHCP client is itself
|
||
a probe.** Waiting for the address to disappear is waiting out precisely that silent window. The
|
||
watchdog therefore flags a DHCP guest unhealthy on `pgrep -x dhclient` alone, while the lease is
|
||
still live and everything else still looks perfect.
|
||
|
||
`internal/guestnet`, built on the wg-tunnel/storage watchdog loop shape:
|
||
|
||
- **Probes** (four fixed-shape `pct exec` argvs, no shell anywhere, no guest data interpolated):
|
||
address, default route, `/etc/network/interfaces` mode, dhclient liveness. Parsers are pinned to
|
||
output captured live from guest 9201 on 2026-07-21 — including the literal backslash `ip -o`
|
||
emits and the docker-bridge routes that must not read as a default route.
|
||
- **Heals** with the incident's restored invocation, verbatim, logged at INFO before it runs:
|
||
`pct exec <vmid> -- dhclient -pf /run/dhclient.eth0.pid -lf /var/lib/dhcp/dhclient.eth0.leases eth0`
|
||
- **Dampers**, because this runs a privileged command inside a customer's container: two CONSECUTIVE
|
||
bad probes before any heal (one blip is not a diagnosis), ≥10 min between heals per guest, ≤3
|
||
heals/hour, and an observe-only window while the guest (or the agent) has been up under 3 minutes.
|
||
- **Refuses to act** on a static guest (dhclient must never fight a static config — a static guest
|
||
missing its address is reported loudly and left to R-50), on an unknown interface mode, on a guest
|
||
it cannot probe, and when the guest list cannot be ownership-proven. The guest source is the
|
||
pool-verified one (`ListLXC` ∩ the felhom pool, audit A1) — never a bare `ListLXC`, which under a
|
||
broad token would run dhclient inside a co-tenant's container.
|
||
- **A failed PROBE never reads as a dead client.** `pgrep` exits 1 with empty stderr when there is
|
||
no match; anything on stderr means the probe itself failed, which reads as unknown. Otherwise a
|
||
missing `pgrep` would heal forever.
|
||
- **Healthy cycles log a Debug line** (v0.91.2's lesson, one day old): if the quiet path is silent,
|
||
"no alarms" and "never probed" are the same evidence, and an inert watchdog is indistinguishable
|
||
from a working one.
|
||
- **Deliberately NOT in the `errc` fan-out** — a watchdog over customer guests must never be able to
|
||
terminate the agent. A test asserts that, because joining the fan-out would also make the shutdown
|
||
drain bound off by one.
|
||
|
||
**Config `guest_net` is this repo's first DEFAULT-ON feature gate, and the inversion is deliberate.**
|
||
Every other gate defaults to false because those features reach outward (an offsite endpoint, an OOB
|
||
tunnel) and enrolling a box by an update would be wrong. This one looks only INWARD at guests the
|
||
agent already owns, and the failure it prevents exists on every box today. A watchdog that must be
|
||
remembered per box is a watchdog that is missing on the box that needed it. Opting out is the
|
||
explicit act: `"guest_net": {"disable": true}`.
|
||
|
||
**Naming deviation from the spec, deliberate:** TASK-D called the report block `WireGuestNet`. In
|
||
this repo `Wire*` is the DOWN direction (`WireDesiredState` / `WirePBSDR` — what the hub sends), and
|
||
UP-direction report stanzas are `*Status`. It ships as `GuestNetStatus` so it is not the one report
|
||
block named against the convention.
|
||
|
||
- Wiring is asserted from `package main` by an AST walk (construct + `SetGuestNetReporter` + the
|
||
started goroutine) — the v0.91.0 defect was exactly a seam whose caller was never written, with
|
||
every unit test green. Red-proof: un-wiring both lines fails the test with both reasons named.
|
||
- Red-proof for the detection itself: reverting `classify` to IP-presence-only makes the July-20
|
||
fixture report **"healthy"** and records **zero** heals — the 80-minute silent window, reproduced.
|
||
- **Ships a sudoers change** (`configs/felhom-agent.sudoers` MUST be deployed with the binary).
|
||
TASK-D assumed none was needed; live verification on felhom-pve proved otherwise — the first sweep
|
||
logged `dhclient liveness probe failed: sudo: a password is required` and correctly reported
|
||
`state=unknown` rather than acting blind. The existing grant covered only lanresolver's address
|
||
read. `FELHOM_GUESTNET` adds four fixed vectors (route, interfaces, pgrep, and the heal), every
|
||
argument after the numeric vmid a literal, so no hub or guest input can widen it. The address read
|
||
is not duplicated — it stays FELHOM_DNSMASQ's.
|
||
- Four `guestnet-*` capability rows so a host missing that sudoers file is VISIBLE as degraded
|
||
rather than silently watchdog-less. Non-critical on purpose: a missing grant must not page an
|
||
operator for every box on rollout day (the R-50b amber-fleet lesson).
|
||
- `var version` in main.go was stale at `0.89.0` (three releases behind); builds set it via ldflags,
|
||
but `go run` and any forgotten `-X` reported a version that had not existed for days.
|
||
|
||
## v0.91.2 — a healthy credential probe is observable (2026-07-21)
|
||
|
||
The probe logged only on failure, so a healthy one was silent — which makes "no `auth_failed`"
|
||
indistinguishable from "never probed", and leaves the leg impossible to demonstrate as running. That
|
||
is precisely how v0.91.0 shipped it inert without anyone noticing. Adds a Debug line on success
|
||
naming the storage: free in normal operation, one log level away when it matters.
|
||
|
||
## v0.91.1 — wire the credential probe (v0.91.0 shipped the seam inert) (2026-07-21)
|
||
|
||
**Supersedes v0.91.0; that artifact is materially incomplete — do not vouch it.**
|
||
|
||
v0.91.0 built the `pbs.AuthSink` seam and the `NoteAuthResult` consumer, and `main.go` never called
|
||
`SetAuthSink`. The reporter deliberately skips probing when no sink is attached, so the whole
|
||
auth-honesty leg was silently inert: no probe, no `auth_failed`, no self-heal — and nothing failed,
|
||
because every unit test injected the sink directly.
|
||
|
||
Caught during STOP-1 live verification by checking the wiring rather than trusting it. **Exactly the
|
||
same class as the controller v0.154.0 defect the day before: a table test over a seam proves the
|
||
seam, not the caller.** The published 0.91.0 artifact is left in place and superseded rather than
|
||
overwritten — a published version must stay immutable.
|
||
|
||
- `main.go`: `pbsReporter.SetAuthSink(pdMgr)` in the pbsdr bridge block.
|
||
- `TestLiveReporter_NoSinkMeansNoProbe` pins the no-sink-no-probe contract, so the inert case stays a
|
||
documented behaviour rather than an accident nobody notices twice. The wiring itself is proven by
|
||
the live STOP-1/STOP-2 evidence, which is the only thing that can prove it.
|
||
|
||
## v0.91.0 — the DR tier can no longer be `applied` and dead at the same time (R-39 fleet fix + R-50b(a)) (2026-07-21)
|
||
|
||
**Requires hub >= v0.68.0** for the re-arm signal. Hub v0.68.0 is safe for 0.90.0 agents (they drop
|
||
the unknown descriptor key), but the guarantees below need THIS agent. **MinAgent → 0.91.0 is an
|
||
operator manifest save, sequenced after the fleet has self-updated — not a code change.**
|
||
|
||
### What was broken
|
||
|
||
Three compounding defects let a box report `applied` while every PBS request 401'd:
|
||
|
||
1. **The re-key was invisible.** An ep0 re-issue rotates the SECRET of an existing token, so
|
||
`token_id`, `fingerprint`, `datastore` and `namespace` all come back byte-identical. The agent
|
||
re-applies on the descriptor's CONTENT HASH, so a converged box short-circuited and never consumed
|
||
the fresh secret. Proof from the N100: `consumed-failed.json` carried a hash byte-identical to the
|
||
`marker.json` written two minutes before the re-issue.
|
||
2. **The agent could not read its own credential.** It WRITES
|
||
`/etc/pve/priv/storage/<id>.pw` through the root wrapper, but that directory is `0700 root:www-data`
|
||
and the wrapper had no read verb — so `pbsTargetsFromPVE` got "permission denied" every cycle,
|
||
logged a Warn and skipped the datastore. The one loop that could have caught the 401 was blind **by
|
||
construction**.
|
||
3. **Nothing probed authentication.** A snapshot list that fails with 401 looked exactly like "PBS is
|
||
busy".
|
||
|
||
### The fix
|
||
|
||
- **`WirePBSDR.SecretGeneration`** — field-exact with the hub's descriptor. Because `descriptorHash`
|
||
marshals this struct, the hub's monotonic mint counter is what finally moves the hash and re-arms a
|
||
converged agent.
|
||
- **Wrapper `read` verb** (+ exactly ONE sudoers line, + a `pbsdr-read` capability row). Prints one
|
||
secret to STDOUT and nothing else: no network, no mutation, no logging of the value, and the secret
|
||
never rides argv (sudo logs argv). Traversal is refused three times over — the id grammar admits no
|
||
slash, `val_sdir` pins the directory, and the RESOLVED path is prefix-asserted.
|
||
- **`pbs.ProbeAuth`** — `GET /version`, the cheapest authenticated question, with a distinct
|
||
`ErrUnauthorized` sentinel. `/version` needs no datastore, namespace or privilege, so a 401 there
|
||
means the CREDENTIAL is bad — not that an ACL is narrow. **403 is deliberately NOT treated as
|
||
unauthorized**: re-keying a too-narrow token would mint credentials forever without fixing anything.
|
||
- **The probe runs on the 15-minute collect path** (not the 6 h verify cadence) and its verdict
|
||
becomes a LOUD `auth_failed` state the hub's pbsdrheal escalates to a fresh mint. **A transport
|
||
error is UNKNOWN, never a rejection** — otherwise every network blip would burn a credential.
|
||
Recovery is self-clearing.
|
||
- **`readPBSSecret`** now prefers a directly-readable file and falls back to the wrapper, so a box
|
||
with its own agent-owned secret dir needs no sudo at all.
|
||
- **R-50b(a):** the report carries the installed wrapper's sha256, so drift against the vouched
|
||
manifest value is finally answerable. Empty = unknown, never drift.
|
||
|
||
### Tests
|
||
|
||
Scenario A (a re-key re-arms a converged agent) with the hash mechanism asserted separately; the
|
||
unknown-field compat direction; auth_failed loud/ignored-for-other-storage/none-before-descriptor/
|
||
self-clearing; and the wrapper `read` verb executed under real bash — traversal refusals, secret to
|
||
stdout only, missing-file refusal, side-effect freedom.
|
||
|
||
**Three red-proofs run at the assertion level.** Removing `SecretGeneration` makes the re-arm test
|
||
fail with `consume calls=1, want 2`. Swallowing the probe result leaves `State:applied
|
||
AuthFailed:false` — the July-18 shape exactly. Deleting the id charset guard alone does **not** open
|
||
a traversal hole (readlink + the prefix assertion still catch it), so the isolating red-proof removes
|
||
the charset guard AND the prefix assertion and shows the out-of-tree secret printed.
|
||
|
||
## docs — the workflow moved to DooPlex-local execution (2026-07-19)
|
||
|
||
**Docs only, no version bump, no code change.** Claude Code now runs on DooPlex (192.168.0.180,
|
||
Debian 13, `kisfenyo`) instead of the Windows workstation, working directly in
|
||
`/mnt/5_hdd/felhom.eu/git/felhom-agent`.
|
||
|
||
- `CLAUDE.md` — the build/deploy table is now local-first: build on DooPlex, then **one** hop
|
||
`scp /tmp/felhom-agent-<v> felhom-pve:/tmp/`. The old path went build-server → Windows box →
|
||
felhom-pve, needing `cygpath -w` for the local scp path; that two-hop detour and its CRLF hazard
|
||
are gone (recorded in the new "Legacy: Windows workstation" note, not deleted).
|
||
- **New clean-tree gate** before any build: `git status --porcelain` empty AND `HEAD` ==
|
||
`origin/main` — the CC working tree is now the tree that gets built. An unpushed change does not
|
||
exist.
|
||
- `MSYS_NO_PATHCONV=1` is no longer needed for `pct` (it was an MSYS path-mangling workaround);
|
||
`ssh felhom-pve` is plain now. Selftests run locally on DooPlex against the demo API.
|
||
- Workspace-root pointer updated to `/mnt/5_hdd/felhom.eu/git/CLAUDE.md`.
|
||
|
||
Historical Windows references in past CHANGELOG entries and `PLAN.md` are left untouched.
|
||
|
||
## build tooling — the golden bakes EVERY infra image, asked from the controller (2026-07-19)
|
||
|
||
**No agent version bump: `configs/build-golden.sh` only (v2.0.0 → v2.1.0). Effective at the NEXT
|
||
golden build — the current golden is NOT rebuilt for this.**
|
||
|
||
- **The bug, observed live twice.** Enabling Megosztás on a fresh box pulled `felhom-samba` from the
|
||
registry with zero feedback: minutes of silent nothing. Cause: this script carried its own
|
||
hand-maintained array of three image tags, with a comment instructing the reader to keep it in sync
|
||
with the controller's `internal/infra` constants. It drifted the moment a fourth stack was added —
|
||
`felhom-samba` was never added here, so the golden baked **3 of 4**.
|
||
- **The fix is structural, not another copy.** The list now comes from the controller image the bake
|
||
just pulled: `docker run --rm <controller> --print-infra-images` (backed by `infra.Images()`, which
|
||
derives from the pins themselves). The golden therefore bakes exactly what **that** controller
|
||
version will request, and the two cannot disagree by construction. A controller-side test parses
|
||
the const block out of the source and fails if a pin is added without reaching `Images()`.
|
||
- **Ordering fix that this exposed.** `docker logout` + `config.json` removal ran immediately after
|
||
the controller pull. `felhom-samba` lives on the **same private registry**, so the infra loop would
|
||
have 401'd. The logout moved to **after** the loop, with an added hard assertion that no credential
|
||
remains in the guest before it is archived — the credential is still never baked.
|
||
- **Fallback, loudly.** A controller older than v0.147.0 has no `--print-infra-images`; the bake falls
|
||
back to the historical 3-image list and prints three WARN lines saying felhom-samba will not be
|
||
baked and Megosztás will pull at runtime. The fallback is exactly the drift-prone thing this change
|
||
removes, so it announces itself rather than passing silently.
|
||
- **ROADMAP:** golden **≥ 0.147.x** carries all four infra images.
|
||
|
||
## v0.90.1 — R-39 hotfix: PBS reconcile must not pass `--server` to `pvesm set` (2026-07-18)
|
||
|
||
**Config-only fix (wrapper + red-proof); the Go binary is unchanged.** Ship the wrapper with the
|
||
next agent deploy — `configs/felhom-pbs-apply` is a shipped artifact, not a build input.
|
||
Green: `go build ./... && go vet ./... && go test ./...` all pass.
|
||
|
||
- **The defect.** The `reconcile` verb built its argv as
|
||
`args=(--server "$server" --fingerprint "$fp")`. PVE treats a PBS storage's `server` as a
|
||
**create-only** parameter and rejects the ENTIRE `pvesm set` call —
|
||
`can't change value of fixed parameter 'server'` — **even when the value passed is byte-identical
|
||
to the stored one**. So `reconcile` could never succeed against an existing entry; it exited 255
|
||
every time.
|
||
|
||
- **Why that was severe rather than cosmetic.** The agent consumes the hub's **one-time** PBS token
|
||
secret BEFORE invoking the wrapper. A wrapper failure therefore *burned* the credential: each hub
|
||
"Re-issue PBS credentials" minted a fresh secret, the agent consumed it, the wrapper rejected the
|
||
apply, and the storage entry stayed pinned to the **revoked** one. Observable end state: the PBS DR
|
||
tier authenticating **401 Unauthorized** indefinitely while the agent reported
|
||
`pbsdr: converged state=applied`. Live-diagnosed on the N100 demo host during the 2026-07-18
|
||
rehearsal wrap (`felhom.eu/documentation/tests/VALIDATION-n100-rehearsal-2026-07-18.md`, finding
|
||
F2 / ROADMAP **R-39**).
|
||
|
||
- **The fix.** Drop `--server` from the reconcile argv. The server address is immutable by
|
||
construction — relocating a PBS endpoint requires a fresh `create` — so there was never anything
|
||
for `reconcile` to reconcile there. `--fingerprint` (and `--password` when a secret is fed) remain,
|
||
which is the mutable identity the verb exists to push. Proven on the live box before committing:
|
||
`pvesm set felhom-pbs --server <same> --fingerprint <same>` → rejected;
|
||
the same call without `--server` → **rc 0**.
|
||
|
||
- **Red-proof** `TestReconcileNeverPassesServerToPvesmSet` (`internal/pbsdr/manager_test.go`):
|
||
isolates the `reconcile)` block from the shipped wrapper and asserts no `--server` reaches
|
||
`pvesm set`, plus that `--fingerprint` is still pushed (so the verb can't be hollowed out).
|
||
Verified RED against the unfixed wrapper and GREEN after the fix. Two traps the proof handles
|
||
explicitly: the pattern is **line-ending tolerant** (`\r?\n`) because this repo is cloned on
|
||
Windows and an `\n`-only pattern would match nothing and pass **vacuously**; and comment lines are
|
||
**stripped before matching**, because the WHY note above the fix necessarily quotes the very flag
|
||
the test forbids.
|
||
|
||
- **NOT fixed here (deliberately, and each still open):**
|
||
1. **R-39's primary half** — the agent re-applies on a change of the *descriptor hash*
|
||
(`manager.go` ~L235), but a hub credential re-issue leaves the descriptor byte-identical
|
||
(same `token_id`, same `fingerprint`; only the side-table secret rotates) and bumps only the
|
||
generation. So a converged agent still ignores a fresh secret. This wrapper fix means the apply
|
||
now *succeeds* once the agent is made to re-apply; it does not make it re-apply.
|
||
2. **The verify-loop read** — `pbs: cannot read token secret … permission denied`: the non-root
|
||
agent reads `/etc/pve/priv/storage/<id>.pw` directly, a path it can only ever *write* through
|
||
the root wrapper (`/etc/pve/priv` is `0700 root:www-data`, and sudoers exposes only
|
||
`create|reconcile|grant` — there is no read verb). The loop is therefore permanently blind to
|
||
the failure it exists to catch.
|
||
Both ride the spec'd R-39 agent train.
|
||
|
||
## v0.90.0 — agent train: guest RAM resize (R-24) + fast-tick-until-convergence (R-28) (2026-07-17)
|
||
|
||
**MinAgent coupling:** felhom-controller v0.143.0 gates its guest-memory-resize UI on this agent
|
||
(`FeatureGuestMemoryResize`, MinAgent 0.90.0). One train, one floor raise (Viktor's ruling).
|
||
Green: `go build ./... && go vet ./... && go test ./...` all pass.
|
||
|
||
- **Item 1 — guest RAM resize (R-24, controller-direct)** (`internal/localapi/guestmemory.go`, new):
|
||
a new self-scoped local-API surface — `GET /guest/memory` (current allocation, live usage, and the
|
||
enforced bounds, all agent-computed in MB) and `POST /guest/memory` (bounded resize). The agent is
|
||
the security boundary: every bound is recomputed FRESH per request from a live read (the UI's numbers
|
||
are decoration). Ruled bounds — min **2048 MB**; max **host_total − 2048 MB** (host reserve); a SHRINK
|
||
is refused below **max(2048, usage + 512 MB)** with `below_usage_floor` (stop apps first). Applies via
|
||
the PVE API `SetConfig` — a **live cgroup apply, no reboot** (Phase-0 PROVEN on the nested demo box:
|
||
maxmem moves with the guest running, `/proc/meminfo` ripples via lxcfs). Verify-after-apply: the new
|
||
maxmem is re-read and must equal the target before success is claimed (a pending/reboot outcome is a
|
||
502, never a false success). Refusals return a machine `code` (below_min/above_max/below_usage_floor)
|
||
+ fresh bounds at 412; SetConfig is never called on a refusal path. Single-flight per host. Optional
|
||
`Options.Memory` (nil → 503 "not configured"); a NEW narrow `MemoryOps` interface (does not touch the
|
||
shared `GuestAPI`). Wired in `main.go` from the existing proxmox client. Memory only; cores stay
|
||
observation.
|
||
- **Item 2 — fast-tick-until-first-convergence (R-28)** (`internal/fasttick/`, new): the agent-plane
|
||
immediacy SECONDARY. While ANY desired-state item is unapplied — most importantly the **pre-tunnel
|
||
WG-registration window a hub poke cannot reach** — it pulses the SAME out-of-band report trigger the
|
||
watchdog/poke use, every **30 s**, and **self-disarms emergently** the instant everything converges.
|
||
State-based by ruling: nothing to journal, no timer to leak. Four cached sources (no exec/network per
|
||
tick): desired-generation == 0 (never fetched); reconcile `LastResult` actionable drift
|
||
(`Planned − Pending > 0` — a destructive `pending_signature` is EXPECTED, excluded); pbsdr
|
||
`waiting_secret` ONLY (the LOUD `consumed_failed`/`verify_failed` are excluded so a stuck box never
|
||
hammers); wgtunnel block-desired-but-not-operational. Supporting seams: `reconcile.Engine.LastResult()`
|
||
(mutex-recorded per pass) and `wgtunnel.Manager.TunnelConvergence()` (cached snapshot refreshed at each
|
||
Apply — the fast-tick never execs `wg`/`systemctl`). A perma-unconverged box fast-ticks at ~2 small
|
||
reports/min (bounded, documented; timers were deliberately rejected).
|
||
- **Guests-0/0 (diagnosis passenger):** REFUTED the pool-membership hypothesis on the live nested box —
|
||
guest 9201 IS a pool member, the agent token sees it (VM.Audit comes from the `/pool/felhom` grant),
|
||
hub reports 1/1. The observed 0/0 was the legitimate **pre-provision reporting window** (no guest
|
||
existed yet); the existing `PoolAddVMID` re-assert (bringup.go:498) already covers the known
|
||
restore-over-existing edge (campaign-2 R2). Item 2's fast-tick is exactly the window's mitigation.
|
||
No code change.
|
||
- **Tests + red-proofs (run-fail-restore):** localapi memory (grow/shrink, the three refusals each with
|
||
a **SetConfig-count == 0** assertion, cross-guest 403, fresh-bounds-per-request, verify-not-reflected
|
||
502, nil-config 503) — red-proofs (i) floor guard removed → below_usage_floor fails; (ii) max guard
|
||
removed → above_max fails. fasttick (pulse-while-unconverged, silent-when-converged, the ruled
|
||
disarm-on-convergence, full-channel non-blocking drop, first-reason) — red-proof (iii) always-pulse →
|
||
the silent + disarm tests fail. `Engine.LastResult` effect + pre-first-run ok=false. All restored green.
|
||
|
||
## v0.89.0 — agent train: PBS-DR self-grant (R-22) + escrow config live-reload + agent-plane poke listener (2026-07-16)
|
||
|
||
Three bundled agent-plane items. Green: `go build ./... && go vet ./... && go test ./...` all pass.
|
||
|
||
- **Item 1 — pbsdr self-grant (R-22, closes the F4 root cause from `tests/VALIDATION-n100`)**
|
||
(`internal/pbsdr/manager.go`): on a NON-DEFAULT PBS storage id the agent token holds no ACL on
|
||
`/storage/<id>` yet, so the reconcile tick's token-auth pre-check `StorageEntry` (GET
|
||
`/storage/<id>`) 403s and — pre-fix — aborted BEFORE the root-run wrapper `grant` that creates
|
||
that very ACL: a permanent self-deadlock that bit every legacy/descriptor-provisioned id (the
|
||
demo's `felhom-offsite`). Fix: on an `IsForbidden` pre-check error ONLY, run
|
||
`felhom-pbs-apply grant <id>` now (root, no secret, no pre-existing entry — `pveum acl modify` on a
|
||
path is unconditional), re-read once, then flow the normal adoption/create path. The pre-check is
|
||
KEPT (once the ACL exists it short-circuits the happy path cheaply); every other error stays
|
||
transient. Grant-then-still-failing surfaces `verify_failed` loudly, never a silent retry storm.
|
||
Red-proof: `TestSelfGrant_PreCheck403DoesNotAbortBeforeGrant` (fake 403 on the pre-check → pre-fix
|
||
never reaches the grant, calls=[] → FAIL; fixed → self-grant + converge, no secret consumed).
|
||
- **Item 2 — escrow config live-reload** (`internal/localapi/escrow_ceremony.go`,
|
||
`cmd/felhom-agent/main.go`): the pbsdr bridge seeds `escrow.pbs_storage_id` into agent.json on DR
|
||
convergence, but `/escrow/preflight` read a daemon-start snapshot and stayed red until a service
|
||
restart. New late-bound `EscrowCeremonyConfig.CurrentPBSStorageID` resolver (mirrors the existing
|
||
`DRConfigured func() bool` pattern) re-reads config from disk at preflight time — exactly what the
|
||
ceremony subprocess itself loads, so the row flips green in-process, no restart. Falls back to the
|
||
boot snapshot on read error / all-env config. Red-proof:
|
||
`TestEscrowPreflight_PBSStorageIDLiveReload` (seed after boot → pre-fix stays red → FAIL).
|
||
- **Item 3 — agent-plane poke listener (Direction-2a, `SPIKE-immediate-sync-transport-2026-07-16`)**
|
||
(new `internal/poke`, `internal/wgtunnel.LoadAssignedAddr`, `cmd/felhom-agent/main.go`,
|
||
`internal/hub/loop.go`): a CONTENTLESS UDP poke relayed hub → ep0 forced-command → wg0-origin →
|
||
this listener fires ONE immediate desired-state cycle (the hub control loop's out-of-band report
|
||
trigger — the SAME channel the storage watchdog uses; fan-in, cap-1 coalescing). The socket binds
|
||
EXCLUSIVELY to the box's WG /32 (`registered.json`; never 0.0.0.0, never the LAN iface) on the
|
||
fixed port **51822**; payload is ignored entirely (a forged/replayed poke costs at most one extra
|
||
debounced tick); leading-edge debounce (`DebounceWindow`) coalesces a burst into ≤1 tick. Enabled
|
||
whenever `wg_tunnel.enabled`; a lost poke is harmless (the 15-min cycle still reconciles). This is
|
||
the FIRST concrete slice of the OOB/mutual-repair arc (R-13) — the listener+trigger only, nothing
|
||
more. Red-proofs: `TestBindConfinement` (wildcard bind → bound to `::` → FAIL) and
|
||
`TestDebounceCoalescesBurst` (guard removed → 10 fires for 10 pokes → FAIL). Port registered in
|
||
REUSE.md as a shared cross-repo constant (hub sender + ep0 `felhom-poke`).
|
||
|
||
## v0.88.0 — controller-driven escrow ceremony: --output=json + localapi job + one-shot R claim (2026-07-13)
|
||
|
||
The agent half of the customer-facing recovery-code wizard (controller v0.127.0; every mechanism
|
||
validated by felhom.eu SPIKE-controller-escrow-2026-07-13 — PTY-under-no-TTY, fixed-argv sudoers
|
||
5/5 refusals, R pipe round-trip, env_reset, 2.3–2.4 s timings). Operator ruling F1 (2026-07-13):
|
||
R transiting the CF tunnel once at display is an accepted, documented risk (threat-model paragraph
|
||
in RUNBOOK-escrow-ceremony.md).
|
||
|
||
- **`--output=json` machine mode** (`cmd/felhom-agent/main.go`): `runSelftestEscrowCreate`'s body
|
||
extracted into the shared `escrowCeremony()` core; text mode stays BYTE-IDENTICAL (banner, R
|
||
block, exit codes 0/1/2, upload-fail-after-R order). json mode emits ONE
|
||
`escrow.CeremonyOutput` object on stdout (version 1: recovery_code, key_fingerprint,
|
||
entropy_bits, blob/identity sizes, restic_pw_sealed, uploaded), every human line to stderr, no
|
||
partial JSON on failure; `--offline`/`--paperkey` are refused in json mode (print-oriented).
|
||
- **The ONE fixed argv** (`internal/escrow/ceremony.go`): `escrow.CeremonyBinary` +
|
||
`escrow.CeremonyArgs()` — the single source shared by the localapi exec, the capability
|
||
manifest entry, and (byte-identically) the new `FELHOM_ESCROW` sudoers alias
|
||
(`configs/felhom-agent.sudoers`). `TestEscrowCeremonyArgvPinned` +
|
||
`TestManifestCoveredBySudoers` transitively lock runner == manifest == sudoers; never build the
|
||
argv with flag helpers, never normalize `--`→`-` (spike §2.2).
|
||
- **localapi ceremony endpoints** (`internal/localapi/escrow_ceremony.go`, `withGuest`-wrapped):
|
||
`POST /escrow/ceremony` (single-flight 409; detached job, 60 s timeout; `sudo -n` + the fixed
|
||
argv; stdout parsed then zeroed — SECRET-BEARING, never logged), `GET /escrow/ceremony/status`
|
||
(non-secret summary + stderr-tail failure detail ≤500; R structurally absent from the job
|
||
struct), `POST /escrow/ceremony/claim` (ONE-SHOT: 200 `{recovery_code}` once → holder zeroed;
|
||
410 on re-claim; **TTL 10 min** → `unclaimed_void`, active AfterFunc belt + lazy check),
|
||
`GET /escrow/preflight` (storage id, DR tier, age, hub target, staged-secret informational,
|
||
`sudo -n -l` grant list-probe). Crash-safety is IN-MEMORY BY DESIGN — an agent restart loses R
|
||
safely (re-run supersedes); no journal, deliberately.
|
||
- **Capability** `escrow-ceremony` (Critical, `GatedBy: pbs_dr` EXPLICIT — non-pbsdr name by
|
||
decision): list-mode probe of the shared argv; inactive (never red) while the DR tier is off.
|
||
- Tests: one-shot claim + double-claim 410, TTL void + zeroed holder, R-substring absent from
|
||
every status/snapshot payload (incl. the serialized job struct), single-flight, supersede on
|
||
re-run, failure taxonomy (exit/unparseable/version), preflight truth table, argv pin. §10
|
||
red-proofs demonstrated (see felhom.eu REPORT).
|
||
|
||
## v0.87.0 — SystemDisks device-mapper walk: legacy-boot hosts get a working drive wizard (IA finding 2, MEDIUM) (2026-07-13)
|
||
|
||
On a legacy-boot PVE (LVM root, no mounted ESP) `SystemDisks` resolved NOTHING — `wholeDiskOf`
|
||
stops at `/dev/mapper/pve-root` — so the all-system fail-safe classified EVERY disk system and
|
||
the drive wizard could never offer a candidate (the IA validation's hot-added 5 GB disk stayed
|
||
invisible on the drill box). Operator ruling 2026-07-13: walk the root's backing device through
|
||
`/sys/block/<dev>/slaves` recursively down to physical disks (dm AND md; topology, never VG
|
||
names); those + any mounted-ESP holder are system; the all-system fail-safe returns to being the
|
||
WALK-FAILURE error case only. SAFETY DIRECTION: the outcome made impossible is a root-backing
|
||
disk classified candidate — per-branch conservatism means ANY unresolvable slave fails the whole
|
||
walk (`ok=false` → all-system, the unchanged code path).
|
||
|
||
- **`physicalDisksOf` + `walkSlaves`** (`internal/storage/role.go`): symlink-canonicalize →
|
||
fast-path `wholeDiskOf` (raw disks/partitions, unchanged) → recursive sysfs slaves walk for
|
||
virtual devices, with cycle/depth guard; non-`/dev` sources (ZFS datasets, NFS, overlay) stay
|
||
unwalkable → fail-safe. Partition slaves resolve via the existing regexes (`sda3 → /dev/sda`);
|
||
exotic partition names the regexes don't cover (e.g. `md0p1`) fail the walk → fail-safe.
|
||
- **`HostReader.BlockSlaves(name)`** — the ONE new seam method (root-free: sysfs is
|
||
world-readable): slaves list + hasDir; production reads `/sys/block/<name>/slaves`; every test
|
||
fake mirrors it. Live-probed (§3): physical disks carry an EMPTY slaves dir — hasDir alone is
|
||
not resolution; a virtual device with an empty/unlistable slaves dir fails the walk.
|
||
- **Behavior deltas:** legacy-boot LVM hosts now resolve (`{root's physical parents}`, ok=true)
|
||
— the wizard lives (still behind the untouched data-bearing/claim guards); EFI(+LVM) hosts are
|
||
byte-identical (ESP and walk agree on the same disk — §3 felhom-pve transcript + fixture);
|
||
md-raid roots mark BOTH members system. NEW protective delta: a host whose ROOT topology
|
||
cannot be fully grounded is now all-system even if an ESP resolved (pre-fix it silently
|
||
trusted the ESP alone; ruling's per-disk conservatism).
|
||
- **Tests** (`role_walk_test.go`, fake-sysfs fixtures mirroring the §3 transcripts): the
|
||
SIGNATURE table (root-backing disk(s) ALWAYS in the system set across legacy-LVM / md-raid /
|
||
EFI+raw / EFI+LVM / nested dm-on-md — never weaken), dead-wizard-lives (A), dangling-slave
|
||
fail-safe through the REAL sysKnown=false path (D), cycle guard, empty-slaves. Red-proofs
|
||
recorded: A pre-fix resolver → dead wizard reproduced; B walk-returns-dm-node → SIGNATURE
|
||
VIOLATION naming the missing disk; D conservatism removed → "scratch classified candidate
|
||
while walk incomplete".
|
||
- Format/mkfs paths, data-bearing guards, wizard UI: UNTOUCHED. MinAgent/controller coupling:
|
||
none (agent-internal classification).
|
||
|
||
## v0.86.0 — DR-tier-by-default: capability `inactive` state + F-3 provision-parent ownership (2026-07-12)
|
||
|
||
Agent half of the DR-tier-by-default batch (DRILL-day0-vm-2026-07-12; operator decisions: DR
|
||
capability is BAKED on every install, activation is a hub flag, disabled ≠ degraded).
|
||
|
||
- **Capability `inactive` state** (`internal/capability`): a third status next to ok/degraded —
|
||
a config-GATED capability whose plumbing is HEALTHY (binary present, sudo granted) but whose
|
||
feature is off reports `inactive` / reason `disabled by configuration`. The 3 `pbsdr-*` entries
|
||
are gated (`GatedBy=GatePBSDR`, applied by the stable name prefix in `Manifest()`); broken
|
||
plumbing (binary missing / grant denied) stays DEGRADED even with the gate off — an un-migrated
|
||
pre-v1.15.0 box must never look deliberately disabled. `Summarize` counts only real degraded
|
||
(inactive never error-logs); the startup self-check logs an `inactive` count and now runs AFTER
|
||
the pbsdr gate wiring so its snapshot matches the first report.
|
||
- **`pbsdr.Manager.DRConfigured()`** — the gate's answer: true when the last-seen descriptor was
|
||
enabled (any live state except `disabled`); before the first desired-state fetch it falls back
|
||
to the persisted converged marker, so an applied box never flaps to inactive across a restart.
|
||
- **F-3 — provision parent-dir ownership** (`internal/provision/backhalf.go`): a ROOT-run
|
||
provision (the Day-0 one-shot) now chowns the just-created `guests/` + `guests/<vmid>/` PARENT
|
||
dirs to the state-dir's owner (`chown --reference`, NON-recursive — the bootstrap leaf stays
|
||
the mapped guest-root's). Previously they were left root:root 0700 → the non-root daemon's
|
||
lanresolver got "permission denied" (drill live-fix now also applied to felhom-pve, which had
|
||
the same latent state; Peti's host unreachable — deferred). A daemon-run (non-root) provision
|
||
skips it (`geteuid` seam).
|
||
- Tests + red-proofs: gate-off-healthy→inactive / gate-off-broken→degraded / gate-on→ok /
|
||
exactly-pbsdr-gated; DRConfigured lifecycle (incl. marker-across-restart + disabled-wins);
|
||
root-run parent chown issued, non-root not, never recursive. All three mutations proven red.
|
||
- Shipping note: `configs/felhom-pbs-apply` already lives in this repo — host-install v1.15.0
|
||
(felhom.eu) now ships it like the mkfs/selfupdate wrappers (drill F-7); no publish change here.
|
||
|
||
## v0.85.0 — the boot/recovery plane: F12 ordering-cycle fix + F11/F10/F9/F2/F1 + appliance self-heal (2026-07-12)
|
||
|
||
Fixes the findings CAMPAIGN-3 (`felhom.eu/documentation/audits/CAMPAIGN-3-2026-07-11.md`) raised
|
||
around the NAS automount lifecycle — the data plane held, the reboot/recovery plane did not.
|
||
|
||
- **F12 (CRITICAL) — boot ordering cycle.** `internal/storage/netmount.go`: BOTH rendered units drop
|
||
`After=/Wants=network-online.target`. The `.mount` keeps `_netdev` (the correct, sufficient network
|
||
ordering — systemd classes it remote-fs); the `.automount` gets NO network relation (a trigger
|
||
needs none, and it must stay orderable before local-fs without dragging the network into the
|
||
transaction). The literal ordering had closed the cycle
|
||
networking→local-fs→automount→network-online→networking, which systemd broke by deleting an
|
||
arbitrary job — one boot lost networking entirely (host dark 7 h), the next lost the automount.
|
||
**Installed-unit migration:** `MigrateNetworkUnits` — a general template-drift reconcile
|
||
(SHA-256 content compare of each marker-owned unit vs a fresh render of its reconstructed spec;
|
||
rewrite + one batched `daemon-reload`; idempotent). Runs at daemon startup before the reassert
|
||
sweep and at the head of `EnsureNetworkMount`, so pre-0.85 units carrying the cycle are repaired,
|
||
not just future adds.
|
||
- **F11 (HIGH) — read the right unit.** `internal/storage/netreassert.go` `netReassertClassify`: the
|
||
re-arm decision is driven ONLY by the host `/proc/mounts` fstype at the mountpoint. Active
|
||
nfs4/cifs → skip-active (inherited by fresh namespaces); anything else → re-arm. The `.automount`
|
||
unit's own state is never consulted (an armed trigger always reports active — the mis-skip trap).
|
||
- **F10 (CRITICAL) — re-arm for real.** A `.mount`/`.automount` left `failed`/start-limit-hit (the
|
||
campaign's unexport→idle-timeout→access×5) is `reset-failed` FIRST (new sudoers verb) — without it
|
||
the `enable --now` is refused by the start limit and the share stays dead across every boot. The
|
||
failed-state read is unprivileged (`systemctl is-failed`, seam-injected).
|
||
- **F9 (HIGH) — say what you did.** The pass enumerates by marker-owned unit files on disk (not
|
||
enablement/runtime state) and logs an INFO verdict line for EVERY share
|
||
(reasserted / reset-failed+rearmed / skip-active / skip-foreign / error) — an empty-looking sweep
|
||
over N shares is now impossible.
|
||
- **F2/F1 (LOW) — zero residue.** `RemoveNetworkMount` (and, via the same path, every verify-fail
|
||
rollback) reset-failed's the pair before removing the files (no `not-found failed` residue) and
|
||
`rmdir`s the now-empty mountpoint (F1 — the campaign's 10 stub-shaped leftovers). rmdir ONLY: a
|
||
non-empty dir is left in place with a WARN (fail-safe; never `rm -rf`).
|
||
- **The hook can never take a guest down (F10/rc255).** `cmd/felhom-agent/main.go` `runHookPhase`:
|
||
every guest-hook phase runs recover-wrapped under a hard timeout and returns cleanly (a panic →
|
||
logged to the PVE task log, swallowed; an overrun → abandoned). The installed wrapper snippet no
|
||
longer `exec`s — it runs the binary as a child, `|| true`, and `exit 0` (the shell belt).
|
||
- **Appliance-mode node self-heal (Part 6).** New `internal/selfheal` package: a minimal
|
||
check/remedy registry gated on `deployment_mode`. One heal ships — host networking recovery
|
||
(F12-class defense in depth): Healthy ⇔ networking.service active AND a default route; the remedy
|
||
(`systemctl start networking.service`, new sudoers verb, ≤3 attempts, 10/30/60 s backoff, terminal
|
||
give-up logged) runs ONLY on `deployment_mode:"appliance"`. A byo host runs the check + WARNs; the
|
||
remedy is structurally unreachable (the Manager gates before any exec). Absent/unknown mode → byo
|
||
(fail-safe). Give-up is log-only (no natural HostReport field; the report schema is out of scope) —
|
||
the ERROR lands in the always-DEBUG applog ring for a hub bundle-pull.
|
||
- **Config:** top-level `deployment_mode` (`internal/config`, `+FELHOM_AGENT_DEPLOYMENT_MODE` overlay,
|
||
`IsAppliance()` fail-safe-to-byo). NOT overloaded onto `Privileged.Mode`.
|
||
- **Sudoers (loud, per the no-widening rule):** two narrow additions —
|
||
`systemctl reset-failed -- mnt-felhom*` (F10) and `rmdir /mnt/felhom-drives/*` (F1, fail-safe:
|
||
rmdir refuses a non-empty dir) in FELHOM_NETMOUNT; `systemctl start networking.service` as the new
|
||
FELHOM_SELFHEAL alias (appliance self-heal; the grant alone cannot harm — starting networking is
|
||
what boot should have done; the remedy is ALSO code-gated on appliance). Capability manifest gains
|
||
the three representative probes.
|
||
- Tests: F12 render (no network-online, `_netdev` present) + reconcile (drift rewritten once,
|
||
idempotent, batched reload, round-trip) + foreign-unit ignore; reassert fstype table with the
|
||
automount-state-ignored red-proof, reset-failed+rearm, per-unit verdict count; hook rc-0 under
|
||
panic/timeout; zero-residue (reset-failed + rmdir, never rm -rf); selfheal state machine +
|
||
byo-never-invokes + absent⇒byo. All green.
|
||
|
||
## v0.84.0 — ReassertNetworkMounts: NAS automount survives guest reboots (RCA fix 1) (2026-07-11)
|
||
|
||
Agent half of the RCA fix pair (controller v0.117.0). Source:
|
||
`felhom.eu/documentation/audits/AUDIT-nas-cwa-rca-2026-07-11.md` — a fresh guest namespace inherits
|
||
REAL submounts (ext4/nfs4) but NOT an idle autofs trigger, so after any guest reboot an idle NAS
|
||
share silently degrades to a local stub inside the guest. The heal (re-create the automount → the
|
||
fresh trigger-mount event propagates live into running guests) was live-proven in the RCA
|
||
remediation; this release makes it automatic.
|
||
|
||
- **`internal/storage/netreassert.go`** `SudoHostOps.ReassertNetworkAutomounts`: per configured
|
||
network mount, the §8 decision table — real nfs/nfs4/cifs mounted → skip (inherited); `autofs`
|
||
trigger at the path → **stop + enable --now the `.automount`** (existing FELHOM_NETMOUNT sudoers
|
||
verbs; there is NO restart grant); neither → skip (removed/orphan states owned by add/remove).
|
||
Idempotent; per-share errors never stop the pass. Unit enumeration factored into
|
||
`networkUnitEntries()` (shared with `ListNetworkMounts`, behavior unchanged).
|
||
- **`internal/localapi/netreassert.go`** `Server.ReassertNetworkMounts`: the daemon leg — runs the
|
||
host-global pass once, then best-effort verifies each RUNNING guest actually sees each share path
|
||
(`GuestSeesMount`; the RCA's masking lesson). Type-asserted capability (the lean
|
||
`NetworkStorageOps` interface and its fakes stay unchanged — the main.go
|
||
`ReassertEnrolledMounts` pattern). Wired at agent startup after `ReassertGuestBinds`;
|
||
deliberately NOT in the 20 s ticker (an idle trigger is healthy and must not be churned).
|
||
- **`internal/guesthook/netreassert.go`** + `PhasePostStart`: the hook leg — PVE runs the hookscript
|
||
as root, so `guest-hook <vmid> post-start` re-arms triggers with DIRECT systemctl and verifies
|
||
via `GuestSeesPath` (hook-process mirror of GuestSeesMount). Non-fatal by contract (stderr → PVE
|
||
task log; exit 0 always); 30 s bound. The installed wrapper snippet already forwards all phases —
|
||
no snippet re-install needed.
|
||
- Tests + red-proofs: §8 table (red-proof: always-rearm shape → FAIL "nfs → rearmed, want
|
||
skip-active" — the live-mount churn the table prevents); rearm emits EXACTLY stop+enable-now on
|
||
the right unit; active mount → ZERO systemctl calls; idempotent double-pass; hook wiring
|
||
(red-proof: PhasePostStart case removed → FAIL "got []"); daemon leg verifies running guests
|
||
only; invisible-share verify is WARN-only (non-fatal).
|
||
|
||
## v0.83.0 — observability pass: always-DEBUG capture ring + GET /debug/logs + heartbeat log-pull + gap-fill sweep (2026-07-11)
|
||
|
||
Agent half of the cross-repo observability task (controller v0.116.0 + hub v0.46.0). Motivating
|
||
incident: a refused NAS verify on an `info`-level box left NOTHING readable remotely — the agent
|
||
logged only to host journald and the new features emitted few lines.
|
||
|
||
- **Capture layer** (`internal/log`): `applog.New` now returns `(logger, *Ring)` — a slog fan-out
|
||
where stderr keeps the configured level (journald unchanged) and a ~1000-entry ring handler is
|
||
FIXED at `LevelDebug`, so flow detail exists for remote pulls without a config flip. The ring is
|
||
an io.Writer fed by a stdlib TextHandler; entries are parsed (time/level/message-verbatim).
|
||
Red-proof: ring gated at the emit level → capture-at-info test FAILS ("ring holds 1, want 2").
|
||
- **`GET /debug/logs`** (local API, same token-auth/self-scoping wrap as siblings): the ring as
|
||
JSON `{entries, total}`; `?raw=1` plain text; 503 when unwired. Plus a request-level DEBUG
|
||
middleware (method/path/status/duration — never bodies) wrapping the whole mux.
|
||
- **Heartbeat log-pull** (the report-channel logtail.go pattern mirrored): the control envelope
|
||
gains `log_tail_requested: bool` (additive); when set, the NEXT heartbeat carries
|
||
`log_tail: {collected_at, lines[]}` (newest-kept, 128 KB cap). Consume-once both ends: local
|
||
pending drains onto the carrying push; a FAILED push leaves the hub request pending → the next
|
||
envelope re-arms (retry proven in tests; red-proof: drain removed → tail ships every cycle →
|
||
FAIL). Serving a pull logs `operator log pull served` (INFO — customer-visible transparency).
|
||
- **Gap-fill sweep** (entry, decisions, outcome+duration, errors): netverify (job start, trigger
|
||
outcome, /proc/mounts verdict, journal byte-count, classification code, rollback outcome,
|
||
duration), netstorage add (pre-probe pass verdict, creds staged/removed — path only), netmount
|
||
Ensure/Remove (per-unit install/enable/remove-step results), signedjobs (jobs fetched ids+
|
||
duration, op received class/host/expiry — never signatures), selfupdate executor (invariants
|
||
passed, download sha-match+duration), disks (assign/eject/decommission outcome INFO),
|
||
ReassertGuestBinds (pass summary), controller-swap (pre-pull verify, negative health verdict),
|
||
desired syncer + hub loop (per-exchange DEBUG with durations).
|
||
- **S7 log-sequence smoke**: a full fake NAS add at emit level info must leave the ordered phase
|
||
markers in the ring (red-proof: dropping the /proc/mounts verdict line → FAIL naming the phase).
|
||
- No new sudoers grants, no journald scraping, no streaming — pull-only. Demo-deploy only — NOT
|
||
published (Peti stays 0.81.0; this reaches him with the next publish train).
|
||
|
||
## v0.82.0 — local-API version channel: X-Felhom-Agent-Version on every response (2026-07-11)
|
||
|
||
The controller's capability detection upgrades from route-probing to version comparison: the
|
||
local-API mux is wrapped so EVERY response (any route, any status, including auth failures) carries
|
||
`X-Felhom-Agent-Version` = the build version (`localapi.Options.AgentVersion`, wired from
|
||
main.version). The controller (v0.115.0) reads it passively from ordinary traffic and compares it
|
||
against a per-feature MinAgent table; header-less (≤0.81.0) agents keep working — the controller
|
||
falls back to the v0.114.0 route probe unchanged. No new routes, no envelope changes, no sudoers
|
||
changes. Test: header asserted on authed/unauthed/404 responses (red-proof: wrap dropped → fails);
|
||
empty version emits NO header. Demo-deploy only — NOT published (Peti stays 0.81.0 = the live
|
||
fallback path).
|
||
|
||
## v0.81.0 — NAS verify-before-commit: retry=0 + detached verify job + journal classification (2026-07-11)
|
||
|
||
Implements the agent half of the "NAS verify-before-commit" task on the SPIKE-nas-verify-2026-07-11
|
||
evidence (felhom.eu/documentation/audits/, commit b57f6c1). `POST /netstorage/add` no longer succeeds
|
||
blind — a bogus share can no longer sit at "Készenlét" forever.
|
||
|
||
- **`retry=0` in the production NFS option string** (`netmount.go mountOptions`; Q4-vi): a dead-NAS
|
||
on-demand access now fails clean in ~4 s (ENODEV) instead of wedging the accessing app until
|
||
systemd's 90 s start cap; verify failures classify as "No route to host" instead of a
|
||
diagnostic-free timeout. SMB string unchanged (retry is a mount.nfs option). Already-installed
|
||
units are NOT rewritten (pre-customer; re-add re-creates them).
|
||
- **`internal/storage/netverify.go`** — `ClassifyNetVerifyFailure`: pure, table-driven journal
|
||
classification on the Q4 VERBATIM substrings → `unreachable | nfs_export | smb_auth | smb_share |
|
||
timeout | mount_failed`. `nfs_export` deliberately MERGES not-found/not-permitted (NFSv4 returns
|
||
the identical string for both — Q4 ii≡iii). Everything exits rc=32, so classification is
|
||
string-based by design.
|
||
- **Detached verify job** (`localapi/netverifyjob.go`, the formatjob shape but IN-MEMORY single
|
||
slot): add = decode → role-gate → SYNC fast-fail (full spec validation + 2 s TCP pre-probe; an
|
||
unreachable server is refused with NOTHING installed) → stage SMB creds → EnsureNetworkMount →
|
||
detached verify (trigger read through the automount; mount success judged from **/proc/mounts
|
||
only** — never readability, a 0700 export EACCES is a good mount) → on failure: journal-classify
|
||
+ **auto-rollback** (RemoveNetworkMount + creds file). `GET /netstorage/verify-status` reports the
|
||
slot; phase `none` after an agent restart is the controller's rollback signal (Scenario F —
|
||
deliberately not persisted). Single-flight: a second add while verifying is a 409.
|
||
- **Journal access is UNPRIVILEGED** (`journalctl -u <unit> -n 20 -o cat`, no sudo, no new sudoers
|
||
grant): requires the felhom-agent user in the `systemd-journal` group (host-install ≥ v1.12.0
|
||
successor adds it; existing hosts: `usermod -aG systemd-journal felhom-agent`). Journal
|
||
unavailable degrades to `mount_failed` + a hint, still rolled back.
|
||
- New exported helpers: `storage.NetworkMountedAt` (autofs trigger ≠ mounted),
|
||
`storage.NetworkEndpointReachable` (the 2 s pre-probe). REUSE.md updated.
|
||
- Tests: classifier table on the live Q4 strings, §8 truth table, rollback effects, pre-probe
|
||
zero-install, single-flight + no-job shape. Red-proof outcomes recorded in REPORT.md.
|
||
|
||
## v0.80.0 — PBS DR tier SLICE 2: the apply-bridge (2026-07-10)
|
||
|
||
Consumes hub v0.44.0's `pbs_dr` desired-state descriptor (slice 1): hub tick → the box grows the
|
||
pbs storage entry + K, hands-free. Laws encoded (spike 00afadc + the offsite bridge precedent),
|
||
each red-proof-verified (REPORT.md):
|
||
|
||
- **`internal/pbsdr`** — the bridge (wgtunnel Loop shape; raw-consumer seam; report stanza
|
||
`pbs_dr`). Flow: **adoption probe first** (existing healthy entry → grant + marker, NO consume;
|
||
**tenancy identity is entry-owned** — a descriptor naming a different namespace never repoints a
|
||
live entry: the demo's manual `felhom-offsite` case) → fresh path: **verify-pin-BEFORE-consume**
|
||
(bare TLS dial pinned to the descriptor fingerprint, `pbs.ProbeFingerprint`) → consume
|
||
(`POST /api/v1/hosts/{id}/pbs/consume-token`, PLURAL /hosts/ — the slice-1 route; typed
|
||
`ErrNoPBSSecret`) → wrapper `create` (secret on STDIN; `--encryption-key autogen` → K born; .enc/
|
||
.pw placed into `backup.pbs_secret_dir` when overridden — the spike §4 escrow-path flag) →
|
||
`grant` (the Part-0-evidenced dual-grant, exactly) → post-apply active probe → **seed
|
||
`escrow.pbs_storage_id`** (bare ceremony one-liner; an operator-set different value is never
|
||
clobbered; unknown config keys preserved) → descriptor-hash marker. **Consumed-but-failed is
|
||
LOUD**: persistent `consumed_failed` report state, no silent retry — recovery only via the hub
|
||
Re-issue (a fresh staged secret).
|
||
- **`configs/felhom-pbs-apply`** (root wrapper, `/usr/local/sbin`, the guarded-mkfs shape) + ONE
|
||
pinned sudoers alias `FELHOM_PBSDR` (create/reconcile/grant). **THE SET-ONLY LAW**: no deletion
|
||
verb exists (entry deletion destroys K = un-decryptable backups); re-apply is `pvesm set`-only —
|
||
grep-gated by `TestSetOnlyLaw` over the shipped file. Secret via wrapper stdin (sudo logs argv).
|
||
- Wire: `WirePBSDR` on `WireDesiredState` (field-exact with hub `pbsDRDescriptor`, pinned by
|
||
`TestWireFieldNames`); nil-safe on pre-v0.44.0 hubs. Report: `PBSDRStatus` stanza (adopted/
|
||
applied/waiting_secret/verify_failed/consumed_failed/disabled) via the `PBSDRReporter` seam.
|
||
- Capability manifest: `pbsdr-create`/`pbsdr-reconcile`/`pbsdr-grant` (non-critical, the
|
||
selfupdate rationale). `config.Config.SourcePath` records the loaded file for the escrow seed.
|
||
- Part 0 (recorded in REPORT.md): status reads ride `FelhomAgentBase` (`Datastore.Audit@/`); the
|
||
WRITE path 403s without `FelhomAgentStore` on `/storage/<id>` → the grant op = the §4b dual-grant
|
||
exactly. Demo's purged grants re-asserted; token vzdump to the PBS entry live-proven OK.
|
||
|
||
## v0.79.0 — SLICE 3: escrow upload carries sha256 of the sealed restic password (2026-07-09)
|
||
|
||
The hub-verified escrow auto-confirm chain, agent third: the escrow-create ceremony now records WHICH
|
||
offsite repo password the blob covers — as a non-reversible sha256 (a 256-bit random secret's hash is safe
|
||
to store/serve; the password itself is never logged or uploaded).
|
||
|
||
- `internal/escrow.HashResticPassword` — the CANONICAL hasher: sha256 hex over the TRIMMED password string
|
||
(exactly the value `AttachResticPassword` seals into the blob). **Pinned cross-repo test vector**
|
||
(`TestHashResticPassword_PinnedVector`, same vector asserted in felhom-controller) so the two hashers can
|
||
never drift silently.
|
||
- `cmd/felhom-agent`: `escrowUploadRequest` gains `restic_pw_sha256,omitempty` — set only when a staged
|
||
password was folded in (no staged file → field OMITTED → the hub stores NULL → the controller stays
|
||
pending; correct, the blob doesn't cover the key). `TestEscrowUploadContract` updated (the hub mirrors it
|
||
in the same commit-pair) + asserts the omitted-when-unstaged behavior.
|
||
- Ceremony flow (create / self-verify / R-banner / staged-file wipe) otherwise untouched.
|
||
|
||
## v0.78.0 — fork-4 hygiene: DELETE /escrow/stage-secret (staged-secret wipe) (2026-07-09)
|
||
|
||
Part of the offsite-provisioning hardening bundle (pairs with controller v0.107.0 + hub v0.39.0). The staged
|
||
restic repo password was wiped only by the escrow-create ceremony; a confirm WITHOUT a fresh ceremony (the
|
||
password already escrowed — the live e2e's Option A) left the 0600 staged file behind indefinitely.
|
||
|
||
- `internal/localapi`: `DELETE /escrow/stage-secret` (withGuest) — removes the staged file (+ any stale
|
||
`.tmp` partial). **Idempotent**: absent file → clean 200 `{removed:false}`. The controller calls it
|
||
whenever `EscrowState` flips to `escrowed`. No value ever logged (nothing to log — it's a removal).
|
||
- Test: stage → wipe (file GONE) → re-wipe idempotent → 401 unauthenticated.
|
||
|
||
## v0.77.0 — fork-4: escrow the offsite restic repo password under R (2026-07-09)
|
||
|
||
Makes the restic-offsite repo password recoverable at DR by riding the existing customer-recovery-code (R)
|
||
zero-knowledge escrow (age-under-R, alongside the IdentityBundle), validated by
|
||
`felhom.eu/documentation/audits/SPIKE-restic-password-custody-2026-07-09.md`. Additive; the PBS-K escrow
|
||
path is untouched.
|
||
|
||
- `internal/escrow/identity.go`: `IdentityBundle` gains `ResticRepoPassword` (`restic_repo_password,omitempty`)
|
||
— rides the existing `WrapIdentityBundle`/`UnwrapIdentityBundle` age-under-R path (self-verified by
|
||
`escrow.Create`). Added `AttachResticPassword` (mirrors `AttachWGKey`), `StagedResticPasswordPath`, and
|
||
`WipeStagedResticPassword`. Pre-fork-4 blobs lack the field and CANNOT be retro-fitted (R never retained)
|
||
— the controller's atomicity gate ensures no offsite ciphertext exists until the key is escrowed.
|
||
- `internal/localapi`: `POST /escrow/stage-secret` (`withGuest`, `scopedFromBody`) transiently stages the
|
||
controller-pushed restic password (0600, atomic tmp+rename, **never logged** — field name only, value
|
||
never echoed), overwritten on re-push. Stage path injectable via `Options.EscrowStagePath` (default the
|
||
canonical `StagedResticPasswordPath`) for testability.
|
||
- `cmd/felhom-agent/main.go` (`runSelftestEscrowCreate`): the escrow-create ceremony auto-injects the staged
|
||
password into the `IdentityBundle` (mirrors the WG-key auto-inject) and **wipes** the staging file after a
|
||
successful create. The ceremony stays operator-invoked (`--selftest=escrow-create`).
|
||
- Tests: `IdentityBundle` round-trip carries `ResticRepoPassword` byte-exact + not-in-blob + wrong-R fails
|
||
closed; `AttachResticPassword` (missing/present/empty); stage endpoint stages 0600 + non-secret ack +
|
||
cross-guest 403 + **value-not-in-log**.
|
||
- NOT yet live-validated — the supervised escrow ceremony (enable→stage→escrow-create→confirm→gated run) is
|
||
the operator-run follow-up.
|
||
|
||
## v0.76.0 — restore-test full-fidelity verification (GL-5b / go-live G12) (2026-07-08)
|
||
|
||
Closes GL-5 finding #2's mirror image: the restore-test's live-source-config bind-override path
|
||
tripped the same PVE drop-unlisted-mountpoints rule DR did, so it boot-verified scratch guests
|
||
WITHOUT their storage mpN — weaker verification than it claimed, and GL-6 will lean on it.
|
||
**0.75.0 was superseded unpublished; this is the Day-0 manifest bump target.**
|
||
|
||
- **Archive-derived params:** the restore-test now derives its restore params exactly like DR
|
||
bring-up — `ExtractArchiveConfig` + `drRestoreOverrides` (the ARCHIVE is the object under test,
|
||
not any live guest's current config): rootfs explicit, every storage mpN passed through (its
|
||
content genuinely extracted — full fidelity; the added runtime IS the verification), structural
|
||
binds → throwaway stand-ins. Unreadable archive config / unknown bind topology → refuse UP FRONT
|
||
(never restore a partial guest to "verify" it).
|
||
- **Mount-parity assert (the non-hollow core):** pre-start, the restored scratch's mpN set is
|
||
compared against the archive's — a missing, mispathed, undersized, or extra mpN FAILS the test
|
||
naming the delta. PVE rule (b) can never regress into a green light again. `MountParity`
|
||
("ok"|"mismatch") + `MountInventory` ride the result and the hub wire record (additive keys — an
|
||
older hub ignores them).
|
||
- Dead code deleted with its tests: `bindMountOverrides`, `archiveVMID` + the volid regexes (the
|
||
restore-test was their only caller; a reachable dead lookalike is how the next bug happens).
|
||
`throwawayVolumeOverride`/`rootfsSizeGB`/`mountPathOf` live on under `drRestoreOverrides`.
|
||
- DR bring-up, 4d, KeepMAC, grows, caps: untouched (bringup.go had zero line changes — the §13
|
||
DR re-run trigger did not fire).
|
||
- Tests: Scenario A (exact 5-param derivation from the archive + parity ok + inventory + teardown),
|
||
B (dropped mp0 → FAIL naming it, never started, still torn down; RED-PROOF: with the parity
|
||
assert removed the run passes silently — run→fail→revert), C (extract-failure + unknown-topology
|
||
refusals, no restore attempted), D (provision/DR bring-up tests all green, byte-untouched);
|
||
`mountParity` pure-function matrix (parity/missing/mispathed/undersized/extra/trivial).
|
||
- Live validation + publish: REPORT.md (full-fidelity runtime measured vs the ~3m data-less shape;
|
||
`AGENT_VERSION`/`AGENT_SHA256` recorded verbatim for the operator's manifest bump).
|
||
|
||
## v0.75.0 — DR bring-up structural bind overrides + real-bind swap (GL-5 / go-live G8) (2026-07-08)
|
||
|
||
Implements the verdict of `felhom.eu/documentation/audits/SPIKE-dr-bindmount-source-2026-07-07.md`:
|
||
`bring-up -mode dr` of a customer archive FAILED outright because the archive carries the two
|
||
structural host-bind mountpoints (mp8 parent bind, mp9 bootstrap bind) that a `pct restore` under
|
||
the privsep token cannot recreate ("restoring 'mp8' to bind mount is only possible for root") —
|
||
bring-up passed no `MountOverrides`. Provision was never affected (the golden has no mp8/mp9; the
|
||
back-half adds them) — that asymmetry was the bug.
|
||
|
||
- **DR restore overrides** (`internal/reconcile/bringup.go`): ModeDRGuestLoss synthesizes throwaway
|
||
1G-volume overrides for the two PLATFORM-CONSTANT mpN (mirroring backhalf.go's values; the spike's
|
||
whole point — no archive parse for the LAYOUT) via the shared `throwawayVolumeOverride` format
|
||
helper (extracted from `bindMountOverrides`; the restore-test's is-a-bind FILTER reads live
|
||
configs, which DR by definition has none of). ModeProvision passes nil — regression-contract test.
|
||
- **LIVE-DISCOVERED: PVE's explicit-params restore is ALL-OR-NOTHING** (neither half was in the
|
||
spike — it never ran an override restore). (a) mpN params without an explicit `rootfs` → HTTP 500
|
||
"mount points configured, but 'rootfs' not set" (same rule restoretest.go:211 documents).
|
||
(b) **Mountpoints NOT named in the params are silently DROPPED** — the first live run came up
|
||
boot+running WITHOUT its mp0/mp1 data volumes (2m56s; the customer's world did not ride along).
|
||
Fix: NEW `Client.ExtractArchiveConfig` (GET `/nodes/{node}/vzdump/extractconfig` — **answers 200
|
||
under the scoped agent token**, verified live; PBS keys stay server-side, the spike's candidate-1
|
||
rejection holds) + `drRestoreOverrides` derives the COMPLETE param set from the archive's own
|
||
embedded config: explicit rootfs, every storage-backed mpN passed through (size + path + backup
|
||
preserved → vzrestore extracts its content), structural binds → throwaways. Unknown bind mpN /
|
||
unparseable size / unreadable config → clean refusal before any restore. Snapshot sections never
|
||
shadow the current config.
|
||
- **Step 4d — real-bind swap** (DR only, pre-start): mp9 bootstrap host dir created (idempotent —
|
||
a same-host guest-loss still has bootstrap.json there, untouched), then mp8/mp9 set to the REAL
|
||
binds via the host runner (`pct set` — bind mounts are root@pam-only, hence NOT the API; sudoers
|
||
already allowlists both shapes), one slot per call so a failure names the exact mpN and rolls
|
||
back per the committed/launched envelope (C2 — never a silent half-wired success). The displaced
|
||
throwaway volumes (PVE parks them as `unusedN`) are deleted via one config PUT (`delete=`);
|
||
a scoped-token refusal logs the residue LOUDLY + warns in the result instead of widening
|
||
privileges. NEW `proxmox.GuestConfig.Unused()`.
|
||
- **Engine seam:** `EngineOptions.HostRunner` + `StateDir` (DR refuses up front on an API-only
|
||
engine); the bring-up selftest wires the same ExecRunner shape as the back-half and removes the
|
||
scratch vmid's mp9 host dir at teardown (never a real drive's bind source).
|
||
- Tests (non-hollow): Scenario A exact-override + exact-swap-command + unusedN-delete asserts;
|
||
provision-nil regression; C2 mid-swap rollback; C3 older-archive-without-mp9; DR-without-runner
|
||
refusal; extract-failure refusal; 403-residue warn; archive rootfs parse (snapshot sections never
|
||
shadow). Red-proofs: override synthesis reverted → A FAILS; unconditional overrides → B FAILS
|
||
(both run→fail→revert).
|
||
- Live validation (campaign-2-precedent scratch DR into vmid 9310 from a real 9201 local archive,
|
||
auto-teardown, guest 9201 untouched): pre-teardown `pct config` shows mp0 200G + mp1 50G restored
|
||
(7m23s — content genuinely extracted), mp8/mp9 = the REAL binds (exact back-half values), rootfs
|
||
32G explicit, ZERO unusedN; boot+running; teardown clean incl. the scratch mp9 host dir. On
|
||
v0.74.0 the same op failed at the restore POST. Full evidence: REPORT.md.
|
||
|
||
## v0.74.0 — pool membership re-asserted after restore-over-existing (campaign-2 R2) (2026-07-07)
|
||
|
||
Closes campaign-2 finding **R2** (`felhom.eu/documentation/tests/CAMPAIGN-2-2026-07-07.md`). Pool
|
||
membership is what lets the pool-scoped `FelhomAgentGuest` token reach a guest (the grant applies only
|
||
to pool MEMBERS). `pct restore --pool` sets membership at CREATE, but a restore **over an existing
|
||
VMID** (the host-loss/finale path) never re-applies it — and no code re-added a guest to the pool — so
|
||
every destroy-restore silently dropped membership, and the NEXT `--selftest=restore-test`/DR 403'd on
|
||
`VM.Audit`/`VM.Allocate`. That empty-pool state is the true root cause of the campaign's "R1" (the
|
||
bind-mount restore failure was a symptom: with no config-read, restore-test's *existing, correct*
|
||
bind-mount neutralization never ran).
|
||
|
||
- **`Client.PoolAddVMID(ctx, pool, vmid)`** (`internal/proxmox/mutate.go`): `PUT /pools/{pool}`
|
||
`vms={vmid}` — PVE-additive (merge, not replace), idempotent (already-member swallowed), needs
|
||
`Pool.Allocate` (the token has it). Sync — no UPID.
|
||
- **bring-up re-asserts** (`internal/reconcile/bringup.go`): after liveness is proven, if `spec.Pool
|
||
!= ""`, call `PoolAddVMID`. A pool-add hiccup is surfaced as a LOUD warning + `res.StartWarnings`
|
||
but must NOT flip a healthy running guest's verdict (membership matters for the NEXT op).
|
||
- **B3 (scratch teardown 403) — diagnosed, no code:** `restoretest.go` already passes
|
||
`Pool: DefaultPool` for scratch restores, so the campaign's `VM.Allocate` teardown 403 was a
|
||
CASCADE of the bind-mount restore failing (half-built guest outside any pool), not an independent
|
||
gap. A re-run on the healed pool confirms.
|
||
- Tests: `PoolAddVMID` PUT shape + idempotency + real-error-surfaces + validation (`pool_test.go`);
|
||
bring-up re-asserts when `Pool!=""` (red-proof: pre-fix no-call → FAIL, demonstrated), no-pool→no-call,
|
||
pool-add failure warns-but-passes (liveness wins, guest kept). Role/ACL untouched — the fault was
|
||
membership, not privileges.
|
||
|
||
## v0.73.0 — F2 mount-role fallback: enrolled user-data drives are ejectable/decommissionable again (2026-07-06)
|
||
|
||
Closes campaign finding **F2** (`felhom.eu/documentation/audits/CAMPAIGN-nomercy-2026-07-06.md` +
|
||
RERUN addendum). `roleForMountPath` (`internal/localapi/disks.go`) resolved a mount's protection role
|
||
ONLY from the PVE storage view (`Observe`), but a **bind-mounted RAW enrolled user-data drive is not
|
||
a PVE storage** → no MountPath match → fail-safe `RoleSystem` → the eject/decommission role gates
|
||
403'd **every user-data drive in the standard topology** (journal: `where=/mnt/teszt_enroll
|
||
role=system`). Customers could not eject or decommission their own drives.
|
||
|
||
- **Fallback** mirrors `durableIDForMount`'s Impl-2b: after the MountPath loop misses (on a SUCCESSFUL
|
||
Observe), resolve the mount's backing device from the host mount table and classify **device-keyed**
|
||
— non-`/dev` source (NAS) → system; device on the same whole disk as a KNOWN target → THAT target's
|
||
role (containment, via new `storage.SameWholeDisk`, whole-disk granularity so a protected disk can't
|
||
be ejected through this path); else `RoleForRawDevice`. **Fail-safe preserved**: an Observe error
|
||
returns system BEFORE the fallback (a blind containment pass could label a backup drive user-data —
|
||
permissive), and a mount-table-read failure or an absent/NAS mount → system.
|
||
- Scope: **only** `roleForMountPath`; gates, `deviceRole`, `DecommissionExecutor`, `classify.go`,
|
||
`ReassertGuestBinds` untouched; the `deviceRole`/`roleForMountPath` unification is deferred.
|
||
- Tests (`f2_role_fallback_test.go`): A1 ejectable + decommission effects; B1 containment→403; B2
|
||
system-disk→403; C1/C2 fail-safe; C3 Observe-error skips the fallback. Three red-proofs demonstrated
|
||
(pre-fix→A1 FAIL, containment-skip→B1 FAIL, fallback-on-error→C3 FAIL). Existing RoleGated/
|
||
Decommission tests green unmodified.
|
||
|
||
## v0.72.0 — OOB operator access: rendered operator /32 + dedicated felhom-sshd + port-adaptive belt + oob health (TASK H1) (2026-07-05)
|
||
|
||
The agent half of the merged E1+H1 operator-SSH-access feature (hub half = felhom-hub v0.35.0).
|
||
Provenance: `felhom.eu/documentation/audits/SPIKE-{felhom-sshd,oob-wg-operator-peer}-2026-07-05.md`.
|
||
Live-validated on felhom-pve + the dev endpoint (both spikes' key probes re-run as acceptance).
|
||
|
||
- **Operator /32 rendered into wg-felhom** (`internal/wgtunnel/manager.go` `renderConf`/`allowedIPsLine`):
|
||
`oob_peer_ip` from the desired-state block is appended to AllowedIPs, deterministically SORTED
|
||
(byte-stable conf-hash — no per-tick flap). RENDERED, not a runtime `wg set`, so it survives the
|
||
agent's self-heal ([OF-1]; live-proven: tunnel stopped → self-heal → operator SSH still works).
|
||
- **Dedicated felhom-sshd** (`internal/felhomsshd/`): a SECOND sshd on a claimed non-22 port
|
||
(`[8822,2222,8022,62222]`, LOUD-fail on exhaustion [SF-4]), own config/host-key/AuthorizedKeysFile
|
||
(`/etc/felhom-sshd/authorized_keys/%u`, outside ~/.ssh [SF-3])/unit — COEXISTS with the customer's
|
||
:22 (never touched). Config: render→`sshd -t`→**reload** (never restart-on-change [SF-2]); operator
|
||
authorized_keys from the hub block; `reset-failed`-then-restart heal with a 10-min cooldown, NEVER
|
||
restarting onto an invalid config. `configs/felhom-sshd.service` SAFE — **no `RuntimeDirectory=`**
|
||
[SF-1].
|
||
- **Port-adaptive belt** (`internal/felhomsshd/belt.go` + `configs/felhom-oob.nft`): a STATIC
|
||
`inet felhom_oob` table; the agent mutates ONLY its SETS — `@operator_ips` + `@ssh_port` [trap 4] —
|
||
so felhom-sshd's port is reachable ONLY from the operator `/32` over `wg-felhom` (off-tunnel +
|
||
box↔box dropped at the host; :22 untouched). Idempotent; a nil/unfetched block never empties it
|
||
(no operator lockout).
|
||
- **OOB health** (`internal/felhomsshd/health.go`): the additive `oob` heartbeat stanza
|
||
(`felhom_sshd_active/port/reachable/config_invalid/operator_peer_configured/operator_key_configured/
|
||
wg_handshake_age_s`) — reaches the hub over HTTPS even with felhom-sshd/tunnel down. `reachable` =
|
||
a listener check (the belt blocks a dial); `operator_*_configured` from persistent state
|
||
(belt/authorized_keys), accurate immediately after a restart.
|
||
- **Sudoers**: `FELHOM_SSHD` (config/authkeys install + `sshd -t/-T` + scoped systemctl) + `FELHOM_OOB`
|
||
(nft SET-element ops only — never a rule grant). `oob.enabled` config DEFAULT FALSE.
|
||
|
||
## v0.71.0 — management-plane break-glass: privsep-dir watchdog + mgmt_plane health (TASK G1) (2026-07-05)
|
||
|
||
Prerequisite for the felhom-sshd OOB feature (H1). Closes the lockout from
|
||
`felhom.eu/documentation/audits/SPIKE-felhom-sshd-2026-07-05.md` §8: a second sshd's
|
||
`RuntimeDirectory=sshd` removed the SHARED `/run/sshd` privsep dir and took the stock sshd on :22 down
|
||
too (sessions reset right after SSH2_MSG_KEXINIT) — a management lockout on a healthy box. Three
|
||
independent layers; this repo ships the host artifacts + the agent reporter (hub vault + surfacing are
|
||
the felhom.eu half).
|
||
|
||
- **Host artifacts** (`configs/`, installed by felhom-host-install): `felhom-privsep.tmpfiles`
|
||
(`d /run/sshd 0755 root root -` — layer 1, boot-persistent, owned by no unit's lifecycle);
|
||
`felhom-mgmt-watchdog.sh` (layer 2 heal action — stat-first recreate `/run/sshd`, `reset-failed`
|
||
ssh ONLY when `failed`, write an RFC3339 heal-marker; NEVER restarts the stock sshd, NEVER touches a
|
||
healthy dir — shellcheck-clean); `.service` (oneshot) + `.timer` (~60s, `Persistent`). The healer is
|
||
**agent-INDEPENDENT** — it self-corrects with felhom-agent stopped (the whole point). **No unit
|
||
declares `RuntimeDirectory=`** (that IS the incident cause; the installer refuses any that does).
|
||
- **Go** (`internal/mgmtplane/`): a read-only `Reporter` (os.Stat `/run/sshd` + read the heal-marker +
|
||
a short TCP dial to sshd:22) producing the additive `mgmt_plane` heartbeat stanza
|
||
(`{privsep_dir_ok, sshd_reachable, healed_recently, privsep_healed_at}` — `omitempty`, the
|
||
SelfUpdatePending precedent, no hub-schema change). Wired always-on via
|
||
`Collector.SetMgmtPlaneReporter`. The hub raises a warning on a new `privsep_healed_at` so a
|
||
recurring clobber surfaces BEFORE a lockout (complements host_staleness). Touches host `/run` + the
|
||
stock sshd only — no guests.
|
||
|
||
Closes the update asymmetry: the root-adjacent agent was updated by manual SSH binary-replace while
|
||
the lower-stakes controller already auto-updates. Design provenance:
|
||
`felhom.eu/documentation/audits/SPIKE-agent-selfupdate-2026-07-05.md` (SF-findings binding). Core
|
||
principle (as with the controller swap): the thing that performs rollback is never the thing being
|
||
updated — here systemd + an ~80-line root wrapper.
|
||
|
||
- **Trust model:** an update is an operator-signed `agent_update` op (new `reconcile.ClassAgentUpdate`,
|
||
always Destructive) through the existing signed-jobs pipeline; the signed params pin version + sha256,
|
||
so the sha is the ONLY integrity root (hub = dumb transport, Gitea = dumb storage — neither can
|
||
substitute a binary). `felhom-opsign -op agent_update -agent-version <v> -sha256 <hex>`.
|
||
- **Host artifacts** (`configs/`): `felhom-selfupdate-guarded` (apply/commit/rollback — root re-verify,
|
||
path confinement, same-fs assert, `.prev`, atomic mv, pending marker, detached restart [SF-6];
|
||
rollback pending-guarded [SF-1]); `felhom-agent-rollback.service` (OnFailure oneshot);
|
||
`felhom-agent-limits.conf` ([Unit]-only drop-in [SF-3] with the spike's tuned 120s/4 [SF-2] +
|
||
OnFailure=); `FELHOM_SELFUPDATE` sudoers alias.
|
||
- **Go** (`internal/selfupdate/`): `Executor` (download → verify vs signed sha → `sudo -n` wrapper
|
||
apply; job completed after verify+download, before apply); `Manager` (startup dwell → `commit`;
|
||
version-mismatch → no-commit + WARN + marker-left; report seam `SelfUpdatePending()`). Wired as a
|
||
3rd executor-chain element + a `MaybeCommit` goroutine after core init. `SelfUpdateConfig`
|
||
(url_template/creds/dwell, Token redacted). Additive report fields `selfupdate_pending(+version)`
|
||
(omitempty → cross-repo golden contract byte-stable, no hub change). 3 non-critical capability probes.
|
||
- **felhom.eu:** `felhom-host-install.sh` installs the wrapper + rollback unit + drop-in on day-0
|
||
(self-update from birth); agent README "Self-update" section.
|
||
- Tests: executor (happy / sha-mismatch + companion / bad-params / wrapper-fail), gate ride-along
|
||
(agent_update rides the LOCKED gate: pinned-key executes, non-pinned + retarget rejected), commit
|
||
(dwell-commit / version-mismatch / no-pending / shutdown), opsign, classify. Wrapper covered by
|
||
shellcheck + the live crash-rollback drill. Full `go test ./...` green.
|
||
|
||
## v0.69.0 — S5: host-loss DR — recovered WG-key install + directive→restore-PLAN (safe halves) (2026-07-04)
|
||
|
||
The two safe, non-destructive mechanical links for host-loss DR (the destructive in-place restore is
|
||
a separate operator-present, STOP-gated drill).
|
||
|
||
- **`internal/wgtunnel.InstallRecoveredKey`** — writes an escrow-recovered WG private key (32-byte
|
||
base64, re-encoded canonical) to the key file so the tunnel re-establishes with the SAME
|
||
identity/pubkey (→ the same hub `/32`), no fresh keygen. **CREATE-ONLY** — refuses if a key file
|
||
exists (a present key may be a live identity); value never logged. Wired into `--selftest=identity-
|
||
consume -install-wg-key` (opt-in; after `UnwrapIdentityBundle`, installs `bundle.WGPrivateKey`;
|
||
pre-S3 blob with no WG key → logged fallback to fresh keygen + re-register, which keeps the /32).
|
||
- **`internal/dr`** (new) — consumes the host_loss `restore_directive` (was logged-and-ignored) into
|
||
an inspectable **RestorePlan** via the `desired.Syncer.AddConsumer` raw seam: per guest →
|
||
{vmid, archive, target storage, sizing}; per drive → {durable_id → expected mount}; + the offsite
|
||
PBS coord. **Derive-and-surface only** — the `Consumer` has NO restore/destroy dependency, so
|
||
"execute nothing" is structural. `guest_loss`/absent → no plan. Recipe fetched on-demand (rare
|
||
directive) via a fresh `Collect`.
|
||
- Tests + red-proofs: WG install (same pubkey/no-keygen; present-key refuse — red-proofed against
|
||
allow-overwrite); plan (host_loss builds; guest_loss/absent/nil-recipe → none — red-proofed against
|
||
a relaxed mode gate). No secrets on argv/stdout/logs (field names only).
|
||
|
||
## v0.68.0 — S4.1: tier-aware restore-task deadline (unattended offsite restore-test) (2026-07-04)
|
||
|
||
The offsite restore-test couldn't complete on the scheduler path because a WAN restore of a large
|
||
guest exceeds the restore-task wait's 10-minute default — the wait expired mid-restore, teardown
|
||
then fired against a still-restoring (not-yet-pool-associated) scratch guest, and it leaked. Make
|
||
the restore-task wait **tier-aware**.
|
||
|
||
- **`internal/reconcile`**: `RestoreTestSpec.RestoreTaskTimeout` (0 → the 10m `WaitOptions` default);
|
||
the restore-task `WaitTask` now passes it. Local-tier restores are UNCHANGED (10m — a local
|
||
restore hanging that long is a genuine fault).
|
||
- **`internal/config`**: `BackupConfig.RestoreTestPBSRestoreTimeoutSeconds` + accessor
|
||
`RestoreTestPBSRestoreTimeout()` (positive as-is, else **120m** — for an unattended nightly test a
|
||
false timeout is worse than a slow pass; very large guests may need more).
|
||
- **`cmd/felhom-agent`**: `restoreTaskTimeout(cfg, tier)` sets the field to the configured PBS
|
||
timeout only when `SourceTier=="pbs"` (both the scheduler + selftest spec builds), else 0.
|
||
- Tests: tier-aware `WaitOptions.Timeout` (pbs→120m, local→0; red-proofed against `WaitOptions{}`) +
|
||
the accessor contract. The "grant scratch-band `VM.Allocate`" follow-up was diagnosed, not
|
||
blind-applied — the scratch guest is restored INTO `/pool/felhom` (whose ACL already grants
|
||
`VM.Allocate`), so the earlier teardown 403 was a *consequence* of the timeout (a not-yet-pooled,
|
||
still-restoring guest), not a missing grant. **No ACL/host-install change.** (Live confirmation of
|
||
the phantom in REPORT.)
|
||
|
||
## v0.67.0 — S4: namespace-aware PBS client (per-customer offsite tenancy) (2026-07-04)
|
||
|
||
Phase-1 live probe on felhom-hetzner proved the offsite tenancy path (backup/restore/list/isolation
|
||
all green over the tunnel with a per-customer DatastoreBackup token) but surfaced that the agent's
|
||
PBS client was **namespace-unaware**: `Snapshots` hit the datastore root (403 for a scoped token)
|
||
and `Verify` was whole-datastore (needs Datastore.Verify ~ admin). Operator-approved small change to
|
||
make the client namespace-scoped so a properly-isolated token services its own tenant.
|
||
|
||
- **`internal/pbs`**: `Config.Namespace` (+ `Client.namespace`). `Snapshots` appends `?ns=<ns>`
|
||
(lists ONLY the tenant's namespace); `Verify` sends `ns=<ns>` (verifies ONLY that namespace —
|
||
Phase-1-confirmed to work with a **DatastoreBackup** token on its own ns, no Datastore.Verify /
|
||
admin widening). Root-ns clients (Namespace="") are unchanged → whole-datastore (the DooPlex
|
||
`felhom-pbs` n100 path). Test `TestClient_NamespaceScoping` pins both (red-proofed).
|
||
- **`internal/proxmox`**: `Storage.Namespace` (parsed from the PVE `/storage` config key `namespace`).
|
||
- **`cmd/felhom-agent`**: `pbsTargetsFromPVE` threads `s.Namespace` into the PBS client, so a PBS
|
||
storage configured with a namespace is verified/reported scoped to it automatically.
|
||
|
||
**Confirmed minimal tenant ACL (recorded live 2026-07-04, felhom-hetzner):** `DatastoreBackup` on
|
||
`/datastore/felhom-offsite/<ns>` (the namespace path — NOT `/ns/<ns>`) granted to **BOTH** the user
|
||
`felhom@pbs` **and** the token `felhom@pbs!<ns>` — PBS privsep tokens = intersection(user, token),
|
||
so both are required; isolation holds because each token's ACL is only its own ns (cross-ns
|
||
list/backup → 403, proven). No token exceeds DatastoreBackup; no admin on the endpoint for the box.
|
||
|
||
## v0.66.0 — S4 agent half: endpoint v4-pin + re-resolve watchdog + FELHOM_WG Critical flips (2026-07-04)
|
||
|
||
The two agent items S4 needs before offsite backups ride the tunnel (the tenancy + storage weight is
|
||
runbook-side). No new sudoers grants; no wire/JSON change.
|
||
|
||
- **`internal/wgtunnel` — v4-pin (doc 06 §4.2).** `renderConf` now takes the **pre-resolved IPv4
|
||
literal** and writes `Endpoint = <ip>:<port>` — never the DNS name, never an AAAA. A new `Resolver`
|
||
seam (`net.DefaultResolver.LookupNetIP(ctx, "ip4", …)` — A records only) resolves in the Manager;
|
||
multiple A records → the numerically **lowest** (deterministic fleet-wide). renderConf stays pure
|
||
(no DNS/IO inside). The resolved IP is cached: **steady-state Apply hits the cache — zero DNS, zero
|
||
execs** (the load-bearing negative). DNS failure on (re)resolve → keep the last-applied conf + throttled
|
||
ERROR — **never a teardown** (teardown stays revocation-only).
|
||
- **`internal/wgtunnel` — re-resolve watchdog (doc 06 §4.2, the slice-3 promise).** New
|
||
`Manager.Watchdog` (loop-driven only, so Apply's zero-exec steady state is untouched): when the
|
||
handshake age exceeds `wg_tunnel.stale_after_seconds` (default **180**) it re-resolves; **IP changed
|
||
→ re-render + restart** (endpoint re-IP recovery); IP unchanged → no churn, one throttled warn
|
||
(endpoint merely down). The staleness read is the existing `wg show … latest-handshakes` (never
|
||
`dump`).
|
||
- **`internal/config`** — `WGTunnelConfig.StaleAfterSeconds` (default 180 via `WithDefaults`).
|
||
- **`internal/capability` — FELHOM_WG Critical flips (S4).** Backups ride the tunnel now, so
|
||
`wg-conf-install`, `wg-enable`, `wg-restart`, `wg-handshake-read` are **Critical=true**
|
||
(operator-alert-worthy on degradation); `wg-tools-install` (one-time) + `wg-disable` (deliberate
|
||
revocation) stay non-critical. `TestWGCapabilityCriticality` pins the exact set (red-proofed).
|
||
- Tests: v4-pin golden (A literal, AAAA/dns_name refused), watchdog (healthy=no-DNS negative,
|
||
stale+re-IP restarts, stale+same-IP no-churn+throttle, resolver-failure keeps conf, initial-resolve-
|
||
failure no-teardown+recovery). Red-proofs a/b/d all fire.
|
||
|
||
## v0.65.0 — S3.1 offsite-tunnel client MTU 1420 → 1280 (resolve the CGNAT-smoke MTU open decision) (2026-07-04)
|
||
|
||
One-constant fix closing `06 §4.3`'s OPEN DECISION. The 2026-07-04 CGNAT smoke test found the
|
||
shipped interface MTU 1420 **silently black-holes bulk TCP** on any path below ~1480 B (mobile
|
||
~1400, DS-Lite ~1452): the WG handshake and ping stay healthy (small packets) while the PBS TLS
|
||
page — and, at S4, the backup itself — drops. "Looks green, loses backups." Must be safe before
|
||
S4 flows bulk TCP over the tunnel.
|
||
|
||
- **`internal/wgtunnel/manager.go`**: new `const clientMTU = 1280` (the RFC 8200 IPv6-minimum link
|
||
MTU — every path carries ≥1280; outer = 1280+60 v4 / +80 v6, fits every realistic path);
|
||
`renderConf` emits `MTU = %d` from it. **Fleet-wide, family-agnostic, permanent** — decouples
|
||
the fix from the v4/v6 endpoint-resolution question (§4.2). **Client-only by construction:** the
|
||
interface MTU caps box→PBS and the advertised MSS (=MTU−40) caps PBS→box, so the endpoint's `wg0`
|
||
is deliberately untouched (zero live-endpoint risk). Rejected: auto-probe / per-connection-type
|
||
(fragile moving part optimizing throughput, a non-metric here) and MSS-clamp (no forwarded flows).
|
||
- **`internal/wgtunnel/manager_test.go`**: golden pins exact `MTU = 1280`; red-proofed (flip const
|
||
→ 1420 fails the golden on the MTU line — non-vacuous).
|
||
- **`internal/hub/report.go`**: stale "MTU 1420" comment → 1280 (still NOT a wire field).
|
||
- No wire/JSON-golden change (MTU is a client-derived constant, never on the wire); no endpoint,
|
||
hub, controller, key, or desired-state change.
|
||
|
||
## v0.64.0 — S3 offsite WG tunnel: keygen + registration + agent-managed wg-quick@wg-felhom + escrow join (2026-07-04)
|
||
|
||
The agent half of doc 06 §3.3 (felhom.eu S1/S2 built the endpoint + hub half). **DEFAULT OFF —
|
||
the safety gate:** `wg_tunnel.enabled` defaults to false; a v0.64.0 rollout without explicit
|
||
config is a no-op (no keygen, no registration, no report stanza). Enabled explicitly on
|
||
felhom-pve only; the default flips when the production endpoint exists.
|
||
|
||
- **`internal/wgtunnel`** (new; the lanresolver host-service shape): pure-Go keygen
|
||
(x/crypto curve25519; key 0600 in 0700 StateDir/wg; corrupt file = refuse, NEVER overwrite —
|
||
it may be the escrowed identity; stored form canonical-clamped — x/crypto clamps derivation
|
||
internally, so the STORED bytes are the property that matters, red-proof-anchored); one-shot
|
||
registration (`POST /hosts/{id}/wg`, marker-gated, exponential backoff cap 15 m);
|
||
desired-state consumption via the new `desired.Syncer.AddConsumer` raw seam (panic-contained);
|
||
conf render golden-tested (MTU 1420, AllowedIPs = pbs_tunnel_ip/32, keepalive 25 — doc 06 §4
|
||
client constants; all inputs strictly validated); hash-gated apply (steady state = ZERO
|
||
execs), restart-not-reload on conf change, self-heal enable, adopt-lost-marker,
|
||
re-key-on-mismatch. **Revocation semantics (doc 06 §3.5 completed): block absent from a
|
||
PRESENT desired-state → disable + marker KEPT + never re-register** (the operator re-adds the
|
||
peer using the reported pubkey); absent DATA (failed fetch) is never a teardown signal.
|
||
- **`internal/hub`**: `WireDesiredState.Wireguard` (field-exact with the S2 cross-repo golden,
|
||
copied byte-identical + decode test), `RegisterWG` client (typed errors, token-free),
|
||
`HostReport.Wireguard` status stanza `{pubkey, registered, active, last_handshake_age_s,
|
||
assigned_ip}` via the collector's `WireguardReporter` seam.
|
||
- **Sudoers/capabilities**: `Cmnd_Alias FELHOM_WG` (fixed-path conf install, enable/restart/
|
||
disable, `wg show wg-felhom latest-handshakes` — the ONLY wg read; `wg show … dump` is
|
||
FORBIDDEN, its interface line carries the PRIVATE KEY) + 6 manifest entries (Critical=false
|
||
until S4 makes the tunnel load-bearing).
|
||
- **Escrow**: `IdentityBundle.WGPrivateKey` (omitempty) + escrow-create auto-inject when the key
|
||
file exists (`escrow.AttachWGKey`; field NAME only in logs). Honest limit: pre-S3 blobs cannot
|
||
be retro-fitted (R is never retained) — S5 DR falls back to fresh-key re-registration, which
|
||
keeps the box's /32 (hub S2 re-key-in-place).
|
||
- `--selftest=wgtunnel` single-shot for supervised bring-up.
|
||
- Live-validated on felhom-pve (agent restart tolerance, host reboot with unit persistence,
|
||
revocation drill, 30-min keepalive soak, escrow inject) — see REPORT.md. Five red-proofs run
|
||
+ reverted. GOTCHA learned: the hub envelope's `poll_interval_seconds` (hub-side constant
|
||
900 s) silently overrides the agent's configured cadence on the FIRST cycle — the agent-side
|
||
`poll_seconds` is only the pre-first-heartbeat default.
|
||
|
||
## configs: build-golden.sh v2.0.0 — mandatory controller tag + bootstrap .path unit (B5 + B1) (2026-07-03)
|
||
|
||
Golden-bake script only — **no agent code, no binary, no agent version bump** (docs/config
|
||
precedent). Resolves drill findings B5 + B1
|
||
(`felhom.eu/documentation/audits/DRILL-day0-cleanroom-2026-07-03.md` §9) and closes the stale
|
||
backlog note `FOLLOWUP-golden-default-controller-tag.md`.
|
||
|
||
- **B5 — `CONTROLLER_IMAGE` (arg 6) is now MANDATORY** — the hand-bumped default rotted twice
|
||
(0.43.0 → 0.85.1 → stale again; 0.85.1 predates the v0.86.0 floor-honoring code, which is why
|
||
every fresh install needed the manual D.1b update). No 6th arg → die with usage (red-proofed:
|
||
exits 1 before any `pct` op). Auto-resolving "latest" was rejected — it could bake an unvouched
|
||
tag.
|
||
- **B1 — baked `felhom-controller-bootstrap.path` unit** (`PathExists=/etc/felhom-bootstrap/
|
||
bootstrap.json`, enabled next to the service): the service's `ConditionPathExists` is evaluated
|
||
only at boot, but the agent back-half hot-plugs the bootstrap mount into the already-running
|
||
guest — the path unit starts the service when the file APPEARS, so the controller deploys with
|
||
NO reboot (isolated systemd proof + full Day-0 proof on the clean-room drill VM; the installer's
|
||
v1.9.1 post-provision reboot is now a redundant belt, kept). `RemainAfterExit=yes` on the service
|
||
prevents re-trigger loops; the service itself stays unchanged.
|
||
- `GOLDEN_SCRIPT_VERSION` (2.0.0) + a `[golden]` provenance line (script version + baked controller
|
||
tag) now open every bake transcript.
|
||
- Baked + published with controller **0.98.3**: golden `0.98.3` /
|
||
sha256 `b9a02ef1b6f02b9b58babc4c6aad9cf6c053ebdfba116c78c8e7830de757fd01` (Gitea generic package,
|
||
201 + sha round-trip). Clean-room validation: bake integrity (mp0+mp1 included), isolated
|
||
hot-plug proof, full local-golden Day-0 install (first boot = 0.98.3, self-update reports
|
||
up-to-date, app deploy OK) — evidence:
|
||
`felhom.eu/documentation/audits/DRILL-golden-098-2026-07-03.md`.
|
||
|
||
## v0.63.0 — B3 + B2: fresh-install fixes — token reload-on-miss + guesthook snippets dir (2026-07-03)
|
||
|
||
The two agent-side gaps the Day-0 clean-room drill surfaced
|
||
(`felhom.eu/documentation/audits/DRILL-day0-cleanroom-2026-07-03.md` findings B3/B2). Both are
|
||
fresh-install path fixes; no behavior change on a warm box.
|
||
|
||
- **B3 — `TokenStore.Lookup` reload-on-miss** (localapi/tokenstore.go): provisioning is a SEPARATE
|
||
one-shot process (`--selftest=provision`) that Mints the new guest's token into the shared
|
||
append-only JSONL, while the long-lived daemon serves Lookup from an index built once at open —
|
||
so the daemon 401'd every token minted after it started (the drill's
|
||
`POST /controller/swap: HTTP 401`, until a manual `systemctl restart felhom-agent`). Lookup now
|
||
re-reads the file ONCE on a miss (`reloadLocked()`, factored from `load()`; full re-read is
|
||
idempotent under `apply`'s last-write-wins) and re-checks. An append-only size short-circuit
|
||
bounds the cost: an unknown token on an unchanged file is one `stat`, no re-read — never a reload
|
||
loop. Fix is entirely behind the `TokenAuthority` seam (no server change). Fail-closed on an
|
||
unreadable store; missing file loads as empty. Tests (tokenstore_test.go): cross-process-mint
|
||
coherence (red-proofed: pre-fix shape returns (0,false)), exactly-once reload bound +
|
||
size short-circuit, cross-process re-mint rotation coherence, deleted-file no-crash
|
||
(linux-only; windows can't unlink an open handle).
|
||
- **B2 — `guesthook.InstallSnippet` ensures the snippets dir** (guesthook/install.go): a fresh PVE
|
||
has no `/var/lib/vz/snippets` and `install` (without `-D`) won't create parents — the pre-start
|
||
self-heal hook silently failed to install on every freshly-bootstrapped box (warn-only in the
|
||
back-half). A fenced `mkdir -p /var/lib/vz/snippets` now precedes the install; **sudoers gains
|
||
exactly that one grant** (FELHOM_GUESTHOOK — configs/felhom-agent.sudoers must ship WITH this
|
||
binary, as always). Test: mkdir-precedes-install argv assertion (red-proofed: pre-fix has no
|
||
mkdir op).
|
||
- Guide follow-through (felhom.eu, separate commit): D.1b's "restart the agent first" step drops
|
||
once B3 is live-verified.
|
||
|
||
## v0.62.0 — A1: pool-membership ownership check for the stale-lock reaper (2026-07-03)
|
||
|
||
Implements audit finding A1 (`AUDIT-blast-radius-hostroot-localapi-2026-07-02` §A) per the spike
|
||
verdict (`SPIKE-a1-pool-membership-read-2026-07-03` — enumeration is pool-filtered under the scoped
|
||
token, so this is defense-in-depth: a future broad-token deployment can no longer re-arm the reaper
|
||
against co-tenant guests). Companion: host-install **v1.9.0** (`Pool.Audit` added to
|
||
`FelhomAgentGuest`) — **rescope BEFORE deploying this agent**, else the reaper fail-safes (skips)
|
||
until the ACL catches up.
|
||
|
||
- **`Client.Pool`** (proxmox/query.go): `GET /pools/{name}` → `PoolInfo{PoolID, Members[]{VMID,Type}}`.
|
||
Requires `Pool.Audit` at `/pool/{name}`; `Pool.Allocate` does NOT satisfy the read (spike T2).
|
||
- **`staleLockController.Guests()`** (localapi/stalelock.go): now returns `ListLXC ∩ pool members`
|
||
(nonzero-vmid, non-storage entries only). Ownership is PROVEN via the pool registry, never assumed
|
||
from enumeration scope. A pool-read failure returns a wrapped error ("pool membership read
|
||
(pool=felhom): …") that rides the existing "guest list unavailable — skipping recovery" guard —
|
||
fail-safe: NO unlock/snapshot-delete/start on ANY guest, never a fallback to the unfiltered list.
|
||
One new INFO line per scan: `stale-lock: scanning pool guests` (pool, listed, scanned) — emitted
|
||
by the controller (the unchanged `StaleLockController` seam can't carry the pre-intersect count).
|
||
- **`NewStaleLockController`** gains `(pool string, logger *slog.Logger)`; main.go threads
|
||
`reconcile.DefaultPool`.
|
||
- **Capability surfacing**: the hub-report prober is now a composed closure — the sudo manifest
|
||
probe + one `pve:pool-read` status (non-critical; degraded ⇒ reaper is fail-safed, visible on the
|
||
report, no operator page). Composed in main.go; `internal/capability/` untouched.
|
||
- **`--selftest`**: new "pool read" line (pool id + member count + guest members).
|
||
- **Tests** (stalelock_pool_test.go, driving the REAL controller over a broad-token-shaped fake):
|
||
`TestStaleLock_ForeignGuestNotReaped` (red-proved: intersect removed ⇒ FAILS with
|
||
`pct unlock 5000` recorded), `TestStaleLock_PoolGuestStillReaped` (anti-over-filter),
|
||
`TestStaleLock_PoolReadFails_SkipsAll` (red-proved: fallback-to-unfiltered ⇒ FAILS with mutations
|
||
recorded), `TestStaleLockController_GuestsIntersect` (storage-member + empty-pool edges). The 9
|
||
existing Server-level stalelock tests pass unmodified.
|
||
|
||
## docs — CLAUDE.md refresh: stable orientation, complete layout (2026-07-03)
|
||
|
||
No code change, no version bump. Deleted the version-pinned "Current: v0.31.0" narrative and the
|
||
per-slice history (stale by 30 versions — current state lives in CONTEXT.md/CHANGELOG top); Layout
|
||
completed with the 8 missing packages (capability, desired, escrow, guesthook, lanresolver, localapi,
|
||
provision, signedjobs + cmd/felhom-opsign — verified against the tree); build/deploy compressed to a
|
||
summary table pointing at the `felhom-build-deploy` skill (commands verified live on felhom-pve:
|
||
non-root `felhom-agent` user, `/usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json`).
|
||
Load-bearing Proxmox-model rules kept verbatim. Standing rule: no version-pinned state in CLAUDE.md.
|
||
|
||
## docs — REUSE.md introduced (2026-07-03)
|
||
|
||
Cross-repo reuse-map rollout (docs-only, no code change, no version bump). New `REUSE.md` at the
|
||
repo root: curated map of canonical helpers (48 rows — exec/sudoers surface, format-safety guards,
|
||
durable-id seams, stores, local-API plumbing), patterns, dangerous lookalikes (req.Device TOCTOU,
|
||
raw mkfs, uuid:-vs-byid: scheme confusion, MemoryNonceStore, pool-blind stale-lock scan…), test
|
||
seams, extension points, and observed duplication (7 clusters, NOT fixed). Every entry code-verified
|
||
at file+symbol; cited paths machine-checked by `felhom.eu/scripts/reuse_refs_check.py` (green).
|
||
CLAUDE.md gains the "See REUSE.md before writing new code" pointer + the same-commit maintenance rule.
|
||
|
||
## v0.61.0 — blast-radius audit fixes B1 + D1 + D2 + D3 (2026-07-03)
|
||
|
||
Four LOW/INFO fixes from `AUDIT-blast-radius-hostroot-localapi-2026-07-02.md` — the "second gate must
|
||
mirror the first" batch. Each shipped with a non-hollow test AND a companion red-proof (test shown
|
||
failing on the pre-fix implementation). A1 (stale-lock pool-membership) is deliberately NOT here — it
|
||
needs a pool-read spike (the role lacks `Pool.Audit`); C1/C2/A2/B2–B5/E1/E2 deferred.
|
||
|
||
- **B1 (LOW) — random temp staging for root-installed scripts.** `internal/guesthook/install.go`
|
||
(`InstallSnippet`) and `internal/localapi/intermediary.go` (`installSharedParentUnit`, new shared
|
||
`stageTemp`) staged root-executed scripts through FIXED, predictable `/tmp` names via `os.WriteFile`
|
||
(no O_EXCL, follows symlinks) — a local TOCTOU into a root-run PVE hookscript / boot script. Both now
|
||
use the `os.CreateTemp` random-name pattern `lanresolver` already used. `configs/felhom-agent.sudoers`
|
||
(`FELHOM_GUESTHOOK`/`FELHOM_INTERMEDIARY`) install-SOURCE grants became globs
|
||
(`/tmp/felhom-guest-hook-*.sh`, `/tmp/felhom-shared-parent-*.{sh,service}`; destinations stay pinned);
|
||
`internal/capability/manifest.go` representative vectors updated to match. Tests:
|
||
`TestInstallSnippet_RandomTempName`, `TestInstallSharedParent_RandomTempName` (fake runner records the
|
||
install source: random pattern, two calls differ, content + cleanup asserted).
|
||
- **D1 (LOW) — the guarded mkfs wrapper now mirrors the classifier's member/RO classes.**
|
||
`configs/felhom-mkfs-guarded.sh` re-checked only system-disk / LVM-PV (PATH-dependent `command -v
|
||
pvs`) / foreign-mount — a bypassed agent could mkfs a ZFS/mdraid/LUKS/swap member or a read-only
|
||
disk. Added (additive; nothing removed/reordered): `/sys/block/<disk>/ro` == 1 → die; an
|
||
lsblk-FSTYPE loop over the whole disk refusing exactly `claim.go`'s `memberFSTypes`
|
||
(LVM2_member/zfs_member/linux_raid_member/crypto_LUKS/swap — the FSTYPE catch works with pvs
|
||
absent); pvs resolved via absolute candidates (/usr/sbin/pvs, /sbin/pvs). Validated by the new
|
||
`scripts/mkfs-guarded-harness.sh` on felhom-pve: throwaway loop devices + PATH-shimmed lsblk + a
|
||
recorder bind-mounted over mkfs.ext4 in a private mount namespace (no real mkfs possible) — fixed
|
||
wrapper 8/8 (incl. plain-blank-disk still formats); pre-fix wrapper red-proof: 7/8 hostile fixtures
|
||
reached mkfs.
|
||
- **D2 (INFO) — `classifyClaim` empty-lsblk fail-safe.** A successful-but-empty `lsblk`
|
||
(`{"blockdevices":[]}`) skipped the member/mount loop and returned `unclaimed`.
|
||
`internal/storage/claim.go` now refuses when the node tree is empty OR the target whole-disk is
|
||
absent from it (undeterminable topology ⇒ claimed). Tests: `TestClassifyClaim_EmptyNodesRefused`,
|
||
`TestClassifyClaim_TargetAbsentFromTree`.
|
||
- **D3 (INFO) — blank-format anti-retarget (AGENT-001's benign-branch twin).** `handleDiskFormat`'s
|
||
blank branch formatted the mutable caller-supplied `req.Device` with no durable-id binding — a /dev
|
||
re-enumeration between inspect and mkfs could format a data-bearing disk that inherited the node.
|
||
The blank branch (`internal/localapi/disks.go`) now derives the device's durable id (no durable id ⇒
|
||
409 refuse — path-only formats are not permitted), re-resolves it via the new
|
||
`antiRetargetResolveBlank` (`wipe_reresolve.go`: shared `antiRetargetResolveExpect` core; the blank
|
||
variant asserts the device is STILL !DataBearing), and formats the RE-RESOLVED device. The format
|
||
job record (`formatjob.go`) carries `blank`; restart recovery re-checks blank jobs with the blank
|
||
variant (durable-id-bound, fail-safe refuse). Confirmed/data-bearing branch untouched. Tests:
|
||
`TestFormatBlankPath_AntiRetarget_{ReassignedDataBearingRefused,ReassignedDifferentDiskRefused,UnresolvableRefused,SameBlankProceeds}`,
|
||
`TestFormat_Blank_{FormatsReresolvedDeviceNotCallerPath,ReresolveRefusalNoMkfs,NoDurableIDRefused}`.
|
||
- Deploy note: the host's `/etc/sudoers.d/felhom-agent` MUST be updated together with the v0.61.0
|
||
binary (the old fixed-name grants deny the new random-name installs, and vice versa).
|
||
|
||
## v0.60.0 — proof-of-launch destroy gating + restore-test band-advance (campaign F1/F2) (2026-07-02)
|
||
|
||
Fixes the pool-effects campaign's HIGH finding (F1, `CAMPAIGN-pool-effects-2026-07-01.md`): the bring-up
|
||
compensating rollback and the restore-test teardown fired `DestroyLXC` on the target vmid even when
|
||
`RestoreLXC` failed synchronously WITHOUT creating anything (PVE refusing a pre-existing vmid the
|
||
pool-blind duplicate guard / band scan couldn't see) — destroying a guest the transaction never made.
|
||
Only the pool ACL's 403 saved the non-pool subset; an in-pool pre-existing guest would have been
|
||
destroyed, and any broad-token deployment re-arms the bug. Root cause: SameTxnCreated/scratch provenance
|
||
was ASSUMED, never verified against proof-of-launch. The fix makes a `RestoreLXC` UPID the sole destroy
|
||
authorization, in ALL THREE destroy paths — the pool ACL is defense-in-depth again, not the guard.
|
||
|
||
- **F1a `internal/reconcile/bringup.go` `runBringUp`:** the compensating-rollback defer is gated on
|
||
`launched` (set only after the restore POST is accepted). A synchronous restore failure (no UPID)
|
||
closes the owning entry terminal-failed WITHOUT any destroy. The pre-restore `OpStarted` journal
|
||
append is kept (crash-safety); `rollbackBringUp` is now only ever called launch-proven.
|
||
- **F1b `internal/reconcile/restoretest.go` `runScratchTest`:** same `launched` gate on
|
||
`teardownScratch` — a synchronous restore refusal never destroys the picked band vmid.
|
||
- **F1c `internal/reconcile/recover.go` `Recover`:** the no-UPID "POST never confirmed → abandon
|
||
fail-safe" check now runs BEFORE the Scratch/Rollback dispatch — a no-UPID Scratch/Rollback entry is
|
||
abandoned (marked failed, NO destroy) instead of destroy-by-vmid-existence. Recover is now safe by
|
||
DESIGN, not by the pool-blind "already gone" accident the campaign observed.
|
||
- **F2 `restoretest.go` `RunRestoreTest` band-advance:** a band vmid PVE refuses with "already exists"
|
||
(an invisible squatter — new `pveAlreadyExists`, mirrors `pveConfigLock`, never misclassifies a real
|
||
restore failure) is skipped and the next free band vmid tried (bounded by the band width;
|
||
`pickScratchVMID` gained an exclude set). A fully-occupied band → `Skipped` (scheduler raises no
|
||
"backup unrestorable" alert), never FAIL — one squatter no longer permanently breaks the restore-test.
|
||
- **Accepted residual (by design):** a crash in the one-statement window between obtaining the UPID and
|
||
journaling it leaks a half-built guest Recover won't destroy — cleanable, and vastly preferable to
|
||
destroying an innocent guest.
|
||
- Tests: red-proof companions verified (gates reverted → `TestRunBringUp_NoLaunchNoDestroy`,
|
||
`TestRunRestoreTest_RestoreNoLaunchNoTeardown`, `TestRecover_{BringUp,Scratch}NoUPIDAbandoned` all
|
||
fail with the innocent-guest destroy); no-regression `…LaunchedTaskFailureStillTearsDown` + rollback
|
||
table now includes an explicit restore-task-failure case; F2 advance + squatter-full-band-skips tests.
|
||
Live-validated on felhom-pve (provision onto existing 9001 → no destroy armed; restore-test advances
|
||
past a 990000 decoy). `go build/vet/test ./...` clean.
|
||
|
||
## v0.59.0 — report backing device + capacity for a registry-sourced drive in /disks (2026-07-01)
|
||
|
||
Completes the `/disks` representation for a registry-sourced (raw, no-PVE-storage) drive: the agent-view
|
||
showed "—" for the device and no size bar, because the union row never populated `backing_device` or
|
||
`total_bytes`/`used_bytes` (Observe drives get those from `pvesm status`, which a raw drive has none of).
|
||
|
||
- **`internal/localapi/disks.go` handleDisks registry union:** resolve `BackingDevice` from the fs-UUID
|
||
(`storage.ByUUIDDevicePath`) and read capacity via `statfsCapacity` (new build-tagged
|
||
`capacity_linux.go` = `syscall.Statfs` on the mount; `capacity_other.go` = no-op for dev builds).
|
||
- `go build/vet/test ./...` clean (Linux + Windows dev). Live: the registry drive now shows its device +
|
||
size in the agent-view, matching the Observe-sourced drives.
|
||
|
||
## v0.58.0 — report GuestPath/BoundUnderParent for a registry-sourced drive in /disks (2026-07-01)
|
||
|
||
Last piece of first-class raw-drive support: the `/disks` union row for a registry-sourced drive (Impl-2a
|
||
— a drive with no PVE storage) omitted `GuestPath` + `BoundUnderParent`, so the controller read it as
|
||
"Leválasztva" (disconnected) even though it was mounted + bound + live in the guest.
|
||
|
||
- **`internal/localapi/disks.go` handleDisks registry union:** populate `GuestPath` (`StablePathForRaw`)
|
||
+ `BoundUnderParent` (`boundUnderParent`) on the registry row, identical to the Observe path — so a
|
||
registry-only drive reports its true bound/active state. `go build/vet/test ./...` clean.
|
||
|
||
## v0.57.0 — re-assert a RAW drive's guest-bind (ReassertGuestBinds mount-table fallback) (2026-07-01)
|
||
|
||
Completes the raw-drive durability the v0.56.0 fix started. `ReassertGuestBinds` (the startup / drive-
|
||
returned reconcile that re-binds an enrolled drive's felhom-data under the shared parent so it's live in
|
||
the guest) built its durable-id→mount map from `Observe()` only — so a RAW enrolled drive was never found
|
||
("enrolled drive not present"), and its in-guest bind was not re-asserted after a reboot or a watchdog
|
||
re-mount (the drive would show "Leválasztva" in the controller).
|
||
|
||
- **`internal/localapi/disks.go` `ReassertGuestBinds`:** augment the durable-id→mount map from the mount
|
||
table — each raw `/mnt/<name>` mount → its device fs-UUID (`HostReader.Mounts`+`ResolveUUID`), skipping
|
||
the `/mnt/felhom-drives` bind (AttachDrive wants the raw path). Observe entries still win. Also: an
|
||
Observe failure is no longer fatal (fall through to the mount-table scan) so raw drives re-assert even
|
||
if the PVE view is momentarily unavailable.
|
||
- **Wiring fix (latent):** `buildLocalAPIServer` never passed `Options.HostReader`, so the local-API
|
||
server's `host` was nil in production — the v0.56.0 `durableIDForMount` raw fallback (and the role gate's
|
||
host classification) silently no-op'd. Now wired to `storage.NewProcHostReader()`. This is what makes
|
||
the v0.56.0 + v0.57.0 raw-mount resolutions actually fire live.
|
||
- `go build/vet/test ./...` clean. With v0.56.0 (guest-bind now RECORDED for raw drives) this closes the
|
||
reboot/reconnect guest-bind durability gap for raw drives end-to-end.
|
||
|
||
## v0.56.0 — record intent/guest-bind for a RAW enrolled drive (durableIDForMount fallback) (2026-07-01)
|
||
|
||
Surfaced by the first live raw enrollment (Impl-2b): a raw drive is not a PVE storage, so
|
||
`durableIDForMount` (Observe-based) returned "" for it → the enroll's intent + guest-bind recording
|
||
logged "durable-id unresolved" and silently skipped. Result: the drive mounted + bound + usable, but was
|
||
NOT intent-tracked (so `RegistryKnownTargets` — which gates on intent ≠ new — didn't health-track it) and
|
||
its guest-bind wasn't persisted.
|
||
|
||
- **`internal/localapi/disks.go` `durableIDForMount`:** after the Observe lookup, fall back to resolving
|
||
the fs-UUID directly from the mount table — the device mounted at `where` (via `HostReader.Mounts`) →
|
||
its by-uuid identity (`HostReader.ResolveUUID`) → `uuid:<fs-uuid>` (the SAME scheme Observe derives, so
|
||
intent keys stay consistent). Fixes intent recording (enroll/eject) AND guest-bind recording for raw
|
||
drives; the PVE-storage path is unchanged.
|
||
- Test `TestDurableIDForMount_RawFallback` (+ red-proof: Observe-only → ""). `go build/vet/test ./...` clean.
|
||
- **(Residual noted here fixed in v0.57.0:** `ReassertGuestBinds` raw-drive guest-bind re-assert.)
|
||
|
||
## v0.55.0 — raw-device discovery + registry-sourced drive tracking (Impl-2a) (2026-07-01)
|
||
|
||
Agent backend for drive enrollment (SPIKE-drive-enrollment §SQ1/SQ4/SQ5). Makes raw (non-PVE-storage)
|
||
drives (a) discoverable for enrollment and (b) health-tracked WITHOUT being a PVE storage — so a drive
|
||
enrolled the new way isn't enrolled-but-untracked (the 3b-fix false-detach class). No mkfs here (Impl-1
|
||
owns it); the controller wizard rewiring is Impl-2b.
|
||
|
||
- **`GET /disks/candidates`** (`internal/storage/candidates.go` + `internal/localapi/disks.go`):
|
||
enumerates host whole-disks from `/sys/block`, runs the Impl-1 unclaimed filter, and returns the free
|
||
ones with probe info (size/model/FS/data-bearing/durable-id preview), split into `initialize` (all
|
||
unclaimed) and `attach` (the subset carrying a mountable ext4/xfs FS). Fail-safe carries through (a
|
||
device not provably unclaimed is omitted).
|
||
- **`RegistryKnownTargets`** (`internal/storage/registry_known.go`): the watchdog's known-DRIVE set now
|
||
comes from the intent registry + Felhom `.mount` units, NOT `Observe()` (PVE storages). A unit is
|
||
tracked iff its intent ≠ `new` (enrolled/ejected/decommissioned; the watchdog's existing IntentReader
|
||
gate still decides re-mount). `main.go` swaps the watchdog `KnownTargets` source. **`Observe()` is
|
||
KEPT** for real PVE storages (local/local-lvm/pbs) — reports + the `/disks` view.
|
||
- **`handleDisks` union:** additive + deduped-by-mount-path — appends registry drives Observe doesn't
|
||
surface (a registry-only drive now appears in the agent-view) without dropping any Observe row (can't
|
||
regress the current view).
|
||
- **Existing-drive migration** (`ReconcileExistingDrives`, idempotent, at agent start): records each
|
||
currently-mounted Felhom-unit drive as `enrolled` so the registry-sourced `Known()` tracks it without
|
||
its legacy PVE dir-storage. Does NOT create/remove PVE storages.
|
||
- **Tests:** RegistryKnownTargets (enrolled tracked / new excluded / ejected tracked) + the **red-proof**
|
||
(Observe-based `Known()` misses a drive with no PVE storage; the registry provider tracks it);
|
||
migration idempotency; candidates init/attach split. `go build`/`vet`/`test ./...` clean. Watchdog/
|
||
HostLiveness/Remounter unchanged (only the source injected).
|
||
|
||
## v0.54.0 — format-safety foundation: unclaimed-disk guard + guarded-mkfs wrapper (2026-07-01)
|
||
|
||
Impl-1 (SPIKE-drive-enrollment-2026-07-01). Hardens the destructive `Format`/mkfs path BEFORE the
|
||
enrollment feature: today `Format` delegates authorization to its caller and only checks `DataBearing`
|
||
(has-data), which is insufficient — the OS disk is data-bearing yet catastrophic — and the sudoers
|
||
permits `mkfs /dev/*`. Two independent, layered guards:
|
||
|
||
- **Part A — mandatory unclaimed-disk guard inside `Format` (the primary safety).** New
|
||
`internal/storage/claim.go`: `classifyClaim` (pure) + `gatherClaimFacts` refuse to format any device
|
||
not provably UNCLAIMED — reusing `SystemDisks` (OS disk) + lsblk FSTYPE (LVM2_member/zfs_member/
|
||
linux_raid_member/crypto_LUKS/swap) + foreign mounts + read-only + authoritative `pvs`/`zpool`. A
|
||
Felhom-owned mount under `/mnt/felhom-drives` is NOT a foreign claim (re-init stays allowed; the
|
||
DataBearing wipe-confirm still gates data loss). **FAIL-SAFE: any read error / undeterminable topology
|
||
→ CLAIMED → refuse.** The guard is in `Format` (lowest layer), not the handler, so no caller can
|
||
bypass it. Read-only sudoers additions: `pvs`, `zpool status` (in `FELHOM_DISK`).
|
||
- **Part B — guarded-mkfs wrapper below the agent.** `configs/felhom-mkfs-guarded.sh` (root, 0755) is
|
||
now the ONLY mkfs path the sudoers allows (`FELHOM_FORMAT` no longer allowlists raw `mkfs.*`). It
|
||
re-checks the cheap catastrophic cases (system disk / LVM PV / foreign mount) and refuses — so even
|
||
an agent bug/compromise can't mkfs the OS disk. `Format` execs the wrapper (`<device> <fstype>`) via
|
||
`Binaries.MkfsGuarded`.
|
||
- **Tests:** `claim_test.go` — table-driven `classifyClaim` (every claim signal + fail-safe + the two
|
||
allow cases) incl. the **red-proof** (a claimed, non-data-bearing OS disk: removing the isSystem check
|
||
flips it to allowed → test fails, proving the guard adds safety beyond `DataBearing`); Format-guard
|
||
integration tests (refuses system disk / LVM member, allows unclaimed → wrapper invoked); capability
|
||
manifest updated (mkfs sample → the wrapper). `go build`/`vet`/`test ./...` clean.
|
||
- **Deferred to Impl-3:** a raw disk passed through to ANOTHER VM looks unused to the host — a host-level
|
||
filter can't detect it; the operator gate (shared-box mode) closes that. Impl-1 closes everything
|
||
host-visible (a strict improvement over today's no-guard state).
|
||
|
||
## v0.53.0 — restore guests INTO the felhom pool (pool-scoped-ACL enabler) (2026-07-01)
|
||
|
||
Colleague-safety batch #4 phase b (agent half). Enables the agent token to be scoped from `/` to
|
||
`/pool/felhom` + `/storage/<targets>` (real blast-radius containment on a shared host) by making every
|
||
restore allocate the guest INTO the pool — the only way a fresh vmid authorizes under a pool-scoped
|
||
token. Grounded by `felhom.eu/documentation/audits/SPIKE-pool-scoped-acl-2026-07-01.md` (PASS).
|
||
|
||
- **`internal/proxmox/mutate.go`:** `RestoreLXCOptions` gains `Pool string`; `RestoreLXC` sends
|
||
`pool=<p>` only when non-empty (`pct restore --pool`). Omit-when-empty (a broad-token restore needs
|
||
no pool) — unit-tested + red-proofed.
|
||
- **`internal/reconcile`:** new `const DefaultPool = "felhom"` (single source of truth); `BringUpSpec`
|
||
gains `Pool`, threaded to the bring-up restore. **Both restore sites** now pool the guest: the
|
||
provision/DR bring-up (`Pool: spec.Pool`, set to `DefaultPool` by the CLI) AND the **restore-test**
|
||
scratch guest (`Pool: DefaultPool`) — the latter closes SPIKE residual #2 (a pool-scoped token would
|
||
otherwise 403 on the out-of-pool scratch guest).
|
||
- **`cmd/felhom-agent/main.go`:** both `BringUpSpec` literals (bring-up/DR + provision) set
|
||
`Pool: reconcile.DefaultPool`.
|
||
- **No ACL/priv change in the agent** — that ships in the host-install script (v1.6.0). The pool param
|
||
is INERT until the token is granted `Pool.Allocate` at `/pool/felhom` and the pool exists; publishing
|
||
is therefore safe ahead of the coordinated ACL swap.
|
||
- New tests: `proxmox.TestRestoreLXC_PoolParam` (set → `pool=felhom`; empty → omitted),
|
||
`reconcile.TestRestoreSitesUsePool` (both restore sites carry `DefaultPool`). Both red-proofed
|
||
(unconditional `Set` → omit test fails; drop either site's `Pool` → both-sites test fails). `go
|
||
build`/`vet`/`test ./...` clean.
|
||
|
||
## v0.52.0 — operator-opt-in CPU/RAM cap for the provisioned guest (`-cores` / `-memory`) (2026-07-01)
|
||
|
||
Colleague-safety batch #3. So a trial appliance guest on a colleague's SHARED production Proxmox does
|
||
not pressure his existing guests, the operator can now cap the guest's CPU cores + RAM **at provision
|
||
time, before the guest's first boot** (the peak container-pull moment). Pure CLI→spec plumbing — the
|
||
reconcile engine already applied the cap; this only wires the flags to it.
|
||
|
||
- **`cmd/felhom-agent/main.go`:** new `-cores N` / `-memory M` (MiB) flags for
|
||
`--selftest=bring-up|provision` (0 = keep the golden's baked size). They flow through `bringUpSizing`
|
||
(now carries `Cores`/`MemoryMB`) into the `reconcile.BringUpSpec{Cores,MemoryMB}` built by BOTH
|
||
`runSelftestBringUp` and `runSelftestProvision`. The `-selftest` usage string documents them.
|
||
- **No engine change.** `internal/reconcile/bringup.go` already carries `BringUpSpec.Cores/MemoryMB`
|
||
(0 = leave as restored) and `buildBringUpConfig` already emits `cores`/`memory` into the SAME
|
||
coalesced config PUT as the identity reset — which runs BEFORE `e.api.Start`, so the cap lands
|
||
pre-boot. Rejected the `pct set`-post-provision alternative (runs after boot = an uncapped window;
|
||
bypasses the token/audit; second config source of truth).
|
||
- **Omit-when-zero guarantee:** an unset cap (0) emits NEITHER `cores` NOR `memory`, so an uncapped
|
||
provision keeps the golden defaults (no regression for the normal single-purpose box) and can never
|
||
shrink the guest to 0 cores. New pure-function test `TestBuildBringUpConfig_ResourceCaps` asserts
|
||
both the set (`cores=2`,`memory=4096`) and the absent-when-unset cases; a red-proof (unconditional
|
||
emit) was run and confirmed to fail the omit assertion, then reverted.
|
||
- **Deploy dependency:** a FRESH host-install `--cores`/`--memory` (felhom.eu script v1.4.0) requires
|
||
the hub artifact manifest to serve **agent ≥ v0.52.0**, else the old agent rejects the unknown flag.
|
||
The flags are opt-in, so nobody hits this until they intentionally cap.
|
||
- `go build` / `go vet` / `go test ./...` clean.
|
||
|
||
## v0.51.0 — local vzdump retention default (`--prune-backups keep-last=3`) (2026-06-30)
|
||
|
||
The PREVENTIVE counterpart to the hub's host_disk + storage_fill detectors: the agent's periodic local
|
||
whole-guest vzdump now prunes its own old archives, so a box can't refill its own root via its own backups
|
||
(the felhom-pve incident's root cause — that vzdump carried no retention, ~18 dumps piled under
|
||
`/var/lib/vz/dump`).
|
||
|
||
- **`internal/proxmox/mutate.go`:** `VzdumpOptions.PruneBackups` → passed as PVE's `--prune-backups` on
|
||
the vzdump POST (vmid+storage scoped, so PVE prunes only THIS guest's archives on THIS storage).
|
||
- **`internal/backup/runner.go`:** `NewBackupRunner` gains a `retention` arg; the backup applies it via
|
||
`localPruneSpec` ONLY when the target is a **non-PBS** storage (resolved via `ListStorage`) — PBS offsite
|
||
retention is a separate lifecycle and is never pruned by the per-run flag. **Fail-safe:** if the target
|
||
type can't be confirmed (lookup error / not found) the run SKIPS pruning rather than risk pruning PBS
|
||
(the detectors remain the safety net). Only the periodic local-API runner sets retention; the
|
||
restore-test / selftest runners pass "".
|
||
- **`internal/config/config.go`:** `backup.local_backup_retention` (keep-last N) with `KeepLast()` clamped
|
||
to **≥1** (0/unset/negative → default 3) — a mis-config can NEVER prune the just-made backup —
|
||
+ `PruneBackupsSpec()` → `keep-last=N`. Wired into the local-API backup runner (`main.go`).
|
||
- **Seeding:** `felhom.eu scripts/felhom-host-install.sh` seeds `local_backup_retention: 3` in the agent
|
||
config; the code default also protects any box where it is unset (KeepLast → 3) from day 0.
|
||
- F2-b stale-vzdump-lock recovery untouched.
|
||
- Tests: the local vzdump carries `--prune-backups keep-last=3` (+ companion: no-retention runner emits no
|
||
prune); **PBS is never pruned** (+ companion: same retention on a local target IS applied);
|
||
fail-safe-on-unknown-target; the **keep-last≥1 clamp** companion. `go build/vet/test ./...` green.
|
||
|
||
## v0.50.0 — NAS network storage Part A1: NFS/SMB automount foundation (2026-06-30)
|
||
|
||
Agent foundation of the validated `SPIKE-nas-storage-2026-06-29.md` (verdict READY): a customer NAS can
|
||
serve **bulk media** to a media app. The agent mounts a NAS share **host-side** under
|
||
`/mnt/felhom-drives/<name>` via a systemd `.automount` (+ `.mount`) pair; it propagates into guest 9201 for
|
||
free through the existing shared `mp8` bind (no new mountpoint, no restart). A NAS is a **distinct storage
|
||
class** — it carries **no durable-id** and never enters the drive enroll/eject/decommission/wipe/SMART/
|
||
watchdog machinery. **Bulk-media class only; STOP before A2 (controller registry/UI) + B (restic-SFTP).**
|
||
|
||
- **`internal/storage/netmount.go` (NEW).** `NetworkMountSpec` + the locked SPIKE recipe:
|
||
- **NFS (preferred):** `What=server:/export`, `Type=nfs4`, `Options=vers=4.1,soft,timeo=50,retrans=2,noatime,_netdev`.
|
||
`soft` is the failure-isolation knob (clean EIO, never a `df`/guest wedge); a default `hard` mount is
|
||
never emitted. The `+100000` uid mapping is the **export's** job (`anonuid=101000`), so the client mount
|
||
carries no uid.
|
||
- **SMB (fallback):** `What=//server/share`, `Type=cifs`,
|
||
`Options=vers=3.0,credentials=<0600 file>,uid=<+100000>,gid=<+100000>,forceuid,forcegid,file_mode=0664,dir_mode=0775,_netdev`
|
||
(plain octal modes, never setgid 2775). **The +100000 rule** (container uid/gid N = host N+100000): a
|
||
container uid 1000 renders `uid=101000` so the guest sees its native id and reads+writes; a naïve `+0`
|
||
lands as `nobody:nogroup` (not writable) — the documented trap, asserted by a companion test.
|
||
- **`.automount` with `TimeoutIdleSec`** (on-demand + idle-unmount): an idle NAS reboot is a non-event.
|
||
- **per-share liveness** (`ListNetworkMounts`): TCP-probes the NAS endpoint (2049/445) + reads
|
||
`/proc/mounts` — it never `stat`s the (possibly EIO/D-state) mountpoint, so a black-holed NAS cannot
|
||
wedge a list. Health `ok | idle | unreachable`, scoped to the affected share, never box-wide.
|
||
- **role gate** `NetworkMountRole`: network storage is **bulk-userdata only** — confined to the
|
||
`/mnt/felhom-drives` namespace; any other target is refused (most-protected).
|
||
- Full validation (`ValidateNetworkMountSpec`) before any unit is rendered: share name (safe segment),
|
||
server, NFS export (absolute, no traversal) / SMB share name, uid/gid range, creds path.
|
||
- **Drive-machinery bypass (Scenario D).** `parseFelhomMountUnit` (the host-reboot drive re-assert's
|
||
classifier) explicitly refuses any unit carrying the network marker, so a NAS mount is never given a
|
||
durable-id, SMART-probed, or re-asserted as a drive. Companion red-proof: the same by-uuid-shaped unit
|
||
with the drive marker DOES parse — the guard is the discriminator, not luck.
|
||
- **`internal/localapi/netstorage.go` (NEW).** Self-scoped endpoints `POST /netstorage/add`,
|
||
`GET /netstorage`, `POST /netstorage/remove`. SMB credentials are written **out-of-band** to a 0600 file
|
||
the agent owns (never in git, never in a plaintext registry, never logged). Role-gated to the user-data
|
||
namespace.
|
||
- **sudoers:** new narrow `FELHOM_NETMOUNT` alias (install/enable/disable/stop the `.automount` + remove the
|
||
felhom mount-unit files; the `.mount` half reuses `FELHOM_MOUNT`, the mountpoint mkdir reuses
|
||
`FELHOM_INTERMEDIARY`). `visudo -cf` clean.
|
||
- **config:** `privileged.smb_creds_dir` (default `/var/lib/felhom-agent/smb-creds`).
|
||
- **Runtime deps:** `mount.nfs` (nfs-common) + `mount.cifs` (cifs-utils) present on the host (confirmed live).
|
||
- Tests: exact NFS/SMB option-set string-asserts + the +100000 companion; validation matrix; role gate;
|
||
unit round-trip + health; the drive-machinery guard + companion; Ensure/Remove command sequences.
|
||
|
||
## v0.49.0 — reboot-during-backup stale-lock recovery (F2-b) + shared-parent script redeploy fix (F2-a) (2026-06-30)
|
||
|
||
Closes the two host-reboot findings from `TESTRUN-fullstack-2026-06-29.md`.
|
||
|
||
- **F2-b — startup stale-lock recovery (`internal/localapi/stalelock.go`, NEW).** A host reboot DURING a
|
||
vzdump backup leaves the guest with a `snapshot-delete`/`backup` lock + a dangling `vzdump` snapshot;
|
||
`onboot:1` then can't start the locked CT → the customer box stays DOWN until a human runs `pct unlock`.
|
||
The agent now self-heals at startup (`Server.RecoverStaleLockedGuests`, called alongside
|
||
`ReassertGuestBinds`/`RecoverFormatJob`): for each guest carrying a backup lock, **only when no vzdump is
|
||
genuinely in-flight** (the load-bearing invariant — at startup the agent's own backup loop hasn't run, so
|
||
the lock is stale; the guard fails SAFE if it can't confirm), it `pct unlock`s → deletes the dangling
|
||
`vzdump` snapshot (API + WaitTask) → starts the CT **iff** `onboot` and not already running. Scope is
|
||
strictly the two vzdump locks; `migrate`/`disk`/`create`/… are left untouched. Idempotent.
|
||
- **`internal/proxmox`:** new reads `GuestConfig.Lock()`/`OnBoot()`, `Client.ListSnapshots`,
|
||
`Client.ListRunningTasks`, and the `Snapshot` type. Reads + snapshot-delete + start go through the API
|
||
token; only `pct unlock` shells out (no API equivalent).
|
||
- **sudoers + capability manifest:** new narrow grant `FELHOM_STALELOCK = /usr/sbin/pct unlock [0-9]*`
|
||
and Critical capability `stalelock-unlock` (a stuck-locked guest = customer box down). `visudo -cf` clean;
|
||
covered by the manifest↔sudoers build gate.
|
||
- Tests: recovery sequence + companions — no-lock touches nothing; non-backup lock left alone; onboot=0
|
||
unlocked-but-not-started; delsnapshot only when a snapshot exists; **invariant guard** (live backup →
|
||
not cleared; unconfirmable → fail-safe); already-running → not restarted.
|
||
|
||
- **F2-a — shared-parent boot script never redeployed (`internal/localapi/intermediary.go`).** The host's
|
||
`/mnt/felhom-drives` was still in root's `shared:1` peer group (so every drive bind DOUBLED) because the
|
||
live boot script predated the v0.36.6 `make-private` fix. Root cause: `EnsureSharedParent` gated the
|
||
(re)install on the **unit** file only, so a script-only change never deployed. Fixed: the new
|
||
`sharedParentInstallStale` helper compares **both** the script and the unit (missing or differing →
|
||
reinstall). Boot-time-only — it rewrites the on-disk script; it does NOT churn the live mount (the live
|
||
bind/make-private/make-shared stays guarded on `!isHostMountpoint`). Verified empirically on the host:
|
||
the correct `bind → make-private → make-shared` sequence gives the parent its own group + no doubling.
|
||
- Tests: a stale-script/current-unit case triggers reinstall (the F2-a regression); both-current is a
|
||
no-op; missing files are stale; a content guard asserts the shipped script keeps `make-private`.
|
||
|
||
- **Live-caught fixes (same version, found during felhom-pve validation):** PVE 9.x rejects
|
||
`GET /nodes/{node}/tasks?running=1` (HTTP 400 "property not defined in schema") — the invariant guard
|
||
now uses `?source=active`. And the unprivileged-LXC start emits a benign `WARNINGS: 1` (systemd-nesting)
|
||
advisory that false-failed the recovery's start — `Start` now uses `AllowWarnings` (matching the
|
||
restore-test's start step).
|
||
- **§D supervised reboot — both findings live-validated.** F2-a: after reboot `/mnt/felhom-drives` came
|
||
up as its OWN peer group (`shared:94`, not `shared:1`) with no doubling; guest sees both drives, apps
|
||
healthy. F2-b: a reboot with the exact stale state (induced `snapshot-delete` lock + a real dangling
|
||
`vzdump` snapshot) reproduced the stuck symptom (pve-guests "CT is locked (snapshot-delete)" → start
|
||
failed) and the agent auto-recovered — unlock → removed the real dangling snapshot → started the CT.
|
||
Zero spurious operator pages on the reboots. See
|
||
`felhom.eu/documentation/audits/TESTRUN-fullstack-2026-06-29.md`.
|
||
- Version `0.48.0 → 0.49.0`.
|
||
|
||
## v0.48.0 — report the served local-API leaf fingerprint (hub-side re-key detection, Part A) (2026-06-29)
|
||
|
||
The agent now rides its **served leaf fingerprint** on every host report so the hub can detect an
|
||
agent re-key fleet-wide (the last self-health leg — `host_leaf_changed`, hub v0.22.0).
|
||
|
||
- **`internal/hub/report.go`:** new `HostReport.LeafFingerprint string` (`leaf_fingerprint`) — the
|
||
SHA-256 of the leaf the agent currently serves. Empty when the local API is disabled (no leaf) → the
|
||
hub treats "" as unknown, never an alert. Not a secret.
|
||
- **`internal/hub/collect.go` + `cmd/felhom-agent/main.go`:** `Collector.SetLeafFingerprint(fp)` threads
|
||
the `fp` from `EnsureLeaf` (the SAME value the loud LOADED/REGENERATED log reports) into every report,
|
||
next to `Capabilities`.
|
||
- Tests: the report includes the fp when set, `""` when unset (local API disabled); golden + contract +
|
||
field-names tests updated (cross-repo golden mirrors `leaf_fingerprint`). Version `0.47.0 → 0.48.0`.
|
||
|
||
## v0.47.0 — controller-swap verify hardening: reject a crash-looping no-healthcheck image (F1) (2026-06-29)
|
||
|
||
Closes F1 from the no-mercy testrun: a controller image with **no HEALTHCHECK that crash-loops** could
|
||
land a single "Running" inspect poll → the swap marked it healthy → **no rollback** (alpine tagged as
|
||
the controller passed in ~4 s, then `Restarting (0)`). The real controller image has a healthcheck so
|
||
the live severity is low, but the rollback safety net had a hole.
|
||
|
||
- **`internal/localapi/controllerswap.go`:** `controllerHealthy` now also reads `{{.RestartCount}}` (a
|
||
4th `docker inspect -f` field) — `running && RestartCount>0` → not-ok (a process that has already
|
||
crash-restarted isn't stably up, regardless of healthcheck). It also signals `needsDwell` for the
|
||
no-healthcheck (`none`) case. `verify` adds a **stability dwell**: a no-healthcheck image must report
|
||
ok on `verifyDwell` (=3) **consecutive** polls before it's accepted; a real `healthy` result is
|
||
trusted immediately (Docker already gated it). Any not-ok resets the dwell. Timeout → existing
|
||
rollback path runs. No change to writeImage, the sudoers grants (the `*` in `docker inspect -f *`
|
||
spans the extended template — confirmed live), or the state-file/rollback orchestration.
|
||
- Tests: F1 **red-proof** (`RestartCount>0` → verify false; companion: rc=0+dwell=1 verifies → the rc
|
||
check is what blocks it); the **dwell** (single ok then crash → verify false; companion dwell=1
|
||
accepts it); a real `healthy` image verifies promptly (no false rollback). Existing
|
||
`RollbackOnUnhealthy` / `HealthyWithNoHealthcheck` stay green. Version `0.46.0 → 0.47.0`.
|
||
|
||
## v0.46.0 — leaf lifecycle: signal + loud-log a regenerated leaf (prevention, Part B.1) (2026-06-29)
|
||
|
||
Makes an accidental local-API leaf **regeneration** (the 2026-06-28 root→non-root migration class —
|
||
moving `/var/lib/felhom-agent` aside silently minted a new leaf → every controller's pin invalidated
|
||
for days) **visible immediately** instead of silent.
|
||
|
||
- **`EnsureLeaf` now returns `generated bool`** (`internal/localapi/cert.go`): false = an existing
|
||
pair was LOADED (stable fingerprint), true = a fresh leaf was GENERATED.
|
||
- **Loud call-site (`cmd/felhom-agent/main.go`):** a load logs `INFO local-api leaf LOADED`; a
|
||
regeneration logs **`WARN local-api leaf REGENERATED — any previously issued bootstrap pins are now
|
||
INVALID; controllers will fail the pin check until re-bootstrapped`** (with the new fingerprint).
|
||
- No new sudo/capability surface — pure return + log change. The companion install-script preservation
|
||
(`--preserve-state-from` + the populated-host guard) lives in `felhom.eu/scripts/felhom-host-install.sh`.
|
||
- Tests: `EnsureLeaf` first call `generated==true`, second `generated==false` AND **same fingerprint**
|
||
(persistence keeps the pin stable). Version `0.45.0 → 0.46.0`.
|
||
|
||
# Changelog
|
||
|
||
All notable changes to **felhom-agent** are recorded here. Update on every code
|
||
change that gets pushed.
|
||
|
||
## v0.45.0 — controller-swap under non-root: stdin `tee` write + narrow sudoers grants (Option A) (2026-06-29)
|
||
|
||
Restores fleet controller-swap / managed auto-update under the **non-root** agent — the one capability
|
||
the 2026-06-29 sudoers audit deliberately left broken because the old write vector needed arbitrary
|
||
in-guest execution. Mechanics spike-proven
|
||
(`felhom.eu/documentation/audits/SPIKE-controllerswap-narrow-grants-2026-06-29.md`, GO). **No controller
|
||
change** — the swap endpoint contract is unchanged; only the agent's internal write mechanism + the
|
||
allowlist.
|
||
|
||
- **`writeImage` no longer shells out.** Was `GuestExec("bash","-c","printf '%s\n' '<img>' > <file>")`
|
||
(the swap's only interpolated/shell vector). Now
|
||
`GuestExecStdin(strings.NewReader(img+"\n"), "tee", "/etc/felhom-controller-image")` — the image ref
|
||
is piped on **stdin** into an in-guest `tee`; no shell, no interpolation. The trailing `\n` keeps the
|
||
on-disk bytes byte-identical to the golden's `printf '%s\n'`, and the bootstrap reads `IMAGE=$(cat …)`
|
||
(newline-stripping), so the write is consumed identically. `ValidControllerImage` still gates upstream.
|
||
- **New stdin seam (no fenced-runner bypass):** `proxmox.Runner.RunStdin` / `ExecRunner.RunStdin` (Run
|
||
with `cmd.Stdin`), `GuestBinder.GuestExecStdin`, and `GuestExecutor.GuestExecStdin` — the swap routes
|
||
stdin through the SAME `sudo -n` fenced runner as every other privileged op.
|
||
- **`FELHOM_CONTROLLERSWAP` sudoers alias (5 narrow, auditable grants):** `cat <fixed file>`,
|
||
`docker image inspect *`, `docker inspect -f *`, `systemctl restart <fixed unit>`,
|
||
`tee <FIXED image file>`. **No general `pct exec`, no `bash -c`** — the spike's negative controls
|
||
(arbitrary exec, `tee` to any other path, `docker rm`, `rm -rf`) stay denied. The 5 are added to the
|
||
v0.44.0 **capability manifest** (Critical — a silently-broken fleet auto-update is operator-alert-worthy),
|
||
so the self-probe watches them and the build-test asserts grant↔code coverage (companion red-proof:
|
||
dropping the `tee` grant fails the gate — demonstrated red→green on the real file).
|
||
- Existing swap tests (happy / rollback-on-unhealthy / image-absent / no-healthcheck / bad-image /
|
||
single-flight) pass over the new write path; a new test asserts the write is stdin-`tee` with exact
|
||
`image\n` and **no** shell vector. Version `0.44.0 → 0.45.0`.
|
||
|
||
## v0.44.0 — privileged-capability self-probe (build-time manifest test + runtime probe + hub snapshot) (2026-06-29)
|
||
|
||
The agent now self-checks the `sudo -n` grants it depends on, so a missing allowlist entry (the
|
||
2026-06-28 cutover class: lxc-info/make-private/…) is caught LOUD — in CI at build time and on the
|
||
host at runtime — instead of surfacing days later as user-visible breakage. **First slice of agent
|
||
self-health; the controller↔agent channel check is a separate later task.**
|
||
|
||
- **`internal/capability` (NEW):** a `Manifest()` of the required `(binary, representative-arg)`
|
||
vectors (seeded from the 2026-06-29 audit — the OK + CLOSED rows; the SURFACED/DEFERRED rows
|
||
`pct exec *`/`pct create`/`mount UUID`/`sensors` are deliberately excluded). `Prober.Probe` lists
|
||
each against the live policy with `sudo -n -l -- <binary> <args>` (a policy LIST — **never
|
||
executes**, safe for mkfs/pct entries) via a DIRECT runner, plus an `os.Stat` existence check,
|
||
mapping to `ok` / `degraded` ("sudo policy denied" | "binary not found"). A total sudo failure
|
||
(drop-in missing) collapses to ONE aggregate signal. Serve-degraded: the probe never blocks
|
||
startup, panics, or errors.
|
||
- **Build-time gate (`manifest_test.go`):** parses `configs/felhom-agent.sudoers`, translates each
|
||
glob to a regex, and asserts **every manifest vector is covered by a grant** — exactly what would
|
||
have caught the dropped `lxc-info`/`make-private` lines in CI. Includes a **red-proof**: with the
|
||
`lxc-info` line removed from an in-memory copy, the check FAILS for `guest-init-pid` (and passes
|
||
on the real file) — proving the gate is not hollow.
|
||
- **Runtime wiring:** `Probe` runs once at startup (INFO `capabilities self-check N/N ok`, plus an
|
||
ERROR per degraded capability naming the gated feature) and on every hub-report cycle; the snapshot
|
||
rides the report as the new non-nil `HostReport.Capabilities []capability.Status` (golden +
|
||
contract test updated; cross-repo hub copy mirrors it).
|
||
- **No allowlist change**; the live host is post-audit complete, so the probe reports N/N ok — itself
|
||
a live proof the probe agrees with the fixed sudoers. Version `0.43.0 → 0.44.0`.
|
||
|
||
## (sudoers completeness audit, folded into v0.44.0) — close non-root allowlist gaps (2026-06-29)
|
||
|
||
A full audit of every privileged command the agent shells via `sudo -n` against
|
||
`configs/felhom-agent.sudoers`, closing the read-only/fixed-vector gaps left by the 2026-06-28
|
||
root→non-root cutover. **Sudoers-only change — no Go change, no version bump** (the file is fetched
|
||
canonically by the host-install script). Root cause of the multi-drive "attach one, the other drops"
|
||
symptom (audit `felhom.eu/documentation/audits/SPIKE-multidrive-mutual-exclusion-2026-06-29.md`): the
|
||
allowlist was incomplete, so several `sudo -n` calls were denied under the non-root user.
|
||
|
||
- **`lxc-info -n [0-9]* -p -H` → FELHOM_INTERMEDIARY (THE root-cause fix).** `guestInitPID`
|
||
(`intermediary.go:256`) shells this to resolve the guest init PID for
|
||
`GuestSeesMount`→`bound_under_parent`. It was absent from the allowlist → `sudo -n` denied → empty
|
||
PID → **every external drive reported absent** → the controller drive-gate stopped each drive's apps
|
||
(flapping). With the grant, `bound_under_parent` reports truthfully and the gate quiesces.
|
||
- **`mount --make-private /mnt/felhom-drives` → FELHOM_INTERMEDIARY.** `EnsureSharedParent`
|
||
(`intermediary.go:110`) calls it to isolate the shared parent's peer group on first setup; the
|
||
allowlist had only `--make-shared`, so the parent stayed in root's peer group and host submounts
|
||
"doubled". Guarded by a mountpoint check (never re-churns a live parent).
|
||
- **`systemctl restart dnsmasq` → FELHOM_DNSMASQ.** The v0.29.x LAN-DNS fix switched `reload`→`restart`
|
||
(`lanresolver.go restartDnsmasq`) but the allowlist still only permitted `reload` → split-horizon
|
||
DNS self-heal was silently denied under non-root. Added alongside the retained `reload`.
|
||
- **`pct set [0-9]* -onboot 1` → FELHOM_PROVISION.** The provision back-half (`backhalf.go`, F3
|
||
auto-start) sets onboot; only `-mp[0-9]*` was allowed → denied under non-root.
|
||
- **`pct reboot [0-9]*` → FELHOM_GUESTHOOK.** `RebootGuest` (`disks.go:448`, the enroll "activate
|
||
pending binds" fallback) was unmatched.
|
||
|
||
**Surfaced for operator decision (NOT added — would require arbitrary root-in-guest):** `GuestExec`'s
|
||
general `pct exec [0-9]* -- <…>` (controller-swap self-update / Phase-2 managed updates) runs variable
|
||
vectors incl. `bash -c "<interpolated>"` — granting it = arbitrary execution. **Controller-swap is
|
||
currently broken under the non-root agent** until narrow per-vector grants are decided. **Deferred:**
|
||
`sensors -j` (defined-but-unwired AND `lm-sensors` not installed on the host — no live caller, path
|
||
unverifiable). **Not added (no daemon caller):** `pct create …` (CreateGoldenLXC, maintenance/broad),
|
||
`mount UUID=… …` (MountUSBByUUID, legacy/unreferenced). Full audit table in `REPORT.md`.
|
||
|
||
## v0.43.0 — canonical systemd unit + binary published to Gitea (BUNDLE slice) (2026-06-28)
|
||
|
||
Day-0 no longer needs a hand-installed agent. The agent binary is now PUBLISHED to Gitea as a generic
|
||
package and the host-bootstrap script fetches → verifies (sha256 vs the hub-vouched manifest) →
|
||
installs it. This commit adds the **canonical systemd unit** (was hand-made per host) and the publish
|
||
tooling; the binary itself is a version-only rebuild (no behavioural change).
|
||
|
||
- **`configs/felhom-agent.service` (NEW, canonical):** `User=felhom-agent`/`Group=felhom-agent` (the
|
||
documented non-root production model — README "Process model"; `privileged.mode: "sudo"` + the
|
||
narrow sudoers allowlist), `ExecStart=/usr/local/bin/felhom-agent --config
|
||
/etc/felhom-agent/agent.json`, `After=network-online.target pve-cluster.service pveproxy.service`,
|
||
`Restart=on-failure`, `StateDirectory=felhom-agent`. **Deliberately NO sandboxing**, with the reasons
|
||
documented inline:
|
||
- `NoNewPrivileges` is NOT set — it would block the setuid `sudo` the agent needs for every host-root
|
||
op (mount/format/pct/dnsmasq), silently killing all privileged capability.
|
||
- NO mount-namespacing hardening (`ProtectHome`/`ProtectSystem`/`PrivateTmp`/…) — any of those give the
|
||
unit a PRIVATE mount namespace, and the intermediary-mount drive model relies on `mount
|
||
--make-shared`/`--bind` propagating into the running guest; in a private namespace every drive
|
||
enrollment would silently break. The agent shares the host mount namespace; the sudoers allowlist is
|
||
the security boundary.
|
||
- **`scripts/publish-agent.sh` (NEW):** builds (optional) + PUTs the binary to
|
||
`/api/packages/admin/generic/felhom-agent/<ver>/felhom-agent` (Gitea generic), prints
|
||
`AGENT_VERSION` + `AGENT_SHA256`, and does a GET round-trip (re-fetch + sha256 re-check) to prove the
|
||
artifact is fetchable + intact. Pinned to a version (never `:latest`); idempotent (delete-then-PUT);
|
||
asserts the binary's `--version` matches the publish version. Creds via `GITEA_USER/GITEA_TOKEN`
|
||
(falls back to `REGISTRY_USER/REGISTRY_TOKEN`).
|
||
- **`configs/build-golden.sh`:** after the vzdump archive is produced, computes its sha256 and PUTs it
|
||
to `/api/packages/admin/generic/felhom-golden/<golden-version>/golden.tar.zst` (`<golden-version>` =
|
||
the baked controller version), printing `GOLDEN_VERSION` + `GOLDEN_SHA256`. Opt-in (only when the
|
||
Gitea creds are set); the local-golden auto-discovery stays as a fallback.
|
||
- **`configs/felhom-agent.sudoers` (latent bug fix):** escaped the commas in the `lvs -o
|
||
lv_name\,data_percent\,metadata_percent` and `lsblk -o NAME\,FSTYPE\,PTTYPE\,MOUNTPOINT` argument
|
||
lists. Sudoers treats a bare comma as a command separator, so `visudo -cf` REJECTED the file — it had
|
||
never been visudo-validated live because the demo host ran the agent as root+`direct` (sudoers
|
||
unused). The escaped commas still match the agent's real comma-bearing args. Surfaced by the BUNDLE
|
||
live install (the host-install script `visudo -cf`-validates before installing).
|
||
- **`cmd/felhom-agent/main.go`:** `version` 0.42.0 → 0.43.0.
|
||
- The operator records the printed agent + golden version+sha256 in the hub (Configs → "Day-0
|
||
artifacts"); the host-bootstrap script verifies fetched artifacts against those before installing.
|
||
- `go build/vet/test ./...` green.
|
||
|
||
## build-golden.sh — default controller image bumped to current; golden rebuilt at 0.85.1 (2026-06-27)
|
||
|
||
**Operational + a default fix (no agent binary change — version stays v0.42.0).**
|
||
|
||
- `configs/build-golden.sh`: the `CONTROLLER_IMAGE` default (positional arg 6) was a stale
|
||
`…/felhom-controller:0.43.0` — an argument-less golden build baked a wildly old controller, so fresh
|
||
Day-0 boxes started old (the demo started at 0.77). Bumped the default to the **current**
|
||
`…/felhom-controller:0.85.1` so the worst case (no explicit arg) is merely "current", not ancient.
|
||
- **Always pass the controller version explicitly at each rebuild** — this default only bounds the
|
||
worst case. A future `make golden` that resolves the latest pullable tag would remove the need for a
|
||
hand-bumped default (Observation, not this task).
|
||
- **Golden rebuilt at 0.85.1** on `felhom-pve` with the image passed **explicitly**
|
||
(`build-golden.sh 9100 … gitea.dooplex.hu/admin/felhom-controller:0.85.1`). New archive volid:
|
||
`local:backup/vzdump-lxc-9100-2026_06_27-11_42_51.tar.zst` (rootfs 32G + Docker-data 16G + user-data
|
||
8G, all in the archive; mp0+mp1 inclusion confirmed in the vzdump log).
|
||
- **Baked-image verify (cheap, mandatory):** in the build guest `/etc/felhom-controller-image` =
|
||
`…:0.85.1` and `docker images` showed it baked (379 MB). New Day-0 provisions now ship current.
|
||
- The host-bootstrap script auto-discovers the newest golden, so it picks up this rebuild
|
||
automatically. The real demo 9201 was **not** re-provisioned (it is the Phase-2 floor test box).
|
||
|
||
## v0.42.0 — agentic controller update: in-guest image swap + rollback (Phase 1) (2026-06-26)
|
||
|
||
The host agent now owns the in-guest controller image **swap** — the new-architecture replacement for
|
||
the controller's dead in-container `docker compose` self-update. The controller pre-pulls the target
|
||
image (shared docker socket, its own registry token) then asks the agent to swap; the agent — external
|
||
to the controller container, so it survives the controller being killed mid-swap — does the rest and
|
||
**rolls back** if the new controller doesn't come up healthy.
|
||
|
||
- **New local-API routes** (`internal/localapi/controllerswap.go`, token-scoped via `withGuest`):
|
||
- `POST /controller/swap {image}` → **202** `{status:"swapping", previous_image, target_image}`, then
|
||
async: record previous (crash-safety state file `/var/lib/felhom-agent/controller-swap-<vmid>.json`)
|
||
→ confirm the target image is present in the guest (else abort, **no swap**) → write
|
||
`/etc/felhom-controller-image` → `systemctl restart felhom-controller-bootstrap.service` → poll the
|
||
new controller to **healthy** (`docker inspect`, ≤90s) → **roll back** to the previous image + restart
|
||
if it doesn't (the guest is never left without a controller). Single-flight per guest (409 if busy).
|
||
Image ref is strict-validated (`gitea.dooplex.hu/admin/felhom-controller:<semver>`) before any action.
|
||
- `GET /controller/swap/status` → `{state: swapping|done|failed, current, previous, target, error}`.
|
||
- **`GuestBinder.GuestExec`** (`internal/localapi/guestbind.go`): the one `pct exec` seam the swap
|
||
composes over (cat/inspect/write/restart), reusing the fenced root runner.
|
||
- **`--selftest=controller-swap -vmid -image <ref>`**: exercise the primitive directly (the target image
|
||
must already be pulled in the guest).
|
||
- Wired `ControllerSwap: guestBinder` into the local-API server (`cmd/felhom-agent/main.go`).
|
||
- Tests (`controllerswap_test.go`): happy swap, **rollback-on-unhealthy** (+ companion red-proof:
|
||
dropping the rollback leaves the guest on the bad image and fails the test), image-absent no-swap,
|
||
no-healthcheck-running, bad-image 400, single-flight 409.
|
||
|
||
## v0.41.0 — provisioned customer guests auto-start after a host reboot (`onboot:1`) (2026-06-24)
|
||
|
||
**F3 fix.** The provision back-half now sets **`onboot:1`** on the customer guest, so after a host
|
||
reboot/power-cut the customer's whole home-server (controller + apps) comes back **on its own** —
|
||
previously every provisioned guest inherited the golden's `--onboot 0` and stayed **stopped** until a
|
||
manual `pct start` (confirmed live in the stable-path/sys-drive restart campaign, Phase 4.1). The new
|
||
step is a fatal `pct set <vmid> -onboot 1` placed right after the config-mount attach (`backhalf.go`),
|
||
mirroring the config-mount/parent-bind `pct set` ops. **No `startup`/boot-order/delay** — the v0.75
|
||
mountpoint-gate already covers the drive-bind race at boot (Phase 4.4), so the controller won't write
|
||
app data onto the rootfs while the agent re-binds drives.
|
||
|
||
The **golden stays `onboot:0`** (`build-golden.sh` unchanged): a template must not auto-start, and
|
||
`onboot` is a per-guest property the back-half is the right place to set. Unit-tested
|
||
(`TestProvision_SetsOnbootOne` asserts the exact `pct set … -onboot 1` invocation, with a red-proof
|
||
against removing the call). The pre-existing demo guest 9201 (provisioned pre-fix) was remediated
|
||
non-destructively with `pct set 9201 -onboot 1`. **Capstone live-validated (2026-06-24):** destroyed +
|
||
re-provisioned 9201 through the real provision chain with v0.41.0 → fresh `pct config` showed `onboot: 1`
|
||
with no manual set; a subsequent **felhom-pve host reboot** brought 9201 back **running with no manual
|
||
`pct start`** (the `onboot:0` scratch guests correctly stayed stopped), controller + base infra healthy,
|
||
drives re-bound at stable, sys_drive separate — the exact Phase-4.1 failure now passes.
|
||
|
||
## v0.40.0 — third CT volume: SSD user-data (`/mnt/sys_drive`, mp1) baked + `-sysdata-grow` (2026-06-23)
|
||
|
||
**The third golden volume.** Extends the OS/Docker-data split (v0.29.x) to a **three-volume layout**:
|
||
rootfs + Docker-data (`mp0`) + **SSD user-data (`mp1` @ `/mnt/sys_drive`, `backup=1`)** — the
|
||
controller's `system_data_path`. Until now `/mnt/sys_drive` was a plain directory on the 32 GB OS
|
||
rootfs, so the controller correctly warned that SSD app data (`<sys_drive>/felhom-data`) lands on the
|
||
OS drive. Baking it as its own thin volume clears that warning with **zero controller change** (the
|
||
controller already auto-discovers `<sys_drive>/felhom-data` and warns via `system.IsMountPoint`); the
|
||
`mp` under the guest's `/mnt` reaches the controller container through the existing
|
||
`-v /mnt:/mnt:rslave` bind.
|
||
|
||
- **`configs/build-golden.sh`** — `pct create` gains
|
||
`--mp1 ${ROOTFS_STORAGE}:${GOLDEN_SYSDATA_GB},mp=/mnt/sys_drive,backup=1` (new env
|
||
`GOLDEN_SYSDATA_GB=8`, near-empty; provision grows it). The resilience guards are mirrored for `mp1`:
|
||
a `findmnt /mnt/sys_drive` separate-mount assertion, and the vzdump-inclusion guard now aborts if
|
||
**either** `mp0` **or** `mp1` is EXCLUDED (the B3 trap — extra mountpoints default `backup=0`). The
|
||
golden does NOT pre-create `felhom-data`; the controller does once it's a real mountpoint.
|
||
- **`internal/reconcile/bringup.go`** — `const DefaultSysDataMount = "mp1"`; `BringUpSpec` gains
|
||
`SysDataGrowGB int` + `SysDataMount string`; a new **"4c"** grow block (online, grow-only `ResizeLXC`,
|
||
its own task) mirrors the "4b" Docker-data grow. `0 = skip` (separateness comes from the golden, not
|
||
the grow — the warning clears regardless of size).
|
||
- **`cmd/felhom-agent/main.go`** — `-sysdata-grow` / `-sysdata-mount` flags (mirror
|
||
`-datavol-grow`/`-datavol-mount`); `bringUpSizing` carries them into all three bring-up/provision call
|
||
sites; `--selftest=provision` help text updated.
|
||
- **Static volume, NOT an enrolled drive.** `/mnt/sys_drive` is part of the baked golden layout; it
|
||
never enrolls/ejects/decommissions and is deliberately kept off the drive-intent machinery.
|
||
`freeMountSlot` auto-skips the baked `mp0`/`mp1` so enrolled drives never collide.
|
||
- Tests: `TestRunBringUp_StorageSplit_SysDataGrow` (asserts `ResizeLXC(vmid,"mp1","+42G")`) +
|
||
`…_SysDataGrowZeroNoResize` (0 → no mp1 resize). RUNBOOK-provisioning-storage.md extended to the
|
||
three-volume layout (default ~512 GB SSD: 32 rootfs + 200 docker-data + 50 user-data).
|
||
|
||
## v0.39.0 — DR recipe completion: live PBS coord + drop the two unfillable drive fields (2026-06-16)
|
||
|
||
**DR-recipe agent-half completion.** A live eyeball of the demo recipe (v0.38.0) found three host-half
|
||
problems; all three are resolved here. No behavior change outside the recipe path.
|
||
|
||
- **PBS coord now resolved LIVE each collect.** New `internal/pbs/live_reporter.go` —
|
||
`LiveSnapshotReporter` implements `hub.PBSReporter` by doing the cheap `Client.Snapshots()` list
|
||
itself, with **last-known-good fallback**, instead of reading only the verify-loop's `SnapshotStore`.
|
||
Previously the recipe's `pbs` block was omitted whenever the store was empty — which a one-shot
|
||
collect (`--selftest=hub`) and the first ~6 h window of every daemon after a restart always saw (the
|
||
verify loop populates the store on its own 6 h cadence). The restore SOURCE must not depend on a
|
||
maintenance cadence. Per-datastore: a live error/timeout → that datastore's last-known-good; a
|
||
successful (even empty) response is authoritative and updates the shared store. Targets-resolution
|
||
failure → the full LKG aggregate. Bounded by `DefaultLiveSnapshotTimeout` (8 s) so a hung PBS never
|
||
stalls the heartbeat. List only — it never triggers a `Verify`. The verify loop keeps Recording into
|
||
the SAME store (shared last-known-good); both use one hoisted `pbsTargets` closure.
|
||
- `SnapshotStore.Get(datastore)` added (per-datastore LKG copy) — the only `SnapshotStore` change.
|
||
- Wired into the collector in BOTH `runDaemon` and `runSelftestHub` (the selftest built its own
|
||
collector with a `nil` reporter — that is why the live `--selftest=hub` showed `pbs_snapshots:[]`).
|
||
- Intended side effect: `report.pbs_snapshots` is now live too (fresher hub PBS view).
|
||
- **`drives[].role` DROPPED from the v1 host-half shape.** A drive's purpose is a hub/operator-owned
|
||
manifest concept, not cleanly derivable host-side (both demo externals are `content=backup`, yet one
|
||
is the primary data drive and the other holds no apps). Deferred until the hub/operator stamps it.
|
||
- **`drives[].restic_repo_coord` DROPPED from the v1 host-half shape.** It named a backup tier that does
|
||
not exist — cross-drive backup is rsync to the SAME internal SSD; there is no offsite/second-failure-
|
||
domain bulk copy. RESERVED for a future tier (see the BACKLOG note in REPORT). v1 drive shape is now
|
||
`{durable_id, mount_path, intent, fs_type?, total_bytes}` — identifiers/intent/size only.
|
||
- The hub reads drives as `json.RawMessage`, so dropping fields needs NO hub struct change — only
|
||
golden + test sync. Cross-repo golden (`host-report.golden.json` here + the hub's copy) re-pinned and
|
||
verified **byte-identical** (sha256 `57f2a5e7…18b2f2b5` — manual checksum-diff discipline): the hub copy
|
||
previously lacked the `dr_recipe` section entirely; it is now a verbatim copy of the agent golden.
|
||
- Tests: new `internal/pbs/live_reporter_test.go` (T1 coord-present-without-prior-verify [load-bearing] +
|
||
inline bare-store companion, T2 error→LKG fallback, T3 success-warms-store, T4 targets-error→aggregate,
|
||
T5 bounded-by-timeout, T6 empty-success-authoritative); `TestDRRecipeHostHalf_V1DriveShape` (drive
|
||
object carries neither `role` nor `restic_repo_coord`); `TestBuildDRRecipeHostHalf` /
|
||
`TestHostReport_ContractMatchesGolden` updated to the v1 drive shape. Each companion was demonstrated
|
||
to FAIL on the pre-fix/mutated code, then reverted (see REPORT).
|
||
|
||
## v0.38.0 — DR recipe: emit the secret-free storage/guest/PBS half in the host-report (2026-06-16)
|
||
|
||
**DR recipe slice (agent half).** Additive `dr_recipe` section on the host-report — the agent half of the
|
||
secret-free reconstruction recipe (`SPIKE-dr-recipe-2026-06-16.md`) that complements escrow (keys) +
|
||
PBS/restic (bytes): the non-secret SCAFFOLDING an operator must rebuild before the PBS bytes can land.
|
||
The hub assembles it with the controller's app half into one customer recipe.
|
||
|
||
- `internal/hub/dr_recipe.go` — `DRRecipeHostHalf{recipe_version, guests[], pbs, drives[], pve_storage[]}`
|
||
built by the pure `BuildDRRecipeHostHalf(guests, targets, pbs)` from facts the report ALREADY collects
|
||
(no new privileged reads): `guests[]` = each guest's sizing (`GuestSpec`, skip status-unknown);
|
||
`drives[]` = the user-data external drives (usb/local-dir with a `uuid:` durable-id + mount path) with
|
||
`{durable_id, role, mount_path, intent, total_bytes}`; `pve_storage[]` = every storage target
|
||
`{name, type, content}` (the `storage.cfg` scaffolding); `pbs` = the latest snapshot's coordinates
|
||
`{repo_id (the pbs storage id), namespace, latest_snapshot_id}`. Wired into `Collect()` after the facts
|
||
are gathered; `HostReport.DRRecipe` (always set, never null).
|
||
- **BOUNDARY (the Phase-1 lesson):** every field is an identifier / intent / size / coordinate — NEVER a
|
||
key, password, token, hash, or `ENC:` value. The PBS encryption key stays in escrow; the access token in
|
||
identity-escrow; the restic password in escrow — the recipe names only the `repo_id`/`namespace`/
|
||
`durable_id`/`restic_repo_coord` the restore TARGETS. `recipe_version=1`; read is ignore-unknown
|
||
(forward-compat). The wire shape is pinned in the cross-repo golden (`host-report.golden.json` here +
|
||
the hub's copy — keep them byte-identical; manual checksum-diff on any change).
|
||
- Tests: `TestBuildDRRecipeHostHalf` (drives = only user-data; pve_storage = all; pbs = latest; guests
|
||
skip nil-spec), `..._NoPBS` (omitted, non-nil slices), `TestDRRecipeHostHalf_NoSecrets` (the lighter
|
||
boundary mirror — serialized half carries NO credential-shaped key; the load-bearing version is on the
|
||
controller emitter), and the `dr_recipe` key-set added to `TestHostReport_ContractMatchesGolden`.
|
||
|
||
## v0.37.0 — host-reboot remount re-resolves enrolled drives by filesystem UUID (2026-06-16)
|
||
|
||
**TASK A — close out the reboot story (agent half).** On a host reboot the kernel can re-enumerate block
|
||
devices and move a drive's node (felhom-usb `/dev/sdb`→`/dev/sdc`), and a `.mount` unit left `disabled`
|
||
by a prior detach never auto-mounts at boot — so an enrolled drive could stay unmounted (or, with any
|
||
node-trusting remount, mount the WRONG device). Root cause pinned LIVE: felhom-usb's systemd mount unit
|
||
was `disabled` (no `multi-user.target.wants` symlink) while felhom-flash's was `enabled`; `What=` was
|
||
already correct (by-UUID), but nothing re-asserted the unit at startup.
|
||
|
||
- `storage.ResolveStorageDevice(durableID)` — resolves the enrolled `uuid:<fs-uuid>` storage scheme to its
|
||
CURRENT backing `/dev` node by re-scanning `/dev/disk/by-uuid` (never a cached node); errors if the UUID
|
||
is genuinely absent so a caller skips a gone drive instead of fail-mounting a stale node.
|
||
- `storage.parseFelhomMountUnit` — pure inverse of `renderMountUnit` (Name/UUID/Where/Type/Options) keyed
|
||
on a `Managed by felhom-agent` marker; ignores any foreign `.mount` unit.
|
||
- `(*SudoHostOps).ReassertEnrolledMounts(ctx)` — at startup (BEFORE binding into the guest) and on the
|
||
periodic 20s tick: for each enrolled `.mount` unit, re-resolve by UUID and re-run `EnsureMount`
|
||
(idempotent `systemctl enable --now`) — re-enables a disabled unit AND mounts the CURRENT device by
|
||
UUID, so a `/dev/sdX` reshuffle is a no-op. Skips ONLY the durable steady state (mounted AND enabled),
|
||
via the pure `shouldReassertMount`; a **mounted-but-DISABLED** unit (the exact live felhom-usb bug — it
|
||
serves now but a reboot would not auto-mount it) is still re-asserted to re-create the wants-symlink.
|
||
Enabled-state is read with a privilege-free `os.Lstat` of the `multi-user.target.wants` symlink
|
||
(`unitEnabled`) — no `systemctl is-enabled` subprocess, no new sudoers entry. An absent UUID is skipped
|
||
(re-asserts on a later tick).
|
||
- Wired in `main.go` ahead of `ReassertGuestBinds` so mounts are live before the guest binds re-assert.
|
||
- Tests (Linux, seam the device-resolution): `TestResolveStorageDevice_ToleratesDeviceLetterMove`
|
||
(UUID symlink moved sdb→sdc → resolves sdc; companion asserts the cached enroll-time node differs from
|
||
the freshly-resolved one — a node-based remount would target the wrong device), `..._AbsentAndScheme`
|
||
(absent UUID errors; only the `uuid:` scheme resolvable), `TestParseFelhomMountUnit` (render→parse
|
||
round-trip + rejects a foreign unit), `TestShouldReassertMount` (the four mounted/enabled combos — pins
|
||
the mounted-but-disabled re-assert), `TestUnitEnabled` (wants-symlink detection).
|
||
|
||
**TASK A2 — verdict: enrolling a NEW drive does NOT need an LXC restart.** The enroll path lands on the
|
||
live intermediary-mount `AttachDrive` (`/disks/guest-attach` → `handleDiskGuestAttach` → `AttachDrive`,
|
||
"no pct, no reboot") under the single shared parent — unbounded named live slots — NOT the legacy
|
||
`RebootGuest` branch. The operator's pre-created-slot-pool idea is therefore unnecessary.
|
||
|
||
## v0.36.7 — isolate the shared parent only on CREATE (no peer-group churn) (2026-06-15)
|
||
|
||
Follow-up to v0.36.6: make-private+make-shared must run ONLY when the self-bind is first created, not on
|
||
every reconcile — re-doing it churns the peer-group id and ORPHANS the guest`s already-established slave
|
||
(propagation silently dies, guest sees empty). Guarded on the mountpoint check; on a fresh boot it runs
|
||
once before pve-guests so the guest slaves the right group.
|
||
|
||
## v0.36.6 — shared parent gets its OWN peer group (make-private first) — ROOT CAUSE of double-bind (2026-06-15)
|
||
|
||
The shared-parent self-bind INHERITED the root mount`s shared peer group (`/mnt/felhom-drives` was
|
||
`shared:1` same as `/`), so every drive bind under it propagated back via the root peer and DOUBLED
|
||
(2 stacked binds per drive — the real cause behind v0.36.3-.5). EnsureSharedParent + the boot script now
|
||
`make-private` (detach from the root group) BEFORE `make-shared` (own group whose only slave is the
|
||
guest), so a drive bind propagates to the guest exactly once.
|
||
|
||
## v0.36.5 — AttachDrive normalizes to exactly one bind (2026-06-15)
|
||
|
||
AttachDrive now COUNTS the binds at a stable path (countHostMounts) and normalizes to exactly one: it is
|
||
a no-op only when there is exactly ONE bind the guest sees; otherwise it strips ALL existing binds
|
||
(bounded loop) and lays down one fresh bind. This converges a stacked double-bind to one — the old
|
||
umount-one+mount-one force-rebind never did. Caught when a double-bind survived a guest reboot.
|
||
|
||
## v0.36.4 — serialize AttachDrive/DetachDrive (no double-bind race) (2026-06-15)
|
||
|
||
A mutex on GuestBinder serializes AttachDrive/DetachDrive so a controller-triggered reconnect and the
|
||
agent`s periodic reconcile can no longer both pass the isHostMountpoint check and double-bind the same
|
||
stable path (a TOCTOU race observed live as 2 stacked binds during rapid eject/reconnect).
|
||
|
||
## v0.36.3 — DetachDrive loop-umounts stacked binds (2026-06-15)
|
||
|
||
DetachDrive now umounts ALL stacked binds at a stable path (bounded loop), not just one layer — so an
|
||
eject/detach fully detaches even if more than one bind accumulated (operator bind on top, or a rare
|
||
attach race), keeping the fail-close intact. Caught in the E13 rapid eject/reconnect sweep.
|
||
|
||
## v0.36.2 — eject also keeps the raw mounted (reconnectable) (2026-06-15)
|
||
|
||
Extends v0.36.1 to EJECT: eject now DetachDrive`s the bind under the parent but LEAVES the raw
|
||
/mnt/<name> mounted (consistent with decommission), so the H1 disconnect→reconnect roundtrip re-binds
|
||
cleanly on a non-removable drive. Physical removal is the separate "remove from system" action. Test:
|
||
eject calls DetachDrive + does NOT unmount the raw.
|
||
|
||
## v0.36.1 — decommission keeps the raw mounted (re-enrollable) (2026-06-15)
|
||
|
||
Fix caught in the E10 acceptance test: the self-serve decommission unmounted the RAW /mnt/<name> host
|
||
mount, which orphaned a non-removable drive (no re-plug) so a one-click re-enroll bound an empty dir. On
|
||
the intermediary model decommission is now a LOGICAL retire — it DetachDrive`s the bind under the parent
|
||
(drive invisible to the guest) but LEAVES the raw mounted, so re-enroll re-binds cleanly. Physical
|
||
removal stays the separate "remove from system" action. Test updated.
|
||
|
||
## v0.36.0 — guest boot-id on /disks (deterministic guest-reboot recreate) (2026-06-15)
|
||
|
||
The agent now emits `guest_boot_id` on GET /disks: `<host-btime>-<guest-init-starttime>` — changes on
|
||
every guest boot (host reboot OR guest reboot) but is STABLE across a controller-only restart. The
|
||
controller persists the last-seen value and DETERMINISTICALLY recreates drive-backed apps when it
|
||
changes (replacing the fragile timed state-sample that could miss an app stopped at the sample instant).
|
||
`GuestBootID` reads `/proc/stat` btime + field 22 of `/proc/<init-pid>/stat` (parsed after the last
|
||
`)` so a comm with spaces/parens does not break it).
|
||
|
||
## v0.35.1 — shared-parent unit: run before pve-guests on host boot (2026-06-15)
|
||
|
||
Fix for the host-reboot ordering (the shared-parent oneshot never ran before pve-guests on the live
|
||
host, so the guest bound a not-yet-shared parent → private bind → propagation broken). The unit now uses
|
||
`WantedBy=pve-guests.service` (pve-guests PULLS IT IN + Before= orders it first) instead of the
|
||
unreliable `WantedBy=multi-user.target`, and drops `DefaultDependencies=no`. `EnsureSharedParent`
|
||
reinstalls the unit when its content differs (so the fix deploys on the next agent start/reconcile).
|
||
|
||
## v0.35.0 — intermediary mount: guest-reboot re-propagation (load-bearing) (2026-06-15)
|
||
|
||
Fix for the guest-reboot gap (caught in the live demo migration). A guest's parent bind is
|
||
NON-RECURSIVE, so on a guest reboot it does NOT carry the pre-existing drive submount, and mount
|
||
propagation only delivers events created AFTER the bind exists — so an enrolled drive is bound on the
|
||
HOST but INVISIBLE in the fresh guest namespace until re-bound. Without this, every guest reboot left
|
||
the apps on empty dirs.
|
||
|
||
- `AttachDrive` now takes `vmid` and checks GUEST visibility (`GuestSeesMount`, reading
|
||
`/proc/<guest-init-pid>/mountinfo`): if the host has the bind but the guest doesn't see it
|
||
(post-reboot), it FORCE re-binds (umount + mount) to re-fire propagation into the current guest ns.
|
||
- A periodic reconcile (20s ticker in main) re-runs `ReassertGuestBinds`, so a guest reboot self-heals
|
||
without an agent restart. `EnsureSharedParent` skips the unit re-install when already present (cheap
|
||
on repeat).
|
||
- `/disks` `BoundUnderParent` now reflects GUEST visibility (not the host mount) — the accurate signal
|
||
the controller's drive-absent gate keys on to stop/restart apps across a guest reboot.
|
||
|
||
## v0.34.0 — intermediary mount model: shared-parent + host-side attach/detach + reconcile (2026-06-15)
|
||
|
||
The drive hot-swap re-architecture (SPIKE-intermediary-mount). Replaces the per-drive `pct set -mpN`
|
||
bind (which needed a guest reboot to activate and bricked the guest when a drive was absent at boot)
|
||
with a SINGLE permanent parent bind `/mnt/felhom-drives` plus host-side swaps underneath it.
|
||
|
||
- `internal/localapi/intermediary.go` — `GuestBinder.EnsureSharedParent` (mkdir + self-bind +
|
||
`--make-shared` + installs/enables a `felhom-shared-parent.service` ordered **Before=pve-guests** so
|
||
the guest's parent bind inherits the shared peer group as `slave`); `AttachDrive` (`mount --bind
|
||
/mnt/<name>/felhom-data /mnt/felhom-drives/<name>` — propagates into the RUNNING guest live, no pct,
|
||
no reboot; confined to felhom-data; the stable dir stays host-root-owned = fail-closed); `DetachDrive`
|
||
(`umount`, leaving the bare fail-closed dir); `StablePathForRaw`/`DriveNameFromRaw`; `isHostMountpoint`.
|
||
- `ReassertGuestBinds` is now a pure HOST-SIDE reconcile: for each enrolled+present drive ensure its
|
||
felhom-data is bound under the parent (no guest-config read, no slot, no reboot) — fixes F9 and
|
||
drive-reconnect for free. Runs at startup (ensures the shared parent first).
|
||
- `handleDiskGuestAttach` uses `AttachDrive` (returns the stable `guest_path`); eject + decommission
|
||
call `DetachDrive`. Legacy `AttachBind`/`DetachBind` retained for the transition (decommission still
|
||
`--delete`s any lingering legacy mp).
|
||
- `/disks` reporting adds `GuestPath` (the stable `/mnt/felhom-drives/<name>` the controller repoints
|
||
HDD_PATH to) and `BoundUnderParent` (live-in-guest signal for the controller's drive-absent gate).
|
||
- Provision adds the one permanent parent bind (`-mp8 /mnt/felhom-drives,mp=/mnt/felhom-drives`).
|
||
|
||
Tests (non-hollow + companions): `TestGuestAttach_BindsUnderParent` (uses AttachDrive not legacy pct),
|
||
`TestReassertGuestBinds_RestoresMissingBind` (host-side reconcile, legacy AttachBind never called),
|
||
`TestStablePathForRaw_DriveName`, `TestDisks_GuestPathAndBoundUnderParent`. Sudoers: new
|
||
`FELHOM_INTERMEDIARY` alias (mount/umount under /mnt/felhom-drives, the unit install, the parent bind).
|
||
|
||
## v0.33.0 — C1 net: pre-start self-heal hook + decommission mp-delete (2026-06-15)
|
||
|
||
The transitional defense for the C1 brick (B3 critical bug) ahead of the intermediary-mount
|
||
re-architecture (which makes C1 structural). Two independent nets:
|
||
|
||
- **Pre-start self-heal hook** (`internal/guesthook`): a PVE `pre-start` hookscript runs
|
||
`felhom-agent guest-hook <vmid> <phase>` which, for every BIND mountpoint whose source path is
|
||
missing, creates an empty **host-root-owned** placeholder dir so the bind succeeds and the guest
|
||
always boots — fail-closed (host uid 0 is unmapped in the unprivileged-LXC userns, so the guest
|
||
can't write to the placeholder; a returning drive shadows it). It CREATES rather than DELETEs
|
||
because `pct set --delete` in pre-start would take the config lock the start task already holds
|
||
(dead-times-out → still bricks); the heal logic is in unit-tested Go, the wrapper just delegates.
|
||
Installed + registered per-guest by the provision back-half (`InstallSnippet`/`Register`).
|
||
- **Decommission mp-delete** (`GuestBinder.DetachBind` + `handleDiskDecommission`): decommission now
|
||
runs `pct set <vmid> --delete mpN` on the slot binding the drive (lock-safe on the running guest),
|
||
so its now-missing source can't brick the next reboot. The old handler unmounted but left the dead
|
||
`mpN` in config — the exact B3 C1 bug. Eject keeps its mp (temporary; the hook covers a
|
||
reboot-while-ejected).
|
||
|
||
Tests (non-hollow, each with a companion that fails the pre-fix/trivial impl):
|
||
`internal/guesthook/heal_test.go` (selector ignores storage volumes + present binds, heals only the
|
||
absent one; "return nothing"/"return all" both fail) and `TestDecommission_DeletesGuestMount`
|
||
(asserts the correct slot is `--delete`d; pre-fix never calls DetachBind → fails).
|
||
|
||
Sudoers: new `FELHOM_GUESTHOOK` alias (snippet install, `pct set --hookscript`, `pct set --delete mpN`).
|
||
|
||
## v0.32.0 — self-serve decommission + intent-aware re-assert (B2a) (2026-06-14)
|
||
|
||
Customer-self-serve storage decommission (no operator signature; non-destructive — never formats),
|
||
plus the load-bearing fix that keeps a decommissioned drive from auto-rebinding into the guest.
|
||
|
||
- **`POST /disks/decommission`** (`internal/localapi/disks.go` `handleDiskDecommission`, route in
|
||
`server.go`) — mirrors `handleDiskEject` exactly: `withGuest` self-scoping, `scopedFromBody`, and the
|
||
same **user-data role gate** (`roleForMountPath` must be `RoleUserData`, else 403; fail-safe-to-
|
||
protected on ambiguity) so a compromised controller can't decommission system/backup storage. It
|
||
records a PERMANENT `IntentDecommissioned`, prunes the `GuestBindStore` entry (hygiene), and unmounts
|
||
(so the drive is physically removable). It **NEVER** calls any format/mkfs path — the data stays on
|
||
the drive. The operator-signed `DecommissionExecutor` + `reconcile.Classify` classification are
|
||
untouched (the absent-drive/DR route).
|
||
- **`ReassertGuestBinds` is now intent-aware** (THE correctness fix): the startup re-assert skips any
|
||
durable-id whose intent is not `enrolled`, so a decommissioned- (or ejected-) but-still-present drive
|
||
is never auto-rebound into the guest on agent restart. A nil intent store falls back to legacy
|
||
bind-all (matching the watchdog's nil-intent rule). Covers both the self-serve and the operator-
|
||
signed decommission paths (both land on `IntentDecommissioned`).
|
||
- **`GuestBindStore.Remove(vmid, durableID)`** (`internal/localapi/guestbindstore.go`) — idempotent
|
||
(absent = no-op), atomic tmp+rename like `Record`; drops the vmid key when its set empties. Re-enroll
|
||
re-`Record`s via the existing `recordGuestBind` on guest-attach, so Remove doesn't break re-commission.
|
||
- `IntentRecorder` extended with `SetDecommissioned` + `Get` (both already on `*storage.IntentStore`).
|
||
- Non-hollow tests (`internal/localapi/decommission_test.go`): role-gate refuses system/backup (403,
|
||
no unmount); decommission sets intent + removes the bind + unmounts + never formats; intent-aware
|
||
re-assert does NOT rebind a decommissioned-but-present drive (companion: enrolled DOES rebind; the
|
||
intent-blind pre-fix code fails this); re-commission re-records; `Remove` idempotency + persistence.
|
||
|
||
## v0.31.0 — live-drive F9 + F20-BUG2 + F20-BUG3 (disk bind/wipe) (2026-06-14)
|
||
|
||
The last live-drive findings, all disk/`localapi`-side, implemented + deployed on `felhom-pve` and
|
||
validated live on guest 9201 (approach: attach-to-existing, no re-provision — see the audit fixspec).
|
||
|
||
- **F9 — guest data-drive bind survives a re-provision** (`4cd1d02`). The in-guest bind (`pct set -mpN`)
|
||
is config state a destroy+re-provision drops, and nothing restored it → a re-provisioned guest came up
|
||
with its enrolled HDD unattached. New `GuestBindStore` (durable-id-keyed, per guest, recorded at
|
||
guest-attach) + `ReassertGuestBinds` on agent startup re-adds any bind a guest is missing — only when
|
||
the durable-id still resolves to a present drive (a swapped/absent disk is never auto-bound), idempotent.
|
||
Plus `DiskInfo.GuestAttached` — the missing "bound into THIS guest" signal (vs mere host presence;
|
||
resolves the F2 `hdd_configured` disagreement). **Live-proven:** dropped the bind, restarted the agent
|
||
(real trigger) → re-attached with no manual call; reboot activated it; an HDD app then deployed onto
|
||
the drive with data on `/dev/sdb1`.
|
||
- **F20-BUG2 — one wipe durable-id scheme** (`a2a76e7`). `/disks` advertised only `durable_id` (`uuid:`,
|
||
used for assign), but the wipe gate resolves `byid:`/`byuuid:` → confirming a wipe with the advertised
|
||
id was a `binding_mismatch`. New `DiskInfo.WipeDurableID` via a shared `s.deviceDurableID` seam used by
|
||
BOTH the list and the gate, so the id the customer copies is the id the gate accepts. **Live-proven:** a
|
||
confirmed wipe using `/api/disks`'s `wipe_durable_id` is accepted (no mismatch).
|
||
- **F20-BUG3 — format runs detached; survives a request deadline AND an agent restart** (`4777f8a`). mkfs
|
||
ran under the HTTP request context, so a client deadline SIGKILLed it mid-write → corrupt disk. Now mkfs
|
||
runs off `s.baseCtx` via a persisted `formatJob` record; the handler still returns the synchronous
|
||
result (backward-compatible) but a dropped request no longer kills it. New `GET /disks/format/status`;
|
||
`RecoverFormatJob` on startup re-runs an interrupted durable-id-bound format (re-resolved; anti-retarget
|
||
— a blank/path-bound or unresolvable job is not auto-re-run). **Live-proven on the 916 GB felhom-usb:** a
|
||
2 s client timeout left a ~30 s mkfs running to a clean ext4 (the live-drive corruption is gone); an
|
||
agent restart mid-format was recovered + completed to a clean fs.
|
||
|
||
|
||
|
||
**Security fix (from the 2026-06-13 deep-sweep audit).** The inline customer-confirmed wipe in
|
||
`internal/localapi/disks.go` `handleDiskFormat` inspected and gate-bound the device by its durable id
|
||
but then ran `mkfs` on the caller-supplied mutable `/dev` path (`req.Device`). A USB re-enumeration
|
||
reassigning that `/dev` node to a different physical disk between inspection and `mkfs` (a
|
||
classify→mkfs TOCTOU) could wipe the wrong drive.
|
||
|
||
- New `internal/localapi/wipe_reresolve.go`: `antiRetargetResolve` (injected-deps, unit-tested) mirrors
|
||
`signedjobs.WipeExecutor.Execute` — resolve the confirmed durable id → current device, re-derive the
|
||
device's durable id and require an exact match, re-inspect (still data-bearing), and return the
|
||
re-resolved device. `(*Server).reresolveDurableForWipe` wires the real storage funcs.
|
||
- `handleDiskFormat` now formats the **re-resolved** device, never `req.Device`; any refusal →
|
||
`409 Conflict`, no `mkfs`. Injectable `reresolveWipe` seam on `Server` (defaults to the real path).
|
||
- Tests: `wipe_reresolve_test.go` covers happy-path, empty/gone/blank, re-inspect-error, and the core
|
||
`retarget-mismatch-refused` case. Round-trip safe for legitimate wipes (`DeviceDurableID` ↔
|
||
`ResolveDurableDevice` schemes match). Agent-only deploy; no golden rebake. See `AGENT-001-FIX-NOTES.md`.
|
||
|
||
## v0.29.1 — lanresolver: RESTART dnsmasq on change (not reload) — fixes stale split-horizon IP (2026-06-13)
|
||
|
||
**Bug:** after a guest's DHCP IP moved (e.g. the v0.29.0 9201 re-provision: .151 → .141), the LAN
|
||
split-horizon resolver kept answering the OLD IP, so LAN clients (via Pi-hole's conditional forward to
|
||
the host dnsmasq) resolved `*.demo-felhom.eu` to the dead IP. Root cause: `lanresolver.Manager` updated
|
||
the per-customer drop-in (`address=/<domain>/<ip>`) correctly but then ran `systemctl reload dnsmasq`
|
||
(SIGHUP) — and **dnsmasq's SIGHUP does NOT re-read its config files** (`/etc/dnsmasq.d/*.conf`); it only
|
||
clears the cache + re-reads `/etc/hosts`/addn-hosts. So the changed `address=` directive never took
|
||
effect until a restart. **Fix:** `reload()` → `restartDnsmasq()` (`systemctl restart dnsmasq`) for every
|
||
config-drop-in change (ReconcileGuest IP change, EnsureDnsmasq base change, Remove/decommission). Restart
|
||
is sub-second and the records carry local-ttl 0, so downstream forwarders don't cache a stale answer.
|
||
(Live: after the fix + a one-time host dnsmasq restart + a Pi-hole cache flush, `*.demo-felhom.eu`
|
||
resolves to the live guest IP again; future IP moves now self-heal on the loop's next tick.)
|
||
|
||
## v0.29.0 — OS / Docker-data storage split: golden + provision (2026-06-13)
|
||
|
||
Phase 1 of the storage-split slice (Phase 2 = felhom-controller v0.58.0 prevention layer). The
|
||
controller guest's OS rootfs and Docker data are carved onto separate `local-lvm` volumes for
|
||
RESILIENCE — an isolated OS rootfs stays bootable + agent-recoverable if the Docker volume fills.
|
||
|
||
- **`configs/build-golden.sh` — split baked in:** `--rootfs ${ROOTFS_STORAGE}:${OS_SIZE_GB}` (default
|
||
**32**, was hardcoded 8) **plus** `--mp0 ${ROOTFS_STORAGE}:${GOLDEN_DOCKER_GB},mp=/var/lib/docker,backup=1`
|
||
(default 16). The baked controller + infra images land on the data volume and travel inside the
|
||
golden archive (no empty-volume shadowing, no deploy-time pull). `backup=1` is MANDATORY — extra LXC
|
||
mountpoints default to `backup=0` = EXCLUDED from vzdump (spike B3), which would drop the images from
|
||
the archive entirely. The script now also bakes Docker **log rotation** into `daemon.json`
|
||
(`max-size 10m`, `max-file 3` — prevention layer 2D), asserts `/var/lib/docker` is a separate mount,
|
||
and **aborts if vzdump excludes mp0**.
|
||
- **`internal/reconcile/bringup.go` — sized provision:** `GuestMount` gains `Backup` (emits `,backup=1`
|
||
— closes the spike-B3/B5 silent-DB-loss trap at the mount builder). `BringUpSpec` gains
|
||
`DataVolGrowGB` + `DataVolMount` (default `mp0`): provision GROWS the golden-carried Docker-data
|
||
volume online to the per-customer target (grow-only, spike B4) rather than attaching a fresh empty
|
||
volume that would shadow the baked images. Plus `RootfsGrowGB` for the OS rootfs.
|
||
- **CLI seam:** `--selftest=bring-up|provision` gain `-rootfs-grow` / `-datavol-grow` / `-datavol-mount`
|
||
flags. Per-customer sizing source = flags now, the slice-10 hub storage manifest later.
|
||
- **`RUNBOOK-provisioning-storage.md`** (new): the split provisioning procedure + fresh-PVE-install
|
||
thin-pool carving knobs (`hdsize`/`maxroot`/`maxvz`, spike B4) + the per-customer sizing seam.
|
||
- Tests: `buildBringUpConfig` backup=1 emission; bring-up issues rootfs + data-volume resizes.
|
||
|
||
## (no version) — storage OS/data-split spike findings (2026-06-13)
|
||
|
||
Investigation only — **no code changed**. Findings report: `REPORT-storage-split-spike.md` (gates the
|
||
provisioning spec for splitting the controller guest's OS rootfs from its Docker/data onto separate
|
||
`local-lvm` volumes). Proven on a throwaway unprivileged LXC (9300, since destroyed): Docker `data-root`
|
||
on a second `local-lvm` mountpoint works (overlayfs/ext4, no idmap issue, reboot-survives); the
|
||
move-then-verify migration is safe (copy-not-move). **Key finding:** additional LXC mountpoints are
|
||
**excluded from vzdump by default** — they need `backup=1` set **and a CT restart** — so the docker-data
|
||
mount must be attached with `,backup=1` or named-volume DBs silently fall out of PBS. The exact seam is
|
||
`internal/reconcile/bringup.go:313` (`buildConfigParams`), which today builds `mpN` without a `backup=`
|
||
flag; `GuestMount` should carry the flag. Per-customer sizes belong in the slice-10 hub storage manifest
|
||
(marked at `bringup.go:49-50`); the golden rootfs is hardcoded `8` at `configs/build-golden.sh:40`.
|
||
|
||
## v0.28.0 — backup re-target → felhom-pbs (offsite DR) + operator-signed decommission (2026-06-12)
|
||
|
||
**Whole-guest backup now defaults to the offsite PBS tier (real DR).** `BackupConfig.BackupTarget()`
|
||
returns the configured `backup.local_backup_target` or, when empty, the new default `felhom-pbs` — a
|
||
PBS datastore on SEPARATE HARDWARE (the DooPlex box), so a host disk/hardware failure no longer takes
|
||
the backups with it. The target stays fully configurable (set `local_backup_target` to `local`/other
|
||
to override); no call site hardcodes it. All `NewBackupRunner` sites (restore-test scheduler, local-API,
|
||
`--selftest=backup`/`restore-test`) route through `BackupTarget()`.
|
||
|
||
Proven live on demo-felhom before the re-point (PHASE 0 gate):
|
||
- snapshot-mode `vzdump → felhom-pbs` still fires the `create storage snapshot 'vzdump'` marker, so the
|
||
8B.2 early-resume/quiesce signal survives a PBS target (the marker is mode-driven, not target-driven);
|
||
- the restore-test enumerates PBS backups through the SAME generic `StorageContent`
|
||
(`/nodes/<node>/storage/felhom-pbs/content` returns `content:"backup"` + ctime/vmid/volid), so
|
||
`PickRestoreCandidate`/`latestArchive` need NO PBS-client change;
|
||
- `pct restore` from a PBS volid round-trips cleanly (storage.cfg encryption key applied transparently);
|
||
- PBS gotchas (`ignore-verified`, node-from-UPID, privsep) touch only the verify-API path, not vzdump/restore.
|
||
|
||
**Operator-signed `decommission` now reachable (slice 10 P3 completion).** The previously-unreachable
|
||
`IntentDecommissioned` state (no production caller) is now reached ONLY via a gate-VERIFIED operator
|
||
signature — never customer-confirmable, distinct from a safe eject. New `internal/signedjobs`
|
||
`DecommissionExecutor` (op `decommission`, classified destructive in `reconcile.Classify`) calls
|
||
`IntentStore.SetDecommissioned`, keyed by the drive's STORAGE durable-id (the watchdog's key, e.g.
|
||
`uuid:<fs-uuid>` — NOT the device-level `byid:/byuuid:` scheme `storage_wipe` uses), so the recorded
|
||
intent actually gates future remounts. New `ExecutorChain` lets the signed-jobs runner serve both
|
||
`storage_wipe` and `decommission`; the runner wiring moved below the intent-store open in `main.go`.
|
||
`felhom-opsign` builds decommission params from `-durable-id`. No controller/customer UI — the operator
|
||
path is hub jobs-queue → signed-jobs runner.
|
||
|
||
**Restore-test now boot-verifies slice-10 enrolled guests (bind-mount mountpoints).** A guest whose
|
||
data drive is a host BIND mount (slice-10 P2 `mp0`) could not be vzrestore'd by the privsep token
|
||
("restoring 'mpN' to bind mount is only possible for root") — so the restore-test failed for every
|
||
enrolled guest, regardless of backup tier (surfaced during the felhom-pbs live validation). The
|
||
restore-test now reads the SOURCE guest config (vmid parsed from the archive volid — PBS `ct/<vmid>/`
|
||
and vzdump `vzdump-lxc-<vmid>-` forms) and passes `RestoreLXCOptions.MountOverrides` that neutralize
|
||
each bind-mount `mpN` to a throwaway 1G volume on the restore storage (needs no root; the boot-verify
|
||
doesn't need the drive's data, and the host paths would otherwise collide). Storage-backed mountpoints
|
||
are restored normally; best-effort (an unreadable source config restores as-is). `proxmox.RestoreLXC`
|
||
gained `MountOverrides`. Verified live: restore-test from felhom-pbs of bind-mounted guest 9201 →
|
||
boot+running PASS.
|
||
|
||
## v0.27.0 — slice 10 P3: self-heal watchdog reconcile + 4-state intent model (2026-06-12)
|
||
|
||
The storage watchdog goes from detect-only → detect-and-reconcile: the agent autonomously re-mounts an
|
||
enrolled external drive that dropped out-of-band (the colleague's Proxmox unmount), gated by a persisted
|
||
INTENT model so it never auto-adopts an unknown drive or fights an official eject.
|
||
|
||
- **`internal/storage/intent.go` — `IntentStore`** — durable, **durable-id-keyed** (UUID/WWN, never
|
||
sdX/path), atomic-write 4-state model: `new` (not recorded → never auto-mount), `enrolled` (desired
|
||
mounted → reconcile drift), `ejected` (intentional unmount → leave alone), `decommissioned`
|
||
(permanent). `OnAbsent` clears `ejected`→`enrolled` so a replug auto-mounts (the replug rule).
|
||
Records intent ONLY through the official enroll/eject paths — an out-of-band unmount records nothing
|
||
and is healed. Tests cover the states, persistence, the replug rule, and the reconcile gate.
|
||
- **`watchdog.go` — intent-gated reconcile + flapping guard (3C)** — the re-mount candidate (device
|
||
present, not mounted) now fires ONLY for an `enrolled` drive (via `IntentReader`); a present→absent
|
||
transition (device gone) calls `OnAbsent`. Exponential backoff (`debounce·2^fails`) + an alert after
|
||
4 failed cycles + a hard stop after 8 (no infinite loop). Failure = "still not present a full backoff
|
||
window after we dispatched" (a slow async re-mount isn't miscounted). Tests: colleague-unmount→
|
||
reconciled; ejected/new/decommissioned→left alone; ejected→absent→replug→auto-mount; flapping→caps.
|
||
- **`internal/localapi`** — `POST /disks/guest-attach` records `enrolled`; `POST /disks/eject` records
|
||
`ejected` (BEFORE unmount, while the durable-id still resolves) via the new `IntentRecorder`. `main.go`
|
||
opens one `IntentStore` (`<StateDir>/drive-intents.json`) shared by the watchdog + local API; open
|
||
failure degrades to ungated legacy remount (logged).
|
||
|
||
## v0.26.0 — slice 10 P2 activation: guest-reboot endpoint (user-triggered drive activation) (2026-06-12)
|
||
|
||
A drive enrolled into a RUNNING unprivileged guest can't be live-activated (proven: `pct set` won't
|
||
hot-apply; `/proc/<pid>/root` bind → mount-locking refusal; `nsenter -m` loses the host source). So the
|
||
bind activates at the next guest boot. This adds the user-triggered restart path.
|
||
|
||
- **`POST /guest/reboot` (`internal/localapi`)** — self-scoped (vmid from token). Runs `pct reboot
|
||
<vmid>` **detached** (it blocks ~30s until the guest is back) and returns **202** immediately, so the
|
||
calling controller gets a clean response before the reboot takes it down (the agent is host-side and
|
||
survives). `GuestBinder.RebootGuest` over the fenced runner. Tests: `TestGuestReboot_Accepted`
|
||
(202 + RebootGuest invoked for the token's vmid), `TestGuestReboot_CrossGuest403` (body vmid mismatch
|
||
refused, no reboot). Pairs with controller v0.49.0 (pending-activation detection + "Újraindítás most").
|
||
|
||
## v0.25.0 — slice 10 P2: bind enrolled user-data drives into the guest (passthrough) (2026-06-12)
|
||
|
||
External user-data drives are mounted on the HOST but were never passed INTO the guest (diagnosed
|
||
Branch A), so apps silently wrote to the rootfs and the controller couldn't see them. This adds the
|
||
guest passthrough. Spike-proven on 9201 first (see REPORT / the usb-passthrough-spike findings):
|
||
`pct set` **bind form** (host path, never `storage:size`), `chown` to the guest base (idmap not clean
|
||
for mixed-ownership data), `shared:49` propagation host↔guest automatic.
|
||
|
||
- **`POST /disks/guest-attach` (`internal/localapi`)** — self-scoped (vmid from token). Binds an
|
||
enrolled drive's **felhom-data namespace** into the guest at `/mnt/<name>` (**Model A**: the
|
||
felhom-data dir is the bind source mounted AT `/mnt/<name>`, so only Felhom's namespace crosses into
|
||
the guest — the customer's other data on the drive never does). Idempotent (returns the existing slot
|
||
if already bound); picks the lowest free `mpN`; validates `where` is `/mnt/<name>` (no traversal).
|
||
- **`GuestBinder` (`internal/localapi/guestbind.go`)** — the host-root steps over the fenced
|
||
`proxmox.Runner` (same pattern as the provision back-half's bind): `mkdir -p <drive>/felhom-data` →
|
||
`chown 100000:100000` the namespace ROOT (not -R; per-app subdirs are chowned at deploy) → `pct set
|
||
<vmid> -mpN <drive>/felhom-data,mp=/mnt/<name>` (RW bind). The namespace is created fresh + uniformly
|
||
owned, which sidesteps the drive's pre-existing mixed-ownership data entirely.
|
||
- **Tests** — `TestGuestAttach_*`: free-slot selection (mp0 when mp9 taken), idempotency (no re-bind +
|
||
`already:true`), bad-path rejection (traversal/non-/mnt/multi-component), not-configured 503.
|
||
|
||
Pairs with felhom-controller P2C (enroll triggers attach) + the golden's `/mnt:rslave` controller bind
|
||
(P2B). Self-heal reconcile (P3) and dual-role (P4) follow.
|
||
|
||
## v0.24.0 — role-gate the eject path (system/backup mounts are unmount-protected at the agent) (2026-06-12)
|
||
|
||
Closes the eject gap in the storage-authorization redesign: `POST /disks/eject` now **refuses to
|
||
unmount a system or backup storage**, enforced at the agent — not just hidden in the controller UI.
|
||
A direct API call (or a compromised controller) trying to `eject {where:"/var/lib/vz"}` or the PBS
|
||
mount is refused 403; only `user-data` mounts are ejectable.
|
||
|
||
- **`handleDiskEject` (`internal/localapi/disks.go`)** — before `Unmount`, resolves the AUTHORITATIVE
|
||
protection role of the storage mounted at `where` (the agent's own storage-view + host-topology
|
||
classification, never the caller's claim) via the new `roleForMountPath`. Refuses (403, no
|
||
`Unmount`) unless the role is `user-data`. **Fails SAFE**: an unresolvable mount (view error or no
|
||
storage target at that path) → treated as protected → refused (the same most-protected-on-ambiguity
|
||
default the wipe gate uses). Mirrors the wipe path's "protected — eject refused by role" logging.
|
||
- **`roleForMountPath` + `hostReader` seam** — `roleForMountPath` keys `RoleForStorage` on the mount
|
||
path (the eject input), mirroring `deviceRole`. `Options.HostReader` (optional; defaults to the
|
||
production `*storage.ProcHostReader`) injects the root-free topology reader so the role-gate is unit-
|
||
testable. `handleDisks`/`deviceRole` now share the same seam.
|
||
- **Tests** — `TestEject_RoleGated` asserts a `system` and a `backup` mount are refused with **no
|
||
`Unmount`**, a `user-data` mount ejects, and an unresolvable mount fails safe to refused (the same
|
||
non-hollowness the wipe tests use). `TestEject_UnmountAndDependents` updated to a user-data target.
|
||
|
||
## v0.23.0 — device-ROLE classification + tiered storage-wipe gate (system/backup operator-only, user-data customer-confirmable) (2026-06-11)
|
||
|
||
The storage-authorization redesign (agent half). The gate's destructive-wipe path is now **tiered by
|
||
the device's protection ROLE**, which the agent classifies from its OWN inspection — never the
|
||
caller's claim (the storage analog of classify.go's data-bearing verdict).
|
||
|
||
- **`internal/storage/role.go`** — `DeviceRole` (`system` | `backup` | `user-data`) + the
|
||
authoritative classifier. `RoleForStorage` (storage-view targets) and `RoleForRawDevice` (a raw
|
||
device, e.g. a fresh disk in the init flow) map a device to its tier via `SystemDisks` (the
|
||
whole-disks backing `/`, `/boot`, `/boot/efi`, root-free reads). Rules: `pbs` → backup; `lvmthin` /
|
||
builtin `local` / nfs / cifs / unknown → system; `usb` / `local-dir` on a **non-system external
|
||
device** → user-data. **Fail-safe**: any ambiguity (system disks unknown, or an unrecognizable
|
||
device topology) → **system** (most-protected) — never silently user-data.
|
||
- **`GET /disks`** — each `DiskInfo` now carries `role`. The controller drives the UI from it
|
||
(system/backup get a lock + no destructive controls; user-data is customer-manageable).
|
||
- **Gate tier (`reconcile`)** — new `CustomerConfirmable` disposition + `Gate.AuthorizeStorageWipe`:
|
||
- role=**user-data** → **customer-confirmable**: allowed iff the request carries an explicit
|
||
customer confirmation **bound to the device's durable id** (the agent re-resolves the durable id
|
||
and matches; a confirmation for one disk can't wipe another). **No operator signature.** A
|
||
user-data drive is already within the in-guest controller's blast radius (it bind-mounts `/mnt`),
|
||
so customer-confirmation adds no new reach. Recorded in the **audit log** with the durable id
|
||
(`AuditRecord.DurableID`).
|
||
- role=**system**/**backup** → unchanged **operator-signature** (`pending_signature`). The
|
||
`confirmed` flag is **IGNORED** — a compromised controller asserting `confirmed:true` on a
|
||
protected device is refused **by role**. Every other destructive class (`guest_destroy`,
|
||
`decommission`, `restore_overwrite`, `key_rotation`) keeps operator-signature exactly as before.
|
||
- **`POST /disks/format`** — accepts `confirmed` + `durable_id` (inert for system/backup). The
|
||
data-bearing path tiers by role: user-data customer-confirmed → `mkfs`; user-data unconfirmed →
|
||
403 `needs_confirmation` (+ the durable id to confirm against, NOT an opsign command); system/backup
|
||
→ 403 with the operator-signature pending op (as before). Blank devices stay benign `mkfs`.
|
||
- **Tests** — `role_test.go` (demo-storage mapping + fail-safe), `storage_wipe_test.go` (the gate
|
||
refuses a `confirmed` wipe on system/backup → no exec; durable-id mismatch / missing-durable
|
||
refused; unknown role fails safe), and the localapi format-handler branches (user-data confirmed →
|
||
mkfs; user-data unconfirmed → needs_confirmation, no opsign; confirmed-but-protected → still refused).
|
||
|
||
Pairs with the controller's lockout + type-to-confirm UX + drive-list restyle.
|
||
|
||
## v0.22.0 — expose durable_id in GET /disks (enable controller-side guided storage) (2026-06-11)
|
||
|
||
One-line, read-only addition: `localapi.DiskInfo` gains `durable_id` (mapped from
|
||
`StorageTarget.DurableID`, e.g. `"uuid:<fs-uuid>"` for usb/local-dir). The de-privileged controller
|
||
cannot read a device's fs UUID itself, yet `POST /disks/assign` mounts strictly by UUID — so without
|
||
this it could not complete the guided init/attach flows. The controller strips the `uuid:` prefix to
|
||
get the assign key. No new privilege, no behaviour change to format/assign/eject or the data-bearing
|
||
gate. Pairs with `felhom-controller` v0.43.0 (the storage-management UI rebuild).
|
||
|
||
## v0.21.0 — agent-managed split-horizon LAN resolver (internal/lanresolver) (2026-06-11)
|
||
|
||
LAN clients can now reach their guest **directly** at the same public hostname with the same real
|
||
wildcard cert (no Cloudflare hairpin), via a host-side dnsmasq the agent manages. The host is the
|
||
stable anchor (static LAN IP); the guest stays DHCP/ephemeral and the agent tracks its live IP.
|
||
|
||
- **`internal/lanresolver`** — renders a dnsmasq base drop-in (bind to the host LAN IP, no-resolv,
|
||
upstreams) + a per-customer drop-in `local=/<domain>/` + `address=/<domain>/<guest-ip>`. The proven
|
||
two-line shape: `local=` makes dnsmasq authoritative for the zone so **AAAA returns NODATA** (no
|
||
Cloudflare-AAAA split-brain — the guest has only link-local v6), `address=` is the wildcard A; all
|
||
other names (and their AAAA) forward upstream unchanged.
|
||
- **`Manager`** ensures dnsmasq present (apt) + the base config + enabled, discovers the guest's live
|
||
IPv4 (`pct exec <vmid> -- ip -4 -o addr show dev eth0`) and domain (read from the guest controller's
|
||
pulled `controller.yaml` — the v2 bootstrap omits it), writes drop-ins **write-if-changed**, and
|
||
**reloads** (not restarts) dnsmasq. Tolerates the early-boot pre-lease window (empty IP → skip+retry,
|
||
never a blank record). Logs IP transitions.
|
||
- **`Loop`** — a 7th daemon goroutine: every interval (default 300s) it enumerates provisioned guests
|
||
(`/var/lib/felhom-agent/guests/<vmid>/`) and reconciles each, so the resolver follows DHCP IP changes.
|
||
Config `lan_resolver.{enable,host_ip,upstreams,interval_seconds}` (host_ip defaults to the local-API
|
||
bridge IP). `--selftest=lanresolver -vmid N`.
|
||
- **`configs/felhom-agent.sudoers`** — new `FELHOM_DNSMASQ` alias (apt install dnsmasq; install
|
||
felhom-*.conf drop-ins; systemctl enable/reload dnsmasq; rm felhom-*.conf; the two FIXED `pct exec`
|
||
reads). The agent never touches `/etc/resolv.conf` (host's own resolution unaffected).
|
||
- **Box-down robustness** is a documented **router config** (DNS = [host-IP primary, upstream
|
||
secondary]) so a box reboot degrades to the Cloudflare path, not total DNS loss — see REPORT install step.
|
||
- Spiked live on felhom-pve first (`:53` free, host IP static `192.168.0.162`, host DNS intact, full
|
||
loop from a real LAN client returned the guest IP + AAAA NODATA + the real wildcard cert `200 0`).
|
||
|
||
## v0.20.0 — golden: stacks-dir bind + per-guest hostname/CT name + bake base-infra images (2026-06-11)
|
||
|
||
Lockstep with `felhom-controller` v0.41.0 + a golden rebake. Changes in `configs/build-golden.sh` and
|
||
the provision path; no change to the proxmox/authz/token fences.
|
||
|
||
- **Section-G mount fix (the load-bearing one):** the in-guest controller writes app/infra compose
|
||
stacks under `/opt/docker/stacks` *inside its container*, but the baked controller-bootstrap `docker run`
|
||
never bind-mounted that path. So `docker compose up` (run by the GUEST daemon over the shared socket)
|
||
resolved every relative bind source on the guest filesystem — silently creating empty dirs — which
|
||
broke **every** bind-mounted stack (base infra AND customer apps like immich/nextcloud). The bootstrap
|
||
unit now `mkdir -p /opt/docker/stacks` and adds a **same-path host bind**
|
||
`-v /opt/docker/stacks:/opt/docker/stacks` (a named volume would NOT fix this). Empirically confirmed on
|
||
guest 9201 before writing the fix.
|
||
- **Per-guest container hostname (3A):** the bootstrap unit derives `customer.id` from
|
||
`/etc/felhom-bootstrap/bootstrap.json` with a portable `sed` parse (NO jq in the golden) and passes
|
||
`--hostname <customer-id>` to `docker run`, so the controller's `os.Hostname()` (its hub-reported
|
||
hostname) is the customer id, not the Docker container ID. Fail-safe: no parse → no `--hostname`.
|
||
- **Per-guest CT/LXC name (3B):** `--selftest=provision` now defaults `-hostname` to the (DNS-safe
|
||
sanitized) `-customer-id` when not given, so the bring-up's existing `SetConfig hostname` step
|
||
(`bringup.go`) names the CT meaningfully (e.g. `demo-felhom`) instead of inheriting the golden's
|
||
`felhom-golden`. New `sanitizeHostname` (lowercase, collapse invalid → `-`, trim, ≤63).
|
||
- **Bake base-infra images:** the golden now also pulls the three PINNED, PUBLIC base-infra images
|
||
(`traefik:v3.6.7`, `cloudflare/cloudflared:2026.6.0`, `gtstef/filebrowser:1.3.3-stable`) into its Docker
|
||
storage so the controller's first-boot bring-up is OFFLINE-capable. A hard gate (`docker manifest
|
||
inspect`) fails the bake early on a bad pin. Tags MUST match the controller's `internal/infra` constants.
|
||
|
||
## v0.19.0 — bootstrap contract v2: agent relays the hub retrieval passphrase (no host key in the guest) (2026-06-11)
|
||
|
||
Lockstep with `felhom-controller` v0.40.0. Fixes the onboarding 401: a freshly provisioned guest's
|
||
controller used to come up with the agent's **host** hub key baked in, which the hub's `/api/v1/report`
|
||
(customer-scoped auth) rejects. The agent now bakes a **v2 bootstrap** carrying only what the controller
|
||
needs to **pull** its own config from the hub — the agent never touches the customer-scoped key or CF
|
||
tokens.
|
||
|
||
### Changed — bootstrap contract `v1 → v2` (`internal/provision`)
|
||
- `SchemaV1 → SchemaV2 = "felhom.bootstrap/v2"`. **`DocCustomer`** drops `name`/`domain`/`email` (keeps
|
||
`id`). **`DocHub`** drops `api_key`/`host_id`, adds **`retrieval_password`** (the customer's hub
|
||
retrieval passphrase — SECRET). `DocLocalAPI` unchanged. The contract is byte-compatible with the
|
||
controller's `internal/bootstrap.Bootstrap` (cross-repo round-trip verified).
|
||
- `backhalf.go`: renders the v2 Doc; validation now requires `customer.id` + `hub.url` +
|
||
`hub.retrieval_password` (was `customer.id` + `customer.domain`). Write/0600/chown/`pct set` unchanged.
|
||
- `cmd/felhom-agent/main.go` `--selftest=provision`: **new required `-hub-password`** flag (the customer's
|
||
hub retrieval passphrase; the customer must already exist in the hub). Stops baking `cfg.Hub.APIKey` /
|
||
`cfg.Hub.HostID`. `-customer-domain/-name/-email` still accepted (bring-up may use them) but NOT baked.
|
||
|
||
### Changed — `configs/build-golden.sh`
|
||
- Default `CONTROLLER_IMAGE` bumped off the stale `:v0.35.0` → `:0.40.0` (matches the registry's no-`v`
|
||
tag convention; latent footgun fixed).
|
||
|
||
### Tests
|
||
- `doc_test.go`/`backhalf_test.go` updated to the v2 shape (assert no `api_key`/`host_id`,
|
||
`retrieval_password` present, `customer` carries only `id`). `go build ./... && go test ./...` green.
|
||
|
||
## v0.18.0 — slice 10D: DR capstone — identity escrow + restore-mode consumption (agent side) (2026-06-10)
|
||
|
||
The agent half of the slice-10 DR capstone (closes slice 10). Grounded by both 10-series spikes
|
||
(escrow-consumption + identity-restore). The hub half (recovery-mode toggle, re-enroll + credential
|
||
rotation, directive serving) is hub v0.11.0. **Operator-side rotation model (locked):** the hub holds
|
||
no Cloudflare write-power; the destructive tunnel/PBS rotation is the operator's step from a trusted
|
||
environment (same spirit as 10B).
|
||
|
||
### Added (`internal/escrow`)
|
||
- **Identity escrow** (`identity.go`): `WrapIdentity`/`UnwrapIdentity` (+ `…Bundle`) wrap the
|
||
`{tunnel_token, pbs_token}` bundle under the SAME recovery code `R` via **`age`** (scrypt +
|
||
ChaCha20-Poly1305 — a vetted passphrase-AEAD, not hand-rolled), reusing the K-escrow pty mechanism
|
||
(passphrase via the tty, data via files; `R`/tokens never logged). Same two-factor, zero-knowledge
|
||
shape as the K-escrow. A **wrong R fails closed** (no bundle). `age` is a runtime dep for the
|
||
identity path (analogous to proxmox-backup-client for K).
|
||
- **`escrow.Create`** gains an optional `IdentityBundle` → also emits an `IdentityBlob` under the same
|
||
R (additive; the K-escrow + 10C `Consume` paths are byte-unchanged). Self-verifies the identity
|
||
round-trip before shipping.
|
||
- **`--selftest=escrow-create -identity-bundle <file> -directive <file>`** — also wrap + upload the
|
||
identity blob + the **non-secret** DR directive (pbs repo/ns, expected key fingerprint, tunnel id).
|
||
- **`--selftest=identity-consume -blob <file> -keydest <file>`** (R via `FELHOM_RECOVERY_CODE`) —
|
||
recover the identity bundle through the real code; tokens written 0600, never logged.
|
||
|
||
### Tests
|
||
- identity bundle round-trips (wrap→unwrap byte-identical; blob is opaque ciphertext); wrong R fails
|
||
closed + the blob stays retryable; input validation. K-escrow/10C tests byte-unchanged (additive).
|
||
(age integration tests gated to a host with the `age` CLI.)
|
||
|
||
## v0.17.0 — slice 10C: escrow consumption (productionize the spike) (2026-06-10)
|
||
|
||
Turns the throwaway 10C spike harness into a real, tested **`Consume`** path: recover the PBS key
|
||
`K` from an R-wrapped escrow blob, **gate it on the expected fingerprint**, and install it for the
|
||
restore. The spike already proved the crypto + real-data restore; this bakes its findings into
|
||
production code. **Agent-only** — 10C *reads* the four inputs as parameters (so it stays
|
||
standalone-testable); 10D sources blob/fingerprint/PBS-connection from the hub and prompts for R.
|
||
**Zero-knowledge holds**: the hub serves everything except **R** (by hand from the customer), so a
|
||
hub compromise alone still can't decrypt.
|
||
|
||
### Added
|
||
- **`escrow.Consume(ctx, blob, R, expectedFingerprint, keyDest)`** — the consumption contract:
|
||
1. **Unwrap** the blob (a copy — F-C6: the input blob is read-only → a failed Consume is
|
||
**retryable**) with `R`; a **wrong R fails closed** at the scrypt KDF (F-C3) → a clear,
|
||
R-free error, **nothing written**.
|
||
2. **Fingerprint gate (F-C4)** — `KeyFingerprint(recovered)` must equal the expected (the hub
|
||
knows it); a mismatch **fails fast + loud, no install, no restore attempted**.
|
||
3. **Atomic install (F-C2)** at `keyDest` (`0600`, write-temp-sibling→rename); any failure leaves
|
||
**no partial install**. The recovered key lives only in a `0700` tempdir that is always removed.
|
||
**Secret discipline:** `R` and key bytes are never logged/persisted (only fingerprint prefixes);
|
||
`K` is never mutated.
|
||
- **`--selftest=escrow-consume`** (`-blob -fingerprint -keydest`, R via env `FELHOM_RECOVERY_CODE`
|
||
to keep it off the command line) — invokes the real `Consume` live (the spike's S3 via the
|
||
production path, not a harness).
|
||
|
||
### Tests (non-hollow)
|
||
- valid → key installed + `KeyFingerprint(dest) == expected` + `0600` + blob byte-unchanged;
|
||
**wrong R** → error, **no file at dest**, blob unchanged; **fingerprint mismatch** → fail fast,
|
||
**no install** (the gate runs before any restore); input validation; format-tolerant fingerprint
|
||
compare (no empty-fingerprint gate-bypass); atomic-install permissions (integration tests gated to
|
||
a host with `proxmox-backup-client`).
|
||
|
||
## v0.16.0 — slice 10B: operator-signed destructive completion (offline key + signing CLI) (2026-06-10)
|
||
|
||
The security centerpiece: a destructive op runs ONLY on a verified, operator-signed authorization
|
||
— signature valid against a **pinned** operator pubkey (never the hub's or the blob's), nonce
|
||
unseen + durably burned, in-window, host-bound, and **resource-bound to a DURABLE device id** that
|
||
execution re-resolves + re-inspects. Decision (a): **offline operator key + signing CLI**,
|
||
hardware-key-ready (`sk-`/YubiKey via ssh-keygen). The key floor holds: the signing key is NOT in
|
||
the hub and NOT in the agent. Concrete consumer: this **closes the 8C data-bearing-wipe
|
||
`pending_signature` gap**. Pairs with hub v0.10.0.
|
||
|
||
### Added
|
||
- **`cmd/felhom-opsign`** — the operator's offline signing CLI. Builds the canonical `OpBlob` by
|
||
**reusing `authz.CanonicalBlob`** (the exact production path the verifier authenticates over — so
|
||
signer + verifier can never drift) and signs it with **`ssh-keygen -Y sign -n felhom-op-v1`**
|
||
(hardware-ready). Output: a `{op_blob_b64, sig_armored}` envelope to hand to the hub jobs queue
|
||
(optional `--upload`). Touches ONLY the operator's signing key.
|
||
- **`authz.CanonicalBlob`** — promoted to production (was test-only) so the CLI + verifier share one
|
||
canonical-bytes source; params canonicalized (sorted keys, compact).
|
||
- **`internal/storage` durable device identity** (`durable_device.go`): `DeviceDurableID` (derive a
|
||
stable `byid:`(wwn/serial)/`byuuid:` id from the world-readable udev symlinks — no privilege, no
|
||
subprocess) + `ResolveDurableDevice` (re-resolve to the current `/dev` path; a path-only/unknown
|
||
scheme is REFUSED). The resource-level anti-retarget.
|
||
- **`internal/signedjobs`** (new): the queue consumer. `Runner` fetches each opaque job → runs it
|
||
through the **gate** (the LOCKED authz pipeline) → on all-pass hands the verified op to an
|
||
`Executor`; the order is **verify → nonce-burn (durable, in Verify) → execute → clear job**. The
|
||
**`WipeExecutor`** is the 8C consumer: resolve the signed durable id → **re-derive + match**
|
||
(anti-retarget) → **re-inspect (8C classifier)** the device is still the data-bearing target →
|
||
`mkfs`. A vanished/changed/non-data-bearing device or a path-only binding is refused **even with a
|
||
valid signature**. Wired as a second `EnvelopeObserver` (runs on `HasSignedOps`).
|
||
- **`hub.Client.Jobs` / `CompleteJob`** + `hub.MultiObserver`; the 8C format refusal now **surfaces
|
||
the bound op** (op + durable id + host) in its 403 `pending_op` + a `felhom-opsign …` hint.
|
||
|
||
### Pinning / rotation
|
||
- Operator pubkeys are pinned via `authz.signers` (config, trusted path — provision/agent config,
|
||
NEVER hub-alone), **multiple** keys (KeyID selects; role-scoped), so a backup/rotation key exists
|
||
without a flag-day. Unchanged from the slice-4 verifier wiring; 10B activates the execute path.
|
||
|
||
### Tests (real crypto, non-hollow)
|
||
- `signedjobs` runner over the **real** gate+verifier (in-Go minted SSHSIGs): valid → executor runs
|
||
once + job cleared; **replay** (nonce burned) / **non-pinned signer** / **expired** / **retarget**
|
||
(other host) / **forged sig** / **no pinned signer** → all rejected, **executor never called**;
|
||
malformed envelope cleared.
|
||
- `WipeExecutor`: valid → `mkfs` runs; **path-only**, **durable-id mismatch**, **device gone**,
|
||
**re-inspect non-data-bearing**, **not-probed** → all refused, `Format` not called.
|
||
- `storage` durable: wwn-preference, uuid-fallback, path-only/traversal refusal, round-trip,
|
||
missing-device error (symlink tests gated to Linux — the agent's OS).
|
||
|
||
## v0.15.0 — slice 10A: hub desired-state serving — the "Down" channel (2026-06-10)
|
||
|
||
The agent half of slice 10A. The control envelope (`hub.ControlEnvelope`) stops being "reserved — ignored" and becomes the live **Down channel**: a cheap change-notification on every heartbeat. The agent caches the hub's desired-state + its generation; only when **`DesiredGeneration` advances** does it fetch the full state (the heartbeat stays light, the heavy state moves on change). The engine then reconciles **benign** deltas and the gate marks an explicit **destructive** delta `pending_signature` (no signer in 10A → never executed; signed execution is 10B). Pairs with hub v0.9.0.
|
||
|
||
### Added / changed
|
||
- **`internal/reconcile`**: `DesiredGuest.Decommission` — the canonical **destructive desired-state delta** (an EXPLICIT flag, not "absent from the list", so a partial hub list can never mass-destroy). The planner emits `ActionDecommission` → `ClassDecommission` → Destructive → the gate refuses it `pending_signature`. `Reconcile` now counts a `pending_signature` refusal as **`Result.Pending`** (expected, logged INFO) rather than a failure; any other refusal stays a real failure. `ActionDecommission` has **no executor** (slice 10B) — a defensive guard refuses to run it. New **`CachingProvider`** (thread-safe DesiredState + generation cache; `Desired`/`Update`/`Generation`) — the production `DesiredProvider`, replacing `EmptyProvider` in the daemon engine (empty until the hub serves intent → cold-start is a live no-op, unchanged).
|
||
- **`internal/hub`**: the **`ControlEnvelope`** fields are now active (DesiredGeneration drives the fetch, HasSignedOps noted). New wire types **`DesiredStateResponse`** + **`WireDesiredState`** (guests + forward-compat `restore_directive` (10D) / `pbs_namespace` / opaque `storage_manifest`+`backup_policy`) + **`WireDesiredGuest`** (vmid/run/spec/description/decommission). New **`Client.FetchDesiredState`** (GET `/api/v1/hosts/{host_id}/desired-state`, self-scoped to the client's own host). New **`EnvelopeObserver`** loop seam + `SetEnvelopeObserver` — the loop hands the envelope to the sync layer each cycle (hub does not import reconcile/desired).
|
||
- **`internal/desired`** (new): the **`Syncer`** — implements `hub.EnvelopeObserver`, fetches desired-state on a generation advance, maps the wire shape to the reconcile domain, and updates the `CachingProvider`. Caches the **fetched** generation (robust to a generation that advanced mid-fetch); a fetch failure keeps the last-known state. `restore_directive` is carried + logged, not acted on (10D). Wired in `cmd/felhom-agent` (daemon): provider → engine, syncer → loop.
|
||
|
||
### Tests
|
||
- reconcile: a desired-state with one benign + one decommission delta → **benign applied, destructive gated pending (not executed)**; `Plan` emits decommission-only for a decommissioned guest + classifies Destructive; `CachingProvider` update/isolation.
|
||
- desired: **fetch-once-on-advance** (no re-fetch on an unchanged generation), fetch-failure-keeps-cache, caches-the-fetched-generation.
|
||
- hub client: `FetchDesiredState` hits the self-scoped path with the bearer + decodes (incl. `restore_directive`); a 403 is a typed `HTTPError`.
|
||
- loop: the cycle notifies the observer + adopts `PollIntervalSeconds`; a report error skips the observer.
|
||
- cross-repo golden: `testdata/desired-state.golden.json` + `control-envelope.golden.json` decode + key-set guard, **byte-identical** with felhom.eu/hub.
|
||
|
||
## v0.14.0 — slice 9: host metrics to the controller (`GET /host/metrics` + CPU-temp collector) (2026-06-10)
|
||
|
||
The de-privileged controller (slice 8C) sees only its own cgroup, so it can't read host health itself. Slice 9 **re-serves** the slice-4 collector's host + per-storage view to the customer over the local API, plus the one missing collector — CPU/chassis temperature — so the customer sees their box's health in the controller. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot). Assumption: **one customer per host** (the home-server model); if a host ever serves multiple customers, host-wide CPU/mem would leak cross-customer load → revisit then.
|
||
|
||
### Added / changed
|
||
- **CPU/chassis-temp collector** (`internal/hub/cputemp.go`): `SysfsTempReader` reads the CPU package temperature straight from sysfs — hwmon (`coretemp`/`k10temp`/`zenpower`/`cpu_thermal`, preferring the `Package id 0` input) then the thermal zones (preferring `x86_pkg_temp`/`coretemp`/`cpu-thermal`, falling back to `acpitz`). **No external binary, no privilege** (sysfs nodes are world-readable), so the root-CLI fence is untouched. **Graceful-null**: a missing sensor, an unsupported board, an implausible reading (outside 5–150 °C), or any read error all degrade to `null` ("n/a") — a missing sensor never fails the report. Wired into the collector via the new `TempReader` seam (nil-safe).
|
||
- **`HostMetrics.CPUTempC *int` (`cpu_temp_c`)** — new nullable wire field on the **shared** `HostMetrics` struct (same nullable contract as the disk `SmartSummary.TemperatureC`). It rides the **hub report too** (operator freebie) → cross-repo host-report golden updated.
|
||
- **`Collector.HostMetricsNow(ctx)`** — a fresh `NodeStatus` + CPU-temp read returning just the host block, the source for the local API (current cpu%/temp, not the 15-min snapshot). `Collect()` now also populates `cpu_temp_c` on the hub report. `Collector.SetTempReader` injects a fake in tests.
|
||
- **`GET /host/metrics`** (`internal/localapi/host_metrics.go`): host-wide health (cpu%/mem/load/uptime/`cpu_temp_c`) + per-storage capacity targets (total/used/fraction, thin-pool, SMART temp+wear). Token-authed via `withGuest` (host-wide data; cross-guest `?vmid=` still 403). Best-effort on storage (a view error still returns the host block). Served only when the `HostMetrics` provider (the shared collector) is wired — else 503 "not configured". Wired in `buildLocalAPIServer`.
|
||
|
||
### Tests
|
||
- `cputemp_test.go`: a fake `/sys` layout proves hwmon package-preference, hwmon first-input fallback, thermal-zone-by-type selection over a non-CPU hwmon, **graceful-null on a sensorless host** (no error), and rejection of implausible (0 m°C) readings.
|
||
- `hostmetrics_test.go`: `HostMetricsNow` populates the temp, gracefully nulls it, hard-errors on `NodeStatus` failure; `Collect()` carries the temp.
|
||
- `host_metrics_test.go` (localapi): populated host+storage with a valid token; `cpu_temp_c:null` serializes; **401 without a token** (collector never invoked); 403 on a cross-guest `?vmid=`; 503 when not configured.
|
||
|
||
## v0.13.0 — slice 8B.2: quiesce downtime optimization (`snapshotted` phase) (2026-06-10)
|
||
|
||
The agent half of slice 8B.2. In snapshot mode, vzdump only needs the app-stopped state captured at
|
||
the **storage-snapshot moment**; after that it reads from the snapshot and the app can resume. The
|
||
agent now emits a **`snapshotted`** phase on `GET /backup/status` when the snapshot is taken, so the
|
||
controller (v0.38.0) resumes its app early — app downtime drops from *whole-backup* to
|
||
*until-snapshot* with no loss of app-consistency. Validated Phase-0 first on PVE 9.2.2: the marker is
|
||
`INFO: create storage snapshot 'vzdump'`; downtime ~24s→~1s for a 934 MB guest.
|
||
|
||
### Added / changed (`internal/backup` + `internal/localapi`)
|
||
- **`BackupRunner.BackupWithSnapshotHook(ctx, vmid, onSnapshot)`** — while the vzdump runs, a watcher
|
||
tails the task log (`TaskLogTail`) for the **`create storage snapshot`** marker and fires
|
||
`onSnapshot` **once**. The marker only appears in snapshot mode (stop/downgraded takes no storage
|
||
snapshot), and the watcher also bails on `backup mode: stop` — so it never fires in stop mode.
|
||
(`Backup` keeps its signature for the scheduler/selftest; both share one body.)
|
||
- **`/backup/status` phase `snapshotted`** (between `running` and `done`): `handleBackup` passes the
|
||
hook → `markSnapshotted` flips the running job to `snapshotted`. `done`/`failed` semantics unchanged.
|
||
|
||
### Tests
|
||
- localapi: snapshot mode → phase reaches `snapshotted` before `done` (gated fake holds the backup
|
||
open); stop mode → `snapshotted` **never** emitted (stays running → done). runner: the watcher
|
||
fires `onSnapshot` on the marker; in stop-mode log it never fires. `snapshotWatchInterval` is a
|
||
package var so tests run fast.
|
||
|
||
## v0.12.0 — slice 8C Phase A: disk endpoints + data-bearing classifier gate + mkfs executor (2026-06-10)
|
||
|
||
The agent half of slice 8C, Phase A (additive). Adds the host disk-management endpoints the
|
||
controller's disk UI drives — with the **8C security invariant**: the agent decides
|
||
data-bearing-ness by **inspecting the actual device** (agent-internal evidence), NEVER from the
|
||
caller's claim. A compromised controller asserting "this drive is blank" cannot wipe a data-bearing
|
||
drive. (Controller rewire + disk-subsystem retirement + de-privilege are Phases B/C, `felhom-controller`.)
|
||
|
||
### Added
|
||
- **`internal/storage` — `mkfs` executor + data-bearing inspection.** `SudoHostOps.Format(device,
|
||
fstype)` (device-pinned, `ValidateBlockDevice`+`ValidateFSType`, narrow `FELHOM_FORMAT` sudoers —
|
||
`mkfs.ext4 -F` / `mkfs.xfs -f` on a `/dev/*` path the agent fine-validates first).
|
||
`SudoHostOps.InspectDevice(device)` → `DeviceProbe` (filesystem signature via `blkid -p`, partition
|
||
table / partitions / mount via `lsblk -J`). **`DeviceProbe.DataBearing()` is conservative**: any
|
||
signature / partition table / partition / mount — OR a probe that did not read cleanly — is
|
||
data-bearing (fail-safe; an unreadable device is never called blank).
|
||
- **`internal/localapi` — the §6 disk endpoints**, all self-scoped (token→guest; cross-guest 403):
|
||
- `GET /disks` — host drives + a **data-bearing flag** (UI hint). Read-only/benign.
|
||
- `POST /disks/assign` — attach a drive as a mount (benign, additive → `EnsureMount`). Self-serve.
|
||
- `POST /disks/eject` — safe-unmount (benign, data preserved) + the **dependent guests** that
|
||
mount it (so the controller can warn which apps lose that storage).
|
||
- `POST /disks/format` — **the security centerpiece**: the agent **inspects the device itself**;
|
||
blank → benign → `mkfs`; **data-bearing → ClassStorageWipe → the slice-4 gate → refused
|
||
`pending_signature`** (the operator-signed completion is slice 10). The caller's claim is
|
||
ignored — only a device the agent reads as blank is formatted.
|
||
- `storageGateAdapter` bridges the format path to the slice-4 reversibility gate (no new gate/crypto).
|
||
|
||
### Tests
|
||
- localapi (security matrix): blank device → **mkfs called, gate not consulted**; a **data-bearing
|
||
device → 403, mkfs NEVER called**, gate consulted (`pending_signature`); an **ambiguous/unprobed
|
||
device → treated destructive** (fail-safe); even a gate that *allows* does not format data-bearing
|
||
in 8C; assign → `EnsureMount`; eject → `Unmount` + dependent guests; cross-guest → 403; bad
|
||
device/fstype → 400; unconfigured → 503.
|
||
- storage: `ValidateBlockDevice`/`ValidateFSType` (whitelist + injection rejection); `InspectDevice`
|
||
blank/filesystem/partition-table/mounted/failed-probe-fail-safe; `Format` invokes the right `mkfs.*`.
|
||
|
||
## v0.11.0 — slice 8B: app-consistent backup — /backup/due policy + /backup/status phases (2026-06-10)
|
||
|
||
The agent half of slice 8B (doc 03 §8). Turns the 8A thin backup stubs into the real policy the
|
||
in-guest controller's quiesce loop drives (controller half: `felhom-controller` v0.36.0). No hub
|
||
change. The downtime optimization (`vzdump --mode snapshot` + a `snapshotted` phase) is the 8B.2
|
||
fast-follow; the hub-served per-guest policy is slice 10.
|
||
|
||
### Changed (`internal/localapi`)
|
||
- **`GET /backup/due`** — real **cadence** policy (replaces the 8A "never backed up" stub): a guest
|
||
is due when no **successful** backup is recorded OR the newest one is older than the agent-local
|
||
cadence (`backup.backup_cadence_seconds`, default 24h). A successful `POST /backup` flips due to
|
||
**false** for the window, so the controller won't re-quiesce in a loop. A failed backup does not
|
||
satisfy the cadence. Returns `age_seconds` for diagnosis.
|
||
- **`GET /backup/status`** — real **phases** `idle | running | done | failed` + the job id, so the
|
||
controller can poll a backup to completion (was: just the latest stored backup).
|
||
- **`POST /backup`** — returns a **job id** + `running` phase; tracks the in-flight job and is
|
||
**single-flight per guest** (a second POST while one runs returns the same job — no concurrent
|
||
vzdump). On completion the job transitions done/failed and the result is recorded to the store.
|
||
- Config: `backup.backup_cadence_seconds` + `BackupCadence()`; the local-API server takes the cadence.
|
||
|
||
### Tests
|
||
- `/backup/due`: due when stale / no backup, **not due within the window after a success**, due again
|
||
past the cadence, **a failed backup does not count**. `/backup/status`: running→done and
|
||
running→failed (gated fake to observe the running phase). `POST /backup` single-flight (one vzdump
|
||
for concurrent POSTs). All still self-scoped (token→guest).
|
||
|
||
## v0.10.0 — slice 8A: agent local-API server + provisioning back-half (2026-06-10)
|
||
|
||
The host-agent half of slice 8A (doc 03 §6). Adds the per-guest **local API** the in-guest
|
||
controller calls over the bridge, and the **provisioning back-half** that follows the slice-7
|
||
bring-up front half. Grounded by `felhom.eu/documentation/tests/slice8a-channel-deploy-spike-findings.md`
|
||
(commit `4a81a96` — channel + deploy plumbing proven; the 5 gotchas resolved here). Controller half
|
||
is `felhom-controller` v0.35.0. No hub change.
|
||
|
||
### Added
|
||
- **`internal/localapi`** — the HTTPS local-API server (doc 03 §6), the **per-guest authorization
|
||
gate**. Serves a **persisted self-signed leaf** with a **stable SHA-256 fingerprint** (generated
|
||
once; a fresh cert each boot would invalidate every baked bootstrap pin). The **7 §6 endpoints**,
|
||
all **self-scoped to the caller's own guest**: `GET /storage` (this guest's mpN mounts + fast/slow
|
||
class from the slice-5/7 storage view), `POST /snapshot`, `POST /rollback`, `POST /backup`
|
||
(enqueued, crash-consistent — the app-consistent quiesce loop is 8B), `GET /backup/due` (thin in
|
||
8A), `GET /backup/status`, `GET /restore-test/status`.
|
||
- **Token store** (`tokenstore.go`): durable, crash-safe per-guest token→guest map that persists
|
||
only a **SHA-256 hash** of each token (the plaintext exists transiently at mint→write-to-mount,
|
||
then is discarded), last-write-wins per guest, fsync'd append-only JSONL (mirrors the nonce store).
|
||
- **Self-scoping**: the VMID is resolved ONLY from the token; an explicit `vmid` (query/body) that
|
||
disagrees → **403 and the proxmox op is never issued for the other guest**; absent/unknown → 401.
|
||
- **`internal/provision`** — the back-half: mint the per-guest token → render the stable
|
||
**`bootstrap.json`** contract (schema `felhom.bootstrap/v1`; **no registry credential** — the
|
||
controller image is baked into the golden) → write it `0600` → **`chown 100000:100000`** (the
|
||
unprivileged-LXC mapped guest-root, spike gotcha 1) → attach a **read-only bind mount** via
|
||
`pct set`. Host-side only (F3 — the agent never enters the guest; **no `pct exec`**). The token
|
||
plaintext is never logged and never returned.
|
||
- **`--selftest=provision`** — the full chain on-demand: bring-up (provision) front half + the
|
||
back half; keeps the guest for the golden's baked controller-bootstrap unit to deploy.
|
||
- **`config.LocalAPIConfig`** (`local_api`) — enable + bridge `listen_addr` + cert/key paths + token
|
||
store path. The server is an optional 6th daemon goroutine, disabled cleanly when unconfigured or
|
||
on a token-store/cert failure (the daemon still reports/reconciles).
|
||
- **`configs/build-golden.sh`** now **bakes the controller image** (pulled once on the trusted build
|
||
host, then `docker logout` — no cred baked) + a **controller-bootstrap unit** that deploys the
|
||
**baked** image from the config mount on boot (no login/pull at deploy).
|
||
- **`configs/felhom-localapi-firewall.example`** — host firewall narrowing of the local-API port to
|
||
the guest bridge subnet (nft/iptables/PVE variants; defense-in-depth — the token stays the gate).
|
||
- **`configs/felhom-agent.sudoers`** — a narrow `FELHOM_PROVISION` alias (`chown 100000:100000` +
|
||
`pct set` bind-mount, both confined to the agent-owned `/var/lib/felhom-agent/guests/*` path) for
|
||
the non-root least-privilege deployment.
|
||
|
||
### Security / design notes
|
||
- The local-API leaf is pinned by **leaf-cert SHA-256** (decision: consistency with the agent's
|
||
PVE/PBS pinning); the fingerprint is baked into each guest's bootstrap.
|
||
- The back-half's host-root ops (chown + bind-mount attach) are **NOT** added to `proxmox.Privileged`
|
||
(which is fenced to its 3 exceptions) — they live in `internal/provision` and run through the shared
|
||
`Runner` (direct as root, or `sudo -n` with the new sudoers alias). This is the per-guest
|
||
provisioning host-root surface, host-side and F3-compliant.
|
||
|
||
### Tests
|
||
- localapi: self-scoping (cross-guest snapshot/rollback/backup → 403, op never issued for the other
|
||
guest; own-guest uses the token's VMID), 401 paths, `/storage` class mapping, `/backup` enqueue,
|
||
the thin `/backup/due`, status scoping; the token store persists only the hash (plaintext never on
|
||
disk), last-write-wins, survives reopen, uniqueness; the leaf fingerprint is stable across reload.
|
||
- provision: writes `0600` + chowns + attaches the bind mount with the right args; the **token never
|
||
appears in the Result**; the cross-repo `bootstrap.json` contract key-set is pinned.
|
||
|
||
## v0.9.0 — slice 7 close-out: PBS recovery-code escrow creation (2026-06-10)
|
||
|
||
The first code that touches the PBS client encryption key `K` and introduces the customer recovery
|
||
code `R`. Default posture is **zero-knowledge**: Felhom holds an opaque `R`-wrapped blob (cannot
|
||
open it), the customer holds `R`. Grounded by `felhom.eu/documentation/tests/slice7-escrow-spike-findings.md`
|
||
(round-trip proven on a throwaway: the `R`-recovered key restores a real encrypted snapshot). Hub
|
||
opaque storage is the `felhom.eu` half (hub v0.8.0); consumption/serving is slice 10.
|
||
|
||
### Secret discipline (overriding)
|
||
`R` is `crypto/rand`, ≥128 bits, surfaced **exactly once** and **never** logged/persisted/committed;
|
||
the wrap pty's echo is discarded so `R` can't leak. `K` is read by location, **never modified** (the
|
||
live key file is byte-unchanged — Wrap operates on a copy), never logged.
|
||
|
||
### Added
|
||
- **`internal/escrow`** — `Create` generates `R` (10 EFF-wordlist words ≈ 129 bits), wraps `K` under
|
||
`R` via the **PBS-native** `proxmox-backup-client key change-passphrase --kdf scrypt`, and
|
||
**self-verifies** the blob recovers `K` (fingerprint match) before shipping. The wrap is driven
|
||
over a **stdlib pty** (`x/sys/unix`; spike F-A1 — the command is TTY-only) with **output discarded**
|
||
(F-A2 — the pty echoes the passphrase). Opt-in outputs: **(b)** `R`-wrapped offline copy (two-factor,
|
||
no extra trust) and **(a)** raw paperkey (single-factor, unrevocable — loud caveat).
|
||
- **`--selftest=escrow-create`** (`-storage`, `-paperkey`, `-offline`, `-upload`): surfaces `R` once
|
||
to stdout (never the logger), prints the opaque blob's size/fingerprint/posture, and with
|
||
`-upload` PUTs the blob to the hub (`/api/v1/hosts/{host_id}/escrow`, per-host key).
|
||
- Config: `escrow` section (`posture` default `zero_knowledge`, `pbs_storage_id`); `PBSEncKeyPath`
|
||
helper (the `<id>.enc` key K).
|
||
- Runtime dependency on the `proxmox-backup-client` CLI (the PBS key+passphrase KDF).
|
||
|
||
### Tests
|
||
- `R` entropy ≥128 / 10-word format / uniqueness; integration round-trip (wrap→unwrap fingerprint
|
||
match, **wrong-`R` fails**, **live `K` byte-unchanged**, blob ≠ plaintext key) guarded to
|
||
linux+`proxmox-backup-client`; the agent→hub wire-contract key-set (mirrors the hub's).
|
||
- **Live-validated** (demo): `escrow-create` → `R` (10 words) surfaced once, blob 383 B opaque,
|
||
self-verify ok, **live `K` sha256 unchanged**, exact `R` absent from stderr/journal.
|
||
|
||
## v0.8.0 — slice 7 Phase 1: unified bring-up reconcile job (provision + guest-loss DR) (2026-06-09)
|
||
|
||
The shared FRONT HALF of provision and guest-loss DR, as a journaled reconcile job mirroring the
|
||
slice-6 restore-test's crash-safety — but it KEEPS the guest on success and applies a
|
||
scenario-specific identity policy. Agent-only; no hub/wire change (the new guest auto-appears in
|
||
the host-report via `ListLXC`). Grounded by the slice-7 bring-up spike findings (commit `3342993`):
|
||
F1 (restore preserves the archived MAC → provision reset is unconditional), F3 (SSH host keys do
|
||
not auto-regenerate → a baked golden first-boot unit, not an agent guest-internal op), F4 (the
|
||
transient PVE config-lock 500 → bounded retry).
|
||
|
||
### Added
|
||
- **`reconcile.RunBringUp`** (`bringup.go`) — `BringUpSpec` (Mode `provision`|`dr_guest_loss`,
|
||
Archive, VMID, RestoreStorage, Hostname, Cores/MemoryMB, RootfsGrowGB, Mounts, KeepMAC,
|
||
BootTimeout) → `BringUpResult` (VMID, AssignedMAC, Pass, Verified, StartWarnings/Recognized).
|
||
Sequence (each mutation preceded by journaling the owning entry): restore → identity reset →
|
||
size → attach mounts → start LINK-UP. **Verdict is liveness (`waitRunning`), never the start
|
||
exitstatus** (reuses the v0.7.0 WARNINGS surface). **Success KEEPS the guest** (no teardown).
|
||
- **Scenario-specific identity reset** (doc 03 §9): *provision* → fresh MAC unconditionally
|
||
(`PUT net0` with `hwaddr` omitted → PVE regenerates, F1) + hostname; machine-id + SSH host keys
|
||
regenerate guest-side on first boot (golden bake + the new unit) — the agent does NOT touch
|
||
guest internals. *dr_guest_loss* → preserve continuity (keep hostname; keep MAC unless
|
||
`KeepMAC=false`); never resets restic/tunnel/hub identity.
|
||
- **Compensating rollback** — any mid-flight failure destroys the just-created guest
|
||
(`ClassGuestDestroy`, benign via `Provenance{SameTxnCreated:true}`, gated); on teardown failure
|
||
the entry is left in-flight for `Recover`. New journal flag **`Rollback`** + `Recover`'s
|
||
`recoverBringUp` reap a half-built guest left by a mid-job crash (idempotent, via `ListLXC`).
|
||
- **F4 config-lock retry** — steps 3+5 coalesced into ONE `PUT config` (net0+hostname+cores+
|
||
memory+mpN); rootfs grow stays its own call. `setConfigWithLockRetry` retries ONLY the transient
|
||
PVE config-lock 500 (`pveConfigLock`: 500 + "can't lock file"/"got timeout"); any other error
|
||
fails immediately — never retried.
|
||
- **`--selftest=bring-up`** (`-mode provision|dr -archive -vmid -hostname [-keep]`) — runs the real
|
||
journaled job (after a `Recover`), then tears the guest down unless `-keep`.
|
||
- **`configs/build-golden.sh`** — the validated golden recipe as a script, incl. the F3
|
||
first-boot `felhom-regen-hostkeys.service` unit (Condition-gated: fires on provision, no-ops on
|
||
DR). The slice-7 spike archive (which lacks the unit) is superseded.
|
||
|
||
### Deferred (stated, not built)
|
||
- Provisioning BACK HALF (controller deploy, bootstrap, per-guest token mint) → **slice 8**.
|
||
- Host-loss DR + PBS escrow consumption → **slice 10**.
|
||
- The SOURCE of a `BringUpSpec` (hub desired-state: which archive/VMID/mounts) → **slice 10**;
|
||
this job takes the spec as input. `GuestMount` is defined minimally (no hub coupling).
|
||
|
||
### Tests
|
||
- provision happy path (fresh MAC = net0 without hwaddr, hostname, coalesced sizing+mount, rootfs
|
||
grow separate, started, **guest NOT destroyed**); compensating rollback at each step (restore /
|
||
config / start-task / waitRunning — asserts the guest WAS destroyed); DR continuity (MAC kept,
|
||
hostname not reset) + DR `KeepMAC=false` resets MAC; liveness verdict (warnings+running pass /
|
||
not-running fail); F4 (lock-500→retry→proceed; non-lock-500→fail without retry); owning entry
|
||
journaled BEFORE restore; reserved/existing VMID refused; `Recover` rolls back / clean.
|
||
|
||
### Live-validated (demo-felhom)
|
||
- provision: fresh MAC + hostname; **SSH host keys regenerated by the baked golden unit** (agent
|
||
issued no `ssh-keygen`), machine-id unique, Docker runs, clean DHCP lease → torn down.
|
||
- dr: continuity preserved (hostname + host keys kept). Recover: a killed mid-restore left an
|
||
orphan; the re-run's `Recover` rolled it back (idempotent).
|
||
- **Live caught a bug, then fixed:** the host-key unit's `ExecStart` was `/usr/sbin/ssh-keygen`
|
||
(203/EXEC); on Debian 13 it is `/usr/bin/ssh-keygen` — corrected in `build-golden.sh`, golden
|
||
rebuilt, re-validated. (Mocked unit tests couldn't surface this; the live run did.)
|
||
|
||
## v0.7.0 — restore-test: verdict is liveness, not start-task exitstatus (2026-06-09)
|
||
|
||
Fixes a correctness bug found by the live hub-enrollment runbook: the self-restore-test reported
|
||
`pass:false` on **every** modern-distro guest. PVE's guest-start task exits `"WARNINGS: 1"` for the
|
||
benign systemd-nesting advisory (`WARN: Systemd 257 detected. You may need to enable nesting.`), and
|
||
`WaitTask` treated any non-`"OK"` exitstatus as a hard failure — so the verdict was decided by an
|
||
advisory exit code instead of by observed liveness, *before* the real boot check ran. A crying-wolf
|
||
test got it disabled on the demo host; this re-enables it. **Single bump (0.6.0→0.7.0) covering the
|
||
agent's part of both task phases**; the wire fields below are consumed by hub from **v0.7.5**.
|
||
|
||
Design invariant (in code): **warning classification affects *visibility only*; pass/fail is
|
||
liveness-only.** A wrong/stale recognizer can at worst over-notice a benign warning — it can never
|
||
false-fail and never hide a real warning.
|
||
|
||
### Added
|
||
- **`proxmox.WaitOptions.AllowWarnings`** — opt-in per call. When set, a task that completes
|
||
`"WARNINGS: N"` is success with the `TaskStatus` (ExitStatus intact) returned so the caller can
|
||
read/surface it. Default (`false`) keeps **every existing caller strict** (vzdump/restore/destroy
|
||
warnings can be meaningful — relaxing them is a future per-call decision with evidence). Any
|
||
non-WARNINGS non-OK exit is still a `*TaskError`.
|
||
- **`reconcile.RestoreTestResult.StartWarnings` / `.WarningsRecognized`** + a version-free recognizer
|
||
(`benignWarningAnchor = "enable nesting"`, case-insensitive substring — contains no systemd version
|
||
number, so it can't rot back into the bug at systemd 258+). `extractWarningLines` pulls `WARN…`
|
||
lines from the start-task log.
|
||
- **`reconcile.GuestAPI.TaskLogTail`** — the engine fetches the start task's log to surface warnings.
|
||
- **`hub.RestoreTest.warnings` / `.warnings_recognized`** wire fields (`omitempty`), populated by
|
||
`ToHubRestoreTest`. Additive: the deployed v0.7.4 hub ignores them; hub v0.7.5 consumes them
|
||
(passed-with-warnings INFO, or WARN when not recognized). Cross-repo golden updated with the hub side.
|
||
|
||
### Changed
|
||
- **Restore-test start step** (`reconcile/restoretest.go`) now waits with `AllowWarnings:true`,
|
||
surfaces any start warnings, and **continues to `waitRunning` as the verdict** — boot+running is the
|
||
pass, exactly as before; a real (non-WARNINGS) start-task error still fails. The restore and
|
||
scratch-teardown WaitTasks stay strict.
|
||
- **Restore-test scheduler logging** distinguishes a clean pass, *passed-with-recognized-warnings*
|
||
(INFO), and *passed-with-unrecognized-warnings* (WARN) — nothing silent.
|
||
|
||
### Tests
|
||
- `WaitTask`: AllowWarnings accepts `WARNINGS` (status returned intact); AllowWarnings still fails a
|
||
real error; default still fails on `WARNINGS` (existing callers unaffected).
|
||
- Restore-test (engine, mock proxmox): start-with-warnings + running → **pass** with warnings
|
||
surfaced+recognized; unrecognized warning + running → pass, not-recognized; **not-running → fail
|
||
regardless of warnings** (verdict is liveness); teardown still runs.
|
||
- **Regression guard:** the `"enable nesting"` recognizer matches the advisory for systemd 256–300,
|
||
proving it's version-independent and can't silently rot back into the false-fail.
|
||
|
||
## v0.6.0 — slice 6 Phase B: PBS offsite tier (verify + PBS-API client + reporting) (2026-06-09)
|
||
|
||
Completes slice 6. The PBS spike (felhom.eu phase5-pbs-spike-findings.md) proved backup-to-PBS
|
||
and restore-from-PBS reuse Phase A UNCHANGED (PBS is just a storage target + a volid), and the
|
||
operator token needs no widening. So the only new agent code is the **verify capability + a
|
||
small PBS-API client + PBSSnapshot reporting**. Escrow + host-loss DR stay slices 7/10.
|
||
|
||
### Added
|
||
- **`internal/pbs` — the PBS-API client** (the agent's SECOND privileged external surface,
|
||
slice-1 discipline): TLS **fingerprint-pinned** to the PBS leaf cert (a spoofed PBS →
|
||
rejected, mirroring the PVE pin), **token auth** (`PBSAPIToken=<id>:<secret>`; id from the
|
||
storage `username`, secret read at runtime from `/etc/pve/priv/storage/<id>.pw` — referenced
|
||
by location, never logged/committed), typed, no shell. Methods: `Verify` (POST
|
||
`/admin/datastore/<ds>/verify` → UPID), `Snapshots` (incl. the `verification` field),
|
||
`TaskStatus`/`WaitVerify` (node extracted from the UPID — `localhost` returns "unknown", the
|
||
spike B4 gotcha), `NodeFromUPID`.
|
||
- **The verify maintenance loop** (`pbs/verify.go`) — the cheap, key-free, ciphertext-level
|
||
integrity check (§8) on its OWN cadence (default 6h, the 5th daemon goroutine). It is a
|
||
reporting/maintenance task like the slice-5 watchdog: it does NOT go through the reconcile
|
||
gate/journal. Each cycle: trigger verify → poll task → re-list snapshots → record
|
||
per-snapshot `verify_state`. A failed verify is logged loudly.
|
||
- **`PBSSnapshot` reporting** — filled the stub (`namespace`/`backup_type`/`backup_id`/
|
||
`backup_time`(RFC3339)/`size_bytes`/`owner`/`protected`/`encrypted` (from `files[].crypt-mode`)
|
||
/`verify_state` (ok|failed|**none** until verified)/`verify_upid`). New `PBSReporter`
|
||
collector seam + an in-memory `SnapshotStore`. Cross-repo golden (both repos, byte-identical)
|
||
+ bidirectional key-set tests; hub `handler.go` parses `pbs_snapshots` and logs a **failed
|
||
verify `[WARN]`** (loudest offsite-DR signal).
|
||
- **Truthful backup mode** (`backup/runner.go`) — `Backup.mode` now reflects the ACTUAL vzdump
|
||
mode read from the task log (`backup mode: <x>`), since PVE may downgrade snapshot→stop for a
|
||
stopped guest (spike B1); falls back to the requested mode if unparseable.
|
||
- **proxmox**: `Storage.Username` (parsed from the pbs storage config — the token id).
|
||
- **config** `BackupConfig.{PBSVerifyCadenceSeconds, PBSSecretDir}` (cadence 0→6h, <0 disabled).
|
||
- **`--selftest=pbs-verify`** — discover pbs storages → verify each → print the PBSSnapshot
|
||
records (covers the runbook's verify + list). Standalone on the host.
|
||
|
||
### Notes
|
||
- Backup/restore-to-PBS reuse Phase A with no change (the restore-test runs with
|
||
`source_tier="pbs"` when fed a pbs volid). Zero-knowledge holds: verify is ciphertext-level,
|
||
the encryption key is never read here, and the PBS server has no client key (spike B6).
|
||
- Daemon runs cleanly with no pbs storage / verify disabled. `go test -race` covers the new
|
||
goroutine. Slice-3/4/5/6A surfaces, goldens, and adversarial tests intact.
|
||
|
||
## v0.6.0-rc1 — slice 6 Phase A: backup + the self-restore-test (local target) (2026-06-09)
|
||
|
||
Phase A of the backup/restore slice (doc 03 §8) — the agent's guest-level backup layer and
|
||
the **self-restore-test**, which closes "a backup you haven't restored isn't a backup".
|
||
Everything here is BENIGN (backup, restore-to-NEW, scratch teardown): reuses the slice-4
|
||
classifier/gate/journal — no new destructive class, no new crypto. Local target only; PBS is
|
||
Phase B. Restore is to a NEW guest only (no overwrite). Backups are crash-consistent only
|
||
(app-consistency needs the controller quiesce, slice 8) — marked so in the report.
|
||
|
||
### Added
|
||
- **proxmox** (`mutate.go`/`query.go`): `DestroyLXC` (DELETE …/lxc/{vmid}?purge=1&destroy-
|
||
unreferenced-disks=1 → UPID; the scratch-teardown primitive); `VzdumpOptions.Notes` →
|
||
`notes-template` (verified on PVE 9.2.2); `LatestBackupVolID` (resolve a produced archive
|
||
from the backup-storage listing — the task status carries no result volid).
|
||
- **reconcile self-restore-test** (`restoretest.go`) — `Engine.RunRestoreTest`: pick a free
|
||
scratch VMID (configured band, excludes 9999; full band → skip, never out-of-band) →
|
||
**journal a Scratch-owned entry BEFORE any mutation** → restore-to-new → benign net
|
||
**link-down** SetConfig (so the clone can't conflict with a running source's MAC/IP; this
|
||
is test-safety, NOT slice-7 identity reset) → boot → verify **reaches `running`** → ALWAYS
|
||
teardown (defer; benign `ClassGuestDestroy` + agent-tagged-scratch provenance, gated). Runs
|
||
on the scratch VMID's queue lane. Reuses the journal/gate; result feeds the report.
|
||
- **Crash-safe recovery** (`recover.go`): a Scratch journal entry is resolved by TEARDOWN,
|
||
not by re-checking the restore sub-task's UPID — special-cased BEFORE the generic path
|
||
(else the restore task's OK would mark it succeeded while the guest leaks). `Recover` now
|
||
destroys a leaked scratch guest (idempotent: already-gone → clean; list-unreadable → left
|
||
in-flight for a later pass). `JournalEntry.Scratch` flag; `RecoverResult.ScratchClean/
|
||
ScratchDestroyed`. GuestAPI gains `RestoreLXC`/`DestroyLXC`/`GuestStatus`.
|
||
- **`internal/backup` package**: `BackupRunner.Backup` (vzdump + archive/size resolve +
|
||
bulk-volume gap — a mountpoint is UNCOVERED unless it carries an explicit `backup=1`, so
|
||
an unset `backup=` is reported uncovered too, the safe DR direction); `PickRestoreCandidate`
|
||
(newest backup); an in-memory `Store` (latest-backup-per-target + latest-restore-test)
|
||
implementing the hub `BackupReporter`/`RestoreTestReporter` seams; a cadence `Scheduler`
|
||
(default 24h; the fourth daemon goroutine; disabled cleanly when off/misconfigured).
|
||
- **hub report** (`report.go`): filled the `Backup` + `RestoreTest` stubs (`PBSSnapshot`
|
||
stays a Phase-B stub); collector `BackupReporter`/`RestoreTestReporter` seams. Cross-repo
|
||
golden updated in BOTH repos (byte-identical) + bidirectional key-set tests for
|
||
`backups[0]`/`restore_tests[0]`. Hub `handler.go` parses + persists them (report_json; no
|
||
new columns) and logs a **FAILED restore-test prominently** (the loudest DR signal).
|
||
- **config** `BackupConfig` (local target, restore storage, restore-test cadence, scratch
|
||
VMID band 990000–990009 default) + accessors + env overlay + cadence-gated validation.
|
||
- **`--selftest=backup -vmid N`** (one-shot backup → print the Backup record) and
|
||
**`--selftest=restore-test [-archive volid]`** (Recover-then restore→boot→verify→teardown,
|
||
print the RestoreTest record). Standalone on the Proxmox host.
|
||
|
||
### Notes
|
||
- The daemon runs cleanly with the cadence off or misconfigured (logs + disables, never
|
||
crashes); a leaked scratch guest from a mid-test crash is reaped by `engine.Recover` on
|
||
restart. `go test -race` covers the new scheduler goroutine.
|
||
- Slice-3/4/5 exported surfaces, goldens, and adversarial tests intact. Version bumps to
|
||
**v0.6.0** when Phase B (PBS) lands.
|
||
|
||
## v0.5.1 — slice 5 live-validation prep: durable_id mis-id fix + re-mount UUID memory (2026-06-09)
|
||
|
||
Two correctness fixes surfaced while preparing the live USB validation on `demo-felhom`
|
||
(a real 1TB USB HDD, sdb1, ext4). Both are DR-load-bearing — exactly the "false-id →
|
||
re-attach the wrong disk" failure mode the slice warned about.
|
||
|
||
### Fixed
|
||
- **Unmounted dir-storage no longer inherits the ROOT filesystem's UUID** (`observe.go`).
|
||
Previously, when a removable dir-storage was unmounted, the observer fell through to the
|
||
*containing* mount (root) for the backing device, so its `durable_id` became
|
||
`uuid:<root-uuid>` — a catastrophic DR mis-id (the hub would re-attach the wrong disk).
|
||
Now the backing device/UUID/`durable_id` are derived ONLY from the target's OWN
|
||
mountpoint; an unmounted target reports no device and a stable `store:<name>` durable_id,
|
||
never another filesystem's UUID. (Removed the `containingMountDevice` root-fallthrough.)
|
||
- **Watchdog remembers the fs-UUID observed while attached** (`watchdog.go`) so a re-mount
|
||
works even after the known-set cache refreshes mid-drop (an unmounted target can't resolve
|
||
its own UUID). The re-mount key is backfilled from this memory — aligning with doc 03 §7's
|
||
"sourced from the existing definition, no hub manifest needed": the agent learns the UUID
|
||
while the target is attached, then re-mounts by it on return.
|
||
|
||
### Tests
|
||
- Observer: an unmounted dir-storage asserts NO `uuid:` durable_id and no backing device.
|
||
- Watchdog: a drop where the cache lost the UUID still re-mounts using the remembered UUID.
|
||
|
||
## v0.5.0 — slice 5 Phase B: the host-root surface (mounts + SMART + grow + destructive gate) (2026-06-09)
|
||
|
||
The write surface — the agent's first step outside its Proxmox API token into OS-root.
|
||
Isolated behind a narrow, argument-validated, adversarially-tested seam, exactly like the
|
||
slice-4 gate. Completes slice 5 (Phase A = read-only observe/report/watchdog at v0.5.0-rc1).
|
||
|
||
### Added
|
||
- **`HostOps` seam + `SudoHostOps`** (`internal/storage/hostops.go`) — the one privileged
|
||
host surface: persistent mounts via **systemd `.mount` units keyed by fs-UUID** (enabled to
|
||
survive reboot), detach (stop+disable), SMART, and thin-pool metadata. Shells out via the
|
||
fenced Runner (`sudo -n`, **fixed arg vectors, no shell**); a fake backs the tests (no real
|
||
root in the suite). `NoopHostOps` is the safe fallback when the surface is unavailable.
|
||
- **The argument validator** (`internal/storage/validate.go`) — the security boundary:
|
||
`ValidateUUID` (strict hex), `ValidateMountPath` (absolute, no traversal, no metacharacters),
|
||
`ValidateSMARTDevice` (raw-disk whitelist), `ValidateLVMName`, and an in-process
|
||
`systemdEscapePath` (no `systemd-escape` shell-out). **Every argument is validated BEFORE a
|
||
command is constructed.** Headline test (`validate_test.go`): an adversarial matrix of
|
||
shell metacharacters / `../` traversal / malformed inputs is rejected with **zero exec**.
|
||
- **SMART** (`internal/storage/smart.go`) — parses `smartctl -a -j` into `StorageTarget.smart`:
|
||
**SATA** (reallocated/pending/offline-uncorrectable, temp, power-on-hours) **and NVMe**
|
||
(critical_warning, media_errors, percentage_used, temp), degrading to `UNKNOWN` for devices
|
||
with no SMART (USB-SATA bridges). **`lvs`** fills the lvmthin thin-pool **metadata** fill
|
||
(the value Phase A left null). Wired into the Observer's enrichment (Observe only, not the
|
||
watchdog's fast Known path).
|
||
- **Watchdog re-mount response** (`internal/storage/watchdog.go`) — on a known mount-backed
|
||
target's device returning **unmounted** (a new `DevicePresent` liveness probe), the watchdog
|
||
**dispatches a benign by-UUID re-mount off the poll path** (a goroutine, never under the
|
||
lock), rate-limited per target to the debounce window. The mount is routed through the gate
|
||
as benign (`gateRemounter` in `main.go`, so `storage` stays decoupled from `reconcile`).
|
||
- **Disk-grow executor** (`internal/reconcile`) — `ActionResize` (benign `ClassResize`), planned
|
||
**grow-only** (desired DiskBytes > actual → `pct resize rootfs +<n>M`; a shrink is refused,
|
||
never silently grown) + a defensive executor guard (size must start with `+`). New
|
||
`proxmox.Client.ResizeLXC` (API; `VM.Config.Disk`+`Datastore.AllocateSpace`; async→UPID).
|
||
Built + fixture-tested; **unfed** live (no hub spec until slice 10).
|
||
- **Destructive storage ops through the slice-4 gate** (`internal/reconcile/storage_ops.go`) —
|
||
`IntentForStorageMount` (benign) and `IntentForStorageDestructive` (`ClassStorageWipe`/
|
||
`ClassDecommission`). Host/target-scoped: the op binds on the storage **target identity**
|
||
(carried in `target.guest_id`). Reuses the existing verifier/role-scoping/binding/audit — no
|
||
new gate, no new crypto. Storage cases added to the adversarial matrix (`storage_test.go`):
|
||
unsigned wipe → `pending_signature`; "wipe A" signature vs "wipe B" → `binding_mismatch`;
|
||
valid → accepted. **Inert** live.
|
||
- **`--selftest=storage` [`-watch <dur>`]** — the live USB-runbook harness: an observe pass
|
||
(full table incl. SMART + thin-pool data+metadata), and a bounded watchdog window with the
|
||
re-mount response live. Runs standalone on the Proxmox host (no hub).
|
||
- **`configs/felhom-agent.sudoers`** — the documented narrow allowlist (install unit / systemctl
|
||
manage / smartctl / lvs), with the agent-side fine validation noted.
|
||
- **Config**: `privileged.{unit_dir,stage_dir,systemctl,install,smartctl,lvs}` (paths must match
|
||
the sudoers entries).
|
||
|
||
### Notes
|
||
- Daemon still runs cleanly with no removable storage / no signers / no hub manifest, and a
|
||
missing/declined sudoers entry degrades with a warning (SMART→UNKNOWN, mount→logged error),
|
||
not a crash. `go test -race` passes (the watchdog re-mount dispatches off the poll path).
|
||
- Slice-3/4 + Phase-A exported surfaces, goldens, and adversarial tests intact. `authz`
|
||
untouched. The destructive-storage executor + grow are built/tested but unfed live until
|
||
slice 10.
|
||
|
||
## v0.5.0-rc1 — slice 5 Phase A: storage observe + report + watchdog (read-only, live) (2026-06-09)
|
||
|
||
Phase A of the storage slice (doc 03 §7). Read-only and live: the agent now observes every
|
||
host storage target, reports it into the host-report's `storage_targets` (previously an empty
|
||
stub), and runs a fast-poll watchdog that pushes a disconnect to the hub in seconds. No
|
||
host-root writes this phase (mounts/SMART/grow/destructive-gate are Phase B). The hub-owned
|
||
desired manifest (class/role/policy/creds) is not served until slice 10, so reconcile against
|
||
it is built-but-unfed — this phase ships only the genuinely-useful read-only footprint.
|
||
|
||
### Added
|
||
- **`internal/storage` package** (new):
|
||
- **`StorageTarget` wire contract** (`internal/hub/report.go`) — filled the slice-3 stub:
|
||
`name`/`type`/`durable_id`/`state`/`reachable`, usage (`total`/`used`/`avail`/
|
||
`used_fraction`), `content`, `mount_path`/`backing_device`, `class_hint` (rotational HINT
|
||
— never authoritative; class is hub-owned), `role` (empty until slice 10), a `thin_pool`
|
||
sub-object (lvmthin data fill; metadata fill is Phase B/`lvs`), and a `smart` sub-object
|
||
(`UNKNOWN` until Phase B). Cross-repo golden kept byte-identical with `felhom.eu/hub` and
|
||
guarded by the bidirectional key-set test (`contract_test.go`).
|
||
- **`durable_id` derivation** (`durableid.go`) — deterministic per type (the DR-load-bearing
|
||
re-attach key): fs-UUID (usb/local-dir), `server:export` (nfs/cifs), `repo+fingerprint`
|
||
(pbs), `vg/pool` (lvmthin); never empty (falls back to a stable store id).
|
||
- **`HostReader` seam + `ProcHostReader`** (`hostread.go`) — non-privileged `/proc/mounts`,
|
||
`/dev/disk/by-uuid`, `/sys/.../rotational` + `removable` reads. Root-free by construction.
|
||
- **`Observer`** (`observe.go`) — builds `[]hub.StorageTarget` from `ListStorage`/`NodeStorage`
|
||
joined with host reads; surfaces the lvmthin thin-pool data fill prominently (warns ≥85%).
|
||
- **Storage watchdog** (`watchdog.go`) — a third daemon goroutine fast-polling the *known*
|
||
target set (a defined Proxmox storage and/or a previously-seen one) for
|
||
`attached↔disconnected` transitions; on a transition it triggers an immediate, **debounced**
|
||
out-of-band host-report. Only flags a *known* target's change (never a never-attached
|
||
device); coalesces flaps within the debounce window (leading + trailing edge).
|
||
`CachingKnownTargets` rate-limits the Proxmox-derived known set; `HostLiveness` probes
|
||
device/mount presence (local) + a reachability dial (network), all non-privileged.
|
||
- **Proxmox `Storage` type** (`internal/proxmox/types.go`) — additive parse-only config fields
|
||
(`server`/`export`/`share`/`datastore`/`fingerprint`/`vgname`/`thinpool`) feeding durable_id.
|
||
- **Collector `StorageObserver` seam** (`internal/hub/collect.go`) — populates `storage_targets`
|
||
via the observer; a nil observer or an observe error degrades to empty (never sinks the
|
||
heartbeat). Hub does not import storage (storage imports hub for the wire type).
|
||
- **Out-of-band report trigger** (`internal/hub/loop.go`) — `Loop.SetTrigger`: a watchdog
|
||
signal runs one extra collect→report immediately without disturbing the regular cadence.
|
||
- **`StorageConfig`** (`internal/config`) — watchdog interval / debounce / known-refresh knobs
|
||
(all optional; package defaults otherwise).
|
||
- **Hub ingest** (`felhom.eu/hub`) — `hostReportPayload` now parses `storage_targets`
|
||
(full mirror struct), persists them via `report_json`, counts + warns on disconnected
|
||
targets, and has its own half of the bidirectional golden key-set test.
|
||
|
||
### Notes
|
||
- The daemon still runs cleanly with no removable storage, no signers, and no hub manifest —
|
||
the watchdog finds nothing to flag; storage reporting is best-effort.
|
||
- `proxmox`/`hub`/`authz`/`reconcile` exported surfaces + their golden/adversarial tests are
|
||
intact. No host-root writes, no destructive paths, no SMART this phase (all Phase B).
|
||
- Version: **v0.5.0-rc1** at the Phase-A checkpoint; **v0.5.0** when Phase B lands.
|
||
|
||
## v0.4.0 — slice 4 Phase B: reversibility gate + signed-op consuming layer (2026-06-08)
|
||
|
||
The security core of slice 4: hub-supplied intent stops being trusted for destructive
|
||
change. Layered in front of the per-guest queue's executor — **every** mutation now
|
||
passes the gate. Reuses `internal/authz` for all crypto (untouched surface). Inert
|
||
this slice: no destructive deltas are served until slice 10, so the destructive path is
|
||
classified, gated, and adversarially tested but not wired to live execution.
|
||
|
||
### Added
|
||
- **Classifier (`classify.go`, doc 03 §4)** — benign vs destructive by **provenance +
|
||
data-bearing-ness, NOT by verb**. The `OpClass` vocabulary (seeded by the committed
|
||
slice-2 `op_blob.json`: `guest_destroy`) is the agent-side contract slice 10 matches.
|
||
Destroy/overwrite of customer data is destructive UNLESS **agent-internal**
|
||
provenance (same-journaled-transaction create → compensating rollback, or
|
||
agent-tagged scratch) makes it benign. `Provenance` is journal-recorded and **never
|
||
populated from the hub** (its zero value is the only thing an external intent may
|
||
carry). Unknown op class fails safe → destructive.
|
||
- **Reversibility gate (`gate.go`)** — `Gate.Authorize(intent, signed)`: benign →
|
||
allowed unsigned; destructive → requires a verified, role-authorized, action-bound
|
||
operator signature, else refused **`pending_signature`**, never executed. Every
|
||
decision is written to an `AuditSink` (audit is a signal, never the guard).
|
||
- **Signed-op consuming layer over `authz`** — verifies via `authz.Verifier.Verify`
|
||
(the locked pipeline, untouched), then enforces on the `VerifiedOp`:
|
||
- **Role-scoping (doc 04 §4)** — recovery key authorizes key-rotation re-pins ONLY;
|
||
operational key authorizes ordinary destructive ops + planned rotation.
|
||
- **Op-to-action binding** — verified `op` + host + guest + `params` must match the
|
||
gated action (a signature for guest X / op A can't authorize guest Y / op B);
|
||
params compared semantically (key-order/whitespace independent).
|
||
- **Signed-job orchestration (`job.go`)** — `RunSignedJob`: idempotency dedupe (the
|
||
op nonce as the journal key — a redelivered completed op is skipped, not re-run),
|
||
gate authorization, then journal-wrapped execution via an injected
|
||
`DestructiveExecutor` (nil this slice — authorized destructive ops are inert, no
|
||
executor wired until 6/7).
|
||
- **Crash-recovery consumer (`recover.go`, Note 1 / doc 03 §10)** — `Engine.Recover`
|
||
consumes the journal's `InFlight()` at startup: an op that crashed AFTER the Proxmox
|
||
POST and BEFORE its terminal record (`OpTaskRunning`, nonce already consumed) is NOT
|
||
covered by idempotency dedupe — only this resume-or-rollback resolves it (re-read the
|
||
task via the new `TaskStatusOnce`, record the real outcome; a no-task-id op is
|
||
abandoned fail-safe). Landed together with the signed-op executor, as Note 1 required.
|
||
- **Daemon wiring** — `runDaemon` builds the verifier from `config.Authz.Signers` (a
|
||
bad key / missing nonce-store path is a fatal misconfig; **no signers = nil verifier**,
|
||
the common slice-4 state), constructs the gate (+ `SlogAudit`), runs `Recover` before
|
||
issuing any mutation, and routes every reconcile action through the gate.
|
||
|
||
### Changed
|
||
- **Memory comparison canonicalized (Note 2)** — `desiredMemoryMiB` makes the
|
||
desired↔actual memory compare in the same MiB unit that is then written, so a
|
||
non-MiB-aligned `MemoryBytes` converges in one pass instead of re-issuing SetConfig
|
||
forever (the numeric cousin of the description-newline normalization). Test proves
|
||
convergence. Slice 10 should still serve MiB-aligned specs at the source.
|
||
|
||
### Tests (the security proof — each independently rejected)
|
||
- **Adversarial matrix** via the REAL `authz.Verifier` with in-test-minted SSHSIGs
|
||
(framing replicated in reconcile's test binary; production authz untouched, no signing
|
||
added to the verify-only package): unsigned destructive **job** → pending_signature;
|
||
unsigned destructive **desired-state delta** → pending_signature (distrusts hub
|
||
desired state, not just jobs); forged/unknown signer → `ErrUnknownSigner`; expired →
|
||
`ErrExpired`; **replayed nonce across an agent restart** (durable `FileNonceStore`) →
|
||
`ErrReplay`; wrong host → `ErrTarget`; wrong guest / wrong op / wrong params →
|
||
binding_mismatch; **recovery key on ordinary destructive** → role_denied;
|
||
**hub-supplied "scratch" tag ignored** → still destructive → refused; **valid + role +
|
||
target + fresh nonce → accepted**, and a second presentation → `ErrReplay` (nonce
|
||
consumed).
|
||
- Classifier (benign/destructive/provenance/key-rotation/fail-safe), role-scoping,
|
||
params binding, crash-recovery (resume OK / fail / still-running / no-task rollback /
|
||
unreadable / one-shot key applied on resume), signed-job idempotency (execute once,
|
||
dedupe redelivery, refused-not-executed, no-executor-inert, executor-error).
|
||
- Full module **race-clean** (`go test -race`) + vet clean on the Linux build server.
|
||
|
||
## v0.4.0-rc1 — slice 4 Phase A: reconcile engine (structural; runs live, unfed) (2026-06-08)
|
||
|
||
The agent-side control core's structural half. **Checkpoint marker** — `-rc1` is the
|
||
Phase-A push; awaiting validation before Phase B (the reversibility gate + signed-op
|
||
consuming layer) lands the final **v0.4.0**. Runs LIVE but UNFED: with no desired-state
|
||
provider until slice 10, the live engine computes an empty action set and performs
|
||
**zero mutations**.
|
||
|
||
### Added
|
||
- **`internal/reconcile`** package — the engine, the per-guest serializer, the
|
||
desired-state model, the normalization layer, and the durable op journal:
|
||
- **Per-guest serializer (`Queue`, doc 03 §10)** — the single choke point ALL
|
||
mutation sources funnel through. Same-vmid jobs run strictly one-at-a-time in
|
||
submit order; independent vmids run in parallel. Each vmid is a cond-var FIFO lane
|
||
(unbounded, non-blocking, order-preserving); graceful drain on `Close`.
|
||
- **Desired-state model + `DesiredProvider` seam** — `DesiredGuest` (per-field
|
||
optional: run-state / `*hub.GuestSpec` / `*description`), `DesiredState`. The only
|
||
live provider is **`EmptyProvider`** (slice 4 has no source); `StaticProvider`
|
||
feeds fixtures. The seam is where slice 10's hub-serving plugs in — no hub/local
|
||
source invented here.
|
||
- **Normalization layer (`FieldNormalizers`)** — reconcile compares *normalized*
|
||
desired-vs-actual so Proxmox round-trip quirks don't read as drift. `description`'s
|
||
trailing newline is the first registered case; the registry takes more (boolean
|
||
coercion, list ordering) as discovered. `normDesc` **promoted** out of
|
||
`cmd/felhom-agent/main.go` to **`reconcile.NormDescription`**; the `--selftest=task`
|
||
description round-trip now uses that shared helper (one source of truth for the quirk).
|
||
- **Plan engine (`Plan`, pure function)** — computes the minimal **benign** action set
|
||
(`Start`/`Stop`/`SetConfig`) for guests present in both desired and actual, with
|
||
normalized comparison, deterministic vmid ordering, config-before-run-state. Skips
|
||
provision (desired-absent-in-actual, slice 7) and destroy (actual-absent-in-desired,
|
||
gated, slice 10); never writes a config it couldn't first read (`SpecKnown`). Disk
|
||
(rootfs grow) intentionally not reconciled here.
|
||
- **Reconcile engine (`Engine`)** — reads desired+actual, plans, dispatches each action
|
||
onto the shared queue. Every Proxmox op handled per the mutate.go contract: non-empty
|
||
UPID → `WaitTask` + assert `exitstatus`; empty UPID → clean **synchronous** success
|
||
(slice-4 proven). Per-action failures are counted, not fatal (other guests still
|
||
converge).
|
||
- **Operation journal (`Journal`)** — durable fsync'd append-only JSONL mirroring
|
||
`authz.FileNonceStore`: records each op's lifecycle (started → task_running →
|
||
succeeded/failed) with its Proxmox task id (crash mid-op is detected and re-checkable
|
||
on restart via `InFlight()`), plus an **idempotency-key store** (`AlreadyApplied`) so
|
||
a one-shot op never re-runs across retries/restarts. Reconcile actions carry no
|
||
idempotency key (convergent — must re-run on real drift).
|
||
- **Daemon wiring (`runDaemon`)** — reconcile runs alongside the hub loop on the poll
|
||
cadence, **sharing the per-guest queue**. Journal path is a `journal.log` sibling of the
|
||
nonce store. The daemon runs cleanly with **no desired state and no signers** (reconcile
|
||
is a logged live no-op; a journal-open failure degrades to journal-less, never crashes).
|
||
|
||
### Tests
|
||
- Serializer: same-guest serialized (max-concurrency 1, submit order preserved) and
|
||
different-guests parallel (cross-waiting jobs both complete — would deadlock if not);
|
||
error propagation; drain-pending-on-close; submit-after-close.
|
||
- Normalization: description round-trip; unknown-field identity; extensibility seam
|
||
(synthetic boolean-coercion + list-ordering normalizers).
|
||
- Plan: run-state start/stop, spec drift (cores/memory), disk-not-reconciled,
|
||
description-newline-not-drift, unmanaged fields, spec-unknown skips config keeps
|
||
run-state, desired-absent skipped, combined ordering, empty-desired no-op, deterministic
|
||
vmid order.
|
||
- Engine: empty-provider zero mutations; async start (WaitTask); synchronous SetConfig
|
||
(no WaitTask); WaitTask failure + POST error counted failed; list error = pass failure.
|
||
- Journal: lifecycle latest-wins; in-flight survives restart; idempotency dedupe across
|
||
restart; failed key not applied; torn-trailing-line skipped.
|
||
- Full module **race-clean** (`go test -race`) on the Linux build server; vet clean.
|
||
|
||
### Not in this phase (Phase B)
|
||
- The benign/destructive classifier, the reversibility gate, and the signed-op consuming
|
||
layer over `internal/authz` (doc 03 §4 / doc 04) — added next, in front of the queue's
|
||
executor, landing **v0.4.0**.
|
||
|
||
## v0.3.2 — SetConfig selftest extension (slice-4 pre-check) (2026-06-08)
|
||
|
||
The gate before slice 4: prove `SetConfig` works live under the scoped token before
|
||
reconcile is built on it. **Self-gated live run PASSED** on `demo-felhom`/guest 9999.
|
||
|
||
### Added
|
||
- **Reversible `SetConfig` step appended to `--selftest=task`** (`cmd/felhom-agent/main.go`,
|
||
`selftestSetConfig`): read `GuestConfig` → write a `description` marker
|
||
(`felhom-selftest <RFC3339>`) → verify it landed → restore the original value (or
|
||
`delete` the key if it was absent) → verify the restore. Handles PVE's dual-mode
|
||
`SetConfig` return per the `mutate.go` contract: empty UPID = synchronous success
|
||
(printed `synchronous`); non-empty UPID = `WaitTask` + assert `exitstatus=OK`.
|
||
The existing snapshot → rollback → delete-snapshot steps are unchanged. First live
|
||
exercise of the **`VM.Config.*`** privilege cluster.
|
||
- **`normDesc` / `extraString` helpers** — `extraString` decodes a string-valued key
|
||
from `GuestConfig.Extra` (raw JSON); `normDesc` strips the trailing newline PVE
|
||
appends to `description` on read, so a written value round-trips equal.
|
||
|
||
### Finding (live)
|
||
- The LXC `description` write returned **synchronous (empty UPID)** — PVE applied it
|
||
inline, no task. The agent's dual-mode `SetConfig` modeling is correct: the
|
||
empty-string path is real and must not be treated as an error.
|
||
- PVE **appends a trailing `\n` to `description`** on read (stored URL-encoded as
|
||
`%0A`). A naive exact-match reconcile would see perpetual drift — slice-4 reconcile
|
||
must normalize `description` comparisons (hence `normDesc`).
|
||
|
||
### Ops
|
||
- Standing operator token (`felhom-agent@pve!agent`, privsep) **rotated** during this
|
||
run (the prior secret was not retrievable); role + both user/token ACL rows
|
||
re-confirmed at `/`. New secret stored out-of-band, **not persisted to the repo**.
|
||
Guest 9999 left pristine (stopped, no `description`, no leftover snapshot). Version → 0.3.2.
|
||
|
||
## Docs + live validation — no version bump (2026-06-08)
|
||
|
||
### Changed
|
||
- **Reflowed `CLAUDE.md`** — removed hard mid-paragraph line wraps (prose, list items, blockquotes now single-line, soft-wrapped); code blocks and tables untouched; rendered output unchanged.
|
||
- **Unified the REPORT/CHANGELOG convention** in `CLAUDE.md`: `CHANGELOG.md` is the cumulative log (newest on top); `REPORT.md` is overwritten with the most-recent implementation/validation only. Added an explicit **no-secrets** rule (never write tokens/passwords/keys into committed files; reference them as stored out-of-band).
|
||
|
||
### Added
|
||
- **`REPORT.md`** rewritten for the live `--selftest=task` validation on the demo host (`demo-felhom`): snapshot → rollback → delete-snapshot on guest 9999, each polled to `exitstatus=OK` under the `felhom-agent@pve!agent` privsep token (UPIDs name the token actor — privsep path genuinely exercised); 16-privilege `FelhomAgent` role + both user & token ACLs confirmed; `--selftest=read` clean. Closes the slice-1 "mutating ops unit-tested only" gap; `WaitTask` async foundation validated live → **slice 4 unblocked**. (Token secret stored out-of-band, not in the repo.)
|
||
|
||
## v0.3.1 — slice-3 validation follow-ups (2026-06-08)
|
||
|
||
### Changed
|
||
- **Collector keeps the known run-status on a `GuestConfig` failure** (`internal/hub/collect.go`):
|
||
previously a per-guest config-read error forced `status="unknown"`; now the run-status from
|
||
`ListLXC` is preserved (only the `spec` is dropped). An empty status is still normalized to
|
||
`unknown` (wire value is always `running|stopped|unknown`). Test renamed to
|
||
`TestCollect_GuestConfigFailureKeepsStatusOmitsSpec` and asserts the preserved `running` + nil spec.
|
||
- **`--selftest` usage** error string now reads `(want read|task|hub)`.
|
||
|
||
### Added
|
||
- **Cross-repo contract fixture** `internal/hub/testdata/host-report.golden.json` +
|
||
`TestHostReport_ContractMatchesGolden` — compares the marshaled `HostReport` field-name sets
|
||
(top level + `host` + `guests[0]`) against the golden, failing on any json-tag drift. The file is
|
||
**kept byte-identical** with felhom-hub's copy (duplicated contract until a shared types module;
|
||
revisit when slices 5/6 populate the empty collections). Version → 0.3.1.
|
||
|
||
## v0.3.0 — hub client + host-report + first daemon loop (slice 3) (2026-06-08)
|
||
|
||
The agent's first daemon: a periodic read-only host-report POSTed to the hub (the
|
||
heartbeat). No Proxmox mutations, no desired-state/signed-op consumption, no
|
||
storage/backup collection yet — those are slices 4/5/6.
|
||
|
||
### Added
|
||
- **`internal/hub`** package:
|
||
- **`HostReport`** wire contract (`report.go`) shared field-for-field with the hub
|
||
ingest: host metrics, guests (`vmid` + spec), `cloudflared` status, and the
|
||
`storage_targets`/`backups`/`restore_tests`/`pbs_snapshots`/`audit_tail`
|
||
collections **defined but emitted empty** (typed `[]`, slices 5/6 fill them).
|
||
- **`Collector`** (`collect.go`) builds the report from a read-only `proxmoxReader`
|
||
(adapted to the real `internal/proxmox` surface — node held by the client, value
|
||
returns, `proxmox.Guest`) + a `CloudflaredProber`. Partial-failure policy: a
|
||
failed `NodeStatus` is a hard error (skip the POST); a failed per-guest
|
||
`GuestConfig` degrades that guest to `status="unknown"` (spec omitted) but still
|
||
sends; a cloudflared probe failure → `"unknown"`, never fatal.
|
||
- **`CloudflaredProber`** + `SystemctlProber` (`systemctl is-active cloudflared`;
|
||
read-only — NOT a Privileged/root op; tunnel management is a later slice).
|
||
- **`Client`** (`client.go`): `POST /api/v1/host-report` with
|
||
`Authorization: Bearer <key>`, standard TLS (system roots or optional `ca_file`;
|
||
verification always on). Typed `*TransportError` / `*HTTPError`; the bearer token
|
||
never appears in any error.
|
||
- **`Loop`** (`loop.go`): the daemon — immediate first report then tick; adopts the
|
||
hub's `poll_interval_seconds` clamped to [60,3600]; resilient (a collect/report
|
||
error is logged and the loop continues); clean shutdown on context cancel.
|
||
- **`ControlEnvelope`**: only `poll_interval_seconds` is acted on; `blocked` /
|
||
`desired_generation` / `has_signed_ops` are parsed-but-ignored (logged at most)
|
||
pending reconcile (slice 4).
|
||
- **Config**: `HubConfig` (url/host_id/api_key/poll_seconds/timeout_seconds/ca_file),
|
||
`FELHOM_AGENT_HUB_*` env overlay, `HubConfig.Validate()` (mode-aware — proxmox-only
|
||
`--selftest=read|task` still runs without hub config), `WithDefaults()`, and
|
||
`Redacted()` now also blanks the hub key. `configs/agent.example.json` gains `hub`
|
||
(and `authz`) blocks.
|
||
- **`cmd/felhom-agent`**: the no-`--selftest` mode is now the **daemon** (poll loop);
|
||
added **`--selftest=hub`** (one collect+report, prints the report + envelope).
|
||
Version 0.2.0 → 0.3.0.
|
||
|
||
### Tests
|
||
- Report serialization (field names; empty collections are `[]` not `null`; spec
|
||
omitted when unknown); client (Bearer header, non-2xx→`*HTTPError`,
|
||
transport→`*TransportError`, **token never in error**); collector (host mapping,
|
||
guest spec, per-guest failure degrades-but-still-reports, NodeStatus hard error,
|
||
cloudflared error→unknown); loop (immediate first report, continuation after an
|
||
injected error, interval adoption + clamp); config (hub validate/redact/env).
|
||
|
||
### Notes
|
||
- `internal/proxmox` and `internal/authz` were **not touched** — no new proxmox
|
||
surface was needed (`ListLXC` already exposes status/maxmem/maxdisk; `GuestConfig`
|
||
exposes cores). The task's `proxmoxReader` sketch (node-arg/pointer/`LXC`) was
|
||
adapted to the real exports as instructed.
|
||
- **Defined-but-empty** this slice: `storage_targets`, `backups`, `restore_tests`,
|
||
`pbs_snapshots`, `audit_tail` (slices 5/6). **Parsed-but-ignored**: the envelope's
|
||
`blocked`/`desired_generation`/`has_signed_ops` (slice 4).
|
||
|
||
## v0.2.0 — `authz` signed-op verifier (slice 2) (2026-06-08)
|
||
|
||
Production form of the Phase-4 signing primitive: a key-type-agnostic SSHSIG
|
||
verifier for operator-signed destructive ops, with the full anti-replay/
|
||
authorization pipeline and a durable, crash-safe nonce store. What slice 4
|
||
(reconcile) will call to gate destructive desired-state deltas. No hub, no signing
|
||
CLI, no reconcile loop.
|
||
|
||
### Added
|
||
- **`internal/authz` — `Verifier`**: `New(signers, store, hostID)` + `Verify(blob,
|
||
sigArmored) (*VerifiedOp, error)`. Runs the LOCKED pipeline (order is
|
||
load-bearing): parse armor → namespace → parse pubkey → allow-list (by key
|
||
**material**, `pub.Marshal()` equality, not key_id) → crypto verify (over the
|
||
**raw received bytes**, never re-canonicalized) → parse blob → target → time
|
||
window → **nonce recorded LAST**. Each post-crypto stage rejects even with a
|
||
valid signature.
|
||
- **SSHSIG framing** (`sshsig.go`) via `golang.org/x/crypto/ssh` — `pem.Decode` →
|
||
strip 6-byte magic → `ssh.Unmarshal` → `ssh.ParsePublicKey` → recompute signed
|
||
data with the named hash → `pub.Verify` (dispatches on key algorithm). No
|
||
hand-rolled crypto. Key-type-agnostic: ed25519 / **sk-ssh-ed25519 (FIDO2)** /
|
||
rsa / ecdsa via the one path.
|
||
- **Fixed namespace** `felhom-op-v1` (package constant, never caller-supplied).
|
||
- **`OpBlob`** (corrected `host_id`/`guest_id` json tags) + **`VerifiedOp`** (op,
|
||
host/guest, params, key_id, matched signer). key_id is advisory/audit only —
|
||
never an authz input.
|
||
- **Typed errors**: `ErrMalformed, ErrNamespace, ErrUnknownSigner, ErrBadSignature,
|
||
ErrTarget, ErrExpired, ErrNotYetValid, ErrReplay` (errors.Is-friendly).
|
||
- **`NonceStore`** + two impls: `MemoryNonceStore` (tests) and **`FileNonceStore`**
|
||
— durable, crash-safe (fsync'd append log, replayed into an index on open,
|
||
periodic compaction, expiry-only pruning). A nonce is fsync'd to disk before
|
||
`SeenOrRecord` returns false; replay protection survives restart; I/O failure
|
||
fails safe (reports seen=true). Target generalization: host_id matched strictly,
|
||
guest_id surfaced for the caller to route.
|
||
- **Config**: `AuthzConfig` (nonce-store path + pinned operator `signers` tagged
|
||
`operational`/`recovery` with a key_id, as authorized_keys lines).
|
||
- **Version 0.2.0.**
|
||
|
||
### Tests
|
||
- Real OpenSSH interop via a committed `ssh-keygen -Y sign` vector (hermetic CI);
|
||
per-stage rejection (each with an otherwise-valid sig); the headline
|
||
**invalid-sig-does-not-burn-the-nonce** invariant; replay; **persistence across
|
||
restart**; synthetic **sk-ssh-ed25519** through the unchanged path; byte-exactness
|
||
(a re-serialized blob fails crypto — not re-canonicalized).
|
||
|
||
### Notes / corrections to the Phase-4 reference
|
||
- §7's `Target` lacked json tags (`host_id`/`guest_id`) — fixed.
|
||
- The doc paired "Go 1.24.4 / x/crypto v0.52.0", but v0.52.0 declares `go 1.25.0`
|
||
and does **not** build on Go 1.24. Resolved by upgrading the build server to
|
||
go1.26.0 (backward-compatible; felhom-controller/hub unaffected); the module is
|
||
`go 1.25.0` on x/crypto v0.52.0.
|
||
- Free function → constructed `Verifier`; returns the full `VerifiedOp`; typed
|
||
errors; clock-skew tolerance added; durable nonce store is the net-new work.
|
||
- **Shared-contract dependency flagged** (not built): the hub and the `felhom-sign`
|
||
CLI must emit byte-identical canonical JSON or signatures won't verify; a shared
|
||
canonicalizer both import would be the right home.
|
||
|
||
## v0.1.0 — Scaffold + `proxmox` interaction layer (slice 1) (2026-06-08)
|
||
|
||
First slice: stand up the host-agent project and its foundation — the typed
|
||
Proxmox interaction layer every other module will call. No reconcile loop, hub
|
||
client, signing, or storage/backup orchestration yet (later slices).
|
||
|
||
### Added
|
||
- **Project scaffold**: module `gitea.dooplex.hu/admin/felhom-agent`, binary
|
||
`felhom-agent` (`cmd/felhom-agent/`), Go 1.24, zero external dependencies
|
||
(pure stdlib). `--version` flag; `version` var overridable via
|
||
`-ldflags "-X main.version=<v>"`.
|
||
- **`internal/proxmox` — API backend (`Client`)**: hand-rolled REST client over
|
||
`https://<host>:8006/api2/json` with `PVEAPIToken` auth. Typed read ops
|
||
(`Version`, `Nodes`, `NodeStatus`, `ListLXC`, `GuestStatus`, `GuestConfig`,
|
||
`ListStorage`, `NodeStorage`, `StorageContent`) and async mutating ops
|
||
returning a UPID (`RestoreLXC` — the primary create path, `Vzdump`, `Snapshot`,
|
||
`Rollback`, `DeleteSnapshot`, `SetConfig`, `Start`, `Stop`).
|
||
- **`WaitTask`**: polls `GET /nodes/{node}/tasks/{upid}/status` until stopped, then
|
||
asserts `exitstatus == "OK"` (authorization can surface at task execution, not
|
||
the POST — phase1-2 §1.3). Exponential backoff (1s→5s cap), context
|
||
cancellation + timeout. `*APIError` parses the offending privilege from a 403;
|
||
`*TaskError` parses it from a failed task exitstatus + log tail.
|
||
- **`internal/proxmox` — fenced root-CLI backend (`Privileged`)**: limited to the
|
||
three proven OS-root exceptions only — `CreateGoldenLXC` (keyctl `pct create`),
|
||
`MountUSBByUUID`, `SMART`, `Sensors`; each cites why it can't be the API. Fence
|
||
is structural (Client never shells out, Privileged never makes an HTTP call) and
|
||
asserted in tests.
|
||
- **TLS trust**: SHA-256 leaf-cert pinning (the host serves a self-signed cert) or
|
||
a CA file; an explicitly-named `insecure_skip_verify` that is off by default. No
|
||
blanket verification disable.
|
||
- **`internal/config`**: JSON config file + `FELHOM_AGENT_*` env overrides; the
|
||
token secret is never logged (`Redacted()`).
|
||
- **`internal/log`**: slog setup (text, stderr, configurable level).
|
||
- **`cmd/felhom-agent --selftest`**: read-only health report against a live host
|
||
(version/nodes/status/guests/storage); `--selftest=task --vmid N` exercises
|
||
`WaitTask` on a reversible snapshot→rollback→delete op (gated; default selftest
|
||
mutates nothing).
|
||
- **Tests**: unit tests with a mock HTTP transport + mock runner (UPID parse,
|
||
`WaitTask` running→OK / failed-403 / timeout / ctx-cancel, 403→privilege error,
|
||
response decoding against shapes captured live from `demo-felhom`, config
|
||
redaction, and the API-vs-root routing fence).
|
||
|
||
### Notes
|
||
- Types are grounded in the spike findings
|
||
(`felhom.eu/documentation/proxmox-platform.md`, `tests/phase{0,1-2,3}-findings.md`)
|
||
and the exact JSON shapes captured live from `demo-felhom` (PVE 9.2.2).
|
||
- Verified: `go build/vet/test` green on Go 1.24.4 (build server) and a live
|
||
read-only `--selftest` against the demo host with TLS fingerprint pinning.
|
||
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
|
||
token) is provisioned out-of-band; the agent only consumes the token.
|