Compare commits

...

124 Commits

Author SHA1 Message Date
admin dd7cdc09e7 R-366 slice 2 (decision 168): the host report carries the archives the restore-test skipped as another key's
gates / gates (push) Successful in 1m6s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin 7b0a8b234b R-105 (decision 169): remove the escrow-create -directive flag and the upload's directive field
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin ce1a4b4758 R-812 option A: the Proxmox package lane (layer pve) + the /etc/pve write gate
The wrapper gains layer "pve" (slow lane): the host's Proxmox userspace
packages only — origin "Proxmox Debian Repository", never a kernel / boot /
firmware / microcode name (R14), no removal, no undo, a new package only from
an allow-list; authority = a signed os_pve_step or the root-owned ring-0 mark.
The night leg runs it in ring 0 after a healthy host step; ring 1 only by a
signed job (PVEStepExecutor). While it runs, the agent's own /etc/pve writes
(every non-GET API call, pct config verbs, pvesm, pveum, felhom-pbs-apply)
wait on internal/pvegate. Health = the host rule + unchanged container ids +
pveversion reads the installed pve-manager.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin ac90169a5d R-861 (a) A1 + (b) B2: the image ref goes to a root verb that checks it; the agent's in-guest tee grant is gone; felhom-op's pct lines are exact (09 §3 decision 165)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:28 +02:00
admin 154d6dcaa9 Shared rule file: no hub image build or deploy in a session the operator does not attend (09 §3 decision 162)
gates / gates (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 09:55:02 +02:00
admin 2f7072050f REPORT: released and delivered (2026-10-07)
gates / gates (push) Successful in 1m4s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 09:31:58 +02:00
admin 85e799f360 CHANGELOG: v0.150.0 released (R-528, R-894, R-330)
gates / gates (push) Successful in 46s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 08:59:08 +02:00
admin 3a72a4811b Shared rule file: rule 11 — every helper prompt carries the brief's fences in full (09 §3 decision 160)
gates / gates (push) Successful in 54s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 08:52:36 +02:00
admin a165d53c88 REPORT: the second burn-down night
gates / gates (push) Successful in 49s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 00:14:50 +02:00
admin adaf86ad57 R-330: the agent sends SMART 187/188/199 (raw; omitted when unknown)
gates / gates (push) Successful in 53s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 22:13:04 +02:00
admin 74b5eae5b0 R-894: after a restart the agent remembers the last backup per tier
gates / gates (push) Successful in 35s
An unreadable storage right after an agent restart read the off-site tier
DUE (the in-memory record was empty). The newest success per tier is now
kept on disk and read ONLY when the storage cannot be read: fresh -> not
due, older than the cadence -> due, none -> due (unknown) as before. A
storage that answers stays the ground truth.

Ships with v0.150.0 after the 2026-10-07 read-back; nothing delivered.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 20:33:27 +02:00
admin 7e82f325b8 CHANGELOG: the memory-kill check (unreleased, to ship as v0.150.0 after the night read-back)
gates / gates (push) Successful in 50s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:36:09 +02:00
admin acccb66bd3 R-528 (09 decision 157): after a Docker engine step the wrapper proves the engine reports a memory kill
felhom-os-apply: a docker-layer apply runs oom_check() after health_after and reports
"oom_check": {result pass|fail|error, oom_killed, oom_event, exit_code, image, detail}.
One throwaway container (the controller's image, --pull never, --network none, 64m cap,
label felhom.oomcheck=1) runs dd bs=200M; pass only with OOMKilled=true AND the oom event
(read after a 2 s settle, --until = guest epoch + 1: measured on demo-hp, an --until taken
right after the run missed the event). docker rm -f always runs in a finally; every call
is bounded (<= 90 s). It never changes the step's outcome or health. New wrapper-only mode
"oom-check" (docker layer) runs the check alone; check_guest etc. still apply.

Agent: WrapperReport/Report gain OOMCheck (json:"oom_check"), copied unchanged in runLayer
and in the kept-copy path.

Tests: 9 wrapper tests + 2 Go tests, each red-proofed (audits/readback-2026-10-07/F/red-*.txt).
Also: test_felhom_os_apply.py's `if __name__` sat mid-file, so 11 tests (UnsentReport,
SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran as a script or from
TestWrapperSuite; moved to the end (they pass).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:34:55 +02:00
admin de812bc027 The shared rule file (09 §3 decision 152), identical to the other copies; no code change
gates / gates (push) Successful in 44s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 16:06:46 +02:00
admin 3e8ebeb96c Instruction files kept true (09 §3 decision 150): stale gate lists, paths and facts corrected; no code change
gates / gates (push) Successful in 46s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 13:43:57 +02:00
admin cefdc731a4 REPORT: the operator's ten answers (2026-10-06)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 12:28:58 +02:00
admin e56dcb8a4c CHANGELOG/v0.149.0 released (shas)
gates / gates (push) Successful in 37s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:50:21 +02:00
admin f277e619e2 CHANGELOG: unreleased — R-856 GET /host/crash-guard
gates / gates (push) Successful in 38s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:40 +02:00
admin 386f51edc6 R-856: GET /host/crash-guard — the host crash guard's last-boot record for the controller (09 decision 143)
The controller waits ~15 minutes with app mails after a crash boot of the host; it learns of the
crash boot from this route. Reads /var/lib/felhom-crash-guard/state.json (read-only, no Proxmox call)
and passes present/last_boot_at/last_boot_unclean/tripped through; a missing, unreadable or garbled
file answers 200 present:false. Guest-token authed like every sibling route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:01 +02:00
admin 130e3ed882 CHANGELOG: unreleased — R-444 weekly guest disk trim, R-99 runbook pointer
gates / gates (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:25:22 +02:00
admin be398f92e8 R-99: the PBS phantom WARN names the cleanup runbook (09 §3 decision 140)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin ee71abd1d4 R-444: weekly guest disk trim (pct fstrim) outside the night, under the heavy-op gate
Operator ruling 09 §3 decision 139. One exact sudoers rule FELHOM_FSTRIM
(`/usr/sbin/pct ^fstrim [0-9]+$`) + manifest entry guest-fstrim; new
internal/fstrim job: due Wednesday from 10:00 host-local, starts only
10:00-20:59, holds backup.InFlight (busy -> deferred to the next hourly
tick), failed trim retried at most 3x per week, bytes parsed from the
"(N bytes) trimmed" lines, last result per guest persisted in
<state_dir>/guest-disk-trim.json and reported as guest_disk_trim.
Config opt-out: "disk_trim": {"disable": true}.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin 37e98f452b REPORT: the burn-down night (2026-10-06)
gates / gates (push) Successful in 45s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 02:13:10 +02:00
admin b2b82ae828 bundle test: the ISO first-boot files the installer names under KEPT are not installer-written (go test red since felhom.eu 85de3f9b); CHANGELOG unreleased (R-426 decoys)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:25 +02:00
admin b78a0ff3ac R-426: decoys for the shared reuse-refs, instructions and observations gates
COVERS "reuse-refs", "instructions", "observations": the three shared
felhom.eu scripts run against a scratch clone of THIS repo (in a scratch
workspace symlinking the sibling clones they reach across to), so the
plant is in the agent's own REUSE.md / CLAUDE.md / REPORT.md. Convicted:
a missing cited .go and .md path, a version literal in CLAUDE.md's
effective text, R-419's prose-only Observations note. Passed: the real
files, the version inside an HTML comment, both genuine markers.
DECOY_SHARED_DIR lets a red-proof judge a mutated copy of the shared
scripts without editing the felhom.eu clone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 96047453cb R-426: decoys for the release-complete gate
COVERS "release-complete": the working-tree gate runs in a scratch clone
whose origin is a scratch bare repo, against the fake Gitea. Convicted:
the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated
commit, a tag with no package, and no-tag wins over a registry 500.
Inconclusive: a registry 500. Passed: the genuine release, an
`## Unreleased` heading above it, a newer version named only in prose or
under `###` (the withdrawn sweep decoy, now asserted the right way round),
a tag only origin has (the shallow-CI shape), and a LOCAL-only tag BY
DESIGN (CI's fresh clone and the published gate's converse probe see it).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 4bf5db6875 R-426: decoy suite for the published gate, against a fake Gitea
scripts/test_gate_decoys.py (new; COVERS "published"): an http.server on
127.0.0.1 stands in for Gitea through the gate's existing GITEA_BASE
seam, proxies stripped, so no case reaches the real registry. Facts
convicted: a tag whose package 404s, a tag tree without the configs, a
package one patch past the newest tag (never tagged), a patch-gap orphan,
a missing package that lexical sorting would drop out of the retention
window. Inconclusive, never a pass: tags api 500, a non-JSON 200, Gitea
unreachable. Passed: a clean registry, a non-semver tag, a version older
than the retention window (BY DESIGN).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 87977ff40a CHANGELOG/v0.148.0 released (shas)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 00:37:07 +02:00
admin 861d32a4b4 CHANGELOG: unreleased — R-349 agent_sha256, R-25 agent half (burn-down night)
gates / gates (push) Successful in 43s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:51 +02:00
admin 769c4c3cf2 R-25 (agent half): the format answer carries the NEW filesystem's UUID, bound to the durable id
After mkfs the agent re-resolves the bound durable id, requires it to name the
device it just formatted, reads the superblock back (blkid -p, requested fstype)
and returns fs_uuid in POST /disks/format and GET /disks/format/status. Anything
unverified returns "" — never a path-resolved guess. The controller half
(mount fs_uuid instead of re-resolving the UUID from the /dev path) is owed
in felhom-controller.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:11 +02:00
admin 64f704d0f7 R-349: the host report carries the sha256 of the RUNNING agent binary
A hand-built proof binary and the published artifact share a version string
but not their bytes, so no version check could see the divergence. The agent
now reports agent_sha256 (hash of /proc/self/exe, once per process; empty =
unknown) beside agent_version, the same mechanism as host.wrapper_sha256.
The hub half (compare against the vouched agent_sha256, surface drift) is
owed in the hub repo.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:11 +02:00
admin 208fac8027 CHANGELOG/REPORT: v0.147.0 released (shas)
gates / gates (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 18:58:37 +02:00
admin f1b9b41214 R-124 recipe root namespace as PBS spells it; R-118 no root size for an absent drive; R-269 rotated-out token rejected at once; R-317 dnsmasq install probed by its unit (burn-down round 2)
gates / gates (push) Successful in 47s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 18:56:13 +02:00
admin d83316326e R-291 retention record source, R-348 restart comment (no binary change; burn-down)
gates / gates (push) Successful in 22s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 16:56:18 +02:00
admin e06ed97fa8 agent v0.146.1 REPORT
gates / gates (push) Successful in 22s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 13:00:22 +02:00
admin e4b5cf9693 R-880: build-step-bundle.py — the transition bundle for a release whose bundle adds paths (an installed felhom-os-apply refuses unknown paths, R16)
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:23:55 +02:00
admin faa3cad92e agent v0.146.1 CHANGELOG (released)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:14:07 +02:00
admin fdd87178d2 R-861 review fixes: the signed update hands the A/B wrapper a root-owned copy of the hashed bytes; mount units accept no Wants/Requires/Before and no continuation lines; the escrow read walks the path without following any symlink
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:13:42 +02:00
admin 0342c7bb57 agent v0.146.0 CHANGELOG (released)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:01:25 +02:00
admin 6ab1e7c56c R-861: narrow the agent's root grants — exact sudo patterns, felhom-priv-apply content checker, fixed hook/parent files in the bundle, signed self-update verified as root, escrow root reads pinned
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:00:43 +02:00
admin 61345790ed REPORT: the 2026-10-05 catch-up session
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 10:29:12 +02:00
admin 7c4b8e599f v0.145.0: CHANGELOG (released 894da35c…, bundle 78c00adc…)
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 09:23:48 +02:00
admin 56ef1d6655 v0.145.0 code: the OS wrapper repairs dpkg's update journal by itself after a power cut (R-876); restore-test first check 30 min after start (R-874); neutral "sent late" text (R-875)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 09:23:23 +02:00
admin 78c890e4bb REPORT: the 2026-10-05 night-fixes session
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 08:25:49 +02:00
admin d48f1bbb23 v0.144.1: CHANGELOG (released 6ccd521d…, bundle e89a9ddf…)
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:49:29 +02:00
admin 8401a30917 v0.144.1 code: the wrapper survives a dead reader (BrokenPipe) so a killed pass still keeps its report; the agent looks for kept copies every 5 min (R-868, measured live)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:49:04 +02:00
admin d1b6004458 v0.144.0: CHANGELOG (released f18093c3…, bundle 6acf42fe…)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:15:07 +02:00
admin ca78c17b29 v0.144.0 code: R8 measures the real download (R-865); an OS pass's report survives a killed agent (R-868); the debug pass runs from the saved block when the hub is away (R-866)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:14:32 +02:00
admin c8d12f1f2a REPORT + CONTEXT: 2026-10-04 night (R-840 / R-860)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 20:39:47 +02:00
admin dc9164c5af CHANGELOG v0.143.0 (R-840); build-golden.sh 3.2.0: GOLDEN_GUEST_PKGS — the approved guest release at bake time, first-night count
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 20:08:16 +02:00
admin c9fa2e717b R-840: the config bundle — a signed agent_config_update brings a box's root-owned files (sudoers, wrappers, units)
gates / gates (push) Successful in 18s
felhom-os-apply gains mode 'bundle' (signed, verified by the wrapper itself against the root-owned
signers file — or, when that file is missing, only the installer's pinned key, which it then creates)
and --install-bundle (the installer's root entry). BUNDLE_FILES is the one table of paths; every check
(visudo, sh/bash -n, python, unit sections, RuntimeDirectory guard, nft -c, the route itself) runs
before the first write; a failed write or self-check puts every previous copy back. The trust root is
never a bundle path (R17). scripts/build-config-bundle.py builds it reproducibly; release-agent.sh
publishes it beside the binary. The agent reports the bundle record in system.config_bundle.
felhom-opsign signs agent_config_update. 43 wrapper tests (22 mutants red), Go executor tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 19:36:58 +02:00
admin 0af1187e03 CHANGELOG + REPORT: v0.142.1 released (R-858)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 18:21:46 +02:00
admin 495003051b R-858: after a Docker engine step the wrapper restarts ONLY the containers that mount the docker socket (controller, traefik); the health rule fails when the controller cannot reach Docker from inside its container (ruling 95)
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 18:21:24 +02:00
admin 42af3ab9bc build-golden.sh 3.1.0: live-restore on in the golden (fail-closed assertion), GOLDEN_DOCKER_PKGS pins the approved Docker engine set
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 17:30:19 +02:00
admin f24dce5b95 CHANGELOG + REPORT: v0.142.0 released
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 16:15:15 +02:00
admin b1746c25af Docker engine slow lane (live-restore once by reload; ring-0 pending-docker under a root-owned ring-0 mark; ring 1 and undo only by a signed os_docker_step the wrapper re-verifies against a root-owned signers file; same-container-id health), the version report (facts mode -> host report system stanza, R-852), guest restart scan every pass (R-849), the crash guard (kernel.panic=10, the 3rd unclean stop in 60 min stays off, 24 h re-arm)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 16:14:36 +02:00
admin 2e2e8f56b8 CHANGELOG + REPORT: v0.141.1 released (host reboot-needed fixes)
gates / gates (push) Successful in 17s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 13:39:35 +02:00
admin a6bc3f1197 os-apply: the host restart scan no longer hides lxc-start (skip ':/lxc/' not 'lxc'), and the host scans on every pass so a reboot clears 'reboot needed' (reboot_scanned reaches the hub) — both found live on demo-felhom
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 13:39:14 +02:00
admin 3bf77c3423 CHANGELOG + REPORT: v0.141.0 released (host fast lane, true tunnel status, fast leg)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 13:19:02 +02:00
admin cfba0d022a contract: the desired-state golden gains host_release (byte-identical with hub v0.131.0); TestOSUpdateGolden_Decodes checks it
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 13:18:43 +02:00
admin c3d08b4821 OS updates host fast lane + true tunnel status + fast leg: wrapper host layer (R12 appliance proof from the root-owned install record, R14 kernel/boot/firmware refused), select pending-fast, one call per layer, host-side version checks, restart scan only after an install, reboot-needed for PID 1/lxc-start; the leg runs the host step after a healthy guest step; GuestTunnelProber reads the cloudflared container + its readiness check (R-841)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 12:54:02 +02:00
admin a55eedcf2c CHANGELOG + REPORT: v0.140.0 released (OS updates, guest fast lane)
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 11:30:14 +02:00
admin 9cac3462bb osupdate: the health baseline is the start of the leg (inventory reading merged with the apply's own) — an app that stops during the run fails it (found live on demo-hp)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 11:22:03 +02:00
admin 1bb8608e88 os-apply: tell an UPDATED conffile from a KEPT one (dpkg's two shapes, measured live on demo-hp)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 11:08:43 +02:00
admin b84e0dd1bd selftest flag accepts os-update and wgtunnel (both dispatched, both refused); a test pins every dispatched mode
gates / gates (push) Successful in 19s
Found live 2026-10-04: the OS leg's debug action could not run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 11:04:03 +02:00
admin 23a8ef3de4 OS updates, guest fast lane (11 §8 step 2): felhom-os-apply wrapper (R1-R13 refusals, repair first, snapshot.debian.org fallback), FELHOM_OSAPPLY sudoers, the OS leg after the primary backup, hub os_update block + os-report, --selftest=os-update
gates / gates (push) Successful in 18s
No automatic undo: a customer guest cannot be snapshotted (R-837, measured).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 10:44:29 +02:00
admin 596238cc2e CHANGELOG + REPORT: v0.139.0 released (R-834)
gates / gates (push) Successful in 17s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 08:52:57 +02:00
admin 475bdce7e4 DR bring-up refuses beside a live original (R-834): source guest present, drives bind, or unreadable config
gates / gates (push) Successful in 18s
The DR route keeps onboot 1, binds the real drives and starts the guest: right on a replaced
host, a second box on the same drives beside a live original. The restore-test's no-host-bind
half is now pinned too (measured safe live on demo-hp 2026-10-04).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 08:52:35 +02:00
admin d766666ff8 CHANGELOG: v0.138.0 vouched with golden 0.283.1
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 11:40:20 +02:00
admin a4c09a7c11 REPORT: v0.138.0 (R-727)
gates / gates (push) Successful in 16s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 11:00:45 +02:00
admin 904dc20466 CHANGELOG: v0.138.0 released (R-727)
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 10:34:25 +02:00
admin e1b8269be0 restore test takes only this box's archives (R-727): an archive encrypted with another key is another box's
gates / gates (push) Successful in 16s
Red-proof RP39. Released as v0.138.0 by release-agent.sh.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 10:33:52 +02:00
admin 5c68c869b6 docs: v0.137.0 vouched with golden 0.276.0 (2026-09-28)
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-28 09:45:08 +02:00
admin 728d12b1a0 CHANGELOG: v0.137.0 released (tag + package verified by download), not vouched
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 14:01:26 +02:00
admin dd81866b16 REPORT: v0.136.0 and v0.137.0
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 14:01:02 +02:00
admin 3ef095fb71 agent: a guest outside the agent's ACL is not a known guest (R-689, v0.136.0 regression)
gates / gates (push) Successful in 14s
PVE answers 403 permission denied, not "does not exist", for a vmid outside the felhom pool;
v0.136.0 turned that into a lookup failure and the local tier read UNKNOWN every evaluation
(measured on demo-hp). Such an archive is skipped. Red-proofed; verified read-only on demo-hp
with the pre-release binary.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 14:00:42 +02:00
admin 7c986915ca CHANGELOG: v0.136.0 released (tag + package verified by download), not vouched
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 13:46:22 +02:00
admin 16dbc83221 agent: the restore test takes only archives of a guest that still exists (R-689, second half)
gates / gates (push) Successful in 14s
Measured on demo-hp right after v0.135.0: with the golden skipped the pick fell to a leftover
archive of guest 9100, deleted in August. "does not exist" skips it; any other lookup error
makes the tier unknown. Red-proofed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 13:45:56 +02:00
admin 9555a7f93b REPORT: v0.135.0 released and signed-delivered to both demo hosts
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 13:02:10 +02:00
admin 9ff937d8fb CHANGELOG: v0.135.0 released (tag + package verified by download), not vouched
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 12:16:17 +02:00
admin d4be12ca95 agent: the restore test proves only backups of a guest (R-689)
gates / gates (push) Successful in 13s
The golden template in local:backup/ was picked as the newest settled archive on demo-hp and
failed every 6 h. Candidates are now vzdump-<type>-<vmid> files or PBS ct|vm/<vmid> snapshots
with a reported vmid. Red-proofed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 12:15:42 +02:00
admin 7403c2a838 REPORT + CONTEXT: v0.133.0 and v0.134.0 delivered to both demo boxes; restore test back on
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 04:55:26 +02:00
admin 309e368731 CHANGELOG: v0.134.0 released (tag + package verified by download), not vouched
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 22:49:19 +02:00
admin 0722b2cdb0 agent: a whole-box backup that cannot fit its local target is skipped with a reason before anything starts (R-685)
gates / gates (push) Successful in 13s
Free space is read from GET /nodes/<node>/storage — GET /storage carries no usage (found live, before release).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 22:48:51 +02:00
admin 4fe2f81a32 CHANGELOG + REPORT + CONTEXT: v0.133.0 released (tag + package verified by download), not delivered
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:25:47 +02:00
admin 9bdb4dae8f v0.133.0: a restore-test can never fill a box's disk; leftovers retried on a timer (R-672, R-673)
gates / gates (push) Successful in 13s
Space preflight before anything is created (uncompressed size from the vzdump log / PBS
snapshot, x1.2 + 5 GiB, thin metadata, off the tested guest's pool when another storage
is eligible, unknown refuses, reported as a non-pass result). Failed scratch teardown and
the stale-lock sweep retried every 10 min (the sweep under the heavy-op gate). A thin
pool crossing 90% requests an immediate host report. Six red-proofs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:23:15 +02:00
admin d9864a94bf CHANGELOG: v0.132.0 vouched for Day-0 installs with golden 0.246.0 (operator decision)
gates / gates (push) Successful in 16s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 12:25:49 +02:00
admin 77cd70f7c0 REPORT: agent v0.132.0 - slow crash loop, signed delivery, live proof with the production window
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 11:30:45 +02:00
admin 1030abd7d6 CHANGELOG: v0.132.0 released (tag + package verified by download)
gates / gates (push) Successful in 11s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:27:16 +02:00
admin 18d03bd437 v0.132.0: the slow crash-loop counter (R-539, operator ruling 3 of 2026-09-16)
gates / gates (push) Successful in 12s
Beside the unchanged 3-in-15-minutes brake, a second counter: restarts the
supervisor performed in the last 24 hours. At the fifth the heartbeat stanza
sets slow_crashloop_since (moving at most once per 24 h), slow_crashloop and
restarts_24h; hub v0.117.0 mints controller_slow_crashloop (warning,
operator-only) when the timestamp moves. It never stops restarting.

Persisted per guest (tmp+rename, 0600) so an agent restart or reboot does not
reset it - unlike the fast record, whose reason for staying in memory (a
persisted give-up outliving the fix) does not apply to a counter that only
warns. Deliberate kills count. The startup line prints the new limits.

Red-proofs seen failing: no counter; the once-per-24h guard removed ('the
operator would be mailed per restart'); the save removed ('Restarts24h:1'
after an agent restart). Negative control: restarts 7 h apart never raise it.
go build/vet/test ./... green, 30 packages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:26:42 +02:00
admin e98b857684 REPORT: v0.131.0 supervisor + per-tier status, delivery and live validation
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 11:37:07 +02:00
admin dcdeb3d16d CHANGELOG: v0.131.0 released (tag + package verified by download)
gates / gates (push) Successful in 12s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:50:47 +02:00
admin 610804b98d v0.131.0: controller supervisor (R-523); per-tier backup status + tier storage presence (R-517/R-518)
gates / gates (push) Successful in 11s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:49:44 +02:00
admin 4586f0f7f6 re-run CI against a register that now carries R-421
gates / gates (push) Successful in 9s
The earlier run convicted correctly: instructions_gate found this repo citing R-421 while
felhom.eu's OPEN-ITEMS.md did not yet have the row. My ordering, not the gate's fault - the register
lives in felhom.eu, so a repo citing a new row must be pushed after it.
2026-09-01 12:45:52 +02:00
admin 205e22babe decoy sweep: no gate changed here, and that is the result (R-421)
gates / gates (push) Failing after 12s
All 29 gate scripts across the four repos were read and DECOYED - the label constructed without the
fact, the gate run, the verdict recorded. 16 were fooled. None of them were in this repo.

A decoy that nobody would write proves nothing, so the attempts that turned out illegitimate were
WITHDRAWN rather than counted. Both of this repo were withdrawn, and both are named in the audit.

The gates here that could not be given a plausible decoy are listed BY NAME in
felhom.eu/scripts/decoy_coverage_gate.py EXEMPT (R-426) as UNTESTED - not as sound. A gate nobody
tried to fool is UNKNOWN, and calling it sound would be the same confident guess this sweep exists
to find.

Survey table: felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md
2026-09-01 12:38:58 +02:00
admin 058b945064 gate 11: register the shared observations gate
gates / gates (push) Successful in 10s
felhom.eu/scripts/observations_gate.py, invoked across the workspace like
reuse_refs_check.py and instructions_gate.py. This repo's REPORT.md has no
observations section today, so the gate passes quietly - it is registered for
the session that writes one.
2026-08-23 13:53:20 +02:00
admin 40d857b527 CHANGELOG: the fleet runs the published bytes, not the proof build (R-349)
gates / gates (push) Successful in 7s
Both boxes were first given a hand build: same source, same version
string, different bytes (256e0829 vs the published a56a92a7), because
release-agent.sh builds with -trimpath -buildvcs=false and a hand build
does not.

Nothing would have corrected it. The boxes already reported 0.130.0, so
self-update saw the vouched version as installed and would have done
nothing, forever. Every version check in the system compares the STRING.

Both reinstalled from the downloaded package; both now report a56a92a7.
2026-08-20 12:53:37 +02:00
admin 7ae6990bac v0.130.0 released: tag + package published, heading now claims it
gates / gates (push) Successful in 8s
sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3
tag v0.130.0 at 7569f34, 14,141,158 bytes.

Reproducible: rebuilding with -trimpath -buildvcs=false matches the
published artifact byte for byte (R-186's property, checked not assumed).

The heading said UNRELEASED while the fix was hand-installed on demo-hp
only -- publishing then would have pushed it onto the control box through
self-update. release-complete convicted on the release heading and was
right to; the answer was to stop claiming a release, not to bypass it.

NOT VOUCHED by this commit. Vouching is the separate operator act.
2026-08-20 12:47:25 +02:00
admin 7569f34aeb CHANGELOG: correct a claim this session made and then disproved
gates / gates (push) Successful in 8s
The v0.130.0 draft said the 388 descriptors already stuck on ep0 would
persist until the PBS proxy restarted. Measured within the hour: they
clear when the AGENT restarts. demo-hp released its 199 in one second
(415 -> 216 fd); demo-felhom released the remaining 203 (220 -> 17 fd in
under two seconds). 17 is precisely ep0's t0 baseline of 2026-08-18.

They were held on both sides. Closing either side ends them. ep0 was
read-only throughout and its proxy PID never changed.
2026-08-20 12:39:12 +02:00
admin ede49b610d R-344: restore the idle-connection timeout our hand-rolled transports lost
gates / gates (push) Successful in 7s
Every client here pins TLS, so none can use http.DefaultTransport and each
hand-rolls its own. A composite literal takes IdleConnTimeout ZERO, which
means retain idle connections forever -- not "use a sane default".
pbsTargetsFromPVE builds a fresh pbs.Client every cycle and drops the
previous one, and an abandoned http.Transport does not close its
connections. One stranded socket per cycle, on both sides, forever.

Measured: 388 established connections on ep0 over 46 h, 194 per box, zero
closed in a 31-minute window. pvestatd and proxmox-backup-client made
162,404 requests in the same window and leaked none.

New leaf package internal/httpx owns the default (90s, http.DefaultTransport's
own value) and NewTransport, which returns a FRESH transport per call and
treats <=0 as "use the default", never "no timeout". pbs.Config gains
IdleConnTimeout for tests only.

hub and proxmox carried the same missing default and are corrected here for
consistency. Neither contributed to the ep0 leak -- both are built once per
process and neither talks to ep0:8007.

Tests count connections SERVER-side and model the abandonment, so they pin
the consequence rather than the field. Two red-proofs, both seen failing:
removing the timeout -> "still holds 5 open connection(s), want 0";
DisableKeepAlives -> "3 sequential requests over 3 connection(s), want 1"
(the leak test PASSES under that one -- it is the worse-fix guard that
catches it).

Not released: hand-installed on demo-hp only so demo-felhom stays the
control. CHANGELOG heading stays UNRELEASED until the publish is authorised.
2026-08-20 11:09:39 +02:00
admin f17ed11599 REPORT: agent v0.129.0 — the retained-package recovery class, released and deployed
gates / gates (push) Successful in 21s
2026-08-12 18:49:51 +02:00
admin 1db56bf837 v0.129.0 — a correct code for an earlier package stops being called wrong (R-311)
gates / gates (push) Successful in 14s
Yesterday's drill proved a retained escrow package opens a set-aside store and
restores planted files byte-identical, while this agent answered the customer's
correct code with "the recovery code did not open the sealed bundle". Nothing had
ever tried the retained packages, so a correct-but-earlier code and a mistype were
genuinely indistinguishable.

OffsiteKeyRecoverer gains an optional FetchRetained, consulted ONLY after the
current package refuses, so the ordinary recovery pays nothing for it and cannot
fail because of it. A match returns ErrCodeOpensRetained wrapped in a
RetainedOpenedError carrying the supersession date - no material, no code, no
password. The local API answers 422: a FIFTH status added to the R-224 switch,
never a restructuring of it.

Fail-safe in every direction. Nil fetcher, a hub too old for the route (404 is a
clean "none"), a transport failure, a malformed package: each leaves the original
refusal standing. Attempts bounded at 6 because each unwrap is ~1s of scrypt.

Seven tests with REAL age crypto - the two situations are indistinguishable AT
THE UNWRAP, so a faked unwrap would prove nothing. Red-proof asserted applied:
remove the retained lookup and the fail-closed wrong-code error returns, which is
the lie in those exact words.
2026-08-12 18:40:00 +02:00
admin 53d047a6c1 Two guards, one number: bound the published check to the retention it must live with
gates / gates (push) Successful in 17s
Gates only. No release, no version bump, no binary published; the agent stays
v0.128.0 at 28ba8593b8 and nothing on a customer's machine changes.

THE COUPLING DEFECT. The registry stopped serving 0.120.0 and older while
check-published-versions.py demanded every tag still be downloadable. Both rules
are sensible and together they are impossible, so CI went red at a commit whose
own run had been GREEN the day before -- and would have gone red again at the
next publish when 0.121.0 was evicted. scripts/retention-policy.json is now THE
number and both readers take it from there.

WHAT CI NO LONGER COVERS, and it prints this on every run rather than leaving it
to be discovered: a released version older than the retention window is no longer
asserted downloadable. Its git tag and its config tree ARE still asserted -- only
the binary's presence is dropped. A missing policy file is INCONCLUSIVE (exit 2),
never silently unbounded.

THE NUMBER IS NOT A LOCATED RULING and the file says so in its own header. Ten is
what the registry demonstrably holds; no register row records a prune, R-210 says
"Nothing was deleted; this is a list, not an action" and concerns local Docker
images, and container packages hold 19 each. The principled bound is the hub's
vouched min_agent floor -- nothing can install below it -- and that is the
recorded follow-up.

check-release-complete.py is the tag half as a machine. release-agent.sh already
warned that "a released version without a git tag 404s a box mid-install, as
root" and the step was still missed, so this is a gate and not a reminder. Legs
1-2 need no network and run in --fast, so the pre-push hook is the earliest
catch. Red-proved by repointing the CHANGELOG head at an unreleased v0.129.0:
both legs convicted and each named its fix command.

Three controls run: green at 10 naming what it dropped; widened to 11 the evicted
version re-enters and convicts; policy removed gives INCONCLUSIVE naming the path.
2026-08-09 19:05:05 +02:00
admin 28ba8593b8 v0.128.0 — R-221: the escrow seed is asserted every tick, not remembered once
gates / gates (push) Failing after 28s
A rebuilt box could not run the escrow ceremony AT ALL, with no way forward from inside the product.
This was the only open item blocking a customer from something we promise them.

MECHANISM, established at file:line rather than assumed. The preflight refuses on
escrow.pbs_storage_id; the pbsdr bridge writes that key; and it wrote it from exactly one place —
finishConverged, reached only on the paths that actually converge.

The marker and the key live in different places and die at different times. The marker is host-side
(<agent-state>/pbsdr/marker.json). The key is in agent.json, which step_agent_config renders from
`base = {}` unless an explicit --preserve-from is given (felhom.eu/scripts/felhom-host-install.sh:
2396 the step, :2449 the render, :2579 the O_TRUNC write; the flag :1246, defaulting empty at :256)
— AND THE RENDER NEVER WRITES AN escrow SECTION AT ALL (grep over the whole heredoc: zero hits). So
a rebuild keeps the marker and takes the key: same descriptor, same hash, early return, and the seed
never runs again into a config that no longer has it.

A rebuild is only the case that was measured. The same hole opens for a hand-edited or restored
config, which is the honest reason this is a seam fix rather than an installer fix: the seed must be
a thing the loop ASSERTS, not a thing it did once.

Apply now re-asserts the seed BEFORE the idempotent early return. seedEscrowStorageID is unchanged
and still never clobbers a different existing value — an operator's own choice outranks the
descriptor's, with a warning naming both.

THE EARLY RETURN IS KEPT. It stops a converged box re-running Proxmox operations every 60s, and
TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls asserts ZERO recorded runner calls on that
tick, so a "fix" that simply deleted the return would fail. Cost: one small file read plus a JSON
parse per tick, no exec, no network, early-returning once the value matches.

A seed failure can never un-converge the box: Warn plus a message on the published status, exactly
as finishConverged does it — no marker write, no state change.

Tests drive the REAL Apply with a real temp-dir agent.json and a call-recording runner; calling
seedEscrowStorageID directly cannot see the early return, which IS the defect. Production wiring
(pbsdr.NewManager(..., cfg.SourcePath, ...)) is asserted by walking main.go's AST, not by
strings.Contains, which a commented-out call also satisfies.

Red-proofs, each with the mutation asserted applied: removing the new call makes Scenario A fail on
today's tree (it did, with the intended message); removing the early return makes the
zero-Proxmox-calls assertion fail (it did).

go build / go vet / go test ./... green (29 packages), run separately from this commit.
2026-08-08 16:29:13 +02:00
admin 6981450110 docs: a comment claimed the hub reads a field it has no field for (R-260)
gates / gates (push) Successful in 26s
Comment-only; no behaviour, no wire change, no version bump, nothing to rebuild.

HostReport.SelfUpdatePending / SelfUpdatePendingVersion carried "The hub reads an absent field as
pending=false, the correct default." The hub has NO FIELD for either, so it reads nothing — present
or absent — and encoding/json discards them on arrival. The sentence described an intent rather than
the code and read as settled long enough that a class sweep had to find it.

The emission is correct and stays: the agent reports the truth and the fault is entirely in the
receiving. The missing consumer is R-264 (OPEN). felhom.eu/scripts/wire_contract_gate.py now refuses
any NEW field of this shape and records the existing ones as reasoned allowlist entries.
2026-08-08 08:46:17 +02:00
admin 703db166e7 v0.127.0: a mount Felhom itself made is not 'something else' (R-220)
gates / gates (push) Successful in 8s
After a rebuild the customer's own drives could not be re-attached: candidates
returned initialize:[] attach:[] while both drives sat there, and the deploy
refused with 'choose an attached drive from the list' — a list that was empty.
Measured live three times.

Mechanism: enrolment mounts a drive TWICE, at /mnt/felhom-drives/<name> and at
the raw /mnt/<name> it creates on the host. The host survives a guest rebuild;
the controller's registry does not. So classifyClaim saw a mount outside the
managed prefix and concluded 'claimed by something else' — about our own mount.

The fix is CORROBORATED, not a widened prefix: a non-managed mountpoint is
forgiven only when the SAME device is also mounted under the managed path, a
pairing only our enrolment produces. A disk another system uses — /srv/data,
/media/x, even /mnt/someone-elses-disk — has no counterpart and is STILL
refused, with its own test and a red-proof showing an over-wide fix offering it
for formatting.

Read from /proc/mounts deliberately: the lsblk invocation is pinned verbatim in
configs/felhom-agent.sudoers, so using the plural MOUNTPOINTS would have coupled
this to a sudoers rollout. /proc/mounts is world-readable — no sudo, no new
allowlisted command, no config change.

Fail-safe: an unreadable mount table corroborates NOTHING, so the device
classifies exactly as before. 'Could not corroborate' must never read as 'ours'.

29 packages ok, vet clean, agent gates OK.
2026-08-06 12:55:28 +02:00
admin aa74294a7d docs: felhom-agent CLAUDE.md becomes a core plus path-scoped rules (R-229 leg b)
gates / gates (push) Successful in 8s
175 -> 99 effective lines. New .claude/rules/{proxmox,localapi,backup,storage}.md alongside the
existing health-checks.md. The release section points at the felhom-build-deploy skill rather than
restating a table that drifts from the script; the layout section's per-package annotations moved
into the rule file for their area instead of being deleted.

Kept in the core because it is the only part re-injected after /compact: the root-CLI fence and its
three exceptions, the destructive-op gate, prove-ownership (audit A1), the gate entry point, the F9
live-validation fence, and the checklist.

health-checks.md overlaps localapi.md and storage.md on three globs -- deliberate, both load,
stated in each file. Go build/vet/test green and unchanged.
2026-08-06 11:28:27 +02:00
admin 5b2666e3a2 docs: R-168 is CLOSED — the "CI is still owed" sentence was stale (R-229 part 2)
gates / gates (push) Successful in 11s
Corrected in all four instruction files across all four repos. Found while confirming this
session own push by run ID, which is precisely the check that catches it.

In felhom-agent/CLAUDE.md the sentence contradicted the same file release section, which
already said R-168 mails the failure -- a contradiction inside one instruction file, the exact
class the R-229 work exists to find.

REPORT.md deliberately NOT overwritten in the sibling repos: a one-line docs correction must not
destroy the record of their last real implementation.
2026-08-06 11:02:59 +02:00
admin 062a7027ab docs: remove expired and contradictory blocks from CLAUDE.md (R-229)
gates / gates (push) Successful in 8s
Surgical corrections only; the file is deliberately NOT restructured (deferred).

Deleted the expired TEMPORARY block. It read "felhom-pve is at a remote site
(until ~2026-08-02) ... Delete this block on return" and was still being read as
current fact on 2026-08-06, four days past its own deadline, while
felhom-controller/CLAUDE.md asserted the opposite. The location-independence fact
worth keeping (localapi binds 169.254.253.1:8443 on vmbr9 since the R-50 island
migration) moved to an HTML comment.

Every component version literal is gone from effective text, including the
--version reading and the go.mod Go directive. Versions change several times a
day; ask the hub's /hosts + /configs or the box.

The drill-VM claim and the host addresses now point at
documentation/operations/nodes.md, which already stated both correctly. This
file's drill-VM claim was the correct one -- confirmed by qm list on demo-hp.

The R-115/R-188/R-186 release narratives moved to an HTML comment and to the
felhom-build-deploy skill; the directives stayed (never hand-roll the build; the
build -> tag -> publish -> push order; reproducible -trimpath -buildvcs=false).

The health-check block-I/O rule became .claude/rules/health-checks.md, scoped to
the five packages where health checks are written. It had been duplicated from
felhom.eu/CLAUDE.md with a note explaining that that file does not load in an
agent-only session -- correct reasoning, made obsolete by path-scoped rules.

agent_gates.py registers the shared instructions gate.

Docs only -- no Go, no version bump, nothing built or deployed.
Ledger: felhom.eu/documentation/audits/LEDGER-instruction-trim-2026-08-06.md

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JJc8sAGRWmavP3rMtdpkr2
2026-08-06 09:38:38 +02:00
admin a2e914f683 v0.126.0: a fetch failure is not a wrong recovery code (R-224)
gates / gates (push) Successful in 7s
A hub the agent could not reach was reported to the customer as a bad recovery
code. Measured live 2026-08-05 (CAMPAIGN-11 F3): hub firewalled off, a CORRECT
current code, and the customer told it did not open their package — in 0.0556s
against ~1.0s for a real unseal. No unseal was attempted.

The discriminator existed here and this boundary threw it away: recover.go
fails at four distinguishable points and the local-api handler had cases for
two, with a default answering 'the recovery code did not open the sealed
bundle, OR the bundle could not be fetched'.

escrow.ErrBundleFetch now joins the fetch leg and the handler routes it to 502
with its own words — the code was NOT used. 502 not 4xx: the request was not
bad, an upstream dependency failed. Four situations, four statuses: 502 fetch /
400 fetched-and-refused / 404 no bundle / 409 predates the field. The
controller classifies on the STATUS and never parses the sentence.

A GREEN TEST NAMED THIS DEFECT AND DID NOT PREVENT IT.
TestRecoverOffsiteRepoPassword_FetchErrorIsDistinct has said since v0.125.0
that the operator must not be sent to re-read their code because the hub was
unreachable — and passed throughout, because it asserted this package's error
STRING one layer below the merge, and a string is not something a caller can
branch on. Re-pointed at the sentinel, with a consequence-level twin asserting
the status.

Red-proofs: removing the %w join fails the sentinel test; deleting the handler
case makes fetch and wrong-code both answer 400 with the wrong-code sentence.

29 packages ok, vet clean, agent gates OK.
2026-08-06 07:55:15 +02:00
admin 0404f60e6a pre-push: refuse a push from a clone outside the felhom workspace (R-204 rider)
gates / gates (push) Successful in 10s
The workspace root is already documented (workspace-CLAUDE.md, the workspace-root
CLAUDE.md 'stay inside it') and work drifted into a home directory anyway. A rule
that has failed once as a reminder is not fixed by writing it down again, so it is
now asserted where it can bite.

A push is the right trigger: throwaway clones under /tmp for probes and red-proofs
never push, so nothing legitimate breaks. Symlinks are resolved on both sides; an
absent workspace root SKIPS the check rather than failing it, so this cannot brick
a legitimate clone on another machine. The only bypass is the documented
--no-verify, whose use is already reportable.

Identical in all four repos.
2026-08-05 10:46:39 +02:00
admin 3f5f61b716 docs: R-199 links 6-8 — CONTEXT + REPORT (proven live on demo-felhom)
gates / gates (push) Successful in 7s
2026-08-04 13:56:38 +02:00
admin 6d7904786c agent v0.125.0: open the sealed bundle, return one field (R-199 links 7-8)
gates / gates (push) Successful in 7s
Link 7's only production caller was a --selftest reading R from an env var. Link 8 did not
exist: that selftest writes the whole bundle JSON and its success message named
"tunnel_token + pbs_token" -- accurate when written, a misstatement since v0.77.0 sealed the
offsite repository password into the same bundle. It now names what THIS bundle carried and
what it did not.

POST /escrow/recover-offsite-password: the controller supplies R, the agent fetches this
host's own blob from the hub (self-scoped by the per-host key), unseals it, and returns ONLY
the offsite restic repository password plus its sha256. Not the tunnel token, not the PBS
token, not the WG key -- the controller is a trust tier down and needs none of them.

R: in memory for one call, cleared on every path, never on disk, never in argv, never logged,
never echoed. A test redirects TMPDIR and asserts the tree is EMPTY afterwards -- emptiness
rather than a content scan, because a content scan is defeated by a later call overwriting the
leaked file, which is how the first version of that test passed its own red-proof while R sat
on disk.

Three distinct outcomes: no blob (404), a bundle that opens but predates the field (409), a
code that does not open it (400, fail-closed at the KDF, nothing written).

The wiring is asserted by an AST walk from func main() to the Options field, not by grep.
2026-08-04 13:41:12 +02:00
admin 856a127cd6 v0.124.1: the repair record must survive the probe that did NOT feed the hub (R-190)
gates / gates (push) Successful in 6s
v0.124.0's transition record never reached the hub, and only the live run showed
it. The capability reported degraded for "one cycle" — the probe call that did the
repair. But probeAll is invoked independently by the self-check log and by the
collector building a host report. On the demo box the repairing call was the log's
(09:39:34, journal shows the self-repair and degraded=1) and the report three
seconds later found the grant present and sent ok. The agent's journal had the
record; the hub had nothing. That is the silence R-190 is about, re-created inside
its own mitigation, with every unit test green.

Fixed with a latch on TIME, not call count: a confirmed repair reports for 20
minutes, which exceeds the 900s report interval, so at least one report must carry
it. It clears on its own and is per tier.

Two hollow tests caught and fixed on the way — one asserting a value it built
itself, one asserting the latch helper rather than the path consuming it (its
red-proof duly passed). The decisions now live in storeGrantHealthyVerdict and
storeGrantRepairedVerdict and the tests call those.
2026-08-04 09:44:56 +02:00
admin 257c4d85c0 v0.124.0: a lost storage grant repairs itself, and says that it was lost (R-190)
gates / gates (push) Successful in 7s
R-190 is a grant that worked at 04:44 on 2026-08-03 and was gone by 09:24, with a
reinstall, logged pveum activity and cluster-log entries all ruled out. The cause
is open; the resilience need not wait for it.

Everything needed already existed and had only ever been called once: the root
wrapper's `grant` verb, its sudoers vector (`grant *`, any storage id — confirmed,
not assumed), and the exact command. The verb had only ever run at storage
creation — the "built but never wired" shape in a verb rather than a seam.

The probe now runs that wrapper on a missing grant and re-reads ONCE to confirm,
the pbsdr R-22 shape including its restraint.

The record is the half that matters. A repair leaving only "ok" behind destroys
the only evidence a permission vanished, so a recurring loss becomes undetectable
— worse than the fault. A confirmed repair therefore reports DEGRADED for exactly
one cycle with the explanation in Feature, because that is the field the hub puts
in the operator's email (Reason does not travel). Nothing new was built: the hub's
existing ok->degraded->ok edge is the channel, so one loss produces one alert pair.
No wire change, no hub change, no new event type.

Bounded at one attempt per tier per hour: a storage can be unreadable for reasons
an ACL cannot fix, and re-granting every cycle is a repair loop wearing a fix's
clothes. A failed repair never masks the fault.
2026-08-04 09:38:27 +02:00
admin 72161f6cf0 REPORT: correct the manifest commit hash (311dc06)
gates / gates (push) Successful in 7s
2026-08-03 19:04:44 +02:00
admin 03b58cec0a REPORT + CONTEXT + REUSE: R-185 closed, with the corrected root cause
gates / gates (push) Successful in 7s
The installer defect was NOT PVE_STORAGES as the row and the task assumed: the
create arm of configure_backup_target grants, the Scenario-F reuse arm did not.
Also records the measured trap (an ungranted path answers with INHERITED
privileges, not empty and not 403), the deviation from the spec's suggested
Prober generalisation in favour of the existing poolReadStatus precedent, the
hollow test caught before it shipped, and that demo-hp carried the same drift and
was fixed.
2026-08-03 19:04:15 +02:00
admin fe14bc62c0 v0.123.0: a tier the box cannot READ now says so (R-185)
gates / gates (push) Successful in 7s
The missing grant is one command; the silence was the defect. On demo-felhom the
agent's token has FelhomAgentStore on local, local-lvm and felhom-pbs — and not
on felhom-backup, the storage the same installer configured as
local_backup_target. That storage answers {"data":[]} through the token while
root sees three archives.

An empty listing is what a FORBIDDEN tier and a NEWBORN tier both return, so
pickForThisRun skipped it as "no settled archive yet" and the tier was never
restore-testable on that box. The permission question, unlike the listing, has a
definite answer, so it is asked directly: Client.Permissions reads
/access/permissions as the agent's OWN token, and one capability.Status per
configured tier reports it — composed around the sudo prober, the way the
pool-read check already is.

Measured first, because the obvious reading is wrong: an ungranted path answers
neither empty nor 403, but with the privileges inherited from the box-wide grant
(Sys.Audit, SDN.Use, Datastore.Audit). Checking for Datastore.Audit would report
a blinded storage healthy — red-proved. The probe tests for
Datastore.AllocateSpace.

The probed set comes from the box's own config, never a fixed list: a hardcoded
probe list is the defect reproduced inside the fix. Critical, because the hub
alerts only on critical — except the "local" fallback target, which is reported
but does not page. It never looks at content, so it cannot alarm on a newborn
tier; it never reports ok when it could not ask. Status wire shape unchanged, so
no hub change.
2026-08-03 18:53:47 +02:00
admin 0b28eae7bb REPORT: R-189/R-188/R-186 — live evidence, the three sha values, and the observations
gates / gates (push) Successful in 6s
Scenario A proven on demo-felhom against the exact observation that filed R-189:
a 675 s offsite restore-test passed, the agent was restarted 11 seconds later
(inside the reporting window), and the hub's very next report carried
'1 restore-tests' where the same sequence produced 0 this morning. The hub's own
database holds the archive, the tier, the pass and the ORIGINAL test time, with
the run mechanics deliberately zero.

Also records the property the validation surfaced: the state holds one proof per
tier, so proving an older archive re-arms a newer one — confirmed live after the
defaults were restored.
2026-08-03 16:58:35 +02:00
admin 7581f8140a v0.122.0: three ways the signals lied about themselves (R-189, R-188, R-186)
gates / gates (push) Successful in 7s
All three are the reporting and release path misreporting its own work. No
customer machine, no backup, no restore, no data. The restore-test itself and
when it runs are unchanged.

R-189 — a passing restore-test no longer vanishes on a restart. restore_tests[]
came only from the in-memory store, whose comment ("lost on restart; the cadence
re-populates") was true under a timer and stopped being true when R-86 made the
agent refuse to re-test a proven archive: the proof is then not repeated for a
whole archive generation. Observed live — a 14.5 GB offsite PASS reached no
host-report because the agent was restarted 2m43s later. RestoreTestState now
carries tier + verified beside the archive and renders reportable entries; the
collector merges them, one per tier, newest by TestedAt. It refuses to lie: a
record missing archive-or-tier produces no entry, and run mechanics are not
re-invented. Only successes are persisted, and the asymmetry is now written where
it will be read.

R-188 — a correct release stops emailing a failure. Only the tag PUSH moved
(build -> tag locally -> publish -> push tag): the push wakes CI, and a tag
visible before its package made the gate correctly fail a correct release about
half the time. The old order's invariant is asserted directly instead — the gate
now refuses a published version with no tag, as a bounded probe that prints its
own coverage, because the package listing api is still 401 without a token.

R-186 — a released binary can be verified by rebuilding it. -trimpath
-buildvcs=false: same source, same bytes, tag or no tag. Measured. publish-agent's
fallback also forced CGO_ENABLED=0 and produced a 74 KB different binary for the
same version; both paths now build identically. CLAUDE.md records the command.
2026-08-03 16:40:18 +02:00
admin 3d0a1d615d REPORT: restart proof, R-188/R-189, and the restored defaults
gates / gates (push) Successful in 7s
2026-08-03 15:36:46 +02:00
admin 77e2cc4583 CONTEXT: v0.121.1 (a quiet evaluation is audible) + the live proof
gates / gates (push) Successful in 6s
2026-08-03 15:32:23 +02:00
admin cd1b087db7 REPORT: R-86 live results (14.5 GB offsite restore-test, due-triggered, 635 s)
gates / gates (push) Successful in 7s
2026-08-03 15:27:33 +02:00
admin 53d0c6bfc4 v0.121.1: 'nothing is due' must be AUDIBLE (R-86 + standing rule 3)
gates / gates (push) Successful in 6s
Before R-86 every tick ran a heavy restore-test, so the scheduler was audible by
construction. After it, 'nothing is due' is the NORMAL outcome — and it was
logged at DEBUG, which journald drops. An empty journal would then be equally
consistent with a healthy loop and a dead goroutine: the shape the R-88 watcher
was retired for, re-created by making the quiet path the common one.

A not-due evaluation now logs one INFO line naming every tier's verdict (four
lines a day at the 6h default), and an unlistable tier reads UNKNOWN with its
error in that same line, so a lookup failure can never present as 'nothing due'.

Red-proved through the scheduler's own tick, not the helper.
2026-08-03 15:26:54 +02:00
190 changed files with 22125 additions and 1079 deletions
+46
View File
@@ -0,0 +1,46 @@
---
paths: ["internal/backup/**", "internal/pbs/**", "internal/pbsdr/**", "internal/dr/**"]
---
# Backup, PBS and DR
`internal/backup/` is the vzdump runner, restore-test scheduler and report store. `internal/pbs/` is
the fingerprint-pinned PBS-API client plus the verify maintenance loop. `internal/pbsdr/` and
`internal/dr/` carry the DR tier and recipe halves.
## The three PBS laws
1. **Set-only.** `pvesm remove` **DELETES the encryption key**. Re-apply configuration; never remove
and re-add a PBS storage to change it.
2. **Secret on stdin.** A token secret is passed on stdin, never as an argv the process table shows.
3. **Verify the pin BEFORE consuming the secret.** A fingerprint check after the secret has been sent
protects nothing.
## Verify is server-side, and its default skips the work
The agent drives verification **remotely** via the PBS API; `proxmox-backup-client` has **no** verify
subcommand. `POST .../verify` defaults to **`ignore-verified=true`, which SKIPS already-verified
snapshots** — send `ignore-verified=false` to actually re-read and detect corruption. A verify that
skipped everything reports success.
## Presence is not success
A timestamp recording an **attempt** is not evidence of a **result**. Where a status field travels
beside a timestamp, the verdict must consult **both** — or the timestamp must record only successes.
Ask of any timestamp: *what exactly must have happened for this to be set?* If the answer is "we
tried", it cannot answer "did it work".
**Corollary:** when a verdict changes which field it counts from, the alarm text changes with it.
Leaving a message reading `last run 8h ago` while alarming on a six-day-old **success** turns a true
alarm into one the operator dismisses.
## Prune is server-side now
`DatastoreBackup` carries **no** `Datastore.Prune`. Boxes set `keep_last: 0` and the off-site endpoint
runs the prune jobs. **Box tokens stay write-only — never widen that grant** (R-89).
<!--
The ignore-verified default is the sharpest instance of the "absent log line" class in this repo: a
verify that silently skipped every snapshot completes fast, exits clean, and reports the same shape
as one that read every byte.
-->
+27
View File
@@ -0,0 +1,27 @@
---
paths: ["internal/capability/**", "internal/storage/**", "internal/localapi/**", "internal/hub/**", "internal/guesthook/**"]
---
# A health check issues no block I/O
No `statfs`, no `getdents`, no read, write or `fsync` — **not even behind a timeout**.
A probe that touches a wedged device enters uninterruptible sleep, survives `SIGKILL`, and cannot be
recovered until the device returns or the host reboots — so `systemctl restart` hangs too. A timeout
protects the caller's control flow and nothing else: the blocked thread remains.
**Liveness is decided from `/proc` and the kernel's own state**, never by reading or writing the
filesystem.
<!--
Measured, R-117 spike §6.3 (felhom.eu/documentation/audits/SPIKE-r117-bind-liveness-2026-07-30.md):
a probe stayed in D state 3m50s after kill -9; a buffered write with no fsync blocked too (O_CREAT
needs journal access); and statfs/getdents returned HEALTHY on a namespace that EIOs every byte —
fast, and wrong.
This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that
felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before
path-scoped rules existed. Deliberate scoped copies now live in felhom.eu/.claude/rules/hub.md and
felhom-controller/.claude/rules/gates.md (hub.md's comment names them); none is the single source. This
file is the copy that loads exactly where agent health checks are written. (2026-08-06; corrected 2026-10-06)
-->
+44
View File
@@ -0,0 +1,44 @@
---
paths: ["internal/localapi/**", "internal/authz/**", "internal/guesthook/**"]
---
# Local API, authz and guest hooks — the per-guest blast radius
`internal/localapi/` is the narrow per-guest local API: token store, disks/format, guest binds,
controller swap, stale-lock recovery, pinned self-signed leaf. `internal/authz/` is the operator
signed-op verifier (SSHSIG) plus the durable nonce store. `internal/guesthook/` installs the
pre-start self-heal hookscript.
> **Overlap note:** `health-checks.md` also matches `internal/localapi/**` and
> `internal/guesthook/**`. That is deliberate — both rules apply there and both load. Neither
> supersedes the other.
## Scoping is the whole security property
This API is reachable **from inside a customer guest**. Every route must be scoped to the guest that
called it — a route that can name another guest's id has escaped its blast radius. Fail **safe to
protected**: an unrecognised or unresolvable caller gets less access, never more.
## Replay protection must survive a restart
**`authz.MemoryNonceStore` on a real host is a defect** — replay protection dies on restart. Use
`authz.FileNonceStore`. The memory store exists for tests.
## The token is a hash on disk, plaintext only at mint
The store keeps **hashes**. The plaintext token exists in exactly one place, `bootstrap.json` on the
PVE host — so a "read the token" step means reading that file, and a lost token is re-minted, never
recovered.
## Binds can brick guest boot
| Do not | Because | Use |
|---|---|---|
| `GuestBinder.AttachBind`/`DetachBind` (per-drive `pct set -mpN`) | legacy model; a missing bind source can **brick guest boot** (C1) | `AttachDrive`/`DetachDrive` (intermediary model) |
| `isHostMountpoint` to reconcile bind state | a boolean cannot converge stacked double-binds (the `/mnt` doubling bug) | `countHostMounts` normalization inside `AttachDrive` |
<!--
Why fail-safe-to-protected rather than fail-closed: this API also carries the recovery paths. A hard
refusal on an unresolvable caller would make a half-broken guest unrecoverable through the very
interface built to recover it. Less access, never none.
-->
+44
View File
@@ -0,0 +1,44 @@
---
paths: ["internal/proxmox/**", "internal/reconcile/**", "internal/signedjobs/**"]
---
# Proxmox — the API contract, and how destructive work is gated
`internal/proxmox/` is the API-first `Client` plus the fenced root-CLI `Privileged`.
`internal/reconcile/` is the reconcile engine, reversibility gate, op journal and crash recovery.
`internal/signedjobs/` holds the operator-signed destructive executors (wipe, decommission).
## A 200 on the POST is not success
**Every mutating op is async**: it returns a **UPID**, and `WaitTask` must assert
`exitstatus == "OK"`. Authorization can fail at *task execution* long after the HTTP call returned
200. Treating the POST's status as the result is how a failed destroy reads as a successful one.
## The privsep token gotcha
A `--privsep 1` token's rights are the **intersection** of the backing user's permissions **and** the
token's own ACLs. The role must be granted on **both** or every call 403s. The same intersection rule
bites on PBS (`token ∩ user`).
## TLS
**SHA-256 leaf-cert pinning** against the self-signed host cert. **No insecure default**, ever. The
pin is the raw leaf-DER sha — the SAN is never checked, so a cert rotation changes the pin and the
agent must be re-pinned.
## The destructive path — never the direct call
| Do not | Because | Use |
|---|---|---|
| `Client.DestroyLXC` / `Vzdump` / `SetConfig` ad-hoc | skips classification, signature, per-guest serialization, crash recovery | `reconcile.Engine` paths / `RunSignedJob`; queue via `Queue.Submit` |
| add a method to `proxmox.Privileged` | breaks the 3-exception root-CLI fence (`routing_test.go`) | `proxmox.Runner` + a new sudoers `Cmnd_Alias` + `validate.go`-style checks |
| treat `ListLXC` output as "guests we own" | audit A1 — pre-v0.62.0 the stale-lock reaper did exactly this, contained only by the pool-scoped token | intersect with `Client.Pool` membership (`staleLockController.Guests()`); **fail safe on read failure** |
Full trap table: `REUSE.md` §3. Every guest joins the `felhom` pool — `VM.Audit` comes from the
`/pool` grant, not from a per-guest ACL.
<!--
The fence is not stylistic. It is what makes this component auditable: two types, one of which can
only speak HTTP and one of which can only shell out, with a test asserting neither crosses. A single
convenience method on Privileged that also makes an HTTP call would end that property silently.
-->
+49
View File
@@ -0,0 +1,49 @@
---
paths: ["internal/storage/**", "internal/escrow/**"]
---
# Storage and escrow — format safety and zero-knowledge recovery
`internal/storage/` is the storage observer, durable IDs, role/claim classifiers, `SudoHostOps` and
the watchdog. `internal/escrow/` is the PBS-key escrow with its zero-knowledge recovery code.
> **Overlap note:** `health-checks.md` also matches `internal/storage/**`. Deliberate — both rules
> apply there and both load.
## Never format the device you inspected
**AGENT-001 is a TOCTOU:** acting on the caller's `req.Device` (or any remembered `/dev` path) after
inspection lets `/dev` re-enumeration retarget the node to a **different physical disk**. Format the
**re-resolved** device — `Server.reresolveWipe` / `reresolveBlank`.
**Never exec raw `mkfs.*`** (including `Binaries.MkfsExt4`/`MkfsXfs`): sudoers no longer allowlists
raw mkfs, and going direct bypasses the claim filter and the wrapper's re-checks. Use
`SudoHostOps.Format`, which routes through `felhom-mkfs-guarded`.
## The two durable-ID schemes refuse each other
They are not interchangeable, and each returns a `binding_mismatch` for the other's scheme:
| Purpose | Scheme | Resolver |
|---|---|---|
| wipe confirmation | `byid:` / `byuuid:` | `ResolveDurableDevice`, `DiskInfo.WipeDurableID` |
| enrolled-storage remount | `uuid:` | `ResolveStorageDevice` |
Using `DiskInfo.DurableID` (a `uuid:`) as a wipe-confirmation id is F20-BUG2.
## Drive data is never taken by force
Plain `umount` only — **never `-l`, never `-f`**, and never any format operation under
`/mnt/felhom-drives`.
## Escrow is zero-knowledge, and a fetch failure is not a wrong code
The server holds no client key; a no-key restore fails with `missing key`. **A fetch failure must
never be reported as a wrong recovery code** — that told a customer their correct code was bad, in
hundredths of a second, when checking a code actually takes about one. Distinguish "we could not
reach the store" from "the code did not match", always.
<!--
The escrow recovery-code "flake" was a REAL defect, not a flake. "Known flake, re-run" needs evidence
before it is said out loud — that phrase cost this project a real finding once.
-->
+75
View File
@@ -0,0 +1,75 @@
---
unconditional: true
---
# Unprompted work — rules for any session without a task file
> Goal sessions, nightly sessions, "work the register" sessions. **A session that starts from
> `/goal` or a standing brief inherits these rules exactly as it inherits the gates.** They are the
> part of `PROMPT-TEMPLATE.md` that a task file used to carry and a goal does not. Same wording lives
> in `felhom.eu`, `felhom-controller`, `felhom-agent` and `app-catalog-felhom.eu` `.claude/rules/`, and in the workspace
> root's unversioned `.claude/rules/`; change all five or none.
## 1. What you may pick up on your own
- A register row **you or another CC session filed**, with owner CC, at P3 or a bounded P2, that
needs **no operator decision**, touches **no customer data by design**, and introduces **no
mechanism nobody has measured**. Smallest first.
- A defect you find while exercising the product, filed as a row **before** you fix it — **unless it is small**:
a small finding is fixed in the session and never filed (the size rule, `OPEN-ITEMS.md` „How a row is filed").
- Hygiene: register compression, stale citations, rows with no owner, documents that contradict
live source.
**Not yours, ever, without a task file or an operator word:** money; anything that changes risk to
customer data; anything that changes a promise the product makes to a customer; anything that
reverses a documented design decision (`documentation/architecture/` — a design decision is not a
defect, R-370); anything on DooPlex or ep0; baking or vouching a golden; promoting a
catalog version; a new external dependency; **a hub image build or hub deploy in a session the operator does not
attend** (operator ruling 2026-10-07, `09` §3 decision 162).
## 2. When you may decide instead of ask (operator grant, 2026-09-14)
You may take a decision yourself when **all** of these hold: the architecture folder and the register
give a clear direction; your choice follows that direction; it is reversible without customer-data
risk; and you can write it in the `09-update-architecture.md` §3 shape — one answerable sentence, the
options, what each costs, why this one. **Then record it** as a dated decision in `CONTEXT.md` and
the owning architecture document, tagged *decided by CC unattended — operator may reverse*, and put
it **first** in the morning note. A decision you cannot write in that shape is one you do not take.
## 3. The discipline a task file used to carry
1. **Baselines first.** Read each repo's `main` hash and version from live source before touching it.
2. **Read the architecture document for the area, and name it** in the report, before any claim.
3. **Red-proof every correctness fix.** A test never seen failing has not been shown to test anything.
4. **Live-validate on a Tier-0 box** through the endpoints the UI invokes. `demo-hp` is `ssh hp`.
Throwaway apps only; the standing apps and `bentopdf` stay.
5. **Evidence off the machine at the end of each phase**, before any revert (R-320).
6. **One release per repo per session**, with a CHANGELOG entry (controller: with its `MinAgent`
line), REPORT overwritten, floor raised to deliver it. **No golden unless a drill or fresh install
needs one** (the waiver, R-468). **No `--no-verify`.**
7. **An enumerated gap becomes a row in the same session — or, if it is small, is fixed in it** (the size rule).
Prose is not a record.
8. **Hungarian text is searched with ASCII fragments**, with a positive and a negative control.
9. **Never leave a half-state.** If time runs out, revert to clean and say what was reverted.
10. **Teardown, three layers, stated** — machine, host, hub — or "provisioned nothing".
11. **Every helper prompt carries the brief's fences in full** (operator ruling 2026-10-07). A helper session (a
subagent, a fork, a workflow agent) gets the brief's fence list word for word — every protected machine, every
„no", every delivery and Docker limit — not a summary and not „the usual fences". Earned on 2026-10-06 night: two
helpers whose prompts carried only part of the fences ran `docker volume prune` on the bench and a Docker-using
gate on DooPlex.
## 4. The morning note
One screen, plain language, in this order: **decisions you took** (§2) first; what you exercised;
what broke and whether you fixed it; rows opened and closed with the register size before and after;
what needs the operator, each with what happens if they do nothing. No file paths, no function
names, no row numbers as the subject of a sentence.
## 5. Instruction files
**Instruction files (`CLAUDE.md`, `.claude/rules/*`) are kept true by the session that finds them wrong**
(operator ruling 2026-10-06, `09` §3 decision 150). A session MAY, without asking: correct a stale fact (a command, a
count, a version, a path, a description of what a gate does), add a fact it proved, and remove a reference to something
that no longer exists. Each edit is named in the report (file, line, before, after, why). A session MAY NOT, without the
operator's word: loosen a safety rule, a fence, a „never", a protected machine, a secret rule, or a review step; or
remove a rule. When in doubt, it is a rule change, and it goes to the operator. If Claude Code's own permission check
asks before such an edit, wait for the operator's click; if it refuses, record that and file the exact line.
+35
View File
@@ -29,6 +29,41 @@ root=$(git rev-parse --show-toplevel 2>/dev/null) || {
}
cd "$root" || exit 1
# ── WORKSPACE-ROOT ASSERTION (2026-08-05, R-204 rider) ───────────────────────────────────────────
# Refuse a push from a clone outside the felhom workspace.
#
# WHY THIS IS A HOOK AND NOT A LINE IN A DOCUMENT: the workspace root is ALREADY written down, in
# documentation/runbooks/workspace-CLAUDE.md and in the workspace-root CLAUDE.md ("stay inside it"),
# and work drifted into a home directory anyway. A rule that has failed once as a reminder is not
# fixed by writing it down again — it has to be asserted where it can bite.
#
# A PUSH IS THE RIGHT TRIGGER, deliberately: throwaway clones under /tmp for probes and red-proofs
# never push, so nothing legitimate breaks. Reads and builds elsewhere stay unaffected.
#
# Symlinks are resolved on BOTH sides before comparison, so a symlinked path neither falsely passes
# nor falsely fails. If the workspace root does not exist on this machine the check is SKIPPED, not
# failed — this hook must not brick a legitimate clone on a different host.
#
# The only bypass is the documented `git push --no-verify`, whose use is already reportable.
FELHOM_WORKSPACE_ROOT=/mnt/5_hdd/felhom.eu
if [ -d "$FELHOM_WORKSPACE_ROOT" ]; then
ws_real=$(cd "$FELHOM_WORKSPACE_ROOT" 2>/dev/null && pwd -P) || ws_real=""
root_real=$(pwd -P) || root_real=""
if [ -n "$ws_real" ] && [ -n "$root_real" ]; then
case "$root_real/" in
"$ws_real"/*) : ;; # inside the workspace — proceed
*)
echo "pre-push: PUSH REFUSED - this clone is OUTSIDE the felhom workspace." >&2
echo " clone: $root_real" >&2
echo " expected: under $ws_real (repos live in $ws_real/git/<repo>)" >&2
echo " Work in the workspace clone, or bypass with 'git push --no-verify'" >&2
echo " and state that you did in the session report." >&2
exit 1
;;
esac
fi
fi
if ! command -v python3 >/dev/null 2>&1; then
echo "pre-push: FAIL - python3 not found, so the gates CANNOT run. This is a failure, never a" >&2
echo " pass by default. Install python3, or push with --no-verify and say so." >&2
+3
View File
@@ -9,3 +9,6 @@
# go
/vendor/
# Python bytecode written by configs/test_felhom_os_apply.py
configs/__pycache__/
+1263
View File
File diff suppressed because it is too large Load Diff
+88 -166
View File
@@ -1,189 +1,111 @@
# CLAUDE.md — `felhom-agent`
> Loads when Claude Code touches this repo. Stable orientation only — **current state lives in
> `CONTEXT.md` and the top of `CHANGELOG.md`**, never here. Cross-repo orientation: workspace-root
> `/mnt/5_hdd/felhom.eu/git/CLAUDE.md`.
> Stable orientation only — **current state lives in `CONTEXT.md` and the top of `CHANGELOG.md`**,
> never here. Cross-repo conventions (artifact taxonomy, access, clean-tree gate, secrets,
> CHANGELOG/REPORT): workspace-root `/mnt/5_hdd/felhom.eu/git/CLAUDE.md`. Path-scoped detail:
> `.claude/rules/`.
## What this repo is
`felhom-agent` is the operator-tier **host agent** that runs on each Proxmox host and owns **all**
Proxmox interaction: provision/restore guests, host storage, backup/restore orchestration, the hub
control loop, and a narrow per-guest local API. It is the **most privilege-sensitive** component.
The operator-tier **host agent**, one per Proxmox host, owning **all** Proxmox interaction:
provision/restore guests, host storage, backup/restore orchestration, the hub control loop, and a
narrow per-guest local API. It is the **most privilege-sensitive component in the system**.
- Renamed former `proxmox-controller` repo.
- **Distinct from `felhom-controller`** — that is the *in-guest* controller (Docker-only, no Proxmox
creds). Do not confuse them.
- Control plane, not data plane: if the agent dies, apps keep serving; only management degrades.
- Renamed from `proxmox-controller`.
- **Distinct from `felhom-controller`** — that is the *in-guest* controller, Docker-only, holding no
Proxmox credentials. Do not confuse them.
- **Control plane, not data plane:** if the agent dies, apps keep serving; only management degrades.
- Pure Go stdlib + `golang.org/x/crypto`. No web frameworks.
## Read before writing code
## Doing X → read Y
- **`REUSE.md`** — canonical helpers, format-safety guards, traps, seams. Check it first; update it
in the same commit that changes a shared helper or pattern.
- `CONTEXT.md` (current state + open threads) and the top `CHANGELOG.md` entry (authoritative history).
- Design doc: `felhom.eu/documentation/architecture/03-host-agent.md` (locked). Platform facts:
`felhom.eu/documentation/proxmox-platform.md` + `tests/phase{0,1-2,3,4}-findings.md`.
| Doing | Read |
|---|---|
| writing any new code | `REUSE.md` — helpers, format-safety guards, traps, seams |
| needing current state / open threads | `CONTEXT.md` + the top `CHANGELOG.md` entry |
| Proxmox, reconcile or signed jobs | loads itself: `.claude/rules/proxmox.md` |
| local API, authz or guest hooks | loads itself: `.claude/rules/localapi.md` |
| backup, PBS or DR | loads itself: `.claude/rules/backup.md` |
| storage or escrow | loads itself: `.claude/rules/storage.md` |
| writing a health check | loads itself: `.claude/rules/health-checks.md` |
| **release, build, publish, deploy, verify a version** | the **`felhom-build-deploy`** skill — **never hand-roll it** |
| writing or reviewing a test, fixing a bug | the **`felhom-testing`** skill |
| host addresses, break-glass, node facts | `felhom.eu/documentation/operations/nodes.md` — never restate them |
| which box may I break | `felhom.eu/documentation/runbooks/target-selection.md` |
| what version is live anywhere | ask the hub (`/hosts`, `/configs`) or the box — **never a doc** |
| the authoritative design | `felhom.eu/documentation/architecture/03-host-agent.md` (locked) |
## Layout (verified against the tree)
## The root-CLI fence — API-first, exactly three exceptions
```
cmd/felhom-agent/ main + flags + --selftest modes + the daemon entry
cmd/felhom-opsign/ offline operator signing CLI (SSHSIG)
internal/authz/ operator signed-op verifier (SSHSIG) + durable FileNonceStore
internal/backup/ vzdump backup runner + restore-test scheduler + report store
internal/capability/ live sudo-policy capability probe (degradation visibility)
internal/config/ JSON config + FELHOM_AGENT_* env overlay; secrets redacted (Redacted())
internal/desired/ hub desired-state syncer (envelope observer)
internal/escrow/ PBS-key escrow (zero-knowledge recovery code)
internal/guesthook/ pre-start self-heal hookscript install
internal/hub/ daemon: HostReport collector + Bearer client + resilient Loop
internal/lanresolver/ split-horizon DNS on guest IP change (dnsmasq RESTART, not reload)
internal/localapi/ per-guest local API: token store, disks/format, guest binds, controller swap,
stale-lock recovery, pinned self-signed leaf
internal/log/ slog setup
internal/pbs/ PBS-API client (fingerprint-pinned) + verify maintenance loop
internal/provision/ guest bootstrap back-half (token mint → bootstrap.json → pct bind)
internal/proxmox/ API-first Client + fenced root-CLI Privileged + UPID WaitTask
internal/reconcile/ reconcile engine + reversibility gate + op journal + crash recovery
internal/signedjobs/ operator-signed destructive executors (wipe, decommission)
internal/storage/ storage observer + durable ids + role/claim classifiers + SudoHostOps + watchdog
```
## Build / run
- Module `gitea.dooplex.hu/admin/felhom-agent`; binary `felhom-agent` (`cmd/felhom-agent/`).
- **Pure Go stdlib + `golang.org/x/crypto` only** — no web frameworks. `go.mod` directive go 1.25.0;
DooPlex (192.168.0.180, where CC runs) has the Go toolchain and is on the same LAN as the demo
host — build and run live tests locally.
- Version via `-ldflags "-X main.version=<v>"`; `--version` flag. Bump on meaningful changes + CHANGELOG entry.
- **Full build/deploy/publish runbook: use the `felhom-build-deploy` skill.** Summary:
> **Clean-tree gate before any build:** `git status --porcelain` must be empty and
> `git rev-parse HEAD` must equal `git rev-parse origin/main` in the repo being built. An unpushed
> change does not exist — never build a dirty or unpushed tree. The `git pull` in the build step
> stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from
> elsewhere).
> **RELEASING IS ONE COMMAND, AND IT PUBLISHES (R-115).** There used to be a raw `go build` line
> here and a *separate* "Publish" row, so publishing was a step someone had to remember — and it was
> **forgotten three times in five days**, the last leaving agent v0.120.0 deployed on both demo hosts
> and undownloadable, where a documented-path reinstall would have silently downgraded them while
> reporting success. Do not hand-roll the build: the script also creates the `v<version>` git TAG
> that `felhom-host-install.sh` fetches this version's sixteen config files from (R-183), and it
> verifies by an **independent download** rather than trusting the publish step's own output.
> `scripts/publish-agent.sh` still exists and is still correct — the release script CALLS it rather
> than reimplementing it.
| Step | Where | One-liner |
|---|---|---|
| **Release** (build + tag + publish + verify) | DooPlex (local) | `GITEA_USER=admin GITEA_TOKEN=<tok> scripts/release-agent.sh <ver>` — refuses a dirty/unpushed tree and refuses to re-release an existing version |
| Copy | local → felhom-pve | `scp /tmp/felhom-agent-<v> felhom-pve:/tmp/` (one hop) |
| Deploy | felhom-pve | backup `.bak-<old>` → `install -m0755` → `systemctl restart felhom-agent` (non-root `felhom-agent` user, config `/etc/felhom-agent/agent.json`) |
| Ship configs | felhom-pve | sudoers (`/etc/sudoers.d/felhom-agent`) + guarded-mkfs wrapper WITH the binary when `configs/` changed |
| **Vouch** | hub operator UI | Configs → Day-0 artifacts. **Deliberately NOT automated** — vouching is what points machines at a version, and it stays your act (prove-then-vouch) |
| Verify | felhom-pve | `felhom-agent --version` + journal (clean ReassertGuestBinds, no capability degradation) |
## Proxmox model (the load-bearing rules)
This is in the core because breaching it is how this component stops being auditable.
- **API-first** via a scoped `FelhomAgent` token. Raw root-CLI is **fenced to exactly 3 exceptions**:
keyctl `pct create` (golden image), USB mount/fstab, SMART/sensors. `Client` never shells out;
`Privileged` never makes HTTP calls (asserted by `routing_test.go`). Keep that fence.
- **Every mutating op is async** → returns a UPID → `WaitTask` asserts `exitstatus == "OK"`. A 200 on
the POST is **not** success; authorization can fail at task execution.
- **TLS:** SHA-256 leaf-cert pinning (self-signed host cert). No insecure default.
- **Privsep token gotcha:** a `--privsep 1` token's rights = intersection of the backing user's perms
AND the token's ACLs — the role must be granted on **both**, or every call 403s.
- Destructive ops go through the reconcile gate / signed-jobs path — never call `Client.DestroyLXC`/
`Vzdump`/`SetConfig` ad-hoc (REUSE.md §3).
keyctl `pct create` (golden image), USB mount/fstab, SMART/sensors.
- **`Client` never shells out; `Privileged` never makes HTTP calls** — asserted by `routing_test.go`.
Adding a method to `proxmox.Privileged` breaks the fence; use `proxmox.Runner` plus a new sudoers
`Cmnd_Alias` and `validate.go`-style checks (`REUSE.md` §3).
- **Destructive ops go through the reconcile gate / signed-jobs path.** Never call
`Client.DestroyLXC` / `Vzdump` / `SetConfig` ad-hoc — that skips classification, signature,
per-guest serialization and crash recovery.
- **Ownership must be PROVEN, never assumed.** A raw `ListLXC` list is not "guests the agent owns";
intersect with `Client.Pool` membership and fail safe on a read failure (audit A1).
## Demo host (for live tests)
## Gates — ONE entry point
Node **`demo-felhom`**, API `https://192.168.0.162:8006`. SSH alias `felhom-pve` (root@pam) —
available to CC as plain `ssh felhom-pve`. A **second demo node `demo-hp`** (HP t740, node name
`felhom-host`, `ssh demo-hp` — no baked key; break-glass root via hub `host_recovery/demo-hp-bb76ea` +
`sshpass`) is the **designated drill+build VM host** per the 2026-07-25 operator ruling, and that ruling
is **realized** — it hosts drill VM `300` (`drill-r50`), so **start there**, not on DooPlex. (The
historical golden-bake `drill.qcow2` still lives on DooPlex and is a bake fixture, not a drill target.)
**Which box is safe to break, and what may be done to each:
`felhom.eu/documentation/runbooks/target-selection.md`** — read it before any destructive test. Both
nodes + the break-glass recipe: `felhom.eu/documentation/operations/nodes.md`. The agent pins the served leaf cert — verify the
fingerprint still matches before a live run. Selftest modes (run locally on DooPlex, pointed at the
demo API): `--selftest[=read|task|hub|storage|backup|restore-test|pbs-verify]`; no flag = the daemon.
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs every
gate in its `GATES` table (that table is the list); the shared ones — `reuse_refs_check`,
`instructions_gate`, `observations_gate` — are the copies in `felhom.eu/scripts/`, never copied into
this repo (a copy recreates the drift they detect; an absent sibling clone FAILS). `--fast` selects the
gates touching no network and no container runtime, and skips `published` (network), naming it. **A missing gate is a FAILURE, never a skip.**
> **TEMPORARY — felhom-pve is at a remote site (until ~2026-08-02).** The home-LAN literal
> `192.168.0.162` is NOT reachable from DooPlex for the duration. Access via Tailscale:
> felhom-pve = 100.70.170.35; the `Host felhom-pve` entry in `~/.ssh/config` on DooPlex already
> points there (the direct-LAN path stays available as `Host felhom-pve-lan`). Delete this block on
> return. All documented `ssh felhom-pve` / `pct exec` workflows are unchanged. Path is **direct**
> (not DERP), ~37 ms rtt per hop. At the remote site the host is on **DHCP**; re-check its address
> rather than trusting one written here (`ip -br addr show vmbr0` — it read `192.168.0.162/24` on
> 2026-07-30, and `felhom-pve-lan` from DooPlex is still `No route to host`). Details + findings:
> `felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md`
>
> **The "agent does not run at the remote site" warning this block used to carry is RETRACTED
> (2026-07-30) — it was true before R-50 and is false now.** `localapi` no longer binds a LAN literal:
> since the R-50 island migration (2026-07-25) it binds `169.254.253.1:8443` on `vmbr9`, which is
> location-independent by design, and `proxmox.endpoint` is `https://127.0.0.1:8006`. Verified live:
> `systemctl is-active felhom-agent` → `active`, `felhom-agent --version` → 0.115.0, and the per-guest
> local API answered `GET /disks` over the island. No config edit and no Viktor GO are outstanding.
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is
**per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS
when this clone is unarmed. `git push --no-verify` bypasses it deliberately; **say so in the session
report when you use it** — CI re-runs the same entry point on every push and **emails the operator on
failure**, so a bypass is noticed even though it is not blocked (R-168, CLOSED 2026-08-02).
> **Legacy: Windows workstation.** Until 2026-07-19 CC ran on Windows 11; `pct` commands over SSH
> needed `export MSYS_NO_PATHCONV=1`, and every remote command used
> `SSH=/c/Windows/System32/OpenSSH/ssh.exe`. Agent deploy was a two-hop copy via the Windows box
> (`cygpath -w` for the local scp path; CRLF hazard on config files).
<!--
WHY ONE ENTRY POINT (2026-08-02, R-29): a census of all gates across the four repos found every check
a CLAUDE.md names was passing, and two of the four nobody is told to run were failing. This repo was
the extreme case — nothing ran against it at all, and 90 cited paths were checked by no one.
-->
## Live validation — the fence
Exercise the **SERVER-SIDE PIPELINE** a real user triggers, end-to-end. **The forbidden shortcut is
BYPASSING it** — the F9 episode was a raw guest-attach with hand-set state, and it proved nothing.
`claude-in-chrome` is NOT available on DooPlex. Invoking the exact endpoint the UI invokes is an
acceptable proxy — **say which method was used**. Low-level mechanism tests where the direct call IS
the mechanism are exempt.
## Conventions
### Trunk-based — no branches
All shippable work commits **directly to `main`**; `main` equals what is deployed.
- Report-only artifacts (audits, findings, fixspecs) → `felhom.eu/documentation/` (`audits/`, `backlog/`).
- Risky/supervised fixes are spec'd, then implemented **during the supervised session, on `main`**.
- Unattended escape hatch: if a fix can't be cleanly verified/shipped, revert + report — never park on a branch.
> **In every repository where you make a change, update both files in that repo:**
> - **`CHANGELOG.md`** — cumulative log, newest on top.
> - **`REPORT.md`** — **overwrite** with the most recent implementation/validation summary only.
>
> **Never write secrets** into any committed file — reference them as "stored out-of-band".
- Code quality: verify generated code for bugs/edge cases; add debug logging; **ask rather than
guess** when you'd otherwise invent input/output.
- **A health check issues no block I/O** — no `statfs`, no `getdents`, no read, write or `fsync`, **not
even behind a timeout**. Liveness is decided from `/proc` and kernel state. The full rule + the
measurement lives in `felhom.eu/CLAUDE.md` "Code quality rules"; it is repeated here because health
checks are written in THIS repo and that file does not load in an agent-only session. R-117 spike §6.3.
- Update `REUSE.md` if you added/changed/deprecated a shared helper or pattern (same commit).
- **Run `python3 scripts/agent_gates.py` from the repo root after ANY change in this repo.** It is
the ONE entry point for this repo's gates. Today it runs one — `reuse_refs_check` over this
repo's `REUSE.md` — and it exists at one gate on purpose: a census on 2026-08-02 found that every
check a `CLAUDE.md` names was passing and two of the four nobody is told to run were failing, and
this repo was the extreme case, with nothing running against it at all and 90 cited paths checked
by no one. It grows when the agent grows a second check. `--fast` selects the gates that touch no
network and no container runtime; today that is all of them. A missing gate is a FAILURE, never a
skip. **The shared `reuse_refs_check.py` lives in `felhom.eu/scripts/` and is never copied here**
— a copy would recreate the drift it detects; an absent sibling clone FAILS the gate.
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It
is per-clone — switch it on once with `git config core.hooksPath .githooks`, and a manual run
WARNS when this clone is unarmed. `git push --no-verify` bypasses it deliberately; **say so in the
session report when you use it.** Both facts are why CI is still owed (`OPEN-ITEMS.md` R-168).
- Testing doctrine (non-hollow tests, red-proofs, seams): use the `felhom-testing` skill.
- **Logging**: the slog logger fans out to journald (configured level) + the always-DEBUG `applog.Ring`
(remote pulls) — English, keys-never-values, durations on outcomes; full rules in
- **Trunk-based — no branches.** All shippable work commits directly to `main`; `main` equals what is
deployed. Report-only artifacts (audits, findings, fixspecs) go to `felhom.eu/documentation/`.
- **Unattended escape hatch:** if a fix cannot be cleanly verified and shipped, **revert and report**
— never park it on a branch.
- **Logging**: the slog logger fans out to journald (configured level) plus the always-DEBUG
`applog.Ring` (remote pulls). English, keys-never-values, durations on outcomes. Full rules:
`felhom.eu/documentation/runbooks/logging-conventions.md`.
- Update `REUSE.md` in the same commit that adds, changes or deprecates a shared helper or pattern.
### Live validation
## End-of-session checklist
Exercise the SERVER-SIDE PIPELINE a real user triggers, end-to-end. The forbidden shortcut is
BYPASSING it (the F9 episode: raw guest-attach + hand-set state). Invoking the exact endpoint the UI
invokes is an acceptable proxy when a browser isn't available — say which method was used. Low-level
mechanism tests where the direct call IS the mechanism are exempt.
- **`CHANGELOG.md`** (cumulative, newest on top) and **`REPORT.md`** (overwritten with this run only)
— in every repo touched.
- **`CONTEXT.md`** — decisions, state, what is next.
- **`REUSE.md`** — if a shared helper or pattern moved.
- **A finding goes in `felhom.eu/documentation/backlog/OPEN-ITEMS.md` first**, never only in a report
or an audit.
- **Confirm your own last push's CI run went green, by run ID** — CI mails on failure, which is a PUSH
signal; this is the PULL check that catches a lost or unread mail. An unchecked green is an
assumption, not an observation.
## Workflow & artifacts
- Implement **`TASK.md` / `TASK-*.md`** specs (when placed as `TASK.md` or told to), then push +
CHANGELOG + REPORT.md.
- **`RUNBOOK-*.md`** — an operational procedure. CC executes the steps it has access and capability
for, including live validation on the demo Proxmox host (CC has root@felhom-pve SSH + the
felhom-agent token). Mark a step HUMAN only when it genuinely needs physical presence, a real-world
decision, or credentials CC truly lacks. Judgment still applies: confirm before irreversible ops on
real customer data — demo scratch guests are fair game.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `felhom.eu/scripts/decoy_coverage_gate.py` (run by felhom.eu's `repo_gates.py`, for all four repos) refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list.
+120
View File
@@ -1,10 +1,122 @@
# CONTEXT — felhom-agent working state
> **2026-10-04 night — v0.143.0 RELEASED + vouched (R-840, decision 96): the config bundle.** `felhom-os-apply` mode
> `bundle` (signed `agent_config_update`, verified by the wrapper itself; trust files never bundle paths) +
> `--install-bundle` (installer 1.31.0); `BUNDLE_FILES` is the one table; `scripts/build-config-bundle.py`;
> `release-agent.sh` publishes it. A box whose `felhom-os-apply` predates 0.143.0 needs ONE by-hand bootstrap
> (`felhom.eu/scripts/felhom-bundle-bootstrap.sh`) — done on both demo boxes; Tester 2 waits for the operator (R-862).
> Both demo boxes: agent 0.143.0, bundle 0.143.0 (record `/etc/felhom/config-bundle.json`). R-861 found: the sudoers is
> root-equivalent. `build-golden.sh` 3.2.0 (`GOLDEN_GUEST_PKGS`). Runbook `felhom.eu/documentation/runbooks/config-bundle.md`.
> **2026-09-25 night — v0.133.0 AND v0.134.0 DELIVERED to both demo boxes (CC-signed `agent_update`, ruling 1);
> restore test back ON (the `-1` config kept as `agent.json.night-0925-off`). v0.134.0 = R-685:** `backup/runner.go`
> `spaceFits` — a vzdump to a LOCAL target needs free ≥ newest archive of that guest × 1.25 + 1 GiB (PVE prunes
> only after success); a shortfall is a named skip (`BackupSkipNoSpacePrefix`), fail-open on PBS / first backup /
> unknown usage. **Free space comes from `NodeStorage` (`GET /nodes/<n>/storage`) — `ListStorage` (`GET /storage`)
> has NO usage**; the first build read it and a real vzdump started in its own live test (aborted, no archive).
> demo-hp: `local_backup_retention` 1 (operator option A, saved `agent.json.pre-a4-retention`). Peti's box: nothing.
> **2026-09-24 — v0.133.0 RELEASED, NOT DELIVERED (R-672, R-673).** Restore-test space preflight
> (`reconcile/restoretest_space.go`, `internal/restorespace`): uncompressed size from the vzdump log / PBS size,
> × 1.2 + 5 GiB, thin metadata, off the tested guest's pool, unknown refuses, reported as `skipped` non-pass.
> Janitor (`cmd/felhom-agent/janitor.go`) every 10 min: `Engine.RetryScratchTeardown` + stale-lock sweep under
> the heavy-op gate. Thin pool ≥ 90 % → immediate report (hub v0.124.0 alarms). **Operator rulings 2026-09-24
> (evening):** the scheduled restore-test is OFF on both demo hosts (`backup.restore_test_eval_interval_seconds:
> -1` — note: 0 means the 6 h DEFAULT, only a negative disables) until v0.133.0 is delivered there; saved configs
> `/etc/felhom-agent/agent.json.pre-r672`. demo-hp 9201 was repaired (stop, fsck, start; two Redis AOF tails cut).
> Snapshot of the current state + open threads. Authoritative history lives in `CHANGELOG.md` (top
> entry = current); the end-of-task detail lives in `REPORT.md`.
## R-199 (v0.125.0) — links 6–8 of the recovery chain, assembled and walked
`POST /escrow/recover-offsite-password` (pinned local API, `withGuest`): the controller supplies the
customer's recovery code, the agent fetches THIS host's own sealed blob from the hub
(`hub.Client.FetchIdentityEscrow` → `GET /hosts/{id}/escrow`, hub >= v0.94.0, self-scoped by the
per-host key), unseals it via `escrow.OffsiteKeyRecoverer`, and returns **only** the offsite restic
repository password plus its sha256.
**Rules that must not erode:**
- **Only that field.** Not the tunnel token, not the PBS token, not the WG key — the controller is a
trust tier down and needs none of them. Narrowing cost nothing and is not recoverable later.
- **The unseal stays in the agent.** `age` is an agent runtime dependency (`/usr/bin/age` — hardcoded,
no config override; 1.2.1 on demo-felhom) and is deliberately absent from the controller image.
- **R:** in memory for one call, cleared on the success path AND every failure path, never on disk,
never in argv, never logged at any level including inside an error, never echoed. Verified live: 0
log lines, 0 files, 0 leftover `felhom-idesc-*` dirs, with a positive control proving the search worked.
- **Three distinct outcomes**, not one generic failure: no blob (404), a bundle that opens but predates
the field (409 — pre-fork-4, cannot be retro-fitted), a code that does not open it (400 — fail-closed
at age's KDF, nothing written).
- **The wiring is pinned by an AST walk** (`cmd/felhom-agent/escrow_recover_wiring_test.go`):
`main` → `runDaemon` → `buildLocalAPIServer`, an `escrow.OffsiteKeyRecoverer` constructed there, the
`Options.EscrowRecovery` field present, and the fetcher calling the DAEMON's own `hubClient` (the
self-scoping that makes cross-host retrieval impossible is a property of WHICH key is used).
Links 6 and 7 were two of this project's six built-but-never-wired instances.
**Proven live on demo-felhom 2026-08-04:** recovered sha256 == on-disk sha256 == the hub's stored hash.
A wrong code five minutes earlier failed closed. **The chain stops at link 8** — nothing installs a
recovered password, reopens a repository, or restores a file.
**§8.6, fixed while here:** `runSelftestIdentityConsume`'s success line used to recite
"tunnel_token + pbs_token", which became a misstatement when v0.77.0 sealed the repository password
into the same bundle — anyone reading it would conclude the password was not there. It now names what
THIS bundle carried and what it did not.
## Current
- **2026-08-03 — v0.123.0 (R-185): a tier the box cannot READ now says so.** The agent's token had
`FelhomAgentStore` on `local`, `local-lvm`, `felhom-pbs` and **not** on `felhom-backup` — the
storage both demo boxes configure as `local_backup_target`. That storage answered `{"data":[]}`
through the token while root listed three archives, and `pickForThisRun` skipped it as *"no settled
archive yet"* — **which is what a brand-new tier reports**, so the host tier was never
restore-testable and nothing said so.
- **The permission question is asked directly**, because unlike the listing it has a definite
answer: `Client.Permissions` reads `/access/permissions?path=/storage/<target>` **as the agent's
own token**, and `storeGrantStatuses` emits one `capability.Status` per configured tier. It
composes AROUND the sudo prober, the way `poolReadStatus` already does — an API read does not
belong inside a sudo-policy probe. `Status`'s wire shape is untouched, so the hub's critical
degraded alert applies with **no hub change**.
- **MEASURED FIRST, and the obvious reading is wrong:** an ungranted path answers neither empty nor
403 — it carries the privileges INHERITED from the box-wide `/` grant
(`Sys.Audit, SDN.Use, Datastore.Audit`). Checking path-presence, or `Datastore.Audit`, reports a
blinded storage HEALTHY. The probe tests **`Datastore.AllocateSpace`**; re-measure before ever
changing that constant (`storeGrantRequiredPriv`, red-proved).
- **The probed set comes from `BackupTiers()`, never a fixed list** — a hardcoded probe list is the
defect reproduced inside the fix. Critical, EXCEPT the `local` fallback target (reported, but it
does not page). It never consults content, so it cannot alarm on a newborn tier; it never reports
ok when it could not ask.
- **LIVE:** degraded observed on the still-blind box (hub emailed `agent_capability_degraded`) →
grant applied on **both** demo boxes → token lists 3 and 4 archives → `ok=70 total=70 degraded=0`
and `degraded → ok` at the hub → **the host tier became a due-check candidate for the first time**,
correctly picking the 08-02 archive (08-03 had not settled 24 h).
- **The installer's real defect was NOT `PVE_STORAGES`** — see `felhom.eu` CONTEXT S-22: Case A
grants, the Scenario-F reuse arm did not. Fixed in installer **1.24.0** with a gate.
- **2026-08-03 — v0.122.0 (R-189 · R-188 · R-186): three signals that lied about their own work.**
None touches data; all three cost attention, which every other signal depends on.
- **R-189 — a passing restore-test no longer vanishes on a restart.** `restore_tests[]` came only
from the in-memory `backup.Store` (*"lost on restart; the cadence re-populates"* — true under a
timer, FALSE since R-86, because the agent will not re-test a proven archive). **Observed live:**
a 14.5 GB offsite PASS at 15:25:14, agent restarted 2 m 43 s later, hub logged `0 restore-tests`
twice. `RestoreTestState` now stores `tier` + `verified` beside the archive (v3 shape; v1/v2
still read, and a record missing archive-or-tier is NOT reported), exposes
`ProvenRestoreTests`, and `Collector.SetProvenRestoreTests` merges it — **one entry per tier,
newest by `TestedAt` wins**, so a fresh failure beats a stored success and a tier never appears
twice. Wiring pinned by an AST test: the method this replaces (`Snapshot`) claimed a
"host-report gauge" in its doc comment and had **no caller** for weeks.
- **ONLY SUCCESSES ARE PERSISTED, and the reason is now in the code:** a success *suppresses*
future work (a proven archive is never re-tested, so a lost proof leaves the box quietly less
tested than it believes); a failure *causes* future work and heals itself at the next evaluation.
- **R-188 — the release stopped emailing false failures.** Only the tag PUSH moved (build → tag
locally → publish → push tag): the push is what wakes CI, and a tag visible before its package
made the gate correctly fail a correct release ~half the time. The old order's invariant is now
asserted directly — `check-published-versions.py` refuses a **published version with no tag**, as
a bounded, printed probe (the package listing api is still 401 without a token, re-measured).
- **R-186 — a released binary is verifiable.** `-trimpath -buildvcs=false`: same source → same
bytes whether or not the tag exists. Measured. `publish-agent.sh`'s fallback also forced
`CGO_ENABLED=0` and built a **74 KB different** binary for the same version — both paths now
identical. The verification command is in `CLAUDE.md`.
- **2026-08-03 — v0.121.0 (R-86): the restore-test follows the BACKUP, not the clock.** The ticker is
now only the **evaluation interval**; a tier is **DUE** when its newest archive that has settled for
`settle` (default 24 h) **has not been proven**. Daily tier → proved daily on yesterday's archive;
@@ -26,6 +138,14 @@
able to make a starting backup record a failure — F-A1), and the candidate picker skips archives
failing `archivePlausiblyComplete` (a phantom would be due forever and fail forever).
- New read-only `--selftest=restore-test-due` prints the per-tier verdict + its cost.
- **v0.121.1 — a quiet evaluation is AUDIBLE.** "Nothing is due" is now the NORMAL outcome, and at
DEBUG it was silent: an empty journal would have been equally consistent with a healthy loop and
a dead goroutine (standing rule 3 — the shape the R-88 watcher was retired for). A not-due
evaluation logs ONE INFO line naming every tier's verdict; an unlistable tier reads `UNKNOWN`
with its error in that same line.
- **PROVEN LIVE 2026-08-03 on demo-felhom:** due-triggered offsite restore-test of a 14.5 GB
encrypted PBS archive — restored, booted, verified, scratch destroyed, **635 s**; the state then
named that archive, a second evaluation ran nothing, and an agent restart ran nothing.
- **R-185 (filed, NOT fixed here):** on demo-felhom the agent token has no ACL on
`/storage/felhom-backup`, so its content listing comes back EMPTY (root sees 3 archives) — the
host tier has never been restore-testable there, and the due-check cannot distinguish that from
+8 -7
View File
@@ -44,13 +44,14 @@ unnoticed until a user hit them. `internal/capability` makes that loud:
(`HostCapabilityChecker`) alerts the operator on a Critical capability going degraded. Serve-degraded
— the probe never blocks startup. (Next self-health slice: the controller↔agent channel check.)
**Controller-swap under non-root (v0.45.0).** The agent-owned controller image swap
(`internal/localapi/controllerswap.go`) no longer shells out: `writeImage` pipes the image ref on
**stdin** into an in-guest `tee /etc/felhom-controller-image` (via `GuestExecStdin` →
`Runner.RunStdin`, the same fenced `sudo -n` runner) — no `bash -c`, no interpolation. Its 5 narrow
grants live in the `FELHOM_CONTROLLERSWAP` sudoers alias (all read-only or fixed-target; the `tee`
target is the FIXED image path, content stdin-fed) and in the capability manifest (Critical), so a
dropped grant is a build failure + a live degraded signal. No general `pct exec` is granted.
**Controller-swap under non-root (v0.45.0; the write since R-861 (a) A1).** The agent-owned controller image swap
(`internal/localapi/controllerswap.go`) no longer shells out. The write goes on **stdin** to the ROOT verb
`felhom-priv-apply controller-image <vmid>` (`GuestBinder.WriteControllerImage` → `Runner.RunStdin`, the same fenced
`sudo -n` runner), which re-checks the ref against our registry + repository + an x.y.z tag and writes
`/etc/felhom-controller-image` inside the guest itself; the agent has no in-guest `tee` grant any more (before, a
compromised agent could feed any image — sudo cannot see stdin). Its grants live in the `FELHOM_CONTROLLERSWAP` sudoers
alias (read-only or fixed-target) and in the capability manifest (Critical), so a dropped grant is a build failure + a
live degraded signal. No general `pct exec` is granted.
## The `storage` package — observe + watchdog (slice 5)
+7 -99
View File
@@ -1,100 +1,8 @@
# REPORT — releasing publishes, and an unreleasable version fails CI (R-115, R-183)
# REPORT — v0.150.0 released and delivered (2026-10-07)
**Date:** 2026-08-03 · **Repo:** `felhom-agent` · **NO VERSION BUMP** — the agent stays **v0.120.0**,
no Go code changed, nothing was built or deployed.
## What changed
| File | |
|---|---|
| `scripts/release-agent.sh` | **new** — THE release path: build → tag → publish → verify by independent download |
| `scripts/check-published-versions.py` | **new** — the R-115 gate |
| `scripts/agent_gates.py` | registers the gate as **not `--fast`** (it needs network) |
| `.gitea/workflows/gates.yml` | CI now runs the **full** gate set, not `--fast` |
| `CLAUDE.md` | the raw `go build` line is replaced by the release script; a **Vouch** row replaces the old Publish row |
## Why
Publishing was a step someone had to remember and was **forgotten three times in five days** —
R-111's seventeen stranded releases, 0.114.0, and 0.120.0, which sat deployed on both demo hosts and
undownloadable, so a documented-path reinstall would have silently downgraded them to the pre-merge
agent **while reporting success**. R-111's own closing line named this leg and closed SHIPPED without
it; it recurred the same afternoon. A note is not a mechanism.
The script also **tags**, because `felhom-host-install.sh` now fetches the agent's sixteen config
files from `raw/tag/v<version>/` (R-183). A released version with no tag 404s a box mid-install, as
root, on a virgin machine. Tag and package are two halves of one release.
It **verifies by downloading what it just published** and comparing the sha to what it built. The
publish step's own success is a report on its own write; a fetch returning the right bytes is a
different claim, and it is the one that matters.
It **does not vouch** — that points machines at a version and stays the operator's act.
## The gate's invariant — not the one specified, and the reason was measured
The task's §8.4 asked for *"the version the hub tells machines to install must be downloadable"*.
**CI cannot see that**, measured rather than assumed (P-C):
| Endpoint | Anonymous |
|---|---|
| Gitea package **download** | **200** (and **404** for a fake version — it discriminates) |
| Gitea **tags** api | **200** |
| Gitea package **listing** api | **401** — token required |
| Hub `/api/v1/artifacts/<customer>` | **401** — per-customer passphrase required |
So a credential-free gate can ask *"is this version installable"* but not *"which version is
vouched"*. Adding an operator credential to CI to close that is the operator's call, not a gate
author's. The implemented invariant — **every `v<semver>` tag must have a downloadable package and a
tag tree that serves the agent's configs** — needs no credential and **catches all three recorded
instances**, because the release script creates the tag and publishes in one act.
**What it does not catch, stated rather than assumed away:** the hub vouching a version that was
never released at all. Nothing here can see that; it belongs at vouch time in the hub. → **R-184**.
## Proof
| Check | Result |
|---|---|
| `go build ./... && go vet ./...` | OK |
| `go test ./...` | **29 packages ok, rc=0** (read separately from any commit) |
| `agent_gates.py --fast` | `published` correctly **SKIPPED** — the pre-push hook must not fail because Gitea blinked |
| `agent_gates.py` (full) | `reuse-refs` OK, `published` OK |
| release script: re-release guard | `ERROR: tag v0.120.0 already exists — releasing over it would make one version name two binaries`, rc=1 |
| release script: clean-tree guard | `ERROR: working tree is dirty — commit and push first`, rc=1 |
### Red-proof F — both directions
- **A tagged-but-unpublished version** (`v9.9.9` created for the purpose): gate **rc=1**,
`binary NOT downloadable (HTTP 404 …)`. This is the R-115 shape exactly.
- **The gate deregistered from the entry point**, same bad state: `agent_gates.py` → **rc=0, "all
agent gates OK"**. Restored → **rc=1, CONVICTED: published**. The guard is what catches it, not
something else.
### Scenario F measured on REAL CI, not inferred
Runs **69** and **70** are on the **same commit** `0db7766`:
| run | state of the repo | CI |
|---|---|---|
| 69 | no `v9.9.9` | **success** |
| 70 | `v9.9.9` tagged, not published | **failure** |
Same code, same workflow, one variable — so the gate demonstrably RUNS in CI and fails for exactly
the R-115 condition. This also retrospectively explains runs 67/68, which were red in the window when
`v9.9.9` first existed. **One deliberate CI failure e-mail reached the operator — that was this
proof, not an incident.**
I could not read CI's own step log to attribute those runs directly: the Gitea jobs endpoint requires
an API token, and the only credential available on this host (`~/.docker/config.json`) is a registry
password, which the API rejects. The controlled before/after above replaced that log rather than an
assumption standing in for it.
`v9.9.9` was deleted afterwards; `git ls-remote --tags` shows only `v0.120.0`.
## Tag convention
`v<semver>`, at the commit the binary was built from. `v0.120.0` was created retroactively at
`cd6e267` — the commit that produced the published binary (sha `a7763d31b55b5ce7…`). `configs/` is
byte-identical between that commit and `main`, so nothing about the sixteen fetched files depends on
the choice; `cd6e267` is tagged because it is the honest one.
On the operator's word (`09` §3 decision 161). `scripts/release-agent.sh 0.150.0`: sha `a23d1c90…`, bundle `88456b38…`,
tag `v0.150.0` = `3a72a48`, verified by download. Vouched with golden 0.301.0 and MinAgent 0.131.0 (unchanged). No bundle
path added (26 → 26), so no step bundle. Signed `agent_update` → demo-hp, demo-felhom, Tester 1 on 0.150.0 (07:06–07:07Z);
signed `agent_config_update` → `BUNDLE DONE written=1 same=24 self-check=ok`, capability probe 68/68 on all three
(07:21Z). Tester 2 not touched. Carries R-528 (the memory-kill check), R-894 (the last backup per tier on disk), R-330
(SMART counters on the wire). Evidence: `felhom.eu/documentation/audits/readback-2026-10-07/delivery/`.
+24 -7
View File
@@ -14,8 +14,15 @@
| `Privileged` (CreateGoldenLXC/MountUSBByUUID/SMART/Sensors) | internal/proxmox/privileged.go | methods on `*Privileged` | the 3 fenced root-CLI exceptions ONLY | Do NOT add methods — fence is structural (`routing_test.go` asserts it) |
| `SudoHostOps.run` | internal/storage/hostops.go | `run(ctx, name, args...) error` | allowlisted exec with stderr-wrapped error | Every arg pre-validated via validate.go before this is called |
| `Prober.Probe` | internal/capability/probe.go | `Probe(ctx) []Status` | live sudo-policy capability check (`sudo -n -l --`) | Needs a DIRECT runner (never the sudo-prefixing one — double-sudo); never executes probed cmds. v0.86.0: config-gated caps (`Capability.GatedBy` + `Prober.GateActive`) report `inactive`/"disabled by configuration" ONLY when healthy — broken plumbing stays degraded; the pbsdr-* gate answers from `pbsdr.Manager.DRConfigured` (marker-backed across restarts) |
| `stageTemp` | internal/localapi/intermediary.go | `stageTemp(pattern, content) (path, err)` | random-named temp before a root `install` (audit B1) | Fixed /tmp names are a TOCTOU — sudoers globs expect `/tmp/felhom-*-*.ext` |
| `guesthook.InstallSnippet` / `Register` | internal/guesthook/install.go | `InstallSnippet(ctx, runner) error` | pre-start self-heal hook install (C1 net) | Same random-temp+install pattern; snippet delegates to the agent binary (no shell logic). Issues `mkdir -p /var/lib/vz/snippets` FIRST (v0.63.0, B2 — fresh boxes lack the dir; sudoers grants exactly that argv) |
| ~~`stageTemp`~~ (REMOVED v0.146.0, R-861) | — | — | — | Nothing the agent writes is `install`ed where root reads it any more: use `felhom-priv-apply` (below) or ship a fixed file in the bundle |
| `felhom-priv-apply` (v0.146.0, R-861) | configs/felhom-priv-apply | `felhom-priv-apply unit <name> \| dnsmasq <tmp> <name> \| wg \| sshd-config \| sshd-key \| controller-image <vmid>` (the last reads the ref on stdin, R-861 (a) A1) | ANY agent-rendered file a root program reads (systemd unit, dnsmasq drop-in, wg-quick conf, OOB sshd) — fixed source + destination, CONTENT checked against the agent's own renderers | A new renderer needs a verb + a contract test (`internal/privapplytest.Check`) feeding its REAL output; never a new `install` sudoers line |
| `privapplytest.Check` | internal/privapplytest/check.go | `Check(t, verb, name, content) string` | the Go↔root-checker contract: a renderer's real output must read `OK` | Skips without python3; one call per rendered shape + one refused control |
| `BUNDLE_FILES` + `Bundle` (mode `bundle`, `--install-bundle`; mode `agent_update` v0.146.0, R-861) | configs/felhom-os-apply | the ONE table of root-owned paths + the installer of them | ANY new root-owned file the installer writes (sudoers line, wrapper, unit) — add it to the table, never a new installer fetch (R-840) | The builder (`scripts/build-config-bundle.py`) and the installer read the same table; `test_every_root_file_the_installer_writes_is_in_the_bundle` fails on a path the bundle lacks. Trust files (`/etc/felhom/os-trust.json`, `operator-signers`) are NEVER bundle paths (R17) |
| `osupdate.ConfigUpdateExecutor` | internal/osupdate/bundle.go | signed op `agent_config_update` {agent_version, bundle_sha256} | delivering the bundle to an installed box | a courier only: the root wrapper re-verifies signature, host, nonce and sha itself |
| `osupdate.Leg.SendUnsent` / `lockPass` (v0.144.0, R-868) | internal/osupdate/unsent.go | `(ctx) int` | an OS-pass report the agent never sent (killed mid-pass): the wrapper keeps `report-<run>-<layer>-apply.json` beside the plan; the agent deletes it once the hub has it | any new caller that runs an apply pass must hold `lockPass` (flock, across processes) — the sender must never take a running pass's copy |
| `osupdate.LoadSavedBlock` (v0.144.0, R-866) | internal/osupdate/leg.go | `(planDir) (block, savedAt, ok)` | the hub's newest os_update block as the daemon last received it (`os-update-block.json`) | the debug pass uses it ONLY when the hub cannot be reached, and says so in its header; no saved block → no pass |
| `dpkg_state()` / `DPKG_STATE_SCRIPT` (v0.145.0, R-876) | configs/felhom-os-apply | `audit, journal = self.dpkg_state()` | dpkg's state in ONE call: `--audit` AND the update journal | never gate a repair on `--audit` alone — a crash leaves only the journal (measured); keep it one call (R-845) |
| `guesthook.SnippetReady` / `Register` (v0.146.0, R-861) | internal/guesthook/install.go | `SnippetReady(path) error` | is the pre-start hook (a FIXED file from the bundle) in place — register only then | The agent never installs the hook (Proxmox runs it as root); a missing hookscript stops a guest start, so never `Register` without `SnippetReady` |
### Disk / format safety (role gates, durable IDs, format guards)
@@ -59,8 +66,8 @@
|---|---|---|---|---|
| `IntentStore` (`Get/SetEnrolled/SetEjected/SetDecommissioned/OnAbsent`) | internal/storage/intent.go | `OpenIntentStore(path)` | drive intent (4-state self-heal) | Keyed by durable-id only; `OnAbsent` is the ONLY ejected→enrolled path; refuses empty ids |
| `GuestBindStore` (`Record/Remove/Guests`) | internal/localapi/guestbindstore.go | `OpenGuestBindStore(path)` | per-guest enrolled binds (F9 re-assert) | Same tmp+rename 0600 pattern as IntentStore |
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) <-chan error` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank |
| `TokenStore.Mint` / `Lookup` | internal/localapi/tokenstore.go | `Mint(vmid) (plaintext, error)` | per-guest local-API tokens | Only the SHA-256 hash persists (fsync'd append log); constant-time compare on lookup; plaintext returned exactly once. Lookup RELOADS the file once on a miss (v0.63.0, B3): the one-shot provisioner mints into the same file the daemon indexes — cross-process coherence without a restart; append-only size check bounds the re-read |
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) (*formatJob, <-chan error)` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank; on success `job.FSUUID` = the new filesystem's UUID, read back only when the durable id still resolves to the formatted device (R-25) — read it only after `done` delivers |
| `TokenStore.Mint` / `Lookup` | internal/localapi/tokenstore.go | `Mint(vmid) (plaintext, error)` | per-guest local-API tokens | Only the SHA-256 hash persists (fsync'd append log); constant-time compare on lookup; plaintext returned exactly once. Lookup stats the file on EVERY call and reloads BEFORE answering when the append-only log grew (R-269; was reload-on-miss only, v0.63.0 B3, which let a token rotated out by another process keep authorizing as a map hit): the one-shot provisioner mints into the same file the daemon indexes — cross-process coherence both ways without a restart; unchanged size = no re-read. Pinned by `TestTokenStore_RotatedOutTokenRejectedFirst` |
| `FileNonceStore.SeenOrRecord` | internal/authz/noncestore.go | `SeenOrRecord(nonce, exp) bool` | durable anti-replay | fsync'd before returning false; prune only after exp |
| `Journal` (`Append/Latest/InFlight/AlreadyApplied`) | internal/reconcile/journal.go | `OpenJournal(path)` | op journal + idempotency + crash recovery | `Recover` consumes `InFlight()`; scratch entries special-cased |
@@ -74,18 +81,23 @@
| `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) |
| `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure |
| `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` |
| `Server.ControllerSupervisorTick` + `ControllerParkedMarker` | internal/localapi/controllersupervisor.go | `ControllerSupervisorTick(ctx)` | R-523: restart a provisioned guest's not-running controller via its bootstrap unit | Two-sweep confirm; honours swapInFlight, the host-side park marker, guest lock + vzdump; 3 restarts/15 min → 30 min pause; record rides the report as `controller_supervisor` (the hub mints the events — the agent has no event channel) |
| `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path |
### Proxmox client / hub / PBS / provisioning
| Symbol | File | Short signature | Use for | Gotchas |
|---|---|---|---|---|
| `pvegate.Write` / `pvegate.Step` | internal/pvegate/pvegate.go | `Write(ctx) (release, waited, err)` / `Step(ctx) (end, err)` | R-812 option A: keep the agent's own /etc/pve writes out of a Proxmox package step (pmxcfs restarts) | Already wired at the two chokepoints — `Client.doBody` (every non-GET) and `ExecRunner.RunStdin` (`WritesEtcPVE`: pct config verbs, pvesm, pveum, felhom-pbs-apply create/reconcile). A new root CLI that writes /etc/pve goes into `WritesEtcPVE`, never its own lock. Never take `Step` around anything but the wrapper call (`Leg.runPVE`) — a `Write` inside a `Step` deadlocks until its context ends. |
| `Client.WaitTask` | internal/proxmox/task.go | `WaitTask(ctx, upid, opts) (TaskStatus, error)` | asserting EVERY mutating op | POST 200 ≠ success; authz can fail at task exec; `AllowWarnings` opt-in |
| `Client.Pool` | internal/proxmox/query.go | `Pool(ctx, name) (PoolInfo, error)` | felhom-pool membership (the ownership registry, A1) | Needs `Pool.Audit` at `/pool/<name>` (host-install v1.9.0+); `Pool.Allocate` does NOT satisfy the read; members can be storages (type `storage`, vmid 0) — filter them |
| `Client` mutate wrappers (`RestoreLXC/Vzdump/DestroyLXC/Snapshot/Rollback/SetConfig/ResizeLXC/Start/Stop`) | internal/proxmox/mutate.go | return `(upid, error)` | all API mutations | Async → always pair with WaitTask; route via gate/queue, not ad-hoc |
| `Client.PoolAddVMID` | internal/proxmox/mutate.go | `PoolAddVMID(ctx, pool, vmid) error` | re-assert pool membership after a restore-over-existing (campaign-2 R2) | SYNC (no UPID, don't WaitTask); PVE `PUT /pools` is additive (merge, not replace) — `delete=1` removes; idempotent (already-member swallowed); needs `Pool.Allocate` at `/pool/<pool>`. `pct restore --pool` sets membership only at CREATE — a restore over an existing vmid drops it, so bring-up re-asserts post-restore |
| `reconcile.PreflightRestoreSpace` + `restorespace.Provider` (v0.133.0) | internal/reconcile/restoretest_space.go, internal/restorespace/restorespace.go | `PreflightRestoreSpace(ctx, space, policy, archive, rawCfg, configured) SpaceVerdict` | ANY step that restores or copies a guest onto a storage — size it first | The restored size is UNCOMPRESSED (vzdump log "Total bytes written" / PBS snapshot size) — never the archive FILE (6.9 GB file → 22.6 GB restore, R-672); an unknown refuses; eligibility needs Datastore.AllocateSpace on `/storage/<id>` specifically (the `/` grant answers every path) |
| `Engine.RetryScratchTeardown` + the daemon janitor (v0.133.0) | internal/reconcile/restoretest_retry.go, cmd/felhom-agent/janitor.go | `RetryScratchTeardown(ctx) ScratchRetryResult` | retrying a leftover on a TIMER instead of only at start | Never `Recover` on a timer — it also resolves generic in-flight ops; a periodic sweep that unlocks guests holds the one-heavy-op gate (`InFlight.TryAcquire`) |
| `TLSConfig.build` / `normalizeFingerprint` | internal/proxmox/tls.go | `build() (*tls.Config, error)` | PVE leaf-cert SHA-256 pinning | No insecure default |
| `pinnedTLS` | internal/pbs/pin.go | `pinnedTLS(fingerprint) (*tls.Config, error)` | PBS leaf pinning | Same model as PVE; 64-hex fingerprint normalized |
| `httpx.NewTransport` | internal/httpx/transport.go | `NewTransport(tlsCfg, idleConnTimeout) *http.Transport` | **EVERY** hand-rolled `http.Transport` in this repo — pbs, hub and proxmox all pin TLS, so none can use `http.DefaultTransport` | **R-344: never inline `&http.Transport{TLSClientConfig: ...}` again.** A composite literal takes `IdleConnTimeout` **zero, which means retain idle connections FOREVER** — `http.DefaultTransport` sets 90s and a literal does not inherit it. Combined with a client rebuilt per cycle and dropped (`pbsTargetsFromPVE`), that stranded **388 sockets on ep0 in 46 h**, held open on BOTH sides. `idleConnTimeout <= 0` means **use the default**, never "no timeout". Returns a **FRESH** transport every call — a shared one would pool connections across differently pinned endpoints. Pinned by `internal/pbs/client_leak_test.go` (server-side connection counting) + `internal/httpx/transport_test.go` |
| `hub.Client.Report` | internal/hub/client.go | `Report(ctx, *HostReport) (*ControlEnvelope, error)` | the heartbeat | Typed `TransportError`/`HTTPError`, never contain the bearer token |
| `hub.Loop` + `MultiObserver` | internal/hub/loop.go | `NewLoop(...)`; `MultiObserver(obs...)` | resilient report loop + envelope fan-out | Errors logged, loop continues; interval clamped 60–3600 s |
| `provision.BackHalf.Provision` | internal/provision/backhalf.go | `Provision(ctx, Input) (Result, error)` | guest bootstrap back-half | mint→render→0600 write→chown 100000:100000→`pct set` ro bind→onboot; token NEVER logged/returned. Bootstrap `local_api.endpoint` = the caller's `cfg.LocalAPI.ListenAddr` (main.go) — moving the agent bind to the island moves the guest dial for free (R-50, no template) |
@@ -110,7 +122,7 @@
| Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only |
| Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set |
| Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction |
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`; R-856 `crashGuardStatePath` — GET /host/crash-guard's state file, internal/localapi/crashguard.go) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
| `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) |
| `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs |
| `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s |
@@ -146,13 +158,18 @@
| `storage.HostOps` | internal/storage/hostops.go | `*SudoHostOps` (prod), `NoopHostOps` (degraded) | fakes in internal/storage/observe_test.go, watchdog_test.go |
| `storage.HostReader` | internal/storage/hostread.go | `*ProcHostReader` | `fakeHostReader` internal/localapi/disks_test.go; internal/storage/role_test.go. v0.87.0: `BlockSlaves(name)` lists `/sys/block/<name>/slaves` (root-free) — backs the `SystemDisks` dm/md walk (`physicalDisksOf`/`walkSlaves`, role.go); per-branch conservatism: an unresolvable slave fails the WHOLE walk → all-system fail-safe. NEVER weaken the signature test `TestSystemDisks_WalkTopologies` (root-backing disk always in the system set). |
| `localapi.DiskOps` / `StorageGate` / `GuestAttacher` / `GuestLister` | internal/localapi/disks.go | `*storage.SudoHostOps`; `storageGateAdapter` (cmd/felhom-agent/main.go); `*GuestBinder`; `*proxmox.Client` | `fakeDiskOps`/`fakeGate`/`fakeGuestAttacher`/`fakeGuestList` internal/localapi/disks_test.go |
| `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring |
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,t)` / `ProvenArchive(target)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. |
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
| `backup.BackupSuccessState` | internal/backup/backup_state.go | `RecordBackupSuccess(target, b)` / `LastKnownSuccess(target, vmid)` | Newest SUCCESSFUL backup per tier+guest, persisted (atomic tmp+rename) — the due-check's fallback when the tier's storage cannot be read after a restart (R-894) | **Read ONLY when the storage cannot be read** — a storage that answers is the ground truth (R-84), and an archive absent there must make the tier due even when this file remembers one. A saved copy older than the cadence still reads due. Only successes are written. |
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
| `localapi.StaleLockController` | internal/localapi/stalelock.go | `*staleLockController` (Client + Runner + pool) | `fakeStaleLock` (Server-level) stalelock_test.go; `fakeStaleLockAPI` (controller-level, tests the A1 pool intersect) stalelock_pool_test.go |
| `localapi.GuestExecutor` | internal/localapi/controllerswap.go | `*GuestBinder` (pct exec) | `fakeGuestExec` internal/localapi/controllerswap_test.go |
| `localapi.GuestExecutor` | internal/localapi/controllerswap.go | `*GuestBinder` (pct exec; the image write via `felhom-priv-apply controller-image`) | `fakeGuestExec` internal/localapi/controllerswap_test.go |
| `guestnet.Runner` / `guestnet.GuestSource` (R-54, v0.92.0) | internal/guestnet/{probe,watchdog}.go | `*proxmox.ExecRunner`; the POOL-VERIFIED `localapi.StaleLockController.Guests` (ListLXC ∩ felhom pool, audit A1) | `scriptedRunner` + `fakeGuests` internal/guestnet/watchdog_test.go. **Never wire a bare `ListLXC` here** — under a broad token that would run dhclient inside a co-tenant's container. Every assertion is an exec COUNT, and the load-bearing ones are the negatives: a static guest, an unprobeable guest, a boot-race guest and an unproven guest list must record **zero** heal calls |
| `guestnet.Watchdog.SetDampers` / `now` (clock seam) | internal/guestnet/watchdog.go | config `guest_net.*`; `now` defaults to `time.Now` | tests advance a manual clock (the storage-watchdog pattern) and assert the heal ceilings EXACTLY — ≥10 min apart, ≤3/hour, and ≤30 over a scripted 10 hours of permanent failure. A damper with no test is a comment |
| `hub.GuestNetReporter` (R-54) | internal/hub/collect.go | `*guestnet.Watchdog` (`GuestNetStatus`) | internal/hub/collect_guestnet_test.go asserts the stanza through the PRODUCTION `Collect` path AND that the `guest_net` key is ABSENT from the wire when no reporter is wired — an always-present empty stanza would make "not wired" and "found nothing" the same signal, which is the shape v0.91.0 hid behind |
@@ -0,0 +1,188 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"strings"
"testing"
)
// Scenario H — THE SEAM IS WIRED IN THE PRODUCTION PATH, proven by walking the AST rather than by
// grepping for a string.
//
// WHY THIS TEST EXISTS AND WHY IT IS AN AST WALK. This project's built-but-never-wired count is six,
// and links 6 and 7 of the recovery chain were TWO of them: `UnwrapIdentityBundle` sat in the tree
// for two months with no caller but a `--selftest`, and the hub's blob-serving endpoints have no
// client to this day. The fix must not become the seventh. `strings.Contains` on the file would pass
// against a commented-out line, a line inside a test helper, or a line in dead code behind a flag
// nobody sets — so this resolves the call graph instead: `Options{EscrowRecovery: …}` must be
// constructed inside a function that `runDaemon` reaches, and `runDaemon` must be reached by `main`.
func parseMain(t *testing.T) (*token.FileSet, *ast.File) {
t.Helper()
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, parser.ParseComments)
if err != nil {
t.Fatalf("parsing main.go: %v", err)
}
return fset, f
}
// callsWithin returns the set of function names called (directly, by identifier or selector) inside
// the named top-level function.
func callsWithin(f *ast.File, fnName string) map[string]bool {
out := map[string]bool{}
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != fnName || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
ce, ok := n.(*ast.CallExpr)
if !ok {
return true
}
switch fn := ce.Fun.(type) {
case *ast.Ident:
out[fn.Name] = true
case *ast.SelectorExpr:
if x, ok := fn.X.(*ast.Ident); ok {
out[x.Name+"."+fn.Sel.Name] = true
}
out[fn.Sel.Name] = true
}
return true
})
}
return out
}
// TestEscrowRecoveryIsWiredIntoTheDaemon asserts the whole chain from func main() to the field.
func TestEscrowRecoveryIsWiredIntoTheDaemon(t *testing.T) {
_, f := parseMain(t)
// 1. main() reaches runDaemon.
if !callsWithin(f, "main")["runDaemon"] {
t.Fatal("func main() does not call runDaemon — the daemon path this test asserts is not the live one")
}
// 2. runDaemon reaches buildLocalAPIServer.
if !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
t.Fatal("runDaemon does not call buildLocalAPIServer — the local API is not built on the daemon path")
}
// 3. Inside buildLocalAPIServer, a localapi.Options composite literal carries EscrowRecovery, and
// an escrow.OffsiteKeyRecoverer is constructed there.
var optionsHasField, recovererConstructed bool
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
cl, ok := n.(*ast.CompositeLit)
if !ok {
return true
}
sel, ok := cl.Type.(*ast.SelectorExpr)
if !ok {
return true
}
pkg, _ := sel.X.(*ast.Ident)
if pkg == nil {
return true
}
switch pkg.Name + "." + sel.Sel.Name {
case "localapi.Options":
for _, el := range cl.Elts {
kv, ok := el.(*ast.KeyValueExpr)
if !ok {
continue
}
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "EscrowRecovery" {
optionsHasField = true
}
}
case "escrow.OffsiteKeyRecoverer":
recovererConstructed = true
}
return true
})
}
if !recovererConstructed {
t.Error("no escrow.OffsiteKeyRecoverer is constructed in buildLocalAPIServer — links 6→8 have no " +
"production assembly point (the built-but-never-wired shape, seventh instance)")
}
if !optionsHasField {
t.Error("localapi.Options in buildLocalAPIServer carries no EscrowRecovery field — the recoverer " +
"exists and the route would answer 503 forever")
}
}
// The hub fetch must be the DAEMON's own hub client, not a freshly constructed one with different
// credentials — the self-scoping that makes cross-host retrieval impossible is a property of WHICH
// key is used.
func TestEscrowRecoveryUsesTheDaemonHubClient(t *testing.T) {
fset, f := parseMain(t)
var fetchUsesHubClient bool
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
ce, ok := n.(*ast.CallExpr)
if !ok {
return true
}
sel, ok := ce.Fun.(*ast.SelectorExpr)
if !ok || sel.Sel.Name != "FetchIdentityEscrow" {
return true
}
if x, ok := sel.X.(*ast.Ident); ok && x.Name == "hubClient" {
fetchUsesHubClient = true
} else {
t.Errorf("FetchIdentityEscrow at %s is called on something other than the injected hub client",
fset.Position(ce.Pos()))
}
return true
})
}
if !fetchUsesHubClient {
t.Fatal("the recoverer's fetcher does not call hubClient.FetchIdentityEscrow — either the fetch is " +
"not wired, or it uses a client whose credentials are not this host's")
}
}
// The route itself must be registered on the local API. A handler with no route is the same defect
// one layer down, and it has shipped here before.
func TestRecoverRouteIsRegistered(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "../../internal/localapi/server.go", nil, 0)
if err != nil {
t.Fatalf("parsing localapi/server.go: %v", err)
}
var registered bool
ast.Inspect(f, func(n ast.Node) bool {
ce, ok := n.(*ast.CallExpr)
if !ok || len(ce.Args) < 2 {
return true
}
sel, ok := ce.Fun.(*ast.SelectorExpr)
if !ok || sel.Sel.Name != "HandleFunc" {
return true
}
lit, ok := ce.Args[0].(*ast.BasicLit)
if !ok {
return true
}
if strings.Contains(lit.Value, "/escrow/recover-offsite-password") {
registered = true
}
return true
})
if !registered {
t.Fatal("POST /escrow/recover-offsite-password is not registered on the local API mux — the handler " +
"exists and nothing can reach it")
}
}
+53
View File
@@ -0,0 +1,53 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"testing"
)
// R-444: the weekly trim has the guestnet shape (component + reporter seam + goroutine), so its wiring is asserted
// from the AST like TestMainWiresGuestNetWatchdog — a unit-green trim job that main.go never starts is the inert-seam
// defect. It must also share the ONE heavy-op gate (heavyOps), or it could run beside a backup.
func TestMainWiresGuestDiskTrim(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, 0)
if err != nil {
t.Fatalf("parse main.go: %v", err)
}
var constructedWithGate, reporterWired, started bool
ast.Inspect(f, func(n ast.Node) bool {
switch node := n.(type) {
case *ast.CallExpr:
if fn, ok := node.Fun.(*ast.SelectorExpr); ok {
switch fn.Sel.Name {
case "New":
if pkg, ok := fn.X.(*ast.Ident); ok && pkg.Name == "fstrim" && len(node.Args) >= 3 {
if id, ok := node.Args[2].(*ast.Ident); ok && id.Name == "heavyOps" {
constructedWithGate = true
}
}
case "SetGuestDiskTrimReporter":
reporterWired = true
}
}
case *ast.GoStmt:
if sel, ok := node.Call.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "Run" {
if id, ok := sel.X.(*ast.Ident); ok && id.Name == "diskTrim" {
started = true
}
}
}
return true
})
if !constructedWithGate {
t.Error("main.go never calls fstrim.New(..., heavyOps, ...) — no trim job, or one outside the heavy-op gate")
}
if !reporterWired {
t.Error("main.go never calls collector.SetGuestDiskTrimReporter — the guest_disk_trim stanza never reaches the hub")
}
if !started {
t.Error("main.go never starts the trim job with `go diskTrim.Run(ctx)`")
}
}
+81
View File
@@ -0,0 +1,81 @@
package main
import (
"context"
"fmt"
"log/slog"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// janitorInterval is how often the leftovers of an interrupted restore-test or backup are retried
// (R-672 rule 3, R-673). Both used to be resolved ONLY at agent start: on 2026-09-24 a failed scratch
// teardown kept a full thin pool full for 2.5 h, and a stale `snapshot-delete` lock blocked 9201's
// whole-box backups for five hours — each cleared within a minute of an agent restart.
const janitorInterval = 10 * time.Minute
// janitorDeps are the janitor's seams (tests drive one pass with fakes).
type janitorDeps struct {
retryScratch func(ctx context.Context) reconcile.ScratchRetryResult
staleLocks func(ctx context.Context) // localapi Server.RecoverStaleLockedGuests; nil when the local API is off
heavy *backup.InFlight
record func(hub.RestoreTest)
now func() time.Time
logger *slog.Logger
}
// janitorPass is one pass. The stale-lock sweep runs only while holding the one-heavy-operation gate, so
// no agent backup can START between its "no vzdump is running" check and its unlock (at start-up the
// sweep ran before the backup loop existed; on a timer that ordering must be made, not assumed). A busy
// gate skips the sweep this pass — the next pass retries.
func janitorPass(ctx context.Context, d janitorDeps) {
if d.retryScratch != nil {
r := d.retryScratch(ctx)
if r.Examined > 0 {
d.logger.Info("janitor: restore-test scratch retry pass", "examined", r.Examined,
"destroyed", r.Destroyed, "already_gone", r.Clean, "failed", r.Failed)
}
for _, vmid := range r.GaveUp {
// The operator is told through the existing restore-test failure path: the hub raises
// restore_test_failed (operator) once per distinct archive — this record's archive names
// the stuck scratch guest.
if d.record != nil {
d.record(hub.RestoreTest{
SourceArchive: fmt.Sprintf("scratch-teardown:%d", vmid),
ScratchVMID: vmid,
Pass: false,
Error: fmt.Sprintf("restore-test scratch guest %d could not be torn down after %d retries — it holds its disks; remove it by hand (pct destroy %d) after checking what keeps it busy",
vmid, reconcile.MaxTeardownTries, vmid),
TestedAt: d.now().UTC().Format(time.RFC3339),
})
}
}
}
if d.staleLocks != nil {
release, busy, ok := d.heavy.TryAcquire("stale-lock-sweep")
if !ok {
d.logger.Info("janitor: stale-lock sweep deferred — a heavy operation is in flight", "busy", busy)
return
}
defer release()
d.staleLocks(ctx)
}
}
// runJanitor runs janitorPass every janitorInterval until ctx ends.
func runJanitor(ctx context.Context, d janitorDeps) {
d.logger.Info("janitor: starting (restore-test scratch retry + stale-lock sweep)", "interval", janitorInterval)
t := time.NewTicker(janitorInterval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
janitorPass(ctx, d)
}
}
}
+60
View File
@@ -0,0 +1,60 @@
package main
import (
"context"
"io"
"log/slog"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-672 / R-673 (v0.133.0): one janitor pass, driven with fakes.
func quiet() *slog.Logger { return slog.New(slog.NewTextHandler(io.Discard, nil)) }
// A scratch the engine gave up on reaches the hub as a failed restore-test record naming it — the
// existing operator path (restore_test_failed). Never a pass.
func TestJanitor_GaveUpIsReportedAsAFailure(t *testing.T) {
var got []hub.RestoreTest
janitorPass(context.Background(), janitorDeps{
retryScratch: func(context.Context) reconcile.ScratchRetryResult {
return reconcile.ScratchRetryResult{Examined: 1, Failed: 1, GaveUp: []int{990000}}
},
heavy: &backup.InFlight{}, record: func(r hub.RestoreTest) { got = append(got, r) },
now: time.Now, logger: quiet(),
})
if len(got) != 1 || got[0].Pass || got[0].ScratchVMID != 990000 || !strings.Contains(got[0].Error, "990000") {
t.Fatalf("records = %+v — want one FAILED record naming scratch 990000", got)
}
}
// R-673: the stale-lock sweep runs only while holding the one-heavy-operation gate, so no agent backup can
// start between its "no vzdump running" check and its unlock.
//
// COMPANION RED-PROOF (REPORT): drop the TryAcquire → "the sweep ran while a backup held the gate".
func TestJanitor_StaleLockSweepWaitsForTheHeavyGate(t *testing.T) {
heavy := &backup.InFlight{}
swept := 0
d := janitorDeps{staleLocks: func(context.Context) { swept++ }, heavy: heavy, now: time.Now, logger: quiet()}
release, _, ok := heavy.TryAcquire("backup")
if !ok {
t.Fatal("setup")
}
janitorPass(context.Background(), d)
if swept != 0 {
t.Fatal("the sweep ran while a backup held the gate")
}
release()
janitorPass(context.Background(), d)
if swept != 1 {
t.Fatalf("swept %d times with the gate free — want 1", swept)
}
if _, _, ok := heavy.TryAcquire("after"); !ok {
t.Fatal("the sweep did not release the gate")
}
}
+845 -82
View File
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,31 @@
package main
import (
"context"
"io"
"log/slog"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/config"
)
// R-861 (agent v0.146.0): the root escrow ceremony builds a file path from the storage id; an id that is a path is
// refused before anything is read. RED-PROOF: drop the pbsStorageIDRe check → the "../" ids reach the key stat and
// come back as a "setup" error instead of "usage".
func TestEscrowCeremony_StorageIDIsNeverAPath(t *testing.T) {
lg := slog.New(slog.NewTextHandler(io.Discard, nil))
for _, id := range []string{"../../../etc/shadow", "a/b", "/etc/pve/priv/x", ".hidden", ""} {
cfg := config.Default()
cfg.Escrow.PBSStorageID = "" // the flag decides here
_, e := escrowCeremony(context.Background(), cfg, lg, escrowCeremonyOpts{storage: id})
if e == nil || e.kind != "usage" {
t.Errorf("storage id %q was not refused as usage (got %+v)", id, e)
}
}
cfg := config.Default()
cfg.Backup.PBSSecretDir = t.TempDir()
_, e := escrowCeremony(context.Background(), cfg, lg, escrowCeremonyOpts{storage: "felhom-pbs"})
if e == nil || e.kind != "setup" {
t.Fatalf("control: a plain id must pass the check and fail later on the missing key (setup), got %+v", e)
}
}
@@ -0,0 +1,55 @@
package main
import (
"context"
"errors"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
)
type r866Fetcher struct{ err error }
func (f r866Fetcher) FetchDesiredState(context.Context) (*hub.DesiredStateResponse, error) {
if f.err != nil {
return nil, f.err
}
r := &hub.DesiredStateResponse{}
r.DesiredState.OSUpdate = &hub.WireOSUpdate{Ring: 1, Enabled: true}
return r, nil
}
// R-866 (v0.144.0). THE NIGHT'S SHAPE (A3, Tester 1 box, hub blocked): `selftest=os-update: desired state: hub:
// transport error … connect: invalid argument` — the debug pass could not run at all. Now it runs from the block the
// daemon saved, and its header says so.
// COMPANION RED-PROOF: return at once on a fetch error in selftestOSBlock → "the pass did not run from the saved block".
func TestR866_DebugPassUsesTheSavedBlockWhenTheHubIsAway(t *testing.T) {
dir := t.TempDir()
leg := &osupdate.Leg{PlanDir: dir}
r := &hub.DesiredStateResponse{}
r.DesiredState.OSUpdate = &hub.WireOSUpdate{Ring: 0, Enabled: true}
leg.OnDesiredState(context.Background(), r) // the daemon received a block and saved it
away := r866Fetcher{err: errors.New("hub: transport error: connect: invalid argument")}
b, src, ok := selftestOSBlock(context.Background(), away, dir)
if !ok || b == nil || b.Ring != 0 {
t.Fatalf("the pass did not run from the saved block: ok=%v block=%+v src=%q", ok, b, src)
}
if !strings.HasPrefix(src, "SAVED(") || !strings.Contains(src, "hub unreachable") {
t.Fatalf("the header must say the block is the saved one: %q", src)
}
// the hub reachable: its block wins, and the header says "hub"
b, src, ok = selftestOSBlock(context.Background(), r866Fetcher{}, dir)
if !ok || b.Ring != 1 || src != "hub" {
t.Fatalf("hub block not used: %+v %q", b, src)
}
}
// No hub and nothing saved: the pass does not run on a guessed block.
func TestR866_NoHubNoSavedBlockDoesNotRun(t *testing.T) {
_, why, ok := selftestOSBlock(context.Background(), r866Fetcher{err: errors.New("down")}, t.TempDir())
if ok || !strings.Contains(why, "no saved block") {
t.Fatalf("ok=%v why=%q", ok, why)
}
}
+74
View File
@@ -0,0 +1,74 @@
package main
import (
"go/ast"
"testing"
)
// R-894 — the on-disk backup record is WIRED on the daemon path (the built-but-never-wired class).
// main → runDaemon → buildLocalAPIServer, and inside it the localapi.Options literal carries
// LastKnownBackups built by backup.NewBackupSuccessState. An AST walk, not a string match, for the
// reasons in escrow_recover_wiring_test.go.
//
// COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this
// fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored.
func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
_, f := parseMain(t)
if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one")
}
var field, built bool
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
cl, ok := n.(*ast.CompositeLit)
if !ok {
return true
}
sel, ok := cl.Type.(*ast.SelectorExpr)
if !ok {
return true
}
if pkg, _ := sel.X.(*ast.Ident); pkg == nil || pkg.Name+"."+sel.Sel.Name != "localapi.Options" {
return true
}
for _, el := range cl.Elts {
kv, ok := el.(*ast.KeyValueExpr)
if !ok {
continue
}
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "LastKnownBackups" {
field = true
if callsIn(kv.Value)["backup.NewBackupSuccessState"] {
built = true
}
}
}
return true
})
}
if !field {
t.Fatal("localapi.Options in buildLocalAPIServer has no LastKnownBackups field")
}
if !built {
t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState")
}
}
func callsIn(n ast.Node) map[string]bool {
out := map[string]bool{}
ast.Inspect(n, func(n ast.Node) bool {
if ce, ok := n.(*ast.CallExpr); ok {
if fn, ok := ce.Fun.(*ast.SelectorExpr); ok {
if x, ok := fn.X.(*ast.Ident); ok {
out[x.Name+"."+fn.Sel.Name] = true
}
}
}
return true
})
return out
}
@@ -111,3 +111,46 @@ func parseMainForWiring(t *testing.T) *ast.File {
}
return f
}
// R-189 Scenario I — the DURABLE proof source must actually be wired into the collector.
//
// This test exists because the method it feeds is the project's own cautionary tale:
// `RestoreTestState.Snapshot` carried the doc comment "for the host-report gauge" from the day it
// was written and **had no caller at all** — a seam built, documented and never connected, found
// only when a live restore-test's PASS reached no host-report. The fix must not become the next
// instance, so the wiring is asserted rather than trusted.
//
// AST, not grep: a commented-out call still contains the string (proven yesterday, when commenting
// out the tier-picker line failed this test while a `strings.Contains` check would have passed).
func TestMainWiresTheDurableRestoreTestProof(t *testing.T) {
f := parseMainForWiring(t)
var wired, feedsState bool
ast.Inspect(f, func(n ast.Node) bool {
call, ok := n.(*ast.CallExpr)
if !ok {
return true
}
sel, ok := call.Fun.(*ast.SelectorExpr)
if !ok || sel.Sel.Name != "SetProvenRestoreTests" {
return true
}
wired = true
// ...and it must be fed the PERSISTED state, not the in-memory store.
if len(call.Args) == 1 {
if id, ok := call.Args[0].(*ast.Ident); ok && id.Name == "rtState" {
feedsState = true
}
}
return true
})
if !wired {
t.Error("main.go never calls collector.SetProvenRestoreTests — the persisted proof would never " +
"reach the hub, which is the R-189 defect exactly: a passing restore-test that vanishes on restart")
}
if wired && !feedsState {
t.Error("collector.SetProvenRestoreTests is not fed rtState — the in-memory store is the thing " +
"that does NOT survive a restart, so wiring it here would fix nothing")
}
}
+39
View File
@@ -0,0 +1,39 @@
package main
import (
"os"
"regexp"
"strings"
"testing"
)
// Every mode the dispatcher (`switch selftest.mode`) runs must be ACCEPTED by the --selftest flag. Found live
// 2026-10-04: --selftest=os-update had a dispatch case and a function but the flag's allow-list refused it, so the
// debug action could not run. Red-proof: drop the "os-update" case from selftestFlag.Set and this fails.
func TestSelftestFlag_AcceptsEveryDispatchedMode(t *testing.T) {
src, err := os.ReadFile("main.go")
if err != nil {
t.Fatal(err)
}
s := string(src)
i := strings.Index(s, "switch selftest.mode {")
if i < 0 {
t.Fatal("dispatch switch not found")
}
block := s[i:]
block = block[:strings.Index(block, "\n\t}\n")]
modes := regexp.MustCompile(`(?m)^\tcase "([a-z-]+)":`).FindAllStringSubmatch(block, -1)
if len(modes) < 5 {
t.Fatalf("parsed only %d dispatch cases — the parser is wrong", len(modes))
}
var bad []string
for _, m := range modes {
var f selftestFlag
if err := f.Set(m[1]); err != nil {
bad = append(bad, m[1])
}
}
if len(bad) > 0 {
t.Fatalf("dispatched but refused by --selftest: %v", bad)
}
}
+417
View File
@@ -0,0 +1,417 @@
package main
import (
"context"
"errors"
"go/ast"
"io"
"log/slog"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
)
// R-185 — a tier the box cannot READ must say so.
//
// THE OBSERVATION (demo-felhom, 2026-08-03, reproduced at the start of this session): root lists
// three archives on `felhom-backup`; the agent's own token gets `{"data":[]}` from the same
// endpoint; and `local`, which has the grant, lists through that same token. The token is the
// variable, not the storage.
//
// The defect is NOT the missing grant — that is one command. It is that an empty content listing is
// what a FORBIDDEN tier and a NEWBORN tier both return, so the box could not tell them apart and
// said nothing. These tests pin the distinction.
// permAnswer is the shape /access/permissions really returns, taken from the live measurement:
// an UNGRANTED path answers with the privileges inherited from the box-wide grant — NOT empty, and
// NOT a 403.
var (
permGranted = map[string]int{"Datastore.Allocate": 1, "Datastore.AllocateSpace": 1}
permUngranted = map[string]int{"Sys.Audit": 1, "SDN.Use": 1, "Datastore.Audit": 1}
)
// probeWith calls the PRODUCTION decision with a permissions answer. **Naming the seam:** everything
// below is true up to `storeGrantVerdict`; that the live call feeds it the real API answer is what
// Part 0's measurement established and what the live run on the box demonstrates. An earlier draft
// of this file re-implemented the branch here — it passed, and would have kept passing while
// production diverged, which is the hollow shape this project keeps catching in its own tests.
func probeWith(privs map[string]int, targetID string, critical bool) capability.Status {
return storeGrantVerdict(targetID, critical, privs, nil)
}
// ── SCENARIO A — a forbidden storage is REPORTED, not passed over ────────────────────────────
//
// COMPANION RED-PROOF (observed 2026-08-03): delete the store-grant probes from `probeAll` in
// main.go — i.e. restore `append(capProber.Probe(ctx), poolReadStatus(ctx, px))` — and
// TestMainWiresTheStoreGrantProbe fails with "main.go never calls storeGrantStatuses". That is
// today's behaviour on the live box: complete silence about a tier it cannot read.
func TestStoreGrant_ForbiddenStorageIsDegradedAndNamed(t *testing.T) {
s := probeWith(permUngranted, "felhom-backup", true)
if s.Status != capability.StatusDegraded {
t.Fatalf("a storage the agent may not read must be DEGRADED, not %q — silence is the defect", s.Status)
}
if !s.Critical {
t.Fatal("it must be CRITICAL: the hub alerts only on critical, so a non-critical entry is the same silence with extra steps")
}
if !strings.Contains(s.Reason, "felhom-backup") {
t.Fatalf("the reason must NAME the storage — 'a grant is missing' costs a diagnosis at 07:00; got %q", s.Reason)
}
if !strings.Contains(s.Reason, "FelhomAgentStore") {
t.Fatalf("the reason must name the ROLE to grant, so the fix is in the alert; got %q", s.Reason)
}
}
// THE TRAP THE LIVE MEASUREMENT CAUGHT, pinned so it cannot be re-introduced: the ungranted answer
// is not empty and not a 403 — it carries the INHERITED box-wide privileges. A probe that asked
// "did the path come back?" or "does it have Datastore.Audit?" would report the blinded storage
// healthy.
func TestStoreGrant_InheritedPrivilegesAreNotAGrant(t *testing.T) {
if len(permUngranted) == 0 {
t.Fatal("fixture wrong: the ungranted answer is NOT empty — that is the whole trap")
}
if permUngranted["Datastore.Audit"] != 1 {
t.Fatal("fixture wrong: the ungranted path DOES carry Datastore.Audit, inherited box-wide")
}
if s := probeWith(permUngranted, "felhom-backup", true); s.Status != capability.StatusDegraded {
t.Fatalf("checking for the wrong privilege reports a blinded storage healthy; got %q", s.Status)
}
// ...and the privilege actually checked is the one whose absence was measured to blind listing.
if storeGrantRequiredPriv != "Datastore.AllocateSpace" {
t.Fatalf("the probed privilege changed to %q — re-measure before trusting it", storeGrantRequiredPriv)
}
}
// ── SCENARIO B — a newborn tier is still silent ──────────────────────────────────────────────
//
// A storage the agent IS allowed to read but which simply holds no archives yet is HEALTHY. The
// probe must not look at content at all, or every freshly provisioned box alarms and the signal dies.
//
// COMPANION RED-PROOF (observed): make the probe degrade on an empty content listing instead of on
// the permission — a granted-but-empty storage then reports degraded, i.e. every newborn box alarms.
func TestStoreGrant_GrantedButEmptyIsHealthy(t *testing.T) {
s := probeWith(permGranted, "felhom-pbs", true)
if s.Status != capability.StatusOK {
t.Fatalf("a readable tier is healthy whether or not it holds archives yet; got %q (%s)", s.Status, s.Reason)
}
if s.Reason != "" {
t.Fatalf("a healthy probe carries no reason; got %q", s.Reason)
}
}
// ── SCENARIO C — the two states are distinguishable at a glance ──────────────────────────────
func TestStoreGrant_ForbiddenAndNewbornAreDistinguishable(t *testing.T) {
forbidden := probeWith(permUngranted, "felhom-backup", true)
newborn := probeWith(permGranted, "felhom-pbs", true)
if forbidden.Status == newborn.Status {
t.Fatalf("the two states must differ — today both read as 'no settled archive yet'; got %q for both", forbidden.Status)
}
if forbidden.Name == newborn.Name {
t.Fatalf("each tier needs its own capability id, or one tier's fault hides another's; got %q twice", forbidden.Name)
}
}
// §8.3, weighed once and pinned: a box with NO dedicated target ("local" — host-install's own
// DEGRADED fallback) must not turn an ordinary configuration into an operator page. It is still
// probed and still reported; only the paging differs.
func TestStoreGrant_TheFallbackTargetIsNotCritical(t *testing.T) {
if storeGrantCritical("local") {
t.Fatal("a box whose backup target is the 'local' fallback must not page the operator about " +
"an ordinary, documented configuration")
}
for _, dedicated := range []string{"felhom-backup", "felhom-pbs", "some-nvme"} {
if !storeGrantCritical(dedicated) {
t.Fatalf("a DEDICATED target that cannot be read is user-facing and must be critical; %q was not", dedicated)
}
}
// The fallback is still reported — silence for it would be the original defect, scoped smaller.
if s := probeWith(permUngranted, "local", storeGrantCritical("local")); s.Status != capability.StatusDegraded {
t.Fatalf("the fallback target must still report degraded when unreadable; got %q", s.Status)
}
}
// A probe that cannot ask must never answer "ok" — unknown reported as healthy is worse than no
// probe, because it looks like coverage.
func TestStoreGrant_UnreachablePVEIsDegradedNotOK(t *testing.T) {
s := storeGrantStatus(context.Background(), nil, "felhom-backup", true, nil)
if s.Status != capability.StatusDegraded {
t.Fatalf("an unaskable probe must be DEGRADED, never ok; got %q", s.Status)
}
if s.Reason == "" {
t.Fatal("it must say why it could not ask")
}
}
// ── SCENARIO H — the seam ────────────────────────────────────────────────────────────────────
//
// This project's "built but never wired" count reached six last week. The fix for a SILENCE must not
// itself be silent. AST, not grep: a commented-out call still contains the string.
func TestMainWiresTheStoreGrantProbe(t *testing.T) {
f := parseMainForWiring(t)
var wired bool
ast.Inspect(f, func(n ast.Node) bool {
call, ok := n.(*ast.CallExpr)
if !ok {
return true
}
if id, ok := call.Fun.(*ast.Ident); ok && id.Name == "storeGrantStatuses" {
wired = true
}
return true
})
if !wired {
t.Error("main.go never calls storeGrantStatuses — the probe would exist and report to nobody, " +
"which is precisely the silence R-185 is about")
}
}
// ── R-190 — the grant repairs itself, and the repair is VISIBLE ──────────────────────────────
//
// R-190 is a storage grant that demonstrably worked at 04:44 on 2026-08-03 and was gone by 09:24,
// with a host reinstall, logged `pveum` activity and cluster-log entries all ruled out. The cause is
// open; the resilience is not conditional on it.
//
// The half that matters is the RECORD. R-190's own words: the probe sees the state, nothing sees the
// transition. A self-repair that leaves only "ok" behind destroys the only evidence a loss happened,
// so a recurring loss becomes undetectable forever — strictly worse than the fault it fixes.
// fakeRepairRunner records wrapper invocations and can be made to fail.
type fakeRepairRunner struct {
calls [][]string
fail bool
}
func (f *fakeRepairRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
f.calls = append(f.calls, append([]string{name}, args...))
if f.fail {
return nil, []byte("pveum: refused"), errors.New("exit status 2")
}
return nil, nil, nil
}
func newRepairer(f *fakeRepairRunner) *storeGrantRepairer {
return &storeGrantRepairer{run: f.Run, log: slog.New(slog.NewTextHandler(io.Discard, nil))}
}
// ── SCENARIO F — the repair is BOUNDED ───────────────────────────────────────────────────────
//
// COMPANION RED-PROOF (observed 2026-08-04): make mayAttempt always return true (drop the
// storeGrantRepairMinInterval check) →
//
// --- FAIL: TestGrantRepair_IsBounded
// storegrant_test.go: a repair must not run on every cycle; 5 cycles produced 5 attempt(s)
//
// which is a re-grant every report cycle, forever, against a fault an ACL cannot fix. Restored.
func TestGrantRepair_IsBounded(t *testing.T) {
f := &fakeRepairRunner{}
r := newRepairer(f)
// Jittered, so the series never lands exactly on the interval boundary — a perfectly regular
// series is how a threshold test passes its own mutation, which has happened here before.
base := time.Date(2026, 8, 4, 9, 17, 43, 0, time.UTC)
offsets := []time.Duration{0, 13*time.Minute + 7*time.Second, 27*time.Minute + 51*time.Second,
41*time.Minute + 19*time.Second, 55*time.Minute + 3*time.Second}
attempts := 0
for _, off := range offsets {
if r.mayAttempt("felhom-backup", base.Add(off)) {
attempts++
}
}
if attempts != 1 {
t.Fatalf("a repair must not run on every cycle; %d cycles produced %d attempt(s) within %s",
len(offsets), attempts, storeGrantRepairMinInterval)
}
// ...and once the interval has genuinely passed, it may try again — a bound is not a ban.
if !r.mayAttempt("felhom-backup", base.Add(storeGrantRepairMinInterval+2*time.Minute+11*time.Second)) {
t.Fatal("after the interval a repair must be allowed again — otherwise one failure disables the repair forever")
}
// A DIFFERENT tier is not throttled by this one's attempt.
if !r.mayAttempt("felhom-pbs", base.Add(time.Minute)) {
t.Fatal("the bound must be per tier — one tier's attempt must not suppress another's")
}
}
// A nil repairer (or one with no runner) never attempts, and never panics.
func TestGrantRepair_NilIsSafe(t *testing.T) {
var r *storeGrantRepairer
if r.mayAttempt("felhom-backup", time.Now()) {
t.Fatal("a nil repairer must never claim an attempt")
}
if (&storeGrantRepairer{}).mayAttempt("felhom-backup", time.Now()) {
t.Fatal("a repairer with no runner must never claim an attempt")
}
}
// The repair calls the EXISTING wrapper verb, with the storage id — no new privileged surface.
func TestGrantRepair_CallsTheExistingWrapperVerb(t *testing.T) {
f := &fakeRepairRunner{}
r := newRepairer(f)
if err := r.repair(context.Background(), "felhom-backup"); err != nil {
t.Fatalf("repair should succeed with a healthy runner: %v", err)
}
if len(f.calls) != 1 {
t.Fatalf("exactly one wrapper invocation expected; got %d", len(f.calls))
}
got := f.calls[0]
want := []string{"/usr/local/sbin/felhom-backup-target-apply", "grant", "felhom-backup"}
if len(got) != len(want) {
t.Fatalf("wrapper argv = %v, want %v", got, want)
}
for i := range want {
if got[i] != want[i] {
t.Fatalf("wrapper argv = %v, want %v — the sudoers vector is `grant *`; anything else is a policy change", got, want)
}
}
}
// A repair that FAILS must surface the failure, not swallow it (Scenario E's precondition).
func TestGrantRepair_FailureIsReturned(t *testing.T) {
f := &fakeRepairRunner{fail: true}
if err := newRepairer(f).repair(context.Background(), "felhom-backup"); err == nil {
t.Fatal("a failed wrapper run must return its error — a repair that cannot run must never read as done")
}
}
// ── SCENARIO D (the half that matters) — the REPAIR MUST BE VISIBLE ──────────────────────────
//
// A repair that leaves only "ok" behind is worse than the fault: the tier works, and the fact that a
// permission vanished is gone with it. R-190 exists because nothing saw the transition.
//
// The channel is the hub's EXISTING ok→degraded→ok edge (§8.5) — nothing new was built. That only
// works if the agent deliberately reports ONE degraded cycle after repairing, and if the explanation
// rides the field the hub actually puts in the operator's e-mail. The hub's message is built from the
// capability NAME and FEATURE (`internal/monitor/host_capability.go` emitTransition) — **not** from
// Reason — so the Feature must carry it.
//
// COMPANION RED-PROOF (observed 2026-08-04): after a successful repair, report ok instead —
//
// s.Status = capability.StatusOK; s.Feature unchanged
//
// → --- FAIL: TestGrantRepair_ARepairedGrantIsReportedAsATransition
//
// storegrant_test.go: a self-repair must still report DEGRADED for one cycle so the hub raises
// its edge; got "ok" — the loss would be invisible
//
// i.e. exactly the silence R-190 is about. Restored.
func TestGrantRepair_ARepairedGrantIsReportedAsATransition(t *testing.T) {
// THE PRODUCTION verdict, not a copy of it. An earlier draft of this test built the Status
// itself and asserted its own construction — it would have passed while production reported ok,
// which is precisely the silence being guarded against.
if pre := probeWith(permUngranted, "felhom-backup", true); pre.Status != capability.StatusDegraded {
t.Fatalf("precondition: a missing grant is degraded; got %q", pre.Status)
}
s := storeGrantRepairedVerdict("felhom-backup", true)
if s.Status != capability.StatusDegraded {
t.Fatalf("a self-repair must still report DEGRADED for one cycle so the hub raises its edge; "+
"got %q — the loss would be invisible", s.Status)
}
// The hub e-mails the FEATURE text. If the explanation is not there, the operator is told a
// capability was degraded and never learns it repaired itself or that anything vanished.
for _, want := range []string{"MISSING", "RESTORED", "felhom-backup", "R-190"} {
if !strings.Contains(s.Feature, want) {
t.Fatalf("the Feature text is what the hub puts in the operator's e-mail; it must contain %q. Got: %s", want, s.Feature)
}
}
if !s.Critical {
t.Fatal("the transition must be CRITICAL or the hub does not alert on it at all")
}
}
// ── SCENARIO H — the seam ────────────────────────────────────────────────────────────────────
//
// The wrapper's `grant` verb is itself a "built but never wired" example: it exists, is
// sudoers-permitted for any id, and had only ever been called at storage CREATION. The repair must
// not become the seventh instance. AST, not grep — a commented-out call still contains the string.
func TestMainWiresTheGrantRepair(t *testing.T) {
f := parseMainForWiring(t)
var built, passed bool
ast.Inspect(f, func(n ast.Node) bool {
switch node := n.(type) {
case *ast.CompositeLit:
if id, ok := node.Type.(*ast.Ident); ok && id.Name == "storeGrantRepairer" {
built = true
}
case *ast.CallExpr:
if id, ok := node.Fun.(*ast.Ident); ok && id.Name == "storeGrantStatuses" && len(node.Args) == 4 {
if a, ok := node.Args[3].(*ast.Ident); ok && a.Name == "grantRepairer" {
passed = true
}
}
}
return true
})
if !built {
t.Error("main.go never constructs a storeGrantRepairer — nothing would ever repair a lost grant")
}
if !passed {
t.Error("storeGrantStatuses is not passed the repairer — the probe would detect the loss and " +
"leave it, which is v0.123.0's behaviour and not R-190's mitigation")
}
}
// The transition must survive a probe that is NOT the one feeding the hub.
//
// MEASURED LIVE 2026-08-04, and this test exists because the first implementation failed it in
// production while every unit test passed: `probeAll` is called independently by the self-check LOG
// and by the collector building a host-report. The repairing call was the log's; the report three
// seconds later found the grant present and reported `ok`. The agent's journal had the record and the
// hub had nothing — the exact silence R-190 is about, re-created inside its own mitigation.
//
// COMPANION RED-PROOF (observed): delete the `recentlyRepaired` branch from the healthy path →
//
// --- FAIL: TestGrantRepair_TransitionSurvivesALaterProbe
// storegrant_test.go: a probe AFTER the repair must still report the transition; got "ok" —
// the host-report would carry ok and the operator would never learn the grant vanished
//
// Restored.
func TestGrantRepair_TransitionSurvivesALaterProbe(t *testing.T) {
r := newRepairer(&fakeRepairRunner{})
// Jittered, never landing on the window boundary.
repairedAt := time.Date(2026, 8, 4, 9, 39, 34, 0, time.UTC)
r.noteRepaired("felhom-backup", repairedAt)
// The DECISION a later probe makes — the production function, not the helper it calls. An
// earlier draft asserted `recentlyRepaired` directly and its red-proof PASSED, because removing
// the latch's USE left the helper untouched.
healthy := probeWith(permGranted, "felhom-backup", true)
if healthy.Status != capability.StatusOK {
t.Fatalf("precondition: a granted tier is ok; got %q", healthy.Status)
}
got := storeGrantHealthyVerdict("felhom-backup", true,
healthy, r.recentlyRepaired("felhom-backup", repairedAt.Add(3*time.Second)))
if got.Status != capability.StatusDegraded {
t.Fatalf("a probe AFTER the repair must still report the transition; got %q — the host-report "+
"would carry ok and the operator would never learn the grant vanished", got.Status)
}
if !strings.Contains(got.Feature, "RESTORED") {
t.Fatalf("the later probe must carry the explanation into the hub's e-mail; got: %s", got.Feature)
}
// Outside the window it reports plain ok again.
late := storeGrantHealthyVerdict("felhom-backup", true,
healthy, r.recentlyRepaired("felhom-backup", repairedAt.Add(storeGrantRepairReportWindow+time.Minute)))
if late.Status != capability.StatusOK {
t.Fatalf("outside the window a healthy tier reports ok; got %q — a permanent degraded state "+
"would be its own false alarm", late.Status)
}
if !r.recentlyRepaired("felhom-backup", repairedAt.Add(14*time.Minute+37*time.Second)) {
t.Fatal("the latch must outlast the 900s hub report interval, or the record never reaches the hub")
}
// ...and it clears on its own rather than latching a box degraded forever.
if r.recentlyRepaired("felhom-backup", repairedAt.Add(storeGrantRepairReportWindow+time.Minute+7*time.Second)) {
t.Fatal("the latch must clear — a permanent degraded state would be its own false alarm")
}
// It is per tier.
if r.recentlyRepaired("felhom-pbs", repairedAt.Add(time.Second)) {
t.Fatal("one tier's repair must not latch another tier's status")
}
// The window MUST exceed the report interval — the property, asserted rather than assumed.
if storeGrantRepairReportWindow <= 15*time.Minute {
t.Fatalf("the report window (%s) must exceed the 900s hub report interval, or a transition can "+
"be missed entirely", storeGrantRepairReportWindow)
}
}
+13 -1
View File
@@ -43,7 +43,7 @@ func main() {
func run() error {
var (
op = flag.String("op", "", "op class to sign, e.g. storage_wipe | guest_destroy | decommission | agent_update")
op = flag.String("op", "", "op class to sign, e.g. storage_wipe | guest_destroy | decommission | agent_update | os_docker_step | os_pve_step | agent_config_update")
host = flag.String("host", "", "target host_id (anti-retarget — the op runs ONLY on this host)")
guest = flag.String("guest", "", "target guest_id (\"\" = host-scoped op)")
keyID = flag.String("key-id", "", "key id of the signing key (must match a pinned agent signer)")
@@ -52,6 +52,7 @@ func run() error {
fstype = flag.String("fstype", "ext4", "for storage_wipe: the filesystem to mkfs after wipe")
agentVer = flag.String("agent-version", "", "for agent_update: the target agent version (e.g. 0.70.1)")
sha256Hex = flag.String("sha256", "", "for agent_update: the pinned lowercase-hex sha256 of the target binary")
bundleSHA = flag.String("bundle-sha256", "", "for agent_config_update: the pinned sha256 of felhom-config-bundle.json (R-840)")
keyFile = flag.String("key", "", "operator signing key (ssh private key / sk- key handle) for ssh-keygen -Y sign")
ttl = flag.Duration("ttl", 30*time.Minute, "validity window from now (issued_at..expires_at)")
nonce = flag.String("nonce", "", "explicit nonce (default: a fresh 128-bit random nonce)")
@@ -95,6 +96,17 @@ func run() error {
}
pj, _ := json.Marshal(map[string]string{"version": *agentVer, "sha256": *sha256Hex})
params = string(pj)
case "agent_config_update":
// R-840: the box's root-owned files. The ROOT wrapper verifies this signature itself and refuses a bundle
// whose sha256 is not exactly this one.
if *agentVer == "" || *bundleSHA == "" {
return fmt.Errorf("agent_config_update needs -agent-version and -bundle-sha256 (the pinned bundle hash)")
}
if !isHex64(*bundleSHA) {
return fmt.Errorf("agent_config_update -bundle-sha256 must be 64 lowercase hex chars (got %d)", len(*bundleSHA))
}
pj, _ := json.Marshal(map[string]string{"agent_version": *agentVer, "bundle_sha256": *bundleSHA})
params = string(pj)
default:
params = "{}"
}
+68 -4
View File
@@ -61,7 +61,7 @@ set -euo pipefail
# Script provenance — logged into every bake transcript next to the baked controller tag, so an
# archive can always be traced to the script that produced it. Bump on any behavior change.
GOLDEN_SCRIPT_VERSION="3.0.0"
GOLDEN_SCRIPT_VERSION="3.2.0"
VMID="${1:-9100}"
TEMPLATE="${2:-local:vztmpl/debian-13-standard_13.1-2_amd64.tar.zst}"
@@ -112,7 +112,15 @@ for i in $(seq 1 30); do
if pct exec "$VMID" -- getent hosts download.docker.com >/dev/null 2>&1; then break; fi
sleep 1
done
pct exec "$VMID" -- bash -c '
# v3.1.0 (`11` §5.8): GOLDEN_DOCKER_PKGS pins the APPROVED Docker engine set (the hub's newest Docker release, all six
# "name=version"); without it the newest stable set is installed and the bake log says so.
GOLDEN_DOCKER_PKGS="${GOLDEN_DOCKER_PKGS:-}"
if [[ -n "$GOLDEN_DOCKER_PKGS" ]]; then
echo "[golden] Docker engine set PINNED to the approved release: $GOLDEN_DOCKER_PKGS"
else
echo "[golden] WARNING: GOLDEN_DOCKER_PKGS not set — installing the newest stable Docker set, not an approved one"
fi
pct exec "$VMID" -- env GOLDEN_DOCKER_PKGS="$GOLDEN_DOCKER_PKGS" bash -c '
set -e
export DEBIAN_FRONTEND=noninteractive
apt-get update -qq
@@ -122,7 +130,54 @@ pct exec "$VMID" -- bash -c '
echo "deb [signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/debian trixie stable" \
> /etc/apt/sources.list.d/docker.list
apt-get update -qq
apt-get install -y -qq docker-ce docker-ce-cli containerd.io >/dev/null
if [ -n "$GOLDEN_DOCKER_PKGS" ]; then
apt-get install -y -qq $GOLDEN_DOCKER_PKGS >/dev/null
else
apt-get install -y -qq docker-ce docker-ce-cli containerd.io >/dev/null
fi
dpkg-query -W containerd.io docker-buildx-plugin docker-ce docker-ce-cli docker-ce-rootless-extras docker-compose-plugin 2>/dev/null | sed "s/^/ installed: /"
'
# v3.2.0 (Part F of the R-840 brief, `11` §5.3): GOLDEN_GUEST_PKGS = the newest APPROVED guest release, as
# "name=version …" (the hub's os_releases row for layer guest, IN FORCE — never a cancelled test approval). The bake
# brings every package the template HAS to exactly that version, under felhom-os-apply's rules: never a package the
# template lacks (--only-upgrade), never newer than approved, never a removal or a new package (a simulation is checked
# first and the bake FAILS on either). Empty = no approved guest release in force: the template's versions stay, and
# the box's first night installs whatever release is approved then. Either way the bake PRINTS the first-night count:
# how many installed packages are older than the approved version (target 0).
GOLDEN_GUEST_PKGS="${GOLDEN_GUEST_PKGS:-}"
pct exec "$VMID" -- env GOLDEN_GUEST_PKGS="$GOLDEN_GUEST_PKGS" bash -c '
set -e
export DEBIAN_FRONTEND=noninteractive
if [ -z "$GOLDEN_GUEST_PKGS" ]; then
echo "[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a"
echo "[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): $(apt list --upgradable 2>/dev/null | grep -c /)"
exit 0
fi
want=""
for nv in $GOLDEN_GUEST_PKGS; do
n=${nv%%=*}; v=${nv#*=}
cur=$(dpkg-query -W -f="\${Version}" "$n" 2>/dev/null) || continue # not in the template: never added
dpkg --compare-versions "$cur" lt "$v" && want="$want $n=$v"
done
if [ -n "$want" ]; then
sim=$(apt-get -s install --only-upgrade -o Dpkg::Options::=--force-confold $want)
if echo "$sim" | grep -q "^Remv "; then echo "[golden] FATAL: the approved guest set would REMOVE a package"; echo "$sim" | grep "^Remv "; exit 1; fi
for p in $(echo "$sim" | awk "/^Inst /{print \$2}"); do
dpkg-query -W "$p" >/dev/null 2>&1 || { echo "[golden] FATAL: the approved guest set would ADD $p - not in the template"; exit 1; }
done
apt-get install -y -qq --only-upgrade -o Dpkg::Options::=--force-confold -o Dpkg::Options::=--force-confdef $want >/dev/null
echo "[golden] approved guest release installed: $(echo $want | wc -w) package(s) brought to the approved version"
else
echo "[golden] approved guest release: the template already runs every approved version"
fi
left=0
for nv in $GOLDEN_GUEST_PKGS; do
n=${nv%%=*}; v=${nv#*=}
cur=$(dpkg-query -W -f="\${Version}" "$n" 2>/dev/null) || continue
dpkg --compare-versions "$cur" lt "$v" && { left=$((left+1)); echo " still older: $n $cur < $v"; }
done
echo "[golden] first-night count vs the approved guest release: $left (target 0)"
[ "$left" -eq 0 ] || { echo "[golden] FATAL: $left package(s) stayed older than the approved release"; exit 1; }
'
echo "[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …"
# containerd-snapshotter (Docker 28+/29 default) keeps the IMAGE content store under
@@ -135,9 +190,12 @@ echo "[golden] baking daemon.json: classic overlay2 driver (containerd-snapshott
# layout (a container's `df /` reports the single volume, phase-0 spike). Since v3.0.0 /var/lib/docker
# is a BIND of <volume>/docker rather than the mp0 mount itself, wired immediately below; data-root
# still needs no override because the path is unchanged. Log caps kill the most common runaway.
# v3.1.0: "live-restore": true (`09` decision 87) — a Docker engine update then restarts no app (`11` C5). A box made
# from this golden never needs the agent's one-time live-restore-on step. NEVER removed by a plain restart (R-835).
pct exec "$VMID" -- bash -c 'mkdir -p /etc/docker; cat > /etc/docker/daemon.json <<JSON
{
"features": { "containerd-snapshotter": false },
"live-restore": true,
"log-driver": "json-file",
"log-opts": { "max-size": "10m", "max-file": "3" }
}
@@ -175,6 +233,8 @@ pct exec "$VMID" -- bash -c 'systemctl restart docker; sleep 3; docker run --rm
# Guard: the image store MUST be on the data volume now. /var/lib/containerd holding the images would
# mean containerd-snapshotter is still on (the split would leave images on the rootfs).
pct exec "$VMID" -- bash -c 'drv=$(docker info 2>/dev/null | sed -n "s/.*Storage Driver: //p"); [ "$drv" = "overlay2" ] || { echo "[golden] FATAL: storage driver is $drv, expected overlay2 — images would not land on the data volume"; exit 1; }'
# v3.1.0 ASSERTION: live-restore is ON in the running daemon (decision 87), or the bake fails closed.
pct exec "$VMID" -- bash -c 'lr=$(docker info --format "{{.LiveRestoreEnabled}}" 2>/dev/null); [ "$lr" = "true" ] && echo " live-restore: on" || { echo "[golden] FATAL: live-restore is $lr, expected true (decision 87)"; exit 1; }'
# ASSERTION 1 (RETARGETED v3.0.0, not removed). /var/lib/docker must be a real mount — now the V-c
# bind of <volume>/docker rather than the mp0 mount itself. Still fails closed on the same failure:
# if the bind did not take, Docker's data-root silently sits on the OS rootfs and the golden ships
@@ -289,7 +349,11 @@ mount --make-rshared /mnt
# Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged,
# no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave
# view, and the docker socket. The controller reaches the agent's local API for disk management.
docker run -d --name felhom-controller --restart unless-stopped "${HOSTNAME_ARGS[@]}" \
# R-523: `always`, not `unless-stopped`. It covers ONE extra case only — a Docker daemon restart after
# the container was stopped by hand. Neither policy restarts a container that `docker kill`/`docker
# stop` ended (measured 2026-09-15, Docker 29.8.0, evidence-p1fixes-2026-09-15/A1); the host agent's
# controller supervisor (felhom-agent v0.131.0, internal/localapi/controllersupervisor.go) covers that.
docker run -d --name felhom-controller --restart always "${HOSTNAME_ARGS[@]}" \
-e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \
-v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \
-v felhom-controller-data:/opt/docker/felhom-controller \
+8
View File
@@ -0,0 +1,8 @@
# /etc/felhom/crash-guard.conf — read by /usr/local/sbin/felhom-crash-guard (`11` §5.9).
# The LIMIT-th unclean stop within WINDOW_MINUTES leaves the box off. Decided by CC unattended — operator may reverse.
LIMIT=3
WINDOW_MINUTES=60
# kernel.panic while armed: seconds after a crash before the kernel restarts the box.
PANIC_SECONDS=10
# A tripped guard re-arms after this many hours of normal running (or `felhom-crash-guard rearm`).
REARM_HOURS=24
+116 -107
View File
@@ -1,30 +1,38 @@
# felhom-agent sudoers allowlist — the NARROW host-root surface (slice 5 Phase B, doc 03 §3/§7).
# felhom-agent sudoers allowlist — the NARROW host-root surface (slice 5 Phase B, doc 03 §3/§7; narrowed R-861).
#
# Install as a drop-in: /etc/sudoers.d/felhom-agent (mode 0440, root:root), validated with
# `visudo -cf`. The agent runs as the non-root `felhom-agent` service user and shells out via
# `sudo -n` with FIXED argument vectors (no shell). The fine-grained validation is done IN
# the agent BEFORE exec (internal/storage/validate.go): UUIDs against a strict hex regex,
# mount paths confined+traversal-checked, SMART devices whitelisted to raw disks, LVM names
# charset-checked. These sudoers wildcards are the COARSE allowlist; the agent is the fine
# gate, so a wildcard can never be abused by a value the agent didn't already validate.
# Install as a drop-in: /etc/sudoers.d/felhom-agent (mode 0440, root:root), validated with `visudo -cf`. It rides the
# signed config bundle (R-840). The agent runs as the non-root `felhom-agent` user and shells out via `sudo -n` with
# FIXED argument vectors (no shell).
#
# Binary paths MUST match the agent config (privileged.systemctl/install/smartctl/lvs). Adjust
# for your distro (Debian/PVE shown). A missing/declined entry degrades the agent with a
# warning (SMART→UNKNOWN, mount→logged error), it does not crash.
# R-861 (agent v0.146.0) — EXACT PATTERNS, NOT GLOBS. A sudoers `*` in the ARGUMENTS also matches spaces, so
# `pct set [0-9]* -onboot 1` matched `pct set 100 --dev0 /dev/sda -onboot 1` (a raw host disk for the guest), and
# `mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*` matched a `..` path onto /etc. Every argument list that varies
# is now a sudo regular expression (`^...$`, sudo >= 1.9.10; Debian 13 / PVE 9 ship 1.9.16): one value per slot, a
# fixed character set, no `..`, no extra argument. Lines with no variable part stay literal. The patterns are pinned
# by the capability manifest (every real call must match: TestManifestCoveredBySudoers) and by injection cases that
# must NOT match (configs/test_sudoers_patterns.py, and live with `sudo -l -U felhom-agent` on the demo boxes).
#
# R-861 — NO FILE THE AGENT WROTE IS INSTALLED WHERE ROOT READS IT. The `install` lines are gone: a mount unit, a
# dnsmasq drop-in, the WireGuard config and the OOB sshd config + key go through `felhom-priv-apply`, a root wrapper
# from the bundle that checks the CONTENT against the agent's own renderers; the guest pre-start hook and the shared
# drive parent are FIXED files that come with the bundle itself; the agent binary is replaced only by an
# operator-signed agent_update that `felhom-os-apply` verifies as root (`felhom-selfupdate-guarded apply` is no longer
# here). The two remaining root runs of agent code (FELHOM_ESCROW, the guest hook) therefore run only a signed binary.
#
# Binary paths MUST match the agent config (privileged.systemctl/install/smartctl/lvs). A missing/declined entry
# degrades the agent with a warning (SMART→UNKNOWN, mount→logged error), it does not crash; the capability probe
# reports it to the hub.
Cmnd_Alias FELHOM_MOUNT = \
/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/* /etc/systemd/system/*.mount, \
/usr/local/sbin/felhom-priv-apply ^unit mnt-[A-Za-z0-9_.\\-]+\.(mount|automount)$, \
/usr/bin/systemctl daemon-reload, \
/usr/bin/systemctl enable --now -- *.mount, \
/usr/bin/systemctl disable -- *.mount, \
/usr/bin/systemctl stop -- *.mount
/usr/bin/systemctl ^enable --now -- mnt-[A-Za-z0-9_.\\-]+\.mount$, \
/usr/bin/systemctl ^disable -- mnt-[A-Za-z0-9_.\\-]+\.mount$, \
/usr/bin/systemctl ^stop -- mnt-[A-Za-z0-9_.\\-]+\.mount$
Cmnd_Alias FELHOM_DISK = \
/usr/sbin/smartctl -a -j /dev/sd[a-z]*, \
/usr/sbin/smartctl -a -j /dev/nvme[0-9]*n[0-9]*, \
/usr/sbin/smartctl -a -j /dev/vd[a-z]*, \
/usr/sbin/smartctl -a -j /dev/hd[a-z]*, \
/usr/sbin/lvs --reportformat json --units b -o lv_name\,data_percent\,metadata_percent -- *, \
/usr/sbin/smartctl ^-a -j /dev/(sd[a-z]+|nvme[0-9]+n[0-9]+|vd[a-z]+|hd[a-z]+)$, \
/usr/sbin/lvs ^--reportformat json --units b -o lv_name\,data_percent\,metadata_percent -- [A-Za-z0-9_.+-]+(/[A-Za-z0-9_.+-]+)?$, \
/usr/sbin/pvs --reportformat json --noheadings -o pv_name, \
/usr/sbin/zpool status -P
@@ -35,9 +43,9 @@ Cmnd_Alias FELHOM_DISK = \
# (the wildcard only ever names a path the agent itself created), and the bootstrap file the agent
# writes there is the only thing these touch. ':' is escaped per sudoers grammar.
Cmnd_Alias FELHOM_PROVISION = \
/usr/bin/chown -R 100000\:100000 /var/lib/felhom-agent/guests/*, \
/usr/sbin/pct set [0-9]* -mp[0-9]* /var/lib/felhom-agent/guests/*, \
/usr/sbin/pct set [0-9]* -onboot 1
/usr/bin/chown ^-R 100000\:100000 /var/lib/felhom-agent/guests/[0-9]+(/bootstrap)?$, \
/usr/sbin/pct ^set [0-9]+ -mp[0-9]+ /var/lib/felhom-agent/guests/[0-9]+/bootstrap\,mp\=/[A-Za-z0-9/_.-]+(\,ro\=1)?$, \
/usr/sbin/pct ^set [0-9]+ -onboot 1$
# Disk inspection + format (slice 8C + Impl-1). blkid/lsblk read the device's data-bearing evidence
# (the agent decides data-bearing-ness from THIS, never the caller's claim). Format goes ONLY through
@@ -45,66 +53,54 @@ Cmnd_Alias FELHOM_PROVISION = \
# mkfs the OS disk — the wrapper re-checks the catastrophic cases (system disk / LVM PV / foreign mount)
# as root and refuses, and the agent's unclaimed-disk filter (claim.go) is the primary guard above it.
Cmnd_Alias FELHOM_FORMAT = \
/usr/sbin/blkid -p -o export /dev/*, \
/usr/bin/lsblk -J -o NAME\,FSTYPE\,PTTYPE\,MOUNTPOINT /dev/*, \
/usr/local/sbin/felhom-mkfs-guarded /dev/* *
/usr/sbin/blkid ^-p -o export /dev/[^ ]+$, \
/usr/bin/lsblk ^-J -o NAME\,FSTYPE\,PTTYPE\,MOUNTPOINT /dev/[^ ]+$, \
/usr/local/sbin/felhom-mkfs-guarded ^/dev/[^ ]+ (ext4|xfs)$
# LAN split-horizon resolver (internal/lanresolver): the agent manages a host-side dnsmasq that
# answers *.<customer-domain> with each guest's live LAN IP. install only ever writes felhom-*.conf
# drop-ins (from agent-written /tmp temp files); the two `pct exec` reads are FIXED command vectors
# answers *.<customer-domain> with each guest's live LAN IP. A felhom-*.conf drop-in reaches /etc/dnsmasq.d
# only through felhom-priv-apply, which allows exactly the lines the resolver renders (R-861: a `dhcp-script=` would
# run as root); the two `pct exec` reads are FIXED command vectors
# (the guest's eth0 IPv4 + the controller's pulled controller.yaml for the domain) — NOT a general
# `pct exec`. systemctl is scoped to the dnsmasq unit only. The agent never edits /etc/resolv.conf.
Cmnd_Alias FELHOM_DNSMASQ = \
/usr/bin/apt-get install -y -q dnsmasq, \
/usr/bin/install -m 0644 /tmp/felhom-resolver-*.conf /etc/dnsmasq.d/felhom-*.conf, \
/usr/local/sbin/felhom-priv-apply ^dnsmasq /tmp/felhom-resolver-[0-9]+\.conf felhom-[a-z0-9][a-z0-9._-]*\.conf$, \
/usr/bin/systemctl enable --now dnsmasq, \
/usr/bin/systemctl reload dnsmasq, \
/usr/bin/systemctl restart dnsmasq, \
/usr/bin/rm -f /etc/dnsmasq.d/felhom-*.conf, \
/usr/sbin/pct exec [0-9]* -- ip -4 -o addr show dev eth0, \
/usr/sbin/pct exec [0-9]* -- docker exec felhom-controller cat /opt/docker/felhom-controller/controller.yaml
/usr/bin/rm ^-f /etc/dnsmasq\.d/felhom-[a-z0-9][a-z0-9._-]*\.conf$, \
/usr/sbin/pct ^exec [0-9]+ -- ip -4 -o addr show dev eth0$, \
/usr/sbin/pct ^exec [0-9]+ -- docker exec felhom-controller cat /opt/docker/felhom-controller/controller\.yaml$
# Guest mountpoint lifecycle (intermediary-mount re-architecture + C1 net). The pre-start self-heal hook
# wrapper is installed once into the PVE snippets dir (from an agent-written /tmp file) and registered
# per-guest; decommission/eject DELETE the dead mountpoint slot so a missing bind source can't brick the
# guest at next boot (the B3 C1 fix). The agent fine-validates the vmid (numeric) + slot (mp[0-9]+) and
# the snippet path is fixed — the wildcards are the coarse allowlist. The install SOURCE is a
# random-named agent temp (os.CreateTemp, audit B1 — a fixed /tmp name was a local TOCTOU), hence the
# glob; the DESTINATION stays pinned. The `mkdir -p` creates the snippets dir on a FRESH box —
# `install` won't create parents, so without it the hook install failed silently on Day-0 boxes
# (B2, DRILL-day0-cleanroom-2026-07-03; fixed agent v0.63.0).
# Guest mountpoint lifecycle (intermediary-mount re-architecture + C1 net). The pre-start self-heal hook is a FIXED file
# from the config bundle (/var/lib/vz/snippets/felhom-guest-hook.sh — R-861: the agent no longer installs it from /tmp;
# Proxmox runs it as root at every guest start). The agent only registers it per guest, deletes a dead mountpoint slot
# (the B3 C1 fix) and reboots a guest to activate binds — each with an exact vmid / slot.
Cmnd_Alias FELHOM_GUESTHOOK = \
/usr/bin/mkdir -p /var/lib/vz/snippets, \
/usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-*.sh /var/lib/vz/snippets/felhom-guest-hook.sh, \
/usr/sbin/pct set [0-9]* --hookscript local\:snippets/felhom-guest-hook.sh, \
/usr/sbin/pct set [0-9]* --delete mp[0-9]*, \
/usr/sbin/pct reboot [0-9]*
/usr/sbin/pct ^set [0-9]+ --hookscript local\:snippets/felhom-guest-hook\.sh$, \
/usr/sbin/pct ^set [0-9]+ --delete mp[0-9]+$, \
/usr/sbin/pct ^reboot [0-9]+$
# Intermediary mount model (the drive hot-swap re-architecture). The agent keeps a SHARED host parent
# /mnt/felhom-drives (self-bind + make-shared + a boot-persistence systemd unit) and binds/unbinds each
# drive's felhom-data namespace UNDERNEATH it so the change propagates into the running guest live (no
# pct, no reboot). The agent fine-validates the drive name + confines paths before any exec; the trailing
# `*` (matching the comma-laden mp spec) mirrors the existing FELHOM_PROVISION pattern.
# `lxc-info -n <vmid> -p -H` resolves the guest init PID for the GuestSeesMount / bound_under_parent check
# (a READ — the drive-gate's "is the drive live in the guest?" signal); WITHOUT it the non-root agent gets
# an empty PID and reports every drive absent (multi-drive flapping, audit 2026-06-29). `make-private`
# isolates the parent's peer group on FIRST setup only (EnsureSharedParent guards on mountpoint, so it
# never re-churns a live parent); without it the parent stays in root's group and submounts double.
# /mnt/felhom-drives (self-bind + make-shared) and binds/unbinds each drive's felhom-data namespace UNDERNEATH it so the
# change propagates into the running guest live. The boot-persistence script + unit are FIXED files from the config
# bundle (R-861: the agent no longer installs them from /tmp); the agent only enables the unit. A drive name is one
# path segment that cannot start with a dot (no `..`); `lxc-info -n <vmid> -p -H` resolves the guest init PID for the
# GuestSeesMount check; `make-private` isolates the parent's peer group on FIRST setup only.
Cmnd_Alias FELHOM_INTERMEDIARY = \
/usr/bin/mkdir -p /mnt/felhom-drives, \
/usr/bin/mkdir -p /mnt/felhom-drives/*, \
/usr/bin/mkdir -p /mnt/*/felhom-data, \
/usr/bin/chown 100000\:100000 /mnt/*/felhom-data, \
/usr/bin/mkdir ^-p /mnt/felhom-drives/[A-Za-z0-9_-][A-Za-z0-9_.-]*$, \
/usr/bin/mkdir ^-p /mnt/[A-Za-z0-9_-][A-Za-z0-9_.-]*/felhom-data$, \
/usr/bin/chown ^100000\:100000 /mnt/[A-Za-z0-9_-][A-Za-z0-9_.-]*/felhom-data$, \
/usr/bin/mount --bind /mnt/felhom-drives /mnt/felhom-drives, \
/usr/bin/mount --make-shared /mnt/felhom-drives, \
/usr/bin/mount --make-private /mnt/felhom-drives, \
/usr/bin/mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*, \
/usr/bin/umount /mnt/felhom-drives/*, \
/usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-*.sh /usr/local/sbin/felhom-shared-parent.sh, \
/usr/bin/install -m 0644 -- /tmp/felhom-shared-parent-*.service /etc/systemd/system/felhom-shared-parent.service, \
/usr/bin/mount ^--bind /mnt/[A-Za-z0-9_-][A-Za-z0-9_.-]*/felhom-data /mnt/felhom-drives/[A-Za-z0-9_-][A-Za-z0-9_.-]*$, \
/usr/bin/umount ^/mnt/felhom-drives/[A-Za-z0-9_-][A-Za-z0-9_.-]*$, \
/usr/bin/systemctl enable felhom-shared-parent.service, \
/usr/bin/lxc-info -n [0-9]* -p -H, \
/usr/sbin/pct set [0-9]* -mp8 /mnt/felhom-drives*
/usr/bin/lxc-info ^-n [0-9]+ -p -H$, \
/usr/sbin/pct ^set [0-9]+ -mp8 /mnt/felhom-drives\,mp\=/mnt/felhom-drives$
# Controller-swap / managed auto-update (Option A, non-root). The agent owns the in-guest controller
# image SWAP (it survives the controller being killed mid-swap): read the baked image ref, check the
@@ -115,15 +111,17 @@ Cmnd_Alias FELHOM_INTERMEDIARY = \
# docker inspect -f * — container running/health/image (read-only; `*` spans the -f template
# + container across spaces, spike-confirmed)
# systemctl restart <fixed unit> — re-run the golden's bootstrap (the only state change)
# tee <FIXED image file> — WRITE the ref; content is fed on STDIN (no shell, no interpolation),
# the agent strict-validates the ref (controllerImageRe) before the write.
# felhom-priv-apply controller-image <vmid> — WRITE the ref (R-861 (a) A1, `09` §3 decision 165): the ref goes on
# STDIN to the ROOT wrapper, which requires our registry + repository + an x.y.z tag
# and writes the guest file itself. The agent's own `tee` grant is GONE: before, a
# compromised agent could hand the guest's bootstrap ANY image (sudo cannot see stdin).
# Validated GO: felhom.eu/documentation/audits/SPIKE-controllerswap-narrow-grants-2026-06-29.md.
Cmnd_Alias FELHOM_CONTROLLERSWAP = \
/usr/sbin/pct exec [0-9]* -- cat /etc/felhom-controller-image, \
/usr/sbin/pct exec [0-9]* -- docker image inspect *, \
/usr/sbin/pct exec [0-9]* -- docker inspect -f *, \
/usr/sbin/pct exec [0-9]* -- systemctl restart felhom-controller-bootstrap.service, \
/usr/sbin/pct exec [0-9]* -- tee /etc/felhom-controller-image
/usr/sbin/pct ^exec [0-9]+ -- cat /etc/felhom-controller-image$, \
/usr/sbin/pct ^exec [0-9]+ -- docker image inspect gitea\.dooplex\.hu/admin/felhom-controller\:[0-9]+\.[0-9]+\.[0-9]+$, \
/usr/sbin/pct ^exec [0-9]+ -- docker inspect -f .+ (felhom-controller|cloudflared)$, \
/usr/sbin/pct ^exec [0-9]+ -- systemctl restart felhom-controller-bootstrap\.service$, \
/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$
# Stale-lock recovery (F2-b, v0.49.0). A host reboot DURING a vzdump backup leaves the guest with a
# `snapshot-delete`/`backup` lock + `onboot:1` then can't start it → the customer box stays DOWN. The
@@ -131,7 +129,15 @@ Cmnd_Alias FELHOM_CONTROLLERSWAP = \
# with no API equivalent (snapshot-delete + start go through the API token); the agent fine-validates the
# vmid (numeric) before exec — the `[0-9]*` is the coarse allowlist.
Cmnd_Alias FELHOM_STALELOCK = \
/usr/sbin/pct unlock [0-9]*
/usr/sbin/pct ^unlock [0-9]+$
# Weekly guest disk trim (R-444, operator ruling `09` §3 decision 139). A thin pool only ever grows from blocks the
# guest has FREED: `fstrim` inside the unprivileged container is refused (FITRIM: Operation not permitted), so the host
# trims the guest's mounts. Measured on demo-hp 2026-10-06: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %,
# apps kept answering. ONE exact pattern: a vmid and nothing else — no `--ignore-mountpoints`, no second argument
# (pinned: TestSudoersFstrimRuleIsExact). The agent runs it on a weekly daytime timer under the heavy-op gate.
Cmnd_Alias FELHOM_FSTRIM = \
/usr/sbin/pct ^fstrim [0-9]+$
# Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS
# leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and
@@ -160,8 +166,9 @@ Cmnd_Alias FELHOM_SCRATCH_TEARDOWN = \
# into the guest through the existing shared bind (an unprivileged LXC cannot mount NFS/CIFS itself).
# A NAS is NOT a drive — no durable-id, no SMART, no wipe; these grants only install/enable/remove the
# unit pair. The agent fine-validates every value (share name, server, export, uid/gid, creds path) before
# any unit is rendered (internal/storage/netmount.go ValidateNetworkMountSpec); the trailing globs are the
# COARSE allowlist. The `.mount` install/enable/disable/stop reuse FELHOM_MOUNT; this alias adds the
# any unit is rendered (internal/storage/netmount.go ValidateNetworkMountSpec); the unit FILE reaches
# /etc/systemd/system only through `felhom-priv-apply unit` (FELHOM_MOUNT), which requires nosuid,nodev on a network
# share (R-861). The `.mount` enable/disable/stop reuse FELHOM_MOUNT; this alias adds the
# `.automount` variants + the unit-file removal. The unit FILE name is the systemd-escaped mountpoint,
# which always begins `mnt-felhom` (the mountpoint is /mnt/felhom-drives/<name>), so the rm glob is scoped
# to felhom mount units only. mkdir of the mountpoint reuses FELHOM_INTERMEDIARY's /mnt/felhom-drives/*.
@@ -176,51 +183,46 @@ Cmnd_Alias FELHOM_SCRATCH_TEARDOWN = \
# behind (the campaign accumulated 10 stub-shaped leftovers). rmdir ONLY (never rm -rf): it refuses
# a non-empty dir, so unexpected data is preserved, not destroyed — a fail-safe grant.
Cmnd_Alias FELHOM_NETMOUNT = \
/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/* /etc/systemd/system/*.automount, \
/usr/bin/systemctl enable --now -- *.automount, \
/usr/bin/systemctl disable -- *.automount, \
/usr/bin/systemctl stop -- *.automount, \
/usr/bin/systemctl reset-failed -- mnt-felhom*, \
/usr/bin/rmdir /mnt/felhom-drives/*, \
/usr/bin/rm -f /etc/systemd/system/mnt-felhom*
/usr/bin/systemctl ^enable --now -- mnt-[A-Za-z0-9_.\\-]+\.automount$, \
/usr/bin/systemctl ^disable -- mnt-[A-Za-z0-9_.\\-]+\.automount$, \
/usr/bin/systemctl ^stop -- mnt-[A-Za-z0-9_.\\-]+\.automount$, \
/usr/bin/systemctl ^reset-failed -- mnt-felhom[A-Za-z0-9_.\\-]*\.(mount|automount)$, \
/usr/bin/rmdir ^/mnt/felhom-drives/[A-Za-z0-9_-][A-Za-z0-9_.-]*$, \
/usr/bin/rm ^-f /etc/systemd/system/mnt-felhom[A-Za-z0-9_.\\-]*\.(mount|automount)$
# Offsite WG tunnel (S3, doc 06 §3.3). The agent manages wg-quick@wg-felhom as an agent-managed
# host service (the dnsmasq/lanresolver shape): conf staged in the agent-owned StateDir (never
# /tmp), installed 0600 to the FIXED destination, unit enable/restart/disable. The ONLY wg read
# /tmp), installed 0600 to the FIXED destination by `felhom-priv-apply wg`, which refuses any key renderConf never
# writes (R-861: PostUp/PreUp run as root under wg-quick), unit enable/restart/disable. The ONLY wg read
# is `latest-handshakes` — `wg show <if> dump` is FORBIDDEN everywhere (its interface line
# carries the PRIVATE KEY; the S1 session-log incident). Both install paths are FIXED (no glob):
# the agent has exactly one tunnel conf to manage.
# carries the PRIVATE KEY; the S1 session-log incident). Source and destination are fixed in the wrapper.
Cmnd_Alias FELHOM_WG = \
/usr/bin/apt-get install -y -q wireguard-tools, \
/usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf, \
/usr/local/sbin/felhom-priv-apply wg, \
/usr/bin/systemctl enable --now wg-quick@wg-felhom, \
/usr/bin/systemctl restart wg-quick@wg-felhom, \
/usr/bin/systemctl disable --now wg-quick@wg-felhom, \
/usr/bin/wg show wg-felhom latest-handshakes
# Agent self-update (TASK D1, SPIKE-agent-selfupdate-2026-07-05). The agent downloads the
# operator-SIGNED binary (sha256 pinned in the signed op — neither hub nor Gitea compromise can
# substitute it), verifies the sha in-process, then hands off to the guarded wrapper, which
# RE-verifies the sha as root, confines the staged path to /var/lib/felhom-agent/selfupdate/,
# performs the A/B flip (atomic same-fs rename, .prev retained) and schedules a detached restart.
# The apply args are a COARSE glob (spike S4b: sudoers fnmatch makes a [a-f0-9]* sha pattern
# first-char-only anyway) — the wrapper's own sha re-verify + path confinement is the real gate.
# `rollback` is normally run by felhom-agent-rollback.service (root, OnFailure=), not via sudo;
# granting it here keeps the verb probe-able (capability self-check) and operator-invokable.
# Agent self-update (TASK D1; R-861). The A/B flip (`felhom-selfupdate-guarded apply`) is NO LONGER the agent's: the
# agent hands the operator-SIGNED agent_update to felhom-os-apply (FELHOM_OSAPPLY, mode agent_update), which verifies
# the signature as root and only then runs the flip. Until v0.146.0 the agent passed the sha itself, so a compromised
# agent could install any binary — the binary FELHOM_ESCROW and the guest hook run as root. `commit` (clear the pending
# marker) and `rollback` (pending-guarded revert, normally run by felhom-agent-rollback.service) stay.
Cmnd_Alias FELHOM_SELFUPDATE = \
/usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/* *, \
/usr/local/sbin/felhom-selfupdate-guarded commit, \
/usr/local/sbin/felhom-selfupdate-guarded rollback
# Dedicated OOB sshd (TASK H1). The agent manages felhom-sshd like wg-felhom/dnsmasq: it RENDERS the
# config (Port from its claim) + the operator's authorized_keys, validates with `sshd -t`, and reloads
# (never restart-on-change [SF-2]). Both install SOURCES are the agent-owned staged files under
# StateDir; both DESTINATIONS are FIXED. `sshd -t/-T` are the validate/discover reads. The
# (never restart-on-change [SF-2]). Both files reach /etc/felhom-sshd only through felhom-priv-apply (R-861): the config
# must be the ONE template with only the Port varying (an AuthorizedKeysFile the agent owns + `StrictModes no` would be
# a root login), the key file one plain public key without options. `sshd -t/-T` are the validate/discover reads. The
# systemctl verbs are SCOPED to felhom-sshd only. reset-failed precedes a deliberate restart [SF-5].
# NOTHING here can touch the stock sshd, :22, or /etc/ssh.
Cmnd_Alias FELHOM_SSHD = \
/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config, \
/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/authorized_keys.felhom-op /etc/felhom-sshd/authorized_keys/felhom-op, \
/usr/local/sbin/felhom-priv-apply sshd-config, \
/usr/local/sbin/felhom-priv-apply sshd-key, \
/usr/sbin/sshd -t -f /var/lib/felhom-agent/felhom-sshd/sshd_config, \
/usr/sbin/sshd -t -f /etc/felhom-sshd/sshd_config, \
/usr/sbin/sshd -T -f /etc/felhom-sshd/sshd_config, \
@@ -270,8 +272,8 @@ Cmnd_Alias FELHOM_OOB = \
/usr/sbin/nft list set inet felhom_oob ssh_port, \
/usr/sbin/nft flush set inet felhom_oob operator_ips, \
/usr/sbin/nft flush set inet felhom_oob ssh_port, \
/usr/sbin/nft add element inet felhom_oob operator_ips *, \
/usr/sbin/nft add element inet felhom_oob ssh_port *
/usr/sbin/nft ^add element inet felhom_oob operator_ips \{ [0-9.]+(/[0-9]+)? \}$, \
/usr/sbin/nft ^add element inet felhom_oob ssh_port \{ [0-9]+ \}$
# Escrow ceremony (controller-driven, TASK 2026-07-13; mechanics validated by
# SPIKE-controller-escrow-2026-07-13). ONE fixed argv — sudoers matches the argument vector
@@ -299,10 +301,17 @@ Cmnd_Alias FELHOM_SELFHEAL = \
# argument after the numeric vmid is a literal, so the grant cannot be widened by anything the guest or
# the hub says. The address read is deliberately NOT duplicated here — it is already FELHOM_DNSMASQ's,
# and the same command must not be granted twice under two names.
Cmnd_Alias FELHOM_GUESTNET = \
/usr/sbin/pct exec [0-9]* -- ip route show default, \
/usr/sbin/pct exec [0-9]* -- cat /etc/network/interfaces, \
/usr/sbin/pct exec [0-9]* -- pgrep -x dhclient, \
/usr/sbin/pct exec [0-9]* -- dhclient -pf /run/dhclient.eth0.pid -lf /var/lib/dhcp/dhclient.eth0.leases eth0
# OS updates, guest fast lane (`11-os-updates.md` §5.4.1, agent v0.140.0). The ONLY entry: the root wrapper with
# one plan file in the agent's own os/ dir. Every safety rule (no removal, no downgrade, no new or unlisted package,
# Debian origin only, the box's own customer guest only) lives in the wrapper, red-proved per rule
# (configs/test_felhom_os_apply.py). The agent gets NO apt grant of its own.
Cmnd_Alias FELHOM_OSAPPLY = \
/usr/local/sbin/felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN
Cmnd_Alias FELHOM_GUESTNET = \
/usr/sbin/pct ^exec [0-9]+ -- ip route show default$, \
/usr/sbin/pct ^exec [0-9]+ -- cat /etc/network/interfaces$, \
/usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \
/usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_FSTRIM, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
+235
View File
@@ -0,0 +1,235 @@
#!/usr/bin/python3
# felhom-crash-guard — a crashed host restarts by itself, but not forever (`09` decision 88, R-851, `11` §5.9).
#
# Install as /usr/local/sbin/felhom-crash-guard (0755 root:root), with felhom-crash-guard.service (boot / clean-stop)
# and felhom-crash-guard-check.timer (hourly re-arm check). Python 3, standard library only.
# Tests: configs/test_felhom_crash_guard.py (temp dirs; nothing real is touched).
#
# WHAT IT DOES
# boot early at every boot. Was the previous boot ended CLEANLY? (the clean-stop marker exists). If not, this
# boot follows an UNCLEAN stop — a kernel crash, a power cut or a hard reset (they cannot be told apart
# on these boxes: measured 2026-10-04 on demo-hp, efi_pstore is on yet saved NOTHING for a real panic;
# the journal and `last` show only "no shutdown"). It records the unclean boot, counts those in the last
# WINDOW_MINUTES, and sets kernel.panic:
# - fewer than LIMIT-1 recent unclean boots → kernel.panic = PANIC_SECONDS (a crash restarts the box);
# - LIMIT-1 or more → the guard TRIPS: kernel.panic = 0, so the LIMIT-th crash within the window
# leaves the box OFF (operator's own words: "if it crashes 3 times within one hour, it stays off").
# A tripped guard stays tripped across further boots until it re-arms.
# clean-stop ExecStop of the service: writes the clean-stop marker during an orderly shutdown or reboot.
# check hourly: a tripped guard re-arms after REARM_HOURS of normal running (since the trip AND since boot).
# rearm the operator re-arms by hand (`felhom-crash-guard rearm`).
# status prints the state.
# The state is /var/lib/felhom-crash-guard/state.json (0644: the non-root agent reads it into its host report).
# Before the service runs (very early boot) the kernel default kernel.panic = 0 applies, so a crash THAT early leaves
# the box off — the safe side: a box that cannot reach userspace must not loop.
import json
import os
import sys
import time
CONF = "/etc/felhom/crash-guard.conf"
STATE_DIR = "/var/lib/felhom-crash-guard"
DEFAULTS = {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10, "REARM_HOURS": 24}
class Env:
"""Paths and clock; tests replace them."""
def __init__(self, conf=CONF, state_dir=STATE_DIR, panic_path="/proc/sys/kernel/panic",
uptime_path="/proc/uptime", boot_id_path="/proc/sys/kernel/random/boot_id"):
self.conf, self.state_dir = conf, state_dir
self.panic_path, self.uptime_path, self.boot_id_path = panic_path, uptime_path, boot_id_path
def now(self):
return time.time()
def log(self, line):
print(line, file=sys.stderr, flush=True)
try:
import subprocess
subprocess.run(["logger", "-t", "felhom-crash-guard", line], timeout=10)
except Exception:
pass
def iso(t):
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(t))
def parse_iso(s):
import calendar
return calendar.timegm(time.strptime(s, "%Y-%m-%dT%H:%M:%SZ"))
def load_conf(env):
c = dict(DEFAULTS)
try:
for line in open(env.conf):
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
k, v = (x.strip() for x in line.split("=", 1))
if k in c and v.isdigit() and int(v) >= (1 if k != "PANIC_SECONDS" else 1):
c[k] = int(v)
except OSError:
pass
return c
def state_path(env):
return os.path.join(env.state_dir, "state.json")
def marker_path(env):
return os.path.join(env.state_dir, "clean-stop")
def load_state(env):
try:
with open(state_path(env)) as f:
s = json.load(f)
return s if isinstance(s, dict) else None
except (OSError, ValueError):
return None
def save_state(env, s):
os.makedirs(env.state_dir, mode=0o755, exist_ok=True)
tmp = state_path(env) + ".tmp"
with open(tmp, "w") as f:
json.dump(s, f, indent=2, sort_keys=True)
f.write("\n")
os.chmod(tmp, 0o644)
os.replace(tmp, state_path(env))
def set_panic(env, seconds):
with open(env.panic_path, "w") as f:
f.write(f"{seconds}\n")
def read(path, default=""):
try:
with open(path) as f:
return f.read().strip()
except OSError:
return default
def summarize(s, c, now):
window = c["WINDOW_MINUTES"] * 60
times = [parse_iso(t) for t in s.get("unclean_boots", [])]
after = parse_iso(s["rearmed_at"]) if s.get("rearmed_at") else 0
# a re-arm starts a fresh window (or the next unclean boot would trip again at once); the history stays
s["unclean_boots_in_window"] = sum(1 for t in times if now - t <= window and t > after)
s["unclean_boots_24h"] = sum(1 for t in times if now - t <= 86400)
s["config"] = c
s["updated_at"] = iso(now)
def boot(env):
c = load_conf(env)
now = env.now()
try:
up = float(read(env.uptime_path, "0").split()[0])
except (ValueError, IndexError):
up = 0.0
boot_at = now - up
prev = load_state(env)
first = prev is None
s = prev or {"version": 1, "unclean_boots": [], "tripped": False}
clean = os.path.exists(marker_path(env))
unclean = (not first) and (not clean)
try:
os.remove(marker_path(env))
except OSError:
pass
# keep 7 days of history (the 24 h figure and the operator's view), drop older
s["unclean_boots"] = [t for t in s.get("unclean_boots", []) if now - parse_iso(t) <= 7 * 86400]
if unclean:
s["unclean_boots"].append(iso(boot_at))
s["last_boot_at"] = iso(boot_at)
s["last_boot_unclean"] = unclean
s["boot_id"] = read(env.boot_id_path, "unknown")
summarize(s, c, now)
if not s.get("tripped") and s["unclean_boots_in_window"] >= c["LIMIT"] - 1:
s["tripped"], s["tripped_at"] = True, iso(now)
s["tripped_reason"] = (f"{s['unclean_boots_in_window']} unclean boots within {c['WINDOW_MINUTES']} minutes — "
f"the next crash leaves the box off (limit {c['LIMIT']})")
env.log(f"crash-guard: TRIPPED: {s['tripped_reason']}")
panic = 0 if s.get("tripped") else c["PANIC_SECONDS"]
set_panic(env, panic)
s["kernel_panic"] = panic
s["armed"] = not s.get("tripped")
save_state(env, s)
env.log(f"crash-guard: boot first={first} unclean={unclean} in-window={s['unclean_boots_in_window']} "
f"tripped={s.get('tripped')} kernel.panic={panic}")
return 0
def clean_stop(env):
os.makedirs(env.state_dir, mode=0o755, exist_ok=True)
with open(marker_path(env), "w") as f:
f.write(iso(env.now()) + "\n")
env.log("crash-guard: clean stop recorded")
return 0
def rearm(env, by):
c = load_conf(env)
now = env.now()
s = load_state(env) or {"version": 1, "unclean_boots": []}
was = bool(s.get("tripped"))
s["tripped"] = False
s["armed"] = True
s["rearmed_at"], s["rearmed_by"] = iso(now), by
if was:
s["last_trip"] = {"at": s.get("tripped_at"), "reason": s.get("tripped_reason")}
s.pop("tripped_at", None)
s.pop("tripped_reason", None)
summarize(s, c, now)
set_panic(env, c["PANIC_SECONDS"])
s["kernel_panic"] = c["PANIC_SECONDS"]
save_state(env, s)
env.log(f"crash-guard: RE-ARMED by {by} (was tripped: {was}); kernel.panic={c['PANIC_SECONDS']}")
return 0
def check(env):
c = load_conf(env)
now = env.now()
s = load_state(env)
if not s:
return 0
if s.get("tripped"):
since = max(parse_iso(s["tripped_at"]), parse_iso(s.get("last_boot_at", s["tripped_at"])))
if now - since >= c["REARM_HOURS"] * 3600:
return rearm(env, f"timer ({c['REARM_HOURS']} h of normal running)")
summarize(s, c, now)
save_state(env, s)
return 0
def main(argv, env=None):
env = env or Env()
cmd = argv[1] if len(argv) == 2 else ""
if cmd == "boot":
return boot(env)
if cmd == "clean-stop":
return clean_stop(env)
if cmd == "check":
return check(env)
if cmd == "rearm":
return rearm(env, "operator")
if cmd == "status":
print(json.dumps(load_state(env), indent=2, sort_keys=True))
return 0
print("usage: felhom-crash-guard boot|clean-stop|check|rearm|status", file=sys.stderr)
return 2
if __name__ == "__main__":
if os.geteuid() != 0:
print("felhom-crash-guard: must run as root", file=sys.stderr)
sys.exit(2)
sys.exit(main(sys.argv))
+7
View File
@@ -0,0 +1,7 @@
# Hourly: a tripped crash guard re-arms after REARM_HOURS of normal running (felhom-crash-guard check).
[Unit]
Description=Felhom crash guard re-arm check
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/felhom-crash-guard check
+9
View File
@@ -0,0 +1,9 @@
[Unit]
Description=Felhom crash guard re-arm check (hourly)
[Timer]
OnBootSec=15min
OnUnitActiveSec=1h
[Install]
WantedBy=timers.target
+19
View File
@@ -0,0 +1,19 @@
# felhom-crash-guard — a crashed host restarts by itself, with a limit (`09` decision 88, R-851, `11` §5.9).
# Starts early at boot (sets kernel.panic for THIS boot); its ExecStop writes the clean-stop marker during an orderly
# shutdown or reboot. A boot that finds no marker followed a crash, a power cut or a hard reset.
[Unit]
Description=Felhom crash guard (restart after a kernel crash, with a limit)
DefaultDependencies=no
After=local-fs.target
Before=sysinit.target shutdown.target
Conflicts=shutdown.target
RequiresMountsFor=/var/lib
[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=/usr/local/sbin/felhom-crash-guard boot
ExecStop=/usr/local/sbin/felhom-crash-guard clean-stop
[Install]
WantedBy=sysinit.target
+4
View File
@@ -0,0 +1,4 @@
#!/bin/sh
# felhom-agent guest pre-start self-heal hook (C1 net). PVE calls: <script> <vmid> <phase>.
/usr/local/bin/felhom-agent guest-hook "$1" "$2" || true
exit 0
+5 -3
View File
@@ -16,8 +16,10 @@ Cmnd_Alias FELHOM_OP_REPAIR = \
/usr/bin/systemctl reset-failed felhom-sshd, \
/usr/bin/systemctl restart felhom-sshd, \
/usr/sbin/pct list, \
/usr/sbin/pct start [0-9]*, \
/usr/sbin/pct stop [0-9]*, \
/usr/sbin/pct unlock [0-9]*
/usr/sbin/pct ^start [0-9]+$, \
/usr/sbin/pct ^stop [0-9]+$, \
/usr/sbin/pct ^unlock [0-9]+$
# R-861 (b) B2 (`09` §3 decision 165, hygiene): one numeric vmid per pct verb, anchored — the old glob `[0-9]*` also
# matched spaces, so `pct stop 9201 --skiplock 1` passed. Pinned by TestFelhomOpSudoersPctIsExact.
felhom-op ALL=(root) NOPASSWD: FELHOM_OP_REPAIR
+1743
View File
File diff suppressed because it is too large Load Diff
+468
View File
@@ -0,0 +1,468 @@
#!/usr/bin/python3
"""felhom-priv-apply — the ROOT half of every file the agent writes into a root-read place (R-861, `03` §3.1).
Install as /usr/local/sbin/felhom-priv-apply (0755 root:root) — it rides the signed config bundle (R-840).
WHY IT EXISTS. Until agent v0.146.0 the agent's sudoers let it `install` a file it had written itself into a place a
root program reads: a systemd .mount unit (a bind mount of an agent-owned directory over /etc/sudoers.d is a root
shell), a dnsmasq drop-in (`dhcp-script=` runs as root), the WireGuard config (`PostUp=` runs as root) and the OOB
sshd config (`AuthorizedKeysFile` + `StrictModes no`). A compromised agent PROCESS was therefore root on its host.
Now the agent stages the file and this wrapper — root-owned, delivered only by an operator-signed bundle — checks
the CONTENT against the exact grammar the agent's own renderers produce, and refuses anything else. The agent can no
longer name the destination: each verb has a fixed source and a fixed (or strictly named) destination.
Verbs (each one sudoers line, exact-match pattern):
unit <name> /var/lib/felhom-agent/units/<name> -> /etc/systemd/system/<name> (.mount | .automount)
dnsmasq <tmp> <name> /tmp/felhom-resolver-<digits>.conf -> /etc/dnsmasq.d/felhom-<...>.conf
wg /var/lib/felhom-agent/wg/wg-felhom.conf -> /etc/wireguard/wg-felhom.conf (0600)
sshd-config /var/lib/felhom-agent/felhom-sshd/sshd_config -> /etc/felhom-sshd/sshd_config
sshd-key /var/lib/felhom-agent/felhom-sshd/authorized_keys.felhom-op -> /etc/felhom-sshd/authorized_keys/felhom-op
controller-image <vmid> the ref on STDIN -> /etc/felhom-controller-image INSIDE guest <vmid> (R-861 (a) A1): only
our registry + our repository + an x.y.z tag; the agent no longer has a `tee` grant
--self-check prints "felhom-priv-apply ok verbs=..." (the bundle's self-check)
Exit codes: 0 installed (or already identical), 2 usage, 3 refused (content or source), 4 install failed.
Every refusal is logged to the journal (tag felhom-priv-apply) with its rule; file CONTENT is never logged.
Pinned by configs/test_felhom_priv_apply.py (one test per rule, red-proofs in the R-861 audit).
"""
import ipaddress
import os
import re
import stat
import subprocess
import sys
AGENT_USER = "felhom-agent"
STATE = "/var/lib/felhom-agent"
UNITS_SRC = STATE + "/units"
UNIT_DIR = "/etc/systemd/system"
DNSMASQ_DIR = "/etc/dnsmasq.d"
WG_SRC, WG_DEST = STATE + "/wg/wg-felhom.conf", "/etc/wireguard/wg-felhom.conf"
SSHD_SRC, SSHD_DEST = STATE + "/felhom-sshd/sshd_config", "/etc/felhom-sshd/sshd_config"
KEY_SRC, KEY_DEST = STATE + "/felhom-sshd/authorized_keys.felhom-op", "/etc/felhom-sshd/authorized_keys/felhom-op"
MAX_BYTES = 64 * 1024
VERBS = ("unit", "dnsmasq", "wg", "sshd-config", "sshd-key", "controller-image")
# R-861 (a) A1 (`09` §3 decision 165): the SAME pattern as the agent's controllerImageRe (internal/localapi/
# controllerswap.go) — a compromised agent cannot hand the guest's bootstrap any other image. Pinned by
# configs/test_felhom_priv_apply.py ControllerImage.
CONTROLLER_IMAGE_RE = re.compile(r"^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$")
CONTROLLER_IMAGE_FILE = "/etc/felhom-controller-image"
CONTROLLER_IMAGE_MAX = 256
VMID_RE = re.compile(r"^[0-9]{1,9}$")
UNIT_NAME_RE = re.compile(r"^mnt-[A-Za-z0-9_.\\-]+\.(mount|automount)$")
DNSMASQ_TMP_RE = re.compile(r"^/tmp/felhom-resolver-[0-9]+\.conf$")
DNSMASQ_NAME_RE = re.compile(r"^felhom-[a-z0-9][a-z0-9._-]*\.conf$")
SEG = r"[A-Za-z0-9_-][A-Za-z0-9_.-]*"
WHERE_RE = re.compile(r"^/mnt/(felhom-drives/)?" + SEG + r"$")
UUID_RE = re.compile(r"^[A-Fa-f0-9]{4,}(-[A-Fa-f0-9]+){0,4}$")
HOST_RE = re.compile(r"^[A-Za-z0-9._:-]{1,255}$")
NET_PATH_RE = re.compile(r"^[A-Za-z0-9._/@+-]{1,512}$")
OPT_RE = re.compile(r"^[A-Za-z0-9_.:/@+-]+(=[A-Za-z0-9_.:/@+-]+)?$")
LOCAL_TYPES = {"ext4", "xfs", "btrfs", "exfat", "vfat", "ntfs3", "ntfs"}
NET_TYPES = {"nfs", "nfs4", "cifs"}
# Options that turn a device mount into something else, or let set-uid/device files act on the host.
FORBIDDEN_OPTS = {"bind", "rbind", "move", "rmove", "remount", "suid", "dev", "user", "users", "owner", "group",
"x-mount.mkdir", "helper"}
DESC_RE = re.compile(r"^[^\x00-\x1f\x7f]{0,200}$")
WG_KEY_RE = re.compile(r"^[A-Za-z0-9+/]{42}[AEIMQUYcgkosw480]=$")
KEY_LINE_RE = re.compile(r"^(ssh-ed25519|ssh-rsa|ecdsa-sha2-nistp(256|384|521)|sk-ssh-ed25519@openssh\.com) "
r"[A-Za-z0-9+/]+={0,3}( [ -~]{0,200})?$")
class Refused(Exception):
def __init__(self, rule, reason):
super().__init__(reason)
self.rule, self.reason = rule, reason
class Host:
"""Every filesystem / process effect, so the tests can play the box in memory."""
def agent_uid(self):
import pwd
return pwd.getpwnam(AGENT_USER).pw_uid
def read_source(self, path):
"""The staged file: a REGULAR file owned by the agent, never a symlink, at most MAX_BYTES."""
try:
fd = os.open(path, os.O_RDONLY | os.O_NOFOLLOW | os.O_CLOEXEC)
except OSError as e:
raise Refused("P1", f"cannot open the staged file {path}: {e.strerror}")
try:
st = os.fstat(fd)
if not stat.S_ISREG(st.st_mode):
raise Refused("P1", f"{path} is not a regular file")
if st.st_uid != self.agent_uid():
raise Refused("P1", f"{path} is not owned by {AGENT_USER}")
if st.st_size > MAX_BYTES:
raise Refused("P1", f"{path} is larger than {MAX_BYTES} bytes")
with os.fdopen(fd, "rb") as f:
fd = -1
return f.read(MAX_BYTES + 1)
finally:
if fd >= 0:
os.close(fd)
def read_dest(self, path):
try:
with open(path, "rb") as f:
return f.read()
except OSError:
return None
def install(self, dest, data, mode):
"""Atomic, root-owned: a temp file beside the destination, fsync, rename."""
d = os.path.dirname(dest)
os.makedirs(d, mode=0o755, exist_ok=True)
tmp = os.path.join(d, f".{os.path.basename(dest)}.felhom-new.{os.getpid()}")
fd = os.open(tmp, os.O_WRONLY | os.O_CREAT | os.O_EXCL | os.O_NOFOLLOW, 0o600)
try:
with os.fdopen(fd, "wb") as f:
f.write(data)
f.flush()
os.fchown(f.fileno(), 0, 0)
os.fchmod(f.fileno(), mode)
os.fsync(f.fileno())
os.replace(tmp, dest)
except BaseException:
try:
os.remove(tmp)
except OSError:
pass
raise
def read_stdin(self, limit):
return sys.stdin.buffer.read(limit + 1)
def write_guest_image(self, vmid, data):
"""As root: `pct exec <vmid> -- tee <the fixed file>` with the checked ref on stdin (no shell)."""
r = subprocess.run(["/usr/sbin/pct", "exec", str(vmid), "--", "tee", CONTROLLER_IMAGE_FILE],
input=data, stdout=subprocess.DEVNULL, stderr=subprocess.PIPE, timeout=60)
if r.returncode != 0:
raise OSError(f"pct exec {vmid} tee exited {r.returncode}: {r.stderr.decode(errors='replace').strip()[:200]}")
def log(self, line):
print(line, file=sys.stderr)
try:
subprocess.run(["logger", "-t", "felhom-priv-apply", "--", line], timeout=10, check=False)
except (OSError, subprocess.SubprocessError):
pass
def systemd_escape_path(path):
"""`systemd-escape --path`: strip the slashes at both ends, `/` -> `-`, every byte outside [A-Za-z0-9:_.] (and a
leading `.`) -> `\\xNN`."""
p = path.strip("/")
out = []
for i, ch in enumerate(p):
if ch == "/":
out.append("-")
elif (ch.isascii() and (ch.isalnum() or ch in ":_.")) and not (i == 0 and ch == "."):
out.append(ch)
else:
out.extend("\\x%02x" % b for b in ch.encode())
return "".join(out)
def text_of(data, what):
if len(data) > MAX_BYTES:
raise Refused("P1", f"{what} is too large")
try:
text = data.decode("utf-8")
except UnicodeDecodeError:
raise Refused("P2", f"{what} is not UTF-8 text")
if "\x00" in text or "\r" in text:
raise Refused("P2", f"{what} carries a NUL or CR byte")
return text
def parse_ini(text, what):
"""[Section] / Key=Value / comments / blank lines. A key outside a section or a repeated key is refused."""
sections, cur = {}, None
for n, raw in enumerate(text.split("\n"), 1):
line = raw.strip()
if line.endswith("\\"):
# systemd joins a line ending in a backslash with the next one; this parser does not. Refused, so the two
# can never read the same bytes differently (review 2026-10-05).
raise Refused("U2", f"{what}: line {n} ends with a backslash (a continuation)")
if not line or line.startswith("#") or line.startswith(";"):
continue
m = re.match(r"^\[([A-Za-z]+)\]$", line)
if m:
cur = m.group(1)
if cur in sections:
raise Refused("U2", f"{what}: section [{cur}] twice")
sections[cur] = {}
continue
if cur is None or "=" not in line:
raise Refused("U2", f"{what}: line {n} is not Key=Value inside a section")
k, v = line.split("=", 1)
k, v = k.strip(), v.strip()
if k in sections[cur]:
raise Refused("U2", f"{what}: {cur}.{k} given twice")
sections[cur][k] = v
return sections
# ---------- the unit verb ----------
def check_unit(name, text):
if not UNIT_NAME_RE.match(name) or "/" in name:
raise Refused("U1", f"unit name {name!r} is not mnt-<escaped path>.mount|.automount")
kind = "automount" if name.endswith(".automount") else "mount"
s = parse_ini(text, name)
body = "Automount" if kind == "automount" else "Mount"
# [Unit] holds ONLY what the renderers write: Description, and After=local-fs-pre.target on a local mount. A
# Wants=/Requires=/Before= naming any unit would start it with the mount (Wants=reboot.target — found by review
# 2026-10-05), so none of them is accepted.
allowed = {"Unit": {"Description", "After"},
body: {"Where", "TimeoutIdleSec"} if kind == "automount" else {"What", "Where", "Type", "Options"},
"Install": {"WantedBy"}}
for sec, keys in s.items():
if sec not in allowed:
raise Refused("U2", f"{name}: section [{sec}] is not allowed")
bad = set(keys) - allowed[sec]
if bad:
raise Refused("U2", f"{name}: [{sec}] key(s) {sorted(bad)} not allowed")
u = s.get("Unit", {})
if not DESC_RE.match(u.get("Description", "")):
raise Refused("U2", f"{name}: Description has control characters")
if "After" in u and u["After"] != "local-fs-pre.target":
raise Refused("U2", f"{name}: After= may only be local-fs-pre.target")
inst = s.get("Install", {})
if inst and inst.get("WantedBy") != "multi-user.target":
raise Refused("U2", f"{name}: WantedBy must be multi-user.target")
m = s.get(body)
if not m or "Where" not in m:
raise Refused("U3", f"{name}: no [{body}] Where=")
where = m["Where"]
if not WHERE_RE.match(where):
raise Refused("U3", f"{name}: Where={where} is not /mnt/<name> or /mnt/felhom-drives/<name>")
if systemd_escape_path(where) + "." + kind != name:
raise Refused("U3", f"{name}: the unit name does not match Where={where}")
if kind == "automount":
t = m.get("TimeoutIdleSec", "")
if t and not re.match(r"^[0-9]{1,6}$", t):
raise Refused("U2", f"{name}: TimeoutIdleSec must be seconds")
return
what, typ = m.get("What", ""), m.get("Type", "")
opts = [o for o in m.get("Options", "").split(",") if o]
net = False
mu = re.match(r"^/dev/disk/by-uuid/(.+)$", what)
if mu:
if not UUID_RE.match(mu.group(1)):
raise Refused("U4", f"{name}: What= is not a filesystem UUID")
if typ and typ not in LOCAL_TYPES:
raise Refused("U4", f"{name}: Type={typ} is not a local filesystem")
else:
net = True
if typ not in NET_TYPES:
raise Refused("U4", f"{name}: What= is neither /dev/disk/by-uuid/<uuid> nor a network source with Type=nfs/nfs4/cifs")
if typ == "cifs":
mm = re.match(r"^//([^/]+)/(.+)$", what)
else:
mm = re.match(r"^([^/:][^:]*):(/.*)$", what)
if not mm or not HOST_RE.match(mm.group(1)) or not NET_PATH_RE.match(mm.group(2)) or ".." in mm.group(2).split("/"):
raise Refused("U4", f"{name}: What= is not a clean {typ} source")
if not where.startswith("/mnt/felhom-drives/"):
raise Refused("U3", f"{name}: a network share mounts only under /mnt/felhom-drives/")
for o in opts:
if not OPT_RE.match(o):
raise Refused("U5", f"{name}: mount option {o!r} has characters a mount option never needs")
if o.split("=", 1)[0].lower() in FORBIDDEN_OPTS or o.lower().startswith("x-mount."):
raise Refused("U5", f"{name}: mount option {o.split('=', 1)[0]!r} is not allowed")
if net and not {"nosuid", "nodev"} <= set(opts):
# A network server is outside the box: a set-uid file on it must never run as root here.
raise Refused("U5", f"{name}: a network share must carry nosuid,nodev")
# ---------- dnsmasq ----------
def _ip(v, v6=True):
try:
a = ipaddress.ip_address(v)
except ValueError:
return False
return v6 or a.version == 4
DOMAIN_RE = re.compile(r"^[A-Za-z0-9]([A-Za-z0-9-]{0,62})(\.[A-Za-z0-9]([A-Za-z0-9-]{0,62}))*$")
def check_dnsmasq(text):
for n, raw in enumerate(text.split("\n"), 1):
line = raw.strip()
if not line or line.startswith("#"):
continue
if line in ("bind-interfaces", "no-resolv"):
continue
k, _, v = line.partition("=")
if k == "listen-address" and _ip(v, v6=False):
continue
if k == "server" and (_ip(v) or (v.count("#") == 1 and _ip(v.split("#")[0]) and v.split("#")[1].isdigit())):
continue
m = re.match(r"^/([^/]+)/$", v)
if k == "local" and m and DOMAIN_RE.match(m.group(1)):
continue
m = re.match(r"^/([^/]+)/([^/]+)$", v)
if k == "address" and m and DOMAIN_RE.match(m.group(1)) and _ip(m.group(2), v6=False):
continue
raise Refused("D1", f"dnsmasq line {n} ({k or line[:20]!r}) is not one the resolver writes")
# ---------- WireGuard ----------
def check_wg(text):
s = parse_ini(text, "wg-felhom.conf")
if set(s) != {"Interface", "Peer"}:
raise Refused("W1", "wg-felhom.conf must hold exactly [Interface] and [Peer]")
i, p = s["Interface"], s["Peer"]
if set(i) - {"PrivateKey", "Address", "MTU"} or set(p) - {"PublicKey", "Endpoint", "AllowedIPs", "PersistentKeepalive"}:
raise Refused("W1", "wg-felhom.conf carries a key the agent never writes (PostUp/PreUp/... run as root)")
if not WG_KEY_RE.match(i.get("PrivateKey", "")) or not WG_KEY_RE.match(p.get("PublicKey", "")):
raise Refused("W2", "a WireGuard key is not 32 bytes of base64")
try:
a = ipaddress.ip_network(i.get("Address", ""), strict=False)
if a.version != 4 or a.prefixlen != 32:
raise ValueError
if not (1280 <= int(i.get("MTU", "1280")) <= 1500):
raise ValueError
host, _, port = p.get("Endpoint", "").rpartition(":")
if ipaddress.ip_address(host).version != 4 or not (1 <= int(port) <= 65535):
raise ValueError
for n in p.get("AllowedIPs", "").split(","):
if ipaddress.ip_network(n.strip(), strict=True).prefixlen != 32:
raise ValueError
if not (0 <= int(p.get("PersistentKeepalive", "25")) <= 3600):
raise ValueError
except ValueError:
raise Refused("W2", "an Address/MTU/Endpoint/AllowedIPs/PersistentKeepalive value is not what the agent renders")
# ---------- OOB sshd ----------
def render_sshd(port):
"""Byte-identical to felhomsshd.renderConfig (internal/felhomsshd/config.go) — pinned by a Go test."""
return ("# felhom OOB sshd — agent-managed (H1); DO NOT EDIT\n"
f"Port {port}\n"
"ListenAddress 0.0.0.0\n"
"ListenAddress ::\n"
"HostKey /etc/felhom-sshd/ssh_host_ed25519_key\n"
"PidFile /run/felhom-sshd.pid\n"
"AuthorizedKeysFile /etc/felhom-sshd/authorized_keys/%u\n"
"PasswordAuthentication no\n"
"PermitRootLogin prohibit-password\n"
"PubkeyAuthentication yes\n"
"KbdInteractiveAuthentication no\n"
"UsePAM yes\n"
"AllowUsers root felhom-op\n"
"X11Forwarding no\n"
"Subsystem sftp internal-sftp\n")
def check_sshd(text):
m = re.search(r"^Port ([0-9]{1,5})$", text, re.M)
if not m or not (1 <= int(m.group(1)) <= 65535) or int(m.group(1)) == 22:
raise Refused("S1", "sshd_config has no Port (or claims :22, the household's sshd)")
if text != render_sshd(int(m.group(1))):
raise Refused("S1", "sshd_config differs from the one fixed template (only the Port may vary)")
def check_key(text):
lines = [l for l in text.split("\n") if l.strip()]
if len(lines) > 1:
raise Refused("S2", "felhom-op's authorized_keys holds more than one key")
if lines and not KEY_LINE_RE.match(lines[0]):
raise Refused("S2", "the key line is not a plain public key (no options such as command= or from=)")
# ---------- main ----------
def plan(argv):
"""(verb, source, dest, mode, checker) for an argv, or Refused("A1")."""
if not argv or argv[0] not in VERBS:
raise Refused("A1", "usage: felhom-priv-apply unit <name> | dnsmasq <tmp> <name> | wg | sshd-config | sshd-key")
v, rest = argv[0], argv[1:]
if v == "unit" and len(rest) == 1:
if not UNIT_NAME_RE.match(rest[0]):
raise Refused("U1", f"unit name {rest[0]!r} is not mnt-<escaped path>.mount|.automount")
return v, os.path.join(UNITS_SRC, rest[0]), os.path.join(UNIT_DIR, rest[0]), 0o644, lambda t: check_unit(rest[0], t)
if v == "dnsmasq" and len(rest) == 2:
if not DNSMASQ_TMP_RE.match(rest[0]) or not DNSMASQ_NAME_RE.match(rest[1]):
raise Refused("D2", "dnsmasq wants /tmp/felhom-resolver-<digits>.conf and felhom-<name>.conf")
return v, rest[0], os.path.join(DNSMASQ_DIR, rest[1]), 0o644, check_dnsmasq
if v == "wg" and not rest:
return v, WG_SRC, WG_DEST, 0o600, check_wg
if v == "sshd-config" and not rest:
return v, SSHD_SRC, SSHD_DEST, 0o644, check_sshd
if v == "sshd-key" and not rest:
return v, KEY_SRC, KEY_DEST, 0o644, check_key
raise Refused("A1", f"wrong arguments for {v}")
def controller_image(rest, host):
"""R-861 (a) A1: read the ref on stdin, check it, write it INSIDE the guest as root."""
try:
if len(rest) != 1 or not VMID_RE.match(rest[0]):
raise Refused("A1", "usage: felhom-priv-apply controller-image <vmid> (the ref on stdin)")
raw = host.read_stdin(CONTROLLER_IMAGE_MAX)
if len(raw) > CONTROLLER_IMAGE_MAX:
raise Refused("I1", f"the image ref is longer than {CONTROLLER_IMAGE_MAX} bytes")
try:
text = raw.decode("ascii")
except UnicodeDecodeError:
raise Refused("I1", "the image ref is not ASCII")
ref = text[:-1] if text.endswith("\n") else text
if not CONTROLLER_IMAGE_RE.match(ref) or "\n" in ref:
raise Refused("I1", "the image ref is not gitea.dooplex.hu/admin/felhom-controller:<x.y.z>")
except Refused as e:
host.log(f"felhom-priv-apply: REFUSED [{e.rule}] controller-image {' '.join(rest)[:40]}: {e.reason}")
return 2 if e.rule == "A1" else 3
vmid = int(rest[0])
try:
host.write_guest_image(vmid, (ref + "\n").encode())
except (OSError, subprocess.SubprocessError) as e:
host.log(f"felhom-priv-apply: FAILED controller-image {vmid}: {e}")
return 4
host.log(f"felhom-priv-apply: WROTE controller-image {vmid} {ref}")
return 0
def main(argv, host=None):
host = host or Host()
if argv == ["--self-check"]:
print("felhom-priv-apply ok verbs=" + ",".join(VERBS))
return 0
if len(argv) >= 3 and argv[0] == "--check":
# CHECK ONLY (tests and the Go contract tests): `--check <verb> [<name>] <file>` validates <file> as that
# verb would and installs nothing. Not in sudoers. Prints OK or the rule; never the content.
verb, file = argv[1], argv[-1]
try:
_, _, _, _, checker = plan([verb] + argv[2:-1] if verb != "dnsmasq" else
[verb, "/tmp/felhom-resolver-1.conf", argv[2]])
with open(file, "rb") as f:
checker(text_of(f.read(), file))
except Refused as e:
print(f"REFUSED [{e.rule}] {e.reason}")
return 3
print("OK")
return 0
if argv and argv[0] == "controller-image":
return controller_image(argv[1:], host)
try:
verb, src, dest, mode, checker = plan(argv)
data = host.read_source(src)
checker(text_of(data, src))
except Refused as e:
host.log(f"felhom-priv-apply: REFUSED [{e.rule}] {' '.join(argv)[:160]}: {e.reason}")
return 2 if e.rule == "A1" else 3
if host.read_dest(dest) == data:
host.log(f"felhom-priv-apply: SAME {verb} {dest}")
return 0
try:
host.install(dest, data, mode)
except OSError as e:
host.log(f"felhom-priv-apply: FAILED {verb} {dest}: {e}")
return 4
host.log(f"felhom-priv-apply: INSTALLED {verb} {dest} ({len(data)} bytes)")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+11 -3
View File
@@ -24,6 +24,11 @@ set -u
BIN=/usr/local/bin/felhom-agent
PREV=$BIN.prev
STAGING=/var/lib/felhom-agent/selfupdate
# R-861 (agent v0.146.1): `apply` takes ONLY the root-owned copy felhom-os-apply writes after it has verified the
# operator's signature and hashed exactly those bytes (mode agent_update). The agent cannot call `apply` any more (it
# left the sudoers), and the agent's own staging dir is no longer accepted: a file in a directory the agent owns can be
# swapped between this script's sha check and its copy.
ROOT_STAGING=/var/lib/felhom-os-apply/agent-update
PENDING=$STAGING/pending.json
UNIT=felhom-agent.service
@@ -42,11 +47,14 @@ apply)
log "refusing apply: usage: apply <staged> <sha256>"
exit 2
fi
# Root-side path confinement: the staged binary MUST live in the agent's staging dir.
# Root-side path confinement: the staged binary MUST be felhom-os-apply's root-owned copy (R-861).
case "$staged" in
"$STAGING"/*) ;;
*) log "refusing apply: staged path outside $STAGING: $staged"; exit 1 ;;
"$ROOT_STAGING"/*) ;;
*) log "refusing apply: staged path outside $ROOT_STAGING: $staged"; exit 1 ;;
esac
if [ -L "$staged" ] || [ "$(stat -c %u "$staged" 2>/dev/null)" != "0" ]; then
log "refusing apply: $staged is a symlink or not root-owned"; exit 1
fi
case "$staged" in
*..*) log "refusing apply: staged path contains '..'"; exit 1 ;;
esac
+13
View File
@@ -0,0 +1,13 @@
[Unit]
Description=Felhom stable drive parent (shared bind for live drive hot-swap)
After=local-fs.target
Before=pve-guests.service
ConditionPathExists=/usr/local/sbin/felhom-shared-parent.sh
[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=/usr/local/sbin/felhom-shared-parent.sh
[Install]
WantedBy=pve-guests.service multi-user.target
+15
View File
@@ -0,0 +1,15 @@
#!/bin/sh
# felhom stable drive parent: a SHARED bind so the agent can swap backing drives underneath it and the
# guest sees the change live (no restart). MUST run before pve-guests so the guest's parent bind inherits
# the shared peer group (slave). Installed + enabled by felhom-agent. Idempotent.
set -e
mkdir -p /mnt/felhom-drives
# Isolate + share ONLY when first creating the self-bind (a fresh boot). The self-bind inherits the root
# mount's shared peer group, so make-private detaches it (else binds under it DOUBLE via the root peer),
# then make-shared gives it its own group whose only slave is the guest's parent bind. Re-running this on
# an existing parent would churn the peer-group id and orphan the guest's slave — so guard on mountpoint.
if ! mountpoint -q /mnt/felhom-drives; then
mount --bind /mnt/felhom-drives /mnt/felhom-drives
mount --make-private /mnt/felhom-drives
mount --make-shared /mnt/felhom-drives
fi
+736
View File
@@ -0,0 +1,736 @@
#!/usr/bin/env python3
"""Tests for the config bundle (R-840, `11` §5.4.2): felhom-os-apply's `bundle` mode and `--install-bundle`, and
scripts/build-config-bundle.py. An in-memory host plays the files; nothing real is written or run. Each refusal has a
test; each test names the rule it pins. Red-proof: `audits/r840-config-bundle-2026-10-04/partB/redproof.txt`.
Run: python3 configs/test_felhom_config_bundle.py (also run by internal/osupdate's Go test)
"""
import sys
sys.dont_write_bytecode = True # importing the builder must not leave scripts/__pycache__ behind
import base64
import contextlib
import hashlib
import importlib.machinery
import importlib.util
import io
import json
import os
import pathlib
import re
import stat as statmod
import unittest
HERE = pathlib.Path(__file__).resolve().parent
REPO = HERE.parent
_loader = importlib.machinery.SourceFileLoader("osapply", os.environ.get("OSAPPLY_UNDER_TEST", str(HERE / "felhom-os-apply")))
_spec = importlib.util.spec_from_loader("osapply", _loader)
osapply = importlib.util.module_from_spec(_spec)
_loader.exec_module(osapply)
_bl = importlib.machinery.SourceFileLoader("bundlebuild", str(REPO / "scripts" / "build-config-bundle.py"))
_bs = importlib.util.spec_from_loader("bundlebuild", _bl)
builder = importlib.util.module_from_spec(_bs)
_bl.exec_module(builder)
PLAN = "/var/lib/felhom-agent/os/plan-b1.json"
BUNDLE = "/var/lib/felhom-agent/os/bundle-0.143.0.json"
HOST = "demo-hp-bb76ea"
INSTALLER = REPO.parent / "felhom.eu" / "scripts" / "felhom-host-install.sh"
class St:
def __init__(self, mode, uid):
self.st_mode, self.st_uid, self.st_size = mode, uid, 100
class Box:
"""An in-memory host. files: path -> bytes; modes/uids per path; dirs: a set."""
def __init__(self, bundle_bytes, signed, oob=False, signers=True):
self.files, self.modes, self.uids = {}, {}, {}
self.dirs = {osapply.OOB_DIR} if oob else set()
self.put(osapply.TRUST_FILE, json.dumps({"host_id": HOST, "ring0_slow_lane": False}).encode(), 0o644)
if signers:
self.put(osapply.TRUST_SIGNERS, b'felhom-op-1 namespaces="felhom-op-v1" ssh-ed25519 AAAA felhom-op-1\n', 0o644)
self.put("/proc/sys/kernel/panic", b"0\n", 0o644)
self.plan = {"release_id": "bundle-0.143.0", "layer": "host", "mode": "bundle", "bundle": BUNDLE, "signed": signed}
self.put(PLAN, json.dumps(self.plan).encode(), 0o600, uid=999)
self.put(BUNDLE, bundle_bytes, 0o600, uid=999)
self.sig_rc, self.nonces, self.clock = 0, {}, 1791115200.0 # 2026-10-04T12:00:00Z
self.calls, self.logs, self.writes = [], [], []
self.visudo_fail = False # the WHOLE sudoers (`visudo -c`) after install
self.sudo_l = " (root) NOPASSWD: /usr/local/sbin/felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json\n"
self.guard = {"armed": True, "kernel_panic": 10}
def put(self, p, data, mode, uid=0):
self.files[p], self.modes[p], self.uids[p] = data, mode, uid
# Runner interface
def now(self):
return self.clock
def log(self, line):
self.logs.append(line)
def agent_uid(self):
return 999
def verify_sig(self, signers, key_id, ns, blob, sig):
self.verified = (signers, key_id, ns)
return self.sig_rc
def read_nonces(self):
return dict(self.nonces)
def write_nonces(self, d):
self.nonces = dict(d)
def read_file(self, p):
if p not in self.files:
raise OSError("no such file")
return self.files[p].decode()
def read_bytes(self, p):
if p not in self.files:
raise OSError("no such file")
return self.files[p]
def read_staged_once(self, p, owner_uid, limit):
if p not in self.files:
raise OSError("no such file")
if self.uids[p] != owner_uid:
raise osapply.Refused("R19", f"{p} is not a regular file owned by felhom-agent")
self.staged_reads = getattr(self, "staged_reads", 0) + 1
return self.files[p]
def stat(self, p):
if p not in self.files:
raise OSError("no such file")
return St(statmod.S_IFREG | self.modes[p], self.uids[p])
def lexists(self, p):
return p in self.files
def isdir(self, p):
return p in self.dirs
def put_file(self, p, data, mode):
self.writes.append(p)
self.put(p, data, mode)
def remove(self, p):
self.writes.append("rm " + p)
del self.files[p]
def list_dir(self, p):
return sorted({k[len(p) + 1:].split("/")[0] for k in self.files if k.startswith(p + "/")})
def rmtree(self, p):
for k in [k for k in self.files if k.startswith(p + "/")]:
del self.files[k]
def check_content(self, kind, data):
return (1, f"{kind}: syntax error") if b"BROKEN-SYNTAX" in data else (0, "")
def host(self, argv, timeout=600, stdin=None):
self.calls.append(argv)
if argv[:2] == ["visudo", "-c"]:
return (1, "", "parse error") if self.visudo_fail else (0, "ok", "")
if argv[:2] == ["sudo", "-n"]:
return 0, self.sudo_l, ""
if argv[-2:] == ["/usr/local/sbin/felhom-priv-apply", "--self-check"]:
body = self.files.get("/usr/local/sbin/felhom-priv-apply", b"")
return (0, "felhom-priv-apply ok verbs=unit\n", "") if b"VERBS" in body else (1, "", "boom")
if argv[-1] == "--self-check":
body = self.files.get("/usr/local/sbin/felhom-os-apply", b"")
return (0, "felhom-os-apply ok bundle-format=1 files=22\n", "") if b"BUNDLE_OP" in body else (1, "", "boom")
if argv[-1] == "/usr/local/sbin/felhom-selfupdate-guarded":
return 2, "", "felhom-selfupdate-guarded: usage: ...\n"
if argv[-1] == "status" and argv[0].endswith("felhom-crash-guard"):
return 0, json.dumps(self.guard), ""
if argv[:3] == ["systemctl", "enable", "--now"] and "felhom-crash-guard.service" in argv:
self.put("/proc/sys/kernel/panic", f"{self.guard['kernel_panic']}\n".encode(), 0o644)
return 0, "", ""
def signed_job(sha, version="0.143.0", host=HOST, op="agent_config_update", nonce="b1",
issued="2026-10-04T11:50:00Z", expires="2026-10-04T12:30:00Z"):
blob = json.dumps({"expires_at": expires, "issued_at": issued, "key_id": "felhom-op-1", "nonce": nonce, "op": op,
"params": {"agent_version": version, "bundle_sha256": sha},
"target": {"guest_id": "", "host_id": host}}, sort_keys=True).encode()
return {"blob_b64": base64.b64encode(blob).decode(), "sig": "-----BEGIN SSH SIGNATURE-----\nx\n-----END SSH SIGNATURE-----\n"}
def real_bundle(version="0.143.0"):
data = builder.build(version)
return data, hashlib.sha256(data).hexdigest()
def edited_bundle(edit):
"""The real bundle, with edit(files_list) applied and every sha recomputed — a SIGNED bundle with bad content."""
b = json.loads(builder.build("0.143.0"))
edit(b["files"])
for e in b["files"]:
e["sha256"] = hashlib.sha256(base64.b64decode(e["content_b64"])).hexdigest()
data = (json.dumps(b, indent=1, sort_keys=True) + "\n").encode()
return data, hashlib.sha256(data).hexdigest()
def replace_content(files, path, fn):
for e in files:
if e["path"] == path:
e["content_b64"] = base64.b64encode(fn(base64.b64decode(e["content_b64"]))).decode()
def run(box, argv=None, environ=None):
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
rc = osapply.main(argv or ["felhom-os-apply", "--plan", PLAN], runner=box, environ=environ or {})
line = [l for l in buf.getvalue().splitlines() if l.startswith("OSAPPLY-REPORT ")][-1]
return rc, json.loads(line[len("OSAPPLY-REPORT "):])
def box_files(box):
return {p: box.files[p] for p in box.files if p in osapply.BUNDLE_DESTS}
class Install(unittest.TestCase):
def test_fresh_box_gets_every_file_and_a_record(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
b = rep["bundle"]
# every path but the four OOB ones (no belt on this box)
self.assertEqual(len(b["written"]), len(osapply.BUNDLE_FILES) - 4, b)
self.assertEqual(len(b["skipped"]), 4)
self.assertEqual(box.files["/etc/sudoers.d/felhom-agent"], (REPO / "configs" / "felhom-agent.sudoers").read_bytes())
self.assertEqual(box.modes["/etc/sudoers.d/felhom-agent"], 0o440)
rec = json.loads(box.files[osapply.BUNDLE_RECORD])
self.assertEqual((rec["agent_version"], rec["bundle_sha256"], rec["authority"]), ("0.143.0", sha, "signed"))
self.assertIn("b1", box.nonces, "the job is consumed")
self.assertEqual(b["self_check"]["crash_guard"], {"armed": True, "kernel_panic": 10})
self.assertIn(["systemctl", "daemon-reload"], box.calls)
def test_sudoers_is_written_after_every_wrapper(self):
"""Order: a referenced wrapper is in place before the sudoers line that allows it."""
data, sha = edited_bundle(lambda f: f.reverse()) # the bundle lists the sudoers FIRST; the wrapper must reorder
box = Box(data, signed_job(sha))
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
w = [p for p in box.writes if p in osapply.BUNDLE_DESTS]
self.assertEqual(w[-1], "/etc/sudoers.d/felhom-agent")
self.assertLess(w.index("/usr/local/sbin/felhom-os-apply"), w.index("/etc/sudoers.d/felhom-agent"))
def test_identical_box_writes_nothing(self):
"""The demo boxes' case: hand-copied files equal to the release → 0 written, all 'same'."""
data, sha = real_bundle()
box = Box(data, signed_job(sha))
for dest, src, mode, _, policy in osapply.BUNDLE_FILES:
if policy != "oob":
box.put(dest, (REPO / "configs" / src).read_bytes(), mode)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["bundle"]["written"], [])
self.assertEqual(rep["bundle"]["same"], len(osapply.BUNDLE_FILES) - 5) # 4 oob skipped + crash-guard.conf kept
self.assertEqual(rep["bundle"]["kept"], ["/etc/felhom/crash-guard.conf"])
def test_a_wrong_mode_is_rewritten(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put("/etc/sudoers.d/felhom-agent", (REPO / "configs" / "felhom-agent.sudoers").read_bytes(), 0o644)
rc, rep = run(box)
self.assertIn("/etc/sudoers.d/felhom-agent", rep["bundle"]["written"])
self.assertEqual(box.modes["/etc/sudoers.d/felhom-agent"], 0o440)
def test_tuned_crash_guard_conf_is_kept(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put("/etc/felhom/crash-guard.conf", b"LIMIT=5\n", 0o644)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(box.files["/etc/felhom/crash-guard.conf"], b"LIMIT=5\n")
def test_oob_files_only_on_a_box_with_the_belt(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha), oob=True)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertIn("/etc/sudoers.d/felhom-op", rep["bundle"]["written"])
self.assertEqual(rep["bundle"]["skipped"], [])
def test_previous_copies_are_kept(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put("/usr/local/sbin/felhom-pbs-apply", b"#!/bin/bash\necho old\n", 0o755)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(box.files[rep["bundle"]["prev_dir"] + "/usr/local/sbin/felhom-pbs-apply"], b"#!/bin/bash\necho old\n")
class Refusals(unittest.TestCase):
"""Each: refused, and NOTHING on the box changed."""
def refused(self, box, code):
before = dict(box_files(box))
rc, rep = run(box)
self.assertEqual(rc, 2, rep)
self.assertEqual(rep["refused"]["code"], code, rep)
self.assertEqual(box_files(box), before, "a refusal changed a file")
self.assertNotIn(osapply.BUNDLE_RECORD, box.files)
return rep
def test_wrong_sha_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job("0" * 64))
self.refused(box, "R18")
self.assertEqual(box.nonces, {}, "a wrong sha must not burn the job")
def test_bad_signature_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.sig_rc = 255
self.refused(box, "R3")
def test_other_op_is_refused(self):
data, sha = real_bundle()
self.refused(Box(data, signed_job(sha, op="agent_update")), "R3")
def test_other_host_is_refused(self):
data, sha = real_bundle()
self.refused(Box(data, signed_job(sha, host="demo-felhom-8363b5")), "R3")
def test_expired_job_is_refused(self):
data, sha = real_bundle()
self.refused(Box(data, signed_job(sha, expires="2026-10-04T11:55:00Z")), "R3")
def test_replayed_job_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.nonces = {"b1": box.clock + 600}
self.refused(box, "R3")
def test_version_mismatch_is_refused(self):
data, sha = real_bundle()
self.refused(Box(data, signed_job(sha, version="0.142.1")), "R18")
def test_sudoers_failing_visudo_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/etc/sudoers.d/felhom-agent", lambda c: c + b"BROKEN-SYNTAX\n"))
self.refused(Box(data, signed_job(sha)), "R18")
def test_sudoers_dropping_the_route_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/etc/sudoers.d/felhom-agent",
lambda c: c.replace(b"/usr/local/sbin/felhom-os-apply --plan", b"/bin/true --plan")))
self.refused(Box(data, signed_job(sha)), "R18")
def test_wrapper_without_bundle_mode_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/usr/local/sbin/felhom-os-apply",
lambda c: c.replace(b'BUNDLE_OP = "agent_config_update"', b'BUNDLE_OPX = 1')))
self.refused(Box(data, signed_job(sha)), "R18")
def test_python_syntax_error_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/usr/local/sbin/felhom-crash-guard", lambda c: c + b"\ndef (\n"))
self.refused(Box(data, signed_job(sha)), "R18")
def test_unit_with_runtime_directory_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/etc/systemd/system/felhom-mgmt-watchdog.service",
lambda c: c + b"RuntimeDirectory=sshd\n"))
self.refused(Box(data, signed_job(sha)), "R18")
def test_agent_unit_not_as_the_agent_user_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/etc/systemd/system/felhom-agent.service",
lambda c: c.replace(b"User=felhom-agent", b"User=root")))
self.refused(Box(data, signed_job(sha)), "R18")
def test_a_bundle_that_changes_a_signer_is_refused(self):
"""R17: the trust root is not a bundle's to change — not even a signed one."""
def add(f):
f.append({"path": osapply.TRUST_SIGNERS, "content_b64": base64.b64encode(b"evil-key\n").decode()})
data, sha = edited_bundle(add)
box = Box(data, signed_job(sha))
self.refused(box, "R17")
self.assertIn(b"felhom-op-1", box.files[osapply.TRUST_SIGNERS])
def test_a_path_outside_the_table_is_refused(self):
def add(f):
f.append({"path": "/etc/shadow", "content_b64": base64.b64encode(b"root::0:0\n").decode()})
data, sha = edited_bundle(add)
self.refused(Box(data, signed_job(sha)), "R16")
def test_content_not_matching_its_sha_is_refused(self):
b = json.loads(builder.build("0.143.0"))
b["files"][0]["content_b64"] = base64.b64encode(b"#!/bin/bash\nexit 0\n").decode()
data = json.dumps(b).encode()
self.refused(Box(data, signed_job(hashlib.sha256(data).hexdigest())), "R18")
def test_bundle_outside_the_plan_dir_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.plan["bundle"] = "/tmp/bundle-0.143.0.json"
box.files[PLAN] = json.dumps(box.plan).encode()
self.refused(box, "R1")
def test_bundle_not_owned_by_the_agent_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.uids[BUNDLE] = 0
self.refused(box, "R1")
def test_no_trust_file_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
del box.files[osapply.TRUST_FILE]
self.refused(box, "R3")
class SelfCheckUndo(unittest.TestCase):
def assert_restored(self, box, before, rep):
self.assertTrue(rep["bundle"]["rolled_back"], rep)
self.assertEqual(box_files(box), before, "the previous files must be back, byte for byte")
self.assertNotIn(osapply.BUNDLE_RECORD, box.files)
def with_old_files(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put("/usr/local/sbin/felhom-pbs-apply", b"#!/bin/bash\necho old\n", 0o755)
box.put("/etc/sudoers.d/felhom-agent", b"# old sudoers\n", 0o440)
return box
def test_route_missing_after_install_puts_everything_back(self):
box = self.with_old_files()
before = dict(box_files(box))
box.sudo_l = " (root) NOPASSWD: /bin/true\n"
rc, rep = run(box)
self.assertEqual(rc, 3, rep)
self.assert_restored(box, before, rep)
self.assertEqual(box.calls[-1], ["visudo", "-c"], "the undo re-checks the whole sudoers")
def test_visudo_failing_after_install_puts_everything_back(self):
box = self.with_old_files()
before = dict(box_files(box))
box.visudo_fail = True
rc, rep = run(box)
self.assertEqual(rc, 3, rep)
self.assert_restored(box, before, rep)
def test_crash_guard_disagreeing_with_kernel_panic_puts_everything_back(self):
box = self.with_old_files()
before = dict(box_files(box))
box.guard = {"armed": True, "kernel_panic": 10}
box.host_orig = box.host
def host(argv, timeout=600, stdin=None):
if argv[:3] == ["systemctl", "enable", "--now"]:
box.calls.append(argv)
return 0, "", "" # the unit "started" but kernel.panic stayed 0
return box.host_orig(argv, timeout, stdin)
box.host = host
rc, rep = run(box)
self.assertEqual(rc, 3, rep)
self.assert_restored(box, before, rep)
class TrustBootstrap(unittest.TestCase):
def test_missing_signers_verifies_against_the_pinned_key_and_creates_it(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha), signers=False)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(box.verified[0], osapply.PINNED_SIGNERS, "verified against the pinned key, nothing else")
self.assertTrue(rep["bundle"]["signers_created"])
self.assertEqual(box.files[osapply.TRUST_SIGNERS], osapply.pinned_signers_line().encode())
self.assertEqual(box.modes[osapply.TRUST_SIGNERS], 0o644)
def test_missing_signers_and_a_job_the_pinned_key_did_not_sign(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha), signers=False)
box.sig_rc = 255
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R3"))
self.assertNotIn(osapply.TRUST_SIGNERS, box.files)
def test_present_signers_are_never_touched(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put(osapply.TRUST_SIGNERS, b'felhom-op-2 namespaces="felhom-op-v1" ssh-ed25519 BBBB felhom-op-2\n', 0o644)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(box.verified[0], osapply.TRUST_SIGNERS)
self.assertFalse(rep["bundle"]["signers_created"])
self.assertIn(b"felhom-op-2", box.files[osapply.TRUST_SIGNERS])
def test_agent_writable_signers_are_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.uids[osapply.TRUST_SIGNERS] = 999
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R3"))
def test_pinned_operator_key_equals_the_installers(self):
"""The bootstrap key is exactly the one felhom-host-install.sh pins (OPERATOR_KEY_OPERATIONAL_*)."""
text = INSTALLER.read_text()
kid = re.search(r'^OPERATOR_KEY_OPERATIONAL_ID="([^"]+)"', text, re.M).group(1)
line = re.search(r'^OPERATOR_KEY_OPERATIONAL_LINE="([^"]+)"', text, re.M).group(1)
self.assertEqual((osapply.PINNED_OPERATOR_KEY_ID, osapply.PINNED_OPERATOR_KEY_LINE), (kid, line))
# and the file format is the installer's printf, byte for byte
self.assertIn("printf '%s namespaces=\"felhom-op-v1\" %s\\n'", text)
class InstallerEntry(unittest.TestCase):
def test_installer_installs_without_a_signature(self):
data, sha = real_bundle()
box = Box(data, None)
box.put("/root/bundle.json", data, 0o600)
rc, rep = run(box, ["felhom-os-apply", "--install-bundle", "/root/bundle.json", "--sha256", sha])
self.assertEqual(rc, 0, rep)
self.assertEqual(json.loads(box.files[osapply.BUNDLE_RECORD])["authority"], "installer")
def test_installer_entry_checks_the_sha(self):
data, sha = real_bundle()
box = Box(data, None)
box.put("/root/bundle.json", data, 0o600)
rc, rep = run(box, ["felhom-os-apply", "--install-bundle", "/root/bundle.json", "--sha256", "1" * 64])
self.assertEqual((rc, rep["refused"]["code"]), (2, "R18"))
def test_installer_entry_is_refused_through_sudo(self):
data, sha = real_bundle()
box = Box(data, None)
box.put("/root/bundle.json", data, 0o600)
rc, rep = run(box, ["felhom-os-apply", "--install-bundle", "/root/bundle.json", "--sha256", sha], {"SUDO_UID": "999"})
self.assertEqual((rc, rep["refused"]["code"]), (2, "R1"))
self.assertNotIn(osapply.BUNDLE_RECORD, box.files)
def test_self_check_answers(self):
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
rc = osapply.main(["felhom-os-apply", "--self-check"], runner=Box(b"", None))
self.assertEqual(rc, 0)
self.assertIn("bundle-format=1", buf.getvalue())
class Facts(unittest.TestCase):
def test_state_reports_drift_against_the_record(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
run(box)
box.put("/usr/local/sbin/felhom-pbs-apply", b"#!/bin/bash\necho by hand\n", 0o755)
a = osapply.Apply(box, PLAN)
st = osapply.Bundle(a).state()
self.assertEqual(st["version"], "0.143.0")
self.assertEqual(st["drift"], ["/usr/local/sbin/felhom-pbs-apply"])
self.assertTrue(st["signers_present"])
def test_state_without_a_record_says_none(self):
box = Box(b"", None)
st = osapply.Bundle(osapply.Apply(box, PLAN)).state()
self.assertEqual(st["version"], "none")
self.assertNotIn("drift", st)
self.assertEqual(st["live"]["/etc/sudoers.d/felhom-agent"], "absent")
class Builder(unittest.TestCase):
def test_reproducible(self):
self.assertEqual(builder.build("0.143.0"), builder.build("0.143.0"))
def test_every_source_exists_and_every_dest_is_unique(self):
dests = [e[0] for e in osapply.BUNDLE_FILES]
self.assertEqual(len(dests), len(set(dests)))
for _, src, *_ in osapply.BUNDLE_FILES:
self.assertTrue((REPO / "configs" / src).is_file(), src)
def test_every_root_file_the_installer_writes_is_in_the_bundle(self):
"""One source of truth: a felhom root-owned path the installer names must be a bundle path, a trust file, or a
path the AGENT itself writes at run time (named here, with why). Scope is a regex over the WHOLE installer."""
text = INSTALLER.read_text()
found = set(re.findall(r"(/usr/local/sbin/felhom-[a-z-]+|/etc/systemd/system/felhom-[a-z.-]+|"
r"/etc/sudoers\.d/felhom-[a-z-]+|/etc/felhom-oob\.nft|/etc/tmpfiles\.d/felhom-[a-z.-]+|"
r"/etc/felhom/[a-z.-]+)", text))
agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"}
trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD}
# Written by the appliance ISO's first boot (felhom.eu scripts/iso/felhom-bootstrap.sh), never by the installer:
# since installer 1.32.0 (R-275) the uninstall only NAMES them under KEPT.
iso_writes = {"/etc/felhom/.bootstrap-done", "/etc/felhom/appliance-pairing-code"}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust | iso_writes)
self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them")
# the limits drop-in is named through $AGENT_UNIT in the installer
self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS)
# ---------- R-861 (agent v0.146.0): the agent binary only by an operator-signed agent_update, checked as root ----------
STAGED = "/var/lib/felhom-agent/selfupdate/felhom-agent-0.146.0"
NEW_BIN = b"\x7fELF the new agent"
def update_job(sha, version="0.146.0", op="agent_update", nonce="u1", host=HOST):
blob = json.dumps({"expires_at": "2026-10-04T12:30:00Z", "issued_at": "2026-10-04T11:50:00Z", "key_id": "felhom-op-1",
"nonce": nonce, "op": op, "params": {"sha256": sha, "version": version},
"target": {"guest_id": "", "host_id": host}}, sort_keys=True).encode()
return {"blob_b64": base64.b64encode(blob).decode(), "sig": "-----BEGIN SSH SIGNATURE-----\nx\n-----END SSH SIGNATURE-----\n"}
def update_box(job, staged=STAGED, content=NEW_BIN):
box = Box(b"{}", None)
box.put(PLAN, json.dumps({"release_id": "agent-0.146.0", "layer": "host", "mode": "agent_update",
"signed": job, "staged": staged}).encode(), 0o600, uid=999)
box.put(STAGED, content, 0o755, uid=999)
return box
def wrapper_calls(box):
return [c for c in box.calls if c and c[0] == osapply.SELFUPDATE_WRAPPER and c[1:2] == ["apply"]]
class AgentUpdate(unittest.TestCase):
"""RED-PROOF: make agent_update skip verify_signed → test_a_bad_signature_never_reaches_the_wrapper fails."""
def test_signed_update_flips_and_burns_the_nonce(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha))
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
root_copy = osapply.SELFUPDATE_ROOT_DIR + "/felhom-agent-0.146.0"
# the wrapper gets the ROOT-OWNED copy of the bytes that were hashed — never the agent's path (review 2026-10-05)
self.assertEqual(wrapper_calls(box), [[osapply.SELFUPDATE_WRAPPER, "apply", root_copy, sha]])
self.assertIn(root_copy, box.writes)
self.assertNotIn(root_copy, box.files, "the root copy is removed after the flip")
self.assertEqual(box.staged_reads, 1, "the agent's file is read exactly once")
self.assertIn("u1", box.nonces)
self.assertEqual(rep["agent_update"]["version"], "0.146.0")
def test_a_staged_file_the_agent_does_not_own_is_refused(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha))
box.uids[STAGED] = 0
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R19"), rep)
self.assertEqual(wrapper_calls(box), [])
def test_a_bad_signature_never_reaches_the_wrapper(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha))
box.sig_rc = 1
rc, rep = run(box)
self.assertTrue(rep.get("refused"), f"a job whose signature does not verify was NOT refused: {rep}")
self.assertEqual((rc, rep["refused"]["code"]), (2, "R3"), rep)
self.assertEqual(wrapper_calls(box), [])
def test_a_staged_binary_the_signature_does_not_pin_is_refused(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha), content=b"\x7fELF something the agent put there")
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R19"), rep)
self.assertEqual(wrapper_calls(box), [])
self.assertNotIn("u1", box.nonces, "a refused job keeps its nonce (the operator fixes the file, not the key)")
def test_another_staging_path_is_refused(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha), staged="/tmp/felhom-agent-0.146.0")
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R19"), rep)
self.assertEqual(wrapper_calls(box), [])
def test_a_bundle_job_is_not_an_agent_update(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha, op="agent_config_update"))
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R3"), rep)
def test_another_hosts_job_and_a_replay_are_refused(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha, host="tester-1-d70be4"))
self.assertEqual(run(box)[1]["refused"]["code"], "R3")
box2 = update_box(update_job(sha))
box2.nonces["u1"] = 1999999999
self.assertEqual(run(box2)[1]["refused"]["code"], "R3")
self.assertEqual(wrapper_calls(box) + wrapper_calls(box2), [])
def test_a_failed_flip_keeps_the_nonce(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha))
orig = box.host
box.host = lambda argv, timeout=600, stdin=None: (1, "", "refusing apply: sha mismatch") if argv[:1] == [osapply.SELFUPDATE_WRAPPER] else orig(argv, timeout, stdin)
rc, rep = run(box)
self.assertEqual(rc, 3, rep)
self.assertNotIn("u1", box.nonces)
class SelfupdateWrapperConfinement(unittest.TestCase):
"""R-861 (v0.146.1): the A/B wrapper takes only felhom-os-apply's root-owned copy, never the agent's staging dir
(a file there can be swapped between the wrapper's sha check and its copy). The path check runs before anything
is touched, so the real script can be run here unprivileged.
RED-PROOF: point ROOT_STAGING back at /var/lib/felhom-agent/selfupdate → this fails (the path is accepted and the
script goes on to `staged file missing`)."""
def test_the_agents_staging_dir_is_refused(self):
import subprocess
sha = "0" * 64
p = subprocess.run(["sh", str(HERE / "felhom-selfupdate-guarded"), "apply",
"/var/lib/felhom-agent/selfupdate/felhom-agent-0.146.1", sha], capture_output=True, text=True)
self.assertEqual(p.returncode, 1, p.stderr)
self.assertIn("outside /var/lib/felhom-os-apply/agent-update", p.stderr)
_sb = importlib.machinery.SourceFileLoader("stepbuild", str(REPO / "scripts" / "build-step-bundle.py"))
_ss = importlib.util.spec_from_loader("stepbuild", _sb)
stepbuild = importlib.util.module_from_spec(_ss)
_sb.exec_module(stepbuild)
NEW_IN_0146 = {"/usr/local/sbin/felhom-priv-apply", "/var/lib/vz/snippets/felhom-guest-hook.sh",
"/usr/local/sbin/felhom-shared-parent.sh", "/etc/systemd/system/felhom-shared-parent.service"}
class StepBundle(unittest.TestCase):
"""R-880 (agent v0.146.1): an INSTALLED wrapper checks an incoming bundle's paths against its OWN table (R16), so a
release that adds paths needs a step bundle: the boxes' current bundle with only felhom-os-apply replaced.
RED-PROOF: deliver the full bundle to the old table → R16 (test_the_full_bundle_is_refused_by_an_old_table)."""
def old_world(self):
"""The base bundle an older wrapper (no R-861 paths) installed, and that wrapper's table."""
full = json.loads(builder.build("0.145.0"))
full["files"] = [e for e in full["files"] if e["path"] not in NEW_IN_0146]
old_wrapper = b'# the v0.145.0 wrapper stands in here\nBUNDLE_OP = "agent_config_update"\n'
for e in full["files"]:
if e["path"] == "/usr/local/sbin/felhom-os-apply":
e["content_b64"], e["sha256"] = base64.b64encode(old_wrapper).decode(), hashlib.sha256(old_wrapper).hexdigest()
base = (json.dumps(full, indent=1, sort_keys=True) + "\n").encode()
old_dests = {k: v for k, v in osapply.BUNDLE_DESTS.items() if k not in NEW_IN_0146}
return base, old_dests
def parse_with_table(self, data, dests, version):
saved = osapply.BUNDLE_DESTS
osapply.BUNDLE_DESTS = dests
try:
return osapply.Bundle(osapply.Apply(Box(b"{}", None), "")).parse(data, hashlib.sha256(data).hexdigest(), version)
finally:
osapply.BUNDLE_DESTS = saved
def test_the_full_bundle_is_refused_by_an_old_table(self):
_, old_dests = self.old_world()
full = builder.build("0.146.1")
with self.assertRaises(osapply.Refused) as cm:
self.parse_with_table(full, old_dests, "0.146.1")
self.assertEqual(cm.exception.code, "R16")
def test_the_step_bundle_is_accepted_by_the_old_table_and_changes_only_the_wrapper(self):
base, old_dests = self.old_world()
new_wrapper = (HERE / "felhom-os-apply").read_bytes()
step = stepbuild.build_step(base, "0.146.1-step1", new_wrapper)
ver, files = self.parse_with_table(step, old_dests, "0.146.1-step1")
self.assertEqual(ver, "0.146.1-step1")
b, s_ = json.loads(base), json.loads(step)
self.assertEqual(sorted(e["path"] for e in b["files"]), sorted(e["path"] for e in s_["files"]), "the paths must not change")
changed = [e["path"] for e, f in zip(sorted(b["files"], key=lambda x: x["path"]), sorted(s_["files"], key=lambda x: x["path"]))
if e != f]
self.assertEqual(changed, ["/usr/local/sbin/felhom-os-apply"], "exactly the wrapper changes")
installed = dict((d, c) for d, c, *_ in files)
self.assertEqual(installed["/usr/local/sbin/felhom-os-apply"], new_wrapper)
# and the NEW wrapper (now installed) knows every path the release's full bundle names
self.assertTrue({e["path"] for e in json.loads(builder.build("0.146.1"))["files"]} <= set(osapply.BUNDLE_DESTS))
def test_a_step_version_must_carry_a_suffix(self):
base, _ = self.old_world()
with self.assertRaises(SystemExit):
stepbuild.build_step(base, "0.146.1", b"x")
if __name__ == "__main__":
unittest.main()
+143
View File
@@ -0,0 +1,143 @@
#!/usr/bin/python3
"""Tests for felhom-crash-guard (`11` §5.9). Temp dirs only; nothing real is touched. Red-proof seam: CRASHGUARD_UNDER_TEST."""
import importlib.machinery
import importlib.util
import json
import os
import pathlib
import tempfile
import unittest
HERE = pathlib.Path(__file__).resolve().parent
_loader = importlib.machinery.SourceFileLoader("crashguard", os.environ.get("CRASHGUARD_UNDER_TEST", str(HERE / "felhom-crash-guard")))
_spec = importlib.util.spec_from_loader("crashguard", _loader)
cg = importlib.util.module_from_spec(_spec)
_loader.exec_module(cg)
T0 = 1791115200.0 # 2026-10-04T12:00:00Z
class FakeEnv(cg.Env):
def __init__(self, d):
super().__init__(conf=os.path.join(d, "conf"), state_dir=os.path.join(d, "state"),
panic_path=os.path.join(d, "panic"), uptime_path=os.path.join(d, "uptime"),
boot_id_path=os.path.join(d, "bootid"))
self.t = T0
self.logs = []
open(self.panic_path, "w").write("0\n")
open(self.uptime_path, "w").write("20.00 10.00\n")
def now(self):
return self.t
def log(self, line):
self.logs.append(line)
def panic(self):
return int(open(self.panic_path).read())
def state(self):
return json.load(open(os.path.join(self.state_dir, "state.json")))
class Guard(unittest.TestCase):
def setUp(self):
self.d = tempfile.TemporaryDirectory()
self.e = FakeEnv(self.d.name)
def tearDown(self):
self.d.cleanup()
def crash_boot(self, minutes_later):
self.e.t += minutes_later * 60
cg.main(["x", "boot"], self.e) # no clean-stop before it: an unclean stop
def clean_reboot(self, minutes_later):
cg.main(["x", "clean-stop"], self.e)
self.e.t += minutes_later * 60
cg.main(["x", "boot"], self.e)
def test_first_boot_is_not_a_crash_and_arms(self):
cg.main(["x", "boot"], self.e)
s = self.e.state()
self.assertFalse(s["last_boot_unclean"])
self.assertEqual(self.e.panic(), 10)
self.assertTrue(s["armed"])
def test_clean_reboots_never_count(self):
cg.main(["x", "boot"], self.e)
for _ in range(5):
self.clean_reboot(1)
s = self.e.state()
self.assertEqual(s["unclean_boots_in_window"], 0)
self.assertEqual(self.e.panic(), 10)
def test_third_crash_in_an_hour_leaves_the_box_off(self):
# operator's words: "if it crashes 3 times within one hour, it stays off" — after crash 2 the guard trips,
# so crash 3 (kernel.panic = 0) does not restart the box.
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.assertEqual(self.e.panic(), 10, "one crash: still restarts")
self.crash_boot(5)
s = self.e.state()
self.assertTrue(s["tripped"], s)
self.assertEqual(self.e.panic(), 0, "after the 2nd crash boot the 3rd crash must leave the box off")
self.assertIn("2 unclean boots within 60 minutes", s["tripped_reason"])
def test_crashes_spread_over_more_than_the_window_do_not_trip(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(61)
self.assertFalse(self.e.state()["tripped"])
self.assertEqual(self.e.panic(), 10)
def test_tripped_stays_tripped_across_boots(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(5)
self.clean_reboot(30) # the operator switched it on; even a clean boot keeps the trip
self.assertTrue(self.e.state()["tripped"])
self.assertEqual(self.e.panic(), 0)
def test_rearms_after_24h_of_normal_running(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(5)
self.e.t += 23 * 3600
cg.main(["x", "check"], self.e)
self.assertTrue(self.e.state()["tripped"], "not before 24 h")
self.e.t += 3600
cg.main(["x", "check"], self.e)
s = self.e.state()
self.assertFalse(s["tripped"])
self.assertEqual(self.e.panic(), 10)
self.assertIn("timer", s["rearmed_by"])
def test_operator_rearm_starts_a_fresh_window(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(5)
self.e.t += 60
cg.main(["x", "rearm"], self.e)
s = self.e.state()
self.assertFalse(s["tripped"])
self.assertEqual(s["rearmed_by"], "operator")
self.assertEqual(s["unclean_boots_24h"], 2, "the history stays")
self.crash_boot(5)
self.assertFalse(self.e.state()["tripped"], "one crash after a re-arm must not trip at once")
def test_config_numbers_are_read(self):
open(self.e.conf, "w").write("LIMIT=2\nPANIC_SECONDS=30\n")
cg.main(["x", "boot"], self.e)
self.assertEqual(self.e.panic(), 30)
self.crash_boot(1)
self.assertTrue(self.e.state()["tripped"], "LIMIT=2: the first crash boot trips")
def test_state_is_world_readable_for_the_agent(self):
cg.main(["x", "boot"], self.e)
mode = os.stat(os.path.join(self.e.state_dir, "state.json")).st_mode & 0o777
self.assertEqual(mode, 0o644)
if __name__ == "__main__":
unittest.main()
File diff suppressed because it is too large Load Diff
+380
View File
@@ -0,0 +1,380 @@
#!/usr/bin/env python3
"""Tests for felhom-priv-apply (R-861, `03` §3.1). An in-memory host plays the files; nothing real is written or run.
Each refusal rule has a test that feeds it the ATTACK it exists for; each accepted shape is a file the agent really
renders (the Go contract tests feed the live renderers too). Red-proof: audits/hub-safety-2026-10-05/partF/.
Run: python3 configs/test_felhom_priv_apply.py (also run by internal/privapply's Go test)
"""
import sys
sys.dont_write_bytecode = True
import importlib.machinery
import importlib.util
import os
import pathlib
import unittest
HERE = pathlib.Path(__file__).resolve().parent
_loader = importlib.machinery.SourceFileLoader("privapply", os.environ.get("PRIVAPPLY_UNDER_TEST", str(HERE / "felhom-priv-apply")))
_spec = importlib.util.spec_from_loader("privapply", _loader)
pa = importlib.util.module_from_spec(_spec)
_loader.exec_module(pa)
AGENT_UID = 999
class FakeHost:
def __init__(self):
self.src, self.dest, self.logs, self.writes = {}, {}, [], []
self.owner, self.kind = {}, {}
def stage(self, path, text, uid=AGENT_UID, kind="file"):
self.src[path] = text.encode() if isinstance(text, str) else text
self.owner[path], self.kind[path] = uid, kind
def agent_uid(self):
return AGENT_UID
def read_source(self, path):
# mirrors Host.read_source's refusals: absent, not a regular file (a symlink is refused by O_NOFOLLOW), owner, size
if path not in self.src:
raise pa.Refused("P1", f"cannot open the staged file {path}")
if self.kind[path] != "file":
raise pa.Refused("P1", f"{path} is not a regular file")
if self.owner[path] != AGENT_UID:
raise pa.Refused("P1", f"{path} is not owned by felhom-agent")
if len(self.src[path]) > pa.MAX_BYTES:
raise pa.Refused("P1", f"{path} is larger than {pa.MAX_BYTES} bytes")
return self.src[path]
def read_dest(self, path):
return self.dest.get(path)
def install(self, dest, data, mode):
self.writes.append((dest, mode))
self.dest[dest] = data
def log(self, line):
self.logs.append(line)
# controller-image (R-861 (a) A1): the ref arrives on stdin and is written INSIDE the guest by root.
stdin = b""
guest_writes = None
def read_stdin(self, limit):
return self.stdin[:limit + 1]
def write_guest_image(self, vmid, data):
if self.guest_writes is None:
self.guest_writes = []
self.guest_writes.append((vmid, data))
LOCAL_UNIT = """# Managed by felhom-agent — do not edit by hand.
[Unit]
Description=Felhom storage mount 91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae
After=local-fs-pre.target
[Mount]
What=/dev/disk/by-uuid/91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae
Where=/mnt/hdd_1
Type=ext4
[Install]
WantedBy=multi-user.target
"""
NET_MOUNT = """# felhom network storage — do not edit by hand.
[Unit]
Description=Felhom network storage media (nfs)
[Mount]
What=nas.lan:/volume1/media
Where=/mnt/felhom-drives/media
Type=nfs4
Options=vers=4.1,soft,timeo=50,retrans=2,noatime,_netdev,retry=0,nosuid,nodev
"""
NET_AUTOMOUNT = """# felhom network storage — do not edit by hand.
[Unit]
Description=Felhom network storage automount media (nfs)
[Automount]
Where=/mnt/felhom-drives/media
TimeoutIdleSec=600
[Install]
WantedBy=multi-user.target
"""
NET_NAME = "mnt-felhom\\x2ddrives-media.mount"
DNS_BASE = """# felhom split-horizon resolver — host base config (agent-managed; DO NOT EDIT)
bind-interfaces
listen-address=192.168.0.104
listen-address=127.0.0.1
no-resolv
server=1.1.1.1
server=9.9.9.9
"""
DNS_GUEST = """# felhom split-horizon DNS — customer demo-hp (agent-managed; DO NOT EDIT)
local=/enkisfelhom.hu/
address=/enkisfelhom.hu/192.168.0.138
"""
K = "AAECAwQFBgcICQoLDA0ODxAREhMUFRYXGBkaGxwdHh8=" # base64 of 32 bytes
WG = f"""# felhom offsite tunnel — agent-managed (S3); DO NOT EDIT
[Interface]
PrivateKey = {K}
Address = 10.77.0.3/32
MTU = 1280
[Peer]
PublicKey = {K}
Endpoint = 49.12.1.2:51820
AllowedIPs = 10.77.0.1/32, 10.77.0.250/32
PersistentKeepalive = 25
"""
KEYLINE = "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIL8z0qCNgA3x2xxAB0Qj5ro8waFjGZ8Ta/sWB63tlLw+ felhom-op-1\n"
def run(host, *argv):
return pa.main(list(argv), host=host)
class Accepts(unittest.TestCase):
"""What the agent really writes is accepted and installed, root-owned, at the fixed destination."""
def test_local_unit_installed(self):
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/mnt-hdd_1.mount", LOCAL_UNIT)
self.assertEqual(run(h, "unit", "mnt-hdd_1.mount"), 0)
self.assertEqual(h.writes, [("/etc/systemd/system/mnt-hdd_1.mount", 0o644)])
def test_local_unit_without_type(self): # the N100's unit has no Type= (autodetect)
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/mnt-hdd_1.mount", LOCAL_UNIT.replace("Type=ext4\n", ""))
self.assertEqual(run(h, "unit", "mnt-hdd_1.mount"), 0)
def test_network_pair(self):
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/" + NET_NAME, NET_MOUNT)
h.stage("/var/lib/felhom-agent/units/mnt-felhom\\x2ddrives-media.automount", NET_AUTOMOUNT)
self.assertEqual(run(h, "unit", NET_NAME), 0)
self.assertEqual(run(h, "unit", "mnt-felhom\\x2ddrives-media.automount"), 0)
def test_identical_is_not_rewritten(self):
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/mnt-hdd_1.mount", LOCAL_UNIT)
h.dest["/etc/systemd/system/mnt-hdd_1.mount"] = LOCAL_UNIT.encode()
self.assertEqual(run(h, "unit", "mnt-hdd_1.mount"), 0)
self.assertEqual(h.writes, [])
def test_dnsmasq_both_dropins(self):
h = FakeHost()
h.stage("/tmp/felhom-resolver-123.conf", DNS_BASE)
self.assertEqual(run(h, "dnsmasq", "/tmp/felhom-resolver-123.conf", "felhom-resolver-base.conf"), 0)
h.stage("/tmp/felhom-resolver-124.conf", DNS_GUEST)
self.assertEqual(run(h, "dnsmasq", "/tmp/felhom-resolver-124.conf", "felhom-demo-hp.conf"), 0)
self.assertIn(("/etc/dnsmasq.d/felhom-demo-hp.conf", 0o644), h.writes)
def test_wg_installed_0600(self):
h = FakeHost()
h.stage(pa.WG_SRC, WG)
self.assertEqual(run(h, "wg"), 0)
self.assertEqual(h.writes, [(pa.WG_DEST, 0o600)])
def test_sshd_template_and_key(self):
h = FakeHost()
h.stage(pa.SSHD_SRC, pa.render_sshd(8822))
h.stage(pa.KEY_SRC, KEYLINE)
self.assertEqual(run(h, "sshd-config"), 0)
self.assertEqual(run(h, "sshd-key"), 0)
h.stage(pa.KEY_SRC, "") # clearing the operator login is allowed
self.assertEqual(run(h, "sshd-key"), 0)
def test_escape_matches_systemd(self):
self.assertEqual(pa.systemd_escape_path("/mnt/felhom-drives/media"), "mnt-felhom\\x2ddrives-media")
self.assertEqual(pa.systemd_escape_path("/mnt/hdd_1"), "mnt-hdd_1")
self.assertEqual(pa.systemd_escape_path("/mnt/.x"), "mnt-.x") # a dot is escaped only at the very start
class Refuses(unittest.TestCase):
"""Each rule, with the attack it exists for. Nothing is written on a refusal."""
def refused(self, h, argv, rule):
rc = run(h, *argv)
self.assertIn(rc, (2, 3), f"{argv} was accepted")
self.assertEqual(h.writes, [], f"{argv} wrote something")
self.assertTrue(any(f"[{rule}]" in l for l in h.logs), f"{argv}: rule {rule} not logged: {h.logs}")
def unit(self, text, name="mnt-hdd_1.mount"):
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/" + name, text)
return h, ["unit", name]
def test_U1_name_outside_mnt(self):
h = FakeHost()
self.refused(h, ["unit", "etc-sudoers.d.mount"], "U1")
self.refused(h, ["unit", "../../etc/x.mount"], "U1")
self.refused(h, ["unit", "mnt-x.service"], "U1")
def test_U2_service_section(self):
self.refused(*self.unit(LOCAL_UNIT + "\n[Service]\nExecStart=/bin/sh -c id\n"), "U2")
def test_U2_wants_starts_another_unit(self): # review 2026-10-05: Wants=reboot.target would reboot the host
for extra in ("Wants=reboot.target", "Requires=felhom-agent-rollback.service", "Before=pve-guests.service"):
self.refused(*self.unit(LOCAL_UNIT.replace("After=local-fs-pre.target", "After=local-fs-pre.target\n" + extra)), "U2")
self.refused(*self.unit(LOCAL_UNIT.replace("After=local-fs-pre.target", "After=poweroff.target")), "U2")
def test_U2_continuation_line(self):
t = LOCAL_UNIT.replace("Description=Felhom storage mount 91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae",
"Description=Felhom storage mount \\")
self.refused(*self.unit(t), "U2")
self.refused(*self.unit(LOCAL_UNIT.replace("# Managed by felhom-agent", "# comment \\\n# Managed by felhom-agent")), "U2")
def test_U2_unknown_key(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Type=ext4", "Type=ext4\nDirectoryMode=0777")), "U2")
def test_U3_bind_over_sudoers_dir(self):
# the R-861 attack: mount an agent-owned directory over /etc/sudoers.d
t = LOCAL_UNIT.replace("Where=/mnt/hdd_1", "Where=/etc/sudoers.d")
self.refused(*self.unit(t, "mnt-hdd_1.mount"), "U3")
def test_U3_name_must_match_where(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Where=/mnt/hdd_1", "Where=/mnt/other")), "U3")
def test_U3_traversal_in_where(self):
# the name passes U1 and equals the escaped Where — ONLY the Where rule stops a mount at /mnt/../etc = /etc
self.assertEqual(pa.systemd_escape_path("/mnt/../etc") + ".mount", "mnt-..-etc.mount")
self.refused(*self.unit(LOCAL_UNIT.replace("Where=/mnt/hdd_1", "Where=/mnt/../etc"), "mnt-..-etc.mount"), "U3")
def test_U4_what_is_an_agent_directory(self):
t = LOCAL_UNIT.replace("What=/dev/disk/by-uuid/91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae", "What=/var/lib/felhom-agent/evil")
self.refused(*self.unit(t), "U4")
def test_U4_tmpfs(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Type=ext4", "Type=tmpfs")), "U4")
def test_U5_bind_option(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Type=ext4", "Type=ext4\nOptions=bind")), "U5")
def test_U5_network_without_nosuid(self):
t = NET_MOUNT.replace(",nosuid,nodev", "")
self.refused(*self.unit(t, NET_NAME), "U5")
def test_U5_suid_option(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Type=ext4", "Type=ext4\nOptions=suid,dev")), "U5")
def test_U3_network_outside_drives(self):
t = NET_MOUNT.replace("/mnt/felhom-drives/media", "/mnt/media")
self.refused(*self.unit(t, "mnt-media.mount"), "U3")
def test_D1_dhcp_script(self): # runs as root
h = FakeHost()
h.stage("/tmp/felhom-resolver-1.conf", DNS_BASE + "dhcp-script=/var/lib/felhom-agent/x.sh\n")
self.refused(h, ["dnsmasq", "/tmp/felhom-resolver-1.conf", "felhom-x.conf"], "D1")
def test_D1_conf_dir_and_log_file(self):
for extra in ("conf-dir=/var/lib/felhom-agent\n", "log-facility=/etc/sudoers.d/x\n", "user=root\n"):
h = FakeHost()
h.stage("/tmp/felhom-resolver-1.conf", DNS_GUEST + extra)
self.refused(h, ["dnsmasq", "/tmp/felhom-resolver-1.conf", "felhom-x.conf"], "D1")
def test_D2_paths(self):
h = FakeHost()
self.refused(h, ["dnsmasq", "/etc/shadow", "felhom-x.conf"], "D2")
self.refused(h, ["dnsmasq", "/tmp/felhom-resolver-1.conf", "../sudoers.d/x.conf"], "D2")
def test_W1_postup(self): # wg-quick runs PostUp as root
h = FakeHost()
h.stage(pa.WG_SRC, WG.replace("MTU = 1280", "MTU = 1280\nPostUp = /bin/sh -c id"))
self.refused(h, ["wg"], "W1")
def test_W2_values(self):
h = FakeHost()
h.stage(pa.WG_SRC, WG.replace("AllowedIPs = 10.77.0.1/32, 10.77.0.250/32", "AllowedIPs = 0.0.0.0/0"))
self.refused(h, ["wg"], "W2")
def test_S1_sshd_strictmodes(self): # an AuthorizedKeysFile the agent owns + StrictModes no = root login
h = FakeHost()
h.stage(pa.SSHD_SRC, pa.render_sshd(8822) + "StrictModes no\n")
self.refused(h, ["sshd-config"], "S1")
h2 = FakeHost()
h2.stage(pa.SSHD_SRC, pa.render_sshd(8822).replace("/etc/felhom-sshd/authorized_keys/%u", "/var/lib/felhom-agent/k"))
self.refused(h2, ["sshd-config"], "S1")
def test_S1_port_22(self):
h = FakeHost()
h.stage(pa.SSHD_SRC, pa.render_sshd(22))
self.refused(h, ["sshd-config"], "S1")
def test_S2_key_options_and_two_keys(self):
h = FakeHost()
h.stage(pa.KEY_SRC, 'command="/bin/sh" ' + KEYLINE)
self.refused(h, ["sshd-key"], "S2")
h2 = FakeHost()
h2.stage(pa.KEY_SRC, KEYLINE + KEYLINE)
self.refused(h2, ["sshd-key"], "S2")
def test_P1_symlink_owner_size(self):
h = FakeHost()
h.stage(pa.WG_SRC, WG, kind="symlink")
self.refused(h, ["wg"], "P1")
h2 = FakeHost()
h2.stage(pa.WG_SRC, WG, uid=0)
self.refused(h2, ["wg"], "P1")
h3 = FakeHost()
h3.stage(pa.WG_SRC, "#" * (pa.MAX_BYTES + 1))
self.refused(h3, ["wg"], "P1")
def test_A1_usage(self):
h = FakeHost()
self.refused(h, ["install", "/etc/shadow"], "A1")
self.refused(h, ["wg", "/etc/shadow"], "A1")
class ControllerImage(unittest.TestCase):
"""R-861 (a) A1 (`09` §3 decision 165): the agent can no longer `tee` any image ref into the guest. The root verb
reads the ref on stdin, requires our registry + our repository + an x.y.z tag, and writes the guest file itself.
RED-PROOF: on the pre-A1 wrapper `controller-image` is not a verb (A1 usage, rc 2) — the accepted case fails."""
def go(self, ref, *argv):
h = FakeHost()
h.stdin = ref.encode() if isinstance(ref, str) else ref
return h, run(h, *(argv or ("controller-image", "9201")))
def test_our_controller_ref_is_written_in_the_guest(self):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0\n")
self.assertEqual(rc, 0, h.logs)
self.assertEqual(h.guest_writes, [(9201, b"gitea.dooplex.hu/admin/felhom-controller:0.301.0\n")])
def test_a_foreign_image_is_refused(self):
for ref in ("docker.io/library/alpine:latest\n", "alpine\n",
"gitea.dooplex.hu/admin/felhom-controller:latest\n",
"gitea.dooplex.hu/admin/other:0.1.0\n",
"evil.example/admin/felhom-controller:0.301.0\n",
"gitea.dooplex.hu/admin/felhom-controller:0.301.0\nalpine\n",
"gitea.dooplex.hu/admin/felhom-controller:0.301.0 x\n",
"", "\n"):
h, rc = self.go(ref)
self.assertEqual(rc, 3, f"{ref!r} was accepted")
self.assertFalse(h.guest_writes, f"{ref!r} wrote the guest file")
self.assertTrue(any("[I1]" in l for l in h.logs), h.logs)
def test_oversize_stdin_is_refused(self):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0" + " " * 300)
self.assertEqual(rc, 3)
self.assertFalse(h.guest_writes)
def test_vmid_must_be_numeric(self):
for argv in (("controller-image", "9201;id"), ("controller-image", "-1"), ("controller-image",),
("controller-image", "9201", "9202")):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0\n", *argv)
self.assertIn(rc, (2, 3), argv)
self.assertFalse(h.guest_writes, argv)
def test_self_check_names_the_verb(self):
import io, contextlib
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
pa.main(["--self-check"])
self.assertIn("controller-image", buf.getvalue())
if __name__ == "__main__":
unittest.main(verbosity=2)
@@ -241,3 +241,23 @@ func TestNewestArchiveTime_DistinctPhantomsEachAnnounced(t *testing.T) {
t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String())
}
}
// R-99 (`09` §3 decision 140): the WARN for a PBS phantom ends with the cleanup runbook, so whoever sees it knows the
// one sanctioned way to remove it; a tiny archive on a dir storage is not a PBS phantom and gets no pointer.
// RED-PROOF: drop the `msg += phantomCleanupPointer` line → "the PBS phantom WARN does not end with the runbook pointer".
func TestRejectedArchiveWarnNamesTheCleanupRunbook(t *testing.T) {
var buf bytes.Buffer
r := runnerWithContent(t, &buf, []proxmox.StorageContent{phantomEntry(), goodPBSEntry()})
if _, _, err := r.NewestArchiveTime(context.Background(), 9201); err != nil {
t.Fatal(err)
}
const want = "INCOMPLETE archive when computing tier freshness — it is not a successful backup — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
if !strings.Contains(buf.String(), want) {
t.Errorf("the PBS phantom WARN does not end with the runbook pointer:\n%s", buf.String())
}
local := phantomEntry()
local.Format, local.VolID = "tar.zst", "local:backup/vzdump-lxc-9201-2026_07_28-05_31_14.tar.zst"
if got := rejectedArchiveMessage(local); strings.Contains(got, "pbs-phantom-cleanup") {
t.Errorf("a dir-storage archive got the PBS runbook pointer: %s", got)
}
}
+131
View File
@@ -0,0 +1,131 @@
package backup
import (
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// BackupSuccessState persists the newest SUCCESSFUL whole-guest backup per tier and guest (R-894).
//
// Why it exists. The due-check (`localapi` handleBackupDue) asks the tier's storage when a backup last
// landed (R-84) and falls back to the in-memory record when the storage cannot be read. The in-memory
// record is empty after an agent restart (Store, R-348), so "storage unreadable" right after a restart
// read as "no record — DUE". Measured 2026-10-05 on demo-hp: the agent restarted at 04:57, the off-site
// storage answered "Can't connect" at 06:25, the 7-day tier — last copy 2026-10-01 — read DUE, the
// controller asked, and vzdump failed. This file is the last known copy the fallback reads instead.
//
// It is read ONLY when the storage cannot be read. A storage that answers is the ground truth and wins,
// in both directions: an archive found there counts, and an archive absent there is absent even when
// this file remembers a success (a pruned or deleted archive must make the tier due — the same reason
// R-84 chose the storage over a persisted record). Pinned by
// TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers.
//
// Only SUCCESSES are written (the RestoreTestState rule): a failure must stay due and be retried, so a
// record of a failure has no reader.
type BackupSuccessState struct {
path string
mu sync.Mutex
last map[string]savedSuccess // key(target, vmid) → the newest success
}
type savedSuccess struct {
target string
vmid int
at time.Time
}
// backupSuccessJSON is one entry on disk.
type backupSuccessJSON struct {
Target string `json:"target"`
VMID int `json:"vmid"`
StartedAt string `json:"started_at"`
}
func backupStateKey(target string, vmid int) string { return target + "/" + strconv.Itoa(vmid) }
// NewBackupSuccessState opens (or creates) the state at path. A missing or unreadable file degrades to
// "nothing known" — the pre-R-894 behaviour, which is DUE — and never wedges the daemon.
func NewBackupSuccessState(path string) *BackupSuccessState {
s := &BackupSuccessState{path: path, last: map[string]savedSuccess{}}
data, err := os.ReadFile(path)
if err != nil {
return s
}
var entries []backupSuccessJSON
if json.Unmarshal(data, &entries) != nil {
return s
}
for _, e := range entries {
t, perr := time.Parse(time.RFC3339, e.StartedAt)
if perr != nil {
continue // one unreadable entry must not lose the others
}
s.last[backupStateKey(e.Target, e.VMID)] = savedSuccess{target: e.Target, vmid: e.VMID, at: t.UTC()}
}
return s
}
// RecordBackupSuccess saves b when it is a success newer than the one on file. target is the tier the
// job ran on (the due-check's key); a failure or an unparseable time is ignored.
func (s *BackupSuccessState) RecordBackupSuccess(target string, b hub.Backup) error {
if s == nil || !b.Success {
return nil
}
t, err := time.Parse(time.RFC3339, b.StartedAt)
if err != nil {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
k := backupStateKey(target, b.VMID)
if old, ok := s.last[k]; ok && !t.After(old.at) {
return nil
}
s.last[k] = savedSuccess{target: target, vmid: b.VMID, at: t.UTC()}
return s.saveLocked()
}
// LastKnownSuccess returns the newest saved success for this tier and guest (ok=false = none on file).
func (s *BackupSuccessState) LastKnownSuccess(target string, vmid int) (time.Time, bool) {
if s == nil {
return time.Time{}, false
}
s.mu.Lock()
defer s.mu.Unlock()
e, ok := s.last[backupStateKey(target, vmid)]
return e.at, ok
}
func (s *BackupSuccessState) saveLocked() error {
entries := make([]backupSuccessJSON, 0, len(s.last))
for _, e := range s.last {
entries = append(entries, backupSuccessJSON{Target: e.target, VMID: e.vmid, StartedAt: e.at.Format(time.RFC3339)})
}
// Deterministic file content (Go's map order is random).
sort.Slice(entries, func(i, j int) bool {
if entries[i].Target != entries[j].Target {
return entries[i].Target < entries[j].Target
}
return entries[i].VMID < entries[j].VMID
})
data, err := json.MarshalIndent(entries, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(s.path), 0o755); err != nil {
return err
}
tmp := s.path + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, s.path)
}
+59
View File
@@ -0,0 +1,59 @@
package backup
import (
"os"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-894: the on-disk newest success per tier survives a restart (a new state from the same file).
func TestBackupSuccessState_SurvivesRestart(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
s := NewBackupSuccessState(path)
at := time.Date(2026, 10, 1, 20, 15, 0, 0, time.UTC)
if err := s.RecordBackupSuccess("felhom-pbs", hub.Backup{VMID: 9201, Success: true, StartedAt: at.Format(time.RFC3339)}); err != nil {
t.Fatal(err)
}
got, ok := NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 9201)
if !ok || !got.Equal(at) {
t.Fatalf("after a restart the saved copy must read back; got %v ok=%v", got, ok)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("local", 9201); ok {
t.Fatal("another tier must not borrow this tier's copy")
}
}
// Only a NEWER success replaces the saved one; failures and unparseable times are ignored.
func TestBackupSuccessState_KeepsNewestSuccessOnly(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
s := NewBackupSuccessState(path)
newer := time.Date(2026, 10, 5, 0, 0, 0, 0, time.UTC)
older := newer.Add(-48 * time.Hour)
for _, b := range []hub.Backup{
{VMID: 1, Success: true, StartedAt: newer.Format(time.RFC3339)},
{VMID: 1, Success: true, StartedAt: older.Format(time.RFC3339)}, // older: ignored
{VMID: 1, Success: false, StartedAt: newer.Add(time.Hour).Format(time.RFC3339)}, // failure: ignored
{VMID: 1, Success: true, StartedAt: "not-a-time"}, // unparseable: ignored
} {
if err := s.RecordBackupSuccess("t", b); err != nil {
t.Fatal(err)
}
}
if got, _ := NewBackupSuccessState(path).LastKnownSuccess("t", 1); !got.Equal(newer) {
t.Fatalf("want the newest success %v, got %v", newer, got)
}
}
// A corrupt file degrades to "nothing known" (the pre-R-894 DUE answer), never a crash.
func TestBackupSuccessState_CorruptFileIsEmpty(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
if err := os.WriteFile(path, []byte("{not json"), 0o600); err != nil {
t.Fatal(err)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("t", 1); ok {
t.Fatal("a corrupt file must read as nothing known")
}
}
+35 -12
View File
@@ -2,6 +2,7 @@ package backup
import (
"context"
"fmt"
"encoding/json"
"errors"
"io"
@@ -21,14 +22,18 @@ type fakeBackupAPI struct {
vzdumpErr error
waitErr error
cfg proxmox.GuestConfig
goneGuests map[int]bool // R-689: vmids whose config lookup answers "does not exist"
aclGuests map[int]bool // R-689: vmids outside the token's ACL — PVE answers 403 "permission denied"
cfgErr error
content []proxmox.StorageContent
contentErr error
storages []proxmox.Storage // returned by ListStorage (the local-prune scope gate)
storageErr error
vzdumps []proxmox.VzdumpOptions
logLines []string // returned by TaskLogTail (e.g. "INFO: backup mode: stop")
waitGate chan struct{} // if non-nil, WaitTask blocks until closed (8B.2 watcher timing)
storages []proxmox.Storage // returned by ListStorage (the local-prune scope gate) — DEFINITIONS: like
// production's GET /storage, it never carries usage; ListStorage strips Avail/Used (R-685's live lesson)
nodeStorages []proxmox.Storage // returned by NodeStorage (GET /nodes/{node}/storage — WITH usage)
storageErr error
vzdumps []proxmox.VzdumpOptions
logLines []string // returned by TaskLogTail (e.g. "INFO: backup mode: stop")
waitGate chan struct{} // if non-nil, WaitTask blocks until closed (8B.2 watcher timing)
}
func (f *fakeBackupAPI) Vzdump(_ context.Context, o proxmox.VzdumpOptions) (string, error) {
@@ -41,13 +46,30 @@ func (f *fakeBackupAPI) WaitTask(_ context.Context, _ string, _ proxmox.WaitOpti
}
return proxmox.TaskStatus{Status: "stopped", ExitStatus: "OK"}, f.waitErr
}
func (f *fakeBackupAPI) GuestConfig(_ context.Context, _ int) (proxmox.GuestConfig, error) {
func (f *fakeBackupAPI) GuestConfig(_ context.Context, vmid int) (proxmox.GuestConfig, error) {
if f.aclGuests[vmid] {
return proxmox.GuestConfig{}, fmt.Errorf("proxmox: GET /nodes/n/lxc/%d/config -> HTTP 403: permission denied at /vms/%d (missing privilege VM.Audit)", vmid, vmid)
}
if f.goneGuests[vmid] { // R-689: PVE's answer for a deleted guest
return proxmox.GuestConfig{}, fmt.Errorf("proxmox: GET /nodes/n/lxc/%d/config -> HTTP 500: Configuration file 'nodes/n/lxc/%d.conf' does not exist", vmid, vmid)
}
return f.cfg, f.cfgErr
}
func (f *fakeBackupAPI) StorageContent(_ context.Context, _ string) ([]proxmox.StorageContent, error) {
return f.content, f.contentErr
}
func (f *fakeBackupAPI) ListStorage(_ context.Context) ([]proxmox.Storage, error) {
out := make([]proxmox.Storage, len(f.storages))
for i, s := range f.storages {
s.Avail, s.Used, s.Total = 0, 0, 0 // GET /storage has no usage — a fake that had it hid R-685's defect
out[i] = s
}
return out, f.storageErr
}
func (f *fakeBackupAPI) NodeStorage(_ context.Context) ([]proxmox.Storage, error) {
if f.nodeStorages != nil {
return f.nodeStorages, f.storageErr
}
return f.storages, f.storageErr
}
func (f *fakeBackupAPI) TaskLogTail(_ context.Context, _ string, _ int) ([]string, error) {
@@ -137,13 +159,14 @@ func TestBackup_VzdumpFailureReturnsFailedRecord(t *testing.T) {
func TestPickRestoreCandidate_NewestOrEmpty(t *testing.T) {
const big = 4 << 30 // a plausible whole-guest archive
api := &fakeBackupAPI{content: []proxmox.StorageContent{
{VolID: "a", Content: "backup", CTime: 10, Size: big},
{VolID: "b", Content: "backup", CTime: 99, Size: big},
// R-689: real vzdump names with their vmid — only a backup OF A GUEST is a candidate.
{VolID: "local:backup/vzdump-lxc-9001-a.tar.zst", VMID: 9001, Content: "backup", CTime: 10, Size: big},
{VolID: "local:backup/vzdump-lxc-9001-b.tar.zst", VMID: 9001, Content: "backup", CTime: 99, Size: big},
{VolID: "iso", Content: "iso", CTime: 999, Size: big}, // not a backup → ignored
}}
r := NewBackupRunner(api, "local", "", "", "", quiet())
vol, err := r.PickRestoreCandidate(context.Background())
if err != nil || vol != "b" {
if err != nil || vol != "local:backup/vzdump-lxc-9001-b.tar.zst" {
t.Fatalf("pick = %q,%v want newest 'b'", vol, err)
}
// no backups → "".
@@ -163,12 +186,12 @@ func TestPickRestoreCandidate_NewestOrEmpty(t *testing.T) {
// `pick = "phantom" want the newest COMPLETE archive 'real'`.
func TestPickRestoreCandidate_SkipsImplausibleArchives(t *testing.T) {
api := &fakeBackupAPI{content: []proxmox.StorageContent{
{VolID: "real", Content: "backup", CTime: 10, Size: 4 << 30},
{VolID: "phantom", Content: "backup", CTime: 99, Size: 1}, // newest, and impossible
{VolID: "felhom-pbs:backup/ct/9001/real", VMID: 9001, Content: "backup", CTime: 10, Size: 4 << 30},
{VolID: "felhom-pbs:backup/ct/9001/phantom", VMID: 9001, Content: "backup", CTime: 99, Size: 1}, // newest, and impossible
}}
r := NewBackupRunner(api, "local", "", "", "", quiet())
vol, err := r.PickRestoreCandidate(context.Background())
if err != nil || vol != "real" {
if err != nil || vol != "felhom-pbs:backup/ct/9001/real" {
t.Fatalf("pick = %q,%v want the newest COMPLETE archive 'real'", vol, err)
}
}
+60
View File
@@ -0,0 +1,60 @@
package backup
import (
"context"
"sort"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// ForeignKeyLedger (R-366 slice 2, `09` §3 decision 168) holds, per backup tier, the whole-guest archives the
// restore-test pick skipped because another key wrote them (R-727 — an earlier install of this box). Since R-727 the
// skip was one INFO log line per archive and nothing else, so after a reinstall the operator was never told that the
// box's older whole-guest copies are unreadable to it. The host report carries this ledger; the hub raises ONE
// operator event when it changes.
//
// It reports nil until a tier has been evaluated since the agent started, so a restart does not read as "the set
// changed to empty" (the hub keeps its last state for an absent field).
type ForeignKeyLedger struct {
mu sync.Mutex
byTarget map[string]hub.ForeignKeyArchives
}
// NewForeignKeyLedger builds an empty ledger.
func NewForeignKeyLedger() *ForeignKeyLedger {
return &ForeignKeyLedger{}
}
func (l *ForeignKeyLedger) set(target string, n int, oldest, newest int64) {
l.mu.Lock()
defer l.mu.Unlock()
if l.byTarget == nil {
l.byTarget = map[string]hub.ForeignKeyArchives{}
}
e := hub.ForeignKeyArchives{Target: target, Count: n}
if n > 0 {
e.Oldest = time.Unix(oldest, 0).UTC().Format(time.RFC3339)
e.Newest = time.Unix(newest, 0).UTC().Format(time.RFC3339)
}
l.byTarget[target] = e
}
// ForeignKeyArchives implements hub.ForeignKeyArchiveReporter: nil before any evaluation; otherwise the tiers that
// hold such archives (`Tiers` empty, never nil, when none do), sorted by tier.
func (l *ForeignKeyLedger) ForeignKeyArchives(context.Context) *hub.ForeignKeyArchivesStanza {
l.mu.Lock()
defer l.mu.Unlock()
if l.byTarget == nil {
return nil
}
out := []hub.ForeignKeyArchives{}
for _, e := range l.byTarget {
if e.Count > 0 {
out = append(out, e)
}
}
sort.Slice(out, func(i, j int) bool { return out[i].Target < out[j].Target })
return &hub.ForeignKeyArchivesStanza{Tiers: out}
}
@@ -0,0 +1,69 @@
package backup
import (
"context"
"encoding/json"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-366 slice 2 (`09` §3 decision 168) — the restore-test's skip of an archive written with another key stops being
// silent: the pick records, per tier, how many it skipped and their time range, and the host report carries it.
//
// COMPANION RED-PROOF (observed): remove the `r.foreign.set(...)` call from PickSettledRestoreCandidateOn → this fails
// with "after one evaluation the ledger must report felhom-pbs: 2 archives …; got []". Restored.
func TestR366_PickRecordsArchivesWrittenWithAnotherKey(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T17:27:32Z", Content: "backup", VMID: 9201, Size: 4774114206, CTime: 1789579652, Encrypted: earlierBox2},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z", Content: "backup", VMID: 9201, Size: 20811501236, CTime: 1789595994, Encrypted: earlierBox1},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey},
},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
l := NewForeignKeyLedger()
r.SetForeignKeyLedger(l)
if got := l.ForeignKeyArchives(context.Background()); got != nil {
t.Fatalf("before any evaluation the ledger must be nil (the hub keeps its state); got %v", got)
}
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err != nil {
t.Fatal(err)
}
st := l.ForeignKeyArchives(context.Background())
want := hub.ForeignKeyArchives{Target: "felhom-pbs", Count: 2, Oldest: "2026-09-16T17:27:32Z", Newest: "2026-09-16T21:59:54Z"}
var got []hub.ForeignKeyArchives
if st != nil {
got = st.Tiers
}
if len(got) != 1 || got[0] != want {
t.Fatalf("after one evaluation the ledger must report felhom-pbs: 2 archives 2026-09-16T17:27:32Z…21:59:54Z; got %v", got)
}
}
// Evaluated and none found → the stanza with `tiers: []`; not evaluated → no stanza at all. Never a null on the wire.
func TestR366_EvaluatedWithNoneIsAnEmptyList(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey}},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
l := NewForeignKeyLedger()
r.SetForeignKeyLedger(l)
b, _ := json.Marshal(hub.HostReport{ForeignKeyArchives: l.ForeignKeyArchives(context.Background())})
if strings.Contains(string(b), "foreign_key_archives") {
t.Fatalf("not evaluated must omit the stanza; got %s", b)
}
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err != nil {
t.Fatal(err)
}
b, _ = json.Marshal(hub.HostReport{ForeignKeyArchives: l.ForeignKeyArchives(context.Background())})
if !strings.Contains(string(b), `"foreign_key_archives":{"tiers":[]}`) {
t.Fatalf("evaluated with none must report tiers: []; got %s", b)
}
}
+36
View File
@@ -0,0 +1,36 @@
package backup
import (
"context"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-672 (v0.133.0): a restore-test the SPACE preflight refused is the test's RESULT — recorded for the
// hub as pass=false with the reason, never dropped (a band skip still is) and never a pass.
//
// COMPANION RED-PROOF (REPORT): the scheduler's pre-v0.133.0 `if res.Skipped { return }` → "a space
// refusal never reached the host report".
func TestR672_SpaceSkipIsReportedNotDropped(t *testing.T) {
store := NewStore()
rt := &fakeRTRunner{res: reconcile.RestoreTestResult{Archive: "vol", Skipped: true,
SkipReason: "skipped: not enough space on local-lvm: restoring 21.1 GiB (vzdump log) needs 30.3 GiB free, has 21.6 GiB"}}
s := NewScheduler(SchedulerOptions{
Runner: rt, Pick: func(context.Context) (string, error) { return "vol", nil }, Store: store,
Spec: func(context.Context, string) reconcile.RestoreTestSpec {
return reconcile.RestoreTestSpec{RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009}
},
Cadence: time.Hour, Logger: quiet(),
})
s.tick(context.Background())
got := store.RestoreTests(context.Background())
if len(got) != 1 {
t.Fatalf("a space refusal never reached the host report: %+v", got)
}
if got[0].Pass || !got[0].Skipped || !strings.HasPrefix(got[0].Error, "skipped: not enough space") {
t.Fatalf("record = %+v — want pass=false, skipped, the reason as the error", got[0])
}
}
+83
View File
@@ -0,0 +1,83 @@
package backup
import (
"context"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-685 (v0.134.0) — a whole-box backup that cannot fit its LOCAL target is a named skip BEFORE anything
// starts, with the numbers in the record's Error, never a vzdump that fills the disk and fails.
const gib = int64(1) << 30
func spaceAPI(avail int64, lastArchive int64, typ string) *fakeBackupAPI {
api := &fakeBackupAPI{vzdumpUPID: "UPID:vzdump:1",
storages: []proxmox.Storage{{Storage: "local", Type: typ, Content: "backup", Avail: avail}}}
if lastArchive > 0 {
api.content = []proxmox.StorageContent{{VolID: "local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst",
Content: "backup", VMID: 9201, Size: lastArchive, CTime: 1790280000}}
}
return api
}
// TestR685_BackupThatCannotFitIsSkipped — demo-hp's shape: an 8.2 GB archive, 4 GiB free. No vzdump is
// started; the record says why, with the numbers, under a stable prefix.
//
// COMPANION RED-PROOF (REPORT.md): drop the spaceFits call from backup() — this test fails at "a vzdump
// was started on a target that cannot hold it".
func TestR685_BackupThatCannotFitIsSkipped(t *testing.T) {
api := spaceAPI(4*gib, 8182759056, "dir")
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
rec, err := r.Backup(context.Background(), 9201)
if len(api.vzdumps) != 0 {
t.Fatalf("a vzdump was started on a target that cannot hold it: %+v", api.vzdumps)
}
if err == nil || rec.Success || !strings.HasPrefix(rec.Error, BackupSkipNoSpacePrefix) {
t.Fatalf("want a named skip, got err=%v rec=%+v", err, rec)
}
for _, want := range []string{"4.0 GiB free", "7.6 GiB", "10.5 GiB"} {
if !strings.Contains(rec.Error, want) {
t.Errorf("the reason must carry the numbers (%q missing): %s", want, rec.Error)
}
}
}
// TestR685_BackupThatFitsRuns — the same archive with 16 GiB free (demo-hp after tonight's prune) runs.
func TestR685_BackupThatFitsRuns(t *testing.T) {
api := spaceAPI(16*gib, 8182759056, "dir")
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
_, _ = r.Backup(context.Background(), 9201)
if len(api.vzdumps) != 1 {
t.Fatalf("a backup that fits must run, vzdumps=%d", len(api.vzdumps))
}
}
// TestR685_FailsOpen — never refuse on what is not KNOWN: a PBS target, a first backup (no archive to
// size from), an unknown free figure, or a storage list that cannot be read.
func TestR685_FailsOpen(t *testing.T) {
cases := map[string]*fakeBackupAPI{
"pbs target": spaceAPI(1*gib, 8*gib, "pbs"),
"first backup": spaceAPI(1*gib, 0, "dir"),
"avail unknown": spaceAPI(0, 8*gib, "dir"),
"storage list error": func() *fakeBackupAPI {
a := spaceAPI(1*gib, 8*gib, "dir")
a.storageErr = context.DeadlineExceeded
return a
}(),
}
for name, api := range cases {
target := "local"
if name == "pbs target" {
api.storages[0].Storage = "felhom-pbs"
target = "felhom-pbs"
}
r := NewBackupRunner(api, target, proxmox.ModeSnapshot, "", "", quiet())
_, _ = r.Backup(context.Background(), 9201)
if len(api.vzdumps) != 1 {
t.Errorf("%s: the preflight must fail OPEN, but no vzdump ran", name)
}
}
}
+112
View File
@@ -0,0 +1,112 @@
package backup
import (
"context"
"fmt"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-689 (v0.135.0) — demo-hp keeps its golden template in `local:backup/`. It is content "backup",
// 654 MB and so "plausibly complete", and it was the newest SETTLED entry: the restore test picked it
// every 6 h and failed extractconfig with a 403 (measured 2026-09-24 10:36, 09-25 04:57 and 10:57),
// while the guest's own archive — younger than the 24 h settle — went untested and nothing was proven.
//
// COMPANION RED-PROOF (REPORT.md): drop the guestBackupArchive call from PickSettledRestoreCandidateOn —
// this test then picks `local:backup/felhom-golden-0.236.0.tar.zst`.
func TestR689_TheRestoreTestNeverPicksTheGolden(t *testing.T) {
const day = int64(86400)
now := int64(1790370000) // 2026-09-25 ~19:00Z
api := &fakeBackupAPI{content: []proxmox.StorageContent{
// the guest's real archive, settled (older than the cutoff below)
{VolID: "local:backup/vzdump-lxc-9201-2026_09_22-21_59_25.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 3*day},
// the golden: newer, settled, big, and NOT a backup of a guest
{VolID: "local:backup/felhom-golden-0.236.0.tar.zst", Content: "backup", Size: 654115664, CTime: now - 2*day},
// a hand-copied tarball that PVE happens to attribute to a vmid — the name is not a vzdump's
{VolID: "local:backup/copy-of-9201.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 2*day},
}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now-day, 0).UTC())
if err != nil {
t.Fatal(err)
}
if got != "local:backup/vzdump-lxc-9201-2026_09_22-21_59_25.tar.zst" {
t.Fatalf("picked %q — the restore test must prove a backup OF A GUEST", got)
}
}
func TestR689_GuestBackupArchiveShapes(t *testing.T) {
for _, c := range []struct {
e proxmox.StorageContent
ok bool
}{
{proxmox.StorageContent{VolID: "local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst", VMID: 9201}, true},
{proxmox.StorageContent{VolID: "local:backup/vzdump-qemu-300-2026_09_24-21_59_25.vma.zst", VMID: 300}, true},
{proxmox.StorageContent{VolID: "felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z", VMID: 9201}, true},
{proxmox.StorageContent{VolID: "felhom-pbs:backup/vm/300/2026-07-28T05:31:14Z", VMID: 300}, true},
{proxmox.StorageContent{VolID: "local:backup/felhom-golden-0.236.0.tar.zst"}, false},
{proxmox.StorageContent{VolID: "local:backup/vzdump-lxc-9201-x.tar.zst", VMID: 9202}, false}, // vmid disagrees with the name
{proxmox.StorageContent{VolID: "felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z"}, false}, // no vmid reported
} {
if ok, why := guestBackupArchive(c.e); ok != c.ok {
t.Errorf("%s vmid=%d: ok=%v (%s), want %v", c.e.VolID, c.e.VMID, ok, why, c.ok)
}
}
}
// R-689 (v0.136.0) — the measured demo-hp shape right after v0.135.0: the golden (skipped), a leftover archive
// of guest 9100 deleted in August (settled), and today's archive of 9201 (not settled yet). The pick must be
// NOTHING — never the deleted guest's archive. With 9201's archive settled, that one.
//
// COMPANION RED-PROOF (REPORT.md): drop the known-guest check — the pick is the 9100 leftover.
func TestR689_AnArchiveOfADeletedGuestIsNeverPicked(t *testing.T) {
const day = int64(86400)
now := int64(1790476000)
api := &fakeBackupAPI{goneGuests: map[int]bool{9100: true}, content: []proxmox.StorageContent{
{VolID: "local:backup/felhom-golden-0.236.0.tar.zst", Content: "backup", Size: 654115664, CTime: now - 14*day},
{VolID: "local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst", Content: "backup", VMID: 9100, Size: 656970239, CTime: now - 37*day},
{VolID: "local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 7*3600},
}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now-day, 0).UTC())
if err != nil || got != "" {
t.Fatalf("picked %q err=%v — a deleted guest's archive proves nothing about this box", got, err)
}
got, _, _ = r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now, 0).UTC())
if got != "local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst" {
t.Fatalf("with 9201's archive settled the pick is %q", got)
}
}
// Any OTHER lookup failure is not "the guest is gone": the tier must read UNKNOWN (an error), never
// "nothing to prove".
func TestR689_AGuestLookupFailureIsUnknownNotEmpty(t *testing.T) {
api := &fakeBackupAPI{cfgErr: fmt.Errorf("proxmox: connection refused"), content: []proxmox.StorageContent{
{VolID: "local:backup/vzdump-lxc-9201-x.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: 10},
}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Time{}); err == nil {
t.Fatal("a failed guest lookup read as a clean answer")
}
}
// v0.137.0 — THE MEASURED ANSWER: the agent's token sees only its pool, so for the deleted guest PVE says 403
// "permission denied at /vms/9100", not "does not exist" (demo-hp, right after v0.136.0 — the local tier read
// UNKNOWN). Such a guest is not one this agent manages: its archive is skipped, the tier is not an error.
//
// COMPANION RED-PROOF (REPORT.md): drop the "permission denied" case — the pick errors.
func TestR689_AGuestOutsideTheAgentsACLIsNotAKnownGuest(t *testing.T) {
const day = int64(86400)
now := int64(1790476000)
api := &fakeBackupAPI{aclGuests: map[int]bool{9100: true}, content: []proxmox.StorageContent{
{VolID: "local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst", Content: "backup", VMID: 9100, Size: 656970239, CTime: now - 37*day},
{VolID: "local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 7*3600},
}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now-day, 0).UTC())
if err != nil || got != "" {
t.Fatalf("picked %q err=%v — want nothing and no error (the only settled archive is not ours)", got, err)
}
}
+74
View File
@@ -0,0 +1,74 @@
package backup
import (
"context"
"errors"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
const (
thisBoxKey = "de:51:7a:18:cb:39:22:30:2c:84:f5:8b:d1:91:4b:7e:81:bb:69:b8:89:0f:57:ac:d3:59:e1:1a:62:25:11:2c"
earlierBox1 = "6b:ca:5f:3f:ca:0f:e2:3f:fb:24:62:89:bf:e7:64:59:9a:41:c5:e6:e3:9f:3f:5f:e1:71:7b:a1:9d:24:67:82"
earlierBox2 = "fe:3d:db:95:d4:df:ab:e1:7d:4a:89:fa:2b:07:53:6a:e4:d2:85:95:d1:90:27:4b:d9:c6:92:20:95:04:e5:d4"
)
// R-727 (v0.138.0) — the 2026-09-30 shape, measured on a fresh box for a returning customer: the PBS
// namespace held two archives of earlier boxes (same guest 9201, same token) and this box's own, which was not
// settled yet. The old picker chose the earlier box's newest settled archive and failed `wrong key`.
// The CONSEQUENCE asserted: no archive of another box is ever picked; with this box's archive settled it is picked.
// COMPANION RED-PROOF: remove the `ownKey != "" && !EqualFold(...)` skip → the first case picks 2026-09-16T21:59:54Z.
func TestR727_TheRestoreTestTakesOnlyThisBoxsArchives(t *testing.T) {
day := int64(86400)
now := int64(1790740000) // 2026-09-30 ~04:00Z
own := proxmox.StorageContent{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey}
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T17:27:32Z", Content: "backup", VMID: 9201, Size: 4774114206, CTime: 1789579652, Encrypted: earlierBox2},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z", Content: "backup", VMID: 9201, Size: 20811501236, CTime: 1789595994, Encrypted: earlierBox1},
own,
},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
// 1. The night of 2026-09-30: this box's own archive is ~6 h old, not settled (cutoff 24 h) — nothing to prove.
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Unix(now-day, 0).UTC())
if err != nil {
t.Fatal(err)
}
if got != "" {
t.Fatalf("picked %q — an archive of ANOTHER box is never this box's proof (R-727)", got)
}
// 2. A day later this box's own archive is settled — it is the one picked.
got, _, err = r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Unix(now+day, 0).UTC())
if err != nil {
t.Fatal(err)
}
if got != own.VolID {
t.Fatalf("picked %q, want this box's own %q", got, own.VolID)
}
}
// An unencrypted storage (a local dir) holds only this box's vzdumps — no key filter applies.
func TestR727_UnencryptedStorageIsNotFiltered(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "local", Type: "dir"}},
content: []proxmox.StorageContent{{VolID: "local:backup/vzdump-lxc-9201-2026_09_29-21_27_05.tar.zst", Content: "backup", VMID: 9201, Size: 955425507, CTime: 1790710025}},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
if got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Time{}); err != nil || got == "" {
t.Fatalf("got %q err %v", got, err)
}
}
// A storage-list failure makes the tier UNKNOWN (an error), never "nothing to prove".
func TestR727_KeyLookupFailureIsUnknown(t *testing.T) {
api := &fakeBackupAPI{storageErr: errors.New("proxmox: GET /storage -> HTTP 500"), content: []proxmox.StorageContent{{VolID: "felhom-pbs:backup/ct/9201/x", Content: "backup", VMID: 9201}}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err == nil {
t.Fatal("a failed key lookup must surface as an error (tier UNKNOWN)")
}
}
+63
View File
@@ -0,0 +1,63 @@
package backup
import (
"context"
"fmt"
"sync/atomic"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-874 (v0.145.0). THE MEASURED SHAPE (Part F spike, Tester 2): power-on sessions of ~1.5 h and ~5 min against a
// 6 h evaluation ticker that restarts at every start — no restore-test ever evaluated. Now the first evaluation runs
// FirstEval after start.
// COMPANION RED-PROOF: drop the first-evaluation timer in Run (back to the bare ticker) → "no evaluation within".
func TestR874_FirstEvaluationAfterStart(t *testing.T) {
var picks int32
s := NewScheduler(SchedulerOptions{
Runner: &fakeRTRunner{res: reconcile.RestoreTestResult{Pass: true, Verified: "boot+running"}},
Pick: func(context.Context) (string, error) {
atomic.AddInt32(&picks, 1)
return fmt.Sprintf("local:backup/vzdump-lxc-9201-%d.tar.zst", atomic.LoadInt32(&picks)), nil
},
Store: NewStore(), Spec: (&specSpy{}).build,
Cadence: 6 * time.Hour, FirstEval: 30 * time.Millisecond, Logger: quiet(),
})
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
go func() { _ = s.Run(ctx); close(done) }()
deadline := time.Now().Add(3 * time.Second)
for atomic.LoadInt32(&picks) == 0 && time.Now().Before(deadline) {
time.Sleep(10 * time.Millisecond)
}
cancel()
<-done
if atomic.LoadInt32(&picks) == 0 {
t.Fatal("no evaluation within 3 s of start (FirstEval 30 ms) — a box with short sessions never gets a restore-test")
}
}
// The earned restraint stays: an agent that restarts before FirstEval never evaluates (a crash loop does not
// hammer a failing tier).
func TestR874_CrashLoopNeverEvaluates(t *testing.T) {
var picks int32
for i := 0; i < 5; i++ { // five quick "restarts"
s := NewScheduler(SchedulerOptions{
Runner: &fakeRTRunner{res: reconcile.RestoreTestResult{Pass: true}},
Pick: func(context.Context) (string, error) { atomic.AddInt32(&picks, 1); return "x", nil },
Store: NewStore(), Spec: (&specSpy{}).build,
Cadence: 6 * time.Hour, FirstEval: 200 * time.Millisecond, Logger: quiet(),
})
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Millisecond)
_ = s.Run(ctx)
cancel()
}
if n := atomic.LoadInt32(&picks); n != 0 {
t.Fatalf("a restart before FirstEval evaluated %d time(s)", n)
}
if DefaultFirstEval != 30*time.Minute {
t.Fatalf("DefaultFirstEval = %s, the documented 30 min", DefaultFirstEval)
}
}
+26
View File
@@ -140,3 +140,29 @@ func (s *Scheduler) evaluateTier(ctx context.Context, target string, cutoff time
func (s *Scheduler) EvaluateDueTier(ctx context.Context, target string) DueVerdict {
return s.evaluateTier(ctx, target, s.settleCutoff())
}
// verdictSummary renders one compact line of per-tier verdicts for the "nothing due" log.
//
// It re-evaluates rather than threading the verdicts out of pickForThisRun, and that is a
// deliberate trade: this runs only on the path where NOTHING is due, so the cost is one extra
// storage listing per tier on an otherwise idle evaluation (measured 18 ms local / 392 ms offsite,
// R-86 Part 1.4), and in exchange the logging path cannot drift from the deciding path by holding a
// stale copy of it. If that cost ever matters, pass the verdicts in — do not let the two diverge.
func (s *Scheduler) verdictSummary(ctx context.Context) string {
out := ""
for _, v := range s.EvaluateDue(ctx) {
if out != "" {
out += "; "
}
switch {
case v.Err != nil:
out += v.Target + ": UNKNOWN (" + v.Err.Error() + ")"
default:
out += v.Target + ": " + v.Reason
}
}
if out == "" {
return "no tiers configured"
}
return out
}
+146 -11
View File
@@ -4,8 +4,10 @@ import (
"context"
"errors"
"fmt"
"log/slog"
"os"
"path/filepath"
"strings"
"testing"
"time"
@@ -124,8 +126,8 @@ func dailyArchives(tier string, n int) []archiveStub {
// COMPANION RED-PROOF (observed 2026-08-03). In Scheduler.evaluateTier, the per-archive comparison
// was replaced by the naive age rule:
//
// - if ok && proven == archive { … not due … }
// + if s.now().Sub(landed) < s.settle { … not due … } // and the proven-archive check deleted
// - if ok && proven == archive { … not due … }
// - if s.now().Sub(landed) < s.settle { … not due … } // and the proven-archive check deleted
//
// and the picker cutoff was removed (`cutoff := time.Time{}`), i.e. exactly "is the newest archive
// old enough". Result:
@@ -194,12 +196,13 @@ func TestDue_WeeklyTierIsProvedOncePerArchive(t *testing.T) {
// COMPANION RED-PROOF (observed 2026-08-03): revert the state to per-tier TIME by making
// ProvenArchive ignore the stored archive —
//
// - if !ok || p.Archive == "" { return "", false }
// + return "", false // per-tier time only, the pre-R-86 state
// - if !ok || p.Archive == "" { return "", false }
// - return "", false // per-tier time only, the pre-R-86 state
//
// → --- FAIL: TestDue_RestartRunsNothing
// restoretest_due_test.go:226: an agent restart must not trigger a restore-test; 2 restart(s)
// produced 4 run(s)
//
// restoretest_due_test.go:226: an agent restart must not trigger a restore-test; 2 restart(s)
// produced 4 run(s)
//
// Four: the same already-proven archive re-tested on EVERY evaluation after EVERY restart, which is
// today's behaviour with the ticker's phase reset by the deploy. Restored.
@@ -256,12 +259,13 @@ func TestDue_NewSettledArchiveMakesAProvedTierDueAgain(t *testing.T) {
//
// COMPANION RED-PROOF (observed 2026-08-03): give credit on failure in Scheduler.tick —
//
// - if rt.Pass && s.rtState != nil && target != "" {
// + if s.rtState != nil && target != "" {
// - if rt.Pass && s.rtState != nil && target != "" {
// - if s.rtState != nil && target != "" {
//
// → --- FAIL: TestDue_FailingTierIsRetriedAndNeverProven
// restoretest_due_test.go: a failing tier must keep being retried; got 1 run(s) over 3
// evaluations
//
// restoretest_due_test.go: a failing tier must keep being retried; got 1 run(s) over 3
// evaluations
//
// A single failure would have retired the archive as proven — a permanently broken DR tier looking
// freshly verified, which is the loudest signal this system produces going silent. Restored.
@@ -438,7 +442,7 @@ func TestRestoreTestState_ArchiveRoundTrips(t *testing.T) {
path := filepath.Join(t.TempDir(), "rt.json")
now := time.Now().UTC().Truncate(time.Second)
st := NewRestoreTestState(path)
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", now); err != nil {
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", "pbs", "boot+running", now); err != nil {
t.Fatal(err)
}
re := NewRestoreTestState(path)
@@ -456,3 +460,134 @@ func TestRestoreTestState_ArchiveRoundTrips(t *testing.T) {
func writeFileForTest(path, content string) error {
return os.WriteFile(path, []byte(content), 0o600)
}
// Standing rule 3: an absent log line is not evidence. "Nothing is due" is now the NORMAL outcome of
// an evaluation, so it must produce a POSITIVE observable naming each tier's verdict — otherwise a
// quiet journal is equally consistent with a healthy loop and a dead goroutine.
//
// COMPANION RED-PROOF (observed 2026-08-03): drop the summary back to a bare
// `s.logger.Debug("backup: restore-test not due this evaluation")` and this fails with
// "a not-due evaluation must name each tier's verdict; got \"\"" — i.e. nothing at INFO at all.
func TestDue_NothingDueStillNamesEveryTiersVerdict(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{
"local": {{volid: "local:backup/a.tar.zst", landed: day0}},
"felhom-pbs": nil, // no archive at all
}}
h := newDueHarness(t, day0.AddDate(0, 0, 1), 24*time.Hour, true, []string{"local", "felhom-pbs"}, ts)
// Prove the local tier so NOTHING is due.
if err := h.st.RecordSuccess("local", "local:backup/a.tar.zst", "local", "boot+running", h.clock); err != nil {
t.Fatal(err)
}
// Assert what the SCHEDULER emits on a real evaluation, not what a helper returns — a helper
// test would pass against a tick that never calls it.
var logbuf strings.Builder
h.s.logger = slog.New(slog.NewTextHandler(&logbuf, &slog.HandlerOptions{Level: slog.LevelInfo}))
h.s.tick(context.Background())
got := logbuf.String()
for _, want := range []string{"local", "felhom-pbs", "already proven", "no settled archive"} {
if !strings.Contains(got, want) {
t.Fatalf("a not-due evaluation must name each tier's verdict; got %q (missing %q)", got, want)
}
}
}
// A tier whose storage cannot be listed must say UNKNOWN in that same line — a lookup failure that
// reads as "nothing due" is the silence this rule exists to prevent.
func TestDue_VerdictSummaryNamesAnUnknownTier(t *testing.T) {
ts := &tierStorage{
archives: map[string][]archiveStub{"local": nil},
err: map[string]error{"felhom-pbs": errors.New("storage unreachable")},
}
h := newDueHarness(t, day0, 24*time.Hour, true, []string{"local", "felhom-pbs"}, ts)
got := h.s.verdictSummary(context.Background())
if !strings.Contains(got, "UNKNOWN") || !strings.Contains(got, "storage unreachable") {
t.Fatalf("an unlistable tier must read as UNKNOWN with its error; got %q", got)
}
}
// ── R-189 — the persisted proof must be REPORTABLE, and must refuse to lie ───────────────────
//
// A proof held only in the in-memory store dies with the process, and under per-archive due-ness the
// agent will not repeat the work. So the persisted record has to be able to become a host-report
// entry — without inventing anything it does not know.
//
// COMPANION RED-PROOF (observed 2026-08-03): drop the `reportable()` filter from
// ProvenRestoreTests, so a pre-R-189 record (archive but no tier) is emitted →
//
// --- FAIL: TestProvenRestoreTests_RefusesToReportWhatItCannotDescribe
// restoretest_due_test.go: a record with no TIER must not be reported (the hub keys its
// per-tier proof on it); got [{... SourceTier: ...}]
//
// Restored.
func TestProvenRestoreTests_RefusesToReportWhatItCannotDescribe(t *testing.T) {
path := filepath.Join(t.TempDir(), "rt.json")
// v1 (a bare time), v2 (archive, no tier) and v3 (complete) side by side — every shape this
// file has ever had, which is what a real box carries after two upgrades.
legacy := `{
"old-v1": "2026-07-30T02:11:07Z",
"old-v2": {"archive":"felhom-backup:backup/vzdump-lxc-9201-a.tar.zst","proven_at":"2026-08-01T04:41:58Z"},
"felhom-pbs": {"archive":"felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z","tier":"pbs","verified":"boot+running","proven_at":"2026-08-03T13:25:14Z"}
}`
if err := writeFileForTest(path, legacy); err != nil {
t.Fatal(err)
}
got := NewRestoreTestState(path).ProvenRestoreTests(context.Background())
if len(got) != 1 {
t.Fatalf("only the record that can be described honestly may be reported; got %d: %+v", len(got), got)
}
e := got[0]
if e.SourceTier != "pbs" {
t.Fatalf("a record with no TIER must not be reported (the hub keys its per-tier proof on it); got %+v", got)
}
if e.SourceArchive != "felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z" || !e.Pass {
t.Fatalf("the reported entry must be the stored proof, unchanged; got %+v", e)
}
if e.TestedAt != "2026-08-03T13:25:14Z" {
t.Fatalf("the entry must carry the time the run passed, not now(); got %q", e.TestedAt)
}
if e.Verified != "boot+running" {
t.Fatalf("what the run verified must survive the round trip; got %q", e.Verified)
}
// Run mechanics are NOT invented: an absent duration is not a claim, a fabricated one would be.
if e.DurationSeconds != 0 || e.ScratchVMID != 0 {
t.Fatalf("the re-report must not invent run mechanics it never stored; got duration=%v scratch=%d",
e.DurationSeconds, e.ScratchVMID)
}
// The legacy records still serve the DUE-check, which is a separate question from reporting.
if _, ok := NewRestoreTestState(path).ProvenArchive("old-v2"); !ok {
t.Fatal("a v2 record must still answer the due-check even though it cannot be reported")
}
}
// A tier proved through the SCHEDULER (not by hand) lands in the state complete enough to report —
// the production path, not a hand-built fixture.
func TestScheduler_ProofIsRecordedReportably(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{"felhom-pbs": {{volid: "felhom-pbs:backup/ct/9201/w0", landed: day0}}}}
h := newDueHarness(t, day0.AddDate(0, 0, 1).Add(97*time.Minute), 24*time.Hour, true, []string{"felhom-pbs"}, ts)
// The fake runner echoes the spec's tier; give the spec a tier the way main.go does.
h.s.spec = func(_ context.Context, archive string) reconcile.RestoreTestSpec {
return reconcile.RestoreTestSpec{RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009, SourceTier: "pbs"}
}
h.s.tick(context.Background())
got := h.st.ProvenRestoreTests(context.Background())
if len(got) != 1 {
t.Fatalf("a scheduled pass must leave a REPORTABLE proof; got %d: %+v", len(got), got)
}
if got[0].SourceTier != "pbs" || got[0].SourceArchive != "felhom-pbs:backup/ct/9201/w0" {
t.Fatalf("the proof must name the tier and the archive the run used; got %+v", got[0])
}
}
// A FAILED run leaves nothing to report — the asymmetry of §8.1, asserted rather than assumed.
func TestScheduler_AFailureLeavesNoPersistedProof(t *testing.T) {
ts := &tierStorage{archives: map[string][]archiveStub{"felhom-pbs": {{volid: "felhom-pbs:backup/ct/9201/w0", landed: day0}}}}
h := newDueHarness(t, day0.AddDate(0, 0, 1).Add(97*time.Minute), 24*time.Hour, false, []string{"felhom-pbs"}, ts)
h.s.tick(context.Background())
if got := h.st.ProvenRestoreTests(context.Background()); len(got) != 0 {
t.Fatalf("a FAILED run must persist nothing — a failing tier is retried, and a stored failure "+
"would outlive the fault; got %+v", got)
}
}
+99 -13
View File
@@ -1,12 +1,15 @@
package backup
import (
"context"
"encoding/json"
"os"
"path/filepath"
"sort"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// RestoreTestState persists the last SUCCESSFUL restore-test per backup tier.
@@ -49,16 +52,45 @@ type RestoreTestState struct {
last map[string]provenTier // target id → what was last PROVEN on that tier
}
// provenTier is one tier's proof: the archive that passed, and when it passed.
// provenTier is one tier's proof: the archive that passed, which tier it was, what was verified,
// and when.
//
// R-189 added `Tier` and `Verified`. Until then this record could answer the DUE-check but could not
// be REPORTED, and being reportable is what closes R-189: a proof held only in the in-memory result
// store vanishes on restart, and under per-archive due-ness the box will not repeat the work, so the
// hub can stay ignorant of a real success until the next archive generation.
//
// `Tier` is stored rather than derived because it is known for certain at proof time (the run's own
// spec used it to choose the restore timeout) and deriving it later would need a storage-type lookup
// at report-building time — a network call that can fail, on a path where failing means mis-labelling
// a proof. Store what you knew when you knew it.
type provenTier struct {
Archive string // volid of the archive that PASSED; "" = a legacy record with no archive
At time.Time // when that run passed (UTC)
Archive string // volid of the archive that PASSED; "" = a legacy record with no archive
Tier string // "local" | "pbs" — as the run reported it; "" = pre-R-189 record
Verified string // what the run verified (e.g. "boot+running"); "" = pre-R-189 record
At time.Time // when that run passed (UTC)
}
// provenTierJSON is the on-disk shape (R-86). The legacy shape was a bare RFC3339 STRING per
// target; both are read, only this one is written — see NewRestoreTestState.
// reportable reports whether this record can be re-reported to the hub as a restore-test result.
//
// It needs BOTH the archive and the tier: the hub keys its edge-triggered failure state on the
// archive and its per-tier proof lookup on the tier, so an entry missing either is not a usable
// proof — and emitting one anyway would be a report the hub cannot act on, dressed as evidence.
// A pre-R-189 record is therefore silently not reported; the tier's next real proof fills it in.
func (p provenTier) reportable() bool { return p.Archive != "" && p.Tier != "" }
// provenTierJSON is the on-disk shape. Two older shapes are read and neither is written:
//
// v1 (pre-R-86) "<target>": "<RFC3339>" — a time, no archive
// v2 (R-86) "<target>": {archive, proven_at} — due-check usable, not reportable
// v3 (R-189) "<target>": {archive, tier, verified, …} — both
//
// Fields absent in an older file unmarshal to "", which is exactly the "no usable proof" signal the
// readers above test for — the migration needs no version number because the absence IS the answer.
type provenTierJSON struct {
Archive string `json:"archive"`
Tier string `json:"tier,omitempty"`
Verified string `json:"verified,omitempty"`
ProvenAt string `json:"proven_at"`
}
@@ -99,21 +131,31 @@ func NewRestoreTestState(path string) *RestoreTestState {
if perr != nil {
continue
}
s.last[target] = provenTier{Archive: cur.Archive, At: t.UTC()}
s.last[target] = provenTier{Archive: cur.Archive, Tier: cur.Tier, Verified: cur.Verified, At: t.UTC()}
}
return s
}
// RecordSuccess stamps a tier as proven at t, naming the ARCHIVE that passed. Only call this for a
// PASSING restore-test — the archive is what makes the tier not-due, so recording one for a failed
// run would retire the archive unproven.
func (s *RestoreTestState) RecordSuccess(target, archive string, t time.Time) error {
// RecordSuccess stamps a tier as proven at t, naming the ARCHIVE that passed, the TIER the run
// reported, and what it verified. Only call this for a PASSING restore-test — the archive is what
// makes the tier not-due, so recording one for a failed run would retire the archive unproven.
//
// ONLY SUCCESSES ARE PERSISTED, AND THE ASYMMETRY IS DELIBERATE (R-189 §8.1). Say it here because
// the next reader will notice failures are absent and try to "fix" it:
//
// a SUCCESS suppresses future work — a proven archive is never re-tested, so a lost proof leaves
// the system quietly less tested than it believes. It must survive a restart.
//
// a FAILURE causes future work — a failing tier stays due and is retried at the next evaluation,
// so a lost failure heals itself within one interval. Persisting it would do the opposite of
// helping: a healed tier would keep reporting a failure that is no longer true.
func (s *RestoreTestState) RecordSuccess(target, archive, tier, verified string, t time.Time) error {
if target == "" {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
s.last[target] = provenTier{Archive: archive, At: t.UTC()}
s.last[target] = provenTier{Archive: archive, Tier: tier, Verified: verified, At: t.UTC()}
return s.saveLocked()
}
@@ -138,7 +180,13 @@ func (s *RestoreTestState) ProvenArchive(target string) (string, bool) {
return p.Archive, true
}
// Snapshot returns a copy of the last-proven TIMES — for the host-report gauge.
// Snapshot returns a copy of the last-proven TIMES.
//
// It carried the comment "for the host-report gauge" from the day it was written and **had no caller
// at all** until R-189 — a seam built and never wired, and an invariant asserted in a comment with
// nothing pinning it, in one method. The host report is now fed by ProvenRestoreTests below, which
// carries the archive and the tier that a bare timestamp cannot. This stays for callers that want
// only the times; if it acquires none, delete it rather than let it claim a purpose again.
func (s *RestoreTestState) Snapshot() map[string]time.Time {
s.mu.Lock()
defer s.mu.Unlock()
@@ -149,6 +197,41 @@ func (s *RestoreTestState) Snapshot() map[string]time.Time {
return out
}
// ProvenRestoreTests renders the persisted proofs as host-report entries — the R-189 fix.
//
// It satisfies hub.RestoreTestReporter's shape, so the collector can merge these with the in-memory
// results. What it emits is a RE-REPORT of a run that really happened, not a synthesis:
//
// - `Pass` is true because ONLY successes are stored (RecordSuccess is the sole writer);
// - `SourceArchive`, `SourceTier`, `Verified` and `TestedAt` are the values that run reported;
// - the run mechanics (scratch VMID, duration, warnings) are NOT re-invented. An absent duration
// is not a claim; a fabricated one would be.
//
// A record that cannot be reported honestly is omitted rather than padded — see provenTier.reportable.
// **A tier with no usable proof produces NO entry**: an unproven tier reading as proven would be a
// worse defect than the one this fixes.
func (s *RestoreTestState) ProvenRestoreTests(context.Context) []hub.RestoreTest {
s.mu.Lock()
defer s.mu.Unlock()
out := make([]hub.RestoreTest, 0, len(s.last))
for _, p := range s.last {
if !p.reportable() {
continue
}
out = append(out, hub.RestoreTest{
SourceArchive: p.Archive,
SourceTier: p.Tier,
Pass: true,
Verified: p.Verified,
TestedAt: p.At.UTC().Format(time.RFC3339),
})
}
// Deterministic order: the report is compared byte-wise by the contract test, and Go's map
// iteration is randomised.
sort.Slice(out, func(i, j int) bool { return out[i].SourceTier < out[j].SourceTier })
return out
}
// OldestFirst orders targets by "least recently proven first"; never-proven sorts FIRST.
//
// This is the operator's 2026-07-26 ruling (Option 1): self-balancing, no new config knob, and it
@@ -185,7 +268,10 @@ func (s *RestoreTestState) OldestFirst(targets []string) []string {
func (s *RestoreTestState) saveLocked() error {
raw := make(map[string]provenTierJSON, len(s.last))
for target, p := range s.last {
raw[target] = provenTierJSON{Archive: p.Archive, ProvenAt: p.At.UTC().Format(time.RFC3339)}
raw[target] = provenTierJSON{
Archive: p.Archive, Tier: p.Tier, Verified: p.Verified,
ProvenAt: p.At.UTC().Format(time.RFC3339),
}
}
data, err := json.MarshalIndent(raw, "", " ")
if err != nil {
+5 -5
View File
@@ -303,11 +303,11 @@ func TestOldestFirst_Ordering(t *testing.T) {
t.Fatalf("unexpected: %v", got)
}
}
_ = st.RecordSuccess("local", "local:backup/a.tar.zst", now)
_ = st.RecordSuccess("local", "local:backup/a.tar.zst", "local", "boot+running", now)
if got := st.OldestFirst([]string{"local", "felhom-pbs"}); got[0] != "felhom-pbs" {
t.Fatalf("a never-proven tier must sort before a proven one; got %v", got)
}
_ = st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/b", now.Add(time.Hour))
_ = st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/b", "pbs", "boot+running", now.Add(time.Hour))
if got := st.OldestFirst([]string{"local", "felhom-pbs"}); got[0] != "local" {
t.Fatalf("the least recently proven must sort first; got %v", got)
}
@@ -318,8 +318,8 @@ func TestOldestFirst_Ordering(t *testing.T) {
func TestOldestFirst_DeterministicOnTies(t *testing.T) {
st := NewRestoreTestState(filepath.Join(t.TempDir(), "rt.json"))
now := time.Now().UTC()
_ = st.RecordSuccess("b-tier", "b:archive", now)
_ = st.RecordSuccess("a-tier", "a:archive", now)
_ = st.RecordSuccess("b-tier", "b:archive", "local", "boot+running", now)
_ = st.RecordSuccess("a-tier", "a:archive", "local", "boot+running", now)
for i := 0; i < 20; i++ {
if got := st.OldestFirst([]string{"b-tier", "a-tier"}); got[0] != "a-tier" {
t.Fatalf("tie-break must be deterministic; iteration %d gave %v", i, got)
@@ -334,7 +334,7 @@ func TestRestoreTestState_PersistenceAndCorruption(t *testing.T) {
now := time.Now().UTC().Truncate(time.Second)
st := NewRestoreTestState(path)
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", now); err != nil {
if err := st.RecordSuccess("felhom-pbs", "felhom-pbs:backup/ct/9201/x", "pbs", "boot+running", now); err != nil {
t.Fatal(err)
}
reopened := NewRestoreTestState(path)
+218 -1
View File
@@ -5,6 +5,7 @@ import (
"fmt"
"log/slog"
"sort"
"strconv"
"strings"
"sync"
"time"
@@ -22,6 +23,10 @@ type BackupAPI interface {
StorageContent(ctx context.Context, store string) ([]proxmox.StorageContent, error)
// ListStorage enumerates storages (name+type) — used to scope local-only retention (never prune PBS).
ListStorage(ctx context.Context) ([]proxmox.Storage, error)
// NodeStorage is GET /nodes/{node}/storage — the storages WITH live usage (avail/used). R-685's space
// preflight reads free space HERE: ListStorage (GET /storage) is the cluster DEFINITIONS and carries no
// usage at all — measured live 2026-09-24 on demo-hp, where reading it let a backup through.
NodeStorage(ctx context.Context) ([]proxmox.Storage, error)
// TaskLogTail reads trailing task-log lines — used to read the ACTUAL vzdump mode
// (PVE may downgrade a requested snapshot to stop for a stopped guest — spike B1).
TaskLogTail(ctx context.Context, upid string, limit int) ([]string, error)
@@ -61,8 +66,13 @@ type BackupRunner struct {
// due-check is served from the local-API handler goroutines.
rejectedMu sync.Mutex
rejected map[string]struct{}
// foreign (R-366 slice 2) records, per tier, the archives the pick skipped as another key's. nil = not wired.
foreign *ForeignKeyLedger
}
// SetForeignKeyLedger wires the R-366 slice-2 ledger the host report reads.
func (r *BackupRunner) SetForeignKeyLedger(l *ForeignKeyLedger) { r.foreign = l }
// NewBackupRunner builds a runner. mode defaults to snapshot (works for a stopped guest and
// for lvm-thin); the caller may pass ModeStop for storages without snapshot support. retention is the
// per-run prune spec ("keep-last=N", or "" to never prune) — only the periodic local backup sets it.
@@ -170,6 +180,16 @@ func (r *BackupRunner) backup(ctx context.Context, vmid int, onSnapshot func())
rec.UncoveredVolumes = []string{}
}
// R-685 (v0.134.0): will the new archive FIT on a local target? Asked before anything runs, so a
// target that cannot hold it is a named SKIP with the numbers, not a nightly "No space left on
// device" that only the vzdump log explains (demo-hp, every night from 2026-09-23 — R-684).
if ok, why := r.spaceFits(ctx, vmid); !ok {
rec.Error = BackupSkipNoSpacePrefix + why
rec.DurationSeconds = time.Since(start).Seconds()
r.logger.Warn("backup SKIPPED by the space preflight (R-685) — nothing was started", "vmid", vmid, "target", r.target, "reason", why)
return rec, fmt.Errorf("backup: %s", rec.Error)
}
upid, err := r.api.Vzdump(ctx, proxmox.VzdumpOptions{
VMID: vmid, Storage: r.target, Mode: r.mode, Notes: r.notes,
PruneBackups: r.localPruneSpec(ctx), // local target → keep-last=N; PBS/unknown → "" (no prune)
@@ -217,6 +237,56 @@ func (r *BackupRunner) backup(ctx context.Context, vmid int, onSnapshot func())
return rec, nil
}
// BackupSkipNoSpacePrefix starts a backup record's Error when the space preflight refused (R-685) — a
// stable prefix the controller's page and the hub can key on.
const BackupSkipNoSpacePrefix = "skipped: not enough space: "
// Space preflight margins (R-685): the new archive is predicted as the newest archive of this guest on the
// target × backupSpaceGrowth, plus backupSpaceFloorBytes of headroom for the host. MEASURED 2026-09-24:
// demo-hp 9201's archives grew 5.8 → 6.2 → 6.9 → 7.6 GB in four nights (+10 % a night at worst), so 1.25
// covers two nights' growth. PVE prunes old archives only AFTER a successful backup, so the free space
// must hold the new archive while every kept one still exists.
const (
backupSpaceGrowth = 1.25
backupSpaceFloorBytes = int64(1) << 30
)
// spaceFits answers whether a new archive of vmid fits on a LOCAL (non-PBS) target. It FAILS OPEN — a
// backup is the thing being protected, so an unreadable storage, an unknown type or a first backup (no
// previous archive to size from) proceeds and says so; only a POSITIVE "it does not fit" refuses.
func (r *BackupRunner) spaceFits(ctx context.Context, vmid int) (bool, string) {
// NodeStorage, never ListStorage: only the node view carries avail (see BackupAPI.NodeStorage).
stores, err := r.api.NodeStorage(ctx)
if err != nil {
r.logger.Warn("backup: space preflight could not read storage usage — proceeding (fail-open)", "target", r.target, "err", err)
return true, ""
}
var st *proxmox.Storage
for i := range stores {
if stores[i].Storage == r.target {
st = &stores[i]
break
}
}
if st == nil || st.Type == "pbs" || st.Avail <= 0 {
return true, "" // PBS dedups and has its own lifecycle; an unknown avail never refuses
}
_, last, err := r.latestArchive(ctx, vmid)
if err != nil || last <= 0 {
r.logger.Info("backup: space preflight has no previous archive to size from — proceeding", "vmid", vmid, "target", r.target)
return true, ""
}
need := int64(float64(last)*backupSpaceGrowth) + backupSpaceFloorBytes
if st.Avail >= need {
r.logger.Info("backup: space preflight passed", "vmid", vmid, "target", r.target, "last_archive_bytes", last, "need_bytes", need, "avail_bytes", st.Avail)
return true, ""
}
return false, fmt.Sprintf("%s has %s free; the last archive of guest %d was %s, so a new one needs about %s (old archives are removed only after a successful backup)",
r.target, humanGiB(st.Avail), vmid, humanGiB(last), humanGiB(need))
}
func humanGiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
// watchForSnapshot polls the running backup's task log until it sees the storage-snapshot marker
// (→ onSnapshot once) or the requested mode is reported as `stop` (→ downgraded; the marker will
// never come, so stop watching) or ctx is cancelled (backup finished). Best-effort: a log-read
@@ -292,12 +362,67 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
if err != nil {
return "", time.Time{}, err
}
// R-727 (v0.138.0): on an ENCRYPTED storage, only archives written with THIS storage's key are this box's.
// Measured 2026-09-30 on a fresh box for a returning customer: the PBS namespace still held two archives
// of earlier boxes (same guest id 9201, same token), the newest settled one was an earlier box's, and the
// test failed `wrong key` every evaluation. The archive carries no host id; its key fingerprint is the
// discriminator (PVE's content `encrypted`, the storage's `encryption-key`). A lookup failure returns
// the error — the tier reads UNKNOWN, never "nothing to prove".
ownKey, err := r.storageKeyFingerprint(ctx, target)
if err != nil {
return "", time.Time{}, fmt.Errorf("reading the key fingerprint of storage %s: %w", target, err)
}
var best string
var bestCTime int64 = -1
known := map[int]bool{} // vmid → the guest exists on this node (asked once per vmid per pick)
var foreignN int // R-366 slice 2: archives skipped as another key's, and their time range
var foreignMin, foreignMax int64
for _, e := range contents {
if e.Content != "backup" {
continue
}
// R-689 (v0.135.0): only a backup OF A GUEST is a restore-test candidate. demo-hp keeps its golden
// template in `local:backup/` — content "backup", 654 MB, plausibly complete — and it was picked as
// the newest settled archive every 6 h and failed extractconfig (403) each time, while the guest's
// real archive went untested.
if ok, why := guestBackupArchive(e); !ok {
r.noteNotAGuestBackupOnce(e, why)
continue
}
if ownKey != "" && !strings.EqualFold(e.Encrypted, ownKey) {
r.noteNotAGuestBackupOnce(e, fmt.Sprintf("written by another box (key %s, this box's key %s) — not this box's proof", shortFP(e.Encrypted), shortFP(ownKey)))
foreignN++
if foreignMin == 0 || e.CTime < foreignMin {
foreignMin = e.CTime
}
if e.CTime > foreignMax {
foreignMax = e.CTime
}
continue
}
// R-689 (v0.136.0): … OF A GUEST THAT STILL EXISTS here. Measured on demo-hp 2026-09-27 right after
// v0.135.0: with the golden skipped, the pick fell to `vzdump-lxc-9100-2026_08_21…`, a leftover of a
// guest deleted in August — proving nothing about any guest this box runs. "Does not exist" skips the
// archive; any OTHER lookup failure is returned, so the tier reads UNKNOWN, never "nothing to prove".
if _, seen := known[e.VMID]; !seen {
_, err := r.api.GuestConfig(ctx, e.VMID)
switch {
case err == nil:
known[e.VMID] = true
case strings.Contains(err.Error(), "does not exist"), strings.Contains(err.Error(), "permission denied"):
// v0.137.0: PVE answers 403 "permission denied at /vms/<id>" — not "does not exist" — for a guest
// outside the agent's ACL (the `felhom` pool). Measured on demo-hp after v0.136.0: the deleted
// guest 9100's archive made the local tier UNKNOWN every evaluation. A guest the agent cannot
// read is not one it manages; its archive is not a candidate.
known[e.VMID] = false
default:
return "", time.Time{}, fmt.Errorf("checking whether guest %d still exists: %w", e.VMID, err)
}
}
if !known[e.VMID] {
r.noteNotAGuestBackupOnce(e, fmt.Sprintf("guest %d does not exist on this node or is not one this agent manages", e.VMID))
continue
}
if !notAfter.IsZero() && e.CTime > notAfter.Unix() {
continue // not settled yet — a newer archive is not a reason to re-prove an older one
}
@@ -309,6 +434,9 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
bestCTime, best = e.CTime, e.VolID
}
}
if r.foreign != nil {
r.foreign.set(target, foreignN, foreignMin, foreignMax)
}
if best == "" {
return "", time.Time{}, nil
}
@@ -407,10 +535,25 @@ func (r *BackupRunner) warnRejectedArchiveOnce(e proxmox.StorageContent, why str
if seen {
return
}
r.logger.Warn("backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup",
r.logger.Warn(rejectedArchiveMessage(e),
"target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
}
// phantomCleanupPointer names the runbook that removes a PBS phantom (R-99, `09` §3 decision 140: a leftover of an
// aborted upload is deleted on the backup server, by a runbook, when one is seen — never automatically).
const phantomCleanupPointer = " — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
// rejectedArchiveMessage is the WARN text for a rejected archive. Only a PBS entry (format pbs-ct / pbs-vm) gets the
// cleanup pointer: the runbook deletes on a PBS datastore, and a tiny archive on a dir storage is not a PBS phantom.
// Pinned by TestRejectedArchiveWarnNamesTheCleanupRunbook.
func rejectedArchiveMessage(e proxmox.StorageContent) string {
msg := "backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup"
if strings.HasPrefix(e.Format, "pbs-") {
msg += phantomCleanupPointer
}
return msg
}
// demo-felhom in a single afternoon of deploys (2026-07-26).
//
// Asking the STORAGE rather than persisting the store is deliberate:
@@ -520,5 +663,79 @@ func ToHubRestoreTest(res reconcile.RestoreTestResult, testedAt time.Time) hub.R
if res.Err != nil {
rt.Error = res.Err.Error()
}
if res.SkipReason != "" { // R-672: the space preflight refused — reported, never a pass
rt.Pass = false
rt.Skipped = true
rt.Error = res.SkipReason
}
return rt
}
// guestBackupArchive reports whether a storage entry is a whole-guest backup of a known guest — a
// `vzdump-<type>-<vmid>-…` file on a dir storage, or a `backup/{ct,vm}/<vmid>/<time>` snapshot on a PBS
// datastore — whose vmid the storage itself reports. Anything else in a backup content type (a golden
// template, a hand-copied tarball) is not a backup of a guest and is never restore-tested (R-689).
// Pure, so the rule is unit-tested without a storage.
func guestBackupArchive(e proxmox.StorageContent) (bool, string) {
if e.VMID <= 0 {
return false, "not a backup of a guest (the storage reports no vmid)"
}
vol := e.VolID
if i := strings.Index(vol, ":"); i >= 0 {
vol = vol[i+1:]
}
vol = strings.TrimPrefix(vol, "backup/")
vmid := strconv.Itoa(e.VMID)
switch {
case strings.HasPrefix(vol, "vzdump-lxc-"+vmid+"-"), strings.HasPrefix(vol, "vzdump-qemu-"+vmid+"-"):
return true, ""
case strings.HasPrefix(vol, "ct/"+vmid+"/"), strings.HasPrefix(vol, "vm/"+vmid+"/"):
return true, ""
}
return false, "not a vzdump archive or a PBS snapshot of guest " + vmid
}
// noteNotAGuestBackupOnce logs, once per volid, that a backup-content entry is not a restore-test
// candidate because it is not a backup of a guest (R-689). INFO, not WARN: a golden template kept in
// the backup directory is the operator's, and not a fault.
func (r *BackupRunner) noteNotAGuestBackupOnce(e proxmox.StorageContent, why string) {
r.rejectedMu.Lock()
if r.rejected == nil {
r.rejected = map[string]struct{}{}
}
_, seen := r.rejected[e.VolID]
if !seen {
r.rejected[e.VolID] = struct{}{}
}
r.rejectedMu.Unlock()
if !seen {
r.logger.Info("backup: restore-test skips an entry that is not a backup of a guest",
"target", r.target, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
}
}
// storageKeyFingerprint returns the named storage's client-side encryption key fingerprint ("" when the
// storage is not encrypted — a local dir holds only this box's own vzdumps).
func (r *BackupRunner) storageKeyFingerprint(ctx context.Context, target string) (string, error) {
sts, err := r.api.ListStorage(ctx)
if err != nil {
return "", err
}
for _, st := range sts {
if st.Storage == target {
return strings.TrimSpace(st.EncryptionKey), nil
}
}
return "", nil
}
// shortFP is the first 8 bytes of a key fingerprint, for a log line.
func shortFP(fp string) string {
if fp == "" {
return "none"
}
if len(fp) > 23 {
return fp[:23] + "…"
}
return fp
}
+51 -9
View File
@@ -60,12 +60,21 @@ type Scheduler struct {
// R-85 tier rotation. All optional: without them the scheduler behaves exactly as before
// (single tier via `pick`), which keeps every existing caller and test working untouched.
tiers []string // configured tier target ids, primary first
tierPick TierPicker // newest archive on a named tier
rtState *RestoreTestState // persisted last-successful-per-tier (drives oldest-first)
inFlight *InFlight // shared with the backup path — Scenario F
tiers []string // configured tier target ids, primary first
tierPick TierPicker // newest archive on a named tier
rtState *RestoreTestState // persisted last-successful-per-tier (drives oldest-first)
inFlight *InFlight // shared with the backup path — Scenario F
firstEval time.Duration // R-874: the first evaluation after start
}
// DefaultFirstEval (R-874): the first due-ness evaluation runs 30 minutes after the agent starts, then every
// cadence. MEASURED need (2026-10-05 Part F spike): a box whose power-on sessions are all shorter than the 6 h
// interval (Tester 2: ~1.5 h and ~5 min) NEVER evaluated, because the ticker restarts at each start. 30 minutes
// keeps the earned restraint below — a crash-looping agent restarts far more often than that and still never
// evaluates — while a box that stays on for half an hour gets its due test. Pinned by
// TestR874_FirstEvaluationAfterStart and TestR874_CrashLoopNeverEvaluates.
const DefaultFirstEval = 30 * time.Minute
// SchedulerOptions configures a Scheduler.
type SchedulerOptions struct {
Runner RestoreTestRunner
@@ -81,6 +90,8 @@ type SchedulerOptions struct {
// 0 → no settle requirement (any archive is a candidate).
Settle time.Duration
Logger *slog.Logger
// FirstEval (R-874, v0.145.0) is when the FIRST evaluation runs after start; 0 → DefaultFirstEval.
FirstEval time.Duration
// R-85 (all optional — omit for the pre-R-85 single-tier behaviour):
// Tiers are the configured tier target ids (primary first); TierPick resolves an archive on a
@@ -110,6 +121,12 @@ func NewScheduler(opts SchedulerOptions) *Scheduler {
tierPick: opts.TierPick,
rtState: opts.State,
inFlight: opts.InFlight,
firstEval: func() time.Duration {
if opts.FirstEval > 0 {
return opts.FirstEval
}
return DefaultFirstEval
}(),
}
}
@@ -120,8 +137,9 @@ func NewScheduler(opts SchedulerOptions) *Scheduler {
// trigger any more: its phase is the process's uptime, and agent deploys reset it, which is exactly
// the defect R-86 removes. What decides that a test happens is `EvaluateDue`.
//
// It still does NOT evaluate immediately on start — the first evaluation is one interval in. That
// is an EARNED restraint, kept deliberately: a restore is heavy, agent restarts are routine, and a
// It still does NOT evaluate immediately on start. v0.145.0 (R-874): the first evaluation is
// firstEval (30 min) in, then every interval — it was one full interval in, which a box with short
// power-on sessions never reached. The restraint itself is EARNED and kept: a restore is heavy, agent restarts are routine, and a
// crash-loop that evaluated at start would hammer a permanently-failing tier as fast as it could
// restart. Due-ness does not expire while we wait, so the only cost is up to one interval of
// latency on a tier that just became due. On-demand runs use `--selftest=restore-test`.
@@ -135,6 +153,16 @@ func (s *Scheduler) Run(ctx context.Context) error {
}
s.logger.Info("backup: restore-test scheduler starting (per-archive due-check)",
"eval_interval", s.cadence, "settle", s.settle)
first := time.NewTimer(s.firstEval)
defer first.Stop()
select {
case <-ctx.Done():
s.logger.Info("backup: restore-test scheduler shutting down", "reason", ctx.Err())
return nil
case <-first.C:
s.logger.Info("backup: restore-test first evaluation after start (R-874)", "after", s.firstEval)
s.tick(ctx)
}
t := time.NewTicker(s.cadence)
defer t.Stop()
for {
@@ -180,7 +208,15 @@ func (s *Scheduler) tick(ctx context.Context) {
return
}
if archive == "" {
s.logger.Debug("backup: restore-test not due this evaluation")
// A POSITIVE OBSERVABLE, at INFO, and this is not noise — it is standing rule 3.
//
// Before R-86 every tick ran a heavy restore-test, so the scheduler was audible by
// construction. Now "nothing is due" is the NORMAL outcome, and at DEBUG it is silent: an
// empty journal would be equally consistent with a healthy loop and with a dead goroutine,
// which is the exact shape the R-88 watcher was retired for. One line per evaluation is four
// lines a day at the 6h default, and it names each tier's verdict so the answer to "why did
// nothing run last night?" is in the log rather than in a re-derivation.
s.logger.Info("backup: restore-test evaluated — nothing due", "verdicts", s.verdictSummary(ctx))
return
}
@@ -200,9 +236,11 @@ func (s *Scheduler) tick(ctx context.Context) {
spec := s.spec(ctx, archive)
spec.Archive = archive
res := s.runner.RunRestoreTest(ctx, spec)
if res.Skipped {
if res.Skipped && res.SkipReason == "" {
return // already logged by the engine (no free scratch VMID)
}
// R-672: a SPACE refusal is the test's result — reported (pass=false, the reason as the error),
// never dropped and never a pass. It earns no rotation credit, so the tier stays due.
rt := ToHubRestoreTest(res, s.now())
s.store.RecordRestoreTest(rt)
// Rotation credit is given ONLY on success. A failing tier must keep sorting first, or a tier
@@ -210,7 +248,11 @@ func (s *Scheduler) tick(ctx context.Context) {
if rt.Pass && s.rtState != nil && target != "" {
// R-86: the ARCHIVE is recorded, not merely the time — that is what makes the tier
// not-due until a NEWER archive settles, and what makes a proof survive a restart.
if err := s.rtState.RecordSuccess(target, archive, s.now()); err != nil {
// R-189: the TIER and what was VERIFIED go with it, so the proof can be RE-REPORTED after a
// restart. Both come from the run's own result, never re-derived — `rt.SourceTier` is what
// this run was actually judged as, and deriving it later would need a storage lookup that
// can fail on the one path where failing means mislabelling a proof.
if err := s.rtState.RecordSuccess(target, archive, rt.SourceTier, rt.Verified, s.now()); err != nil {
s.logger.Warn("backup: could not persist the restore-test proof state", "target", target, "err", err)
}
}
+23 -2
View File
@@ -10,8 +10,29 @@ import (
// Store holds the agent's LATEST backup result per target and the latest restore-test
// result — the point-in-time state the host-report surfaces. It is updated by the backup
// runner + the restore-test scheduler/selftest and read by the collector via the hub
// BackupReporter / RestoreTestReporter seams. In-memory (lost on restart; the cadence
// re-populates) and mutex-guarded for the concurrent collector vs scheduler access.
// BackupReporter / RestoreTestReporter seams. In-memory and mutex-guarded for the concurrent
// collector vs scheduler access.
//
// **"lost on restart; the cadence re-populates" — that sentence used to be here and it is now
// FALSE for restore-tests (R-189, 2026-08-03).** It was true while a timer re-tested every tier
// daily. Under R-86's per-archive due-check the agent will NOT re-test an archive it has already
// proven, so a proof lost to a restart is not repeated until the next archive generation — a week on
// the offsite tier — and the hub reports that tier unproven throughout. Observed, not predicted: a
// real 14.5 GB offsite restore passed, the agent was restarted 2 m 43 s later for a deploy, and two
// consecutive host-reports carried `0 restore-tests`.
//
// The durable half is `RestoreTestState` (on disk, per tier, with the archive) and the collector
// merges the two — see hub.ProvenRestoreTestReporter. This store remains the ONLY place a FAILURE is
// recorded, and that asymmetry is deliberate: a failing tier stays due and is retried, so a lost
// failure heals itself, while a lost success leaves the system quietly less tested than it believes.
// Backups are NOT unaffected (corrected 2026-10-05, R-348): byTarget is in memory too, so after a restart the
// reported backup LIST reads 0 until the next backup of each tier runs (daily local, weekly offsite) — measured
// 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What
// is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/
// deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84).
// The due-check's fallback for an UNREADABLE storage no longer reads this store alone (R-894): the newest
// success per tier is also on disk (BackupSuccessState), so a restart followed by an unreachable storage
// reads the last known copy, not "never".
type Store struct {
mu sync.Mutex
byTarget map[string]hub.Backup // latest backup per target id
+17 -13
View File
@@ -81,10 +81,8 @@ var manifest = []Capability{
{"drives-mkdir-sub", "per-drive stable dir create", "/usr/bin/mkdir", []string{"-p", "/mnt/felhom-drives/felhom-usb"}, false, ""},
{"drives-mkdir-data", "felhom-data namespace create", "/usr/bin/mkdir", []string{"-p", "/mnt/felhom-usb/felhom-data"}, false, ""},
{"drives-chown-data", "felhom-data guest-root chown", "/usr/bin/chown", []string{"100000:100000", "/mnt/felhom-usb/felhom-data"}, false, ""},
{"parent-script-install", "shared-parent boot script install", "/usr/bin/install", []string{"-m", "0755", "--", "/tmp/felhom-shared-parent-123456789.sh", "/usr/local/sbin/felhom-shared-parent.sh"}, false, ""},
{"parent-unit-install", "shared-parent boot unit install", "/usr/bin/install", []string{"-m", "0644", "--", "/tmp/felhom-shared-parent-123456789.service", "/etc/systemd/system/felhom-shared-parent.service"}, false, ""},
{"parent-unit-enable", "shared-parent boot-persistence enable", "/usr/bin/systemctl", []string{"enable", "felhom-shared-parent.service"}, false, ""},
{"parent-bind-mp8", "parent bind into guest at provision", "/usr/sbin/pct", []string{"set", "9201", "-mp8", "/mnt/felhom-drives"}, false, ""},
{"parent-bind-mp8", "parent bind into guest at provision", "/usr/sbin/pct", []string{"set", "9201", "-mp8", "/mnt/felhom-drives,mp=/mnt/felhom-drives"}, false, ""},
// ---- Disk inspect / format gate (Critical: the data-bearing classifier + format) ----
{"disk-blkid", "disk data-bearing classify (format gate)", "/usr/sbin/blkid", []string{"-p", "-o", "export", "/dev/sda"}, true, ""},
@@ -95,11 +93,11 @@ var manifest = []Capability{
{"disk-lvs", "thin-pool usage read", "/usr/sbin/lvs", []string{"--reportformat", "json", "--units", "b", "-o", "lv_name,data_percent,metadata_percent", "--", "pve/data"}, false, ""},
// ---- Storage mount units (watchdog re-mount) ----
{"mount-unit-install", "fs-UUID mount unit install", "/usr/bin/install", []string{"-o", "root", "-g", "root", "-m", "0644", "--", "/var/lib/felhom-agent/units/felhom-x.mount", "/etc/systemd/system/felhom-x.mount"}, false, ""},
{"mount-unit-install", "fs-UUID mount unit install (root content check, R-861)", "/usr/local/sbin/felhom-priv-apply", []string{"unit", "mnt-felhom\\x2dx.mount"}, false, ""},
{"mount-daemon-reload", "systemd reload after unit write", "/usr/bin/systemctl", []string{"daemon-reload"}, false, ""},
{"mount-unit-enable", "mount unit enable", "/usr/bin/systemctl", []string{"enable", "--now", "--", "felhom-x.mount"}, false, ""},
{"mount-unit-disable", "mount unit disable", "/usr/bin/systemctl", []string{"disable", "--", "felhom-x.mount"}, false, ""},
{"mount-unit-stop", "mount unit stop", "/usr/bin/systemctl", []string{"stop", "--", "felhom-x.mount"}, false, ""},
{"mount-unit-enable", "mount unit enable", "/usr/bin/systemctl", []string{"enable", "--now", "--", "mnt-felhom\\x2dx.mount"}, false, ""},
{"mount-unit-disable", "mount unit disable", "/usr/bin/systemctl", []string{"disable", "--", "mnt-felhom\\x2dx.mount"}, false, ""},
{"mount-unit-stop", "mount unit stop", "/usr/bin/systemctl", []string{"stop", "--", "mnt-felhom\\x2dx.mount"}, false, ""},
// ---- Network storage re-arm + cleanup (CAMPAIGN-3 F10/F1) ----
{"netmount-reset-failed", "NAS automount re-arm after start-limit (F10)", "/usr/bin/systemctl", []string{"reset-failed", "--", "mnt-felhom\\x2ddrives-media.automount"}, false, ""},
@@ -110,19 +108,21 @@ var manifest = []Capability{
// ---- Provisioning back-half ----
{"provision-chown", "bootstrap mount guest-root chown", "/usr/bin/chown", []string{"-R", "100000:100000", "/var/lib/felhom-agent/guests/9201"}, false, ""},
{"provision-config-mount", "bootstrap config bind mount", "/usr/sbin/pct", []string{"set", "9201", "-mp0", "/var/lib/felhom-agent/guests/9201"}, false, ""},
{"provision-config-mount", "bootstrap config bind mount", "/usr/sbin/pct", []string{"set", "9201", "-mp0", "/var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1"}, false, ""},
{"provision-onboot", "customer guest autostart (onboot)", "/usr/sbin/pct", []string{"set", "9201", "-onboot", "1"}, false, ""},
// ---- Pre-start self-heal hook + guest lifecycle ----
{"guesthook-install", "pre-start hook snippet install", "/usr/bin/install", []string{"-m", "0755", "--", "/tmp/felhom-guest-hook-123456789.sh", "/var/lib/vz/snippets/felhom-guest-hook.sh"}, false, ""},
{"guesthook-register", "pre-start hook register", "/usr/sbin/pct", []string{"set", "9201", "--hookscript", "local:snippets/felhom-guest-hook.sh"}, false, ""},
{"guesthook-delete-mp", "dead mountpoint slot delete (C1 net)", "/usr/sbin/pct", []string{"set", "9201", "--delete", "mp0"}, false, ""},
{"guest-reboot", "enroll activate-binds reboot", "/usr/sbin/pct", []string{"reboot", "9201"}, false, ""},
// ---- LAN split-horizon resolver (dnsmasq) ----
{"dnsmasq-install", "dnsmasq package install", "/usr/bin/apt-get", []string{"install", "-y", "-q", "dnsmasq"}, false, ""},
{"dnsmasq-write", "dnsmasq drop-in write", "/usr/bin/install", []string{"-m", "0644", "/tmp/felhom-resolver-x.conf", "/etc/dnsmasq.d/felhom-x.conf"}, false, ""},
{"dnsmasq-write", "dnsmasq drop-in write (root content check, R-861)", "/usr/local/sbin/felhom-priv-apply", []string{"dnsmasq", "/tmp/felhom-resolver-123456789.conf", "felhom-x.conf"}, false, ""},
{"dnsmasq-enable", "dnsmasq enable", "/usr/bin/systemctl", []string{"enable", "--now", "dnsmasq"}, false, ""},
// ---- OS updates, guest fast lane (`11` §5.4.1; the wrapper holds every rule) ----
{"osapply-run", "OS update wrapper (guest fast lane)", "/usr/local/sbin/felhom-os-apply", []string{"--plan", "/var/lib/felhom-agent/os/plan-x.json"}, false, ""},
{"dnsmasq-reload", "dnsmasq reload", "/usr/bin/systemctl", []string{"reload", "dnsmasq"}, false, ""},
{"dnsmasq-restart", "dnsmasq restart (LAN-DNS self-heal)", "/usr/bin/systemctl", []string{"restart", "dnsmasq"}, false, ""},
{"dnsmasq-rm", "dnsmasq drop-in remove (decommission)", "/usr/bin/rm", []string{"-f", "/etc/dnsmasq.d/felhom-x.conf"}, false, ""},
@@ -145,19 +145,24 @@ var manifest = []Capability{
{"controllerswap-image-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "image", "inspect", "gitea.dooplex.hu/admin/felhom-controller:0.0.0"}, true, ""},
{"controllerswap-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "inspect", "-f", "{{.State.Running}}", "felhom-controller"}, true, ""},
{"controllerswap-restart", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "systemctl", "restart", "felhom-controller-bootstrap.service"}, true, ""},
{"controllerswap-write", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "tee", "/etc/felhom-controller-image"}, true, ""},
// R-861 (a) A1 (decision 165): the write goes through the ROOT verb that checks the ref; the agent has no `tee` grant.
{"controllerswap-write", "controller-swap / managed auto-update", "/usr/local/sbin/felhom-priv-apply", []string{"controller-image", "9201"}, true, ""},
// ---- Stale-lock recovery (FELHOM_STALELOCK, v0.49.0; Critical: a guest stuck behind a stale
// reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ----
{"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""},
// ---- Weekly guest disk trim (FELHOM_FSTRIM, R-444). NON-critical: a missing grant means the thin pool is not
// reclaimed this week (the trim job WARNs per guest and the report shows the failure), not a serving outage. ----
{"guest-fstrim", "weekly guest disk trim (thin-pool reclaim, R-444)", "/usr/sbin/pct", []string{"fstrim", "9201"}, false, ""},
// ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite
// backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the
// conf install, unit enable/restart and the handshake read gate the backup path. apt-install
// (one-time bootstrap) and disable (revocation, a deliberate teardown) stay non-critical. The
// handshake read is the ONLY wg invocation (never `dump`). ----
{"wg-tools-install", "wireguard-tools package install", "/usr/bin/apt-get", []string{"install", "-y", "-q", "wireguard-tools"}, false, ""},
{"wg-conf-install", "wg-felhom conf install", "/usr/bin/install", []string{"-o", "root", "-g", "root", "-m", "0600", "--", "/var/lib/felhom-agent/wg/wg-felhom.conf", "/etc/wireguard/wg-felhom.conf"}, true, ""},
{"wg-conf-install", "wg-felhom conf install (root content check, R-861)", "/usr/local/sbin/felhom-priv-apply", []string{"wg"}, true, ""},
{"wg-enable", "wg-quick@wg-felhom enable", "/usr/bin/systemctl", []string{"enable", "--now", "wg-quick@wg-felhom"}, true, ""},
{"wg-restart", "wg-quick@wg-felhom restart (conf change)", "/usr/bin/systemctl", []string{"restart", "wg-quick@wg-felhom"}, true, ""},
{"wg-disable", "wg-quick@wg-felhom disable (revocation)", "/usr/bin/systemctl", []string{"disable", "--now", "wg-quick@wg-felhom"}, false, ""},
@@ -189,7 +194,6 @@ var manifest = []Capability{
// operator-driven op, not a steady-state serving path — a degraded grant means "can't
// self-update" (fall back to a manual SSH deploy), not a serving outage. The apply repr uses a
// staging-dir path + a placeholder sha (list-mode never runs it). ----
{"selfupdate-apply", "agent self-update apply (A/B flip)", "/usr/local/sbin/felhom-selfupdate-guarded", []string{"apply", "/var/lib/felhom-agent/selfupdate/felhom-agent-0.0.0", "0000000000000000000000000000000000000000000000000000000000000000"}, false, ""},
{"selfupdate-commit", "agent self-update commit", "/usr/local/sbin/felhom-selfupdate-guarded", []string{"commit"}, false, ""},
{"selfupdate-rollback", "agent self-update rollback", "/usr/local/sbin/felhom-selfupdate-guarded", []string{"rollback"}, false, ""},
}
+23 -6
View File
@@ -56,8 +56,10 @@ func parseSudoersEntries(t *testing.T, text string) []string {
var cur strings.Builder
for i := 0; i < len(raw); i++ {
c := raw[i]
if c == '\\' && i+1 < len(raw) {
cur.WriteByte(raw[i+1]) // unescape: keep the next char literally (\, → , ; \: → :)
if c == '\\' && i+1 < len(raw) && strings.IndexByte(",:=", raw[i+1]) >= 0 {
// unescape the sudoers grammar escapes only (\, → , ; \: → : ; \= → =). Every other backslash is kept: in
// a regex entry (R-861) `\.` `\{` `\\` mean exactly what sudo's regex engine reads.
cur.WriteByte(raw[i+1])
i++
continue
}
@@ -108,10 +110,24 @@ func globToRegex(pat string) *regexp.Regexp {
return regexp.MustCompile(b.String())
}
// entryRegex compiles one sudoers entry. R-861 (v0.146.0): an entry whose ARGUMENTS are a sudo regular expression
// (`^...$`, sudo >= 1.9.10) is matched as one — the binary literally, then a space, then the arguments joined by spaces
// (sudo's own rule). Every other entry is an fnmatch glob (globToRegex). The parser has already undone the sudoers
// escapes (`\,` `\:` `\=` `\\`), which leaves a valid RE2 pattern.
func entryRegex(e string) *regexp.Regexp {
if i := strings.IndexByte(e, ' '); i > 0 {
bin, args := e[:i], e[i+1:]
if strings.HasPrefix(args, "^") && strings.HasSuffix(args, "$") {
return regexp.MustCompile("^" + regexp.QuoteMeta(bin) + " " + args[1:])
}
}
return globToRegex(e)
}
// matchesAny reports whether cmdline matches at least one sudoers entry pattern.
func matchesAny(cmdline string, entries []string) bool {
for _, e := range entries {
if globToRegex(e).MatchString(cmdline) {
if entryRegex(e).MatchString(cmdline) {
return true
}
}
@@ -184,7 +200,8 @@ func TestRedProof_DroppedGrantFailsCheck(t *testing.T) {
}
// TestRedProof_DroppedControllerSwapTeeFailsCheck is the companion red-proof for the v0.45.0
// FELHOM_CONTROLLERSWAP grants: with the `tee /etc/felhom-controller-image` line removed, the
// FELHOM_CONTROLLERSWAP grants: with the write grant removed (since R-861 (a) A1 the `felhom-priv-apply controller-image`
// line; before it, an agent `tee /etc/felhom-controller-image`), the
// controllerswap-write capability MUST be reported uncovered. Proves the build gate watches the new
// swap write grant (so dropping it can't ship a non-root agent that silently can't auto-update).
func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
@@ -194,7 +211,7 @@ func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
}
var kept []string
for _, ln := range strings.Split(string(data), "\n") {
if strings.Contains(ln, "tee /etc/felhom-controller-image") {
if strings.Contains(ln, "felhom-priv-apply ^controller-image") { // R-861 (a) A1: the write's grant
continue
}
kept = append(kept, ln)
@@ -213,7 +230,7 @@ func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
}
cmdline := write.Binary + " " + strings.Join(write.ReprArgs, " ")
if matchesAny(cmdline, entries) {
t.Errorf("red-proof FAILED: controllerswap-write still matches after dropping the tee grant")
t.Errorf("red-proof FAILED: controllerswap-write still matches after dropping its grant")
}
if full := parseSudoersEntries(t, string(data)); !matchesAny(cmdline, full) {
t.Errorf("controllerswap-write should be covered by the real sudoers")
+138
View File
@@ -0,0 +1,138 @@
package capability
import (
"os"
"testing"
)
// R-861 (agent v0.146.0): the attacks the old globs let through, each as the exact argv a compromised agent would
// send. NONE may match any entry of the new sudoers. The same list ran against the REAL sudo 1.9.16 (a throwaway
// container, and `sudo -l -U felhom-agent` on both demo boxes after the bundle): audits/hub-safety-2026-10-05/partF/.
//
// RED-PROOF: run this list against the v0.145.0 sudoers (git show v0.145.0:configs/felhom-agent.sudoers) → most of
// these MATCH (recorded in the same evidence folder).
var r861Injections = []string{
// a raw host disk for a guest (pct options smuggled through a vmid glob)
"/usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1",
"/usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives",
"/usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda",
"/usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda",
// a bind mount over /etc through traversal
"/usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x",
"/usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d",
"/usr/bin/umount /mnt/felhom-drives/x /",
"/usr/bin/chown 100000:100000 /mnt/a/felhom-data /etc/shadow",
"/usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount",
// root-read files the agent writes: gone as `install` lines
"/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount",
"/usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh",
"/usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh",
"/usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf",
"/usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf",
"/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config",
// the unsigned binary flip
"/usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000",
// enabling or removing anything that is not ours
"/usr/bin/systemctl enable --now -- mnt-hdd_1.mount evil.service",
"/usr/bin/systemctl enable --now -- etc-sudoers.d.mount",
"/usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd",
"/usr/bin/rm -f /etc/dnsmasq.d/felhom-x.conf /etc/shadow",
"/usr/bin/rmdir /mnt/felhom-drives/x /etc",
// nftables commands chained after a set element
"/usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ; flush ruleset",
// extra options to read-only tools
"/usr/sbin/smartctl -a -j /dev/sda -s off",
"/usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x",
"/usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller",
"/usr/sbin/pct unlock 9201 --whatever",
// the checker with a path it must never take
"/usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount",
"/usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf",
"/usr/local/sbin/felhom-priv-apply wg /etc/shadow",
// R-861 (a) A1 (decision 165): the agent wrote ANY image ref into the guest by `tee` — now only the root verb may
"/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image",
"/usr/local/sbin/felhom-priv-apply controller-image 9201 9202",
"/usr/local/sbin/felhom-priv-apply controller-image 9201;id",
}
func TestSudoersRefusesTheR861Injections(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
for _, c := range r861Injections {
if matchesAny(c, entries) {
t.Errorf("the sudoers still allows: %s", c)
}
}
}
// R-444: the weekly trim's grant is ONE exact shape — `pct fstrim <vmid>` — and nothing smuggled after it.
// The manifest entry (guest-fstrim) proves the real call is still allowed (TestManifestCoveredBySudoers); this
// pins the other direction. RED-PROOF: write the rule as the glob `/usr/sbin/pct fstrim [0-9]*` → every decoy
// below with a trailing argument matches (the glob's `*` eats spaces).
func TestSudoersFstrimRuleIsExact(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
if !matchesAny("/usr/sbin/pct fstrim 9201", entries) {
t.Fatal("the sudoers does not allow `pct fstrim 9201` — the weekly trim cannot run")
}
for _, c := range []string{
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints",
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints 1",
"/usr/sbin/pct fstrim 9201; x",
"/usr/sbin/pct fstrim 9201 9202",
"/usr/sbin/pct fstrim 92a1",
"/usr/sbin/pct fstrim ",
"/usr/sbin/pct fstrim -- 9201",
"/usr/sbin/pct destroy 9201",
"/usr/sbin/pct destroy 9201 --purge",
} {
if matchesAny(c, entries) {
t.Errorf("the sudoers allows a command the trim rule must not: %q", c)
}
}
}
// R-861 (a) A1: the managed controller update still has its route — the root verb, one numeric vmid.
func TestSudoersAllowsTheControllerImageVerb(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
if !matchesAny("/usr/local/sbin/felhom-priv-apply controller-image 9201", parseSudoersEntries(t, string(data))) {
t.Fatal("the sudoers does not allow `felhom-priv-apply controller-image 9201` — a managed controller update cannot write its image")
}
}
// R-861 (b) B2 (decision 165, hygiene): felhom-op's `pct start|stop|unlock` grants are ONE numeric vmid each. The old
// glob `[0-9]*` eats spaces, so `pct stop 9201 --skiplock 1` and two vmids matched.
// RED-PROOF: on the pre-B2 felhom-op.sudoers (`/usr/sbin/pct stop [0-9]*`) the decoys match.
func TestFelhomOpSudoersPctIsExact(t *testing.T) {
data, err := os.ReadFile("../../configs/felhom-op.sudoers")
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
for _, ok := range []string{"/usr/sbin/pct start 9201", "/usr/sbin/pct stop 9201", "/usr/sbin/pct unlock 9201", "/usr/sbin/pct list"} {
if !matchesAny(ok, entries) {
t.Errorf("felhom-op lost a repair verb: %s", ok)
}
}
for _, bad := range []string{
"/usr/sbin/pct stop 9201 --skiplock 1",
"/usr/sbin/pct start 9201 9202",
"/usr/sbin/pct unlock 9201 --whatever",
"/usr/sbin/pct start 92a1",
"/usr/sbin/pct stop ",
"/usr/sbin/pct destroy 9201",
} {
if matchesAny(bad, entries) {
t.Errorf("felhom-op's sudoers allows %q", bad)
}
}
}
+31
View File
@@ -33,6 +33,7 @@ type Config struct {
LANResolver LANResolverConfig `json:"lan_resolver"`
WGTunnel WGTunnelConfig `json:"wg_tunnel"`
GuestNet GuestNetConfig `json:"guest_net"`
DiskTrim DiskTrimConfig `json:"disk_trim"`
OOB OOBConfig `json:"oob"`
SelfUpdate SelfUpdateConfig `json:"selfupdate"`
LogLevel string `json:"log_level"` // debug|info|warn|error (default info)
@@ -139,6 +140,16 @@ func (w WGTunnelConfig) WithDefaults() WGTunnelConfig {
return w
}
// DiskTrimConfig configures the R-444 weekly guest disk trim (internal/fstrim). DEFAULT-ON, like GuestNetConfig and for
// the same reason: it only acts on guests the agent already owns, and the operator ruled every box trims (`09` §3
// decision 139). Opting out is the explicit act: `"disk_trim": {"disable": true}`.
type DiskTrimConfig struct {
Disable bool `json:"disable"`
}
// Enabled reports whether the weekly trim should run.
func (d DiskTrimConfig) Enabled() bool { return !d.Disable }
// GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet).
//
// **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other
@@ -345,6 +356,12 @@ type BackupConfig struct {
// between an archive settling and its proof, and the retry rate of a tier whose restore-test
// keeps failing. See defaultRestoreTestEvalInterval for the measurement it was chosen from.
RestoreTestEvalIntervalSeconds int `json:"restore_test_eval_interval_seconds"`
// RestoreTestSpaceFactor / RestoreTestSpaceReserveGiB are the restore-test's space margin (R-672,
// v0.133.0): a test starts only when the target storage has free ≥ restored × factor + reserve,
// `restored` being the UNCOMPRESSED size. 0/unset → 1.2 and 5 GiB. A test config may raise them to
// watch the refusal (the brief's live case a).
RestoreTestSpaceFactor float64 `json:"restore_test_space_factor,omitempty"`
RestoreTestSpaceReserveGiB float64 `json:"restore_test_space_reserve_gib,omitempty"`
// RestoreTestSettleSeconds is how long an archive must have sat on its tier before it is a
// restore-test candidate (R-86); 0 → default (24h), negative → 0 (no settle requirement).
// Restore-testing an archive a backup is still writing proves nothing about the backup that
@@ -615,6 +632,20 @@ func (b BackupConfig) RestoreTestEvalInterval() time.Duration {
}
}
// RestoreTestSpace returns the restore-test's space margin (R-672): factor (≥ 1) and reserve bytes.
// Unset or out-of-range → 1.2 and 5 GiB.
func (b BackupConfig) RestoreTestSpace() (factor float64, reserveBytes int64) {
factor = b.RestoreTestSpaceFactor
if factor < 1 {
factor = 1.2
}
reserveBytes = int64(b.RestoreTestSpaceReserveGiB * float64(1<<30))
if reserveBytes <= 0 {
reserveBytes = 5 << 30
}
return factor, reserveBytes
}
// RestoreTestSettle returns how long an archive must have sat before it is a restore-test
// candidate (R-86): a positive value as-is, negative → 0 (no settle requirement), 0 → the default.
//
+48 -2
View File
@@ -7,9 +7,11 @@ import (
"encoding/hex"
"encoding/json"
"fmt"
"io"
"os"
"path/filepath"
"strings"
"syscall"
)
// Slice 10D.1 — IDENTITY escrow. The K-escrow (above) wraps the PBS *encryption key* via the
@@ -71,7 +73,7 @@ func HashResticPassword(pw string) string {
// (non-empty) but NEVER logged by callers — log the field NAME only (mirrors AttachWGKey). A missing file
// is a clean no-attach (pre-fork-4 behavior, byte-compatible bundle).
func AttachResticPassword(b *IdentityBundle, stagePath string) (bool, error) {
raw, err := os.ReadFile(stagePath)
raw, err := readStagedNoFollow(stagePath)
if err != nil {
if os.IsNotExist(err) {
return false, nil
@@ -92,7 +94,7 @@ func AttachResticPassword(b *IdentityBundle, stagePath string) (bool, error) {
// no-attach (pre-S3 behavior, byte-compatible bundle); a corrupt one is an error (the operator
// should know their escrow would silently lack a live identity).
func AttachWGKey(b *IdentityBundle, keyPath string) (bool, error) {
raw, err := os.ReadFile(keyPath)
raw, err := readStagedNoFollow(keyPath)
if err != nil {
if os.IsNotExist(err) {
return false, nil
@@ -180,3 +182,47 @@ func UnwrapIdentityBundle(ctx context.Context, blob []byte, recoveryCode string)
}
return b, nil
}
// readStagedNoFollow reads a file the AGENT staged, for the escrow ceremony that runs as ROOT (FELHOM_ESCROW). R-861
// (agent v0.146.0/0.146.1): both files live in the agent's own directory, so a compromised agent could put a SYMLINK
// there — at the file OR at any directory on the way (review 2026-10-05) — to a root-only file, and the root ceremony
// would seal that file into the blob and hand the agent R. So the path is walked from "/" one component at a time with
// openat(O_NOFOLLOW): no symlink anywhere, the last a regular file of at most 4 KiB. Once a directory is open, renaming
// it does not redirect the walk. A missing file keeps its os.IsNotExist meaning. Pinned by TestAttach_RefusesASymlink*.
func readStagedNoFollow(path string) ([]byte, error) {
if !filepath.IsAbs(path) {
return nil, fmt.Errorf("%s is not an absolute path", path)
}
clean := filepath.Clean(path)
parts := strings.Split(strings.TrimPrefix(clean, "/"), "/")
dirfd, err := syscall.Open("/", syscall.O_RDONLY|syscall.O_DIRECTORY|syscall.O_CLOEXEC, 0)
if err != nil {
return nil, err
}
for i, part := range parts {
last := i == len(parts)-1
flags := syscall.O_RDONLY | syscall.O_NOFOLLOW | syscall.O_CLOEXEC
if !last {
flags |= syscall.O_DIRECTORY
}
fd, err := syscall.Openat(dirfd, part, flags, 0)
syscall.Close(dirfd)
if err != nil {
return nil, &os.PathError{Op: "open", Path: clean, Err: err}
}
dirfd = fd
}
f := os.NewFile(uintptr(dirfd), clean)
defer f.Close()
fi, err := f.Stat()
if err != nil {
return nil, err
}
if !fi.Mode().IsRegular() {
return nil, fmt.Errorf("%s is not a regular file", path)
}
if fi.Size() > 4096 {
return nil, fmt.Errorf("%s is larger than 4 KiB", path)
}
return io.ReadAll(io.LimitReader(f, 4097))
}
+63
View File
@@ -0,0 +1,63 @@
package escrow
import (
"os"
"path/filepath"
"testing"
)
// R-861 (agent v0.146.0): the root escrow ceremony reads two files from the AGENT's directory. A symlink there to a
// root-only file must never be read (it would be sealed under R and handed to the agent).
// RED-PROOF (audits/hub-safety-2026-10-05/partF/red-proof.txt): use os.ReadFile in readStagedNoFollow → this fails.
func TestAttach_RefusesASymlink(t *testing.T) {
d := t.TempDir()
secret := filepath.Join(d, "root-only")
if err := os.WriteFile(secret, []byte("ROOT-ONLY-CANARY"), 0o600); err != nil {
t.Fatal(err)
}
link := filepath.Join(d, "restic_repo_password")
if err := os.Symlink(secret, link); err != nil {
t.Fatal(err)
}
var b IdentityBundle
if ok, err := AttachResticPassword(&b, link); err == nil || ok || b.ResticRepoPassword != "" {
t.Fatalf("a symlinked staged file was read: ok=%v err=%v value-set=%v", ok, err, b.ResticRepoPassword != "")
}
if ok, err := AttachWGKey(&b, link); err == nil || ok {
t.Fatalf("a symlinked wg key was read: ok=%v err=%v", ok, err)
}
// control: the real staged file is still read; a missing one is still a clean no-attach
real := filepath.Join(d, "real")
_ = os.WriteFile(real, []byte("pw\n"), 0o600)
if ok, err := AttachResticPassword(&b, real); err != nil || !ok || b.ResticRepoPassword != "pw" {
t.Fatalf("control: the plain staged file was not read: %v %v", ok, err)
}
if ok, err := AttachResticPassword(&b, filepath.Join(d, "absent")); err != nil || ok {
t.Fatalf("control: a missing file must stay a clean no-attach: %v %v", ok, err)
}
}
// Review 2026-10-05: a symlinked DIRECTORY on the way must stop the read too (O_NOFOLLOW alone guards only the last
// component). RED-PROOF: open the full path with O_NOFOLLOW only → this fails.
func TestAttach_RefusesASymlinkedDirectory(t *testing.T) {
d := t.TempDir()
secretDir := filepath.Join(d, "root-only-dir")
_ = os.Mkdir(secretDir, 0o700)
_ = os.WriteFile(filepath.Join(secretDir, "private.key"), []byte("AAECAwQFBgcICQoLDA0ODxAREhMUFRYXGBkaGxwdHh8=\n"), 0o600)
agentDir := filepath.Join(d, "agent")
_ = os.Mkdir(agentDir, 0o700)
if err := os.Symlink(secretDir, filepath.Join(agentDir, "wg")); err != nil {
t.Fatal(err)
}
var b IdentityBundle
if ok, err := AttachWGKey(&b, filepath.Join(agentDir, "wg", "private.key")); err == nil || ok || b.WGPrivateKey != "" {
t.Fatalf("a key behind a symlinked directory was read: ok=%v err=%v", ok, err)
}
// control: the same key under a REAL directory is read
_ = os.Remove(filepath.Join(agentDir, "wg"))
_ = os.Mkdir(filepath.Join(agentDir, "wg"), 0o700)
_ = os.WriteFile(filepath.Join(agentDir, "wg", "private.key"), []byte("AAECAwQFBgcICQoLDA0ODxAREhMUFRYXGBkaGxwdHh8=\n"), 0o600)
if ok, err := AttachWGKey(&b, filepath.Join(agentDir, "wg", "private.key")); err != nil || !ok {
t.Fatalf("control: a key under a real directory was not read: %v %v", ok, err)
}
}
+210
View File
@@ -0,0 +1,210 @@
package escrow
import (
"context"
"errors"
"fmt"
)
// R-199 links 6→8 — fetch this host's own sealed identity blob, open it with the customer's recovery
// code R, and hand back EXACTLY ONE field: the offsite restic repository password.
//
// WHY ONLY ONE FIELD. The bundle also carries the Cloudflare tunnel token, the PBS access token and
// the WG private key (see IdentityBundle). The caller in this flow — the in-guest controller, one
// trust tier down — needs none of them, and returning them would widen the blast radius of a
// controller compromise for no gain. Narrowing costs nothing here and is not recoverable later.
//
// WHY R NEVER TOUCHES DISK. `UnwrapIdentity` stages the BLOB and the recovered plaintext in a
// `MkdirTemp` that it removes, and feeds R through the pty; R itself is never written. This wrapper
// keeps that property: it takes R as an argument, passes it straight through, and holds no copy.
// Callers must clear their own reference (the `R = ""` discipline in cmd/felhom-agent).
//
// The errors below are DISTINCT on purpose. "could not fetch", "no blob", "wrong code" and "the blob
// predates the field" are FOUR different situations for the operator and only one of them is a fault.
//
// ⚠ THERE WERE THREE, AND THE FOURTH WAS THE DEFECT (R-224, 2026-08-06). This comment said "three"
// and named "no blob", "wrong code" and "predates the field" — while a FAILED FETCH was wrapped as an
// anonymous error and fell through the caller's `default` branch into the wrong-code message. So a
// hub that could not be reached was reported to the customer as a bad recovery code.
//
// Measured live on 2026-08-05 (CAMPAIGN-11 F3): with the hub REJECTed at the appliance's firewall and
// a CORRECT current recovery code, the customer was told the code did not open their package — in
// 0.0556 s, when a real unseal costs ~1 s of scrypt. The agent's own log carried the truth the whole
// time (`escrow: fetching the sealed bundle: hub: transport error: … no route to host`) and the HTTP
// boundary threw it away.
//
// The discriminator therefore has to be a VALUE, not a log line — that is what ErrBundleFetch is.
var (
// ErrBundleFetch — the sealed bundle could not be FETCHED (the hub refused, was unreachable, or
// the transport failed). **The recovery code was never used**, so nothing about it is known and
// nothing may be said about it. Wraps the underlying cause for the operator log; carries no secret.
ErrBundleFetch = errors.New("escrow: the sealed bundle could not be fetched")
// ErrNoEscrowBlob — the hub holds no sealed bundle for this host. Not a fault: no ceremony has run.
ErrNoEscrowBlob = errors.New("escrow: the hub holds no sealed identity bundle for this host (no ceremony has run)")
// ErrNoResticPassword — the bundle opened, but carries no repository password. Real and expected
// for a pre-fork-4 blob (agent < v0.77.0, 2026-07-09): the field did not exist and CANNOT be
// retro-fitted, because R is never retained. Distinguished from a wrong code so the operator is
// not sent hunting for a mistyped recovery code that was typed correctly.
ErrNoResticPassword = errors.New("escrow: the recovered bundle carries NO offsite repository password (a pre-fork-4 blob — the field did not exist when it was sealed and cannot be retro-fitted)")
// ErrCodeOpensRetained — the code did NOT open the package the hub currently holds, and DID open a
// RETAINED (earlier) one. R-311.
//
// ⚠ THIS IS NOT A FAILURE OF THE CUSTOMER'S. It is the single most important distinction on this
// path, because until 2026-08-12 it was indistinguishable from a mistype and was reported as one.
// The screen could only say "it may be a typo, or it may be an older code, and we cannot tell them
// apart from here" — and it could not tell them apart because NOTHING EVER LOOKED. Now something
// looks, so the sentence can stop hedging.
//
// It carries no material and no code: only WHICH earlier package opened, by its supersession date,
// which is the one fact the customer needs to recognise it.
ErrCodeOpensRetained = errors.New("escrow: the recovery code did not open the CURRENT sealed package, but it DID open a retained earlier one")
)
// RetainedMatch says which retained package a code opened. Returned inside RetainedOpenedError; it
// carries no secret — not the code, not the bundle, not the repository password.
type RetainedMatch struct {
// SupersededAt is when this package stopped being the current one (hub-supplied, RFC3339-ish).
// It is what the recovery screen shows so the customer can recognise which code they are holding.
SupersededAt string
// KeyFingerprint is the escrow key fingerprint of that package — operator-log material only.
KeyFingerprint string
// Index is the hub's position label within ONE response. Not durable; do not persist it.
Index int
// HasResticPassword is false when the retained package opened but carries no repository password
// (a pre-fork-4 seal). The code is still CORRECT; the history behind it still cannot be reopened.
// Collapsing this into "recoverable" would repeat R-202's mistake on a new surface.
HasResticPassword bool
}
// RetainedOpenedError wraps ErrCodeOpensRetained with the match. Callers classify with errors.Is on
// the sentinel and read the detail with errors.As.
type RetainedOpenedError struct {
Match RetainedMatch
}
func (e *RetainedOpenedError) Error() string {
return ErrCodeOpensRetained.Error() + " (superseded_at=" + e.Match.SupersededAt + ")"
}
func (e *RetainedOpenedError) Unwrap() error { return ErrCodeOpensRetained }
// BlobFetcher yields this host's own opaque identity-escrow blob. present=false is a clean "none".
// An interface-free func field keeps this package free of any dependency on the hub client.
type BlobFetcher func(ctx context.Context) (blob []byte, present bool, err error)
// RetainedBlob is one retained sealed package as the recoverer sees it: opaque bytes plus the labels
// needed to name it. No secret.
type RetainedBlob struct {
Blob []byte
SupersededAt string
KeyFingerprint string
Index int
}
// RetainedFetcher yields this host's RETAINED sealed packages, newest-superseded first. An empty
// slice is a clean "none". R-311.
type RetainedFetcher func(ctx context.Context) (blobs []RetainedBlob, unopenable int, err error)
// OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R.
type OffsiteKeyRecoverer struct {
Fetch BlobFetcher
// FetchRetained is OPTIONAL and consulted ONLY after the current package has refused the code.
// nil keeps the pre-R-311 behaviour exactly: a refusal stays a refusal. That is deliberate — an
// agent wired without it must not behave differently from one that has no retained packages.
FetchRetained RetainedFetcher
// MaxRetainedTried bounds the scrypt work a single wrong code can cost. Each attempt is ~1 s of
// KDF by design, so an unbounded loop over a long supersession history would turn one wrong code
// into a minutes-long hang on the customer's screen. 0 means the built-in default.
MaxRetainedTried int
}
// defaultMaxRetainedTried — six attempts is ~6 s worst case, which is a slow screen and not a hang.
const defaultMaxRetainedTried = 6
// RecoverOffsiteRepoPassword fetches, unseals and extracts. It returns ONLY the repository password.
//
// A WRONG RECOVERY CODE FAILS CLOSED at the scrypt KDF inside UnwrapIdentity — `age -d` exits
// non-zero and emits no plaintext, so there is no partial result and nothing is written anywhere.
// That property is the crypto's, not a check here, which is why this function has no "validate R"
// step to get wrong.
//
// NOTHING IS LOGGED BY THIS FUNCTION and no error it returns contains R, the password, or blob bytes.
func (r OffsiteKeyRecoverer) RecoverOffsiteRepoPassword(ctx context.Context, recoveryCode string) (string, error) {
if r.Fetch == nil {
return "", fmt.Errorf("escrow: recoverer has no blob fetcher configured")
}
if recoveryCode == "" {
return "", fmt.Errorf("escrow: the recovery code is required")
}
blob, present, err := r.Fetch(ctx)
if err != nil {
// R-224: joined with ErrBundleFetch so the caller can classify by VALUE. The cause stays
// wrapped for the operator log; neither carries a secret. Before this, the fetch failure was
// an anonymous error and the local-api handler's `default` branch reported it to the customer
// as a wrong recovery code.
return "", fmt.Errorf("%w: %w", ErrBundleFetch, err)
}
if !present || len(blob) == 0 {
return "", ErrNoEscrowBlob
}
bundle, err := UnwrapIdentityBundle(ctx, blob, recoveryCode)
if err != nil {
// R-311 — BEFORE calling this a wrong code, ask whether it is the RIGHT code for an EARLIER
// package. The engine fails closed identically either way, so the two are indistinguishable
// from the unwrap alone; the only way to tell is to try. Until this existed nobody tried, and
// the screen said so out loud ("innen nem tudjuk megkülönböztetni őket") — a true sentence
// about our own incuriosity, read by the customer as a statement about their code.
if m, ok := r.tryRetained(ctx, recoveryCode); ok {
return "", &RetainedOpenedError{Match: m}
}
return "", err // the fail-closed "the recovery code did not unwrap…" message; no secret in it
}
if bundle.ResticRepoPassword == "" {
return "", ErrNoResticPassword
}
return bundle.ResticRepoPassword, nil
}
// tryRetained reports whether the code opens one of this host's RETAINED packages, and which.
//
// FAILURE HERE IS SILENT AND MEANS "NO", NEVER "YES" and never a different verdict for the caller. A
// hub that cannot answer, a route an older hub does not have, a malformed blob — each leaves the
// original refusal standing, unchanged. That is the fail-safe direction: the worst outcome of this
// function breaking is the behaviour we had before it existed.
//
// NOTHING IS LOGGED HERE and no return value carries the code, a bundle or a password.
func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode string) (RetainedMatch, bool) {
if r.FetchRetained == nil {
return RetainedMatch{}, false
}
blobs, _, err := r.FetchRetained(ctx)
if err != nil || len(blobs) == 0 {
return RetainedMatch{}, false
}
limit := r.MaxRetainedTried
if limit <= 0 {
limit = defaultMaxRetainedTried
}
for i, rb := range blobs {
if i >= limit {
break
}
if len(rb.Blob) == 0 {
continue
}
bundle, uerr := UnwrapIdentityBundle(ctx, rb.Blob, recoveryCode)
if uerr != nil {
continue // this one is not the customer's; try the next
}
return RetainedMatch{
SupersededAt: rb.SupersededAt,
KeyFingerprint: rb.KeyFingerprint,
Index: rb.Index,
// A retained package can itself predate the repository-password field. The code is still
// correct and must be told so — but the history behind it still cannot be reopened, and
// saying otherwise would be a promise this path cannot keep.
HasResticPassword: bundle.ResticRepoPassword != "",
}, true
}
return RetainedMatch{}, false
}
+230
View File
@@ -0,0 +1,230 @@
package escrow
import (
"context"
"errors"
"fmt"
"testing"
)
// R-311 — a correct code for an EARLIER package must stop being reported as a wrong code.
//
// These use REAL age crypto, like the R-199 tests beside them, because the whole point is that the
// two situations are indistinguishable AT THE UNWRAP: both fail closed on the current package. A
// faked unwrap would prove nothing about the thing that was actually broken.
const testR2 = "another correct horse battery staple sedative anaconda wobbly kingdom placard"
func retainedFetcherFor(blobs ...RetainedBlob) RetainedFetcher {
return func(context.Context) ([]RetainedBlob, int, error) { return blobs, 0, nil }
}
// THE ONE THAT MATTERS. The customer holds the code for a package we superseded. Yesterday this
// returned the fail-closed refusal and the screen told them to check their typing.
//
// RED-PROOF: remove the `if m, ok := r.tryRetained(...)` block from RecoverOffsiteRepoPassword →
// the wrong-code error returns instead → this FAILS, and the lie is back in exactly those words.
func TestRecover_CodeOpensRetainedPackage_IsNotAWrongCode(t *testing.T) {
ensureAge(t)
const oldPW = "aaaa567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "cccc567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
retained := sealBundle(t, IdentityBundle{ResticRepoPassword: oldPW}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{
Blob: retained, SupersededAt: "2026-08-12 15:18:55", KeyFingerprint: "7e:a6:af", Index: 0,
}),
}.RecoverOffsiteRepoPassword(context.Background(), testR) // the OLD code
if err == nil {
t.Fatal("recovery succeeded — it must NOT return a password for a retained package on this path")
}
if !errors.Is(err, ErrCodeOpensRetained) {
t.Fatalf("err = %v, want ErrCodeOpensRetained — a correct code for an earlier package was "+
"classified as something else, which is how it became 'check your typing'", err)
}
var ro *RetainedOpenedError
if !errors.As(err, &ro) {
t.Fatalf("err does not carry a RetainedOpenedError: %v", err)
}
if ro.Match.SupersededAt != "2026-08-12 15:18:55" {
t.Errorf("SupersededAt = %q — the screen needs this date to name the package", ro.Match.SupersededAt)
}
if !ro.Match.HasResticPassword {
t.Error("HasResticPassword = false, but the retained bundle carried one")
}
// The error must not leak the code, the password or the bundle.
for _, secret := range []string{testR, oldPW} {
if containsStr(err.Error(), secret) {
t.Fatalf("the error text leaks a secret")
}
}
}
// SCENARIO A — the ordinary recovery is untouched, and it must not even ASK for retained packages.
// If the current package opens, the customer is not in this story at all.
//
// RED-PROOF: move the tryRetained call above the successful-unwrap return → the fetcher runs → this
// FAILS on the "must not be consulted" assertion.
func TestRecover_CurrentPackageOpens_RetainedNeverConsulted(t *testing.T) {
ensureAge(t)
const pw = "bbbb567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
current := sealBundle(t, IdentityBundle{ResticRepoPassword: pw}, testR)
consulted := false
got, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
consulted = true
return nil, 0, nil
},
}.RecoverOffsiteRepoPassword(context.Background(), testR)
if err != nil {
t.Fatalf("the ordinary recovery broke: %v", err)
}
if got != pw {
t.Fatalf("recovered password is not the sealed one")
}
if consulted {
t.Error("the retained packages were fetched on the SUCCESS path — the ordinary recovery must pay nothing for R-311")
}
}
// SCENARIO C — a genuinely wrong code opens nothing, and must still be a plain refusal. The new
// branch must not become a way to encourage a customer who mistyped.
//
// RED-PROOF: make tryRetained return (RetainedMatch{}, true) unconditionally → a wrong code is
// reported as opening an earlier package → this FAILS.
func TestRecover_WrongCode_StaysAPlainRefusal(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "cccc567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
retained := sealBundle(t, IdentityBundle{ResticRepoPassword: "dddd567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: retained, SupersededAt: "2026-08-01 00:00:00"}),
}.RecoverOffsiteRepoPassword(context.Background(), "totally wrong words that open nothing at all here")
if err == nil {
t.Fatal("a wrong code succeeded")
}
if errors.Is(err, ErrCodeOpensRetained) {
t.Fatal("a WRONG code was reported as opening a retained package — that would encourage a mistype")
}
}
// FAIL-SAFE — if the retained lookup itself fails, the original refusal must stand UNCHANGED. The
// worst outcome of this feature breaking is the behaviour we had before it.
//
// RED-PROOF: make tryRetained propagate the fetch error instead of returning false → the customer
// gets a new, unexplained failure mode → this FAILS.
func TestRecover_RetainedFetchFails_OriginalRefusalStands(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "eeee567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
return nil, 0, fmt.Errorf("hub exploded")
},
}.RecoverOffsiteRepoPassword(context.Background(), testR2)
if err == nil {
t.Fatal("expected a refusal")
}
if errors.Is(err, ErrCodeOpensRetained) {
t.Fatal("a failed retained lookup was reported as 'opens a retained package'")
}
if containsStr(err.Error(), "hub exploded") {
t.Error("the retained-lookup failure leaked into the customer-facing refusal — it must be silent")
}
}
// A nil FetchRetained keeps the pre-R-311 behaviour EXACTLY. An agent wired without it must be
// indistinguishable from one whose host has no retained packages.
//
// RED-PROOF: remove the `if r.FetchRetained == nil` guard → nil-deref panic → this FAILS.
func TestRecover_NilRetainedFetcher_IsPreR311Behaviour(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "ffff567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
_, err := OffsiteKeyRecoverer{Fetch: fetcherFor(current)}.RecoverOffsiteRepoPassword(context.Background(), testR2)
if err == nil {
t.Fatal("expected a refusal")
}
if errors.Is(err, ErrCodeOpensRetained) {
t.Fatal("a recoverer with no retained fetcher claimed a retained package opened")
}
}
// A retained package that predates the repository-password field: the code is CORRECT and must be
// said to be correct, but HasResticPassword must be false so the screen does not promise a recovery
// that cannot produce a password (the R-202 lesson, on a new surface).
//
// RED-PROOF: hardcode HasResticPassword: true → this FAILS.
func TestRecover_RetainedOpensButPredatesTheField(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "1111567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
// No ResticRepoPassword at all — the pre-fork-4 shape.
retained := sealBundle(t, IdentityBundle{TunnelToken: "T", PBSToken: "P"}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: retained, SupersededAt: "2026-08-04 07:20:08"}),
}.RecoverOffsiteRepoPassword(context.Background(), testR)
if !errors.Is(err, ErrCodeOpensRetained) {
t.Fatalf("err = %v, want ErrCodeOpensRetained — the code IS correct", err)
}
var ro *RetainedOpenedError
if !errors.As(err, &ro) {
t.Fatalf("no RetainedOpenedError: %v", err)
}
if ro.Match.HasResticPassword {
t.Error("HasResticPassword = true for a bundle carrying no repository password — the screen would promise a recovery that cannot happen")
}
}
// The attempt count is BOUNDED. Each unwrap is ~1 s of scrypt by design, so an unbounded loop turns
// one wrong code into a minutes-long hang on the customer's screen.
//
// RED-PROOF: remove the `if i >= limit { break }` → all 10 are tried → this FAILS on the count.
func TestRecover_RetainedAttemptsAreBounded(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "2222567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
junk := sealBundle(t, IdentityBundle{ResticRepoPassword: "3333567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
tried := 0
blobs := make([]RetainedBlob, 0, 10)
for i := 0; i < 10; i++ {
blobs = append(blobs, RetainedBlob{Blob: junk, SupersededAt: "2026-08-01 00:00:00", Index: i})
}
rec := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
tried++
return blobs, 0, nil
},
MaxRetainedTried: 2,
}
// A code that opens NEITHER the current package nor any retained one.
if _, err := rec.RecoverOffsiteRepoPassword(context.Background(), "a code that opens nothing whatsoever in this test"); err == nil {
t.Fatal("expected a refusal")
}
if tried != 1 {
t.Errorf("the retained list was fetched %d times, want exactly 1", tried)
}
}
func containsStr(hay, needle string) bool {
return len(needle) > 0 && len(hay) >= len(needle) && (func() bool {
for i := 0; i+len(needle) <= len(hay); i++ {
if hay[i:i+len(needle)] == needle {
return true
}
}
return false
})()
}
+267
View File
@@ -0,0 +1,267 @@
package escrow
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
)
// R-199 links 6→8, with REAL crypto (age is present on the build/demo host; ensureAge skips
// elsewhere). These are the unit half of the session's question — "is the repository password
// actually recoverable from the sealed bundle" — and the live half is the same equality on hardware.
const testR = "correct horse battery staple sedative anaconda wobbly kingdom placard yodel"
func sealBundle(t *testing.T, b IdentityBundle, r string) []byte {
t.Helper()
blob, err := WrapIdentityBundle(context.Background(), b, r)
if err != nil {
t.Fatalf("WrapIdentityBundle: %v", err)
}
return blob
}
func fetcherFor(blob []byte) BlobFetcher {
return func(context.Context) ([]byte, bool, error) { return blob, true, nil }
}
// Scenario A (unit) — the recovered repository password is BYTE-IDENTICAL to the sealed one, and it
// is the REPOSITORY password rather than some other field of a bundle that also parses.
//
// RED-PROOF: return bundle.PBSToken (or TunnelToken, or WGPrivateKey) instead of
// bundle.ResticRepoPassword → a plausible-looking bundle yields a non-matching key → this FAILS.
// That mutation is the shape of the bug that would otherwise ship silently, because every one of
// those fields is a non-empty string that looks like a secret.
func TestRecoverOffsiteRepoPassword_ReturnsTheRepositoryPassword(t *testing.T) {
ensureAge(t)
const repoPW = "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
blob := sealBundle(t, IdentityBundle{
TunnelToken: "TUNNEL-TOKEN-NOT-THE-ANSWER",
PBSToken: "PBS-TOKEN-NOT-THE-ANSWER",
WGPrivateKey: "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA=",
ResticRepoPassword: repoPW,
}, testR)
got, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), testR)
if err != nil {
t.Fatalf("recover: %v", err)
}
if got != repoPW {
t.Fatalf("the recovered key is not the sealed repository password (len %d vs %d) — a different "+
"field of the bundle was returned", len(got), len(repoPW))
}
// Belt: it must not be any of the OTHER fields, so a future refactor cannot satisfy the check
// above by coincidence.
for _, other := range []string{"TUNNEL-TOKEN-NOT-THE-ANSWER", "PBS-TOKEN-NOT-THE-ANSWER", "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA="} {
if got == other {
t.Fatalf("the recoverer returned the wrong bundle field")
}
}
}
// Scenario B — a WRONG recovery code fails closed, the failure names no secret, and nothing is
// written. The fail-closed property is the crypto's (age's scrypt KDF), which is why there is no
// validation step here to get wrong — the test pins that it stays that way.
func TestRecoverOffsiteRepoPassword_WrongCodeFailsClosed(t *testing.T) {
ensureAge(t)
const repoPW = "ffffffffffffffffffffffffffffffffffffffffffffffffffffffffffffffff"
blob := sealBundle(t, IdentityBundle{TunnelToken: "t", PBSToken: "p", ResticRepoPassword: repoPW}, testR)
got, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), "not the recovery code at all")
if err == nil {
t.Fatal("a wrong recovery code MUST fail — a plausible-but-wrong bundle is the one outcome the design forbids")
}
if got != "" {
t.Fatalf("a failed unseal returned %d bytes — there must be no partial result", len(got))
}
// The error may name the step; it may never name a secret.
for _, secret := range []string{repoPW, testR, "not the recovery code at all"} {
if strings.Contains(err.Error(), secret) {
t.Fatalf("the failure message leaked a secret: %v", err)
}
}
}
// A bundle with no repository password is its OWN answer, not a wrong-code error. Sealed before
// fork-4 (agent < v0.77.0) the field did not exist; sending the operator to re-check a correctly
// typed recovery code would be the wrong instruction.
func TestRecoverOffsiteRepoPassword_PreForkFourBundle(t *testing.T) {
ensureAge(t)
blob := sealBundle(t, IdentityBundle{TunnelToken: "t", PBSToken: "p"}, testR)
_, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), testR)
if !errors.Is(err, ErrNoResticPassword) {
t.Fatalf("a pre-fork-4 bundle must report its own error, got %v", err)
}
}
// Scenario D at this layer — no blob is a clean, distinguishable answer.
func TestRecoverOffsiteRepoPassword_NoBlob(t *testing.T) {
rec := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) { return nil, false, nil }}
_, err := rec.RecoverOffsiteRepoPassword(context.Background(), testR)
if !errors.Is(err, ErrNoEscrowBlob) {
t.Fatalf("absent blob must yield ErrNoEscrowBlob, got %v", err)
}
}
// Scenario F — R persists NOWHERE. TMPDIR is redirected into the test's own directory, the unseal is
// run for real, and the whole tree is then walked: no file may contain R (or the recovered password),
// and the staging directory the unseal creates must be gone.
//
// RED-PROOF: write R to a temp file anywhere in the flow (e.g. add
// `os.WriteFile(filepath.Join(work,"r"), []byte(recoveryCode), 0o600)` inside UnwrapIdentity before
// its defer removes the dir — or simply drop that defer and let the plaintext staging survive) → the
// walk finds it → this FAILS.
func TestRecoverOffsiteRepoPassword_RLeavesNoTrace(t *testing.T) {
ensureAge(t)
const repoPW = "1111111111111111111111111111111111111111111111111111111111111111"
tmp := t.TempDir()
t.Setenv("TMPDIR", tmp) // os.MkdirTemp honours this — every staging dir lands under the walk
const wrongR = "wrong code entirely"
blob := sealBundle(t, IdentityBundle{TunnelToken: "t", PBSToken: "p", ResticRepoPassword: repoPW}, testR)
if _, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), testR); err != nil {
t.Fatalf("recover: %v", err)
}
// A failed unseal must leave nothing either — exercise both paths before walking.
_, _ = (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), wrongR)
// THE PRIMARY ASSERTION IS EMPTINESS, not content. A content scan alone is defeatable by a later
// call OVERWRITING the leaked file with a different secret — which is exactly how the first
// version of this test passed its own red-proof while R sat on disk. Nothing in this test writes
// under TMPDIR, so after both calls the tree must contain no files at all.
var survivors []string
err := filepath.Walk(tmp, func(path string, info os.FileInfo, err error) error {
if err != nil || info == nil || info.IsDir() || path == tmp {
return nil
}
survivors = append(survivors, strings.TrimPrefix(path, tmp))
return nil
})
if err != nil {
t.Fatal(err)
}
if len(survivors) > 0 {
t.Fatalf("the unseal left %d file(s) behind under TMPDIR: %v — R, the sealed blob and the "+
"recovered plaintext all pass through there and none of them may outlive the call", len(survivors), survivors)
}
// Defence in depth: any secret that DOES appear anywhere is named, for every code used.
_ = filepath.Walk(tmp, func(path string, info os.FileInfo, err error) error {
if err != nil || info == nil || info.IsDir() {
return nil
}
body, rerr := os.ReadFile(path)
if rerr != nil {
return nil
}
for label, secret := range map[string]string{"R": testR, "a wrong R": wrongR, "the repository password": repoPW} {
if strings.Contains(string(body), secret) {
t.Errorf("%s survived on disk at %s", label, path)
}
}
return nil
})
// And the staging directories are gone, not merely free of secrets.
entries, _ := os.ReadDir(tmp)
for _, e := range entries {
if e.IsDir() && strings.HasPrefix(e.Name(), "felhom-idesc-") {
t.Fatalf("an unseal staging directory survived: %s", e.Name())
}
}
}
// A fetch failure surfaces as a fetch failure, not as a wrong-code error — the operator must not be
// sent to re-read their recovery code because the hub was unreachable.
//
// ⚠ THIS TEST WAS GREEN THROUGHOUT THE DEFECT IT DESCRIBES (R-224, 2026-08-06). Its sentence is
// exactly right and it did not prevent anything, for two reasons worth keeping:
//
// 1. **It asserted the MECHANISM, one layer below the consequence.** It checked this package's error
// STRING. The merge happened one layer up, in the local-api handler's `default` branch, which
// answered a fetch failure with "the recovery code did not open the sealed bundle". The customer
// never sees this string; they see that one. The project's own rule — prefer the test that asserts
// the CONSEQUENCE (does the customer get blamed?) over the one that asserts the MECHANISM (is the
// error distinct here?) — names this case precisely.
// 2. **It asserted on TEXT.** `strings.Contains(err.Error(), …)` cannot be consumed by a caller, so
// it pinned something no production code could branch on. The distinction it checked was real and
// unusable.
//
// It now asserts the SENTINEL, which is what the handler branches on, and its consequence-level twin
// lives in `internal/localapi/escrow_recover_class_test.go` where the status is asserted.
func TestRecoverOffsiteRepoPassword_FetchErrorIsDistinct(t *testing.T) {
rec := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) {
return nil, false, errors.New("hub: connection refused")
}}
_, err := rec.RecoverOffsiteRepoPassword(context.Background(), testR)
if err == nil || !errors.Is(err, ErrBundleFetch) {
t.Fatalf("a fetch failure must classify as ErrBundleFetch, got %v", err)
}
if errors.Is(err, ErrNoEscrowBlob) || errors.Is(err, ErrNoResticPassword) {
t.Fatal("a transport failure must not masquerade as a content verdict")
}
}
// ── R-224 — A FAILED FETCH IS NOT A WRONG CODE ──────────────────────────────────────────────────
//
// CAMPAIGN-11 F3 measured the consequence of these two being indistinguishable: with the hub
// firewalled off and a CORRECT current recovery code, the customer was told the code did not open
// their package, in 0.0556 s — no unseal was attempted at all.
//
// The pair below is the whole point. Asserting only the first would pass with a `return ErrBundleFetch`
// stuck on every error path, which is the same defect pointing the other way.
func TestRecoverOffsiteRepoPassword_FetchFailureIsClassifiedAsFetch(t *testing.T) {
boom := errors.New("hub: transport error: dial tcp 37.191.56.193:443: connect: no route to host")
r := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) { return nil, false, boom }}
_, err := r.RecoverOffsiteRepoPassword(context.Background(), testR)
if err == nil {
t.Fatal("a failing fetch must return an error")
}
// RED-PROOF: drop the `%w: %w` join in RecoverOffsiteRepoPassword (return the bare wrapped cause,
// as it was before R-224) → this FAILS, and the local-api handler falls back to the wrong-code
// message exactly as it did on 2026-08-05.
if !errors.Is(err, ErrBundleFetch) {
t.Fatalf("a failed fetch must classify as ErrBundleFetch, got %v", err)
}
// The underlying cause survives for the operator log.
if !errors.Is(err, boom) {
t.Fatalf("the fetch cause must stay wrapped for the operator, got %v", err)
}
// And it must NOT be mistaken for either of the bundle-content situations.
if errors.Is(err, ErrNoEscrowBlob) || errors.Is(err, ErrNoResticPassword) {
t.Fatalf("a transport failure is neither of the bundle-content errors: %v", err)
}
}
// The other half: a genuinely wrong code must NOT classify as a fetch failure, or the fix trades one
// misattribution for its mirror image and the customer is told the hub is down when they mistyped.
func TestRecoverOffsiteRepoPassword_WrongCodeIsNotAFetchFailure(t *testing.T) {
ensureAge(t)
blob := sealBundle(t, IdentityBundle{ResticRepoPassword: "0123456789abcdef"}, testR)
r := OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}
_, err := r.RecoverOffsiteRepoPassword(context.Background(),
"wrong horse battery staple sedative anaconda wobbly kingdom placard yodel")
if err == nil {
t.Fatal("a wrong recovery code must fail closed")
}
if errors.Is(err, ErrBundleFetch) {
t.Fatalf("a wrong code must NOT classify as a fetch failure, got %v", err)
}
}
// A clean "the hub holds nothing" keeps its own identity too — it is not a fetch failure, and the
// customer must not be told the hub was unreachable when it answered perfectly well.
func TestRecoverOffsiteRepoPassword_AbsentBlobIsNotAFetchFailure(t *testing.T) {
r := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) { return nil, false, nil }}
_, err := r.RecoverOffsiteRepoPassword(context.Background(), testR)
if !errors.Is(err, ErrNoEscrowBlob) {
t.Fatalf("an absent blob must stay ErrNoEscrowBlob, got %v", err)
}
if errors.Is(err, ErrBundleFetch) {
t.Fatalf("an absent blob is not a fetch FAILURE, got %v", err)
}
}
+2
View File
@@ -27,6 +27,8 @@ const (
PortFile = ConfDir + "/port"
// PidFile is the instance pidfile (NOT a RuntimeDirectory — that is the G1 incident cause).
PidFile = "/run/felhom-sshd.pid"
// PrivApply is the root content checker that installs the config and felhom-op's key (R-861, agent v0.146.0).
PrivApply = "/usr/local/sbin/felhom-priv-apply"
// Unit is the systemd unit name.
Unit = "felhom-sshd"
// OperatorUser is the default operator login (scoped sudo; key in AuthKeysDir only).
+5 -2
View File
@@ -143,7 +143,8 @@ func (m *Manager) Apply(ctx context.Context, block *hub.WireWireguard) (int, err
m.logger.Error("felhomsshd: staged config failed sshd -t — NOT installing", "err", err, "stderr", strings.TrimSpace(string(errOut)))
return port, err
}
if _, errOut, err := m.runner.Run(ctx, "install", "-o", "root", "-g", "root", "-m", "0644", "--", m.stagedConfPath(), ConfPath); err != nil {
// R-861 (v0.146.0): the root checker installs it, and only if it is renderConfig's template for some port.
if _, errOut, err := m.runner.Run(ctx, PrivApply, "sshd-config"); err != nil {
m.logger.Error("felhomsshd: config install failed", "err", err, "stderr", strings.TrimSpace(string(errOut)))
return port, err
}
@@ -181,7 +182,9 @@ func (m *Manager) applyAuthorizedKeys(ctx context.Context, sshKey string) {
m.logger.Error("felhomsshd: staging authorized_keys", "err", err)
return
}
if _, errOut, err := m.runner.Run(ctx, "install", "-o", "root", "-g", "root", "-m", "0644", "--", staged, AuthKeysUserPath); err != nil {
// R-861: the root checker installs it — one plain public key, no options (command=, from=, …), or empty.
_ = staged
if _, errOut, err := m.runner.Run(ctx, PrivApply, "sshd-key"); err != nil {
m.logger.Error("felhomsshd: authorized_keys install failed", "err", err, "stderr", strings.TrimSpace(string(errOut)))
return
}
@@ -0,0 +1,26 @@
package felhomsshd
import (
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/privapplytest"
)
// R-861: renderConfig is byte-identical to the checker's template (only the Port varies); a changed line is refused.
// RED-PROOF: change one directive in renderConfig → this fails (the two templates drifted).
func TestPrivApply_AcceptsTheRenderedConfig(t *testing.T) {
for _, port := range []int{2222, 8822, 60022} {
conf, err := renderConfig(port)
if err != nil {
t.Fatal(err)
}
if got := privapplytest.Check(t, "sshd-config", "", conf); got != "OK" {
t.Errorf("port %d: %s", port, got)
}
}
conf, _ := renderConfig(8822)
if got := privapplytest.Check(t, "sshd-config", "", conf+"StrictModes no\n"); !strings.HasPrefix(got, "REFUSED") {
t.Fatalf("control: an extra directive was not refused: %s", got)
}
}
+346
View File
@@ -0,0 +1,346 @@
// Package fstrim is the weekly guest disk trim (R-444, operator ruling `09` §3 decision 139).
//
// Why: a thin pool only ever grows from blocks a guest has already FREED — `fstrim` inside the unprivileged container
// is refused (FITRIM: Operation not permitted), and nothing else on the box gives the blocks back. A full thin pool
// takes every guest on the host read-only, so the pool can reach 100 % from deleted data alone. Measured on demo-hp
// 2026-10-06 09:14Z: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %, 18/18 app probes 200, max 1.1 s
// (audits/ten-answers-2026-10-06/r444-measure.txt).
//
// The rule, each part pinned by a test in fstrim_test.go:
// - Weekly: a guest is DUE from Wednesday 10:00 local until it has been trimmed once since then (a box that was off
// on Wednesday catches up at its next eligible hour).
// - Daytime only: a trim starts only between 10:00 and 20:59 local — never in the night window (01:00–06:59) where
// the backups and the restore-tests run (TestEligibleHourNeverInTheNight).
// - Never beside a backup, a restore-test or another heavy operation: the pass holds the host-wide one-heavy-op gate
// (backup.InFlight) for its whole run; a busy gate DEFERS the pass to the next hourly tick.
// - A failed trim is retried at the next eligible hour, at most MaxAttemptsPerWeek times in one week.
// - The last result per guest (time, bytes, ok/fail) is persisted, so a restart neither loses it nor re-trims.
//
// The command is the ONE exact sudoers shape `pct fstrim <vmid>` (FELHOM_FSTRIM). Only guests from the pool-verified
// source (ListLXC ∩ the felhom pool, audit A1) and only RUNNING ones are trimmed.
package fstrim
import (
"context"
"encoding/json"
"fmt"
"log/slog"
"os"
"path/filepath"
"regexp"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// Schedule. The weekday/hours are fixed on purpose (one sentence the operator can read on the System page).
const (
Weekday = time.Wednesday
StartHour = 10 // first eligible local hour (inclusive)
EndHour = 21 // first NOT-eligible local hour (exclusive): last start is 20:59
MaxAttemptsPerWeek = 3
// TickInterval is how often the job looks; a deferred or failed pass is therefore retried the next hour.
TickInterval = time.Hour
// FirstTickDelay lets the agent settle after a start before the first look.
FirstTickDelay = 5 * time.Minute
// PerGuestTimeout bounds one `pct fstrim` (measured 24.4 s for 84 GiB).
PerGuestTimeout = 30 * time.Minute
)
// ScheduleText is the human description carried on the host report.
const ScheduleText = "weekly, due Wednesday from 10:00 host-local time; starts only 10:00-20:59; never beside a backup or restore-test"
// Runner runs a host command (proxmox.ExecRunner in production, through `sudo -n`).
type Runner interface {
Run(ctx context.Context, name string, args ...string) (stdout, stderr []byte, err error)
}
// GuestSource yields the guests this agent OWNS (the pool-verified source, never a bare ListLXC).
type GuestSource interface {
Guests(ctx context.Context) ([]proxmox.Guest, error)
}
// Gate is the host-wide one-heavy-operation gate (*backup.InFlight).
type Gate interface {
TryAcquire(what string) (release func(), busy string, ok bool)
}
// GateName is what the gate reports as busy while a trim runs.
const GateName = "guest-fstrim"
// Record is one guest's last trim attempt, as persisted.
type Record struct {
LastAttemptAt time.Time `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt time.Time `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
// Attempts counts the attempts since the current week's due time (reset by the first attempt of a new week).
Attempts int `json:"attempts"`
}
// Trimmer is the weekly job.
type Trimmer struct {
runner Runner
guests GuestSource
gate Gate
statePath string
logger *slog.Logger
loc *time.Location
now func() time.Time
mu sync.Mutex
records map[int]Record
}
// New builds the job and loads the persisted state. A missing state file is an empty state; a corrupt one is logged
// and treated as empty (the cost is one extra trim, never a missed one).
func New(runner Runner, guests GuestSource, gate Gate, statePath string, logger *slog.Logger) *Trimmer {
if logger == nil {
logger = slog.Default()
}
t := &Trimmer{runner: runner, guests: guests, gate: gate, statePath: statePath, logger: logger,
loc: time.Local, now: time.Now, records: map[int]Record{}}
t.load()
return t
}
func (t *Trimmer) load() {
data, err := os.ReadFile(t.statePath)
if err != nil {
if !os.IsNotExist(err) {
t.logger.Warn("fstrim: state read failed — starting empty", "path", t.statePath, "err", err)
}
return
}
var raw map[string]Record
if err := json.Unmarshal(data, &raw); err != nil {
t.logger.Warn("fstrim: state file corrupt — starting empty", "path", t.statePath, "err", err)
return
}
for k, r := range raw {
if id, err := strconv.Atoi(k); err == nil && id > 0 {
t.records[id] = r
}
}
}
func (t *Trimmer) saveLocked() error {
raw := make(map[string]Record, len(t.records))
for id, r := range t.records {
raw[strconv.Itoa(id)] = r
}
data, err := json.MarshalIndent(raw, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(t.statePath), 0o755); err != nil {
return err
}
tmp := t.statePath + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, t.statePath)
}
// EligibleHour reports whether a trim may START at local time lt.
func EligibleHour(lt time.Time) bool {
h := lt.Hour()
return h >= StartHour && h < EndHour
}
// weekAnchor is the most recent Wednesday StartHour:00 at or before lt (same location as lt).
func weekAnchor(lt time.Time) time.Time {
daysBack := (int(lt.Weekday()) - int(Weekday) + 7) % 7
d := lt.AddDate(0, 0, -daysBack)
a := time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
if a.After(lt) {
d = d.AddDate(0, 0, -7)
a = time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
}
return a
}
// due reports whether a guest with record r (ok=false: none) is due at local time lt.
func due(r Record, has bool, lt time.Time) bool {
if !has {
return true
}
anchor := weekAnchor(lt)
if r.LastAttemptAt.Before(anchor) {
return true // not tried this week
}
return !r.OK && r.Attempts < MaxAttemptsPerWeek
}
// Run looks every TickInterval until ctx ends. It never returns an error: a failed trim is a reported fact.
func (t *Trimmer) Run(ctx context.Context) {
t.logger.Info("fstrim: weekly guest disk trim starting", "schedule", ScheduleText)
timer := time.NewTimer(FirstTickDelay)
defer timer.Stop()
for {
select {
case <-ctx.Done():
return
case <-timer.C:
t.Pass(ctx)
timer.Reset(TickInterval)
}
}
}
// Pass is one look: outside the daytime window it does nothing; otherwise it trims every due, running, owned guest
// while holding the heavy-op gate.
func (t *Trimmer) Pass(ctx context.Context) {
lt := t.now().In(t.loc)
if !EligibleHour(lt) {
t.logger.Debug("fstrim: outside the daytime window — not looking", "local", lt.Format("Mon 15:04"))
return
}
guests, err := t.guests.Guests(ctx)
if err != nil {
t.logger.Warn("fstrim: owned-guest list unavailable — skipping this pass", "err", err)
return
}
owned := make(map[int]bool, len(guests))
var todo []int
t.mu.Lock()
for _, g := range guests {
owned[g.VMID] = true
r, has := t.records[g.VMID]
if !due(r, has, lt) {
continue
}
if g.Status != "running" {
t.logger.Info("fstrim: guest not running — trimmed when it runs", "vmid", g.VMID, "status", g.Status)
continue
}
todo = append(todo, g.VMID)
}
// A guest the agent no longer owns has no result to report.
pruned := false
for id := range t.records {
if !owned[id] {
delete(t.records, id)
pruned = true
}
}
if pruned {
if err := t.saveLocked(); err != nil {
t.logger.Warn("fstrim: state save failed", "err", err)
}
}
t.mu.Unlock()
if len(todo) == 0 {
return
}
sort.Ints(todo)
release, busy, ok := t.gate.TryAcquire(GateName)
if !ok {
t.logger.Info("fstrim: deferred — a heavy operation is in flight; retrying next hour", "busy", busy, "due_guests", len(todo))
return
}
defer release()
for _, vmid := range todo {
if ctx.Err() != nil {
return
}
t.trimOne(ctx, vmid, lt)
}
}
var trimmedLine = regexp.MustCompile(`\((\d+) bytes\) trimmed`)
// ParseTrimmed sums the "(N bytes) trimmed" lines of `pct fstrim` output and counts them (one per mount point), e.g.
// `/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed`.
func ParseTrimmed(out string) (bytes int64, mounts int) {
for _, m := range trimmedLine.FindAllStringSubmatch(out, -1) {
n, err := strconv.ParseInt(m[1], 10, 64)
if err != nil {
continue
}
bytes += n
mounts++
}
return bytes, mounts
}
// GiB renders bytes as "30.1 GiB".
func GiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
func (t *Trimmer) trimOne(ctx context.Context, vmid int, lt time.Time) {
start := t.now()
cctx, cancel := context.WithTimeout(ctx, PerGuestTimeout)
stdout, stderr, err := t.runner.Run(cctx, "pct", "fstrim", strconv.Itoa(vmid))
cancel()
dur := t.now().Sub(start)
bytes, mounts := ParseTrimmed(string(stdout) + "\n" + string(stderr))
t.mu.Lock()
prev, has := t.records[vmid]
r := Record{LastAttemptAt: start.UTC(), OK: err == nil, BytesTrimmed: bytes, Mounts: mounts,
DurationSeconds: float64(dur.Round(100*time.Millisecond)) / float64(time.Second), LastOKAt: prev.LastOKAt}
if has && !prev.LastAttemptAt.Before(weekAnchor(lt)) {
r.Attempts = prev.Attempts + 1
} else {
r.Attempts = 1
}
if err == nil {
r.LastOKAt = start.UTC()
} else {
msg := strings.TrimSpace(err.Error() + ": " + strings.TrimSpace(string(stderr)))
if len(msg) > 300 {
msg = msg[:300]
}
r.Error = msg
}
t.records[vmid] = r
saveErr := t.saveLocked()
t.mu.Unlock()
if err == nil {
t.logger.Info(fmt.Sprintf("fstrim: guest %d trimmed %s in %.1fs", vmid, GiB(bytes), r.DurationSeconds),
"vmid", vmid, "bytes_trimmed", bytes, "mounts", mounts, "duration_s", r.DurationSeconds)
if mounts == 0 {
t.logger.Warn("fstrim: pct fstrim succeeded but reported no trimmed mount — output not understood",
"vmid", vmid, "stdout", strings.TrimSpace(string(stdout)))
}
} else {
t.logger.Warn(fmt.Sprintf("fstrim: guest %d trim FAILED after %.1fs", vmid, r.DurationSeconds),
"vmid", vmid, "attempt", r.Attempts, "max_attempts_per_week", MaxAttemptsPerWeek, "err", r.Error)
}
if saveErr != nil {
t.logger.Warn("fstrim: state save failed — the result will not survive a restart", "path", t.statePath, "err", saveErr)
}
}
// GuestDiskTrimStatus implements hub.GuestDiskTrimReporter: a pure read of the persisted results (never runs pct).
func (t *Trimmer) GuestDiskTrimStatus(context.Context) *hub.GuestDiskTrimStatus {
t.mu.Lock()
defer t.mu.Unlock()
out := &hub.GuestDiskTrimStatus{Schedule: ScheduleText}
ids := make([]int, 0, len(t.records))
for id := range t.records {
ids = append(ids, id)
}
sort.Ints(ids)
for _, id := range ids {
r := t.records[id]
g := hub.GuestDiskTrim{VMID: id, LastAttemptAt: r.LastAttemptAt.UTC().Format(time.RFC3339), OK: r.OK,
BytesTrimmed: r.BytesTrimmed, Mounts: r.Mounts, DurationSeconds: r.DurationSeconds, Error: r.Error}
if !r.LastOKAt.IsZero() {
g.LastOKAt = r.LastOKAt.UTC().Format(time.RFC3339)
}
out.Guests = append(out.Guests, g)
}
return out
}
+267
View File
@@ -0,0 +1,267 @@
package fstrim
import (
"bytes"
"context"
"encoding/json"
"errors"
"log/slog"
"path/filepath"
"reflect"
"strings"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// The real `pct fstrim 9201` output measured on demo-hp 2026-10-06 (audits/ten-answers-2026-10-06/r444-measure.txt).
const measuredOut = "/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed\n" +
"/var/lib/lxc/9201/rootfs/var/lib/felhom: 53.9 GiB (57865633792 bytes) trimmed\n"
const measuredBytes = int64(32277680128 + 57865633792)
type fakeRunner struct {
mu sync.Mutex
calls [][]string
out string
err error
onRun func()
}
func (f *fakeRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
f.mu.Lock()
f.calls = append(f.calls, append([]string{name}, args...))
f.mu.Unlock()
if f.onRun != nil {
f.onRun()
}
if f.err != nil {
return nil, []byte("mount busy"), f.err
}
return []byte(f.out), nil, nil
}
type fakeGuests struct {
g []proxmox.Guest
err error
}
func (f fakeGuests) Guests(context.Context) ([]proxmox.Guest, error) { return f.g, f.err }
// A Wednesday 10:30 in a fixed zone (CEST-like), so the tests do not depend on the machine's zone.
var zone = time.FixedZone("CEST", 2*3600)
func at(day, hour, min int) time.Time { return time.Date(2026, 10, day, hour, min, 0, 0, zone) } // 2026-10-07 = Wednesday
func newT(t *testing.T, r Runner, g GuestSource, gate Gate, now *time.Time) (*Trimmer, *bytes.Buffer, string) {
t.Helper()
var logs bytes.Buffer
path := filepath.Join(t.TempDir(), "guest-disk-trim.json")
tr := New(r, g, gate, path, slog.New(slog.NewTextHandler(&logs, &slog.HandlerOptions{Level: slog.LevelDebug})))
tr.loc = zone
tr.now = func() time.Time { return *now }
return tr, &logs, path
}
func running(ids ...int) fakeGuests {
var g []proxmox.Guest
for _, id := range ids {
g = append(g, proxmox.Guest{VMID: id, Status: "running", Type: "lxc"})
}
return fakeGuests{g: g}
}
func TestParseTrimmedTheMeasuredOutput(t *testing.T) {
b, m := ParseTrimmed(measuredOut)
if b != measuredBytes || m != 2 {
t.Fatalf("ParseTrimmed = %d bytes over %d mounts, want %d over 2", b, m, measuredBytes)
}
if b, m := ParseTrimmed("something else\n"); b != 0 || m != 0 {
t.Fatalf("unrelated output parsed as %d/%d", b, m)
}
if got := GiB(measuredBytes); got != "84.0 GiB" {
t.Fatalf("GiB = %q", got)
}
}
// The night window (01:00–06:59) must never be eligible, and the daytime window is exactly 10:00–20:59.
func TestEligibleHourNeverInTheNight(t *testing.T) {
for h := 0; h < 24; h++ {
lt := time.Date(2026, 10, 7, h, 30, 0, 0, zone)
got := EligibleHour(lt)
if h >= 1 && h <= 6 && got {
t.Errorf("hour %02d is in the night window and must not be eligible", h)
}
if want := h >= 10 && h <= 20; got != want {
t.Errorf("EligibleHour(%02d:30) = %v, want %v", h, got, want)
}
}
}
func TestWeekAnchorIsTheLastWednesdayTen(t *testing.T) {
cases := map[time.Time]time.Time{
at(7, 10, 0): at(7, 10, 0), // Wednesday 10:00 itself
at(7, 9, 59): time.Date(2026, 9, 30, 10, 0, 0, 0, zone), // before 10:00 Wednesday → the previous week
at(8, 15, 0): at(7, 10, 0), // Thursday
at(13, 20, 0): at(7, 10, 0), // next Tuesday
at(14, 11, 0): at(14, 10, 0), // next Wednesday
}
for in, want := range cases {
if got := weekAnchor(in); !got.Equal(want) {
t.Errorf("weekAnchor(%s) = %s, want %s", in.Format("Mon 01-02 15:04"), got.Format("Mon 01-02 15:04"), want.Format("Mon 01-02 15:04"))
}
}
}
// The consequence: on Wednesday 10:30 a running owned guest is trimmed with the ONE exact argv, the bytes are parsed,
// the positive log line is written, the result is persisted, and the host report carries it.
func TestPassTrimsADueGuestAndReportsIt(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
tr, logs, path := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9201"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("runner calls = %q, want %q", r.calls, want)
}
if !strings.Contains(logs.String(), "fstrim: guest 9201 trimmed 84.0 GiB in ") {
t.Fatalf("no positive per-guest log line:\n%s", logs.String())
}
st := tr.GuestDiskTrimStatus(context.Background())
if st == nil || st.Schedule != ScheduleText || len(st.Guests) != 1 {
t.Fatalf("report stanza = %+v", st)
}
g := st.Guests[0]
if g.VMID != 9201 || !g.OK || g.BytesTrimmed != measuredBytes || g.Mounts != 2 || g.LastOKAt == "" || g.LastAttemptAt == "" {
t.Fatalf("report guest = %+v", g)
}
// Persisted: a NEW Trimmer over the same file (an agent restart) still has it and does not trim again this week.
now = at(8, 11, 0)
r2 := &fakeRunner{out: measuredOut}
tr2 := New(r2, running(9201), &backup.InFlight{}, path, slog.New(slog.NewTextHandler(&bytes.Buffer{}, nil)))
tr2.loc, tr2.now = zone, func() time.Time { return now }
if st2 := tr2.GuestDiskTrimStatus(context.Background()); len(st2.Guests) != 1 || st2.Guests[0].BytesTrimmed != measuredBytes {
t.Fatalf("result lost over a restart: %+v", st2)
}
tr2.Pass(context.Background())
if len(r2.calls) != 0 {
t.Fatalf("trimmed again in the same week after a restart: %q", r2.calls)
}
// Next week it is due again.
now = at(14, 10, 5)
tr2.Pass(context.Background())
if len(r2.calls) != 1 {
t.Fatalf("not trimmed in the next week: %q", r2.calls)
}
}
func TestPassNeverRunsInTheNight(t *testing.T) {
for _, h := range []int{1, 3, 6, 9, 21, 23} {
now := at(7, h, 15)
r := &fakeRunner{out: measuredOut}
tr, _, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Errorf("trimmed at %02d:15: %q", h, r.calls)
}
}
}
// A backup (or restore-test) holding the heavy-op gate DEFERS the trim; the next hour, gate free, it runs. And while
// a trim runs, the gate is held, so a backup cannot start beside it.
func TestPassDefersToAHeavyOperationAndRetriesNextHour(t *testing.T) {
now := at(7, 10, 30)
gate := &backup.InFlight{}
release, _, _ := gate.TryAcquire("backup:9201")
var busyDuringTrim string
r := &fakeRunner{out: measuredOut}
r.onRun = func() { busyDuringTrim = gate.Busy() }
tr, logs, _ := newT(t, r, running(9201), gate, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Fatalf("trimmed beside a running backup: %q", r.calls)
}
if !strings.Contains(logs.String(), "fstrim: deferred") || !strings.Contains(logs.String(), "backup:9201") {
t.Fatalf("the deferral is not logged with what holds the gate:\n%s", logs.String())
}
release()
now = now.Add(time.Hour)
tr.Pass(context.Background())
if len(r.calls) != 1 {
t.Fatalf("not retried the next hour: %q", r.calls)
}
if busyDuringTrim != GateName {
t.Fatalf("the heavy-op gate was %q during the trim, want %q", busyDuringTrim, GateName)
}
if gate.Busy() != "" {
t.Fatalf("the gate was not released after the pass: %q", gate.Busy())
}
}
func TestFailedTrimWarnsIsRecordedAndRetriedAtMostThreeTimes(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{err: errors.New("exit status 255")}
tr, logs, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
for i := 0; i < 6; i++ {
tr.Pass(context.Background())
now = now.Add(time.Hour)
}
if len(r.calls) != MaxAttemptsPerWeek {
t.Fatalf("attempts in one week = %d, want %d", len(r.calls), MaxAttemptsPerWeek)
}
if !strings.Contains(logs.String(), "level=WARN") || !strings.Contains(logs.String(), "fstrim: guest 9201 trim FAILED") {
t.Fatalf("no WARN for the failure:\n%s", logs.String())
}
g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]
if g.OK || g.LastOKAt != "" || !strings.Contains(g.Error, "exit status 255") || !strings.Contains(g.Error, "mount busy") {
t.Fatalf("failed result not recorded as a failure: %+v", g)
}
// A success later keeps a clean record.
r.err = nil
r.out = measuredOut
now = at(14, 10, 10)
tr.Pass(context.Background())
if g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]; !g.OK || g.Error != "" || g.BytesTrimmed != measuredBytes {
t.Fatalf("success after failure: %+v", g)
}
}
func TestOnlyRunningOwnedGuestsAndAFailedListActsOnNothing(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
g := fakeGuests{g: []proxmox.Guest{{VMID: 9201, Status: "stopped"}, {VMID: 9202, Status: "running"}}}
tr, _, _ := newT(t, r, g, &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9202"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("calls = %q, want only the running guest", r.calls)
}
r2 := &fakeRunner{out: measuredOut}
tr2, logs, _ := newT(t, r2, fakeGuests{err: errors.New("pool read 403")}, &backup.InFlight{}, &now)
tr2.Pass(context.Background())
if len(r2.calls) != 0 || !strings.Contains(logs.String(), "owned-guest list unavailable") {
t.Fatalf("a failed ownership read must act on nothing: calls %q", r2.calls)
}
}
func TestReportJSONShape(t *testing.T) {
now := at(7, 10, 30)
tr, _, _ := newT(t, &fakeRunner{out: measuredOut}, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
b, err := json.Marshal(tr.GuestDiskTrimStatus(context.Background()))
if err != nil {
t.Fatal(err)
}
for _, k := range []string{`"schedule":`, `"guests":[{"vmid":9201`, `"last_attempt_at":"2026-10-07T08:30:00Z"`, `"ok":true`,
`"bytes_trimmed":90143313920`, `"mounts":2`, `"duration_seconds":`, `"last_ok_at":"2026-10-07T08:30:00Z"`} {
if !strings.Contains(string(b), k) {
t.Errorf("report JSON lacks %s: %s", k, b)
}
}
if strings.Contains(string(b), `"error"`) {
t.Errorf("an ok result must omit error: %s", b)
}
}
+16 -25
View File
@@ -38,35 +38,26 @@ const snippetBody = `#!/bin/sh
exit 0
`
// InstallSnippet writes the pre-start hook wrapper into the PVE snippets dir (idempotent, root-owned,
// executable). The agent runs as a non-root service user, so it writes an agent-writable temp file then
// `install`s it host-root (same pattern as the bootstrap mount + dnsmasq drop-ins). Safe to call repeatedly.
// The temp file is a RANDOM-named os.CreateTemp (audit B1): a fixed, predictable /tmp name could be
// pre-created by another local user and rewritten between our write and root's install (TOCTOU into a
// root-executed hookscript). The final mode comes from `install -m`, so the 0600 temp is fine.
func InstallSnippet(ctx context.Context, runner proxmox.Runner) error {
f, err := os.CreateTemp("", "felhom-guest-hook-*.sh")
// SnippetReady (R-861, agent v0.146.0) reports whether the pre-start hook is in place: a regular file at path whose
// content is exactly snippetBody. The hook is a FIXED, ROOT-OWNED file that arrives with the signed config bundle
// (configs/felhom-guest-hook.sh, pinned byte-identical by TestSnippetEqualsTheBundle). The agent no longer installs it:
// until v0.146.0 it `install`ed it from /tmp, and Proxmox runs a hookscript as root at every guest start — so the
// install grant was a root shell for a compromised agent. A missing or different hook is an error the caller logs; it
// then does NOT register the hook (a guest whose hookscript is missing does not start).
func SnippetReady(path string) error {
fi, err := os.Lstat(path)
if err != nil {
return fmt.Errorf("guesthook: create temp snippet: %w", err)
return fmt.Errorf("guesthook: %s is missing — it arrives with the signed config bundle (agent_config_update): %w", path, err)
}
tmp := f.Name()
defer os.Remove(tmp)
if _, err := f.WriteString(snippetBody); err != nil {
f.Close()
return fmt.Errorf("guesthook: write temp snippet: %w", err)
if !fi.Mode().IsRegular() {
return fmt.Errorf("guesthook: %s is not a regular file", path)
}
if err := f.Close(); err != nil {
return fmt.Errorf("guesthook: close temp snippet: %w", err)
b, err := os.ReadFile(path)
if err != nil {
return fmt.Errorf("guesthook: read %s: %w", path, err)
}
// Ensure the snippets dir exists FIRST (B2, DRILL-day0-cleanroom-2026-07-03): a fresh PVE has
// no /var/lib/vz/snippets, and `install` (without -D) won't create the parent — the whole
// hook install silently failed on a freshly-bootstrapped box. Fenced root op like the install
// itself; idempotent.
if _, stderr, err := runner.Run(ctx, "mkdir", "-p", SnippetDir); err != nil {
return fmt.Errorf("guesthook: ensure snippets dir %s: %w: %s", SnippetDir, err, string(stderr))
}
if _, stderr, err := runner.Run(ctx, "install", "-m", "0755", "--", tmp, SnippetPath); err != nil {
return fmt.Errorf("guesthook: install snippet to %s: %w: %s", SnippetPath, err, string(stderr))
if string(b) != snippetBody {
return fmt.Errorf("guesthook: %s differs from this agent's hook — the next config bundle replaces it", path)
}
return nil
}
+24 -90
View File
@@ -4,7 +4,7 @@ import (
"context"
"io"
"os"
"regexp"
"path/filepath"
"testing"
)
@@ -29,100 +29,34 @@ func (r *recordingRunner) RunStdin(ctx context.Context, _ io.Reader, name string
return r.Run(ctx, name, args...)
}
// TestInstallSnippet_RandomTempName is the audit-B1 negative test: the staged install SOURCE must be a
// RANDOM os.CreateTemp name (felhom-guest-hook-<random>.sh), never the fixed, pre-creatable
// /tmp/felhom-guest-hook.sh (a local TOCTOU into a root-executed hookscript), and two consecutive
// installs must stage through DIFFERENT paths.
func TestInstallSnippet_RandomTempName(t *testing.T) {
r := &recordingRunner{}
if err := InstallSnippet(context.Background(), r); err != nil {
t.Fatalf("InstallSnippet #1: %v", err)
// R-861 (agent v0.146.0): the hook file comes with the signed config bundle; the agent never installs it, only checks.
// RED-PROOF (audits/hub-safety-2026-10-05/partF/red-proof.txt): make SnippetReady accept any content → the
// "differs" case fails.
func TestSnippetReady(t *testing.T) {
d := t.TempDir()
p := filepath.Join(d, "felhom-guest-hook.sh")
if err := SnippetReady(p); err == nil {
t.Fatal("a missing hook read as ready — the guest would get a hookscript that does not exist")
}
if err := InstallSnippet(context.Background(), r); err != nil {
t.Fatalf("InstallSnippet #2: %v", err)
_ = os.WriteFile(p, []byte("#!/bin/sh\nid > /tmp/x\n"), 0o755)
if err := SnippetReady(p); err == nil {
t.Fatal("a hook with other content read as ready")
}
var installs [][]string
for _, call := range r.calls {
if call[0] == "install" {
installs = append(installs, call)
}
_ = os.WriteFile(p, []byte(snippetBody), 0o755)
if err := SnippetReady(p); err != nil {
t.Fatalf("the bundle's hook was not accepted: %v", err)
}
if len(installs) != 2 {
t.Fatalf("expected 2 install calls, got %d: %v", len(installs), r.calls)
}
randomName := regexp.MustCompile(`felhom-guest-hook-[^/\\]+\.sh$`)
fixedName := regexp.MustCompile(`felhom-guest-hook\.sh$`)
var srcs []string
for i, call := range installs {
// install -m 0755 -- <src> <dest>
if len(call) != 6 {
t.Fatalf("call %d: unexpected vector %v", i, call)
}
src, dest := call[4], call[5]
if dest != SnippetPath {
t.Errorf("call %d: dest = %q, want %q", i, dest, SnippetPath)
}
if !randomName.MatchString(src) {
t.Errorf("call %d: source %q does not match the random felhom-guest-hook-*.sh pattern", i, src)
}
if fixedName.MatchString(src) {
t.Errorf("call %d: source %q is the FIXED predictable temp name (B1 TOCTOU)", i, src)
}
srcs = append(srcs, src)
}
if srcs[0] == srcs[1] {
t.Errorf("two consecutive installs staged through the SAME source path %q — must be random per call", srcs[0])
}
// Non-hollow: the staged file must actually carry the snippet body at install time.
for i, c := range r.srcContent {
if c != snippetBody {
t.Errorf("call %d: staged content is not the snippet body (got %d bytes)", i, len(c))
}
}
// And the temp is cleaned up after.
for _, src := range srcs {
if _, err := os.Stat(src); err == nil {
t.Errorf("staged temp %q left behind (defer os.Remove missing)", src)
}
l := filepath.Join(d, "link.sh")
_ = os.Symlink(p, l)
if err := SnippetReady(l); err == nil {
t.Fatal("a symlink read as the hook")
}
}
// B2 Scenario D (DRILL-day0-cleanroom-2026-07-03): on a fresh PVE, /var/lib/vz/snippets does not
// exist and `install` (no -D) cannot create it — the drill saw
// `install: cannot create regular file … No such file or directory` and the guest silently got no
// pre-start self-heal hook. InstallSnippet must therefore issue a `mkdir -p <SnippetDir>` fenced op
// BEFORE the `install` op. Pre-fix wrong outcome: no mkdir call at all — only the doomed install.
func TestInstallSnippet_EnsuresSnippetsDirFirst(t *testing.T) {
r := &recordingRunner{}
if err := InstallSnippet(context.Background(), r); err != nil {
t.Fatalf("InstallSnippet: %v", err)
}
mkdirIdx, installIdx := -1, -1
for i, call := range r.calls {
switch call[0] {
case "mkdir":
if mkdirIdx == -1 {
mkdirIdx = i
want := []string{"mkdir", "-p", SnippetDir}
if len(call) != 3 || call[1] != want[1] || call[2] != want[2] {
t.Errorf("mkdir vector = %v, want %v (the sudoers fence matches exactly this argv)", call, want)
}
}
case "install":
if installIdx == -1 {
installIdx = i
}
}
}
if mkdirIdx == -1 {
t.Fatalf("no `mkdir -p %s` op issued — on a fresh box the snippet install fails ENOENT (B2); calls: %v", SnippetDir, r.calls)
}
if installIdx == -1 {
t.Fatalf("no install op issued; calls: %v", r.calls)
}
if mkdirIdx > installIdx {
t.Fatalf("mkdir (call %d) must PRECEDE install (call %d) — order: %v", mkdirIdx, installIdx, r.calls)
// The bundle's copy is byte-identical to the body the agent checks against.
func TestSnippetEqualsTheBundle(t *testing.T) {
got, err := os.ReadFile(filepath.Join("..", "..", "configs", "felhom-guest-hook.sh"))
if err != nil || string(got) != snippetBody {
t.Fatalf("configs/felhom-guest-hook.sh differs from snippetBody (err %v)", err)
}
}
+59
View File
@@ -0,0 +1,59 @@
// Package httpx holds the one HTTP-transport default this repo may not lose.
//
// Every client here pins TLS — PBS and PVE by leaf-cert SHA-256, the hub by an optional CA file —
// so none of them can use http.DefaultTransport and each hand-rolls its own. Hand-rolling silently
// discards DefaultTransport's settings, and one of them is load-bearing:
//
// Transport: &http.Transport{TLSClientConfig: tlsCfg} // IdleConnTimeout == 0 == NO timeout
//
// A zero IdleConnTimeout means idle keep-alive connections are retained FOREVER, not "use a sane
// default". Combined with a client that is rebuilt on a schedule and dropped (pbsTargetsFromPVE
// builds a fresh pbs.Client per cycle), every cycle strands one connection that nothing will ever
// close: the abandoned Transport becomes unreachable but its persistConn read-loop goroutine keeps
// the socket alive, and an unreachable Transport does not close its connections.
//
// Measured cost, live: 388 established connections accumulated on ep0's PBS proxy between
// 2026-08-18 09:51:22Z and 2026-08-20 08:02:13Z — 194 from each of the two boxes, held open on BOTH
// sides, one per agent poll cycle, on a proxy whose descriptor ceiling is 65536. See R-344 and
// felhom.eu/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md.
package httpx
import (
"crypto/tls"
"net/http"
"time"
)
// DefaultIdleConnTimeout is how long an idle keep-alive connection is retained before it is closed.
//
// It is 90s because that is http.DefaultTransport's own value: the fix for R-344 restores a
// standard-library default rather than inventing a number, so there is nothing here to tune and
// nothing to justify. It is comfortably shorter than every cadence that drives these clients (the
// 15-minute live-snapshot collect and the 6-hour verify loop), so a connection abandoned by one
// cycle is closed long before the next.
const DefaultIdleConnTimeout = 90 * time.Second
// NewTransport builds a FRESH *http.Transport pinned to tlsCfg, with the idle-connection timeout
// applied.
//
// Fresh, never shared: each caller pins a different endpoint, and a shared transport would pool
// connections across differently pinned servers. Reusing http.DefaultTransport for the same reason
// is not an option — it would drop the pin entirely.
//
// idleConnTimeout <= 0 means USE THE DEFAULT. It deliberately does not mean "no timeout": no-timeout
// is the bug this package exists to prevent, and an unset field must never be able to reintroduce
// it. Callers pass their configured value straight through; only tests pass a short one.
//
// Only IdleConnTimeout is set. The other DefaultTransport settings this transport also lacks
// (MaxIdleConns, TLSHandshakeTimeout, ExpectContinueTimeout) are deliberately left alone: none of
// them accumulates anything, every client bounds its whole request with http.Client.Timeout, and
// widening the change would have made the R-344 measurement unattributable.
func NewTransport(tlsCfg *tls.Config, idleConnTimeout time.Duration) *http.Transport {
if idleConnTimeout <= 0 {
idleConnTimeout = DefaultIdleConnTimeout
}
return &http.Transport{
TLSClientConfig: tlsCfg,
IdleConnTimeout: idleConnTimeout,
}
}
+74
View File
@@ -0,0 +1,74 @@
package httpx
import (
"crypto/tls"
"net/http"
"testing"
"time"
)
// TestNewTransport_ZeroMeansDefaultNeverForever is the whole point of this package.
//
// http.Transport's zero IdleConnTimeout means "retain idle connections FOREVER". Any code path that
// can reach that zero reintroduces R-344, so an unset, zero or negative value must all land on the
// default. If someone later "simplifies" NewTransport by passing the argument straight through,
// this fails.
func TestNewTransport_ZeroMeansDefaultNeverForever(t *testing.T) {
for _, tc := range []struct {
name string
in time.Duration
want time.Duration
}{
{"zero", 0, DefaultIdleConnTimeout},
{"negative", -time.Hour, DefaultIdleConnTimeout},
{"explicit short value (tests)", 50 * time.Millisecond, 50 * time.Millisecond},
{"explicit long value", time.Hour, time.Hour},
} {
t.Run(tc.name, func(t *testing.T) {
got := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, tc.in).IdleConnTimeout
if got != tc.want {
t.Fatalf("IdleConnTimeout = %v, want %v", got, tc.want)
}
if got == 0 {
t.Fatal("IdleConnTimeout is 0 — that is 'never expire', which is the R-344 defect itself")
}
})
}
}
// TestDefaultIdleConnTimeout_MatchesTheStandardLibrary pins the number to its justification.
//
// 90s is not a tuned value; it is what http.DefaultTransport uses. Reading it off the standard
// library rather than hardcoding 90 means the constant cannot drift away from the reason given for
// it in the package doc.
func TestDefaultIdleConnTimeout_MatchesTheStandardLibrary(t *testing.T) {
std, ok := http.DefaultTransport.(*http.Transport)
if !ok {
t.Skip("http.DefaultTransport is not an *http.Transport in this Go build")
}
if DefaultIdleConnTimeout != std.IdleConnTimeout {
t.Fatalf("DefaultIdleConnTimeout = %v but http.DefaultTransport uses %v — the doc comment's justification no longer holds",
DefaultIdleConnTimeout, std.IdleConnTimeout)
}
}
// TestNewTransport_IsFreshEveryCall guards the pooling property the pinning relies on.
//
// Each caller pins a DIFFERENT endpoint. A shared transport would pool connections across
// differently pinned servers, so returning a package-level singleton would be a security change
// dressed as a tidy-up.
func TestNewTransport_IsFreshEveryCall(t *testing.T) {
a := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
b := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
if a == b {
t.Fatal("NewTransport returned the SAME transport twice — connections would be pooled across differently pinned endpoints")
}
}
// TestNewTransport_KeepsTheTLSConfig — the transport gains a field; it must lose nothing.
func TestNewTransport_KeepsTheTLSConfig(t *testing.T) {
cfg := &tls.Config{MinVersion: tls.VersionTLS12, InsecureSkipVerify: true} //nolint:gosec // test only
if got := NewTransport(cfg, 0).TLSClientConfig; got != cfg {
t.Fatalf("TLSClientConfig = %p, want the config passed in (%p) — the pin would be dropped", got, cfg)
}
}
+25
View File
@@ -0,0 +1,25 @@
package hub
import (
"os"
"path/filepath"
"testing"
)
// R-840: the agent reports the bundle record itself — "none" when no bundle ever reached the box (so an old wrapper
// cannot hide that), "unknown" when it cannot be read, else the version and sha the root wrapper recorded.
func TestReadBundleRecord(t *testing.T) {
dir := t.TempDir()
p := filepath.Join(dir, "config-bundle.json")
if got := string(readBundleRecord(p)); got != `{"version":"none"}` {
t.Fatalf("absent: %s", got)
}
os.WriteFile(p, []byte("{broken"), 0o644)
if got := string(readBundleRecord(p)); got != `{"version":"unknown"}` {
t.Fatalf("broken: %s", got)
}
os.WriteFile(p, []byte(`{"format":1,"agent_version":"0.143.0","bundle_sha256":"abc","installed_at":"2026-10-04T20:00:00Z","authority":"signed","files":{"/x":"y"}}`), 0o644)
if got := string(readBundleRecord(p)); got != `{"authority":"signed","bundle_sha256":"abc","installed_at":"2026-10-04T20:00:00Z","version":"0.143.0"}` {
t.Fatalf("record: %s", got)
}
}
+144 -2
View File
@@ -15,6 +15,8 @@ import (
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/config"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
const reportPath = "/api/v1/host-report"
@@ -49,8 +51,11 @@ func NewClient(cfg config.HubConfig, logger *slog.Logger) (*Client, error) {
tlsCfg.RootCAs = pool
}
hc := &http.Client{
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
Transport: &http.Transport{TLSClientConfig: tlsCfg},
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
// R-344, consistency only: this client is built ONCE per process, so it never accumulated
// and contributed nothing to the ep0 leak. It carried the same missing default, which over
// a tunnel is how one idle connection survives long enough to fail on next use.
Transport: httpx.NewTransport(tlsCfg, 0),
}
return newClient(cfg.URL, cfg.APIKey, cfg.HostID, hc, logger), nil
}
@@ -307,3 +312,140 @@ func tail(b []byte, max int) string {
}
return s
}
// IdentityEscrowResponse mirrors GET /api/v1/hosts/{host_id}/escrow (hub >= v0.94.0, R-199).
// Present=false is a CLEAN answer, not a fault: the host simply has no sealed bundle yet.
type IdentityEscrowResponse struct {
HostID string `json:"host_id"`
Present bool `json:"present"`
IdentityEscrowB64 string `json:"identity_escrow_b64"`
}
// FetchIdentityEscrow reads back THIS host's own opaque identity-escrow blob (R-199 link 6 — the
// mirror of UploadEscrow, self-scoped server-side by the per-host key). The bytes are ciphertext: they
// are useless without the customer's recovery code R, which neither the hub nor this agent ever holds.
//
// It is the ONLY retrieval this client performs, and it is deliberately narrow — no directive, no
// K-escrow, no key rotation. The operator-driven DR path (recovery-mode re-enroll) is a different
// endpoint with a different gate and is not reached from here.
//
// Errors are typed (transport vs HTTP) and never include the bearer token. The BLOB is never logged —
// only its length.
func (c *Client) FetchIdentityEscrow(ctx context.Context) (*IdentityEscrowResponse, error) {
if c.hostID == "" {
return nil, fmt.Errorf("hub: FetchIdentityEscrow requires a configured host_id")
}
url := c.baseURL + "/api/v1/hosts/" + c.hostID + "/escrow"
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, fmt.Errorf("hub: building escrow-fetch request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+c.apiKey)
req.Header.Set("Accept", "application/json")
resp, err := c.hc.Do(req)
if err != nil {
return nil, &TransportError{Err: err}
}
defer resp.Body.Close()
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, &HTTPError{StatusCode: resp.StatusCode, BodyTail: tail(raw, 256)}
}
var out IdentityEscrowResponse
if err := json.Unmarshal(raw, &out); err != nil {
return nil, fmt.Errorf("hub: decoding escrow fetch: %w", err)
}
return &out, nil
}
// RetainedEscrowPackage is one RETAINED (superseded) sealed identity package. The blob is ciphertext
// and is useless without R. `SupersededAt` is the only thing here a human ever sees — it is what lets
// the recovery screen name WHICH earlier package a code belongs to.
type RetainedEscrowPackage struct {
Index int `json:"index"`
SupersededAt string `json:"superseded_at"`
KeyFingerprint string `json:"key_fingerprint"`
IdentityEscrowB64 string `json:"identity_escrow_b64"`
}
// RetainedEscrowResponse mirrors GET /api/v1/hosts/{host_id}/escrow/retained (hub >= v0.103.0, R-311).
//
// UnopenableCount is NOT noise. It counts retained packages the hub holds whose key material is absent
// (every pre-v0.93.0 row): on a box with those and nothing else, a perfectly correct old recovery code
// opens nothing, and the reason is a defect of ours. A caller that ignores this number will tell such a
// customer their code is wrong — the exact failure this whole chain exists to stop.
type RetainedEscrowResponse struct {
HostID string `json:"host_id"`
Count int `json:"count"`
UnopenableCount int `json:"unopenable_count"`
TruncatedCount int `json:"truncated_count"`
Packages []RetainedEscrowPackage `json:"packages"`
}
// FetchRetainedIdentityEscrow reads back THIS host's RETAINED sealed identity packages (R-311 —
// the retained siblings of FetchIdentityEscrow, self-scoped server-side by the same per-host key).
//
// SEPARATE FROM FetchIdentityEscrow ON PURPOSE. The ordinary recovery must not pay for this call, and
// must not fail because of it: the current package is tried first and alone, and this is reached only
// after that has refused. A hub too old to know this route answers 404, which is a CLEAN "none" here
// and must never be reported as a failed recovery.
func (c *Client) FetchRetainedIdentityEscrow(ctx context.Context) (*RetainedEscrowResponse, error) {
if c.hostID == "" {
return nil, fmt.Errorf("hub: FetchRetainedIdentityEscrow requires a configured host_id")
}
url := c.baseURL + "/api/v1/hosts/" + c.hostID + "/escrow/retained"
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, fmt.Errorf("hub: building retained-escrow request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+c.apiKey)
req.Header.Set("Accept", "application/json")
resp, err := c.hc.Do(req)
if err != nil {
return nil, &TransportError{Err: err}
}
defer resp.Body.Close()
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
if resp.StatusCode == http.StatusNotFound {
// A hub older than v0.103.0 has no such route. That is "no retained packages", not a fault —
// returning an error here would turn an old hub into a failed recovery on a box whose current
// package simply did not open.
return &RetainedEscrowResponse{HostID: c.hostID}, nil
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, &HTTPError{StatusCode: resp.StatusCode, BodyTail: tail(raw, 256)}
}
var out RetainedEscrowResponse
if err := json.Unmarshal(raw, &out); err != nil {
return nil, fmt.Errorf("hub: decoding retained escrow fetch: %w", err)
}
return &out, nil
}
// PostOSReport sends the OS-update leg's report after every run (hub v0.130.0): POST /api/v1/hosts/{id}/os-report.
// Per-host key, self-scoped on the hub. Errors are typed like RegisterWG's and never include the bearer.
func (c *Client) PostOSReport(ctx context.Context, body []byte) error {
if c.hostID == "" {
return fmt.Errorf("hub: PostOSReport requires a configured host_id")
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL+"/api/v1/hosts/"+c.hostID+"/os-report", bytes.NewReader(body))
if err != nil {
return fmt.Errorf("hub: building os-report request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+c.apiKey)
req.Header.Set("Content-Type", "application/json")
resp, err := c.hc.Do(req)
if err != nil {
return &TransportError{Err: err}
}
defer resp.Body.Close()
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 64<<10))
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return &HTTPError{StatusCode: resp.StatusCode, BodyTail: tail(raw, 256)}
}
return nil
}
+97 -31
View File
@@ -2,45 +2,111 @@ package hub
import (
"context"
"os/exec"
"fmt"
"strings"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// CloudflaredProber reports the cloudflared tunnel service health. It is a
// READ-ONLY probe: the agent does NOT manage or restart cloudflared in this slice
// (that is the tunnel-management slice — this is the seam for it). Injectable so
// tests use a fake and never exec.
// Tunnel states the agent reports (R-841, agent v0.141.0). THREE, never two: a probe that could not ask is
// `unknown`, which the hub never shows as up or down and never alarms on (R-96 rule 3).
const (
TunnelRunning = "running" // the cloudflared container runs AND its readiness check says CONNECTED
TunnelNotRunning = "not_running" // stopped / exited / absent, or running but NOT connected (Detail says which)
TunnelUnknown = "unknown" // the probe could not ask (guest down, pct/sudo error, health still starting)
)
// CloudflaredProber reports the box's tunnel. Injectable so tests use a fake and never exec.
type CloudflaredProber interface {
// Status returns one of: "active" | "inactive" | "failed" | "unknown".
Status(ctx context.Context) (string, error)
// Status returns one of the Tunnel* states and a short detail (why not_running / why unknown).
Status(ctx context.Context) (status, detail string)
}
// SystemctlProber runs `systemctl is-active cloudflared`. This is NOT a Privileged
// (root-CLI) op — `is-active` is non-root readable and is not one of the three
// proven root exceptions, so it does not go through internal/proxmox.Privileged.
type SystemctlProber struct {
Unit string // defaults to "cloudflared"
// GuestTunnelProber reads the REAL tunnel: the `cloudflared` container in the box's own customer guest.
//
// Before v0.141.0 the agent ran `systemctl is-active cloudflared` on the HOST — a unit that does not exist (cloudflared
// is a guest container, `11-os-updates.md` C8), so every box reported `inactive` (R-841).
//
// It uses ONLY the existing sudoers line `pct exec [0-9]* -- docker inspect -f *` (03 §3): the container's state, exit
// code and Docker health status. The health status comes from the compose health check controller v0.292.0 adds
// (`cloudflared tunnel --metrics localhost:20241 ready` → /ready: 200 only with ≥ 1 connection). Measured 2026-10-04:
// with a wrong token the container stays "running" while /ready answers 503 — so the container state alone would lie.
// A container with no health check (an older controller) is judged on its state alone, and Detail says so.
type GuestTunnelProber struct {
Runner proxmox.Runner
// Guests returns the box's customer guest vmids (running pool guests that bind /mnt/felhom-drives).
Guests func(ctx context.Context) ([]int, error)
}
// Status maps `systemctl is-active` output to the report vocabulary. systemctl
// exits non-zero for inactive/failed, so the output string is authoritative over
// the exit code; any exec error (binary missing, etc.) maps to "unknown".
func (p SystemctlProber) Status(ctx context.Context) (string, error) {
unit := p.Unit
if unit == "" {
unit = "cloudflared"
const tunnelInspect = `{{.State.Status}}|{{.State.ExitCode}}|{{if .State.Health}}{{.State.Health.Status}}{{else}}none{{end}}`
// Status probes every customer guest and reports the worst state (normally there is exactly one guest).
func (p GuestTunnelProber) Status(ctx context.Context) (string, string) {
if p.Runner == nil || p.Guests == nil {
return TunnelUnknown, "no probe wired"
}
out, _ := exec.CommandContext(ctx, "systemctl", "is-active", unit).Output()
switch strings.TrimSpace(string(out)) {
case "active":
return "active", nil
case "failed":
return "failed", nil
case "inactive", "deactivating", "activating":
return "inactive", nil
case "":
return "unknown", nil // no output → systemctl/exec problem
default:
return "unknown", nil
vmids, err := p.Guests(ctx)
if err != nil {
return TunnelUnknown, "could not list the customer guest: " + err.Error()
}
if len(vmids) == 0 {
return TunnelUnknown, "no running customer guest"
}
worst, wdetail := "", ""
rank := map[string]int{TunnelRunning: 0, TunnelUnknown: 1, TunnelNotRunning: 2}
for _, v := range vmids {
out, errOut, err := p.Runner.Run(ctx, "/usr/sbin/pct", "exec", fmt.Sprint(v), "--", "docker", "inspect", "-f", tunnelInspect, "cloudflared")
st, d := ClassifyTunnel(string(out), string(errOut), err)
if len(vmids) > 1 {
d = fmt.Sprintf("guest %d: %s", v, d)
}
if worst == "" || rank[st] > rank[worst] {
worst, wdetail = st, d
}
}
return worst, wdetail
}
// ClassifyTunnel maps one `docker inspect` answer to a state. Pure; pinned by TestClassifyTunnel.
func ClassifyTunnel(stdout, stderr string, err error) (string, string) {
out := strings.TrimSpace(stdout)
if err != nil || out == "" {
if strings.Contains(stderr, "No such object") || strings.Contains(stderr, "No such container") {
return TunnelNotRunning, "no cloudflared container in the guest"
}
return TunnelUnknown, "could not ask the guest: " + firstLine(stderr, err)
}
parts := strings.Split(out, "|")
if len(parts) != 3 {
return TunnelUnknown, "unreadable docker answer: " + out
}
state, code, health := parts[0], parts[1], parts[2]
if state != "running" {
return TunnelNotRunning, fmt.Sprintf("container %s, exit code %s", state, code)
}
switch health {
case "healthy":
return TunnelRunning, "connected"
case "unhealthy":
return TunnelNotRunning, "container running but the tunnel is NOT connected (cloudflared /ready fails)"
case "starting":
return TunnelUnknown, "container running, readiness check still starting"
case "none":
return TunnelRunning, "container running (no readiness check on this controller — connection not checked)"
}
return TunnelUnknown, "unknown health state " + health
}
func firstLine(stderr string, err error) string {
s := strings.TrimSpace(stderr)
if i := strings.IndexByte(s, '\n'); i >= 0 {
s = s[:i]
}
if s == "" && err != nil {
s = err.Error()
}
if len(s) > 160 {
s = s[:160]
}
return s
}
+67
View File
@@ -0,0 +1,67 @@
package hub
import (
"context"
"errors"
"io"
"strings"
"testing"
)
// R-841: the three states from one `docker inspect` answer. Red-proof: map "unhealthy" to running (the container
// state alone — what a plain "is it running" probe would say) and the "running but not connected" case fails.
func TestClassifyTunnel(t *testing.T) {
cases := []struct {
name, out, errOut string
err error
want string
detail string
}{
{"connected", "running|0|healthy\n", "", nil, TunnelRunning, "connected"},
{"running but not connected", "running|0|unhealthy\n", "", nil, TunnelNotRunning, "NOT connected"},
{"stopped", "exited|137|unhealthy\n", "", nil, TunnelNotRunning, "exit code 137"},
{"absent", "", "Error: No such object: cloudflared", errors.New("exit status 1"), TunnelNotRunning, "no cloudflared container"},
{"still starting", "running|0|starting\n", "", nil, TunnelUnknown, "starting"},
{"no health check (older controller)", "running|0|none\n", "", nil, TunnelRunning, "connection not checked"},
{"guest not running", "", "CT 9201 not running", errors.New("exit status 255"), TunnelUnknown, "could not ask"},
{"sudo refused", "", "sudo: a password is required", errors.New("exit status 1"), TunnelUnknown, "could not ask"},
}
for _, c := range cases {
st, d := ClassifyTunnel(c.out, c.errOut, c.err)
if st != c.want || !strings.Contains(d, c.detail) {
t.Errorf("%s: got %q (%s), want %q (…%s…)", c.name, st, d, c.want, c.detail)
}
}
}
type tunnelRunner struct {
calls []string
out map[string]string
}
func (r *tunnelRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
line := name + " " + strings.Join(args, " ")
r.calls = append(r.calls, line)
return []byte(r.out[args[1]]), nil, nil
}
func (r *tunnelRunner) RunStdin(ctx context.Context, _ io.Reader, name string, args ...string) ([]byte, []byte, error) {
return r.Run(ctx, name, args...)
}
// The probe uses EXACTLY the existing sudoers shape `pct exec <vmid> -- docker inspect -f <tmpl> cloudflared`, and
// with no customer guest it is unknown, never down.
func TestGuestTunnelProber(t *testing.T) {
r := &tunnelRunner{out: map[string]string{"9201": "running|0|healthy"}}
p := GuestTunnelProber{Runner: r, Guests: func(context.Context) ([]int, error) { return []int{9201}, nil }}
if st, d := p.Status(context.Background()); st != TunnelRunning || d != "connected" {
t.Fatalf("got %q %q", st, d)
}
want := "/usr/sbin/pct exec 9201 -- docker inspect -f " + tunnelInspect + " cloudflared"
if len(r.calls) != 1 || r.calls[0] != want {
t.Fatalf("command = %q, want %q", r.calls, want)
}
none := GuestTunnelProber{Runner: r, Guests: func(context.Context) ([]int, error) { return nil, nil }}
if st, _ := none.Status(context.Background()); st != TunnelUnknown {
t.Fatalf("no guest → %q, want unknown", st)
}
}
+271 -36
View File
@@ -4,10 +4,13 @@ import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"log/slog"
"os"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
@@ -47,6 +50,22 @@ type RestoreTestReporter interface {
RestoreTests(ctx context.Context) []RestoreTest
}
// ProvenRestoreTestReporter is the DURABLE half of the restore-test signal (R-189).
//
// RestoreTestReporter above is backed by an in-memory store whose own comment used to read "lost on
// restart; the cadence re-populates". That was true while a timer re-tested every tier daily. It
// stopped being true on 2026-08-03: under per-archive due-ness the agent will not re-test an archive
// it has already proven, so a proof lost to a restart is not repeated for a whole archive generation
// — a week on the offsite tier — and the hub reports the tier unproven the entire time.
//
// Observed, not predicted: a real 14.5 GB offsite restore passed at 15:25:14, the agent was restarted
// 2 m 43 s later for a deploy, and the hub logged `0 restore-tests` on the next two reports.
//
// (*backup.RestoreTestState).ProvenRestoreTests satisfies this. nil → the merge is a no-op.
type ProvenRestoreTestReporter interface {
ProvenRestoreTests(ctx context.Context) []RestoreTest
}
// PBSReporter is the slice-6-Phase-B seam the pbs verify loop plugs into (same pattern).
// Returns the agent's latest-known PBS snapshot inventory + verify-state. nil → empty.
type PBSReporter interface {
@@ -71,30 +90,43 @@ type GuestNetReporter interface {
GuestNetStatus(ctx context.Context) *GuestNetStatus
}
// GuestDiskTrimReporter is the R-444 seam the weekly trim job plugs into (same consumer-side pattern — hub does not
// import fstrim). nil (feature not wired) → no guest_disk_trim stanza.
type GuestDiskTrimReporter interface {
GuestDiskTrimStatus(ctx context.Context) *GuestDiskTrimStatus
}
// Collector builds a HostReport from read-only sources. All deps are behind narrow
// interfaces for unit testing.
type Collector struct {
px proxmoxReader
cf CloudflaredProber
storage StorageObserver
backups BackupReporter
restoreTests RestoreTestReporter
pbs PBSReporter
temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp)
capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty)
leafFP string // v0.48.0: served local-API leaf fp (static per process; "" when local API disabled)
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
oob OOBReporter // H1: operator-access health (nil → stanza omitted)
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
hostID string
agentVersion string
logger *slog.Logger
now func() time.Time
px proxmoxReader
cf CloudflaredProber
storage StorageObserver
backups BackupReporter
restoreTests RestoreTestReporter
provenTests ProvenRestoreTestReporter
pbs PBSReporter
temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp)
capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty)
leafFP string // v0.48.0: served local-API leaf fp (static per process; "" when local API disabled)
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
diskTrim GuestDiskTrimReporter // R-444: weekly guest disk trim (nil → stanza omitted)
foreignKey ForeignKeyArchiveReporter // R-366 slice 2 (nil → omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
oob OOBReporter // H1: operator-access health (nil → stanza omitted)
system SystemReporter // R-852: the box versions (nil → API fields only)
bundleRecordPath string // R-840: test seam; "" = BundleRecordPath
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
hostID string
agentVersion string
selfSHA func() string // R-349: sha256 of the running binary; default runningBinarySHA256
logger *slog.Logger
now func() time.Time
}
// NewCollector builds a collector. hostID echoes config.Hub.HostID; agentVersion is
@@ -113,6 +145,7 @@ func NewCollector(px proxmoxReader, cf CloudflaredProber, storage StorageObserve
temp: SysfsTempReader{}, // slice 9: real sysfs reader by default; tests inject a fake
hostID: hostID,
agentVersion: agentVersion,
selfSHA: runningBinarySHA256,
logger: logger,
now: func() time.Time { return time.Now().UTC() },
}
@@ -178,6 +211,17 @@ func (c *Collector) SetPBSDRReporter(p PBSDRReporter) *Collector {
return c
}
// ControllerSupervisorReporter is the R-523 seam (satisfied by *localapi.Server).
type ControllerSupervisorReporter interface {
ControllerSupervisorStatus(ctx context.Context) *ControllerSupervisorStatus
}
// SetControllerSupervisorReporter wires the R-523 controller supervisor as a report source (nil-safe).
func (c *Collector) SetControllerSupervisorReporter(r ControllerSupervisorReporter) *Collector {
c.ctrlSup = r
return c
}
// SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza
// omitted). Returns the collector for chaining.
func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
@@ -185,6 +229,24 @@ func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
return c
}
// ForeignKeyArchiveReporter is the R-366 slice-2 seam: the restore-test's ledger of archives written with another
// key. nil = not evaluated yet (the stanza is omitted and the hub keeps its state).
type ForeignKeyArchiveReporter interface {
ForeignKeyArchives(ctx context.Context) *ForeignKeyArchivesStanza
}
// SetForeignKeyArchiveReporter wires the restore-test's foreign-key ledger (R-366 slice 2; nil-safe → omitted).
func (c *Collector) SetForeignKeyArchiveReporter(r ForeignKeyArchiveReporter) *Collector {
c.foreignKey = r
return c
}
// SetGuestDiskTrimReporter wires the R-444 weekly trim job as a report source (nil-safe → stanza omitted).
func (c *Collector) SetGuestDiskTrimReporter(r GuestDiskTrimReporter) *Collector {
c.diskTrim = r
return c
}
// SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side
// pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report.
type SelfUpdateReporter interface {
@@ -219,6 +281,69 @@ type OOBReporter interface {
}
// SetOOBReporter wires the operator-access health source (H1; nil-safe → stanza omitted).
// SystemReporter reads the box's versions (R-852): the customer guest's vmid and the wrapper's raw facts.
type SystemReporter interface {
SystemFacts(ctx context.Context) (vmid int, facts json.RawMessage, err error)
}
// SetSystemReporter wires the facts read (agent v0.142.0). Without it the stanza carries the Proxmox API fields only.
func (c *Collector) SetSystemReporter(r SystemReporter) *Collector {
c.system = r
return c
}
// BundleRecordPath is the root-owned record felhom-os-apply writes after a config bundle installs (R-840).
const BundleRecordPath = "/etc/felhom/config-bundle.json"
// readBundleRecord returns the record's summary: version/sha/installed_at, "none" when the file is absent, "unknown"
// when it cannot be read or parsed (never a guess).
func readBundleRecord(path string) json.RawMessage {
if path == "" {
path = BundleRecordPath
}
b, err := os.ReadFile(path)
if os.IsNotExist(err) {
return json.RawMessage(`{"version":"none"}`)
}
var rec struct {
AgentVersion string `json:"agent_version"`
BundleSHA256 string `json:"bundle_sha256"`
InstalledAt string `json:"installed_at"`
Authority string `json:"authority"`
}
if err != nil || json.Unmarshal(b, &rec) != nil || rec.AgentVersion == "" {
return json.RawMessage(`{"version":"unknown"}`)
}
out, _ := json.Marshal(map[string]string{"version": rec.AgentVersion, "bundle_sha256": rec.BundleSHA256,
"installed_at": rec.InstalledAt, "authority": rec.Authority})
return out
}
func unknownIfEmpty(s string) string {
if strings.TrimSpace(s) == "" {
return "unknown"
}
return s
}
// systemInfo builds the `system` stanza. Never fatal: a failed facts read is FactsError, the API fields stay.
func (c *Collector) systemInfo(ctx context.Context, ns proxmox.NodeStatus) *SystemInfo {
si := &SystemInfo{PVEVersion: unknownIfEmpty(ns.PVEVersion), KernelVersion: unknownIfEmpty(ns.KVersion),
ReadAt: c.now().Format(time.RFC3339)}
si.ConfigBundle = readBundleRecord(c.bundleRecordPath)
if c.system == nil {
si.FactsError = "no facts reader wired"
return si
}
vmid, f, err := c.system.SystemFacts(ctx)
si.VMID, si.Facts = vmid, f
if err != nil {
si.FactsError = err.Error()
c.logger.Debug("hub: system facts unavailable", "err", err)
}
return si
}
func (c *Collector) SetOOBReporter(o OOBReporter) *Collector {
c.oob = o
return c
@@ -241,6 +366,7 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
HostID: c.hostID,
ReportedAt: c.now().Format(time.RFC3339),
AgentVersion: c.agentVersion,
AgentSHA256: c.agentSHA256(),
Host: host,
Guests: c.collectGuests(ctx),
// storage_targets populated this slice (slice 5) via the observer; the rest stay
@@ -251,10 +377,11 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
PBSSnapshots: c.collectPBSSnapshots(ctx),
AuditTail: []AuditEntry{},
Cloudflared: Cloudflared{Status: c.cloudflaredStatus(ctx)},
Cloudflared: c.cloudflared(ctx),
Capabilities: c.capabilities(ctx),
LeafFingerprint: c.leafFP,
Addresses: c.collectAddresses(),
System: c.systemInfo(ctx, ns),
}
// DR recipe host-half — derived from the just-collected guest/storage/PBS facts (no new reads).
// Secret-free by construction (identifiers/intents/sizes/coordinates only).
@@ -268,10 +395,22 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.pbsdr != nil {
report.PBSDR = c.pbsdr.PBSDRStatus(ctx)
}
// R-523: controller supervisor record (nil reporter = not wired → stanza omitted).
if c.ctrlSup != nil {
report.ControllerSupervisor = c.ctrlSup.ControllerSupervisorStatus(ctx)
}
// R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted).
if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
}
// R-444: the last weekly trim result per guest (nil reporter = not wired → stanza omitted).
if c.diskTrim != nil {
report.GuestDiskTrim = c.diskTrim.GuestDiskTrimStatus(ctx)
}
// R-366 slice 2: archives the restore-test skipped as another key's (nil → not evaluated yet → omitted).
if c.foreignKey != nil {
report.ForeignKeyArchives = c.foreignKey.ForeignKeyArchives(ctx)
}
// D1: agent self-update pending status (nil reporter → pending=false, the steady state).
if c.selfUpdate != nil {
report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending()
@@ -330,8 +469,28 @@ const pbsWrapperPath = "/usr/local/sbin/felhom-pbs-apply"
// unreadable file yields "", which the hub reads as UNKNOWN rather than as drift — a host that
// legitimately has no DR wrapper must not light up amber. The file is 0755, so no privilege is
// needed to read it.
func pbsWrapperSHA256() string {
f, err := os.Open(pbsWrapperPath)
func pbsWrapperSHA256() string { return fileSHA256(pbsWrapperPath) }
// selfExePath is the running binary as the kernel holds it. /proc/self/exe, not the installed path:
// after an A/B flip the file at /usr/local/bin/felhom-agent may already be the NEXT binary while this
// process still runs the old one, and the report must describe what runs (R-349). Test seam.
var selfExePath = "/proc/self/exe"
// runningBinarySHA256 hashes the running binary ONCE per process — the bytes cannot change under a
// running process, and re-hashing ~20 MB every report cycle buys nothing. A failed read is cached as
// "" (UNKNOWN); it never fails the report.
var runningBinarySHA256 = sync.OnceValue(func() string { return fileSHA256(selfExePath) })
func (c *Collector) agentSHA256() string {
if c.selfSHA == nil {
return ""
}
return c.selfSHA()
}
// fileSHA256 is the hex sha256 of a file's bytes, or "" when it cannot be read.
func fileSHA256(path string) string {
f, err := os.Open(path)
if err != nil {
return ""
}
@@ -427,16 +586,90 @@ func (c *Collector) collectBackups(ctx context.Context) []Backup {
return []Backup{}
}
// collectRestoreTests merges the in-memory result with the PERSISTED per-tier proofs (R-189).
//
// The rule is ONE ENTRY PER TIER, NEWEST WINS, and it falls out of what each source means rather
// than from a preference between them:
//
// - the in-memory store holds this process's latest run, pass OR fail. A failure exists nowhere
// else and must always reach the hub — a failing tier is retried at the next evaluation, so its
// record is short-lived by design;
// - the persisted state holds the last SUCCESS per tier and survives a restart.
//
// Comparing by TestedAt gives the right answer in every case without special-casing: a fresh failure
// beats an older stored success (the failure is the news), a stored success beats a stale in-memory
// entry after a restart, and a tier proved twice never appears twice — two entries for one tier would
// read at the hub as two tests.
//
// A tier with no usable persisted proof contributes NOTHING. Reporting an unproven tier as proven
// would be a worse defect than the one this closes.
func (c *Collector) collectRestoreTests(ctx context.Context) []RestoreTest {
if c.restoreTests == nil {
return []RestoreTest{}
out := []RestoreTest{}
if c.restoreTests != nil {
if r := c.restoreTests.RestoreTests(ctx); r != nil {
out = append(out, r...)
}
}
if r := c.restoreTests.RestoreTests(ctx); r != nil {
return r
if c.provenTests == nil {
return out
}
return []RestoreTest{}
// Index what we already have by tier, keeping the newest per tier.
best := map[string]int{} // tier → index into out
for i, rt := range out {
if rt.SourceTier == "" {
continue // untiered entry: never deduped, never overwritten — we cannot say what it is
}
if j, seen := best[rt.SourceTier]; !seen || newerRestoreTest(rt, out[j]) {
best[rt.SourceTier] = i
}
}
for _, p := range c.provenTests.ProvenRestoreTests(ctx) {
if p.SourceTier == "" {
continue // not usable as a per-tier proof; the state layer already filters these
}
i, seen := best[p.SourceTier]
if !seen {
out = append(out, p)
best[p.SourceTier] = len(out) - 1
continue
}
if newerRestoreTest(p, out[i]) {
out[i] = p
}
}
return out
}
// newerRestoreTest reports whether a was tested after b. An unparseable or absent timestamp is
// treated as OLDER, so a malformed entry can never displace a good one.
func newerRestoreTest(a, b RestoreTest) bool {
ta, aok := parseRestoreTestedAt(a.TestedAt)
tb, bok := parseRestoreTestedAt(b.TestedAt)
if !aok {
return false
}
if !bok {
return true
}
return ta.After(tb)
}
func parseRestoreTestedAt(s string) (time.Time, bool) {
t, err := time.Parse(time.RFC3339, s)
if err != nil {
return time.Time{}, false
}
return t.UTC(), true
}
// SetProvenRestoreTests wires the durable proof source. It is a setter rather than a constructor
// argument because the persisted state is opened later in main() than the collector is built; the
// same shape as the other late-wired seams here. **The wiring is asserted by an AST test** — the
// method it feeds carried a doc comment naming a "host-report gauge" for weeks with no caller at
// all, and this fix must not become the next instance of that.
func (c *Collector) SetProvenRestoreTests(p ProvenRestoreTestReporter) { c.provenTests = p }
// collectPBSSnapshots reads the latest PBS snapshot inventory via the seam (nil → empty).
func (c *Collector) collectPBSSnapshots(ctx context.Context) []PBSSnapshot {
if c.pbs == nil {
@@ -448,16 +681,18 @@ func (c *Collector) collectPBSSnapshots(ctx context.Context) []PBSSnapshot {
return []PBSSnapshot{}
}
func (c *Collector) cloudflaredStatus(ctx context.Context) string {
func (c *Collector) cloudflared(ctx context.Context) Cloudflared {
if c.cf == nil {
return "unknown"
return Cloudflared{Status: TunnelUnknown, Detail: "no probe wired"}
}
st, err := c.cf.Status(ctx)
if err != nil || st == "" {
c.logger.Warn("hub: cloudflared probe failed", "err", err)
return "unknown"
st, d := c.cf.Status(ctx)
if st == "" {
st = TunnelUnknown
}
return st
if st == TunnelUnknown {
c.logger.Debug("hub: tunnel probe could not decide", "detail", d)
}
return Cloudflared{Status: st, Detail: d}
}
func percent(used, total int64) float64 {
+64
View File
@@ -0,0 +1,64 @@
package hub
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
)
// R-349: the report carries the sha256 of the binary that is RUNNING, so the hub can tell a
// hand-built proof binary from the vouched artifact of the same version string. The consequence
// asserted: the wire field equals the hash of this very test binary's bytes (read independently via
// os.Executable, a different channel from /proc/self/exe), and it is on the wire as agent_sha256.
func TestCollect_AgentSHA256IsTheRunningBinary(t *testing.T) {
exe, err := os.Executable()
if err != nil {
t.Skipf("os.Executable: %v", err)
}
raw, err := os.ReadFile(exe)
if err != nil {
t.Fatalf("read own binary: %v", err)
}
sum := sha256.Sum256(raw)
want := hex.EncodeToString(sum[:])
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
if r.AgentSHA256 != want {
t.Fatalf("agent_sha256 = %q, want the running binary's %q", r.AgentSHA256, want)
}
b, err := json.Marshal(r)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(b), `"agent_sha256":"`+want+`"`) {
t.Fatalf("agent_sha256 not on the wire: %s", b)
}
}
// An unreadable binary is UNKNOWN (empty, omitted) — never a made-up hash, never a failed report.
func TestFileSHA256_UnreadableIsEmpty(t *testing.T) {
if got := fileSHA256(filepath.Join(t.TempDir(), "absent")); got != "" {
t.Fatalf("absent file hashed to %q, want empty", got)
}
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
c.selfSHA = func() string { return "" }
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect must not fail on an unreadable binary: %v", err)
}
b, _ := json.Marshal(r)
if strings.Contains(string(b), "agent_sha256") {
t.Fatalf("empty agent_sha256 must be omitted: %s", b)
}
}
+54
View File
@@ -0,0 +1,54 @@
package hub
import (
"context"
"encoding/json"
"testing"
)
// R-444: the guest_disk_trim stanza must reach a report built through the PRODUCTION collect path, be absent from
// the wire when the job is not wired, and carry the keys the hub's System page reads.
type fakeDiskTrim struct{ st *GuestDiskTrimStatus }
func (f fakeDiskTrim) GuestDiskTrimStatus(context.Context) *GuestDiskTrimStatus { return f.st }
func TestCollect_GuestDiskTrim(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.150.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ := json.Marshal(r)
var m map[string]any
_ = json.Unmarshal(b, &m)
if _, ok := m["guest_disk_trim"]; ok {
t.Fatalf("guest_disk_trim on the wire with no reporter wired: %s", b)
}
c.SetGuestDiskTrimReporter(fakeDiskTrim{st: &GuestDiskTrimStatus{Schedule: "weekly", Guests: []GuestDiskTrim{{
VMID: 9201, LastAttemptAt: "2026-10-07T08:30:00Z", OK: true, BytesTrimmed: 90143313920, Mounts: 2,
DurationSeconds: 24.4, LastOKAt: "2026-10-07T08:30:00Z",
}}}})
r, err = c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ = json.Marshal(r)
m = nil
_ = json.Unmarshal(b, &m)
dt, ok := m["guest_disk_trim"].(map[string]any)
if !ok || dt["schedule"] != "weekly" {
t.Fatalf("guest_disk_trim missing or wrong on the wire: %s", b)
}
g := dt["guests"].([]any)[0].(map[string]any)
for _, k := range []string{"vmid", "last_attempt_at", "ok", "bytes_trimmed", "mounts", "duration_seconds", "last_ok_at"} {
if _, ok := g[k]; !ok {
t.Fatalf("guest_disk_trim.guests[0] lacks %q: %v", k, g)
}
}
if g["bytes_trimmed"] != float64(90143313920) || g["ok"] != true {
t.Fatalf("values did not survive the round trip: %v", g)
}
}
+2 -2
View File
@@ -16,7 +16,7 @@ func (f fakeGuestNet) GuestNetStatus(context.Context) *GuestNetStatus { return f
func TestCollect_GuestNetOmittedWhenReporterNil(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.92.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.92.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
@@ -38,7 +38,7 @@ func TestCollect_GuestNetOmittedWhenReporterNil(t *testing.T) {
func TestCollect_GuestNetPopulatedWhenWired(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.92.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.92.0", quietLogger())
c.SetGuestNetReporter(fakeGuestNet{st: &GuestNetStatus{
CheckedAt: "2026-07-21T10:00:00Z",
Guests: []GuestNetGuest{{
+4 -4
View File
@@ -12,7 +12,7 @@ func (f fakeMgmtPlane) MgmtPlaneStatus(context.Context) *MgmtPlaneStatus { retur
func TestCollect_MgmtPlaneOmittedWhenReporterNil(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.71.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.71.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
@@ -24,7 +24,7 @@ func TestCollect_MgmtPlaneOmittedWhenReporterNil(t *testing.T) {
func TestCollect_MgmtPlanePopulatedWhenWired(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.71.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.71.0", quietLogger())
c.SetMgmtPlaneReporter(fakeMgmtPlane{st: &MgmtPlaneStatus{
PrivsepDirOK: true, SshdReachable: true, HealedRecently: true, PrivsepHealedAt: "2026-07-05T16:42:17Z",
}})
@@ -47,7 +47,7 @@ func (f fakeOOB) OOBStatus(context.Context) *OOBStatus { return f.st }
func TestCollect_OOBOmittedWhenNil(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.72.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.72.0", quietLogger())
r, _ := c.Collect(context.Background())
if r.OOB != nil {
t.Fatalf("no reporter → oob omitted, got %+v", r.OOB)
@@ -56,7 +56,7 @@ func TestCollect_OOBOmittedWhenNil(t *testing.T) {
func TestCollect_OOBPopulatedWhenWired(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.72.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.72.0", quietLogger())
c.SetOOBReporter(fakeOOB{st: &OOBStatus{FelhomSshdActive: true, FelhomSshdPort: 8822, Reachable: true}})
r, _ := c.Collect(context.Background())
if r.OOB == nil || r.OOB.FelhomSshdPort != 8822 || !r.OOB.Reachable {
+9 -9
View File
@@ -33,7 +33,7 @@ func TestCollect_StorageTargetsFromObserver(t *testing.T) {
obs := fakeObserver{targets: []StorageTarget{
{Name: "local-lvm", Type: StorageTypeLVMThin, State: StorageStateAttached, Reachable: true},
}}
c := NewCollector(px, fakeProber{status: "active"}, obs, nil, nil, nil, "h", "0.5.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, obs, nil, nil, nil, "h", "0.5.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
@@ -45,7 +45,7 @@ func TestCollect_StorageTargetsFromObserver(t *testing.T) {
func TestCollect_StorageObserverErrorDegradesToEmpty(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{err: errors.New("proxmox down")}, nil, nil, nil, "h", "0.5.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{err: errors.New("proxmox down")}, nil, nil, nil, "h", "0.5.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("a storage observe error must not sink the heartbeat: %v", err)
@@ -64,7 +64,7 @@ func TestCollect_HostAndGuests(t *testing.T) {
},
cfg: map[int]proxmox.GuestConfig{100: {Cores: 2, Memory: 2048}},
}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "demo-host-01", "0.3.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "demo-host-01", "0.3.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
@@ -88,7 +88,7 @@ func TestCollect_HostAndGuests(t *testing.T) {
if g.Spec.Cores != 2 || g.Spec.MemoryBytes != 2147483648 || g.Spec.DiskBytes != 21474836480 {
t.Errorf("spec = %+v", g.Spec)
}
if r.Cloudflared.Status != "active" {
if r.Cloudflared.Status != "running" || r.Cloudflared.Detail != "connected" {
t.Errorf("cloudflared = %q", r.Cloudflared.Status)
}
}
@@ -104,7 +104,7 @@ func TestCollect_GuestConfigFailureKeepsStatusOmitsSpec(t *testing.T) {
cfg: map[int]proxmox.GuestConfig{100: {Cores: 2}},
cfgErr: map[int]error{200: errors.New("config read failed")},
}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.3.1", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.3.1", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("a per-guest failure must NOT fail the whole report: %v", err)
@@ -125,7 +125,7 @@ func TestCollect_GuestConfigFailureKeepsStatusOmitsSpec(t *testing.T) {
func TestCollect_NodeStatusFailureIsHardError(t *testing.T) {
px := &fakePx{node: "n", nsErr: errors.New("proxmox down")}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
if _, err := c.Collect(context.Background()); err == nil {
t.Fatal("NodeStatus failure must be a hard error (no useful report)")
}
@@ -133,7 +133,7 @@ func TestCollect_NodeStatusFailureIsHardError(t *testing.T) {
func TestCollect_CloudflaredProbeErrorIsUnknown(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{err: errors.New("no systemctl")}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
c := NewCollector(px, fakeProber{status: "", detail: "could not ask"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("cloudflared failure must not be fatal: %v", err)
@@ -153,7 +153,7 @@ func TestCollect_LeafFingerprint(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
const fp = "60b5974d586f5f3c8ec41eb998d0f07406178219c36bf6d3ff377570279d8245"
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.48.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.48.0", quietLogger())
c.SetLeafFingerprint(fp)
r, err := c.Collect(context.Background())
if err != nil {
@@ -164,7 +164,7 @@ func TestCollect_LeafFingerprint(t *testing.T) {
}
// Companion: no SetLeafFingerprint (local API disabled) → empty, never a fabricated value.
c2 := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.48.0", quietLogger())
c2 := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.48.0", quietLogger())
r2, _ := c2.Collect(context.Background())
if r2.LeafFingerprint != "" {
t.Fatalf("unset leaf_fingerprint = %q, want empty", r2.LeafFingerprint)
+9 -7
View File
@@ -54,11 +54,13 @@ const (
DRReasonNoPBSStorage = "no_pbs_storage_observed"
)
// PBSRootNamespace is how the recipe spells PBS's root namespace. The PBS API spells it as the EMPTY
// string (and `pct restore --ns root` would name a namespace that does not exist) — "root" is a display
// convention this wire has always used, kept here so the field's meaning did not change under R-106.
// Only a box with no `namespace` line in its pbs storage.cfg stanza ever emits it.
const PBSRootNamespace = "root"
// PBSRootNamespace is how the recipe spells PBS's root namespace: the EMPTY string, PBS's own spelling (R-124,
// agent v0.147.0). It used to be the display word "root", which no PBS namespace is named — an operator pasting it
// into `proxmox-backup-client … --ns root` during a real recovery got a failure. An empty namespace is ambiguous on
// its own, so READ IT WITH namespace_state: resolved + "" = the root namespace (pass no --ns, or --ns ""); unknown +
// "" = the agent could not tell. Only a box with no `namespace` line in its pbs storage.cfg stanza emits it.
// Pinned by TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown and TestR124_RootNamespaceOnTheWireIsPBSSpelling.
const PBSRootNamespace = ""
// DRRecipeHostHalf is the agent-emitted half (guest/drive/storage/PBS scaffolding). Derived entirely
// from facts the report already collects — no new privileged reads.
@@ -123,8 +125,8 @@ type DRPBSCoord struct {
RepoID string `json:"repo_id"` // the PVE pbs storage id (e.g. "felhom-pbs") — not a token
// Namespace is the PBS namespace the restore targets, resolved from the pbs storage's storage.cfg
// stanza — the same field `vzdump --storage <pbs>` makes PVE read, so the recipe cannot disagree
// with the backup that produced the snapshot. PBSRootNamespace when the box has no namespace
// configured; "" when NamespaceState is unknown.
// with the backup that produced the snapshot. PBSRootNamespace ("", PBS's spelling, R-124) when the box has no
// namespace configured; also "" when NamespaceState is unknown — consult NamespaceState.
//
// R-106: this used to come from the listed snapshot's own `ns`, which PBS does not echo per item once
// the request is already namespace-scoped via `?ns=` (internal/pbs/client.go). The field was
+41 -5
View File
@@ -63,8 +63,8 @@ func TestBuildDRRecipeHostHalf(t *testing.T) {
t.Error("felhom-flash (local-dir user-data drive) missing from drives")
}
// pbs: latest snapshot's coords + the pbs storage id as repo_id.
if h.PBS == nil || h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != "root" || h.PBS.LatestSnapshotID != "9201" {
t.Errorf("pbs coord = %+v, want repo felhom-pbs/root/9201", h.PBS)
if h.PBS == nil || h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != PBSRootNamespace || h.PBS.LatestSnapshotID != "9201" {
t.Errorf("pbs coord = %+v, want repo felhom-pbs, the root namespace (\"\", R-124), snapshot 9201", h.PBS)
}
}
@@ -252,7 +252,7 @@ func TestDRRecipe_PBSNamespaceIsThePerCustomerOne(t *testing.T) {
}
// TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown: a box with a pbs storage and NO namespace line is
// genuinely in the root namespace. That is an answer, not a gap — it must read resolved/"root", so the
// genuinely in the root namespace. That is an answer, not a gap — it must read resolved/"" (PBS's spelling, R-124), so the
// honest root case is never confused with "I could not tell".
func TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown(t *testing.T) {
h := BuildDRRecipeHostHalf(nil,
@@ -384,7 +384,7 @@ func TestCollectDRRecipe_ProductionPath(t *testing.T) {
obs := fakeObserver{targets: capturedDemoFelhomTargets()}
pbsRep := fakePBSReporter{snaps: capturedDemoFelhomSnapshots()}
c := NewCollector(px, fakeProber{status: "active"}, obs, nil, nil, pbsRep, "h", "0.118.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, obs, nil, nil, pbsRep, "h", "0.118.0", quietLogger())
c.SetBackupTargetResolver(func() ConfiguredBackupTarget {
return ConfiguredBackupTarget{StorageID: "felhom-backup", Known: true}
})
@@ -409,7 +409,7 @@ func TestCollectDRRecipe_ProductionPath(t *testing.T) {
// test that would have caught shipping the seam without wiring it.
func TestCollectDRRecipe_UnwiredSeamReportsUnknown(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{targets: capturedDemoFelhomTargets()},
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{targets: capturedDemoFelhomTargets()},
nil, nil, nil, "h", "0.118.0", quietLogger())
r, err := c.Collect(context.Background())
@@ -453,3 +453,39 @@ func assertNoSecretKeys(t *testing.T, jsonBytes []byte) {
}
walk("<root>", v)
}
// R-124: on the WIRE the root namespace is PBS's own spelling — an empty string, present (not omitted), beside
// namespace_state "resolved". The display word "root" names no PBS namespace, and `--ns root` fails in a recovery.
// RED-PROOF: set PBSRootNamespace back to "root" → this test fails.
func TestR124_RootNamespaceOnTheWireIsPBSSpelling(t *testing.T) {
h := BuildDRRecipeHostHalf(nil,
[]StorageTarget{{Name: "felhom-pbs", Type: StorageTypePBS, Content: "backup", PBSNamespace: ""}},
capturedDemoFelhomSnapshots(),
ConfiguredBackupTarget{StorageID: "felhom-pbs", Known: true})
b, err := json.Marshal(h.PBS)
if err != nil {
t.Fatal(err)
}
var m map[string]any
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
ns, present := m["namespace"]
if !present {
t.Fatalf("namespace key missing from %s — an omitted key reads as 'unknown', not 'root'", b)
}
if ns != "" {
t.Fatalf("root namespace on the wire = %q, want \"\" (PBS's spelling; no namespace is named %q)", ns, ns)
}
if m["namespace_state"] != DRStateResolved {
t.Fatalf("namespace_state = %v, want %q beside the empty root namespace", m["namespace_state"], DRStateResolved)
}
// A configured namespace still passes through unchanged.
h2 := BuildDRRecipeHostHalf(nil,
[]StorageTarget{{Name: "felhom-pbs", Type: StorageTypePBS, Content: "backup", PBSNamespace: "demo-felhom"}},
capturedDemoFelhomSnapshots(),
ConfiguredBackupTarget{StorageID: "felhom-pbs", Known: true})
if h2.PBS.Namespace != "demo-felhom" {
t.Fatalf("configured namespace = %q, want demo-felhom", h2.PBS.Namespace)
}
}
+4 -4
View File
@@ -16,7 +16,7 @@ func intp(v int) *int { return &v }
// HostMetricsNow returns a fresh host block with cpu% from NodeStatus and the temp from the reader.
func TestHostMetricsNow_PopulatesTemp(t *testing.T) {
px := &fakePx{node: "demo-felhom", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
SetTempReader(fakeTemp{c: intp(46)})
h, err := c.HostMetricsNow(context.Background())
if err != nil {
@@ -36,7 +36,7 @@ func TestHostMetricsNow_PopulatesTemp(t *testing.T) {
// A missing temp sensor gracefully nulls cpu_temp_c without failing the host read.
func TestHostMetricsNow_GracefulNullTemp(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
SetTempReader(fakeTemp{c: nil})
h, err := c.HostMetricsNow(context.Background())
if err != nil {
@@ -50,7 +50,7 @@ func TestHostMetricsNow_GracefulNullTemp(t *testing.T) {
// A NodeStatus failure is a hard error (no useful host view).
func TestHostMetricsNow_NodeStatusErrorIsHard(t *testing.T) {
px := &fakePx{node: "n", nsErr: errors.New("proxmox down")}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger())
if _, err := c.HostMetricsNow(context.Background()); err == nil {
t.Fatal("NodeStatus failure must be a hard error")
}
@@ -59,7 +59,7 @@ func TestHostMetricsNow_NodeStatusErrorIsHard(t *testing.T) {
// Collect() (the hub report) also carries the temp now — the operator freebie.
func TestCollect_HostReportCarriesTemp(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
SetTempReader(fakeTemp{c: intp(51)})
r, err := c.Collect(context.Background())
if err != nil {
+2 -2
View File
@@ -56,7 +56,7 @@ func (f *fakePx) GuestConfig(ctx context.Context, vmid int) (proxmox.GuestConfig
// fakeProber is a fake CloudflaredProber.
type fakeProber struct {
status string
err error
detail string
}
func (p fakeProber) Status(ctx context.Context) (string, error) { return p.status, p.err }
func (p fakeProber) Status(ctx context.Context) (string, string) { return p.status, p.detail }
+36
View File
@@ -0,0 +1,36 @@
package hub
import (
"encoding/json"
"os"
"testing"
)
// The os_update block is a cross-repo contract: testdata/desired-state-osupdate.golden.json is byte-identical with
// felhom.eu/hub/internal/api/testdata (the hub's TestOSUpdate_DesiredBlockMatchesTheGolden proves the hub SERVES
// it). Here: the agent DECODES every field. A renamed json tag on either side fails one of the two tests.
func TestOSUpdateGolden_Decodes(t *testing.T) {
raw, err := os.ReadFile("testdata/desired-state-osupdate.golden.json")
if err != nil {
t.Fatal(err)
}
var resp DesiredStateResponse
if err := json.Unmarshal(raw, &resp); err != nil {
t.Fatal(err)
}
o := resp.DesiredState.OSUpdate
if o == nil || o.Ring != 1 || !o.Enabled || o.Release == nil {
t.Fatalf("os_update = %+v", o)
}
r := o.Release
if r.ID != "os-guest-20261004-120000" || r.Snapshot != "20261004T120000Z" || len(r.Packages) != 2 ||
r.Packages[1].Name != "openssl" || r.Packages[1].Version != "3.5.7-1~deb13u3" || r.Packages[1].Origin != "Debian-Security" {
t.Fatalf("release = %+v", r)
}
// v0.141.0: the host layer's own approved set (`11` §8 step 3).
h := o.HostRelease
if h == nil || h.ID != "os-host-20261004-120000" || h.Snapshot != "20261004T120000Z" || len(h.Packages) != 1 ||
h.Packages[0].Name != "libssl3t64" || h.Packages[0].Origin != "Debian-Security" {
t.Fatalf("host_release = %+v", h)
}
}
+165 -4
View File
@@ -18,6 +18,16 @@ type HostReport struct {
HostID string `json:"host_id"` // echoes config.Hub.HostID
ReportedAt string `json:"reported_at"` // RFC3339, agent clock
AgentVersion string `json:"agent_version"`
// AgentSHA256 is the sha256 of the binary this process is RUNNING (read through /proc/self/exe,
// once per process), R-349. The version string cannot tell a hand-built proof binary from the
// published, vouched artifact of the same version — same source, different bytes (`-trimpath
// -buildvcs=false` in release-agent.sh) — so self-update sees "already installed" and never
// corrects it. Reporting the bytes lets the hub compare against the vouched agent_sha256, the
// same mechanism host.wrapper_sha256 is for the PBS wrapper (R-50b(a)).
//
// Empty = unreadable, which the hub must treat as UNKNOWN, never as drift. Pinned by
// TestCollect_AgentSHA256IsTheRunningBinary.
AgentSHA256 string `json:"agent_sha256,omitempty"`
Host HostMetrics `json:"host"`
Guests []Guest `json:"guests"`
@@ -68,8 +78,14 @@ type HostReport struct {
// report is stored opaquely hub-side, so these additive fields need no hub-schema change.
// Both are `omitempty` (the Wireguard precedent): in the steady state (no update in flight)
// they are absent — which keeps the cross-repo host-report golden contract byte-stable without
// a hub change. They appear only while an update is pending. The hub reads an absent field as
// pending=false, the correct default.
// a hub change. They appear only while an update is pending.
//
// ⚠ CORRECTED 2026-08-08 (R-260). This comment used to end "The hub reads an absent field as
// pending=false, the correct default." THE HUB HAS NO FIELD FOR EITHER OF THESE, so it reads
// nothing — present or absent — and encoding/json discards them on arrival. The sentence
// described an intent, not the code, and it read as settled for long enough that a sweep had to
// find it. The emission is correct and stays; the missing consumer is tracked as R-264, and
// `felhom.eu/scripts/wire_contract_gate.py` now refuses any NEW field of this shape.
SelfUpdatePending bool `json:"selfupdate_pending,omitempty"`
SelfUpdatePendingVersion string `json:"selfupdate_pending_version,omitempty"`
@@ -83,6 +99,12 @@ type HostReport struct {
// hub-schema change and are absent when the reporter is not wired.
MgmtPlane *MgmtPlaneStatus `json:"mgmt_plane,omitempty"`
// System is the box's versions for the hub's System page (agent v0.142.0, R-852, `09` decision 89): Proxmox and the
// running kernel from the Proxmox API, and the wrapper's read-only facts (host Debian, next-boot kernel, held
// packages, taint, the crash guard; guest Debian, Docker engine, containerd, live-restore). A value nobody could
// read is "unknown", never empty and never guessed. The hub v0.132.0 consumes it (hosts + System pages).
System *SystemInfo `json:"system,omitempty"`
// PBSDR is the PBS-DR-tier bridge status stanza (slice 2). Present only when the pbsdr
// consumer is wired. `consumed_failed` is the LOUD persistent state: the one-time token
// secret was consumed but the apply failed afterwards — the secret is burned, the bridge
@@ -102,6 +124,17 @@ type HostReport struct {
// on HostReport would have been the only report block named against that convention.
GuestNet *GuestNetStatus `json:"guest_net,omitempty"`
// GuestDiskTrim is the weekly guest disk trim stanza (R-444, `09` §3 decision 139): the schedule and, per owned
// guest, the LAST trim result as persisted by the agent (it survives a restart). Present only when the trim job is
// wired; an empty `guests` list means the job runs and no guest has been trimmed yet. No secret.
GuestDiskTrim *GuestDiskTrimStatus `json:"guest_disk_trim,omitempty"`
// ForeignKeyArchives (R-366 slice 2, `09` §3 decision 168): per backup tier, the whole-guest archives the
// restore-test SKIPPED because they were written with another key (an earlier install of this box). This box
// cannot open them; the hub turns a CHANGE of this list into one operator event. Absent = not evaluated yet since
// the agent started (the hub keeps its last state); `tiers: []` = evaluated, none found.
ForeignKeyArchives *ForeignKeyArchivesStanza `json:"foreign_key_archives,omitempty"`
// LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent
// mirror of the controller's report log_tails channel. Present ONLY on the heartbeat
// right after the control envelope requested it (log_tail_requested); consume-once on
@@ -118,6 +151,39 @@ type HostReport struct {
// HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent
// when the feature is not wired (pre-H1) — additive, no hub-schema change.
OOB *OOBStatus `json:"oob,omitempty"`
// ControllerSupervisor (R-523, v0.131.0) is the in-guest controller supervisor's per-guest record:
// how many times the agent restarted a dead controller, when last and why, whether it gave up
// (crash-loop pause) and whether the operator parked it. The hub's ControllerSupervisorChecker
// mints `controller_restarted_by_agent` when last_restart_at MOVES and `controller_crashloop` when
// crashloop_since MOVES — timestamps, not counters, because the record is in-memory and an agent
// restart zeroes the counter. `omitempty`: absent when not wired, so the cross-repo golden stays
// byte-stable. The hub parser is pinned by hub/internal/monitor/controller_supervisor_test.go
// against the JSON TestControllerSupervisorStanza_WireShape pins here.
ControllerSupervisor *ControllerSupervisorStatus `json:"controller_supervisor,omitempty"`
}
// ControllerSupervisorStatus is the R-523 stanza. Carries no secret.
type ControllerSupervisorStatus struct {
Guests []ControllerSupervisorGuest `json:"guests"`
}
// ControllerSupervisorGuest is one supervised guest.
type ControllerSupervisorGuest struct {
VMID int `json:"vmid"`
RestartsTotal int `json:"restarts_total"`
LastRestartAt string `json:"last_restart_at,omitempty"` // RFC3339
LastReason string `json:"last_reason,omitempty"`
Crashloop bool `json:"crashloop"`
CrashloopSince string `json:"crashloop_since,omitempty"` // RFC3339; the last crash-loop, kept after it ends
Parked bool `json:"parked"`
// R-539 (v0.132.0) — the SLOW crash loop. Restarts24h counts restarts the supervisor performed in
// the last 24 hours (persisted, so an agent restart does not reset it); SlowCrashloop is true while
// the last raise is under 24 hours old; SlowCrashloopSince is the raise itself, which the hub keys on
// MOVING (hub v0.117.0 controller_slow_crashloop). It moves at most once per 24 hours.
Restarts24h int `json:"restarts_24h"`
SlowCrashloop bool `json:"slow_crashloop"`
SlowCrashloopSince string `json:"slow_crashloop_since,omitempty"` // RFC3339
}
// PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States:
@@ -153,6 +219,27 @@ type GuestNetGuest struct {
Message string `json:"message,omitempty"`
}
// GuestDiskTrimStatus is the R-444 weekly trim stanza. `schedule` is a plain description of when the job runs (local
// time of the host); `guests` holds one entry per owned guest that has had at least one trim attempt.
type GuestDiskTrimStatus struct {
Schedule string `json:"schedule"`
Guests []GuestDiskTrim `json:"guests,omitempty"`
}
// GuestDiskTrim is one guest's LAST trim attempt. `ok` with `last_attempt_at` is the verdict of that attempt — never
// read the time alone as success; `last_ok_at` is the last attempt that succeeded ("" = never). `bytes_trimmed` is
// the sum of the "(N bytes) trimmed" lines `pct fstrim` printed, over `mounts` mount points.
type GuestDiskTrim struct {
VMID int `json:"vmid"`
LastAttemptAt string `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt string `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
}
type PBSDRStatus struct {
State string `json:"state"`
StorageID string `json:"storage_id,omitempty"`
@@ -202,6 +289,20 @@ type WireguardStatus struct {
AssignedIP string `json:"assigned_ip,omitempty"` // from the marker, e.g. "10.77.0.2/32"
}
// SystemInfo is the `system` stanza (see HostReport.System).
type SystemInfo struct {
PVEVersion string `json:"pve_version"` // GET /nodes/{node}/status pveversion
KernelVersion string `json:"kernel_version"` // GET /nodes/{node}/status kversion
VMID int `json:"vmid,omitempty"` // the customer guest the facts read
Facts json.RawMessage `json:"facts,omitempty"`
FactsError string `json:"facts_error,omitempty"`
ReadAt string `json:"read_at"`
// ConfigBundle is the box's root-owned config bundle record (R-840, agent v0.143.0), read by the agent itself from
// /etc/felhom/config-bundle.json (0644): {"version":"none"} on a box no bundle reached, so an OLD wrapper cannot
// hide it. The wrapper's facts carry the same record plus the drift (files changed by hand).
ConfigBundle json.RawMessage `json:"config_bundle,omitempty"`
}
// HostMetrics is the host block, sourced from proxmox NodeStatus.
type HostMetrics struct {
Node string `json:"node"`
@@ -250,9 +351,10 @@ type GuestSpec struct {
DiskBytes int64 `json:"disk_bytes"`
}
// Cloudflared is the tunnel service health (read-only probe this slice).
// Cloudflared is the box's tunnel (R-841, agent v0.141.0): the cloudflared container in the customer guest.
type Cloudflared struct {
Status string `json:"status"` // active | inactive | failed | unknown
Status string `json:"status"` // running | not_running | unknown (TunnelRunning …)
Detail string `json:"detail,omitempty"` // why not_running / unknown, or "connected"
}
// The following element types are declared now so the empty collections above are
@@ -358,6 +460,17 @@ type SmartSummary struct {
ReallocatedSectors *int `json:"reallocated_sectors"`
PendingSectors *int `json:"pending_sectors"`
OfflineUncorrectable *int `json:"offline_uncorrectable"`
// R-330 (disk health Phase 2): three more SATA raw counters. omitempty + pointer: absent (an
// older agent, an NVMe/USB device, or a drive that does not report the attribute) is OMITTED —
// unknown, never a zero (S-39). Wire only: no verdict reads them yet.
// 187 Reported_Uncorrect — the failing drive's most telling counter (normalized 1 vs thresh 0,
// raw 1001) while SMART still said PASSED.
// 188 Command_Timeout — some vendors PACK several counters into the 48-bit raw value, so the
// number is carried as reported and must not be compared across vendors.
// 199 UDMA_CRC_Error_Count — cabling / link errors, not the medium.
ReportedUncorrect *int64 `json:"reported_uncorrect,omitempty"`
CommandTimeout *int64 `json:"command_timeout,omitempty"`
UDMACRCErrors *int64 `json:"udma_crc_errors,omitempty"`
// NVMe attributes.
CriticalWarning *int `json:"critical_warning"`
@@ -431,6 +544,10 @@ type RestoreTest struct {
// mount layout, not just booted. Additive — a hub that predates them ignores the unknown keys.
MountParity string `json:"mount_parity,omitempty"`
MountInventory []string `json:"mount_inventory,omitempty"`
// Skipped (R-672, v0.133.0): the test did NOT run — the space preflight refused, and Error says
// why ("skipped: not enough space on …"). Pass is false. A hub that predates the key reads a
// failed test with that error, which is the honest reading.
Skipped bool `json:"skipped,omitempty"`
}
// PBSSnapshot is one PBS (offsite) snapshot's inventory + integrity state (doc 03 §8, slice
@@ -505,6 +622,35 @@ type WireDesiredState struct {
RestoreDirective *WireRestoreDirective `json:"restore_directive,omitempty"` // slice 10D (forward-compat)
Wireguard *WireWireguard `json:"wireguard,omitempty"` // S3 (doc 06 §3.2; golden-pinned)
PBSDR *WirePBSDR `json:"pbs_dr,omitempty"` // PBS DR tier (slice 2 consumer)
OSUpdate *WireOSUpdate `json:"os_update,omitempty"` // OS updates, guest fast lane (agent v0.140.0)
}
// WireOSUpdate is the hub-OWNED OS-update block (hub v0.130.0, `11-os-updates.md` §5.3), merged into the served
// document at read time. Ring 0 installs every pending Debian / Debian-Security fix; ring 1 installs exactly the
// newest approved release. Absent (older hub) → the agent treats the box as ring 1, ON, no release: it reports
// and installs nothing. Golden: testdata/desired-state-osupdate.golden.json (byte-identical with the hub's).
type WireOSUpdate struct {
Ring int `json:"ring"`
Enabled bool `json:"enabled"`
Release *WireOSRelease `json:"release,omitempty"`
// HostRelease is the newest approved HOST release (hub v0.131.0, `11` §8 step 3) — a separate set: a version
// approved for the guest is not approved for the host by that fact alone.
HostRelease *WireOSRelease `json:"host_release,omitempty"`
}
// WireOSRelease is an approved version set; Snapshot is the approval time (YYYYMMDDTHHMMSSZ) the wrapper uses
// for snapshot.debian.org when Debian has already replaced a version (decision 79).
type WireOSRelease struct {
ID string `json:"id"`
Snapshot string `json:"snapshot"`
Packages []WireOSPackage `json:"packages"`
}
// WireOSPackage is one approved name=version and its origin ("Debian" | "Debian-Security").
type WireOSPackage struct {
Name string `json:"name"`
Version string `json:"version"`
Origin string `json:"origin"`
}
// WirePBSDR is the hub's PBS-DR-tier descriptor (PBS DR slice 1, hub/internal/web/pbsdr.go
@@ -586,3 +732,18 @@ type WireRestoreDirective struct {
Archive string `json:"archive,omitempty"` // source archive/snapshot to restore from
VMID int `json:"vmid,omitempty"`
}
// ForeignKeyArchivesStanza wraps the per-tier list so "evaluated, none" (`tiers: []`) differs from "not evaluated"
// (the stanza absent) without a null on the wire.
type ForeignKeyArchivesStanza struct {
Tiers []ForeignKeyArchives `json:"tiers"`
}
// ForeignKeyArchives is one tier's count of archives written with another key (R-366 slice 2): the count and the
// newest/oldest archive time (RFC3339, UTC). No key material — the fingerprints stay on the box.
type ForeignKeyArchives struct {
Target string `json:"target"`
Count int `json:"count"`
Oldest string `json:"oldest"`
Newest string `json:"newest"`
}
+213
View File
@@ -0,0 +1,213 @@
package hub
import (
"context"
"testing"
"time"
)
// R-189 — a passing restore-test must survive an agent restart and reach the hub.
//
// THE OBSERVATION THIS EXISTS FOR (2026-08-03, demo-felhom): a real 14.5 GB offsite restore-test
// PASSED at 15:25:14; the agent was restarted 2 m 43 s later for a deploy; the hub logged
// `0 restore-tests` on the next two host-reports. The in-memory store's own comment said "lost on
// restart; the cadence re-populates", which was true under a timer and stopped being true when R-86
// made the agent refuse to re-test an archive it has already proven.
//
// Timestamps here carry JITTER (odd minutes and seconds, not round hours) — yesterday a test was
// hollow because a perfectly regular series landed exactly on a threshold and passed under the
// mutation it was meant to catch.
type fakeLatest struct{ tests []RestoreTest }
func (f *fakeLatest) RestoreTests(context.Context) []RestoreTest { return f.tests }
type fakeProven struct{ tests []RestoreTest }
func (f *fakeProven) ProvenRestoreTests(context.Context) []RestoreTest { return f.tests }
func rt(tier, archive string, pass bool, at time.Time) RestoreTest {
return RestoreTest{
SourceArchive: archive, SourceTier: tier, Pass: pass,
Verified: "boot+running", TestedAt: at.UTC().Format(time.RFC3339),
}
}
// mergeCollector builds a Collector with only the two restore-test seams wired — the merge is what
// is under test, not the rest of the collection.
func mergeCollector(latest, proven []RestoreTest) *Collector {
c := &Collector{}
if latest != nil {
c.restoreTests = &fakeLatest{tests: latest}
}
if proven != nil {
c.provenTests = &fakeProven{tests: proven}
}
return c
}
func findTier(got []RestoreTest, tier string) (RestoreTest, int) {
var hit RestoreTest
n := 0
for _, e := range got {
if e.SourceTier == tier {
hit, n = e, n+1
}
}
return hit, n
}
// ── SCENARIO A — a proof survives a restart and reaches the hub ──────────────────────────────
//
// COMPANION RED-PROOF (observed 2026-08-03): delete the `c.provenTests` merge from
// collectRestoreTests (return the in-memory slice as it used to) →
//
// --- FAIL: TestMerge_ProofSurvivesARestart
// restoretest_merge_test.go: after a restart the persisted proof must be reported; got 0 entr(ies)
//
// which is exactly the live observation: `0 restore-tests`. Restored.
func TestMerge_ProofSurvivesARestart(t *testing.T) {
provenAt := time.Date(2026, 8, 3, 13, 25, 14, 0, time.UTC) // the real run's timestamp
// After a restart the in-memory store is EMPTY — this is the whole point.
c := mergeCollector([]RestoreTest{}, []RestoreTest{
rt("pbs", "felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z", true, provenAt),
})
got := c.collectRestoreTests(context.Background())
if len(got) != 1 {
t.Fatalf("after a restart the persisted proof must be reported; got %d entr(ies): %+v", len(got), got)
}
e := got[0]
if e.SourceArchive != "felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z" {
t.Fatalf("the entry must name the archive that was proven — the hub keys on it; got %q", e.SourceArchive)
}
if e.SourceTier != "pbs" || !e.Pass {
t.Fatalf("the entry must be a PASS on the tier it was proven on; got tier=%q pass=%v", e.SourceTier, e.Pass)
}
if e.TestedAt != provenAt.Format(time.RFC3339) {
t.Fatalf("the entry must carry the ORIGINAL test time, not now(); got %q", e.TestedAt)
}
}
// ── SCENARIO B — the report does not invent a pass ───────────────────────────────────────────
//
// COMPANION RED-PROOF (observed): make the state layer emit an entry for an unproven tier (drop the
// `reportable()` filter in ProvenRestoreTests, so a legacy record with no archive is emitted) — the
// equivalent at this layer is a proven-source that returns an entry for a tier nothing proved, which
// this test injects directly and the assertion below rejects.
func TestMerge_NeverInventsAPassForAnUnprovenTier(t *testing.T) {
// Nothing proven anywhere: no in-memory result, no persisted proof.
c := mergeCollector([]RestoreTest{}, []RestoreTest{})
if got := c.collectRestoreTests(context.Background()); len(got) != 0 {
t.Fatalf("a tier with no proof must produce NO entry — an unproven tier reading as proven is "+
"worse than the defect being fixed; got %+v", got)
}
// And an entry the state layer could not describe (no tier) is never promoted into a proof.
c2 := mergeCollector([]RestoreTest{}, []RestoreTest{
{SourceArchive: "local:backup/x.tar.zst", SourceTier: "", Pass: true,
TestedAt: time.Date(2026, 8, 1, 4, 41, 58, 0, time.UTC).Format(time.RFC3339)},
})
if got := c2.collectRestoreTests(context.Background()); len(got) != 0 {
t.Fatalf("a persisted record with no tier is not a usable proof and must be dropped; got %+v", got)
}
}
// ── SCENARIO C — a fresh in-memory result wins, and never duplicates ─────────────────────────
//
// COMPANION RED-PROOF (observed 2026-08-03): remove the de-duplication (append every persisted entry
// unconditionally) →
//
// --- FAIL: TestMerge_NewerWinsAndNeverDuplicatesATier
// restoretest_merge_test.go: one entry per tier; got 2 for "pbs" — the hub would read two tests
//
// Restored.
func TestMerge_NewerWinsAndNeverDuplicatesATier(t *testing.T) {
lastWeek := time.Date(2026, 7, 27, 19, 55, 41, 0, time.UTC) // jittered, from the real box
fiveMinAgo := time.Date(2026, 8, 3, 13, 25, 14, 0, time.UTC)
c := mergeCollector(
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/new", true, fiveMinAgo)},
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/old", true, lastWeek)},
)
got := c.collectRestoreTests(context.Background())
e, n := findTier(got, "pbs")
if n != 1 {
t.Fatalf("one entry per tier; got %d for \"pbs\" — the hub would read two tests: %+v", n, got)
}
if e.SourceArchive != "felhom-pbs:backup/ct/9201/new" {
t.Fatalf("the NEWER result must win; got %q tested %q", e.SourceArchive, e.TestedAt)
}
// ...and the older-in-memory / newer-persisted direction, which is the post-restart case.
c2 := mergeCollector(
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/old", true, lastWeek)},
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/new", true, fiveMinAgo)},
)
e2, n2 := findTier(c2.collectRestoreTests(context.Background()), "pbs")
if n2 != 1 || e2.SourceArchive != "felhom-pbs:backup/ct/9201/new" {
t.Fatalf("newest must win regardless of which source it came from; got %d entr(ies), archive %q", n2, e2.SourceArchive)
}
}
// ── SCENARIO D — a failure still reaches the hub ─────────────────────────────────────────────
//
// The merge must not mask a failure with an older stored success. A failing tier is retried at the
// next evaluation and its record lives ONLY in memory, so losing it here would silence the loudest
// DR signal this system produces.
func TestMerge_AFailureIsStillReported(t *testing.T) {
provenLastWeek := time.Date(2026, 7, 27, 19, 55, 41, 0, time.UTC)
failedJustNow := time.Date(2026, 8, 3, 13, 41, 7, 0, time.UTC)
c := mergeCollector(
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/new", false, failedJustNow)},
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/old", true, provenLastWeek)},
)
e, n := findTier(c.collectRestoreTests(context.Background()), "pbs")
if n != 1 {
t.Fatalf("one entry per tier; got %d: %+v", n, c.collectRestoreTests(context.Background()))
}
if e.Pass {
t.Fatalf("a FAILURE newer than the stored proof must be what is reported — masking it would "+
"silence the loudest DR signal there is; got pass=%v archive=%q", e.Pass, e.SourceArchive)
}
}
// Two different tiers are both reported — the merge is per tier, not a single slot.
func TestMerge_BothTiersSurvive(t *testing.T) {
c := mergeCollector(
[]RestoreTest{rt("local", "felhom-backup:backup/vzdump-lxc-9201-a.tar.zst", true,
time.Date(2026, 8, 3, 4, 44, 50, 0, time.UTC))},
[]RestoreTest{rt("pbs", "felhom-pbs:backup/ct/9201/x", true,
time.Date(2026, 8, 2, 5, 12, 33, 0, time.UTC))},
)
got := c.collectRestoreTests(context.Background())
if _, n := findTier(got, "local"); n != 1 {
t.Fatalf("the in-memory tier must survive the merge; got %+v", got)
}
if _, n := findTier(got, "pbs"); n != 1 {
t.Fatalf("the persisted tier must survive the merge; got %+v", got)
}
}
// A malformed timestamp must never displace a good entry — "unparseable" is not "newest".
func TestMerge_MalformedTimestampNeverWins(t *testing.T) {
good := rt("pbs", "felhom-pbs:backup/ct/9201/good", true, time.Date(2026, 8, 3, 13, 25, 14, 0, time.UTC))
bad := RestoreTest{SourceArchive: "felhom-pbs:backup/ct/9201/bad", SourceTier: "pbs", Pass: true, TestedAt: "not-a-time"}
c := mergeCollector([]RestoreTest{good}, []RestoreTest{bad})
e, n := findTier(c.collectRestoreTests(context.Background()), "pbs")
if n != 1 || e.SourceArchive != "felhom-pbs:backup/ct/9201/good" {
t.Fatalf("an unparseable timestamp must not displace a good entry; got %d entr(ies), archive %q", n, e.SourceArchive)
}
}
// A nil proven-source leaves the pre-R-189 behaviour exactly as it was.
func TestMerge_NilProvenSourceIsANoOp(t *testing.T) {
only := rt("local", "felhom-backup:backup/x.tar.zst", true, time.Date(2026, 8, 3, 4, 44, 50, 0, time.UTC))
c := mergeCollector([]RestoreTest{only}, nil)
got := c.collectRestoreTests(context.Background())
if len(got) != 1 || got[0].SourceArchive != only.SourceArchive {
t.Fatalf("a nil durable source must not change anything; got %+v", got)
}
}
@@ -0,0 +1,24 @@
{
"generation": 1,
"desired_state": {
"os_update": {
"ring": 1,
"enabled": true,
"release": {
"id": "os-guest-20261004-120000",
"snapshot": "20261004T120000Z",
"packages": [
{"name": "libc6", "version": "2.41-12+deb13u4", "origin": "Debian"},
{"name": "openssl", "version": "3.5.7-1~deb13u3", "origin": "Debian-Security"}
]
},
"host_release": {
"id": "os-host-20261004-120000",
"snapshot": "20261004T120000Z",
"packages": [
{"name": "libssl3t64", "version": "3.5.7-1~deb13u3", "origin": "Debian-Security"}
]
}
}
}
}
+122
View File
@@ -0,0 +1,122 @@
package lanresolver
import (
"context"
"io"
"log/slog"
"os"
"path/filepath"
"strings"
"sync"
"testing"
)
// recRunner records every privileged command EnsureDnsmasq would run and succeeds — nothing reaches
// apt, systemctl or the root checker.
type recRunner struct {
mu sync.Mutex
calls []string
}
func (r *recRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
r.mu.Lock()
defer r.mu.Unlock()
r.calls = append(r.calls, strings.Join(append([]string{name}, args...), " "))
return nil, nil, nil
}
func (r *recRunner) RunStdin(ctx context.Context, _ io.Reader, name string, args ...string) ([]byte, []byte, error) {
return r.Run(ctx, name, args...)
}
func (r *recRunner) installed() bool {
for _, c := range r.calls {
if strings.HasPrefix(c, "apt-get install") && strings.HasSuffix(c, " dnsmasq") {
return true
}
}
return false
}
// fixtureRoot builds a fake host root holding exactly the given relative files and points the REAL
// probe at it for the test's duration.
func fixtureRoot(t *testing.T, files ...string) {
t.Helper()
root := t.TempDir()
for _, f := range files {
p := filepath.Join(root, f)
if err := os.MkdirAll(filepath.Dir(p), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(p, nil, 0o644); err != nil {
t.Fatal(err)
}
}
prev := hostRoot
hostRoot = root
t.Cleanup(func() { hostRoot = prev })
}
func ensure(t *testing.T) *recRunner {
t.Helper()
r := &recRunner{}
m := NewManager(r, "192.0.2.10", []string{"1.1.1.1"}, slog.New(slog.NewTextHandler(io.Discard, nil)))
if err := m.EnsureDnsmasq(context.Background()); err != nil {
t.Fatalf("EnsureDnsmasq: %v", err)
}
return r
}
// R-317: a host with `dnsmasq-base` (the /usr/sbin/dnsmasq binary) but WITHOUT the `dnsmasq` package
// (the service unit) must get the package installed — else the following `systemctl enable --now
// dnsmasq` hits a unit that does not exist and LAN name resolution silently never comes up.
//
// RED-PROOF: probe "usr/sbin/dnsmasq" instead of the unit paths in dnsmasqUnitInstalled → this fails
// with "install was skipped".
func TestEnsureDnsmasq_BinaryWithoutUnitInstalls(t *testing.T) {
fixtureRoot(t, "usr/sbin/dnsmasq")
r := ensure(t)
if !r.installed() {
t.Fatalf("install was skipped on a dnsmasq-base-only host (binary present, unit absent) — "+
"the enable that follows targets a missing unit (R-317). calls: %q", r.calls)
}
}
func TestEnsureDnsmasq_UnitPresentSkipsInstall(t *testing.T) {
for _, unit := range []string{"usr/lib/systemd/system/dnsmasq.service", "lib/systemd/system/dnsmasq.service"} {
t.Run(unit, func(t *testing.T) {
fixtureRoot(t, "usr/sbin/dnsmasq", unit)
if r := ensure(t); r.installed() {
t.Fatalf("apt-get install ran although the dnsmasq unit is present at %s: %q", unit, r.calls)
}
})
}
}
func TestEnsureDnsmasq_NothingPresentInstalls(t *testing.T) {
fixtureRoot(t)
if r := ensure(t); !r.installed() {
t.Fatalf("install skipped on a host with no dnsmasq at all: %q", r.calls)
}
}
// Production wiring for the hostRoot seam: the shipped probe resolves against the real root and asks
// about the unit the `dnsmasq` package owns — never the dnsmasq-base binary.
func TestEnsureDnsmasq_ProductionProbeIsTheUnit(t *testing.T) {
if hostRoot != "/" {
t.Fatalf("hostRoot default = %q, want \"/\" — the production probe would look in the wrong tree", hostRoot)
}
var sawUsrLib bool
for _, p := range dnsmasqUnitPaths {
full := filepath.Join(hostRoot, p)
if strings.HasSuffix(full, "/sbin/dnsmasq") || strings.HasSuffix(full, "/bin/dnsmasq") {
t.Errorf("probe path %s is the dnsmasq-base binary, not the dnsmasq unit (R-317)", full)
}
if full == "/usr/lib/systemd/system/dnsmasq.service" {
sawUsrLib = true
}
}
if !sawUsrLib {
t.Errorf("probe paths %q miss /usr/lib/systemd/system/dnsmasq.service (dpkg -S: owned by dnsmasq)", dnsmasqUnitPaths)
}
}
+35 -3
View File
@@ -33,6 +33,8 @@ const (
DropinDir = "/etc/dnsmasq.d"
// BaseDropin holds the host-wide listen/upstream config (one per host).
BaseDropin = "felhom-resolver-base.conf"
// PrivApply is the root content checker that installs a drop-in (R-861).
PrivApply = "/usr/local/sbin/felhom-priv-apply"
)
// RenderBase returns the host-wide dnsmasq drop-in: bind to the host LAN IP (+ loopback), no-resolv,
@@ -99,11 +101,35 @@ func NewManager(runner proxmox.Runner, hostIP string, upstreams []string, logger
}
}
// hostRoot is the filesystem root the install probe resolves against: "/" in production; a test
// points it at a fixture tree so the REAL probe runs against files it controls.
var hostRoot = "/"
// dnsmasqUnitPaths are where the `dnsmasq` package ships its systemd unit (Debian; /lib is the
// pre-usrmerge spelling). R-317: probe the UNIT, never /usr/sbin/dnsmasq — that binary belongs to
// `dnsmasq-base`, so a host carrying dnsmasq-base without dnsmasq used to skip the install and then
// `systemctl enable --now dnsmasq` failed against a unit that is not there (resolver never up).
// Pinned by TestEnsureDnsmasq_BinaryWithoutUnitInstalls.
var dnsmasqUnitPaths = []string{
"usr/lib/systemd/system/dnsmasq.service",
"lib/systemd/system/dnsmasq.service",
}
// dnsmasqUnitInstalled reports whether the dnsmasq service unit (the `dnsmasq` package) is present.
func dnsmasqUnitInstalled() bool {
for _, p := range dnsmasqUnitPaths {
if _, err := os.Stat(filepath.Join(hostRoot, p)); err == nil {
return true
}
}
return false
}
// EnsureDnsmasq makes dnsmasq present + enabled and writes the host base config. Idempotent: it
// installs the package only when absent, and writes the base drop-in only when its content changes.
func (m *Manager) EnsureDnsmasq(ctx context.Context) error {
if _, err := os.Stat("/usr/sbin/dnsmasq"); err != nil { // metadata read, no privilege needed
m.logger.Info("lanresolver: dnsmasq absent — installing")
if !dnsmasqUnitInstalled() { // metadata read, no privilege needed
m.logger.Info("lanresolver: dnsmasq service unit absent — installing")
if out, errOut, ierr := m.runner.Run(ctx, "apt-get", "install", "-y", "-q", "dnsmasq"); ierr != nil {
return fmt.Errorf("install dnsmasq: %s: %w", strings.TrimSpace(string(errOut))+string(out), ierr)
}
@@ -263,7 +289,13 @@ func (m *Manager) writeFileIfChanged(ctx context.Context, path, content, mode st
return false, fmt.Errorf("write temp: %w", err)
}
tmp.Close()
if _, errOut, err := m.runner.Run(ctx, "install", "-m", mode, tmpName, path); err != nil {
// R-861 (v0.146.0): a dnsmasq drop-in reaches /etc/dnsmasq.d only through the root checker, which allows exactly
// the lines RenderBase/RenderGuestDropin write (a `dhcp-script=` would run as root). Mode is fixed (0644) there.
_ = mode
if filepath.Dir(path) != DropinDir {
return false, fmt.Errorf("drop-in %s is not in %s", path, DropinDir)
}
if _, errOut, err := m.runner.Run(ctx, PrivApply, "dnsmasq", tmpName, filepath.Base(path)); err != nil {
return false, fmt.Errorf("install %s: %s: %w", path, strings.TrimSpace(string(errOut)), err)
}
return true, nil

Some files were not shown because too many files have changed in this diff Show More