Compare commits

..

60 Commits

Author SHA1 Message Date
admin d03ab7f1f5 R-836: the kernel lane — one-shot boot through the ESP flag, boot good / one self-revert, night step on a told night
gates / gates (push) Successful in 44s
Wrapper layer kernel (stage / reboot / boot / good / revert / cancel / status;
R20-R23), the two GRUB generators in the bundle (option C on the one-shot
entry), the agent's night step and after-boot judge (host health rule + hub
reached, 20 min measured), the signed os_kernel_step (stage only).
Red-proofs: felhom.eu audits/kernel-lane-2026-10-07/A/redproof.txt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 15:18:24 +02:00
admin b17d1c597d REPORT/CHANGELOG: 2026-10-07 day
gates / gates (push) Successful in 59s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 14:21:01 +02:00
admin 0ae01dfb52 CHANGELOG: v0.151.0 released (R-861, R-812 A, R-366, R-105)
gates / gates (push) Successful in 58s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 13:14:47 +02:00
admin dd7cdc09e7 R-366 slice 2 (decision 168): the host report carries the archives the restore-test skipped as another key's
gates / gates (push) Successful in 1m6s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin 7b0a8b234b R-105 (decision 169): remove the escrow-create -directive flag and the upload's directive field
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin ce1a4b4758 R-812 option A: the Proxmox package lane (layer pve) + the /etc/pve write gate
The wrapper gains layer "pve" (slow lane): the host's Proxmox userspace
packages only — origin "Proxmox Debian Repository", never a kernel / boot /
firmware / microcode name (R14), no removal, no undo, a new package only from
an allow-list; authority = a signed os_pve_step or the root-owned ring-0 mark.
The night leg runs it in ring 0 after a healthy host step; ring 1 only by a
signed job (PVEStepExecutor). While it runs, the agent's own /etc/pve writes
(every non-GET API call, pct config verbs, pvesm, pveum, felhom-pbs-apply)
wait on internal/pvegate. Health = the host rule + unchanged container ids +
pveversion reads the installed pve-manager.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin ac90169a5d R-861 (a) A1 + (b) B2: the image ref goes to a root verb that checks it; the agent's in-guest tee grant is gone; felhom-op's pct lines are exact (09 §3 decision 165)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:28 +02:00
admin 154d6dcaa9 Shared rule file: no hub image build or deploy in a session the operator does not attend (09 §3 decision 162)
gates / gates (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 09:55:02 +02:00
admin 2f7072050f REPORT: released and delivered (2026-10-07)
gates / gates (push) Successful in 1m4s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 09:31:58 +02:00
admin 85e799f360 CHANGELOG: v0.150.0 released (R-528, R-894, R-330)
gates / gates (push) Successful in 46s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 08:59:08 +02:00
admin 3a72a4811b Shared rule file: rule 11 — every helper prompt carries the brief's fences in full (09 §3 decision 160)
gates / gates (push) Successful in 54s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 08:52:36 +02:00
admin a165d53c88 REPORT: the second burn-down night
gates / gates (push) Successful in 49s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 00:14:50 +02:00
admin adaf86ad57 R-330: the agent sends SMART 187/188/199 (raw; omitted when unknown)
gates / gates (push) Successful in 53s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 22:13:04 +02:00
admin 74b5eae5b0 R-894: after a restart the agent remembers the last backup per tier
gates / gates (push) Successful in 35s
An unreadable storage right after an agent restart read the off-site tier
DUE (the in-memory record was empty). The newest success per tier is now
kept on disk and read ONLY when the storage cannot be read: fresh -> not
due, older than the cadence -> due, none -> due (unknown) as before. A
storage that answers stays the ground truth.

Ships with v0.150.0 after the 2026-10-07 read-back; nothing delivered.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 20:33:27 +02:00
admin 7e82f325b8 CHANGELOG: the memory-kill check (unreleased, to ship as v0.150.0 after the night read-back)
gates / gates (push) Successful in 50s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:36:09 +02:00
admin acccb66bd3 R-528 (09 decision 157): after a Docker engine step the wrapper proves the engine reports a memory kill
felhom-os-apply: a docker-layer apply runs oom_check() after health_after and reports
"oom_check": {result pass|fail|error, oom_killed, oom_event, exit_code, image, detail}.
One throwaway container (the controller's image, --pull never, --network none, 64m cap,
label felhom.oomcheck=1) runs dd bs=200M; pass only with OOMKilled=true AND the oom event
(read after a 2 s settle, --until = guest epoch + 1: measured on demo-hp, an --until taken
right after the run missed the event). docker rm -f always runs in a finally; every call
is bounded (<= 90 s). It never changes the step's outcome or health. New wrapper-only mode
"oom-check" (docker layer) runs the check alone; check_guest etc. still apply.

Agent: WrapperReport/Report gain OOMCheck (json:"oom_check"), copied unchanged in runLayer
and in the kept-copy path.

Tests: 9 wrapper tests + 2 Go tests, each red-proofed (audits/readback-2026-10-07/F/red-*.txt).
Also: test_felhom_os_apply.py's `if __name__` sat mid-file, so 11 tests (UnsentReport,
SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran as a script or from
TestWrapperSuite; moved to the end (they pass).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:34:55 +02:00
admin de812bc027 The shared rule file (09 §3 decision 152), identical to the other copies; no code change
gates / gates (push) Successful in 44s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 16:06:46 +02:00
admin 3e8ebeb96c Instruction files kept true (09 §3 decision 150): stale gate lists, paths and facts corrected; no code change
gates / gates (push) Successful in 46s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 13:43:57 +02:00
admin cefdc731a4 REPORT: the operator's ten answers (2026-10-06)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 12:28:58 +02:00
admin e56dcb8a4c CHANGELOG/v0.149.0 released (shas)
gates / gates (push) Successful in 37s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:50:21 +02:00
admin f277e619e2 CHANGELOG: unreleased — R-856 GET /host/crash-guard
gates / gates (push) Successful in 38s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:40 +02:00
admin 386f51edc6 R-856: GET /host/crash-guard — the host crash guard's last-boot record for the controller (09 decision 143)
The controller waits ~15 minutes with app mails after a crash boot of the host; it learns of the
crash boot from this route. Reads /var/lib/felhom-crash-guard/state.json (read-only, no Proxmox call)
and passes present/last_boot_at/last_boot_unclean/tripped through; a missing, unreadable or garbled
file answers 200 present:false. Guest-token authed like every sibling route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:01 +02:00
admin 130e3ed882 CHANGELOG: unreleased — R-444 weekly guest disk trim, R-99 runbook pointer
gates / gates (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:25:22 +02:00
admin be398f92e8 R-99: the PBS phantom WARN names the cleanup runbook (09 §3 decision 140)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin ee71abd1d4 R-444: weekly guest disk trim (pct fstrim) outside the night, under the heavy-op gate
Operator ruling 09 §3 decision 139. One exact sudoers rule FELHOM_FSTRIM
(`/usr/sbin/pct ^fstrim [0-9]+$`) + manifest entry guest-fstrim; new
internal/fstrim job: due Wednesday from 10:00 host-local, starts only
10:00-20:59, holds backup.InFlight (busy -> deferred to the next hourly
tick), failed trim retried at most 3x per week, bytes parsed from the
"(N bytes) trimmed" lines, last result per guest persisted in
<state_dir>/guest-disk-trim.json and reported as guest_disk_trim.
Config opt-out: "disk_trim": {"disable": true}.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin 37e98f452b REPORT: the burn-down night (2026-10-06)
gates / gates (push) Successful in 45s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 02:13:10 +02:00
admin b2b82ae828 bundle test: the ISO first-boot files the installer names under KEPT are not installer-written (go test red since felhom.eu 85de3f9b); CHANGELOG unreleased (R-426 decoys)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:25 +02:00
admin b78a0ff3ac R-426: decoys for the shared reuse-refs, instructions and observations gates
COVERS "reuse-refs", "instructions", "observations": the three shared
felhom.eu scripts run against a scratch clone of THIS repo (in a scratch
workspace symlinking the sibling clones they reach across to), so the
plant is in the agent's own REUSE.md / CLAUDE.md / REPORT.md. Convicted:
a missing cited .go and .md path, a version literal in CLAUDE.md's
effective text, R-419's prose-only Observations note. Passed: the real
files, the version inside an HTML comment, both genuine markers.
DECOY_SHARED_DIR lets a red-proof judge a mutated copy of the shared
scripts without editing the felhom.eu clone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 96047453cb R-426: decoys for the release-complete gate
COVERS "release-complete": the working-tree gate runs in a scratch clone
whose origin is a scratch bare repo, against the fake Gitea. Convicted:
the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated
commit, a tag with no package, and no-tag wins over a registry 500.
Inconclusive: a registry 500. Passed: the genuine release, an
`## Unreleased` heading above it, a newer version named only in prose or
under `###` (the withdrawn sweep decoy, now asserted the right way round),
a tag only origin has (the shallow-CI shape), and a LOCAL-only tag BY
DESIGN (CI's fresh clone and the published gate's converse probe see it).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 4bf5db6875 R-426: decoy suite for the published gate, against a fake Gitea
scripts/test_gate_decoys.py (new; COVERS "published"): an http.server on
127.0.0.1 stands in for Gitea through the gate's existing GITEA_BASE
seam, proxies stripped, so no case reaches the real registry. Facts
convicted: a tag whose package 404s, a tag tree without the configs, a
package one patch past the newest tag (never tagged), a patch-gap orphan,
a missing package that lexical sorting would drop out of the retention
window. Inconclusive, never a pass: tags api 500, a non-JSON 200, Gitea
unreachable. Passed: a clean registry, a non-semver tag, a version older
than the retention window (BY DESIGN).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 87977ff40a CHANGELOG/v0.148.0 released (shas)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 00:37:07 +02:00
admin 861d32a4b4 CHANGELOG: unreleased — R-349 agent_sha256, R-25 agent half (burn-down night)
gates / gates (push) Successful in 43s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:51 +02:00
admin 769c4c3cf2 R-25 (agent half): the format answer carries the NEW filesystem's UUID, bound to the durable id
After mkfs the agent re-resolves the bound durable id, requires it to name the
device it just formatted, reads the superblock back (blkid -p, requested fstype)
and returns fs_uuid in POST /disks/format and GET /disks/format/status. Anything
unverified returns "" — never a path-resolved guess. The controller half
(mount fs_uuid instead of re-resolving the UUID from the /dev path) is owed
in felhom-controller.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:11 +02:00
admin 64f704d0f7 R-349: the host report carries the sha256 of the RUNNING agent binary
A hand-built proof binary and the published artifact share a version string
but not their bytes, so no version check could see the divergence. The agent
now reports agent_sha256 (hash of /proc/self/exe, once per process; empty =
unknown) beside agent_version, the same mechanism as host.wrapper_sha256.
The hub half (compare against the vouched agent_sha256, surface drift) is
owed in the hub repo.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:11 +02:00
admin 208fac8027 CHANGELOG/REPORT: v0.147.0 released (shas)
gates / gates (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 18:58:37 +02:00
admin f1b9b41214 R-124 recipe root namespace as PBS spells it; R-118 no root size for an absent drive; R-269 rotated-out token rejected at once; R-317 dnsmasq install probed by its unit (burn-down round 2)
gates / gates (push) Successful in 47s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 18:56:13 +02:00
admin d83316326e R-291 retention record source, R-348 restart comment (no binary change; burn-down)
gates / gates (push) Successful in 22s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 16:56:18 +02:00
admin e06ed97fa8 agent v0.146.1 REPORT
gates / gates (push) Successful in 22s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 13:00:22 +02:00
admin e4b5cf9693 R-880: build-step-bundle.py — the transition bundle for a release whose bundle adds paths (an installed felhom-os-apply refuses unknown paths, R16)
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:23:55 +02:00
admin faa3cad92e agent v0.146.1 CHANGELOG (released)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:14:07 +02:00
admin fdd87178d2 R-861 review fixes: the signed update hands the A/B wrapper a root-owned copy of the hashed bytes; mount units accept no Wants/Requires/Before and no continuation lines; the escrow read walks the path without following any symlink
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:13:42 +02:00
admin 0342c7bb57 agent v0.146.0 CHANGELOG (released)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:01:25 +02:00
admin 6ab1e7c56c R-861: narrow the agent's root grants — exact sudo patterns, felhom-priv-apply content checker, fixed hook/parent files in the bundle, signed self-update verified as root, escrow root reads pinned
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:00:43 +02:00
admin 61345790ed REPORT: the 2026-10-05 catch-up session
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 10:29:12 +02:00
admin 7c4b8e599f v0.145.0: CHANGELOG (released 894da35c…, bundle 78c00adc…)
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 09:23:48 +02:00
admin 56ef1d6655 v0.145.0 code: the OS wrapper repairs dpkg's update journal by itself after a power cut (R-876); restore-test first check 30 min after start (R-874); neutral "sent late" text (R-875)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 09:23:23 +02:00
admin 78c890e4bb REPORT: the 2026-10-05 night-fixes session
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 08:25:49 +02:00
admin d48f1bbb23 v0.144.1: CHANGELOG (released 6ccd521d…, bundle e89a9ddf…)
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:49:29 +02:00
admin 8401a30917 v0.144.1 code: the wrapper survives a dead reader (BrokenPipe) so a killed pass still keeps its report; the agent looks for kept copies every 5 min (R-868, measured live)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:49:04 +02:00
admin d1b6004458 v0.144.0: CHANGELOG (released f18093c3…, bundle 6acf42fe…)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:15:07 +02:00
admin ca78c17b29 v0.144.0 code: R8 measures the real download (R-865); an OS pass's report survives a killed agent (R-868); the debug pass runs from the saved block when the hub is away (R-866)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:14:32 +02:00
admin c8d12f1f2a REPORT + CONTEXT: 2026-10-04 night (R-840 / R-860)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 20:39:47 +02:00
admin dc9164c5af CHANGELOG v0.143.0 (R-840); build-golden.sh 3.2.0: GOLDEN_GUEST_PKGS — the approved guest release at bake time, first-night count
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 20:08:16 +02:00
admin c9fa2e717b R-840: the config bundle — a signed agent_config_update brings a box's root-owned files (sudoers, wrappers, units)
gates / gates (push) Successful in 18s
felhom-os-apply gains mode 'bundle' (signed, verified by the wrapper itself against the root-owned
signers file — or, when that file is missing, only the installer's pinned key, which it then creates)
and --install-bundle (the installer's root entry). BUNDLE_FILES is the one table of paths; every check
(visudo, sh/bash -n, python, unit sections, RuntimeDirectory guard, nft -c, the route itself) runs
before the first write; a failed write or self-check puts every previous copy back. The trust root is
never a bundle path (R17). scripts/build-config-bundle.py builds it reproducibly; release-agent.sh
publishes it beside the binary. The agent reports the bundle record in system.config_bundle.
felhom-opsign signs agent_config_update. 43 wrapper tests (22 mutants red), Go executor tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 19:36:58 +02:00
admin 0af1187e03 CHANGELOG + REPORT: v0.142.1 released (R-858)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 18:21:46 +02:00
admin 495003051b R-858: after a Docker engine step the wrapper restarts ONLY the containers that mount the docker socket (controller, traefik); the health rule fails when the controller cannot reach Docker from inside its container (ruling 95)
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 18:21:24 +02:00
admin 42af3ab9bc build-golden.sh 3.1.0: live-restore on in the golden (fail-closed assertion), GOLDEN_DOCKER_PKGS pins the approved Docker engine set
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 17:30:19 +02:00
admin f24dce5b95 CHANGELOG + REPORT: v0.142.0 released
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 16:15:15 +02:00
admin b1746c25af Docker engine slow lane (live-restore once by reload; ring-0 pending-docker under a root-owned ring-0 mark; ring 1 and undo only by a signed os_docker_step the wrapper re-verifies against a root-owned signers file; same-container-id health), the version report (facts mode -> host report system stanza, R-852), guest restart scan every pass (R-849), the crash guard (kernel.panic=10, the 3rd unclean stop in 60 min stays off, 24 h re-arm)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 16:14:36 +02:00
admin 2e2e8f56b8 CHANGELOG + REPORT: v0.141.1 released (host reboot-needed fixes)
gates / gates (push) Successful in 17s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 13:39:35 +02:00
125 changed files with 13271 additions and 851 deletions
+3 -2
View File
@@ -21,6 +21,7 @@ fast, and wrong.
This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that
felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before
path-scoped rules existed. The single source is now felhom.eu/CLAUDE.md "Code quality rules"; this
file is the scoped copy that loads exactly where health checks are written. (2026-08-06)
path-scoped rules existed. Deliberate scoped copies now live in felhom.eu/.claude/rules/hub.md and
felhom-controller/.claude/rules/gates.md (hub.md's comment names them); none is the single source. This
file is the copy that loads exactly where agent health checks are written. (2026-08-06; corrected 2026-10-06)
-->
+75
View File
@@ -0,0 +1,75 @@
---
unconditional: true
---
# Unprompted work — rules for any session without a task file
> Goal sessions, nightly sessions, "work the register" sessions. **A session that starts from
> `/goal` or a standing brief inherits these rules exactly as it inherits the gates.** They are the
> part of `PROMPT-TEMPLATE.md` that a task file used to carry and a goal does not. Same wording lives
> in `felhom.eu`, `felhom-controller`, `felhom-agent` and `app-catalog-felhom.eu` `.claude/rules/`, and in the workspace
> root's unversioned `.claude/rules/`; change all five or none.
## 1. What you may pick up on your own
- A register row **you or another CC session filed**, with owner CC, at P3 or a bounded P2, that
needs **no operator decision**, touches **no customer data by design**, and introduces **no
mechanism nobody has measured**. Smallest first.
- A defect you find while exercising the product, filed as a row **before** you fix it — **unless it is small**:
a small finding is fixed in the session and never filed (the size rule, `OPEN-ITEMS.md` „How a row is filed").
- Hygiene: register compression, stale citations, rows with no owner, documents that contradict
live source.
**Not yours, ever, without a task file or an operator word:** money; anything that changes risk to
customer data; anything that changes a promise the product makes to a customer; anything that
reverses a documented design decision (`documentation/architecture/` — a design decision is not a
defect, R-370); anything on DooPlex or ep0; baking or vouching a golden; promoting a
catalog version; a new external dependency; **a hub image build or hub deploy in a session the operator does not
attend** (operator ruling 2026-10-07, `09` §3 decision 162).
## 2. When you may decide instead of ask (operator grant, 2026-09-14)
You may take a decision yourself when **all** of these hold: the architecture folder and the register
give a clear direction; your choice follows that direction; it is reversible without customer-data
risk; and you can write it in the `09-update-architecture.md` §3 shape — one answerable sentence, the
options, what each costs, why this one. **Then record it** as a dated decision in `CONTEXT.md` and
the owning architecture document, tagged *decided by CC unattended — operator may reverse*, and put
it **first** in the morning note. A decision you cannot write in that shape is one you do not take.
## 3. The discipline a task file used to carry
1. **Baselines first.** Read each repo's `main` hash and version from live source before touching it.
2. **Read the architecture document for the area, and name it** in the report, before any claim.
3. **Red-proof every correctness fix.** A test never seen failing has not been shown to test anything.
4. **Live-validate on a Tier-0 box** through the endpoints the UI invokes. `demo-hp` is `ssh hp`.
Throwaway apps only; the standing apps and `bentopdf` stay.
5. **Evidence off the machine at the end of each phase**, before any revert (R-320).
6. **One release per repo per session**, with a CHANGELOG entry (controller: with its `MinAgent`
line), REPORT overwritten, floor raised to deliver it. **No golden unless a drill or fresh install
needs one** (the waiver, R-468). **No `--no-verify`.**
7. **An enumerated gap becomes a row in the same session — or, if it is small, is fixed in it** (the size rule).
Prose is not a record.
8. **Hungarian text is searched with ASCII fragments**, with a positive and a negative control.
9. **Never leave a half-state.** If time runs out, revert to clean and say what was reverted.
10. **Teardown, three layers, stated** — machine, host, hub — or "provisioned nothing".
11. **Every helper prompt carries the brief's fences in full** (operator ruling 2026-10-07). A helper session (a
subagent, a fork, a workflow agent) gets the brief's fence list word for word — every protected machine, every
„no", every delivery and Docker limit — not a summary and not „the usual fences". Earned on 2026-10-06 night: two
helpers whose prompts carried only part of the fences ran `docker volume prune` on the bench and a Docker-using
gate on DooPlex.
## 4. The morning note
One screen, plain language, in this order: **decisions you took** (§2) first; what you exercised;
what broke and whether you fixed it; rows opened and closed with the register size before and after;
what needs the operator, each with what happens if they do nothing. No file paths, no function
names, no row numbers as the subject of a sentence.
## 5. Instruction files
**Instruction files (`CLAUDE.md`, `.claude/rules/*`) are kept true by the session that finds them wrong**
(operator ruling 2026-10-06, `09` §3 decision 150). A session MAY, without asking: correct a stale fact (a command, a
count, a version, a path, a description of what a gate does), add a fact it proved, and remove a reference to something
that no longer exists. Each edit is named in the report (file, line, before, after, why). A session MAY NOT, without the
operator's word: loosen a safety rule, a fence, a „never", a protected machine, a secret rule, or a review step; or
remove a rule. When in doubt, it is a rule change, and it goes to the operator. If Claude Code's own permission check
asks before such an edit, wait for the operator's click; if it refuses, record that and file the exact line.
+428
View File
@@ -1,3 +1,431 @@
## Unreleased — part of v0.152.0: the kernel lane (R-836; `09` §3 decisions 164, 172; `11` §5.11) (2026-10-07)
**Delivery: agent binary, then the STEP bundle `0.152.0-step1`, then the bundle `0.152.0`** — the bundle ADDS two paths
(the GRUB generators), and an installed `felhom-os-apply` refuses a path its own table lacks (R16, R-880).
A new kernel boots ONCE; if it crashes the box comes back on the old kernel by itself; it becomes the default only after
a healthy boot; a booted-but-unhealthy kernel is reverted ONCE by the agent with no person. Built on the spike's
candidate 2 (`audits/kernel-spike-2026-10-07/`), with option C on the one-shot entry.
- `configs/felhom-grub-oneshot.sh` → `/etc/grub.d/01_felhom_oneshot` (bundle): reads `felhom_next` from a GRUB env block
on the ESP (`EFI/felhom/oneshot.env`), clears and saves it BEFORE the menu, and sets the default to that kernel's
one-shot entry only when the name is an installed kernel. No vfat ESP → prints nothing.
- `configs/felhom-grub-oneshot-entries.sh` → `/etc/grub.d/42_felhom_oneshot` (bundle): one entry per installed kernel,
id `felhom-oneshot-<ver>`, the normal entry plus `softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10`
(option C). Sorted after `10_linux`: never entry 0, never the default.
- `configs/felhom-os-apply`: layer `kernel` (lane slow; an appliance; authority = a signed `os_kernel_step` or the
root-owned ring-0 mark). Modes: `apply` STAGES (select `pending-kernel` or a signed `listed` set; `expect_kver` = the
kernel the household was told about): pins the GRUB default to the RUNNING kernel in
`/etc/default/grub.d/zz-felhom-kernel-default.cfg` and proves it from grub.cfg, installs, proves the default did not
move and the one-shot entry exists, writes the flag; never reboots. `kernel-reboot` (a staged step only),
`kernel-boot` (judging | fell_back | self_reverted | revert_failed), `kernel-good` (the new kernel becomes the
default, proved), `kernel-revert` (ONE per step; refused when the default is not the old kernel), `kernel-cancel`,
`kernel-status`. New refusals: R20 (the box cannot do a one-shot: not UEFI, no vfat ESP, a separate /boot, GRUB
without fat/loadenv, the generators missing, a hand pin), R21 (the crash guard tripped or an unclean boot in its
window), R22 (the phase does not allow the mode; never two steps within 20 h), R23 (not exactly one newer kernel, or
not the one signed / told). State `/var/lib/felhom-kernel/state.json`. Facts carry `kernel_lane`; the next-boot
kernel reads the flag and the grub.cfg default. Tests: `KernelLane` (27), red-proof
`felhom.eu/documentation/audits/kernel-lane-2026-10-07/A/redproof.txt`.
- `configs/felhom-crash-guard` unchanged; `KernelStepCannotLeaveTheBoxOff` (3 tests) pins that a step's planned reboot,
one crash and one self-revert add ONE unclean boot (a panic before userspace adds none), so the box cannot stay off.
- `internal/osupdate/kernel.go`: the night leg ends with the kernel step — after a healthy host step (and a healthy
Proxmox step when one ran), trigger `night` only, on a night the hub's `os_update.kernel` block marks `tonight` (the
household was mailed the day before — no mail, no step). Ring 0 stages + reboots; ring 1 reboots only a kernel a
signed `os_kernel_step` staged (`KernelStepExecutor`: stage only, under the heavy-op gate). The hub hears `staged`
BEFORE the reboot. At every start `KernelAfterBoot`: on the new kernel it JUDGES the boot — `KernelVerdict` = the
host health rule (`11` §8.2) AND the box reached the hub (the `judging` report itself) — for 20 minutes (measured:
everything healthy 68 s after the reboot on demo-felhom, 272 s on demo-hp; under the hub's 30-minute `host_stale`).
Healthy → `kernel-good`, outcome `applied`; not healthy → outcome `health_failed`, then ONE `kernel-revert`.
Tests: `TestKernel*` (13); red-proofs in the same file.
- `internal/hub`: `WireOSUpdate.Kernel` {kver, tonight, notified_at}. `internal/reconcile`: `os_kernel_step` is
destructive-class. `cmd/felhom-opsign`: the op is listed.
## v0.151.0 — the agent can no longer hand the guest any image; the Proxmox package lane; the other-key archives reported; the DR directive retired (R-861, R-812 A, R-366, R-105; `09` §3 163, 165, 168, 169) (2026-10-07)
Released by `scripts/release-agent.sh`: binary sha256 `0464354f2cdf452a7c5d2a74d9191fe91415fcfa244154480d26d5b30e10b194`
config bundle sha256 `bacd1d175a9392bc1755341d01a106abbd72aba48ab575e7a32dd819d8f5da4c` (tag `v0.151.0` = `dd7cdc0`).
**Deliver the binary FIRST, then the bundle:** the bundle's sudoers removes the `tee` grant the 0.150.0 binary still uses.
No path added (26 → 26), so no step bundle.
### Part of v0.151.0 — the agent can no longer hand the guest any image; felhom-op's pct lines are exact (R-861 (a) A1, (b) B2; `09` §3 decision 165)
**Delivery order: agent binary FIRST, then the config bundle.** The new sudoers drops the agent's in-guest `tee`
grant; an older binary still calls `tee`, so a bundle that lands before the binary would stop managed controller
updates (and the old binary's capability probe would read `controllerswap-write` degraded). No bundle path is added
(`felhom-priv-apply` and `/etc/sudoers.d/felhom-op` are already bundle files), so no step bundle.
- `configs/felhom-priv-apply`: new verb `controller-image <vmid>` — reads the ref on stdin (≤ 256 bytes, ASCII, one
optional trailing newline), requires `^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$` (the agent's
own `controllerImageRe`), then runs `pct exec <vmid> -- tee /etc/felhom-controller-image` AS ROOT; refusal rule `I1`
(rc 3), a bad vmid `A1` (rc 2); listed in `--self-check`.
- `configs/felhom-agent.sudoers` `FELHOM_CONTROLLERSWAP`: `pct ^exec [0-9]+ -- tee /etc/felhom-controller-image$`
REMOVED; `/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$` added.
- `internal/localapi`: `GuestExecutor.GuestExecStdin` replaced by `WriteControllerImage`; `GuestBinder.WriteControllerImage`
pipes `ref\n` to the verb through the fenced runner; the swap's `writeImage` calls it. Capability `controllerswap-write`
now probes the verb.
- `configs/felhom-op.sudoers` (B2, hygiene): `pct start|stop|unlock [0-9]*` → `pct ^start [0-9]+$` etc. (the glob's `*`
matched spaces: `pct stop 9201 --skiplock 1` passed).
- Tests: `ControllerImage` (5, `configs/test_felhom_priv_apply.py`), `TestSudoersRefusesTheR861Injections` (+3 lines),
`TestSudoersAllowsTheControllerImageVerb`, `TestFelhomOpSudoersPctIsExact`, `TestR861_WriteControllerImageUsesTheRootVerb`,
`TestControllerSwap_WriteViaRootVerb_NoShell`. Red-proofs: `felhom.eu/documentation/audits/day-2026-10-07/C/`.
- `README.md`: the controller-swap paragraph described the removed `tee` path — corrected.
### Part of v0.151.0 — the Proxmox package lane (R-812 option A, `09` §3 decision 163)
**MinAgent impact: none** (a new layer; an older hub ignores the pve report). **The bundle carries the new
`felhom-os-apply` — deliver it with the binary** (signed `agent_update`, then signed `agent_config_update`).
- `configs/felhom-os-apply`: new layer `pve`, lane `slow` only — the host's Proxmox USERSPACE packages: origin `Proxmox Debian Repository` only (R2), never a kernel / boot / firmware / microcode name (R14, `HOST_SLOW_RE` — the kernel is R-836's lane), no removal (R4), no undo (R5), a new package only from `PVE_NEW_ALLOW` (`proxmox-firewall-data`, measured on demo-felhom; R6 otherwise), an appliance only (R12), authority = a signed `os_pve_step` or the root-owned ring-0 mark (R3). Select `pending-pve` (ring 0): installed Proxmox-origin packages with a pending upgrade. The report carries `pve_manager` (pveversion after the step).
- `internal/pvegate` (new): the agent's own writes to /etc/pve wait while a pve step runs (pmxcfs restarts); the step waits for writes in flight (bounded, 2 min — then it fails and does not run). Wired at `proxmox.Client.doBody` (every non-GET) and `ExecRunner.RunStdin` (`WritesEtcPVE`: pct config verbs, pvesm, pveum, felhom-pbs-apply create/reconcile).
- `internal/osupdate`: `LayerPVE`; the night leg runs the pve step in ring 0 after a healthy host step (an appliance; ring 1 never in the night leg); `PVEHealthVerdict` = the host rule + every running container keeps its id + pveversion reads the installed pve-manager; the pve report carries Proxmox userspace only (the hub's candidate set). `PVEStepExecutor` (signed `os_pve_step`, ring 1, under the heavy-op gate and the /etc/pve gate); `reconcile.ClassOSPVEStep` (destructive-class); `felhom-opsign -op os_pve_step` (params by `-params`).
- Tests: wrapper `PVELane` (17; red first — the `pve-manager` plan was refused R12 on the old code), `pvegate` (5), `TestPVEGate_*` + `TestWritesEtcPVE`, `TestPVE_*`, `TestPVEHealthVerdict`, `TestPVEStepExecutor_*`. Red-proofs: `felhom.eu/documentation/audits/day-2026-10-07/B/`.
### Part of v0.151.0
- R-366 slice 2 (`09` §3 decision 168): the restore-test pick records, per tier, the archives it skipped as written with another key (count, oldest, newest — no key material) in a `ForeignKeyLedger`; the host report carries it as `foreign_key_archives.tiers` (the stanza absent until a tier was evaluated since start, `tiers: []` when none — no null on the wire, the report contract forbids it). The hub turns a change into one operator line. Tests `TestR366_PickRecordsArchivesWrittenWithAnotherKey`, `TestR366_EvaluatedWithNoneIsAnEmptyList` (red-proved, `felhom.eu/documentation/audits/day-2026-10-07/E/`).
- R-105 option A (`09` §3 decision 169): the `--selftest=escrow-create -directive <file>` flag and the escrow upload's `directive` field are removed — nothing read the directive; the DR path reads the recipe, tenantsync and the escrow blob. The hub ignores a `directive` from an older agent.
## v0.150.0 — the Docker step proves the engine reports a memory kill; after a restart the agent remembers the last backup per tier; three more SMART counters on the wire (R-528, R-894, R-330; `09` §3 decisions 157, 161) (2026-10-07)
Released by `scripts/release-agent.sh`: binary sha256 `a23d1c9085bc7fd4fc48fe0327f6504aa83e6331510fb4a3d23e042dddb26f9c`
config bundle sha256 `88456b386d9b1027bd22861cac8c23df004bf9fd9f67644d6595bfca8c94498e` (tag `v0.150.0` = `3a72a48`).
**The bundle carries the new `felhom-os-apply` (the memory-kill check) — deliver it with the binary:** signed
`agent_update`, then signed `agent_config_update`. No path added (26 → 26), so no step bundle.
### Part of v0.150.0 (2026-10-06 night, later) — after a restart the agent remembers the last backup per tier (R-894); three more SMART counters on the wire (R-330)
Ships with the memory-kill check below as v0.150.0, AFTER the 2026-10-07 night read-back. Nothing delivered tonight.
- **The defect (measured 2026-10-05 on demo-hp):** the agent restarted at 04:57; at 06:25 the off-site storage answered *Can't connect*; the per-tier backup record is in memory only, so the due-check fell back to an EMPTY record and the 7-day tier (last copy 4 days old) read DUE; the controller asked and vzdump failed.
- New `internal/backup/backup_state.go` `BackupSuccessState`: the newest SUCCESSFUL backup per tier and guest, on disk (`<oob state dir>/backup-success-state.json`, atomic tmp+rename, 0600). Only successes are written; a corrupt file reads as nothing known.
- `internal/localapi` `handleBackupDue`: when the tier's storage CANNOT be read, the saved copy stands in for the in-memory record. A fresh copy → not due („… (storage unreadable — age from the last success saved on disk)"); a copy older than the cadence → DUE; no copy → the old answer (DUE, age unknown). A storage that answers stays the ground truth: an archive absent there is due even when the file remembers one.
- Wired in `buildLocalAPIServer` (`LastKnownBackups`); the local API's backup job saves each success.
- Tests: `TestBackupDue_R894_*` (restart = a new server and a new state from the same file; fresh / old / none / storage answers / failed backup not saved), `TestBackupSuccessState_*`, `TestR894_LastKnownBackupsIsWiredIntoTheDaemon` (AST). Four red-proofs observed (`felhom.eu/documentation/audits/night-burndown-2026-10-06/s4/`).
- **R-330 (disk health Phase 2, the wire only):** the SMART summary carries three more SATA raw counters — `reported_uncorrect` (187), `command_timeout` (188, carried as the vendor reports it; some pack several counters), `udma_crc_errors` (199). Pointer + omitempty: an attribute the drive does not report is OMITTED (unknown), never 0. No verdict reads them yet. Tests `TestParseSMART_R330_*` (two red-proofs, `felhom.eu/documentation/audits/night-burndown-2026-10-06/r330/`).
### Part of v0.150.0 (2026-10-06 night) — the Docker step proves the engine reports a memory kill (`09` §3 decision 157, R-528)
To be released as v0.150.0 with its config bundle AFTER the 2026-10-07 night read-back (the night of 2026-10-06 runs v0.149.0 on purpose).
- `configs/felhom-os-apply`: after a docker-layer APPLY (after `health_after`) the wrapper runs `oom_check()`: a throwaway container from the image the running controller uses (`--pull never`, `--network none`, no volume, label `felhom.oomcheck=1`, 64 MB cap) asks for one 200 MB block; „pass" only when `OOMKilled=true` AND the `oom` event; it waits 2 s and reads the events window to the guest's epoch + 1 (measured: a window closed in the same second missed the event); the container is always removed. Reported as `oom_check`; it never changes the step's outcome or health. A wrapper-only mode `oom-check` runs the check alone (no apt, no engine change), by hand as root.
- `internal/osupdate`: `WrapperReport` and `Report` carry `oom_check` verbatim, on the normal pass and on the kept-copy path (R-868).
- `configs/test_felhom_os_apply.py`: its `unittest.main()` sat in the middle of the file, so 11 tests (UnsentReport, SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran — moved to the end; all pass.
- Tests: the OOMCheck class (pass, OOMKilled=false, no event, unreadable image, removal on an inspect error, not on other layers or in health mode, the events window after the settle wait, mode oom-check alone and its refusals); TestDocker_OOMCheckReachesTheHubUnchanged, TestR868_KeptCopyCarriesTheOOMCheck. 12 red-proofs in `felhom.eu/documentation/audits/readback-2026-10-07/F/`.
### Part of v0.150.0 (2026-10-06 evening) — the shared rule file (`09` §3 decision 152); no code change
- `.claude/rules/unprompted-work.md` added, byte-identical to the copies in felhom.eu, felhom-controller, app-catalog-felhom.eu and the workspace root (checked with `diff` against the controller's copy and one md5 across all five). Its copies line names five copies.
### Part of v0.150.0 (2026-10-06 afternoon) — instruction files kept true (`09` §3 decision 150); no code change
- `CLAUDE.md` „Gates — ONE entry point": the runner runs every gate in its `GATES` table (five: three shared, `published`, `release-complete`); `--fast` skips `published` (network). It said two gates and „all of them".
- `CLAUDE.md`: the decoy gate and its audit are named with their `felhom.eu/` prefix (they do not exist in this repo).
- `.claude/rules/health-checks.md` (comment): the health-check rule's copies live in felhom.eu `hub.md` and the controller's `gates.md`; it named felhom.eu `CLAUDE.md` „Code quality rules", which holds no such rule.
## v0.149.0 — a weekly disk trim of each customer guest, the crash-boot fact for the controller, the phantom WARN names its runbook (R-444, R-856, R-99; operator rulings `09` §3 139, 143, 140) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `6bcae9c2eb5d97e8285316583870059835793893299e291891a53a4ce505585f`
config bundle sha256 `e182c82dcf4a67faa3bcb74dbe4ffa7b06e0b27dc8451cb7574d6339ce91ad66` (tag `v0.149.0` = `f277e61`).
**The bundle carries the new sudoers rule for the trim (`FELHOM_FSTRIM`) — deliver it with the binary:** signed
`agent_update`, then signed `agent_config_update`.
- R-856 (`09` §3 decision 143): new local-API route `GET /host/crash-guard` — passes the host crash guard's last-boot record (present, last_boot_at, last_boot_unclean, tripped) from /var/lib/felhom-crash-guard/state.json to the controller, which waits ~15 min with app mails after a crash boot. Read-only, no Proxmox call, guest-token authed; a missing/unreadable/garbled file answers 200 present:false (never an error page). An older agent answers 404, which the controller reads as unknown (normal 90 s grace) — no controller MinAgent raise needed.
- R-444 (`09` §3 decision 139): weekly guest disk trim. New sudoers alias FELHOM_FSTRIM with ONE exact rule `/usr/sbin/pct ^fstrim [0-9]+$` (rides the signed config bundle; decoys pinned by TestSudoersFstrimRuleIsExact) and capability guest-fstrim (non-critical). New internal/fstrim job: each owned RUNNING guest gets `pct fstrim <vmid>` once a week - due Wednesday from 10:00 host-local, starts only 10:00-20:59 (never the 01:00-06:59 night), holds the one-heavy-op gate so it never runs beside a backup or restore-test (busy -> deferred to the next hourly tick; a box that was off catches up at its next daytime hour); a failed trim WARNs and is retried at most 3 times that week; bytes parsed from `pct fstrim`'s "(N bytes) trimmed" lines; positive log `fstrim: guest N trimmed X GiB in Ys`; last result per guest persisted in <state_dir>/guest-disk-trim.json and reported as the new omitempty host-report stanza `guest_disk_trim`. Opt-out: agent.json "disk_trim": {"disable": true}.
- R-99 (`09` §3 decision 140): the agent's WARN for a PBS archive below the 1 MiB plausibility floor now ends with the pointer to the sanctioned cleanup (`documentation/runbooks/pbs-phantom-cleanup.md`); detection only — nothing is deleted automatically. Dir-storage archives keep the old text.
## unreleased
- R-426: scripts/test_gate_decoys.py (new) — the published gate judged against a fake Gitea (127.0.0.1, via GITEA_BASE; never the real registry): 11 cases; COVERS published — `felhom-agent/published` leaves the decoy-coverage EXEMPT list.
- R-426: release-complete gate — 10 decoy cases (scratch clone + scratch bare origin + fake Gitea); COVERS release-complete — `felhom-agent/release-complete` leaves the decoy-coverage EXEMPT list.
- R-426: the shared reuse-refs/instructions/observations gates get agent-side decoys (9 cases on a scratch clone of this repo); COVERS reuse-refs, instructions, observations — three `felhom-agent/*` entries leave the decoy-coverage EXEMPT list.
- **Fixed without a row:** `configs/test_felhom_config_bundle.py` read the two ISO first-boot files that installer 1.32.0 now NAMES under KEPT (R-275) as files the installer writes — `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 (felhom.eu `85de3f9b`); they are listed with why. Test only; agent v0.148.0's code is unaffected.
## v0.148.0 — the host report names the running binary's sha; the format answer carries the new filesystem's UUID (burn-down night: R-349, R-25 agent halves) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `3e68a0870e0e2ce262cb4819294edddeb0a73e8c558a31611a20329a6d9ee283`
config bundle sha256 `a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de` (tag `v0.148.0` = `861d32a`).
Delivery order as for v0.147.0: signed `agent_update`, then signed `agent_config_update`.
- **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from
`/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub
comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved.
- **R-25 (agent half):** `POST /disks/format` and `GET /disks/format/status` return `fs_uuid`, the new filesystem's UUID
read back after mkfs only when the bound durable id still resolves to the formatted device and the superblock is the
requested type (empty = not verified); `DeviceProbe` gains `FSUUID` from blkid. Tests `TestFormat_*FSUUID*`; three
red-proofs. The controller half (mount that UUID) is a controller change.
## v0.147.0 — the recovery recipe spells the root namespace the way PBS does; a removed drive no longer shows the root disk's size; a rotated-out token stops at once; the dnsmasq check looks at the right package (burn-down round 2: R-124, R-118, R-269, R-317) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `642c4d196c48671c14ff653118303c5903af1670b7abb550edeaf73e701cd5b8`,
config bundle sha256 `326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d` (tag `v0.147.0` = `f1b9b41`).
Delivery order as for v0.146.1: signed `agent_update`, then signed `agent_config_update`.
MinAgent impact: none (the controller needs nothing new from this agent). Config bundle content unchanged from v0.146.1.
- **R-124 (operator ruling 2026-10-05: fix it):** the DR recipe's `pbs.namespace` for a box in PBS's ROOT namespace is
now `""` — PBS's own spelling — beside `namespace_state: resolved`; it used to be the word `root`, which no namespace
is named, so `--ns root` failed in a recovery. `hub.PBSRootNamespace`; `TestR124_RootNamespaceOnTheWireIsPBSSpelling`
(red-proof: back to "root" → FAIL). Runbook: `felhom.eu runbooks/ep0-datastore-copy.md` step 2 says how to read it
(and to treat a recorded `root` from older agents as empty). The hub stores the recipe raw; its fixture follows.
- **R-118:** the local API's drive list reads a drive's capacity only while its DEVICE is present — with the device gone
the bare mountpoint is a directory on the root filesystem, whose size was reported as the drive's.
`TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity` (red-proof convicts).
- **R-269:** the token store re-reads its shared file whenever it has grown, BEFORE answering — so a token rotated out by
another process stops authorizing on its next use (it used to keep working until an unrelated miss). One `stat` per
call. `TestTokenStore_RotatedOutTokenRejectedFirst` (red-proof convicts).
- **R-317:** the LAN resolver decides whether to install `dnsmasq` by its service UNIT, not by `/usr/sbin/dnsmasq` (which
the `dnsmasq-base` package also ships). Same `apt-get install` command; no sudoers change. `TestEnsureDnsmasq_*`
(red-proof convicts). Red-proofs: `felhom.eu/documentation/audits/burndown2-2026-10-05/agent-red-proofs.txt`,
`r124-red-proof.txt`.
Also in this release (no binary effect; from burn-down round 1):
- **R-291:** `scripts/retention-policy.json` names where its 10 comes from — the R-267 newest-10 prune of generic
packages, established 2026-08-10 (R-287) — instead of „observed, no located ruling"; the non-existent
`registry-retention.md` reader is dropped. `check-published-versions.py` still reads 10 (checked).
- **R-348:** `internal/backup/store.go` no longer says backups are „unaffected" by a restart: the reported backup list
reads 0 until the next backup runs; only the hub's verdict (7-day look-back) is unaffected.
## v0.146.1 — R-861 review fixes: the signed update flips a root-owned copy; no Wants=/continuations in mount units; the escrow read follows no symlink anywhere (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `badd6c9a2e40c8bfe856d2d1a203443b21b7eb92ecc35d6090ab44518d4d082a`,
config bundle sha256 `42333e969028867ad8142335e6c1bc4040eec231de0d8d330c2d4b2cf7bc3442`. **Supersedes v0.146.0, which
was released but never vouched or delivered to any box.** The same order applies: signed `agent_update` first, then the
signed `agent_config_update`.
**Delivery needs a STEP bundle (R-880, found while delivering).** An installed `felhom-os-apply` checks an incoming
bundle's paths against its OWN table (R16), so every box on the v0.145.0 bundle REFUSES the v0.146.1 bundle (it adds 4
paths). `scripts/build-step-bundle.py` builds the transition: the box's current bundle with ONLY `felhom-os-apply`
replaced (same paths — the old wrapper accepts it), published as bundle version `0.146.1-step1`; then the release's own
bundle. Order on a box: `agent_update` 0.146.1 → `agent_config_update` 0.146.1-step1 → `agent_config_update` 0.146.1.
Tests `StepBundle` (the R16 refusal reproduced; the step accepted; exactly one file changed). Tooling only — not in the
binary or the bundle.
A background security review of the v0.146.0 commit found three holes in the new code; each is fixed and red-proved
(`felhom.eu/documentation/audits/hub-safety-2026-10-05/partF/red-proof.txt`, S1–S3):
- **S1 — a race in the signed update.** `felhom-os-apply` hashed the agent's staged file and then let the A/B wrapper copy
it BY PATH; the agent owns that directory and could swap the file in between. Now the root step reads the file ONCE
(`read_staged_once`: O_NOFOLLOW, fstat, owner, size), hashes those bytes, writes them to a root-owned directory
(`/var/lib/felhom-os-apply/agent-update/`) and hands ONLY that copy to `felhom-selfupdate-guarded apply`, which now
refuses any other directory, a symlink, or a file not owned by root. Tests: `AgentUpdate` (+1),
`SelfupdateWrapperConfinement`.
- **S2 — an allowlist escape in `felhom-priv-apply`.** `[Unit]` accepted `Wants=`/`Requires=`/`Before=` naming any unit, so
a mount unit could start e.g. `reboot.target`. `[Unit]` now holds only `Description` and `After=local-fs-pre.target`
(what the renderers write), and any line ending in a backslash (a systemd continuation this parser would read
differently) is refused. Tests `test_U2_wants_starts_another_unit`, `test_U2_continuation_line`.
- **S3 — a path traversal in the escrow read.** `O_NOFOLLOW` guards only the last component; a symlinked DIRECTORY in the
agent's own state dir still redirected the root read. `readStagedNoFollow` now walks the path from `/` with
`openat(O_NOFOLLOW)` per component. Test `TestAttach_RefusesASymlinkedDirectory`.
## v0.146.0 — the agent's root grants narrowed: exact sudo patterns, a root content checker, fixed files from the bundle, the signed update checked as root (R-861) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `b860af465076041e07f35fed1b12d64ae2b2985d8995f0ce167418d39c2b00d5`,
config bundle sha256 `161c737e523aa7910cf32ce41b83f989569bee55b8c5938e7211c92aef68548e`. **Order on a box: the signed
`agent_update` FIRST (the old bundle still grants the old flip), then the signed `agent_config_update`.** Between the
two (minutes) the new agent's checker calls are refused and retried; nothing is lost. After the bundle, an agent BELOW
0.146.0 cannot update itself on that box any more (the unsigned flip grant is gone) — deliver both together.
Design: `felhom.eu/documentation/architecture/03-host-agent.md` §3.1 (new). Measured before the change (real sudo
1.9.16, a throwaway container): the v0.145.0 sudoers let **23 of 29** attack command lines through; v0.146.0 lets
**0** through and still allows all **64** commands the agent's capability check uses.
- **Exact patterns.** A sudoers `*` in the arguments also matches spaces: `pct set [0-9]* -onboot 1` matched
`pct set 100 --dev0 /dev/sda -onboot 1` (a raw host disk for a guest), `mount --bind /mnt/*/felhom-data
/mnt/felhom-drives/*` matched a `..` path onto `/etc/sudoers.d`, `nft add element … *` took a chained `; flush
ruleset`. Every varying argument list is now a sudo regex (`^…$`): one value per slot, a fixed character set, no
`..`, no extra argument. `TestSudoersRefusesTheR861Injections` (29 attacks) + `TestManifestCoveredBySudoers`
(regex-aware now).
- **`felhom-priv-apply`** (new root wrapper, in the bundle). A systemd mount/automount unit, a dnsmasq drop-in, the
WireGuard config and the OOB sshd config + felhom-op key reach their root-read places only through it: fixed source,
fixed destination, CONTENT checked against what the agent's renderers write (no `[Service]`, `Where=` only
`/mnt/<name>` or `/mnt/felhom-drives/<name>` and equal to the unit name, no `bind`/`suid`; a network share must carry
`nosuid,nodev`; no `dhcp-script=`; no `PostUp=`; the sshd config only the one template with its Port). Its 30 tests
(`configs/test_felhom_priv_apply.py`) + Go contract tests feeding each renderer's real output
(`internal/privapplytest`). Pre-flight: every live file on both demo boxes reads OK.
- **NFS/SMB options gain `nosuid,nodev`** (a set-uid file on a server outside the box never acts on the host).
- **Fixed files from the bundle.** The guest pre-start hook (`/var/lib/vz/snippets/felhom-guest-hook.sh`, run as root at
every guest start) and the shared drive parent script + unit are bundle files now (byte-identical to the agent's
constants, pinned). The agent no longer installs them from `/tmp`; it checks them (`guesthook.SnippetReady`,
`ensureSharedParentBoot`) and only registers / enables.
- **The signed update is checked as root.** `felhom-os-apply` mode `agent_update` verifies the operator signature
(root-owned signers, this host, the window, the nonce), re-hashes the staged binary against the SIGNED sha, then runs
the A/B flip; `felhom-selfupdate-guarded apply` is no longer in the agent's sudoers. 7 tests (`AgentUpdate`).
- **The root escrow run reads no path from the agent's config.** As root it pins the PVE secret dir and the WireGuard
state dir to their defaults, refuses a storage id that is a path, and reads its two staged files without following a
symlink (`readStagedNoFollow`) — before, a symlink in the agent's own directory sealed any root file into the blob.
- **Not narrowed here (named in `03` §3.1):** `FELHOM_CONTROLLERSWAP` stays guest-scoped (a compromised agent can run a
chosen controller image in the guest — the household's data, not host root); `FELHOM_ESCROW` still hands the agent R
by design (the agent relays the ceremony); the mkfs / pbs-apply / backup-target wrappers keep a coarse argument and
their own checks.
- Red-proofs F1–F9: `felhom.eu/documentation/audits/hub-safety-2026-10-05/partF/red-proof.txt` (F1's first run did NOT
convict — the name rule masked it — and the test now uses the pair only the Where rule stops).
## v0.145.0 — the OS update repairs itself after a power cut; a short-session box gets restore-tested; "sent late" (R-876, R-874, R-875) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `894da35c7b9e1ac78885690b78352b634e6831e7d99b321573c75d878db8886e`,
config bundle sha256 `78c00adce662d2d966b2ac50ebde46cde1ae225f0107c6a7c02b70ec8ce80c4f`. The wrapper changed: a box
needs the signed `agent_update` AND the signed `agent_config_update`.
- **R-876.** After a crash during an install, `dpkg --audit` can read clean while dpkg's update journal
(`/var/lib/dpkg/updates/`) is not — and apt refuses every install until `dpkg --configure -a` (measured on demo-hp
2026-10-05: every later pass failed until a person typed it). The wrapper now reads `--audit` and the journal in ONE
`sh -c` call (`DPKG_STATE_SCRIPT`) — a clean pass still costs one call (R-845's speed, pinned) — and repairs when
either shows something; as a belt, when apt itself says "dpkg was interrupted", it repairs and retries the install
ONCE. `REPAIR` now logs `journal=N`; a journal still not empty after the repair refuses (R13). Tests
`CrashLeftTheJournal` (the measured shape, the speed, the belt); 3 red-proofs.
- **R-874.** The restore-test's first due-check runs 30 minutes after the agent starts (`DefaultFirstEval`), then every
interval; a box whose power-on sessions are shorter than the 6 h interval never evaluated. A crash-looping agent
restarting faster than 30 minutes still never evaluates (the earned restraint, pinned).
- **R-875.** A kept report's reason is neutral — "sent late — kept on the box until the hub could take it" — the copy
cannot tell a killed agent from an absent hub.
- Red-proofs: `felhom.eu/documentation/audits/catchup-2026-10-05/part{C,D}/`.
## v0.144.1 — a killed pass really keeps its report: the wrapper survives a dead reader; the agent looks again every 5 minutes (R-868, measured live) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `6ccd521d47e64999017e8eb5bc613d724543cfdc5ef9b13bcae9e3ea8c53b8f3`,
config bundle sha256 `e89a9ddfb767e177f8874d56f3dcd3bfd47409d157ff830d362653333bf815e8`. The wrapper changed again:
a box needs the signed `agent_update` AND the signed `agent_config_update`.
- **Found live on demo-hp 2026-10-05 05:45 UTC with v0.144.0** (the night's A5 shape: kill -9 of the pass and the
daemon while apt-get ran): apt finished all 13 packages, but the wrapper's next log line went to a stderr pipe no
process read any more → `BrokenPipeError` → the wrapper died before it saved its report copy (journal: `PLAN
upgrade=13`, then nothing; no copy; the hub got nothing). v0.144.0's mechanism was right and never reached.
`Runner.log` and the final `OSAPPLY-REPORT` line now survive a dead reader (the journal still gets every line).
Test `AgentDiesMidPass` drives the REAL `log()` into a pipe that breaks while apt-get runs.
- **Also found live:** the restarted daemon looked for kept copies ~7 s before the orphaned wrapper wrote one. The
daemon now looks at start and every 5 minutes (`Leg.SendUnsentLoop`); `TestR868_ACopyWrittenAfterTheStartIsSentByTheLoop`.
- Red-proofs: `felhom.eu/documentation/audits/night-fixes-2026-10-05/partD/r868-brokenpipe-red-proof.txt`.
- A second agent release in one session, against "one release per repo": recorded as `09` decision 108 (operator may
reverse) — the alternative was to ship a fix proven not to work.
## v0.144.0 — R8 measures the real download; an OS pass reports even when its agent was killed; the debug pass runs with the hub away (R-865, R-868, R-866) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `f18093c3466749ec4cd47f83f97a401160a24bad1e704f1183051adad14db928`,
config bundle `felhom-config-bundle.json` sha256 `6acf42fe46df5223384d767801cb2bf73238ba4dab82debb7811ca2f790591d8`.
The wrapper `felhom-os-apply` changed, so a box needs BOTH the signed `agent_update` and the signed `agent_config_update`.
- **R-865.** `download_bytes` runs `apt-get --print-uris` WITHOUT `-s`: with `-s` apt prints the simulation and no URI
list, so R8 summed 0 B and only its 500 MB floor ever applied. `--print-uris` alone downloads nothing (measured on
9202: the archive cache and the versions unchanged). The test fake now answers like real apt (with `-s`: no URIs),
and `test_R8_counts_the_real_download` / `test_download_bytes_never_simulates` pin it.
- **R-868.** The wrapper writes every apply pass's report to `<plan dir>/report-<run>-<layer>-apply.json` before it
prints it (root writes into the agent's dir: the dir opened O_NOFOLLOW and checked to be the agent's own, the file
created O_EXCL|O_NOFOLLOW, 0600, handed to the agent). The plan now carries `run_id`, `trigger`, `ring`, echoed in
the report. The agent deletes the copy once the hub has the report; a copy left on disk (the agent was killed, or
the hub was away) is sent at the agent's start and before every pass (`Leg.SendUnsent`), then deleted. A pass lock
(flock on `pass.lock`, across the daemon and a selftest) keeps the sender off a pass that is still running.
- **R-866.** The daemon saves the hub's newest os_update block (`os-update-block.json`); `--selftest=os-update` uses it
when the hub cannot be reached and says so in its header (`block=SAVED(<time>; hub unreachable: …)`); with no hub
and nothing saved it does not run.
- Tests: `configs/test_felhom_os_apply.py` (UnsentReport, SaveReportOnDisk — real files, symlink cases), Go
`TestR868_*`, `TestR866_*`. Red-proofs: `felhom.eu/documentation/audits/night-fixes-2026-10-05/part{C,D}/`.
## v0.143.0 — the config bundle: a signed route for a box's root-owned files (R-840, decision 96) (2026-10-04)
Released by `scripts/release-agent.sh`: binary sha256 `41c0d3060013dfda795262147454248149bee0888c0935170fdb85de6e7a35da`,
config bundle `felhom-config-bundle.json` sha256 `8d7273cf5313ef62b867cb6f831c631923a436452d6f90b8ff7f0771170396ba`.
- **The bundle.** Every root-owned file the installer's step 5 writes (sudoers ×2, the five wrappers, the crash guard
and its units, the agent and rollback units, the start-limit drop-in, the mgmt watchdog, the OOB belt's files) as ONE
reproducible JSON file, built by `scripts/build-config-bundle.py` from `BUNDLE_FILES` in `configs/felhom-os-apply`
(one table) and published beside the binary. The installer (1.31.0) installs the same file.
- **The route.** A signed `agent_config_update` {agent_version, bundle_sha256} (`felhom-opsign -op agent_config_update
-bundle-sha256 …`). The agent is the courier (downloads, checks the sha, hands over); `felhom-os-apply` mode `bundle`
verifies it ITSELF: the operator signature against the root-owned `/etc/felhom/operator-signers` (or, when that file
is missing, ONLY the installer's pinned key, after which it creates the file with exactly that key), the host binding,
the window, its own nonce; the bundle sha; every path in `BUNDLE_FILES` (R16) and never a trust file (R17); every
content check before the first write (visudo, sh/bash -n, python, unit sections, the RuntimeDirectory guard, User=,
nft -c, and that the route itself survives). Atomic per file, previous copies kept under
`/var/lib/felhom-os-apply/bundle-prev/`; a self-check after (visudo -c, `sudo -l` lists the route, the new wrapper's
`--self-check`, the self-update wrapper's usage, the crash guard's status = kernel.panic); any failure puts every
previous copy back. A newly installed crash guard is started (`enable --now`: kernel.panic for this boot, no reboot).
Record `/etc/felhom/config-bundle.json`; the agent reports it as `system.config_bundle`, the facts mode adds drift.
- **Bootstrap.** A box whose `felhom-os-apply` predates 0.143.0 cannot take the first bundle by the route (nothing on it
can write a root file from a signed job): `felhom.eu/scripts/felhom-bundle-bootstrap.sh` is the one by-hand step.
- **`build-golden.sh` 3.2.0** (not part of the binary): `GOLDEN_GUEST_PKGS` brings the template to exactly the
approved guest release (only installed packages, never newer, never a removal or a new package) and prints the
first-night count.
- Tests: `configs/test_felhom_config_bundle.py` (43; 22 of 22 mutants red), `internal/osupdate/bundle_test.go`,
`internal/hub/bundle_record_test.go`. Live: `felhom.eu/documentation/audits/r840-config-bundle-2026-10-04/partB/`.
## v0.142.1 — a Docker step no longer leaves the controller and traefik blind (R-858, `09` decision 95)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.142.1` (`4950030`), sha256
> `003f882a59abfc10021a1f97a4744ab78ab1e916b1392dcef132e7067f3c62bd`, verified by download. Not vouched at release time.
**MinAgent impact:** none. A same-day patch release, by the operator's ruling 95 after a live incident.
- **The incident.** v0.142.0's Docker step on demo-felhom (14:13 UTC) restarted dockerd, which recreates
`/run/docker.sock`. live-restore kept every container running — including `felhom-controller` and `traefik`, which
bind-mount the socket FILE and so kept the deleted inode. The controller could not reach Docker for 1 h 44 min (hub:
DOWN, "docker not reachable"); its own health check stayed "healthy", so the step's health rule PASSED.
- **The fix.** After a Docker step that installed something, the wrapper restarts ONLY the containers that mount
`/var/run/docker.sock` or `/run/docker.sock` (`os-apply: SOCKET-USERS restarted=…`; the apps and the engine untouched).
The guest health reading gains `controller_docker_ok` (`docker exec felhom-controller docker version`), and the health
rule fails when it is false — the consequence, not the mechanism.
- Proven live on demo-hp BEFORE the release (signed undo to 29.7.2): the two restarted, guest / controller / traefik on
the same socket inode, healthy. Red-proofs: `felhom.eu/documentation/audits/os-docker-crash-2026-10-04/partE-incident/`.
## v0.142.0 — the Docker engine slow lane, the version report, the crash guard (`11` §5.8, §5.9; `09` decisions 87–89)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.142.0` (`b1746c2`), sha256
> `7beb32224d6495e9561acfd3ad8a48393a799520f196080011cb27fceb6d1de6`, verified by download. Not vouched at release time.
**MinAgent impact:** none. **Needs hub v0.132.0** for the System page, the Docker approval and the crash events; an older
hub stores the new `system` stanza unread. **Needs the new root files** on an installed box (R-840): the wrapper,
`/etc/felhom/os-trust.json`, `/etc/felhom/operator-signers` and the crash guard — the installer 1.30.0 writes them; the
demo boxes got them by hand.
- **Docker `live-restore` ON** (decision 87). Wrapper mode `live-restore-on`: merge `"live-restore": true` into the
guest's `/etc/docker/daemon.json` and `systemctl reload docker` — never a restart (R-835). An invalid daemon.json is
left alone (R16); a reload that does not enable it puts the old file back. The leg runs it once before a Docker step;
`--selftest=live-restore -vmid N` runs it by hand. Measured: demo-hp 24 containers, demo-felhom 5 — the same ids after.
- **The Docker engine slow lane** (`11` §5.8). Wrapper layer `docker`, lane `slow` only, the six Docker packages only,
origin `Docker CE` only (R2; a Docker package in a fast-lane plan is refused). **R3 — the wrapper checks the authority
itself**, never the agent's config: a ring-1 step or ANY undo needs a signed `os_docker_step` it verifies with
`ssh-keygen -Y verify` against the ROOT-owned `/etc/felhom/operator-signers` (namespace `felhom-op-v1`, the blob's
`key_id`), bound to `/etc/felhom/os-trust.json` `host_id`, inside its time window, never replayed (a root-owned nonce
file), with exactly the signed packages and undo flag; an unsigned ring-0 step needs that file's
`"ring0_slow_lane": true` (the demo boxes only, set by hand). **R15:** live-restore must be on. An undo may downgrade
(`--allow-downgrades`) only inside a signed job. The leg: ring 0 runs the Docker step at night after a healthy guest
and host step (`select pending-docker`); ring 1 never does — only `DockerStepExecutor` (signed job, heavy-op gate).
**Health:** the guest rule + every container running at the start has the SAME id after + the engine reports the
installed version; a changed id is `health_failed`. Measured ring 0: 29.7.x → 29.8.2 on both demo boxes, every id kept.
- **The version report** (R-852, decision 89). Wrapper mode `facts` (read-only): host Debian, running and next-boot
kernel (`next_entry` > saved default > newest installed, by dpkg order), held packages (R-848), kernel taint (oops,
warn), `kernel.panic`, the crash guard state; guest Debian, Docker engine, containerd, live-restore. The host report
gains `system {pve_version, kernel_version, vmid, facts, facts_error}`, read at most every 10 min (~2 s);
`--selftest=os-facts -vmid N`. A value nobody could read is `unknown`.
- **R-849:** the guest is scanned for "restart needed" on every pass too, so a guest restart clears it.
- **The crash guard** (decision 88, R-851): `configs/felhom-crash-guard` + `felhom-crash-guard.service` (early boot;
its ExecStop writes a clean-stop marker) + an hourly re-arm timer + `/etc/felhom/crash-guard.conf`. A boot without
the marker followed an unclean stop (a crash, a power cut or a hard reset — pstore saved nothing for a real panic on
demo-hp, so they cannot be told apart). Armed: `kernel.panic = 10`. After the 2nd unclean boot within 60 min it
TRIPS (`kernel.panic = 0`), so the 3rd crash within the hour leaves the box off; it re-arms after 24 h of normal
running or `felhom-crash-guard rearm`. State in `/var/lib/felhom-crash-guard/state.json` (0644; read by facts).
- **After the tag (main only, not shipped to boxes): `configs/build-golden.sh` 3.1.0** — the golden's `daemon.json`
carries `"live-restore": true` with a fail-closed assertion, and `GOLDEN_DOCKER_PKGS` pins the approved Docker engine
set (all six `name=version`); without it the bake log warns that the set is the newest, not an approved one.
- Tests: wrapper 76 (DockerLane, LiveRestore, Facts, RealSignatureCheck with a throwaway key), crash guard 9, Go leg +
executor; red-proofs `felhom.eu/documentation/audits/os-docker-crash-2026-10-04/partB/agent-redproofs.txt` (17 caught).
## v0.141.1 — "reboot needed" is true on the host (found live on demo-felhom, 2026-10-04)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.141.1` (`a6bc3f1`), sha256
> `b712f577099fd2d374f648df1e825874821302fe1b46e93c7b4b70044fbe84b5`, verified by download. Not vouched at release time.
**MinAgent impact:** none. **Pairs with hub v0.131.1** (reads `reboot_scanned`); an older hub ignores the field.
A patch release in the same session as v0.141.0 — a deliberate exception to "one release per repo": v0.141.0's host
"reboot needed" was wrong in two ways, and it feeds an operator alarm.
- **The scan hid `lxc-start`.** The host scan skipped every process whose cgroup line contains `lxc`, to leave out the
guests' own processes. `lxc-start` lives in `0::/lxc.monitor/<vmid>`, so it was skipped too. Measured: after a
108-package host pass (libc6 included) `lxc-start` mapped 20 deleted files and the report said `reboot_needed:
false`. The pattern is now `:/lxc/` (`RESTART_SKIP_CGROUP`), pinned by a test that runs `grep` against the measured
cgroup lines.
- **A reboot never cleared it.** v0.141.0 scanned only after an install (R-845). The host now scans on EVERY pass (it
is local, no `pct exec`); the guest still scans only after an install. The report carries `reboot_scanned`.
- Red-proofs: `felhom.eu/documentation/audits/os-host-lane-2026-10-04/partB/live-defects-redproofs.txt`.
## v0.141.0 — OS updates: the host fast lane (`11` §8 step 3); the tunnel status is true (R-841); the leg is fast (R-845)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.141.0` (`cfba0d0`), sha256
+7 -7
View File
@@ -52,11 +52,11 @@ This is in the core because breaching it is how this component stops being audit
## Gates — ONE entry point
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs this
repo's gates — `reuse_refs_check` and `instructions_gate`, both the **shared** copies in
`felhom.eu/scripts/`, never copied into this repo (a copy recreates the drift they detect; an absent
sibling clone FAILS). `--fast` selects the gates touching no network and no container runtime; today
that is all of them. **A missing gate is a FAILURE, never a skip.**
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs every
gate in its `GATES` table (that table is the list); the shared ones — `reuse_refs_check`,
`instructions_gate`, `observations_gate` — are the copies in `felhom.eu/scripts/`, never copied into
this repo (a copy recreates the drift they detect; an absent sibling clone FAILS). `--fast` selects the
gates touching no network and no container runtime, and skips `published` (network), naming it. **A missing gate is a FAILURE, never a skip.**
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is
**per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS
@@ -104,8 +104,8 @@ the mechanism are exempt.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
lacks. `felhom.eu/scripts/decoy_coverage_gate.py` (run by felhom.eu's `repo_gates.py`, for all four repos) refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
decoys withdrawn as illegitimate: `felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list.
+8
View File
@@ -1,5 +1,13 @@
# CONTEXT — felhom-agent working state
> **2026-10-04 night — v0.143.0 RELEASED + vouched (R-840, decision 96): the config bundle.** `felhom-os-apply` mode
> `bundle` (signed `agent_config_update`, verified by the wrapper itself; trust files never bundle paths) +
> `--install-bundle` (installer 1.31.0); `BUNDLE_FILES` is the one table; `scripts/build-config-bundle.py`;
> `release-agent.sh` publishes it. A box whose `felhom-os-apply` predates 0.143.0 needs ONE by-hand bootstrap
> (`felhom.eu/scripts/felhom-bundle-bootstrap.sh`) — done on both demo boxes; Tester 2 waits for the operator (R-862).
> Both demo boxes: agent 0.143.0, bundle 0.143.0 (record `/etc/felhom/config-bundle.json`). R-861 found: the sudoers is
> root-equivalent. `build-golden.sh` 3.2.0 (`GOLDEN_GUEST_PKGS`). Runbook `felhom.eu/documentation/runbooks/config-bundle.md`.
> **2026-09-25 night — v0.133.0 AND v0.134.0 DELIVERED to both demo boxes (CC-signed `agent_update`, ruling 1);
> restore test back ON (the `-1` config kept as `agent.json.night-0925-off`). v0.134.0 = R-685:** `backup/runner.go`
+8 -7
View File
@@ -44,13 +44,14 @@ unnoticed until a user hit them. `internal/capability` makes that loud:
(`HostCapabilityChecker`) alerts the operator on a Critical capability going degraded. Serve-degraded
— the probe never blocks startup. (Next self-health slice: the controller↔agent channel check.)
**Controller-swap under non-root (v0.45.0).** The agent-owned controller image swap
(`internal/localapi/controllerswap.go`) no longer shells out: `writeImage` pipes the image ref on
**stdin** into an in-guest `tee /etc/felhom-controller-image` (via `GuestExecStdin` →
`Runner.RunStdin`, the same fenced `sudo -n` runner) — no `bash -c`, no interpolation. Its 5 narrow
grants live in the `FELHOM_CONTROLLERSWAP` sudoers alias (all read-only or fixed-target; the `tee`
target is the FIXED image path, content stdin-fed) and in the capability manifest (Critical), so a
dropped grant is a build failure + a live degraded signal. No general `pct exec` is granted.
**Controller-swap under non-root (v0.45.0; the write since R-861 (a) A1).** The agent-owned controller image swap
(`internal/localapi/controllerswap.go`) no longer shells out. The write goes on **stdin** to the ROOT verb
`felhom-priv-apply controller-image <vmid>` (`GuestBinder.WriteControllerImage` → `Runner.RunStdin`, the same fenced
`sudo -n` runner), which re-checks the ref against our registry + repository + an x.y.z tag and writes
`/etc/felhom-controller-image` inside the guest itself; the agent has no in-guest `tee` grant any more (before, a
compromised agent could feed any image — sudo cannot see stdin). Its grants live in the `FELHOM_CONTROLLERSWAP` sudoers
alias (read-only or fixed-target) and in the capability manifest (Critical), so a dropped grant is a build failure + a
live degraded signal. No general `pct exec` is granted.
## The `storage` package — observe + watchdog (slice 5)
+8 -8
View File
@@ -1,9 +1,9 @@
# REPORT — 2026-10-04: v0.141.0, the host fast lane, the true tunnel status, the fast leg
# REPORT — v0.151.0 released and delivered (2026-10-07 day)
Full session report: `felhom.eu/REPORT-os-host-lane-2026-10-04.md`.
- Tunnel (R-841): the agent reads the guest's cloudflared container and its health check; three states.
- Host fast lane: the wrapper gains the host layer (R12 appliance proof from the root-owned install record, R14 no
kernel/boot/firmware); the leg runs the host step after a healthy guest step; host health rule.
- Speed (R-845): one call per layer instead of one per package; measured before/after in the session report.
- Tests green; red-proofs in the audit folder.
On the operator's word (`09` §3 decisions 163, 165, 168, 169). sha `0464354f…`, bundle `bacd1d17…`, tag `v0.151.0` = `dd7cdc0`.
Delivered binary first, then the bundle (its sudoers drops the `tee` grant 0.150.0 used), to demo-hp, demo-felhom and
Tester 1 — probe 68/68 on each. **Vouch refused by the hub** (`golden_behind_fleet`): new installs keep 0.150.0 until the
weekly golden. Carries: the Proxmox package lane + `/etc/pve` write gate (R-812 A — proven on demo-felhom: 65 packages,
70 s, healthy), the `controller-image` root verb (R-861 a — on demo-hp an `alpine` ref is refused; a managed swap not yet
seen), anchored felhom-op lines (B2), the other-key archive ledger (R-366), the `-directive` flag removed (R-105). Evidence
`felhom.eu/documentation/audits/day-2026-10-07/`. Shared rule file: decision 162 line added.
+18 -6
View File
@@ -14,8 +14,15 @@
| `Privileged` (CreateGoldenLXC/MountUSBByUUID/SMART/Sensors) | internal/proxmox/privileged.go | methods on `*Privileged` | the 3 fenced root-CLI exceptions ONLY | Do NOT add methods — fence is structural (`routing_test.go` asserts it) |
| `SudoHostOps.run` | internal/storage/hostops.go | `run(ctx, name, args...) error` | allowlisted exec with stderr-wrapped error | Every arg pre-validated via validate.go before this is called |
| `Prober.Probe` | internal/capability/probe.go | `Probe(ctx) []Status` | live sudo-policy capability check (`sudo -n -l --`) | Needs a DIRECT runner (never the sudo-prefixing one — double-sudo); never executes probed cmds. v0.86.0: config-gated caps (`Capability.GatedBy` + `Prober.GateActive`) report `inactive`/"disabled by configuration" ONLY when healthy — broken plumbing stays degraded; the pbsdr-* gate answers from `pbsdr.Manager.DRConfigured` (marker-backed across restarts) |
| `stageTemp` | internal/localapi/intermediary.go | `stageTemp(pattern, content) (path, err)` | random-named temp before a root `install` (audit B1) | Fixed /tmp names are a TOCTOU — sudoers globs expect `/tmp/felhom-*-*.ext` |
| `guesthook.InstallSnippet` / `Register` | internal/guesthook/install.go | `InstallSnippet(ctx, runner) error` | pre-start self-heal hook install (C1 net) | Same random-temp+install pattern; snippet delegates to the agent binary (no shell logic). Issues `mkdir -p /var/lib/vz/snippets` FIRST (v0.63.0, B2 — fresh boxes lack the dir; sudoers grants exactly that argv) |
| ~~`stageTemp`~~ (REMOVED v0.146.0, R-861) | — | — | — | Nothing the agent writes is `install`ed where root reads it any more: use `felhom-priv-apply` (below) or ship a fixed file in the bundle |
| `felhom-priv-apply` (v0.146.0, R-861) | configs/felhom-priv-apply | `felhom-priv-apply unit <name> \| dnsmasq <tmp> <name> \| wg \| sshd-config \| sshd-key \| controller-image <vmid>` (the last reads the ref on stdin, R-861 (a) A1) | ANY agent-rendered file a root program reads (systemd unit, dnsmasq drop-in, wg-quick conf, OOB sshd) — fixed source + destination, CONTENT checked against the agent's own renderers | A new renderer needs a verb + a contract test (`internal/privapplytest.Check`) feeding its REAL output; never a new `install` sudoers line |
| `privapplytest.Check` | internal/privapplytest/check.go | `Check(t, verb, name, content) string` | the Go↔root-checker contract: a renderer's real output must read `OK` | Skips without python3; one call per rendered shape + one refused control |
| `BUNDLE_FILES` + `Bundle` (mode `bundle`, `--install-bundle`; mode `agent_update` v0.146.0, R-861) | configs/felhom-os-apply | the ONE table of root-owned paths + the installer of them | ANY new root-owned file the installer writes (sudoers line, wrapper, unit) — add it to the table, never a new installer fetch (R-840) | The builder (`scripts/build-config-bundle.py`) and the installer read the same table; `test_every_root_file_the_installer_writes_is_in_the_bundle` fails on a path the bundle lacks. Trust files (`/etc/felhom/os-trust.json`, `operator-signers`) are NEVER bundle paths (R17) |
| `osupdate.ConfigUpdateExecutor` | internal/osupdate/bundle.go | signed op `agent_config_update` {agent_version, bundle_sha256} | delivering the bundle to an installed box | a courier only: the root wrapper re-verifies signature, host, nonce and sha itself |
| `osupdate.Leg.SendUnsent` / `lockPass` (v0.144.0, R-868) | internal/osupdate/unsent.go | `(ctx) int` | an OS-pass report the agent never sent (killed mid-pass): the wrapper keeps `report-<run>-<layer>-apply.json` beside the plan; the agent deletes it once the hub has it | any new caller that runs an apply pass must hold `lockPass` (flock, across processes) — the sender must never take a running pass's copy |
| `osupdate.LoadSavedBlock` (v0.144.0, R-866) | internal/osupdate/leg.go | `(planDir) (block, savedAt, ok)` | the hub's newest os_update block as the daemon last received it (`os-update-block.json`) | the debug pass uses it ONLY when the hub cannot be reached, and says so in its header; no saved block → no pass |
| `dpkg_state()` / `DPKG_STATE_SCRIPT` (v0.145.0, R-876) | configs/felhom-os-apply | `audit, journal = self.dpkg_state()` | dpkg's state in ONE call: `--audit` AND the update journal | never gate a repair on `--audit` alone — a crash leaves only the journal (measured); keep it one call (R-845) |
| `guesthook.SnippetReady` / `Register` (v0.146.0, R-861) | internal/guesthook/install.go | `SnippetReady(path) error` | is the pre-start hook (a FIXED file from the bundle) in place — register only then | The agent never installs the hook (Proxmox runs it as root); a missing hookscript stops a guest start, so never `Register` without `SnippetReady` |
### Disk / format safety (role gates, durable IDs, format guards)
@@ -59,8 +66,8 @@
|---|---|---|---|---|
| `IntentStore` (`Get/SetEnrolled/SetEjected/SetDecommissioned/OnAbsent`) | internal/storage/intent.go | `OpenIntentStore(path)` | drive intent (4-state self-heal) | Keyed by durable-id only; `OnAbsent` is the ONLY ejected→enrolled path; refuses empty ids |
| `GuestBindStore` (`Record/Remove/Guests`) | internal/localapi/guestbindstore.go | `OpenGuestBindStore(path)` | per-guest enrolled binds (F9 re-assert) | Same tmp+rename 0600 pattern as IntentStore |
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) <-chan error` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank |
| `TokenStore.Mint` / `Lookup` | internal/localapi/tokenstore.go | `Mint(vmid) (plaintext, error)` | per-guest local-API tokens | Only the SHA-256 hash persists (fsync'd append log); constant-time compare on lookup; plaintext returned exactly once. Lookup RELOADS the file once on a miss (v0.63.0, B3): the one-shot provisioner mints into the same file the daemon indexes — cross-process coherence without a restart; append-only size check bounds the re-read |
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) (*formatJob, <-chan error)` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank; on success `job.FSUUID` = the new filesystem's UUID, read back only when the durable id still resolves to the formatted device (R-25) — read it only after `done` delivers |
| `TokenStore.Mint` / `Lookup` | internal/localapi/tokenstore.go | `Mint(vmid) (plaintext, error)` | per-guest local-API tokens | Only the SHA-256 hash persists (fsync'd append log); constant-time compare on lookup; plaintext returned exactly once. Lookup stats the file on EVERY call and reloads BEFORE answering when the append-only log grew (R-269; was reload-on-miss only, v0.63.0 B3, which let a token rotated out by another process keep authorizing as a map hit): the one-shot provisioner mints into the same file the daemon indexes — cross-process coherence both ways without a restart; unchanged size = no re-read. Pinned by `TestTokenStore_RotatedOutTokenRejectedFirst` |
| `FileNonceStore.SeenOrRecord` | internal/authz/noncestore.go | `SeenOrRecord(nonce, exp) bool` | durable anti-replay | fsync'd before returning false; prune only after exp |
| `Journal` (`Append/Latest/InFlight/AlreadyApplied`) | internal/reconcile/journal.go | `OpenJournal(path)` | op journal + idempotency + crash recovery | `Recover` consumes `InFlight()`; scratch entries special-cased |
@@ -81,6 +88,8 @@
| Symbol | File | Short signature | Use for | Gotchas |
|---|---|---|---|---|
| `pvegate.Write` / `pvegate.Step` | internal/pvegate/pvegate.go | `Write(ctx) (release, waited, err)` / `Step(ctx) (end, err)` | R-812 option A: keep the agent's own /etc/pve writes out of a Proxmox package step (pmxcfs restarts) | Already wired at the two chokepoints — `Client.doBody` (every non-GET) and `ExecRunner.RunStdin` (`WritesEtcPVE`: pct config verbs, pvesm, pveum, felhom-pbs-apply create/reconcile). A new root CLI that writes /etc/pve goes into `WritesEtcPVE`, never its own lock. Never take `Step` around anything but the wrapper call (`Leg.runPVE`) — a `Write` inside a `Step` deadlocks until its context ends. |
| `osupdate.KernelVerdict` / `Leg.KernelAfterBoot` | internal/osupdate/kernel.go | `KernelVerdict(before, after, tunnel, hubReached) (ok, why)` | R-836: THE one-shot-boot rule — the host health rule (`HostHealthVerdict`) AND the box reached the hub since the boot | The only judge of a new kernel. Never reboot the host from Go: every reboot is the wrapper's (`kernel-reboot`, `kernel-revert`), and only for a staged step. A new kernel-lane state lives in the wrapper's `/var/lib/felhom-kernel/state.json`, never in the agent's own files (the agent can write those). |
| `Client.WaitTask` | internal/proxmox/task.go | `WaitTask(ctx, upid, opts) (TaskStatus, error)` | asserting EVERY mutating op | POST 200 ≠ success; authz can fail at task exec; `AllowWarnings` opt-in |
| `Client.Pool` | internal/proxmox/query.go | `Pool(ctx, name) (PoolInfo, error)` | felhom-pool membership (the ownership registry, A1) | Needs `Pool.Audit` at `/pool/<name>` (host-install v1.9.0+); `Pool.Allocate` does NOT satisfy the read; members can be storages (type `storage`, vmid 0) — filter them |
| `Client` mutate wrappers (`RestoreLXC/Vzdump/DestroyLXC/Snapshot/Rollback/SetConfig/ResizeLXC/Start/Stop`) | internal/proxmox/mutate.go | return `(upid, error)` | all API mutations | Async → always pair with WaitTask; route via gate/queue, not ad-hoc |
@@ -114,7 +123,7 @@
| Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only |
| Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set |
| Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction |
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`; R-856 `crashGuardStatePath` — GET /host/crash-guard's state file, internal/localapi/crashguard.go) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
| `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) |
| `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs |
| `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s |
@@ -150,15 +159,18 @@
| `storage.HostOps` | internal/storage/hostops.go | `*SudoHostOps` (prod), `NoopHostOps` (degraded) | fakes in internal/storage/observe_test.go, watchdog_test.go |
| `storage.HostReader` | internal/storage/hostread.go | `*ProcHostReader` | `fakeHostReader` internal/localapi/disks_test.go; internal/storage/role_test.go. v0.87.0: `BlockSlaves(name)` lists `/sys/block/<name>/slaves` (root-free) — backs the `SystemDisks` dm/md walk (`physicalDisksOf`/`walkSlaves`, role.go); per-branch conservatism: an unresolvable slave fails the WHOLE walk → all-system fail-safe. NEVER weaken the signature test `TestSystemDisks_WalkTopologies` (root-backing disk always in the system set). |
| `localapi.DiskOps` / `StorageGate` / `GuestAttacher` / `GuestLister` | internal/localapi/disks.go | `*storage.SudoHostOps`; `storageGateAdapter` (cmd/felhom-agent/main.go); `*GuestBinder`; `*proxmox.Client` | `fakeDiskOps`/`fakeGate`/`fakeGuestAttacher`/`fakeGuestList` internal/localapi/disks_test.go |
| `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring |
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
| `backup.BackupSuccessState` | internal/backup/backup_state.go | `RecordBackupSuccess(target, b)` / `LastKnownSuccess(target, vmid)` | Newest SUCCESSFUL backup per tier+guest, persisted (atomic tmp+rename) — the due-check's fallback when the tier's storage cannot be read after a restart (R-894) | **Read ONLY when the storage cannot be read** — a storage that answers is the ground truth (R-84), and an archive absent there must make the tier due even when this file remembers one. A saved copy older than the cadence still reads due. Only successes are written. |
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
| `localapi.StaleLockController` | internal/localapi/stalelock.go | `*staleLockController` (Client + Runner + pool) | `fakeStaleLock` (Server-level) stalelock_test.go; `fakeStaleLockAPI` (controller-level, tests the A1 pool intersect) stalelock_pool_test.go |
| `localapi.GuestExecutor` | internal/localapi/controllerswap.go | `*GuestBinder` (pct exec) | `fakeGuestExec` internal/localapi/controllerswap_test.go |
| `localapi.GuestExecutor` | internal/localapi/controllerswap.go | `*GuestBinder` (pct exec; the image write via `felhom-priv-apply controller-image`) | `fakeGuestExec` internal/localapi/controllerswap_test.go |
| `guestnet.Runner` / `guestnet.GuestSource` (R-54, v0.92.0) | internal/guestnet/{probe,watchdog}.go | `*proxmox.ExecRunner`; the POOL-VERIFIED `localapi.StaleLockController.Guests` (ListLXC ∩ felhom pool, audit A1) | `scriptedRunner` + `fakeGuests` internal/guestnet/watchdog_test.go. **Never wire a bare `ListLXC` here** — under a broad token that would run dhclient inside a co-tenant's container. Every assertion is an exec COUNT, and the load-bearing ones are the negatives: a static guest, an unprobeable guest, a boot-race guest and an unproven guest list must record **zero** heal calls |
| `guestnet.Watchdog.SetDampers` / `now` (clock seam) | internal/guestnet/watchdog.go | config `guest_net.*`; `now` defaults to `time.Now` | tests advance a manual clock (the storage-watchdog pattern) and assert the heal ceilings EXACTLY — ≥10 min apart, ≤3/hour, and ≤30 over a scripted 10 hours of permanent failure. A damper with no test is a comment |
| `hub.GuestNetReporter` (R-54) | internal/hub/collect.go | `*guestnet.Watchdog` (`GuestNetStatus`) | internal/hub/collect_guestnet_test.go asserts the stanza through the PRODUCTION `Collect` path AND that the `guest_net` key is ABSENT from the wire when no reporter is wired — an always-present empty stanza would make "not wired" and "found nothing" the same signal, which is the shape v0.91.0 hid behind |
+53
View File
@@ -0,0 +1,53 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"testing"
)
// R-444: the weekly trim has the guestnet shape (component + reporter seam + goroutine), so its wiring is asserted
// from the AST like TestMainWiresGuestNetWatchdog — a unit-green trim job that main.go never starts is the inert-seam
// defect. It must also share the ONE heavy-op gate (heavyOps), or it could run beside a backup.
func TestMainWiresGuestDiskTrim(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, 0)
if err != nil {
t.Fatalf("parse main.go: %v", err)
}
var constructedWithGate, reporterWired, started bool
ast.Inspect(f, func(n ast.Node) bool {
switch node := n.(type) {
case *ast.CallExpr:
if fn, ok := node.Fun.(*ast.SelectorExpr); ok {
switch fn.Sel.Name {
case "New":
if pkg, ok := fn.X.(*ast.Ident); ok && pkg.Name == "fstrim" && len(node.Args) >= 3 {
if id, ok := node.Args[2].(*ast.Ident); ok && id.Name == "heavyOps" {
constructedWithGate = true
}
}
case "SetGuestDiskTrimReporter":
reporterWired = true
}
}
case *ast.GoStmt:
if sel, ok := node.Call.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "Run" {
if id, ok := sel.X.(*ast.Ident); ok && id.Name == "diskTrim" {
started = true
}
}
}
return true
})
if !constructedWithGate {
t.Error("main.go never calls fstrim.New(..., heavyOps, ...) — no trim job, or one outside the heavy-op gate")
}
if !reporterWired {
t.Error("main.go never calls collector.SetGuestDiskTrimReporter — the guest_disk_trim stanza never reaches the hub")
}
if !started {
t.Error("main.go never starts the trim job with `go diskTrim.Run(ctx)`")
}
}
+309 -110
View File
@@ -22,6 +22,7 @@ import (
"os/exec"
"os/signal"
"path/filepath"
"regexp"
"strconv"
"strings"
"sync"
@@ -37,6 +38,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
"gitea.dooplex.hu/admin/felhom-agent/internal/fasttick"
"gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd"
"gitea.dooplex.hu/admin/felhom-agent/internal/fstrim"
"gitea.dooplex.hu/admin/felhom-agent/internal/guesthook"
"gitea.dooplex.hu/admin/felhom-agent/internal/guestnet"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
@@ -44,11 +46,11 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/localapi"
applog "gitea.dooplex.hu/admin/felhom-agent/internal/log"
"gitea.dooplex.hu/admin/felhom-agent/internal/mgmtplane"
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
"gitea.dooplex.hu/admin/felhom-agent/internal/pbs"
"gitea.dooplex.hu/admin/felhom-agent/internal/pbsdr"
"gitea.dooplex.hu/admin/felhom-agent/internal/poke"
"gitea.dooplex.hu/admin/felhom-agent/internal/provision"
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/restorespace"
@@ -133,39 +135,38 @@ func main() {
return
}
var (
cfgPath string
selftest selftestFlag
vmid int
watch time.Duration
archive string
mode string
hostname string
keep bool
rootfsGrow int
dataVolGrow int
dataVolMount string
sysDataGrow int
sysDataMount string
cores int
memoryMB int
pbsStorage string
paperkey bool
offline bool
upload bool
custID string
custDomain string
custName string
custEmail string
hubPassword string
blobPath string
expectedFP string
keyDest string
installWGKey bool
idBundlePath string
directivePath string
swapImage string
outputMode string
showVersion bool
cfgPath string
selftest selftestFlag
vmid int
watch time.Duration
archive string
mode string
hostname string
keep bool
rootfsGrow int
dataVolGrow int
dataVolMount string
sysDataGrow int
sysDataMount string
cores int
memoryMB int
pbsStorage string
paperkey bool
offline bool
upload bool
custID string
custDomain string
custName string
custEmail string
hubPassword string
blobPath string
expectedFP string
keyDest string
installWGKey bool
idBundlePath string
swapImage string
outputMode string
showVersion bool
)
flag.StringVar(&cfgPath, "config", envOr("FELHOM_AGENT_CONFIG", "/etc/felhom-agent/agent.json"), "path to the agent config file (JSON)")
flag.Var(&selftest, "selftest", "run a self-test and exit: bare/`read` = read-only queries; `task` = reversible mutating exercise (needs -vmid); `hub` = one collect+report; `storage` = observe storage (+ -watch); `backup` = one-shot backup of -vmid; `restore-test` = restore→boot→verify→teardown of -archive (or newest backup); `restore-test-due` = READ-ONLY: print the per-tier due verdict the scheduler would act on, with its cost; `pbs-verify` = trigger a PBS verify + print snapshot records; `bring-up` = restore→reset identity→size→start link-up of -archive into -vmid (needs -mode/-archive/-vmid; optional -cores/-memory cap; tears down unless -keep); `provision` = full slice-8A chain: bring-up provision + mint token + populate bootstrap config mount (needs -archive/-vmid/-customer-id/-hub-password; optional -rootfs-grow/-datavol-grow/-cores/-memory (-sysdata-grow is deprecated: folded into -datavol-grow); keeps the guest)")
@@ -196,7 +197,6 @@ func main() {
flag.StringVar(&keyDest, "keydest", "", "for --selftest=escrow-consume: where to install the recovered key (0600)")
flag.BoolVar(&installWGKey, "install-wg-key", false, "for --selftest=identity-consume: ALSO install the recovered wg_private_key into wgtunnel's key file (S5 DR; create-only, refuses to overwrite)")
flag.StringVar(&idBundlePath, "identity-bundle", "", "for --selftest=escrow-create: a 0600 JSON file {tunnel_token,pbs_token} to ALSO escrow under R (10D)")
flag.StringVar(&directivePath, "directive", "", "for --selftest=escrow-create: a JSON file with the non-secret DR directive (pbs repo/ns, expected fingerprint, tunnel id)")
flag.StringVar(&custID, "customer-id", "", "for --selftest=provision: the customer id — the hub config-pull target, baked into the guest's bootstrap")
flag.StringVar(&hubPassword, "hub-password", "", "for --selftest=provision: the customer's hub RETRIEVAL PASSPHRASE (SECRET) — baked into bootstrap.json so the controller pulls its config (and the customer-scoped hub key) from the hub. The customer must already exist in the hub.")
flag.StringVar(&swapImage, "image", "", "for --selftest=controller-swap: the target controller image ref (gitea.dooplex.hu/admin/felhom-controller:<semver>) — must already be pulled in the guest")
@@ -245,6 +245,10 @@ func main() {
os.Exit(runSelftestRestoreTestDue(context.Background(), cfg, logger))
case "os-update":
os.Exit(runSelftestOSUpdate(context.Background(), cfg, logger, vmid))
case "os-facts":
os.Exit(runSelftestFacts(context.Background(), cfg, logger, vmid))
case "live-restore":
os.Exit(runSelftestLiveRestore(context.Background(), cfg, logger, vmid))
case "pbs-verify":
os.Exit(runSelftestPBSVerify(context.Background(), cfg, logger))
case "lanresolver":
@@ -263,7 +267,7 @@ func main() {
SysDataGrowGB: sysDataGrow, SysDataMount: sysDataMount, Cores: cores, MemoryMB: memoryMB},
}))
case "escrow-create":
os.Exit(runSelftestEscrowCreate(context.Background(), cfg, logger, pbsStorage, paperkey, offline, upload, idBundlePath, directivePath, outputMode))
os.Exit(runSelftestEscrowCreate(context.Background(), cfg, logger, pbsStorage, paperkey, offline, upload, idBundlePath, outputMode))
case "escrow-consume":
os.Exit(runSelftestEscrowConsume(context.Background(), logger, blobPath, expectedFP, keyDest))
case "identity-consume":
@@ -786,6 +790,9 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
pbsTargets := pbsTargetsFromPVE(cfg, px, logger)
pbsReporter := pbs.NewLiveSnapshotReporter(pbsTargets, pbsStore, pbs.DefaultLiveSnapshotTimeout, logger)
collector := hub.NewCollector(px, newTunnelProber(cfg, px), observer, backupStore, backupStore, pbsReporter, cfg.Hub.HostID, version, logger)
// R-366 slice 2: the restore-test's ledger of archives written with another key → the host report.
foreignKeys := backup.NewForeignKeyLedger()
collector.SetForeignKeyArchiveReporter(foreignKeys)
collector.SetBackupTargetResolver(primaryBackupTargetOf(cfg)) // R-109: the recipe names the live target
// Privileged-capability self-check (v0.44.0): probe the sudoers grants the non-root agent
// depends on. The probe runs `sudo -n -l` LITERALLY (a policy LIST, never executing the
@@ -845,6 +852,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// OS updates, guest fast lane (agent v0.140.0, `11-os-updates.md` §8 step 2): the leg consumes the hub's
// os_update block and runs after each successful primary whole-guest backup (wired on the local API below).
osLeg := newOSLeg(cfg, client, px, logger)
collector.SetSystemReporter(&factsReporter{leg: osLeg, guest: firstGuest(px)}) // R-852: the versions
desiredSyncer.AddConsumer(osLeg)
// S5: consume a host_loss restore_directive into an inspectable restore PLAN (derive + surface,
// execute nothing). The recipe is fetched on-demand (rare directive) via a fresh Collect.
@@ -863,6 +871,18 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
logger.Info("felhom-agent daemon starting",
"version", version, "host_id", cfg.Hub.HostID, "hub_url", cfg.Hub.URL,
"interval_s", hcfg.PollSeconds) // hub key intentionally not logged
// R-868 (v0.144.0): an OS pass whose agent was killed kept its report on disk — send it now (a pass that starts
// first sends it itself; the pass lock keeps the two apart).
// v0.144.1: and again every 5 minutes — the wrapper of a killed pass can finish AFTER the restart (measured).
go osLeg.SendUnsentLoop(ctx, 5*time.Minute, func(n int) {
logger.Info("osupdate: sent kept report(s)", "count", n)
})
// R-836 (`09` §3 decision 172): what became of a kernel step across this boot; on a one-shot boot of a new kernel,
// judge it (the host health rule + the hub reached) for KernelJudgeWait, then make it the default or revert ONCE.
// The wait: measured 2026-10-07 (`audits/kernel-lane-2026-10-07/B/`) — every container healthy 68 s after the
// reboot on demo-felhom and 272 s on demo-hp (the hub reached at 63 s / 189 s); 20 minutes leaves room for a slow
// network and stays under the hub's 30-minute host_stale. On a box without the kernel lane (an older wrapper, a BYO host) the check is refused and logged.
go osLeg.KernelAfterBoot(ctx, 0, osupdate.KernelJudge{Wait: osupdate.DefaultKernelJudgeWait})
// Reconcile (slice 4) runs alongside the hub loop, sharing the per-guest queue
// (doc 03 §10). At slice 4 the desired-state provider is empty (no hub serving
@@ -1004,7 +1024,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// with the local API so a backup and a restore-test can never run together.
rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json"))
heavyOps := &backup.InFlight{}
scheduler := buildRestoreTestScheduler(cfg, px, engine, backupStore, rtState, heavyOps, logger)
scheduler := buildRestoreTestScheduler(cfg, px, engine, backupStore, rtState, heavyOps, foreignKeys, logger)
// R-189: the host report's restore_tests[] must survive an agent restart. The in-memory store
// holds only this process's latest run, and under per-archive due-ness the agent will not
// re-test an archive it has already proven — so without this the hub can report a tier unproven
@@ -1087,7 +1107,46 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// recurring clobber before it locks the box out. Port 22 for G1 (H1 passes the felhom-sshd port).
collector.SetMgmtPlaneReporter(mgmtplane.NewReporter(mgmtplane.DefaultPrivsepDir, mgmtplane.DefaultHealMarker, mgmtplane.DefaultSshdPort))
jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec}, cfg.Hub.HostID, logger)
// Agent v0.142.0: a signed Docker engine step (`11` §5.8) — ring 1 and every undo; under the heavy-op gate.
dockerExec := osupdate.DockerStepExecutor{Leg: osLeg, Guest: firstGuest(px),
Gate: func(ctx context.Context) (func(), error) {
release, busy, ok := heavyOps.TryAcquire("os-docker-step")
if !ok {
return nil, fmt.Errorf("busy: %s", busy)
}
return release, nil
}}
// R-812 option A: a signed Proxmox package step (`11` §5.10) — ring 1; under the heavy-op gate and the /etc/pve gate.
pveExec := osupdate.PVEStepExecutor{Leg: osLeg, Guest: firstGuest(px),
Gate: func(ctx context.Context) (func(), error) {
release, busy, ok := heavyOps.TryAcquire("os-pve-step")
if !ok {
return nil, fmt.Errorf("busy: %s", busy)
}
return release, nil
}}
// R-836 (`09` §3 decision 172): a signed kernel step STAGES a kernel on a ring-1 box (install + the one-shot flag,
// never a reboot — the night leg reboots it on a night the household was told about); under the heavy-op gate.
kernelExec := osupdate.KernelStepExecutor{Leg: osLeg, Guest: firstGuest(px),
Gate: func(ctx context.Context) (func(), error) {
release, busy, ok := heavyOps.TryAcquire("os-kernel-step")
if !ok {
return nil, fmt.Errorf("busy: %s", busy)
}
return release, nil
}}
// Agent v0.143.0 (R-840): the config bundle — the box's root-owned files by a signed job; the wrapper verifies it.
bundleExec := osupdate.ConfigUpdateExecutor{Leg: osLeg, URLTemplate: suCfg.URLTemplate, Username: suCfg.Username, Token: suCfg.Token,
// The capability probe confirms from the agent's side: `sudo -l` lists every command the new sudoers grants.
AfterInstall: func(ctx context.Context) {
ok, total, degraded := capability.Summarize(probeAll(ctx))
names := make([]string, 0, len(degraded))
for _, d := range degraded {
names = append(names, d.Name)
}
logger.Warn("osupdate: capability probe after the config bundle", "ok", ok, "total", total, "degraded", strings.Join(names, ","))
}}
jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec, dockerExec, pveExec, kernelExec, bundleExec}, cfg.Hub.HostID, logger)
loop.SetEnvelopeObserver(hub.MultiObserver(desiredSyncer, jobsRunner))
// Controller-driven escrow ceremony (v0.88.0): static config facts + the LATE-BOUND DR gate —
@@ -1127,7 +1186,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
return
case <-time.After(90 * time.Second):
}
_, _ = osLeg.Run(ctx, vmid, "night")
_ = osLeg.Run(ctx, vmid, "night")
})
}
if localTokens != nil {
@@ -1447,6 +1506,25 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
}
go runJanitor(ctx, jd)
}
// R-444 (`09` §3 decision 139): the weekly guest disk trim — `pct fstrim <vmid>` of every owned, running guest,
// Wednesday from 10:00 local, daytime only, under the one-heavy-op gate. Not part of the errc fan-out: a trim job
// must never be able to bring the agent down.
if cfg.DiskTrim.Enabled() {
dtMode := proxmox.RunnerMode(cfg.Privileged.Mode)
if dtMode == "" {
dtMode = proxmox.RunnerSudo
}
dtRunner := &proxmox.ExecRunner{Mode: dtMode, SudoPath: cfg.Privileged.SudoPath}
dtGuests := localapi.NewStaleLockController(px, dtRunner, reconcile.DefaultPool, logger)
if dtGuests != nil {
diskTrim := fstrim.New(dtRunner, dtGuests, heavyOps,
filepath.Join(cfg.OOB.WithDefaults().StateDir, "guest-disk-trim.json"), logger)
collector.SetGuestDiskTrimReporter(diskTrim)
go diskTrim.Run(ctx)
}
} else {
logger.Info("fstrim: weekly guest disk trim disabled by config (disk_trim.disable)")
}
if lanLoop != nil {
lanServers = 1
go func() { errc <- lanLoop.Run(ctx) }()
@@ -1648,7 +1726,7 @@ func primaryBackupTargetOf(cfg config.Config) func() hub.ConfiguredBackupTarget
// disables the cadence (returns a scheduler that just waits) when the cadence is off or the
// scratch band / restore storage is invalid — a misconfig must not crash the daemon, and the
// machinery still works on-demand via --selftest=restore-test.
func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *reconcile.Engine, store *backup.Store, rtState *backup.RestoreTestState, inFlight *backup.InFlight, logger *slog.Logger) *backup.Scheduler {
func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *reconcile.Engine, store *backup.Store, rtState *backup.RestoreTestState, inFlight *backup.InFlight, foreign *backup.ForeignKeyLedger, logger *slog.Logger) *backup.Scheduler {
// R-86: this is the EVALUATION interval, not the trigger. What decides a test happens is the
// per-archive due-check in internal/backup/restoretest_due.go.
cadence := cfg.Backup.RestoreTestEvalInterval()
@@ -1667,6 +1745,9 @@ func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *re
min, max := cfg.Backup.ScratchBand()
target := cfg.Backup.BackupTarget()
runner := backup.NewBackupRunner(px, target, "", "felhom restore-test", "", logger)
if foreign != nil {
runner.SetForeignKeyLedger(foreign) // R-366 slice 2: the pick records archives written with another key
}
// Every configured tier is a rotation candidate, not just the primary.
cfgTiers, _ := cfg.Backup.BackupTiers() // warnings already logged where the tiers are armed
tierIDs := make([]string, 0, len(cfgTiers))
@@ -1849,12 +1930,15 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
Store: store,
Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
Tokens: tokens,
BackupCadence: cfg.Backup.BackupCadence(),
// R-894: the newest success per tier on disk — the due-check's fallback when the storage cannot be
// read right after a restart. Same state dir as restore-test-state.json.
LastKnownBackups: backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json")),
Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
Tokens: tokens,
BackupCadence: cfg.Backup.BackupCadence(),
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
Disks: hostOps,
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
@@ -1869,7 +1953,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
ConfigPath: cfg.SourcePath,
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
SmbCredsDir: cfg.Privileged.SmbCredsDir,
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
@@ -2209,7 +2293,7 @@ func runSelftestRestoreTestDue(ctx context.Context, cfg config.Config, logger *s
return 1
}
rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json"))
sched := buildRestoreTestScheduler(cfg, px, nil, backup.NewStore(), rtState, &backup.InFlight{}, logger)
sched := buildRestoreTestScheduler(cfg, px, nil, backup.NewStore(), rtState, &backup.InFlight{}, nil, logger)
fmt.Printf("eval_interval=%s settle=%s\n", cfg.Backup.RestoreTestEvalInterval(), cfg.Backup.RestoreTestSettle())
start := time.Now()
@@ -2635,7 +2719,6 @@ type escrowCeremonyOpts struct {
offline bool
upload bool
identityBundlePath string
directivePath string
}
// escrowCeremonyOutcome is the shared core's result. R is the ONLY secret; Sum mirrors the
@@ -2666,6 +2749,9 @@ func (e *escrowCeremonyErr) Error() string { return e.err.Error() }
// restic password auto-attach), Create (R + self-verified blob), wipe the staged secret, upload
// when asked. It PRINTS NOTHING — the output-mode shells own every byte of stdout/stderr. R is
// returned for the caller to surface exactly once; escrow.Create never logs it and neither do we.
// pbsStorageIDRe is a PVE storage id (letters, digits, '-', '_', '.'; starts with a letter) — never a path (R-861).
var pbsStorageIDRe = regexp.MustCompile(`^[A-Za-z][A-Za-z0-9_.-]{0,63}$`)
func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger, opts escrowCeremonyOpts) (escrowCeremonyOutcome, *escrowCeremonyErr) {
var out escrowCeremonyOutcome
storage := opts.storage
@@ -2675,17 +2761,26 @@ func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger,
if storage == "" {
return out, &escrowCeremonyErr{kind: "usage", err: fmt.Errorf("selftest=escrow-create requires -storage <pbs-storage-id> (or escrow.pbs_storage_id)")}
}
// R-861 (v0.146.0): this runs as ROOT through FELHOM_ESCROW, but agent.json is owned by the agent user. So the
// paths a root run reads never come from it: the PVE secret dir and the WireGuard state dir are the fixed defaults,
// and the storage id is a plain PVE id (no slash, no dot-dot) — a crafted id or dir would read another root file.
if os.Geteuid() == 0 {
cfg.Backup.PBSSecretDir = ""
cfg.WGTunnel.StateDir = ""
}
if !pbsStorageIDRe.MatchString(storage) {
return out, &escrowCeremonyErr{kind: "usage", err: fmt.Errorf("selftest=escrow-create: storage id %q is not a plain PVE storage id", storage)}
}
out.Storage = storage
keyPath := cfg.Backup.PBSEncKeyPath(storage)
if _, err := os.Stat(keyPath); err != nil {
return out, &escrowCeremonyErr{kind: "setup", err: fmt.Errorf("PBS key for %q not found (%s): %v", storage, keyPath, err)}
}
// Slice 10D.1: optionally ALSO wrap the identity bundle under the same R, and carry the non-secret
// directive for the hub. The bundle file is a 0600 secret (tunnel/pbs tokens); the directive is
// non-secret (pbs repo/ns, expected fingerprint, tunnel id).
// Slice 10D.1: optionally ALSO wrap the identity bundle under the same R. The bundle file is a 0600 secret
// (tunnel/pbs tokens). The non-secret "directive" that used to ride along is retired (R-105, `09` §3 decision 169:
// nothing read it; the DR path reads the recipe, tenantsync and the escrow blob).
var identity *escrow.IdentityBundle
var directive json.RawMessage
if opts.identityBundlePath != "" {
raw, err := os.ReadFile(opts.identityBundlePath)
if err != nil {
@@ -2696,11 +2791,6 @@ func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger,
return out, &escrowCeremonyErr{kind: "setup", err: fmt.Errorf("identity bundle is not valid JSON {tunnel_token,pbs_token}: %v", err)}
}
identity = &b
if opts.directivePath != "" {
if d, err := os.ReadFile(opts.directivePath); err == nil && json.Valid(d) {
directive = d
}
}
}
// S3: auto-inject the offsite WG private key into the escrowed identity when the key file
// exists — a NEW escrow run should always capture the live tunnel identity. Field NAME only
@@ -2779,7 +2869,7 @@ func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger,
ResticPwSealed: resticStaged,
}
if opts.upload {
if err := uploadEscrowBlob(ctx, cfg, res, directive, resticPwSHA256); err != nil {
if err := uploadEscrowBlob(ctx, cfg, res, resticPwSHA256); err != nil {
// R is minted and the blob self-verified — only the hub leg failed. kind "upload" lets
// the text shell keep the pre-extraction order (R surfaced, THEN the failure).
return out, &escrowCeremonyErr{kind: "upload", err: err}
@@ -2836,7 +2926,7 @@ func printEscrowTextRBlock(out *escrowCeremonyOutcome) {
// NOTHING else there; every human/info line goes to stderr; failures exit non-zero with no
// partial JSON. This is the controller-driven ceremony's parse surface (spike §2.3: the text
// banner is positionally brittle).
func runSelftestEscrowCreate(ctx context.Context, cfg config.Config, logger *slog.Logger, storage string, paperkey, offline, upload bool, identityBundlePath, directivePath, outputMode string) int {
func runSelftestEscrowCreate(ctx context.Context, cfg config.Config, logger *slog.Logger, storage string, paperkey, offline, upload bool, identityBundlePath, outputMode string) int {
switch outputMode {
case "", "text", "json":
default:
@@ -2854,7 +2944,7 @@ func runSelftestEscrowCreate(ctx context.Context, cfg config.Config, logger *slo
out, cerr := escrowCeremony(ctx, cfg, logger, escrowCeremonyOpts{
storage: storage, paperkey: paperkey, offline: offline, upload: upload,
identityBundlePath: identityBundlePath, directivePath: directivePath,
identityBundlePath: identityBundlePath,
})
if cerr != nil {
switch cerr.kind {
@@ -3037,20 +3127,19 @@ type escrowUploadRequest struct {
BlobB64 string `json:"blob_b64"` // base64 of the opaque R-wrapped blob (ciphertext)
KeyFingerprint string `json:"key_fingerprint"` // for operator display only
Posture string `json:"posture"` // e.g. "zero_knowledge"
// Slice 10D.1 — optional DR bundle (identity escrow + non-secret directive). Omitted in slice-7.
IdentityBlobB64 string `json:"identity_blob_b64,omitempty"`
DirectiveJSON json.RawMessage `json:"directive,omitempty"`
CreatedAt string `json:"created_at"` // RFC3339
// Slice 10D.1 — optional identity escrow. Omitted in slice-7. (The `directive` is retired — R-105.)
IdentityBlobB64 string `json:"identity_blob_b64,omitempty"`
CreatedAt string `json:"created_at"` // RFC3339
// SLICE 3 — sha256 hex of the offsite restic repo password sealed in the identity blob (present only
// when a staged password was folded in). Non-reversible hash of a 256-bit random secret — safe to
// store/serve; lets the controller VERIFY "the escrow covers the CURRENT key" and auto-confirm.
ResticPwSHA256 string `json:"restic_pw_sha256,omitempty"`
}
// uploadEscrowBlob PUTs the opaque blob (and, for 10D, the identity blob + non-secret directive) to
// uploadEscrowBlob PUTs the opaque blob (and, for 10D, the identity blob) to
// the hub, authed with the per-host key. The hub stores ciphertext + non-secret fields; no usable
// secret leaves the agent.
func uploadEscrowBlob(ctx context.Context, cfg config.Config, res escrow.CreateResult, directive json.RawMessage, resticPwSHA256 string) error {
func uploadEscrowBlob(ctx context.Context, cfg config.Config, res escrow.CreateResult, resticPwSHA256 string) error {
if cfg.Hub.URL == "" || cfg.Hub.HostID == "" || cfg.Hub.APIKey == "" {
return fmt.Errorf("hub not configured (url/host_id/api_key)")
}
@@ -3063,7 +3152,6 @@ func uploadEscrowBlob(ctx context.Context, cfg config.Config, res escrow.CreateR
}
if len(res.IdentityBlob) > 0 {
upReq.IdentityBlobB64 = base64.StdEncoding.EncodeToString(res.IdentityBlob)
upReq.DirectiveJSON = directive
}
body, _ := json.Marshal(upReq)
url := strings.TrimRight(cfg.Hub.URL, "/") + "/api/v1/hosts/" + cfg.Hub.HostID + "/escrow"
@@ -3498,10 +3586,12 @@ func (f *selftestFlag) Set(v string) error {
f.mode = "controller-swap"
case "os-update":
f.mode = "os-update"
case "os-facts", "live-restore": // agent v0.142.0
f.mode = v
case "wgtunnel": // dispatched since S3 but refused here until 2026-10-04 (TestSelftestFlag_AcceptsEveryDispatchedMode)
f.mode = "wgtunnel"
default:
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|restore-test-due|pbs-verify|lanresolver|wgtunnel|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap|os-update)", v)
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|restore-test-due|pbs-verify|lanresolver|wgtunnel|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap|os-update|os-facts|live-restore)", v)
}
return nil
}
@@ -3548,39 +3638,99 @@ func runSelftestOSUpdate(ctx context.Context, cfg config.Config, logger *slog.Lo
return 1
}
leg := newOSLeg(cfg, client, px, logger)
resp, err := client.FetchDesiredState(ctx)
if err != nil {
fmt.Fprintln(os.Stderr, "selftest=os-update: desired state:", err)
blk, source, ok := selftestOSBlock(ctx, client, leg.PlanDir)
if !ok {
fmt.Fprintln(os.Stderr, "selftest=os-update:", source)
return 1
}
leg.SetBlock(resp.DesiredState.OSUpdate)
leg.SetBlock(blk)
b := leg.Block()
fmt.Printf("=== felhom-agent %s selftest=os-update vmid=%d ring=%d enabled=%v guest-release=%v host-release=%v appliance=%v ===\n",
version, vmid, b.Ring, b.Enabled, b.Release != nil, b.HostRelease != nil, leg.Appliance)
fmt.Printf("=== felhom-agent %s selftest=os-update vmid=%d ring=%d enabled=%v guest-release=%v host-release=%v appliance=%v block=%s ===\n",
version, vmid, b.Ring, b.Enabled, b.Release != nil, b.HostRelease != nil, leg.Appliance, source)
start := time.Now()
g, h := leg.Run(ctx, vmid, "debug")
for _, rep := range []osupdate.Report{g, h} {
pass := leg.Run(ctx, vmid, "debug")
worst := pass.Guest
for i, rep := range []osupdate.Report{pass.Guest, pass.Host, pass.Docker} {
if rep.Layer == "" {
fmt.Println(" host step: skipped (see the log line above)")
fmt.Printf(" %s step: skipped (see the log line above)\n", []string{"guest", "host", "docker"}[i])
continue
}
printJSON("os-update report ("+rep.Layer+")", map[string]any{"run_id": rep.RunID, "ring": rep.Ring, "release_id": rep.ReleaseID,
"mode": rep.Mode, "outcome": rep.Outcome, "healthy": rep.Healthy, "health_reason": rep.HealthReason,
"upgraded": rep.Upgraded, "pending": len(rep.Pending), "not_covered": rep.NotCovered,
"restart_needed": rep.RestartNeeded, "reboot_needed": rep.RebootNeeded, "wrapper_seconds": rep.PassSeconds, "refused": rep.Refused})
"restart_needed": rep.RestartNeeded, "reboot_needed": rep.RebootNeeded, "wrapper_seconds": rep.PassSeconds,
"refused": rep.Refused, "docker_engine": rep.DockerEngine, "authority": rep.Authority})
if !(rep.Outcome == "applied" || rep.Outcome == "nothing" || rep.Outcome == "inventory" || rep.Outcome == "skipped") {
worst = rep
}
}
fmt.Printf(" pass took %s\n", time.Since(start).Round(100*time.Millisecond))
rep := g
if h.Layer != "" && !(h.Outcome == "applied" || h.Outcome == "nothing" || h.Outcome == "inventory") {
rep = h
}
switch rep.Outcome {
switch worst.Outcome {
case "applied", "nothing", "inventory", "skipped":
return 0
}
return 1
}
// selftestOSBlock is the debug pass's os_update block (R-866, v0.144.0): the hub's, fetched fresh; when the hub cannot
// be reached, the block the daemon last saved — named in the selftest's first line, so a pass with the hub away can
// be exercised by hand (the daemon's own leg already ran from its last block; the selftest stopped). No saved block
// and no hub → not run (ok=false), never a guessed block.
func selftestOSBlock(ctx context.Context, f interface {
FetchDesiredState(context.Context) (*hub.DesiredStateResponse, error)
}, planDir string) (*hub.WireOSUpdate, string, bool) {
resp, err := f.FetchDesiredState(ctx)
if err == nil {
return resp.DesiredState.OSUpdate, "hub", true
}
saved, at, ok := osupdate.LoadSavedBlock(planDir)
if !ok {
return nil, fmt.Sprintf("desired state: %v — and no saved block (the daemon saves one when the hub sends it)", err), false
}
return saved, fmt.Sprintf("SAVED(%s; hub unreachable: %v)", at.UTC().Format(time.RFC3339), err), true
}
// runSelftestFacts prints the versions the host report carries (R-852, agent v0.142.0) — read-only.
//
// sudo -u felhom-agent felhom-agent --config … --selftest=os-facts -vmid 9201
func runSelftestFacts(ctx context.Context, cfg config.Config, logger *slog.Logger, vmid int) int {
if vmid <= 0 {
fmt.Fprintln(os.Stderr, "selftest=os-facts: -vmid is required")
return 2
}
px, _ := newProxmoxClient(cfg)
leg := newOSLeg(cfg, nil, px, logger)
start := time.Now()
f, err := leg.Facts(ctx, vmid)
if err != nil {
fmt.Fprintln(os.Stderr, "selftest=os-facts:", err)
return 1
}
var v any
_ = json.Unmarshal(f, &v)
printJSON(fmt.Sprintf("facts (vmid %d, %s)", vmid, time.Since(start).Round(100*time.Millisecond)), v)
return 0
}
// runSelftestLiveRestore is the ONE-TIME live-restore step (`09` decision 87) as a debug action — the night leg does
// the same before a ring-0 Docker step. It prints the container ids before and after (they must not change).
//
// sudo -u felhom-agent felhom-agent --config … --selftest=live-restore -vmid 9202
func runSelftestLiveRestore(ctx context.Context, cfg config.Config, logger *slog.Logger, vmid int) int {
if vmid <= 0 {
fmt.Fprintln(os.Stderr, "selftest=live-restore: -vmid is required")
return 2
}
px, _ := newProxmoxClient(cfg)
leg := newOSLeg(cfg, nil, px, logger)
if err := leg.EnsureLiveRestore(ctx, time.Now().UTC().Format("20060102T150405Z"), vmid); err != nil {
fmt.Fprintln(os.Stderr, "selftest=live-restore:", err)
return 1
}
fmt.Println("live-restore: on (see the wrapper's LIVE-RESTORE line above for the container ids)")
return 0
}
// newTunnelProber reads the box's REAL tunnel (R-841, agent v0.141.0): the cloudflared container in each running
// customer guest — a guest that binds /mnt/felhom-drives, the same rule the OS wrapper's R10 uses — through the
// existing `pct exec [0-9]* -- docker inspect -f *` sudoers line.
@@ -3591,31 +3741,80 @@ func newTunnelProber(cfg config.Config, px *proxmox.Client) hub.CloudflaredProbe
}
return hub.GuestTunnelProber{
Runner: &proxmox.ExecRunner{Mode: mode, SudoPath: cfg.Privileged.SudoPath},
Guests: func(ctx context.Context) ([]int, error) {
if px == nil {
return nil, fmt.Errorf("no proxmox client")
}
gs, err := px.ListLXC(ctx)
if err != nil {
return nil, err
}
var out []int
for _, g := range gs {
if g.Status != "running" {
continue
}
gc, err := px.GuestConfig(ctx, g.VMID)
if err != nil {
continue
}
for _, v := range gc.MountPoints() {
if src, _, _ := strings.Cut(v, ","); src == "/mnt/felhom-drives" {
out = append(out, g.VMID)
break
}
}
}
return out, nil
},
Guests: customerGuests(px),
}
}
// customerGuests lists the running guests that bind /mnt/felhom-drives — the box's customer guest(s), the same rule the
// OS wrapper's R10 uses. Shared by the tunnel probe, the facts read and the signed Docker step.
func customerGuests(px *proxmox.Client) func(ctx context.Context) ([]int, error) {
return func(ctx context.Context) ([]int, error) {
if px == nil {
return nil, fmt.Errorf("no proxmox client")
}
gs, err := px.ListLXC(ctx)
if err != nil {
return nil, err
}
var out []int
for _, g := range gs {
if g.Status != "running" {
continue
}
gc, err := px.GuestConfig(ctx, g.VMID)
if err != nil {
continue
}
for _, v := range gc.MountPoints() {
if src, _, _ := strings.Cut(v, ","); src == "/mnt/felhom-drives" {
out = append(out, g.VMID)
break
}
}
}
return out, nil
}
}
// firstGuest is the single customer guest (an error when there is none).
func firstGuest(px *proxmox.Client) func(ctx context.Context) (int, error) {
f := customerGuests(px)
return func(ctx context.Context) (int, error) {
v, err := f(ctx)
if err != nil {
return 0, err
}
if len(v) == 0 {
return 0, fmt.Errorf("no running customer guest")
}
return v[0], nil
}
}
// factsReporter feeds the host report's `system` stanza (R-852, agent v0.142.0) from the wrapper's read-only facts
// mode, at most every 10 minutes (each read is ~2 s of pct exec; the host reports every 15 min).
type factsReporter struct {
leg *osupdate.Leg
guest func(ctx context.Context) (int, error)
mu sync.Mutex
at time.Time
vmid int
facts json.RawMessage
err error
}
func (f *factsReporter) SystemFacts(ctx context.Context) (int, json.RawMessage, error) {
f.mu.Lock()
defer f.mu.Unlock()
if !f.at.IsZero() && time.Since(f.at) < 10*time.Minute {
return f.vmid, f.facts, f.err
}
f.at = time.Now()
f.vmid, f.err = f.guest(ctx)
if f.err != nil {
f.facts = nil
return 0, nil, f.err
}
f.facts, f.err = f.leg.Facts(ctx, f.vmid)
return f.vmid, f.facts, f.err
}
@@ -0,0 +1,31 @@
package main
import (
"context"
"io"
"log/slog"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/config"
)
// R-861 (agent v0.146.0): the root escrow ceremony builds a file path from the storage id; an id that is a path is
// refused before anything is read. RED-PROOF: drop the pbsStorageIDRe check → the "../" ids reach the key stat and
// come back as a "setup" error instead of "usage".
func TestEscrowCeremony_StorageIDIsNeverAPath(t *testing.T) {
lg := slog.New(slog.NewTextHandler(io.Discard, nil))
for _, id := range []string{"../../../etc/shadow", "a/b", "/etc/pve/priv/x", ".hidden", ""} {
cfg := config.Default()
cfg.Escrow.PBSStorageID = "" // the flag decides here
_, e := escrowCeremony(context.Background(), cfg, lg, escrowCeremonyOpts{storage: id})
if e == nil || e.kind != "usage" {
t.Errorf("storage id %q was not refused as usage (got %+v)", id, e)
}
}
cfg := config.Default()
cfg.Backup.PBSSecretDir = t.TempDir()
_, e := escrowCeremony(context.Background(), cfg, lg, escrowCeremonyOpts{storage: "felhom-pbs"})
if e == nil || e.kind != "setup" {
t.Fatalf("control: a plain id must pass the check and fail later on the missing key (setup), got %+v", e)
}
}
@@ -0,0 +1,55 @@
package main
import (
"context"
"errors"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
)
type r866Fetcher struct{ err error }
func (f r866Fetcher) FetchDesiredState(context.Context) (*hub.DesiredStateResponse, error) {
if f.err != nil {
return nil, f.err
}
r := &hub.DesiredStateResponse{}
r.DesiredState.OSUpdate = &hub.WireOSUpdate{Ring: 1, Enabled: true}
return r, nil
}
// R-866 (v0.144.0). THE NIGHT'S SHAPE (A3, Tester 1 box, hub blocked): `selftest=os-update: desired state: hub:
// transport error … connect: invalid argument` — the debug pass could not run at all. Now it runs from the block the
// daemon saved, and its header says so.
// COMPANION RED-PROOF: return at once on a fetch error in selftestOSBlock → "the pass did not run from the saved block".
func TestR866_DebugPassUsesTheSavedBlockWhenTheHubIsAway(t *testing.T) {
dir := t.TempDir()
leg := &osupdate.Leg{PlanDir: dir}
r := &hub.DesiredStateResponse{}
r.DesiredState.OSUpdate = &hub.WireOSUpdate{Ring: 0, Enabled: true}
leg.OnDesiredState(context.Background(), r) // the daemon received a block and saved it
away := r866Fetcher{err: errors.New("hub: transport error: connect: invalid argument")}
b, src, ok := selftestOSBlock(context.Background(), away, dir)
if !ok || b == nil || b.Ring != 0 {
t.Fatalf("the pass did not run from the saved block: ok=%v block=%+v src=%q", ok, b, src)
}
if !strings.HasPrefix(src, "SAVED(") || !strings.Contains(src, "hub unreachable") {
t.Fatalf("the header must say the block is the saved one: %q", src)
}
// the hub reachable: its block wins, and the header says "hub"
b, src, ok = selftestOSBlock(context.Background(), r866Fetcher{}, dir)
if !ok || b.Ring != 1 || src != "hub" {
t.Fatalf("hub block not used: %+v %q", b, src)
}
}
// No hub and nothing saved: the pass does not run on a guessed block.
func TestR866_NoHubNoSavedBlockDoesNotRun(t *testing.T) {
_, why, ok := selftestOSBlock(context.Background(), r866Fetcher{err: errors.New("down")}, t.TempDir())
if ok || !strings.Contains(why, "no saved block") {
t.Fatalf("ok=%v why=%q", ok, why)
}
}
+74
View File
@@ -0,0 +1,74 @@
package main
import (
"go/ast"
"testing"
)
// R-894 — the on-disk backup record is WIRED on the daemon path (the built-but-never-wired class).
// main → runDaemon → buildLocalAPIServer, and inside it the localapi.Options literal carries
// LastKnownBackups built by backup.NewBackupSuccessState. An AST walk, not a string match, for the
// reasons in escrow_recover_wiring_test.go.
//
// COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this
// fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored.
func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
_, f := parseMain(t)
if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one")
}
var field, built bool
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
cl, ok := n.(*ast.CompositeLit)
if !ok {
return true
}
sel, ok := cl.Type.(*ast.SelectorExpr)
if !ok {
return true
}
if pkg, _ := sel.X.(*ast.Ident); pkg == nil || pkg.Name+"."+sel.Sel.Name != "localapi.Options" {
return true
}
for _, el := range cl.Elts {
kv, ok := el.(*ast.KeyValueExpr)
if !ok {
continue
}
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "LastKnownBackups" {
field = true
if callsIn(kv.Value)["backup.NewBackupSuccessState"] {
built = true
}
}
}
return true
})
}
if !field {
t.Fatal("localapi.Options in buildLocalAPIServer has no LastKnownBackups field")
}
if !built {
t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState")
}
}
func callsIn(n ast.Node) map[string]bool {
out := map[string]bool{}
ast.Inspect(n, func(n ast.Node) bool {
if ce, ok := n.(*ast.CallExpr); ok {
if fn, ok := ce.Fun.(*ast.SelectorExpr); ok {
if x, ok := fn.X.(*ast.Ident); ok {
out[x.Name+"."+fn.Sel.Name] = true
}
}
}
return true
})
return out
}
+13 -1
View File
@@ -43,7 +43,7 @@ func main() {
func run() error {
var (
op = flag.String("op", "", "op class to sign, e.g. storage_wipe | guest_destroy | decommission | agent_update")
op = flag.String("op", "", "op class to sign, e.g. storage_wipe | guest_destroy | decommission | agent_update | os_docker_step | os_pve_step | os_kernel_step | agent_config_update")
host = flag.String("host", "", "target host_id (anti-retarget — the op runs ONLY on this host)")
guest = flag.String("guest", "", "target guest_id (\"\" = host-scoped op)")
keyID = flag.String("key-id", "", "key id of the signing key (must match a pinned agent signer)")
@@ -52,6 +52,7 @@ func run() error {
fstype = flag.String("fstype", "ext4", "for storage_wipe: the filesystem to mkfs after wipe")
agentVer = flag.String("agent-version", "", "for agent_update: the target agent version (e.g. 0.70.1)")
sha256Hex = flag.String("sha256", "", "for agent_update: the pinned lowercase-hex sha256 of the target binary")
bundleSHA = flag.String("bundle-sha256", "", "for agent_config_update: the pinned sha256 of felhom-config-bundle.json (R-840)")
keyFile = flag.String("key", "", "operator signing key (ssh private key / sk- key handle) for ssh-keygen -Y sign")
ttl = flag.Duration("ttl", 30*time.Minute, "validity window from now (issued_at..expires_at)")
nonce = flag.String("nonce", "", "explicit nonce (default: a fresh 128-bit random nonce)")
@@ -95,6 +96,17 @@ func run() error {
}
pj, _ := json.Marshal(map[string]string{"version": *agentVer, "sha256": *sha256Hex})
params = string(pj)
case "agent_config_update":
// R-840: the box's root-owned files. The ROOT wrapper verifies this signature itself and refuses a bundle
// whose sha256 is not exactly this one.
if *agentVer == "" || *bundleSHA == "" {
return fmt.Errorf("agent_config_update needs -agent-version and -bundle-sha256 (the pinned bundle hash)")
}
if !isHex64(*bundleSHA) {
return fmt.Errorf("agent_config_update -bundle-sha256 must be 64 lowercase hex chars (got %d)", len(*bundleSHA))
}
pj, _ := json.Marshal(map[string]string{"agent_version": *agentVer, "bundle_sha256": *bundleSHA})
params = string(pj)
default:
params = "{}"
}
+63 -3
View File
@@ -61,7 +61,7 @@ set -euo pipefail
# Script provenance — logged into every bake transcript next to the baked controller tag, so an
# archive can always be traced to the script that produced it. Bump on any behavior change.
GOLDEN_SCRIPT_VERSION="3.0.0"
GOLDEN_SCRIPT_VERSION="3.2.0"
VMID="${1:-9100}"
TEMPLATE="${2:-local:vztmpl/debian-13-standard_13.1-2_amd64.tar.zst}"
@@ -112,7 +112,15 @@ for i in $(seq 1 30); do
if pct exec "$VMID" -- getent hosts download.docker.com >/dev/null 2>&1; then break; fi
sleep 1
done
pct exec "$VMID" -- bash -c '
# v3.1.0 (`11` §5.8): GOLDEN_DOCKER_PKGS pins the APPROVED Docker engine set (the hub's newest Docker release, all six
# "name=version"); without it the newest stable set is installed and the bake log says so.
GOLDEN_DOCKER_PKGS="${GOLDEN_DOCKER_PKGS:-}"
if [[ -n "$GOLDEN_DOCKER_PKGS" ]]; then
echo "[golden] Docker engine set PINNED to the approved release: $GOLDEN_DOCKER_PKGS"
else
echo "[golden] WARNING: GOLDEN_DOCKER_PKGS not set — installing the newest stable Docker set, not an approved one"
fi
pct exec "$VMID" -- env GOLDEN_DOCKER_PKGS="$GOLDEN_DOCKER_PKGS" bash -c '
set -e
export DEBIAN_FRONTEND=noninteractive
apt-get update -qq
@@ -122,7 +130,54 @@ pct exec "$VMID" -- bash -c '
echo "deb [signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/debian trixie stable" \
> /etc/apt/sources.list.d/docker.list
apt-get update -qq
apt-get install -y -qq docker-ce docker-ce-cli containerd.io >/dev/null
if [ -n "$GOLDEN_DOCKER_PKGS" ]; then
apt-get install -y -qq $GOLDEN_DOCKER_PKGS >/dev/null
else
apt-get install -y -qq docker-ce docker-ce-cli containerd.io >/dev/null
fi
dpkg-query -W containerd.io docker-buildx-plugin docker-ce docker-ce-cli docker-ce-rootless-extras docker-compose-plugin 2>/dev/null | sed "s/^/ installed: /"
'
# v3.2.0 (Part F of the R-840 brief, `11` §5.3): GOLDEN_GUEST_PKGS = the newest APPROVED guest release, as
# "name=version …" (the hub's os_releases row for layer guest, IN FORCE — never a cancelled test approval). The bake
# brings every package the template HAS to exactly that version, under felhom-os-apply's rules: never a package the
# template lacks (--only-upgrade), never newer than approved, never a removal or a new package (a simulation is checked
# first and the bake FAILS on either). Empty = no approved guest release in force: the template's versions stay, and
# the box's first night installs whatever release is approved then. Either way the bake PRINTS the first-night count:
# how many installed packages are older than the approved version (target 0).
GOLDEN_GUEST_PKGS="${GOLDEN_GUEST_PKGS:-}"
pct exec "$VMID" -- env GOLDEN_GUEST_PKGS="$GOLDEN_GUEST_PKGS" bash -c '
set -e
export DEBIAN_FRONTEND=noninteractive
if [ -z "$GOLDEN_GUEST_PKGS" ]; then
echo "[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a"
echo "[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): $(apt list --upgradable 2>/dev/null | grep -c /)"
exit 0
fi
want=""
for nv in $GOLDEN_GUEST_PKGS; do
n=${nv%%=*}; v=${nv#*=}
cur=$(dpkg-query -W -f="\${Version}" "$n" 2>/dev/null) || continue # not in the template: never added
dpkg --compare-versions "$cur" lt "$v" && want="$want $n=$v"
done
if [ -n "$want" ]; then
sim=$(apt-get -s install --only-upgrade -o Dpkg::Options::=--force-confold $want)
if echo "$sim" | grep -q "^Remv "; then echo "[golden] FATAL: the approved guest set would REMOVE a package"; echo "$sim" | grep "^Remv "; exit 1; fi
for p in $(echo "$sim" | awk "/^Inst /{print \$2}"); do
dpkg-query -W "$p" >/dev/null 2>&1 || { echo "[golden] FATAL: the approved guest set would ADD $p - not in the template"; exit 1; }
done
apt-get install -y -qq --only-upgrade -o Dpkg::Options::=--force-confold -o Dpkg::Options::=--force-confdef $want >/dev/null
echo "[golden] approved guest release installed: $(echo $want | wc -w) package(s) brought to the approved version"
else
echo "[golden] approved guest release: the template already runs every approved version"
fi
left=0
for nv in $GOLDEN_GUEST_PKGS; do
n=${nv%%=*}; v=${nv#*=}
cur=$(dpkg-query -W -f="\${Version}" "$n" 2>/dev/null) || continue
dpkg --compare-versions "$cur" lt "$v" && { left=$((left+1)); echo " still older: $n $cur < $v"; }
done
echo "[golden] first-night count vs the approved guest release: $left (target 0)"
[ "$left" -eq 0 ] || { echo "[golden] FATAL: $left package(s) stayed older than the approved release"; exit 1; }
'
echo "[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …"
# containerd-snapshotter (Docker 28+/29 default) keeps the IMAGE content store under
@@ -135,9 +190,12 @@ echo "[golden] baking daemon.json: classic overlay2 driver (containerd-snapshott
# layout (a container's `df /` reports the single volume, phase-0 spike). Since v3.0.0 /var/lib/docker
# is a BIND of <volume>/docker rather than the mp0 mount itself, wired immediately below; data-root
# still needs no override because the path is unchanged. Log caps kill the most common runaway.
# v3.1.0: "live-restore": true (`09` decision 87) — a Docker engine update then restarts no app (`11` C5). A box made
# from this golden never needs the agent's one-time live-restore-on step. NEVER removed by a plain restart (R-835).
pct exec "$VMID" -- bash -c 'mkdir -p /etc/docker; cat > /etc/docker/daemon.json <<JSON
{
"features": { "containerd-snapshotter": false },
"live-restore": true,
"log-driver": "json-file",
"log-opts": { "max-size": "10m", "max-file": "3" }
}
@@ -175,6 +233,8 @@ pct exec "$VMID" -- bash -c 'systemctl restart docker; sleep 3; docker run --rm
# Guard: the image store MUST be on the data volume now. /var/lib/containerd holding the images would
# mean containerd-snapshotter is still on (the split would leave images on the rootfs).
pct exec "$VMID" -- bash -c 'drv=$(docker info 2>/dev/null | sed -n "s/.*Storage Driver: //p"); [ "$drv" = "overlay2" ] || { echo "[golden] FATAL: storage driver is $drv, expected overlay2 — images would not land on the data volume"; exit 1; }'
# v3.1.0 ASSERTION: live-restore is ON in the running daemon (decision 87), or the bake fails closed.
pct exec "$VMID" -- bash -c 'lr=$(docker info --format "{{.LiveRestoreEnabled}}" 2>/dev/null); [ "$lr" = "true" ] && echo " live-restore: on" || { echo "[golden] FATAL: live-restore is $lr, expected true (decision 87)"; exit 1; }'
# ASSERTION 1 (RETARGETED v3.0.0, not removed). /var/lib/docker must be a real mount — now the V-c
# bind of <volume>/docker rather than the mp0 mount itself. Still fails closed on the same failure:
# if the bind did not take, Docker's data-root silently sits on the OS rootfs and the golden ships
+8
View File
@@ -0,0 +1,8 @@
# /etc/felhom/crash-guard.conf — read by /usr/local/sbin/felhom-crash-guard (`11` §5.9).
# The LIMIT-th unclean stop within WINDOW_MINUTES leaves the box off. Decided by CC unattended — operator may reverse.
LIMIT=3
WINDOW_MINUTES=60
# kernel.panic while armed: seconds after a crash before the kernel restarts the box.
PANIC_SECONDS=10
# A tripped guard re-arms after this many hours of normal running (or `felhom-crash-guard rearm`).
REARM_HOURS=24
+108 -106
View File
@@ -1,30 +1,38 @@
# felhom-agent sudoers allowlist — the NARROW host-root surface (slice 5 Phase B, doc 03 §3/§7).
# felhom-agent sudoers allowlist — the NARROW host-root surface (slice 5 Phase B, doc 03 §3/§7; narrowed R-861).
#
# Install as a drop-in: /etc/sudoers.d/felhom-agent (mode 0440, root:root), validated with
# `visudo -cf`. The agent runs as the non-root `felhom-agent` service user and shells out via
# `sudo -n` with FIXED argument vectors (no shell). The fine-grained validation is done IN
# the agent BEFORE exec (internal/storage/validate.go): UUIDs against a strict hex regex,
# mount paths confined+traversal-checked, SMART devices whitelisted to raw disks, LVM names
# charset-checked. These sudoers wildcards are the COARSE allowlist; the agent is the fine
# gate, so a wildcard can never be abused by a value the agent didn't already validate.
# Install as a drop-in: /etc/sudoers.d/felhom-agent (mode 0440, root:root), validated with `visudo -cf`. It rides the
# signed config bundle (R-840). The agent runs as the non-root `felhom-agent` user and shells out via `sudo -n` with
# FIXED argument vectors (no shell).
#
# Binary paths MUST match the agent config (privileged.systemctl/install/smartctl/lvs). Adjust
# for your distro (Debian/PVE shown). A missing/declined entry degrades the agent with a
# warning (SMART→UNKNOWN, mount→logged error), it does not crash.
# R-861 (agent v0.146.0) — EXACT PATTERNS, NOT GLOBS. A sudoers `*` in the ARGUMENTS also matches spaces, so
# `pct set [0-9]* -onboot 1` matched `pct set 100 --dev0 /dev/sda -onboot 1` (a raw host disk for the guest), and
# `mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*` matched a `..` path onto /etc. Every argument list that varies
# is now a sudo regular expression (`^...$`, sudo >= 1.9.10; Debian 13 / PVE 9 ship 1.9.16): one value per slot, a
# fixed character set, no `..`, no extra argument. Lines with no variable part stay literal. The patterns are pinned
# by the capability manifest (every real call must match: TestManifestCoveredBySudoers) and by injection cases that
# must NOT match (configs/test_sudoers_patterns.py, and live with `sudo -l -U felhom-agent` on the demo boxes).
#
# R-861 — NO FILE THE AGENT WROTE IS INSTALLED WHERE ROOT READS IT. The `install` lines are gone: a mount unit, a
# dnsmasq drop-in, the WireGuard config and the OOB sshd config + key go through `felhom-priv-apply`, a root wrapper
# from the bundle that checks the CONTENT against the agent's own renderers; the guest pre-start hook and the shared
# drive parent are FIXED files that come with the bundle itself; the agent binary is replaced only by an
# operator-signed agent_update that `felhom-os-apply` verifies as root (`felhom-selfupdate-guarded apply` is no longer
# here). The two remaining root runs of agent code (FELHOM_ESCROW, the guest hook) therefore run only a signed binary.
#
# Binary paths MUST match the agent config (privileged.systemctl/install/smartctl/lvs). A missing/declined entry
# degrades the agent with a warning (SMART→UNKNOWN, mount→logged error), it does not crash; the capability probe
# reports it to the hub.
Cmnd_Alias FELHOM_MOUNT = \
/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/* /etc/systemd/system/*.mount, \
/usr/local/sbin/felhom-priv-apply ^unit mnt-[A-Za-z0-9_.\\-]+\.(mount|automount)$, \
/usr/bin/systemctl daemon-reload, \
/usr/bin/systemctl enable --now -- *.mount, \
/usr/bin/systemctl disable -- *.mount, \
/usr/bin/systemctl stop -- *.mount
/usr/bin/systemctl ^enable --now -- mnt-[A-Za-z0-9_.\\-]+\.mount$, \
/usr/bin/systemctl ^disable -- mnt-[A-Za-z0-9_.\\-]+\.mount$, \
/usr/bin/systemctl ^stop -- mnt-[A-Za-z0-9_.\\-]+\.mount$
Cmnd_Alias FELHOM_DISK = \
/usr/sbin/smartctl -a -j /dev/sd[a-z]*, \
/usr/sbin/smartctl -a -j /dev/nvme[0-9]*n[0-9]*, \
/usr/sbin/smartctl -a -j /dev/vd[a-z]*, \
/usr/sbin/smartctl -a -j /dev/hd[a-z]*, \
/usr/sbin/lvs --reportformat json --units b -o lv_name\,data_percent\,metadata_percent -- *, \
/usr/sbin/smartctl ^-a -j /dev/(sd[a-z]+|nvme[0-9]+n[0-9]+|vd[a-z]+|hd[a-z]+)$, \
/usr/sbin/lvs ^--reportformat json --units b -o lv_name\,data_percent\,metadata_percent -- [A-Za-z0-9_.+-]+(/[A-Za-z0-9_.+-]+)?$, \
/usr/sbin/pvs --reportformat json --noheadings -o pv_name, \
/usr/sbin/zpool status -P
@@ -35,9 +43,9 @@ Cmnd_Alias FELHOM_DISK = \
# (the wildcard only ever names a path the agent itself created), and the bootstrap file the agent
# writes there is the only thing these touch. ':' is escaped per sudoers grammar.
Cmnd_Alias FELHOM_PROVISION = \
/usr/bin/chown -R 100000\:100000 /var/lib/felhom-agent/guests/*, \
/usr/sbin/pct set [0-9]* -mp[0-9]* /var/lib/felhom-agent/guests/*, \
/usr/sbin/pct set [0-9]* -onboot 1
/usr/bin/chown ^-R 100000\:100000 /var/lib/felhom-agent/guests/[0-9]+(/bootstrap)?$, \
/usr/sbin/pct ^set [0-9]+ -mp[0-9]+ /var/lib/felhom-agent/guests/[0-9]+/bootstrap\,mp\=/[A-Za-z0-9/_.-]+(\,ro\=1)?$, \
/usr/sbin/pct ^set [0-9]+ -onboot 1$
# Disk inspection + format (slice 8C + Impl-1). blkid/lsblk read the device's data-bearing evidence
# (the agent decides data-bearing-ness from THIS, never the caller's claim). Format goes ONLY through
@@ -45,66 +53,54 @@ Cmnd_Alias FELHOM_PROVISION = \
# mkfs the OS disk — the wrapper re-checks the catastrophic cases (system disk / LVM PV / foreign mount)
# as root and refuses, and the agent's unclaimed-disk filter (claim.go) is the primary guard above it.
Cmnd_Alias FELHOM_FORMAT = \
/usr/sbin/blkid -p -o export /dev/*, \
/usr/bin/lsblk -J -o NAME\,FSTYPE\,PTTYPE\,MOUNTPOINT /dev/*, \
/usr/local/sbin/felhom-mkfs-guarded /dev/* *
/usr/sbin/blkid ^-p -o export /dev/[^ ]+$, \
/usr/bin/lsblk ^-J -o NAME\,FSTYPE\,PTTYPE\,MOUNTPOINT /dev/[^ ]+$, \
/usr/local/sbin/felhom-mkfs-guarded ^/dev/[^ ]+ (ext4|xfs)$
# LAN split-horizon resolver (internal/lanresolver): the agent manages a host-side dnsmasq that
# answers *.<customer-domain> with each guest's live LAN IP. install only ever writes felhom-*.conf
# drop-ins (from agent-written /tmp temp files); the two `pct exec` reads are FIXED command vectors
# answers *.<customer-domain> with each guest's live LAN IP. A felhom-*.conf drop-in reaches /etc/dnsmasq.d
# only through felhom-priv-apply, which allows exactly the lines the resolver renders (R-861: a `dhcp-script=` would
# run as root); the two `pct exec` reads are FIXED command vectors
# (the guest's eth0 IPv4 + the controller's pulled controller.yaml for the domain) — NOT a general
# `pct exec`. systemctl is scoped to the dnsmasq unit only. The agent never edits /etc/resolv.conf.
Cmnd_Alias FELHOM_DNSMASQ = \
/usr/bin/apt-get install -y -q dnsmasq, \
/usr/bin/install -m 0644 /tmp/felhom-resolver-*.conf /etc/dnsmasq.d/felhom-*.conf, \
/usr/local/sbin/felhom-priv-apply ^dnsmasq /tmp/felhom-resolver-[0-9]+\.conf felhom-[a-z0-9][a-z0-9._-]*\.conf$, \
/usr/bin/systemctl enable --now dnsmasq, \
/usr/bin/systemctl reload dnsmasq, \
/usr/bin/systemctl restart dnsmasq, \
/usr/bin/rm -f /etc/dnsmasq.d/felhom-*.conf, \
/usr/sbin/pct exec [0-9]* -- ip -4 -o addr show dev eth0, \
/usr/sbin/pct exec [0-9]* -- docker exec felhom-controller cat /opt/docker/felhom-controller/controller.yaml
/usr/bin/rm ^-f /etc/dnsmasq\.d/felhom-[a-z0-9][a-z0-9._-]*\.conf$, \
/usr/sbin/pct ^exec [0-9]+ -- ip -4 -o addr show dev eth0$, \
/usr/sbin/pct ^exec [0-9]+ -- docker exec felhom-controller cat /opt/docker/felhom-controller/controller\.yaml$
# Guest mountpoint lifecycle (intermediary-mount re-architecture + C1 net). The pre-start self-heal hook
# wrapper is installed once into the PVE snippets dir (from an agent-written /tmp file) and registered
# per-guest; decommission/eject DELETE the dead mountpoint slot so a missing bind source can't brick the
# guest at next boot (the B3 C1 fix). The agent fine-validates the vmid (numeric) + slot (mp[0-9]+) and
# the snippet path is fixed — the wildcards are the coarse allowlist. The install SOURCE is a
# random-named agent temp (os.CreateTemp, audit B1 — a fixed /tmp name was a local TOCTOU), hence the
# glob; the DESTINATION stays pinned. The `mkdir -p` creates the snippets dir on a FRESH box —
# `install` won't create parents, so without it the hook install failed silently on Day-0 boxes
# (B2, DRILL-day0-cleanroom-2026-07-03; fixed agent v0.63.0).
# Guest mountpoint lifecycle (intermediary-mount re-architecture + C1 net). The pre-start self-heal hook is a FIXED file
# from the config bundle (/var/lib/vz/snippets/felhom-guest-hook.sh — R-861: the agent no longer installs it from /tmp;
# Proxmox runs it as root at every guest start). The agent only registers it per guest, deletes a dead mountpoint slot
# (the B3 C1 fix) and reboots a guest to activate binds — each with an exact vmid / slot.
Cmnd_Alias FELHOM_GUESTHOOK = \
/usr/bin/mkdir -p /var/lib/vz/snippets, \
/usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-*.sh /var/lib/vz/snippets/felhom-guest-hook.sh, \
/usr/sbin/pct set [0-9]* --hookscript local\:snippets/felhom-guest-hook.sh, \
/usr/sbin/pct set [0-9]* --delete mp[0-9]*, \
/usr/sbin/pct reboot [0-9]*
/usr/sbin/pct ^set [0-9]+ --hookscript local\:snippets/felhom-guest-hook\.sh$, \
/usr/sbin/pct ^set [0-9]+ --delete mp[0-9]+$, \
/usr/sbin/pct ^reboot [0-9]+$
# Intermediary mount model (the drive hot-swap re-architecture). The agent keeps a SHARED host parent
# /mnt/felhom-drives (self-bind + make-shared + a boot-persistence systemd unit) and binds/unbinds each
# drive's felhom-data namespace UNDERNEATH it so the change propagates into the running guest live (no
# pct, no reboot). The agent fine-validates the drive name + confines paths before any exec; the trailing
# `*` (matching the comma-laden mp spec) mirrors the existing FELHOM_PROVISION pattern.
# `lxc-info -n <vmid> -p -H` resolves the guest init PID for the GuestSeesMount / bound_under_parent check
# (a READ — the drive-gate's "is the drive live in the guest?" signal); WITHOUT it the non-root agent gets
# an empty PID and reports every drive absent (multi-drive flapping, audit 2026-06-29). `make-private`
# isolates the parent's peer group on FIRST setup only (EnsureSharedParent guards on mountpoint, so it
# never re-churns a live parent); without it the parent stays in root's group and submounts double.
# /mnt/felhom-drives (self-bind + make-shared) and binds/unbinds each drive's felhom-data namespace UNDERNEATH it so the
# change propagates into the running guest live. The boot-persistence script + unit are FIXED files from the config
# bundle (R-861: the agent no longer installs them from /tmp); the agent only enables the unit. A drive name is one
# path segment that cannot start with a dot (no `..`); `lxc-info -n <vmid> -p -H` resolves the guest init PID for the
# GuestSeesMount check; `make-private` isolates the parent's peer group on FIRST setup only.
Cmnd_Alias FELHOM_INTERMEDIARY = \
/usr/bin/mkdir -p /mnt/felhom-drives, \
/usr/bin/mkdir -p /mnt/felhom-drives/*, \
/usr/bin/mkdir -p /mnt/*/felhom-data, \
/usr/bin/chown 100000\:100000 /mnt/*/felhom-data, \
/usr/bin/mkdir ^-p /mnt/felhom-drives/[A-Za-z0-9_-][A-Za-z0-9_.-]*$, \
/usr/bin/mkdir ^-p /mnt/[A-Za-z0-9_-][A-Za-z0-9_.-]*/felhom-data$, \
/usr/bin/chown ^100000\:100000 /mnt/[A-Za-z0-9_-][A-Za-z0-9_.-]*/felhom-data$, \
/usr/bin/mount --bind /mnt/felhom-drives /mnt/felhom-drives, \
/usr/bin/mount --make-shared /mnt/felhom-drives, \
/usr/bin/mount --make-private /mnt/felhom-drives, \
/usr/bin/mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*, \
/usr/bin/umount /mnt/felhom-drives/*, \
/usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-*.sh /usr/local/sbin/felhom-shared-parent.sh, \
/usr/bin/install -m 0644 -- /tmp/felhom-shared-parent-*.service /etc/systemd/system/felhom-shared-parent.service, \
/usr/bin/mount ^--bind /mnt/[A-Za-z0-9_-][A-Za-z0-9_.-]*/felhom-data /mnt/felhom-drives/[A-Za-z0-9_-][A-Za-z0-9_.-]*$, \
/usr/bin/umount ^/mnt/felhom-drives/[A-Za-z0-9_-][A-Za-z0-9_.-]*$, \
/usr/bin/systemctl enable felhom-shared-parent.service, \
/usr/bin/lxc-info -n [0-9]* -p -H, \
/usr/sbin/pct set [0-9]* -mp8 /mnt/felhom-drives*
/usr/bin/lxc-info ^-n [0-9]+ -p -H$, \
/usr/sbin/pct ^set [0-9]+ -mp8 /mnt/felhom-drives\,mp\=/mnt/felhom-drives$
# Controller-swap / managed auto-update (Option A, non-root). The agent owns the in-guest controller
# image SWAP (it survives the controller being killed mid-swap): read the baked image ref, check the
@@ -115,15 +111,17 @@ Cmnd_Alias FELHOM_INTERMEDIARY = \
# docker inspect -f * — container running/health/image (read-only; `*` spans the -f template
# + container across spaces, spike-confirmed)
# systemctl restart <fixed unit> — re-run the golden's bootstrap (the only state change)
# tee <FIXED image file> — WRITE the ref; content is fed on STDIN (no shell, no interpolation),
# the agent strict-validates the ref (controllerImageRe) before the write.
# felhom-priv-apply controller-image <vmid> — WRITE the ref (R-861 (a) A1, `09` §3 decision 165): the ref goes on
# STDIN to the ROOT wrapper, which requires our registry + repository + an x.y.z tag
# and writes the guest file itself. The agent's own `tee` grant is GONE: before, a
# compromised agent could hand the guest's bootstrap ANY image (sudo cannot see stdin).
# Validated GO: felhom.eu/documentation/audits/SPIKE-controllerswap-narrow-grants-2026-06-29.md.
Cmnd_Alias FELHOM_CONTROLLERSWAP = \
/usr/sbin/pct exec [0-9]* -- cat /etc/felhom-controller-image, \
/usr/sbin/pct exec [0-9]* -- docker image inspect *, \
/usr/sbin/pct exec [0-9]* -- docker inspect -f *, \
/usr/sbin/pct exec [0-9]* -- systemctl restart felhom-controller-bootstrap.service, \
/usr/sbin/pct exec [0-9]* -- tee /etc/felhom-controller-image
/usr/sbin/pct ^exec [0-9]+ -- cat /etc/felhom-controller-image$, \
/usr/sbin/pct ^exec [0-9]+ -- docker image inspect gitea\.dooplex\.hu/admin/felhom-controller\:[0-9]+\.[0-9]+\.[0-9]+$, \
/usr/sbin/pct ^exec [0-9]+ -- docker inspect -f .+ (felhom-controller|cloudflared)$, \
/usr/sbin/pct ^exec [0-9]+ -- systemctl restart felhom-controller-bootstrap\.service$, \
/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$
# Stale-lock recovery (F2-b, v0.49.0). A host reboot DURING a vzdump backup leaves the guest with a
# `snapshot-delete`/`backup` lock + `onboot:1` then can't start it → the customer box stays DOWN. The
@@ -131,7 +129,15 @@ Cmnd_Alias FELHOM_CONTROLLERSWAP = \
# with no API equivalent (snapshot-delete + start go through the API token); the agent fine-validates the
# vmid (numeric) before exec — the `[0-9]*` is the coarse allowlist.
Cmnd_Alias FELHOM_STALELOCK = \
/usr/sbin/pct unlock [0-9]*
/usr/sbin/pct ^unlock [0-9]+$
# Weekly guest disk trim (R-444, operator ruling `09` §3 decision 139). A thin pool only ever grows from blocks the
# guest has FREED: `fstrim` inside the unprivileged container is refused (FITRIM: Operation not permitted), so the host
# trims the guest's mounts. Measured on demo-hp 2026-10-06: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %,
# apps kept answering. ONE exact pattern: a vmid and nothing else — no `--ignore-mountpoints`, no second argument
# (pinned: TestSudoersFstrimRuleIsExact). The agent runs it on a weekly daytime timer under the heavy-op gate.
Cmnd_Alias FELHOM_FSTRIM = \
/usr/sbin/pct ^fstrim [0-9]+$
# Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS
# leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and
@@ -160,8 +166,9 @@ Cmnd_Alias FELHOM_SCRATCH_TEARDOWN = \
# into the guest through the existing shared bind (an unprivileged LXC cannot mount NFS/CIFS itself).
# A NAS is NOT a drive — no durable-id, no SMART, no wipe; these grants only install/enable/remove the
# unit pair. The agent fine-validates every value (share name, server, export, uid/gid, creds path) before
# any unit is rendered (internal/storage/netmount.go ValidateNetworkMountSpec); the trailing globs are the
# COARSE allowlist. The `.mount` install/enable/disable/stop reuse FELHOM_MOUNT; this alias adds the
# any unit is rendered (internal/storage/netmount.go ValidateNetworkMountSpec); the unit FILE reaches
# /etc/systemd/system only through `felhom-priv-apply unit` (FELHOM_MOUNT), which requires nosuid,nodev on a network
# share (R-861). The `.mount` enable/disable/stop reuse FELHOM_MOUNT; this alias adds the
# `.automount` variants + the unit-file removal. The unit FILE name is the systemd-escaped mountpoint,
# which always begins `mnt-felhom` (the mountpoint is /mnt/felhom-drives/<name>), so the rm glob is scoped
# to felhom mount units only. mkdir of the mountpoint reuses FELHOM_INTERMEDIARY's /mnt/felhom-drives/*.
@@ -176,51 +183,46 @@ Cmnd_Alias FELHOM_SCRATCH_TEARDOWN = \
# behind (the campaign accumulated 10 stub-shaped leftovers). rmdir ONLY (never rm -rf): it refuses
# a non-empty dir, so unexpected data is preserved, not destroyed — a fail-safe grant.
Cmnd_Alias FELHOM_NETMOUNT = \
/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/* /etc/systemd/system/*.automount, \
/usr/bin/systemctl enable --now -- *.automount, \
/usr/bin/systemctl disable -- *.automount, \
/usr/bin/systemctl stop -- *.automount, \
/usr/bin/systemctl reset-failed -- mnt-felhom*, \
/usr/bin/rmdir /mnt/felhom-drives/*, \
/usr/bin/rm -f /etc/systemd/system/mnt-felhom*
/usr/bin/systemctl ^enable --now -- mnt-[A-Za-z0-9_.\\-]+\.automount$, \
/usr/bin/systemctl ^disable -- mnt-[A-Za-z0-9_.\\-]+\.automount$, \
/usr/bin/systemctl ^stop -- mnt-[A-Za-z0-9_.\\-]+\.automount$, \
/usr/bin/systemctl ^reset-failed -- mnt-felhom[A-Za-z0-9_.\\-]*\.(mount|automount)$, \
/usr/bin/rmdir ^/mnt/felhom-drives/[A-Za-z0-9_-][A-Za-z0-9_.-]*$, \
/usr/bin/rm ^-f /etc/systemd/system/mnt-felhom[A-Za-z0-9_.\\-]*\.(mount|automount)$
# Offsite WG tunnel (S3, doc 06 §3.3). The agent manages wg-quick@wg-felhom as an agent-managed
# host service (the dnsmasq/lanresolver shape): conf staged in the agent-owned StateDir (never
# /tmp), installed 0600 to the FIXED destination, unit enable/restart/disable. The ONLY wg read
# /tmp), installed 0600 to the FIXED destination by `felhom-priv-apply wg`, which refuses any key renderConf never
# writes (R-861: PostUp/PreUp run as root under wg-quick), unit enable/restart/disable. The ONLY wg read
# is `latest-handshakes` — `wg show <if> dump` is FORBIDDEN everywhere (its interface line
# carries the PRIVATE KEY; the S1 session-log incident). Both install paths are FIXED (no glob):
# the agent has exactly one tunnel conf to manage.
# carries the PRIVATE KEY; the S1 session-log incident). Source and destination are fixed in the wrapper.
Cmnd_Alias FELHOM_WG = \
/usr/bin/apt-get install -y -q wireguard-tools, \
/usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf, \
/usr/local/sbin/felhom-priv-apply wg, \
/usr/bin/systemctl enable --now wg-quick@wg-felhom, \
/usr/bin/systemctl restart wg-quick@wg-felhom, \
/usr/bin/systemctl disable --now wg-quick@wg-felhom, \
/usr/bin/wg show wg-felhom latest-handshakes
# Agent self-update (TASK D1, SPIKE-agent-selfupdate-2026-07-05). The agent downloads the
# operator-SIGNED binary (sha256 pinned in the signed op — neither hub nor Gitea compromise can
# substitute it), verifies the sha in-process, then hands off to the guarded wrapper, which
# RE-verifies the sha as root, confines the staged path to /var/lib/felhom-agent/selfupdate/,
# performs the A/B flip (atomic same-fs rename, .prev retained) and schedules a detached restart.
# The apply args are a COARSE glob (spike S4b: sudoers fnmatch makes a [a-f0-9]* sha pattern
# first-char-only anyway) — the wrapper's own sha re-verify + path confinement is the real gate.
# `rollback` is normally run by felhom-agent-rollback.service (root, OnFailure=), not via sudo;
# granting it here keeps the verb probe-able (capability self-check) and operator-invokable.
# Agent self-update (TASK D1; R-861). The A/B flip (`felhom-selfupdate-guarded apply`) is NO LONGER the agent's: the
# agent hands the operator-SIGNED agent_update to felhom-os-apply (FELHOM_OSAPPLY, mode agent_update), which verifies
# the signature as root and only then runs the flip. Until v0.146.0 the agent passed the sha itself, so a compromised
# agent could install any binary — the binary FELHOM_ESCROW and the guest hook run as root. `commit` (clear the pending
# marker) and `rollback` (pending-guarded revert, normally run by felhom-agent-rollback.service) stay.
Cmnd_Alias FELHOM_SELFUPDATE = \
/usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/* *, \
/usr/local/sbin/felhom-selfupdate-guarded commit, \
/usr/local/sbin/felhom-selfupdate-guarded rollback
# Dedicated OOB sshd (TASK H1). The agent manages felhom-sshd like wg-felhom/dnsmasq: it RENDERS the
# config (Port from its claim) + the operator's authorized_keys, validates with `sshd -t`, and reloads
# (never restart-on-change [SF-2]). Both install SOURCES are the agent-owned staged files under
# StateDir; both DESTINATIONS are FIXED. `sshd -t/-T` are the validate/discover reads. The
# (never restart-on-change [SF-2]). Both files reach /etc/felhom-sshd only through felhom-priv-apply (R-861): the config
# must be the ONE template with only the Port varying (an AuthorizedKeysFile the agent owns + `StrictModes no` would be
# a root login), the key file one plain public key without options. `sshd -t/-T` are the validate/discover reads. The
# systemctl verbs are SCOPED to felhom-sshd only. reset-failed precedes a deliberate restart [SF-5].
# NOTHING here can touch the stock sshd, :22, or /etc/ssh.
Cmnd_Alias FELHOM_SSHD = \
/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config, \
/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/authorized_keys.felhom-op /etc/felhom-sshd/authorized_keys/felhom-op, \
/usr/local/sbin/felhom-priv-apply sshd-config, \
/usr/local/sbin/felhom-priv-apply sshd-key, \
/usr/sbin/sshd -t -f /var/lib/felhom-agent/felhom-sshd/sshd_config, \
/usr/sbin/sshd -t -f /etc/felhom-sshd/sshd_config, \
/usr/sbin/sshd -T -f /etc/felhom-sshd/sshd_config, \
@@ -270,8 +272,8 @@ Cmnd_Alias FELHOM_OOB = \
/usr/sbin/nft list set inet felhom_oob ssh_port, \
/usr/sbin/nft flush set inet felhom_oob operator_ips, \
/usr/sbin/nft flush set inet felhom_oob ssh_port, \
/usr/sbin/nft add element inet felhom_oob operator_ips *, \
/usr/sbin/nft add element inet felhom_oob ssh_port *
/usr/sbin/nft ^add element inet felhom_oob operator_ips \{ [0-9.]+(/[0-9]+)? \}$, \
/usr/sbin/nft ^add element inet felhom_oob ssh_port \{ [0-9]+ \}$
# Escrow ceremony (controller-driven, TASK 2026-07-13; mechanics validated by
# SPIKE-controller-escrow-2026-07-13). ONE fixed argv — sudoers matches the argument vector
@@ -307,9 +309,9 @@ Cmnd_Alias FELHOM_OSAPPLY = \
/usr/local/sbin/felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json
Cmnd_Alias FELHOM_GUESTNET = \
/usr/sbin/pct exec [0-9]* -- ip route show default, \
/usr/sbin/pct exec [0-9]* -- cat /etc/network/interfaces, \
/usr/sbin/pct exec [0-9]* -- pgrep -x dhclient, \
/usr/sbin/pct exec [0-9]* -- dhclient -pf /run/dhclient.eth0.pid -lf /var/lib/dhcp/dhclient.eth0.leases eth0
/usr/sbin/pct ^exec [0-9]+ -- ip route show default$, \
/usr/sbin/pct ^exec [0-9]+ -- cat /etc/network/interfaces$, \
/usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \
/usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_FSTRIM, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
+235
View File
@@ -0,0 +1,235 @@
#!/usr/bin/python3
# felhom-crash-guard — a crashed host restarts by itself, but not forever (`09` decision 88, R-851, `11` §5.9).
#
# Install as /usr/local/sbin/felhom-crash-guard (0755 root:root), with felhom-crash-guard.service (boot / clean-stop)
# and felhom-crash-guard-check.timer (hourly re-arm check). Python 3, standard library only.
# Tests: configs/test_felhom_crash_guard.py (temp dirs; nothing real is touched).
#
# WHAT IT DOES
# boot early at every boot. Was the previous boot ended CLEANLY? (the clean-stop marker exists). If not, this
# boot follows an UNCLEAN stop — a kernel crash, a power cut or a hard reset (they cannot be told apart
# on these boxes: measured 2026-10-04 on demo-hp, efi_pstore is on yet saved NOTHING for a real panic;
# the journal and `last` show only "no shutdown"). It records the unclean boot, counts those in the last
# WINDOW_MINUTES, and sets kernel.panic:
# - fewer than LIMIT-1 recent unclean boots → kernel.panic = PANIC_SECONDS (a crash restarts the box);
# - LIMIT-1 or more → the guard TRIPS: kernel.panic = 0, so the LIMIT-th crash within the window
# leaves the box OFF (operator's own words: "if it crashes 3 times within one hour, it stays off").
# A tripped guard stays tripped across further boots until it re-arms.
# clean-stop ExecStop of the service: writes the clean-stop marker during an orderly shutdown or reboot.
# check hourly: a tripped guard re-arms after REARM_HOURS of normal running (since the trip AND since boot).
# rearm the operator re-arms by hand (`felhom-crash-guard rearm`).
# status prints the state.
# The state is /var/lib/felhom-crash-guard/state.json (0644: the non-root agent reads it into its host report).
# Before the service runs (very early boot) the kernel default kernel.panic = 0 applies, so a crash THAT early leaves
# the box off — the safe side: a box that cannot reach userspace must not loop.
import json
import os
import sys
import time
CONF = "/etc/felhom/crash-guard.conf"
STATE_DIR = "/var/lib/felhom-crash-guard"
DEFAULTS = {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10, "REARM_HOURS": 24}
class Env:
"""Paths and clock; tests replace them."""
def __init__(self, conf=CONF, state_dir=STATE_DIR, panic_path="/proc/sys/kernel/panic",
uptime_path="/proc/uptime", boot_id_path="/proc/sys/kernel/random/boot_id"):
self.conf, self.state_dir = conf, state_dir
self.panic_path, self.uptime_path, self.boot_id_path = panic_path, uptime_path, boot_id_path
def now(self):
return time.time()
def log(self, line):
print(line, file=sys.stderr, flush=True)
try:
import subprocess
subprocess.run(["logger", "-t", "felhom-crash-guard", line], timeout=10)
except Exception:
pass
def iso(t):
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(t))
def parse_iso(s):
import calendar
return calendar.timegm(time.strptime(s, "%Y-%m-%dT%H:%M:%SZ"))
def load_conf(env):
c = dict(DEFAULTS)
try:
for line in open(env.conf):
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
k, v = (x.strip() for x in line.split("=", 1))
if k in c and v.isdigit() and int(v) >= (1 if k != "PANIC_SECONDS" else 1):
c[k] = int(v)
except OSError:
pass
return c
def state_path(env):
return os.path.join(env.state_dir, "state.json")
def marker_path(env):
return os.path.join(env.state_dir, "clean-stop")
def load_state(env):
try:
with open(state_path(env)) as f:
s = json.load(f)
return s if isinstance(s, dict) else None
except (OSError, ValueError):
return None
def save_state(env, s):
os.makedirs(env.state_dir, mode=0o755, exist_ok=True)
tmp = state_path(env) + ".tmp"
with open(tmp, "w") as f:
json.dump(s, f, indent=2, sort_keys=True)
f.write("\n")
os.chmod(tmp, 0o644)
os.replace(tmp, state_path(env))
def set_panic(env, seconds):
with open(env.panic_path, "w") as f:
f.write(f"{seconds}\n")
def read(path, default=""):
try:
with open(path) as f:
return f.read().strip()
except OSError:
return default
def summarize(s, c, now):
window = c["WINDOW_MINUTES"] * 60
times = [parse_iso(t) for t in s.get("unclean_boots", [])]
after = parse_iso(s["rearmed_at"]) if s.get("rearmed_at") else 0
# a re-arm starts a fresh window (or the next unclean boot would trip again at once); the history stays
s["unclean_boots_in_window"] = sum(1 for t in times if now - t <= window and t > after)
s["unclean_boots_24h"] = sum(1 for t in times if now - t <= 86400)
s["config"] = c
s["updated_at"] = iso(now)
def boot(env):
c = load_conf(env)
now = env.now()
try:
up = float(read(env.uptime_path, "0").split()[0])
except (ValueError, IndexError):
up = 0.0
boot_at = now - up
prev = load_state(env)
first = prev is None
s = prev or {"version": 1, "unclean_boots": [], "tripped": False}
clean = os.path.exists(marker_path(env))
unclean = (not first) and (not clean)
try:
os.remove(marker_path(env))
except OSError:
pass
# keep 7 days of history (the 24 h figure and the operator's view), drop older
s["unclean_boots"] = [t for t in s.get("unclean_boots", []) if now - parse_iso(t) <= 7 * 86400]
if unclean:
s["unclean_boots"].append(iso(boot_at))
s["last_boot_at"] = iso(boot_at)
s["last_boot_unclean"] = unclean
s["boot_id"] = read(env.boot_id_path, "unknown")
summarize(s, c, now)
if not s.get("tripped") and s["unclean_boots_in_window"] >= c["LIMIT"] - 1:
s["tripped"], s["tripped_at"] = True, iso(now)
s["tripped_reason"] = (f"{s['unclean_boots_in_window']} unclean boots within {c['WINDOW_MINUTES']} minutes — "
f"the next crash leaves the box off (limit {c['LIMIT']})")
env.log(f"crash-guard: TRIPPED: {s['tripped_reason']}")
panic = 0 if s.get("tripped") else c["PANIC_SECONDS"]
set_panic(env, panic)
s["kernel_panic"] = panic
s["armed"] = not s.get("tripped")
save_state(env, s)
env.log(f"crash-guard: boot first={first} unclean={unclean} in-window={s['unclean_boots_in_window']} "
f"tripped={s.get('tripped')} kernel.panic={panic}")
return 0
def clean_stop(env):
os.makedirs(env.state_dir, mode=0o755, exist_ok=True)
with open(marker_path(env), "w") as f:
f.write(iso(env.now()) + "\n")
env.log("crash-guard: clean stop recorded")
return 0
def rearm(env, by):
c = load_conf(env)
now = env.now()
s = load_state(env) or {"version": 1, "unclean_boots": []}
was = bool(s.get("tripped"))
s["tripped"] = False
s["armed"] = True
s["rearmed_at"], s["rearmed_by"] = iso(now), by
if was:
s["last_trip"] = {"at": s.get("tripped_at"), "reason": s.get("tripped_reason")}
s.pop("tripped_at", None)
s.pop("tripped_reason", None)
summarize(s, c, now)
set_panic(env, c["PANIC_SECONDS"])
s["kernel_panic"] = c["PANIC_SECONDS"]
save_state(env, s)
env.log(f"crash-guard: RE-ARMED by {by} (was tripped: {was}); kernel.panic={c['PANIC_SECONDS']}")
return 0
def check(env):
c = load_conf(env)
now = env.now()
s = load_state(env)
if not s:
return 0
if s.get("tripped"):
since = max(parse_iso(s["tripped_at"]), parse_iso(s.get("last_boot_at", s["tripped_at"])))
if now - since >= c["REARM_HOURS"] * 3600:
return rearm(env, f"timer ({c['REARM_HOURS']} h of normal running)")
summarize(s, c, now)
save_state(env, s)
return 0
def main(argv, env=None):
env = env or Env()
cmd = argv[1] if len(argv) == 2 else ""
if cmd == "boot":
return boot(env)
if cmd == "clean-stop":
return clean_stop(env)
if cmd == "check":
return check(env)
if cmd == "rearm":
return rearm(env, "operator")
if cmd == "status":
print(json.dumps(load_state(env), indent=2, sort_keys=True))
return 0
print("usage: felhom-crash-guard boot|clean-stop|check|rearm|status", file=sys.stderr)
return 2
if __name__ == "__main__":
if os.geteuid() != 0:
print("felhom-crash-guard: must run as root", file=sys.stderr)
sys.exit(2)
sys.exit(main(sys.argv))
+7
View File
@@ -0,0 +1,7 @@
# Hourly: a tripped crash guard re-arms after REARM_HOURS of normal running (felhom-crash-guard check).
[Unit]
Description=Felhom crash guard re-arm check
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/felhom-crash-guard check
+9
View File
@@ -0,0 +1,9 @@
[Unit]
Description=Felhom crash guard re-arm check (hourly)
[Timer]
OnBootSec=15min
OnUnitActiveSec=1h
[Install]
WantedBy=timers.target
+19
View File
@@ -0,0 +1,19 @@
# felhom-crash-guard — a crashed host restarts by itself, with a limit (`09` decision 88, R-851, `11` §5.9).
# Starts early at boot (sets kernel.panic for THIS boot); its ExecStop writes the clean-stop marker during an orderly
# shutdown or reboot. A boot that finds no marker followed a crash, a power cut or a hard reset.
[Unit]
Description=Felhom crash guard (restart after a kernel crash, with a limit)
DefaultDependencies=no
After=local-fs.target
Before=sysinit.target shutdown.target
Conflicts=shutdown.target
RequiresMountsFor=/var/lib
[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=/usr/local/sbin/felhom-crash-guard boot
ExecStop=/usr/local/sbin/felhom-crash-guard clean-stop
[Install]
WantedBy=sysinit.target
+39
View File
@@ -0,0 +1,39 @@
#!/bin/sh
# /etc/grub.d/42_felhom_oneshot — the kernel lane's one-shot ENTRIES (R-836, `09` §3 decision 172, `11` §5.11).
# Installed by the config bundle (felhom-os-apply BUNDLE_FILES), 0755 root. update-grub runs it.
#
# One menu entry per installed Proxmox kernel, id `felhom-oneshot-<version>`, booted ONLY when 01_felhom_oneshot found
# the flag naming it. It is the normal entry plus option C (decision 172): softlockup_panic=1 hardlockup_panic=1
# hung_task_panic=1 panic=10 — a lockup the kernel can detect becomes a panic, and a panic restarts the box in 10 s into
# the default (the old kernel). A true dead freeze still needs a person (spike candidate 3 failed on all three boxes).
# It sorts AFTER 10_linux, so it is never entry 0 and never the default.
#
# No vfat ESP at /boot/efi → prints nothing (no flag can name these entries).
set -e
prefix="/usr"
exec_prefix="/usr"
datarootdir="/usr/share"
. "$datarootdir/grub/grub-mkconfig_lib"
esp_uuid=$(findmnt -n -o UUID,FSTYPE /boot/efi 2>/dev/null | awk '$2 == "vfat" { print $1 }')
[ -n "$esp_uuid" ] || exit 0
kernels=$(ls /boot/vmlinuz-*-pve 2>/dev/null | sed 's#^/boot/vmlinuz-##' | grep -E '^[0-9]+\.[0-9]+\.[0-9]+-[0-9]+-pve$' || true)
[ -n "$kernels" ] || exit 0
case "${GRUB_DEVICE}" in
/dev/mapper/*|/dev/dm-*|"") root_arg="root=${GRUB_DEVICE}" ;;
*) if [ -n "${GRUB_DEVICE_UUID}" ]; then root_arg="root=UUID=${GRUB_DEVICE_UUID}"; else root_arg="root=${GRUB_DEVICE}"; fi ;;
esac
[ -n "${GRUB_DEVICE}" ] || root_arg="root=$(findmnt -n -o SOURCE /)"
rel=$(make_system_path_relative_to_its_root /boot)
prep=$(prepare_grub_to_access_device "$(${grub_probe:-grub-probe} --target=device /boot)" | sed 's/^/ /')
for k in $kernels; do
[ -f "/boot/initrd.img-$k" ] || continue
cat <<EOF
menuentry 'Felhom one-shot: $k' --class proxmox --id felhom-oneshot-$k {
insmod gzio
$prep
echo 'Loading Linux $k (felhom one-shot) ...'
linux $rel/vmlinuz-$k $root_arg ro ${GRUB_CMDLINE_LINUX} ${GRUB_CMDLINE_LINUX_DEFAULT} softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10
initrd $rel/initrd.img-$k
}
EOF
done
+36
View File
@@ -0,0 +1,36 @@
#!/bin/sh
# /etc/grub.d/01_felhom_oneshot — the kernel lane's ONE-SHOT boot (R-836, `09` §3 decisions 164 + 172, `11` §5.11).
# Installed by the config bundle (felhom-os-apply BUNDLE_FILES), 0755 root. update-grub runs it; it prints GRUB script.
#
# At boot, GRUB reads `felhom_next` from an environment block on the ESP (vfat — GRUB can rewrite a file there; it
# cannot on the LVM /boot, R-836), CLEARS it, and — only if it names an installed kernel — boots that kernel's one-shot
# entry (42_felhom_oneshot) instead of the default. The next boot uses the default again whatever happens: a new kernel
# that panics comes back on the old one by itself. felhom-os-apply writes the flag (mode apply) and never the default
# for a new kernel. Measured on the Tester 1 VM, demo-felhom and demo-hp (Secure Boot on):
# `audits/kernel-spike-2026-10-07/` (candidate 2).
#
# No vfat ESP at /boot/efi, or no Proxmox kernel → prints nothing (the box boots exactly as before).
set -e
esp_uuid=$(findmnt -n -o UUID,FSTYPE /boot/efi 2>/dev/null | awk '$2 == "vfat" { print $1 }')
[ -n "$esp_uuid" ] || exit 0
kernels=$(ls /boot/vmlinuz-*-pve 2>/dev/null | sed 's#^/boot/vmlinuz-##' | grep -E '^[0-9]+\.[0-9]+\.[0-9]+-[0-9]+-pve$' || true)
[ -n "$kernels" ] || exit 0
cat <<EOF
# felhom kernel lane: a one-shot kernel named on the ESP, read and cleared before the menu
insmod part_gpt
insmod fat
search --no-floppy --fs-uuid --set=felhom_esp $esp_uuid
if [ -f (\$felhom_esp)/EFI/felhom/oneshot.env ]; then
load_env -f (\$felhom_esp)/EFI/felhom/oneshot.env felhom_next
if [ "\${felhom_next}" ]; then
set felhom_boot="\${felhom_next}"
set felhom_next=
save_env -f (\$felhom_esp)/EFI/felhom/oneshot.env felhom_next
EOF
for k in $kernels; do
printf ' if [ "${felhom_boot}" = "%s" ]; then set default="felhom-oneshot-%s"; fi\n' "$k" "$k"
done
cat <<EOF
fi
fi
EOF
+4
View File
@@ -0,0 +1,4 @@
#!/bin/sh
# felhom-agent guest pre-start self-heal hook (C1 net). PVE calls: <script> <vmid> <phase>.
/usr/local/bin/felhom-agent guest-hook "$1" "$2" || true
exit 0
+5 -3
View File
@@ -16,8 +16,10 @@ Cmnd_Alias FELHOM_OP_REPAIR = \
/usr/bin/systemctl reset-failed felhom-sshd, \
/usr/bin/systemctl restart felhom-sshd, \
/usr/sbin/pct list, \
/usr/sbin/pct start [0-9]*, \
/usr/sbin/pct stop [0-9]*, \
/usr/sbin/pct unlock [0-9]*
/usr/sbin/pct ^start [0-9]+$, \
/usr/sbin/pct ^stop [0-9]+$, \
/usr/sbin/pct ^unlock [0-9]+$
# R-861 (b) B2 (`09` §3 decision 165, hygiene): one numeric vmid per pct verb, anchored — the old glob `[0-9]*` also
# matched spaces, so `pct stop 9201 --skiplock 1` passed. Pinned by TestFelhomOpSudoersPctIsExact.
felhom-op ALL=(root) NOPASSWD: FELHOM_OP_REPAIR
+1674 -44
View File
File diff suppressed because it is too large Load Diff
+468
View File
@@ -0,0 +1,468 @@
#!/usr/bin/python3
"""felhom-priv-apply — the ROOT half of every file the agent writes into a root-read place (R-861, `03` §3.1).
Install as /usr/local/sbin/felhom-priv-apply (0755 root:root) — it rides the signed config bundle (R-840).
WHY IT EXISTS. Until agent v0.146.0 the agent's sudoers let it `install` a file it had written itself into a place a
root program reads: a systemd .mount unit (a bind mount of an agent-owned directory over /etc/sudoers.d is a root
shell), a dnsmasq drop-in (`dhcp-script=` runs as root), the WireGuard config (`PostUp=` runs as root) and the OOB
sshd config (`AuthorizedKeysFile` + `StrictModes no`). A compromised agent PROCESS was therefore root on its host.
Now the agent stages the file and this wrapper — root-owned, delivered only by an operator-signed bundle — checks
the CONTENT against the exact grammar the agent's own renderers produce, and refuses anything else. The agent can no
longer name the destination: each verb has a fixed source and a fixed (or strictly named) destination.
Verbs (each one sudoers line, exact-match pattern):
unit <name> /var/lib/felhom-agent/units/<name> -> /etc/systemd/system/<name> (.mount | .automount)
dnsmasq <tmp> <name> /tmp/felhom-resolver-<digits>.conf -> /etc/dnsmasq.d/felhom-<...>.conf
wg /var/lib/felhom-agent/wg/wg-felhom.conf -> /etc/wireguard/wg-felhom.conf (0600)
sshd-config /var/lib/felhom-agent/felhom-sshd/sshd_config -> /etc/felhom-sshd/sshd_config
sshd-key /var/lib/felhom-agent/felhom-sshd/authorized_keys.felhom-op -> /etc/felhom-sshd/authorized_keys/felhom-op
controller-image <vmid> the ref on STDIN -> /etc/felhom-controller-image INSIDE guest <vmid> (R-861 (a) A1): only
our registry + our repository + an x.y.z tag; the agent no longer has a `tee` grant
--self-check prints "felhom-priv-apply ok verbs=..." (the bundle's self-check)
Exit codes: 0 installed (or already identical), 2 usage, 3 refused (content or source), 4 install failed.
Every refusal is logged to the journal (tag felhom-priv-apply) with its rule; file CONTENT is never logged.
Pinned by configs/test_felhom_priv_apply.py (one test per rule, red-proofs in the R-861 audit).
"""
import ipaddress
import os
import re
import stat
import subprocess
import sys
AGENT_USER = "felhom-agent"
STATE = "/var/lib/felhom-agent"
UNITS_SRC = STATE + "/units"
UNIT_DIR = "/etc/systemd/system"
DNSMASQ_DIR = "/etc/dnsmasq.d"
WG_SRC, WG_DEST = STATE + "/wg/wg-felhom.conf", "/etc/wireguard/wg-felhom.conf"
SSHD_SRC, SSHD_DEST = STATE + "/felhom-sshd/sshd_config", "/etc/felhom-sshd/sshd_config"
KEY_SRC, KEY_DEST = STATE + "/felhom-sshd/authorized_keys.felhom-op", "/etc/felhom-sshd/authorized_keys/felhom-op"
MAX_BYTES = 64 * 1024
VERBS = ("unit", "dnsmasq", "wg", "sshd-config", "sshd-key", "controller-image")
# R-861 (a) A1 (`09` §3 decision 165): the SAME pattern as the agent's controllerImageRe (internal/localapi/
# controllerswap.go) — a compromised agent cannot hand the guest's bootstrap any other image. Pinned by
# configs/test_felhom_priv_apply.py ControllerImage.
CONTROLLER_IMAGE_RE = re.compile(r"^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$")
CONTROLLER_IMAGE_FILE = "/etc/felhom-controller-image"
CONTROLLER_IMAGE_MAX = 256
VMID_RE = re.compile(r"^[0-9]{1,9}$")
UNIT_NAME_RE = re.compile(r"^mnt-[A-Za-z0-9_.\\-]+\.(mount|automount)$")
DNSMASQ_TMP_RE = re.compile(r"^/tmp/felhom-resolver-[0-9]+\.conf$")
DNSMASQ_NAME_RE = re.compile(r"^felhom-[a-z0-9][a-z0-9._-]*\.conf$")
SEG = r"[A-Za-z0-9_-][A-Za-z0-9_.-]*"
WHERE_RE = re.compile(r"^/mnt/(felhom-drives/)?" + SEG + r"$")
UUID_RE = re.compile(r"^[A-Fa-f0-9]{4,}(-[A-Fa-f0-9]+){0,4}$")
HOST_RE = re.compile(r"^[A-Za-z0-9._:-]{1,255}$")
NET_PATH_RE = re.compile(r"^[A-Za-z0-9._/@+-]{1,512}$")
OPT_RE = re.compile(r"^[A-Za-z0-9_.:/@+-]+(=[A-Za-z0-9_.:/@+-]+)?$")
LOCAL_TYPES = {"ext4", "xfs", "btrfs", "exfat", "vfat", "ntfs3", "ntfs"}
NET_TYPES = {"nfs", "nfs4", "cifs"}
# Options that turn a device mount into something else, or let set-uid/device files act on the host.
FORBIDDEN_OPTS = {"bind", "rbind", "move", "rmove", "remount", "suid", "dev", "user", "users", "owner", "group",
"x-mount.mkdir", "helper"}
DESC_RE = re.compile(r"^[^\x00-\x1f\x7f]{0,200}$")
WG_KEY_RE = re.compile(r"^[A-Za-z0-9+/]{42}[AEIMQUYcgkosw480]=$")
KEY_LINE_RE = re.compile(r"^(ssh-ed25519|ssh-rsa|ecdsa-sha2-nistp(256|384|521)|sk-ssh-ed25519@openssh\.com) "
r"[A-Za-z0-9+/]+={0,3}( [ -~]{0,200})?$")
class Refused(Exception):
def __init__(self, rule, reason):
super().__init__(reason)
self.rule, self.reason = rule, reason
class Host:
"""Every filesystem / process effect, so the tests can play the box in memory."""
def agent_uid(self):
import pwd
return pwd.getpwnam(AGENT_USER).pw_uid
def read_source(self, path):
"""The staged file: a REGULAR file owned by the agent, never a symlink, at most MAX_BYTES."""
try:
fd = os.open(path, os.O_RDONLY | os.O_NOFOLLOW | os.O_CLOEXEC)
except OSError as e:
raise Refused("P1", f"cannot open the staged file {path}: {e.strerror}")
try:
st = os.fstat(fd)
if not stat.S_ISREG(st.st_mode):
raise Refused("P1", f"{path} is not a regular file")
if st.st_uid != self.agent_uid():
raise Refused("P1", f"{path} is not owned by {AGENT_USER}")
if st.st_size > MAX_BYTES:
raise Refused("P1", f"{path} is larger than {MAX_BYTES} bytes")
with os.fdopen(fd, "rb") as f:
fd = -1
return f.read(MAX_BYTES + 1)
finally:
if fd >= 0:
os.close(fd)
def read_dest(self, path):
try:
with open(path, "rb") as f:
return f.read()
except OSError:
return None
def install(self, dest, data, mode):
"""Atomic, root-owned: a temp file beside the destination, fsync, rename."""
d = os.path.dirname(dest)
os.makedirs(d, mode=0o755, exist_ok=True)
tmp = os.path.join(d, f".{os.path.basename(dest)}.felhom-new.{os.getpid()}")
fd = os.open(tmp, os.O_WRONLY | os.O_CREAT | os.O_EXCL | os.O_NOFOLLOW, 0o600)
try:
with os.fdopen(fd, "wb") as f:
f.write(data)
f.flush()
os.fchown(f.fileno(), 0, 0)
os.fchmod(f.fileno(), mode)
os.fsync(f.fileno())
os.replace(tmp, dest)
except BaseException:
try:
os.remove(tmp)
except OSError:
pass
raise
def read_stdin(self, limit):
return sys.stdin.buffer.read(limit + 1)
def write_guest_image(self, vmid, data):
"""As root: `pct exec <vmid> -- tee <the fixed file>` with the checked ref on stdin (no shell)."""
r = subprocess.run(["/usr/sbin/pct", "exec", str(vmid), "--", "tee", CONTROLLER_IMAGE_FILE],
input=data, stdout=subprocess.DEVNULL, stderr=subprocess.PIPE, timeout=60)
if r.returncode != 0:
raise OSError(f"pct exec {vmid} tee exited {r.returncode}: {r.stderr.decode(errors='replace').strip()[:200]}")
def log(self, line):
print(line, file=sys.stderr)
try:
subprocess.run(["logger", "-t", "felhom-priv-apply", "--", line], timeout=10, check=False)
except (OSError, subprocess.SubprocessError):
pass
def systemd_escape_path(path):
"""`systemd-escape --path`: strip the slashes at both ends, `/` -> `-`, every byte outside [A-Za-z0-9:_.] (and a
leading `.`) -> `\\xNN`."""
p = path.strip("/")
out = []
for i, ch in enumerate(p):
if ch == "/":
out.append("-")
elif (ch.isascii() and (ch.isalnum() or ch in ":_.")) and not (i == 0 and ch == "."):
out.append(ch)
else:
out.extend("\\x%02x" % b for b in ch.encode())
return "".join(out)
def text_of(data, what):
if len(data) > MAX_BYTES:
raise Refused("P1", f"{what} is too large")
try:
text = data.decode("utf-8")
except UnicodeDecodeError:
raise Refused("P2", f"{what} is not UTF-8 text")
if "\x00" in text or "\r" in text:
raise Refused("P2", f"{what} carries a NUL or CR byte")
return text
def parse_ini(text, what):
"""[Section] / Key=Value / comments / blank lines. A key outside a section or a repeated key is refused."""
sections, cur = {}, None
for n, raw in enumerate(text.split("\n"), 1):
line = raw.strip()
if line.endswith("\\"):
# systemd joins a line ending in a backslash with the next one; this parser does not. Refused, so the two
# can never read the same bytes differently (review 2026-10-05).
raise Refused("U2", f"{what}: line {n} ends with a backslash (a continuation)")
if not line or line.startswith("#") or line.startswith(";"):
continue
m = re.match(r"^\[([A-Za-z]+)\]$", line)
if m:
cur = m.group(1)
if cur in sections:
raise Refused("U2", f"{what}: section [{cur}] twice")
sections[cur] = {}
continue
if cur is None or "=" not in line:
raise Refused("U2", f"{what}: line {n} is not Key=Value inside a section")
k, v = line.split("=", 1)
k, v = k.strip(), v.strip()
if k in sections[cur]:
raise Refused("U2", f"{what}: {cur}.{k} given twice")
sections[cur][k] = v
return sections
# ---------- the unit verb ----------
def check_unit(name, text):
if not UNIT_NAME_RE.match(name) or "/" in name:
raise Refused("U1", f"unit name {name!r} is not mnt-<escaped path>.mount|.automount")
kind = "automount" if name.endswith(".automount") else "mount"
s = parse_ini(text, name)
body = "Automount" if kind == "automount" else "Mount"
# [Unit] holds ONLY what the renderers write: Description, and After=local-fs-pre.target on a local mount. A
# Wants=/Requires=/Before= naming any unit would start it with the mount (Wants=reboot.target — found by review
# 2026-10-05), so none of them is accepted.
allowed = {"Unit": {"Description", "After"},
body: {"Where", "TimeoutIdleSec"} if kind == "automount" else {"What", "Where", "Type", "Options"},
"Install": {"WantedBy"}}
for sec, keys in s.items():
if sec not in allowed:
raise Refused("U2", f"{name}: section [{sec}] is not allowed")
bad = set(keys) - allowed[sec]
if bad:
raise Refused("U2", f"{name}: [{sec}] key(s) {sorted(bad)} not allowed")
u = s.get("Unit", {})
if not DESC_RE.match(u.get("Description", "")):
raise Refused("U2", f"{name}: Description has control characters")
if "After" in u and u["After"] != "local-fs-pre.target":
raise Refused("U2", f"{name}: After= may only be local-fs-pre.target")
inst = s.get("Install", {})
if inst and inst.get("WantedBy") != "multi-user.target":
raise Refused("U2", f"{name}: WantedBy must be multi-user.target")
m = s.get(body)
if not m or "Where" not in m:
raise Refused("U3", f"{name}: no [{body}] Where=")
where = m["Where"]
if not WHERE_RE.match(where):
raise Refused("U3", f"{name}: Where={where} is not /mnt/<name> or /mnt/felhom-drives/<name>")
if systemd_escape_path(where) + "." + kind != name:
raise Refused("U3", f"{name}: the unit name does not match Where={where}")
if kind == "automount":
t = m.get("TimeoutIdleSec", "")
if t and not re.match(r"^[0-9]{1,6}$", t):
raise Refused("U2", f"{name}: TimeoutIdleSec must be seconds")
return
what, typ = m.get("What", ""), m.get("Type", "")
opts = [o for o in m.get("Options", "").split(",") if o]
net = False
mu = re.match(r"^/dev/disk/by-uuid/(.+)$", what)
if mu:
if not UUID_RE.match(mu.group(1)):
raise Refused("U4", f"{name}: What= is not a filesystem UUID")
if typ and typ not in LOCAL_TYPES:
raise Refused("U4", f"{name}: Type={typ} is not a local filesystem")
else:
net = True
if typ not in NET_TYPES:
raise Refused("U4", f"{name}: What= is neither /dev/disk/by-uuid/<uuid> nor a network source with Type=nfs/nfs4/cifs")
if typ == "cifs":
mm = re.match(r"^//([^/]+)/(.+)$", what)
else:
mm = re.match(r"^([^/:][^:]*):(/.*)$", what)
if not mm or not HOST_RE.match(mm.group(1)) or not NET_PATH_RE.match(mm.group(2)) or ".." in mm.group(2).split("/"):
raise Refused("U4", f"{name}: What= is not a clean {typ} source")
if not where.startswith("/mnt/felhom-drives/"):
raise Refused("U3", f"{name}: a network share mounts only under /mnt/felhom-drives/")
for o in opts:
if not OPT_RE.match(o):
raise Refused("U5", f"{name}: mount option {o!r} has characters a mount option never needs")
if o.split("=", 1)[0].lower() in FORBIDDEN_OPTS or o.lower().startswith("x-mount."):
raise Refused("U5", f"{name}: mount option {o.split('=', 1)[0]!r} is not allowed")
if net and not {"nosuid", "nodev"} <= set(opts):
# A network server is outside the box: a set-uid file on it must never run as root here.
raise Refused("U5", f"{name}: a network share must carry nosuid,nodev")
# ---------- dnsmasq ----------
def _ip(v, v6=True):
try:
a = ipaddress.ip_address(v)
except ValueError:
return False
return v6 or a.version == 4
DOMAIN_RE = re.compile(r"^[A-Za-z0-9]([A-Za-z0-9-]{0,62})(\.[A-Za-z0-9]([A-Za-z0-9-]{0,62}))*$")
def check_dnsmasq(text):
for n, raw in enumerate(text.split("\n"), 1):
line = raw.strip()
if not line or line.startswith("#"):
continue
if line in ("bind-interfaces", "no-resolv"):
continue
k, _, v = line.partition("=")
if k == "listen-address" and _ip(v, v6=False):
continue
if k == "server" and (_ip(v) or (v.count("#") == 1 and _ip(v.split("#")[0]) and v.split("#")[1].isdigit())):
continue
m = re.match(r"^/([^/]+)/$", v)
if k == "local" and m and DOMAIN_RE.match(m.group(1)):
continue
m = re.match(r"^/([^/]+)/([^/]+)$", v)
if k == "address" and m and DOMAIN_RE.match(m.group(1)) and _ip(m.group(2), v6=False):
continue
raise Refused("D1", f"dnsmasq line {n} ({k or line[:20]!r}) is not one the resolver writes")
# ---------- WireGuard ----------
def check_wg(text):
s = parse_ini(text, "wg-felhom.conf")
if set(s) != {"Interface", "Peer"}:
raise Refused("W1", "wg-felhom.conf must hold exactly [Interface] and [Peer]")
i, p = s["Interface"], s["Peer"]
if set(i) - {"PrivateKey", "Address", "MTU"} or set(p) - {"PublicKey", "Endpoint", "AllowedIPs", "PersistentKeepalive"}:
raise Refused("W1", "wg-felhom.conf carries a key the agent never writes (PostUp/PreUp/... run as root)")
if not WG_KEY_RE.match(i.get("PrivateKey", "")) or not WG_KEY_RE.match(p.get("PublicKey", "")):
raise Refused("W2", "a WireGuard key is not 32 bytes of base64")
try:
a = ipaddress.ip_network(i.get("Address", ""), strict=False)
if a.version != 4 or a.prefixlen != 32:
raise ValueError
if not (1280 <= int(i.get("MTU", "1280")) <= 1500):
raise ValueError
host, _, port = p.get("Endpoint", "").rpartition(":")
if ipaddress.ip_address(host).version != 4 or not (1 <= int(port) <= 65535):
raise ValueError
for n in p.get("AllowedIPs", "").split(","):
if ipaddress.ip_network(n.strip(), strict=True).prefixlen != 32:
raise ValueError
if not (0 <= int(p.get("PersistentKeepalive", "25")) <= 3600):
raise ValueError
except ValueError:
raise Refused("W2", "an Address/MTU/Endpoint/AllowedIPs/PersistentKeepalive value is not what the agent renders")
# ---------- OOB sshd ----------
def render_sshd(port):
"""Byte-identical to felhomsshd.renderConfig (internal/felhomsshd/config.go) — pinned by a Go test."""
return ("# felhom OOB sshd — agent-managed (H1); DO NOT EDIT\n"
f"Port {port}\n"
"ListenAddress 0.0.0.0\n"
"ListenAddress ::\n"
"HostKey /etc/felhom-sshd/ssh_host_ed25519_key\n"
"PidFile /run/felhom-sshd.pid\n"
"AuthorizedKeysFile /etc/felhom-sshd/authorized_keys/%u\n"
"PasswordAuthentication no\n"
"PermitRootLogin prohibit-password\n"
"PubkeyAuthentication yes\n"
"KbdInteractiveAuthentication no\n"
"UsePAM yes\n"
"AllowUsers root felhom-op\n"
"X11Forwarding no\n"
"Subsystem sftp internal-sftp\n")
def check_sshd(text):
m = re.search(r"^Port ([0-9]{1,5})$", text, re.M)
if not m or not (1 <= int(m.group(1)) <= 65535) or int(m.group(1)) == 22:
raise Refused("S1", "sshd_config has no Port (or claims :22, the household's sshd)")
if text != render_sshd(int(m.group(1))):
raise Refused("S1", "sshd_config differs from the one fixed template (only the Port may vary)")
def check_key(text):
lines = [l for l in text.split("\n") if l.strip()]
if len(lines) > 1:
raise Refused("S2", "felhom-op's authorized_keys holds more than one key")
if lines and not KEY_LINE_RE.match(lines[0]):
raise Refused("S2", "the key line is not a plain public key (no options such as command= or from=)")
# ---------- main ----------
def plan(argv):
"""(verb, source, dest, mode, checker) for an argv, or Refused("A1")."""
if not argv or argv[0] not in VERBS:
raise Refused("A1", "usage: felhom-priv-apply unit <name> | dnsmasq <tmp> <name> | wg | sshd-config | sshd-key")
v, rest = argv[0], argv[1:]
if v == "unit" and len(rest) == 1:
if not UNIT_NAME_RE.match(rest[0]):
raise Refused("U1", f"unit name {rest[0]!r} is not mnt-<escaped path>.mount|.automount")
return v, os.path.join(UNITS_SRC, rest[0]), os.path.join(UNIT_DIR, rest[0]), 0o644, lambda t: check_unit(rest[0], t)
if v == "dnsmasq" and len(rest) == 2:
if not DNSMASQ_TMP_RE.match(rest[0]) or not DNSMASQ_NAME_RE.match(rest[1]):
raise Refused("D2", "dnsmasq wants /tmp/felhom-resolver-<digits>.conf and felhom-<name>.conf")
return v, rest[0], os.path.join(DNSMASQ_DIR, rest[1]), 0o644, check_dnsmasq
if v == "wg" and not rest:
return v, WG_SRC, WG_DEST, 0o600, check_wg
if v == "sshd-config" and not rest:
return v, SSHD_SRC, SSHD_DEST, 0o644, check_sshd
if v == "sshd-key" and not rest:
return v, KEY_SRC, KEY_DEST, 0o644, check_key
raise Refused("A1", f"wrong arguments for {v}")
def controller_image(rest, host):
"""R-861 (a) A1: read the ref on stdin, check it, write it INSIDE the guest as root."""
try:
if len(rest) != 1 or not VMID_RE.match(rest[0]):
raise Refused("A1", "usage: felhom-priv-apply controller-image <vmid> (the ref on stdin)")
raw = host.read_stdin(CONTROLLER_IMAGE_MAX)
if len(raw) > CONTROLLER_IMAGE_MAX:
raise Refused("I1", f"the image ref is longer than {CONTROLLER_IMAGE_MAX} bytes")
try:
text = raw.decode("ascii")
except UnicodeDecodeError:
raise Refused("I1", "the image ref is not ASCII")
ref = text[:-1] if text.endswith("\n") else text
if not CONTROLLER_IMAGE_RE.match(ref) or "\n" in ref:
raise Refused("I1", "the image ref is not gitea.dooplex.hu/admin/felhom-controller:<x.y.z>")
except Refused as e:
host.log(f"felhom-priv-apply: REFUSED [{e.rule}] controller-image {' '.join(rest)[:40]}: {e.reason}")
return 2 if e.rule == "A1" else 3
vmid = int(rest[0])
try:
host.write_guest_image(vmid, (ref + "\n").encode())
except (OSError, subprocess.SubprocessError) as e:
host.log(f"felhom-priv-apply: FAILED controller-image {vmid}: {e}")
return 4
host.log(f"felhom-priv-apply: WROTE controller-image {vmid} {ref}")
return 0
def main(argv, host=None):
host = host or Host()
if argv == ["--self-check"]:
print("felhom-priv-apply ok verbs=" + ",".join(VERBS))
return 0
if len(argv) >= 3 and argv[0] == "--check":
# CHECK ONLY (tests and the Go contract tests): `--check <verb> [<name>] <file>` validates <file> as that
# verb would and installs nothing. Not in sudoers. Prints OK or the rule; never the content.
verb, file = argv[1], argv[-1]
try:
_, _, _, _, checker = plan([verb] + argv[2:-1] if verb != "dnsmasq" else
[verb, "/tmp/felhom-resolver-1.conf", argv[2]])
with open(file, "rb") as f:
checker(text_of(f.read(), file))
except Refused as e:
print(f"REFUSED [{e.rule}] {e.reason}")
return 3
print("OK")
return 0
if argv and argv[0] == "controller-image":
return controller_image(argv[1:], host)
try:
verb, src, dest, mode, checker = plan(argv)
data = host.read_source(src)
checker(text_of(data, src))
except Refused as e:
host.log(f"felhom-priv-apply: REFUSED [{e.rule}] {' '.join(argv)[:160]}: {e.reason}")
return 2 if e.rule == "A1" else 3
if host.read_dest(dest) == data:
host.log(f"felhom-priv-apply: SAME {verb} {dest}")
return 0
try:
host.install(dest, data, mode)
except OSError as e:
host.log(f"felhom-priv-apply: FAILED {verb} {dest}: {e}")
return 4
host.log(f"felhom-priv-apply: INSTALLED {verb} {dest} ({len(data)} bytes)")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+11 -3
View File
@@ -24,6 +24,11 @@ set -u
BIN=/usr/local/bin/felhom-agent
PREV=$BIN.prev
STAGING=/var/lib/felhom-agent/selfupdate
# R-861 (agent v0.146.1): `apply` takes ONLY the root-owned copy felhom-os-apply writes after it has verified the
# operator's signature and hashed exactly those bytes (mode agent_update). The agent cannot call `apply` any more (it
# left the sudoers), and the agent's own staging dir is no longer accepted: a file in a directory the agent owns can be
# swapped between this script's sha check and its copy.
ROOT_STAGING=/var/lib/felhom-os-apply/agent-update
PENDING=$STAGING/pending.json
UNIT=felhom-agent.service
@@ -42,11 +47,14 @@ apply)
log "refusing apply: usage: apply <staged> <sha256>"
exit 2
fi
# Root-side path confinement: the staged binary MUST live in the agent's staging dir.
# Root-side path confinement: the staged binary MUST be felhom-os-apply's root-owned copy (R-861).
case "$staged" in
"$STAGING"/*) ;;
*) log "refusing apply: staged path outside $STAGING: $staged"; exit 1 ;;
"$ROOT_STAGING"/*) ;;
*) log "refusing apply: staged path outside $ROOT_STAGING: $staged"; exit 1 ;;
esac
if [ -L "$staged" ] || [ "$(stat -c %u "$staged" 2>/dev/null)" != "0" ]; then
log "refusing apply: $staged is a symlink or not root-owned"; exit 1
fi
case "$staged" in
*..*) log "refusing apply: staged path contains '..'"; exit 1 ;;
esac
+13
View File
@@ -0,0 +1,13 @@
[Unit]
Description=Felhom stable drive parent (shared bind for live drive hot-swap)
After=local-fs.target
Before=pve-guests.service
ConditionPathExists=/usr/local/sbin/felhom-shared-parent.sh
[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=/usr/local/sbin/felhom-shared-parent.sh
[Install]
WantedBy=pve-guests.service multi-user.target
+15
View File
@@ -0,0 +1,15 @@
#!/bin/sh
# felhom stable drive parent: a SHARED bind so the agent can swap backing drives underneath it and the
# guest sees the change live (no restart). MUST run before pve-guests so the guest's parent bind inherits
# the shared peer group (slave). Installed + enabled by felhom-agent. Idempotent.
set -e
mkdir -p /mnt/felhom-drives
# Isolate + share ONLY when first creating the self-bind (a fresh boot). The self-bind inherits the root
# mount's shared peer group, so make-private detaches it (else binds under it DOUBLE via the root peer),
# then make-shared gives it its own group whose only slave is the guest's parent bind. Re-running this on
# an existing parent would churn the peer-group id and orphan the guest's slave — so guard on mountpoint.
if ! mountpoint -q /mnt/felhom-drives; then
mount --bind /mnt/felhom-drives /mnt/felhom-drives
mount --make-private /mnt/felhom-drives
mount --make-shared /mnt/felhom-drives
fi
+736
View File
@@ -0,0 +1,736 @@
#!/usr/bin/env python3
"""Tests for the config bundle (R-840, `11` §5.4.2): felhom-os-apply's `bundle` mode and `--install-bundle`, and
scripts/build-config-bundle.py. An in-memory host plays the files; nothing real is written or run. Each refusal has a
test; each test names the rule it pins. Red-proof: `audits/r840-config-bundle-2026-10-04/partB/redproof.txt`.
Run: python3 configs/test_felhom_config_bundle.py (also run by internal/osupdate's Go test)
"""
import sys
sys.dont_write_bytecode = True # importing the builder must not leave scripts/__pycache__ behind
import base64
import contextlib
import hashlib
import importlib.machinery
import importlib.util
import io
import json
import os
import pathlib
import re
import stat as statmod
import unittest
HERE = pathlib.Path(__file__).resolve().parent
REPO = HERE.parent
_loader = importlib.machinery.SourceFileLoader("osapply", os.environ.get("OSAPPLY_UNDER_TEST", str(HERE / "felhom-os-apply")))
_spec = importlib.util.spec_from_loader("osapply", _loader)
osapply = importlib.util.module_from_spec(_spec)
_loader.exec_module(osapply)
_bl = importlib.machinery.SourceFileLoader("bundlebuild", str(REPO / "scripts" / "build-config-bundle.py"))
_bs = importlib.util.spec_from_loader("bundlebuild", _bl)
builder = importlib.util.module_from_spec(_bs)
_bl.exec_module(builder)
PLAN = "/var/lib/felhom-agent/os/plan-b1.json"
BUNDLE = "/var/lib/felhom-agent/os/bundle-0.143.0.json"
HOST = "demo-hp-bb76ea"
INSTALLER = REPO.parent / "felhom.eu" / "scripts" / "felhom-host-install.sh"
class St:
def __init__(self, mode, uid):
self.st_mode, self.st_uid, self.st_size = mode, uid, 100
class Box:
"""An in-memory host. files: path -> bytes; modes/uids per path; dirs: a set."""
def __init__(self, bundle_bytes, signed, oob=False, signers=True):
self.files, self.modes, self.uids = {}, {}, {}
self.dirs = {osapply.OOB_DIR} if oob else set()
self.put(osapply.TRUST_FILE, json.dumps({"host_id": HOST, "ring0_slow_lane": False}).encode(), 0o644)
if signers:
self.put(osapply.TRUST_SIGNERS, b'felhom-op-1 namespaces="felhom-op-v1" ssh-ed25519 AAAA felhom-op-1\n', 0o644)
self.put("/proc/sys/kernel/panic", b"0\n", 0o644)
self.plan = {"release_id": "bundle-0.143.0", "layer": "host", "mode": "bundle", "bundle": BUNDLE, "signed": signed}
self.put(PLAN, json.dumps(self.plan).encode(), 0o600, uid=999)
self.put(BUNDLE, bundle_bytes, 0o600, uid=999)
self.sig_rc, self.nonces, self.clock = 0, {}, 1791115200.0 # 2026-10-04T12:00:00Z
self.calls, self.logs, self.writes = [], [], []
self.visudo_fail = False # the WHOLE sudoers (`visudo -c`) after install
self.sudo_l = " (root) NOPASSWD: /usr/local/sbin/felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json\n"
self.guard = {"armed": True, "kernel_panic": 10}
def put(self, p, data, mode, uid=0):
self.files[p], self.modes[p], self.uids[p] = data, mode, uid
# Runner interface
def now(self):
return self.clock
def log(self, line):
self.logs.append(line)
def agent_uid(self):
return 999
def verify_sig(self, signers, key_id, ns, blob, sig):
self.verified = (signers, key_id, ns)
return self.sig_rc
def read_nonces(self):
return dict(self.nonces)
def write_nonces(self, d):
self.nonces = dict(d)
def read_file(self, p):
if p not in self.files:
raise OSError("no such file")
return self.files[p].decode()
def read_bytes(self, p):
if p not in self.files:
raise OSError("no such file")
return self.files[p]
def read_staged_once(self, p, owner_uid, limit):
if p not in self.files:
raise OSError("no such file")
if self.uids[p] != owner_uid:
raise osapply.Refused("R19", f"{p} is not a regular file owned by felhom-agent")
self.staged_reads = getattr(self, "staged_reads", 0) + 1
return self.files[p]
def stat(self, p):
if p not in self.files:
raise OSError("no such file")
return St(statmod.S_IFREG | self.modes[p], self.uids[p])
def lexists(self, p):
return p in self.files
def isdir(self, p):
return p in self.dirs
def put_file(self, p, data, mode):
self.writes.append(p)
self.put(p, data, mode)
def remove(self, p):
self.writes.append("rm " + p)
del self.files[p]
def list_dir(self, p):
return sorted({k[len(p) + 1:].split("/")[0] for k in self.files if k.startswith(p + "/")})
def rmtree(self, p):
for k in [k for k in self.files if k.startswith(p + "/")]:
del self.files[k]
def check_content(self, kind, data):
return (1, f"{kind}: syntax error") if b"BROKEN-SYNTAX" in data else (0, "")
def host(self, argv, timeout=600, stdin=None):
self.calls.append(argv)
if argv[:2] == ["visudo", "-c"]:
return (1, "", "parse error") if self.visudo_fail else (0, "ok", "")
if argv[:2] == ["sudo", "-n"]:
return 0, self.sudo_l, ""
if argv[-2:] == ["/usr/local/sbin/felhom-priv-apply", "--self-check"]:
body = self.files.get("/usr/local/sbin/felhom-priv-apply", b"")
return (0, "felhom-priv-apply ok verbs=unit\n", "") if b"VERBS" in body else (1, "", "boom")
if argv[-1] == "--self-check":
body = self.files.get("/usr/local/sbin/felhom-os-apply", b"")
return (0, "felhom-os-apply ok bundle-format=1 files=22\n", "") if b"BUNDLE_OP" in body else (1, "", "boom")
if argv[-1] == "/usr/local/sbin/felhom-selfupdate-guarded":
return 2, "", "felhom-selfupdate-guarded: usage: ...\n"
if argv[-1] == "status" and argv[0].endswith("felhom-crash-guard"):
return 0, json.dumps(self.guard), ""
if argv[:3] == ["systemctl", "enable", "--now"] and "felhom-crash-guard.service" in argv:
self.put("/proc/sys/kernel/panic", f"{self.guard['kernel_panic']}\n".encode(), 0o644)
return 0, "", ""
def signed_job(sha, version="0.143.0", host=HOST, op="agent_config_update", nonce="b1",
issued="2026-10-04T11:50:00Z", expires="2026-10-04T12:30:00Z"):
blob = json.dumps({"expires_at": expires, "issued_at": issued, "key_id": "felhom-op-1", "nonce": nonce, "op": op,
"params": {"agent_version": version, "bundle_sha256": sha},
"target": {"guest_id": "", "host_id": host}}, sort_keys=True).encode()
return {"blob_b64": base64.b64encode(blob).decode(), "sig": "-----BEGIN SSH SIGNATURE-----\nx\n-----END SSH SIGNATURE-----\n"}
def real_bundle(version="0.143.0"):
data = builder.build(version)
return data, hashlib.sha256(data).hexdigest()
def edited_bundle(edit):
"""The real bundle, with edit(files_list) applied and every sha recomputed — a SIGNED bundle with bad content."""
b = json.loads(builder.build("0.143.0"))
edit(b["files"])
for e in b["files"]:
e["sha256"] = hashlib.sha256(base64.b64decode(e["content_b64"])).hexdigest()
data = (json.dumps(b, indent=1, sort_keys=True) + "\n").encode()
return data, hashlib.sha256(data).hexdigest()
def replace_content(files, path, fn):
for e in files:
if e["path"] == path:
e["content_b64"] = base64.b64encode(fn(base64.b64decode(e["content_b64"]))).decode()
def run(box, argv=None, environ=None):
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
rc = osapply.main(argv or ["felhom-os-apply", "--plan", PLAN], runner=box, environ=environ or {})
line = [l for l in buf.getvalue().splitlines() if l.startswith("OSAPPLY-REPORT ")][-1]
return rc, json.loads(line[len("OSAPPLY-REPORT "):])
def box_files(box):
return {p: box.files[p] for p in box.files if p in osapply.BUNDLE_DESTS}
class Install(unittest.TestCase):
def test_fresh_box_gets_every_file_and_a_record(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
b = rep["bundle"]
# every path but the four OOB ones (no belt on this box)
self.assertEqual(len(b["written"]), len(osapply.BUNDLE_FILES) - 4, b)
self.assertEqual(len(b["skipped"]), 4)
self.assertEqual(box.files["/etc/sudoers.d/felhom-agent"], (REPO / "configs" / "felhom-agent.sudoers").read_bytes())
self.assertEqual(box.modes["/etc/sudoers.d/felhom-agent"], 0o440)
rec = json.loads(box.files[osapply.BUNDLE_RECORD])
self.assertEqual((rec["agent_version"], rec["bundle_sha256"], rec["authority"]), ("0.143.0", sha, "signed"))
self.assertIn("b1", box.nonces, "the job is consumed")
self.assertEqual(b["self_check"]["crash_guard"], {"armed": True, "kernel_panic": 10})
self.assertIn(["systemctl", "daemon-reload"], box.calls)
def test_sudoers_is_written_after_every_wrapper(self):
"""Order: a referenced wrapper is in place before the sudoers line that allows it."""
data, sha = edited_bundle(lambda f: f.reverse()) # the bundle lists the sudoers FIRST; the wrapper must reorder
box = Box(data, signed_job(sha))
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
w = [p for p in box.writes if p in osapply.BUNDLE_DESTS]
self.assertEqual(w[-1], "/etc/sudoers.d/felhom-agent")
self.assertLess(w.index("/usr/local/sbin/felhom-os-apply"), w.index("/etc/sudoers.d/felhom-agent"))
def test_identical_box_writes_nothing(self):
"""The demo boxes' case: hand-copied files equal to the release → 0 written, all 'same'."""
data, sha = real_bundle()
box = Box(data, signed_job(sha))
for dest, src, mode, _, policy in osapply.BUNDLE_FILES:
if policy != "oob":
box.put(dest, (REPO / "configs" / src).read_bytes(), mode)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["bundle"]["written"], [])
self.assertEqual(rep["bundle"]["same"], len(osapply.BUNDLE_FILES) - 5) # 4 oob skipped + crash-guard.conf kept
self.assertEqual(rep["bundle"]["kept"], ["/etc/felhom/crash-guard.conf"])
def test_a_wrong_mode_is_rewritten(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put("/etc/sudoers.d/felhom-agent", (REPO / "configs" / "felhom-agent.sudoers").read_bytes(), 0o644)
rc, rep = run(box)
self.assertIn("/etc/sudoers.d/felhom-agent", rep["bundle"]["written"])
self.assertEqual(box.modes["/etc/sudoers.d/felhom-agent"], 0o440)
def test_tuned_crash_guard_conf_is_kept(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put("/etc/felhom/crash-guard.conf", b"LIMIT=5\n", 0o644)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(box.files["/etc/felhom/crash-guard.conf"], b"LIMIT=5\n")
def test_oob_files_only_on_a_box_with_the_belt(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha), oob=True)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertIn("/etc/sudoers.d/felhom-op", rep["bundle"]["written"])
self.assertEqual(rep["bundle"]["skipped"], [])
def test_previous_copies_are_kept(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put("/usr/local/sbin/felhom-pbs-apply", b"#!/bin/bash\necho old\n", 0o755)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(box.files[rep["bundle"]["prev_dir"] + "/usr/local/sbin/felhom-pbs-apply"], b"#!/bin/bash\necho old\n")
class Refusals(unittest.TestCase):
"""Each: refused, and NOTHING on the box changed."""
def refused(self, box, code):
before = dict(box_files(box))
rc, rep = run(box)
self.assertEqual(rc, 2, rep)
self.assertEqual(rep["refused"]["code"], code, rep)
self.assertEqual(box_files(box), before, "a refusal changed a file")
self.assertNotIn(osapply.BUNDLE_RECORD, box.files)
return rep
def test_wrong_sha_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job("0" * 64))
self.refused(box, "R18")
self.assertEqual(box.nonces, {}, "a wrong sha must not burn the job")
def test_bad_signature_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.sig_rc = 255
self.refused(box, "R3")
def test_other_op_is_refused(self):
data, sha = real_bundle()
self.refused(Box(data, signed_job(sha, op="agent_update")), "R3")
def test_other_host_is_refused(self):
data, sha = real_bundle()
self.refused(Box(data, signed_job(sha, host="demo-felhom-8363b5")), "R3")
def test_expired_job_is_refused(self):
data, sha = real_bundle()
self.refused(Box(data, signed_job(sha, expires="2026-10-04T11:55:00Z")), "R3")
def test_replayed_job_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.nonces = {"b1": box.clock + 600}
self.refused(box, "R3")
def test_version_mismatch_is_refused(self):
data, sha = real_bundle()
self.refused(Box(data, signed_job(sha, version="0.142.1")), "R18")
def test_sudoers_failing_visudo_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/etc/sudoers.d/felhom-agent", lambda c: c + b"BROKEN-SYNTAX\n"))
self.refused(Box(data, signed_job(sha)), "R18")
def test_sudoers_dropping_the_route_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/etc/sudoers.d/felhom-agent",
lambda c: c.replace(b"/usr/local/sbin/felhom-os-apply --plan", b"/bin/true --plan")))
self.refused(Box(data, signed_job(sha)), "R18")
def test_wrapper_without_bundle_mode_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/usr/local/sbin/felhom-os-apply",
lambda c: c.replace(b'BUNDLE_OP = "agent_config_update"', b'BUNDLE_OPX = 1')))
self.refused(Box(data, signed_job(sha)), "R18")
def test_python_syntax_error_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/usr/local/sbin/felhom-crash-guard", lambda c: c + b"\ndef (\n"))
self.refused(Box(data, signed_job(sha)), "R18")
def test_unit_with_runtime_directory_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/etc/systemd/system/felhom-mgmt-watchdog.service",
lambda c: c + b"RuntimeDirectory=sshd\n"))
self.refused(Box(data, signed_job(sha)), "R18")
def test_agent_unit_not_as_the_agent_user_is_refused(self):
data, sha = edited_bundle(lambda f: replace_content(f, "/etc/systemd/system/felhom-agent.service",
lambda c: c.replace(b"User=felhom-agent", b"User=root")))
self.refused(Box(data, signed_job(sha)), "R18")
def test_a_bundle_that_changes_a_signer_is_refused(self):
"""R17: the trust root is not a bundle's to change — not even a signed one."""
def add(f):
f.append({"path": osapply.TRUST_SIGNERS, "content_b64": base64.b64encode(b"evil-key\n").decode()})
data, sha = edited_bundle(add)
box = Box(data, signed_job(sha))
self.refused(box, "R17")
self.assertIn(b"felhom-op-1", box.files[osapply.TRUST_SIGNERS])
def test_a_path_outside_the_table_is_refused(self):
def add(f):
f.append({"path": "/etc/shadow", "content_b64": base64.b64encode(b"root::0:0\n").decode()})
data, sha = edited_bundle(add)
self.refused(Box(data, signed_job(sha)), "R16")
def test_content_not_matching_its_sha_is_refused(self):
b = json.loads(builder.build("0.143.0"))
b["files"][0]["content_b64"] = base64.b64encode(b"#!/bin/bash\nexit 0\n").decode()
data = json.dumps(b).encode()
self.refused(Box(data, signed_job(hashlib.sha256(data).hexdigest())), "R18")
def test_bundle_outside_the_plan_dir_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.plan["bundle"] = "/tmp/bundle-0.143.0.json"
box.files[PLAN] = json.dumps(box.plan).encode()
self.refused(box, "R1")
def test_bundle_not_owned_by_the_agent_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.uids[BUNDLE] = 0
self.refused(box, "R1")
def test_no_trust_file_is_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
del box.files[osapply.TRUST_FILE]
self.refused(box, "R3")
class SelfCheckUndo(unittest.TestCase):
def assert_restored(self, box, before, rep):
self.assertTrue(rep["bundle"]["rolled_back"], rep)
self.assertEqual(box_files(box), before, "the previous files must be back, byte for byte")
self.assertNotIn(osapply.BUNDLE_RECORD, box.files)
def with_old_files(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put("/usr/local/sbin/felhom-pbs-apply", b"#!/bin/bash\necho old\n", 0o755)
box.put("/etc/sudoers.d/felhom-agent", b"# old sudoers\n", 0o440)
return box
def test_route_missing_after_install_puts_everything_back(self):
box = self.with_old_files()
before = dict(box_files(box))
box.sudo_l = " (root) NOPASSWD: /bin/true\n"
rc, rep = run(box)
self.assertEqual(rc, 3, rep)
self.assert_restored(box, before, rep)
self.assertEqual(box.calls[-1], ["visudo", "-c"], "the undo re-checks the whole sudoers")
def test_visudo_failing_after_install_puts_everything_back(self):
box = self.with_old_files()
before = dict(box_files(box))
box.visudo_fail = True
rc, rep = run(box)
self.assertEqual(rc, 3, rep)
self.assert_restored(box, before, rep)
def test_crash_guard_disagreeing_with_kernel_panic_puts_everything_back(self):
box = self.with_old_files()
before = dict(box_files(box))
box.guard = {"armed": True, "kernel_panic": 10}
box.host_orig = box.host
def host(argv, timeout=600, stdin=None):
if argv[:3] == ["systemctl", "enable", "--now"]:
box.calls.append(argv)
return 0, "", "" # the unit "started" but kernel.panic stayed 0
return box.host_orig(argv, timeout, stdin)
box.host = host
rc, rep = run(box)
self.assertEqual(rc, 3, rep)
self.assert_restored(box, before, rep)
class TrustBootstrap(unittest.TestCase):
def test_missing_signers_verifies_against_the_pinned_key_and_creates_it(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha), signers=False)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(box.verified[0], osapply.PINNED_SIGNERS, "verified against the pinned key, nothing else")
self.assertTrue(rep["bundle"]["signers_created"])
self.assertEqual(box.files[osapply.TRUST_SIGNERS], osapply.pinned_signers_line().encode())
self.assertEqual(box.modes[osapply.TRUST_SIGNERS], 0o644)
def test_missing_signers_and_a_job_the_pinned_key_did_not_sign(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha), signers=False)
box.sig_rc = 255
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R3"))
self.assertNotIn(osapply.TRUST_SIGNERS, box.files)
def test_present_signers_are_never_touched(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.put(osapply.TRUST_SIGNERS, b'felhom-op-2 namespaces="felhom-op-v1" ssh-ed25519 BBBB felhom-op-2\n', 0o644)
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
self.assertEqual(box.verified[0], osapply.TRUST_SIGNERS)
self.assertFalse(rep["bundle"]["signers_created"])
self.assertIn(b"felhom-op-2", box.files[osapply.TRUST_SIGNERS])
def test_agent_writable_signers_are_refused(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
box.uids[osapply.TRUST_SIGNERS] = 999
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R3"))
def test_pinned_operator_key_equals_the_installers(self):
"""The bootstrap key is exactly the one felhom-host-install.sh pins (OPERATOR_KEY_OPERATIONAL_*)."""
text = INSTALLER.read_text()
kid = re.search(r'^OPERATOR_KEY_OPERATIONAL_ID="([^"]+)"', text, re.M).group(1)
line = re.search(r'^OPERATOR_KEY_OPERATIONAL_LINE="([^"]+)"', text, re.M).group(1)
self.assertEqual((osapply.PINNED_OPERATOR_KEY_ID, osapply.PINNED_OPERATOR_KEY_LINE), (kid, line))
# and the file format is the installer's printf, byte for byte
self.assertIn("printf '%s namespaces=\"felhom-op-v1\" %s\\n'", text)
class InstallerEntry(unittest.TestCase):
def test_installer_installs_without_a_signature(self):
data, sha = real_bundle()
box = Box(data, None)
box.put("/root/bundle.json", data, 0o600)
rc, rep = run(box, ["felhom-os-apply", "--install-bundle", "/root/bundle.json", "--sha256", sha])
self.assertEqual(rc, 0, rep)
self.assertEqual(json.loads(box.files[osapply.BUNDLE_RECORD])["authority"], "installer")
def test_installer_entry_checks_the_sha(self):
data, sha = real_bundle()
box = Box(data, None)
box.put("/root/bundle.json", data, 0o600)
rc, rep = run(box, ["felhom-os-apply", "--install-bundle", "/root/bundle.json", "--sha256", "1" * 64])
self.assertEqual((rc, rep["refused"]["code"]), (2, "R18"))
def test_installer_entry_is_refused_through_sudo(self):
data, sha = real_bundle()
box = Box(data, None)
box.put("/root/bundle.json", data, 0o600)
rc, rep = run(box, ["felhom-os-apply", "--install-bundle", "/root/bundle.json", "--sha256", sha], {"SUDO_UID": "999"})
self.assertEqual((rc, rep["refused"]["code"]), (2, "R1"))
self.assertNotIn(osapply.BUNDLE_RECORD, box.files)
def test_self_check_answers(self):
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
rc = osapply.main(["felhom-os-apply", "--self-check"], runner=Box(b"", None))
self.assertEqual(rc, 0)
self.assertIn("bundle-format=1", buf.getvalue())
class Facts(unittest.TestCase):
def test_state_reports_drift_against_the_record(self):
data, sha = real_bundle()
box = Box(data, signed_job(sha))
run(box)
box.put("/usr/local/sbin/felhom-pbs-apply", b"#!/bin/bash\necho by hand\n", 0o755)
a = osapply.Apply(box, PLAN)
st = osapply.Bundle(a).state()
self.assertEqual(st["version"], "0.143.0")
self.assertEqual(st["drift"], ["/usr/local/sbin/felhom-pbs-apply"])
self.assertTrue(st["signers_present"])
def test_state_without_a_record_says_none(self):
box = Box(b"", None)
st = osapply.Bundle(osapply.Apply(box, PLAN)).state()
self.assertEqual(st["version"], "none")
self.assertNotIn("drift", st)
self.assertEqual(st["live"]["/etc/sudoers.d/felhom-agent"], "absent")
class Builder(unittest.TestCase):
def test_reproducible(self):
self.assertEqual(builder.build("0.143.0"), builder.build("0.143.0"))
def test_every_source_exists_and_every_dest_is_unique(self):
dests = [e[0] for e in osapply.BUNDLE_FILES]
self.assertEqual(len(dests), len(set(dests)))
for _, src, *_ in osapply.BUNDLE_FILES:
self.assertTrue((REPO / "configs" / src).is_file(), src)
def test_every_root_file_the_installer_writes_is_in_the_bundle(self):
"""One source of truth: a felhom root-owned path the installer names must be a bundle path, a trust file, or a
path the AGENT itself writes at run time (named here, with why). Scope is a regex over the WHOLE installer."""
text = INSTALLER.read_text()
found = set(re.findall(r"(/usr/local/sbin/felhom-[a-z-]+|/etc/systemd/system/felhom-[a-z.-]+|"
r"/etc/sudoers\.d/felhom-[a-z-]+|/etc/felhom-oob\.nft|/etc/tmpfiles\.d/felhom-[a-z.-]+|"
r"/etc/felhom/[a-z.-]+)", text))
agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"}
trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD}
# Written by the appliance ISO's first boot (felhom.eu scripts/iso/felhom-bootstrap.sh), never by the installer:
# since installer 1.32.0 (R-275) the uninstall only NAMES them under KEPT.
iso_writes = {"/etc/felhom/.bootstrap-done", "/etc/felhom/appliance-pairing-code"}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust | iso_writes)
self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them")
# the limits drop-in is named through $AGENT_UNIT in the installer
self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS)
# ---------- R-861 (agent v0.146.0): the agent binary only by an operator-signed agent_update, checked as root ----------
STAGED = "/var/lib/felhom-agent/selfupdate/felhom-agent-0.146.0"
NEW_BIN = b"\x7fELF the new agent"
def update_job(sha, version="0.146.0", op="agent_update", nonce="u1", host=HOST):
blob = json.dumps({"expires_at": "2026-10-04T12:30:00Z", "issued_at": "2026-10-04T11:50:00Z", "key_id": "felhom-op-1",
"nonce": nonce, "op": op, "params": {"sha256": sha, "version": version},
"target": {"guest_id": "", "host_id": host}}, sort_keys=True).encode()
return {"blob_b64": base64.b64encode(blob).decode(), "sig": "-----BEGIN SSH SIGNATURE-----\nx\n-----END SSH SIGNATURE-----\n"}
def update_box(job, staged=STAGED, content=NEW_BIN):
box = Box(b"{}", None)
box.put(PLAN, json.dumps({"release_id": "agent-0.146.0", "layer": "host", "mode": "agent_update",
"signed": job, "staged": staged}).encode(), 0o600, uid=999)
box.put(STAGED, content, 0o755, uid=999)
return box
def wrapper_calls(box):
return [c for c in box.calls if c and c[0] == osapply.SELFUPDATE_WRAPPER and c[1:2] == ["apply"]]
class AgentUpdate(unittest.TestCase):
"""RED-PROOF: make agent_update skip verify_signed → test_a_bad_signature_never_reaches_the_wrapper fails."""
def test_signed_update_flips_and_burns_the_nonce(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha))
rc, rep = run(box)
self.assertEqual(rc, 0, rep)
root_copy = osapply.SELFUPDATE_ROOT_DIR + "/felhom-agent-0.146.0"
# the wrapper gets the ROOT-OWNED copy of the bytes that were hashed — never the agent's path (review 2026-10-05)
self.assertEqual(wrapper_calls(box), [[osapply.SELFUPDATE_WRAPPER, "apply", root_copy, sha]])
self.assertIn(root_copy, box.writes)
self.assertNotIn(root_copy, box.files, "the root copy is removed after the flip")
self.assertEqual(box.staged_reads, 1, "the agent's file is read exactly once")
self.assertIn("u1", box.nonces)
self.assertEqual(rep["agent_update"]["version"], "0.146.0")
def test_a_staged_file_the_agent_does_not_own_is_refused(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha))
box.uids[STAGED] = 0
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R19"), rep)
self.assertEqual(wrapper_calls(box), [])
def test_a_bad_signature_never_reaches_the_wrapper(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha))
box.sig_rc = 1
rc, rep = run(box)
self.assertTrue(rep.get("refused"), f"a job whose signature does not verify was NOT refused: {rep}")
self.assertEqual((rc, rep["refused"]["code"]), (2, "R3"), rep)
self.assertEqual(wrapper_calls(box), [])
def test_a_staged_binary_the_signature_does_not_pin_is_refused(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha), content=b"\x7fELF something the agent put there")
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R19"), rep)
self.assertEqual(wrapper_calls(box), [])
self.assertNotIn("u1", box.nonces, "a refused job keeps its nonce (the operator fixes the file, not the key)")
def test_another_staging_path_is_refused(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha), staged="/tmp/felhom-agent-0.146.0")
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R19"), rep)
self.assertEqual(wrapper_calls(box), [])
def test_a_bundle_job_is_not_an_agent_update(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha, op="agent_config_update"))
rc, rep = run(box)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R3"), rep)
def test_another_hosts_job_and_a_replay_are_refused(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha, host="tester-1-d70be4"))
self.assertEqual(run(box)[1]["refused"]["code"], "R3")
box2 = update_box(update_job(sha))
box2.nonces["u1"] = 1999999999
self.assertEqual(run(box2)[1]["refused"]["code"], "R3")
self.assertEqual(wrapper_calls(box) + wrapper_calls(box2), [])
def test_a_failed_flip_keeps_the_nonce(self):
sha = hashlib.sha256(NEW_BIN).hexdigest()
box = update_box(update_job(sha))
orig = box.host
box.host = lambda argv, timeout=600, stdin=None: (1, "", "refusing apply: sha mismatch") if argv[:1] == [osapply.SELFUPDATE_WRAPPER] else orig(argv, timeout, stdin)
rc, rep = run(box)
self.assertEqual(rc, 3, rep)
self.assertNotIn("u1", box.nonces)
class SelfupdateWrapperConfinement(unittest.TestCase):
"""R-861 (v0.146.1): the A/B wrapper takes only felhom-os-apply's root-owned copy, never the agent's staging dir
(a file there can be swapped between the wrapper's sha check and its copy). The path check runs before anything
is touched, so the real script can be run here unprivileged.
RED-PROOF: point ROOT_STAGING back at /var/lib/felhom-agent/selfupdate → this fails (the path is accepted and the
script goes on to `staged file missing`)."""
def test_the_agents_staging_dir_is_refused(self):
import subprocess
sha = "0" * 64
p = subprocess.run(["sh", str(HERE / "felhom-selfupdate-guarded"), "apply",
"/var/lib/felhom-agent/selfupdate/felhom-agent-0.146.1", sha], capture_output=True, text=True)
self.assertEqual(p.returncode, 1, p.stderr)
self.assertIn("outside /var/lib/felhom-os-apply/agent-update", p.stderr)
_sb = importlib.machinery.SourceFileLoader("stepbuild", str(REPO / "scripts" / "build-step-bundle.py"))
_ss = importlib.util.spec_from_loader("stepbuild", _sb)
stepbuild = importlib.util.module_from_spec(_ss)
_sb.exec_module(stepbuild)
NEW_IN_0146 = {"/usr/local/sbin/felhom-priv-apply", "/var/lib/vz/snippets/felhom-guest-hook.sh",
"/usr/local/sbin/felhom-shared-parent.sh", "/etc/systemd/system/felhom-shared-parent.service"}
class StepBundle(unittest.TestCase):
"""R-880 (agent v0.146.1): an INSTALLED wrapper checks an incoming bundle's paths against its OWN table (R16), so a
release that adds paths needs a step bundle: the boxes' current bundle with only felhom-os-apply replaced.
RED-PROOF: deliver the full bundle to the old table → R16 (test_the_full_bundle_is_refused_by_an_old_table)."""
def old_world(self):
"""The base bundle an older wrapper (no R-861 paths) installed, and that wrapper's table."""
full = json.loads(builder.build("0.145.0"))
full["files"] = [e for e in full["files"] if e["path"] not in NEW_IN_0146]
old_wrapper = b'# the v0.145.0 wrapper stands in here\nBUNDLE_OP = "agent_config_update"\n'
for e in full["files"]:
if e["path"] == "/usr/local/sbin/felhom-os-apply":
e["content_b64"], e["sha256"] = base64.b64encode(old_wrapper).decode(), hashlib.sha256(old_wrapper).hexdigest()
base = (json.dumps(full, indent=1, sort_keys=True) + "\n").encode()
old_dests = {k: v for k, v in osapply.BUNDLE_DESTS.items() if k not in NEW_IN_0146}
return base, old_dests
def parse_with_table(self, data, dests, version):
saved = osapply.BUNDLE_DESTS
osapply.BUNDLE_DESTS = dests
try:
return osapply.Bundle(osapply.Apply(Box(b"{}", None), "")).parse(data, hashlib.sha256(data).hexdigest(), version)
finally:
osapply.BUNDLE_DESTS = saved
def test_the_full_bundle_is_refused_by_an_old_table(self):
_, old_dests = self.old_world()
full = builder.build("0.146.1")
with self.assertRaises(osapply.Refused) as cm:
self.parse_with_table(full, old_dests, "0.146.1")
self.assertEqual(cm.exception.code, "R16")
def test_the_step_bundle_is_accepted_by_the_old_table_and_changes_only_the_wrapper(self):
base, old_dests = self.old_world()
new_wrapper = (HERE / "felhom-os-apply").read_bytes()
step = stepbuild.build_step(base, "0.146.1-step1", new_wrapper)
ver, files = self.parse_with_table(step, old_dests, "0.146.1-step1")
self.assertEqual(ver, "0.146.1-step1")
b, s_ = json.loads(base), json.loads(step)
self.assertEqual(sorted(e["path"] for e in b["files"]), sorted(e["path"] for e in s_["files"]), "the paths must not change")
changed = [e["path"] for e, f in zip(sorted(b["files"], key=lambda x: x["path"]), sorted(s_["files"], key=lambda x: x["path"]))
if e != f]
self.assertEqual(changed, ["/usr/local/sbin/felhom-os-apply"], "exactly the wrapper changes")
installed = dict((d, c) for d, c, *_ in files)
self.assertEqual(installed["/usr/local/sbin/felhom-os-apply"], new_wrapper)
# and the NEW wrapper (now installed) knows every path the release's full bundle names
self.assertTrue({e["path"] for e in json.loads(builder.build("0.146.1"))["files"]} <= set(osapply.BUNDLE_DESTS))
def test_a_step_version_must_carry_a_suffix(self):
base, _ = self.old_world()
with self.assertRaises(SystemExit):
stepbuild.build_step(base, "0.146.1", b"x")
if __name__ == "__main__":
unittest.main()
+200
View File
@@ -0,0 +1,200 @@
#!/usr/bin/python3
"""Tests for felhom-crash-guard (`11` §5.9). Temp dirs only; nothing real is touched. Red-proof seam: CRASHGUARD_UNDER_TEST."""
import importlib.machinery
import importlib.util
import json
import os
import pathlib
import tempfile
import unittest
HERE = pathlib.Path(__file__).resolve().parent
_loader = importlib.machinery.SourceFileLoader("crashguard", os.environ.get("CRASHGUARD_UNDER_TEST", str(HERE / "felhom-crash-guard")))
_spec = importlib.util.spec_from_loader("crashguard", _loader)
cg = importlib.util.module_from_spec(_spec)
_loader.exec_module(cg)
T0 = 1791115200.0 # 2026-10-04T12:00:00Z
class FakeEnv(cg.Env):
def __init__(self, d):
super().__init__(conf=os.path.join(d, "conf"), state_dir=os.path.join(d, "state"),
panic_path=os.path.join(d, "panic"), uptime_path=os.path.join(d, "uptime"),
boot_id_path=os.path.join(d, "bootid"))
self.t = T0
self.logs = []
open(self.panic_path, "w").write("0\n")
open(self.uptime_path, "w").write("20.00 10.00\n")
def now(self):
return self.t
def log(self, line):
self.logs.append(line)
def panic(self):
return int(open(self.panic_path).read())
def state(self):
return json.load(open(os.path.join(self.state_dir, "state.json")))
class Guard(unittest.TestCase):
def setUp(self):
self.d = tempfile.TemporaryDirectory()
self.e = FakeEnv(self.d.name)
def tearDown(self):
self.d.cleanup()
def crash_boot(self, minutes_later):
self.e.t += minutes_later * 60
cg.main(["x", "boot"], self.e) # no clean-stop before it: an unclean stop
def clean_reboot(self, minutes_later):
cg.main(["x", "clean-stop"], self.e)
self.e.t += minutes_later * 60
cg.main(["x", "boot"], self.e)
def test_first_boot_is_not_a_crash_and_arms(self):
cg.main(["x", "boot"], self.e)
s = self.e.state()
self.assertFalse(s["last_boot_unclean"])
self.assertEqual(self.e.panic(), 10)
self.assertTrue(s["armed"])
def test_clean_reboots_never_count(self):
cg.main(["x", "boot"], self.e)
for _ in range(5):
self.clean_reboot(1)
s = self.e.state()
self.assertEqual(s["unclean_boots_in_window"], 0)
self.assertEqual(self.e.panic(), 10)
def test_third_crash_in_an_hour_leaves_the_box_off(self):
# operator's words: "if it crashes 3 times within one hour, it stays off" — after crash 2 the guard trips,
# so crash 3 (kernel.panic = 0) does not restart the box.
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.assertEqual(self.e.panic(), 10, "one crash: still restarts")
self.crash_boot(5)
s = self.e.state()
self.assertTrue(s["tripped"], s)
self.assertEqual(self.e.panic(), 0, "after the 2nd crash boot the 3rd crash must leave the box off")
self.assertIn("2 unclean boots within 60 minutes", s["tripped_reason"])
def test_crashes_spread_over_more_than_the_window_do_not_trip(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(61)
self.assertFalse(self.e.state()["tripped"])
self.assertEqual(self.e.panic(), 10)
def test_tripped_stays_tripped_across_boots(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(5)
self.clean_reboot(30) # the operator switched it on; even a clean boot keeps the trip
self.assertTrue(self.e.state()["tripped"])
self.assertEqual(self.e.panic(), 0)
def test_rearms_after_24h_of_normal_running(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(5)
self.e.t += 23 * 3600
cg.main(["x", "check"], self.e)
self.assertTrue(self.e.state()["tripped"], "not before 24 h")
self.e.t += 3600
cg.main(["x", "check"], self.e)
s = self.e.state()
self.assertFalse(s["tripped"])
self.assertEqual(self.e.panic(), 10)
self.assertIn("timer", s["rearmed_by"])
def test_operator_rearm_starts_a_fresh_window(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(5)
self.e.t += 60
cg.main(["x", "rearm"], self.e)
s = self.e.state()
self.assertFalse(s["tripped"])
self.assertEqual(s["rearmed_by"], "operator")
self.assertEqual(s["unclean_boots_24h"], 2, "the history stays")
self.crash_boot(5)
self.assertFalse(self.e.state()["tripped"], "one crash after a re-arm must not trip at once")
def test_config_numbers_are_read(self):
open(self.e.conf, "w").write("LIMIT=2\nPANIC_SECONDS=30\n")
cg.main(["x", "boot"], self.e)
self.assertEqual(self.e.panic(), 30)
self.crash_boot(1)
self.assertTrue(self.e.state()["tripped"], "LIMIT=2: the first crash boot trips")
def test_state_is_world_readable_for_the_agent(self):
cg.main(["x", "boot"], self.e)
mode = os.stat(os.path.join(self.e.state_dir, "state.json")).st_mode & 0o777
self.assertEqual(mode, 0o644)
class KernelStepCannotLeaveTheBoxOff(unittest.TestCase):
"""R-836 / `11` §5.11 Part B 4: a kernel step's planned reboot, one crash and one self-revert cannot add up to the
box staying off. The wrapper starts a step only when the guard is armed with NO unclean boot in its window
(felhom-os-apply Kernel.check_guard, R21); the planned reboot and the self-revert are orderly (`systemctl reboot` —
the clean-stop marker); so the step adds at most ONE unclean boot, and the box stays off only after the LIMIT-th
(3rd) within the hour. Red-proof: audits/kernel-lane-2026-10-07/A/redproof.txt."""
def setUp(self):
self.d = tempfile.TemporaryDirectory()
self.e = FakeEnv(self.d.name)
cg.main(["x", "boot"], self.e)
s = self.e.state()
self.assertTrue(s["armed"])
self.assertEqual(s["unclean_boots_in_window"], 0, "the wrapper's precondition (R21)")
def tearDown(self):
self.d.cleanup()
def planned(self, minutes):
cg.main(["x", "clean-stop"], self.e)
self.e.t += minutes * 60
cg.main(["x", "boot"], self.e)
def crash(self, minutes):
self.e.t += minutes * 60
cg.main(["x", "boot"], self.e)
def test_planned_reboot_one_crash_one_self_revert(self):
self.planned(2) # the step's one-shot reboot (orderly)
self.crash(3) # the new kernel crashes after the guard ran; the box restarts (panic=10)
self.planned(2) # the self-revert (orderly)
s = self.e.state()
self.assertFalse(s["tripped"], s)
self.assertTrue(s["armed"])
self.assertEqual(self.e.panic(), 10, "the box still restarts after a crash")
self.assertEqual(s["unclean_boots_in_window"], 1, "the step added exactly one unclean boot")
def test_a_panic_before_userspace_is_not_even_counted(self):
# the one-shot kernel panics before the guard's unit runs (measured: rdinit= and init= missing): the planned
# reboot's clean-stop marker is still there when the old kernel boots, so this boot counts as clean.
cg.main(["x", "clean-stop"], self.e)
self.e.t += 120 # the panicking boot: no userspace, the guard never ran
cg.main(["x", "boot"], self.e)
s = self.e.state()
self.assertEqual(s["unclean_boots_in_window"], 0, s)
self.assertEqual(self.e.panic(), 10)
def test_the_box_stays_off_only_after_two_more_crashes_than_the_step_makes(self):
self.planned(2)
self.crash(3) # the step's one crash
self.planned(2) # the self-revert
self.crash(5) # an UNRELATED crash within the hour: the guard trips (the 3rd would leave it off)
s = self.e.state()
self.assertTrue(s["tripped"])
self.assertEqual(s["unclean_boots_in_window"], 2, "two unclean boots: one from the step, one not")
if __name__ == "__main__":
unittest.main()
File diff suppressed because it is too large Load Diff
+380
View File
@@ -0,0 +1,380 @@
#!/usr/bin/env python3
"""Tests for felhom-priv-apply (R-861, `03` §3.1). An in-memory host plays the files; nothing real is written or run.
Each refusal rule has a test that feeds it the ATTACK it exists for; each accepted shape is a file the agent really
renders (the Go contract tests feed the live renderers too). Red-proof: audits/hub-safety-2026-10-05/partF/.
Run: python3 configs/test_felhom_priv_apply.py (also run by internal/privapply's Go test)
"""
import sys
sys.dont_write_bytecode = True
import importlib.machinery
import importlib.util
import os
import pathlib
import unittest
HERE = pathlib.Path(__file__).resolve().parent
_loader = importlib.machinery.SourceFileLoader("privapply", os.environ.get("PRIVAPPLY_UNDER_TEST", str(HERE / "felhom-priv-apply")))
_spec = importlib.util.spec_from_loader("privapply", _loader)
pa = importlib.util.module_from_spec(_spec)
_loader.exec_module(pa)
AGENT_UID = 999
class FakeHost:
def __init__(self):
self.src, self.dest, self.logs, self.writes = {}, {}, [], []
self.owner, self.kind = {}, {}
def stage(self, path, text, uid=AGENT_UID, kind="file"):
self.src[path] = text.encode() if isinstance(text, str) else text
self.owner[path], self.kind[path] = uid, kind
def agent_uid(self):
return AGENT_UID
def read_source(self, path):
# mirrors Host.read_source's refusals: absent, not a regular file (a symlink is refused by O_NOFOLLOW), owner, size
if path not in self.src:
raise pa.Refused("P1", f"cannot open the staged file {path}")
if self.kind[path] != "file":
raise pa.Refused("P1", f"{path} is not a regular file")
if self.owner[path] != AGENT_UID:
raise pa.Refused("P1", f"{path} is not owned by felhom-agent")
if len(self.src[path]) > pa.MAX_BYTES:
raise pa.Refused("P1", f"{path} is larger than {pa.MAX_BYTES} bytes")
return self.src[path]
def read_dest(self, path):
return self.dest.get(path)
def install(self, dest, data, mode):
self.writes.append((dest, mode))
self.dest[dest] = data
def log(self, line):
self.logs.append(line)
# controller-image (R-861 (a) A1): the ref arrives on stdin and is written INSIDE the guest by root.
stdin = b""
guest_writes = None
def read_stdin(self, limit):
return self.stdin[:limit + 1]
def write_guest_image(self, vmid, data):
if self.guest_writes is None:
self.guest_writes = []
self.guest_writes.append((vmid, data))
LOCAL_UNIT = """# Managed by felhom-agent — do not edit by hand.
[Unit]
Description=Felhom storage mount 91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae
After=local-fs-pre.target
[Mount]
What=/dev/disk/by-uuid/91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae
Where=/mnt/hdd_1
Type=ext4
[Install]
WantedBy=multi-user.target
"""
NET_MOUNT = """# felhom network storage — do not edit by hand.
[Unit]
Description=Felhom network storage media (nfs)
[Mount]
What=nas.lan:/volume1/media
Where=/mnt/felhom-drives/media
Type=nfs4
Options=vers=4.1,soft,timeo=50,retrans=2,noatime,_netdev,retry=0,nosuid,nodev
"""
NET_AUTOMOUNT = """# felhom network storage — do not edit by hand.
[Unit]
Description=Felhom network storage automount media (nfs)
[Automount]
Where=/mnt/felhom-drives/media
TimeoutIdleSec=600
[Install]
WantedBy=multi-user.target
"""
NET_NAME = "mnt-felhom\\x2ddrives-media.mount"
DNS_BASE = """# felhom split-horizon resolver — host base config (agent-managed; DO NOT EDIT)
bind-interfaces
listen-address=192.168.0.104
listen-address=127.0.0.1
no-resolv
server=1.1.1.1
server=9.9.9.9
"""
DNS_GUEST = """# felhom split-horizon DNS — customer demo-hp (agent-managed; DO NOT EDIT)
local=/enkisfelhom.hu/
address=/enkisfelhom.hu/192.168.0.138
"""
K = "AAECAwQFBgcICQoLDA0ODxAREhMUFRYXGBkaGxwdHh8=" # base64 of 32 bytes
WG = f"""# felhom offsite tunnel — agent-managed (S3); DO NOT EDIT
[Interface]
PrivateKey = {K}
Address = 10.77.0.3/32
MTU = 1280
[Peer]
PublicKey = {K}
Endpoint = 49.12.1.2:51820
AllowedIPs = 10.77.0.1/32, 10.77.0.250/32
PersistentKeepalive = 25
"""
KEYLINE = "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIL8z0qCNgA3x2xxAB0Qj5ro8waFjGZ8Ta/sWB63tlLw+ felhom-op-1\n"
def run(host, *argv):
return pa.main(list(argv), host=host)
class Accepts(unittest.TestCase):
"""What the agent really writes is accepted and installed, root-owned, at the fixed destination."""
def test_local_unit_installed(self):
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/mnt-hdd_1.mount", LOCAL_UNIT)
self.assertEqual(run(h, "unit", "mnt-hdd_1.mount"), 0)
self.assertEqual(h.writes, [("/etc/systemd/system/mnt-hdd_1.mount", 0o644)])
def test_local_unit_without_type(self): # the N100's unit has no Type= (autodetect)
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/mnt-hdd_1.mount", LOCAL_UNIT.replace("Type=ext4\n", ""))
self.assertEqual(run(h, "unit", "mnt-hdd_1.mount"), 0)
def test_network_pair(self):
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/" + NET_NAME, NET_MOUNT)
h.stage("/var/lib/felhom-agent/units/mnt-felhom\\x2ddrives-media.automount", NET_AUTOMOUNT)
self.assertEqual(run(h, "unit", NET_NAME), 0)
self.assertEqual(run(h, "unit", "mnt-felhom\\x2ddrives-media.automount"), 0)
def test_identical_is_not_rewritten(self):
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/mnt-hdd_1.mount", LOCAL_UNIT)
h.dest["/etc/systemd/system/mnt-hdd_1.mount"] = LOCAL_UNIT.encode()
self.assertEqual(run(h, "unit", "mnt-hdd_1.mount"), 0)
self.assertEqual(h.writes, [])
def test_dnsmasq_both_dropins(self):
h = FakeHost()
h.stage("/tmp/felhom-resolver-123.conf", DNS_BASE)
self.assertEqual(run(h, "dnsmasq", "/tmp/felhom-resolver-123.conf", "felhom-resolver-base.conf"), 0)
h.stage("/tmp/felhom-resolver-124.conf", DNS_GUEST)
self.assertEqual(run(h, "dnsmasq", "/tmp/felhom-resolver-124.conf", "felhom-demo-hp.conf"), 0)
self.assertIn(("/etc/dnsmasq.d/felhom-demo-hp.conf", 0o644), h.writes)
def test_wg_installed_0600(self):
h = FakeHost()
h.stage(pa.WG_SRC, WG)
self.assertEqual(run(h, "wg"), 0)
self.assertEqual(h.writes, [(pa.WG_DEST, 0o600)])
def test_sshd_template_and_key(self):
h = FakeHost()
h.stage(pa.SSHD_SRC, pa.render_sshd(8822))
h.stage(pa.KEY_SRC, KEYLINE)
self.assertEqual(run(h, "sshd-config"), 0)
self.assertEqual(run(h, "sshd-key"), 0)
h.stage(pa.KEY_SRC, "") # clearing the operator login is allowed
self.assertEqual(run(h, "sshd-key"), 0)
def test_escape_matches_systemd(self):
self.assertEqual(pa.systemd_escape_path("/mnt/felhom-drives/media"), "mnt-felhom\\x2ddrives-media")
self.assertEqual(pa.systemd_escape_path("/mnt/hdd_1"), "mnt-hdd_1")
self.assertEqual(pa.systemd_escape_path("/mnt/.x"), "mnt-.x") # a dot is escaped only at the very start
class Refuses(unittest.TestCase):
"""Each rule, with the attack it exists for. Nothing is written on a refusal."""
def refused(self, h, argv, rule):
rc = run(h, *argv)
self.assertIn(rc, (2, 3), f"{argv} was accepted")
self.assertEqual(h.writes, [], f"{argv} wrote something")
self.assertTrue(any(f"[{rule}]" in l for l in h.logs), f"{argv}: rule {rule} not logged: {h.logs}")
def unit(self, text, name="mnt-hdd_1.mount"):
h = FakeHost()
h.stage("/var/lib/felhom-agent/units/" + name, text)
return h, ["unit", name]
def test_U1_name_outside_mnt(self):
h = FakeHost()
self.refused(h, ["unit", "etc-sudoers.d.mount"], "U1")
self.refused(h, ["unit", "../../etc/x.mount"], "U1")
self.refused(h, ["unit", "mnt-x.service"], "U1")
def test_U2_service_section(self):
self.refused(*self.unit(LOCAL_UNIT + "\n[Service]\nExecStart=/bin/sh -c id\n"), "U2")
def test_U2_wants_starts_another_unit(self): # review 2026-10-05: Wants=reboot.target would reboot the host
for extra in ("Wants=reboot.target", "Requires=felhom-agent-rollback.service", "Before=pve-guests.service"):
self.refused(*self.unit(LOCAL_UNIT.replace("After=local-fs-pre.target", "After=local-fs-pre.target\n" + extra)), "U2")
self.refused(*self.unit(LOCAL_UNIT.replace("After=local-fs-pre.target", "After=poweroff.target")), "U2")
def test_U2_continuation_line(self):
t = LOCAL_UNIT.replace("Description=Felhom storage mount 91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae",
"Description=Felhom storage mount \\")
self.refused(*self.unit(t), "U2")
self.refused(*self.unit(LOCAL_UNIT.replace("# Managed by felhom-agent", "# comment \\\n# Managed by felhom-agent")), "U2")
def test_U2_unknown_key(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Type=ext4", "Type=ext4\nDirectoryMode=0777")), "U2")
def test_U3_bind_over_sudoers_dir(self):
# the R-861 attack: mount an agent-owned directory over /etc/sudoers.d
t = LOCAL_UNIT.replace("Where=/mnt/hdd_1", "Where=/etc/sudoers.d")
self.refused(*self.unit(t, "mnt-hdd_1.mount"), "U3")
def test_U3_name_must_match_where(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Where=/mnt/hdd_1", "Where=/mnt/other")), "U3")
def test_U3_traversal_in_where(self):
# the name passes U1 and equals the escaped Where — ONLY the Where rule stops a mount at /mnt/../etc = /etc
self.assertEqual(pa.systemd_escape_path("/mnt/../etc") + ".mount", "mnt-..-etc.mount")
self.refused(*self.unit(LOCAL_UNIT.replace("Where=/mnt/hdd_1", "Where=/mnt/../etc"), "mnt-..-etc.mount"), "U3")
def test_U4_what_is_an_agent_directory(self):
t = LOCAL_UNIT.replace("What=/dev/disk/by-uuid/91d2dc2d-2d28-4929-9bdd-3e11fa2f41ae", "What=/var/lib/felhom-agent/evil")
self.refused(*self.unit(t), "U4")
def test_U4_tmpfs(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Type=ext4", "Type=tmpfs")), "U4")
def test_U5_bind_option(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Type=ext4", "Type=ext4\nOptions=bind")), "U5")
def test_U5_network_without_nosuid(self):
t = NET_MOUNT.replace(",nosuid,nodev", "")
self.refused(*self.unit(t, NET_NAME), "U5")
def test_U5_suid_option(self):
self.refused(*self.unit(LOCAL_UNIT.replace("Type=ext4", "Type=ext4\nOptions=suid,dev")), "U5")
def test_U3_network_outside_drives(self):
t = NET_MOUNT.replace("/mnt/felhom-drives/media", "/mnt/media")
self.refused(*self.unit(t, "mnt-media.mount"), "U3")
def test_D1_dhcp_script(self): # runs as root
h = FakeHost()
h.stage("/tmp/felhom-resolver-1.conf", DNS_BASE + "dhcp-script=/var/lib/felhom-agent/x.sh\n")
self.refused(h, ["dnsmasq", "/tmp/felhom-resolver-1.conf", "felhom-x.conf"], "D1")
def test_D1_conf_dir_and_log_file(self):
for extra in ("conf-dir=/var/lib/felhom-agent\n", "log-facility=/etc/sudoers.d/x\n", "user=root\n"):
h = FakeHost()
h.stage("/tmp/felhom-resolver-1.conf", DNS_GUEST + extra)
self.refused(h, ["dnsmasq", "/tmp/felhom-resolver-1.conf", "felhom-x.conf"], "D1")
def test_D2_paths(self):
h = FakeHost()
self.refused(h, ["dnsmasq", "/etc/shadow", "felhom-x.conf"], "D2")
self.refused(h, ["dnsmasq", "/tmp/felhom-resolver-1.conf", "../sudoers.d/x.conf"], "D2")
def test_W1_postup(self): # wg-quick runs PostUp as root
h = FakeHost()
h.stage(pa.WG_SRC, WG.replace("MTU = 1280", "MTU = 1280\nPostUp = /bin/sh -c id"))
self.refused(h, ["wg"], "W1")
def test_W2_values(self):
h = FakeHost()
h.stage(pa.WG_SRC, WG.replace("AllowedIPs = 10.77.0.1/32, 10.77.0.250/32", "AllowedIPs = 0.0.0.0/0"))
self.refused(h, ["wg"], "W2")
def test_S1_sshd_strictmodes(self): # an AuthorizedKeysFile the agent owns + StrictModes no = root login
h = FakeHost()
h.stage(pa.SSHD_SRC, pa.render_sshd(8822) + "StrictModes no\n")
self.refused(h, ["sshd-config"], "S1")
h2 = FakeHost()
h2.stage(pa.SSHD_SRC, pa.render_sshd(8822).replace("/etc/felhom-sshd/authorized_keys/%u", "/var/lib/felhom-agent/k"))
self.refused(h2, ["sshd-config"], "S1")
def test_S1_port_22(self):
h = FakeHost()
h.stage(pa.SSHD_SRC, pa.render_sshd(22))
self.refused(h, ["sshd-config"], "S1")
def test_S2_key_options_and_two_keys(self):
h = FakeHost()
h.stage(pa.KEY_SRC, 'command="/bin/sh" ' + KEYLINE)
self.refused(h, ["sshd-key"], "S2")
h2 = FakeHost()
h2.stage(pa.KEY_SRC, KEYLINE + KEYLINE)
self.refused(h2, ["sshd-key"], "S2")
def test_P1_symlink_owner_size(self):
h = FakeHost()
h.stage(pa.WG_SRC, WG, kind="symlink")
self.refused(h, ["wg"], "P1")
h2 = FakeHost()
h2.stage(pa.WG_SRC, WG, uid=0)
self.refused(h2, ["wg"], "P1")
h3 = FakeHost()
h3.stage(pa.WG_SRC, "#" * (pa.MAX_BYTES + 1))
self.refused(h3, ["wg"], "P1")
def test_A1_usage(self):
h = FakeHost()
self.refused(h, ["install", "/etc/shadow"], "A1")
self.refused(h, ["wg", "/etc/shadow"], "A1")
class ControllerImage(unittest.TestCase):
"""R-861 (a) A1 (`09` §3 decision 165): the agent can no longer `tee` any image ref into the guest. The root verb
reads the ref on stdin, requires our registry + our repository + an x.y.z tag, and writes the guest file itself.
RED-PROOF: on the pre-A1 wrapper `controller-image` is not a verb (A1 usage, rc 2) — the accepted case fails."""
def go(self, ref, *argv):
h = FakeHost()
h.stdin = ref.encode() if isinstance(ref, str) else ref
return h, run(h, *(argv or ("controller-image", "9201")))
def test_our_controller_ref_is_written_in_the_guest(self):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0\n")
self.assertEqual(rc, 0, h.logs)
self.assertEqual(h.guest_writes, [(9201, b"gitea.dooplex.hu/admin/felhom-controller:0.301.0\n")])
def test_a_foreign_image_is_refused(self):
for ref in ("docker.io/library/alpine:latest\n", "alpine\n",
"gitea.dooplex.hu/admin/felhom-controller:latest\n",
"gitea.dooplex.hu/admin/other:0.1.0\n",
"evil.example/admin/felhom-controller:0.301.0\n",
"gitea.dooplex.hu/admin/felhom-controller:0.301.0\nalpine\n",
"gitea.dooplex.hu/admin/felhom-controller:0.301.0 x\n",
"", "\n"):
h, rc = self.go(ref)
self.assertEqual(rc, 3, f"{ref!r} was accepted")
self.assertFalse(h.guest_writes, f"{ref!r} wrote the guest file")
self.assertTrue(any("[I1]" in l for l in h.logs), h.logs)
def test_oversize_stdin_is_refused(self):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0" + " " * 300)
self.assertEqual(rc, 3)
self.assertFalse(h.guest_writes)
def test_vmid_must_be_numeric(self):
for argv in (("controller-image", "9201;id"), ("controller-image", "-1"), ("controller-image",),
("controller-image", "9201", "9202")):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0\n", *argv)
self.assertIn(rc, (2, 3), argv)
self.assertFalse(h.guest_writes, argv)
def test_self_check_names_the_verb(self):
import io, contextlib
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
pa.main(["--self-check"])
self.assertIn("controller-image", buf.getvalue())
if __name__ == "__main__":
unittest.main(verbosity=2)
@@ -241,3 +241,23 @@ func TestNewestArchiveTime_DistinctPhantomsEachAnnounced(t *testing.T) {
t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String())
}
}
// R-99 (`09` §3 decision 140): the WARN for a PBS phantom ends with the cleanup runbook, so whoever sees it knows the
// one sanctioned way to remove it; a tiny archive on a dir storage is not a PBS phantom and gets no pointer.
// RED-PROOF: drop the `msg += phantomCleanupPointer` line → "the PBS phantom WARN does not end with the runbook pointer".
func TestRejectedArchiveWarnNamesTheCleanupRunbook(t *testing.T) {
var buf bytes.Buffer
r := runnerWithContent(t, &buf, []proxmox.StorageContent{phantomEntry(), goodPBSEntry()})
if _, _, err := r.NewestArchiveTime(context.Background(), 9201); err != nil {
t.Fatal(err)
}
const want = "INCOMPLETE archive when computing tier freshness — it is not a successful backup — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
if !strings.Contains(buf.String(), want) {
t.Errorf("the PBS phantom WARN does not end with the runbook pointer:\n%s", buf.String())
}
local := phantomEntry()
local.Format, local.VolID = "tar.zst", "local:backup/vzdump-lxc-9201-2026_07_28-05_31_14.tar.zst"
if got := rejectedArchiveMessage(local); strings.Contains(got, "pbs-phantom-cleanup") {
t.Errorf("a dir-storage archive got the PBS runbook pointer: %s", got)
}
}
+131
View File
@@ -0,0 +1,131 @@
package backup
import (
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// BackupSuccessState persists the newest SUCCESSFUL whole-guest backup per tier and guest (R-894).
//
// Why it exists. The due-check (`localapi` handleBackupDue) asks the tier's storage when a backup last
// landed (R-84) and falls back to the in-memory record when the storage cannot be read. The in-memory
// record is empty after an agent restart (Store, R-348), so "storage unreadable" right after a restart
// read as "no record — DUE". Measured 2026-10-05 on demo-hp: the agent restarted at 04:57, the off-site
// storage answered "Can't connect" at 06:25, the 7-day tier — last copy 2026-10-01 — read DUE, the
// controller asked, and vzdump failed. This file is the last known copy the fallback reads instead.
//
// It is read ONLY when the storage cannot be read. A storage that answers is the ground truth and wins,
// in both directions: an archive found there counts, and an archive absent there is absent even when
// this file remembers a success (a pruned or deleted archive must make the tier due — the same reason
// R-84 chose the storage over a persisted record). Pinned by
// TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers.
//
// Only SUCCESSES are written (the RestoreTestState rule): a failure must stay due and be retried, so a
// record of a failure has no reader.
type BackupSuccessState struct {
path string
mu sync.Mutex
last map[string]savedSuccess // key(target, vmid) → the newest success
}
type savedSuccess struct {
target string
vmid int
at time.Time
}
// backupSuccessJSON is one entry on disk.
type backupSuccessJSON struct {
Target string `json:"target"`
VMID int `json:"vmid"`
StartedAt string `json:"started_at"`
}
func backupStateKey(target string, vmid int) string { return target + "/" + strconv.Itoa(vmid) }
// NewBackupSuccessState opens (or creates) the state at path. A missing or unreadable file degrades to
// "nothing known" — the pre-R-894 behaviour, which is DUE — and never wedges the daemon.
func NewBackupSuccessState(path string) *BackupSuccessState {
s := &BackupSuccessState{path: path, last: map[string]savedSuccess{}}
data, err := os.ReadFile(path)
if err != nil {
return s
}
var entries []backupSuccessJSON
if json.Unmarshal(data, &entries) != nil {
return s
}
for _, e := range entries {
t, perr := time.Parse(time.RFC3339, e.StartedAt)
if perr != nil {
continue // one unreadable entry must not lose the others
}
s.last[backupStateKey(e.Target, e.VMID)] = savedSuccess{target: e.Target, vmid: e.VMID, at: t.UTC()}
}
return s
}
// RecordBackupSuccess saves b when it is a success newer than the one on file. target is the tier the
// job ran on (the due-check's key); a failure or an unparseable time is ignored.
func (s *BackupSuccessState) RecordBackupSuccess(target string, b hub.Backup) error {
if s == nil || !b.Success {
return nil
}
t, err := time.Parse(time.RFC3339, b.StartedAt)
if err != nil {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
k := backupStateKey(target, b.VMID)
if old, ok := s.last[k]; ok && !t.After(old.at) {
return nil
}
s.last[k] = savedSuccess{target: target, vmid: b.VMID, at: t.UTC()}
return s.saveLocked()
}
// LastKnownSuccess returns the newest saved success for this tier and guest (ok=false = none on file).
func (s *BackupSuccessState) LastKnownSuccess(target string, vmid int) (time.Time, bool) {
if s == nil {
return time.Time{}, false
}
s.mu.Lock()
defer s.mu.Unlock()
e, ok := s.last[backupStateKey(target, vmid)]
return e.at, ok
}
func (s *BackupSuccessState) saveLocked() error {
entries := make([]backupSuccessJSON, 0, len(s.last))
for _, e := range s.last {
entries = append(entries, backupSuccessJSON{Target: e.target, VMID: e.vmid, StartedAt: e.at.Format(time.RFC3339)})
}
// Deterministic file content (Go's map order is random).
sort.Slice(entries, func(i, j int) bool {
if entries[i].Target != entries[j].Target {
return entries[i].Target < entries[j].Target
}
return entries[i].VMID < entries[j].VMID
})
data, err := json.MarshalIndent(entries, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(s.path), 0o755); err != nil {
return err
}
tmp := s.path + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, s.path)
}
+59
View File
@@ -0,0 +1,59 @@
package backup
import (
"os"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-894: the on-disk newest success per tier survives a restart (a new state from the same file).
func TestBackupSuccessState_SurvivesRestart(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
s := NewBackupSuccessState(path)
at := time.Date(2026, 10, 1, 20, 15, 0, 0, time.UTC)
if err := s.RecordBackupSuccess("felhom-pbs", hub.Backup{VMID: 9201, Success: true, StartedAt: at.Format(time.RFC3339)}); err != nil {
t.Fatal(err)
}
got, ok := NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 9201)
if !ok || !got.Equal(at) {
t.Fatalf("after a restart the saved copy must read back; got %v ok=%v", got, ok)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("local", 9201); ok {
t.Fatal("another tier must not borrow this tier's copy")
}
}
// Only a NEWER success replaces the saved one; failures and unparseable times are ignored.
func TestBackupSuccessState_KeepsNewestSuccessOnly(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
s := NewBackupSuccessState(path)
newer := time.Date(2026, 10, 5, 0, 0, 0, 0, time.UTC)
older := newer.Add(-48 * time.Hour)
for _, b := range []hub.Backup{
{VMID: 1, Success: true, StartedAt: newer.Format(time.RFC3339)},
{VMID: 1, Success: true, StartedAt: older.Format(time.RFC3339)}, // older: ignored
{VMID: 1, Success: false, StartedAt: newer.Add(time.Hour).Format(time.RFC3339)}, // failure: ignored
{VMID: 1, Success: true, StartedAt: "not-a-time"}, // unparseable: ignored
} {
if err := s.RecordBackupSuccess("t", b); err != nil {
t.Fatal(err)
}
}
if got, _ := NewBackupSuccessState(path).LastKnownSuccess("t", 1); !got.Equal(newer) {
t.Fatalf("want the newest success %v, got %v", newer, got)
}
}
// A corrupt file degrades to "nothing known" (the pre-R-894 DUE answer), never a crash.
func TestBackupSuccessState_CorruptFileIsEmpty(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
if err := os.WriteFile(path, []byte("{not json"), 0o600); err != nil {
t.Fatal(err)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("t", 1); ok {
t.Fatal("a corrupt file must read as nothing known")
}
}
+60
View File
@@ -0,0 +1,60 @@
package backup
import (
"context"
"sort"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// ForeignKeyLedger (R-366 slice 2, `09` §3 decision 168) holds, per backup tier, the whole-guest archives the
// restore-test pick skipped because another key wrote them (R-727 — an earlier install of this box). Since R-727 the
// skip was one INFO log line per archive and nothing else, so after a reinstall the operator was never told that the
// box's older whole-guest copies are unreadable to it. The host report carries this ledger; the hub raises ONE
// operator event when it changes.
//
// It reports nil until a tier has been evaluated since the agent started, so a restart does not read as "the set
// changed to empty" (the hub keeps its last state for an absent field).
type ForeignKeyLedger struct {
mu sync.Mutex
byTarget map[string]hub.ForeignKeyArchives
}
// NewForeignKeyLedger builds an empty ledger.
func NewForeignKeyLedger() *ForeignKeyLedger {
return &ForeignKeyLedger{}
}
func (l *ForeignKeyLedger) set(target string, n int, oldest, newest int64) {
l.mu.Lock()
defer l.mu.Unlock()
if l.byTarget == nil {
l.byTarget = map[string]hub.ForeignKeyArchives{}
}
e := hub.ForeignKeyArchives{Target: target, Count: n}
if n > 0 {
e.Oldest = time.Unix(oldest, 0).UTC().Format(time.RFC3339)
e.Newest = time.Unix(newest, 0).UTC().Format(time.RFC3339)
}
l.byTarget[target] = e
}
// ForeignKeyArchives implements hub.ForeignKeyArchiveReporter: nil before any evaluation; otherwise the tiers that
// hold such archives (`Tiers` empty, never nil, when none do), sorted by tier.
func (l *ForeignKeyLedger) ForeignKeyArchives(context.Context) *hub.ForeignKeyArchivesStanza {
l.mu.Lock()
defer l.mu.Unlock()
if l.byTarget == nil {
return nil
}
out := []hub.ForeignKeyArchives{}
for _, e := range l.byTarget {
if e.Count > 0 {
out = append(out, e)
}
}
sort.Slice(out, func(i, j int) bool { return out[i].Target < out[j].Target })
return &hub.ForeignKeyArchivesStanza{Tiers: out}
}
@@ -0,0 +1,69 @@
package backup
import (
"context"
"encoding/json"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-366 slice 2 (`09` §3 decision 168) — the restore-test's skip of an archive written with another key stops being
// silent: the pick records, per tier, how many it skipped and their time range, and the host report carries it.
//
// COMPANION RED-PROOF (observed): remove the `r.foreign.set(...)` call from PickSettledRestoreCandidateOn → this fails
// with "after one evaluation the ledger must report felhom-pbs: 2 archives …; got []". Restored.
func TestR366_PickRecordsArchivesWrittenWithAnotherKey(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T17:27:32Z", Content: "backup", VMID: 9201, Size: 4774114206, CTime: 1789579652, Encrypted: earlierBox2},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z", Content: "backup", VMID: 9201, Size: 20811501236, CTime: 1789595994, Encrypted: earlierBox1},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey},
},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
l := NewForeignKeyLedger()
r.SetForeignKeyLedger(l)
if got := l.ForeignKeyArchives(context.Background()); got != nil {
t.Fatalf("before any evaluation the ledger must be nil (the hub keeps its state); got %v", got)
}
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err != nil {
t.Fatal(err)
}
st := l.ForeignKeyArchives(context.Background())
want := hub.ForeignKeyArchives{Target: "felhom-pbs", Count: 2, Oldest: "2026-09-16T17:27:32Z", Newest: "2026-09-16T21:59:54Z"}
var got []hub.ForeignKeyArchives
if st != nil {
got = st.Tiers
}
if len(got) != 1 || got[0] != want {
t.Fatalf("after one evaluation the ledger must report felhom-pbs: 2 archives 2026-09-16T17:27:32Z…21:59:54Z; got %v", got)
}
}
// Evaluated and none found → the stanza with `tiers: []`; not evaluated → no stanza at all. Never a null on the wire.
func TestR366_EvaluatedWithNoneIsAnEmptyList(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey}},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
l := NewForeignKeyLedger()
r.SetForeignKeyLedger(l)
b, _ := json.Marshal(hub.HostReport{ForeignKeyArchives: l.ForeignKeyArchives(context.Background())})
if strings.Contains(string(b), "foreign_key_archives") {
t.Fatalf("not evaluated must omit the stanza; got %s", b)
}
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err != nil {
t.Fatal(err)
}
b, _ = json.Marshal(hub.HostReport{ForeignKeyArchives: l.ForeignKeyArchives(context.Background())})
if !strings.Contains(string(b), `"foreign_key_archives":{"tiers":[]}`) {
t.Fatalf("evaluated with none must report tiers: []; got %s", b)
}
}
+63
View File
@@ -0,0 +1,63 @@
package backup
import (
"context"
"fmt"
"sync/atomic"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-874 (v0.145.0). THE MEASURED SHAPE (Part F spike, Tester 2): power-on sessions of ~1.5 h and ~5 min against a
// 6 h evaluation ticker that restarts at every start — no restore-test ever evaluated. Now the first evaluation runs
// FirstEval after start.
// COMPANION RED-PROOF: drop the first-evaluation timer in Run (back to the bare ticker) → "no evaluation within".
func TestR874_FirstEvaluationAfterStart(t *testing.T) {
var picks int32
s := NewScheduler(SchedulerOptions{
Runner: &fakeRTRunner{res: reconcile.RestoreTestResult{Pass: true, Verified: "boot+running"}},
Pick: func(context.Context) (string, error) {
atomic.AddInt32(&picks, 1)
return fmt.Sprintf("local:backup/vzdump-lxc-9201-%d.tar.zst", atomic.LoadInt32(&picks)), nil
},
Store: NewStore(), Spec: (&specSpy{}).build,
Cadence: 6 * time.Hour, FirstEval: 30 * time.Millisecond, Logger: quiet(),
})
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
go func() { _ = s.Run(ctx); close(done) }()
deadline := time.Now().Add(3 * time.Second)
for atomic.LoadInt32(&picks) == 0 && time.Now().Before(deadline) {
time.Sleep(10 * time.Millisecond)
}
cancel()
<-done
if atomic.LoadInt32(&picks) == 0 {
t.Fatal("no evaluation within 3 s of start (FirstEval 30 ms) — a box with short sessions never gets a restore-test")
}
}
// The earned restraint stays: an agent that restarts before FirstEval never evaluates (a crash loop does not
// hammer a failing tier).
func TestR874_CrashLoopNeverEvaluates(t *testing.T) {
var picks int32
for i := 0; i < 5; i++ { // five quick "restarts"
s := NewScheduler(SchedulerOptions{
Runner: &fakeRTRunner{res: reconcile.RestoreTestResult{Pass: true}},
Pick: func(context.Context) (string, error) { atomic.AddInt32(&picks, 1); return "x", nil },
Store: NewStore(), Spec: (&specSpy{}).build,
Cadence: 6 * time.Hour, FirstEval: 200 * time.Millisecond, Logger: quiet(),
})
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Millisecond)
_ = s.Run(ctx)
cancel()
}
if n := atomic.LoadInt32(&picks); n != 0 {
t.Fatalf("a restart before FirstEval evaluated %d time(s)", n)
}
if DefaultFirstEval != 30*time.Minute {
t.Fatalf("DefaultFirstEval = %s, the documented 30 min", DefaultFirstEval)
}
}
+33 -1
View File
@@ -66,8 +66,13 @@ type BackupRunner struct {
// due-check is served from the local-API handler goroutines.
rejectedMu sync.Mutex
rejected map[string]struct{}
// foreign (R-366 slice 2) records, per tier, the archives the pick skipped as another key's. nil = not wired.
foreign *ForeignKeyLedger
}
// SetForeignKeyLedger wires the R-366 slice-2 ledger the host report reads.
func (r *BackupRunner) SetForeignKeyLedger(l *ForeignKeyLedger) { r.foreign = l }
// NewBackupRunner builds a runner. mode defaults to snapshot (works for a stopped guest and
// for lvm-thin); the caller may pass ModeStop for storages without snapshot support. retention is the
// per-run prune spec ("keep-last=N", or "" to never prune) — only the periodic local backup sets it.
@@ -370,6 +375,8 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
var best string
var bestCTime int64 = -1
known := map[int]bool{} // vmid → the guest exists on this node (asked once per vmid per pick)
var foreignN int // R-366 slice 2: archives skipped as another key's, and their time range
var foreignMin, foreignMax int64
for _, e := range contents {
if e.Content != "backup" {
continue
@@ -384,6 +391,13 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
}
if ownKey != "" && !strings.EqualFold(e.Encrypted, ownKey) {
r.noteNotAGuestBackupOnce(e, fmt.Sprintf("written by another box (key %s, this box's key %s) — not this box's proof", shortFP(e.Encrypted), shortFP(ownKey)))
foreignN++
if foreignMin == 0 || e.CTime < foreignMin {
foreignMin = e.CTime
}
if e.CTime > foreignMax {
foreignMax = e.CTime
}
continue
}
// R-689 (v0.136.0): … OF A GUEST THAT STILL EXISTS here. Measured on demo-hp 2026-09-27 right after
@@ -420,6 +434,9 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
bestCTime, best = e.CTime, e.VolID
}
}
if r.foreign != nil {
r.foreign.set(target, foreignN, foreignMin, foreignMax)
}
if best == "" {
return "", time.Time{}, nil
}
@@ -518,10 +535,25 @@ func (r *BackupRunner) warnRejectedArchiveOnce(e proxmox.StorageContent, why str
if seen {
return
}
r.logger.Warn("backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup",
r.logger.Warn(rejectedArchiveMessage(e),
"target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
}
// phantomCleanupPointer names the runbook that removes a PBS phantom (R-99, `09` §3 decision 140: a leftover of an
// aborted upload is deleted on the backup server, by a runbook, when one is seen — never automatically).
const phantomCleanupPointer = " — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
// rejectedArchiveMessage is the WARN text for a rejected archive. Only a PBS entry (format pbs-ct / pbs-vm) gets the
// cleanup pointer: the runbook deletes on a PBS datastore, and a tiny archive on a dir storage is not a PBS phantom.
// Pinned by TestRejectedArchiveWarnNamesTheCleanupRunbook.
func rejectedArchiveMessage(e proxmox.StorageContent) string {
msg := "backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup"
if strings.HasPrefix(e.Format, "pbs-") {
msg += phantomCleanupPointer
}
return msg
}
// demo-felhom in a single afternoon of deploys (2026-07-26).
//
// Asking the STORAGE rather than persisting the store is deliberate:
+34 -6
View File
@@ -60,12 +60,21 @@ type Scheduler struct {
// R-85 tier rotation. All optional: without them the scheduler behaves exactly as before
// (single tier via `pick`), which keeps every existing caller and test working untouched.
tiers []string // configured tier target ids, primary first
tierPick TierPicker // newest archive on a named tier
rtState *RestoreTestState // persisted last-successful-per-tier (drives oldest-first)
inFlight *InFlight // shared with the backup path — Scenario F
tiers []string // configured tier target ids, primary first
tierPick TierPicker // newest archive on a named tier
rtState *RestoreTestState // persisted last-successful-per-tier (drives oldest-first)
inFlight *InFlight // shared with the backup path — Scenario F
firstEval time.Duration // R-874: the first evaluation after start
}
// DefaultFirstEval (R-874): the first due-ness evaluation runs 30 minutes after the agent starts, then every
// cadence. MEASURED need (2026-10-05 Part F spike): a box whose power-on sessions are all shorter than the 6 h
// interval (Tester 2: ~1.5 h and ~5 min) NEVER evaluated, because the ticker restarts at each start. 30 minutes
// keeps the earned restraint below — a crash-looping agent restarts far more often than that and still never
// evaluates — while a box that stays on for half an hour gets its due test. Pinned by
// TestR874_FirstEvaluationAfterStart and TestR874_CrashLoopNeverEvaluates.
const DefaultFirstEval = 30 * time.Minute
// SchedulerOptions configures a Scheduler.
type SchedulerOptions struct {
Runner RestoreTestRunner
@@ -81,6 +90,8 @@ type SchedulerOptions struct {
// 0 → no settle requirement (any archive is a candidate).
Settle time.Duration
Logger *slog.Logger
// FirstEval (R-874, v0.145.0) is when the FIRST evaluation runs after start; 0 → DefaultFirstEval.
FirstEval time.Duration
// R-85 (all optional — omit for the pre-R-85 single-tier behaviour):
// Tiers are the configured tier target ids (primary first); TierPick resolves an archive on a
@@ -110,6 +121,12 @@ func NewScheduler(opts SchedulerOptions) *Scheduler {
tierPick: opts.TierPick,
rtState: opts.State,
inFlight: opts.InFlight,
firstEval: func() time.Duration {
if opts.FirstEval > 0 {
return opts.FirstEval
}
return DefaultFirstEval
}(),
}
}
@@ -120,8 +137,9 @@ func NewScheduler(opts SchedulerOptions) *Scheduler {
// trigger any more: its phase is the process's uptime, and agent deploys reset it, which is exactly
// the defect R-86 removes. What decides that a test happens is `EvaluateDue`.
//
// It still does NOT evaluate immediately on start — the first evaluation is one interval in. That
// is an EARNED restraint, kept deliberately: a restore is heavy, agent restarts are routine, and a
// It still does NOT evaluate immediately on start. v0.145.0 (R-874): the first evaluation is
// firstEval (30 min) in, then every interval — it was one full interval in, which a box with short
// power-on sessions never reached. The restraint itself is EARNED and kept: a restore is heavy, agent restarts are routine, and a
// crash-loop that evaluated at start would hammer a permanently-failing tier as fast as it could
// restart. Due-ness does not expire while we wait, so the only cost is up to one interval of
// latency on a tier that just became due. On-demand runs use `--selftest=restore-test`.
@@ -135,6 +153,16 @@ func (s *Scheduler) Run(ctx context.Context) error {
}
s.logger.Info("backup: restore-test scheduler starting (per-archive due-check)",
"eval_interval", s.cadence, "settle", s.settle)
first := time.NewTimer(s.firstEval)
defer first.Stop()
select {
case <-ctx.Done():
s.logger.Info("backup: restore-test scheduler shutting down", "reason", ctx.Err())
return nil
case <-first.C:
s.logger.Info("backup: restore-test first evaluation after start (R-874)", "after", s.firstEval)
s.tick(ctx)
}
t := time.NewTicker(s.cadence)
defer t.Stop()
for {
+8 -1
View File
@@ -25,7 +25,14 @@ import (
// merges the two — see hub.ProvenRestoreTestReporter. This store remains the ONLY place a FAILURE is
// recorded, and that asymmetry is deliberate: a failing tier stays due and is retried, so a lost
// failure heals itself, while a lost success leaves the system quietly less tested than it believes.
// Backups are unaffected — their freshness has a ground truth on the storage (R-84).
// Backups are NOT unaffected (corrected 2026-10-05, R-348): byTarget is in memory too, so after a restart the
// reported backup LIST reads 0 until the next backup of each tier runs (daily local, weekly offsite) — measured
// 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What
// is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/
// deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84).
// The due-check's fallback for an UNREADABLE storage no longer reads this store alone (R-894): the newest
// success per tier is also on disk (BackupSuccessState), so a restart followed by an unreachable storage
// reads the last known copy, not "never".
type Store struct {
mu sync.Mutex
byTarget map[string]hub.Backup // latest backup per target id
+14 -13
View File
@@ -81,10 +81,8 @@ var manifest = []Capability{
{"drives-mkdir-sub", "per-drive stable dir create", "/usr/bin/mkdir", []string{"-p", "/mnt/felhom-drives/felhom-usb"}, false, ""},
{"drives-mkdir-data", "felhom-data namespace create", "/usr/bin/mkdir", []string{"-p", "/mnt/felhom-usb/felhom-data"}, false, ""},
{"drives-chown-data", "felhom-data guest-root chown", "/usr/bin/chown", []string{"100000:100000", "/mnt/felhom-usb/felhom-data"}, false, ""},
{"parent-script-install", "shared-parent boot script install", "/usr/bin/install", []string{"-m", "0755", "--", "/tmp/felhom-shared-parent-123456789.sh", "/usr/local/sbin/felhom-shared-parent.sh"}, false, ""},
{"parent-unit-install", "shared-parent boot unit install", "/usr/bin/install", []string{"-m", "0644", "--", "/tmp/felhom-shared-parent-123456789.service", "/etc/systemd/system/felhom-shared-parent.service"}, false, ""},
{"parent-unit-enable", "shared-parent boot-persistence enable", "/usr/bin/systemctl", []string{"enable", "felhom-shared-parent.service"}, false, ""},
{"parent-bind-mp8", "parent bind into guest at provision", "/usr/sbin/pct", []string{"set", "9201", "-mp8", "/mnt/felhom-drives"}, false, ""},
{"parent-bind-mp8", "parent bind into guest at provision", "/usr/sbin/pct", []string{"set", "9201", "-mp8", "/mnt/felhom-drives,mp=/mnt/felhom-drives"}, false, ""},
// ---- Disk inspect / format gate (Critical: the data-bearing classifier + format) ----
{"disk-blkid", "disk data-bearing classify (format gate)", "/usr/sbin/blkid", []string{"-p", "-o", "export", "/dev/sda"}, true, ""},
@@ -95,11 +93,11 @@ var manifest = []Capability{
{"disk-lvs", "thin-pool usage read", "/usr/sbin/lvs", []string{"--reportformat", "json", "--units", "b", "-o", "lv_name,data_percent,metadata_percent", "--", "pve/data"}, false, ""},
// ---- Storage mount units (watchdog re-mount) ----
{"mount-unit-install", "fs-UUID mount unit install", "/usr/bin/install", []string{"-o", "root", "-g", "root", "-m", "0644", "--", "/var/lib/felhom-agent/units/felhom-x.mount", "/etc/systemd/system/felhom-x.mount"}, false, ""},
{"mount-unit-install", "fs-UUID mount unit install (root content check, R-861)", "/usr/local/sbin/felhom-priv-apply", []string{"unit", "mnt-felhom\\x2dx.mount"}, false, ""},
{"mount-daemon-reload", "systemd reload after unit write", "/usr/bin/systemctl", []string{"daemon-reload"}, false, ""},
{"mount-unit-enable", "mount unit enable", "/usr/bin/systemctl", []string{"enable", "--now", "--", "felhom-x.mount"}, false, ""},
{"mount-unit-disable", "mount unit disable", "/usr/bin/systemctl", []string{"disable", "--", "felhom-x.mount"}, false, ""},
{"mount-unit-stop", "mount unit stop", "/usr/bin/systemctl", []string{"stop", "--", "felhom-x.mount"}, false, ""},
{"mount-unit-enable", "mount unit enable", "/usr/bin/systemctl", []string{"enable", "--now", "--", "mnt-felhom\\x2dx.mount"}, false, ""},
{"mount-unit-disable", "mount unit disable", "/usr/bin/systemctl", []string{"disable", "--", "mnt-felhom\\x2dx.mount"}, false, ""},
{"mount-unit-stop", "mount unit stop", "/usr/bin/systemctl", []string{"stop", "--", "mnt-felhom\\x2dx.mount"}, false, ""},
// ---- Network storage re-arm + cleanup (CAMPAIGN-3 F10/F1) ----
{"netmount-reset-failed", "NAS automount re-arm after start-limit (F10)", "/usr/bin/systemctl", []string{"reset-failed", "--", "mnt-felhom\\x2ddrives-media.automount"}, false, ""},
@@ -110,18 +108,17 @@ var manifest = []Capability{
// ---- Provisioning back-half ----
{"provision-chown", "bootstrap mount guest-root chown", "/usr/bin/chown", []string{"-R", "100000:100000", "/var/lib/felhom-agent/guests/9201"}, false, ""},
{"provision-config-mount", "bootstrap config bind mount", "/usr/sbin/pct", []string{"set", "9201", "-mp0", "/var/lib/felhom-agent/guests/9201"}, false, ""},
{"provision-config-mount", "bootstrap config bind mount", "/usr/sbin/pct", []string{"set", "9201", "-mp0", "/var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1"}, false, ""},
{"provision-onboot", "customer guest autostart (onboot)", "/usr/sbin/pct", []string{"set", "9201", "-onboot", "1"}, false, ""},
// ---- Pre-start self-heal hook + guest lifecycle ----
{"guesthook-install", "pre-start hook snippet install", "/usr/bin/install", []string{"-m", "0755", "--", "/tmp/felhom-guest-hook-123456789.sh", "/var/lib/vz/snippets/felhom-guest-hook.sh"}, false, ""},
{"guesthook-register", "pre-start hook register", "/usr/sbin/pct", []string{"set", "9201", "--hookscript", "local:snippets/felhom-guest-hook.sh"}, false, ""},
{"guesthook-delete-mp", "dead mountpoint slot delete (C1 net)", "/usr/sbin/pct", []string{"set", "9201", "--delete", "mp0"}, false, ""},
{"guest-reboot", "enroll activate-binds reboot", "/usr/sbin/pct", []string{"reboot", "9201"}, false, ""},
// ---- LAN split-horizon resolver (dnsmasq) ----
{"dnsmasq-install", "dnsmasq package install", "/usr/bin/apt-get", []string{"install", "-y", "-q", "dnsmasq"}, false, ""},
{"dnsmasq-write", "dnsmasq drop-in write", "/usr/bin/install", []string{"-m", "0644", "/tmp/felhom-resolver-x.conf", "/etc/dnsmasq.d/felhom-x.conf"}, false, ""},
{"dnsmasq-write", "dnsmasq drop-in write (root content check, R-861)", "/usr/local/sbin/felhom-priv-apply", []string{"dnsmasq", "/tmp/felhom-resolver-123456789.conf", "felhom-x.conf"}, false, ""},
{"dnsmasq-enable", "dnsmasq enable", "/usr/bin/systemctl", []string{"enable", "--now", "dnsmasq"}, false, ""},
// ---- OS updates, guest fast lane (`11` §5.4.1; the wrapper holds every rule) ----
@@ -148,19 +145,24 @@ var manifest = []Capability{
{"controllerswap-image-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "image", "inspect", "gitea.dooplex.hu/admin/felhom-controller:0.0.0"}, true, ""},
{"controllerswap-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "inspect", "-f", "{{.State.Running}}", "felhom-controller"}, true, ""},
{"controllerswap-restart", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "systemctl", "restart", "felhom-controller-bootstrap.service"}, true, ""},
{"controllerswap-write", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "tee", "/etc/felhom-controller-image"}, true, ""},
// R-861 (a) A1 (decision 165): the write goes through the ROOT verb that checks the ref; the agent has no `tee` grant.
{"controllerswap-write", "controller-swap / managed auto-update", "/usr/local/sbin/felhom-priv-apply", []string{"controller-image", "9201"}, true, ""},
// ---- Stale-lock recovery (FELHOM_STALELOCK, v0.49.0; Critical: a guest stuck behind a stale
// reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ----
{"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""},
// ---- Weekly guest disk trim (FELHOM_FSTRIM, R-444). NON-critical: a missing grant means the thin pool is not
// reclaimed this week (the trim job WARNs per guest and the report shows the failure), not a serving outage. ----
{"guest-fstrim", "weekly guest disk trim (thin-pool reclaim, R-444)", "/usr/sbin/pct", []string{"fstrim", "9201"}, false, ""},
// ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite
// backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the
// conf install, unit enable/restart and the handshake read gate the backup path. apt-install
// (one-time bootstrap) and disable (revocation, a deliberate teardown) stay non-critical. The
// handshake read is the ONLY wg invocation (never `dump`). ----
{"wg-tools-install", "wireguard-tools package install", "/usr/bin/apt-get", []string{"install", "-y", "-q", "wireguard-tools"}, false, ""},
{"wg-conf-install", "wg-felhom conf install", "/usr/bin/install", []string{"-o", "root", "-g", "root", "-m", "0600", "--", "/var/lib/felhom-agent/wg/wg-felhom.conf", "/etc/wireguard/wg-felhom.conf"}, true, ""},
{"wg-conf-install", "wg-felhom conf install (root content check, R-861)", "/usr/local/sbin/felhom-priv-apply", []string{"wg"}, true, ""},
{"wg-enable", "wg-quick@wg-felhom enable", "/usr/bin/systemctl", []string{"enable", "--now", "wg-quick@wg-felhom"}, true, ""},
{"wg-restart", "wg-quick@wg-felhom restart (conf change)", "/usr/bin/systemctl", []string{"restart", "wg-quick@wg-felhom"}, true, ""},
{"wg-disable", "wg-quick@wg-felhom disable (revocation)", "/usr/bin/systemctl", []string{"disable", "--now", "wg-quick@wg-felhom"}, false, ""},
@@ -192,7 +194,6 @@ var manifest = []Capability{
// operator-driven op, not a steady-state serving path — a degraded grant means "can't
// self-update" (fall back to a manual SSH deploy), not a serving outage. The apply repr uses a
// staging-dir path + a placeholder sha (list-mode never runs it). ----
{"selfupdate-apply", "agent self-update apply (A/B flip)", "/usr/local/sbin/felhom-selfupdate-guarded", []string{"apply", "/var/lib/felhom-agent/selfupdate/felhom-agent-0.0.0", "0000000000000000000000000000000000000000000000000000000000000000"}, false, ""},
{"selfupdate-commit", "agent self-update commit", "/usr/local/sbin/felhom-selfupdate-guarded", []string{"commit"}, false, ""},
{"selfupdate-rollback", "agent self-update rollback", "/usr/local/sbin/felhom-selfupdate-guarded", []string{"rollback"}, false, ""},
}
+23 -6
View File
@@ -56,8 +56,10 @@ func parseSudoersEntries(t *testing.T, text string) []string {
var cur strings.Builder
for i := 0; i < len(raw); i++ {
c := raw[i]
if c == '\\' && i+1 < len(raw) {
cur.WriteByte(raw[i+1]) // unescape: keep the next char literally (\, → , ; \: → :)
if c == '\\' && i+1 < len(raw) && strings.IndexByte(",:=", raw[i+1]) >= 0 {
// unescape the sudoers grammar escapes only (\, → , ; \: → : ; \= → =). Every other backslash is kept: in
// a regex entry (R-861) `\.` `\{` `\\` mean exactly what sudo's regex engine reads.
cur.WriteByte(raw[i+1])
i++
continue
}
@@ -108,10 +110,24 @@ func globToRegex(pat string) *regexp.Regexp {
return regexp.MustCompile(b.String())
}
// entryRegex compiles one sudoers entry. R-861 (v0.146.0): an entry whose ARGUMENTS are a sudo regular expression
// (`^...$`, sudo >= 1.9.10) is matched as one — the binary literally, then a space, then the arguments joined by spaces
// (sudo's own rule). Every other entry is an fnmatch glob (globToRegex). The parser has already undone the sudoers
// escapes (`\,` `\:` `\=` `\\`), which leaves a valid RE2 pattern.
func entryRegex(e string) *regexp.Regexp {
if i := strings.IndexByte(e, ' '); i > 0 {
bin, args := e[:i], e[i+1:]
if strings.HasPrefix(args, "^") && strings.HasSuffix(args, "$") {
return regexp.MustCompile("^" + regexp.QuoteMeta(bin) + " " + args[1:])
}
}
return globToRegex(e)
}
// matchesAny reports whether cmdline matches at least one sudoers entry pattern.
func matchesAny(cmdline string, entries []string) bool {
for _, e := range entries {
if globToRegex(e).MatchString(cmdline) {
if entryRegex(e).MatchString(cmdline) {
return true
}
}
@@ -184,7 +200,8 @@ func TestRedProof_DroppedGrantFailsCheck(t *testing.T) {
}
// TestRedProof_DroppedControllerSwapTeeFailsCheck is the companion red-proof for the v0.45.0
// FELHOM_CONTROLLERSWAP grants: with the `tee /etc/felhom-controller-image` line removed, the
// FELHOM_CONTROLLERSWAP grants: with the write grant removed (since R-861 (a) A1 the `felhom-priv-apply controller-image`
// line; before it, an agent `tee /etc/felhom-controller-image`), the
// controllerswap-write capability MUST be reported uncovered. Proves the build gate watches the new
// swap write grant (so dropping it can't ship a non-root agent that silently can't auto-update).
func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
@@ -194,7 +211,7 @@ func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
}
var kept []string
for _, ln := range strings.Split(string(data), "\n") {
if strings.Contains(ln, "tee /etc/felhom-controller-image") {
if strings.Contains(ln, "felhom-priv-apply ^controller-image") { // R-861 (a) A1: the write's grant
continue
}
kept = append(kept, ln)
@@ -213,7 +230,7 @@ func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
}
cmdline := write.Binary + " " + strings.Join(write.ReprArgs, " ")
if matchesAny(cmdline, entries) {
t.Errorf("red-proof FAILED: controllerswap-write still matches after dropping the tee grant")
t.Errorf("red-proof FAILED: controllerswap-write still matches after dropping its grant")
}
if full := parseSudoersEntries(t, string(data)); !matchesAny(cmdline, full) {
t.Errorf("controllerswap-write should be covered by the real sudoers")
+138
View File
@@ -0,0 +1,138 @@
package capability
import (
"os"
"testing"
)
// R-861 (agent v0.146.0): the attacks the old globs let through, each as the exact argv a compromised agent would
// send. NONE may match any entry of the new sudoers. The same list ran against the REAL sudo 1.9.16 (a throwaway
// container, and `sudo -l -U felhom-agent` on both demo boxes after the bundle): audits/hub-safety-2026-10-05/partF/.
//
// RED-PROOF: run this list against the v0.145.0 sudoers (git show v0.145.0:configs/felhom-agent.sudoers) → most of
// these MATCH (recorded in the same evidence folder).
var r861Injections = []string{
// a raw host disk for a guest (pct options smuggled through a vmid glob)
"/usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1",
"/usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives",
"/usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda",
"/usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda",
// a bind mount over /etc through traversal
"/usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x",
"/usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d",
"/usr/bin/umount /mnt/felhom-drives/x /",
"/usr/bin/chown 100000:100000 /mnt/a/felhom-data /etc/shadow",
"/usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount",
// root-read files the agent writes: gone as `install` lines
"/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount",
"/usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh",
"/usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh",
"/usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf",
"/usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf",
"/usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config",
// the unsigned binary flip
"/usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000",
// enabling or removing anything that is not ours
"/usr/bin/systemctl enable --now -- mnt-hdd_1.mount evil.service",
"/usr/bin/systemctl enable --now -- etc-sudoers.d.mount",
"/usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd",
"/usr/bin/rm -f /etc/dnsmasq.d/felhom-x.conf /etc/shadow",
"/usr/bin/rmdir /mnt/felhom-drives/x /etc",
// nftables commands chained after a set element
"/usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ; flush ruleset",
// extra options to read-only tools
"/usr/sbin/smartctl -a -j /dev/sda -s off",
"/usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x",
"/usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller",
"/usr/sbin/pct unlock 9201 --whatever",
// the checker with a path it must never take
"/usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount",
"/usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf",
"/usr/local/sbin/felhom-priv-apply wg /etc/shadow",
// R-861 (a) A1 (decision 165): the agent wrote ANY image ref into the guest by `tee` — now only the root verb may
"/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image",
"/usr/local/sbin/felhom-priv-apply controller-image 9201 9202",
"/usr/local/sbin/felhom-priv-apply controller-image 9201;id",
}
func TestSudoersRefusesTheR861Injections(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
for _, c := range r861Injections {
if matchesAny(c, entries) {
t.Errorf("the sudoers still allows: %s", c)
}
}
}
// R-444: the weekly trim's grant is ONE exact shape — `pct fstrim <vmid>` — and nothing smuggled after it.
// The manifest entry (guest-fstrim) proves the real call is still allowed (TestManifestCoveredBySudoers); this
// pins the other direction. RED-PROOF: write the rule as the glob `/usr/sbin/pct fstrim [0-9]*` → every decoy
// below with a trailing argument matches (the glob's `*` eats spaces).
func TestSudoersFstrimRuleIsExact(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
if !matchesAny("/usr/sbin/pct fstrim 9201", entries) {
t.Fatal("the sudoers does not allow `pct fstrim 9201` — the weekly trim cannot run")
}
for _, c := range []string{
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints",
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints 1",
"/usr/sbin/pct fstrim 9201; x",
"/usr/sbin/pct fstrim 9201 9202",
"/usr/sbin/pct fstrim 92a1",
"/usr/sbin/pct fstrim ",
"/usr/sbin/pct fstrim -- 9201",
"/usr/sbin/pct destroy 9201",
"/usr/sbin/pct destroy 9201 --purge",
} {
if matchesAny(c, entries) {
t.Errorf("the sudoers allows a command the trim rule must not: %q", c)
}
}
}
// R-861 (a) A1: the managed controller update still has its route — the root verb, one numeric vmid.
func TestSudoersAllowsTheControllerImageVerb(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
if !matchesAny("/usr/local/sbin/felhom-priv-apply controller-image 9201", parseSudoersEntries(t, string(data))) {
t.Fatal("the sudoers does not allow `felhom-priv-apply controller-image 9201` — a managed controller update cannot write its image")
}
}
// R-861 (b) B2 (decision 165, hygiene): felhom-op's `pct start|stop|unlock` grants are ONE numeric vmid each. The old
// glob `[0-9]*` eats spaces, so `pct stop 9201 --skiplock 1` and two vmids matched.
// RED-PROOF: on the pre-B2 felhom-op.sudoers (`/usr/sbin/pct stop [0-9]*`) the decoys match.
func TestFelhomOpSudoersPctIsExact(t *testing.T) {
data, err := os.ReadFile("../../configs/felhom-op.sudoers")
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
for _, ok := range []string{"/usr/sbin/pct start 9201", "/usr/sbin/pct stop 9201", "/usr/sbin/pct unlock 9201", "/usr/sbin/pct list"} {
if !matchesAny(ok, entries) {
t.Errorf("felhom-op lost a repair verb: %s", ok)
}
}
for _, bad := range []string{
"/usr/sbin/pct stop 9201 --skiplock 1",
"/usr/sbin/pct start 9201 9202",
"/usr/sbin/pct unlock 9201 --whatever",
"/usr/sbin/pct start 92a1",
"/usr/sbin/pct stop ",
"/usr/sbin/pct destroy 9201",
} {
if matchesAny(bad, entries) {
t.Errorf("felhom-op's sudoers allows %q", bad)
}
}
}
+11
View File
@@ -33,6 +33,7 @@ type Config struct {
LANResolver LANResolverConfig `json:"lan_resolver"`
WGTunnel WGTunnelConfig `json:"wg_tunnel"`
GuestNet GuestNetConfig `json:"guest_net"`
DiskTrim DiskTrimConfig `json:"disk_trim"`
OOB OOBConfig `json:"oob"`
SelfUpdate SelfUpdateConfig `json:"selfupdate"`
LogLevel string `json:"log_level"` // debug|info|warn|error (default info)
@@ -139,6 +140,16 @@ func (w WGTunnelConfig) WithDefaults() WGTunnelConfig {
return w
}
// DiskTrimConfig configures the R-444 weekly guest disk trim (internal/fstrim). DEFAULT-ON, like GuestNetConfig and for
// the same reason: it only acts on guests the agent already owns, and the operator ruled every box trims (`09` §3
// decision 139). Opting out is the explicit act: `"disk_trim": {"disable": true}`.
type DiskTrimConfig struct {
Disable bool `json:"disable"`
}
// Enabled reports whether the weekly trim should run.
func (d DiskTrimConfig) Enabled() bool { return !d.Disable }
// GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet).
//
// **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other
+48 -2
View File
@@ -7,9 +7,11 @@ import (
"encoding/hex"
"encoding/json"
"fmt"
"io"
"os"
"path/filepath"
"strings"
"syscall"
)
// Slice 10D.1 — IDENTITY escrow. The K-escrow (above) wraps the PBS *encryption key* via the
@@ -71,7 +73,7 @@ func HashResticPassword(pw string) string {
// (non-empty) but NEVER logged by callers — log the field NAME only (mirrors AttachWGKey). A missing file
// is a clean no-attach (pre-fork-4 behavior, byte-compatible bundle).
func AttachResticPassword(b *IdentityBundle, stagePath string) (bool, error) {
raw, err := os.ReadFile(stagePath)
raw, err := readStagedNoFollow(stagePath)
if err != nil {
if os.IsNotExist(err) {
return false, nil
@@ -92,7 +94,7 @@ func AttachResticPassword(b *IdentityBundle, stagePath string) (bool, error) {
// no-attach (pre-S3 behavior, byte-compatible bundle); a corrupt one is an error (the operator
// should know their escrow would silently lack a live identity).
func AttachWGKey(b *IdentityBundle, keyPath string) (bool, error) {
raw, err := os.ReadFile(keyPath)
raw, err := readStagedNoFollow(keyPath)
if err != nil {
if os.IsNotExist(err) {
return false, nil
@@ -180,3 +182,47 @@ func UnwrapIdentityBundle(ctx context.Context, blob []byte, recoveryCode string)
}
return b, nil
}
// readStagedNoFollow reads a file the AGENT staged, for the escrow ceremony that runs as ROOT (FELHOM_ESCROW). R-861
// (agent v0.146.0/0.146.1): both files live in the agent's own directory, so a compromised agent could put a SYMLINK
// there — at the file OR at any directory on the way (review 2026-10-05) — to a root-only file, and the root ceremony
// would seal that file into the blob and hand the agent R. So the path is walked from "/" one component at a time with
// openat(O_NOFOLLOW): no symlink anywhere, the last a regular file of at most 4 KiB. Once a directory is open, renaming
// it does not redirect the walk. A missing file keeps its os.IsNotExist meaning. Pinned by TestAttach_RefusesASymlink*.
func readStagedNoFollow(path string) ([]byte, error) {
if !filepath.IsAbs(path) {
return nil, fmt.Errorf("%s is not an absolute path", path)
}
clean := filepath.Clean(path)
parts := strings.Split(strings.TrimPrefix(clean, "/"), "/")
dirfd, err := syscall.Open("/", syscall.O_RDONLY|syscall.O_DIRECTORY|syscall.O_CLOEXEC, 0)
if err != nil {
return nil, err
}
for i, part := range parts {
last := i == len(parts)-1
flags := syscall.O_RDONLY | syscall.O_NOFOLLOW | syscall.O_CLOEXEC
if !last {
flags |= syscall.O_DIRECTORY
}
fd, err := syscall.Openat(dirfd, part, flags, 0)
syscall.Close(dirfd)
if err != nil {
return nil, &os.PathError{Op: "open", Path: clean, Err: err}
}
dirfd = fd
}
f := os.NewFile(uintptr(dirfd), clean)
defer f.Close()
fi, err := f.Stat()
if err != nil {
return nil, err
}
if !fi.Mode().IsRegular() {
return nil, fmt.Errorf("%s is not a regular file", path)
}
if fi.Size() > 4096 {
return nil, fmt.Errorf("%s is larger than 4 KiB", path)
}
return io.ReadAll(io.LimitReader(f, 4097))
}
+63
View File
@@ -0,0 +1,63 @@
package escrow
import (
"os"
"path/filepath"
"testing"
)
// R-861 (agent v0.146.0): the root escrow ceremony reads two files from the AGENT's directory. A symlink there to a
// root-only file must never be read (it would be sealed under R and handed to the agent).
// RED-PROOF (audits/hub-safety-2026-10-05/partF/red-proof.txt): use os.ReadFile in readStagedNoFollow → this fails.
func TestAttach_RefusesASymlink(t *testing.T) {
d := t.TempDir()
secret := filepath.Join(d, "root-only")
if err := os.WriteFile(secret, []byte("ROOT-ONLY-CANARY"), 0o600); err != nil {
t.Fatal(err)
}
link := filepath.Join(d, "restic_repo_password")
if err := os.Symlink(secret, link); err != nil {
t.Fatal(err)
}
var b IdentityBundle
if ok, err := AttachResticPassword(&b, link); err == nil || ok || b.ResticRepoPassword != "" {
t.Fatalf("a symlinked staged file was read: ok=%v err=%v value-set=%v", ok, err, b.ResticRepoPassword != "")
}
if ok, err := AttachWGKey(&b, link); err == nil || ok {
t.Fatalf("a symlinked wg key was read: ok=%v err=%v", ok, err)
}
// control: the real staged file is still read; a missing one is still a clean no-attach
real := filepath.Join(d, "real")
_ = os.WriteFile(real, []byte("pw\n"), 0o600)
if ok, err := AttachResticPassword(&b, real); err != nil || !ok || b.ResticRepoPassword != "pw" {
t.Fatalf("control: the plain staged file was not read: %v %v", ok, err)
}
if ok, err := AttachResticPassword(&b, filepath.Join(d, "absent")); err != nil || ok {
t.Fatalf("control: a missing file must stay a clean no-attach: %v %v", ok, err)
}
}
// Review 2026-10-05: a symlinked DIRECTORY on the way must stop the read too (O_NOFOLLOW alone guards only the last
// component). RED-PROOF: open the full path with O_NOFOLLOW only → this fails.
func TestAttach_RefusesASymlinkedDirectory(t *testing.T) {
d := t.TempDir()
secretDir := filepath.Join(d, "root-only-dir")
_ = os.Mkdir(secretDir, 0o700)
_ = os.WriteFile(filepath.Join(secretDir, "private.key"), []byte("AAECAwQFBgcICQoLDA0ODxAREhMUFRYXGBkaGxwdHh8=\n"), 0o600)
agentDir := filepath.Join(d, "agent")
_ = os.Mkdir(agentDir, 0o700)
if err := os.Symlink(secretDir, filepath.Join(agentDir, "wg")); err != nil {
t.Fatal(err)
}
var b IdentityBundle
if ok, err := AttachWGKey(&b, filepath.Join(agentDir, "wg", "private.key")); err == nil || ok || b.WGPrivateKey != "" {
t.Fatalf("a key behind a symlinked directory was read: ok=%v err=%v", ok, err)
}
// control: the same key under a REAL directory is read
_ = os.Remove(filepath.Join(agentDir, "wg"))
_ = os.Mkdir(filepath.Join(agentDir, "wg"), 0o700)
_ = os.WriteFile(filepath.Join(agentDir, "wg", "private.key"), []byte("AAECAwQFBgcICQoLDA0ODxAREhMUFRYXGBkaGxwdHh8=\n"), 0o600)
if ok, err := AttachWGKey(&b, filepath.Join(agentDir, "wg", "private.key")); err != nil || !ok {
t.Fatalf("control: a key under a real directory was not read: %v %v", ok, err)
}
}
+2
View File
@@ -27,6 +27,8 @@ const (
PortFile = ConfDir + "/port"
// PidFile is the instance pidfile (NOT a RuntimeDirectory — that is the G1 incident cause).
PidFile = "/run/felhom-sshd.pid"
// PrivApply is the root content checker that installs the config and felhom-op's key (R-861, agent v0.146.0).
PrivApply = "/usr/local/sbin/felhom-priv-apply"
// Unit is the systemd unit name.
Unit = "felhom-sshd"
// OperatorUser is the default operator login (scoped sudo; key in AuthKeysDir only).
+5 -2
View File
@@ -143,7 +143,8 @@ func (m *Manager) Apply(ctx context.Context, block *hub.WireWireguard) (int, err
m.logger.Error("felhomsshd: staged config failed sshd -t — NOT installing", "err", err, "stderr", strings.TrimSpace(string(errOut)))
return port, err
}
if _, errOut, err := m.runner.Run(ctx, "install", "-o", "root", "-g", "root", "-m", "0644", "--", m.stagedConfPath(), ConfPath); err != nil {
// R-861 (v0.146.0): the root checker installs it, and only if it is renderConfig's template for some port.
if _, errOut, err := m.runner.Run(ctx, PrivApply, "sshd-config"); err != nil {
m.logger.Error("felhomsshd: config install failed", "err", err, "stderr", strings.TrimSpace(string(errOut)))
return port, err
}
@@ -181,7 +182,9 @@ func (m *Manager) applyAuthorizedKeys(ctx context.Context, sshKey string) {
m.logger.Error("felhomsshd: staging authorized_keys", "err", err)
return
}
if _, errOut, err := m.runner.Run(ctx, "install", "-o", "root", "-g", "root", "-m", "0644", "--", staged, AuthKeysUserPath); err != nil {
// R-861: the root checker installs it — one plain public key, no options (command=, from=, …), or empty.
_ = staged
if _, errOut, err := m.runner.Run(ctx, PrivApply, "sshd-key"); err != nil {
m.logger.Error("felhomsshd: authorized_keys install failed", "err", err, "stderr", strings.TrimSpace(string(errOut)))
return
}
@@ -0,0 +1,26 @@
package felhomsshd
import (
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/privapplytest"
)
// R-861: renderConfig is byte-identical to the checker's template (only the Port varies); a changed line is refused.
// RED-PROOF: change one directive in renderConfig → this fails (the two templates drifted).
func TestPrivApply_AcceptsTheRenderedConfig(t *testing.T) {
for _, port := range []int{2222, 8822, 60022} {
conf, err := renderConfig(port)
if err != nil {
t.Fatal(err)
}
if got := privapplytest.Check(t, "sshd-config", "", conf); got != "OK" {
t.Errorf("port %d: %s", port, got)
}
}
conf, _ := renderConfig(8822)
if got := privapplytest.Check(t, "sshd-config", "", conf+"StrictModes no\n"); !strings.HasPrefix(got, "REFUSED") {
t.Fatalf("control: an extra directive was not refused: %s", got)
}
}
+346
View File
@@ -0,0 +1,346 @@
// Package fstrim is the weekly guest disk trim (R-444, operator ruling `09` §3 decision 139).
//
// Why: a thin pool only ever grows from blocks a guest has already FREED — `fstrim` inside the unprivileged container
// is refused (FITRIM: Operation not permitted), and nothing else on the box gives the blocks back. A full thin pool
// takes every guest on the host read-only, so the pool can reach 100 % from deleted data alone. Measured on demo-hp
// 2026-10-06 09:14Z: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %, 18/18 app probes 200, max 1.1 s
// (audits/ten-answers-2026-10-06/r444-measure.txt).
//
// The rule, each part pinned by a test in fstrim_test.go:
// - Weekly: a guest is DUE from Wednesday 10:00 local until it has been trimmed once since then (a box that was off
// on Wednesday catches up at its next eligible hour).
// - Daytime only: a trim starts only between 10:00 and 20:59 local — never in the night window (01:00–06:59) where
// the backups and the restore-tests run (TestEligibleHourNeverInTheNight).
// - Never beside a backup, a restore-test or another heavy operation: the pass holds the host-wide one-heavy-op gate
// (backup.InFlight) for its whole run; a busy gate DEFERS the pass to the next hourly tick.
// - A failed trim is retried at the next eligible hour, at most MaxAttemptsPerWeek times in one week.
// - The last result per guest (time, bytes, ok/fail) is persisted, so a restart neither loses it nor re-trims.
//
// The command is the ONE exact sudoers shape `pct fstrim <vmid>` (FELHOM_FSTRIM). Only guests from the pool-verified
// source (ListLXC ∩ the felhom pool, audit A1) and only RUNNING ones are trimmed.
package fstrim
import (
"context"
"encoding/json"
"fmt"
"log/slog"
"os"
"path/filepath"
"regexp"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// Schedule. The weekday/hours are fixed on purpose (one sentence the operator can read on the System page).
const (
Weekday = time.Wednesday
StartHour = 10 // first eligible local hour (inclusive)
EndHour = 21 // first NOT-eligible local hour (exclusive): last start is 20:59
MaxAttemptsPerWeek = 3
// TickInterval is how often the job looks; a deferred or failed pass is therefore retried the next hour.
TickInterval = time.Hour
// FirstTickDelay lets the agent settle after a start before the first look.
FirstTickDelay = 5 * time.Minute
// PerGuestTimeout bounds one `pct fstrim` (measured 24.4 s for 84 GiB).
PerGuestTimeout = 30 * time.Minute
)
// ScheduleText is the human description carried on the host report.
const ScheduleText = "weekly, due Wednesday from 10:00 host-local time; starts only 10:00-20:59; never beside a backup or restore-test"
// Runner runs a host command (proxmox.ExecRunner in production, through `sudo -n`).
type Runner interface {
Run(ctx context.Context, name string, args ...string) (stdout, stderr []byte, err error)
}
// GuestSource yields the guests this agent OWNS (the pool-verified source, never a bare ListLXC).
type GuestSource interface {
Guests(ctx context.Context) ([]proxmox.Guest, error)
}
// Gate is the host-wide one-heavy-operation gate (*backup.InFlight).
type Gate interface {
TryAcquire(what string) (release func(), busy string, ok bool)
}
// GateName is what the gate reports as busy while a trim runs.
const GateName = "guest-fstrim"
// Record is one guest's last trim attempt, as persisted.
type Record struct {
LastAttemptAt time.Time `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt time.Time `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
// Attempts counts the attempts since the current week's due time (reset by the first attempt of a new week).
Attempts int `json:"attempts"`
}
// Trimmer is the weekly job.
type Trimmer struct {
runner Runner
guests GuestSource
gate Gate
statePath string
logger *slog.Logger
loc *time.Location
now func() time.Time
mu sync.Mutex
records map[int]Record
}
// New builds the job and loads the persisted state. A missing state file is an empty state; a corrupt one is logged
// and treated as empty (the cost is one extra trim, never a missed one).
func New(runner Runner, guests GuestSource, gate Gate, statePath string, logger *slog.Logger) *Trimmer {
if logger == nil {
logger = slog.Default()
}
t := &Trimmer{runner: runner, guests: guests, gate: gate, statePath: statePath, logger: logger,
loc: time.Local, now: time.Now, records: map[int]Record{}}
t.load()
return t
}
func (t *Trimmer) load() {
data, err := os.ReadFile(t.statePath)
if err != nil {
if !os.IsNotExist(err) {
t.logger.Warn("fstrim: state read failed — starting empty", "path", t.statePath, "err", err)
}
return
}
var raw map[string]Record
if err := json.Unmarshal(data, &raw); err != nil {
t.logger.Warn("fstrim: state file corrupt — starting empty", "path", t.statePath, "err", err)
return
}
for k, r := range raw {
if id, err := strconv.Atoi(k); err == nil && id > 0 {
t.records[id] = r
}
}
}
func (t *Trimmer) saveLocked() error {
raw := make(map[string]Record, len(t.records))
for id, r := range t.records {
raw[strconv.Itoa(id)] = r
}
data, err := json.MarshalIndent(raw, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(t.statePath), 0o755); err != nil {
return err
}
tmp := t.statePath + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, t.statePath)
}
// EligibleHour reports whether a trim may START at local time lt.
func EligibleHour(lt time.Time) bool {
h := lt.Hour()
return h >= StartHour && h < EndHour
}
// weekAnchor is the most recent Wednesday StartHour:00 at or before lt (same location as lt).
func weekAnchor(lt time.Time) time.Time {
daysBack := (int(lt.Weekday()) - int(Weekday) + 7) % 7
d := lt.AddDate(0, 0, -daysBack)
a := time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
if a.After(lt) {
d = d.AddDate(0, 0, -7)
a = time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
}
return a
}
// due reports whether a guest with record r (ok=false: none) is due at local time lt.
func due(r Record, has bool, lt time.Time) bool {
if !has {
return true
}
anchor := weekAnchor(lt)
if r.LastAttemptAt.Before(anchor) {
return true // not tried this week
}
return !r.OK && r.Attempts < MaxAttemptsPerWeek
}
// Run looks every TickInterval until ctx ends. It never returns an error: a failed trim is a reported fact.
func (t *Trimmer) Run(ctx context.Context) {
t.logger.Info("fstrim: weekly guest disk trim starting", "schedule", ScheduleText)
timer := time.NewTimer(FirstTickDelay)
defer timer.Stop()
for {
select {
case <-ctx.Done():
return
case <-timer.C:
t.Pass(ctx)
timer.Reset(TickInterval)
}
}
}
// Pass is one look: outside the daytime window it does nothing; otherwise it trims every due, running, owned guest
// while holding the heavy-op gate.
func (t *Trimmer) Pass(ctx context.Context) {
lt := t.now().In(t.loc)
if !EligibleHour(lt) {
t.logger.Debug("fstrim: outside the daytime window — not looking", "local", lt.Format("Mon 15:04"))
return
}
guests, err := t.guests.Guests(ctx)
if err != nil {
t.logger.Warn("fstrim: owned-guest list unavailable — skipping this pass", "err", err)
return
}
owned := make(map[int]bool, len(guests))
var todo []int
t.mu.Lock()
for _, g := range guests {
owned[g.VMID] = true
r, has := t.records[g.VMID]
if !due(r, has, lt) {
continue
}
if g.Status != "running" {
t.logger.Info("fstrim: guest not running — trimmed when it runs", "vmid", g.VMID, "status", g.Status)
continue
}
todo = append(todo, g.VMID)
}
// A guest the agent no longer owns has no result to report.
pruned := false
for id := range t.records {
if !owned[id] {
delete(t.records, id)
pruned = true
}
}
if pruned {
if err := t.saveLocked(); err != nil {
t.logger.Warn("fstrim: state save failed", "err", err)
}
}
t.mu.Unlock()
if len(todo) == 0 {
return
}
sort.Ints(todo)
release, busy, ok := t.gate.TryAcquire(GateName)
if !ok {
t.logger.Info("fstrim: deferred — a heavy operation is in flight; retrying next hour", "busy", busy, "due_guests", len(todo))
return
}
defer release()
for _, vmid := range todo {
if ctx.Err() != nil {
return
}
t.trimOne(ctx, vmid, lt)
}
}
var trimmedLine = regexp.MustCompile(`\((\d+) bytes\) trimmed`)
// ParseTrimmed sums the "(N bytes) trimmed" lines of `pct fstrim` output and counts them (one per mount point), e.g.
// `/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed`.
func ParseTrimmed(out string) (bytes int64, mounts int) {
for _, m := range trimmedLine.FindAllStringSubmatch(out, -1) {
n, err := strconv.ParseInt(m[1], 10, 64)
if err != nil {
continue
}
bytes += n
mounts++
}
return bytes, mounts
}
// GiB renders bytes as "30.1 GiB".
func GiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
func (t *Trimmer) trimOne(ctx context.Context, vmid int, lt time.Time) {
start := t.now()
cctx, cancel := context.WithTimeout(ctx, PerGuestTimeout)
stdout, stderr, err := t.runner.Run(cctx, "pct", "fstrim", strconv.Itoa(vmid))
cancel()
dur := t.now().Sub(start)
bytes, mounts := ParseTrimmed(string(stdout) + "\n" + string(stderr))
t.mu.Lock()
prev, has := t.records[vmid]
r := Record{LastAttemptAt: start.UTC(), OK: err == nil, BytesTrimmed: bytes, Mounts: mounts,
DurationSeconds: float64(dur.Round(100*time.Millisecond)) / float64(time.Second), LastOKAt: prev.LastOKAt}
if has && !prev.LastAttemptAt.Before(weekAnchor(lt)) {
r.Attempts = prev.Attempts + 1
} else {
r.Attempts = 1
}
if err == nil {
r.LastOKAt = start.UTC()
} else {
msg := strings.TrimSpace(err.Error() + ": " + strings.TrimSpace(string(stderr)))
if len(msg) > 300 {
msg = msg[:300]
}
r.Error = msg
}
t.records[vmid] = r
saveErr := t.saveLocked()
t.mu.Unlock()
if err == nil {
t.logger.Info(fmt.Sprintf("fstrim: guest %d trimmed %s in %.1fs", vmid, GiB(bytes), r.DurationSeconds),
"vmid", vmid, "bytes_trimmed", bytes, "mounts", mounts, "duration_s", r.DurationSeconds)
if mounts == 0 {
t.logger.Warn("fstrim: pct fstrim succeeded but reported no trimmed mount — output not understood",
"vmid", vmid, "stdout", strings.TrimSpace(string(stdout)))
}
} else {
t.logger.Warn(fmt.Sprintf("fstrim: guest %d trim FAILED after %.1fs", vmid, r.DurationSeconds),
"vmid", vmid, "attempt", r.Attempts, "max_attempts_per_week", MaxAttemptsPerWeek, "err", r.Error)
}
if saveErr != nil {
t.logger.Warn("fstrim: state save failed — the result will not survive a restart", "path", t.statePath, "err", saveErr)
}
}
// GuestDiskTrimStatus implements hub.GuestDiskTrimReporter: a pure read of the persisted results (never runs pct).
func (t *Trimmer) GuestDiskTrimStatus(context.Context) *hub.GuestDiskTrimStatus {
t.mu.Lock()
defer t.mu.Unlock()
out := &hub.GuestDiskTrimStatus{Schedule: ScheduleText}
ids := make([]int, 0, len(t.records))
for id := range t.records {
ids = append(ids, id)
}
sort.Ints(ids)
for _, id := range ids {
r := t.records[id]
g := hub.GuestDiskTrim{VMID: id, LastAttemptAt: r.LastAttemptAt.UTC().Format(time.RFC3339), OK: r.OK,
BytesTrimmed: r.BytesTrimmed, Mounts: r.Mounts, DurationSeconds: r.DurationSeconds, Error: r.Error}
if !r.LastOKAt.IsZero() {
g.LastOKAt = r.LastOKAt.UTC().Format(time.RFC3339)
}
out.Guests = append(out.Guests, g)
}
return out
}
+267
View File
@@ -0,0 +1,267 @@
package fstrim
import (
"bytes"
"context"
"encoding/json"
"errors"
"log/slog"
"path/filepath"
"reflect"
"strings"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// The real `pct fstrim 9201` output measured on demo-hp 2026-10-06 (audits/ten-answers-2026-10-06/r444-measure.txt).
const measuredOut = "/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed\n" +
"/var/lib/lxc/9201/rootfs/var/lib/felhom: 53.9 GiB (57865633792 bytes) trimmed\n"
const measuredBytes = int64(32277680128 + 57865633792)
type fakeRunner struct {
mu sync.Mutex
calls [][]string
out string
err error
onRun func()
}
func (f *fakeRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
f.mu.Lock()
f.calls = append(f.calls, append([]string{name}, args...))
f.mu.Unlock()
if f.onRun != nil {
f.onRun()
}
if f.err != nil {
return nil, []byte("mount busy"), f.err
}
return []byte(f.out), nil, nil
}
type fakeGuests struct {
g []proxmox.Guest
err error
}
func (f fakeGuests) Guests(context.Context) ([]proxmox.Guest, error) { return f.g, f.err }
// A Wednesday 10:30 in a fixed zone (CEST-like), so the tests do not depend on the machine's zone.
var zone = time.FixedZone("CEST", 2*3600)
func at(day, hour, min int) time.Time { return time.Date(2026, 10, day, hour, min, 0, 0, zone) } // 2026-10-07 = Wednesday
func newT(t *testing.T, r Runner, g GuestSource, gate Gate, now *time.Time) (*Trimmer, *bytes.Buffer, string) {
t.Helper()
var logs bytes.Buffer
path := filepath.Join(t.TempDir(), "guest-disk-trim.json")
tr := New(r, g, gate, path, slog.New(slog.NewTextHandler(&logs, &slog.HandlerOptions{Level: slog.LevelDebug})))
tr.loc = zone
tr.now = func() time.Time { return *now }
return tr, &logs, path
}
func running(ids ...int) fakeGuests {
var g []proxmox.Guest
for _, id := range ids {
g = append(g, proxmox.Guest{VMID: id, Status: "running", Type: "lxc"})
}
return fakeGuests{g: g}
}
func TestParseTrimmedTheMeasuredOutput(t *testing.T) {
b, m := ParseTrimmed(measuredOut)
if b != measuredBytes || m != 2 {
t.Fatalf("ParseTrimmed = %d bytes over %d mounts, want %d over 2", b, m, measuredBytes)
}
if b, m := ParseTrimmed("something else\n"); b != 0 || m != 0 {
t.Fatalf("unrelated output parsed as %d/%d", b, m)
}
if got := GiB(measuredBytes); got != "84.0 GiB" {
t.Fatalf("GiB = %q", got)
}
}
// The night window (01:00–06:59) must never be eligible, and the daytime window is exactly 10:00–20:59.
func TestEligibleHourNeverInTheNight(t *testing.T) {
for h := 0; h < 24; h++ {
lt := time.Date(2026, 10, 7, h, 30, 0, 0, zone)
got := EligibleHour(lt)
if h >= 1 && h <= 6 && got {
t.Errorf("hour %02d is in the night window and must not be eligible", h)
}
if want := h >= 10 && h <= 20; got != want {
t.Errorf("EligibleHour(%02d:30) = %v, want %v", h, got, want)
}
}
}
func TestWeekAnchorIsTheLastWednesdayTen(t *testing.T) {
cases := map[time.Time]time.Time{
at(7, 10, 0): at(7, 10, 0), // Wednesday 10:00 itself
at(7, 9, 59): time.Date(2026, 9, 30, 10, 0, 0, 0, zone), // before 10:00 Wednesday → the previous week
at(8, 15, 0): at(7, 10, 0), // Thursday
at(13, 20, 0): at(7, 10, 0), // next Tuesday
at(14, 11, 0): at(14, 10, 0), // next Wednesday
}
for in, want := range cases {
if got := weekAnchor(in); !got.Equal(want) {
t.Errorf("weekAnchor(%s) = %s, want %s", in.Format("Mon 01-02 15:04"), got.Format("Mon 01-02 15:04"), want.Format("Mon 01-02 15:04"))
}
}
}
// The consequence: on Wednesday 10:30 a running owned guest is trimmed with the ONE exact argv, the bytes are parsed,
// the positive log line is written, the result is persisted, and the host report carries it.
func TestPassTrimsADueGuestAndReportsIt(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
tr, logs, path := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9201"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("runner calls = %q, want %q", r.calls, want)
}
if !strings.Contains(logs.String(), "fstrim: guest 9201 trimmed 84.0 GiB in ") {
t.Fatalf("no positive per-guest log line:\n%s", logs.String())
}
st := tr.GuestDiskTrimStatus(context.Background())
if st == nil || st.Schedule != ScheduleText || len(st.Guests) != 1 {
t.Fatalf("report stanza = %+v", st)
}
g := st.Guests[0]
if g.VMID != 9201 || !g.OK || g.BytesTrimmed != measuredBytes || g.Mounts != 2 || g.LastOKAt == "" || g.LastAttemptAt == "" {
t.Fatalf("report guest = %+v", g)
}
// Persisted: a NEW Trimmer over the same file (an agent restart) still has it and does not trim again this week.
now = at(8, 11, 0)
r2 := &fakeRunner{out: measuredOut}
tr2 := New(r2, running(9201), &backup.InFlight{}, path, slog.New(slog.NewTextHandler(&bytes.Buffer{}, nil)))
tr2.loc, tr2.now = zone, func() time.Time { return now }
if st2 := tr2.GuestDiskTrimStatus(context.Background()); len(st2.Guests) != 1 || st2.Guests[0].BytesTrimmed != measuredBytes {
t.Fatalf("result lost over a restart: %+v", st2)
}
tr2.Pass(context.Background())
if len(r2.calls) != 0 {
t.Fatalf("trimmed again in the same week after a restart: %q", r2.calls)
}
// Next week it is due again.
now = at(14, 10, 5)
tr2.Pass(context.Background())
if len(r2.calls) != 1 {
t.Fatalf("not trimmed in the next week: %q", r2.calls)
}
}
func TestPassNeverRunsInTheNight(t *testing.T) {
for _, h := range []int{1, 3, 6, 9, 21, 23} {
now := at(7, h, 15)
r := &fakeRunner{out: measuredOut}
tr, _, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Errorf("trimmed at %02d:15: %q", h, r.calls)
}
}
}
// A backup (or restore-test) holding the heavy-op gate DEFERS the trim; the next hour, gate free, it runs. And while
// a trim runs, the gate is held, so a backup cannot start beside it.
func TestPassDefersToAHeavyOperationAndRetriesNextHour(t *testing.T) {
now := at(7, 10, 30)
gate := &backup.InFlight{}
release, _, _ := gate.TryAcquire("backup:9201")
var busyDuringTrim string
r := &fakeRunner{out: measuredOut}
r.onRun = func() { busyDuringTrim = gate.Busy() }
tr, logs, _ := newT(t, r, running(9201), gate, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Fatalf("trimmed beside a running backup: %q", r.calls)
}
if !strings.Contains(logs.String(), "fstrim: deferred") || !strings.Contains(logs.String(), "backup:9201") {
t.Fatalf("the deferral is not logged with what holds the gate:\n%s", logs.String())
}
release()
now = now.Add(time.Hour)
tr.Pass(context.Background())
if len(r.calls) != 1 {
t.Fatalf("not retried the next hour: %q", r.calls)
}
if busyDuringTrim != GateName {
t.Fatalf("the heavy-op gate was %q during the trim, want %q", busyDuringTrim, GateName)
}
if gate.Busy() != "" {
t.Fatalf("the gate was not released after the pass: %q", gate.Busy())
}
}
func TestFailedTrimWarnsIsRecordedAndRetriedAtMostThreeTimes(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{err: errors.New("exit status 255")}
tr, logs, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
for i := 0; i < 6; i++ {
tr.Pass(context.Background())
now = now.Add(time.Hour)
}
if len(r.calls) != MaxAttemptsPerWeek {
t.Fatalf("attempts in one week = %d, want %d", len(r.calls), MaxAttemptsPerWeek)
}
if !strings.Contains(logs.String(), "level=WARN") || !strings.Contains(logs.String(), "fstrim: guest 9201 trim FAILED") {
t.Fatalf("no WARN for the failure:\n%s", logs.String())
}
g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]
if g.OK || g.LastOKAt != "" || !strings.Contains(g.Error, "exit status 255") || !strings.Contains(g.Error, "mount busy") {
t.Fatalf("failed result not recorded as a failure: %+v", g)
}
// A success later keeps a clean record.
r.err = nil
r.out = measuredOut
now = at(14, 10, 10)
tr.Pass(context.Background())
if g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]; !g.OK || g.Error != "" || g.BytesTrimmed != measuredBytes {
t.Fatalf("success after failure: %+v", g)
}
}
func TestOnlyRunningOwnedGuestsAndAFailedListActsOnNothing(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
g := fakeGuests{g: []proxmox.Guest{{VMID: 9201, Status: "stopped"}, {VMID: 9202, Status: "running"}}}
tr, _, _ := newT(t, r, g, &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9202"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("calls = %q, want only the running guest", r.calls)
}
r2 := &fakeRunner{out: measuredOut}
tr2, logs, _ := newT(t, r2, fakeGuests{err: errors.New("pool read 403")}, &backup.InFlight{}, &now)
tr2.Pass(context.Background())
if len(r2.calls) != 0 || !strings.Contains(logs.String(), "owned-guest list unavailable") {
t.Fatalf("a failed ownership read must act on nothing: calls %q", r2.calls)
}
}
func TestReportJSONShape(t *testing.T) {
now := at(7, 10, 30)
tr, _, _ := newT(t, &fakeRunner{out: measuredOut}, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
b, err := json.Marshal(tr.GuestDiskTrimStatus(context.Background()))
if err != nil {
t.Fatal(err)
}
for _, k := range []string{`"schedule":`, `"guests":[{"vmid":9201`, `"last_attempt_at":"2026-10-07T08:30:00Z"`, `"ok":true`,
`"bytes_trimmed":90143313920`, `"mounts":2`, `"duration_seconds":`, `"last_ok_at":"2026-10-07T08:30:00Z"`} {
if !strings.Contains(string(b), k) {
t.Errorf("report JSON lacks %s: %s", k, b)
}
}
if strings.Contains(string(b), `"error"`) {
t.Errorf("an ok result must omit error: %s", b)
}
}
+16 -25
View File
@@ -38,35 +38,26 @@ const snippetBody = `#!/bin/sh
exit 0
`
// InstallSnippet writes the pre-start hook wrapper into the PVE snippets dir (idempotent, root-owned,
// executable). The agent runs as a non-root service user, so it writes an agent-writable temp file then
// `install`s it host-root (same pattern as the bootstrap mount + dnsmasq drop-ins). Safe to call repeatedly.
// The temp file is a RANDOM-named os.CreateTemp (audit B1): a fixed, predictable /tmp name could be
// pre-created by another local user and rewritten between our write and root's install (TOCTOU into a
// root-executed hookscript). The final mode comes from `install -m`, so the 0600 temp is fine.
func InstallSnippet(ctx context.Context, runner proxmox.Runner) error {
f, err := os.CreateTemp("", "felhom-guest-hook-*.sh")
// SnippetReady (R-861, agent v0.146.0) reports whether the pre-start hook is in place: a regular file at path whose
// content is exactly snippetBody. The hook is a FIXED, ROOT-OWNED file that arrives with the signed config bundle
// (configs/felhom-guest-hook.sh, pinned byte-identical by TestSnippetEqualsTheBundle). The agent no longer installs it:
// until v0.146.0 it `install`ed it from /tmp, and Proxmox runs a hookscript as root at every guest start — so the
// install grant was a root shell for a compromised agent. A missing or different hook is an error the caller logs; it
// then does NOT register the hook (a guest whose hookscript is missing does not start).
func SnippetReady(path string) error {
fi, err := os.Lstat(path)
if err != nil {
return fmt.Errorf("guesthook: create temp snippet: %w", err)
return fmt.Errorf("guesthook: %s is missing — it arrives with the signed config bundle (agent_config_update): %w", path, err)
}
tmp := f.Name()
defer os.Remove(tmp)
if _, err := f.WriteString(snippetBody); err != nil {
f.Close()
return fmt.Errorf("guesthook: write temp snippet: %w", err)
if !fi.Mode().IsRegular() {
return fmt.Errorf("guesthook: %s is not a regular file", path)
}
if err := f.Close(); err != nil {
return fmt.Errorf("guesthook: close temp snippet: %w", err)
b, err := os.ReadFile(path)
if err != nil {
return fmt.Errorf("guesthook: read %s: %w", path, err)
}
// Ensure the snippets dir exists FIRST (B2, DRILL-day0-cleanroom-2026-07-03): a fresh PVE has
// no /var/lib/vz/snippets, and `install` (without -D) won't create the parent — the whole
// hook install silently failed on a freshly-bootstrapped box. Fenced root op like the install
// itself; idempotent.
if _, stderr, err := runner.Run(ctx, "mkdir", "-p", SnippetDir); err != nil {
return fmt.Errorf("guesthook: ensure snippets dir %s: %w: %s", SnippetDir, err, string(stderr))
}
if _, stderr, err := runner.Run(ctx, "install", "-m", "0755", "--", tmp, SnippetPath); err != nil {
return fmt.Errorf("guesthook: install snippet to %s: %w: %s", SnippetPath, err, string(stderr))
if string(b) != snippetBody {
return fmt.Errorf("guesthook: %s differs from this agent's hook — the next config bundle replaces it", path)
}
return nil
}
+24 -90
View File
@@ -4,7 +4,7 @@ import (
"context"
"io"
"os"
"regexp"
"path/filepath"
"testing"
)
@@ -29,100 +29,34 @@ func (r *recordingRunner) RunStdin(ctx context.Context, _ io.Reader, name string
return r.Run(ctx, name, args...)
}
// TestInstallSnippet_RandomTempName is the audit-B1 negative test: the staged install SOURCE must be a
// RANDOM os.CreateTemp name (felhom-guest-hook-<random>.sh), never the fixed, pre-creatable
// /tmp/felhom-guest-hook.sh (a local TOCTOU into a root-executed hookscript), and two consecutive
// installs must stage through DIFFERENT paths.
func TestInstallSnippet_RandomTempName(t *testing.T) {
r := &recordingRunner{}
if err := InstallSnippet(context.Background(), r); err != nil {
t.Fatalf("InstallSnippet #1: %v", err)
// R-861 (agent v0.146.0): the hook file comes with the signed config bundle; the agent never installs it, only checks.
// RED-PROOF (audits/hub-safety-2026-10-05/partF/red-proof.txt): make SnippetReady accept any content → the
// "differs" case fails.
func TestSnippetReady(t *testing.T) {
d := t.TempDir()
p := filepath.Join(d, "felhom-guest-hook.sh")
if err := SnippetReady(p); err == nil {
t.Fatal("a missing hook read as ready — the guest would get a hookscript that does not exist")
}
if err := InstallSnippet(context.Background(), r); err != nil {
t.Fatalf("InstallSnippet #2: %v", err)
_ = os.WriteFile(p, []byte("#!/bin/sh\nid > /tmp/x\n"), 0o755)
if err := SnippetReady(p); err == nil {
t.Fatal("a hook with other content read as ready")
}
var installs [][]string
for _, call := range r.calls {
if call[0] == "install" {
installs = append(installs, call)
}
_ = os.WriteFile(p, []byte(snippetBody), 0o755)
if err := SnippetReady(p); err != nil {
t.Fatalf("the bundle's hook was not accepted: %v", err)
}
if len(installs) != 2 {
t.Fatalf("expected 2 install calls, got %d: %v", len(installs), r.calls)
}
randomName := regexp.MustCompile(`felhom-guest-hook-[^/\\]+\.sh$`)
fixedName := regexp.MustCompile(`felhom-guest-hook\.sh$`)
var srcs []string
for i, call := range installs {
// install -m 0755 -- <src> <dest>
if len(call) != 6 {
t.Fatalf("call %d: unexpected vector %v", i, call)
}
src, dest := call[4], call[5]
if dest != SnippetPath {
t.Errorf("call %d: dest = %q, want %q", i, dest, SnippetPath)
}
if !randomName.MatchString(src) {
t.Errorf("call %d: source %q does not match the random felhom-guest-hook-*.sh pattern", i, src)
}
if fixedName.MatchString(src) {
t.Errorf("call %d: source %q is the FIXED predictable temp name (B1 TOCTOU)", i, src)
}
srcs = append(srcs, src)
}
if srcs[0] == srcs[1] {
t.Errorf("two consecutive installs staged through the SAME source path %q — must be random per call", srcs[0])
}
// Non-hollow: the staged file must actually carry the snippet body at install time.
for i, c := range r.srcContent {
if c != snippetBody {
t.Errorf("call %d: staged content is not the snippet body (got %d bytes)", i, len(c))
}
}
// And the temp is cleaned up after.
for _, src := range srcs {
if _, err := os.Stat(src); err == nil {
t.Errorf("staged temp %q left behind (defer os.Remove missing)", src)
}
l := filepath.Join(d, "link.sh")
_ = os.Symlink(p, l)
if err := SnippetReady(l); err == nil {
t.Fatal("a symlink read as the hook")
}
}
// B2 Scenario D (DRILL-day0-cleanroom-2026-07-03): on a fresh PVE, /var/lib/vz/snippets does not
// exist and `install` (no -D) cannot create it — the drill saw
// `install: cannot create regular file … No such file or directory` and the guest silently got no
// pre-start self-heal hook. InstallSnippet must therefore issue a `mkdir -p <SnippetDir>` fenced op
// BEFORE the `install` op. Pre-fix wrong outcome: no mkdir call at all — only the doomed install.
func TestInstallSnippet_EnsuresSnippetsDirFirst(t *testing.T) {
r := &recordingRunner{}
if err := InstallSnippet(context.Background(), r); err != nil {
t.Fatalf("InstallSnippet: %v", err)
}
mkdirIdx, installIdx := -1, -1
for i, call := range r.calls {
switch call[0] {
case "mkdir":
if mkdirIdx == -1 {
mkdirIdx = i
want := []string{"mkdir", "-p", SnippetDir}
if len(call) != 3 || call[1] != want[1] || call[2] != want[2] {
t.Errorf("mkdir vector = %v, want %v (the sudoers fence matches exactly this argv)", call, want)
}
}
case "install":
if installIdx == -1 {
installIdx = i
}
}
}
if mkdirIdx == -1 {
t.Fatalf("no `mkdir -p %s` op issued — on a fresh box the snippet install fails ENOENT (B2); calls: %v", SnippetDir, r.calls)
}
if installIdx == -1 {
t.Fatalf("no install op issued; calls: %v", r.calls)
}
if mkdirIdx > installIdx {
t.Fatalf("mkdir (call %d) must PRECEDE install (call %d) — order: %v", mkdirIdx, installIdx, r.calls)
// The bundle's copy is byte-identical to the body the agent checks against.
func TestSnippetEqualsTheBundle(t *testing.T) {
got, err := os.ReadFile(filepath.Join("..", "..", "configs", "felhom-guest-hook.sh"))
if err != nil || string(got) != snippetBody {
t.Fatalf("configs/felhom-guest-hook.sh differs from snippetBody (err %v)", err)
}
}
+25
View File
@@ -0,0 +1,25 @@
package hub
import (
"os"
"path/filepath"
"testing"
)
// R-840: the agent reports the bundle record itself — "none" when no bundle ever reached the box (so an old wrapper
// cannot hide that), "unknown" when it cannot be read, else the version and sha the root wrapper recorded.
func TestReadBundleRecord(t *testing.T) {
dir := t.TempDir()
p := filepath.Join(dir, "config-bundle.json")
if got := string(readBundleRecord(p)); got != `{"version":"none"}` {
t.Fatalf("absent: %s", got)
}
os.WriteFile(p, []byte("{broken"), 0o644)
if got := string(readBundleRecord(p)); got != `{"version":"unknown"}` {
t.Fatalf("broken: %s", got)
}
os.WriteFile(p, []byte(`{"format":1,"agent_version":"0.143.0","bundle_sha256":"abc","installed_at":"2026-10-04T20:00:00Z","authority":"signed","files":{"/x":"y"}}`), 0o644)
if got := string(readBundleRecord(p)); got != `{"authority":"signed","bundle_sha256":"abc","installed_at":"2026-10-04T20:00:00Z","version":"0.143.0"}` {
t.Fatalf("record: %s", got)
}
}
+151 -25
View File
@@ -4,10 +4,13 @@ import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"log/slog"
"os"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
@@ -87,32 +90,43 @@ type GuestNetReporter interface {
GuestNetStatus(ctx context.Context) *GuestNetStatus
}
// GuestDiskTrimReporter is the R-444 seam the weekly trim job plugs into (same consumer-side pattern — hub does not
// import fstrim). nil (feature not wired) → no guest_disk_trim stanza.
type GuestDiskTrimReporter interface {
GuestDiskTrimStatus(ctx context.Context) *GuestDiskTrimStatus
}
// Collector builds a HostReport from read-only sources. All deps are behind narrow
// interfaces for unit testing.
type Collector struct {
px proxmoxReader
cf CloudflaredProber
storage StorageObserver
backups BackupReporter
restoreTests RestoreTestReporter
provenTests ProvenRestoreTestReporter
pbs PBSReporter
temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp)
capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty)
leafFP string // v0.48.0: served local-API leaf fp (static per process; "" when local API disabled)
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
oob OOBReporter // H1: operator-access health (nil → stanza omitted)
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
hostID string
agentVersion string
logger *slog.Logger
now func() time.Time
px proxmoxReader
cf CloudflaredProber
storage StorageObserver
backups BackupReporter
restoreTests RestoreTestReporter
provenTests ProvenRestoreTestReporter
pbs PBSReporter
temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp)
capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty)
leafFP string // v0.48.0: served local-API leaf fp (static per process; "" when local API disabled)
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
diskTrim GuestDiskTrimReporter // R-444: weekly guest disk trim (nil → stanza omitted)
foreignKey ForeignKeyArchiveReporter // R-366 slice 2 (nil → omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
oob OOBReporter // H1: operator-access health (nil → stanza omitted)
system SystemReporter // R-852: the box versions (nil → API fields only)
bundleRecordPath string // R-840: test seam; "" = BundleRecordPath
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
hostID string
agentVersion string
selfSHA func() string // R-349: sha256 of the running binary; default runningBinarySHA256
logger *slog.Logger
now func() time.Time
}
// NewCollector builds a collector. hostID echoes config.Hub.HostID; agentVersion is
@@ -131,6 +145,7 @@ func NewCollector(px proxmoxReader, cf CloudflaredProber, storage StorageObserve
temp: SysfsTempReader{}, // slice 9: real sysfs reader by default; tests inject a fake
hostID: hostID,
agentVersion: agentVersion,
selfSHA: runningBinarySHA256,
logger: logger,
now: func() time.Time { return time.Now().UTC() },
}
@@ -214,6 +229,24 @@ func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
return c
}
// ForeignKeyArchiveReporter is the R-366 slice-2 seam: the restore-test's ledger of archives written with another
// key. nil = not evaluated yet (the stanza is omitted and the hub keeps its state).
type ForeignKeyArchiveReporter interface {
ForeignKeyArchives(ctx context.Context) *ForeignKeyArchivesStanza
}
// SetForeignKeyArchiveReporter wires the restore-test's foreign-key ledger (R-366 slice 2; nil-safe → omitted).
func (c *Collector) SetForeignKeyArchiveReporter(r ForeignKeyArchiveReporter) *Collector {
c.foreignKey = r
return c
}
// SetGuestDiskTrimReporter wires the R-444 weekly trim job as a report source (nil-safe → stanza omitted).
func (c *Collector) SetGuestDiskTrimReporter(r GuestDiskTrimReporter) *Collector {
c.diskTrim = r
return c
}
// SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side
// pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report.
type SelfUpdateReporter interface {
@@ -248,6 +281,69 @@ type OOBReporter interface {
}
// SetOOBReporter wires the operator-access health source (H1; nil-safe → stanza omitted).
// SystemReporter reads the box's versions (R-852): the customer guest's vmid and the wrapper's raw facts.
type SystemReporter interface {
SystemFacts(ctx context.Context) (vmid int, facts json.RawMessage, err error)
}
// SetSystemReporter wires the facts read (agent v0.142.0). Without it the stanza carries the Proxmox API fields only.
func (c *Collector) SetSystemReporter(r SystemReporter) *Collector {
c.system = r
return c
}
// BundleRecordPath is the root-owned record felhom-os-apply writes after a config bundle installs (R-840).
const BundleRecordPath = "/etc/felhom/config-bundle.json"
// readBundleRecord returns the record's summary: version/sha/installed_at, "none" when the file is absent, "unknown"
// when it cannot be read or parsed (never a guess).
func readBundleRecord(path string) json.RawMessage {
if path == "" {
path = BundleRecordPath
}
b, err := os.ReadFile(path)
if os.IsNotExist(err) {
return json.RawMessage(`{"version":"none"}`)
}
var rec struct {
AgentVersion string `json:"agent_version"`
BundleSHA256 string `json:"bundle_sha256"`
InstalledAt string `json:"installed_at"`
Authority string `json:"authority"`
}
if err != nil || json.Unmarshal(b, &rec) != nil || rec.AgentVersion == "" {
return json.RawMessage(`{"version":"unknown"}`)
}
out, _ := json.Marshal(map[string]string{"version": rec.AgentVersion, "bundle_sha256": rec.BundleSHA256,
"installed_at": rec.InstalledAt, "authority": rec.Authority})
return out
}
func unknownIfEmpty(s string) string {
if strings.TrimSpace(s) == "" {
return "unknown"
}
return s
}
// systemInfo builds the `system` stanza. Never fatal: a failed facts read is FactsError, the API fields stay.
func (c *Collector) systemInfo(ctx context.Context, ns proxmox.NodeStatus) *SystemInfo {
si := &SystemInfo{PVEVersion: unknownIfEmpty(ns.PVEVersion), KernelVersion: unknownIfEmpty(ns.KVersion),
ReadAt: c.now().Format(time.RFC3339)}
si.ConfigBundle = readBundleRecord(c.bundleRecordPath)
if c.system == nil {
si.FactsError = "no facts reader wired"
return si
}
vmid, f, err := c.system.SystemFacts(ctx)
si.VMID, si.Facts = vmid, f
if err != nil {
si.FactsError = err.Error()
c.logger.Debug("hub: system facts unavailable", "err", err)
}
return si
}
func (c *Collector) SetOOBReporter(o OOBReporter) *Collector {
c.oob = o
return c
@@ -270,6 +366,7 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
HostID: c.hostID,
ReportedAt: c.now().Format(time.RFC3339),
AgentVersion: c.agentVersion,
AgentSHA256: c.agentSHA256(),
Host: host,
Guests: c.collectGuests(ctx),
// storage_targets populated this slice (slice 5) via the observer; the rest stay
@@ -284,6 +381,7 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
Capabilities: c.capabilities(ctx),
LeafFingerprint: c.leafFP,
Addresses: c.collectAddresses(),
System: c.systemInfo(ctx, ns),
}
// DR recipe host-half — derived from the just-collected guest/storage/PBS facts (no new reads).
// Secret-free by construction (identifiers/intents/sizes/coordinates only).
@@ -305,6 +403,14 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
}
// R-444: the last weekly trim result per guest (nil reporter = not wired → stanza omitted).
if c.diskTrim != nil {
report.GuestDiskTrim = c.diskTrim.GuestDiskTrimStatus(ctx)
}
// R-366 slice 2: archives the restore-test skipped as another key's (nil → not evaluated yet → omitted).
if c.foreignKey != nil {
report.ForeignKeyArchives = c.foreignKey.ForeignKeyArchives(ctx)
}
// D1: agent self-update pending status (nil reporter → pending=false, the steady state).
if c.selfUpdate != nil {
report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending()
@@ -363,8 +469,28 @@ const pbsWrapperPath = "/usr/local/sbin/felhom-pbs-apply"
// unreadable file yields "", which the hub reads as UNKNOWN rather than as drift — a host that
// legitimately has no DR wrapper must not light up amber. The file is 0755, so no privilege is
// needed to read it.
func pbsWrapperSHA256() string {
f, err := os.Open(pbsWrapperPath)
func pbsWrapperSHA256() string { return fileSHA256(pbsWrapperPath) }
// selfExePath is the running binary as the kernel holds it. /proc/self/exe, not the installed path:
// after an A/B flip the file at /usr/local/bin/felhom-agent may already be the NEXT binary while this
// process still runs the old one, and the report must describe what runs (R-349). Test seam.
var selfExePath = "/proc/self/exe"
// runningBinarySHA256 hashes the running binary ONCE per process — the bytes cannot change under a
// running process, and re-hashing ~20 MB every report cycle buys nothing. A failed read is cached as
// "" (UNKNOWN); it never fails the report.
var runningBinarySHA256 = sync.OnceValue(func() string { return fileSHA256(selfExePath) })
func (c *Collector) agentSHA256() string {
if c.selfSHA == nil {
return ""
}
return c.selfSHA()
}
// fileSHA256 is the hex sha256 of a file's bytes, or "" when it cannot be read.
func fileSHA256(path string) string {
f, err := os.Open(path)
if err != nil {
return ""
}
+64
View File
@@ -0,0 +1,64 @@
package hub
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
)
// R-349: the report carries the sha256 of the binary that is RUNNING, so the hub can tell a
// hand-built proof binary from the vouched artifact of the same version string. The consequence
// asserted: the wire field equals the hash of this very test binary's bytes (read independently via
// os.Executable, a different channel from /proc/self/exe), and it is on the wire as agent_sha256.
func TestCollect_AgentSHA256IsTheRunningBinary(t *testing.T) {
exe, err := os.Executable()
if err != nil {
t.Skipf("os.Executable: %v", err)
}
raw, err := os.ReadFile(exe)
if err != nil {
t.Fatalf("read own binary: %v", err)
}
sum := sha256.Sum256(raw)
want := hex.EncodeToString(sum[:])
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
if r.AgentSHA256 != want {
t.Fatalf("agent_sha256 = %q, want the running binary's %q", r.AgentSHA256, want)
}
b, err := json.Marshal(r)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(b), `"agent_sha256":"`+want+`"`) {
t.Fatalf("agent_sha256 not on the wire: %s", b)
}
}
// An unreadable binary is UNKNOWN (empty, omitted) — never a made-up hash, never a failed report.
func TestFileSHA256_UnreadableIsEmpty(t *testing.T) {
if got := fileSHA256(filepath.Join(t.TempDir(), "absent")); got != "" {
t.Fatalf("absent file hashed to %q, want empty", got)
}
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
c.selfSHA = func() string { return "" }
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect must not fail on an unreadable binary: %v", err)
}
b, _ := json.Marshal(r)
if strings.Contains(string(b), "agent_sha256") {
t.Fatalf("empty agent_sha256 must be omitted: %s", b)
}
}
+54
View File
@@ -0,0 +1,54 @@
package hub
import (
"context"
"encoding/json"
"testing"
)
// R-444: the guest_disk_trim stanza must reach a report built through the PRODUCTION collect path, be absent from
// the wire when the job is not wired, and carry the keys the hub's System page reads.
type fakeDiskTrim struct{ st *GuestDiskTrimStatus }
func (f fakeDiskTrim) GuestDiskTrimStatus(context.Context) *GuestDiskTrimStatus { return f.st }
func TestCollect_GuestDiskTrim(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.150.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ := json.Marshal(r)
var m map[string]any
_ = json.Unmarshal(b, &m)
if _, ok := m["guest_disk_trim"]; ok {
t.Fatalf("guest_disk_trim on the wire with no reporter wired: %s", b)
}
c.SetGuestDiskTrimReporter(fakeDiskTrim{st: &GuestDiskTrimStatus{Schedule: "weekly", Guests: []GuestDiskTrim{{
VMID: 9201, LastAttemptAt: "2026-10-07T08:30:00Z", OK: true, BytesTrimmed: 90143313920, Mounts: 2,
DurationSeconds: 24.4, LastOKAt: "2026-10-07T08:30:00Z",
}}}})
r, err = c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ = json.Marshal(r)
m = nil
_ = json.Unmarshal(b, &m)
dt, ok := m["guest_disk_trim"].(map[string]any)
if !ok || dt["schedule"] != "weekly" {
t.Fatalf("guest_disk_trim missing or wrong on the wire: %s", b)
}
g := dt["guests"].([]any)[0].(map[string]any)
for _, k := range []string{"vmid", "last_attempt_at", "ok", "bytes_trimmed", "mounts", "duration_seconds", "last_ok_at"} {
if _, ok := g[k]; !ok {
t.Fatalf("guest_disk_trim.guests[0] lacks %q: %v", k, g)
}
}
if g["bytes_trimmed"] != float64(90143313920) || g["ok"] != true {
t.Fatalf("values did not survive the round trip: %v", g)
}
}
+9 -7
View File
@@ -54,11 +54,13 @@ const (
DRReasonNoPBSStorage = "no_pbs_storage_observed"
)
// PBSRootNamespace is how the recipe spells PBS's root namespace. The PBS API spells it as the EMPTY
// string (and `pct restore --ns root` would name a namespace that does not exist) — "root" is a display
// convention this wire has always used, kept here so the field's meaning did not change under R-106.
// Only a box with no `namespace` line in its pbs storage.cfg stanza ever emits it.
const PBSRootNamespace = "root"
// PBSRootNamespace is how the recipe spells PBS's root namespace: the EMPTY string, PBS's own spelling (R-124,
// agent v0.147.0). It used to be the display word "root", which no PBS namespace is named — an operator pasting it
// into `proxmox-backup-client … --ns root` during a real recovery got a failure. An empty namespace is ambiguous on
// its own, so READ IT WITH namespace_state: resolved + "" = the root namespace (pass no --ns, or --ns ""); unknown +
// "" = the agent could not tell. Only a box with no `namespace` line in its pbs storage.cfg stanza emits it.
// Pinned by TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown and TestR124_RootNamespaceOnTheWireIsPBSSpelling.
const PBSRootNamespace = ""
// DRRecipeHostHalf is the agent-emitted half (guest/drive/storage/PBS scaffolding). Derived entirely
// from facts the report already collects — no new privileged reads.
@@ -123,8 +125,8 @@ type DRPBSCoord struct {
RepoID string `json:"repo_id"` // the PVE pbs storage id (e.g. "felhom-pbs") — not a token
// Namespace is the PBS namespace the restore targets, resolved from the pbs storage's storage.cfg
// stanza — the same field `vzdump --storage <pbs>` makes PVE read, so the recipe cannot disagree
// with the backup that produced the snapshot. PBSRootNamespace when the box has no namespace
// configured; "" when NamespaceState is unknown.
// with the backup that produced the snapshot. PBSRootNamespace ("", PBS's spelling, R-124) when the box has no
// namespace configured; also "" when NamespaceState is unknown — consult NamespaceState.
//
// R-106: this used to come from the listed snapshot's own `ns`, which PBS does not echo per item once
// the request is already namespace-scoped via `?ns=` (internal/pbs/client.go). The field was
+39 -3
View File
@@ -63,8 +63,8 @@ func TestBuildDRRecipeHostHalf(t *testing.T) {
t.Error("felhom-flash (local-dir user-data drive) missing from drives")
}
// pbs: latest snapshot's coords + the pbs storage id as repo_id.
if h.PBS == nil || h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != "root" || h.PBS.LatestSnapshotID != "9201" {
t.Errorf("pbs coord = %+v, want repo felhom-pbs/root/9201", h.PBS)
if h.PBS == nil || h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != PBSRootNamespace || h.PBS.LatestSnapshotID != "9201" {
t.Errorf("pbs coord = %+v, want repo felhom-pbs, the root namespace (\"\", R-124), snapshot 9201", h.PBS)
}
}
@@ -252,7 +252,7 @@ func TestDRRecipe_PBSNamespaceIsThePerCustomerOne(t *testing.T) {
}
// TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown: a box with a pbs storage and NO namespace line is
// genuinely in the root namespace. That is an answer, not a gap — it must read resolved/"root", so the
// genuinely in the root namespace. That is an answer, not a gap — it must read resolved/"" (PBS's spelling, R-124), so the
// honest root case is never confused with "I could not tell".
func TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown(t *testing.T) {
h := BuildDRRecipeHostHalf(nil,
@@ -453,3 +453,39 @@ func assertNoSecretKeys(t *testing.T, jsonBytes []byte) {
}
walk("<root>", v)
}
// R-124: on the WIRE the root namespace is PBS's own spelling — an empty string, present (not omitted), beside
// namespace_state "resolved". The display word "root" names no PBS namespace, and `--ns root` fails in a recovery.
// RED-PROOF: set PBSRootNamespace back to "root" → this test fails.
func TestR124_RootNamespaceOnTheWireIsPBSSpelling(t *testing.T) {
h := BuildDRRecipeHostHalf(nil,
[]StorageTarget{{Name: "felhom-pbs", Type: StorageTypePBS, Content: "backup", PBSNamespace: ""}},
capturedDemoFelhomSnapshots(),
ConfiguredBackupTarget{StorageID: "felhom-pbs", Known: true})
b, err := json.Marshal(h.PBS)
if err != nil {
t.Fatal(err)
}
var m map[string]any
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
ns, present := m["namespace"]
if !present {
t.Fatalf("namespace key missing from %s — an omitted key reads as 'unknown', not 'root'", b)
}
if ns != "" {
t.Fatalf("root namespace on the wire = %q, want \"\" (PBS's spelling; no namespace is named %q)", ns, ns)
}
if m["namespace_state"] != DRStateResolved {
t.Fatalf("namespace_state = %v, want %q beside the empty root namespace", m["namespace_state"], DRStateResolved)
}
// A configured namespace still passes through unchanged.
h2 := BuildDRRecipeHostHalf(nil,
[]StorageTarget{{Name: "felhom-pbs", Type: StorageTypePBS, Content: "backup", PBSNamespace: "demo-felhom"}},
capturedDemoFelhomSnapshots(),
ConfiguredBackupTarget{StorageID: "felhom-pbs", Known: true})
if h2.PBS.Namespace != "demo-felhom" {
t.Fatalf("configured namespace = %q, want demo-felhom", h2.PBS.Namespace)
}
}
+99
View File
@@ -18,6 +18,16 @@ type HostReport struct {
HostID string `json:"host_id"` // echoes config.Hub.HostID
ReportedAt string `json:"reported_at"` // RFC3339, agent clock
AgentVersion string `json:"agent_version"`
// AgentSHA256 is the sha256 of the binary this process is RUNNING (read through /proc/self/exe,
// once per process), R-349. The version string cannot tell a hand-built proof binary from the
// published, vouched artifact of the same version — same source, different bytes (`-trimpath
// -buildvcs=false` in release-agent.sh) — so self-update sees "already installed" and never
// corrects it. Reporting the bytes lets the hub compare against the vouched agent_sha256, the
// same mechanism host.wrapper_sha256 is for the PBS wrapper (R-50b(a)).
//
// Empty = unreadable, which the hub must treat as UNKNOWN, never as drift. Pinned by
// TestCollect_AgentSHA256IsTheRunningBinary.
AgentSHA256 string `json:"agent_sha256,omitempty"`
Host HostMetrics `json:"host"`
Guests []Guest `json:"guests"`
@@ -89,6 +99,12 @@ type HostReport struct {
// hub-schema change and are absent when the reporter is not wired.
MgmtPlane *MgmtPlaneStatus `json:"mgmt_plane,omitempty"`
// System is the box's versions for the hub's System page (agent v0.142.0, R-852, `09` decision 89): Proxmox and the
// running kernel from the Proxmox API, and the wrapper's read-only facts (host Debian, next-boot kernel, held
// packages, taint, the crash guard; guest Debian, Docker engine, containerd, live-restore). A value nobody could
// read is "unknown", never empty and never guessed. The hub v0.132.0 consumes it (hosts + System pages).
System *SystemInfo `json:"system,omitempty"`
// PBSDR is the PBS-DR-tier bridge status stanza (slice 2). Present only when the pbsdr
// consumer is wired. `consumed_failed` is the LOUD persistent state: the one-time token
// secret was consumed but the apply failed afterwards — the secret is burned, the bridge
@@ -108,6 +124,17 @@ type HostReport struct {
// on HostReport would have been the only report block named against that convention.
GuestNet *GuestNetStatus `json:"guest_net,omitempty"`
// GuestDiskTrim is the weekly guest disk trim stanza (R-444, `09` §3 decision 139): the schedule and, per owned
// guest, the LAST trim result as persisted by the agent (it survives a restart). Present only when the trim job is
// wired; an empty `guests` list means the job runs and no guest has been trimmed yet. No secret.
GuestDiskTrim *GuestDiskTrimStatus `json:"guest_disk_trim,omitempty"`
// ForeignKeyArchives (R-366 slice 2, `09` §3 decision 168): per backup tier, the whole-guest archives the
// restore-test SKIPPED because they were written with another key (an earlier install of this box). This box
// cannot open them; the hub turns a CHANGE of this list into one operator event. Absent = not evaluated yet since
// the agent started (the hub keeps its last state); `tiers: []` = evaluated, none found.
ForeignKeyArchives *ForeignKeyArchivesStanza `json:"foreign_key_archives,omitempty"`
// LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent
// mirror of the controller's report log_tails channel. Present ONLY on the heartbeat
// right after the control envelope requested it (log_tail_requested); consume-once on
@@ -192,6 +219,27 @@ type GuestNetGuest struct {
Message string `json:"message,omitempty"`
}
// GuestDiskTrimStatus is the R-444 weekly trim stanza. `schedule` is a plain description of when the job runs (local
// time of the host); `guests` holds one entry per owned guest that has had at least one trim attempt.
type GuestDiskTrimStatus struct {
Schedule string `json:"schedule"`
Guests []GuestDiskTrim `json:"guests,omitempty"`
}
// GuestDiskTrim is one guest's LAST trim attempt. `ok` with `last_attempt_at` is the verdict of that attempt — never
// read the time alone as success; `last_ok_at` is the last attempt that succeeded ("" = never). `bytes_trimmed` is
// the sum of the "(N bytes) trimmed" lines `pct fstrim` printed, over `mounts` mount points.
type GuestDiskTrim struct {
VMID int `json:"vmid"`
LastAttemptAt string `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt string `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
}
type PBSDRStatus struct {
State string `json:"state"`
StorageID string `json:"storage_id,omitempty"`
@@ -241,6 +289,20 @@ type WireguardStatus struct {
AssignedIP string `json:"assigned_ip,omitempty"` // from the marker, e.g. "10.77.0.2/32"
}
// SystemInfo is the `system` stanza (see HostReport.System).
type SystemInfo struct {
PVEVersion string `json:"pve_version"` // GET /nodes/{node}/status pveversion
KernelVersion string `json:"kernel_version"` // GET /nodes/{node}/status kversion
VMID int `json:"vmid,omitempty"` // the customer guest the facts read
Facts json.RawMessage `json:"facts,omitempty"`
FactsError string `json:"facts_error,omitempty"`
ReadAt string `json:"read_at"`
// ConfigBundle is the box's root-owned config bundle record (R-840, agent v0.143.0), read by the agent itself from
// /etc/felhom/config-bundle.json (0644): {"version":"none"} on a box no bundle reached, so an OLD wrapper cannot
// hide it. The wrapper's facts carry the same record plus the drift (files changed by hand).
ConfigBundle json.RawMessage `json:"config_bundle,omitempty"`
}
// HostMetrics is the host block, sourced from proxmox NodeStatus.
type HostMetrics struct {
Node string `json:"node"`
@@ -398,6 +460,17 @@ type SmartSummary struct {
ReallocatedSectors *int `json:"reallocated_sectors"`
PendingSectors *int `json:"pending_sectors"`
OfflineUncorrectable *int `json:"offline_uncorrectable"`
// R-330 (disk health Phase 2): three more SATA raw counters. omitempty + pointer: absent (an
// older agent, an NVMe/USB device, or a drive that does not report the attribute) is OMITTED —
// unknown, never a zero (S-39). Wire only: no verdict reads them yet.
// 187 Reported_Uncorrect — the failing drive's most telling counter (normalized 1 vs thresh 0,
// raw 1001) while SMART still said PASSED.
// 188 Command_Timeout — some vendors PACK several counters into the 48-bit raw value, so the
// number is carried as reported and must not be compared across vendors.
// 199 UDMA_CRC_Error_Count — cabling / link errors, not the medium.
ReportedUncorrect *int64 `json:"reported_uncorrect,omitempty"`
CommandTimeout *int64 `json:"command_timeout,omitempty"`
UDMACRCErrors *int64 `json:"udma_crc_errors,omitempty"`
// NVMe attributes.
CriticalWarning *int `json:"critical_warning"`
@@ -563,6 +636,17 @@ type WireOSUpdate struct {
// HostRelease is the newest approved HOST release (hub v0.131.0, `11` §8 step 3) — a separate set: a version
// approved for the guest is not approved for the host by that fact alone.
HostRelease *WireOSRelease `json:"host_release,omitempty"`
// Kernel is the kernel lane's instruction for THIS box (R-836, `09` §3 decision 172, `11` §5.11): the kernel the
// household was told about, and whether tonight is a told night. Nil (an older hub, or no kernel due) = no kernel
// step. The hub sets Tonight only after the household's mail the day before went out — no mail, no step.
Kernel *WireKernelStep `json:"kernel,omitempty"`
}
// WireKernelStep is the hub's kernel-lane instruction (hub osupdates.KernelBlock — field-exact, cross-repo).
type WireKernelStep struct {
Kver string `json:"kver"` // e.g. "7.0.14-22-pve" — the wrapper refuses any other (R23)
Tonight bool `json:"tonight"` // the household was mailed the day before: tonight's leg may reboot
NotifiedAt string `json:"notified_at,omitempty"` // when that mail went out (RFC 3339), for the log
}
// WireOSRelease is an approved version set; Snapshot is the approval time (YYYYMMDDTHHMMSSZ) the wrapper uses
@@ -659,3 +743,18 @@ type WireRestoreDirective struct {
Archive string `json:"archive,omitempty"` // source archive/snapshot to restore from
VMID int `json:"vmid,omitempty"`
}
// ForeignKeyArchivesStanza wraps the per-tier list so "evaluated, none" (`tiers: []`) differs from "not evaluated"
// (the stanza absent) without a null on the wire.
type ForeignKeyArchivesStanza struct {
Tiers []ForeignKeyArchives `json:"tiers"`
}
// ForeignKeyArchives is one tier's count of archives written with another key (R-366 slice 2): the count and the
// newest/oldest archive time (RFC3339, UTC). No key material — the fingerprints stay on the box.
type ForeignKeyArchives struct {
Target string `json:"target"`
Count int `json:"count"`
Oldest string `json:"oldest"`
Newest string `json:"newest"`
}
+122
View File
@@ -0,0 +1,122 @@
package lanresolver
import (
"context"
"io"
"log/slog"
"os"
"path/filepath"
"strings"
"sync"
"testing"
)
// recRunner records every privileged command EnsureDnsmasq would run and succeeds — nothing reaches
// apt, systemctl or the root checker.
type recRunner struct {
mu sync.Mutex
calls []string
}
func (r *recRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
r.mu.Lock()
defer r.mu.Unlock()
r.calls = append(r.calls, strings.Join(append([]string{name}, args...), " "))
return nil, nil, nil
}
func (r *recRunner) RunStdin(ctx context.Context, _ io.Reader, name string, args ...string) ([]byte, []byte, error) {
return r.Run(ctx, name, args...)
}
func (r *recRunner) installed() bool {
for _, c := range r.calls {
if strings.HasPrefix(c, "apt-get install") && strings.HasSuffix(c, " dnsmasq") {
return true
}
}
return false
}
// fixtureRoot builds a fake host root holding exactly the given relative files and points the REAL
// probe at it for the test's duration.
func fixtureRoot(t *testing.T, files ...string) {
t.Helper()
root := t.TempDir()
for _, f := range files {
p := filepath.Join(root, f)
if err := os.MkdirAll(filepath.Dir(p), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(p, nil, 0o644); err != nil {
t.Fatal(err)
}
}
prev := hostRoot
hostRoot = root
t.Cleanup(func() { hostRoot = prev })
}
func ensure(t *testing.T) *recRunner {
t.Helper()
r := &recRunner{}
m := NewManager(r, "192.0.2.10", []string{"1.1.1.1"}, slog.New(slog.NewTextHandler(io.Discard, nil)))
if err := m.EnsureDnsmasq(context.Background()); err != nil {
t.Fatalf("EnsureDnsmasq: %v", err)
}
return r
}
// R-317: a host with `dnsmasq-base` (the /usr/sbin/dnsmasq binary) but WITHOUT the `dnsmasq` package
// (the service unit) must get the package installed — else the following `systemctl enable --now
// dnsmasq` hits a unit that does not exist and LAN name resolution silently never comes up.
//
// RED-PROOF: probe "usr/sbin/dnsmasq" instead of the unit paths in dnsmasqUnitInstalled → this fails
// with "install was skipped".
func TestEnsureDnsmasq_BinaryWithoutUnitInstalls(t *testing.T) {
fixtureRoot(t, "usr/sbin/dnsmasq")
r := ensure(t)
if !r.installed() {
t.Fatalf("install was skipped on a dnsmasq-base-only host (binary present, unit absent) — "+
"the enable that follows targets a missing unit (R-317). calls: %q", r.calls)
}
}
func TestEnsureDnsmasq_UnitPresentSkipsInstall(t *testing.T) {
for _, unit := range []string{"usr/lib/systemd/system/dnsmasq.service", "lib/systemd/system/dnsmasq.service"} {
t.Run(unit, func(t *testing.T) {
fixtureRoot(t, "usr/sbin/dnsmasq", unit)
if r := ensure(t); r.installed() {
t.Fatalf("apt-get install ran although the dnsmasq unit is present at %s: %q", unit, r.calls)
}
})
}
}
func TestEnsureDnsmasq_NothingPresentInstalls(t *testing.T) {
fixtureRoot(t)
if r := ensure(t); !r.installed() {
t.Fatalf("install skipped on a host with no dnsmasq at all: %q", r.calls)
}
}
// Production wiring for the hostRoot seam: the shipped probe resolves against the real root and asks
// about the unit the `dnsmasq` package owns — never the dnsmasq-base binary.
func TestEnsureDnsmasq_ProductionProbeIsTheUnit(t *testing.T) {
if hostRoot != "/" {
t.Fatalf("hostRoot default = %q, want \"/\" — the production probe would look in the wrong tree", hostRoot)
}
var sawUsrLib bool
for _, p := range dnsmasqUnitPaths {
full := filepath.Join(hostRoot, p)
if strings.HasSuffix(full, "/sbin/dnsmasq") || strings.HasSuffix(full, "/bin/dnsmasq") {
t.Errorf("probe path %s is the dnsmasq-base binary, not the dnsmasq unit (R-317)", full)
}
if full == "/usr/lib/systemd/system/dnsmasq.service" {
sawUsrLib = true
}
}
if !sawUsrLib {
t.Errorf("probe paths %q miss /usr/lib/systemd/system/dnsmasq.service (dpkg -S: owned by dnsmasq)", dnsmasqUnitPaths)
}
}
+35 -3
View File
@@ -33,6 +33,8 @@ const (
DropinDir = "/etc/dnsmasq.d"
// BaseDropin holds the host-wide listen/upstream config (one per host).
BaseDropin = "felhom-resolver-base.conf"
// PrivApply is the root content checker that installs a drop-in (R-861).
PrivApply = "/usr/local/sbin/felhom-priv-apply"
)
// RenderBase returns the host-wide dnsmasq drop-in: bind to the host LAN IP (+ loopback), no-resolv,
@@ -99,11 +101,35 @@ func NewManager(runner proxmox.Runner, hostIP string, upstreams []string, logger
}
}
// hostRoot is the filesystem root the install probe resolves against: "/" in production; a test
// points it at a fixture tree so the REAL probe runs against files it controls.
var hostRoot = "/"
// dnsmasqUnitPaths are where the `dnsmasq` package ships its systemd unit (Debian; /lib is the
// pre-usrmerge spelling). R-317: probe the UNIT, never /usr/sbin/dnsmasq — that binary belongs to
// `dnsmasq-base`, so a host carrying dnsmasq-base without dnsmasq used to skip the install and then
// `systemctl enable --now dnsmasq` failed against a unit that is not there (resolver never up).
// Pinned by TestEnsureDnsmasq_BinaryWithoutUnitInstalls.
var dnsmasqUnitPaths = []string{
"usr/lib/systemd/system/dnsmasq.service",
"lib/systemd/system/dnsmasq.service",
}
// dnsmasqUnitInstalled reports whether the dnsmasq service unit (the `dnsmasq` package) is present.
func dnsmasqUnitInstalled() bool {
for _, p := range dnsmasqUnitPaths {
if _, err := os.Stat(filepath.Join(hostRoot, p)); err == nil {
return true
}
}
return false
}
// EnsureDnsmasq makes dnsmasq present + enabled and writes the host base config. Idempotent: it
// installs the package only when absent, and writes the base drop-in only when its content changes.
func (m *Manager) EnsureDnsmasq(ctx context.Context) error {
if _, err := os.Stat("/usr/sbin/dnsmasq"); err != nil { // metadata read, no privilege needed
m.logger.Info("lanresolver: dnsmasq absent — installing")
if !dnsmasqUnitInstalled() { // metadata read, no privilege needed
m.logger.Info("lanresolver: dnsmasq service unit absent — installing")
if out, errOut, ierr := m.runner.Run(ctx, "apt-get", "install", "-y", "-q", "dnsmasq"); ierr != nil {
return fmt.Errorf("install dnsmasq: %s: %w", strings.TrimSpace(string(errOut))+string(out), ierr)
}
@@ -263,7 +289,13 @@ func (m *Manager) writeFileIfChanged(ctx context.Context, path, content, mode st
return false, fmt.Errorf("write temp: %w", err)
}
tmp.Close()
if _, errOut, err := m.runner.Run(ctx, "install", "-m", mode, tmpName, path); err != nil {
// R-861 (v0.146.0): a dnsmasq drop-in reaches /etc/dnsmasq.d only through the root checker, which allows exactly
// the lines RenderBase/RenderGuestDropin write (a `dhcp-script=` would run as root). Mode is fixed (0644) there.
_ = mode
if filepath.Dir(path) != DropinDir {
return false, fmt.Errorf("drop-in %s is not in %s", path, DropinDir)
}
if _, errOut, err := m.runner.Run(ctx, PrivApply, "dnsmasq", tmpName, filepath.Base(path)); err != nil {
return false, fmt.Errorf("install %s: %s: %w", path, strings.TrimSpace(string(errOut)), err)
}
return true, nil
@@ -0,0 +1,22 @@
package lanresolver
import (
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/privapplytest"
)
// R-861: both drop-ins the resolver renders are accepted by the root checker; a dhcp-script line is not.
func TestPrivApply_AcceptsTheRenderedDropins(t *testing.T) {
if got := privapplytest.Check(t, "dnsmasq", BaseDropin, RenderBase("192.168.0.104", []string{"1.1.1.1", "8.8.8.8"})); got != "OK" {
t.Errorf("base drop-in: %s", got)
}
if got := privapplytest.Check(t, "dnsmasq", DropinName("demo-hp"), RenderGuestDropin("demo-hp", "enkisfelhom.hu", "192.168.0.138")); got != "OK" {
t.Errorf("guest drop-in: %s", got)
}
evil := RenderGuestDropin("demo-hp", "enkisfelhom.hu", "192.168.0.138") + "dhcp-script=/var/lib/felhom-agent/x\n"
if got := privapplytest.Check(t, "dnsmasq", "felhom-x.conf", evil); !strings.HasPrefix(got, "REFUSED") {
t.Fatalf("control: dhcp-script was not refused: %s", got)
}
}
@@ -4,7 +4,6 @@ import (
"context"
"encoding/json"
"errors"
"io"
"log/slog"
"os"
"path/filepath"
@@ -54,8 +53,8 @@ func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string
}
return "", errors.New("supExec: unexpected args")
}
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) {
return "", errors.New("supExec: no stdin exec expected")
func (f *supExec) WriteControllerImage(context.Context, int, string) error {
return errors.New("supExec: no image write expected")
}
func (f *supExec) count(vmid int) int {
f.mu.Lock()
+9 -12
View File
@@ -4,7 +4,6 @@ import (
"context"
"encoding/json"
"fmt"
"io"
"log/slog"
"net/http"
"os"
@@ -41,9 +40,10 @@ func ValidControllerImage(ref string) bool { return controllerImageRe.MatchStrin
// faked in tests. The single seam the swap composes over (no hand-rolled pct).
type GuestExecutor interface {
GuestExec(ctx context.Context, vmid int, args ...string) (string, error)
// GuestExecStdin is GuestExec with the command's stdin fed from stdin — the swap write pipes the
// image ref into an in-guest `tee` (no shell vector).
GuestExecStdin(ctx context.Context, vmid int, stdin io.Reader, args ...string) (string, error)
// WriteControllerImage writes the image ref into the guest's /etc/felhom-controller-image through the ROOT
// verb `felhom-priv-apply controller-image <vmid>` (R-861 (a) A1, `09` §3 decision 165), which re-checks the
// ref against our registry + repository + x.y.z. The agent no longer holds a `tee` grant into the guest.
WriteControllerImage(ctx context.Context, vmid int, image string) error
}
// ControllerSwapState is the durable record of a swap (crash-safety + status). Written before the swap
@@ -141,14 +141,11 @@ func (c *ControllerSwapper) imagePresent(ctx context.Context, vmid int, image st
}
func (c *ControllerSwapper) writeImage(ctx context.Context, vmid int, image string) error {
// Non-root path: pipe the image ref into an in-guest `tee` over stdin — no shell, no
// interpolation, no `bash -c` (the only swap vector that would have needed an arbitrary-exec
// grant). The trailing "\n" makes the on-disk bytes byte-identical to the golden's
// `printf '%s\n'`; the bootstrap reads `IMAGE=$(cat …)` so the newline is stripped on read
// (spike SPIKE-controllerswap-narrow-grants-2026-06-29). image is strict-validated
// (controllerImageRe) upstream in Swap; defence-in-depth, the stdin path can't smuggle anyway.
_, err := c.exec.GuestExecStdin(ctx, vmid, strings.NewReader(image+"\n"), "tee", controllerImageFile)
return err
// R-861 (a) A1: the ROOT verb writes the file (it re-checks the ref — a compromised agent cannot hand the guest's
// bootstrap another image). The bytes are `image\n`, byte-identical to the golden's `printf '%s\n'`; the bootstrap
// reads `IMAGE=$(cat …)` so the newline is stripped on read. image is also strict-validated (controllerImageRe)
// upstream in Swap.
return c.exec.WriteControllerImage(ctx, vmid, image)
}
func (c *ControllerSwapper) restartBootstrap(ctx context.Context, vmid int) error {
+14 -28
View File
@@ -23,7 +23,7 @@ type fakeGuestExec struct {
present map[string]bool // images pulled into the guest
good map[string]bool // images that report healthy when running
containerImg string // image the running container currently has
teeStdin []string // raw bytes piped into each `tee` write (the swap's write vector)
teeStdin []string // image refs handed to WriteControllerImage (the root verb, R-861 (a) A1)
failRestart bool
noHealthBlock bool // if set, .State.Health is absent ("none")
restartCount int // F1: .RestartCount reported by docker inspect (a crash-looper has >0)
@@ -75,20 +75,15 @@ func (f *fakeGuestExec) GuestExec(_ context.Context, _ int, args ...string) (str
return "", fmt.Errorf("fake: unexpected exec %v", args)
}
// GuestExecStdin models the swap's write vector: `tee /etc/felhom-controller-image` with the image
// piped on stdin. It records the raw stdin bytes and sets the modeled file content (newline-stripped,
// as the bootstrap's `IMAGE=$(cat …)` read would see it).
func (f *fakeGuestExec) GuestExecStdin(_ context.Context, _ int, stdin io.Reader, args ...string) (string, error) {
// WriteControllerImage models the swap's write: the root verb `felhom-priv-apply controller-image <vmid>` (R-861
// (a) A1). It records the ref and sets the modeled file content, as the bootstrap's `IMAGE=$(cat …)` would read it.
func (f *fakeGuestExec) WriteControllerImage(_ context.Context, _ int, image string) error {
f.mu.Lock()
defer f.mu.Unlock()
f.calls = append(f.calls, args)
b, _ := io.ReadAll(stdin)
if len(args) >= 2 && args[0] == "tee" && args[1] == controllerImageFile {
f.teeStdin = append(f.teeStdin, string(b))
f.imageFile = strings.TrimSpace(string(b))
return string(b), nil // tee echoes stdin to stdout
}
return "", fmt.Errorf("fake: unexpected exec-stdin args=%v stdin=%q", args, string(b))
f.calls = append(f.calls, []string{"felhom-priv-apply", "controller-image", image})
f.teeStdin = append(f.teeStdin, image+"\n")
f.imageFile = image
return nil
}
// wrote reports whether the image was written via the stdin `tee` vector with the exact `image\n`
@@ -162,7 +157,7 @@ func TestControllerSwap_Happy(t *testing.T) {
// The write vector must be the stdin `tee` with byte-identical `image\n` and NO shell — the
// controllerswap.go writeImage rewrite. This would FAIL on the pre-change `bash -c "printf … >"` impl.
func TestControllerSwap_WriteViaStdinTee_NoShell(t *testing.T) {
func TestControllerSwap_WriteViaRootVerb_NoShell(t *testing.T) {
fe := &fakeGuestExec{
imageFile: prevImg,
present: map[string]bool{newImg: true},
@@ -173,22 +168,15 @@ func TestControllerSwap_WriteViaStdinTee_NoShell(t *testing.T) {
t.Fatalf("state = %q, want done", st.State)
}
if !fe.wrote(newImg) {
t.Errorf("expected a tee write of %q+\\n; teeStdin=%q", newImg, fe.teeStdin)
t.Errorf("expected the root verb to write %q; writes=%q", newImg, fe.teeStdin)
}
sawTee := false
for _, c := range fe.calls {
if len(c) >= 2 && c[0] == "tee" {
sawTee = true
if c[1] != controllerImageFile {
t.Errorf("tee target = %q, want fixed %q", c[1], controllerImageFile)
}
if len(c) >= 1 && c[0] == "tee" {
t.Errorf("the swap still uses an in-guest tee (R-861 (a) A1 removed that grant): %v", c)
}
}
if !sawTee {
t.Error("no tee call recorded — writeImage did not use the stdin tee vector")
}
if fe.usedShell() {
t.Errorf("swap used a shell vector (bash/-c/printf) — must be stdin tee only; calls=%v", fe.calls)
t.Errorf("swap used a shell vector (bash/-c/printf); calls=%v", fe.calls)
}
}
@@ -300,9 +288,7 @@ func (s *inspectScript) GuestExec(_ context.Context, _ int, args ...string) (str
}
return "", nil
}
func (s *inspectScript) GuestExecStdin(_ context.Context, _ int, _ io.Reader, _ ...string) (string, error) {
return "", nil
}
func (s *inspectScript) WriteControllerImage(context.Context, int, string) error { return nil }
func fastSwapper(exec GuestExecutor) *ControllerSwapper {
s := NewControllerSwapper(exec, "", discardLogger())
+102
View File
@@ -0,0 +1,102 @@
package localapi
import (
"encoding/json"
"errors"
"io"
"io/fs"
"net/http"
"os"
"time"
)
// GET /host/crash-guard (R-856, `09` §3 decision 143): what the host's crash guard
// (configs/felhom-crash-guard, `11` §5.9) recorded about the most recent HOST boot. The controller
// reads it once after it starts: when the host's last boot followed an UNCLEAN stop, its app mails
// wait ~15 minutes instead of the normal 90 s boot grace.
//
// Read-only and Proxmox-free: the agent reads the guard's state file (root-owned, 0644 — the
// non-root agent can read it) and passes four fields through. Host-wide, token-authed (any valid
// per-guest token sees the host's view, as GET /host/metrics does).
//
// NEVER an error page. A missing file (no guard installed, or no boot recorded yet), an unreadable
// one, or one that does not parse answers 200 with present:false — the controller reads that as
// UNKNOWN and keeps its normal boot grace. Pinned by TestR856_CrashGuard*.
// defaultCrashGuardStatePath is where configs/felhom-crash-guard writes its state (STATE_DIR there).
const defaultCrashGuardStatePath = "/var/lib/felhom-crash-guard/state.json"
// crashGuardStateMax bounds the read; the real file is well under 4 KiB.
const crashGuardStateMax = 1 << 20
// CrashGuardResponse is the data block of GET /host/crash-guard. Field names are the controller's
// agentapi.CrashGuardState (felhom-controller internal/agentapi/crashguard.go) — a wire contract,
// pinned by TestR856_CrashGuardWireMatchesControllerClient.
type CrashGuardResponse struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"` // RFC3339 UTC ("2006-01-02T15:04:05Z")
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// crashGuardFile is the subset of the guard's state.json the route passes through. Every other key
// (armed, boot_id, config, unclean_boots, last_trip, ...) is ignored.
type crashGuardFile struct {
LastBootAt string `json:"last_boot_at"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// readCrashGuardState reads and parses the guard's state file. ok=false on ANY failure (missing,
// unreadable, oversized, not a JSON object, a field of the wrong type); reason says which, for the log.
func readCrashGuardState(path string) (resp CrashGuardResponse, ok bool, reason string) {
f, err := os.Open(path)
if err != nil {
if errors.Is(err, fs.ErrNotExist) {
return resp, false, "no state file"
}
return resp, false, "unreadable: " + err.Error()
}
defer f.Close()
raw, err := io.ReadAll(io.LimitReader(f, crashGuardStateMax+1))
if err != nil {
return resp, false, "read: " + err.Error()
}
if len(raw) > crashGuardStateMax {
return resp, false, "state file too large"
}
var st crashGuardFile
// Unmarshal into a struct fails on a non-object top level (null decodes, so reject it below).
if err := json.Unmarshal(raw, &st); err != nil {
return resp, false, "unparseable: " + err.Error()
}
var probe map[string]json.RawMessage
if err := json.Unmarshal(raw, &probe); err != nil || probe == nil {
return resp, false, "unparseable: not a JSON object"
}
resp = CrashGuardResponse{Present: true, LastBootUnclean: st.LastBootUnclean, Tripped: st.Tripped}
// Normalise to RFC3339 UTC; an unparseable time passes through as-is (the controller reads an
// unparseable boot time as "not this start's boot" → its normal grace).
if t, perr := time.Parse(time.RFC3339, st.LastBootAt); perr == nil {
resp.LastBootAt = t.UTC().Format(time.RFC3339)
} else {
resp.LastBootAt = st.LastBootAt
}
return resp, true, ""
}
func (s *Server) handleCrashGuard(w http.ResponseWriter, r *http.Request, vmid int) {
path := s.crashGuardStatePath
if path == "" {
path = defaultCrashGuardStatePath
}
resp, ok, reason := readCrashGuardState(path)
if !ok {
s.logger.Debug("local-api: /host/crash-guard not present", "vmid", vmid, "reason", reason)
writeOK(w, CrashGuardResponse{Present: false})
return
}
s.logger.Debug("local-api: /host/crash-guard served", "vmid", vmid,
"last_boot_at", resp.LastBootAt, "last_boot_unclean", resp.LastBootUnclean, "tripped", resp.Tripped)
writeOK(w, resp)
}
+205
View File
@@ -0,0 +1,205 @@
package localapi
import (
"encoding/json"
"io"
"log/slog"
"net/http"
"os"
"path/filepath"
"testing"
)
// The shape of /var/lib/felhom-crash-guard/state.json as read on demo-hp on 2026-10-06 (values from
// that read where they matter; lists/objects kept to the same key set).
const crashGuardFixture = `{
"armed": true,
"boot_id": "3f1c0f1e-6a0b-4d7e-9b7a-0c2d4e6f8a1b",
"config": {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10},
"kernel_panic": 10,
"last_boot_at": "2026-10-05T07:56:41Z",
"last_boot_unclean": true,
"last_trip": {},
"rearmed_at": "2026-10-04T14:02:11Z",
"rearmed_by": "operator",
"tripped": false,
"unclean_boots": ["2026-10-05T07:56:41Z"],
"unclean_boots_24h": 1,
"unclean_boots_in_window": 1,
"updated_at": "2026-10-05T07:57:02Z",
"version": 1
}`
// controllerCrashGuardState is a COPY of the controller's wire type, felhom-controller
// controller/internal/agentapi/crashguard.go `CrashGuardState` (commit 8b13a5e) — same field names,
// same tags. If either side renames a key, the contract test below fails.
type controllerCrashGuardState struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
func newCrashGuardServer(t *testing.T, statePath string) http.Handler {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0",
Guests: &fakeGuests{},
Backups: &fakeBackups{},
Store: &fakeStore{},
Storage: fakeStorage{},
Tokens: staticTokens{"A": 8200, "B": 9300},
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatalf("new server: %v", err)
}
srv.crashGuardStatePath = statePath
return srv.Handler()
}
func writeCrashGuardFixture(t *testing.T, body string) string {
t.Helper()
p := filepath.Join(t.TempDir(), "state.json")
if err := os.WriteFile(p, []byte(body), 0o644); err != nil {
t.Fatal(err)
}
return p
}
// getCrashGuard calls the route and decodes the envelope with the CONTROLLER's type.
func getCrashGuard(t *testing.T, h http.Handler, token string) (int, controllerCrashGuardState, string) {
t.Helper()
w := do(t, h, "GET", "/host/crash-guard", token, "")
var env struct {
OK bool `json:"ok"`
Data controllerCrashGuardState `json:"data"`
}
if w.Code == http.StatusOK {
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil {
t.Fatalf("decode %q: %v", w.Body.String(), err)
}
if !env.OK {
t.Fatalf("ok=false: %s", w.Body.String())
}
}
return w.Code, env.Data, w.Body.String()
}
// A present state file (demo-hp's shape) passes the three facts through.
func TestR856_CrashGuardPresentFile(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, body)
}
want := controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: true, Tripped: false}
if st != want {
t.Fatalf("state = %+v, want %+v", st, want)
}
// A tripped, clean boot reads back as such (both bools are carried, not defaulted).
h = newCrashGuardServer(t, writeCrashGuardFixture(t,
`{"last_boot_at":"2026-10-05T09:56:41+02:00","last_boot_unclean":false,"tripped":true,"version":1}`))
_, st, _ = getCrashGuard(t, h, "B")
want = controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: false, Tripped: true}
if st != want {
t.Fatalf("offset time / tripped: state = %+v, want %+v (time normalised to UTC Z)", st, want)
}
}
// No state file (no guard on this host, or no boot recorded yet) → 200 present:false.
func TestR856_CrashGuardMissingFile(t *testing.T) {
h := newCrashGuardServer(t, filepath.Join(t.TempDir(), "absent", "state.json"))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("missing file: got %d, want 200 (%s)", code, body)
}
if st.Present || st.LastBootUnclean || st.Tripped || st.LastBootAt != "" {
t.Fatalf("missing file: state = %+v, want present:false and nothing else", st)
}
}
// A garbled file → 200 present:false, never a 5xx — every shape of garbage.
func TestR856_CrashGuardGarbageFile(t *testing.T) {
for name, body := range map[string]string{
"truncated": crashGuardFixture[:40],
"not json": "this is not json\n",
"empty": "",
"null": "null",
"array": `[{"last_boot_unclean":true}]`,
"wrong type": `{"last_boot_at":"2026-10-05T07:56:41Z","last_boot_unclean":"yes","tripped":false}`,
"lone brace": "{",
} {
t.Run(name, func(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, body))
code, st, raw := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, raw)
}
if st.Present || st.LastBootUnclean {
t.Fatalf("garbage %q read as %+v, want present:false", name, st)
}
})
}
// The path is a directory, not a file: unreadable → present:false, 200.
h := newCrashGuardServer(t, t.TempDir())
if code, st, raw := getCrashGuard(t, h, "A"); code != http.StatusOK || st.Present {
t.Fatalf("directory path: got %d %+v (%s), want 200 present:false", code, st, raw)
}
}
// No / unknown token → 401, like every sibling route; a cross-guest ?vmid= → 403.
func TestR856_CrashGuardRequiresGuestToken(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
for _, tok := range []string{"", "bogus"} {
w := do(t, h, "GET", "/host/crash-guard", tok, "")
if w.Code != http.StatusUnauthorized {
t.Fatalf("token %q: got %d, want 401", tok, w.Code)
}
if json.Valid(w.Body.Bytes()) {
var env struct {
Data controllerCrashGuardState `json:"data"`
}
_ = json.Unmarshal(w.Body.Bytes(), &env)
if env.Data.Present || env.Data.LastBootUnclean {
t.Fatalf("token %q: the refusal leaked the state: %s", tok, w.Body.String())
}
}
}
if w := do(t, h, "GET", "/host/crash-guard?vmid=9300", "A", ""); w.Code != http.StatusForbidden {
t.Fatalf("cross-guest query: got %d, want 403", w.Code)
}
}
// Wire contract: every key the controller's CrashGuardState decodes is emitted under exactly that
// name, and the agent emits no key the controller does not know.
func TestR856_CrashGuardWireMatchesControllerClient(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
w := do(t, h, "GET", "/host/crash-guard", "A", "")
var env struct {
OK bool `json:"ok"`
Data map[string]json.RawMessage `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil || !env.OK {
t.Fatalf("envelope: %v %s", err, w.Body.String())
}
want := []string{"present", "last_boot_at", "last_boot_unclean", "tripped"}
for _, k := range want {
if _, ok := env.Data[k]; !ok {
t.Errorf("agent does not emit %q, which the controller decodes (%s)", k, w.Body.String())
}
}
if len(env.Data) != len(want) {
t.Errorf("agent emits %d keys, controller knows %d: %s", len(env.Data), len(want), w.Body.String())
}
// And the agent's own type agrees with the controller's copy, field for field.
var mine CrashGuardResponse
var theirs controllerCrashGuardState
raw, _ := json.Marshal(env.Data)
_ = json.Unmarshal(raw, &mine)
_ = json.Unmarshal(raw, &theirs)
if (controllerCrashGuardState{mine.Present, mine.LastBootAt, mine.LastBootUnclean, mine.Tripped}) != theirs {
t.Errorf("agent %+v vs controller %+v", mine, theirs)
}
}
+20 -8
View File
@@ -398,9 +398,16 @@ func (s *Server) handleDisks(w http.ResponseWriter, r *http.Request, vmid int) {
di.Smart = &sm
}
}
if total, used, okc := statfsCapacity(d.MountPath); okc {
di.TotalBytes, di.UsedBytes = total, used
di.UsedFraction = float64(used) / float64(total)
// R-118: statfs ONLY while the drive's device is present. With the device gone the raw
// mountpoint reverts to a bare directory on the ROOT filesystem, and statfs would report
// pve-root's size as this drive's (measured: a 4 GB drive advertising 46 GiB). Same trap
// observe.go guards on the Observe path. Absent → capacity left zero (unknown), never root's.
// Pinned by TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity.
if s.devicePresent(d.MountPath) {
if total, used, okc := statfsCapacity(d.MountPath); okc {
di.TotalBytes, di.UsedBytes = total, used
di.UsedFraction = float64(used) / float64(total)
}
}
out = append(out, di)
}
@@ -761,6 +768,11 @@ type FormatResponse struct {
// signature — the customer authorizes the wipe of their own data drive.
NeedsConfirmation bool `json:"needs_confirmation,omitempty"`
DurableID string `json:"durable_id,omitempty"` // the durable id to confirm against (user-data)
// FSUUID (on Formatted) is the UUID of the filesystem the agent just made, verified against the
// bound durable id after mkfs (R-25). The caller mounts THIS — re-resolving a UUID from the /dev
// path later can name another disk if /dev re-enumerated. "" = not verified: the caller must not
// substitute a path-resolved guess silently.
FSUUID string `json:"fs_uuid,omitempty"`
// PendingOp is set on a SYSTEM/BACKUP data-bearing refusal — the exact op the operator must sign.
PendingOp *PendingOp `json:"pending_op,omitempty"`
}
@@ -808,7 +820,7 @@ func (s *Server) handleDiskFormatStatus(w http.ResponseWriter, r *http.Request,
writeOK(w, map[string]any{
"vmid": vmid, "phase": job.Phase, "device": job.Device, "fstype": job.FSType,
"durable_id": job.DurableID, "error": job.Error, "started_at": job.StartedAt, "updated_at": job.UpdatedAt,
"job_id": job.JobID,
"job_id": job.JobID, "fs_uuid": job.FSUUID, // R-25: "" until done + verified
})
}
@@ -871,7 +883,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
"format refused (device may have changed since inspection): "+rerr.Error())
return
}
done := s.startFormatDetached(device, blankDurable, req.FSType, true)
job, done := s.startFormatDetached(device, blankDurable, req.FSType, true)
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
if err == errFormatClientGone {
return // client gone; mkfs continues detached + the job record records the outcome
@@ -880,7 +892,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
writeErr(w, http.StatusBadGateway, "format failed: "+err.Error())
return
}
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, Reason: "blank device formatted " + req.FSType})
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, FSUUID: job.FSUUID, Reason: "blank device formatted " + req.FSType})
return
}
@@ -917,7 +929,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
// F20-BUG3: run the destructive mkfs DETACHED off s.baseCtx (bound durable id recorded for
// restart-recovery), so a request/client deadline can never SIGKILL it mid-write and corrupt the
// disk. We still wait to return the synchronous result (backward-compatible with the controller).
done := s.startFormatDetached(device, deviceDurable, req.FSType, false)
job, done := s.startFormatDetached(device, deviceDurable, req.FSType, false)
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
if err == errFormatClientGone {
return // client gone; the wipe continues detached + survives a restart via the job record
@@ -929,7 +941,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
s.logger.Warn("local-api: USER-DATA data-bearing format — CUSTOMER CONFIRMED (no operator signature)",
"vmid", vmid, "device", device, "durable_id", deviceDurable, "fstype", req.FSType, "why", probe.Reason())
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: true,
Role: string(role), DurableID: deviceDurable, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"})
Role: string(role), DurableID: deviceDurable, FSUUID: job.FSUUID, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"})
return
}
@@ -5,6 +5,7 @@ import (
"encoding/json"
"io"
"log/slog"
"runtime"
"strings"
"testing"
@@ -184,3 +185,39 @@ func TestDisks_DevicePresence_WireFieldIsFalseOnDeviceLoss(t *testing.T) {
t.Fatalf("the drive never reached the wire: %s", body)
}
}
// ── R-118 — an absent drive must not advertise the ROOT filesystem's capacity ───────────────────
// TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity drives the REAL statfsCapacity (no capacity
// seam): the registry drive's mount path is a real, bare temp directory — exactly what /mnt/<name>
// becomes once its device is gone (a plain directory on the host's filesystem). With the device absent
// the row must carry NO capacity; before R-118 the union path statfs'd that bare directory and reported
// the host filesystem's size and usage as the drive's (46 GiB at 9.2 % for a 4 GB drive, measured).
// The present half proves the test is not hollow: the same directory DOES yield capacity when the
// device is there, so a zero on the absent half is the guard's doing, not a statfs failure.
//
// RED-PROOF: drop the `if s.devicePresent(d.MountPath)` guard around statfsCapacity in disks.go → the
// absent subtest fails with "advertises ... bytes".
func TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity(t *testing.T) {
if runtime.GOOS != "linux" {
t.Skip("statfsCapacity is linux-only; production target is linux")
}
bare := t.TempDir()
known := []storage.KnownTarget{
{Name: "cel", Type: hub.StorageTypeUSB, MountPath: bare, DurableID: "uuid:4242", UUID: "4242"},
}
t.Run("absent", func(t *testing.T) {
di := diskByMount(t, presenceServer(t, nil, known, true, false), bare)
if di.TotalBytes != 0 || di.UsedBytes != 0 || di.UsedFraction != 0 {
t.Errorf("absent drive advertises total=%d used=%d frac=%.3f — that is the filesystem UNDER "+
"the bare mountpoint, not the drive (R-118)", di.TotalBytes, di.UsedBytes, di.UsedFraction)
}
})
t.Run("present", func(t *testing.T) {
di := diskByMount(t, presenceServer(t, nil, known, true, true), bare)
if di.TotalBytes <= 0 {
t.Errorf("present drive reports no capacity (total=%d) — the guard over-corrected and the "+
"size bar is gone for every healthy registry drive", di.TotalBytes)
}
})
}
+6
View File
@@ -28,6 +28,7 @@ type fakeDiskOps struct {
unmountCalls []string
candidates []storage.CandidateDisk // returned by ListCandidateDisks
candErr error
afterFormat *storage.DeviceProbe // R-25: when set, InspectDevice returns it once a format ran
}
func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.CandidateDisk, error) {
@@ -35,7 +36,12 @@ func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.Candidate
}
func (f *fakeDiskOps) InspectDevice(_ context.Context, device string) (storage.DeviceProbe, error) {
f.mu.Lock()
p := f.probe
if f.afterFormat != nil && len(f.formatCalls) > 0 {
p = *f.afterFormat
}
f.mu.Unlock()
p.Device = device
return p, f.inspectErr
}
+96
View File
@@ -0,0 +1,96 @@
package localapi
import (
"context"
"encoding/json"
"net/http"
"sync"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/storage"
)
// R-25: the format answer carries the UUID of the filesystem the agent JUST made, verified against the
// bound durable id after mkfs, so the controller mounts that filesystem rather than whatever the /dev
// path resolves to a few requests later. The consequence asserted: the UUID on the wire (and in the
// polled job record) is the new superblock's — and is EMPTY whenever the binding cannot be re-proved.
const newFSUUID = "0fc63daf-8483-4772-8e79-3d69d8477de4"
func confirmedFormat(t *testing.T, d *fakeDiskOps, srv *Server, fj *FormatJobStore) (string, *formatJob) {
t.Helper()
w := do(t, srv.Handler(), "POST", "/disks/format", "A", `{"device":"/dev/sdb1","fstype":"ext4","confirmed":true,"durable_id":"byid:wwn-/dev/sdb1"}`)
if w.Code != http.StatusOK {
t.Fatalf("confirmed format: %d (%s)", w.Code, w.Body.String())
}
var resp struct {
Data FormatResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatalf("decode: %v (%s)", err, w.Body.String())
}
if !resp.Data.Formatted {
t.Fatalf("not formatted: %s", w.Body.String())
}
return resp.Data.FSUUID, waitFormatPhase(t, fj, formatPhaseDone)
}
func confirmedGate() *fakeGate {
return &fakeGate{decision: WipeDecision{Allowed: true, Tier: "customer_confirmable", Reason: "customer_confirmed"}}
}
func TestFormat_ReportsNewFilesystemUUID(t *testing.T) {
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
fj := tempFormatStore(t)
srv := formatServer(t, d, confirmedGate(), fj)
got, job := confirmedFormat(t, d, srv, fj)
if got != newFSUUID {
t.Fatalf("response fs_uuid = %q, want the new filesystem's %q", got, newFSUUID)
}
if job.FSUUID != newFSUUID {
t.Fatalf("job record fs_uuid = %q, want %q (the polled status path)", job.FSUUID, newFSUUID)
}
}
// The node moved between mkfs and the read-back: the bound durable id now resolves elsewhere. The
// UUID must NOT be reported — reading it would name the other disk's filesystem.
func TestFormat_FSUUIDWithheldWhenDurableIDMoved(t *testing.T) {
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
fj := tempFormatStore(t)
srv := formatServer(t, d, confirmedGate(), fj)
var mu sync.Mutex
calls := 0
srv.reresolveWipe = func(_ context.Context, _ string) (string, error) {
mu.Lock()
defer mu.Unlock()
calls++
if calls == 1 {
return "/dev/sdb", nil // the pre-mkfs anti-retarget re-resolve
}
return "/dev/sdc", nil // after mkfs: the durable id now names another node
}
got, job := confirmedFormat(t, d, srv, fj)
if got != "" || job.FSUUID != "" {
t.Fatalf("fs_uuid reported after the durable id moved (response %q, job %q) — must be empty", got, job.FSUUID)
}
if calls < 2 {
t.Fatalf("the post-mkfs re-resolve never ran (calls=%d)", calls)
}
}
// The superblock did not read back as the requested filesystem → not verified → empty.
func TestFormat_FSUUIDWithheldOnFSTypeMismatch(t *testing.T) {
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "xfs", FSUUID: newFSUUID}}
fj := tempFormatStore(t)
srv := formatServer(t, d, confirmedGate(), fj)
got, job := confirmedFormat(t, d, srv, fj)
if got != "" || job.FSUUID != "" {
t.Fatalf("fs_uuid reported for a superblock of the wrong type (response %q, job %q)", got, job.FSUUID)
}
}
+39 -3
View File
@@ -22,6 +22,10 @@ type formatJob struct {
Blank bool `json:"blank,omitempty"` // audit D3: blank (benign) format — recovery re-checks STILL-blank, not data-bearing
Phase string `json:"phase"` // running | done | failed
Error string `json:"error,omitempty"`
// FSUUID is the filesystem UUID of the NEW filesystem, read by the agent right after mkfs on the
// device the bound durable id still resolves to (R-25). "" = not verified (the controller must not
// read that as a UUID). Set only on phase done.
FSUUID string `json:"fs_uuid,omitempty"`
StartedAt string `json:"started_at"`
UpdatedAt string `json:"updated_at"`
}
@@ -96,7 +100,10 @@ func (s *FormatJobStore) save(j *formatJob) error {
// runs to completion and records the outcome. device is the ALREADY anti-retarget-resolved device; the
// record carries durableID so a restart can re-resolve + re-run. blank marks a benign (blank-device)
// format, so restart recovery re-checks STILL-blank rather than data-bearing (audit D3).
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) <-chan error {
//
// The returned job may be read (FSUUID) only AFTER a value arrives on done — the goroutine writes it
// before the send, which is the happens-before edge.
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) (*formatJob, <-chan error) {
base := s.baseCtx
if base == nil {
base = context.Background()
@@ -116,10 +123,39 @@ func (s *Server) startFormatDetached(device, durableID, fstype string, blank boo
ctx, cancel := context.WithTimeout(base, 60*time.Minute)
defer cancel()
err := s.disks.Format(ctx, device, fstype)
if err == nil {
job.FSUUID = s.formattedFSUUID(ctx, device, durableID, fstype)
}
s.finishFormatJob(job, err)
done <- err
}()
return done
return job, done
}
// formattedFSUUID reads the UUID of the filesystem the agent has JUST made (R-25). The caller used to
// re-resolve the UUID from the mutable /dev path afterwards, over separate requests — a re-enumeration
// in that window could hand it ANOTHER disk's filesystem to mount. Here the bound durable id must still
// resolve to the very device that was formatted (and re-derive to the same id), and the superblock must
// carry the fstype that was asked for; anything else returns "" (not verified), never a guess.
func (s *Server) formattedFSUUID(ctx context.Context, device, durableID, fstype string) string {
if durableID == "" || s.reresolveWipe == nil {
return ""
}
// The device now holds a filesystem, so the data-bearing anti-retarget re-resolve is the right one.
now, err := s.reresolveWipe(ctx, durableID)
if err != nil || now != device {
s.logger.Warn("format: new filesystem UUID NOT reported — bound durable id no longer resolves to the formatted device",
"device", device, "durable_id", durableID, "resolves_to", now, "err", err)
return ""
}
probe, err := s.disks.InspectDevice(ctx, device)
if err != nil || !probe.Probed || probe.FSType != fstype || probe.FSUUID == "" {
s.logger.Warn("format: new filesystem UUID NOT reported — superblock did not read back as the requested filesystem",
"device", device, "want_fstype", fstype, "got_fstype", probe.FSType, "has_uuid", probe.FSUUID != "", "err", err)
return ""
}
s.logger.Info("format: new filesystem bound to its durable id", "device", device, "durable_id", durableID, "fs_uuid", probe.FSUUID)
return probe.FSUUID
}
// finishFormatJob updates the persisted record to done/failed.
@@ -172,7 +208,7 @@ func (s *Server) RecoverFormatJob(ctx context.Context) {
return
}
s.logger.Warn("format-job recover: re-running interrupted format detached", "durable_id", job.DurableID, "device", device, "fstype", job.FSType, "blank", job.Blank)
_ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion
_, _ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion
}
// nowFn returns the server clock (testable), defaulting to time.Now.
+10 -9
View File
@@ -3,7 +3,6 @@ package localapi
import (
"context"
"fmt"
"io"
"log/slog"
"strconv"
"strings"
@@ -126,14 +125,16 @@ func (b *GuestBinder) GuestExec(ctx context.Context, vmid int, args ...string) (
return string(out), nil
}
// GuestExecStdin is GuestExec with the in-guest command's stdin fed from stdin. The controller-swap
// write uses it to pipe the image ref into an in-guest `tee` (no shell vector, no interpolation),
// through the same fenced runner so the `sudo -n` prefix stays in one place.
func (b *GuestBinder) GuestExecStdin(ctx context.Context, vmid int, stdin io.Reader, args ...string) (string, error) {
pctArgs := append([]string{"exec", strconv.Itoa(vmid), "--"}, args...)
out, stderr, err := b.runner.RunStdin(ctx, stdin, "pct", pctArgs...)
// privApplyBin is the root content checker (R-861); its `controller-image` verb writes the guest's image file.
const privApplyBin = "/usr/local/sbin/felhom-priv-apply"
// WriteControllerImage pipes `image\n` to `felhom-priv-apply controller-image <vmid>` through the same fenced runner
// (the `sudo -n` prefix stays in one place). The verb checks the ref as root and writes the guest file itself
// (R-861 (a) A1, `09` §3 decision 165). Pinned by TestR861_WriteControllerImageUsesTheRootVerb.
func (b *GuestBinder) WriteControllerImage(ctx context.Context, vmid int, image string) error {
_, stderr, err := b.runner.RunStdin(ctx, strings.NewReader(image+"\n"), privApplyBin, "controller-image", strconv.Itoa(vmid))
if err != nil {
return string(out), fmt.Errorf("pct exec %d %v: %w: %s", vmid, args, err, strings.TrimSpace(string(stderr)))
return fmt.Errorf("felhom-priv-apply controller-image %d: %w: %s", vmid, err, strings.TrimSpace(string(stderr)))
}
return string(out), nil
return nil
}
+20 -50
View File
@@ -121,32 +121,33 @@ func (b *GuestBinder) EnsureSharedParent(ctx context.Context) error {
// only the script (the unit was unchanged), so the earlier unit-only gate never redeployed it —
// leaving hosts running the pre-fix script (no make-private), whose self-bind stays in root's shared
// peer group and DOUBLES every drive bind. Comparing both files closes that deploy gap.
if sharedParentInstallStale(sharedParentUnitPath, sharedParentScriptPath) {
if ierr := b.installSharedParentUnit(ctx); ierr != nil {
b.logger.Warn("shared-parent: boot-persistence (re)install failed (live setup OK; survives until host reboot)", "err", ierr)
}
if err := b.ensureSharedParentBoot(ctx, sharedParentUnitPath, sharedParentScriptPath, sharedParentWantsLink); err != nil {
b.logger.Warn("shared-parent: boot persistence not enabled (live setup OK; survives until host reboot)", "err", err)
}
return nil
}
// stageTemp writes content to a fresh random-named temp file (os.CreateTemp pattern — `*` is replaced
// by a random string) and returns its path. Caller removes it after the privileged `install`.
func stageTemp(pattern, content string) (string, error) {
f, err := os.CreateTemp("", pattern)
if err != nil {
return "", err
// sharedParentWantsLink exists once `systemctl enable felhom-shared-parent.service` ran (WantedBy=pve-guests.service).
const sharedParentWantsLink = "/etc/systemd/system/pve-guests.service.wants/felhom-shared-parent.service"
// ensureSharedParentBoot (R-861, agent v0.146.0) — the boot script and its unit are ROOT-OWNED FIXED files that arrive
// with the signed config bundle (configs/felhom-shared-parent.{sh,service}, byte-identical to the constants above,
// pinned by TestSharedParentFilesEqualTheBundle). The agent NEVER installs them any more: until v0.146.0 it installed
// them from /tmp, and a script a root unit runs at boot was a root shell for a compromised agent. Here it only checks
// them, and enables the unit (a fixed sudoers line) when the bundle has put it in place but nothing has enabled it — a
// fresh install. Missing or different files: one warning, nothing run (the next bundle brings them).
func (b *GuestBinder) ensureSharedParentBoot(ctx context.Context, unitPath, scriptPath, wantsLink string) error {
if sharedParentInstallStale(unitPath, scriptPath) {
return fmt.Errorf("%s or %s is missing or differs from this agent's — it arrives with the signed config bundle (agent_config_update); the agent does not install it (R-861)", scriptPath, unitPath)
}
name := f.Name()
if _, err := f.WriteString(content); err != nil {
f.Close()
os.Remove(name)
return "", err
if _, err := os.Lstat(wantsLink); err == nil {
return nil
}
if err := f.Close(); err != nil {
os.Remove(name)
return "", err
if err := b.run(ctx, "systemctl", "enable", "felhom-shared-parent.service"); err != nil {
return fmt.Errorf("enable unit: %w", err)
}
return name, nil
b.logger.Info("shared-parent: boot-persistence unit enabled (files from the config bundle)")
return nil
}
// sharedParentInstallStale reports whether the on-disk boot script OR unit is missing or differs from
@@ -162,37 +163,6 @@ func sharedParentInstallStale(unitPath, scriptPath string) bool {
return false
}
// installSharedParentUnit writes the script + unit (from agent-written temps) and enables the unit so the
// shared parent is re-established on every host boot before pve-guests. Idempotent. The temps are
// RANDOM-named os.CreateTemp files (audit B1): a fixed, predictable /tmp name could be pre-created by
// another local user and rewritten between our write and root's install (TOCTOU into a root-executed
// boot script). The final modes come from `install -m`, so the 0600 temps are fine.
func (b *GuestBinder) installSharedParentUnit(ctx context.Context) error {
tmpScript, err := stageTemp("felhom-shared-parent-*.sh", sharedParentScript)
if err != nil {
return fmt.Errorf("write temp script: %w", err)
}
defer os.Remove(tmpScript)
if err := b.run(ctx, "install", "-m", "0755", "--", tmpScript, sharedParentScriptPath); err != nil {
return fmt.Errorf("install script: %w", err)
}
tmpUnit, err := stageTemp("felhom-shared-parent-*.service", sharedParentUnit)
if err != nil {
return fmt.Errorf("write temp unit: %w", err)
}
defer os.Remove(tmpUnit)
if err := b.run(ctx, "install", "-m", "0644", "--", tmpUnit, sharedParentUnitPath); err != nil {
return fmt.Errorf("install unit: %w", err)
}
if err := b.run(ctx, "systemctl", "daemon-reload"); err != nil {
return fmt.Errorf("daemon-reload: %w", err)
}
if err := b.run(ctx, "systemctl", "enable", "felhom-shared-parent.service"); err != nil {
return fmt.Errorf("enable unit: %w", err)
}
return nil
}
// AttachDrive binds a drive's felhom-data namespace under the stable parent so it appears live in the
// guest at the returned stable path (via propagation — no pct, no reboot). `where` is the drive's RAW
// host PVE mount (/mnt/<name>); only `<where>/felhom-data` crosses into the guest (confinement). The
+58 -70
View File
@@ -4,94 +4,82 @@ import (
"context"
"io"
"os"
"regexp"
"path/filepath"
"testing"
)
// stagingRecorderRunner is a fake proxmox.Runner recording every call vector and snapshotting each
// install SOURCE file's content at call time (the deferred os.Remove erases it afterwards).
type stagingRecorderRunner struct {
calls [][]string
srcContent map[string]string // install dest → staged source content
}
// R-861 (agent v0.146.0): the shared-parent boot script and unit arrive with the signed config bundle; the agent never
// installs them. It checks them and enables the unit when nothing has.
//
// RED-PROOF (audits/hub-safety-2026-10-05/partF/red-proof.txt): put the old installSharedParentUnit call back (an
// `install` of a /tmp file) → TestSharedParentBoot_NeverInstalls fails.
func (r *stagingRecorderRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
type callRecorder struct{ calls [][]string }
func (r *callRecorder) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
r.calls = append(r.calls, append([]string{name}, args...))
if name == "install" && len(args) >= 2 {
src, dest := args[len(args)-2], args[len(args)-1]
if r.srcContent == nil {
r.srcContent = map[string]string{}
}
b, _ := os.ReadFile(src)
r.srcContent[dest] = string(b)
}
return nil, nil, nil
}
func (r *stagingRecorderRunner) RunStdin(ctx context.Context, _ io.Reader, name string, args ...string) ([]byte, []byte, error) {
func (r *callRecorder) RunStdin(ctx context.Context, _ io.Reader, name string, args ...string) ([]byte, []byte, error) {
return r.Run(ctx, name, args...)
}
// installSources returns the install-call source paths keyed by destination.
func (r *stagingRecorderRunner) installSources() map[string][]string {
out := map[string][]string{}
for _, c := range r.calls {
if c[0] == "install" && len(c) >= 3 {
src, dest := c[len(c)-2], c[len(c)-1]
out[dest] = append(out[dest], src)
}
func bootFiles(t *testing.T, script, unit string) (string, string, string) {
t.Helper()
d := t.TempDir()
sp, up := filepath.Join(d, "felhom-shared-parent.sh"), filepath.Join(d, "felhom-shared-parent.service")
if script != "" {
_ = os.WriteFile(sp, []byte(script), 0o755)
}
return out
if unit != "" {
_ = os.WriteFile(up, []byte(unit), 0o644)
}
return sp, up, filepath.Join(d, "wants-link")
}
// TestInstallSharedParent_RandomTempName is the audit-B1 negative test for the shared-parent boot
// persistence install: both staged install SOURCES (script + unit) must be RANDOM os.CreateTemp names
// (felhom-shared-parent-<random>.sh / .service), never the fixed, pre-creatable /tmp names (a local
// TOCTOU into a root-executed boot script), and two consecutive installs must use DIFFERENT paths.
func TestInstallSharedParent_RandomTempName(t *testing.T) {
r := &stagingRecorderRunner{}
func TestSharedParentBoot_NeverInstalls(t *testing.T) {
for _, c := range []struct{ name, script, unit string }{
{"both missing", "", ""},
{"script differs", "#!/bin/sh\necho old\n", sharedParentUnit},
{"unit missing", sharedParentScript, ""},
} {
r := &callRecorder{}
b := NewGuestBinder(r, nil)
sp, up, link := bootFiles(t, c.script, c.unit)
if err := b.ensureSharedParentBoot(context.Background(), up, sp, link); err == nil {
t.Errorf("%s: no error — the operator would not learn the bundle is missing", c.name)
}
if len(r.calls) != 0 {
t.Errorf("%s: the agent ran %v — it must install nothing (R-861)", c.name, r.calls)
}
}
}
func TestSharedParentBoot_EnablesOnceTheBundleBroughtTheFiles(t *testing.T) {
r := &callRecorder{}
b := NewGuestBinder(r, nil)
if err := b.installSharedParentUnit(context.Background()); err != nil {
t.Fatalf("installSharedParentUnit #1: %v", err)
sp, up, link := bootFiles(t, sharedParentScript, sharedParentUnit)
if err := b.ensureSharedParentBoot(context.Background(), up, sp, link); err != nil {
t.Fatal(err)
}
if err := b.installSharedParentUnit(context.Background()); err != nil {
t.Fatalf("installSharedParentUnit #2: %v", err)
if len(r.calls) != 1 || len(r.calls[0]) != 3 || r.calls[0][0] != "systemctl" || r.calls[0][1] != "enable" ||
r.calls[0][2] != "felhom-shared-parent.service" {
t.Fatalf("want exactly `systemctl enable felhom-shared-parent.service`, got %v", r.calls)
}
_ = os.Symlink(up, link)
r.calls = nil
if err := b.ensureSharedParentBoot(context.Background(), up, sp, link); err != nil || len(r.calls) != 0 {
t.Fatalf("already enabled: want no calls, got %v (%v)", r.calls, err)
}
}
srcs := r.installSources()
cases := []struct {
dest string
random *regexp.Regexp
fixed *regexp.Regexp
content string
}{
{sharedParentScriptPath, regexp.MustCompile(`felhom-shared-parent-[^/\\]+\.sh$`),
regexp.MustCompile(`felhom-shared-parent\.sh$`), sharedParentScript},
{sharedParentUnitPath, regexp.MustCompile(`felhom-shared-parent-[^/\\]+\.service$`),
regexp.MustCompile(`felhom-shared-parent\.service$`), sharedParentUnit},
}
for _, tc := range cases {
got := srcs[tc.dest]
if len(got) != 2 {
t.Fatalf("dest %s: expected 2 install calls, got %d (%v)", tc.dest, len(got), got)
}
for i, src := range got {
if !tc.random.MatchString(src) {
t.Errorf("dest %s call %d: source %q does not match the random temp pattern", tc.dest, i, src)
}
if tc.fixed.MatchString(src) {
t.Errorf("dest %s call %d: source %q is the FIXED predictable temp name (B1 TOCTOU)", tc.dest, i, src)
}
if _, err := os.Stat(src); err == nil {
t.Errorf("dest %s: staged temp %q left behind (defer os.Remove missing)", tc.dest, src)
}
}
if got[0] == got[1] {
t.Errorf("dest %s: two consecutive installs staged through the SAME source %q — must be random per call", tc.dest, got[0])
}
// Non-hollow: the staged file carried the real content at install time.
if r.srcContent[tc.dest] != tc.content {
t.Errorf("dest %s: staged content mismatch (got %d bytes, want %d)", tc.dest, len(r.srcContent[tc.dest]), len(tc.content))
// The bundle's copies are byte-identical to the constants the agent compares against.
func TestSharedParentFilesEqualTheBundle(t *testing.T) {
for file, want := range map[string]string{"felhom-shared-parent.sh": sharedParentScript, "felhom-shared-parent.service": sharedParentUnit} {
got, err := os.ReadFile(filepath.Join("..", "..", "configs", file))
if err != nil || string(got) != want {
t.Errorf("configs/%s differs from the agent's constant (err %v)", file, err)
}
}
}
@@ -0,0 +1,48 @@
package localapi
import (
"context"
"io"
"testing"
)
// stdinRecorder is a proxmox.Runner that records each call and the stdin it was handed.
type stdinRecorder struct {
name string
args []string
stdin string
}
func (r *stdinRecorder) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
r.name, r.args = name, args
return nil, nil, nil
}
func (r *stdinRecorder) RunStdin(_ context.Context, stdin io.Reader, name string, args ...string) ([]byte, []byte, error) {
b, _ := io.ReadAll(stdin)
r.name, r.args, r.stdin = name, args, string(b)
return nil, nil, nil
}
// R-861 (a) A1 (`09` §3 decision 165): the managed controller update writes the guest's image file through the ROOT
// verb, never through an in-guest `tee` the agent could feed any image.
//
// COMPANION RED-PROOF (observed): restore the pre-A1 body (`b.runner.RunStdin(ctx, …, "pct", "exec", vmid, "--",
// "tee", controllerImageFile)`) → this fails with "the image write ran pct …, want felhom-priv-apply". Restored.
func TestR861_WriteControllerImageUsesTheRootVerb(t *testing.T) {
rec := &stdinRecorder{}
b := NewGuestBinder(rec, discardLogger())
const img = "gitea.dooplex.hu/admin/felhom-controller:0.302.0"
if err := b.WriteControllerImage(context.Background(), 9201, img); err != nil {
t.Fatal(err)
}
if rec.name != privApplyBin {
t.Fatalf("the image write ran %s %v, want felhom-priv-apply", rec.name, rec.args)
}
if len(rec.args) != 2 || rec.args[0] != "controller-image" || rec.args[1] != "9201" {
t.Fatalf("argv = %v, want [controller-image 9201] (the sudoers line `^controller-image [0-9]+$`)", rec.args)
}
if rec.stdin != img+"\n" {
t.Fatalf("stdin = %q, want the ref plus one newline", rec.stdin)
}
}
+144
View File
@@ -0,0 +1,144 @@
package localapi
import (
"context"
"io"
"log/slog"
"net/http"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-894 — after an agent restart, an UNREADABLE storage must fall back to the last success saved on
// disk, not to "never". Measured 2026-10-05 on demo-hp: a restart at 04:57, the off-site storage
// unreachable at 06:25, the 7-day tier (last copy 4 days old) read DUE, vzdump failed.
//
// Every test here builds a NEW server and a NEW BackupSuccessState from the same file — that is the
// restart. The in-memory store (fakeStore) is always fresh, as after a real restart.
// r894Server builds a two-tier server whose off-site tier answers the storage listing with lister.
func r894Server(t *testing.T, path string, pbsSvc BackupService) *Server {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: &fakeStore{},
Storage: fakeStorage{targets: []hub.StorageTarget{{Name: "local"}, {Name: "felhom-pbs"}}},
Tokens: staticTokens{"A": 8200},
BackupTiers: []BackupTier{
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: pbsSvc},
},
LastKnownBackups: backup.NewBackupSuccessState(path),
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatal(err)
}
srv.baseCtx = context.Background()
srv.now = func() time.Time { return testNow }
return srv
}
// unreadable is the off-site storage as demo-hp saw it: "Can't connect to 10.77.0.1:8007".
func unreadable() archiveLister {
return archiveLister{fakeBackups: &fakeBackups{}, err: errStorageRead}
}
// THE R-894 CASE, end to end. Agent 1 takes an off-site backup through POST /backup (the fake
// runner's success is 12 h before testNow). The agent restarts. The storage cannot be read. The tier
// must read NOT due, from the copy saved on disk.
//
// COMPANION RED-PROOF (observed): delete the `lookup == archiveUnknown && s.lastKnown != nil` block in
// handleBackupDue → this fails with "after a restart an unreadable storage must fall back to the saved
// copy (12 h old, 7-day tier) — NOT due; got {… Due:true … AgeState:unknown …}". Restored.
func TestBackupDue_R894_RestartThenUnreadableStorage_FreshSavedCopyIsNotDue(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
// Agent 1: a real backup job through the endpoint the controller calls.
first := r894Server(t, path, &fakeBackups{})
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
}
waitFor(t, func() bool {
_, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200)
return ok
})
// Agent 2: a restart (new server, new state from the same file), and the storage is unreachable.
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if got.Due {
t.Fatalf("after a restart an unreadable storage must fall back to the saved copy (12 h old, 7-day tier) — NOT due; got %+v", got)
}
if got.AgeState != AgeStateKnown || got.AgeSecs == nil || *got.AgeSecs != int64((12*time.Hour).Seconds()) {
t.Fatalf("the age must come from the saved copy (12 h, known); got %+v", got)
}
}
// The deliberate rule stays: an unreadable storage must not suppress a backup that IS due. A saved
// copy older than the cadence reads DUE.
//
// COMPANION RED-PROOF (observed): make the fallback answer not-due whenever a saved copy exists
// (`if fromDisk { …Due:false… }` before the cadence check) → this fails with "a saved copy 9 days old
// under a 7-day cadence MUST read due". Restored.
func TestBackupDue_R894_RestartThenUnreadableStorage_OldSavedCopyIsDue(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
st := backup.NewBackupSuccessState(path)
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, 9*24*time.Hour, true)); err != nil {
t.Fatal(err)
}
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("a saved copy 9 days old under a 7-day cadence MUST read due; got %+v", got)
}
if got.AgeState != AgeStateKnown {
t.Fatalf("the age is known (from disk); got %+v", got)
}
}
// No saved copy → the pre-R-894 answer, byte for byte: DUE, age UNKNOWN (never ABSENT — the controller
// fires its window-gate valve only on absent, R-88).
func TestBackupDue_R894_RestartThenUnreadableStorage_NoSavedCopyIsDueUnknown(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due || got.AgeState != AgeStateUnknown || got.AgeSecs != nil {
t.Fatalf("no saved copy + unreadable storage must stay DUE with age unknown; got %+v", got)
}
}
// A storage that ANSWERS is the ground truth: an archive absent there makes the tier due even when the
// file remembers a fresh success (a pruned or deleted copy must be made again).
//
// COMPANION RED-PROOF (observed): drop `lookup == archiveUnknown &&` from the fallback condition → this
// fails with "the storage answered 'no archive' — the saved copy must NOT stand in for it". Restored.
func TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
st := backup.NewBackupSuccessState(path)
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, time.Hour, true)); err != nil {
t.Fatal(err)
}
absent := archiveLister{fakeBackups: &fakeBackups{}, found: false}
got := dueFor(t, r894Server(t, path, absent).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("the storage answered 'no archive' — the saved copy must NOT stand in for it; got %+v", got)
}
}
// A FAILED backup is never saved: it must not make a tier look fresh after a restart.
func TestBackupDue_R894_FailedBackupIsNotSaved(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
first := r894Server(t, path, &fakeBackups{failErr: "could not activate storage 'felhom-pbs'"})
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
}
waitFor(t, func() bool { return len(first.store.Backups(context.Background())) == 1 })
if _, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200); ok {
t.Fatal("a failed backup must never be saved as a success")
}
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("after a failed backup and a restart the tier must still be due; got %+v", got)
}
}
+67 -22
View File
@@ -85,6 +85,13 @@ type BackupStore interface {
RestoreTests(ctx context.Context) []hub.RestoreTest
}
// LastKnownBackupStore (R-894) is the on-disk newest-success-per-tier record. Satisfied by
// *backup.BackupSuccessState.
type LastKnownBackupStore interface {
RecordBackupSuccess(target string, b hub.Backup) error
LastKnownSuccess(target string, vmid int) (time.Time, bool)
}
// StorageView yields the host's observed storage targets (for mapping a mount's storage id →
// fast/slow class). Satisfied by *storage.Observer.
type StorageView interface {
@@ -164,6 +171,10 @@ type Options struct {
// PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg
// that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs.
AfterPrimaryBackup func(ctx context.Context, vmid int)
// LastKnownBackups (R-894) keeps the newest successful backup per tier ON DISK, so the due-check's
// fallback for an UNREADABLE storage after an agent restart is the last known copy, not "never".
// nil = the pre-R-894 behaviour (in-memory record only).
LastKnownBackups LastKnownBackupStore
// Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when
// nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner.
Privileged PrivilegedRunner
@@ -229,7 +240,6 @@ type Options struct {
// POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured"
// (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer.
EscrowRecovery EscrowRecoverer
}
// defaultBackupCadence is the fallback /backup/due window when none is configured.
@@ -279,30 +289,34 @@ type Server struct {
// is the pre-R-82 shape.
tiers []BackupTier
// inFlight (R-85) is shared with the restore-test scheduler so the two never run together.
inFlight *backup.InFlight
inFlight *backup.InFlight
afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none
logger *slog.Logger
now func() time.Time
lastKnown LastKnownBackupStore // R-894: on-disk newest success per tier; nil = none
logger *slog.Logger
now func() time.Time
disks DiskOps // slice 8C (optional)
diskGate StorageGate // slice 8C (optional)
guestList GuestLister // slice 8C (optional)
guestAttach GuestAttacher // slice 10 P2 (optional)
mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional)
memMu sync.Mutex // single-flight around a resize apply (one customer per host)
netStorage NetworkStorageOps // Part A1: NAS network mounts (optional)
netMountRoot string // the user-data namespace root for the network-mount role gate
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
disks DiskOps // slice 8C (optional)
diskGate StorageGate // slice 8C (optional)
guestList GuestLister // slice 8C (optional)
guestAttach GuestAttacher // slice 10 P2 (optional)
mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional)
memMu sync.Mutex // single-flight around a resize apply (one customer per host)
netStorage NetworkStorageOps // Part A1: NAS network mounts (optional)
netMountRoot string // the user-data namespace root for the network-mount role gate
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
// crashGuardStatePath (R-856) is the host crash guard's state file read by GET /host/crash-guard;
// empty = defaultCrashGuardStatePath. A seam: tests point it at a fixture.
crashGuardStatePath string
// escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
// offsite repository password. OPTIONAL — nil (no hub client configured) makes
// POST /escrow/recover-offsite-password answer 503 rather than pretending.
escrowRecovery EscrowRecoverer
intent IntentRecorder // slice 10 P3 (optional)
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
staleLock StaleLockController // F2-b startup stale-lock recovery (optional)
escrowRecovery EscrowRecoverer
intent IntentRecorder // slice 10 P3 (optional)
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
staleLock StaleLockController // F2-b startup stale-lock recovery (optional)
// guestPower (F-REBOOT) is per-guest start-attempt state for the guest-power watchdog.
// Guarded by guestPowerMu in guestpower.go; in-memory on purpose (see guestPowerState).
guestPower map[int]guestPowerState
@@ -475,6 +489,7 @@ func NewServer(o Options) (*Server, error) {
s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence)
s.inFlight = o.InFlight
s.afterPrimaryBackup = o.AfterPrimaryBackup
s.lastKnown = o.LastKnownBackups
if s.backups == nil && len(s.tiers) > 0 {
s.backups = s.tiers[0].Service
}
@@ -520,6 +535,9 @@ func (s *Server) Handler() http.Handler {
// Host metrics (slice 9): host-wide health + per-storage capacity for the customer's monitoring
// view. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot).
mux.HandleFunc("GET /host/metrics", s.withGuest(s.handleHostMetrics))
// R-856 (`09` §3 decision 143): the host crash guard's record of the last HOST boot — the controller
// waits ~15 min with app mails after an unclean one. Read-only; a missing/garbled file = present:false.
mux.HandleFunc("GET /host/crash-guard", s.withGuest(s.handleCrashGuard))
// Disk management (slice 8C) — self-scoped; format routes through the data-bearing classifier+gate.
mux.HandleFunc("GET /disks", s.withGuest(s.handleDisks))
mux.HandleFunc("GET /disks/candidates", s.withGuest(s.handleDiskCandidates))
@@ -899,6 +917,13 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
s.logger.Info("local-api: backup job complete", "vmid", vmid, "target", tier.TargetID, "job", jobID, "archive", b.Archive)
}
s.store.RecordBackup(b)
if b.Success && s.lastKnown != nil {
if err := s.lastKnown.RecordBackupSuccess(tier.TargetID, b); err != nil {
// Not fatal: the backup exists. Only the fallback after a restart loses this copy.
s.logger.Warn("local-api: could not save the backup on disk for the due-check fallback (R-894)",
"vmid", vmid, "target", tier.TargetID, "err", err)
}
}
s.finishJob(key, jobID, b)
// OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate.
if b.Success && tier.Primary && s.afterPrimaryBackup != nil {
@@ -1061,6 +1086,20 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
newest, haveNewest = t, true
unparseable = false // ground truth supersedes an unreadable in-memory timestamp
}
// R-894: the storage could not be read → the last success saved ON DISK stands in for the in-memory
// record a restart emptied. ONLY on archiveUnknown: a storage that answers is the ground truth, and an
// archive absent there must make the tier due even when the file remembers one (a pruned copy).
// A saved copy older than the cadence still reads DUE below — an unreadable storage never suppresses
// a backup that is due.
fromDisk := false
if lookup == archiveUnknown && s.lastKnown != nil {
if saved, ok := s.lastKnown.LastKnownSuccess(tier.TargetID, vmid); ok && (!haveNewest || saved.After(newest)) {
newest, haveNewest, fromDisk = saved, true, true
unparseable = false
s.logger.Info("local-api: backup storage unreadable — due-check uses the last success saved on disk (R-894)",
"vmid", vmid, "target", tier.TargetID, "saved", saved.UTC().Format(time.RFC3339))
}
}
if !haveNewest {
// R-88 Part 2: THREE distinct reasons for a nil age, each with its own state. Only ABSENT is a
// positive claim of "never backed up"; only that one may license the controller to bypass its
@@ -1083,13 +1122,17 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
}
age := s.now().Sub(newest)
ageSecs := int64(age.Seconds())
suffix := ""
if fromDisk {
suffix = " (storage unreadable — age from the last success saved on disk)"
}
if age >= tier.Cadence {
writeOK(w, BackupDueResponse{VMID: vmid, Due: true, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
Reason: "older than cadence", Target: echo})
Reason: "older than cadence" + suffix, Target: echo})
return
}
writeOK(w, BackupDueResponse{VMID: vmid, Due: false, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
Reason: "within cadence window", Target: echo})
Reason: "within cadence window" + suffix, Target: echo})
}
// BackupTiersResponse is GET /backup/tiers (R-82): the tiers this agent serves, primary first.
@@ -1477,4 +1520,6 @@ func writeStatus(w http.ResponseWriter, code int, ok bool, data any, errMsg stri
}
// SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0).
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) { s.afterPrimaryBackup = fn }
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) {
s.afterPrimaryBackup = fn
}
+31 -17
View File
@@ -154,12 +154,15 @@ func (s *TokenStore) Mint(vmid int) (string, error) {
// looks it up; the per-candidate comparison is constant-time to avoid a timing oracle on the
// stored hash. ok is false for an unknown/empty token.
//
// Reload-on-miss (B3): the store FILE is shared across processes — the one-shot provisioner
// (`--selftest=provision`) Mints into it while the long-lived daemon serves Lookup from an index
// built at open. On a miss, re-read the file ONCE and re-check, so a token minted after this
// process started authorizes without a daemon restart (the drill's fresh-install 401). The
// append-only log makes an unchanged file size proof of no new records, so a genuinely unknown
// token costs at most one stat once the index is current — never a reload loop.
// Reload-on-change (B3, R-269): the store FILE is shared across processes — the one-shot
// provisioner (`--selftest=provision`) Mints into it while the long-lived daemon serves Lookup from
// an index built at open. Every Lookup stats the file first and re-reads it when the append-only
// log has grown, BEFORE answering — so a token minted elsewhere authorizes without a restart AND a
// token rotated out elsewhere stops authorizing on its very next presentation. (Before R-269 the
// re-read ran only on a MISS, so a superseded token was a direct map hit and kept authorizing until
// some unrelated miss forced the reload.) An unchanged size is proof of no new records, so the
// steady state costs one stat per call and never a reload loop. Pinned by
// TestTokenStore_RotatedOutTokenRejectedFirst.
func (s *TokenStore) Lookup(token string) (int, bool) {
if token == "" {
return 0, false
@@ -167,23 +170,34 @@ func (s *TokenStore) Lookup(token string) (int, bool) {
want := hashToken(token)
s.mu.Lock()
defer s.mu.Unlock()
// Direct map hit is the common path; the constant-time compare guards against a timing
// side-channel by re-checking the matched key (map lookup itself is not the secret-bearing
// comparison — the hash of a random 256-bit token is not feasibly guessable regardless).
if vmid, ok := s.byHash[want]; ok {
if subtle.ConstantTimeCompare([]byte(want), []byte(s.byVMID[vmid])) == 1 {
return vmid, true
st, statErr := os.Stat(s.path)
if statErr == nil && st.Size() != s.loadedSize {
// The log changed under us (another process minted/rotated): converge first, then answer.
s.reloads++
if err := s.reloadLocked(); err != nil {
return 0, false // unreadable store: fail closed, never crash the auth path
}
return s.matchLocked(want)
}
// Miss: skip the re-read when the append-only log has not grown (nothing new to see).
// A stat error falls through to the reload, which handles a missing file as empty.
if st, err := os.Stat(s.path); err == nil && st.Size() == s.loadedSize {
return 0, false
if vmid, ok := s.matchLocked(want); ok {
return vmid, true
}
if statErr == nil {
return 0, false // file unchanged since the last (re)load: genuinely unknown
}
// Stat failed (e.g. the file vanished): reload, which treats a missing file as empty.
s.reloads++
if err := s.reloadLocked(); err != nil {
return 0, false // unreadable store: fail closed, never crash the auth path
return 0, false
}
return s.matchLocked(want)
}
// matchLocked answers from the in-memory index. Direct map hit is the common path; the
// constant-time compare re-checks the matched key against the guest's CURRENT hash (map lookup
// itself is not the secret-bearing comparison — the hash of a random 256-bit token is not
// feasibly guessable regardless). Caller holds the mutex.
func (s *TokenStore) matchLocked(want string) (int, bool) {
if vmid, ok := s.byHash[want]; ok {
if subtle.ConstantTimeCompare([]byte(want), []byte(s.byVMID[vmid])) == 1 {
return vmid, true
+45
View File
@@ -210,6 +210,51 @@ func TestTokenStore_ReloadOnMiss_RemintCoherence(t *testing.T) {
}
}
// R-269: a token rotated out by ANOTHER process must stop authorizing on its very next
// presentation — with NO intervening lookup of the new token. This is the order the operator hits
// after rotating a leaked token: the leaked one is presented first. RemintCoherence above looks the
// NEW token up first, and that miss is what used to evict the old hash, so it passed while the leaked
// token kept returning HTTP 200 on hardware (2026-08-09) until something unrelated forced a reload.
//
// RED-PROOF: restore the reload-on-MISS-only Lookup (answer a map hit before stat-ing the file) and
// this fails with "rotated-out token still authorizes".
func TestTokenStore_RotatedOutTokenRejectedFirst(t *testing.T) {
path := filepath.Join(t.TempDir(), "tokens.log")
daemon, err := OpenTokenStore(path)
if err != nil {
t.Fatalf("open daemon store: %v", err)
}
defer daemon.Close()
minter, err := OpenTokenStore(path)
if err != nil {
t.Fatalf("open minter store: %v", err)
}
defer minter.Close()
old, err := minter.Mint(130)
if err != nil {
t.Fatalf("mint old: %v", err)
}
if vmid, ok := daemon.Lookup(old); !ok || vmid != 130 { // the daemon has learned the old token
t.Fatalf("old token before rotation: (%d,%v), want (130,true)", vmid, ok)
}
fresh, err := minter.Mint(130) // rotation, written by another process
if err != nil {
t.Fatalf("mint fresh: %v", err)
}
if vmid, ok := daemon.Lookup(old); ok { // the leaked token FIRST
t.Fatalf("rotated-out token still authorizes vmid %d on its first presentation after rotation — "+
"Mint's 'any previous token for this guest is revoked' is false across processes (R-269)", vmid)
}
if vmid, ok := daemon.Lookup(fresh); !ok || vmid != 130 {
t.Fatalf("fresh token after rotation: (%d,%v), want (130,true)", vmid, ok)
}
if vmid, ok := daemon.Lookup(old); ok {
t.Fatalf("rotated-out token authorizes vmid %d after the fresh one was seen", vmid)
}
}
// §8 edge: the store file deleted between open and a miss — reload treats it as empty; Lookup
// fails closed, no crash.
func TestTokenStore_ReloadOnMiss_MissingFile(t *testing.T) {
+157
View File
@@ -0,0 +1,157 @@
package osupdate
import (
"context"
"crypto/sha256"
"encoding/base64"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"path/filepath"
"regexp"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// OpConfigUpdate is the signed op that brings a box's ROOT-OWNED files (sudoers, wrappers, units) to a release's config
// bundle (R-840, `11` §5.4.2). The params pin the agent version and the bundle's sha256; the root wrapper re-verifies
// the signature, the host binding, the nonce and the sha ITSELF — this executor is only the courier.
const OpConfigUpdate = "agent_config_update"
// BundleFileName is the bundle's name beside the binary in the Gitea generic package felhom-agent/<version>/.
const BundleFileName = "felhom-config-bundle.json"
var (
bundleVersionRe = regexp.MustCompile(`^[0-9]+\.[0-9]+\.[0-9]+(-[0-9A-Za-z.]+)?$`)
bundleSHARe = regexp.MustCompile(`^[0-9a-f]{64}$`)
)
// ConfigUpdateParams are the signed params (the wrapper reads the same two names out of the signed blob).
type ConfigUpdateParams struct {
AgentVersion string `json:"agent_version"`
BundleSHA256 string `json:"bundle_sha256"`
}
// BundleURL derives the bundle's URL from the agent binary's URL template (".../felhom-agent/{version}/felhom-agent").
func BundleURL(binaryTemplate, version string) (string, error) {
if !strings.HasSuffix(binaryTemplate, "/felhom-agent") || !strings.Contains(binaryTemplate, "{version}") {
return "", fmt.Errorf("cannot derive the bundle URL from %q (want …/{version}/felhom-agent)", binaryTemplate)
}
t := strings.TrimSuffix(binaryTemplate, "felhom-agent") + BundleFileName
return strings.ReplaceAll(t, "{version}", version), nil
}
// ConfigUpdateExecutor runs a verified agent_config_update (signedjobs.Executor).
type ConfigUpdateExecutor struct {
Leg *Leg
URLTemplate string // the agent binary's template (config.SelfUpdate.URLTemplate)
Username string
Token string
HTTPClient *http.Client
// AfterInstall runs after a bundle installed (the capability re-probe; nil = none).
AfterInstall func(ctx context.Context)
}
// Execute implements signedjobs.Executor.
func (e ConfigUpdateExecutor) Execute(ctx context.Context, op string, params json.RawMessage) error {
if op != OpConfigUpdate {
return signedjobs.ErrNoExecutor
}
so, ok := signedjobs.SignedOpFrom(ctx)
if !ok {
return fmt.Errorf("agent_config_update: no signed envelope in the context — the wrapper could not verify it")
}
var p ConfigUpdateParams
if err := json.Unmarshal(params, &p); err != nil {
return fmt.Errorf("agent_config_update: bad params: %w", err)
}
if !bundleVersionRe.MatchString(p.AgentVersion) || !bundleSHARe.MatchString(p.BundleSHA256) {
return fmt.Errorf("agent_config_update: params must pin agent_version (semver) and bundle_sha256 (64 hex)")
}
lg := e.Leg.log().With("op", OpConfigUpdate, "agent_version", p.AgentVersion)
url, err := BundleURL(e.URLTemplate, p.AgentVersion)
if err != nil {
return fmt.Errorf("agent_config_update: %w", err)
}
dir := e.Leg.PlanDir
if dir == "" {
dir = DefaultPlanDir
}
if err := os.MkdirAll(dir, 0o700); err != nil {
return fmt.Errorf("agent_config_update: plan dir: %w", err)
}
path := filepath.Join(dir, "bundle-"+p.AgentVersion+".json")
start := time.Now()
got, err := e.download(ctx, url, path)
if err != nil {
_ = os.Remove(path)
return fmt.Errorf("agent_config_update: download %s: %w", url, err)
}
defer os.Remove(path)
// The agent's own check is a courtesy (an early, clear error); the wrapper's is the gate.
if got != p.BundleSHA256 {
return fmt.Errorf("agent_config_update: the downloaded bundle's sha256 is %s, the signed job pins %s — nothing installed", got, p.BundleSHA256)
}
lg.Info("osupdate: config bundle downloaded; handing it to the root wrapper", "sha256", got[:16], "duration_ms", time.Since(start).Milliseconds())
runID := "bundle-" + e.Leg.now().UTC().Format("20060102T150405Z")
wr, err := e.Leg.call(ctx, runID, map[string]any{"release_id": "bundle-" + p.AgentVersion, "layer": LayerHost,
"mode": "bundle", "bundle": path,
"signed": map[string]string{"blob_b64": base64.StdEncoding.EncodeToString(so.Blob), "sig": string(so.Sig)}})
if err != nil {
return fmt.Errorf("agent_config_update: %w", err)
}
if wr.refused() {
lg.Warn("osupdate: config bundle REFUSED by the wrapper — nothing changed", "refused", string(wr.Refused))
return fmt.Errorf("agent_config_update: refused: %s", wr.Refused)
}
if wr.failed() {
lg.Error("osupdate: config bundle FAILED — the wrapper put the previous files back", "failed", string(wr.Failed), "bundle", string(wr.Bundle))
return fmt.Errorf("agent_config_update: failed (previous files restored): %s", wr.Failed)
}
lg.Warn("osupdate: config bundle INSTALLED", "bundle", string(wr.Bundle), "pass_seconds", wr.PassSeconds)
if e.AfterInstall != nil {
e.AfterInstall(ctx)
}
return nil
}
func (e ConfigUpdateExecutor) download(ctx context.Context, url, dest string) (string, error) {
hc := e.HTTPClient
if hc == nil {
hc = &http.Client{Timeout: 2 * time.Minute}
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return "", err
}
if e.Username != "" || e.Token != "" {
req.SetBasicAuth(e.Username, e.Token)
}
resp, err := hc.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return "", fmt.Errorf("HTTP %d", resp.StatusCode)
}
f, err := os.OpenFile(dest, os.O_CREATE|os.O_TRUNC|os.O_WRONLY, 0o600)
if err != nil {
return "", err
}
h := sha256.New()
// 4 MB is the wrapper's own limit; read one byte more so an oversized bundle fails there, visibly.
if _, err := io.Copy(io.MultiWriter(f, h), io.LimitReader(resp.Body, 4*1024*1024+1)); err != nil {
f.Close()
return "", err
}
if err := f.Close(); err != nil {
return "", err
}
return hex.EncodeToString(h.Sum(nil)), nil
}
+155
View File
@@ -0,0 +1,155 @@
package osupdate
import (
"context"
"crypto/sha256"
"encoding/base64"
"encoding/hex"
"encoding/json"
"errors"
"io"
"net/http"
"net/http/httptest"
"os"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// bundleWrapper plays felhom-os-apply for mode "bundle": it records the plan and the bundle file's bytes AT CALL TIME
// (the executor deletes the file afterwards), and answers with rep.
type bundleWrapper struct {
t *testing.T
rep string
plans []map[string]any
bodies [][]byte
}
func (b *bundleWrapper) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
if name != WrapperPath || len(args) != 2 || args[0] != "--plan" {
b.t.Fatalf("unexpected command %s %v", name, args)
}
raw, err := os.ReadFile(args[1])
if err != nil {
b.t.Fatal(err)
}
var plan map[string]any
_ = json.Unmarshal(raw, &plan)
b.plans = append(b.plans, plan)
body, _ := os.ReadFile(plan["bundle"].(string))
b.bodies = append(b.bodies, body)
return []byte("OSAPPLY-REPORT " + b.rep + "\n"), nil, nil
}
func serveBundle(t *testing.T, body []byte) (*httptest.Server, string) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/felhom-agent/0.143.0/"+BundleFileName {
http.NotFound(w, r)
return
}
_, _ = w.Write(body)
}))
t.Cleanup(srv.Close)
sum := sha256.Sum256(body)
return srv, hex.EncodeToString(sum[:])
}
func bundleExec(t *testing.T, w *bundleWrapper, srvURL string) (ConfigUpdateExecutor, *bool) {
l, _ := newLeg(t, &fakeWrapper{t: t}, nil)
l.Runner = w
called := false
return ConfigUpdateExecutor{Leg: l, URLTemplate: srvURL + "/felhom-agent/{version}/felhom-agent",
AfterInstall: func(context.Context) { called = true }}, &called
}
func signedCtx() context.Context {
return signedjobs.WithSignedOp(context.Background(), &reconcile.SignedOp{Blob: []byte(`{"op":"agent_config_update"}`), Sig: []byte("SIG")})
}
func params(v, sha string) json.RawMessage {
p, _ := json.Marshal(ConfigUpdateParams{AgentVersion: v, BundleSHA256: sha})
return p
}
// The courier hands the wrapper the bundle bytes it downloaded and the RAW signed envelope (the wrapper verifies both
// itself), then runs the capability probe. Red-proof: drop "signed" from the plan → the plan check below fails.
func TestConfigUpdate_PassesBundleAndEnvelopeToTheWrapper(t *testing.T) {
body := []byte(`{"format":1,"agent_version":"0.143.0","files":[]}`)
srv, sha := serveBundle(t, body)
w := &bundleWrapper{t: t, rep: `{"mode":"bundle","bundle":{"agent_version":"0.143.0","written":["/etc/sudoers.d/felhom-agent"]}}`}
e, called := bundleExec(t, w, srv.URL)
if err := e.Execute(signedCtx(), OpConfigUpdate, params("0.143.0", sha)); err != nil {
t.Fatal(err)
}
p := w.plans[0]
sg, _ := p["signed"].(map[string]any)
if p["mode"] != "bundle" || p["layer"] != "host" || sg == nil || sg["sig"] != "SIG" ||
sg["blob_b64"] != base64.StdEncoding.EncodeToString([]byte(`{"op":"agent_config_update"}`)) {
t.Fatalf("plan = %v", p)
}
if string(w.bodies[0]) != string(body) || !strings.HasSuffix(p["bundle"].(string), "/bundle-0.143.0.json") {
t.Fatalf("the wrapper got %q at %v", w.bodies[0], p["bundle"])
}
if !*called {
t.Fatal("the capability probe must run after an install")
}
if _, err := os.Stat(p["bundle"].(string)); !os.IsNotExist(err) {
t.Fatal("the downloaded bundle must be removed after the call")
}
}
func TestConfigUpdate_WrongShaNeverReachesTheWrapper(t *testing.T) {
srv, _ := serveBundle(t, []byte("tampered"))
w := &bundleWrapper{t: t}
e, called := bundleExec(t, w, srv.URL)
err := e.Execute(signedCtx(), OpConfigUpdate, params("0.143.0", strings.Repeat("a", 64)))
if err == nil || len(w.plans) != 0 || *called {
t.Fatalf("err=%v plans=%d", err, len(w.plans))
}
}
func TestConfigUpdate_RefusedAndFailedAreErrors(t *testing.T) {
for _, rep := range []string{`{"refused":{"code":"R17","reason":"trust root"}}`, `{"failed":{"rc":3},"bundle":{"rolled_back":["/x"]}}`} {
srv, sha := serveBundle(t, []byte("{}"))
w := &bundleWrapper{t: t, rep: rep}
e, called := bundleExec(t, w, srv.URL)
if err := e.Execute(signedCtx(), OpConfigUpdate, params("0.143.0", sha)); err == nil || *called {
t.Fatalf("%s: err=%v called=%v", rep, err, *called)
}
}
}
func TestConfigUpdate_GuardsBeforeAnyDownload(t *testing.T) {
w := &bundleWrapper{t: t}
e, _ := bundleExec(t, w, "http://127.0.0.1:1")
if err := e.Execute(signedCtx(), "agent_update", nil); !errors.Is(err, signedjobs.ErrNoExecutor) {
t.Fatalf("another op must pass through the chain: %v", err)
}
if err := e.Execute(context.Background(), OpConfigUpdate, params("0.143.0", strings.Repeat("a", 64))); err == nil {
t.Fatal("no envelope must refuse")
}
for _, p := range []json.RawMessage{params("0.143", strings.Repeat("a", 64)), params("0.143.0", "ABC"), json.RawMessage(`nope`)} {
if err := e.Execute(signedCtx(), OpConfigUpdate, p); err == nil {
t.Fatalf("bad params %s must refuse", p)
}
}
if len(w.plans) != 0 {
t.Fatal("no wrapper call on a refusal")
}
}
func TestBundleURL(t *testing.T) {
u, err := BundleURL("https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/{version}/felhom-agent", "0.143.0")
if err != nil || u != "https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/0.143.0/felhom-config-bundle.json" {
t.Fatalf("%s %v", u, err)
}
if _, err := BundleURL("https://example/felhom-agent-{version}.bin", "0.143.0"); err == nil {
t.Fatal("an underivable template must refuse")
}
}
func (b *bundleWrapper) RunStdin(ctx context.Context, _ io.Reader, name string, args ...string) ([]byte, []byte, error) {
return b.Run(ctx, name, args...)
}
+91
View File
@@ -0,0 +1,91 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// OpDockerStep is the signed op class of a Docker engine step (`11` §5.8): a ring-1 box takes an approved engine set,
// and every box takes an UNDO, only through it. CC may sign it until the first paying customer (R-530 ruling).
const OpDockerStep = "os_docker_step"
// DockerStepParams are the signed params. The wrapper compares Packages and Undo with the plan byte-for-byte.
type DockerStepParams struct {
ReleaseID string `json:"release_id"`
Packages []Package `json:"packages"`
Undo bool `json:"undo"`
VMID int `json:"vmid,omitempty"`
}
// DockerStepExecutor runs a verified os_docker_step (signedjobs.Executor). Guest finds the box's customer guest when
// the params name none; Gate (optional) takes the host-wide heavy-op gate so a step never runs beside a backup.
type DockerStepExecutor struct {
Leg *Leg
Guest func(ctx context.Context) (int, error)
Gate func(ctx context.Context) (release func(), err error)
}
// Execute implements signedjobs.Executor.
func (e DockerStepExecutor) Execute(ctx context.Context, op string, params json.RawMessage) error {
if op != OpDockerStep {
return signedjobs.ErrNoExecutor
}
so, ok := signedjobs.SignedOpFrom(ctx)
if !ok {
return fmt.Errorf("os_docker_step: no signed envelope in the context — the wrapper could not verify it")
}
var p DockerStepParams
if err := json.Unmarshal(params, &p); err != nil || len(p.Packages) == 0 {
return fmt.Errorf("os_docker_step: params must name the engine set: %v", err)
}
vmid := p.VMID
if vmid == 0 {
if e.Guest == nil {
return fmt.Errorf("os_docker_step: no vmid and no guest finder")
}
v, err := e.Guest(ctx)
if err != nil {
return fmt.Errorf("os_docker_step: find the customer guest: %w", err)
}
vmid = v
}
if e.Gate != nil {
release, err := e.Gate(ctx)
if err != nil {
return fmt.Errorf("os_docker_step: heavy-op gate busy (a backup or restore-test runs): %w", err)
}
defer release()
}
rep := e.Leg.RunDockerSigned(ctx, vmid, p, so.Blob, string(so.Sig))
switch rep.Outcome {
case "applied", "nothing":
if rep.Healthy {
return nil
}
}
return fmt.Errorf("os_docker_step: %s (%s) %s", rep.Outcome, rep.HealthReason, string(rep.Refused))
}
// RunDockerSigned is one signed Docker step (ring 1 or an undo): live-restore first (decision 87, a no-op when on),
// then the docker layer with the signed envelope, which the wrapper verifies itself.
func (l *Leg) RunDockerSigned(ctx context.Context, vmid int, p DockerStepParams, blob []byte, sig string) Report {
unlock := l.lockPass(true)
defer unlock()
l.sendUnsentLocked(ctx) // R-868
runID := l.now().UTC().Format("20060102T150405Z")
lg := l.log().With("run", runID, "vmid", vmid, "trigger", "signed", "release", p.ReleaseID, "undo", p.Undo)
if err := l.EnsureLiveRestore(ctx, runID, vmid); err != nil {
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerDocker, Trigger: "signed", Ring: l.Block().Ring, VMID: vmid,
Mode: "apply", ReleaseID: p.ReleaseID, Outcome: "failed", HealthReason: "live-restore could not be turned on: " + err.Error()})
}
rid := p.ReleaseID
if rid == "" {
rid = "signed-" + runID
}
return l.runLayer(ctx, runID, LayerDocker, vmid, "signed", l.Block(), dockerOpts{releaseID: rid, packages: p.Packages,
undo: p.Undo, signed: map[string]string{"blob_b64": base64.StdEncoding.EncodeToString(blob), "sig": sig}})
}
+49
View File
@@ -0,0 +1,49 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"errors"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// The executor hands the RAW signed bytes to the wrapper (which verifies them itself) and the exact signed package
// list. Red-proof: drop the `signed` field from the docker plan in runLayer and the plan check fails.
func TestDockerStepExecutor_PassesTheSignedEnvelope(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}}, DockerEngine: "29.8.2", Authority: "signed"}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
e := DockerStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
params, _ := json.Marshal(DockerStepParams{ReleaseID: "os-docker-1", Packages: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie", Origin: "Docker CE"}}})
ctx := signedjobs.WithSignedOp(context.Background(), &reconcile.SignedOp{Blob: []byte(`{"op":"os_docker_step"}`), Sig: []byte("SIG")})
if err := e.Execute(ctx, OpDockerStep, params); err != nil {
t.Fatal(err)
}
dp := w.plans[len(w.plans)-1]
sg, _ := dp["signed"].(map[string]any)
if dp["layer"] != "docker" || dp["lane"] != "slow" || dp["release_id"] != "os-docker-1" || sg == nil ||
sg["blob_b64"] != base64.StdEncoding.EncodeToString([]byte(`{"op":"os_docker_step"}`)) || sg["sig"] != "SIG" {
t.Fatalf("docker plan = %v", dp)
}
if calls(w) != "guest:live-restore-on,docker:apply" || len(h.reports) != 1 || h.reports[0].Trigger != "signed" {
t.Fatalf("calls=%s reports=%+v", calls(w), h.reports)
}
}
func TestDockerStepExecutor_RefusesWithoutEnvelopeAndPassesOtherOps(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
e := DockerStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
if err := e.Execute(context.Background(), "agent_update", nil); !errors.Is(err, signedjobs.ErrNoExecutor) {
t.Fatalf("another op must pass through the chain: %v", err)
}
params, _ := json.Marshal(DockerStepParams{Packages: []Package{{Name: "docker-ce", Version: "1"}}})
if err := e.Execute(context.Background(), OpDockerStep, params); err == nil || len(w.plans) != 0 {
t.Fatalf("no envelope must refuse before any wrapper call: err=%v calls=%s", err, calls(w))
}
}
+409
View File
@@ -0,0 +1,409 @@
package osupdate
// The kernel lane (R-836, `09` §3 decisions 164 + 172, `11` §5.11). The root half is felhom-os-apply's layer "kernel"
// (configs/, its own tests); this file decides WHEN and judges the boot:
//
// - the night leg (Run → runKernel): after a healthy host step, on a night the hub marks as told (the household was
// mailed the day before — no mail, no step): ring 0 STAGES the pending kernel (select pending-kernel, the root-owned
// ring-0 mark) and reboots; ring 1 reboots only a step a signed os_kernel_step staged earlier (KernelStepExecutor).
// - after a boot (KernelAfterBoot, at daemon start): the wrapper says what became of the step. On the new kernel the
// agent JUDGES the boot — the host health rule (`11` §8.2: the Proxmox daemons, the guest running and healthy, the
// tunnel) AND the box reaching the hub — for KernelJudgeWait. Healthy → kernel-good (the new kernel becomes the
// default). Not healthy by the deadline → ONE self-revert (kernel-revert: a reboot into the old kernel, still the
// default). A crash on the new kernel needs nothing from the agent: GRUB already boots the old default.
//
// The host is rebooted by this file only through the wrapper (kernel-reboot, kernel-revert), and only for a staged step.
import (
"context"
"encoding/base64"
"encoding/json"
"fmt"
"log/slog"
"regexp"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// OpKernelStep is the signed op class that STAGES a kernel set on a ring-1 box (it never reboots: the night leg does,
// once the household was told). CC may sign it until the first paying customer (R-530 ruling).
const OpKernelStep = "os_kernel_step"
// DefaultKernelJudgeWait is how long a one-shot boot may take to come back healthy before the agent reverts it ONCE.
// Measured 2026-10-07 (`audits/kernel-lane-2026-10-07/B/`): every container healthy 68 s after a reboot on demo-felhom,
// 272 s on demo-hp; 20 minutes stays under the hub's 30-minute host_stale (a box that never comes back alarms after it).
const DefaultKernelJudgeWait = 20 * time.Minute
var kverRE = regexp.MustCompile(`^[0-9]+\.[0-9]+\.[0-9]+-[0-9]+-pve$`)
// KernelView is the wrapper's kernel object (felhom-os-apply Kernel.view).
type KernelView struct {
Running string `json:"running"`
Default string `json:"default"`
Flag *string `json:"flag"`
Phase string `json:"phase"`
From string `json:"from"`
To string `json:"to"`
SelfRevertUsed bool `json:"self_revert_used"`
Reason string `json:"reason"`
VMID int `json:"vmid"` // the customer guest the step was staged for (the health rule's guest)
}
func parseKernel(raw json.RawMessage) KernelView {
var v KernelView
_ = json.Unmarshal(raw, &v)
return v
}
// kernelPlan is one kernel-layer wrapper call.
func kernelPlan(mode string, vmid int, extra map[string]any) map[string]any {
p := map[string]any{"release_id": "kernel", "layer": LayerKernel, "lane": "slow", "vmid": vmid, "mode": mode,
"packages": []Package{}}
for k, v := range extra {
p[k] = v
}
return p
}
// KernelStatus reads the kernel lane's state (read only).
func (l *Leg) KernelStatus(ctx context.Context, vmid int) (KernelView, error) {
wr, err := l.call(ctx, "kstatus"+l.now().UTC().Format("150405"), kernelPlan("kernel-status", vmid, nil))
if err != nil {
return KernelView{}, err
}
if wr.refused() {
return KernelView{}, fmt.Errorf("kernel-status refused: %s", wr.Refused)
}
return parseKernel(wr.Kernel), nil
}
// kernelDue reports whether tonight's leg may take a kernel step, and why not.
func kernelDue(blk hub.WireOSUpdate, trigger string) (bool, string) {
switch {
case trigger != "night":
return false, "a kernel step runs only in the night leg (never a debug pass)"
case blk.Kernel == nil:
return false, "the hub names no kernel step for this box"
case !blk.Enabled:
return false, "OS updates are switched off for this box"
case !kverRE.MatchString(blk.Kernel.Kver):
return false, "the hub's kernel " + blk.Kernel.Kver + " is not a kernel version"
case !blk.Kernel.Tonight:
return false, "the household has not been told about tonight (no mail, no step — `09` §3 decision 172)"
}
return true, ""
}
// runKernel is the night leg's last step. It returns the stage report (Layer "" when nothing ran). On success the box
// is rebooting when it returns.
func (l *Leg) runKernel(ctx context.Context, runID string, vmid int, trigger string, blk hub.WireOSUpdate) Report {
lg := l.log().With("run", runID, "layer", LayerKernel, "vmid", vmid, "trigger", trigger, "ring", blk.Ring)
if ok, why := kernelDue(blk, trigger); !ok {
lg.Info("osupdate: kernel step skipped — " + why)
return Report{}
}
want := blk.Kernel.Kver
st, err := l.KernelStatus(ctx, vmid)
if err != nil {
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: trigger, Ring: blk.Ring, VMID: vmid,
Mode: "apply", Outcome: "failed", HealthReason: "kernel status unreadable: " + err.Error()})
}
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: trigger, Ring: blk.Ring, VMID: vmid, Mode: "apply", ReleaseID: want}
switch {
case st.Phase == "staged" && st.To == want:
lg.Info("osupdate: kernel step — a staged kernel waits for tonight", "from", st.From, "to", st.To)
rep.Outcome, rep.Healthy = "staged", true
case blk.Ring != 0:
lg.Info("osupdate: kernel step skipped — ring 1 boots only a kernel a signed os_kernel_step staged", "phase", st.Phase, "staged", st.To, "want", want)
return Report{}
default:
wr, cerr := l.call(ctx, runID, kernelPlan("apply", vmid, map[string]any{"release_id": "ring0-" + runID,
"select": "pending-kernel", "expect_kver": want, "run_id": runID, "trigger": trigger, "ring": blk.Ring}))
rep.unsent = reportFile(l.planDir(), runID, LayerKernel, "apply")
rep.Kernel = rawOrNil(wr.Kernel)
switch {
case cerr != nil:
rep.Outcome, rep.HealthReason = "failed", cerr.Error()
return l.finish(ctx, lg, rep)
case wr.refused():
rep.Outcome, rep.Refused = "refused", wr.Refused
return l.finish(ctx, lg, rep)
case wr.failed():
rep.Outcome, rep.Refused = "failed", wr.Failed
return l.finish(ctx, lg, rep)
case wr.OutcomeHint == "nothing" || len(wr.Upgraded) == 0:
rep.Outcome, rep.Healthy = "nothing", true
return l.finish(ctx, lg, rep)
}
rep.Outcome, rep.Healthy, rep.Upgraded, rep.Authority, rep.PassSeconds = "staged", true, wr.Upgraded, wr.Authority, wr.PassSeconds
rep.RebootNeeded = true
}
rep = l.finish(ctx, lg, rep) // the hub hears "staged" BEFORE the box goes down
wr, err := l.call(ctx, runID, kernelPlan("kernel-reboot", vmid, nil))
switch {
case err != nil:
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: trigger, Ring: blk.Ring, VMID: vmid,
Mode: "kernel-reboot", ReleaseID: want, Outcome: "failed", HealthReason: "kernel-reboot: " + err.Error()})
case wr.refused() || wr.failed():
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: trigger, Ring: blk.Ring, VMID: vmid,
Mode: "kernel-reboot", ReleaseID: want, Outcome: "refused", Refused: firstRaw(wr.Refused, wr.Failed),
Kernel: rawOrNil(wr.Kernel)})
}
lg.Warn("osupdate: kernel step — the box restarts now for its one-shot boot", "to", want)
return rep
}
func firstRaw(a, b json.RawMessage) json.RawMessage {
if r := rawOrNil(a); r != nil {
return r
}
return rawOrNil(b)
}
// KernelJudge is what KernelAfterBoot needs besides the leg: the hub reachability probe is the "judging" report itself.
type KernelJudge struct {
Wait time.Duration // default DefaultKernelJudgeWait
Poll time.Duration // default 30 s
}
// KernelAfterBoot runs once at daemon start: what became of a kernel step across the boot. On the new kernel it judges
// the boot (blocking up to the wait — run it in a goroutine). vmid 0 = the guest the step recorded (it may not run yet).
func (l *Leg) KernelAfterBoot(ctx context.Context, vmid int, j KernelJudge) Report {
runID := "boot-" + l.now().UTC().Format("20060102T150405Z")
lg := l.log().With("run", runID, "layer", LayerKernel, "vmid", vmid)
wr, err := l.call(ctx, runID, kernelPlan("kernel-boot", vmid, nil))
if err != nil {
lg.Warn("osupdate: kernel after-boot check failed", "err", err)
return Report{}
}
if wr.refused() {
lg.Info("osupdate: kernel after-boot check refused (an older wrapper, or a BYO host)", "refused", string(wr.Refused))
return Report{}
}
v := parseKernel(wr.Kernel)
if vmid <= 0 {
vmid = v.VMID // after a boot the guest may not run yet — the step's own record names it
}
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid, Mode: "kernel-boot",
ReleaseID: v.To, Kernel: rawOrNil(wr.Kernel)}
switch wr.KernelEvent {
case "fell_back":
rep.Outcome, rep.HealthReason = "fell_back", v.Reason
lg.Warn("osupdate: kernel step FELL BACK — the new kernel did not come up; the box runs the old one", "from", v.From, "to", v.To)
return l.finish(ctx, lg, rep)
case "self_reverted":
rep.Outcome, rep.HealthReason = "self_reverted", v.Reason
lg.Warn("osupdate: kernel step SELF-REVERTED — back on the old kernel", "from", v.From, "to", v.To, "reason", v.Reason)
return l.finish(ctx, lg, rep)
case "revert_failed":
rep.Outcome, rep.HealthReason = "revert_failed", v.Reason
lg.Error("osupdate: kernel self-revert came back on the NEW kernel — no second revert; the operator decides", "to", v.To)
return l.finish(ctx, lg, rep)
case "judging":
return l.judgeKernel(ctx, runID, vmid, v, wr.HealthBefore, j, lg)
}
return Report{}
}
// KernelVerdict is THE one-shot boot rule (R-836; pinned by TestKernelVerdict): the host health rule (`11` §8.2 —
// the Proxmox daemons and the agent active, the customer guest running and its own rule passing, the tunnel running)
// AND the box reached the hub since this boot.
func KernelVerdict(before, after *Health, tunnel string, hubReached bool) (bool, string) {
if ok, why := HostHealthVerdict(before, after, tunnel); !ok {
return false, why
}
if !hubReached {
return false, "the box has not reached the hub since the boot"
}
return true, ""
}
func (l *Leg) judgeKernel(ctx context.Context, runID string, vmid int, v KernelView, before *Health, j KernelJudge, lg *slog.Logger) Report {
wait, poll := j.Wait, j.Poll
if wait <= 0 {
wait = DefaultKernelJudgeWait
}
if poll <= 0 {
poll = 30 * time.Second
}
lg.Info("osupdate: kernel step — judging the one-shot boot", "from", v.From, "to", v.To, "wait", wait.String())
start := l.now()
deadline := start.Add(wait)
hubReached := false
var why string
for {
if !hubReached && l.Hub != nil {
// the hub's reachability IS this report reaching it (and the operator sees the box is back on the new kernel)
body, _ := json.Marshal(Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid,
Mode: "kernel-boot", ReleaseID: v.To, Outcome: "judging", Kernel: mustRaw(v)})
rctx, cancel := context.WithTimeout(ctx, 30*time.Second)
if err := l.Hub.PostOSReport(rctx, body); err == nil {
hubReached = true
lg.Info("osupdate: kernel step — the box reached the hub on the new kernel", "after", l.now().Sub(start).Round(time.Second).String())
}
cancel()
}
var h *Health
hr, err := l.call(ctx, runID, kernelPlan("health", vmid, nil))
switch {
case err != nil:
why = "no health reading: " + err.Error()
case hr.refused():
why = "no health reading: " + string(hr.Refused)
default:
h = hr.Health
}
if h != nil {
t := hub.TunnelUnknown
if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx)
}
var ok bool
ok, why = KernelVerdict(before, h, t, hubReached)
if ok {
return l.kernelGood(ctx, runID, vmid, v, start, lg)
}
}
if !l.now().Before(deadline) || ctx.Err() != nil {
break
}
l.sleep(ctx, poll)
}
if ctx.Err() != nil {
lg.Warn("osupdate: kernel judging stopped (the agent is stopping) — the next start judges again", "reason", why)
return Report{}
}
// not healthy by the deadline: tell the hub (best effort), then ONE self-revert into the old kernel
rep := l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid,
Mode: "kernel-revert", ReleaseID: v.To, Outcome: "health_failed", HealthReason: why + " — reverting to " + v.From,
Kernel: mustRaw(v)})
lg.Error("osupdate: kernel step — the one-shot boot is NOT healthy; restarting ONCE into the old kernel", "reason", why,
"waited", wait.String(), "from", v.From, "to", v.To)
wr, err := l.call(ctx, runID, kernelPlan("kernel-revert", vmid, map[string]any{"reason": truncate(why, 280)}))
if err != nil || wr.refused() || wr.failed() {
lg.Error("osupdate: kernel self-revert did not start — the box stays on the new kernel; the operator decides",
"err", err, "refused", string(firstRaw(wr.Refused, wr.Failed)))
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid,
Mode: "kernel-revert", ReleaseID: v.To, Outcome: "revert_failed", Refused: firstRaw(wr.Refused, wr.Failed),
HealthReason: "the self-revert did not start"})
}
return rep
}
func (l *Leg) kernelGood(ctx context.Context, runID string, vmid int, v KernelView, start time.Time, lg *slog.Logger) Report {
wr, err := l.call(ctx, runID, kernelPlan("kernel-good", vmid, nil))
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid, Mode: "kernel-good",
ReleaseID: v.To}
switch {
case err != nil:
rep.Outcome, rep.HealthReason = "failed", "kernel-good: "+err.Error()
case wr.refused() || wr.failed():
rep.Outcome, rep.Refused, rep.HealthReason = "failed", firstRaw(wr.Refused, wr.Failed), "kernel-good did not move the default"
default:
rep.Outcome, rep.Healthy = "applied", true
rep.HealthReason = fmt.Sprintf("healthy %s after the agent started; the new kernel is the default", l.now().Sub(start).Round(time.Second))
}
rep.Kernel = rawOrNil(wr.Kernel)
lg.Info("osupdate: kernel step — "+rep.Outcome, "to", v.To, "reason", rep.HealthReason)
return l.finish(ctx, lg, rep)
}
func mustRaw(v any) json.RawMessage {
b, _ := json.Marshal(v)
return b
}
func truncate(s string, n int) string {
if len(s) <= n {
return s
}
return s[:n]
}
// KernelStepParams are a signed os_kernel_step's params: the exact kernel set (the wrapper compares it with the plan).
type KernelStepParams struct {
ReleaseID string `json:"release_id"`
Packages []Package `json:"packages"`
Kver string `json:"kver"`
VMID int `json:"vmid,omitempty"`
}
// KernelStepExecutor STAGES a verified os_kernel_step (signedjobs.Executor) under the heavy-op gate. It never reboots:
// the night leg reboots a staged kernel on a night the household was told about.
type KernelStepExecutor struct {
Leg *Leg
Guest func(ctx context.Context) (int, error)
Gate func(ctx context.Context) (release func(), err error)
}
// Execute implements signedjobs.Executor.
func (e KernelStepExecutor) Execute(ctx context.Context, op string, params json.RawMessage) error {
if op != OpKernelStep {
return signedjobs.ErrNoExecutor
}
so, ok := signedjobs.SignedOpFrom(ctx)
if !ok {
return fmt.Errorf("os_kernel_step: no signed envelope in the context — the wrapper could not verify it")
}
var p KernelStepParams
if err := json.Unmarshal(params, &p); err != nil || len(p.Packages) == 0 || !kverRE.MatchString(p.Kver) {
return fmt.Errorf("os_kernel_step: params must name the kernel set and its kver: %v", err)
}
vmid := p.VMID
if vmid == 0 {
if e.Guest == nil {
return fmt.Errorf("os_kernel_step: no vmid and no guest finder")
}
v, err := e.Guest(ctx)
if err != nil {
return fmt.Errorf("os_kernel_step: find the customer guest: %w", err)
}
vmid = v
}
if e.Gate != nil {
release, err := e.Gate(ctx)
if err != nil {
return fmt.Errorf("os_kernel_step: heavy-op gate busy (a backup or restore-test runs): %w", err)
}
defer release()
}
rep := e.Leg.StageKernelSigned(ctx, vmid, p, so.Blob, string(so.Sig))
if rep.Outcome == "staged" || rep.Outcome == "nothing" {
return nil
}
return fmt.Errorf("os_kernel_step: %s (%s) %s", rep.Outcome, rep.HealthReason, string(rep.Refused))
}
// StageKernelSigned stages a signed kernel set (ring 1): install + flag, no reboot.
func (l *Leg) StageKernelSigned(ctx context.Context, vmid int, p KernelStepParams, blob []byte, sig string) Report {
unlock := l.lockPass(true)
defer unlock()
l.sendUnsentLocked(ctx) // R-868
runID := l.now().UTC().Format("20060102T150405Z")
rid := p.ReleaseID
if rid == "" {
rid = "signed-" + runID
}
lg := l.log().With("run", runID, "layer", LayerKernel, "vmid", vmid, "trigger", "signed", "release", rid)
wr, err := l.call(ctx, runID, kernelPlan("apply", vmid, map[string]any{"release_id": rid, "select": "listed",
"packages": p.Packages, "expect_kver": p.Kver, "run_id": runID, "trigger": "signed", "ring": l.Block().Ring,
"signed": map[string]string{"blob_b64": base64.StdEncoding.EncodeToString(blob), "sig": sig}}))
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "signed", Ring: l.Block().Ring, VMID: vmid, Mode: "apply",
ReleaseID: rid, Kernel: rawOrNil(wr.Kernel), unsent: reportFile(l.planDir(), runID, LayerKernel, "apply")}
switch {
case err != nil:
rep.Outcome, rep.HealthReason = "failed", err.Error()
case wr.refused():
rep.Outcome, rep.Refused = "refused", wr.Refused
case wr.failed():
rep.Outcome, rep.Refused = "failed", wr.Failed
case wr.OutcomeHint == "nothing" || len(wr.Upgraded) == 0:
rep.Outcome, rep.Healthy = "nothing", true
default:
rep.Outcome, rep.Healthy, rep.Upgraded, rep.Authority, rep.PassSeconds = "staged", true, wr.Upgraded, wr.Authority, wr.PassSeconds
rep.RebootNeeded = true
}
return l.finish(ctx, lg, rep)
}
+314
View File
@@ -0,0 +1,314 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// ---- the kernel lane (R-836, `09` §3 decision 172, `11` §5.11) ----
const kOld, kNew = "7.0.2-6-pve", "7.0.14-22-pve"
func kview(phase string) json.RawMessage {
return mustRaw(KernelView{Running: kOld, Default: kOld, Phase: phase, From: kOld, To: kNew, VMID: 9201})
}
func tonight(ring int) *hub.WireOSUpdate {
return &hub.WireOSUpdate{Ring: ring, Enabled: true, Kernel: &hub.WireKernelStep{Kver: kNew, Tonight: true}}
}
func kernelCalls(w *fakeWrapper) []string {
var m []string
for _, p := range w.plans {
if p["layer"] == LayerKernel {
m = append(m, p["mode"].(string))
}
}
return m
}
// Ring 0, a told night: after the healthy host step the leg stages the pending kernel (select pending-kernel, the
// kernel the household was told about), tells the hub "staged", THEN reboots — the kernel step ends the night.
func TestKernel_Ring0ToldNightStagesThenReboots(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{
"kernel-status": {{Kernel: kview("none")}},
"apply": {{Upgraded: []Package{{Name: "proxmox-kernel-7.0", Version: "7.0.14-22"}}, Authority: "ring0", Kernel: kview("staged")}},
"kernel-reboot": {{Kernel: kview("oneshot")}},
}}
l, h := newLeg(t, w, tonight(0))
p := l.Run(context.Background(), 9201, "night")
if got := strings.Join(kernelCalls(w), ","); got != "kernel-status,apply,kernel-reboot" {
t.Fatalf("kernel calls = %s", got)
}
var ap map[string]any
for _, x := range w.plans {
if x["layer"] == LayerKernel && x["mode"] == "apply" {
ap = x
}
}
if ap["select"] != "pending-kernel" || ap["expect_kver"] != kNew || ap["lane"] != "slow" {
t.Fatalf("stage plan = %v", ap)
}
if p.Kernel.Outcome != "staged" || !p.Kernel.Healthy {
t.Fatalf("kernel report = %+v", p.Kernel)
}
last := h.reports[len(h.reports)-1]
if last.Layer != LayerKernel || last.Outcome != "staged" {
t.Fatalf("the hub must hear 'staged' before the reboot: %+v", h.reports)
}
// the kernel step is the LAST wrapper call of the night
if lp := w.plans[len(w.plans)-1]; lp["layer"] != LayerKernel || lp["mode"] != "kernel-reboot" {
t.Fatalf("the reboot must end the night, last call = %v", lp)
}
}
// No mail, no step: a kernel the household was NOT told about never runs; nor in a debug pass; nor without a block.
// COMPANION RED-PROOF (observed): drop the `!blk.Kernel.Tonight` case in kernelDue → the first sub-case fails.
func TestKernel_NoMailNoStep(t *testing.T) {
cases := map[string]struct {
blk *hub.WireOSUpdate
trigger string
}{
"not told": {&hub.WireOSUpdate{Ring: 0, Enabled: true, Kernel: &hub.WireKernelStep{Kver: kNew, Tonight: false}}, "night"},
"debug pass": {tonight(0), "debug"},
"no block": {&hub.WireOSUpdate{Ring: 0, Enabled: true}, "night"},
"switch off": {&hub.WireOSUpdate{Ring: 0, Enabled: false, Kernel: &hub.WireKernelStep{Kver: kNew, Tonight: true}}, "night"},
"bad kver": {&hub.WireOSUpdate{Ring: 0, Enabled: true, Kernel: &hub.WireKernelStep{Kver: "7.0; reboot", Tonight: true}}, "night"},
}
for name, c := range cases {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, c.blk)
p := l.Run(context.Background(), 9201, c.trigger)
if len(kernelCalls(w)) != 0 || p.Kernel.Layer != "" {
t.Fatalf("%s: a kernel step ran: %v", name, kernelCalls(w))
}
}
}
// The kernel step needs a healthy host step on an appliance, and a healthy Proxmox step when one ran.
func TestKernel_SkippedWithoutHealthyEarlierSteps(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, tonight(0))
l.Appliance = false
l.Run(context.Background(), 9201, "night")
if len(kernelCalls(w)) != 0 {
t.Fatalf("a BYO box took a kernel step: %v", kernelCalls(w))
}
w2 := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerPVE: {Upgraded: []Package{{Name: "pve-manager", Version: "9.2.21"}},
PVEManager: "9.2.2"}}} // pveversion still old → the pve step is unhealthy
l2, _ := newLeg(t, w2, tonight(0))
l2.Run(context.Background(), 9201, "night")
if len(kernelCalls(w2)) != 0 {
t.Fatalf("a kernel step ran after an unhealthy Proxmox step: %v", kernelCalls(w2))
}
}
// Ring 1 reboots only a kernel a signed job staged — never stages one itself in the night leg.
func TestKernel_Ring1RebootsOnlyASignedStage(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-status": {{Kernel: kview("none")}}}}
l, _ := newLeg(t, w, tonight(1))
l.Run(context.Background(), 9201, "night")
if got := strings.Join(kernelCalls(w), ","); got != "kernel-status" {
t.Fatalf("ring 1 without a staged kernel: calls = %s", got)
}
w2 := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-status": {{Kernel: kview("staged")}},
"kernel-reboot": {{Kernel: kview("oneshot")}}}}
l2, _ := newLeg(t, w2, tonight(1))
p := l2.Run(context.Background(), 9201, "night")
if got := strings.Join(kernelCalls(w2), ","); got != "kernel-status,kernel-reboot" || p.Kernel.Outcome != "staged" {
t.Fatalf("ring 1 with a staged kernel: calls = %s report = %+v", got, p.Kernel)
}
}
// A refused stage never reboots.
func TestKernel_RefusedStageNeverReboots(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-status": {{Kernel: kview("none")}},
"apply": {{Refused: json.RawMessage(`{"code":"R20","reason":"/boot/efi is not a mounted vfat ESP"}`)}}}}
l, h := newLeg(t, w, tonight(0))
p := l.Run(context.Background(), 9201, "night")
if got := strings.Join(kernelCalls(w), ","); got != "kernel-status,apply" || p.Kernel.Outcome != "refused" {
t.Fatalf("calls = %s report = %+v", got, p.Kernel)
}
if last := h.reports[len(h.reports)-1]; last.Layer != LayerKernel || last.Outcome != "refused" {
t.Fatalf("the hub must hear the refusal: %+v", last)
}
}
// THE one-shot boot rule: the host rule AND the hub reached. COMPANION RED-PROOF (observed): drop the hubReached
// check in KernelVerdict → the second case fails.
func TestKernelVerdict(t *testing.T) {
if ok, why := KernelVerdict(hostOK(), hostOK(), hub.TunnelRunning, true); !ok {
t.Fatalf("a healthy boot read unhealthy: %s", why)
}
if ok, why := KernelVerdict(hostOK(), hostOK(), hub.TunnelRunning, false); ok || !strings.Contains(why, "hub") {
t.Fatalf("a box that has not reached the hub must not pass: ok=%v %q", ok, why)
}
down := hostOK()
down.GuestRunning = new(bool)
if ok, _ := KernelVerdict(hostOK(), down, hub.TunnelRunning, true); ok {
t.Fatal("a guest that does not run must fail")
}
if ok, _ := KernelVerdict(hostOK(), hostOK(), hub.TunnelUnknown, true); ok {
t.Fatal("an unknown tunnel must fail (the host rule)")
}
}
func judgingLeg(t *testing.T, w *fakeWrapper) (*Leg, *fakeHub) {
if w.kernelRep == nil {
w.kernelRep = map[string][]WrapperReport{}
}
if _, ok := w.kernelRep["kernel-boot"]; !ok {
w.kernelRep["kernel-boot"] = []WrapperReport{{KernelEvent: "judging", Kernel: mustRaw(KernelView{Running: kNew,
Default: kOld, Phase: "judging", From: kOld, To: kNew, VMID: 9201}), HealthBefore: hostOK()}}
}
return newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
}
// A healthy one-shot boot: the hub hears "judging", then kernel-good, then "applied".
func TestKernelAfterBoot_HealthyBecomesTheDefault(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-good": {{Kernel: kview("good")}}}}
l, h := judgingLeg(t, w)
r := l.KernelAfterBoot(context.Background(), 0, KernelJudge{Wait: 10 * time.Minute, Poll: 30 * time.Second})
if got := strings.Join(kernelCalls(w), ","); got != "kernel-boot,health,kernel-good" {
t.Fatalf("calls = %s", got)
}
if r.Outcome != "applied" || !r.Healthy {
t.Fatalf("report = %+v", r)
}
if len(h.reports) != 2 || h.reports[0].Outcome != "judging" || h.reports[1].Outcome != "applied" {
t.Fatalf("hub reports = %+v", h.reports)
}
if w.plans[1]["vmid"] != float64(9201) {
t.Fatalf("the health reading must use the step's own guest, got %v", w.plans[1]["vmid"])
}
}
// An unhealthy one-shot boot: wait the full judge time, tell the hub, then ONE kernel-revert.
// COMPANION RED-PROOF (observed): return before the kernel-revert call in judgeKernel → "calls" fails.
func TestKernelAfterBoot_UnhealthyRevertsOnceAfterTheWait(t *testing.T) {
down := hostOK()
down.GuestRunning = new(bool)
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"health": {{Health: down}},
"kernel-revert": {{Kernel: kview("reverting")}}}}
l, h := judgingLeg(t, w)
start := l.now()
r := l.KernelAfterBoot(context.Background(), 0, KernelJudge{Wait: 10 * time.Minute, Poll: time.Minute})
calls := kernelCalls(w)
if calls[len(calls)-1] != "kernel-revert" || strings.Count(strings.Join(calls, ","), "kernel-revert") != 1 {
t.Fatalf("calls = %v", calls)
}
if waited := l.now().Sub(start); waited < 10*time.Minute {
t.Fatalf("reverted after %s — before the judge wait", waited)
}
if r.Outcome != "health_failed" || !strings.Contains(r.HealthReason, "not running") {
t.Fatalf("report = %+v", r)
}
if last := h.reports[len(h.reports)-1]; last.Outcome != "health_failed" {
t.Fatalf("the hub must hear health_failed before the revert reboot: %+v", h.reports)
}
for _, p := range w.plans {
if p["mode"] == "kernel-revert" && !strings.Contains(p["reason"].(string), "not running") {
t.Fatalf("the revert must carry the reason: %v", p)
}
}
}
// A box that never reaches the hub is not "healthy" — it reverts too.
func TestKernelAfterBoot_NoHubMeansRevert(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-revert": {{Kernel: kview("reverting")}}}}
l, _ := judgingLeg(t, w)
l.Hub = unreachableHub{}
l.KernelAfterBoot(context.Background(), 0, KernelJudge{Wait: 5 * time.Minute, Poll: time.Minute})
if c := kernelCalls(w); c[len(c)-1] != "kernel-revert" {
t.Fatalf("calls = %v", c)
}
}
type unreachableHub struct{}
func (unreachableHub) PostOSReport(context.Context, []byte) error { return context.DeadlineExceeded }
// What kernel-boot found becomes the hub's outcome, with no judging and no reboot.
func TestKernelAfterBoot_FallBackAndRevertResultsAreReported(t *testing.T) {
for ev, want := range map[string]string{"fell_back": "fell_back", "self_reverted": "self_reverted", "revert_failed": "revert_failed"} {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-boot": {{KernelEvent: ev,
Kernel: mustRaw(KernelView{Running: kOld, Default: kOld, Phase: ev, From: kOld, To: kNew, Reason: "r", VMID: 9201})}}}}
l, h := judgingLeg(t, w)
r := l.KernelAfterBoot(context.Background(), 0, KernelJudge{})
if r.Outcome != want || len(h.reports) != 1 || h.reports[0].Outcome != want {
t.Fatalf("%s: report %+v hub %+v", ev, r, h.reports)
}
if got := strings.Join(kernelCalls(w), ","); got != "kernel-boot" {
t.Fatalf("%s: calls = %s", ev, got)
}
}
// nothing to do → nothing reported
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-boot": {{KernelEvent: "none", Kernel: kview("good")}}}}
l, h := judgingLeg(t, w)
if r := l.KernelAfterBoot(context.Background(), 0, KernelJudge{}); r.Layer != "" || len(h.reports) != 0 {
t.Fatalf("an ordinary boot must report nothing: %+v %+v", r, h.reports)
}
}
// The signed executor STAGES (listed + the raw envelope + the kver) and never reboots.
func TestKernelStepExecutor_StagesNeverReboots(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"apply": {{Upgraded: []Package{{Name: "proxmox-kernel-7.0",
Version: "7.0.14-22"}}, Authority: "signed", Kernel: kview("staged")}}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
e := KernelStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
params, _ := json.Marshal(KernelStepParams{ReleaseID: "os-kernel-1", Kver: kNew,
Packages: []Package{{Name: "proxmox-kernel-7.0", Version: "7.0.14-22", Origin: PVEOrigin}}})
ctx := signedjobs.WithSignedOp(context.Background(), &reconcile.SignedOp{Blob: []byte(`{"op":"os_kernel_step"}`), Sig: []byte("SIG")})
if err := e.Execute(ctx, OpKernelStep, params); err != nil {
t.Fatal(err)
}
pp := w.plans[len(w.plans)-1]
sg, _ := pp["signed"].(map[string]any)
if pp["mode"] != "apply" || pp["select"] != "listed" || pp["expect_kver"] != kNew || sg == nil ||
sg["blob_b64"] != base64.StdEncoding.EncodeToString([]byte(`{"op":"os_kernel_step"}`)) {
t.Fatalf("plan = %v", pp)
}
if got := strings.Join(kernelCalls(w), ","); got != "apply" {
t.Fatalf("a signed stage must never reboot: %s", got)
}
if len(h.reports) != 1 || h.reports[0].Outcome != "staged" {
t.Fatalf("hub = %+v", h.reports)
}
if err := e.Execute(context.Background(), OpKernelStep, params); err == nil {
t.Fatal("no envelope must refuse")
}
bad, _ := json.Marshal(KernelStepParams{Kver: "x", Packages: []Package{{Name: "a"}}})
if err := e.Execute(ctx, OpKernelStep, bad); err == nil {
t.Fatal("a bad kver must refuse")
}
if err := e.Execute(ctx, OpPVEStep, params); err != signedjobs.ErrNoExecutor {
t.Fatalf("another op must pass through the chain: %v", err)
}
}
// os_kernel_step is never benign.
func TestKernelStep_IsDestructiveClass(t *testing.T) {
if reconcile.Classify(reconcile.ClassOSKernelStep, reconcile.Provenance{}) != reconcile.Destructive {
t.Fatal("os_kernel_step must be destructive-class (signed, operational key)")
}
}
// A kept stage report (the agent was killed mid-stage) reaches the hub as "staged" with its kernel view.
func TestKernel_KeptStageReportIsStaged(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, tonight(0))
ring := 0
rep := l.reportFromKept(context.Background(), WrapperReport{Layer: LayerKernel, Mode: "apply", Ring: &ring,
Upgraded: []Package{{Name: "proxmox-kernel-7.0", Version: "7.0.14-22"}}, Kernel: kview("staged")}, "/x/report-r-kernel-apply.json")
if rep.Outcome != "staged" || !rep.Healthy || len(rep.Kernel) == 0 {
t.Fatalf("kept = %+v", rep)
}
}
+457 -32
View File
@@ -1,5 +1,7 @@
// Package osupdate is the agent's OS-update leg (`11-os-updates.md` §8 steps 2–3): the customer GUEST's Debian fast
// lane (agent v0.140.0) and, after it in the same pass, the HOST's (agent v0.141.0).
// Package osupdate is the agent's OS-update leg (`11-os-updates.md` §8 steps 2–3, §5.8): the customer GUEST's Debian
// fast lane (agent v0.140.0), after it in the same pass the HOST's (agent v0.141.0), and then — ring 0 only — the
// guest's DOCKER engine set, the slow lane (agent v0.142.0; a ring-1 box takes a Docker step only inside a signed
// operator job, DockerStepExecutor). It also reads the box's versions for the hub's System page (Facts, R-852).
//
// It runs right after the night's successful whole-guest backup, while the backup goroutine still holds the host-wide
// heavy-op gate (so it never overlaps a backup or a restore-test, `11` C10), at most once per night. All root work is
@@ -18,6 +20,7 @@ import (
"log/slog"
"os"
"path/filepath"
"regexp"
"sort"
"strings"
"sync"
@@ -25,6 +28,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
)
// WrapperPath is the pinned sudoers vector (configs/felhom-agent.sudoers FELHOM_OSAPPLY).
@@ -35,10 +39,31 @@ const DefaultPlanDir = "/var/lib/felhom-agent/os"
// Layers.
const (
LayerGuest = "guest"
LayerHost = "host"
LayerGuest = "guest"
LayerHost = "host"
LayerDocker = "docker" // the guest's Docker engine set — slow lane (`11` §5.8)
// LayerPVE is the HOST's Proxmox userspace packages — slow lane (R-812 option A, `09` §3 decision 163, `11` §5.10):
// ring 0 in the night leg after a healthy host step, ring 1 only inside a signed os_pve_step. Never a kernel.
LayerPVE = "pve"
// LayerKernel is the HOST's kernel — the kernel lane (R-836, `09` §3 decision 172, `11` §5.11): a one-shot boot of
// the new kernel through the ESP flag, the default moved only after a healthy boot (kernel.go).
LayerKernel = "kernel"
)
// PVEOrigin is the origin apt prints for download.proxmox.com (the wrapper's PVE_ORIGIN); InstalledPVEOrigin is how
// the wrapper's inventory names the same source.
const (
PVEOrigin = "Proxmox Debian Repository"
InstalledPVEOrigin = "Proxmox"
)
// hostSlowRE mirrors the wrapper's HOST_SLOW_RE: kernel, boot, firmware and microcode names never ride the pve lane.
var hostSlowRE = regexp.MustCompile(`^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr)`)
// DockerNames are the six packages of the Docker engine set (the wrapper's DOCKER_NAMES).
var DockerNames = map[string]bool{"containerd.io": true, "docker-buildx-plugin": true, "docker-ce": true,
"docker-ce-cli": true, "docker-ce-rootless-extras": true, "docker-compose-plugin": true}
// Package is one name=version with its origin.
type Package struct {
Name string `json:"name"`
@@ -58,18 +83,22 @@ type Pending struct {
type Container struct {
State string `json:"state"`
Health string `json:"health"` // healthy | unhealthy | starting | none
ID string `json:"id,omitempty"`
}
// Health is one health reading. Guest layer: DockerOK..Containers. Host layer: HostServices, GuestRunning and the
// guest's own reading in Guest.
type Health struct {
DockerOK bool `json:"docker_ok"`
NetworkOK bool `json:"network_ok"`
Controller string `json:"controller"`
Containers map[string]Container `json:"containers"`
HostServices map[string]string `json:"host_services,omitempty"`
GuestRunning *bool `json:"guest_running,omitempty"`
Guest *Health `json:"guest,omitempty"`
DockerOK bool `json:"docker_ok"`
NetworkOK bool `json:"network_ok"`
Controller string `json:"controller"`
Containers map[string]Container `json:"containers"`
// ControllerDockerOK: the controller reaches the engine from INSIDE its container (R-858, wrapper ≥ v0.142.1;
// nil from an older wrapper = not checked). Its own health check stayed "healthy" while it was blind.
ControllerDockerOK *bool `json:"controller_docker_ok,omitempty"`
HostServices map[string]string `json:"host_services,omitempty"`
GuestRunning *bool `json:"guest_running,omitempty"`
Guest *Health `json:"guest,omitempty"`
}
// WrapperReport is the wrapper's OSAPPLY-REPORT object.
@@ -89,6 +118,29 @@ type WrapperReport struct {
HealthAfter *Health `json:"health_after"`
Health *Health `json:"health"`
PassSeconds float64 `json:"pass_seconds"`
DockerEngine string `json:"docker_engine"`
Authority string `json:"authority"`
Undo bool `json:"undo"`
LiveRestore json.RawMessage `json:"live_restore"`
Facts json.RawMessage `json:"facts"`
Bundle json.RawMessage `json:"bundle"` // the config bundle's result (R-840, mode "bundle")
// OOMCheck (R-528, `09` decision 157): the docker layer's memory-kill check, {result, oom_killed, oom_event,
// exit_code, image, detail}. Carried to the hub UNCHANGED; the agent never reads it.
OOMCheck json.RawMessage `json:"oom_check"`
// PVEManager: the pve layer — pveversion's pve-manager version after the step ("unknown" = unreadable).
PVEManager string `json:"pve_manager"`
// Kernel (R-836): the kernel layer's view {running, default, flag, phase, from, to, …}; KernelEvent what kernel-boot
// found after a boot; OutcomeHint "nothing" when no kernel was pending.
Kernel json.RawMessage `json:"kernel"`
KernelEvent string `json:"kernel_event"`
OutcomeHint string `json:"outcome_hint"`
// R-868 (v0.144.0): the agent's ids, echoed from the plan, so a report kept on disk can be sent without the
// agent process that started the pass. ReleaseID / VMID were always in the report.
RunID string `json:"run_id"`
Trigger string `json:"trigger"`
Ring *int `json:"ring"`
ReleaseID string `json:"release_id"`
VMID int `json:"vmid"`
}
func (w WrapperReport) refused() bool { return len(w.Refused) > 0 && string(w.Refused) != "null" }
@@ -116,6 +168,20 @@ type Report struct {
RebootScanned bool `json:"reboot_scanned,omitempty"` // the pass looked (host: every pass) — a false RebootNeeded then means "not needed"
Refused json.RawMessage `json:"refused,omitempty"`
PassSeconds float64 `json:"pass_seconds,omitempty"`
DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step
Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed
Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade)
// OOMCheck: docker layer — the wrapper's oom_check object, byte-for-byte (R-528; the hub decides approval on it).
// Pinned by TestDocker_OOMCheckReachesTheHubUnchanged and TestR868_KeptCopyCarriesTheOOMCheck.
OOMCheck json.RawMessage `json:"oom_check,omitempty"`
// PVEManager: pve layer — pve-manager's version after the step (the hub's System page; R-812 option A).
PVEManager string `json:"pve_manager,omitempty"`
// Kernel: kernel layer — the wrapper's kernel view, byte for byte (running, default, flag, phase, from, to). The
// outcomes of this layer: staged | applied (the new kernel is the default) | fell_back | health_failed (self-revert
// started) | self_reverted | revert_failed | judging | nothing | refused | failed.
Kernel json.RawMessage `json:"kernel,omitempty"`
unsent string // R-868: the wrapper's kept copy of this pass's report — deleted once the hub has it
}
// Reporter posts a report to the hub (*hub.Client).
@@ -142,14 +208,65 @@ type Leg struct {
block *hub.WireOSUpdate
}
// planFile / reportFile: the plan the agent writes and the copy of the report the wrapper keeps beside it (R-868).
func planFile(dir, runID, layer, mode string) string {
return filepath.Join(dir, fmt.Sprintf("plan-%s-%s-%s.json", runID, layer, mode))
}
func reportFile(dir, runID, layer, mode string) string {
return filepath.Join(dir, fmt.Sprintf("report-%s-%s-%s.json", runID, layer, mode))
}
func (l *Leg) planDir() string {
if l.PlanDir == "" {
return DefaultPlanDir
}
return l.PlanDir
}
// OnDesiredState stores the hub's os_update block (desired.RawConsumer — store only, never block).
func (l *Leg) OnDesiredState(_ context.Context, resp *hub.DesiredStateResponse) {
if resp == nil {
return
}
l.mu.Lock()
defer l.mu.Unlock()
l.block = resp.DesiredState.OSUpdate
l.mu.Unlock()
l.saveBlock(resp.DesiredState.OSUpdate)
}
// SavedBlockFile is the hub's newest os_update block as the daemon last received it (R-866, v0.144.0): the debug
// pass falls back to it when the hub cannot be reached, and says so.
const SavedBlockFile = "os-update-block.json"
type savedBlock struct {
SavedAt time.Time `json:"saved_at"`
Block *hub.WireOSUpdate `json:"block"`
}
func (l *Leg) saveBlock(b *hub.WireOSUpdate) {
dir := l.planDir()
if err := os.MkdirAll(dir, 0o700); err != nil {
return
}
body, _ := json.Marshal(savedBlock{SavedAt: l.now().UTC(), Block: b})
tmp := filepath.Join(dir, SavedBlockFile+".tmp")
if err := os.WriteFile(tmp, body, 0o600); err == nil {
_ = os.Rename(tmp, filepath.Join(dir, SavedBlockFile))
}
}
// LoadSavedBlock reads the block the daemon saved (R-866). ok=false: none saved yet.
func LoadSavedBlock(dir string) (b *hub.WireOSUpdate, savedAt time.Time, ok bool) {
raw, err := os.ReadFile(filepath.Join(dir, SavedBlockFile))
if err != nil {
return nil, time.Time{}, false
}
var s savedBlock
if json.Unmarshal(raw, &s) != nil {
return nil, time.Time{}, false
}
return s.Block, s.SavedAt, true
}
// Block returns the newest os_update block. No block (an older hub, or nothing fetched yet) = ring 1, ON, no
@@ -224,6 +341,9 @@ func HealthVerdict(before, after *Health) (bool, string) {
if after.Controller != "healthy" {
return false, "the controller is " + after.Controller
}
if after.ControllerDockerOK != nil && !*after.ControllerDockerOK {
return false, "the controller cannot reach Docker (it holds an old socket — R-858)"
}
if before == nil {
return true, ""
}
@@ -284,17 +404,87 @@ func HostHealthVerdict(before, after *Health, tunnel string) (bool, string) {
return true, ""
}
// EngineOf is the engine version `docker version` prints for a docker-ce package version: "5:29.8.2-1~debian.13~trixie"
// → "29.8.2" (no epoch, no Debian revision).
func EngineOf(pkgVersion string) string {
v := pkgVersion
if i := strings.Index(v, ":"); i >= 0 {
v = v[i+1:]
}
if i := strings.Index(v, "-"); i >= 0 {
v = v[:i]
}
return v
}
// PVEHealthVerdict is THE Proxmox-package-step health rule (R-812 option A; pinned by TestPVEHealthVerdict): the host
// rule (every host service active, the guest running and healthy, the tunnel running), plus every container running at
// the start still runs as the SAME container (a Proxmox step must not restart the household's apps), plus pveversion
// now reports the pve-manager the step installed (wantPVE "" = pve-manager was not in the step).
func PVEHealthVerdict(before, after *Health, tunnel, wantPVE, gotPVE string) (bool, string) {
if ok, why := HostHealthVerdict(before, after, tunnel); !ok {
return false, why
}
if before != nil && before.Guest != nil && after.Guest != nil {
names := make([]string, 0, len(before.Guest.Containers))
for n := range before.Guest.Containers {
names = append(names, n)
}
sort.Strings(names)
for _, n := range names {
b := before.Guest.Containers[n]
if b.State != "running" || b.ID == "" {
continue
}
if a := after.Guest.Containers[n]; a.ID != b.ID {
return false, n + " is a new container (id changed) — the Proxmox step restarted the household's app"
}
}
}
if wantPVE != "" && gotPVE != wantPVE {
return false, "pveversion reads pve-manager " + gotPVE + ", not " + wantPVE
}
return true, ""
}
// DockerHealthVerdict is THE Docker-step health rule (`11` §5.8; pinned by TestDockerHealthVerdict): the guest rule,
// plus every container running at the start still runs as the SAME container (same id — a changed id means the
// household's apps restarted, which `live-restore` exists to prevent), plus the engine now reports the version the step
// installed (wantEngine "" = no engine change expected).
func DockerHealthVerdict(before, after *Health, wantEngine, gotEngine string) (bool, string) {
if ok, why := HealthVerdict(before, after); !ok {
return false, why
}
if before != nil {
names := make([]string, 0, len(before.Containers))
for n := range before.Containers {
names = append(names, n)
}
sort.Strings(names)
for _, n := range names {
b := before.Containers[n]
if b.State != "running" || b.ID == "" {
continue
}
if a := after.Containers[n]; a.ID != b.ID {
return false, n + " is a new container (id changed) — the engine step restarted it"
}
}
}
if wantEngine != "" && gotEngine != wantEngine {
return false, "the engine is " + gotEngine + ", not " + wantEngine
}
return true, ""
}
// call writes the plan and runs the wrapper once.
func (l *Leg) call(ctx context.Context, runID string, plan map[string]any) (WrapperReport, error) {
dir := l.PlanDir
if dir == "" {
dir = DefaultPlanDir
}
dir := l.planDir()
if err := os.MkdirAll(dir, 0o700); err != nil {
return WrapperReport{}, fmt.Errorf("osupdate: plan dir: %w", err)
}
b, _ := json.Marshal(plan)
path := filepath.Join(dir, fmt.Sprintf("plan-%s-%s-%s.json", runID, plan["layer"], plan["mode"]))
path := planFile(dir, runID, fmt.Sprint(plan["layer"]), fmt.Sprint(plan["mode"]))
if err := os.WriteFile(path, b, 0o600); err != nil {
return WrapperReport{}, fmt.Errorf("osupdate: write plan: %w", err)
}
@@ -320,9 +510,129 @@ func (l *Leg) call(ctx context.Context, runID string, plan map[string]any) (Wrap
return rep, nil // a refusal / failure is IN the report (exit 2 / 3), not an error here
}
// Run is one pass: the guest layer, then (on an appliance, after a good guest step) the host layer. Returns both
// reports (host empty when skipped). trigger is "night" or "debug".
func (l *Leg) Run(ctx context.Context, vmid int, trigger string) (guest Report, host Report) {
// Pass is one leg's reports; an empty Layer means the step did not run.
type Pass struct {
Guest, Host, Docker, PVE, Kernel Report
}
// Run is one pass: the guest layer, then (on an appliance, after a good guest step) the host layer, then (ring 0
// only, after good earlier steps) the Docker engine set. trigger is "night" or "debug".
func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Pass {
unlock := l.lockPass(true)
defer unlock()
l.sendUnsentLocked(ctx) // R-868: a report a killed agent never sent goes first
g, h := l.runFast(ctx, vmid, trigger)
p := Pass{Guest: g, Host: h}
if g.Outcome == "skipped" {
return p
}
blk := l.Block()
okStep := func(r Report) bool {
return (r.Outcome == "applied" || r.Outcome == "nothing" || r.Outcome == "inventory") && r.Healthy
}
lg := l.log().With("run", g.RunID, "vmid", vmid, "trigger", trigger)
switch {
case blk.Ring != 0 || !blk.Enabled:
lg.Info("osupdate: docker step skipped — ring 1 takes an engine set only inside a signed operator job (`11` §5.8)", "ring", blk.Ring, "enabled", blk.Enabled)
case !okStep(g) || (h.Layer != "" && !okStep(h)):
lg.Warn("osupdate: docker step skipped — an earlier step did not end healthy")
default:
if err := l.EnsureLiveRestore(ctx, g.RunID, vmid); err != nil {
p.Docker = l.finish(ctx, lg, Report{RunID: g.RunID, Layer: LayerDocker, Trigger: trigger, Ring: 0, VMID: vmid,
Mode: "apply", Outcome: "failed", HealthReason: "live-restore could not be turned on: " + err.Error()})
} else {
p.Docker = l.runLayer(ctx, g.RunID, LayerDocker, vmid, trigger, blk, dockerOpts{})
}
}
// R-812 option A: the Proxmox package step — ring 0, an appliance, after a HEALTHY host step (it is a host change).
// A Docker step's outcome does not gate it (the Docker set lives in the guest). Pinned by TestPVE_*.
switch {
case blk.Ring != 0 || !blk.Enabled:
lg.Info("osupdate: pve step skipped — ring 1 takes a Proxmox set only inside a signed operator job (`11` §5.10)", "ring", blk.Ring, "enabled", blk.Enabled)
case !l.Appliance || h.Layer == "" || !okStep(h):
lg.Info("osupdate: pve step skipped — no healthy host step this pass (an appliance only)", "appliance", l.Appliance, "host_outcome", h.Outcome)
default:
p.PVE = l.runPVE(ctx, g.RunID, vmid, trigger, blk, dockerOpts{})
}
// R-836 (`09` §3 decision 172): the kernel step ENDS the night — an appliance, after a healthy host step (and a
// healthy Proxmox step when one ran), only on a night the hub marks as told. It reboots the box. Pinned by TestKernel_*.
switch {
case !l.Appliance || h.Layer == "" || !okStep(h):
lg.Info("osupdate: kernel step skipped — no healthy host step this pass (an appliance only)", "appliance", l.Appliance, "host_outcome", h.Outcome)
case p.PVE.Layer != "" && !okStep(p.PVE):
lg.Warn("osupdate: kernel step skipped — the Proxmox step did not end healthy", "pve_outcome", p.PVE.Outcome)
default:
p.Kernel = l.runKernel(ctx, g.RunID, vmid, trigger, blk)
}
return p
}
// pveDrainWait bounds how long a pve step waits for the agent's own /etc/pve writes in flight (pvegate).
var pveDrainWait = 2 * time.Minute
// runPVE runs the pve layer while holding pvegate: the agent's own /etc/pve writes wait until it ends.
func (l *Leg) runPVE(ctx context.Context, runID string, vmid int, trigger string, blk hub.WireOSUpdate, do dockerOpts) Report {
lg := l.log().With("run", runID, "layer", LayerPVE, "vmid", vmid, "trigger", trigger)
dctx, cancel := context.WithTimeout(ctx, pveDrainWait)
end, err := pvegate.Step(dctx)
cancel()
if err != nil {
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerPVE, Trigger: trigger, Ring: blk.Ring, VMID: vmid, Mode: "apply",
ReleaseID: do.releaseID, Outcome: "failed", HealthReason: "an agent write to /etc/pve did not finish in time (pvegate): " + err.Error()})
}
lg.Info("osupdate: pve step holds the /etc/pve write gate — the agent's own writes wait until it ends")
defer func() {
end()
lg.Info("osupdate: pve step released the /etc/pve write gate")
}()
return l.runLayer(ctx, runID, LayerPVE, vmid, trigger, blk, do)
}
// dockerOpts is a signed Docker step (DockerStepExecutor); the zero value is ring 0's unsigned "pending-docker".
type dockerOpts struct {
releaseID string
packages []Package
undo bool
signed map[string]string // blob_b64, sig — the wrapper verifies them ITSELF
}
// EnsureLiveRestore is the ONE-TIME step of `09` decision 87: the wrapper merges `"live-restore": true` into the guest's
// daemon.json and RELOADS docker (never a restart, R-835). A no-op when it is already on.
func (l *Leg) EnsureLiveRestore(ctx context.Context, runID string, vmid int) error {
wr, err := l.call(ctx, runID, map[string]any{"release_id": "live-restore", "layer": LayerGuest, "lane": "fast",
"vmid": vmid, "mode": "live-restore-on", "packages": []Package{}})
if err != nil {
return err
}
if wr.refused() {
return fmt.Errorf("refused: %s", wr.Refused)
}
if wr.failed() {
return fmt.Errorf("failed: %s", wr.Failed)
}
l.log().Info("osupdate: live-restore", "vmid", vmid, "result", string(wr.LiveRestore))
return nil
}
// Facts reads the box's versions through the wrapper's read-only facts mode (R-852): host Debian, kernels, held
// packages, taint, the crash guard; guest Debian, Docker engine, containerd, live-restore. Raw JSON, the wrapper's shape.
func (l *Leg) Facts(ctx context.Context, vmid int) (json.RawMessage, error) {
wr, err := l.call(ctx, "facts"+l.now().UTC().Format("150405"), map[string]any{"release_id": "facts", "layer": LayerHost,
"lane": "fast", "vmid": vmid, "mode": "facts", "packages": []Package{}})
if err != nil {
return nil, err
}
if wr.refused() {
return nil, fmt.Errorf("facts refused: %s", wr.Refused)
}
if len(wr.Facts) == 0 {
return nil, fmt.Errorf("facts: the wrapper returned none (an older wrapper?)")
}
return wr.Facts, nil
}
// runFast is the guest + host fast lane (agent v0.141.x behaviour).
func (l *Leg) runFast(ctx context.Context, vmid int, trigger string) (guest Report, host Report) {
runID := l.now().UTC().Format("20060102T150405Z")
lg := l.log().With("run", runID, "vmid", vmid, "trigger", trigger)
if trigger == "night" && l.StatePath != "" {
@@ -338,7 +648,7 @@ func (l *Leg) Run(ctx context.Context, vmid int, trigger string) (guest Report,
}
}
blk := l.Block()
guest = l.runLayer(ctx, runID, LayerGuest, vmid, trigger, blk)
guest = l.runLayer(ctx, runID, LayerGuest, vmid, trigger, blk, dockerOpts{})
if trigger == "night" && l.StatePath != "" {
_ = os.WriteFile(l.StatePath, []byte(l.now().UTC().Format(time.RFC3339)), 0o600)
}
@@ -348,20 +658,27 @@ func (l *Leg) Run(ctx context.Context, vmid int, trigger string) (guest Report,
case !(guest.Outcome == "applied" || guest.Outcome == "nothing" || guest.Outcome == "inventory") || !guest.Healthy:
lg.Warn("osupdate: host step skipped — the guest step did not end healthy", "guest_outcome", guest.Outcome, "reason", guest.HealthReason)
default:
host = l.runLayer(ctx, runID, LayerHost, vmid, trigger, blk)
host = l.runLayer(ctx, runID, LayerHost, vmid, trigger, blk, dockerOpts{})
}
return guest, host
}
func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigger string, blk hub.WireOSUpdate) Report {
func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigger string, blk hub.WireOSUpdate, do dockerOpts) Report {
rel := hub.WireOSRelease{ID: "ring0-" + runID}
var wire *hub.WireOSRelease
if layer == LayerGuest {
switch layer {
case LayerGuest:
wire = blk.Release
} else {
case LayerHost:
wire = blk.HostRelease
}
if blk.Ring == 1 {
lane := "fast"
if layer == LayerDocker || layer == LayerPVE {
lane = "slow"
if do.signed != nil {
rel = hub.WireOSRelease{ID: do.releaseID}
}
} else if blk.Ring == 1 {
rel = hub.WireOSRelease{}
if wire != nil {
rel = *wire
@@ -371,13 +688,31 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
lg := l.log().With("run", runID, "layer", layer, "vmid", vmid, "ring", blk.Ring, "trigger", trigger)
lg.Info("osupdate: START", "enabled", blk.Enabled, "release", rel.ID)
plan := map[string]any{"release_id": rel.ID, "layer": layer, "lane": "fast", "vmid": vmid, "snapshot": rel.Snapshot,
"packages": []Package{}, "mode": "apply", "select": "listed"}
plan := map[string]any{"release_id": rel.ID, "layer": layer, "lane": lane, "vmid": vmid, "snapshot": rel.Snapshot,
"packages": []Package{}, "mode": "apply", "select": "listed",
"run_id": runID, "trigger": trigger, "ring": blk.Ring} // R-868: echoed into the wrapper's kept copy
if rel.ID == "" {
plan["release_id"] = "none"
}
planned := map[string]bool{}
switch {
case layer == LayerDocker && do.signed != nil:
plan["packages"], plan["signed"] = do.packages, do.signed
if do.undo {
plan["undo"] = true
}
for _, p := range do.packages {
planned[p.Name] = true
}
case layer == LayerDocker:
plan["select"] = "pending-docker" // ring 0: the wrapper checks the box's ROOT-OWNED ring-0 mark itself
case layer == LayerPVE && do.signed != nil:
plan["packages"], plan["signed"] = do.packages, do.signed
for _, p := range do.packages {
planned[p.Name] = true
}
case layer == LayerPVE:
plan["select"] = "pending-pve" // ring 0: the same root-owned mark; the wrapper picks installed Proxmox userspace
case !blk.Enabled:
plan["mode"] = "inventory"
lg.Info("osupdate: switched OFF for this box — reporting only")
@@ -395,6 +730,9 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
}
rep.Mode = plan["mode"].(string)
wr, err := l.call(ctx, runID, plan)
if rep.Mode == "apply" {
rep.unsent = reportFile(l.planDir(), runID, layer, rep.Mode) // the wrapper kept a copy (R-868)
}
switch {
case err != nil:
rep.Outcome, rep.HealthReason = "failed", err.Error()
@@ -406,6 +744,9 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
rep.Outcome, rep.Refused = "failed", wr.Failed
}
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
rep.OOMCheck = rawOrNil(wr.OOMCheck)
rep.PVEManager = wr.PVEManager
if rep.Outcome == "" {
switch {
case rep.Mode == "inventory" && !blk.Enabled:
@@ -423,12 +764,27 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
}
// Health: compare with the start of the pass; give restarted services time (only after an install).
cur := wr.HealthAfter
wantEngine, wantPVE := "", ""
for _, u := range wr.Upgraded {
if u.Name == "docker-ce" {
wantEngine = EngineOf(u.Version)
}
if u.Name == "pve-manager" {
wantPVE = u.Version
}
}
verdict := func(h *Health) (bool, string) {
if layer == LayerHost {
if layer == LayerDocker {
return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine)
}
if layer == LayerHost || layer == LayerPVE {
t := hub.TunnelUnknown
if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx)
}
if layer == LayerPVE {
return PVEHealthVerdict(wr.HealthBefore, h, t, wantPVE, wr.PVEManager)
}
return HostHealthVerdict(wr.HealthBefore, h, t)
}
return HealthVerdict(wr.HealthBefore, h)
@@ -449,7 +805,7 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
break
}
l.sleep(ctx, poll)
hp := map[string]any{"release_id": plan["release_id"], "layer": layer, "lane": "fast", "vmid": vmid, "mode": "health", "packages": []Package{}}
hp := map[string]any{"release_id": plan["release_id"], "layer": layer, "lane": lane, "vmid": vmid, "mode": "health", "packages": []Package{}}
hr, herr := l.call(ctx, runID, hp)
if herr == nil && hr.Health != nil {
cur = hr.Health
@@ -464,10 +820,75 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
rep.Installed, rep.Pending = wr.Installed, wr.Pending
rep.RestartNeeded, rep.DockerRestartNeeded, rep.RebootNeeded = wr.RestartNeeded, wr.DockerRestartNeeded, wr.RebootNeeded
rep.RebootScanned = wr.RebootScanned
rep.NotCovered = notCovered(wr.Pending, blk.Ring, planned)
if layer == LayerDocker {
// the docker report carries the engine set only (the guest report already carries the Debian packages)
rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending)
rep.NotCovered = nil
} else if layer == LayerPVE {
// the pve report carries the Proxmox userspace set only — the hub's candidate is built from it
rep.Installed, rep.Pending = onlyPVE(wr.Installed), onlyPVEPending(wr.Pending)
rep.NotCovered = nil
} else {
rep.NotCovered = notCovered(wr.Pending, blk.Ring, planned)
}
return l.finish(ctx, lg, rep)
}
// rawOrNil: a wrapper field that is absent or JSON null stays out of the hub report (omitempty).
func rawOrNil(m json.RawMessage) json.RawMessage {
if len(m) == 0 || string(m) == "null" {
return nil
}
return m
}
func onlyDocker(in []Package) []Package {
var out []Package
for _, p := range in {
if DockerNames[p.Name] {
out = append(out, p)
}
}
return out
}
// onlyPVE keeps the installed Proxmox-origin packages the pve lane may touch (never a kernel / boot / firmware name).
func onlyPVE(in []Package) []Package {
var out []Package
for _, p := range in {
if (p.Origin == InstalledPVEOrigin || p.Origin == PVEOrigin) && !hostSlowRE.MatchString(p.Name) && !DockerNames[p.Name] {
out = append(out, p)
}
}
return out
}
func onlyPVEPending(in []Pending) []Pending {
var out []Pending
for _, p := range in {
if p.From == "" || hostSlowRE.MatchString(p.Name) || DockerNames[p.Name] {
continue
}
for _, o := range p.Origin {
if o == PVEOrigin {
out = append(out, p)
break
}
}
}
return out
}
func onlyDockerPending(in []Pending) []Pending {
var out []Pending
for _, p := range in {
if DockerNames[p.Name] {
out = append(out, p)
}
}
return out
}
// notCovered lists pending updates no approved release covers: in ring 0 everything outside the fast lane; in ring 1
// also every fast-lane update the release did not name.
func notCovered(pending []Pending, ring int, planned map[string]bool) []string {
@@ -489,7 +910,11 @@ func (l *Leg) finish(ctx context.Context, lg *slog.Logger, rep Report) Report {
rctx, cancel := context.WithTimeout(context.WithoutCancel(ctx), time.Minute)
defer cancel()
if err := l.Hub.PostOSReport(rctx, body); err != nil {
lg.Warn("osupdate: reporting to the hub failed (the run itself is done)", "err", err)
lg.Warn("osupdate: reporting to the hub failed (the run itself is done; the kept copy is sent at the next start or pass)", "err", err)
} else if rep.unsent != "" {
if rerr := os.Remove(rep.unsent); rerr != nil && !os.IsNotExist(rerr) {
lg.Warn("osupdate: could not delete the sent report's kept copy (it may be sent twice)", "path", rep.unsent, "err", rerr)
}
}
}
return rep
+235 -34
View File
@@ -3,6 +3,7 @@ package osupdate
import (
"context"
"encoding/json"
"fmt"
"io"
"os"
"os/exec"
@@ -12,15 +13,19 @@ import (
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
)
// fakeWrapper plays /usr/local/sbin/felhom-os-apply: it reads the plan the leg wrote and answers per layer and mode.
type fakeWrapper struct {
t *testing.T
pending []Pending
applyRep map[string]WrapperReport // per layer
healthSeq map[string][]*Health // per layer: answers to successive "health" calls
plans []map[string]any
t *testing.T
pending []Pending
applyRep map[string]WrapperReport // per layer
healthSeq map[string][]*Health // per layer: answers to successive "health" calls
plans []map[string]any
keep bool // R-868: like the real wrapper, keep an apply report beside the plan
pveGateHeld bool
kernelRep map[string][]WrapperReport // R-836: per kernel-layer mode, successive answers (the last one repeats)
}
func yes() *bool { b := true; return &b }
@@ -48,9 +53,22 @@ func (f *fakeWrapper) Run(_ context.Context, name string, args ...string) ([]byt
f.plans = append(f.plans, plan)
layer := plan["layer"].(string)
ok := guestOK()
if layer == LayerHost {
if layer == LayerHost || layer == LayerPVE || layer == LayerKernel {
ok = hostOK()
}
if layer == LayerKernel {
if seq := f.kernelRep[plan["mode"].(string)]; len(seq) > 0 {
rep := seq[0]
if len(seq) > 1 {
f.kernelRep[plan["mode"].(string)] = seq[1:]
}
out, _ := json.Marshal(rep)
return []byte("OSAPPLY-REPORT " + string(out) + "\n"), []byte("os-apply: DONE rc=0\n"), nil
}
}
if layer == LayerPVE && plan["mode"] == "apply" {
f.pveGateHeld = pvegate.Stepping() // R-812: the /etc/pve write gate must be held while the pve step runs
}
var rep WrapperReport
switch plan["mode"] {
case "inventory":
@@ -72,6 +90,22 @@ func (f *fakeWrapper) Run(_ context.Context, name string, args ...string) ([]byt
rep.Health = ok
}
}
if f.keep && plan["mode"] == "apply" {
kept := rep
kept.Layer, kept.RunID, _ = layer, fmt.Sprint(plan["run_id"]), 0
if tr, ok := plan["trigger"].(string); ok {
kept.Trigger = tr
}
if r, ok := plan["ring"].(float64); ok {
ri := int(r)
kept.Ring = &ri
}
kb, _ := json.Marshal(kept)
dst := filepath.Join(filepath.Dir(args[1]), "report-"+strings.TrimPrefix(filepath.Base(args[1]), "plan-"))
if err := os.WriteFile(dst, kb, 0o600); err != nil {
f.t.Fatal(err)
}
}
out, _ := json.Marshal(rep)
return []byte("OSAPPLY-REPORT " + string(out) + "\n"), []byte("os-apply: DONE rc=0\n"), nil
}
@@ -80,9 +114,13 @@ func (f *fakeWrapper) RunStdin(ctx context.Context, _ io.Reader, name string, ar
return f.Run(ctx, name, args...)
}
type fakeHub struct{ reports []Report }
type fakeHub struct {
reports []Report
bodies [][]byte // the exact bytes posted (R-528: the oom_check object must arrive unchanged)
}
func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error {
h.bodies = append(h.bodies, append([]byte(nil), body...))
var r Report
json.Unmarshal(body, &r)
h.reports = append(h.reports, r)
@@ -134,14 +172,14 @@ func TestRing0_OneCallPerLayer(t *testing.T) {
LayerHost: {Upgraded: []Package{{Name: "openssl", Version: "u3"}}},
}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
g, ho := l.Run(context.Background(), 9201, "night")
g, ho := run2(l, "night")
if g.Outcome != "applied" || !g.Healthy || ho.Outcome != "applied" || !ho.Healthy {
t.Fatalf("guest %+v\nhost %+v", g, ho)
}
if calls(w) != "guest:apply,host:apply" {
t.Fatalf("calls = %s, want one apply per layer, guest first", calls(w))
if calls(w) != "guest:apply,host:apply,guest:live-restore-on,docker:apply,pve:apply" {
t.Fatalf("calls = %s, want one apply per layer, guest first, then live-restore, the ring-0 docker step and the pve step", calls(w))
}
for _, p := range w.plans {
for _, p := range w.plans[:2] {
if p["select"] != "pending-fast" || p["snapshot"] != "" || len(p["packages"].([]any)) != 0 {
t.Fatalf("ring-0 plan = %v", p)
}
@@ -149,7 +187,7 @@ func TestRing0_OneCallPerLayer(t *testing.T) {
if len(g.NotCovered) != 1 || g.NotCovered[0] != "docker-ce" {
t.Fatalf("not covered = %v", g.NotCovered)
}
if len(h.reports) != 2 || h.reports[0].Layer != LayerGuest || h.reports[1].Layer != LayerHost {
if len(h.reports) != 4 || h.reports[0].Layer != LayerGuest || h.reports[1].Layer != LayerHost || h.reports[2].Layer != LayerDocker || h.reports[3].Layer != LayerPVE {
t.Fatalf("hub got %+v", h.reports)
}
}
@@ -163,7 +201,7 @@ func TestRing1_EachLayerItsOwnRelease(t *testing.T) {
gr := &hub.WireOSRelease{ID: "os-g", Snapshot: "20261004T080000Z", Packages: []hub.WireOSPackage{{Name: "libc6", Version: "g-u4", Origin: "Debian"}}}
hr := &hub.WireOSRelease{ID: "os-h", Snapshot: "20261004T090000Z", Packages: []hub.WireOSPackage{{Name: "openssl", Version: "h-u3", Origin: "Debian-Security"}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true, Release: gr, HostRelease: hr})
g, ho := l.Run(context.Background(), 9201, "night")
g, ho := run2(l, "night")
if g.ReleaseID != "os-g" || ho.ReleaseID != "os-h" {
t.Fatalf("release ids %q %q", g.ReleaseID, ho.ReleaseID)
}
@@ -182,7 +220,7 @@ func TestRing1_EachLayerItsOwnRelease(t *testing.T) {
func TestRing1_NoReleaseIsInventory(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
g, ho := l.Run(context.Background(), 9201, "night")
g, ho := run2(l, "night")
if g.Outcome != "nothing" || ho.Outcome != "nothing" || calls(w) != "guest:inventory,host:inventory" {
t.Fatalf("g=%+v h=%+v calls=%s", g, ho, calls(w))
}
@@ -192,7 +230,7 @@ func TestRing1_NoReleaseIsInventory(t *testing.T) {
func TestNoBlock_IsRing1Nothing(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend}
l, _ := newLeg(t, w, nil)
if g, _ := l.Run(context.Background(), 9201, "night"); g.Outcome != "nothing" || g.Ring != 1 {
if g, _ := run2(l, "night"); g.Outcome != "nothing" || g.Ring != 1 {
t.Fatalf("g=%+v calls=%s", g, calls(w))
}
}
@@ -201,7 +239,7 @@ func TestNoBlock_IsRing1Nothing(t *testing.T) {
func TestSwitchOff_ReportsOnly(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: false})
g, ho := l.Run(context.Background(), 9201, "night")
g, ho := run2(l, "night")
if g.Outcome != "inventory" || ho.Outcome != "inventory" || calls(w) != "guest:inventory,host:inventory" || len(h.reports) != 2 {
t.Fatalf("g=%+v h=%+v calls=%s", g, ho, calls(w))
}
@@ -212,8 +250,9 @@ func TestBYO_NoHostPlan(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Appliance = false
_, ho := l.Run(context.Background(), 9201, "night")
if ho.Outcome != "" || calls(w) != "guest:apply" || len(h.reports) != 1 {
_, ho := run2(l, "night")
// the guest (and so its Docker engine) is ours on a BYO box too: only the HOST is the owner's
if ho.Outcome != "" || calls(w) != "guest:apply,guest:live-restore-on,docker:apply" || len(h.reports) != 2 {
t.Fatalf("a BYO box got a host step: host=%+v calls=%s", ho, calls(w))
}
}
@@ -223,7 +262,7 @@ func TestGuestFailure_SkipsTheHost(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{
LayerGuest: {Refused: json.RawMessage(`{"code":"R6","reason":"x"}`)}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
g, ho := l.Run(context.Background(), 9201, "night")
g, ho := run2(l, "night")
if g.Outcome != "refused" || ho.Outcome != "" || calls(w) != "guest:apply" {
t.Fatalf("g=%+v h=%+v calls=%s", g, ho, calls(w))
}
@@ -236,7 +275,7 @@ func TestHealth_FailsAfterTheWait(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}, HealthAfter: bad}},
healthSeq: map[string][]*Health{LayerGuest: {bad, bad, bad, bad, bad, bad, bad, bad}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
g, ho := l.Run(context.Background(), 9201, "night")
g, ho := run2(l, "night")
if g.Outcome != "health_failed" || g.Healthy || !strings.Contains(g.HealthReason, "app was running") || ho.Outcome != "" {
t.Fatalf("g=%+v h=%+v", g, ho)
}
@@ -251,7 +290,7 @@ func TestHealth_RecoversInsideTheWait(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}, HealthAfter: starting}},
healthSeq: map[string][]*Health{LayerGuest: {starting}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
if g, _ := l.Run(context.Background(), 9201, "night"); g.Outcome != "applied" || !g.Healthy {
if g, _ := run2(l, "night"); g.Outcome != "applied" || !g.Healthy {
t.Fatalf("g = %+v", g)
}
}
@@ -261,7 +300,7 @@ func TestHost_TunnelDownFailsTheHostStep(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerHost: {Upgraded: []Package{{Name: "openssl"}}}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Tunnel = fakeTunnel{hub.TunnelNotRunning}
_, ho := l.Run(context.Background(), 9201, "night")
_, ho := run2(l, "night")
if ho.Outcome != "health_failed" || !strings.Contains(ho.HealthReason, "tunnel") {
t.Fatalf("host = %+v", ho)
}
@@ -328,12 +367,12 @@ func TestHostHealthVerdict(t *testing.T) {
func TestOncePerNight(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Run(context.Background(), 9201, "night")
run2(l, "night")
n := len(w.plans)
if g, _ := l.Run(context.Background(), 9201, "night"); g.Outcome != "skipped" || len(w.plans) != n {
if g, _ := run2(l, "night"); g.Outcome != "skipped" || len(w.plans) != n {
t.Fatalf("a second night run in the same night ran: %+v", g)
}
if g, _ := l.Run(context.Background(), 9201, "debug"); g.Outcome == "skipped" {
if g, _ := run2(l, "debug"); g.Outcome == "skipped" {
t.Fatal("the debug action must not be throttled")
}
}
@@ -344,13 +383,16 @@ func TestWrapperSuite(t *testing.T) {
if err != nil {
t.Skip("python3 not available")
}
cmd := exec.Command(py, "-B", "../../configs/test_felhom_os_apply.py")
out, err := cmd.CombinedOutput()
if err != nil {
t.Fatalf("wrapper suite failed: %v\n%s", err, out)
}
if !strings.Contains(string(out), "OK") {
t.Fatalf("wrapper suite did not report OK:\n%s", out)
// the OS wrapper and (agent v0.142.0) the crash guard — both root programs in configs/ with their own suites
for _, suite := range []string{"../../configs/test_felhom_os_apply.py", "../../configs/test_felhom_crash_guard.py", "../../configs/test_felhom_config_bundle.py", "../../configs/test_felhom_priv_apply.py"} {
cmd := exec.Command(py, "-B", suite)
out, err := cmd.CombinedOutput()
if err != nil {
t.Fatalf("%s failed: %v\n%s", suite, err, out)
}
if !strings.Contains(string(out), "OK") {
t.Fatalf("%s did not report OK:\n%s", suite, out)
}
}
}
@@ -362,8 +404,167 @@ func TestHostReport_CarriesRebootScanned(t *testing.T) {
LayerHost: {RebootScanned: true, RebootNeeded: true, RestartNeeded: []string{"lxc-start"}},
}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Run(context.Background(), 9201, "night")
if len(h.reports) != 2 || !h.reports[1].RebootScanned || !h.reports[1].RebootNeeded || h.reports[0].RebootScanned {
run2(l, "night")
if len(h.reports) != 4 || !h.reports[1].RebootScanned || !h.reports[1].RebootNeeded || h.reports[0].RebootScanned {
t.Fatalf("hub got %+v", h.reports)
}
}
// run2 is the guest + host reports of one pass (the tests written before the docker step).
func run2(l *Leg, trigger string) (Report, Report) {
p := l.Run(context.Background(), 9201, trigger)
return p.Guest, p.Host
}
// ---- the Docker step (`11` §5.8, agent v0.142.0) ----
// Ring 1 never takes an engine step in the night leg — only inside a signed operator job. Red-proof: drop the
// `blk.Ring != 0` case in Run and the ring-1 pass makes a docker call.
func TestDocker_Ring1NightLegNeverSteps(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
if p.Docker.Layer != "" || strings.Contains(calls(w), "docker") || strings.Contains(calls(w), "live-restore") {
t.Fatalf("ring 1 took a docker step: %s", calls(w))
}
}
// An unhealthy earlier step skips the docker step.
func TestDocker_SkippedAfterAnUnhealthyStep(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}}},
healthSeq: map[string][]*Health{}}
bad := guestOK()
bad.Controller = "unhealthy"
w.applyRep[LayerGuest] = WrapperReport{Upgraded: []Package{{Name: "libc6"}}, HealthAfter: bad}
w.healthSeq[LayerGuest] = []*Health{bad, bad, bad, bad, bad, bad, bad, bad}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
if p.Docker.Layer != "" || strings.Contains(calls(w), "docker") {
t.Fatalf("docker step ran after an unhealthy guest step: %s", calls(w))
}
}
// The docker plan is the slow lane, pending-docker for ring 0; the report carries only the engine set.
func TestDocker_Ring0PlanAndReport(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}},
Installed: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie", Origin: "Docker"}, {Name: "libc6", Version: "u4", Origin: "Debian"}},
DockerEngine: "29.8.2", Authority: "ring0"}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
var dp map[string]any
for _, x := range w.plans {
if x["layer"] == "docker" {
dp = x
}
}
if dp["layer"] != "docker" || dp["lane"] != "slow" || dp["select"] != "pending-docker" {
t.Fatalf("docker plan = %v", dp)
}
d := p.Docker
if d.Outcome != "applied" || !d.Healthy || d.DockerEngine != "29.8.2" || len(d.Installed) != 1 || d.Installed[0].Name != "docker-ce" {
t.Fatalf("docker report = %+v", d)
}
}
// THE docker health rule. Red-proof: drop the id comparison (or the engine check) in DockerHealthVerdict and a case fails.
func TestDockerHealthVerdict(t *testing.T) {
before := guestOK()
before.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"},
"app": {State: "running", Health: "healthy", ID: "b"}}
same := guestOK()
same.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"},
"app": {State: "running", Health: "healthy", ID: "b"}}
moved := guestOK()
moved.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"},
"app": {State: "running", Health: "healthy", ID: "c"}}
if ok, why := DockerHealthVerdict(before, same, "29.8.2", "29.8.2"); !ok {
t.Fatalf("same ids, right engine: %s", why)
}
if ok, _ := DockerHealthVerdict(before, moved, "29.8.2", "29.8.2"); ok {
t.Fatal("a changed container id passed — live-restore failed and the apps restarted")
}
if ok, _ := DockerHealthVerdict(before, same, "29.8.2", "29.7.2"); ok {
t.Fatal("the engine did not move and the step passed")
}
if EngineOf("5:29.8.2-1~debian.13~trixie") != "29.8.2" {
t.Fatalf("EngineOf = %q", EngineOf("5:29.8.2-1~debian.13~trixie"))
}
}
// A changed id after the step → health_failed (the consequence, not only the verdict).
func TestDocker_ChangedIDIsHealthFailed(t *testing.T) {
before := guestOK()
before.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"}}
after := guestOK()
after.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "z"}}
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1"}}, DockerEngine: "29.8.2",
HealthBefore: before, HealthAfter: after}}, healthSeq: map[string][]*Health{}}
w.healthSeq[LayerDocker] = []*Health{after, after, after, after, after, after, after, after}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
if p.Docker.Outcome != "health_failed" || !strings.Contains(p.Docker.HealthReason, "id changed") {
t.Fatalf("docker = %+v", p.Docker)
}
}
// R-858: a controller that cannot reach Docker fails the health rule even though its own check says healthy.
// Red-proof: drop the ControllerDockerOK check in HealthVerdict and this fails.
func TestHealthVerdict_ControllerBlindToDockerFails(t *testing.T) {
no, yes2 := false, true
after := guestOK()
after.ControllerDockerOK = &no
if ok, why := HealthVerdict(guestOK(), after); ok || !strings.Contains(why, "R-858") {
t.Fatalf("a blind controller passed: %v %q", ok, why)
}
if ok, _ := DockerHealthVerdict(guestOK(), after, "", ""); ok {
t.Fatal("the Docker rule passed a blind controller")
}
after.ControllerDockerOK = &yes2
if ok, why := HealthVerdict(guestOK(), after); !ok {
t.Fatalf("a seeing controller failed: %s", why)
}
if ok, _ := HealthVerdict(guestOK(), guestOK()); !ok {
t.Fatal("an older wrapper (no field) must not fail")
}
}
// R-528 (`09` decision 157): the wrapper's oom_check object reaches the hub's docker report byte-for-byte; the guest
// and host reports carry none. COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in runLayer → "no
// oom_check in the docker report".
const oomCheckWire = `{"detail":"the engine reported the memory kill: OOMKilled=true and the oom event","exit_code":137,"image":"gitea.dooplex.hu/admin/felhom-controller:0.300.0","oom_event":true,"oom_killed":true,"result":"pass"}`
func TestDocker_OOMCheckReachesTheHubUnchanged(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}},
DockerEngine: "29.8.2", Authority: "ring0", OOMCheck: json.RawMessage(oomCheckWire)}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Run(context.Background(), 9201, "night")
found := false
for _, b := range h.bodies {
var m map[string]json.RawMessage
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
var layer string
json.Unmarshal(m["layer"], &layer)
oc, has := m["oom_check"]
if layer != LayerDocker {
if has {
t.Fatalf("the %s report carries an oom_check: %s", layer, oc)
}
continue
}
found = true
if !has {
t.Fatalf("no oom_check in the docker report: %s", b)
}
if string(oc) != oomCheckWire {
t.Fatalf("oom_check changed on the way:\n got %s\nwant %s", oc, oomCheckWire)
}
}
if !found {
t.Fatalf("no docker report posted: %s", calls(w))
}
}
+152
View File
@@ -0,0 +1,152 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// ---- the Proxmox package step (R-812 option A, `09` §3 decision 163, `11` §5.10) ----
// Ring 0: after a healthy host step the leg runs the pve layer — slow lane, select pending-pve — while holding the
// /etc/pve write gate; the report carries only Proxmox userspace packages and pve-manager's version.
//
// COMPANION RED-PROOF (observed): call runLayer instead of runPVE in Run → "the /etc/pve write gate was not held".
func TestPVE_Ring0PlanGateAndReport(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerPVE: {
Upgraded: []Package{{Name: "pve-manager", Version: "9.2.21"}},
Installed: []Package{{Name: "pve-manager", Version: "9.2.21", Origin: "Proxmox"},
{Name: "proxmox-kernel-helper", Version: "9.0.4", Origin: "Proxmox"}, {Name: "libc6", Version: "u4", Origin: "Debian"}},
Pending: []Pending{{Name: "qemu-server", From: "9.0.1", To: "9.0.9", Origin: []string{PVEOrigin}},
{Name: "proxmox-kernel-7.0", From: "7.0.2", To: "7.0.14", Origin: []string{PVEOrigin}},
{Name: "libc6", From: "u3", To: "u4", Origin: []string{"Debian"}}},
PVEManager: "9.2.21", Authority: "ring0"}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
var pp map[string]any
for _, x := range w.plans {
if x["layer"] == LayerPVE {
pp = x
}
}
if pp == nil || pp["lane"] != "slow" || pp["select"] != "pending-pve" {
t.Fatalf("pve plan = %v (calls %s)", pp, calls(w))
}
if !w.pveGateHeld {
t.Fatal("the /etc/pve write gate was not held while the pve step ran")
}
if pvegate.Stepping() {
t.Fatal("the gate must be released after the step")
}
r := p.PVE
if r.Outcome != "applied" || !r.Healthy || r.PVEManager != "9.2.21" {
t.Fatalf("pve report = %+v", r)
}
if len(r.Installed) != 1 || r.Installed[0].Name != "pve-manager" || len(r.Pending) != 1 || r.Pending[0].Name != "qemu-server" {
t.Fatalf("the pve report must carry Proxmox userspace only: installed=%v pending=%v", r.Installed, r.Pending)
}
if h.reports[len(h.reports)-1].Layer != LayerPVE {
t.Fatalf("the hub must get the pve report: %+v", h.reports)
}
}
// Ring 1 never takes a Proxmox step in the night leg.
func TestPVE_Ring1NightLegNeverSteps(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
if p := l.Run(context.Background(), 9201, "night"); p.PVE.Layer != "" || strings.Contains(calls(w), "pve") {
t.Fatalf("ring 1 took a pve step: %s", calls(w))
}
}
// No healthy host step (a BYO box, or an unhealthy host step) → no pve step.
func TestPVE_SkippedWithoutAHealthyHostStep(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Appliance = false
if p := l.Run(context.Background(), 9201, "night"); p.PVE.Layer != "" || strings.Contains(calls(w), "pve") {
t.Fatalf("a BYO box took a pve step: %s", calls(w))
}
w2 := &fakeWrapper{t: t}
l2, _ := newLeg(t, w2, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l2.Tunnel = fakeTunnel{"stopped"} // the host step reads unhealthy
w2.applyRep = map[string]WrapperReport{LayerHost: {Upgraded: []Package{{Name: "libc6", Version: "u4"}}}}
if p := l2.Run(context.Background(), 9201, "night"); p.PVE.Layer != "" || strings.Contains(calls(w2), "pve") {
t.Fatalf("a pve step ran after an unhealthy host step: %s", calls(w2))
}
}
// A write in flight that never finishes makes the pve step give up (failed), never run without the gate.
func TestPVE_GivesUpWhenAWriteDoesNotFinish(t *testing.T) {
old := pveDrainWait
pveDrainWait = 50_000_000 // 50 ms
defer func() { pveDrainWait = old }()
rel, _, _ := pvegate.Write(context.Background())
defer rel()
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
if p.PVE.Outcome != "failed" || strings.Contains(calls(w), "pve") {
t.Fatalf("the pve step must fail without a wrapper call: %+v calls=%s", p.PVE, calls(w))
}
}
// THE pve health rule. COMPANION RED-PROOF (observed): drop the container-id loop or the pve-manager check in
// PVEHealthVerdict → the matching case below fails.
func TestPVEHealthVerdict(t *testing.T) {
before, after := hostOK(), hostOK()
before.Guest.Containers["app"] = Container{State: "running", Health: "healthy", ID: "a1"}
after.Guest.Containers["app"] = Container{State: "running", Health: "healthy", ID: "a1"}
if ok, why := PVEHealthVerdict(before, after, hub.TunnelRunning, "9.2.21", "9.2.21"); !ok {
t.Fatalf("healthy step read unhealthy: %s", why)
}
if ok, _ := PVEHealthVerdict(before, after, hub.TunnelRunning, "9.2.21", "9.2.2"); ok {
t.Fatal("pveversion still on the old pve-manager must fail")
}
after.Guest.Containers["app"] = Container{State: "running", Health: "healthy", ID: "b2"}
if ok, why := PVEHealthVerdict(before, after, hub.TunnelRunning, "", "9.2.2"); ok || !strings.Contains(why, "id changed") {
t.Fatalf("an app restarted by the Proxmox step must fail, got ok=%v %q", ok, why)
}
}
// The signed executor hands the RAW envelope and the exact list to the wrapper's pve layer.
func TestPVEStepExecutor_PassesTheSignedEnvelope(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerPVE: {
Upgraded: []Package{{Name: "pve-manager", Version: "9.2.21"}}, PVEManager: "9.2.21", Authority: "signed"}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
e := PVEStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
params, _ := json.Marshal(PVEStepParams{ReleaseID: "os-pve-1", Packages: []Package{{Name: "pve-manager", Version: "9.2.21", Origin: PVEOrigin}}})
ctx := signedjobs.WithSignedOp(context.Background(), &reconcile.SignedOp{Blob: []byte(`{"op":"os_pve_step"}`), Sig: []byte("SIG")})
if err := e.Execute(ctx, OpPVEStep, params); err != nil {
t.Fatal(err)
}
pp := w.plans[len(w.plans)-1]
sg, _ := pp["signed"].(map[string]any)
if pp["layer"] != LayerPVE || pp["lane"] != "slow" || pp["release_id"] != "os-pve-1" || sg == nil ||
sg["blob_b64"] != base64.StdEncoding.EncodeToString([]byte(`{"op":"os_pve_step"}`)) || sg["sig"] != "SIG" {
t.Fatalf("pve plan = %v", pp)
}
if !w.pveGateHeld || calls(w) != "pve:apply" || len(h.reports) != 1 || h.reports[0].Trigger != "signed" {
t.Fatalf("gate=%v calls=%s reports=%+v", w.pveGateHeld, calls(w), h.reports)
}
if err := e.Execute(context.Background(), OpPVEStep, params); err == nil {
t.Fatal("no envelope must refuse")
}
if err := e.Execute(context.Background(), OpDockerStep, params); err != signedjobs.ErrNoExecutor {
t.Fatalf("another op must pass through the chain: %v", err)
}
}
// os_pve_step is never benign.
func TestPVEStep_IsDestructiveClass(t *testing.T) {
if reconcile.Classify(reconcile.ClassOSPVEStep, reconcile.Provenance{}) != reconcile.Destructive {
t.Fatal("os_pve_step must be destructive-class (signed, operational key)")
}
}
+85
View File
@@ -0,0 +1,85 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// OpPVEStep is the signed op class of a Proxmox package step (R-812 option A, `11` §5.10): a ring-1 box takes an
// approved Proxmox set only through it. No undo in this release. CC may sign it until the first paying customer.
const OpPVEStep = "os_pve_step"
// PVEStepParams are the signed params. The wrapper compares Packages with the plan byte-for-byte.
type PVEStepParams struct {
ReleaseID string `json:"release_id"`
Packages []Package `json:"packages"`
VMID int `json:"vmid,omitempty"`
}
// PVEStepExecutor runs a verified os_pve_step (signedjobs.Executor) under the host-wide heavy-op gate (Gate) and the
// /etc/pve write gate (inside runPVE).
type PVEStepExecutor struct {
Leg *Leg
Guest func(ctx context.Context) (int, error)
Gate func(ctx context.Context) (release func(), err error)
}
// Execute implements signedjobs.Executor.
func (e PVEStepExecutor) Execute(ctx context.Context, op string, params json.RawMessage) error {
if op != OpPVEStep {
return signedjobs.ErrNoExecutor
}
so, ok := signedjobs.SignedOpFrom(ctx)
if !ok {
return fmt.Errorf("os_pve_step: no signed envelope in the context — the wrapper could not verify it")
}
var p PVEStepParams
if err := json.Unmarshal(params, &p); err != nil || len(p.Packages) == 0 {
return fmt.Errorf("os_pve_step: params must name the Proxmox set: %v", err)
}
vmid := p.VMID
if vmid == 0 {
if e.Guest == nil {
return fmt.Errorf("os_pve_step: no vmid and no guest finder")
}
v, err := e.Guest(ctx)
if err != nil {
return fmt.Errorf("os_pve_step: find the customer guest: %w", err)
}
vmid = v
}
if e.Gate != nil {
release, err := e.Gate(ctx)
if err != nil {
return fmt.Errorf("os_pve_step: heavy-op gate busy (a backup or restore-test runs): %w", err)
}
defer release()
}
rep := e.Leg.RunPVESigned(ctx, vmid, p, so.Blob, string(so.Sig))
switch rep.Outcome {
case "applied", "nothing":
if rep.Healthy {
return nil
}
}
return fmt.Errorf("os_pve_step: %s (%s) %s", rep.Outcome, rep.HealthReason, string(rep.Refused))
}
// RunPVESigned is one signed Proxmox step (ring 1): the pve layer with the signed envelope, which the wrapper verifies
// itself, holding the /etc/pve write gate.
func (l *Leg) RunPVESigned(ctx context.Context, vmid int, p PVEStepParams, blob []byte, sig string) Report {
unlock := l.lockPass(true)
defer unlock()
l.sendUnsentLocked(ctx) // R-868
runID := l.now().UTC().Format("20060102T150405Z")
rid := p.ReleaseID
if rid == "" {
rid = "signed-" + runID
}
return l.runPVE(ctx, runID, vmid, "signed", l.Block(), dockerOpts{releaseID: rid, packages: p.Packages,
signed: map[string]string{"blob_b64": base64.StdEncoding.EncodeToString(blob), "sig": sig}})
}
+218
View File
@@ -0,0 +1,218 @@
package osupdate
import (
"context"
"encoding/json"
"os"
"path/filepath"
"strings"
"syscall"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// ── R-868 (v0.144.0): a pass whose agent was killed still reports ─────────────────────────────────────
//
// MEASURED 2026-10-05 02:57 UTC on demo-hp (night drill A5): the debug pass and the agent daemon were kill -9-ed while
// apt-get ran. The root wrapper (its own process under sudo) finished all six packages, but the agent that would have
// read its stdout and posted the report was gone; the hub never learned what the pass installed.
//
// THE MECHANISM: the wrapper writes its apply report to <plan dir>/report-<run>-<layer>-apply.json BEFORE printing
// it (configs/felhom-os-apply save_report). The agent deletes that copy once the hub has the report (finish). A copy
// still on disk is a report nobody sent: SendUnsent posts it — at the agent's start, and before every pass — and then
// deletes it. A pass lock (flock on <plan dir>/pass.lock, released by the kernel when a process dies) keeps the
// sender from picking up the copy of a pass that is still running, also across the daemon and a selftest process.
// Pinned by TestR868_* (unsent_test.go).
// lockPass takes the pass lock. block=false returns ok=false at once when another pass holds it. A lock that cannot
// be opened at all (no plan dir yet) does not stop a pass: the unlock is then a no-op.
func (l *Leg) lockPass(block bool) (unlock func()) {
u, _ := l.tryLockPass(block)
return u
}
func (l *Leg) tryLockPass(block bool) (unlock func(), ok bool) {
dir := l.planDir()
_ = os.MkdirAll(dir, 0o700)
f, err := os.OpenFile(filepath.Join(dir, "pass.lock"), os.O_CREATE|os.O_RDWR, 0o600)
if err != nil {
l.log().Warn("osupdate: pass lock unavailable — continuing without it", "err", err)
return func() {}, true
}
how := syscall.LOCK_EX
if !block {
how |= syscall.LOCK_NB
}
if err := syscall.Flock(int(f.Fd()), how); err != nil {
f.Close()
return func() {}, false
}
return func() { _ = syscall.Flock(int(f.Fd()), syscall.LOCK_UN); f.Close() }, true
}
// SendUnsent posts every report a pass kept on disk and nobody sent (R-868). It skips when a pass runs now (that
// pass sends them first). Called at the agent's start.
func (l *Leg) SendUnsent(ctx context.Context) int {
unlock, ok := l.tryLockPass(false)
if !ok {
l.log().Info("osupdate: a pass is running — its start sends any kept report")
return 0
}
defer unlock()
return l.sendUnsentLocked(ctx)
}
// SendUnsentLoop runs SendUnsent now and then every `every` until ctx ends (v0.144.1). MEASURED live on demo-hp
// 2026-10-05: after a kill -9 the daemon restarted in ~5 s, while the orphaned root wrapper was still installing — its
// copy appeared ~7 s AFTER the start-time sender had looked. One look at start is therefore not enough. A glob of the
// plan dir every few minutes costs nothing; the pass lock keeps it off a running pass. Pinned by
// TestR868_ACopyWrittenAfterTheStartIsSentByTheLoop.
func (l *Leg) SendUnsentLoop(ctx context.Context, every time.Duration, onSent func(int)) {
for {
if n := l.SendUnsent(ctx); n > 0 && onSent != nil {
onSent(n)
}
select {
case <-ctx.Done():
return
case <-time.After(every):
}
}
}
func (l *Leg) sendUnsentLocked(ctx context.Context) int {
if l.Hub == nil {
return 0 // nobody to send to: keep the copies for a process that has the hub
}
files, _ := filepath.Glob(filepath.Join(l.planDir(), "report-*.json"))
sent := 0
for _, f := range files {
b, err := os.ReadFile(f)
if err != nil {
l.log().Warn("osupdate: a kept report cannot be read", "path", f, "err", err)
continue
}
var wr WrapperReport
if err := json.Unmarshal(b, &wr); err != nil || wr.Layer == "" {
l.log().Warn("osupdate: a kept report is not a report — moved aside", "path", f, "err", err)
_ = os.Rename(f, f+".bad")
continue
}
rep := l.reportFromKept(ctx, wr, f)
lg := l.log().With("run", rep.RunID, "layer", rep.Layer, "vmid", rep.VMID, "ring", rep.Ring, "trigger", rep.Trigger)
lg.Info("osupdate: sending a kept report late (the agent was stopped mid-pass or the hub was away, R-868)", "path", f)
before := rep.unsent
_ = l.finish(ctx, lg, rep)
if _, err := os.Stat(before); os.IsNotExist(err) {
sent++
// the pass's plan file is left behind too when the agent was killed inside call()
_ = os.Remove(filepath.Join(filepath.Dir(f), "plan-"+strings.TrimPrefix(filepath.Base(f), "report-")))
}
}
return sent
}
// reportFromKept builds the hub report from a kept wrapper report, as runLayer would have. Health: the copy's own
// before/after reading; when that is not healthy after an install, one fresh reading (services restart after an
// install, and the pass that would have waited for them is gone).
func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string) Report {
ring := 1
if wr.Ring != nil {
ring = *wr.Ring
}
runID, trigger := wr.RunID, wr.Trigger
if runID == "" {
runID = strings.TrimSuffix(strings.TrimPrefix(filepath.Base(path), "report-"), ".json")
}
if trigger == "" {
trigger = "unknown"
}
rep := Report{RunID: runID, Layer: wr.Layer, Trigger: trigger, Mode: wr.Mode, Ring: ring, ReleaseID: wr.ReleaseID,
VMID: wr.VMID, unsent: path}
prefix := "sent late — kept on the box until the hub could take it (R-868, R-875)"
switch {
case wr.refused():
rep.Outcome, rep.Refused, rep.HealthReason = "refused", wr.Refused, prefix
return rep
case wr.failed():
rep.Outcome, rep.Refused = "failed", wr.Failed
case len(wr.Upgraded) == 0:
rep.Outcome = "nothing"
default:
rep.Outcome = "applied"
}
if wr.Layer == LayerKernel {
// a kernel STAGE changes nothing the box runs (the new kernel only boots once, at the night's reboot), so its
// kept copy needs no fresh health reading (R-836)
rep.Kernel = rawOrNil(wr.Kernel)
rep.Upgraded, rep.PassSeconds, rep.Authority = wr.Upgraded, wr.PassSeconds, wr.Authority
if rep.Outcome == "applied" {
rep.Outcome = "staged"
}
rep.Healthy, rep.HealthReason = rep.Outcome == "staged" || rep.Outcome == "nothing", prefix
return rep
}
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
rep.OOMCheck = rawOrNil(wr.OOMCheck)
rep.PVEManager = wr.PVEManager
wantEngine, wantPVE := "", ""
for _, u := range wr.Upgraded {
if u.Name == "docker-ce" {
wantEngine = EngineOf(u.Version)
}
if u.Name == "pve-manager" {
wantPVE = u.Version
}
}
verdict := func(h *Health) (bool, string) {
switch wr.Layer {
case LayerDocker:
return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine)
case LayerHost, LayerPVE:
t := hub.TunnelUnknown
if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx)
}
if wr.Layer == LayerPVE {
return PVEHealthVerdict(wr.HealthBefore, h, t, wantPVE, wr.PVEManager)
}
return HostHealthVerdict(wr.HealthBefore, h, t)
}
return HealthVerdict(wr.HealthBefore, h)
}
ok, why := verdict(wr.HealthAfter)
if !ok && len(wr.Upgraded) > 0 && wr.VMID > 0 {
lane := "fast"
if wr.Layer == LayerDocker || wr.Layer == LayerPVE {
lane = "slow"
}
if hr, err := l.call(ctx, "kept-"+runID, map[string]any{"release_id": "kept", "layer": wr.Layer, "lane": lane,
"vmid": wr.VMID, "mode": "health", "packages": []Package{}}); err == nil && hr.Health != nil {
ok, why = verdict(hr.Health)
}
}
rep.Healthy, rep.HealthReason = ok, prefix
if why != "" {
rep.HealthReason = prefix + ": " + why
}
if !ok && rep.Outcome == "applied" {
rep.Outcome = "health_failed"
}
rep.Installed, rep.Pending = wr.Installed, wr.Pending
rep.RestartNeeded, rep.DockerRestartNeeded, rep.RebootNeeded = wr.RestartNeeded, wr.DockerRestartNeeded, wr.RebootNeeded
rep.RebootScanned = wr.RebootScanned
if wr.Layer == LayerDocker {
rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending)
} else if wr.Layer == LayerPVE {
rep.Installed, rep.Pending = onlyPVE(wr.Installed), onlyPVEPending(wr.Pending)
} else {
planned := map[string]bool{}
for _, u := range wr.Upgraded {
planned[u.Name] = true
}
rep.NotCovered = notCovered(wr.Pending, ring, planned)
}
return rep
}
+173
View File
@@ -0,0 +1,173 @@
package osupdate
import (
"context"
"encoding/json"
"errors"
"os"
"path/filepath"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-868 (v0.144.0). THE NIGHT'S SHAPE (A5, demo-hp 2026-10-05 02:57 UTC): the wrapper finished six packages, the agent
// was kill -9-ed before it read the report; the hub never got it. Here: the wrapper's kept copy and the plan file are
// on disk, a NEW agent process starts — it must send exactly one `applied` report and delete both files.
// COMPANION RED-PROOF: drop the SendUnsent body (return 0) → "the hub got no report".
func TestR868_KilledPassIsReportedAtStart(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
ring := 0
kept := WrapperReport{Mode: "apply", Layer: LayerGuest, RunID: "20261005T025700Z", Trigger: "debug", Ring: &ring, VMID: 9201,
ReleaseID: "ring0-20261005T025700Z", HealthBefore: guestOK(), HealthAfter: guestOK(),
Upgraded: []Package{{Name: "libc6", Version: "u4"}, {Name: "openssl", Version: "u3"}}}
b, _ := json.Marshal(kept)
rp := reportFile(l.PlanDir, kept.RunID, LayerGuest, "apply")
pp := planFile(l.PlanDir, kept.RunID, LayerGuest, "apply")
must(t, os.WriteFile(rp, b, 0o600))
must(t, os.WriteFile(pp, []byte("{}"), 0o600))
if n := l.SendUnsent(context.Background()); n != 1 {
t.Fatalf("sent %d, want 1", n)
}
if len(h.reports) != 1 {
t.Fatalf("the hub got no report (or several): %+v", h.reports)
}
r := h.reports[0]
if r.Outcome != "applied" || !r.Healthy || r.Trigger != "debug" || r.RunID != kept.RunID || r.Ring != 0 || len(r.Upgraded) != 2 || r.VMID != 9201 {
t.Fatalf("report = %+v", r)
}
// R-875 (v0.145.0): neutral — the copy cannot tell a killed agent from an absent hub.
if !strings.HasPrefix(r.HealthReason, "sent late") || strings.Contains(r.HealthReason, "stopped mid-pass") {
t.Fatalf("health_reason = %q, want the neutral \"sent late …\"", r.HealthReason)
}
for _, p := range []string{rp, pp} {
if _, err := os.Stat(p); !os.IsNotExist(err) {
t.Fatalf("%s still on disk after the hub got it", filepath.Base(p))
}
}
// a second start sends nothing again — no duplicate report
if n := l.SendUnsent(context.Background()); n != 0 || len(h.reports) != 1 {
t.Fatalf("sent again: %d, reports %d", n, len(h.reports))
}
}
// A normal pass: the hub gets ONE report per layer and the kept copies are gone after it (nothing resent later).
// COMPANION RED-PROOF: drop the os.Remove(rep.unsent) in finish → the next pass resends → "2 guest reports".
func TestR868_NormalPassLeavesNoCopyAndNoDuplicate(t *testing.T) {
w := &fakeWrapper{t: t, keep: true, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6", Version: "u4"}}}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Run(context.Background(), 9201, "debug")
left, _ := filepath.Glob(filepath.Join(l.PlanDir, "report-*.json"))
if len(left) != 0 {
t.Fatalf("kept copies left after the hub got them: %v", left)
}
l.Run(context.Background(), 9201, "debug") // the next pass sends kept copies first
guest := 0
for _, r := range h.reports {
if r.Layer == LayerGuest && r.Outcome == "applied" {
guest++
}
}
if guest != 2 {
t.Fatalf("%d guest reports for 2 passes (a duplicate or a loss)", guest)
}
}
type failingHub struct{ n int }
func (h *failingHub) PostOSReport(context.Context, []byte) error {
h.n++
return errors.New("hub away")
}
// The hub away: the copy stays, and goes at the next chance.
func TestR868_HubAwayKeepsTheCopy(t *testing.T) {
w := &fakeWrapper{t: t, keep: true, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6", Version: "u4"}}}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Appliance = false
l.Hub = &failingHub{}
l.Run(context.Background(), 9201, "night")
left, _ := filepath.Glob(filepath.Join(l.PlanDir, "report-*.json"))
if len(left) != 2 { // ring 0: the guest step and the Docker step each kept one
t.Fatalf("the copies must stay while the hub is away: %v", left)
}
h := &fakeHub{}
l.Hub = h
if n := l.SendUnsent(context.Background()); n != 2 {
t.Fatalf("sent %d: %+v", n, h.reports)
}
for _, r := range h.reports {
if r.Trigger != "night" || r.Ring != 0 {
t.Fatalf("the kept report lost its ids: %+v", r)
}
}
}
// A pass in progress holds the lock: the sender at start must not take that pass's copy (it would be sent twice).
func TestR868_RunningPassKeepsTheSenderOff(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, nil)
must(t, os.WriteFile(reportFile(l.PlanDir, "r1", LayerGuest, "apply"), []byte(`{"mode":"apply","layer":"guest"}`), 0o600))
unlock := l.lockPass(true)
if n := l.SendUnsent(context.Background()); n != 0 || len(h.reports) != 0 {
t.Fatalf("sent while a pass ran: %d", n)
}
unlock()
if n := l.SendUnsent(context.Background()); n != 1 {
t.Fatalf("not sent after the pass: %d", n)
}
}
// v0.144.1 — THE LIVE SHAPE (demo-hp 2026-10-05 05:45 UTC): the restarted daemon looked at 05:45:09, the orphaned
// wrapper wrote its copy at ~05:45:16. The loop must still send it.
// COMPANION RED-PROOF: make SendUnsentLoop return after the first look → "the late copy was never sent".
func TestR868_ACopyWrittenAfterTheStartIsSentByTheLoop(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, nil)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
sent := make(chan int, 4)
go l.SendUnsentLoop(ctx, 20*time.Millisecond, func(n int) { sent <- n })
time.Sleep(50 * time.Millisecond) // the start-time look found nothing
must(t, os.WriteFile(reportFile(l.PlanDir, "late", LayerGuest, "apply"),
[]byte(`{"mode":"apply","layer":"guest","run_id":"late","trigger":"debug","ring":0,"vmid":9201,"upgraded":[{"name":"openssl","version":"u3"}]}`), 0o600))
select {
case n := <-sent:
if n != 1 || len(h.reports) != 1 || h.reports[0].RunID != "late" || h.reports[0].Outcome == "" {
t.Fatalf("sent %d: %+v", n, h.reports)
}
case <-time.After(3 * time.Second):
t.Fatal("the late copy was never sent")
}
}
func must(t *testing.T, err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
// R-528: a kept docker report (the agent was killed) still carries the oom_check object to the hub unchanged.
// COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in reportFromKept → "oom_check lost".
func TestR868_KeptCopyCarriesTheOOMCheck(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
ring := 0
kept := WrapperReport{Mode: "apply", Layer: LayerDocker, RunID: "20261007T020000Z", Trigger: "night", Ring: &ring, VMID: 9201,
ReleaseID: "ring0-20261007T020000Z", HealthBefore: guestOK(), HealthAfter: guestOK(), DockerEngine: "29.8.2",
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}}, OOMCheck: json.RawMessage(oomCheckWire)}
b, _ := json.Marshal(kept)
must(t, os.WriteFile(reportFile(l.PlanDir, kept.RunID, LayerDocker, "apply"), b, 0o600))
if n := l.SendUnsent(context.Background()); n != 1 || len(h.bodies) != 1 {
t.Fatalf("sent %d, bodies %d", n, len(h.bodies))
}
var m map[string]json.RawMessage
must(t, json.Unmarshal(h.bodies[0], &m))
if string(m["oom_check"]) != oomCheckWire {
t.Fatalf("oom_check lost or changed: %q", m["oom_check"])
}
}
+2 -2
View File
@@ -60,8 +60,8 @@ func TestLiveReporter_CoordPresentWithoutPriorVerify(t *testing.T) {
if h.PBS == nil {
t.Fatal("pbs coord absent despite a reachable PBS — the gap this fixes")
}
if h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != "root" || h.PBS.LatestSnapshotID != "9201" {
t.Errorf("pbs coord = %+v, want felhom-pbs/root/9201", h.PBS)
if h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != hub.PBSRootNamespace || h.PBS.LatestSnapshotID != "9201" {
t.Errorf("pbs coord = %+v, want felhom-pbs, the root namespace (\"\", R-124), 9201", h.PBS)
}
// COMPANION (pre-fix): the bare SnapshotStore (no live read) with an empty store omits pbs.
+34
View File
@@ -0,0 +1,34 @@
// Package privapplytest runs configs/felhom-priv-apply in CHECK-ONLY mode from Go tests (R-861): each renderer's
// real output must be accepted by the root checker, so the two can never drift apart unnoticed. Test-only helper.
package privapplytest
import (
"os"
"os/exec"
"path/filepath"
"runtime"
"strings"
"testing"
)
// Check writes content to a temp file and runs `felhom-priv-apply --check <verb> [name] <file>`. It returns the
// checker's verdict line ("OK" or "REFUSED [rule] …"). Skips when python3 is absent.
func Check(t *testing.T, verb, name, content string) string {
t.Helper()
py, err := exec.LookPath("python3")
if err != nil {
t.Skip("python3 not available")
}
_, here, _, _ := runtime.Caller(0)
wrapper := filepath.Join(filepath.Dir(here), "..", "..", "configs", "felhom-priv-apply")
f := filepath.Join(t.TempDir(), "staged")
if err := os.WriteFile(f, []byte(content), 0o600); err != nil {
t.Fatal(err)
}
args := []string{"-B", wrapper, "--check", verb}
if name != "" {
args = append(args, name)
}
out, _ := exec.Command(py, append(args, f)...).CombinedOutput()
return strings.TrimSpace(string(out))
}
+3 -2
View File
@@ -186,8 +186,9 @@ func (b *BackHalf) Provision(ctx context.Context, in Input) (Result, error) {
// 6. Install + register the pre-start self-heal hook (C1 net): if a data drive is absent at a future
// boot, the hook creates a placeholder for its missing bind source so the guest still starts.
// Best-effort + non-fatal — it's defense-in-depth; a provision must not fail over the hook.
if err := guesthook.InstallSnippet(ctx, b.runner); err != nil {
b.logger.Warn("provision: pre-start hook snippet install failed (non-fatal)", "vmid", in.VMID, "err", err)
// R-861 (v0.146.0): the hook FILE comes with the signed config bundle; the agent only checks it and registers it.
if err := guesthook.SnippetReady(guesthook.SnippetPath); err != nil {
b.logger.Warn("provision: pre-start hook not in place — not registering it (non-fatal)", "vmid", in.VMID, "err", err)
} else if err := guesthook.Register(ctx, b.runner, in.VMID); err != nil {
b.logger.Warn("provision: pre-start hook registration failed (non-fatal)", "vmid", in.VMID, "err", err)
}
+10
View File
@@ -5,6 +5,7 @@ import (
"context"
"encoding/json"
"fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
"io"
"net/http"
"net/url"
@@ -105,6 +106,15 @@ func (c *Client) do(ctx context.Context, method, path string, body io.Reader, ou
// doBody is the single HTTP chokepoint: builds the request, sets auth, executes,
// maps non-2xx to APIError, and decodes the data envelope.
func (c *Client) doBody(ctx context.Context, method, path string, body io.Reader, contentType string, out any) error {
// R-812 option A: every non-GET call may write /etc/pve — it waits while a Proxmox package step restarts pmxcfs
// (pvegate). Pinned by TestPVEGate_ClientWriteWaitsGetDoesNot.
if method != http.MethodGet {
release, _, gerr := pvegate.Write(ctx)
if gerr != nil {
return fmt.Errorf("proxmox: %s %s held back by a Proxmox package step: %w", method, path, gerr)
}
defer release()
}
req, err := http.NewRequestWithContext(ctx, method, c.base+path, body)
if err != nil {
return fmt.Errorf("proxmox: building request: %w", err)

Some files were not shown because too many files have changed in this diff Show More