Compare commits

...

20 Commits

Author SHA1 Message Date
admin f277e619e2 CHANGELOG: unreleased — R-856 GET /host/crash-guard
gates / gates (push) Successful in 38s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:40 +02:00
admin 386f51edc6 R-856: GET /host/crash-guard — the host crash guard's last-boot record for the controller (09 decision 143)
The controller waits ~15 minutes with app mails after a crash boot of the host; it learns of the
crash boot from this route. Reads /var/lib/felhom-crash-guard/state.json (read-only, no Proxmox call)
and passes present/last_boot_at/last_boot_unclean/tripped through; a missing, unreadable or garbled
file answers 200 present:false. Guest-token authed like every sibling route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:01 +02:00
admin 130e3ed882 CHANGELOG: unreleased — R-444 weekly guest disk trim, R-99 runbook pointer
gates / gates (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:25:22 +02:00
admin be398f92e8 R-99: the PBS phantom WARN names the cleanup runbook (09 §3 decision 140)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin ee71abd1d4 R-444: weekly guest disk trim (pct fstrim) outside the night, under the heavy-op gate
Operator ruling 09 §3 decision 139. One exact sudoers rule FELHOM_FSTRIM
(`/usr/sbin/pct ^fstrim [0-9]+$`) + manifest entry guest-fstrim; new
internal/fstrim job: due Wednesday from 10:00 host-local, starts only
10:00-20:59, holds backup.InFlight (busy -> deferred to the next hourly
tick), failed trim retried at most 3x per week, bytes parsed from the
"(N bytes) trimmed" lines, last result per guest persisted in
<state_dir>/guest-disk-trim.json and reported as guest_disk_trim.
Config opt-out: "disk_trim": {"disable": true}.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin 37e98f452b REPORT: the burn-down night (2026-10-06)
gates / gates (push) Successful in 45s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 02:13:10 +02:00
admin b2b82ae828 bundle test: the ISO first-boot files the installer names under KEPT are not installer-written (go test red since felhom.eu 85de3f9b); CHANGELOG unreleased (R-426 decoys)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:25 +02:00
admin b78a0ff3ac R-426: decoys for the shared reuse-refs, instructions and observations gates
COVERS "reuse-refs", "instructions", "observations": the three shared
felhom.eu scripts run against a scratch clone of THIS repo (in a scratch
workspace symlinking the sibling clones they reach across to), so the
plant is in the agent's own REUSE.md / CLAUDE.md / REPORT.md. Convicted:
a missing cited .go and .md path, a version literal in CLAUDE.md's
effective text, R-419's prose-only Observations note. Passed: the real
files, the version inside an HTML comment, both genuine markers.
DECOY_SHARED_DIR lets a red-proof judge a mutated copy of the shared
scripts without editing the felhom.eu clone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 96047453cb R-426: decoys for the release-complete gate
COVERS "release-complete": the working-tree gate runs in a scratch clone
whose origin is a scratch bare repo, against the fake Gitea. Convicted:
the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated
commit, a tag with no package, and no-tag wins over a registry 500.
Inconclusive: a registry 500. Passed: the genuine release, an
`## Unreleased` heading above it, a newer version named only in prose or
under `###` (the withdrawn sweep decoy, now asserted the right way round),
a tag only origin has (the shallow-CI shape), and a LOCAL-only tag BY
DESIGN (CI's fresh clone and the published gate's converse probe see it).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 4bf5db6875 R-426: decoy suite for the published gate, against a fake Gitea
scripts/test_gate_decoys.py (new; COVERS "published"): an http.server on
127.0.0.1 stands in for Gitea through the gate's existing GITEA_BASE
seam, proxies stripped, so no case reaches the real registry. Facts
convicted: a tag whose package 404s, a tag tree without the configs, a
package one patch past the newest tag (never tagged), a patch-gap orphan,
a missing package that lexical sorting would drop out of the retention
window. Inconclusive, never a pass: tags api 500, a non-JSON 200, Gitea
unreachable. Passed: a clean registry, a non-semver tag, a version older
than the retention window (BY DESIGN).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 87977ff40a CHANGELOG/v0.148.0 released (shas)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 00:37:07 +02:00
admin 861d32a4b4 CHANGELOG: unreleased — R-349 agent_sha256, R-25 agent half (burn-down night)
gates / gates (push) Successful in 43s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:51 +02:00
admin 769c4c3cf2 R-25 (agent half): the format answer carries the NEW filesystem's UUID, bound to the durable id
After mkfs the agent re-resolves the bound durable id, requires it to name the
device it just formatted, reads the superblock back (blkid -p, requested fstype)
and returns fs_uuid in POST /disks/format and GET /disks/format/status. Anything
unverified returns "" — never a path-resolved guess. The controller half
(mount fs_uuid instead of re-resolving the UUID from the /dev path) is owed
in felhom-controller.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:11 +02:00
admin 64f704d0f7 R-349: the host report carries the sha256 of the RUNNING agent binary
A hand-built proof binary and the published artifact share a version string
but not their bytes, so no version check could see the divergence. The agent
now reports agent_sha256 (hash of /proc/self/exe, once per process; empty =
unknown) beside agent_version, the same mechanism as host.wrapper_sha256.
The hub half (compare against the vouched agent_sha256, surface drift) is
owed in the hub repo.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:11 +02:00
admin 208fac8027 CHANGELOG/REPORT: v0.147.0 released (shas)
gates / gates (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 18:58:37 +02:00
admin f1b9b41214 R-124 recipe root namespace as PBS spells it; R-118 no root size for an absent drive; R-269 rotated-out token rejected at once; R-317 dnsmasq install probed by its unit (burn-down round 2)
gates / gates (push) Successful in 47s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 18:56:13 +02:00
admin d83316326e R-291 retention record source, R-348 restart comment (no binary change; burn-down)
gates / gates (push) Successful in 22s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 16:56:18 +02:00
admin e06ed97fa8 agent v0.146.1 REPORT
gates / gates (push) Successful in 22s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 13:00:22 +02:00
admin e4b5cf9693 R-880: build-step-bundle.py — the transition bundle for a release whose bundle adds paths (an installed felhom-os-apply refuses unknown paths, R16)
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:23:55 +02:00
admin faa3cad92e agent v0.146.1 CHANGELOG (released)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:14:07 +02:00
39 changed files with 2391 additions and 72 deletions
+91
View File
@@ -1,3 +1,94 @@
## unreleased
- R-856 (`09` §3 decision 143): new local-API route `GET /host/crash-guard` — passes the host crash guard's last-boot record (present, last_boot_at, last_boot_unclean, tripped) from /var/lib/felhom-crash-guard/state.json to the controller, which waits ~15 min with app mails after a crash boot. Read-only, no Proxmox call, guest-token authed; a missing/unreadable/garbled file answers 200 present:false (never an error page). An older agent answers 404, which the controller reads as unknown (normal 90 s grace) — no controller MinAgent raise needed.
- R-444 (`09` §3 decision 139): weekly guest disk trim. New sudoers alias FELHOM_FSTRIM with ONE exact rule `/usr/sbin/pct ^fstrim [0-9]+$` (rides the signed config bundle; decoys pinned by TestSudoersFstrimRuleIsExact) and capability guest-fstrim (non-critical). New internal/fstrim job: each owned RUNNING guest gets `pct fstrim <vmid>` once a week - due Wednesday from 10:00 host-local, starts only 10:00-20:59 (never the 01:00-06:59 night), holds the one-heavy-op gate so it never runs beside a backup or restore-test (busy -> deferred to the next hourly tick; a box that was off catches up at its next daytime hour); a failed trim WARNs and is retried at most 3 times that week; bytes parsed from `pct fstrim`'s "(N bytes) trimmed" lines; positive log `fstrim: guest N trimmed X GiB in Ys`; last result per guest persisted in <state_dir>/guest-disk-trim.json and reported as the new omitempty host-report stanza `guest_disk_trim`. Opt-out: agent.json "disk_trim": {"disable": true}.
- R-99 (`09` §3 decision 140): the agent's WARN for a PBS archive below the 1 MiB plausibility floor now ends with the pointer to the sanctioned cleanup (`documentation/runbooks/pbs-phantom-cleanup.md`); detection only — nothing is deleted automatically. Dir-storage archives keep the old text.
## unreleased
- R-426: scripts/test_gate_decoys.py (new) — the published gate judged against a fake Gitea (127.0.0.1, via GITEA_BASE; never the real registry): 11 cases; COVERS published — `felhom-agent/published` leaves the decoy-coverage EXEMPT list.
- R-426: release-complete gate — 10 decoy cases (scratch clone + scratch bare origin + fake Gitea); COVERS release-complete — `felhom-agent/release-complete` leaves the decoy-coverage EXEMPT list.
- R-426: the shared reuse-refs/instructions/observations gates get agent-side decoys (9 cases on a scratch clone of this repo); COVERS reuse-refs, instructions, observations — three `felhom-agent/*` entries leave the decoy-coverage EXEMPT list.
- **Fixed without a row:** `configs/test_felhom_config_bundle.py` read the two ISO first-boot files that installer 1.32.0 now NAMES under KEPT (R-275) as files the installer writes — `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 (felhom.eu `85de3f9b`); they are listed with why. Test only; agent v0.148.0's code is unaffected.
## v0.148.0 — the host report names the running binary's sha; the format answer carries the new filesystem's UUID (burn-down night: R-349, R-25 agent halves) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `3e68a0870e0e2ce262cb4819294edddeb0a73e8c558a31611a20329a6d9ee283`
config bundle sha256 `a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de` (tag `v0.148.0` = `861d32a`).
Delivery order as for v0.147.0: signed `agent_update`, then signed `agent_config_update`.
- **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from
`/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub
comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved.
- **R-25 (agent half):** `POST /disks/format` and `GET /disks/format/status` return `fs_uuid`, the new filesystem's UUID
read back after mkfs only when the bound durable id still resolves to the formatted device and the superblock is the
requested type (empty = not verified); `DeviceProbe` gains `FSUUID` from blkid. Tests `TestFormat_*FSUUID*`; three
red-proofs. The controller half (mount that UUID) is a controller change.
## v0.147.0 — the recovery recipe spells the root namespace the way PBS does; a removed drive no longer shows the root disk's size; a rotated-out token stops at once; the dnsmasq check looks at the right package (burn-down round 2: R-124, R-118, R-269, R-317) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `642c4d196c48671c14ff653118303c5903af1670b7abb550edeaf73e701cd5b8`,
config bundle sha256 `326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d` (tag `v0.147.0` = `f1b9b41`).
Delivery order as for v0.146.1: signed `agent_update`, then signed `agent_config_update`.
MinAgent impact: none (the controller needs nothing new from this agent). Config bundle content unchanged from v0.146.1.
- **R-124 (operator ruling 2026-10-05: fix it):** the DR recipe's `pbs.namespace` for a box in PBS's ROOT namespace is
now `""` — PBS's own spelling — beside `namespace_state: resolved`; it used to be the word `root`, which no namespace
is named, so `--ns root` failed in a recovery. `hub.PBSRootNamespace`; `TestR124_RootNamespaceOnTheWireIsPBSSpelling`
(red-proof: back to "root" → FAIL). Runbook: `felhom.eu runbooks/ep0-datastore-copy.md` step 2 says how to read it
(and to treat a recorded `root` from older agents as empty). The hub stores the recipe raw; its fixture follows.
- **R-118:** the local API's drive list reads a drive's capacity only while its DEVICE is present — with the device gone
the bare mountpoint is a directory on the root filesystem, whose size was reported as the drive's.
`TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity` (red-proof convicts).
- **R-269:** the token store re-reads its shared file whenever it has grown, BEFORE answering — so a token rotated out by
another process stops authorizing on its next use (it used to keep working until an unrelated miss). One `stat` per
call. `TestTokenStore_RotatedOutTokenRejectedFirst` (red-proof convicts).
- **R-317:** the LAN resolver decides whether to install `dnsmasq` by its service UNIT, not by `/usr/sbin/dnsmasq` (which
the `dnsmasq-base` package also ships). Same `apt-get install` command; no sudoers change. `TestEnsureDnsmasq_*`
(red-proof convicts). Red-proofs: `felhom.eu/documentation/audits/burndown2-2026-10-05/agent-red-proofs.txt`,
`r124-red-proof.txt`.
Also in this release (no binary effect; from burn-down round 1):
- **R-291:** `scripts/retention-policy.json` names where its 10 comes from — the R-267 newest-10 prune of generic
packages, established 2026-08-10 (R-287) — instead of „observed, no located ruling"; the non-existent
`registry-retention.md` reader is dropped. `check-published-versions.py` still reads 10 (checked).
- **R-348:** `internal/backup/store.go` no longer says backups are „unaffected" by a restart: the reported backup list
reads 0 until the next backup runs; only the hub's verdict (7-day look-back) is unaffected.
## v0.146.1 — R-861 review fixes: the signed update flips a root-owned copy; no Wants=/continuations in mount units; the escrow read follows no symlink anywhere (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `badd6c9a2e40c8bfe856d2d1a203443b21b7eb92ecc35d6090ab44518d4d082a`,
config bundle sha256 `42333e969028867ad8142335e6c1bc4040eec231de0d8d330c2d4b2cf7bc3442`. **Supersedes v0.146.0, which
was released but never vouched or delivered to any box.** The same order applies: signed `agent_update` first, then the
signed `agent_config_update`.
**Delivery needs a STEP bundle (R-880, found while delivering).** An installed `felhom-os-apply` checks an incoming
bundle's paths against its OWN table (R16), so every box on the v0.145.0 bundle REFUSES the v0.146.1 bundle (it adds 4
paths). `scripts/build-step-bundle.py` builds the transition: the box's current bundle with ONLY `felhom-os-apply`
replaced (same paths — the old wrapper accepts it), published as bundle version `0.146.1-step1`; then the release's own
bundle. Order on a box: `agent_update` 0.146.1 → `agent_config_update` 0.146.1-step1 → `agent_config_update` 0.146.1.
Tests `StepBundle` (the R16 refusal reproduced; the step accepted; exactly one file changed). Tooling only — not in the
binary or the bundle.
A background security review of the v0.146.0 commit found three holes in the new code; each is fixed and red-proved
(`felhom.eu/documentation/audits/hub-safety-2026-10-05/partF/red-proof.txt`, S1–S3):
- **S1 — a race in the signed update.** `felhom-os-apply` hashed the agent's staged file and then let the A/B wrapper copy
it BY PATH; the agent owns that directory and could swap the file in between. Now the root step reads the file ONCE
(`read_staged_once`: O_NOFOLLOW, fstat, owner, size), hashes those bytes, writes them to a root-owned directory
(`/var/lib/felhom-os-apply/agent-update/`) and hands ONLY that copy to `felhom-selfupdate-guarded apply`, which now
refuses any other directory, a symlink, or a file not owned by root. Tests: `AgentUpdate` (+1),
`SelfupdateWrapperConfinement`.
- **S2 — an allowlist escape in `felhom-priv-apply`.** `[Unit]` accepted `Wants=`/`Requires=`/`Before=` naming any unit, so
a mount unit could start e.g. `reboot.target`. `[Unit]` now holds only `Description` and `After=local-fs-pre.target`
(what the renderers write), and any line ending in a backslash (a systemd continuation this parser would read
differently) is refused. Tests `test_U2_wants_starts_another_unit`, `test_U2_continuation_line`.
- **S3 — a path traversal in the escrow read.** `O_NOFOLLOW` guards only the last component; a symlinked DIRECTORY in the
agent's own state dir still redirected the root read. `readStagedNoFollow` now walks the path from `/` with
`openat(O_NOFOLLOW)` per component. Test `TestAttach_RefusesASymlinkedDirectory`.
## v0.146.0 — the agent's root grants narrowed: exact sudo patterns, a root content checker, fixed files from the bundle, the signed update checked as root (R-861) (2026-10-05) ## v0.146.0 — the agent's root grants narrowed: exact sudo patterns, a root content checker, fixed files from the bundle, the signed update checked as root (R-861) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `b860af465076041e07f35fed1b12d64ae2b2985d8995f0ce167418d39c2b00d5`, Released by `scripts/release-agent.sh`: binary sha256 `b860af465076041e07f35fed1b12d64ae2b2985d8995f0ce167418d39c2b00d5`,
+13 -11
View File
@@ -1,14 +1,16 @@
# REPORT — agent v0.145.0 (2026-10-05, afternoon): the OS update repairs itself after a power cut # REPORT — agent v0.148.0 (2026-10-06, the burn-down night)
Brief: the 2026-10-05 catch-up brief (operator), Parts C (R-874, R-875) and D (R-876). Full session report: Full session report: `felhom.eu/REPORT-burndown3-2026-10-06.md`. Baseline `208fac8` (v0.147.0).
`felhom.eu/REPORT-catchup-2026-10-05.md`. Architecture: `11-os-updates.md` §5.4.1, §8.4–8.5; `09` decisions 117–118.
| Row | Fix | Proof | **Released:** v0.148.0 (tag = `861d32a`; binary sha256 `3e68a087…`, bundle `a6fa4f58…`, verified by download). R-349
|---|---|---| (the host report carries `agent_sha256` — the hub's host pages read „matches vouched” for all three boxes) and R-25's
| R-876 | the wrapper reads `dpkg --audit` AND the update journal in ONE call, repairs on either; belt: repair + retry once when apt says "interrupted" | 3 tests + 3 red-proofs; **live (operator's go): crash mid-unpack on demo-hp → the next pass `REPAIR … journal=1` → `DONE rc=0 upgraded=12`, nobody touched the box** | agent half (the format answer carries the verified `fs_uuid`). **Delivered** by signed `agent_update` (all three on
| R-874 | the restore-test's first due-check 30 min after start | 2 tests + red-proof; live on demo-felhom: passed restore-test at start + 30 min | 0.148.0 by 22:50Z) and `agent_config_update` (root files 0.148.0 by 22:58Z) to demo-hp, demo-felhom, Tester 1. Tester 2:
| R-875 | a kept report's reason is neutral ("sent late …") | test + red-proof | nothing sent (off).
Released by `scripts/release-agent.sh`: v0.145.0 (`894da35c…`, bundle `78c00adc…`), verified by download; signed **On main, unreleased:** R-426 decoys (new `scripts/test_gate_decoys.py`: published against a fake Gitea,
`agent_update` + `agent_config_update` to demo-hp, demo-felhom, tester-1 (71/71 after each bundle); vouched with release-complete, the shared reuse-refs/instructions/observations).
golden 0.295.0, min_agent 0.131.0. `go test ./...` rc 0, Python suites OK, `agent_gates.py` OK.
**Said plainly:** `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 — `configs/test_felhom_config_bundle.py`
read the installer 1.32.0 KEPT names as installer-written files. The v0.148.0 binary was released inside that window;
its code is unaffected (a test-only sibling coupling; CI has no Go). Fixed in `b2b82ae`.
+5 -3
View File
@@ -66,8 +66,8 @@
|---|---|---|---|---| |---|---|---|---|---|
| `IntentStore` (`Get/SetEnrolled/SetEjected/SetDecommissioned/OnAbsent`) | internal/storage/intent.go | `OpenIntentStore(path)` | drive intent (4-state self-heal) | Keyed by durable-id only; `OnAbsent` is the ONLY ejected→enrolled path; refuses empty ids | | `IntentStore` (`Get/SetEnrolled/SetEjected/SetDecommissioned/OnAbsent`) | internal/storage/intent.go | `OpenIntentStore(path)` | drive intent (4-state self-heal) | Keyed by durable-id only; `OnAbsent` is the ONLY ejected→enrolled path; refuses empty ids |
| `GuestBindStore` (`Record/Remove/Guests`) | internal/localapi/guestbindstore.go | `OpenGuestBindStore(path)` | per-guest enrolled binds (F9 re-assert) | Same tmp+rename 0600 pattern as IntentStore | | `GuestBindStore` (`Record/Remove/Guests`) | internal/localapi/guestbindstore.go | `OpenGuestBindStore(path)` | per-guest enrolled binds (F9 re-assert) | Same tmp+rename 0600 pattern as IntentStore |
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) <-chan error` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank | | `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) (*formatJob, <-chan error)` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank; on success `job.FSUUID` = the new filesystem's UUID, read back only when the durable id still resolves to the formatted device (R-25) — read it only after `done` delivers |
| `TokenStore.Mint` / `Lookup` | internal/localapi/tokenstore.go | `Mint(vmid) (plaintext, error)` | per-guest local-API tokens | Only the SHA-256 hash persists (fsync'd append log); constant-time compare on lookup; plaintext returned exactly once. Lookup RELOADS the file once on a miss (v0.63.0, B3): the one-shot provisioner mints into the same file the daemon indexes — cross-process coherence without a restart; append-only size check bounds the re-read | | `TokenStore.Mint` / `Lookup` | internal/localapi/tokenstore.go | `Mint(vmid) (plaintext, error)` | per-guest local-API tokens | Only the SHA-256 hash persists (fsync'd append log); constant-time compare on lookup; plaintext returned exactly once. Lookup stats the file on EVERY call and reloads BEFORE answering when the append-only log grew (R-269; was reload-on-miss only, v0.63.0 B3, which let a token rotated out by another process keep authorizing as a map hit): the one-shot provisioner mints into the same file the daemon indexes — cross-process coherence both ways without a restart; unchanged size = no re-read. Pinned by `TestTokenStore_RotatedOutTokenRejectedFirst` |
| `FileNonceStore.SeenOrRecord` | internal/authz/noncestore.go | `SeenOrRecord(nonce, exp) bool` | durable anti-replay | fsync'd before returning false; prune only after exp | | `FileNonceStore.SeenOrRecord` | internal/authz/noncestore.go | `SeenOrRecord(nonce, exp) bool` | durable anti-replay | fsync'd before returning false; prune only after exp |
| `Journal` (`Append/Latest/InFlight/AlreadyApplied`) | internal/reconcile/journal.go | `OpenJournal(path)` | op journal + idempotency + crash recovery | `Recover` consumes `InFlight()`; scratch entries special-cased | | `Journal` (`Append/Latest/InFlight/AlreadyApplied`) | internal/reconcile/journal.go | `OpenJournal(path)` | op journal + idempotency + crash recovery | `Recover` consumes `InFlight()`; scratch entries special-cased |
@@ -121,7 +121,7 @@
| Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only | | Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only |
| Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set | | Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set |
| Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction | | Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction |
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did | | Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`; R-856 `crashGuardStatePath` — GET /host/crash-guard's state file, internal/localapi/crashguard.go) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
| `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) | | `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) |
| `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs | | `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs |
| `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s | | `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s |
@@ -157,8 +157,10 @@
| `storage.HostOps` | internal/storage/hostops.go | `*SudoHostOps` (prod), `NoopHostOps` (degraded) | fakes in internal/storage/observe_test.go, watchdog_test.go | | `storage.HostOps` | internal/storage/hostops.go | `*SudoHostOps` (prod), `NoopHostOps` (degraded) | fakes in internal/storage/observe_test.go, watchdog_test.go |
| `storage.HostReader` | internal/storage/hostread.go | `*ProcHostReader` | `fakeHostReader` internal/localapi/disks_test.go; internal/storage/role_test.go. v0.87.0: `BlockSlaves(name)` lists `/sys/block/<name>/slaves` (root-free) — backs the `SystemDisks` dm/md walk (`physicalDisksOf`/`walkSlaves`, role.go); per-branch conservatism: an unresolvable slave fails the WHOLE walk → all-system fail-safe. NEVER weaken the signature test `TestSystemDisks_WalkTopologies` (root-backing disk always in the system set). | | `storage.HostReader` | internal/storage/hostread.go | `*ProcHostReader` | `fakeHostReader` internal/localapi/disks_test.go; internal/storage/role_test.go. v0.87.0: `BlockSlaves(name)` lists `/sys/block/<name>/slaves` (root-free) — backs the `SystemDisks` dm/md walk (`physicalDisksOf`/`walkSlaves`, role.go); per-branch conservatism: an unresolvable slave fails the WHOLE walk → all-system fail-safe. NEVER weaken the signature test `TestSystemDisks_WalkTopologies` (root-backing disk always in the system set). |
| `localapi.DiskOps` / `StorageGate` / `GuestAttacher` / `GuestLister` | internal/localapi/disks.go | `*storage.SudoHostOps`; `storageGateAdapter` (cmd/felhom-agent/main.go); `*GuestBinder`; `*proxmox.Client` | `fakeDiskOps`/`fakeGate`/`fakeGuestAttacher`/`fakeGuestList` internal/localapi/disks_test.go | | `localapi.DiskOps` / `StorageGate` / `GuestAttacher` / `GuestLister` | internal/localapi/disks.go | `*storage.SudoHostOps`; `storageGateAdapter` (cmd/felhom-agent/main.go); `*GuestBinder`; `*proxmox.Client` | `fakeDiskOps`/`fakeGate`/`fakeGuestAttacher`/`fakeGuestList` internal/localapi/disks_test.go |
| `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring |
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go | | `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). | | `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. | | `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. | | `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. | | `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
+53
View File
@@ -0,0 +1,53 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"testing"
)
// R-444: the weekly trim has the guestnet shape (component + reporter seam + goroutine), so its wiring is asserted
// from the AST like TestMainWiresGuestNetWatchdog — a unit-green trim job that main.go never starts is the inert-seam
// defect. It must also share the ONE heavy-op gate (heavyOps), or it could run beside a backup.
func TestMainWiresGuestDiskTrim(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, 0)
if err != nil {
t.Fatalf("parse main.go: %v", err)
}
var constructedWithGate, reporterWired, started bool
ast.Inspect(f, func(n ast.Node) bool {
switch node := n.(type) {
case *ast.CallExpr:
if fn, ok := node.Fun.(*ast.SelectorExpr); ok {
switch fn.Sel.Name {
case "New":
if pkg, ok := fn.X.(*ast.Ident); ok && pkg.Name == "fstrim" && len(node.Args) >= 3 {
if id, ok := node.Args[2].(*ast.Ident); ok && id.Name == "heavyOps" {
constructedWithGate = true
}
}
case "SetGuestDiskTrimReporter":
reporterWired = true
}
}
case *ast.GoStmt:
if sel, ok := node.Call.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "Run" {
if id, ok := sel.X.(*ast.Ident); ok && id.Name == "diskTrim" {
started = true
}
}
}
return true
})
if !constructedWithGate {
t.Error("main.go never calls fstrim.New(..., heavyOps, ...) — no trim job, or one outside the heavy-op gate")
}
if !reporterWired {
t.Error("main.go never calls collector.SetGuestDiskTrimReporter — the guest_disk_trim stanza never reaches the hub")
}
if !started {
t.Error("main.go never starts the trim job with `go diskTrim.Run(ctx)`")
}
}
+20
View File
@@ -38,6 +38,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow" "gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
"gitea.dooplex.hu/admin/felhom-agent/internal/fasttick" "gitea.dooplex.hu/admin/felhom-agent/internal/fasttick"
"gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd" "gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd"
"gitea.dooplex.hu/admin/felhom-agent/internal/fstrim"
"gitea.dooplex.hu/admin/felhom-agent/internal/guesthook" "gitea.dooplex.hu/admin/felhom-agent/internal/guesthook"
"gitea.dooplex.hu/admin/felhom-agent/internal/guestnet" "gitea.dooplex.hu/admin/felhom-agent/internal/guestnet"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub" "gitea.dooplex.hu/admin/felhom-agent/internal/hub"
@@ -1479,6 +1480,25 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
} }
go runJanitor(ctx, jd) go runJanitor(ctx, jd)
} }
// R-444 (`09` §3 decision 139): the weekly guest disk trim — `pct fstrim <vmid>` of every owned, running guest,
// Wednesday from 10:00 local, daytime only, under the one-heavy-op gate. Not part of the errc fan-out: a trim job
// must never be able to bring the agent down.
if cfg.DiskTrim.Enabled() {
dtMode := proxmox.RunnerMode(cfg.Privileged.Mode)
if dtMode == "" {
dtMode = proxmox.RunnerSudo
}
dtRunner := &proxmox.ExecRunner{Mode: dtMode, SudoPath: cfg.Privileged.SudoPath}
dtGuests := localapi.NewStaleLockController(px, dtRunner, reconcile.DefaultPool, logger)
if dtGuests != nil {
diskTrim := fstrim.New(dtRunner, dtGuests, heavyOps,
filepath.Join(cfg.OOB.WithDefaults().StateDir, "guest-disk-trim.json"), logger)
collector.SetGuestDiskTrimReporter(diskTrim)
go diskTrim.Run(ctx)
}
} else {
logger.Info("fstrim: weekly guest disk trim disabled by config (disk_trim.disable)")
}
if lanLoop != nil { if lanLoop != nil {
lanServers = 1 lanServers = 1
go func() { errc <- lanLoop.Run(ctx) }() go func() { errc <- lanLoop.Run(ctx) }()
+9 -1
View File
@@ -129,6 +129,14 @@ Cmnd_Alias FELHOM_CONTROLLERSWAP = \
Cmnd_Alias FELHOM_STALELOCK = \ Cmnd_Alias FELHOM_STALELOCK = \
/usr/sbin/pct ^unlock [0-9]+$ /usr/sbin/pct ^unlock [0-9]+$
# Weekly guest disk trim (R-444, operator ruling `09` §3 decision 139). A thin pool only ever grows from blocks the
# guest has FREED: `fstrim` inside the unprivileged container is refused (FITRIM: Operation not permitted), so the host
# trims the guest's mounts. Measured on demo-hp 2026-10-06: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %,
# apps kept answering. ONE exact pattern: a vmid and nothing else — no `--ignore-mountpoints`, no second argument
# (pinned: TestSudoersFstrimRuleIsExact). The agent runs it on a weekly daytime timer under the heavy-op gate.
Cmnd_Alias FELHOM_FSTRIM = \
/usr/sbin/pct ^fstrim [0-9]+$
# Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS # Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS
# leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and # leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and
# a guest joins that pool only when its restore COMPLETES, so a failed restore leaves a pool-less guest # a guest joins that pool only when its restore COMPLETES, so a failed restore leaves a pool-less guest
@@ -304,4 +312,4 @@ Cmnd_Alias FELHOM_GUESTNET = \
/usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \ /usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \
/usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$ /usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_FSTRIM, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
+66 -1
View File
@@ -546,7 +546,10 @@ class Builder(unittest.TestCase):
r"/etc/felhom/[a-z.-]+)", text)) r"/etc/felhom/[a-z.-]+)", text))
agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"} agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"}
trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD} trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust) # Written by the appliance ISO's first boot (felhom.eu scripts/iso/felhom-bootstrap.sh), never by the installer:
# since installer 1.32.0 (R-275) the uninstall only NAMES them under KEPT.
iso_writes = {"/etc/felhom/.bootstrap-done", "/etc/felhom/appliance-pairing-code"}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust | iso_writes)
self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them") self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them")
# the limits drop-in is named through $AGENT_UNIT in the installer # the limits drop-in is named through $AGENT_UNIT in the installer
self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS) self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS)
@@ -667,5 +670,67 @@ class SelfupdateWrapperConfinement(unittest.TestCase):
self.assertEqual(p.returncode, 1, p.stderr) self.assertEqual(p.returncode, 1, p.stderr)
self.assertIn("outside /var/lib/felhom-os-apply/agent-update", p.stderr) self.assertIn("outside /var/lib/felhom-os-apply/agent-update", p.stderr)
_sb = importlib.machinery.SourceFileLoader("stepbuild", str(REPO / "scripts" / "build-step-bundle.py"))
_ss = importlib.util.spec_from_loader("stepbuild", _sb)
stepbuild = importlib.util.module_from_spec(_ss)
_sb.exec_module(stepbuild)
NEW_IN_0146 = {"/usr/local/sbin/felhom-priv-apply", "/var/lib/vz/snippets/felhom-guest-hook.sh",
"/usr/local/sbin/felhom-shared-parent.sh", "/etc/systemd/system/felhom-shared-parent.service"}
class StepBundle(unittest.TestCase):
"""R-880 (agent v0.146.1): an INSTALLED wrapper checks an incoming bundle's paths against its OWN table (R16), so a
release that adds paths needs a step bundle: the boxes' current bundle with only felhom-os-apply replaced.
RED-PROOF: deliver the full bundle to the old table → R16 (test_the_full_bundle_is_refused_by_an_old_table)."""
def old_world(self):
"""The base bundle an older wrapper (no R-861 paths) installed, and that wrapper's table."""
full = json.loads(builder.build("0.145.0"))
full["files"] = [e for e in full["files"] if e["path"] not in NEW_IN_0146]
old_wrapper = b'# the v0.145.0 wrapper stands in here\nBUNDLE_OP = "agent_config_update"\n'
for e in full["files"]:
if e["path"] == "/usr/local/sbin/felhom-os-apply":
e["content_b64"], e["sha256"] = base64.b64encode(old_wrapper).decode(), hashlib.sha256(old_wrapper).hexdigest()
base = (json.dumps(full, indent=1, sort_keys=True) + "\n").encode()
old_dests = {k: v for k, v in osapply.BUNDLE_DESTS.items() if k not in NEW_IN_0146}
return base, old_dests
def parse_with_table(self, data, dests, version):
saved = osapply.BUNDLE_DESTS
osapply.BUNDLE_DESTS = dests
try:
return osapply.Bundle(osapply.Apply(Box(b"{}", None), "")).parse(data, hashlib.sha256(data).hexdigest(), version)
finally:
osapply.BUNDLE_DESTS = saved
def test_the_full_bundle_is_refused_by_an_old_table(self):
_, old_dests = self.old_world()
full = builder.build("0.146.1")
with self.assertRaises(osapply.Refused) as cm:
self.parse_with_table(full, old_dests, "0.146.1")
self.assertEqual(cm.exception.code, "R16")
def test_the_step_bundle_is_accepted_by_the_old_table_and_changes_only_the_wrapper(self):
base, old_dests = self.old_world()
new_wrapper = (HERE / "felhom-os-apply").read_bytes()
step = stepbuild.build_step(base, "0.146.1-step1", new_wrapper)
ver, files = self.parse_with_table(step, old_dests, "0.146.1-step1")
self.assertEqual(ver, "0.146.1-step1")
b, s_ = json.loads(base), json.loads(step)
self.assertEqual(sorted(e["path"] for e in b["files"]), sorted(e["path"] for e in s_["files"]), "the paths must not change")
changed = [e["path"] for e, f in zip(sorted(b["files"], key=lambda x: x["path"]), sorted(s_["files"], key=lambda x: x["path"]))
if e != f]
self.assertEqual(changed, ["/usr/local/sbin/felhom-os-apply"], "exactly the wrapper changes")
installed = dict((d, c) for d, c, *_ in files)
self.assertEqual(installed["/usr/local/sbin/felhom-os-apply"], new_wrapper)
# and the NEW wrapper (now installed) knows every path the release's full bundle names
self.assertTrue({e["path"] for e in json.loads(builder.build("0.146.1"))["files"]} <= set(osapply.BUNDLE_DESTS))
def test_a_step_version_must_carry_a_suffix(self):
base, _ = self.old_world()
with self.assertRaises(SystemExit):
stepbuild.build_step(base, "0.146.1", b"x")
if __name__ == "__main__": if __name__ == "__main__":
unittest.main() unittest.main()
@@ -241,3 +241,23 @@ func TestNewestArchiveTime_DistinctPhantomsEachAnnounced(t *testing.T) {
t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String()) t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String())
} }
} }
// R-99 (`09` §3 decision 140): the WARN for a PBS phantom ends with the cleanup runbook, so whoever sees it knows the
// one sanctioned way to remove it; a tiny archive on a dir storage is not a PBS phantom and gets no pointer.
// RED-PROOF: drop the `msg += phantomCleanupPointer` line → "the PBS phantom WARN does not end with the runbook pointer".
func TestRejectedArchiveWarnNamesTheCleanupRunbook(t *testing.T) {
var buf bytes.Buffer
r := runnerWithContent(t, &buf, []proxmox.StorageContent{phantomEntry(), goodPBSEntry()})
if _, _, err := r.NewestArchiveTime(context.Background(), 9201); err != nil {
t.Fatal(err)
}
const want = "INCOMPLETE archive when computing tier freshness — it is not a successful backup — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
if !strings.Contains(buf.String(), want) {
t.Errorf("the PBS phantom WARN does not end with the runbook pointer:\n%s", buf.String())
}
local := phantomEntry()
local.Format, local.VolID = "tar.zst", "local:backup/vzdump-lxc-9201-2026_07_28-05_31_14.tar.zst"
if got := rejectedArchiveMessage(local); strings.Contains(got, "pbs-phantom-cleanup") {
t.Errorf("a dir-storage archive got the PBS runbook pointer: %s", got)
}
}
+16 -1
View File
@@ -518,10 +518,25 @@ func (r *BackupRunner) warnRejectedArchiveOnce(e proxmox.StorageContent, why str
if seen { if seen {
return return
} }
r.logger.Warn("backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup", r.logger.Warn(rejectedArchiveMessage(e),
"target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why) "target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
} }
// phantomCleanupPointer names the runbook that removes a PBS phantom (R-99, `09` §3 decision 140: a leftover of an
// aborted upload is deleted on the backup server, by a runbook, when one is seen — never automatically).
const phantomCleanupPointer = " — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
// rejectedArchiveMessage is the WARN text for a rejected archive. Only a PBS entry (format pbs-ct / pbs-vm) gets the
// cleanup pointer: the runbook deletes on a PBS datastore, and a tiny archive on a dir storage is not a PBS phantom.
// Pinned by TestRejectedArchiveWarnNamesTheCleanupRunbook.
func rejectedArchiveMessage(e proxmox.StorageContent) string {
msg := "backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup"
if strings.HasPrefix(e.Format, "pbs-") {
msg += phantomCleanupPointer
}
return msg
}
// demo-felhom in a single afternoon of deploys (2026-07-26). // demo-felhom in a single afternoon of deploys (2026-07-26).
// //
// Asking the STORAGE rather than persisting the store is deliberate: // Asking the STORAGE rather than persisting the store is deliberate:
+5 -1
View File
@@ -25,7 +25,11 @@ import (
// merges the two — see hub.ProvenRestoreTestReporter. This store remains the ONLY place a FAILURE is // merges the two — see hub.ProvenRestoreTestReporter. This store remains the ONLY place a FAILURE is
// recorded, and that asymmetry is deliberate: a failing tier stays due and is retried, so a lost // recorded, and that asymmetry is deliberate: a failing tier stays due and is retried, so a lost
// failure heals itself, while a lost success leaves the system quietly less tested than it believes. // failure heals itself, while a lost success leaves the system quietly less tested than it believes.
// Backups are unaffected — their freshness has a ground truth on the storage (R-84). // Backups are NOT unaffected (corrected 2026-10-05, R-348): byTarget is in memory too, so after a restart the
// reported backup LIST reads 0 until the next backup of each tier runs (daily local, weekly offsite) — measured
// 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What
// is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/
// deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84).
type Store struct { type Store struct {
mu sync.Mutex mu sync.Mutex
byTarget map[string]hub.Backup // latest backup per target id byTarget map[string]hub.Backup // latest backup per target id
+4
View File
@@ -151,6 +151,10 @@ var manifest = []Capability{
// reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ---- // reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ----
{"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""}, {"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""},
// ---- Weekly guest disk trim (FELHOM_FSTRIM, R-444). NON-critical: a missing grant means the thin pool is not
// reclaimed this week (the trim job WARNs per guest and the report shows the failure), not a serving outage. ----
{"guest-fstrim", "weekly guest disk trim (thin-pool reclaim, R-444)", "/usr/sbin/pct", []string{"fstrim", "9201"}, false, ""},
// ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite // ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite
// backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the // backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the
// conf install, unit enable/restart and the handshake read gate the backup path. apt-install // conf install, unit enable/restart and the handshake read gate the backup path. apt-install
@@ -63,3 +63,33 @@ func TestSudoersRefusesTheR861Injections(t *testing.T) {
} }
} }
} }
// R-444: the weekly trim's grant is ONE exact shape — `pct fstrim <vmid>` — and nothing smuggled after it.
// The manifest entry (guest-fstrim) proves the real call is still allowed (TestManifestCoveredBySudoers); this
// pins the other direction. RED-PROOF: write the rule as the glob `/usr/sbin/pct fstrim [0-9]*` → every decoy
// below with a trailing argument matches (the glob's `*` eats spaces).
func TestSudoersFstrimRuleIsExact(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
if !matchesAny("/usr/sbin/pct fstrim 9201", entries) {
t.Fatal("the sudoers does not allow `pct fstrim 9201` — the weekly trim cannot run")
}
for _, c := range []string{
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints",
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints 1",
"/usr/sbin/pct fstrim 9201; x",
"/usr/sbin/pct fstrim 9201 9202",
"/usr/sbin/pct fstrim 92a1",
"/usr/sbin/pct fstrim ",
"/usr/sbin/pct fstrim -- 9201",
"/usr/sbin/pct destroy 9201",
"/usr/sbin/pct destroy 9201 --purge",
} {
if matchesAny(c, entries) {
t.Errorf("the sudoers allows a command the trim rule must not: %q", c)
}
}
}
+11
View File
@@ -33,6 +33,7 @@ type Config struct {
LANResolver LANResolverConfig `json:"lan_resolver"` LANResolver LANResolverConfig `json:"lan_resolver"`
WGTunnel WGTunnelConfig `json:"wg_tunnel"` WGTunnel WGTunnelConfig `json:"wg_tunnel"`
GuestNet GuestNetConfig `json:"guest_net"` GuestNet GuestNetConfig `json:"guest_net"`
DiskTrim DiskTrimConfig `json:"disk_trim"`
OOB OOBConfig `json:"oob"` OOB OOBConfig `json:"oob"`
SelfUpdate SelfUpdateConfig `json:"selfupdate"` SelfUpdate SelfUpdateConfig `json:"selfupdate"`
LogLevel string `json:"log_level"` // debug|info|warn|error (default info) LogLevel string `json:"log_level"` // debug|info|warn|error (default info)
@@ -139,6 +140,16 @@ func (w WGTunnelConfig) WithDefaults() WGTunnelConfig {
return w return w
} }
// DiskTrimConfig configures the R-444 weekly guest disk trim (internal/fstrim). DEFAULT-ON, like GuestNetConfig and for
// the same reason: it only acts on guests the agent already owns, and the operator ruled every box trims (`09` §3
// decision 139). Opting out is the explicit act: `"disk_trim": {"disable": true}`.
type DiskTrimConfig struct {
Disable bool `json:"disable"`
}
// Enabled reports whether the weekly trim should run.
func (d DiskTrimConfig) Enabled() bool { return !d.Disable }
// GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet). // GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet).
// //
// **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other // **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other
+346
View File
@@ -0,0 +1,346 @@
// Package fstrim is the weekly guest disk trim (R-444, operator ruling `09` §3 decision 139).
//
// Why: a thin pool only ever grows from blocks a guest has already FREED — `fstrim` inside the unprivileged container
// is refused (FITRIM: Operation not permitted), and nothing else on the box gives the blocks back. A full thin pool
// takes every guest on the host read-only, so the pool can reach 100 % from deleted data alone. Measured on demo-hp
// 2026-10-06 09:14Z: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %, 18/18 app probes 200, max 1.1 s
// (audits/ten-answers-2026-10-06/r444-measure.txt).
//
// The rule, each part pinned by a test in fstrim_test.go:
// - Weekly: a guest is DUE from Wednesday 10:00 local until it has been trimmed once since then (a box that was off
// on Wednesday catches up at its next eligible hour).
// - Daytime only: a trim starts only between 10:00 and 20:59 local — never in the night window (01:00–06:59) where
// the backups and the restore-tests run (TestEligibleHourNeverInTheNight).
// - Never beside a backup, a restore-test or another heavy operation: the pass holds the host-wide one-heavy-op gate
// (backup.InFlight) for its whole run; a busy gate DEFERS the pass to the next hourly tick.
// - A failed trim is retried at the next eligible hour, at most MaxAttemptsPerWeek times in one week.
// - The last result per guest (time, bytes, ok/fail) is persisted, so a restart neither loses it nor re-trims.
//
// The command is the ONE exact sudoers shape `pct fstrim <vmid>` (FELHOM_FSTRIM). Only guests from the pool-verified
// source (ListLXC ∩ the felhom pool, audit A1) and only RUNNING ones are trimmed.
package fstrim
import (
"context"
"encoding/json"
"fmt"
"log/slog"
"os"
"path/filepath"
"regexp"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// Schedule. The weekday/hours are fixed on purpose (one sentence the operator can read on the System page).
const (
Weekday = time.Wednesday
StartHour = 10 // first eligible local hour (inclusive)
EndHour = 21 // first NOT-eligible local hour (exclusive): last start is 20:59
MaxAttemptsPerWeek = 3
// TickInterval is how often the job looks; a deferred or failed pass is therefore retried the next hour.
TickInterval = time.Hour
// FirstTickDelay lets the agent settle after a start before the first look.
FirstTickDelay = 5 * time.Minute
// PerGuestTimeout bounds one `pct fstrim` (measured 24.4 s for 84 GiB).
PerGuestTimeout = 30 * time.Minute
)
// ScheduleText is the human description carried on the host report.
const ScheduleText = "weekly, due Wednesday from 10:00 host-local time; starts only 10:00-20:59; never beside a backup or restore-test"
// Runner runs a host command (proxmox.ExecRunner in production, through `sudo -n`).
type Runner interface {
Run(ctx context.Context, name string, args ...string) (stdout, stderr []byte, err error)
}
// GuestSource yields the guests this agent OWNS (the pool-verified source, never a bare ListLXC).
type GuestSource interface {
Guests(ctx context.Context) ([]proxmox.Guest, error)
}
// Gate is the host-wide one-heavy-operation gate (*backup.InFlight).
type Gate interface {
TryAcquire(what string) (release func(), busy string, ok bool)
}
// GateName is what the gate reports as busy while a trim runs.
const GateName = "guest-fstrim"
// Record is one guest's last trim attempt, as persisted.
type Record struct {
LastAttemptAt time.Time `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt time.Time `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
// Attempts counts the attempts since the current week's due time (reset by the first attempt of a new week).
Attempts int `json:"attempts"`
}
// Trimmer is the weekly job.
type Trimmer struct {
runner Runner
guests GuestSource
gate Gate
statePath string
logger *slog.Logger
loc *time.Location
now func() time.Time
mu sync.Mutex
records map[int]Record
}
// New builds the job and loads the persisted state. A missing state file is an empty state; a corrupt one is logged
// and treated as empty (the cost is one extra trim, never a missed one).
func New(runner Runner, guests GuestSource, gate Gate, statePath string, logger *slog.Logger) *Trimmer {
if logger == nil {
logger = slog.Default()
}
t := &Trimmer{runner: runner, guests: guests, gate: gate, statePath: statePath, logger: logger,
loc: time.Local, now: time.Now, records: map[int]Record{}}
t.load()
return t
}
func (t *Trimmer) load() {
data, err := os.ReadFile(t.statePath)
if err != nil {
if !os.IsNotExist(err) {
t.logger.Warn("fstrim: state read failed — starting empty", "path", t.statePath, "err", err)
}
return
}
var raw map[string]Record
if err := json.Unmarshal(data, &raw); err != nil {
t.logger.Warn("fstrim: state file corrupt — starting empty", "path", t.statePath, "err", err)
return
}
for k, r := range raw {
if id, err := strconv.Atoi(k); err == nil && id > 0 {
t.records[id] = r
}
}
}
func (t *Trimmer) saveLocked() error {
raw := make(map[string]Record, len(t.records))
for id, r := range t.records {
raw[strconv.Itoa(id)] = r
}
data, err := json.MarshalIndent(raw, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(t.statePath), 0o755); err != nil {
return err
}
tmp := t.statePath + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, t.statePath)
}
// EligibleHour reports whether a trim may START at local time lt.
func EligibleHour(lt time.Time) bool {
h := lt.Hour()
return h >= StartHour && h < EndHour
}
// weekAnchor is the most recent Wednesday StartHour:00 at or before lt (same location as lt).
func weekAnchor(lt time.Time) time.Time {
daysBack := (int(lt.Weekday()) - int(Weekday) + 7) % 7
d := lt.AddDate(0, 0, -daysBack)
a := time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
if a.After(lt) {
d = d.AddDate(0, 0, -7)
a = time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
}
return a
}
// due reports whether a guest with record r (ok=false: none) is due at local time lt.
func due(r Record, has bool, lt time.Time) bool {
if !has {
return true
}
anchor := weekAnchor(lt)
if r.LastAttemptAt.Before(anchor) {
return true // not tried this week
}
return !r.OK && r.Attempts < MaxAttemptsPerWeek
}
// Run looks every TickInterval until ctx ends. It never returns an error: a failed trim is a reported fact.
func (t *Trimmer) Run(ctx context.Context) {
t.logger.Info("fstrim: weekly guest disk trim starting", "schedule", ScheduleText)
timer := time.NewTimer(FirstTickDelay)
defer timer.Stop()
for {
select {
case <-ctx.Done():
return
case <-timer.C:
t.Pass(ctx)
timer.Reset(TickInterval)
}
}
}
// Pass is one look: outside the daytime window it does nothing; otherwise it trims every due, running, owned guest
// while holding the heavy-op gate.
func (t *Trimmer) Pass(ctx context.Context) {
lt := t.now().In(t.loc)
if !EligibleHour(lt) {
t.logger.Debug("fstrim: outside the daytime window — not looking", "local", lt.Format("Mon 15:04"))
return
}
guests, err := t.guests.Guests(ctx)
if err != nil {
t.logger.Warn("fstrim: owned-guest list unavailable — skipping this pass", "err", err)
return
}
owned := make(map[int]bool, len(guests))
var todo []int
t.mu.Lock()
for _, g := range guests {
owned[g.VMID] = true
r, has := t.records[g.VMID]
if !due(r, has, lt) {
continue
}
if g.Status != "running" {
t.logger.Info("fstrim: guest not running — trimmed when it runs", "vmid", g.VMID, "status", g.Status)
continue
}
todo = append(todo, g.VMID)
}
// A guest the agent no longer owns has no result to report.
pruned := false
for id := range t.records {
if !owned[id] {
delete(t.records, id)
pruned = true
}
}
if pruned {
if err := t.saveLocked(); err != nil {
t.logger.Warn("fstrim: state save failed", "err", err)
}
}
t.mu.Unlock()
if len(todo) == 0 {
return
}
sort.Ints(todo)
release, busy, ok := t.gate.TryAcquire(GateName)
if !ok {
t.logger.Info("fstrim: deferred — a heavy operation is in flight; retrying next hour", "busy", busy, "due_guests", len(todo))
return
}
defer release()
for _, vmid := range todo {
if ctx.Err() != nil {
return
}
t.trimOne(ctx, vmid, lt)
}
}
var trimmedLine = regexp.MustCompile(`\((\d+) bytes\) trimmed`)
// ParseTrimmed sums the "(N bytes) trimmed" lines of `pct fstrim` output and counts them (one per mount point), e.g.
// `/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed`.
func ParseTrimmed(out string) (bytes int64, mounts int) {
for _, m := range trimmedLine.FindAllStringSubmatch(out, -1) {
n, err := strconv.ParseInt(m[1], 10, 64)
if err != nil {
continue
}
bytes += n
mounts++
}
return bytes, mounts
}
// GiB renders bytes as "30.1 GiB".
func GiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
func (t *Trimmer) trimOne(ctx context.Context, vmid int, lt time.Time) {
start := t.now()
cctx, cancel := context.WithTimeout(ctx, PerGuestTimeout)
stdout, stderr, err := t.runner.Run(cctx, "pct", "fstrim", strconv.Itoa(vmid))
cancel()
dur := t.now().Sub(start)
bytes, mounts := ParseTrimmed(string(stdout) + "\n" + string(stderr))
t.mu.Lock()
prev, has := t.records[vmid]
r := Record{LastAttemptAt: start.UTC(), OK: err == nil, BytesTrimmed: bytes, Mounts: mounts,
DurationSeconds: float64(dur.Round(100*time.Millisecond)) / float64(time.Second), LastOKAt: prev.LastOKAt}
if has && !prev.LastAttemptAt.Before(weekAnchor(lt)) {
r.Attempts = prev.Attempts + 1
} else {
r.Attempts = 1
}
if err == nil {
r.LastOKAt = start.UTC()
} else {
msg := strings.TrimSpace(err.Error() + ": " + strings.TrimSpace(string(stderr)))
if len(msg) > 300 {
msg = msg[:300]
}
r.Error = msg
}
t.records[vmid] = r
saveErr := t.saveLocked()
t.mu.Unlock()
if err == nil {
t.logger.Info(fmt.Sprintf("fstrim: guest %d trimmed %s in %.1fs", vmid, GiB(bytes), r.DurationSeconds),
"vmid", vmid, "bytes_trimmed", bytes, "mounts", mounts, "duration_s", r.DurationSeconds)
if mounts == 0 {
t.logger.Warn("fstrim: pct fstrim succeeded but reported no trimmed mount — output not understood",
"vmid", vmid, "stdout", strings.TrimSpace(string(stdout)))
}
} else {
t.logger.Warn(fmt.Sprintf("fstrim: guest %d trim FAILED after %.1fs", vmid, r.DurationSeconds),
"vmid", vmid, "attempt", r.Attempts, "max_attempts_per_week", MaxAttemptsPerWeek, "err", r.Error)
}
if saveErr != nil {
t.logger.Warn("fstrim: state save failed — the result will not survive a restart", "path", t.statePath, "err", saveErr)
}
}
// GuestDiskTrimStatus implements hub.GuestDiskTrimReporter: a pure read of the persisted results (never runs pct).
func (t *Trimmer) GuestDiskTrimStatus(context.Context) *hub.GuestDiskTrimStatus {
t.mu.Lock()
defer t.mu.Unlock()
out := &hub.GuestDiskTrimStatus{Schedule: ScheduleText}
ids := make([]int, 0, len(t.records))
for id := range t.records {
ids = append(ids, id)
}
sort.Ints(ids)
for _, id := range ids {
r := t.records[id]
g := hub.GuestDiskTrim{VMID: id, LastAttemptAt: r.LastAttemptAt.UTC().Format(time.RFC3339), OK: r.OK,
BytesTrimmed: r.BytesTrimmed, Mounts: r.Mounts, DurationSeconds: r.DurationSeconds, Error: r.Error}
if !r.LastOKAt.IsZero() {
g.LastOKAt = r.LastOKAt.UTC().Format(time.RFC3339)
}
out.Guests = append(out.Guests, g)
}
return out
}
+267
View File
@@ -0,0 +1,267 @@
package fstrim
import (
"bytes"
"context"
"encoding/json"
"errors"
"log/slog"
"path/filepath"
"reflect"
"strings"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// The real `pct fstrim 9201` output measured on demo-hp 2026-10-06 (audits/ten-answers-2026-10-06/r444-measure.txt).
const measuredOut = "/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed\n" +
"/var/lib/lxc/9201/rootfs/var/lib/felhom: 53.9 GiB (57865633792 bytes) trimmed\n"
const measuredBytes = int64(32277680128 + 57865633792)
type fakeRunner struct {
mu sync.Mutex
calls [][]string
out string
err error
onRun func()
}
func (f *fakeRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
f.mu.Lock()
f.calls = append(f.calls, append([]string{name}, args...))
f.mu.Unlock()
if f.onRun != nil {
f.onRun()
}
if f.err != nil {
return nil, []byte("mount busy"), f.err
}
return []byte(f.out), nil, nil
}
type fakeGuests struct {
g []proxmox.Guest
err error
}
func (f fakeGuests) Guests(context.Context) ([]proxmox.Guest, error) { return f.g, f.err }
// A Wednesday 10:30 in a fixed zone (CEST-like), so the tests do not depend on the machine's zone.
var zone = time.FixedZone("CEST", 2*3600)
func at(day, hour, min int) time.Time { return time.Date(2026, 10, day, hour, min, 0, 0, zone) } // 2026-10-07 = Wednesday
func newT(t *testing.T, r Runner, g GuestSource, gate Gate, now *time.Time) (*Trimmer, *bytes.Buffer, string) {
t.Helper()
var logs bytes.Buffer
path := filepath.Join(t.TempDir(), "guest-disk-trim.json")
tr := New(r, g, gate, path, slog.New(slog.NewTextHandler(&logs, &slog.HandlerOptions{Level: slog.LevelDebug})))
tr.loc = zone
tr.now = func() time.Time { return *now }
return tr, &logs, path
}
func running(ids ...int) fakeGuests {
var g []proxmox.Guest
for _, id := range ids {
g = append(g, proxmox.Guest{VMID: id, Status: "running", Type: "lxc"})
}
return fakeGuests{g: g}
}
func TestParseTrimmedTheMeasuredOutput(t *testing.T) {
b, m := ParseTrimmed(measuredOut)
if b != measuredBytes || m != 2 {
t.Fatalf("ParseTrimmed = %d bytes over %d mounts, want %d over 2", b, m, measuredBytes)
}
if b, m := ParseTrimmed("something else\n"); b != 0 || m != 0 {
t.Fatalf("unrelated output parsed as %d/%d", b, m)
}
if got := GiB(measuredBytes); got != "84.0 GiB" {
t.Fatalf("GiB = %q", got)
}
}
// The night window (01:00–06:59) must never be eligible, and the daytime window is exactly 10:00–20:59.
func TestEligibleHourNeverInTheNight(t *testing.T) {
for h := 0; h < 24; h++ {
lt := time.Date(2026, 10, 7, h, 30, 0, 0, zone)
got := EligibleHour(lt)
if h >= 1 && h <= 6 && got {
t.Errorf("hour %02d is in the night window and must not be eligible", h)
}
if want := h >= 10 && h <= 20; got != want {
t.Errorf("EligibleHour(%02d:30) = %v, want %v", h, got, want)
}
}
}
func TestWeekAnchorIsTheLastWednesdayTen(t *testing.T) {
cases := map[time.Time]time.Time{
at(7, 10, 0): at(7, 10, 0), // Wednesday 10:00 itself
at(7, 9, 59): time.Date(2026, 9, 30, 10, 0, 0, 0, zone), // before 10:00 Wednesday → the previous week
at(8, 15, 0): at(7, 10, 0), // Thursday
at(13, 20, 0): at(7, 10, 0), // next Tuesday
at(14, 11, 0): at(14, 10, 0), // next Wednesday
}
for in, want := range cases {
if got := weekAnchor(in); !got.Equal(want) {
t.Errorf("weekAnchor(%s) = %s, want %s", in.Format("Mon 01-02 15:04"), got.Format("Mon 01-02 15:04"), want.Format("Mon 01-02 15:04"))
}
}
}
// The consequence: on Wednesday 10:30 a running owned guest is trimmed with the ONE exact argv, the bytes are parsed,
// the positive log line is written, the result is persisted, and the host report carries it.
func TestPassTrimsADueGuestAndReportsIt(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
tr, logs, path := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9201"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("runner calls = %q, want %q", r.calls, want)
}
if !strings.Contains(logs.String(), "fstrim: guest 9201 trimmed 84.0 GiB in ") {
t.Fatalf("no positive per-guest log line:\n%s", logs.String())
}
st := tr.GuestDiskTrimStatus(context.Background())
if st == nil || st.Schedule != ScheduleText || len(st.Guests) != 1 {
t.Fatalf("report stanza = %+v", st)
}
g := st.Guests[0]
if g.VMID != 9201 || !g.OK || g.BytesTrimmed != measuredBytes || g.Mounts != 2 || g.LastOKAt == "" || g.LastAttemptAt == "" {
t.Fatalf("report guest = %+v", g)
}
// Persisted: a NEW Trimmer over the same file (an agent restart) still has it and does not trim again this week.
now = at(8, 11, 0)
r2 := &fakeRunner{out: measuredOut}
tr2 := New(r2, running(9201), &backup.InFlight{}, path, slog.New(slog.NewTextHandler(&bytes.Buffer{}, nil)))
tr2.loc, tr2.now = zone, func() time.Time { return now }
if st2 := tr2.GuestDiskTrimStatus(context.Background()); len(st2.Guests) != 1 || st2.Guests[0].BytesTrimmed != measuredBytes {
t.Fatalf("result lost over a restart: %+v", st2)
}
tr2.Pass(context.Background())
if len(r2.calls) != 0 {
t.Fatalf("trimmed again in the same week after a restart: %q", r2.calls)
}
// Next week it is due again.
now = at(14, 10, 5)
tr2.Pass(context.Background())
if len(r2.calls) != 1 {
t.Fatalf("not trimmed in the next week: %q", r2.calls)
}
}
func TestPassNeverRunsInTheNight(t *testing.T) {
for _, h := range []int{1, 3, 6, 9, 21, 23} {
now := at(7, h, 15)
r := &fakeRunner{out: measuredOut}
tr, _, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Errorf("trimmed at %02d:15: %q", h, r.calls)
}
}
}
// A backup (or restore-test) holding the heavy-op gate DEFERS the trim; the next hour, gate free, it runs. And while
// a trim runs, the gate is held, so a backup cannot start beside it.
func TestPassDefersToAHeavyOperationAndRetriesNextHour(t *testing.T) {
now := at(7, 10, 30)
gate := &backup.InFlight{}
release, _, _ := gate.TryAcquire("backup:9201")
var busyDuringTrim string
r := &fakeRunner{out: measuredOut}
r.onRun = func() { busyDuringTrim = gate.Busy() }
tr, logs, _ := newT(t, r, running(9201), gate, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Fatalf("trimmed beside a running backup: %q", r.calls)
}
if !strings.Contains(logs.String(), "fstrim: deferred") || !strings.Contains(logs.String(), "backup:9201") {
t.Fatalf("the deferral is not logged with what holds the gate:\n%s", logs.String())
}
release()
now = now.Add(time.Hour)
tr.Pass(context.Background())
if len(r.calls) != 1 {
t.Fatalf("not retried the next hour: %q", r.calls)
}
if busyDuringTrim != GateName {
t.Fatalf("the heavy-op gate was %q during the trim, want %q", busyDuringTrim, GateName)
}
if gate.Busy() != "" {
t.Fatalf("the gate was not released after the pass: %q", gate.Busy())
}
}
func TestFailedTrimWarnsIsRecordedAndRetriedAtMostThreeTimes(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{err: errors.New("exit status 255")}
tr, logs, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
for i := 0; i < 6; i++ {
tr.Pass(context.Background())
now = now.Add(time.Hour)
}
if len(r.calls) != MaxAttemptsPerWeek {
t.Fatalf("attempts in one week = %d, want %d", len(r.calls), MaxAttemptsPerWeek)
}
if !strings.Contains(logs.String(), "level=WARN") || !strings.Contains(logs.String(), "fstrim: guest 9201 trim FAILED") {
t.Fatalf("no WARN for the failure:\n%s", logs.String())
}
g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]
if g.OK || g.LastOKAt != "" || !strings.Contains(g.Error, "exit status 255") || !strings.Contains(g.Error, "mount busy") {
t.Fatalf("failed result not recorded as a failure: %+v", g)
}
// A success later keeps a clean record.
r.err = nil
r.out = measuredOut
now = at(14, 10, 10)
tr.Pass(context.Background())
if g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]; !g.OK || g.Error != "" || g.BytesTrimmed != measuredBytes {
t.Fatalf("success after failure: %+v", g)
}
}
func TestOnlyRunningOwnedGuestsAndAFailedListActsOnNothing(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
g := fakeGuests{g: []proxmox.Guest{{VMID: 9201, Status: "stopped"}, {VMID: 9202, Status: "running"}}}
tr, _, _ := newT(t, r, g, &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9202"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("calls = %q, want only the running guest", r.calls)
}
r2 := &fakeRunner{out: measuredOut}
tr2, logs, _ := newT(t, r2, fakeGuests{err: errors.New("pool read 403")}, &backup.InFlight{}, &now)
tr2.Pass(context.Background())
if len(r2.calls) != 0 || !strings.Contains(logs.String(), "owned-guest list unavailable") {
t.Fatalf("a failed ownership read must act on nothing: calls %q", r2.calls)
}
}
func TestReportJSONShape(t *testing.T) {
now := at(7, 10, 30)
tr, _, _ := newT(t, &fakeRunner{out: measuredOut}, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
b, err := json.Marshal(tr.GuestDiskTrimStatus(context.Background()))
if err != nil {
t.Fatal(err)
}
for _, k := range []string{`"schedule":`, `"guests":[{"vmid":9201`, `"last_attempt_at":"2026-10-07T08:30:00Z"`, `"ok":true`,
`"bytes_trimmed":90143313920`, `"mounts":2`, `"duration_seconds":`, `"last_ok_at":"2026-10-07T08:30:00Z"`} {
if !strings.Contains(string(b), k) {
t.Errorf("report JSON lacks %s: %s", k, b)
}
}
if strings.Contains(string(b), `"error"`) {
t.Errorf("an ok result must omit error: %s", b)
}
}
+43 -2
View File
@@ -10,6 +10,7 @@ import (
"log/slog" "log/slog"
"os" "os"
"strings" "strings"
"sync"
"time" "time"
"gitea.dooplex.hu/admin/felhom-agent/internal/capability" "gitea.dooplex.hu/admin/felhom-agent/internal/capability"
@@ -89,6 +90,12 @@ type GuestNetReporter interface {
GuestNetStatus(ctx context.Context) *GuestNetStatus GuestNetStatus(ctx context.Context) *GuestNetStatus
} }
// GuestDiskTrimReporter is the R-444 seam the weekly trim job plugs into (same consumer-side pattern — hub does not
// import fstrim). nil (feature not wired) → no guest_disk_trim stanza.
type GuestDiskTrimReporter interface {
GuestDiskTrimStatus(ctx context.Context) *GuestDiskTrimStatus
}
// Collector builds a HostReport from read-only sources. All deps are behind narrow // Collector builds a HostReport from read-only sources. All deps are behind narrow
// interfaces for unit testing. // interfaces for unit testing.
type Collector struct { type Collector struct {
@@ -107,6 +114,7 @@ type Collector struct {
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted) pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted) ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted) guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
diskTrim GuestDiskTrimReporter // R-444: weekly guest disk trim (nil → stanza omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false) selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted) mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
oob OOBReporter // H1: operator-access health (nil → stanza omitted) oob OOBReporter // H1: operator-access health (nil → stanza omitted)
@@ -115,6 +123,7 @@ type Collector struct {
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown) backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
hostID string hostID string
agentVersion string agentVersion string
selfSHA func() string // R-349: sha256 of the running binary; default runningBinarySHA256
logger *slog.Logger logger *slog.Logger
now func() time.Time now func() time.Time
} }
@@ -135,6 +144,7 @@ func NewCollector(px proxmoxReader, cf CloudflaredProber, storage StorageObserve
temp: SysfsTempReader{}, // slice 9: real sysfs reader by default; tests inject a fake temp: SysfsTempReader{}, // slice 9: real sysfs reader by default; tests inject a fake
hostID: hostID, hostID: hostID,
agentVersion: agentVersion, agentVersion: agentVersion,
selfSHA: runningBinarySHA256,
logger: logger, logger: logger,
now: func() time.Time { return time.Now().UTC() }, now: func() time.Time { return time.Now().UTC() },
} }
@@ -218,6 +228,12 @@ func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
return c return c
} }
// SetGuestDiskTrimReporter wires the R-444 weekly trim job as a report source (nil-safe → stanza omitted).
func (c *Collector) SetGuestDiskTrimReporter(r GuestDiskTrimReporter) *Collector {
c.diskTrim = r
return c
}
// SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side // SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side
// pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report. // pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report.
type SelfUpdateReporter interface { type SelfUpdateReporter interface {
@@ -337,6 +353,7 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
HostID: c.hostID, HostID: c.hostID,
ReportedAt: c.now().Format(time.RFC3339), ReportedAt: c.now().Format(time.RFC3339),
AgentVersion: c.agentVersion, AgentVersion: c.agentVersion,
AgentSHA256: c.agentSHA256(),
Host: host, Host: host,
Guests: c.collectGuests(ctx), Guests: c.collectGuests(ctx),
// storage_targets populated this slice (slice 5) via the observer; the rest stay // storage_targets populated this slice (slice 5) via the observer; the rest stay
@@ -373,6 +390,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.guestNet != nil { if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx) report.GuestNet = c.guestNet.GuestNetStatus(ctx)
} }
// R-444: the last weekly trim result per guest (nil reporter = not wired → stanza omitted).
if c.diskTrim != nil {
report.GuestDiskTrim = c.diskTrim.GuestDiskTrimStatus(ctx)
}
// D1: agent self-update pending status (nil reporter → pending=false, the steady state). // D1: agent self-update pending status (nil reporter → pending=false, the steady state).
if c.selfUpdate != nil { if c.selfUpdate != nil {
report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending() report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending()
@@ -431,8 +452,28 @@ const pbsWrapperPath = "/usr/local/sbin/felhom-pbs-apply"
// unreadable file yields "", which the hub reads as UNKNOWN rather than as drift — a host that // unreadable file yields "", which the hub reads as UNKNOWN rather than as drift — a host that
// legitimately has no DR wrapper must not light up amber. The file is 0755, so no privilege is // legitimately has no DR wrapper must not light up amber. The file is 0755, so no privilege is
// needed to read it. // needed to read it.
func pbsWrapperSHA256() string { func pbsWrapperSHA256() string { return fileSHA256(pbsWrapperPath) }
f, err := os.Open(pbsWrapperPath)
// selfExePath is the running binary as the kernel holds it. /proc/self/exe, not the installed path:
// after an A/B flip the file at /usr/local/bin/felhom-agent may already be the NEXT binary while this
// process still runs the old one, and the report must describe what runs (R-349). Test seam.
var selfExePath = "/proc/self/exe"
// runningBinarySHA256 hashes the running binary ONCE per process — the bytes cannot change under a
// running process, and re-hashing ~20 MB every report cycle buys nothing. A failed read is cached as
// "" (UNKNOWN); it never fails the report.
var runningBinarySHA256 = sync.OnceValue(func() string { return fileSHA256(selfExePath) })
func (c *Collector) agentSHA256() string {
if c.selfSHA == nil {
return ""
}
return c.selfSHA()
}
// fileSHA256 is the hex sha256 of a file's bytes, or "" when it cannot be read.
func fileSHA256(path string) string {
f, err := os.Open(path)
if err != nil { if err != nil {
return "" return ""
} }
+64
View File
@@ -0,0 +1,64 @@
package hub
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
)
// R-349: the report carries the sha256 of the binary that is RUNNING, so the hub can tell a
// hand-built proof binary from the vouched artifact of the same version string. The consequence
// asserted: the wire field equals the hash of this very test binary's bytes (read independently via
// os.Executable, a different channel from /proc/self/exe), and it is on the wire as agent_sha256.
func TestCollect_AgentSHA256IsTheRunningBinary(t *testing.T) {
exe, err := os.Executable()
if err != nil {
t.Skipf("os.Executable: %v", err)
}
raw, err := os.ReadFile(exe)
if err != nil {
t.Fatalf("read own binary: %v", err)
}
sum := sha256.Sum256(raw)
want := hex.EncodeToString(sum[:])
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
if r.AgentSHA256 != want {
t.Fatalf("agent_sha256 = %q, want the running binary's %q", r.AgentSHA256, want)
}
b, err := json.Marshal(r)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(b), `"agent_sha256":"`+want+`"`) {
t.Fatalf("agent_sha256 not on the wire: %s", b)
}
}
// An unreadable binary is UNKNOWN (empty, omitted) — never a made-up hash, never a failed report.
func TestFileSHA256_UnreadableIsEmpty(t *testing.T) {
if got := fileSHA256(filepath.Join(t.TempDir(), "absent")); got != "" {
t.Fatalf("absent file hashed to %q, want empty", got)
}
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
c.selfSHA = func() string { return "" }
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect must not fail on an unreadable binary: %v", err)
}
b, _ := json.Marshal(r)
if strings.Contains(string(b), "agent_sha256") {
t.Fatalf("empty agent_sha256 must be omitted: %s", b)
}
}
+54
View File
@@ -0,0 +1,54 @@
package hub
import (
"context"
"encoding/json"
"testing"
)
// R-444: the guest_disk_trim stanza must reach a report built through the PRODUCTION collect path, be absent from
// the wire when the job is not wired, and carry the keys the hub's System page reads.
type fakeDiskTrim struct{ st *GuestDiskTrimStatus }
func (f fakeDiskTrim) GuestDiskTrimStatus(context.Context) *GuestDiskTrimStatus { return f.st }
func TestCollect_GuestDiskTrim(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.150.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ := json.Marshal(r)
var m map[string]any
_ = json.Unmarshal(b, &m)
if _, ok := m["guest_disk_trim"]; ok {
t.Fatalf("guest_disk_trim on the wire with no reporter wired: %s", b)
}
c.SetGuestDiskTrimReporter(fakeDiskTrim{st: &GuestDiskTrimStatus{Schedule: "weekly", Guests: []GuestDiskTrim{{
VMID: 9201, LastAttemptAt: "2026-10-07T08:30:00Z", OK: true, BytesTrimmed: 90143313920, Mounts: 2,
DurationSeconds: 24.4, LastOKAt: "2026-10-07T08:30:00Z",
}}}})
r, err = c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ = json.Marshal(r)
m = nil
_ = json.Unmarshal(b, &m)
dt, ok := m["guest_disk_trim"].(map[string]any)
if !ok || dt["schedule"] != "weekly" {
t.Fatalf("guest_disk_trim missing or wrong on the wire: %s", b)
}
g := dt["guests"].([]any)[0].(map[string]any)
for _, k := range []string{"vmid", "last_attempt_at", "ok", "bytes_trimmed", "mounts", "duration_seconds", "last_ok_at"} {
if _, ok := g[k]; !ok {
t.Fatalf("guest_disk_trim.guests[0] lacks %q: %v", k, g)
}
}
if g["bytes_trimmed"] != float64(90143313920) || g["ok"] != true {
t.Fatalf("values did not survive the round trip: %v", g)
}
}
+9 -7
View File
@@ -54,11 +54,13 @@ const (
DRReasonNoPBSStorage = "no_pbs_storage_observed" DRReasonNoPBSStorage = "no_pbs_storage_observed"
) )
// PBSRootNamespace is how the recipe spells PBS's root namespace. The PBS API spells it as the EMPTY // PBSRootNamespace is how the recipe spells PBS's root namespace: the EMPTY string, PBS's own spelling (R-124,
// string (and `pct restore --ns root` would name a namespace that does not exist) — "root" is a display // agent v0.147.0). It used to be the display word "root", which no PBS namespace is named — an operator pasting it
// convention this wire has always used, kept here so the field's meaning did not change under R-106. // into `proxmox-backup-client … --ns root` during a real recovery got a failure. An empty namespace is ambiguous on
// Only a box with no `namespace` line in its pbs storage.cfg stanza ever emits it. // its own, so READ IT WITH namespace_state: resolved + "" = the root namespace (pass no --ns, or --ns ""); unknown +
const PBSRootNamespace = "root" // "" = the agent could not tell. Only a box with no `namespace` line in its pbs storage.cfg stanza emits it.
// Pinned by TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown and TestR124_RootNamespaceOnTheWireIsPBSSpelling.
const PBSRootNamespace = ""
// DRRecipeHostHalf is the agent-emitted half (guest/drive/storage/PBS scaffolding). Derived entirely // DRRecipeHostHalf is the agent-emitted half (guest/drive/storage/PBS scaffolding). Derived entirely
// from facts the report already collects — no new privileged reads. // from facts the report already collects — no new privileged reads.
@@ -123,8 +125,8 @@ type DRPBSCoord struct {
RepoID string `json:"repo_id"` // the PVE pbs storage id (e.g. "felhom-pbs") — not a token RepoID string `json:"repo_id"` // the PVE pbs storage id (e.g. "felhom-pbs") — not a token
// Namespace is the PBS namespace the restore targets, resolved from the pbs storage's storage.cfg // Namespace is the PBS namespace the restore targets, resolved from the pbs storage's storage.cfg
// stanza — the same field `vzdump --storage <pbs>` makes PVE read, so the recipe cannot disagree // stanza — the same field `vzdump --storage <pbs>` makes PVE read, so the recipe cannot disagree
// with the backup that produced the snapshot. PBSRootNamespace when the box has no namespace // with the backup that produced the snapshot. PBSRootNamespace ("", PBS's spelling, R-124) when the box has no
// configured; "" when NamespaceState is unknown. // namespace configured; also "" when NamespaceState is unknown — consult NamespaceState.
// //
// R-106: this used to come from the listed snapshot's own `ns`, which PBS does not echo per item once // R-106: this used to come from the listed snapshot's own `ns`, which PBS does not echo per item once
// the request is already namespace-scoped via `?ns=` (internal/pbs/client.go). The field was // the request is already namespace-scoped via `?ns=` (internal/pbs/client.go). The field was
+39 -3
View File
@@ -63,8 +63,8 @@ func TestBuildDRRecipeHostHalf(t *testing.T) {
t.Error("felhom-flash (local-dir user-data drive) missing from drives") t.Error("felhom-flash (local-dir user-data drive) missing from drives")
} }
// pbs: latest snapshot's coords + the pbs storage id as repo_id. // pbs: latest snapshot's coords + the pbs storage id as repo_id.
if h.PBS == nil || h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != "root" || h.PBS.LatestSnapshotID != "9201" { if h.PBS == nil || h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != PBSRootNamespace || h.PBS.LatestSnapshotID != "9201" {
t.Errorf("pbs coord = %+v, want repo felhom-pbs/root/9201", h.PBS) t.Errorf("pbs coord = %+v, want repo felhom-pbs, the root namespace (\"\", R-124), snapshot 9201", h.PBS)
} }
} }
@@ -252,7 +252,7 @@ func TestDRRecipe_PBSNamespaceIsThePerCustomerOne(t *testing.T) {
} }
// TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown: a box with a pbs storage and NO namespace line is // TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown: a box with a pbs storage and NO namespace line is
// genuinely in the root namespace. That is an answer, not a gap — it must read resolved/"root", so the // genuinely in the root namespace. That is an answer, not a gap — it must read resolved/"" (PBS's spelling, R-124), so the
// honest root case is never confused with "I could not tell". // honest root case is never confused with "I could not tell".
func TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown(t *testing.T) { func TestDRRecipe_PBSNamespaceRootIsResolvedNotUnknown(t *testing.T) {
h := BuildDRRecipeHostHalf(nil, h := BuildDRRecipeHostHalf(nil,
@@ -453,3 +453,39 @@ func assertNoSecretKeys(t *testing.T, jsonBytes []byte) {
} }
walk("<root>", v) walk("<root>", v)
} }
// R-124: on the WIRE the root namespace is PBS's own spelling — an empty string, present (not omitted), beside
// namespace_state "resolved". The display word "root" names no PBS namespace, and `--ns root` fails in a recovery.
// RED-PROOF: set PBSRootNamespace back to "root" → this test fails.
func TestR124_RootNamespaceOnTheWireIsPBSSpelling(t *testing.T) {
h := BuildDRRecipeHostHalf(nil,
[]StorageTarget{{Name: "felhom-pbs", Type: StorageTypePBS, Content: "backup", PBSNamespace: ""}},
capturedDemoFelhomSnapshots(),
ConfiguredBackupTarget{StorageID: "felhom-pbs", Known: true})
b, err := json.Marshal(h.PBS)
if err != nil {
t.Fatal(err)
}
var m map[string]any
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
ns, present := m["namespace"]
if !present {
t.Fatalf("namespace key missing from %s — an omitted key reads as 'unknown', not 'root'", b)
}
if ns != "" {
t.Fatalf("root namespace on the wire = %q, want \"\" (PBS's spelling; no namespace is named %q)", ns, ns)
}
if m["namespace_state"] != DRStateResolved {
t.Fatalf("namespace_state = %v, want %q beside the empty root namespace", m["namespace_state"], DRStateResolved)
}
// A configured namespace still passes through unchanged.
h2 := BuildDRRecipeHostHalf(nil,
[]StorageTarget{{Name: "felhom-pbs", Type: StorageTypePBS, Content: "backup", PBSNamespace: "demo-felhom"}},
capturedDemoFelhomSnapshots(),
ConfiguredBackupTarget{StorageID: "felhom-pbs", Known: true})
if h2.PBS.Namespace != "demo-felhom" {
t.Fatalf("configured namespace = %q, want demo-felhom", h2.PBS.Namespace)
}
}
+36
View File
@@ -18,6 +18,16 @@ type HostReport struct {
HostID string `json:"host_id"` // echoes config.Hub.HostID HostID string `json:"host_id"` // echoes config.Hub.HostID
ReportedAt string `json:"reported_at"` // RFC3339, agent clock ReportedAt string `json:"reported_at"` // RFC3339, agent clock
AgentVersion string `json:"agent_version"` AgentVersion string `json:"agent_version"`
// AgentSHA256 is the sha256 of the binary this process is RUNNING (read through /proc/self/exe,
// once per process), R-349. The version string cannot tell a hand-built proof binary from the
// published, vouched artifact of the same version — same source, different bytes (`-trimpath
// -buildvcs=false` in release-agent.sh) — so self-update sees "already installed" and never
// corrects it. Reporting the bytes lets the hub compare against the vouched agent_sha256, the
// same mechanism host.wrapper_sha256 is for the PBS wrapper (R-50b(a)).
//
// Empty = unreadable, which the hub must treat as UNKNOWN, never as drift. Pinned by
// TestCollect_AgentSHA256IsTheRunningBinary.
AgentSHA256 string `json:"agent_sha256,omitempty"`
Host HostMetrics `json:"host"` Host HostMetrics `json:"host"`
Guests []Guest `json:"guests"` Guests []Guest `json:"guests"`
@@ -114,6 +124,11 @@ type HostReport struct {
// on HostReport would have been the only report block named against that convention. // on HostReport would have been the only report block named against that convention.
GuestNet *GuestNetStatus `json:"guest_net,omitempty"` GuestNet *GuestNetStatus `json:"guest_net,omitempty"`
// GuestDiskTrim is the weekly guest disk trim stanza (R-444, `09` §3 decision 139): the schedule and, per owned
// guest, the LAST trim result as persisted by the agent (it survives a restart). Present only when the trim job is
// wired; an empty `guests` list means the job runs and no guest has been trimmed yet. No secret.
GuestDiskTrim *GuestDiskTrimStatus `json:"guest_disk_trim,omitempty"`
// LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent // LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent
// mirror of the controller's report log_tails channel. Present ONLY on the heartbeat // mirror of the controller's report log_tails channel. Present ONLY on the heartbeat
// right after the control envelope requested it (log_tail_requested); consume-once on // right after the control envelope requested it (log_tail_requested); consume-once on
@@ -198,6 +213,27 @@ type GuestNetGuest struct {
Message string `json:"message,omitempty"` Message string `json:"message,omitempty"`
} }
// GuestDiskTrimStatus is the R-444 weekly trim stanza. `schedule` is a plain description of when the job runs (local
// time of the host); `guests` holds one entry per owned guest that has had at least one trim attempt.
type GuestDiskTrimStatus struct {
Schedule string `json:"schedule"`
Guests []GuestDiskTrim `json:"guests,omitempty"`
}
// GuestDiskTrim is one guest's LAST trim attempt. `ok` with `last_attempt_at` is the verdict of that attempt — never
// read the time alone as success; `last_ok_at` is the last attempt that succeeded ("" = never). `bytes_trimmed` is
// the sum of the "(N bytes) trimmed" lines `pct fstrim` printed, over `mounts` mount points.
type GuestDiskTrim struct {
VMID int `json:"vmid"`
LastAttemptAt string `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt string `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
}
type PBSDRStatus struct { type PBSDRStatus struct {
State string `json:"state"` State string `json:"state"`
StorageID string `json:"storage_id,omitempty"` StorageID string `json:"storage_id,omitempty"`
+122
View File
@@ -0,0 +1,122 @@
package lanresolver
import (
"context"
"io"
"log/slog"
"os"
"path/filepath"
"strings"
"sync"
"testing"
)
// recRunner records every privileged command EnsureDnsmasq would run and succeeds — nothing reaches
// apt, systemctl or the root checker.
type recRunner struct {
mu sync.Mutex
calls []string
}
func (r *recRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
r.mu.Lock()
defer r.mu.Unlock()
r.calls = append(r.calls, strings.Join(append([]string{name}, args...), " "))
return nil, nil, nil
}
func (r *recRunner) RunStdin(ctx context.Context, _ io.Reader, name string, args ...string) ([]byte, []byte, error) {
return r.Run(ctx, name, args...)
}
func (r *recRunner) installed() bool {
for _, c := range r.calls {
if strings.HasPrefix(c, "apt-get install") && strings.HasSuffix(c, " dnsmasq") {
return true
}
}
return false
}
// fixtureRoot builds a fake host root holding exactly the given relative files and points the REAL
// probe at it for the test's duration.
func fixtureRoot(t *testing.T, files ...string) {
t.Helper()
root := t.TempDir()
for _, f := range files {
p := filepath.Join(root, f)
if err := os.MkdirAll(filepath.Dir(p), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(p, nil, 0o644); err != nil {
t.Fatal(err)
}
}
prev := hostRoot
hostRoot = root
t.Cleanup(func() { hostRoot = prev })
}
func ensure(t *testing.T) *recRunner {
t.Helper()
r := &recRunner{}
m := NewManager(r, "192.0.2.10", []string{"1.1.1.1"}, slog.New(slog.NewTextHandler(io.Discard, nil)))
if err := m.EnsureDnsmasq(context.Background()); err != nil {
t.Fatalf("EnsureDnsmasq: %v", err)
}
return r
}
// R-317: a host with `dnsmasq-base` (the /usr/sbin/dnsmasq binary) but WITHOUT the `dnsmasq` package
// (the service unit) must get the package installed — else the following `systemctl enable --now
// dnsmasq` hits a unit that does not exist and LAN name resolution silently never comes up.
//
// RED-PROOF: probe "usr/sbin/dnsmasq" instead of the unit paths in dnsmasqUnitInstalled → this fails
// with "install was skipped".
func TestEnsureDnsmasq_BinaryWithoutUnitInstalls(t *testing.T) {
fixtureRoot(t, "usr/sbin/dnsmasq")
r := ensure(t)
if !r.installed() {
t.Fatalf("install was skipped on a dnsmasq-base-only host (binary present, unit absent) — "+
"the enable that follows targets a missing unit (R-317). calls: %q", r.calls)
}
}
func TestEnsureDnsmasq_UnitPresentSkipsInstall(t *testing.T) {
for _, unit := range []string{"usr/lib/systemd/system/dnsmasq.service", "lib/systemd/system/dnsmasq.service"} {
t.Run(unit, func(t *testing.T) {
fixtureRoot(t, "usr/sbin/dnsmasq", unit)
if r := ensure(t); r.installed() {
t.Fatalf("apt-get install ran although the dnsmasq unit is present at %s: %q", unit, r.calls)
}
})
}
}
func TestEnsureDnsmasq_NothingPresentInstalls(t *testing.T) {
fixtureRoot(t)
if r := ensure(t); !r.installed() {
t.Fatalf("install skipped on a host with no dnsmasq at all: %q", r.calls)
}
}
// Production wiring for the hostRoot seam: the shipped probe resolves against the real root and asks
// about the unit the `dnsmasq` package owns — never the dnsmasq-base binary.
func TestEnsureDnsmasq_ProductionProbeIsTheUnit(t *testing.T) {
if hostRoot != "/" {
t.Fatalf("hostRoot default = %q, want \"/\" — the production probe would look in the wrong tree", hostRoot)
}
var sawUsrLib bool
for _, p := range dnsmasqUnitPaths {
full := filepath.Join(hostRoot, p)
if strings.HasSuffix(full, "/sbin/dnsmasq") || strings.HasSuffix(full, "/bin/dnsmasq") {
t.Errorf("probe path %s is the dnsmasq-base binary, not the dnsmasq unit (R-317)", full)
}
if full == "/usr/lib/systemd/system/dnsmasq.service" {
sawUsrLib = true
}
}
if !sawUsrLib {
t.Errorf("probe paths %q miss /usr/lib/systemd/system/dnsmasq.service (dpkg -S: owned by dnsmasq)", dnsmasqUnitPaths)
}
}
+26 -2
View File
@@ -101,11 +101,35 @@ func NewManager(runner proxmox.Runner, hostIP string, upstreams []string, logger
} }
} }
// hostRoot is the filesystem root the install probe resolves against: "/" in production; a test
// points it at a fixture tree so the REAL probe runs against files it controls.
var hostRoot = "/"
// dnsmasqUnitPaths are where the `dnsmasq` package ships its systemd unit (Debian; /lib is the
// pre-usrmerge spelling). R-317: probe the UNIT, never /usr/sbin/dnsmasq — that binary belongs to
// `dnsmasq-base`, so a host carrying dnsmasq-base without dnsmasq used to skip the install and then
// `systemctl enable --now dnsmasq` failed against a unit that is not there (resolver never up).
// Pinned by TestEnsureDnsmasq_BinaryWithoutUnitInstalls.
var dnsmasqUnitPaths = []string{
"usr/lib/systemd/system/dnsmasq.service",
"lib/systemd/system/dnsmasq.service",
}
// dnsmasqUnitInstalled reports whether the dnsmasq service unit (the `dnsmasq` package) is present.
func dnsmasqUnitInstalled() bool {
for _, p := range dnsmasqUnitPaths {
if _, err := os.Stat(filepath.Join(hostRoot, p)); err == nil {
return true
}
}
return false
}
// EnsureDnsmasq makes dnsmasq present + enabled and writes the host base config. Idempotent: it // EnsureDnsmasq makes dnsmasq present + enabled and writes the host base config. Idempotent: it
// installs the package only when absent, and writes the base drop-in only when its content changes. // installs the package only when absent, and writes the base drop-in only when its content changes.
func (m *Manager) EnsureDnsmasq(ctx context.Context) error { func (m *Manager) EnsureDnsmasq(ctx context.Context) error {
if _, err := os.Stat("/usr/sbin/dnsmasq"); err != nil { // metadata read, no privilege needed if !dnsmasqUnitInstalled() { // metadata read, no privilege needed
m.logger.Info("lanresolver: dnsmasq absent — installing") m.logger.Info("lanresolver: dnsmasq service unit absent — installing")
if out, errOut, ierr := m.runner.Run(ctx, "apt-get", "install", "-y", "-q", "dnsmasq"); ierr != nil { if out, errOut, ierr := m.runner.Run(ctx, "apt-get", "install", "-y", "-q", "dnsmasq"); ierr != nil {
return fmt.Errorf("install dnsmasq: %s: %w", strings.TrimSpace(string(errOut))+string(out), ierr) return fmt.Errorf("install dnsmasq: %s: %w", strings.TrimSpace(string(errOut))+string(out), ierr)
} }
+102
View File
@@ -0,0 +1,102 @@
package localapi
import (
"encoding/json"
"errors"
"io"
"io/fs"
"net/http"
"os"
"time"
)
// GET /host/crash-guard (R-856, `09` §3 decision 143): what the host's crash guard
// (configs/felhom-crash-guard, `11` §5.9) recorded about the most recent HOST boot. The controller
// reads it once after it starts: when the host's last boot followed an UNCLEAN stop, its app mails
// wait ~15 minutes instead of the normal 90 s boot grace.
//
// Read-only and Proxmox-free: the agent reads the guard's state file (root-owned, 0644 — the
// non-root agent can read it) and passes four fields through. Host-wide, token-authed (any valid
// per-guest token sees the host's view, as GET /host/metrics does).
//
// NEVER an error page. A missing file (no guard installed, or no boot recorded yet), an unreadable
// one, or one that does not parse answers 200 with present:false — the controller reads that as
// UNKNOWN and keeps its normal boot grace. Pinned by TestR856_CrashGuard*.
// defaultCrashGuardStatePath is where configs/felhom-crash-guard writes its state (STATE_DIR there).
const defaultCrashGuardStatePath = "/var/lib/felhom-crash-guard/state.json"
// crashGuardStateMax bounds the read; the real file is well under 4 KiB.
const crashGuardStateMax = 1 << 20
// CrashGuardResponse is the data block of GET /host/crash-guard. Field names are the controller's
// agentapi.CrashGuardState (felhom-controller internal/agentapi/crashguard.go) — a wire contract,
// pinned by TestR856_CrashGuardWireMatchesControllerClient.
type CrashGuardResponse struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"` // RFC3339 UTC ("2006-01-02T15:04:05Z")
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// crashGuardFile is the subset of the guard's state.json the route passes through. Every other key
// (armed, boot_id, config, unclean_boots, last_trip, ...) is ignored.
type crashGuardFile struct {
LastBootAt string `json:"last_boot_at"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// readCrashGuardState reads and parses the guard's state file. ok=false on ANY failure (missing,
// unreadable, oversized, not a JSON object, a field of the wrong type); reason says which, for the log.
func readCrashGuardState(path string) (resp CrashGuardResponse, ok bool, reason string) {
f, err := os.Open(path)
if err != nil {
if errors.Is(err, fs.ErrNotExist) {
return resp, false, "no state file"
}
return resp, false, "unreadable: " + err.Error()
}
defer f.Close()
raw, err := io.ReadAll(io.LimitReader(f, crashGuardStateMax+1))
if err != nil {
return resp, false, "read: " + err.Error()
}
if len(raw) > crashGuardStateMax {
return resp, false, "state file too large"
}
var st crashGuardFile
// Unmarshal into a struct fails on a non-object top level (null decodes, so reject it below).
if err := json.Unmarshal(raw, &st); err != nil {
return resp, false, "unparseable: " + err.Error()
}
var probe map[string]json.RawMessage
if err := json.Unmarshal(raw, &probe); err != nil || probe == nil {
return resp, false, "unparseable: not a JSON object"
}
resp = CrashGuardResponse{Present: true, LastBootUnclean: st.LastBootUnclean, Tripped: st.Tripped}
// Normalise to RFC3339 UTC; an unparseable time passes through as-is (the controller reads an
// unparseable boot time as "not this start's boot" → its normal grace).
if t, perr := time.Parse(time.RFC3339, st.LastBootAt); perr == nil {
resp.LastBootAt = t.UTC().Format(time.RFC3339)
} else {
resp.LastBootAt = st.LastBootAt
}
return resp, true, ""
}
func (s *Server) handleCrashGuard(w http.ResponseWriter, r *http.Request, vmid int) {
path := s.crashGuardStatePath
if path == "" {
path = defaultCrashGuardStatePath
}
resp, ok, reason := readCrashGuardState(path)
if !ok {
s.logger.Debug("local-api: /host/crash-guard not present", "vmid", vmid, "reason", reason)
writeOK(w, CrashGuardResponse{Present: false})
return
}
s.logger.Debug("local-api: /host/crash-guard served", "vmid", vmid,
"last_boot_at", resp.LastBootAt, "last_boot_unclean", resp.LastBootUnclean, "tripped", resp.Tripped)
writeOK(w, resp)
}
+205
View File
@@ -0,0 +1,205 @@
package localapi
import (
"encoding/json"
"io"
"log/slog"
"net/http"
"os"
"path/filepath"
"testing"
)
// The shape of /var/lib/felhom-crash-guard/state.json as read on demo-hp on 2026-10-06 (values from
// that read where they matter; lists/objects kept to the same key set).
const crashGuardFixture = `{
"armed": true,
"boot_id": "3f1c0f1e-6a0b-4d7e-9b7a-0c2d4e6f8a1b",
"config": {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10},
"kernel_panic": 10,
"last_boot_at": "2026-10-05T07:56:41Z",
"last_boot_unclean": true,
"last_trip": {},
"rearmed_at": "2026-10-04T14:02:11Z",
"rearmed_by": "operator",
"tripped": false,
"unclean_boots": ["2026-10-05T07:56:41Z"],
"unclean_boots_24h": 1,
"unclean_boots_in_window": 1,
"updated_at": "2026-10-05T07:57:02Z",
"version": 1
}`
// controllerCrashGuardState is a COPY of the controller's wire type, felhom-controller
// controller/internal/agentapi/crashguard.go `CrashGuardState` (commit 8b13a5e) — same field names,
// same tags. If either side renames a key, the contract test below fails.
type controllerCrashGuardState struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
func newCrashGuardServer(t *testing.T, statePath string) http.Handler {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0",
Guests: &fakeGuests{},
Backups: &fakeBackups{},
Store: &fakeStore{},
Storage: fakeStorage{},
Tokens: staticTokens{"A": 8200, "B": 9300},
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatalf("new server: %v", err)
}
srv.crashGuardStatePath = statePath
return srv.Handler()
}
func writeCrashGuardFixture(t *testing.T, body string) string {
t.Helper()
p := filepath.Join(t.TempDir(), "state.json")
if err := os.WriteFile(p, []byte(body), 0o644); err != nil {
t.Fatal(err)
}
return p
}
// getCrashGuard calls the route and decodes the envelope with the CONTROLLER's type.
func getCrashGuard(t *testing.T, h http.Handler, token string) (int, controllerCrashGuardState, string) {
t.Helper()
w := do(t, h, "GET", "/host/crash-guard", token, "")
var env struct {
OK bool `json:"ok"`
Data controllerCrashGuardState `json:"data"`
}
if w.Code == http.StatusOK {
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil {
t.Fatalf("decode %q: %v", w.Body.String(), err)
}
if !env.OK {
t.Fatalf("ok=false: %s", w.Body.String())
}
}
return w.Code, env.Data, w.Body.String()
}
// A present state file (demo-hp's shape) passes the three facts through.
func TestR856_CrashGuardPresentFile(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, body)
}
want := controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: true, Tripped: false}
if st != want {
t.Fatalf("state = %+v, want %+v", st, want)
}
// A tripped, clean boot reads back as such (both bools are carried, not defaulted).
h = newCrashGuardServer(t, writeCrashGuardFixture(t,
`{"last_boot_at":"2026-10-05T09:56:41+02:00","last_boot_unclean":false,"tripped":true,"version":1}`))
_, st, _ = getCrashGuard(t, h, "B")
want = controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: false, Tripped: true}
if st != want {
t.Fatalf("offset time / tripped: state = %+v, want %+v (time normalised to UTC Z)", st, want)
}
}
// No state file (no guard on this host, or no boot recorded yet) → 200 present:false.
func TestR856_CrashGuardMissingFile(t *testing.T) {
h := newCrashGuardServer(t, filepath.Join(t.TempDir(), "absent", "state.json"))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("missing file: got %d, want 200 (%s)", code, body)
}
if st.Present || st.LastBootUnclean || st.Tripped || st.LastBootAt != "" {
t.Fatalf("missing file: state = %+v, want present:false and nothing else", st)
}
}
// A garbled file → 200 present:false, never a 5xx — every shape of garbage.
func TestR856_CrashGuardGarbageFile(t *testing.T) {
for name, body := range map[string]string{
"truncated": crashGuardFixture[:40],
"not json": "this is not json\n",
"empty": "",
"null": "null",
"array": `[{"last_boot_unclean":true}]`,
"wrong type": `{"last_boot_at":"2026-10-05T07:56:41Z","last_boot_unclean":"yes","tripped":false}`,
"lone brace": "{",
} {
t.Run(name, func(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, body))
code, st, raw := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, raw)
}
if st.Present || st.LastBootUnclean {
t.Fatalf("garbage %q read as %+v, want present:false", name, st)
}
})
}
// The path is a directory, not a file: unreadable → present:false, 200.
h := newCrashGuardServer(t, t.TempDir())
if code, st, raw := getCrashGuard(t, h, "A"); code != http.StatusOK || st.Present {
t.Fatalf("directory path: got %d %+v (%s), want 200 present:false", code, st, raw)
}
}
// No / unknown token → 401, like every sibling route; a cross-guest ?vmid= → 403.
func TestR856_CrashGuardRequiresGuestToken(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
for _, tok := range []string{"", "bogus"} {
w := do(t, h, "GET", "/host/crash-guard", tok, "")
if w.Code != http.StatusUnauthorized {
t.Fatalf("token %q: got %d, want 401", tok, w.Code)
}
if json.Valid(w.Body.Bytes()) {
var env struct {
Data controllerCrashGuardState `json:"data"`
}
_ = json.Unmarshal(w.Body.Bytes(), &env)
if env.Data.Present || env.Data.LastBootUnclean {
t.Fatalf("token %q: the refusal leaked the state: %s", tok, w.Body.String())
}
}
}
if w := do(t, h, "GET", "/host/crash-guard?vmid=9300", "A", ""); w.Code != http.StatusForbidden {
t.Fatalf("cross-guest query: got %d, want 403", w.Code)
}
}
// Wire contract: every key the controller's CrashGuardState decodes is emitted under exactly that
// name, and the agent emits no key the controller does not know.
func TestR856_CrashGuardWireMatchesControllerClient(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
w := do(t, h, "GET", "/host/crash-guard", "A", "")
var env struct {
OK bool `json:"ok"`
Data map[string]json.RawMessage `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil || !env.OK {
t.Fatalf("envelope: %v %s", err, w.Body.String())
}
want := []string{"present", "last_boot_at", "last_boot_unclean", "tripped"}
for _, k := range want {
if _, ok := env.Data[k]; !ok {
t.Errorf("agent does not emit %q, which the controller decodes (%s)", k, w.Body.String())
}
}
if len(env.Data) != len(want) {
t.Errorf("agent emits %d keys, controller knows %d: %s", len(env.Data), len(want), w.Body.String())
}
// And the agent's own type agrees with the controller's copy, field for field.
var mine CrashGuardResponse
var theirs controllerCrashGuardState
raw, _ := json.Marshal(env.Data)
_ = json.Unmarshal(raw, &mine)
_ = json.Unmarshal(raw, &theirs)
if (controllerCrashGuardState{mine.Present, mine.LastBootAt, mine.LastBootUnclean, mine.Tripped}) != theirs {
t.Errorf("agent %+v vs controller %+v", mine, theirs)
}
}
+20 -8
View File
@@ -398,9 +398,16 @@ func (s *Server) handleDisks(w http.ResponseWriter, r *http.Request, vmid int) {
di.Smart = &sm di.Smart = &sm
} }
} }
if total, used, okc := statfsCapacity(d.MountPath); okc { // R-118: statfs ONLY while the drive's device is present. With the device gone the raw
di.TotalBytes, di.UsedBytes = total, used // mountpoint reverts to a bare directory on the ROOT filesystem, and statfs would report
di.UsedFraction = float64(used) / float64(total) // pve-root's size as this drive's (measured: a 4 GB drive advertising 46 GiB). Same trap
// observe.go guards on the Observe path. Absent → capacity left zero (unknown), never root's.
// Pinned by TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity.
if s.devicePresent(d.MountPath) {
if total, used, okc := statfsCapacity(d.MountPath); okc {
di.TotalBytes, di.UsedBytes = total, used
di.UsedFraction = float64(used) / float64(total)
}
} }
out = append(out, di) out = append(out, di)
} }
@@ -761,6 +768,11 @@ type FormatResponse struct {
// signature — the customer authorizes the wipe of their own data drive. // signature — the customer authorizes the wipe of their own data drive.
NeedsConfirmation bool `json:"needs_confirmation,omitempty"` NeedsConfirmation bool `json:"needs_confirmation,omitempty"`
DurableID string `json:"durable_id,omitempty"` // the durable id to confirm against (user-data) DurableID string `json:"durable_id,omitempty"` // the durable id to confirm against (user-data)
// FSUUID (on Formatted) is the UUID of the filesystem the agent just made, verified against the
// bound durable id after mkfs (R-25). The caller mounts THIS — re-resolving a UUID from the /dev
// path later can name another disk if /dev re-enumerated. "" = not verified: the caller must not
// substitute a path-resolved guess silently.
FSUUID string `json:"fs_uuid,omitempty"`
// PendingOp is set on a SYSTEM/BACKUP data-bearing refusal — the exact op the operator must sign. // PendingOp is set on a SYSTEM/BACKUP data-bearing refusal — the exact op the operator must sign.
PendingOp *PendingOp `json:"pending_op,omitempty"` PendingOp *PendingOp `json:"pending_op,omitempty"`
} }
@@ -808,7 +820,7 @@ func (s *Server) handleDiskFormatStatus(w http.ResponseWriter, r *http.Request,
writeOK(w, map[string]any{ writeOK(w, map[string]any{
"vmid": vmid, "phase": job.Phase, "device": job.Device, "fstype": job.FSType, "vmid": vmid, "phase": job.Phase, "device": job.Device, "fstype": job.FSType,
"durable_id": job.DurableID, "error": job.Error, "started_at": job.StartedAt, "updated_at": job.UpdatedAt, "durable_id": job.DurableID, "error": job.Error, "started_at": job.StartedAt, "updated_at": job.UpdatedAt,
"job_id": job.JobID, "job_id": job.JobID, "fs_uuid": job.FSUUID, // R-25: "" until done + verified
}) })
} }
@@ -871,7 +883,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
"format refused (device may have changed since inspection): "+rerr.Error()) "format refused (device may have changed since inspection): "+rerr.Error())
return return
} }
done := s.startFormatDetached(device, blankDurable, req.FSType, true) job, done := s.startFormatDetached(device, blankDurable, req.FSType, true)
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil { if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
if err == errFormatClientGone { if err == errFormatClientGone {
return // client gone; mkfs continues detached + the job record records the outcome return // client gone; mkfs continues detached + the job record records the outcome
@@ -880,7 +892,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
writeErr(w, http.StatusBadGateway, "format failed: "+err.Error()) writeErr(w, http.StatusBadGateway, "format failed: "+err.Error())
return return
} }
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, Reason: "blank device formatted " + req.FSType}) writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, FSUUID: job.FSUUID, Reason: "blank device formatted " + req.FSType})
return return
} }
@@ -917,7 +929,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
// F20-BUG3: run the destructive mkfs DETACHED off s.baseCtx (bound durable id recorded for // F20-BUG3: run the destructive mkfs DETACHED off s.baseCtx (bound durable id recorded for
// restart-recovery), so a request/client deadline can never SIGKILL it mid-write and corrupt the // restart-recovery), so a request/client deadline can never SIGKILL it mid-write and corrupt the
// disk. We still wait to return the synchronous result (backward-compatible with the controller). // disk. We still wait to return the synchronous result (backward-compatible with the controller).
done := s.startFormatDetached(device, deviceDurable, req.FSType, false) job, done := s.startFormatDetached(device, deviceDurable, req.FSType, false)
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil { if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
if err == errFormatClientGone { if err == errFormatClientGone {
return // client gone; the wipe continues detached + survives a restart via the job record return // client gone; the wipe continues detached + survives a restart via the job record
@@ -929,7 +941,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
s.logger.Warn("local-api: USER-DATA data-bearing format — CUSTOMER CONFIRMED (no operator signature)", s.logger.Warn("local-api: USER-DATA data-bearing format — CUSTOMER CONFIRMED (no operator signature)",
"vmid", vmid, "device", device, "durable_id", deviceDurable, "fstype", req.FSType, "why", probe.Reason()) "vmid", vmid, "device", device, "durable_id", deviceDurable, "fstype", req.FSType, "why", probe.Reason())
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: true, writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: true,
Role: string(role), DurableID: deviceDurable, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"}) Role: string(role), DurableID: deviceDurable, FSUUID: job.FSUUID, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"})
return return
} }
@@ -5,6 +5,7 @@ import (
"encoding/json" "encoding/json"
"io" "io"
"log/slog" "log/slog"
"runtime"
"strings" "strings"
"testing" "testing"
@@ -184,3 +185,39 @@ func TestDisks_DevicePresence_WireFieldIsFalseOnDeviceLoss(t *testing.T) {
t.Fatalf("the drive never reached the wire: %s", body) t.Fatalf("the drive never reached the wire: %s", body)
} }
} }
// ── R-118 — an absent drive must not advertise the ROOT filesystem's capacity ───────────────────
// TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity drives the REAL statfsCapacity (no capacity
// seam): the registry drive's mount path is a real, bare temp directory — exactly what /mnt/<name>
// becomes once its device is gone (a plain directory on the host's filesystem). With the device absent
// the row must carry NO capacity; before R-118 the union path statfs'd that bare directory and reported
// the host filesystem's size and usage as the drive's (46 GiB at 9.2 % for a 4 GB drive, measured).
// The present half proves the test is not hollow: the same directory DOES yield capacity when the
// device is there, so a zero on the absent half is the guard's doing, not a statfs failure.
//
// RED-PROOF: drop the `if s.devicePresent(d.MountPath)` guard around statfsCapacity in disks.go → the
// absent subtest fails with "advertises ... bytes".
func TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity(t *testing.T) {
if runtime.GOOS != "linux" {
t.Skip("statfsCapacity is linux-only; production target is linux")
}
bare := t.TempDir()
known := []storage.KnownTarget{
{Name: "cel", Type: hub.StorageTypeUSB, MountPath: bare, DurableID: "uuid:4242", UUID: "4242"},
}
t.Run("absent", func(t *testing.T) {
di := diskByMount(t, presenceServer(t, nil, known, true, false), bare)
if di.TotalBytes != 0 || di.UsedBytes != 0 || di.UsedFraction != 0 {
t.Errorf("absent drive advertises total=%d used=%d frac=%.3f — that is the filesystem UNDER "+
"the bare mountpoint, not the drive (R-118)", di.TotalBytes, di.UsedBytes, di.UsedFraction)
}
})
t.Run("present", func(t *testing.T) {
di := diskByMount(t, presenceServer(t, nil, known, true, true), bare)
if di.TotalBytes <= 0 {
t.Errorf("present drive reports no capacity (total=%d) — the guard over-corrected and the "+
"size bar is gone for every healthy registry drive", di.TotalBytes)
}
})
}
+6
View File
@@ -28,6 +28,7 @@ type fakeDiskOps struct {
unmountCalls []string unmountCalls []string
candidates []storage.CandidateDisk // returned by ListCandidateDisks candidates []storage.CandidateDisk // returned by ListCandidateDisks
candErr error candErr error
afterFormat *storage.DeviceProbe // R-25: when set, InspectDevice returns it once a format ran
} }
func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.CandidateDisk, error) { func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.CandidateDisk, error) {
@@ -35,7 +36,12 @@ func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.Candidate
} }
func (f *fakeDiskOps) InspectDevice(_ context.Context, device string) (storage.DeviceProbe, error) { func (f *fakeDiskOps) InspectDevice(_ context.Context, device string) (storage.DeviceProbe, error) {
f.mu.Lock()
p := f.probe p := f.probe
if f.afterFormat != nil && len(f.formatCalls) > 0 {
p = *f.afterFormat
}
f.mu.Unlock()
p.Device = device p.Device = device
return p, f.inspectErr return p, f.inspectErr
} }
+96
View File
@@ -0,0 +1,96 @@
package localapi
import (
"context"
"encoding/json"
"net/http"
"sync"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/storage"
)
// R-25: the format answer carries the UUID of the filesystem the agent JUST made, verified against the
// bound durable id after mkfs, so the controller mounts that filesystem rather than whatever the /dev
// path resolves to a few requests later. The consequence asserted: the UUID on the wire (and in the
// polled job record) is the new superblock's — and is EMPTY whenever the binding cannot be re-proved.
const newFSUUID = "0fc63daf-8483-4772-8e79-3d69d8477de4"
func confirmedFormat(t *testing.T, d *fakeDiskOps, srv *Server, fj *FormatJobStore) (string, *formatJob) {
t.Helper()
w := do(t, srv.Handler(), "POST", "/disks/format", "A", `{"device":"/dev/sdb1","fstype":"ext4","confirmed":true,"durable_id":"byid:wwn-/dev/sdb1"}`)
if w.Code != http.StatusOK {
t.Fatalf("confirmed format: %d (%s)", w.Code, w.Body.String())
}
var resp struct {
Data FormatResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatalf("decode: %v (%s)", err, w.Body.String())
}
if !resp.Data.Formatted {
t.Fatalf("not formatted: %s", w.Body.String())
}
return resp.Data.FSUUID, waitFormatPhase(t, fj, formatPhaseDone)
}
func confirmedGate() *fakeGate {
return &fakeGate{decision: WipeDecision{Allowed: true, Tier: "customer_confirmable", Reason: "customer_confirmed"}}
}
func TestFormat_ReportsNewFilesystemUUID(t *testing.T) {
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
fj := tempFormatStore(t)
srv := formatServer(t, d, confirmedGate(), fj)
got, job := confirmedFormat(t, d, srv, fj)
if got != newFSUUID {
t.Fatalf("response fs_uuid = %q, want the new filesystem's %q", got, newFSUUID)
}
if job.FSUUID != newFSUUID {
t.Fatalf("job record fs_uuid = %q, want %q (the polled status path)", job.FSUUID, newFSUUID)
}
}
// The node moved between mkfs and the read-back: the bound durable id now resolves elsewhere. The
// UUID must NOT be reported — reading it would name the other disk's filesystem.
func TestFormat_FSUUIDWithheldWhenDurableIDMoved(t *testing.T) {
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
fj := tempFormatStore(t)
srv := formatServer(t, d, confirmedGate(), fj)
var mu sync.Mutex
calls := 0
srv.reresolveWipe = func(_ context.Context, _ string) (string, error) {
mu.Lock()
defer mu.Unlock()
calls++
if calls == 1 {
return "/dev/sdb", nil // the pre-mkfs anti-retarget re-resolve
}
return "/dev/sdc", nil // after mkfs: the durable id now names another node
}
got, job := confirmedFormat(t, d, srv, fj)
if got != "" || job.FSUUID != "" {
t.Fatalf("fs_uuid reported after the durable id moved (response %q, job %q) — must be empty", got, job.FSUUID)
}
if calls < 2 {
t.Fatalf("the post-mkfs re-resolve never ran (calls=%d)", calls)
}
}
// The superblock did not read back as the requested filesystem → not verified → empty.
func TestFormat_FSUUIDWithheldOnFSTypeMismatch(t *testing.T) {
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "xfs", FSUUID: newFSUUID}}
fj := tempFormatStore(t)
srv := formatServer(t, d, confirmedGate(), fj)
got, job := confirmedFormat(t, d, srv, fj)
if got != "" || job.FSUUID != "" {
t.Fatalf("fs_uuid reported for a superblock of the wrong type (response %q, job %q)", got, job.FSUUID)
}
}
+39 -3
View File
@@ -22,6 +22,10 @@ type formatJob struct {
Blank bool `json:"blank,omitempty"` // audit D3: blank (benign) format — recovery re-checks STILL-blank, not data-bearing Blank bool `json:"blank,omitempty"` // audit D3: blank (benign) format — recovery re-checks STILL-blank, not data-bearing
Phase string `json:"phase"` // running | done | failed Phase string `json:"phase"` // running | done | failed
Error string `json:"error,omitempty"` Error string `json:"error,omitempty"`
// FSUUID is the filesystem UUID of the NEW filesystem, read by the agent right after mkfs on the
// device the bound durable id still resolves to (R-25). "" = not verified (the controller must not
// read that as a UUID). Set only on phase done.
FSUUID string `json:"fs_uuid,omitempty"`
StartedAt string `json:"started_at"` StartedAt string `json:"started_at"`
UpdatedAt string `json:"updated_at"` UpdatedAt string `json:"updated_at"`
} }
@@ -96,7 +100,10 @@ func (s *FormatJobStore) save(j *formatJob) error {
// runs to completion and records the outcome. device is the ALREADY anti-retarget-resolved device; the // runs to completion and records the outcome. device is the ALREADY anti-retarget-resolved device; the
// record carries durableID so a restart can re-resolve + re-run. blank marks a benign (blank-device) // record carries durableID so a restart can re-resolve + re-run. blank marks a benign (blank-device)
// format, so restart recovery re-checks STILL-blank rather than data-bearing (audit D3). // format, so restart recovery re-checks STILL-blank rather than data-bearing (audit D3).
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) <-chan error { //
// The returned job may be read (FSUUID) only AFTER a value arrives on done — the goroutine writes it
// before the send, which is the happens-before edge.
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) (*formatJob, <-chan error) {
base := s.baseCtx base := s.baseCtx
if base == nil { if base == nil {
base = context.Background() base = context.Background()
@@ -116,10 +123,39 @@ func (s *Server) startFormatDetached(device, durableID, fstype string, blank boo
ctx, cancel := context.WithTimeout(base, 60*time.Minute) ctx, cancel := context.WithTimeout(base, 60*time.Minute)
defer cancel() defer cancel()
err := s.disks.Format(ctx, device, fstype) err := s.disks.Format(ctx, device, fstype)
if err == nil {
job.FSUUID = s.formattedFSUUID(ctx, device, durableID, fstype)
}
s.finishFormatJob(job, err) s.finishFormatJob(job, err)
done <- err done <- err
}() }()
return done return job, done
}
// formattedFSUUID reads the UUID of the filesystem the agent has JUST made (R-25). The caller used to
// re-resolve the UUID from the mutable /dev path afterwards, over separate requests — a re-enumeration
// in that window could hand it ANOTHER disk's filesystem to mount. Here the bound durable id must still
// resolve to the very device that was formatted (and re-derive to the same id), and the superblock must
// carry the fstype that was asked for; anything else returns "" (not verified), never a guess.
func (s *Server) formattedFSUUID(ctx context.Context, device, durableID, fstype string) string {
if durableID == "" || s.reresolveWipe == nil {
return ""
}
// The device now holds a filesystem, so the data-bearing anti-retarget re-resolve is the right one.
now, err := s.reresolveWipe(ctx, durableID)
if err != nil || now != device {
s.logger.Warn("format: new filesystem UUID NOT reported — bound durable id no longer resolves to the formatted device",
"device", device, "durable_id", durableID, "resolves_to", now, "err", err)
return ""
}
probe, err := s.disks.InspectDevice(ctx, device)
if err != nil || !probe.Probed || probe.FSType != fstype || probe.FSUUID == "" {
s.logger.Warn("format: new filesystem UUID NOT reported — superblock did not read back as the requested filesystem",
"device", device, "want_fstype", fstype, "got_fstype", probe.FSType, "has_uuid", probe.FSUUID != "", "err", err)
return ""
}
s.logger.Info("format: new filesystem bound to its durable id", "device", device, "durable_id", durableID, "fs_uuid", probe.FSUUID)
return probe.FSUUID
} }
// finishFormatJob updates the persisted record to done/failed. // finishFormatJob updates the persisted record to done/failed.
@@ -172,7 +208,7 @@ func (s *Server) RecoverFormatJob(ctx context.Context) {
return return
} }
s.logger.Warn("format-job recover: re-running interrupted format detached", "durable_id", job.DurableID, "device", device, "fstype", job.FSType, "blank", job.Blank) s.logger.Warn("format-job recover: re-running interrupted format detached", "durable_id", job.DurableID, "device", device, "fstype", job.FSType, "blank", job.Blank)
_ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion _, _ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion
} }
// nowFn returns the server clock (testable), defaulting to time.Now. // nowFn returns the server clock (testable), defaulting to time.Now.
+6
View File
@@ -294,6 +294,9 @@ type Server struct {
netMountRoot string // the user-data namespace root for the network-mount role gate netMountRoot string // the user-data namespace root for the network-mount role gate
smbCredsDir string // where SMB creds files are written (out-of-band, 0600) smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
// crashGuardStatePath (R-856) is the host crash guard's state file read by GET /host/crash-guard;
// empty = defaultCrashGuardStatePath. A seam: tests point it at a fixture.
crashGuardStatePath string
// escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed // escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the // identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
// offsite repository password. OPTIONAL — nil (no hub client configured) makes // offsite repository password. OPTIONAL — nil (no hub client configured) makes
@@ -520,6 +523,9 @@ func (s *Server) Handler() http.Handler {
// Host metrics (slice 9): host-wide health + per-storage capacity for the customer's monitoring // Host metrics (slice 9): host-wide health + per-storage capacity for the customer's monitoring
// view. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot). // view. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot).
mux.HandleFunc("GET /host/metrics", s.withGuest(s.handleHostMetrics)) mux.HandleFunc("GET /host/metrics", s.withGuest(s.handleHostMetrics))
// R-856 (`09` §3 decision 143): the host crash guard's record of the last HOST boot — the controller
// waits ~15 min with app mails after an unclean one. Read-only; a missing/garbled file = present:false.
mux.HandleFunc("GET /host/crash-guard", s.withGuest(s.handleCrashGuard))
// Disk management (slice 8C) — self-scoped; format routes through the data-bearing classifier+gate. // Disk management (slice 8C) — self-scoped; format routes through the data-bearing classifier+gate.
mux.HandleFunc("GET /disks", s.withGuest(s.handleDisks)) mux.HandleFunc("GET /disks", s.withGuest(s.handleDisks))
mux.HandleFunc("GET /disks/candidates", s.withGuest(s.handleDiskCandidates)) mux.HandleFunc("GET /disks/candidates", s.withGuest(s.handleDiskCandidates))
+31 -17
View File
@@ -154,12 +154,15 @@ func (s *TokenStore) Mint(vmid int) (string, error) {
// looks it up; the per-candidate comparison is constant-time to avoid a timing oracle on the // looks it up; the per-candidate comparison is constant-time to avoid a timing oracle on the
// stored hash. ok is false for an unknown/empty token. // stored hash. ok is false for an unknown/empty token.
// //
// Reload-on-miss (B3): the store FILE is shared across processes — the one-shot provisioner // Reload-on-change (B3, R-269): the store FILE is shared across processes — the one-shot
// (`--selftest=provision`) Mints into it while the long-lived daemon serves Lookup from an index // provisioner (`--selftest=provision`) Mints into it while the long-lived daemon serves Lookup from
// built at open. On a miss, re-read the file ONCE and re-check, so a token minted after this // an index built at open. Every Lookup stats the file first and re-reads it when the append-only
// process started authorizes without a daemon restart (the drill's fresh-install 401). The // log has grown, BEFORE answering — so a token minted elsewhere authorizes without a restart AND a
// append-only log makes an unchanged file size proof of no new records, so a genuinely unknown // token rotated out elsewhere stops authorizing on its very next presentation. (Before R-269 the
// token costs at most one stat once the index is current — never a reload loop. // re-read ran only on a MISS, so a superseded token was a direct map hit and kept authorizing until
// some unrelated miss forced the reload.) An unchanged size is proof of no new records, so the
// steady state costs one stat per call and never a reload loop. Pinned by
// TestTokenStore_RotatedOutTokenRejectedFirst.
func (s *TokenStore) Lookup(token string) (int, bool) { func (s *TokenStore) Lookup(token string) (int, bool) {
if token == "" { if token == "" {
return 0, false return 0, false
@@ -167,23 +170,34 @@ func (s *TokenStore) Lookup(token string) (int, bool) {
want := hashToken(token) want := hashToken(token)
s.mu.Lock() s.mu.Lock()
defer s.mu.Unlock() defer s.mu.Unlock()
// Direct map hit is the common path; the constant-time compare guards against a timing st, statErr := os.Stat(s.path)
// side-channel by re-checking the matched key (map lookup itself is not the secret-bearing if statErr == nil && st.Size() != s.loadedSize {
// comparison — the hash of a random 256-bit token is not feasibly guessable regardless). // The log changed under us (another process minted/rotated): converge first, then answer.
if vmid, ok := s.byHash[want]; ok { s.reloads++
if subtle.ConstantTimeCompare([]byte(want), []byte(s.byVMID[vmid])) == 1 { if err := s.reloadLocked(); err != nil {
return vmid, true return 0, false // unreadable store: fail closed, never crash the auth path
} }
return s.matchLocked(want)
} }
// Miss: skip the re-read when the append-only log has not grown (nothing new to see). if vmid, ok := s.matchLocked(want); ok {
// A stat error falls through to the reload, which handles a missing file as empty. return vmid, true
if st, err := os.Stat(s.path); err == nil && st.Size() == s.loadedSize {
return 0, false
} }
if statErr == nil {
return 0, false // file unchanged since the last (re)load: genuinely unknown
}
// Stat failed (e.g. the file vanished): reload, which treats a missing file as empty.
s.reloads++ s.reloads++
if err := s.reloadLocked(); err != nil { if err := s.reloadLocked(); err != nil {
return 0, false // unreadable store: fail closed, never crash the auth path return 0, false
} }
return s.matchLocked(want)
}
// matchLocked answers from the in-memory index. Direct map hit is the common path; the
// constant-time compare re-checks the matched key against the guest's CURRENT hash (map lookup
// itself is not the secret-bearing comparison — the hash of a random 256-bit token is not
// feasibly guessable regardless). Caller holds the mutex.
func (s *TokenStore) matchLocked(want string) (int, bool) {
if vmid, ok := s.byHash[want]; ok { if vmid, ok := s.byHash[want]; ok {
if subtle.ConstantTimeCompare([]byte(want), []byte(s.byVMID[vmid])) == 1 { if subtle.ConstantTimeCompare([]byte(want), []byte(s.byVMID[vmid])) == 1 {
return vmid, true return vmid, true
+45
View File
@@ -210,6 +210,51 @@ func TestTokenStore_ReloadOnMiss_RemintCoherence(t *testing.T) {
} }
} }
// R-269: a token rotated out by ANOTHER process must stop authorizing on its very next
// presentation — with NO intervening lookup of the new token. This is the order the operator hits
// after rotating a leaked token: the leaked one is presented first. RemintCoherence above looks the
// NEW token up first, and that miss is what used to evict the old hash, so it passed while the leaked
// token kept returning HTTP 200 on hardware (2026-08-09) until something unrelated forced a reload.
//
// RED-PROOF: restore the reload-on-MISS-only Lookup (answer a map hit before stat-ing the file) and
// this fails with "rotated-out token still authorizes".
func TestTokenStore_RotatedOutTokenRejectedFirst(t *testing.T) {
path := filepath.Join(t.TempDir(), "tokens.log")
daemon, err := OpenTokenStore(path)
if err != nil {
t.Fatalf("open daemon store: %v", err)
}
defer daemon.Close()
minter, err := OpenTokenStore(path)
if err != nil {
t.Fatalf("open minter store: %v", err)
}
defer minter.Close()
old, err := minter.Mint(130)
if err != nil {
t.Fatalf("mint old: %v", err)
}
if vmid, ok := daemon.Lookup(old); !ok || vmid != 130 { // the daemon has learned the old token
t.Fatalf("old token before rotation: (%d,%v), want (130,true)", vmid, ok)
}
fresh, err := minter.Mint(130) // rotation, written by another process
if err != nil {
t.Fatalf("mint fresh: %v", err)
}
if vmid, ok := daemon.Lookup(old); ok { // the leaked token FIRST
t.Fatalf("rotated-out token still authorizes vmid %d on its first presentation after rotation — "+
"Mint's 'any previous token for this guest is revoked' is false across processes (R-269)", vmid)
}
if vmid, ok := daemon.Lookup(fresh); !ok || vmid != 130 {
t.Fatalf("fresh token after rotation: (%d,%v), want (130,true)", vmid, ok)
}
if vmid, ok := daemon.Lookup(old); ok {
t.Fatalf("rotated-out token authorizes vmid %d after the fresh one was seen", vmid)
}
}
// §8 edge: the store file deleted between open and a miss — reload treats it as empty; Lookup // §8 edge: the store file deleted between open and a miss — reload treats it as empty; Lookup
// fails closed, no crash. // fails closed, no crash.
func TestTokenStore_ReloadOnMiss_MissingFile(t *testing.T) { func TestTokenStore_ReloadOnMiss_MissingFile(t *testing.T) {
+2 -2
View File
@@ -60,8 +60,8 @@ func TestLiveReporter_CoordPresentWithoutPriorVerify(t *testing.T) {
if h.PBS == nil { if h.PBS == nil {
t.Fatal("pbs coord absent despite a reachable PBS — the gap this fixes") t.Fatal("pbs coord absent despite a reachable PBS — the gap this fixes")
} }
if h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != "root" || h.PBS.LatestSnapshotID != "9201" { if h.PBS.RepoID != "felhom-pbs" || h.PBS.Namespace != hub.PBSRootNamespace || h.PBS.LatestSnapshotID != "9201" {
t.Errorf("pbs coord = %+v, want felhom-pbs/root/9201", h.PBS) t.Errorf("pbs coord = %+v, want felhom-pbs, the root namespace (\"\", R-124), 9201", h.PBS)
} }
// COMPANION (pre-fix): the bare SnapshotStore (no live read) with an empty store omits pbs. // COMPANION (pre-fix): the bare SnapshotStore (no live read) with an empty store omits pbs.
+6
View File
@@ -57,6 +57,10 @@ type DeviceProbe struct {
HasPartitions bool `json:"has_partitions"` // child partitions present (lsblk) HasPartitions bool `json:"has_partitions"` // child partitions present (lsblk)
Mounted bool `json:"mounted"` // currently mounted somewhere Mounted bool `json:"mounted"` // currently mounted somewhere
FSType string `json:"fstype,omitempty"` FSType string `json:"fstype,omitempty"`
// FSUUID is the filesystem UUID blkid read from the on-disk superblock (`blkid -p`, no cache),
// "" when there is none. R-25: the format path reports it so the caller mounts the filesystem the
// agent just made, not whatever a /dev path resolves to later.
FSUUID string `json:"fs_uuid,omitempty"`
} }
// DataBearing is the conservative verdict: any signature / partition table / partition / mount — // DataBearing is the conservative verdict: any signature / partition table / partition / mount —
@@ -423,6 +427,8 @@ func (h *SudoHostOps) InspectDevice(ctx context.Context, device string) (DeviceP
probe.FSType = v probe.FSType = v
case "PTTYPE": case "PTTYPE":
probe.HasPartitionTable = true probe.HasPartitionTable = true
case "UUID":
probe.FSUUID = v
case "USAGE": case "USAGE":
if v != "" { if v != "" {
probe.HasFilesystem = true // filesystem/raid/crypto member = data-bearing probe.HasFilesystem = true // filesystem/raid/crypto member = data-bearing
+14
View File
@@ -208,3 +208,17 @@ func TestFormat_RejectsBadArgs(t *testing.T) {
t.Fatalf("mkfs ran despite invalid input: %v", r.calls) t.Fatalf("mkfs ran despite invalid input: %v", r.calls)
} }
} }
// R-25: the probe carries the superblock's filesystem UUID so the format path can report the new one.
func TestInspect_ReadsFilesystemUUID(t *testing.T) {
r := &scriptedRunner{
outputs: map[string][]byte{
"blkid": []byte("DEVNAME=/dev/sdb\nUUID=0fc63daf-8483-4772-8e79-3d69d8477de4\nTYPE=ext4\nUSAGE=filesystem\n"),
"lsblk": []byte(`{"blockdevices":[{"name":"sdb","fstype":"ext4","pttype":null,"mountpoint":null}]}`),
},
}
p, _ := newSudo(r).InspectDevice(context.Background(), "/dev/sdb")
if p.FSUUID != "0fc63daf-8483-4772-8e79-3d69d8477de4" {
t.Fatalf("FSUUID = %q, want the blkid UUID", p.FSUUID)
}
}
+59
View File
@@ -0,0 +1,59 @@
#!/usr/bin/env python3
"""build-step-bundle.py — the TRANSITION bundle for a release whose bundle ADDS a path (R-880, `11` §5.4.2).
Usage: python3 scripts/build-step-bundle.py <base-bundle.json> <step-version> <out.json> (prints the sha256)
WHY. A box's INSTALLED felhom-os-apply checks every path of an incoming bundle against ITS OWN table (rule R16) — so a
bundle that adds a path (v0.146.1 adds felhom-priv-apply, the guest hook and the shared-parent files, R-861) is refused
by every box still running an older wrapper. The fix is a step: first a bundle the old wrapper accepts that brings ONLY
the new felhom-os-apply (the new table), then the release's own bundle, which the new wrapper accepts.
WHAT IT BUILDS. <base-bundle.json> is the bundle the boxes run now (download it from the package registry, e.g.
felhom-agent/0.145.0/felhom-config-bundle.json, and check its sha against the hub's record). The step bundle is that
bundle with EXACTLY ONE change: the /usr/local/sbin/felhom-os-apply entry's content is replaced by configs/felhom-os-apply
(this tree). Same paths, same modes, same checks, every other byte identical; agent_version is <step-version> (e.g.
0.146.1-step1). The old wrapper verifies it like any bundle (signature, sha, R16, content checks, self-check of the new
wrapper) — nothing about the trust route changes.
Pinned by configs/test_felhom_config_bundle.py (StepBundle): same paths as the base, only the wrapper differs, the
new wrapper's table is a superset of the base's paths.
"""
import base64
import hashlib
import json
import pathlib
import re
import sys
REPO = pathlib.Path(__file__).resolve().parent.parent
OSAPPLY_DEST = "/usr/local/sbin/felhom-os-apply"
def build_step(base_bytes, version, new_osapply_bytes):
if not re.match(r"^[0-9]+\.[0-9]+\.[0-9]+-[0-9A-Za-z.]+$", version):
raise SystemExit(f"build-step-bundle: {version!r} must be a semver with a step suffix, e.g. 0.146.1-step1")
base = json.loads(base_bytes)
files = base.get("files")
if base.get("format") != 1 or not isinstance(files, list):
raise SystemExit("build-step-bundle: the base is not a format-1 bundle")
hit = [e for e in files if e.get("path") == OSAPPLY_DEST]
if len(hit) != 1:
raise SystemExit(f"build-step-bundle: the base has {len(hit)} {OSAPPLY_DEST} entries, want exactly 1")
hit[0]["content_b64"] = base64.b64encode(new_osapply_bytes).decode()
hit[0]["sha256"] = hashlib.sha256(new_osapply_bytes).hexdigest()
base["agent_version"] = version
return (json.dumps(base, indent=1, sort_keys=True) + "\n").encode()
def main(argv):
if len(argv) != 4:
print(__doc__, file=sys.stderr)
return 2
data = build_step(pathlib.Path(argv[1]).read_bytes(), argv[2], (REPO / "configs" / "felhom-os-apply").read_bytes())
pathlib.Path(argv[3]).write_bytes(data)
print(hashlib.sha256(data).hexdigest())
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv))
+7 -10
View File
@@ -10,13 +10,11 @@
"CI went red at a commit whose own run had been green the day before, on a true finding that", "CI went red at a commit whose own run had been green the day before, on a true finding that",
"no one could act on. The red will return at the next publish unless the two read one number.", "no one could act on. The red will return at the next publish unless the two read one number.",
"", "",
"HOW THE NUMBER WAS ARRIVED AT — stated honestly, because it is weaker than it looks.", "HOW THE NUMBER WAS ARRIVED AT. generic_versions_kept is 10 because that is what the registry",
"generic_versions_kept is 10 because that is what the registry demonstrably holds today", "keeps: the deleter was ESTABLISHED on 2026-08-10 (R-287) — the newest-10 prune of generic packages",
"(felhom-agent 0.121.0..0.128.0 = 10 versions, queried 2026-08-09). It is an OBSERVED state,", "run under R-267 (felhom-agent and felhom-golden generic held exactly 10 afterwards). Until then this",
"NOT a ruling anyone has been able to locate: no register row records a package prune, R-210", "file called the 10 an observed state with no located ruling; that is superseded (corrected",
"is WAITING-ON-OPERATOR and says 'Nothing was deleted; this is a list, not an action', and it", "2026-10-05, R-291). Container packages are not pruned by that rule.",
"concerns local Docker images rather than this registry. Container packages currently hold 19",
"each, so there is no uniform ten-per-package cap visible either. See R-287.",
"", "",
"SO THIS FILE IS A FLOOR, NOT A LICENCE. It says: CI may assume nothing older than the newest", "SO THIS FILE IS A FLOOR, NOT A LICENCE. It says: CI may assume nothing older than the newest",
"N generic versions is still downloadable. It does NOT authorise deleting anything, and the", "N generic versions is still downloadable. It does NOT authorise deleting anything, and the",
@@ -36,9 +34,8 @@
], ],
"generic_versions_kept": 10, "generic_versions_kept": 10,
"readers": [ "readers": [
"scripts/check-published-versions.py — bounds its assertion to the newest N versions", "scripts/check-published-versions.py — bounds its assertion to the newest N versions"
"documentation/runbooks/registry-retention.md (felhom.eu) — the prune procedure"
], ],
"recorded": "2026-08-09", "recorded": "2026-08-09",
"recorded_by": "CC, from the registry's observed state; NOT from a located operator ruling" "recorded_by": "CC 2026-08-09; the number's source (the R-267 newest-10 prune, established 2026-08-10 by R-287) recorded 2026-10-05 (R-291)"
} }
+367
View File
@@ -0,0 +1,367 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""test_gate_decoys.py — can this repo's gates be fooled by a LABEL? (R-421, R-426)
The same instrument as `felhom.eu/scripts/test_gate_decoys.py`: a decoy is the LABEL without the
FACT, and a gate that passes on the label alone — or refuses the genuine article — is a live hole.
Every gate is asserted in BOTH directions: the decoy must be convicted, the genuine article passed.
Covered here (the `COVERS` literal is AST-read by `felhom.eu/scripts/decoy_coverage_gate.py`, which
never imports this file):
published check-published-versions.py, against a FAKE Gitea (see below).
release-complete check-release-complete.py, in a scratch clone whose `origin` is a scratch bare
repository, against the same fake Gitea.
reuse-refs, instructions, observations
the three SHARED felhom.eu scripts, run against a scratch clone of THIS repo —
so the decoy is planted in the agent's own REUSE.md / CLAUDE.md / REPORT.md and
coverage is per input, not per script.
NEVER THE REAL GITEA. Both network gates read `GITEA_BASE` from the environment (CI already sets it
to the in-cluster URL); here it points at an `http.server` bound to 127.0.0.1 inside this process,
and every proxy variable is removed from the child's environment so urllib cannot route around it.
A test that asked the real registry would pass or fail on whatever was published that day — the
constant-for-measurement shape — and would reach the network from a hook.
NEVER THE REAL TREE. Every planted file lives in a scratch directory: a workspace that holds a
clone of this repo beside symlinks to the sibling clones the shared scripts reach across to.
Run from the repo root: python3 scripts/test_gate_decoys.py
Exit 0 every decoy judged correctly · 1 a decoy passed or a genuine article was refused.
"""
import http.server
import io
import json
import os
import shutil
import socketserver
import subprocess
import sys
import tempfile
import threading
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
PARENT = os.path.dirname(ROOT)
# DECOY_SHARED_DIR exists for ONE purpose: the red-proof. It lets a mutated COPY of the shared scripts
# be judged without editing the felhom.eu clone. Unset, the suite judges the real shared scripts.
SHARED = os.environ.get("DECOY_SHARED_DIR") or os.path.join(PARENT, "felhom.eu", "scripts")
# ── WHAT THIS FILE COVERS ────────────────────────────────────────────────────────────────────────
# Read by felhom.eu/scripts/decoy_coverage_gate.py, which AST-parses this literal. A gate named here
# MUST have a decoy below that has been seen to fail.
COVERS = {
"published": "against a FAKE Gitea: a tag whose package 404s, a tag whose tree lacks the configs, a package one "
"version past the newest tag (published, never tagged), a patch-gap orphan, a missing package that "
"lexical sorting would drop out of the retention window (0.9.x vs 0.10.x); a tags api answering 500 "
"or a non-JSON 200 is INCONCLUSIVE, never a pass - vs a clean registry, a non-semver tag and a version "
"older than the retention window (not asserted, BY DESIGN) (R-426)",
"release-complete": "the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated commit, a tag "
"with no package; a registry 500 is INCONCLUSIVE - vs the genuine release, an `## Unreleased` "
"heading above it, a newer version named only in prose or under `###`, a tag that only "
"origin has (the shallow-CI shape); a LOCAL-only tag passes BY DESIGN (CI's fresh clone "
"and the published gate's converse probe are what see it) (R-426)",
"reuse-refs": "a cited .go and a cited .md path that do not exist, planted in THIS repo's REUSE.md - vs the "
"real file (R-426)",
"instructions": "a component version literal in THIS repo's CLAUDE.md effective text - vs the same sentence "
"inside an HTML comment (R-426)",
"observations": "R-419 in THIS repo's REPORT.md: an Observations note SAYING it carries no marker - vs the "
"two genuine markers (R-426)",
}
fails = []
ran = 0
def report(name, rc, out, expect_rc, must=()):
global ran
ran += 1
missing = [m for m in must if m not in out]
if rc == expect_rc and not missing:
print(" ok %-62s rc=%d (expected %d)" % (name, rc, expect_rc))
else:
hole = expect_rc != 0 and rc == 0
fails.append("%s: rc=%d expected %d%s; missing %s\n%s" % (
name, rc, expect_rc, " - LIVE HOLE" if hole else "", missing, out[-900:]))
# ── the fake Gitea ───────────────────────────────────────────────────────────────────────────────
class Fake(object):
"""What the fake registry serves. Reset per case."""
def reset(self):
self.tags = [] # tag names, as the tags api lists them
self.packages = set() # versions whose generic package downloads
self.raw = set() # versions whose tag tree serves the probe config
self.tags_status = 200
self.tags_body = None # override bytes for the tags api
self.pkg_status = None # override status for EVERY package request
self.hits = []
FAKE = Fake()
FAKE.reset()
PKG_PREFIX = "/api/packages/admin/generic/felhom-agent/"
RAW_PREFIX = "/admin/felhom-agent/raw/tag/v"
class Handler(http.server.BaseHTTPRequestHandler):
def log_message(self, *a):
pass
def _answer(self, status, body=b""):
self.send_response(status)
self.send_header("Content-Length", str(len(body)))
self.end_headers()
if self.command != "HEAD":
self.wfile.write(body)
def do_GET(self):
p = self.path
FAKE.hits.append(p)
if p.startswith("/api/v1/repos/admin/felhom-agent/tags"):
body = FAKE.tags_body if FAKE.tags_body is not None else \
json.dumps([{"name": t} for t in FAKE.tags]).encode()
return self._answer(FAKE.tags_status, body)
if p.startswith(PKG_PREFIX):
if FAKE.pkg_status is not None:
return self._answer(FAKE.pkg_status)
v = p[len(PKG_PREFIX):].split("/", 1)[0]
return self._answer(200, b"ELF") if v in FAKE.packages else self._answer(404)
if p.startswith(RAW_PREFIX):
v = p[len(RAW_PREFIX):].split("/", 1)[0]
ok = v in FAKE.raw and p.endswith("/configs/felhom-agent.service")
return self._answer(200, b"[Unit]\n") if ok else self._answer(404)
return self._answer(404)
do_HEAD = do_GET
class Server(socketserver.ThreadingMixIn, http.server.HTTPServer):
daemon_threads = True
def child_env(base):
env = {k: v for k, v in os.environ.items() if "proxy" not in k.lower()}
env["GITEA_BASE"] = base
env["NO_PROXY"] = env["no_proxy"] = "127.0.0.1,localhost"
return env
def run(argv, cwd, env=None):
# input="" — a child must never inherit (and block on) this process's stdin
p = subprocess.run(argv, cwd=cwd, env=env, capture_output=True, text=True, input="")
return p.returncode, p.stdout + p.stderr
def sh(argv, cwd):
rc, out = run(argv, cwd)
if rc != 0:
raise SystemExit("setup command failed (%s): %s" % (" ".join(argv), out))
return out.strip()
# ── published ────────────────────────────────────────────────────────────────────────────────────
def published_cases(base):
gate = os.path.join(ROOT, "scripts", "check-published-versions.py")
keep = json.load(io.open(os.path.join(ROOT, "scripts", "retention-policy.json"),
encoding="utf-8"))["generic_versions_kept"]
TAGS = ["v0.150.%d" % i for i in range(3)] # inside any retention window >= 3
VERS = [t[1:] for t in TAGS]
def case(name, setup, expect_rc, must=()):
FAKE.reset()
FAKE.tags = list(TAGS)
FAKE.packages = set(VERS)
FAKE.raw = set(VERS)
setup()
rc, out = run([sys.executable, gate], ROOT, child_env(base))
if not FAKE.hits:
fails.append("published/%s: the gate never asked the fake Gitea - the seam is not wired" % name)
report("published: " + name, rc, out, expect_rc, must)
case("GENUINE: every tag downloadable and serving its configs", lambda: None, 0,
("ALL RELEASED VERSIONS INSTALLABLE",))
case("GENUINE: a non-semver tag is not a release", lambda: FAKE.tags.append("v0.150.2-rc1"), 0,
("ALL RELEASED VERSIONS INSTALLABLE",))
case("FACT: a tag whose package 404s", lambda: FAKE.packages.discard("0.150.1"), 1,
("FAIL v0.150.1", "binary NOT downloadable"))
case("FACT: a tag whose tree does not serve the configs", lambda: FAKE.raw.discard("0.150.2"), 1,
("FAIL v0.150.2", "does not serve"))
case("FACT: published one patch past the newest tag, never tagged",
lambda: FAKE.packages.add("0.150.3"), 1, ("PUBLISHED VERSION(S) WITH NO TAG", "v0.150.3"))
case("FACT: published in a patch GAP between two tags",
lambda: (FAKE.tags.remove("v0.150.1"),), 1, ("v0.150.1 is downloadable", "has no git tag"))
def lexical():
# keep+1 tags: 0.9.0 and 0.10.0..0.10.<keep-1>. By SEMVER the oldest is 0.9.0 (dropped); by
# STRING sort "0.10.0" is the smallest and would be the one dropped - so its missing package
# is convicted only if the window is cut by semver.
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.10.0"}
FAKE.raw = set(t[1:] for t in FAKE.tags)
case("FACT: a missing package lexical sorting would drop (0.10.0 vs 0.9.0)", lexical, 1,
("FAIL v0.10.0",))
def retired():
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.9.0"}
FAKE.raw = set(t[1:] for t in FAKE.tags)
case("BY DESIGN: a version older than the retention window is not asserted", retired, 0,
("NOT ASSERTED", "0.9.0"))
def five_hundred():
FAKE.tags_status = 500
case("INCONCLUSIVE: the tags api answers 500", five_hundred, 2, ("INCONCLUSIVE",))
def html():
FAKE.tags_body = b"<html>sign in</html>"
case("INCONCLUSIVE: the tags api answers a 200 that is not JSON", html, 2, ("INCONCLUSIVE",))
# an unreachable Gitea: a port nothing listens on
s = Server(("127.0.0.1", 0), Handler)
dead = "http://127.0.0.1:%d" % s.server_address[1]
s.server_close()
rc, out = run([sys.executable, gate], ROOT, child_env(dead))
report("published: INCONCLUSIVE: Gitea unreachable", rc, out, 2, ("INCONCLUSIVE", "URLs tried"))
# ── release-complete ─────────────────────────────────────────────────────────────────────────────
def release_cases(base, ws):
bare = os.path.join(ws, "origin.git")
work = os.path.join(ws, "rc-work")
sh(["git", "clone", "-q", "--bare", "--no-tags", "file://" + ROOT, bare], ws)
sh(["git", "clone", "-q", "--no-tags", "file://" + bare, work], ws)
sh(["git", "config", "user.email", "decoy@gate.invalid"], work)
sh(["git", "config", "user.name", "decoy"], work)
# the WORKING-TREE gate, so the file under test is the one being edited, not HEAD's
shutil.copy(os.path.join(ROOT, "scripts", "check-release-complete.py"),
os.path.join(work, "scripts", "check-release-complete.py"))
sh(["git", "add", "scripts/check-release-complete.py"], work)
sh(["git", "commit", "-q", "--allow-empty", "-m", "the gate under test"], work)
base_sha = sh(["git", "rev-parse", "HEAD"], work)
ch = os.path.join(work, "CHANGELOG.md")
original = io.open(ch, encoding="utf-8").read()
V = "9.9.9"
def case(name, top, expect_rc, must=(), tag=None, origin_tag=False, packaged=True, pkg_status=None):
FAKE.reset()
if packaged:
FAKE.packages = {V}
FAKE.pkg_status = pkg_status
try:
io.open(ch, "w", encoding="utf-8").write(top + original)
sh(["git", "commit", "-q", "-am", name], work)
if tag == "head":
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", "HEAD"], work)
elif tag == "unrelated":
empty = sh(["git", "mktree"], work) # stdin is "" — the empty tree, written to this repo
orphan = sh(["git", "commit-tree", "-m", "unrelated", empty], work)
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", orphan], work)
if origin_tag:
sh(["git", "push", "-q", "origin", "HEAD:refs/tags/v" + V], work)
rc, out = run([sys.executable, os.path.join(work, "scripts", "check-release-complete.py")],
work, child_env(base))
report("release-complete: " + name, rc, out, expect_rc, must)
finally:
run(["git", "tag", "-d", "v" + V], work)
run(["git", "push", "-q", "origin", ":refs/tags/v" + V], work)
sh(["git", "reset", "-q", "--hard", base_sha], work)
HEAD = "## v%s — 2026-10-06\n\n- decoy release\n\n" % V
case("GENUINE: tagged at HEAD and published", HEAD, 0,
("newest CHANGELOG version: v9.9.9", "is tagged, placed and published"), tag="head")
case("GENUINE: an `## Unreleased` heading above the release", "## Unreleased\n\n- wip\n\n" + HEAD, 0,
("newest CHANGELOG version: v9.9.9",), tag="head")
case("GENUINE: a newer version named only in prose and under ###",
"The `## v10.0.0` heading is not written yet.\n### v10.0.0 notes\n\n" + HEAD, 0,
("newest CHANGELOG version: v9.9.9",), tag="head")
case("GENUINE: the tag only on origin (the shallow-CI shape)", HEAD, 0,
("exists on origin",), origin_tag=True)
case("BY DESIGN: a LOCAL-only tag passes (CI's fresh clone sees only origin)", HEAD, 0,
("an ancestor of HEAD",), tag="head")
case("FACT: the newest heading has no tag anywhere", HEAD, 1, ("DOES NOT EXIST",))
case("FACT: a tag parked on an unrelated commit", HEAD, 1, ("NOT an ancestor",), tag="unrelated")
case("FACT: tagged, never published", HEAD, 1, ("IS NOT PUBLISHED",), tag="head", packaged=False)
case("INCONCLUSIVE: the registry answers 500", HEAD, 2, ("INCONCLUSIVE",), tag="head", pkg_status=500)
case("FACT beats INCONCLUSIVE: no tag AND the registry answers 500", HEAD, 1, ("DOES NOT EXIST",),
pkg_status=500)
# ── the shared felhom.eu scripts, against THIS repo's inputs ─────────────────────────────────────
def shared_cases(ws):
"""A scratch WORKSPACE: a clone of this repo beside symlinks to the siblings, because the shared
scripts reach across (REUSE.md cites hub paths; instructions_gate reads the workspace CLAUDE.md)."""
for g in ("reuse_refs_check.py", "instructions_gate.py", "observations_gate.py"):
if not os.path.isfile(os.path.join(SHARED, g)):
fails.append("shared gate %s is MISSING beside this clone (tried %s) - a failure, never a skip"
% (g, SHARED))
return
space = os.path.join(ws, "workspace")
os.makedirs(space)
for entry in sorted(os.listdir(PARENT)):
if entry in ("felhom.eu", "felhom-controller", "app-catalog-felhom.eu", "homelab-manifests",
"CLAUDE.md", ".claude-memory"):
os.symlink(os.path.join(PARENT, entry), os.path.join(space, entry))
repo = os.path.join(space, "felhom-agent")
sh(["git", "clone", "-q", "--no-tags", "file://" + ROOT, repo], ws)
# the WORKING-TREE inputs the plants go into, so a case judges today's file
for f in ("REUSE.md", "CLAUDE.md", "REPORT.md"):
shutil.copy(os.path.join(ROOT, f), os.path.join(repo, f))
def case(name, gate, relpath, extra, expect_rc, must=()):
p = os.path.join(repo, relpath)
backup = io.open(p, encoding="utf-8").read()
try:
if extra:
io.open(p, "w", encoding="utf-8").write(backup + extra)
rc, out = run([sys.executable, os.path.join(SHARED, gate), repo], repo)
report(name, rc, out, expect_rc, must)
finally:
io.open(p, "w", encoding="utf-8").write(backup)
case("reuse-refs: GENUINE: this repo's REUSE.md", "reuse_refs_check.py", "REUSE.md", "", 0, ("FAILED 0",))
case("reuse-refs: FACT: a cited .go path that does not exist", "reuse_refs_check.py", "REUSE.md",
u"\n- see `internal/localapi/does_not_exist.go`\n", 1, ("does_not_exist.go",))
case("reuse-refs: FACT: a cited .md path that does not exist", "reuse_refs_check.py", "REUSE.md",
u"\n- see `docs/99-does-not-exist.md`\n", 1, ("99-does-not-exist.md",))
case("instructions: GENUINE: this repo's CLAUDE.md", "instructions_gate.py", "CLAUDE.md", "", 0,
("instructions_gate: OK",))
case("instructions: FACT: a version literal in effective text", "instructions_gate.py", "CLAUDE.md",
u"\nThe agent runs v0.148.0 today.\n", 1, ("v0.148.0",))
case("instructions: GENUINE: the same sentence in an HTML comment", "instructions_gate.py", "CLAUDE.md",
u"\n<!--\nThe agent ran v0.148.0 on 2026-10-06.\n-->\n", 0, ("instructions_gate: OK",))
case("observations: FACT: R-419, prose SAYING it has no marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** It carries no `FILED:` marker and no "
u"`NOT-A-FINDING:` marker, deliberately.\n", 1)
case("observations: GENUINE: a FILED marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** Something broke. **FILED: R-419**\n", 0)
case("observations: GENUINE: a NOT-A-FINDING marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** Odd. **NOT-A-FINDING: my own typo, corrected in "
u"the same minute.**\n", 0)
def main():
srv = Server(("127.0.0.1", 0), Handler)
threading.Thread(target=srv.serve_forever, daemon=True).start()
base = "http://127.0.0.1:%d" % srv.server_address[1]
ws = tempfile.mkdtemp(prefix="agent-decoys-")
print("agent gate decoys — fake Gitea at %s, scratch %s" % (base, ws))
try:
published_cases(base)
release_cases(base, ws)
shared_cases(ws)
finally:
srv.shutdown()
srv.server_close()
shutil.rmtree(ws, ignore_errors=True)
if fails:
print()
for f in fails:
print("FAIL: %s" % f)
return 1
print("\nagent gate decoys OK — %d case(s), every label judged on its fact (R-421)" % ran)
return 0
if __name__ == "__main__":
sys.exit(main())