Compare commits

...

28 Commits

Author SHA1 Message Date
admin dd7cdc09e7 R-366 slice 2 (decision 168): the host report carries the archives the restore-test skipped as another key's
gates / gates (push) Successful in 1m6s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin 7b0a8b234b R-105 (decision 169): remove the escrow-create -directive flag and the upload's directive field
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin ce1a4b4758 R-812 option A: the Proxmox package lane (layer pve) + the /etc/pve write gate
The wrapper gains layer "pve" (slow lane): the host's Proxmox userspace
packages only — origin "Proxmox Debian Repository", never a kernel / boot /
firmware / microcode name (R14), no removal, no undo, a new package only from
an allow-list; authority = a signed os_pve_step or the root-owned ring-0 mark.
The night leg runs it in ring 0 after a healthy host step; ring 1 only by a
signed job (PVEStepExecutor). While it runs, the agent's own /etc/pve writes
(every non-GET API call, pct config verbs, pvesm, pveum, felhom-pbs-apply)
wait on internal/pvegate. Health = the host rule + unchanged container ids +
pveversion reads the installed pve-manager.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin ac90169a5d R-861 (a) A1 + (b) B2: the image ref goes to a root verb that checks it; the agent's in-guest tee grant is gone; felhom-op's pct lines are exact (09 §3 decision 165)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:28 +02:00
admin 154d6dcaa9 Shared rule file: no hub image build or deploy in a session the operator does not attend (09 §3 decision 162)
gates / gates (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 09:55:02 +02:00
admin 2f7072050f REPORT: released and delivered (2026-10-07)
gates / gates (push) Successful in 1m4s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 09:31:58 +02:00
admin 85e799f360 CHANGELOG: v0.150.0 released (R-528, R-894, R-330)
gates / gates (push) Successful in 46s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 08:59:08 +02:00
admin 3a72a4811b Shared rule file: rule 11 — every helper prompt carries the brief's fences in full (09 §3 decision 160)
gates / gates (push) Successful in 54s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 08:52:36 +02:00
admin a165d53c88 REPORT: the second burn-down night
gates / gates (push) Successful in 49s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 00:14:50 +02:00
admin adaf86ad57 R-330: the agent sends SMART 187/188/199 (raw; omitted when unknown)
gates / gates (push) Successful in 53s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 22:13:04 +02:00
admin 74b5eae5b0 R-894: after a restart the agent remembers the last backup per tier
gates / gates (push) Successful in 35s
An unreadable storage right after an agent restart read the off-site tier
DUE (the in-memory record was empty). The newest success per tier is now
kept on disk and read ONLY when the storage cannot be read: fresh -> not
due, older than the cadence -> due, none -> due (unknown) as before. A
storage that answers stays the ground truth.

Ships with v0.150.0 after the 2026-10-07 read-back; nothing delivered.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 20:33:27 +02:00
admin 7e82f325b8 CHANGELOG: the memory-kill check (unreleased, to ship as v0.150.0 after the night read-back)
gates / gates (push) Successful in 50s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:36:09 +02:00
admin acccb66bd3 R-528 (09 decision 157): after a Docker engine step the wrapper proves the engine reports a memory kill
felhom-os-apply: a docker-layer apply runs oom_check() after health_after and reports
"oom_check": {result pass|fail|error, oom_killed, oom_event, exit_code, image, detail}.
One throwaway container (the controller's image, --pull never, --network none, 64m cap,
label felhom.oomcheck=1) runs dd bs=200M; pass only with OOMKilled=true AND the oom event
(read after a 2 s settle, --until = guest epoch + 1: measured on demo-hp, an --until taken
right after the run missed the event). docker rm -f always runs in a finally; every call
is bounded (<= 90 s). It never changes the step's outcome or health. New wrapper-only mode
"oom-check" (docker layer) runs the check alone; check_guest etc. still apply.

Agent: WrapperReport/Report gain OOMCheck (json:"oom_check"), copied unchanged in runLayer
and in the kept-copy path.

Tests: 9 wrapper tests + 2 Go tests, each red-proofed (audits/readback-2026-10-07/F/red-*.txt).
Also: test_felhom_os_apply.py's `if __name__` sat mid-file, so 11 tests (UnsentReport,
SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran as a script or from
TestWrapperSuite; moved to the end (they pass).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:34:55 +02:00
admin de812bc027 The shared rule file (09 §3 decision 152), identical to the other copies; no code change
gates / gates (push) Successful in 44s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 16:06:46 +02:00
admin 3e8ebeb96c Instruction files kept true (09 §3 decision 150): stale gate lists, paths and facts corrected; no code change
gates / gates (push) Successful in 46s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 13:43:57 +02:00
admin cefdc731a4 REPORT: the operator's ten answers (2026-10-06)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 12:28:58 +02:00
admin e56dcb8a4c CHANGELOG/v0.149.0 released (shas)
gates / gates (push) Successful in 37s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:50:21 +02:00
admin f277e619e2 CHANGELOG: unreleased — R-856 GET /host/crash-guard
gates / gates (push) Successful in 38s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:40 +02:00
admin 386f51edc6 R-856: GET /host/crash-guard — the host crash guard's last-boot record for the controller (09 decision 143)
The controller waits ~15 minutes with app mails after a crash boot of the host; it learns of the
crash boot from this route. Reads /var/lib/felhom-crash-guard/state.json (read-only, no Proxmox call)
and passes present/last_boot_at/last_boot_unclean/tripped through; a missing, unreadable or garbled
file answers 200 present:false. Guest-token authed like every sibling route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:01 +02:00
admin 130e3ed882 CHANGELOG: unreleased — R-444 weekly guest disk trim, R-99 runbook pointer
gates / gates (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:25:22 +02:00
admin be398f92e8 R-99: the PBS phantom WARN names the cleanup runbook (09 §3 decision 140)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin ee71abd1d4 R-444: weekly guest disk trim (pct fstrim) outside the night, under the heavy-op gate
Operator ruling 09 §3 decision 139. One exact sudoers rule FELHOM_FSTRIM
(`/usr/sbin/pct ^fstrim [0-9]+$`) + manifest entry guest-fstrim; new
internal/fstrim job: due Wednesday from 10:00 host-local, starts only
10:00-20:59, holds backup.InFlight (busy -> deferred to the next hourly
tick), failed trim retried at most 3x per week, bytes parsed from the
"(N bytes) trimmed" lines, last result per guest persisted in
<state_dir>/guest-disk-trim.json and reported as guest_disk_trim.
Config opt-out: "disk_trim": {"disable": true}.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin 37e98f452b REPORT: the burn-down night (2026-10-06)
gates / gates (push) Successful in 45s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 02:13:10 +02:00
admin b2b82ae828 bundle test: the ISO first-boot files the installer names under KEPT are not installer-written (go test red since felhom.eu 85de3f9b); CHANGELOG unreleased (R-426 decoys)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:25 +02:00
admin b78a0ff3ac R-426: decoys for the shared reuse-refs, instructions and observations gates
COVERS "reuse-refs", "instructions", "observations": the three shared
felhom.eu scripts run against a scratch clone of THIS repo (in a scratch
workspace symlinking the sibling clones they reach across to), so the
plant is in the agent's own REUSE.md / CLAUDE.md / REPORT.md. Convicted:
a missing cited .go and .md path, a version literal in CLAUDE.md's
effective text, R-419's prose-only Observations note. Passed: the real
files, the version inside an HTML comment, both genuine markers.
DECOY_SHARED_DIR lets a red-proof judge a mutated copy of the shared
scripts without editing the felhom.eu clone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 96047453cb R-426: decoys for the release-complete gate
COVERS "release-complete": the working-tree gate runs in a scratch clone
whose origin is a scratch bare repo, against the fake Gitea. Convicted:
the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated
commit, a tag with no package, and no-tag wins over a registry 500.
Inconclusive: a registry 500. Passed: the genuine release, an
`## Unreleased` heading above it, a newer version named only in prose or
under `###` (the withdrawn sweep decoy, now asserted the right way round),
a tag only origin has (the shallow-CI shape), and a LOCAL-only tag BY
DESIGN (CI's fresh clone and the published gate's converse probe see it).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 4bf5db6875 R-426: decoy suite for the published gate, against a fake Gitea
scripts/test_gate_decoys.py (new; COVERS "published"): an http.server on
127.0.0.1 stands in for Gitea through the gate's existing GITEA_BASE
seam, proxies stripped, so no case reaches the real registry. Facts
convicted: a tag whose package 404s, a tag tree without the configs, a
package one patch past the newest tag (never tagged), a patch-gap orphan,
a missing package that lexical sorting would drop out of the retention
window. Inconclusive, never a pass: tags api 500, a non-JSON 200, Gitea
unreachable. Passed: a clean registry, a non-semver tag, a version older
than the retention window (BY DESIGN).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 87977ff40a CHANGELOG/v0.148.0 released (shas)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 00:37:07 +02:00
58 changed files with 4129 additions and 226 deletions
+3 -2
View File
@@ -21,6 +21,7 @@ fast, and wrong.
This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that
felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before
path-scoped rules existed. The single source is now felhom.eu/CLAUDE.md "Code quality rules"; this
file is the scoped copy that loads exactly where health checks are written. (2026-08-06)
path-scoped rules existed. Deliberate scoped copies now live in felhom.eu/.claude/rules/hub.md and
felhom-controller/.claude/rules/gates.md (hub.md's comment names them); none is the single source. This
file is the copy that loads exactly where agent health checks are written. (2026-08-06; corrected 2026-10-06)
-->
+75
View File
@@ -0,0 +1,75 @@
---
unconditional: true
---
# Unprompted work — rules for any session without a task file
> Goal sessions, nightly sessions, "work the register" sessions. **A session that starts from
> `/goal` or a standing brief inherits these rules exactly as it inherits the gates.** They are the
> part of `PROMPT-TEMPLATE.md` that a task file used to carry and a goal does not. Same wording lives
> in `felhom.eu`, `felhom-controller`, `felhom-agent` and `app-catalog-felhom.eu` `.claude/rules/`, and in the workspace
> root's unversioned `.claude/rules/`; change all five or none.
## 1. What you may pick up on your own
- A register row **you or another CC session filed**, with owner CC, at P3 or a bounded P2, that
needs **no operator decision**, touches **no customer data by design**, and introduces **no
mechanism nobody has measured**. Smallest first.
- A defect you find while exercising the product, filed as a row **before** you fix it — **unless it is small**:
a small finding is fixed in the session and never filed (the size rule, `OPEN-ITEMS.md` „How a row is filed").
- Hygiene: register compression, stale citations, rows with no owner, documents that contradict
live source.
**Not yours, ever, without a task file or an operator word:** money; anything that changes risk to
customer data; anything that changes a promise the product makes to a customer; anything that
reverses a documented design decision (`documentation/architecture/` — a design decision is not a
defect, R-370); anything on DooPlex or ep0; baking or vouching a golden; promoting a
catalog version; a new external dependency; **a hub image build or hub deploy in a session the operator does not
attend** (operator ruling 2026-10-07, `09` §3 decision 162).
## 2. When you may decide instead of ask (operator grant, 2026-09-14)
You may take a decision yourself when **all** of these hold: the architecture folder and the register
give a clear direction; your choice follows that direction; it is reversible without customer-data
risk; and you can write it in the `09-update-architecture.md` §3 shape — one answerable sentence, the
options, what each costs, why this one. **Then record it** as a dated decision in `CONTEXT.md` and
the owning architecture document, tagged *decided by CC unattended — operator may reverse*, and put
it **first** in the morning note. A decision you cannot write in that shape is one you do not take.
## 3. The discipline a task file used to carry
1. **Baselines first.** Read each repo's `main` hash and version from live source before touching it.
2. **Read the architecture document for the area, and name it** in the report, before any claim.
3. **Red-proof every correctness fix.** A test never seen failing has not been shown to test anything.
4. **Live-validate on a Tier-0 box** through the endpoints the UI invokes. `demo-hp` is `ssh hp`.
Throwaway apps only; the standing apps and `bentopdf` stay.
5. **Evidence off the machine at the end of each phase**, before any revert (R-320).
6. **One release per repo per session**, with a CHANGELOG entry (controller: with its `MinAgent`
line), REPORT overwritten, floor raised to deliver it. **No golden unless a drill or fresh install
needs one** (the waiver, R-468). **No `--no-verify`.**
7. **An enumerated gap becomes a row in the same session — or, if it is small, is fixed in it** (the size rule).
Prose is not a record.
8. **Hungarian text is searched with ASCII fragments**, with a positive and a negative control.
9. **Never leave a half-state.** If time runs out, revert to clean and say what was reverted.
10. **Teardown, three layers, stated** — machine, host, hub — or "provisioned nothing".
11. **Every helper prompt carries the brief's fences in full** (operator ruling 2026-10-07). A helper session (a
subagent, a fork, a workflow agent) gets the brief's fence list word for word — every protected machine, every
„no", every delivery and Docker limit — not a summary and not „the usual fences". Earned on 2026-10-06 night: two
helpers whose prompts carried only part of the fences ran `docker volume prune` on the bench and a Docker-using
gate on DooPlex.
## 4. The morning note
One screen, plain language, in this order: **decisions you took** (§2) first; what you exercised;
what broke and whether you fixed it; rows opened and closed with the register size before and after;
what needs the operator, each with what happens if they do nothing. No file paths, no function
names, no row numbers as the subject of a sentence.
## 5. Instruction files
**Instruction files (`CLAUDE.md`, `.claude/rules/*`) are kept true by the session that finds them wrong**
(operator ruling 2026-10-06, `09` §3 decision 150). A session MAY, without asking: correct a stale fact (a command, a
count, a version, a path, a description of what a gate does), add a fact it proved, and remove a reference to something
that no longer exists. Each edit is named in the report (file, line, before, after, why). A session MAY NOT, without the
operator's word: loosen a safety rule, a fence, a „never", a protected machine, a secret rule, or a review step; or
remove a rule. When in doubt, it is a rule change, and it goes to the operator. If Claude Code's own permission check
asks before such an edit, wait for the operator's click; if it refuses, record that and file the exact line.
+95
View File
@@ -1,5 +1,100 @@
## Unreleased (2026-10-07) — the agent can no longer hand the guest any image; felhom-op's pct lines are exact (R-861 (a) A1, (b) B2; `09` §3 decision 165)
**Delivery order: agent binary FIRST, then the config bundle.** The new sudoers drops the agent's in-guest `tee`
grant; an older binary still calls `tee`, so a bundle that lands before the binary would stop managed controller
updates (and the old binary's capability probe would read `controllerswap-write` degraded). No bundle path is added
(`felhom-priv-apply` and `/etc/sudoers.d/felhom-op` are already bundle files), so no step bundle.
- `configs/felhom-priv-apply`: new verb `controller-image <vmid>` — reads the ref on stdin (≤ 256 bytes, ASCII, one
optional trailing newline), requires `^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$` (the agent's
own `controllerImageRe`), then runs `pct exec <vmid> -- tee /etc/felhom-controller-image` AS ROOT; refusal rule `I1`
(rc 3), a bad vmid `A1` (rc 2); listed in `--self-check`.
- `configs/felhom-agent.sudoers` `FELHOM_CONTROLLERSWAP`: `pct ^exec [0-9]+ -- tee /etc/felhom-controller-image$`
REMOVED; `/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$` added.
- `internal/localapi`: `GuestExecutor.GuestExecStdin` replaced by `WriteControllerImage`; `GuestBinder.WriteControllerImage`
pipes `ref\n` to the verb through the fenced runner; the swap's `writeImage` calls it. Capability `controllerswap-write`
now probes the verb.
- `configs/felhom-op.sudoers` (B2, hygiene): `pct start|stop|unlock [0-9]*` → `pct ^start [0-9]+$` etc. (the glob's `*`
matched spaces: `pct stop 9201 --skiplock 1` passed).
- Tests: `ControllerImage` (5, `configs/test_felhom_priv_apply.py`), `TestSudoersRefusesTheR861Injections` (+3 lines),
`TestSudoersAllowsTheControllerImageVerb`, `TestFelhomOpSudoersPctIsExact`, `TestR861_WriteControllerImageUsesTheRootVerb`,
`TestControllerSwap_WriteViaRootVerb_NoShell`. Red-proofs: `felhom.eu/documentation/audits/day-2026-10-07/C/`.
- `README.md`: the controller-swap paragraph described the removed `tee` path — corrected.
## Unreleased (2026-10-07) — the Proxmox package lane (R-812 option A, `09` §3 decision 163)
**MinAgent impact: none** (a new layer; an older hub ignores the pve report). **The bundle carries the new
`felhom-os-apply` — deliver it with the binary** (signed `agent_update`, then signed `agent_config_update`).
- `configs/felhom-os-apply`: new layer `pve`, lane `slow` only — the host's Proxmox USERSPACE packages: origin `Proxmox Debian Repository` only (R2), never a kernel / boot / firmware / microcode name (R14, `HOST_SLOW_RE` — the kernel is R-836's lane), no removal (R4), no undo (R5), a new package only from `PVE_NEW_ALLOW` (`proxmox-firewall-data`, measured on demo-felhom; R6 otherwise), an appliance only (R12), authority = a signed `os_pve_step` or the root-owned ring-0 mark (R3). Select `pending-pve` (ring 0): installed Proxmox-origin packages with a pending upgrade. The report carries `pve_manager` (pveversion after the step).
- `internal/pvegate` (new): the agent's own writes to /etc/pve wait while a pve step runs (pmxcfs restarts); the step waits for writes in flight (bounded, 2 min — then it fails and does not run). Wired at `proxmox.Client.doBody` (every non-GET) and `ExecRunner.RunStdin` (`WritesEtcPVE`: pct config verbs, pvesm, pveum, felhom-pbs-apply create/reconcile).
- `internal/osupdate`: `LayerPVE`; the night leg runs the pve step in ring 0 after a healthy host step (an appliance; ring 1 never in the night leg); `PVEHealthVerdict` = the host rule + every running container keeps its id + pveversion reads the installed pve-manager; the pve report carries Proxmox userspace only (the hub's candidate set). `PVEStepExecutor` (signed `os_pve_step`, ring 1, under the heavy-op gate and the /etc/pve gate); `reconcile.ClassOSPVEStep` (destructive-class); `felhom-opsign -op os_pve_step` (params by `-params`).
- Tests: wrapper `PVELane` (17; red first — the `pve-manager` plan was refused R12 on the old code), `pvegate` (5), `TestPVEGate_*` + `TestWritesEtcPVE`, `TestPVE_*`, `TestPVEHealthVerdict`, `TestPVEStepExecutor_*`. Red-proofs: `felhom.eu/documentation/audits/day-2026-10-07/B/`.
## Unreleased (2026-10-07)
- R-366 slice 2 (`09` §3 decision 168): the restore-test pick records, per tier, the archives it skipped as written with another key (count, oldest, newest — no key material) in a `ForeignKeyLedger`; the host report carries it as `foreign_key_archives.tiers` (the stanza absent until a tier was evaluated since start, `tiers: []` when none — no null on the wire, the report contract forbids it). The hub turns a change into one operator line. Tests `TestR366_PickRecordsArchivesWrittenWithAnotherKey`, `TestR366_EvaluatedWithNoneIsAnEmptyList` (red-proved, `felhom.eu/documentation/audits/day-2026-10-07/E/`).
- R-105 option A (`09` §3 decision 169): the `--selftest=escrow-create -directive <file>` flag and the escrow upload's `directive` field are removed — nothing read the directive; the DR path reads the recipe, tenantsync and the escrow blob. The hub ignores a `directive` from an older agent.
## v0.150.0 — the Docker step proves the engine reports a memory kill; after a restart the agent remembers the last backup per tier; three more SMART counters on the wire (R-528, R-894, R-330; `09` §3 decisions 157, 161) (2026-10-07)
Released by `scripts/release-agent.sh`: binary sha256 `a23d1c9085bc7fd4fc48fe0327f6504aa83e6331510fb4a3d23e042dddb26f9c`
config bundle sha256 `88456b386d9b1027bd22861cac8c23df004bf9fd9f67644d6595bfca8c94498e` (tag `v0.150.0` = `3a72a48`).
**The bundle carries the new `felhom-os-apply` (the memory-kill check) — deliver it with the binary:** signed
`agent_update`, then signed `agent_config_update`. No path added (26 → 26), so no step bundle.
### Part of v0.150.0 (2026-10-06 night, later) — after a restart the agent remembers the last backup per tier (R-894); three more SMART counters on the wire (R-330)
Ships with the memory-kill check below as v0.150.0, AFTER the 2026-10-07 night read-back. Nothing delivered tonight.
- **The defect (measured 2026-10-05 on demo-hp):** the agent restarted at 04:57; at 06:25 the off-site storage answered *Can't connect*; the per-tier backup record is in memory only, so the due-check fell back to an EMPTY record and the 7-day tier (last copy 4 days old) read DUE; the controller asked and vzdump failed.
- New `internal/backup/backup_state.go` `BackupSuccessState`: the newest SUCCESSFUL backup per tier and guest, on disk (`<oob state dir>/backup-success-state.json`, atomic tmp+rename, 0600). Only successes are written; a corrupt file reads as nothing known.
- `internal/localapi` `handleBackupDue`: when the tier's storage CANNOT be read, the saved copy stands in for the in-memory record. A fresh copy → not due („… (storage unreadable — age from the last success saved on disk)"); a copy older than the cadence → DUE; no copy → the old answer (DUE, age unknown). A storage that answers stays the ground truth: an archive absent there is due even when the file remembers one.
- Wired in `buildLocalAPIServer` (`LastKnownBackups`); the local API's backup job saves each success.
- Tests: `TestBackupDue_R894_*` (restart = a new server and a new state from the same file; fresh / old / none / storage answers / failed backup not saved), `TestBackupSuccessState_*`, `TestR894_LastKnownBackupsIsWiredIntoTheDaemon` (AST). Four red-proofs observed (`felhom.eu/documentation/audits/night-burndown-2026-10-06/s4/`).
- **R-330 (disk health Phase 2, the wire only):** the SMART summary carries three more SATA raw counters — `reported_uncorrect` (187), `command_timeout` (188, carried as the vendor reports it; some pack several counters), `udma_crc_errors` (199). Pointer + omitempty: an attribute the drive does not report is OMITTED (unknown), never 0. No verdict reads them yet. Tests `TestParseSMART_R330_*` (two red-proofs, `felhom.eu/documentation/audits/night-burndown-2026-10-06/r330/`).
### Part of v0.150.0 (2026-10-06 night) — the Docker step proves the engine reports a memory kill (`09` §3 decision 157, R-528)
To be released as v0.150.0 with its config bundle AFTER the 2026-10-07 night read-back (the night of 2026-10-06 runs v0.149.0 on purpose).
- `configs/felhom-os-apply`: after a docker-layer APPLY (after `health_after`) the wrapper runs `oom_check()`: a throwaway container from the image the running controller uses (`--pull never`, `--network none`, no volume, label `felhom.oomcheck=1`, 64 MB cap) asks for one 200 MB block; „pass" only when `OOMKilled=true` AND the `oom` event; it waits 2 s and reads the events window to the guest's epoch + 1 (measured: a window closed in the same second missed the event); the container is always removed. Reported as `oom_check`; it never changes the step's outcome or health. A wrapper-only mode `oom-check` runs the check alone (no apt, no engine change), by hand as root.
- `internal/osupdate`: `WrapperReport` and `Report` carry `oom_check` verbatim, on the normal pass and on the kept-copy path (R-868).
- `configs/test_felhom_os_apply.py`: its `unittest.main()` sat in the middle of the file, so 11 tests (UnsentReport, SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran — moved to the end; all pass.
- Tests: the OOMCheck class (pass, OOMKilled=false, no event, unreadable image, removal on an inspect error, not on other layers or in health mode, the events window after the settle wait, mode oom-check alone and its refusals); TestDocker_OOMCheckReachesTheHubUnchanged, TestR868_KeptCopyCarriesTheOOMCheck. 12 red-proofs in `felhom.eu/documentation/audits/readback-2026-10-07/F/`.
### Part of v0.150.0 (2026-10-06 evening) — the shared rule file (`09` §3 decision 152); no code change
- `.claude/rules/unprompted-work.md` added, byte-identical to the copies in felhom.eu, felhom-controller, app-catalog-felhom.eu and the workspace root (checked with `diff` against the controller's copy and one md5 across all five). Its copies line names five copies.
### Part of v0.150.0 (2026-10-06 afternoon) — instruction files kept true (`09` §3 decision 150); no code change
- `CLAUDE.md` „Gates — ONE entry point": the runner runs every gate in its `GATES` table (five: three shared, `published`, `release-complete`); `--fast` skips `published` (network). It said two gates and „all of them".
- `CLAUDE.md`: the decoy gate and its audit are named with their `felhom.eu/` prefix (they do not exist in this repo).
- `.claude/rules/health-checks.md` (comment): the health-check rule's copies live in felhom.eu `hub.md` and the controller's `gates.md`; it named felhom.eu `CLAUDE.md` „Code quality rules", which holds no such rule.
## v0.149.0 — a weekly disk trim of each customer guest, the crash-boot fact for the controller, the phantom WARN names its runbook (R-444, R-856, R-99; operator rulings `09` §3 139, 143, 140) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `6bcae9c2eb5d97e8285316583870059835793893299e291891a53a4ce505585f`
config bundle sha256 `e182c82dcf4a67faa3bcb74dbe4ffa7b06e0b27dc8451cb7574d6339ce91ad66` (tag `v0.149.0` = `f277e61`).
**The bundle carries the new sudoers rule for the trim (`FELHOM_FSTRIM`) — deliver it with the binary:** signed
`agent_update`, then signed `agent_config_update`.
- R-856 (`09` §3 decision 143): new local-API route `GET /host/crash-guard` — passes the host crash guard's last-boot record (present, last_boot_at, last_boot_unclean, tripped) from /var/lib/felhom-crash-guard/state.json to the controller, which waits ~15 min with app mails after a crash boot. Read-only, no Proxmox call, guest-token authed; a missing/unreadable/garbled file answers 200 present:false (never an error page). An older agent answers 404, which the controller reads as unknown (normal 90 s grace) — no controller MinAgent raise needed.
- R-444 (`09` §3 decision 139): weekly guest disk trim. New sudoers alias FELHOM_FSTRIM with ONE exact rule `/usr/sbin/pct ^fstrim [0-9]+$` (rides the signed config bundle; decoys pinned by TestSudoersFstrimRuleIsExact) and capability guest-fstrim (non-critical). New internal/fstrim job: each owned RUNNING guest gets `pct fstrim <vmid>` once a week - due Wednesday from 10:00 host-local, starts only 10:00-20:59 (never the 01:00-06:59 night), holds the one-heavy-op gate so it never runs beside a backup or restore-test (busy -> deferred to the next hourly tick; a box that was off catches up at its next daytime hour); a failed trim WARNs and is retried at most 3 times that week; bytes parsed from `pct fstrim`'s "(N bytes) trimmed" lines; positive log `fstrim: guest N trimmed X GiB in Ys`; last result per guest persisted in <state_dir>/guest-disk-trim.json and reported as the new omitempty host-report stanza `guest_disk_trim`. Opt-out: agent.json "disk_trim": {"disable": true}.
- R-99 (`09` §3 decision 140): the agent's WARN for a PBS archive below the 1 MiB plausibility floor now ends with the pointer to the sanctioned cleanup (`documentation/runbooks/pbs-phantom-cleanup.md`); detection only — nothing is deleted automatically. Dir-storage archives keep the old text.
## unreleased
- R-426: scripts/test_gate_decoys.py (new) — the published gate judged against a fake Gitea (127.0.0.1, via GITEA_BASE; never the real registry): 11 cases; COVERS published — `felhom-agent/published` leaves the decoy-coverage EXEMPT list.
- R-426: release-complete gate — 10 decoy cases (scratch clone + scratch bare origin + fake Gitea); COVERS release-complete — `felhom-agent/release-complete` leaves the decoy-coverage EXEMPT list.
- R-426: the shared reuse-refs/instructions/observations gates get agent-side decoys (9 cases on a scratch clone of this repo); COVERS reuse-refs, instructions, observations — three `felhom-agent/*` entries leave the decoy-coverage EXEMPT list.
- **Fixed without a row:** `configs/test_felhom_config_bundle.py` read the two ISO first-boot files that installer 1.32.0 now NAMES under KEPT (R-275) as files the installer writes — `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 (felhom.eu `85de3f9b`); they are listed with why. Test only; agent v0.148.0's code is unaffected.
## v0.148.0 — the host report names the running binary's sha; the format answer carries the new filesystem's UUID (burn-down night: R-349, R-25 agent halves) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `3e68a0870e0e2ce262cb4819294edddeb0a73e8c558a31611a20329a6d9ee283`
config bundle sha256 `a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de` (tag `v0.148.0` = `861d32a`).
Delivery order as for v0.147.0: signed `agent_update`, then signed `agent_config_update`.
- **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from
`/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub
comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved.
+7 -7
View File
@@ -52,11 +52,11 @@ This is in the core because breaching it is how this component stops being audit
## Gates — ONE entry point
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs this
repo's gates — `reuse_refs_check` and `instructions_gate`, both the **shared** copies in
`felhom.eu/scripts/`, never copied into this repo (a copy recreates the drift they detect; an absent
sibling clone FAILS). `--fast` selects the gates touching no network and no container runtime; today
that is all of them. **A missing gate is a FAILURE, never a skip.**
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs every
gate in its `GATES` table (that table is the list); the shared ones — `reuse_refs_check`,
`instructions_gate`, `observations_gate` — are the copies in `felhom.eu/scripts/`, never copied into
this repo (a copy recreates the drift they detect; an absent sibling clone FAILS). `--fast` selects the
gates touching no network and no container runtime, and skips `published` (network), naming it. **A missing gate is a FAILURE, never a skip.**
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is
**per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS
@@ -104,8 +104,8 @@ the mechanism are exempt.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
lacks. `felhom.eu/scripts/decoy_coverage_gate.py` (run by felhom.eu's `repo_gates.py`, for all four repos) refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
decoys withdrawn as illegitimate: `felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list.
+8 -7
View File
@@ -44,13 +44,14 @@ unnoticed until a user hit them. `internal/capability` makes that loud:
(`HostCapabilityChecker`) alerts the operator on a Critical capability going degraded. Serve-degraded
— the probe never blocks startup. (Next self-health slice: the controller↔agent channel check.)
**Controller-swap under non-root (v0.45.0).** The agent-owned controller image swap
(`internal/localapi/controllerswap.go`) no longer shells out: `writeImage` pipes the image ref on
**stdin** into an in-guest `tee /etc/felhom-controller-image` (via `GuestExecStdin` →
`Runner.RunStdin`, the same fenced `sudo -n` runner) — no `bash -c`, no interpolation. Its 5 narrow
grants live in the `FELHOM_CONTROLLERSWAP` sudoers alias (all read-only or fixed-target; the `tee`
target is the FIXED image path, content stdin-fed) and in the capability manifest (Critical), so a
dropped grant is a build failure + a live degraded signal. No general `pct exec` is granted.
**Controller-swap under non-root (v0.45.0; the write since R-861 (a) A1).** The agent-owned controller image swap
(`internal/localapi/controllerswap.go`) no longer shells out. The write goes on **stdin** to the ROOT verb
`felhom-priv-apply controller-image <vmid>` (`GuestBinder.WriteControllerImage` → `Runner.RunStdin`, the same fenced
`sudo -n` runner), which re-checks the ref against our registry + repository + an x.y.z tag and writes
`/etc/felhom-controller-image` inside the guest itself; the agent has no in-guest `tee` grant any more (before, a
compromised agent could feed any image — sudo cannot see stdin). Its grants live in the `FELHOM_CONTROLLERSWAP` sudoers
alias (read-only or fixed-target) and in the capability manifest (Critical), so a dropped grant is a build failure + a
live degraded signal. No general `pct exec` is granted.
## The `storage` package — observe + watchdog (slice 5)
+7 -9
View File
@@ -1,10 +1,8 @@
# REPORT — agent v0.147.0 (2026-10-05, burn-down round 2)
# REPORT — v0.150.0 released and delivered (2026-10-07)
Full session report: `felhom.eu/REPORT-burndown2-2026-10-05.md`. Baseline `d833163` (v0.146.1). Code commit `f1b9b41`
(CI job 1365 success), tag `v0.147.0`, binary sha256 `642c4d19…`, bundle sha256 `326527d0…` (verified by download).
Rows: R-124 (recipe root namespace = ""), R-118 (no root size for an absent drive), R-269 (rotated-out token rejected at
once), R-317 (dnsmasq install probed by its unit). Tests + red-proofs: `felhom.eu/documentation/audits/burndown2-2026-10-05/`.
`go build/vet/test ./...` green; `agent_gates.py --fast` green after the release (release-complete needs the tag).
Delivery: see the session report (vouch, signed jobs per box, the hub System page afterwards).
On the operator's word (`09` §3 decision 161). `scripts/release-agent.sh 0.150.0`: sha `a23d1c90…`, bundle `88456b38…`,
tag `v0.150.0` = `3a72a48`, verified by download. Vouched with golden 0.301.0 and MinAgent 0.131.0 (unchanged). No bundle
path added (26 → 26), so no step bundle. Signed `agent_update` → demo-hp, demo-felhom, Tester 1 on 0.150.0 (07:06–07:07Z);
signed `agent_config_update` → `BUNDLE DONE written=1 same=24 self-check=ok`, capability probe 68/68 on all three
(07:21Z). Tester 2 not touched. Carries R-528 (the memory-kill check), R-894 (the last backup per tier on disk), R-330
(SMART counters on the wire). Evidence: `felhom.eu/documentation/audits/readback-2026-10-07/delivery/`.
+6 -3
View File
@@ -15,7 +15,7 @@
| `SudoHostOps.run` | internal/storage/hostops.go | `run(ctx, name, args...) error` | allowlisted exec with stderr-wrapped error | Every arg pre-validated via validate.go before this is called |
| `Prober.Probe` | internal/capability/probe.go | `Probe(ctx) []Status` | live sudo-policy capability check (`sudo -n -l --`) | Needs a DIRECT runner (never the sudo-prefixing one — double-sudo); never executes probed cmds. v0.86.0: config-gated caps (`Capability.GatedBy` + `Prober.GateActive`) report `inactive`/"disabled by configuration" ONLY when healthy — broken plumbing stays degraded; the pbsdr-* gate answers from `pbsdr.Manager.DRConfigured` (marker-backed across restarts) |
| ~~`stageTemp`~~ (REMOVED v0.146.0, R-861) | — | — | — | Nothing the agent writes is `install`ed where root reads it any more: use `felhom-priv-apply` (below) or ship a fixed file in the bundle |
| `felhom-priv-apply` (v0.146.0, R-861) | configs/felhom-priv-apply | `felhom-priv-apply unit <name> \| dnsmasq <tmp> <name> \| wg \| sshd-config \| sshd-key` | ANY agent-rendered file a root program reads (systemd unit, dnsmasq drop-in, wg-quick conf, OOB sshd) — fixed source + destination, CONTENT checked against the agent's own renderers | A new renderer needs a verb + a contract test (`internal/privapplytest.Check`) feeding its REAL output; never a new `install` sudoers line |
| `felhom-priv-apply` (v0.146.0, R-861) | configs/felhom-priv-apply | `felhom-priv-apply unit <name> \| dnsmasq <tmp> <name> \| wg \| sshd-config \| sshd-key \| controller-image <vmid>` (the last reads the ref on stdin, R-861 (a) A1) | ANY agent-rendered file a root program reads (systemd unit, dnsmasq drop-in, wg-quick conf, OOB sshd) — fixed source + destination, CONTENT checked against the agent's own renderers | A new renderer needs a verb + a contract test (`internal/privapplytest.Check`) feeding its REAL output; never a new `install` sudoers line |
| `privapplytest.Check` | internal/privapplytest/check.go | `Check(t, verb, name, content) string` | the Go↔root-checker contract: a renderer's real output must read `OK` | Skips without python3; one call per rendered shape + one refused control |
| `BUNDLE_FILES` + `Bundle` (mode `bundle`, `--install-bundle`; mode `agent_update` v0.146.0, R-861) | configs/felhom-os-apply | the ONE table of root-owned paths + the installer of them | ANY new root-owned file the installer writes (sudoers line, wrapper, unit) — add it to the table, never a new installer fetch (R-840) | The builder (`scripts/build-config-bundle.py`) and the installer read the same table; `test_every_root_file_the_installer_writes_is_in_the_bundle` fails on a path the bundle lacks. Trust files (`/etc/felhom/os-trust.json`, `operator-signers`) are NEVER bundle paths (R17) |
| `osupdate.ConfigUpdateExecutor` | internal/osupdate/bundle.go | signed op `agent_config_update` {agent_version, bundle_sha256} | delivering the bundle to an installed box | a courier only: the root wrapper re-verifies signature, host, nonce and sha itself |
@@ -88,6 +88,7 @@
| Symbol | File | Short signature | Use for | Gotchas |
|---|---|---|---|---|
| `pvegate.Write` / `pvegate.Step` | internal/pvegate/pvegate.go | `Write(ctx) (release, waited, err)` / `Step(ctx) (end, err)` | R-812 option A: keep the agent's own /etc/pve writes out of a Proxmox package step (pmxcfs restarts) | Already wired at the two chokepoints — `Client.doBody` (every non-GET) and `ExecRunner.RunStdin` (`WritesEtcPVE`: pct config verbs, pvesm, pveum, felhom-pbs-apply create/reconcile). A new root CLI that writes /etc/pve goes into `WritesEtcPVE`, never its own lock. Never take `Step` around anything but the wrapper call (`Leg.runPVE`) — a `Write` inside a `Step` deadlocks until its context ends. |
| `Client.WaitTask` | internal/proxmox/task.go | `WaitTask(ctx, upid, opts) (TaskStatus, error)` | asserting EVERY mutating op | POST 200 ≠ success; authz can fail at task exec; `AllowWarnings` opt-in |
| `Client.Pool` | internal/proxmox/query.go | `Pool(ctx, name) (PoolInfo, error)` | felhom-pool membership (the ownership registry, A1) | Needs `Pool.Audit` at `/pool/<name>` (host-install v1.9.0+); `Pool.Allocate` does NOT satisfy the read; members can be storages (type `storage`, vmid 0) — filter them |
| `Client` mutate wrappers (`RestoreLXC/Vzdump/DestroyLXC/Snapshot/Rollback/SetConfig/ResizeLXC/Start/Stop`) | internal/proxmox/mutate.go | return `(upid, error)` | all API mutations | Async → always pair with WaitTask; route via gate/queue, not ad-hoc |
@@ -121,7 +122,7 @@
| Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only |
| Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set |
| Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction |
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`; R-856 `crashGuardStatePath` — GET /host/crash-guard's state file, internal/localapi/crashguard.go) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
| `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) |
| `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs |
| `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s |
@@ -160,13 +161,15 @@
| `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring |
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
| `backup.BackupSuccessState` | internal/backup/backup_state.go | `RecordBackupSuccess(target, b)` / `LastKnownSuccess(target, vmid)` | Newest SUCCESSFUL backup per tier+guest, persisted (atomic tmp+rename) — the due-check's fallback when the tier's storage cannot be read after a restart (R-894) | **Read ONLY when the storage cannot be read** — a storage that answers is the ground truth (R-84), and an archive absent there must make the tier due even when this file remembers one. A saved copy older than the cadence still reads due. Only successes are written. |
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
| `localapi.StaleLockController` | internal/localapi/stalelock.go | `*staleLockController` (Client + Runner + pool) | `fakeStaleLock` (Server-level) stalelock_test.go; `fakeStaleLockAPI` (controller-level, tests the A1 pool intersect) stalelock_pool_test.go |
| `localapi.GuestExecutor` | internal/localapi/controllerswap.go | `*GuestBinder` (pct exec) | `fakeGuestExec` internal/localapi/controllerswap_test.go |
| `localapi.GuestExecutor` | internal/localapi/controllerswap.go | `*GuestBinder` (pct exec; the image write via `felhom-priv-apply controller-image`) | `fakeGuestExec` internal/localapi/controllerswap_test.go |
| `guestnet.Runner` / `guestnet.GuestSource` (R-54, v0.92.0) | internal/guestnet/{probe,watchdog}.go | `*proxmox.ExecRunner`; the POOL-VERIFIED `localapi.StaleLockController.Guests` (ListLXC ∩ felhom pool, audit A1) | `scriptedRunner` + `fakeGuests` internal/guestnet/watchdog_test.go. **Never wire a bare `ListLXC` here** — under a broad token that would run dhclient inside a co-tenant's container. Every assertion is an exec COUNT, and the load-bearing ones are the negatives: a static guest, an unprobeable guest, a boot-race guest and an unproven guest list must record **zero** heal calls |
| `guestnet.Watchdog.SetDampers` / `now` (clock seam) | internal/guestnet/watchdog.go | config `guest_net.*`; `now` defaults to `time.Now` | tests advance a manual clock (the storage-watchdog pattern) and assert the heal ceilings EXACTLY — ≥10 min apart, ≤3/hour, and ≤30 over a scripted 10 hours of permanent failure. A damper with no test is a comment |
| `hub.GuestNetReporter` (R-54) | internal/hub/collect.go | `*guestnet.Watchdog` (`GuestNetStatus`) | internal/hub/collect_guestnet_test.go asserts the stanza through the PRODUCTION `Collect` path AND that the `guest_net` key is ABSENT from the wire when no reporter is wired — an always-present empty stanza would make "not wired" and "found nothing" the same signal, which is the shape v0.91.0 hid behind |
+53
View File
@@ -0,0 +1,53 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"testing"
)
// R-444: the weekly trim has the guestnet shape (component + reporter seam + goroutine), so its wiring is asserted
// from the AST like TestMainWiresGuestNetWatchdog — a unit-green trim job that main.go never starts is the inert-seam
// defect. It must also share the ONE heavy-op gate (heavyOps), or it could run beside a backup.
func TestMainWiresGuestDiskTrim(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, 0)
if err != nil {
t.Fatalf("parse main.go: %v", err)
}
var constructedWithGate, reporterWired, started bool
ast.Inspect(f, func(n ast.Node) bool {
switch node := n.(type) {
case *ast.CallExpr:
if fn, ok := node.Fun.(*ast.SelectorExpr); ok {
switch fn.Sel.Name {
case "New":
if pkg, ok := fn.X.(*ast.Ident); ok && pkg.Name == "fstrim" && len(node.Args) >= 3 {
if id, ok := node.Args[2].(*ast.Ident); ok && id.Name == "heavyOps" {
constructedWithGate = true
}
}
case "SetGuestDiskTrimReporter":
reporterWired = true
}
}
case *ast.GoStmt:
if sel, ok := node.Call.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "Run" {
if id, ok := sel.X.(*ast.Ident); ok && id.Name == "diskTrim" {
started = true
}
}
}
return true
})
if !constructedWithGate {
t.Error("main.go never calls fstrim.New(..., heavyOps, ...) — no trim job, or one outside the heavy-op gate")
}
if !reporterWired {
t.Error("main.go never calls collector.SetGuestDiskTrimReporter — the guest_disk_trim stanza never reaches the hub")
}
if !started {
t.Error("main.go never starts the trim job with `go diskTrim.Run(ctx)`")
}
}
+92 -65
View File
@@ -38,6 +38,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
"gitea.dooplex.hu/admin/felhom-agent/internal/fasttick"
"gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd"
"gitea.dooplex.hu/admin/felhom-agent/internal/fstrim"
"gitea.dooplex.hu/admin/felhom-agent/internal/guesthook"
"gitea.dooplex.hu/admin/felhom-agent/internal/guestnet"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
@@ -134,39 +135,38 @@ func main() {
return
}
var (
cfgPath string
selftest selftestFlag
vmid int
watch time.Duration
archive string
mode string
hostname string
keep bool
rootfsGrow int
dataVolGrow int
dataVolMount string
sysDataGrow int
sysDataMount string
cores int
memoryMB int
pbsStorage string
paperkey bool
offline bool
upload bool
custID string
custDomain string
custName string
custEmail string
hubPassword string
blobPath string
expectedFP string
keyDest string
installWGKey bool
idBundlePath string
directivePath string
swapImage string
outputMode string
showVersion bool
cfgPath string
selftest selftestFlag
vmid int
watch time.Duration
archive string
mode string
hostname string
keep bool
rootfsGrow int
dataVolGrow int
dataVolMount string
sysDataGrow int
sysDataMount string
cores int
memoryMB int
pbsStorage string
paperkey bool
offline bool
upload bool
custID string
custDomain string
custName string
custEmail string
hubPassword string
blobPath string
expectedFP string
keyDest string
installWGKey bool
idBundlePath string
swapImage string
outputMode string
showVersion bool
)
flag.StringVar(&cfgPath, "config", envOr("FELHOM_AGENT_CONFIG", "/etc/felhom-agent/agent.json"), "path to the agent config file (JSON)")
flag.Var(&selftest, "selftest", "run a self-test and exit: bare/`read` = read-only queries; `task` = reversible mutating exercise (needs -vmid); `hub` = one collect+report; `storage` = observe storage (+ -watch); `backup` = one-shot backup of -vmid; `restore-test` = restore→boot→verify→teardown of -archive (or newest backup); `restore-test-due` = READ-ONLY: print the per-tier due verdict the scheduler would act on, with its cost; `pbs-verify` = trigger a PBS verify + print snapshot records; `bring-up` = restore→reset identity→size→start link-up of -archive into -vmid (needs -mode/-archive/-vmid; optional -cores/-memory cap; tears down unless -keep); `provision` = full slice-8A chain: bring-up provision + mint token + populate bootstrap config mount (needs -archive/-vmid/-customer-id/-hub-password; optional -rootfs-grow/-datavol-grow/-cores/-memory (-sysdata-grow is deprecated: folded into -datavol-grow); keeps the guest)")
@@ -197,7 +197,6 @@ func main() {
flag.StringVar(&keyDest, "keydest", "", "for --selftest=escrow-consume: where to install the recovered key (0600)")
flag.BoolVar(&installWGKey, "install-wg-key", false, "for --selftest=identity-consume: ALSO install the recovered wg_private_key into wgtunnel's key file (S5 DR; create-only, refuses to overwrite)")
flag.StringVar(&idBundlePath, "identity-bundle", "", "for --selftest=escrow-create: a 0600 JSON file {tunnel_token,pbs_token} to ALSO escrow under R (10D)")
flag.StringVar(&directivePath, "directive", "", "for --selftest=escrow-create: a JSON file with the non-secret DR directive (pbs repo/ns, expected fingerprint, tunnel id)")
flag.StringVar(&custID, "customer-id", "", "for --selftest=provision: the customer id — the hub config-pull target, baked into the guest's bootstrap")
flag.StringVar(&hubPassword, "hub-password", "", "for --selftest=provision: the customer's hub RETRIEVAL PASSPHRASE (SECRET) — baked into bootstrap.json so the controller pulls its config (and the customer-scoped hub key) from the hub. The customer must already exist in the hub.")
flag.StringVar(&swapImage, "image", "", "for --selftest=controller-swap: the target controller image ref (gitea.dooplex.hu/admin/felhom-controller:<semver>) — must already be pulled in the guest")
@@ -268,7 +267,7 @@ func main() {
SysDataGrowGB: sysDataGrow, SysDataMount: sysDataMount, Cores: cores, MemoryMB: memoryMB},
}))
case "escrow-create":
os.Exit(runSelftestEscrowCreate(context.Background(), cfg, logger, pbsStorage, paperkey, offline, upload, idBundlePath, directivePath, outputMode))
os.Exit(runSelftestEscrowCreate(context.Background(), cfg, logger, pbsStorage, paperkey, offline, upload, idBundlePath, outputMode))
case "escrow-consume":
os.Exit(runSelftestEscrowConsume(context.Background(), logger, blobPath, expectedFP, keyDest))
case "identity-consume":
@@ -791,6 +790,9 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
pbsTargets := pbsTargetsFromPVE(cfg, px, logger)
pbsReporter := pbs.NewLiveSnapshotReporter(pbsTargets, pbsStore, pbs.DefaultLiveSnapshotTimeout, logger)
collector := hub.NewCollector(px, newTunnelProber(cfg, px), observer, backupStore, backupStore, pbsReporter, cfg.Hub.HostID, version, logger)
// R-366 slice 2: the restore-test's ledger of archives written with another key → the host report.
foreignKeys := backup.NewForeignKeyLedger()
collector.SetForeignKeyArchiveReporter(foreignKeys)
collector.SetBackupTargetResolver(primaryBackupTargetOf(cfg)) // R-109: the recipe names the live target
// Privileged-capability self-check (v0.44.0): probe the sudoers grants the non-root agent
// depends on. The probe runs `sudo -n -l` LITERALLY (a policy LIST, never executing the
@@ -1016,7 +1018,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// with the local API so a backup and a restore-test can never run together.
rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json"))
heavyOps := &backup.InFlight{}
scheduler := buildRestoreTestScheduler(cfg, px, engine, backupStore, rtState, heavyOps, logger)
scheduler := buildRestoreTestScheduler(cfg, px, engine, backupStore, rtState, heavyOps, foreignKeys, logger)
// R-189: the host report's restore_tests[] must survive an agent restart. The in-memory store
// holds only this process's latest run, and under per-archive due-ness the agent will not
// re-test an archive it has already proven — so without this the hub can report a tier unproven
@@ -1108,6 +1110,15 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
}
return release, nil
}}
// R-812 option A: a signed Proxmox package step (`11` §5.10) — ring 1; under the heavy-op gate and the /etc/pve gate.
pveExec := osupdate.PVEStepExecutor{Leg: osLeg, Guest: firstGuest(px),
Gate: func(ctx context.Context) (func(), error) {
release, busy, ok := heavyOps.TryAcquire("os-pve-step")
if !ok {
return nil, fmt.Errorf("busy: %s", busy)
}
return release, nil
}}
// Agent v0.143.0 (R-840): the config bundle — the box's root-owned files by a signed job; the wrapper verifies it.
bundleExec := osupdate.ConfigUpdateExecutor{Leg: osLeg, URLTemplate: suCfg.URLTemplate, Username: suCfg.Username, Token: suCfg.Token,
// The capability probe confirms from the agent's side: `sudo -l` lists every command the new sudoers grants.
@@ -1119,7 +1130,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
}
logger.Warn("osupdate: capability probe after the config bundle", "ok", ok, "total", total, "degraded", strings.Join(names, ","))
}}
jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec, dockerExec, bundleExec}, cfg.Hub.HostID, logger)
jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec, dockerExec, pveExec, bundleExec}, cfg.Hub.HostID, logger)
loop.SetEnvelopeObserver(hub.MultiObserver(desiredSyncer, jobsRunner))
// Controller-driven escrow ceremony (v0.88.0): static config facts + the LATE-BOUND DR gate —
@@ -1479,6 +1490,25 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
}
go runJanitor(ctx, jd)
}
// R-444 (`09` §3 decision 139): the weekly guest disk trim — `pct fstrim <vmid>` of every owned, running guest,
// Wednesday from 10:00 local, daytime only, under the one-heavy-op gate. Not part of the errc fan-out: a trim job
// must never be able to bring the agent down.
if cfg.DiskTrim.Enabled() {
dtMode := proxmox.RunnerMode(cfg.Privileged.Mode)
if dtMode == "" {
dtMode = proxmox.RunnerSudo
}
dtRunner := &proxmox.ExecRunner{Mode: dtMode, SudoPath: cfg.Privileged.SudoPath}
dtGuests := localapi.NewStaleLockController(px, dtRunner, reconcile.DefaultPool, logger)
if dtGuests != nil {
diskTrim := fstrim.New(dtRunner, dtGuests, heavyOps,
filepath.Join(cfg.OOB.WithDefaults().StateDir, "guest-disk-trim.json"), logger)
collector.SetGuestDiskTrimReporter(diskTrim)
go diskTrim.Run(ctx)
}
} else {
logger.Info("fstrim: weekly guest disk trim disabled by config (disk_trim.disable)")
}
if lanLoop != nil {
lanServers = 1
go func() { errc <- lanLoop.Run(ctx) }()
@@ -1680,7 +1710,7 @@ func primaryBackupTargetOf(cfg config.Config) func() hub.ConfiguredBackupTarget
// disables the cadence (returns a scheduler that just waits) when the cadence is off or the
// scratch band / restore storage is invalid — a misconfig must not crash the daemon, and the
// machinery still works on-demand via --selftest=restore-test.
func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *reconcile.Engine, store *backup.Store, rtState *backup.RestoreTestState, inFlight *backup.InFlight, logger *slog.Logger) *backup.Scheduler {
func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *reconcile.Engine, store *backup.Store, rtState *backup.RestoreTestState, inFlight *backup.InFlight, foreign *backup.ForeignKeyLedger, logger *slog.Logger) *backup.Scheduler {
// R-86: this is the EVALUATION interval, not the trigger. What decides a test happens is the
// per-archive due-check in internal/backup/restoretest_due.go.
cadence := cfg.Backup.RestoreTestEvalInterval()
@@ -1699,6 +1729,9 @@ func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *re
min, max := cfg.Backup.ScratchBand()
target := cfg.Backup.BackupTarget()
runner := backup.NewBackupRunner(px, target, "", "felhom restore-test", "", logger)
if foreign != nil {
runner.SetForeignKeyLedger(foreign) // R-366 slice 2: the pick records archives written with another key
}
// Every configured tier is a rotation candidate, not just the primary.
cfgTiers, _ := cfg.Backup.BackupTiers() // warnings already logged where the tiers are armed
tierIDs := make([]string, 0, len(cfgTiers))
@@ -1881,12 +1914,15 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
Store: store,
Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
Tokens: tokens,
BackupCadence: cfg.Backup.BackupCadence(),
// R-894: the newest success per tier on disk — the due-check's fallback when the storage cannot be
// read right after a restart. Same state dir as restore-test-state.json.
LastKnownBackups: backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json")),
Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
Tokens: tokens,
BackupCadence: cfg.Backup.BackupCadence(),
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
Disks: hostOps,
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
@@ -2241,7 +2277,7 @@ func runSelftestRestoreTestDue(ctx context.Context, cfg config.Config, logger *s
return 1
}
rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json"))
sched := buildRestoreTestScheduler(cfg, px, nil, backup.NewStore(), rtState, &backup.InFlight{}, logger)
sched := buildRestoreTestScheduler(cfg, px, nil, backup.NewStore(), rtState, &backup.InFlight{}, nil, logger)
fmt.Printf("eval_interval=%s settle=%s\n", cfg.Backup.RestoreTestEvalInterval(), cfg.Backup.RestoreTestSettle())
start := time.Now()
@@ -2667,7 +2703,6 @@ type escrowCeremonyOpts struct {
offline bool
upload bool
identityBundlePath string
directivePath string
}
// escrowCeremonyOutcome is the shared core's result. R is the ONLY secret; Sum mirrors the
@@ -2726,11 +2761,10 @@ func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger,
return out, &escrowCeremonyErr{kind: "setup", err: fmt.Errorf("PBS key for %q not found (%s): %v", storage, keyPath, err)}
}
// Slice 10D.1: optionally ALSO wrap the identity bundle under the same R, and carry the non-secret
// directive for the hub. The bundle file is a 0600 secret (tunnel/pbs tokens); the directive is
// non-secret (pbs repo/ns, expected fingerprint, tunnel id).
// Slice 10D.1: optionally ALSO wrap the identity bundle under the same R. The bundle file is a 0600 secret
// (tunnel/pbs tokens). The non-secret "directive" that used to ride along is retired (R-105, `09` §3 decision 169:
// nothing read it; the DR path reads the recipe, tenantsync and the escrow blob).
var identity *escrow.IdentityBundle
var directive json.RawMessage
if opts.identityBundlePath != "" {
raw, err := os.ReadFile(opts.identityBundlePath)
if err != nil {
@@ -2741,11 +2775,6 @@ func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger,
return out, &escrowCeremonyErr{kind: "setup", err: fmt.Errorf("identity bundle is not valid JSON {tunnel_token,pbs_token}: %v", err)}
}
identity = &b
if opts.directivePath != "" {
if d, err := os.ReadFile(opts.directivePath); err == nil && json.Valid(d) {
directive = d
}
}
}
// S3: auto-inject the offsite WG private key into the escrowed identity when the key file
// exists — a NEW escrow run should always capture the live tunnel identity. Field NAME only
@@ -2824,7 +2853,7 @@ func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger,
ResticPwSealed: resticStaged,
}
if opts.upload {
if err := uploadEscrowBlob(ctx, cfg, res, directive, resticPwSHA256); err != nil {
if err := uploadEscrowBlob(ctx, cfg, res, resticPwSHA256); err != nil {
// R is minted and the blob self-verified — only the hub leg failed. kind "upload" lets
// the text shell keep the pre-extraction order (R surfaced, THEN the failure).
return out, &escrowCeremonyErr{kind: "upload", err: err}
@@ -2881,7 +2910,7 @@ func printEscrowTextRBlock(out *escrowCeremonyOutcome) {
// NOTHING else there; every human/info line goes to stderr; failures exit non-zero with no
// partial JSON. This is the controller-driven ceremony's parse surface (spike §2.3: the text
// banner is positionally brittle).
func runSelftestEscrowCreate(ctx context.Context, cfg config.Config, logger *slog.Logger, storage string, paperkey, offline, upload bool, identityBundlePath, directivePath, outputMode string) int {
func runSelftestEscrowCreate(ctx context.Context, cfg config.Config, logger *slog.Logger, storage string, paperkey, offline, upload bool, identityBundlePath, outputMode string) int {
switch outputMode {
case "", "text", "json":
default:
@@ -2899,7 +2928,7 @@ func runSelftestEscrowCreate(ctx context.Context, cfg config.Config, logger *slo
out, cerr := escrowCeremony(ctx, cfg, logger, escrowCeremonyOpts{
storage: storage, paperkey: paperkey, offline: offline, upload: upload,
identityBundlePath: identityBundlePath, directivePath: directivePath,
identityBundlePath: identityBundlePath,
})
if cerr != nil {
switch cerr.kind {
@@ -3082,20 +3111,19 @@ type escrowUploadRequest struct {
BlobB64 string `json:"blob_b64"` // base64 of the opaque R-wrapped blob (ciphertext)
KeyFingerprint string `json:"key_fingerprint"` // for operator display only
Posture string `json:"posture"` // e.g. "zero_knowledge"
// Slice 10D.1 — optional DR bundle (identity escrow + non-secret directive). Omitted in slice-7.
IdentityBlobB64 string `json:"identity_blob_b64,omitempty"`
DirectiveJSON json.RawMessage `json:"directive,omitempty"`
CreatedAt string `json:"created_at"` // RFC3339
// Slice 10D.1 — optional identity escrow. Omitted in slice-7. (The `directive` is retired — R-105.)
IdentityBlobB64 string `json:"identity_blob_b64,omitempty"`
CreatedAt string `json:"created_at"` // RFC3339
// SLICE 3 — sha256 hex of the offsite restic repo password sealed in the identity blob (present only
// when a staged password was folded in). Non-reversible hash of a 256-bit random secret — safe to
// store/serve; lets the controller VERIFY "the escrow covers the CURRENT key" and auto-confirm.
ResticPwSHA256 string `json:"restic_pw_sha256,omitempty"`
}
// uploadEscrowBlob PUTs the opaque blob (and, for 10D, the identity blob + non-secret directive) to
// uploadEscrowBlob PUTs the opaque blob (and, for 10D, the identity blob) to
// the hub, authed with the per-host key. The hub stores ciphertext + non-secret fields; no usable
// secret leaves the agent.
func uploadEscrowBlob(ctx context.Context, cfg config.Config, res escrow.CreateResult, directive json.RawMessage, resticPwSHA256 string) error {
func uploadEscrowBlob(ctx context.Context, cfg config.Config, res escrow.CreateResult, resticPwSHA256 string) error {
if cfg.Hub.URL == "" || cfg.Hub.HostID == "" || cfg.Hub.APIKey == "" {
return fmt.Errorf("hub not configured (url/host_id/api_key)")
}
@@ -3108,7 +3136,6 @@ func uploadEscrowBlob(ctx context.Context, cfg config.Config, res escrow.CreateR
}
if len(res.IdentityBlob) > 0 {
upReq.IdentityBlobB64 = base64.StdEncoding.EncodeToString(res.IdentityBlob)
upReq.DirectiveJSON = directive
}
body, _ := json.Marshal(upReq)
url := strings.TrimRight(cfg.Hub.URL, "/") + "/api/v1/hosts/" + cfg.Hub.HostID + "/escrow"
+74
View File
@@ -0,0 +1,74 @@
package main
import (
"go/ast"
"testing"
)
// R-894 — the on-disk backup record is WIRED on the daemon path (the built-but-never-wired class).
// main → runDaemon → buildLocalAPIServer, and inside it the localapi.Options literal carries
// LastKnownBackups built by backup.NewBackupSuccessState. An AST walk, not a string match, for the
// reasons in escrow_recover_wiring_test.go.
//
// COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this
// fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored.
func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
_, f := parseMain(t)
if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one")
}
var field, built bool
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
cl, ok := n.(*ast.CompositeLit)
if !ok {
return true
}
sel, ok := cl.Type.(*ast.SelectorExpr)
if !ok {
return true
}
if pkg, _ := sel.X.(*ast.Ident); pkg == nil || pkg.Name+"."+sel.Sel.Name != "localapi.Options" {
return true
}
for _, el := range cl.Elts {
kv, ok := el.(*ast.KeyValueExpr)
if !ok {
continue
}
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "LastKnownBackups" {
field = true
if callsIn(kv.Value)["backup.NewBackupSuccessState"] {
built = true
}
}
}
return true
})
}
if !field {
t.Fatal("localapi.Options in buildLocalAPIServer has no LastKnownBackups field")
}
if !built {
t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState")
}
}
func callsIn(n ast.Node) map[string]bool {
out := map[string]bool{}
ast.Inspect(n, func(n ast.Node) bool {
if ce, ok := n.(*ast.CallExpr); ok {
if fn, ok := ce.Fun.(*ast.SelectorExpr); ok {
if x, ok := fn.X.(*ast.Ident); ok {
out[x.Name+"."+fn.Sel.Name] = true
}
}
}
return true
})
return out
}
+1 -1
View File
@@ -43,7 +43,7 @@ func main() {
func run() error {
var (
op = flag.String("op", "", "op class to sign, e.g. storage_wipe | guest_destroy | decommission | agent_update | os_docker_step | agent_config_update")
op = flag.String("op", "", "op class to sign, e.g. storage_wipe | guest_destroy | decommission | agent_update | os_docker_step | os_pve_step | agent_config_update")
host = flag.String("host", "", "target host_id (anti-retarget — the op runs ONLY on this host)")
guest = flag.String("guest", "", "target guest_id (\"\" = host-scoped op)")
keyID = flag.String("key-id", "", "key id of the signing key (must match a pinned agent signer)")
+14 -4
View File
@@ -111,15 +111,17 @@ Cmnd_Alias FELHOM_INTERMEDIARY = \
# docker inspect -f * — container running/health/image (read-only; `*` spans the -f template
# + container across spaces, spike-confirmed)
# systemctl restart <fixed unit> — re-run the golden's bootstrap (the only state change)
# tee <FIXED image file> — WRITE the ref; content is fed on STDIN (no shell, no interpolation),
# the agent strict-validates the ref (controllerImageRe) before the write.
# felhom-priv-apply controller-image <vmid> — WRITE the ref (R-861 (a) A1, `09` §3 decision 165): the ref goes on
# STDIN to the ROOT wrapper, which requires our registry + repository + an x.y.z tag
# and writes the guest file itself. The agent's own `tee` grant is GONE: before, a
# compromised agent could hand the guest's bootstrap ANY image (sudo cannot see stdin).
# Validated GO: felhom.eu/documentation/audits/SPIKE-controllerswap-narrow-grants-2026-06-29.md.
Cmnd_Alias FELHOM_CONTROLLERSWAP = \
/usr/sbin/pct ^exec [0-9]+ -- cat /etc/felhom-controller-image$, \
/usr/sbin/pct ^exec [0-9]+ -- docker image inspect gitea\.dooplex\.hu/admin/felhom-controller\:[0-9]+\.[0-9]+\.[0-9]+$, \
/usr/sbin/pct ^exec [0-9]+ -- docker inspect -f .+ (felhom-controller|cloudflared)$, \
/usr/sbin/pct ^exec [0-9]+ -- systemctl restart felhom-controller-bootstrap\.service$, \
/usr/sbin/pct ^exec [0-9]+ -- tee /etc/felhom-controller-image$
/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$
# Stale-lock recovery (F2-b, v0.49.0). A host reboot DURING a vzdump backup leaves the guest with a
# `snapshot-delete`/`backup` lock + `onboot:1` then can't start it → the customer box stays DOWN. The
@@ -129,6 +131,14 @@ Cmnd_Alias FELHOM_CONTROLLERSWAP = \
Cmnd_Alias FELHOM_STALELOCK = \
/usr/sbin/pct ^unlock [0-9]+$
# Weekly guest disk trim (R-444, operator ruling `09` §3 decision 139). A thin pool only ever grows from blocks the
# guest has FREED: `fstrim` inside the unprivileged container is refused (FITRIM: Operation not permitted), so the host
# trims the guest's mounts. Measured on demo-hp 2026-10-06: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %,
# apps kept answering. ONE exact pattern: a vmid and nothing else — no `--ignore-mountpoints`, no second argument
# (pinned: TestSudoersFstrimRuleIsExact). The agent runs it on a weekly daytime timer under the heavy-op gate.
Cmnd_Alias FELHOM_FSTRIM = \
/usr/sbin/pct ^fstrim [0-9]+$
# Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS
# leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and
# a guest joins that pool only when its restore COMPLETES, so a failed restore leaves a pool-less guest
@@ -304,4 +314,4 @@ Cmnd_Alias FELHOM_GUESTNET = \
/usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \
/usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_FSTRIM, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
+5 -3
View File
@@ -16,8 +16,10 @@ Cmnd_Alias FELHOM_OP_REPAIR = \
/usr/bin/systemctl reset-failed felhom-sshd, \
/usr/bin/systemctl restart felhom-sshd, \
/usr/sbin/pct list, \
/usr/sbin/pct start [0-9]*, \
/usr/sbin/pct stop [0-9]*, \
/usr/sbin/pct unlock [0-9]*
/usr/sbin/pct ^start [0-9]+$, \
/usr/sbin/pct ^stop [0-9]+$, \
/usr/sbin/pct ^unlock [0-9]+$
# R-861 (b) B2 (`09` §3 decision 165, hygiene): one numeric vmid per pct verb, anchored — the old glob `[0-9]*` also
# matched spaces, so `pct stop 9201 --skiplock 1` passed. Pinned by TestFelhomOpSudoersPctIsExact.
felhom-op ALL=(root) NOPASSWD: FELHOM_OP_REPAIR
+170 -17
View File
@@ -21,6 +21,11 @@
# or, for an unsigned ring-0 step, the root-owned TRUST_FILE saying `"ring0_slow_lane": true` (set by hand on the demo
# boxes only). The agent's own config is NOT trusted for either: the agent can write it. A Docker step also needs
# `live-restore` ON (R15) — without it every container restarts.
# Agent v0.151.0 (R-812 option A, `09` §3 decision 163) adds the layer "pve": the HOST's Proxmox USERSPACE packages,
# lane "slow" only, origin "Proxmox Debian Repository" only, never a kernel / boot / firmware / microcode name (R14 —
# the kernel is R-836's lane), no removal, no undo, a NEW package only from PVE_NEW_ALLOW, an appliance only (R12), and
# the same authority as the Docker step (R3: a signed `os_pve_step`, or the root-owned ring-0 mark). Select
# "pending-pve" (ring 0): every installed Proxmox-origin package with a pending upgrade, minus HOST_SLOW_RE.
#
# Modes (plan field "mode"):
# inventory `apt-get update`, then report what is installed (with origin), what is pending, and health.
@@ -32,6 +37,14 @@
# held packages, kernel taint, the crash guard; guest Debian, Docker engine, containerd, live-restore.
# live-restore-on (v0.142.0, layer guest) the ONE-TIME step of `09` decision 87: merge `"live-restore": true`
# into the guest's /etc/docker/daemon.json and `systemctl reload docker`. NEVER a restart (R-835).
# oom-check (R-528, `09` decision 157; layer docker, lane slow) ONLY the memory-kill check below: no apt, no engine
# change, no authority needed. Wrapper-only — the agent never writes this plan; run it by hand as root.
# R-528 (`09` decision 157): a docker-layer APPLY also runs `oom_check()` after health_after and reports it as
# "oom_check": {"result": "pass"|"fail"|"error", "oom_killed": bool, "oom_event": bool, "exit_code": int|null,
# "image": str|null, "detail": str}
# — one throwaway container (the controller's own image, no network / volume / port, 64 MB cap) is made to exceed its
# memory; "pass" only when the engine says OOMKilled=true AND emits the `oom` event. It never changes the step's
# outcome or health: the hub decides whether the engine set can be approved. Pinned by the OOMCheck tests.
# Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is
# OSAPPLY-REPORT <one JSON object>
# which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install.
@@ -68,6 +81,16 @@ JOURNAL_MARK = "@@FELHOM-DPKG-JOURNAL@@"
DPKG_STATE_SCRIPT = "dpkg --audit; echo " + JOURNAL_MARK + "; ls -A /var/lib/dpkg/updates 2>/dev/null; true"
# The installer's ROOT-OWNED record (felhom-host-install.sh `state_set mode`); the agent cannot write it.
INSTALL_STATE = "/var/lib/felhom-install/state.json"
# R-528: the memory-kill check (oom_check). One 200 MB block under a 64 MB cap: measured on Docker 29.8.2 to be
# OOM-killed with OOMKilled=true and an `oom` event. Every call is bounded: timeouts (the clock read twice) + the
# settle wait stay within 90 s.
OOMCHECK_PREFIX = "felhom-oomcheck-"
OOMCHECK_SCRIPT = "dd if=/dev/zero of=/dev/null bs=200M count=1"
OOMCHECK_TIMEOUTS = {"image": 10, "clock": 5, "run": 30, "inspect": 10, "events": 10, "rm": 15}
# Measured on demo-hp 9201 (Docker 29.8.2, 2026-10-06, audits/readback-2026-10-07/F/F1, F2): an `--until` taken right
# after the run MISSED the oom event although OOMKilled=true; after a 2 s wait and `--until` = guest epoch + 1 it is
# seen. Pinned by test_events_window_ends_after_the_settle_wait.
OOMCHECK_SETTLE = 2
# Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host
# reboot is needed for them to take effect, and a bad one can stop the box from booting.
HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|"
@@ -82,6 +105,12 @@ HOST_SERVICES = ["pveproxy", "pvedaemon", "pvestatd", "pve-cluster", "felhom-age
DOCKER_NAMES = ("containerd.io", "docker-buildx-plugin", "docker-ce", "docker-ce-cli", "docker-ce-rootless-extras",
"docker-compose-plugin")
DOCKER_ORIGIN = "Docker CE"
# The Proxmox package lane (R-812 option A): the origin apt prints for download.proxmox.com, and the ONLY new packages
# a pve step may add (measured on demo-felhom 2026-10-07: a full upgrade adds proxmox-firewall-data and the kernel;
# the kernel is refused by R14 whatever this list says).
PVE_ORIGIN = "Proxmox Debian Repository"
PVE_NEW_ALLOW = ("proxmox-firewall-data",)
PVE_SIGNED_OP = "os_pve_step"
# ROOT-OWNED trust anchors (the installer writes them; the demo boxes got them by hand, R-840). Never the agent's config.
TRUST_FILE = "/etc/felhom/os-trust.json" # {"host_id": "...", "ring0_slow_lane": false}
TRUST_SIGNERS = "/etc/felhom/operator-signers" # ssh allowed_signers: <key_id> namespaces="felhom-op-v1" <key>
@@ -405,7 +434,7 @@ class Apply:
def check_plan(self, plan):
mode = plan.get("mode", "apply")
if mode not in ("apply", "inventory", "health", "facts", "live-restore-on", "bundle", "agent_update"):
if mode not in ("apply", "inventory", "health", "facts", "live-restore-on", "bundle", "agent_update", "oom-check"):
raise Refused("R11", f"unknown mode {mode!r}")
if mode == "agent_update":
if plan.get("layer") != "host":
@@ -416,18 +445,22 @@ class Apply:
raise Refused("R11", "bundle is a host-layer mode")
return mode, "host", 0, "bundle"
layer = plan.get("layer")
if layer not in ("guest", "host", "docker"):
raise Refused("R12", f"layer {layer!r} is not guest, host or docker")
if layer not in ("guest", "host", "docker", "pve"):
raise Refused("R12", f"layer {layer!r} is not guest, host, docker or pve")
lane = plan.get("lane", "fast")
if layer == "docker" and lane != "slow":
raise Refused("R3", "the Docker engine is the slow lane (`11` §5.8); a fast-lane Docker plan is refused")
if layer != "docker" and lane != "fast":
if layer == "pve" and lane != "slow":
raise Refused("R3", "the Proxmox packages are the slow lane (`11` §5.10); a fast-lane pve plan is refused")
if layer not in ("docker", "pve") and lane != "fast":
raise Refused("R3", f"the {layer} layer has no slow lane in this release (kernel, Proxmox: `11` §8 step 6)")
if mode == "facts" and layer != "host":
raise Refused("R11", "facts is a host-layer mode (it reads the host and the guest)")
if mode == "live-restore-on" and layer != "guest":
raise Refused("R11", "live-restore-on is a guest-layer mode")
if plan.get("undo") and layer != "docker":
if mode == "oom-check" and layer != "docker":
raise Refused("R11", "oom-check is a docker-layer mode (it checks the guest's Docker engine)")
if plan.get("undo") and layer != "docker": # the pve layer has no undo in this release (R-812 option A)
raise Refused("R5", "an undo (downgrade) exists only for the Docker layer, inside a signed job")
vmid = plan.get("vmid")
if not isinstance(vmid, int) or isinstance(vmid, bool) or vmid <= 0:
@@ -438,17 +471,19 @@ class Apply:
if plan.get("allow_new"):
raise Refused("R6", "allow_new is a slow-lane field; the fast lane never adds a package")
select = plan.get("select", "listed")
if select not in ("listed", "pending-fast", "pending-docker"):
if select not in ("listed", "pending-fast", "pending-docker", "pending-pve"):
raise Refused("R11", f"unknown select {select!r}")
if (select == "pending-docker") != (layer == "docker" and select != "listed"):
if select == "pending-docker" or layer == "docker":
raise Refused("R11", f"select {select!r} does not fit layer {layer!r}")
if (select == "pending-pve") != (layer == "pve" and select != "listed"):
raise Refused("R11", f"select {select!r} does not fit layer {layer!r}")
pk = plan.get("packages", [])
if not isinstance(pk, list):
raise Refused("R11", "packages must be a list")
if mode == "apply" and select == "listed" and not pk:
raise Refused("R11", "packages must be a non-empty list in apply mode (select listed)")
if select in ("pending-fast", "pending-docker") and pk:
if select in ("pending-fast", "pending-docker", "pending-pve") and pk:
raise Refused("R11", f"select {select} takes no package list")
seen = set()
for e in pk:
@@ -468,6 +503,12 @@ class Apply:
continue
if n in DOCKER_NAMES:
raise Refused("R2", f"{n} is a Docker package — the slow lane (`11` §5.8), never in a {layer} plan")
if layer == "pve":
if o != PVE_ORIGIN:
raise Refused("R2", f"{n}: origin {o!r} is not {PVE_ORIGIN!r} (the pve layer)")
if HOST_SLOW_RE.match(n):
raise Refused("R14", f"{n} is a kernel / boot / firmware package — never the pve lane (R-836)")
continue
if o not in FAST_ORIGINS:
raise Refused("R2", f"{n}: origin {o!r} is not Debian / Debian-Security (the fast lane, `11` C3)")
if layer == "host" and HOST_SLOW_RE.match(n):
@@ -610,12 +651,13 @@ class Apply:
now = self.r.now()
self.r.write_nonces({k: v for k, v in seen.items() if v > now})
def docker_authority(self, plan):
"""R3 for the docker layer: returns (who, undo). A signed job binds the EXACT package list and the undo flag."""
def docker_authority(self, plan, op_name=SIGNED_OP):
"""R3 for the docker (and, op_name os_pve_step, the pve) layer: returns (who, undo). A signed job binds the EXACT
package list and the undo flag."""
trust = self.load_trust()
signed = plan.get("signed")
if signed:
params = self.verify_signed(signed, trust)
params = self.verify_signed(signed, trust, op_name=op_name)
want = sorted(f"{e.get('name')}={e.get('version')}" for e in params.get("packages") or [])
got = sorted(f"{e['name']}={e['version']}" for e in plan.get("packages", []))
if not want or want != got:
@@ -629,7 +671,7 @@ class Apply:
raise Refused("R3", "an undo (downgrade) needs a signed operator job")
if trust.get("ring0_slow_lane") is True:
return "ring0", False
raise Refused("R3", "a Docker step needs a signed operator job (ring 1) or this box's root-owned ring-0 mark")
raise Refused("R3", f"a {self.layer} slow-lane step needs a signed operator job (ring 1) or this box's root-owned ring-0 mark")
def live_restore(self):
rc, out, _ = self.g(["docker", "info", "--format", "{{.LiveRestoreEnabled}}"], timeout=60)
@@ -775,8 +817,8 @@ class Apply:
# ---------- target helpers ----------
def x(self, argv, timeout=1800):
"""Run in the TARGET layer: the guest via pct exec, or the host directly."""
if self.layer == "host":
"""Run in the TARGET layer: the guest via pct exec, or the host directly (host and pve)."""
if self.layer in ("host", "pve"):
return self.r.host(argv, timeout)
return self.r.guest(self.vmid, argv, timeout) # guest and docker both live in the customer guest
@@ -871,7 +913,7 @@ class Apply:
def restart_needed(self):
"""Processes still mapping deleted files, OUTSIDE containers (C11). Guest: outside docker; host: outside the
LXC guests (the host's /proc shows guest processes too)."""
skip = RESTART_SKIP_CGROUP["host" if self.layer == "host" else "guest"]
skip = RESTART_SKIP_CGROUP["host" if self.layer in ("host", "pve") else "guest"]
script = ('for p in /proc/[0-9]*; do grep -q "(deleted)" $p/maps 2>/dev/null || continue; '
'grep -q "%s" $p/cgroup 2>/dev/null && continue; echo "${p#/proc/} $(cat $p/comm 2>/dev/null)"; done' % skip)
rc, out, _ = self.x(["sh", "-c", script], timeout=120)
@@ -945,13 +987,20 @@ class Apply:
return Bundle(self).from_plan(plan)
if self.mode == "agent_update":
return self.agent_update(plan)
if self.layer == "host":
if self.layer in ("host", "pve"):
self.check_appliance()
self.check_guest(self.vmid)
log = self.r.log
if self.mode == "live-restore-on":
return self.live_restore_on()
if self.mode == "oom-check":
# R-528: the check alone — no apt, no engine change; check_guest above still applies.
self.report["oom_check"] = self.oom_check()
return 0
self.who, self.allow_downgrade = ("fast", False)
if self.layer == "pve" and self.mode == "apply":
self.who, self.allow_downgrade = self.docker_authority(plan, op_name=PVE_SIGNED_OP)
self.report["authority"] = self.who
if self.layer == "docker" and self.mode == "apply":
self.who, self.allow_downgrade = self.docker_authority(plan)
if self.live_restore() != "true":
@@ -964,7 +1013,7 @@ class Apply:
log(f"os-apply: START release={plan.get('release_id')} layer={self.layer}" +
(f":{self.vmid}" if self.layer != "host" else "") +
f" lane={plan.get('lane', 'fast')} mode={self.mode} select={self.select} packages={len(plan.get('packages', []))}" +
(f" authority={self.who}{' UNDO' if self.allow_downgrade else ''}" if self.layer == "docker" else ""))
(f" authority={self.who}{' UNDO' if self.allow_downgrade else ''}" if self.layer in ("docker", "pve") else ""))
if self.apt_lock_held():
raise Refused("R9", f"another apt/dpkg holds the lock on the {self.layer}")
self.report["health_before"] = self.health()
@@ -988,10 +1037,91 @@ class Apply:
if self.layer == "docker":
rc_v, out_v, _ = self.g(["docker", "version", "--format", "{{.Server.Version}}"], timeout=60)
self.report["docker_engine"] = out_v.strip() if rc_v == 0 and out_v.strip() else "unknown"
if self.layer == "pve":
self.report["pve_manager"] = self.pve_manager()
self.report["reboot_scanned"] = "reboot_needed" in self.report
self.report["health_after"] = self.health()
if self.layer == "docker" and self.mode == "apply":
# R-528 (`09` decision 157): does the engine report a memory kill? Reported only — never the outcome.
self.report["oom_check"] = self.oom_check()
return 0
# ---------- R-528: the memory-kill check ----------
def oom_check(self):
"""Run one throwaway container over its memory cap in the guest and read what the engine says about it.
Never raises: any failure becomes result "error". The container is ALWAYS removed (finally), and a failed
removal is named in the detail."""
t = OOMCHECK_TIMEOUTS
res = {"result": "error", "oom_killed": False, "oom_event": False, "exit_code": None, "image": None, "detail": ""}
log = self.r.log
try:
rc, out, err = self.g(["docker", "inspect", "-f", "{{.Config.Image}}", "felhom-controller"], timeout=t["image"])
except Exception as e:
rc, out, err = -1, "", str(e)
img = out.strip() if rc == 0 else ""
if not img or any(c.isspace() for c in img):
res["detail"] = f"the controller's image could not be read (rc={rc}): {(err or out).strip()[:200]}"
log(f"os-apply: OOM-CHECK error — {res['detail']}")
return res
res["image"] = img
name = f"{OOMCHECK_PREFIX}{os.getpid()}-{os.urandom(4).hex()}"
notes = []
try:
t0 = self.guest_epoch()
if t0 is None:
raise RuntimeError("the guest clock could not be read")
rrc, rout, rerr = self.g(["docker", "run", "--name", name, "--pull", "never", "--network", "none",
"--memory", "64m", "--memory-swap", "64m", "--label", "felhom.oomcheck=1",
"--entrypoint", "sh", img, "-c", OOMCHECK_SCRIPT], timeout=t["run"])
self.r.sleep(OOMCHECK_SETTLE) # the engine publishes the oom event a moment after the run returns (F1/F2)
t1 = self.guest_epoch()
if t1 is None:
t1 = t0 + t["run"] + OOMCHECK_SETTLE + 1
irc, iout, ierr = self.g(["docker", "inspect", "-f", "{{.State.OOMKilled}} {{.State.ExitCode}}", name], timeout=t["inspect"])
if irc != 0:
raise RuntimeError(f"the check container could not be inspected (run rc={rrc}: {(rerr or rout).strip()[:120]}; "
f"inspect rc={irc}: {(ierr or iout).strip()[:120]})")
f = iout.split()
res["oom_killed"] = bool(f) and f[0] == "true"
try:
res["exit_code"] = int(f[1]) if len(f) > 1 else None
except ValueError:
res["exit_code"] = None
erc, eout, eerr = self.g(["docker", "events", "--since", str(t0 - 1), "--until", str(t1 + 1),
"--filter", f"container={name}", "--filter", "event=oom",
"--format", "{{.Action}}"], timeout=t["events"])
if erc != 0:
notes.append(f"the event read failed (rc={erc}): {(eerr or eout).strip()[:120]}")
res["oom_event"] = erc == 0 and any(l.strip() == "oom" for l in eout.splitlines())
if res["oom_killed"] and res["oom_event"]:
res["result"] = "pass"
notes.insert(0, "the engine reported the memory kill: OOMKilled=true and the oom event")
else:
res["result"] = "fail"
miss = [w for w, ok in (("OOMKilled=true", res["oom_killed"]), ("the oom event", res["oom_event"])) if not ok]
notes.insert(0, f"the engine did not report the memory kill: missing {' and '.join(miss)} (exit code {res['exit_code']})")
except Exception as e:
res["result"] = "error"
notes.insert(0, f"the check could not finish: {type(e).__name__}: {str(e)[:200]}")
finally:
try:
mrc, mout, merr = self.g(["docker", "rm", "-f", name], timeout=t["rm"])
if mrc != 0 and "no such container" not in (merr + mout).lower():
notes.append(f"the check container {name} could not be removed (rc={mrc}): {(merr or mout).strip()[:120]}")
except Exception as e:
notes.append(f"the check container {name} could not be removed: {type(e).__name__}: {str(e)[:120]}")
res["detail"] = "; ".join(notes)
log(f"os-apply: OOM-CHECK result={res['result']} oom_killed={res['oom_killed']} oom_event={res['oom_event']} "
f"exit={res['exit_code']} image={img} — {res['detail']}")
return res
def guest_epoch(self):
rc, out, _ = self.g(["date", "+%s"], timeout=OOMCHECK_TIMEOUTS["clock"])
try:
return int(out.strip()) if rc == 0 else None
except ValueError:
return None
def dpkg_state(self):
"""`dpkg --audit` AND dpkg's update journal, in ONE call (R-876, agent v0.145.0). A crash in the middle of an
install can leave `/var/lib/dpkg/updates/` non-empty while `--audit` reads clean — measured on demo-hp
@@ -1058,10 +1188,26 @@ class Apply:
return [{"name": p["name"], "version": p["to"], "origin": DOCKER_ORIGIN} for p in pend
if p["from"] is not None and p["name"] in DOCKER_NAMES and self.origin_name(p["origin"]) == {DOCKER_ORIGIN}]
def pending_pve(self):
"""Ring 0 (select pending-pve): the newest pending version of each INSTALLED Proxmox-origin package, never a
kernel / boot / firmware / microcode name (HOST_SLOW_RE, R-836's lane), never another origin."""
rc, pend, remv, _ = self.simulate(["dist-upgrade"])
return [{"name": p["name"], "version": p["to"], "origin": PVE_ORIGIN} for p in pend
if p["from"] is not None and not HOST_SLOW_RE.match(p["name"]) and p["name"] not in DOCKER_NAMES
and self.origin_name(p["origin"]) == {PVE_ORIGIN}]
def pve_manager(self):
"""pveversion's pve-manager version ("unknown" when it cannot be read)."""
rc, out, _ = self.r.host(["pveversion"], 60)
m = re.match(r"^pve-manager/([^/\s]+)", out.strip()) if rc == 0 else None
return m.group(1) if m else "unknown"
def origin_ok(self, origin):
o = self.origin_name(origin)
if self.layer == "docker":
return o == {DOCKER_ORIGIN}
if self.layer == "pve":
return o == {PVE_ORIGIN}
return bool(o & set(FAST_ORIGINS))
def apply(self, plan):
@@ -1070,6 +1216,8 @@ class Apply:
packages = plan["packages"]
elif self.select == "pending-docker":
packages = self.pending_docker()
elif self.select == "pending-pve":
packages = self.pending_pve()
else:
packages = self.pending_fast()
cmp_op = "ne" if self.allow_downgrade else "gt"
@@ -1118,6 +1266,11 @@ class Apply:
raise Refused("R4", f"the plan would remove {', '.join(remv[:5])}")
want = dict(upgrade)
for p in sim:
if p["from"] is None and self.layer == "pve" and p["name"] in PVE_NEW_ALLOW and self.origin_ok(p["origin"]) \
and not HOST_SLOW_RE.match(p["name"]):
self.r.log(f"os-apply: NEW {p['name']}={p['to']} (on the pve lane's allow-list)")
self.report.setdefault("added", []).append({"name": p["name"], "version": p["to"]})
continue
if p["from"] is None:
raise Refused("R6", f"the plan would add a package that is not installed: {p['name']}")
if p["name"] not in want:
@@ -1128,7 +1281,7 @@ class Apply:
raise Refused("R5", f"{p['name']} would be downgraded {p['from']} -> {p['to']}")
if not self.origin_ok(p["origin"]):
raise Refused("R2", f"{p['name']} would come from {p['origin']}, not the {self.layer} layer's origin")
if self.layer == "host" and HOST_SLOW_RE.match(p["name"]):
if self.layer in ("host", "pve") and HOST_SLOW_RE.match(p["name"]):
raise Refused("R14", f"{p['name']} is a kernel / boot / firmware package — the host's slow lane")
need = self.download_bytes(args)
free = self.free_bytes()
+50 -1
View File
@@ -17,6 +17,8 @@ Verbs (each one sudoers line, exact-match pattern):
wg /var/lib/felhom-agent/wg/wg-felhom.conf -> /etc/wireguard/wg-felhom.conf (0600)
sshd-config /var/lib/felhom-agent/felhom-sshd/sshd_config -> /etc/felhom-sshd/sshd_config
sshd-key /var/lib/felhom-agent/felhom-sshd/authorized_keys.felhom-op -> /etc/felhom-sshd/authorized_keys/felhom-op
controller-image <vmid> the ref on STDIN -> /etc/felhom-controller-image INSIDE guest <vmid> (R-861 (a) A1): only
our registry + our repository + an x.y.z tag; the agent no longer has a `tee` grant
--self-check prints "felhom-priv-apply ok verbs=..." (the bundle's self-check)
Exit codes: 0 installed (or already identical), 2 usage, 3 refused (content or source), 4 install failed.
@@ -39,7 +41,14 @@ WG_SRC, WG_DEST = STATE + "/wg/wg-felhom.conf", "/etc/wireguard/wg-felhom.conf"
SSHD_SRC, SSHD_DEST = STATE + "/felhom-sshd/sshd_config", "/etc/felhom-sshd/sshd_config"
KEY_SRC, KEY_DEST = STATE + "/felhom-sshd/authorized_keys.felhom-op", "/etc/felhom-sshd/authorized_keys/felhom-op"
MAX_BYTES = 64 * 1024
VERBS = ("unit", "dnsmasq", "wg", "sshd-config", "sshd-key")
VERBS = ("unit", "dnsmasq", "wg", "sshd-config", "sshd-key", "controller-image")
# R-861 (a) A1 (`09` §3 decision 165): the SAME pattern as the agent's controllerImageRe (internal/localapi/
# controllerswap.go) — a compromised agent cannot hand the guest's bootstrap any other image. Pinned by
# configs/test_felhom_priv_apply.py ControllerImage.
CONTROLLER_IMAGE_RE = re.compile(r"^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$")
CONTROLLER_IMAGE_FILE = "/etc/felhom-controller-image"
CONTROLLER_IMAGE_MAX = 256
VMID_RE = re.compile(r"^[0-9]{1,9}$")
UNIT_NAME_RE = re.compile(r"^mnt-[A-Za-z0-9_.\\-]+\.(mount|automount)$")
DNSMASQ_TMP_RE = re.compile(r"^/tmp/felhom-resolver-[0-9]+\.conf$")
@@ -123,6 +132,16 @@ class Host:
pass
raise
def read_stdin(self, limit):
return sys.stdin.buffer.read(limit + 1)
def write_guest_image(self, vmid, data):
"""As root: `pct exec <vmid> -- tee <the fixed file>` with the checked ref on stdin (no shell)."""
r = subprocess.run(["/usr/sbin/pct", "exec", str(vmid), "--", "tee", CONTROLLER_IMAGE_FILE],
input=data, stdout=subprocess.DEVNULL, stderr=subprocess.PIPE, timeout=60)
if r.returncode != 0:
raise OSError(f"pct exec {vmid} tee exited {r.returncode}: {r.stderr.decode(errors='replace').strip()[:200]}")
def log(self, line):
print(line, file=sys.stderr)
try:
@@ -377,6 +396,34 @@ def plan(argv):
raise Refused("A1", f"wrong arguments for {v}")
def controller_image(rest, host):
"""R-861 (a) A1: read the ref on stdin, check it, write it INSIDE the guest as root."""
try:
if len(rest) != 1 or not VMID_RE.match(rest[0]):
raise Refused("A1", "usage: felhom-priv-apply controller-image <vmid> (the ref on stdin)")
raw = host.read_stdin(CONTROLLER_IMAGE_MAX)
if len(raw) > CONTROLLER_IMAGE_MAX:
raise Refused("I1", f"the image ref is longer than {CONTROLLER_IMAGE_MAX} bytes")
try:
text = raw.decode("ascii")
except UnicodeDecodeError:
raise Refused("I1", "the image ref is not ASCII")
ref = text[:-1] if text.endswith("\n") else text
if not CONTROLLER_IMAGE_RE.match(ref) or "\n" in ref:
raise Refused("I1", "the image ref is not gitea.dooplex.hu/admin/felhom-controller:<x.y.z>")
except Refused as e:
host.log(f"felhom-priv-apply: REFUSED [{e.rule}] controller-image {' '.join(rest)[:40]}: {e.reason}")
return 2 if e.rule == "A1" else 3
vmid = int(rest[0])
try:
host.write_guest_image(vmid, (ref + "\n").encode())
except (OSError, subprocess.SubprocessError) as e:
host.log(f"felhom-priv-apply: FAILED controller-image {vmid}: {e}")
return 4
host.log(f"felhom-priv-apply: WROTE controller-image {vmid} {ref}")
return 0
def main(argv, host=None):
host = host or Host()
if argv == ["--self-check"]:
@@ -396,6 +443,8 @@ def main(argv, host=None):
return 3
print("OK")
return 0
if argv and argv[0] == "controller-image":
return controller_image(argv[1:], host)
try:
verb, src, dest, mode, checker = plan(argv)
data = host.read_source(src)
+4 -1
View File
@@ -546,7 +546,10 @@ class Builder(unittest.TestCase):
r"/etc/felhom/[a-z.-]+)", text))
agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"}
trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust)
# Written by the appliance ISO's first boot (felhom.eu scripts/iso/felhom-bootstrap.sh), never by the installer:
# since installer 1.32.0 (R-275) the uninstall only NAMES them under KEPT.
iso_writes = {"/etc/felhom/.bootstrap-done", "/etc/felhom/appliance-pairing-code"}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust | iso_writes)
self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them")
# the limits drop-in is named through $AGENT_UNIT in the installer
self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS)
+303 -4
View File
@@ -83,7 +83,7 @@ class Fake:
return self.clock
def sleep(self, s):
pass
self.sleeps = getattr(self, "sleeps", []) + [(len(self.calls), s)] # (calls made before it, seconds)
def verify_sig(self, signers, key_id, ns, blob, sig):
self.verified = (signers, key_id, ns, blob, sig)
@@ -196,6 +196,27 @@ class Fake:
return 0, self.engine + "\n", ""
if cmd == "docker" and a[1:3] == ["ps", "-q"]:
return 0, "".join(i + "\n" for i in self.ids), ""
# R-528: the memory-kill check's engine (oom_image None = unreadable; oom_state / oom_event the engine's answer)
if cmd == "date" and a[1:] == ["+%s"]:
# each read is 3 s later than the last, so the order of the reads is visible in the values
self.oom_epochs = getattr(self, "oom_epochs", []) + [int(self.clock) + 3 * len(getattr(self, "oom_epochs", []))]
return 0, f"{self.oom_epochs[-1]}\n", ""
if cmd == "docker" and a[1:4] == ["inspect", "-f", "{{.Config.Image}}"]:
img = getattr(self, "oom_image", "gitea.dooplex.hu/admin/felhom-controller:0.300.0")
return (0, img + "\n", "") if img is not None else (1, "", "Error: No such object: felhom-controller")
if cmd == "docker" and a[1] == "run":
self.oom_runs = getattr(self, "oom_runs", []) + [a]
return 137, "", ""
if cmd == "docker" and a[1:4] == ["inspect", "-f", "{{.State.OOMKilled}} {{.State.ExitCode}}"]:
if getattr(self, "oom_inspect_raises", False):
raise subprocess.TimeoutExpired(a, 10)
return 0, getattr(self, "oom_state", "true 137") + "\n", ""
if cmd == "docker" and a[1] == "events":
self.oom_events_argv = a
return 0, ("oom\n" if getattr(self, "oom_event", True) else ""), ""
if cmd == "docker" and a[1:3] == ["rm", "-f"]:
self.oom_removed = getattr(self, "oom_removed", []) + a[3:]
return 0, a[3] + "\n", ""
if cmd == "docker" and a[1] == "inspect":
mounts = {"aaa111": "/felhom-controller|/var/run/docker.sock;/app/data;", "bbb222": "/app|/data;"}
return 0, mounts.get(a[-1], "/other|;") + "\n", ""
@@ -211,6 +232,8 @@ class Fake:
return (0, self.daemon_json, "") if self.daemon_json is not None else (1, "", "No such file")
if cmd == "cat" and a[1] == "/etc/debian_version":
return 0, "13.7\n", ""
if cmd == "pveversion":
return 0, f"pve-manager/{self.installed.get('pve-manager', '9.2.2')}/abcdef (running kernel: 7.0.14-20-pve)\n", ""
if cmd == "uname":
return 0, "7.0.14-20-pve\n", ""
if cmd == "apt-mark":
@@ -263,7 +286,7 @@ class Fake:
n, v = x.split("=", 1)
if v not in self.avail(n):
return 100, "", f"E: Version '{v}' for '{n}' was not found"
origin = "Docker CE:trixie" if n in osapply.DOCKER_NAMES else SEC if n == "openssl" else DEB
origin = getattr(self, "origins", {}).get(n) or ("Docker CE:trixie" if n in osapply.DOCKER_NAMES else SEC if n == "openssl" else DEB)
out += f"Inst {n} [{self.installed[n]}] ({v} {origin} [amd64])\n"
out += "".join(l + "\n" for l in self.extra_sim)
return 0, out, ""
@@ -997,8 +1020,6 @@ class RealSignatureCheck(unittest.TestCase):
self.assertNotEqual(r.verify_sig(self.signers, "someone-else", "felhom-op-v1", blob, sig), 0)
self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob, ns="other-ns")), 0)
if __name__ == "__main__":
unittest.main()
class UnsentReport(unittest.TestCase):
@@ -1179,3 +1200,281 @@ class CrashLeftTheJournal(unittest.TestCase):
self.assertEqual(rc, 0, rep)
self.assertTrue(any("INTERRUPTED" in l for l in f.logs), f.logs)
self.assertTrue(any(l.startswith("os-apply: REPAIR ") and l.endswith("forced") for l in f.logs), f.logs)
class OOMCheck(unittest.TestCase):
"""R-528 (`09` decision 157): after a Docker engine step the wrapper proves the engine reports a memory kill
(OOMKilled=true AND the `oom` event). Reported only; the hub decides. Red-proofs: audits/readback-2026-10-07/F/."""
def apply(self, **kw):
f = docker_fake(signed=signed_job())
for k, v in kw.items():
setattr(f, k, v)
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
return f, rep
def test_pass_when_oomkilled_and_the_event(self):
f, rep = self.apply()
oc = rep["oom_check"]
self.assertEqual(oc["result"], "pass", oc)
self.assertEqual((oc["oom_killed"], oc["oom_event"], oc["exit_code"]), (True, True, 137))
self.assertEqual(oc["image"], "gitea.dooplex.hu/admin/felhom-controller:0.300.0")
self.assertEqual(sorted(oc), ["detail", "exit_code", "image", "oom_event", "oom_killed", "result"])
run_argv = f.oom_runs[0]
name = run_argv[run_argv.index("--name") + 1]
self.assertTrue(re.match(r"^felhom-oomcheck-[0-9]+-[0-9a-f]{8}$", name), name)
for flag, val in (("--pull", "never"), ("--network", "none"), ("--memory", "64m"), ("--memory-swap", "64m"),
("--label", "felhom.oomcheck=1"), ("--entrypoint", "sh")):
self.assertEqual(run_argv[run_argv.index(flag) + 1], val, flag)
self.assertNotIn("-v", run_argv)
self.assertNotIn("-p", run_argv)
self.assertEqual(run_argv[-3:], ["gitea.dooplex.hu/admin/felhom-controller:0.300.0", "-c", osapply.OOMCHECK_SCRIPT])
self.assertIn(f"container={name}", f.oom_events_argv)
self.assertIn("event=oom", f.oom_events_argv)
self.assertEqual(f.oom_removed, [name], "the check container must be removed")
self.assertTrue(rep["health_after"], "health is read before the check")
# bounded: every call has a timeout and the clock is read twice — the worst case stays within 90 s
self.assertLessEqual(sum(osapply.OOMCHECK_TIMEOUTS.values()) + osapply.OOMCHECK_TIMEOUTS["clock"]
+ osapply.OOMCHECK_SETTLE, 90)
def test_events_window_ends_after_the_settle_wait(self):
# Measured (F1/F2): an --until taken right after the run missed the oom event. The window must end after a
# wait of >= 2 s that comes AFTER the run, at the guest epoch read after that wait, + 1.
f, rep = self.apply()
run_i = next(i for i, c in enumerate(f.calls) if c[-1][:2] == ["docker", "run"])
date_i = [i for i, c in enumerate(f.calls) if c[-1] == ["date", "+%s"]]
waits = [(i, s) for i, s in getattr(f, "sleeps", []) if i > run_i]
self.assertTrue(waits and waits[0][1] >= 2, f"no settle wait after the run: {getattr(f, 'sleeps', None)}")
self.assertTrue(date_i[-1] >= waits[0][0], "the end epoch must be read after the wait")
ev = f.oom_events_argv
since, until = int(ev[ev.index("--since") + 1]), int(ev[ev.index("--until") + 1])
self.assertEqual(until, f.oom_epochs[-1] + 1, "until = the guest epoch read after the wait, + 1")
self.assertGreater(until, f.oom_epochs[0] + 1)
self.assertLess(since, f.oom_epochs[0] + 1)
def test_oomkilled_false_is_fail(self):
f, rep = self.apply(oom_state="false 0")
oc = rep["oom_check"]
self.assertEqual(oc["result"], "fail", oc)
self.assertIn("OOMKilled=true", oc["detail"])
self.assertTrue(oc["oom_event"])
self.assertEqual(len(f.oom_removed), 1)
def test_no_event_is_fail(self):
f, rep = self.apply(oom_event=False)
oc = rep["oom_check"]
self.assertEqual(oc["result"], "fail", oc)
self.assertIn("the oom event", oc["detail"])
self.assertTrue(oc["oom_killed"])
def test_image_unreadable_is_error(self):
f, rep = self.apply(oom_image=None)
oc = rep["oom_check"]
self.assertEqual(oc["result"], "error", oc)
self.assertIsNone(oc["image"])
self.assertIn("image could not be read", oc["detail"])
self.assertFalse(hasattr(f, "oom_runs"), "no container is started without an image")
def test_container_removed_even_when_inspect_raises(self):
f, rep = self.apply(oom_inspect_raises=True)
oc = rep["oom_check"]
self.assertEqual(oc["result"], "error", oc)
self.assertEqual(len(f.oom_removed), 1, "the container must be removed even when inspect raised")
self.assertTrue(f.oom_removed[0].startswith(osapply.OOMCHECK_PREFIX))
self.assertNotIn("failed", rep, "the check never turns the step into a failure")
def test_never_on_guest_or_host_or_in_health_mode(self):
g = Fake()
rc, rep = run(g)
self.assertEqual(rc, 0, rep)
h = Fake()
h.plan["layer"] = "host"
h.plan["packages"] = [{"name": "bash", "version": "5.2.37-2+b10", "origin": "Debian"}]
h.installed["bash"] = "5.2.37-2+b9"
h.live["bash"] = {"5.2.37-2+b10"}
rc_h, rep_h = run(h)
self.assertEqual(rc_h, 0, rep_h)
d = docker_fake(signed=signed_job())
d.plan["mode"] = "health"
rc_d, rep_d = run(d)
self.assertEqual(rc_d, 0, rep_d)
for f, r in ((g, rep), (h, rep_h), (d, rep_d)):
self.assertNotIn("oom_check", r)
self.assertFalse(hasattr(f, "oom_runs"), r.get("layer"))
self.assertFalse(any(c[-1][:2] == ["docker", "run"] for c in f.calls))
def test_mode_oom_check_runs_only_the_check(self):
f = docker_fake() # no authority: the check changes nothing, so it needs none
f.plan["mode"], f.plan["packages"] = "oom-check", []
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["oom_check"]["result"], "pass", rep)
self.assertFalse(any("apt-get" in c[-1] or "dpkg-query" in c[-1] for c in f.calls), f.calls)
self.assertEqual(len(f.oom_removed), 1)
def test_mode_oom_check_keeps_the_refusals(self):
f = docker_fake()
f.plan["mode"], f.plan["packages"] = "oom-check", []
f.files["/etc/pve/lxc/9201.conf"] = "arch: amd64\n"
rc, rep = run(f)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R10"), rep)
self.assertFalse(hasattr(f, "oom_runs"))
g = Fake()
g.plan["mode"] = "oom-check"
rc, rep = run(g)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R11"), rep)
PVE = "Proxmox Debian Repository:stable"
PVE_SET = [{"name": "pve-manager", "version": "9.2.21", "origin": "Proxmox Debian Repository"},
{"name": "libpve-common-perl", "version": "9.1.9", "origin": "Proxmox Debian Repository"}]
def pve_fake(signed=None, ring0=False):
f = Fake()
f.installed.update({"pve-manager": "9.2.2", "libpve-common-perl": "9.1.1", "proxmox-kernel-helper": "9.0.4"})
f.live["pve-manager"] = {"9.2.21", "9.2.2"}
f.live["libpve-common-perl"] = {"9.1.9", "9.1.1"}
f.origins = {"pve-manager": PVE, "libpve-common-perl": PVE, "proxmox-kernel-helper": PVE, "shim-signed": PVE,
"proxmox-firewall-data": PVE}
f.plan = {"release_id": "os-pve-t1", "layer": "pve", "lane": "slow", "vmid": 9201, "mode": "apply",
"packages": [dict(p) for p in PVE_SET]}
if ring0:
f.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": True})
if signed is not None:
f.plan["signed"] = signed
return f
class PVELane(unittest.TestCase):
"""R-812 option A (`09` §3 decision 163): the host's Proxmox USERSPACE packages, slow lane, no kernel / boot /
firmware / microcode (R14), no removal, a new package only from PVE_NEW_ALLOW. Red-proof: audits/day-2026-10-07/B/."""
def refused(self, f, code):
rc, rep = run(f)
self.assertEqual(rc, 2, rep)
self.assertEqual(rep["refused"]["code"], code, rep)
self.assertEqual(f.installed["pve-manager"], "9.2.2", "nothing may be installed on a refusal")
return rep
# THE RED TEST (design-R-812 §5): before the pve layer existed this plan was refused R12.
def test_pve_manager_plan_is_installed_on_the_host(self):
f = pve_fake(signed=signed_job(packages=PVE_SET, op="os_pve_step"))
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(f.installed["pve-manager"], "9.2.21")
self.assertEqual(rep["authority"], "signed")
self.assertEqual(rep["pve_manager"], "9.2.21", "pveversion after the step is reported")
inst = [c for c in f.calls if c[0] == "host" and "install" in c[1] and "-s" not in c[1] and "-f" not in c[1]
and "--print-uris" not in c[1]]
self.assertTrue(inst, "the pve layer installs on the HOST")
self.assertFalse([c for c in f.calls if c[0] == "guest" and "install" in c[2]], "nothing installed in the guest")
self.assertEqual(sorted(rep["health_after"]["host_services"]), sorted(osapply.HOST_SERVICES))
def test_kernel_in_a_pve_plan_is_refused(self):
f = pve_fake(ring0=True)
f.plan["packages"].append({"name": "proxmox-kernel-7.0", "version": "7.0.14-20", "origin": "Proxmox Debian Repository"})
self.refused(f, "R14")
def test_shim_in_a_pve_plan_is_refused(self):
f = pve_fake(ring0=True)
f.plan["packages"].append({"name": "shim-signed", "version": "1.47+pmx1", "origin": "Proxmox Debian Repository"})
self.refused(f, "R14")
def test_kernel_pulled_in_by_the_simulation_is_refused(self):
f = pve_fake(ring0=True)
f.installed["proxmox-kernel-helper"] = "9.0.4"
f.extra_sim = ["Inst proxmox-kernel-helper [9.0.4] (9.0.6 Proxmox Debian Repository:stable [all])"]
self.refused(f, "R6") # not in the plan — refused before the name check; R14 below when it IS in the plan
g = pve_fake(ring0=True)
g.live["proxmox-kernel-helper"] = {"9.0.6"}
g.plan["packages"] = [dict(PVE_SET[0])]
g.extra_sim = ["Inst proxmox-kernel-helper [9.0.4] (9.0.6 Proxmox Debian Repository:stable [all])"]
# a kernel-helper the plan did not name is R6; the R14 name check covers a listed one (test above)
self.refused(g, "R6")
def test_debian_package_in_a_pve_plan_is_refused(self):
f = pve_fake(ring0=True)
f.plan["packages"].append({"name": "libc6", "version": "2.41-12+deb13u4", "origin": "Debian"})
self.refused(f, "R2")
def test_debian_origin_in_the_pve_simulation_is_refused(self):
f = pve_fake(ring0=True)
f.origins["libpve-common-perl"] = DEB
self.refused(f, "R2")
def test_docker_package_in_a_pve_plan_is_refused(self):
f = pve_fake(ring0=True)
f.plan["packages"].append({"name": "docker-ce", "version": "5:29.8.2-1~debian.13~trixie", "origin": "Proxmox Debian Repository"})
self.refused(f, "R2")
def test_pve_in_the_fast_lane_is_refused(self):
f = pve_fake(ring0=True)
f.plan["lane"] = "fast"
self.refused(f, "R3")
def test_no_authority_is_refused(self):
self.refused(pve_fake(), "R3")
def test_docker_signed_op_does_not_authorize_a_pve_step(self):
self.refused(pve_fake(signed=signed_job(packages=PVE_SET, op="os_docker_step")), "R3")
def test_pve_on_a_byo_box_is_refused(self):
f = pve_fake(ring0=True)
f.files[osapply.INSTALL_STATE] = json.dumps({"mode": "byo"})
self.refused(f, "R12")
def test_unlisted_new_package_is_refused(self):
f = pve_fake(ring0=True)
f.extra_sim = ["Inst proxmox-new-thing (1.0 Proxmox Debian Repository:stable [all])"]
self.refused(f, "R6")
def test_allow_listed_new_package_is_accepted(self):
f = pve_fake(ring0=True)
f.extra_sim = ["Inst proxmox-firewall-data (0.1 Proxmox Debian Repository:stable [all])"]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(f.installed["pve-manager"], "9.2.21")
def test_allow_new_is_never_an_open_door_in_the_fast_lane(self):
f = Fake()
f.plan["layer"] = "host"
f.extra_sim = ["Inst proxmox-firewall-data (0.1 Proxmox Debian Repository:stable [all])"]
rc, rep = run(f)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R6"), rep)
def test_undo_is_refused(self):
f = pve_fake(signed=signed_job(packages=PVE_SET, op="os_pve_step"))
f.plan["undo"] = True
self.refused(f, "R5")
def test_ring0_pending_pve_takes_only_installed_proxmox_userspace(self):
f = pve_fake(ring0=True)
f.plan["select"], f.plan["packages"] = "pending-pve", []
f.installed.update({"linux-image-amd64": "6.12.1", "tailscale": "1.102.2"})
f.pending_sim = [
"Inst pve-manager [9.2.2] (9.2.21 Proxmox Debian Repository:stable [amd64])",
"Inst libpve-common-perl [9.1.1] (9.1.9 Proxmox Debian Repository:stable [all])",
"Inst proxmox-kernel-helper [9.0.4] (9.0.6 Proxmox Debian Repository:stable [all])",
"Inst proxmox-kernel-7.0.14-20-pve-signed (7.0.14-20 Proxmox Debian Repository:stable [amd64])",
"Inst proxmox-firewall-data (0.1 Proxmox Debian Repository:stable [all])",
"Inst libc6 [2.41-12+deb13u3] (2.41-12+deb13u4 Debian:13.7/stable [amd64])",
"Inst tailscale [1.102.2] (1.102.5 Tailscale:pkgs.tailscale.com [amd64])",
]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["authority"], "ring0")
self.assertEqual(sorted(u["name"] for u in rep["upgraded"]), ["libpve-common-perl", "pve-manager"],
"pending-pve: installed, Proxmox-origin, never kernel/boot/firmware, never Debian or other origins")
self.assertEqual(f.installed["libc6"], "2.41-12+deb13u3")
def test_select_pending_pve_needs_the_pve_layer(self):
f = Fake()
f.plan["layer"], f.plan["select"], f.plan["packages"] = "host", "pending-pve", []
rc, rep = run(f)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R11"), rep)
if __name__ == "__main__":
unittest.main()
+60
View File
@@ -56,6 +56,18 @@ class FakeHost:
def log(self, line):
self.logs.append(line)
# controller-image (R-861 (a) A1): the ref arrives on stdin and is written INSIDE the guest by root.
stdin = b""
guest_writes = None
def read_stdin(self, limit):
return self.stdin[:limit + 1]
def write_guest_image(self, vmid, data):
if self.guest_writes is None:
self.guest_writes = []
self.guest_writes.append((vmid, data))
LOCAL_UNIT = """# Managed by felhom-agent — do not edit by hand.
[Unit]
@@ -316,5 +328,53 @@ class Refuses(unittest.TestCase):
self.refused(h, ["wg", "/etc/shadow"], "A1")
class ControllerImage(unittest.TestCase):
"""R-861 (a) A1 (`09` §3 decision 165): the agent can no longer `tee` any image ref into the guest. The root verb
reads the ref on stdin, requires our registry + our repository + an x.y.z tag, and writes the guest file itself.
RED-PROOF: on the pre-A1 wrapper `controller-image` is not a verb (A1 usage, rc 2) — the accepted case fails."""
def go(self, ref, *argv):
h = FakeHost()
h.stdin = ref.encode() if isinstance(ref, str) else ref
return h, run(h, *(argv or ("controller-image", "9201")))
def test_our_controller_ref_is_written_in_the_guest(self):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0\n")
self.assertEqual(rc, 0, h.logs)
self.assertEqual(h.guest_writes, [(9201, b"gitea.dooplex.hu/admin/felhom-controller:0.301.0\n")])
def test_a_foreign_image_is_refused(self):
for ref in ("docker.io/library/alpine:latest\n", "alpine\n",
"gitea.dooplex.hu/admin/felhom-controller:latest\n",
"gitea.dooplex.hu/admin/other:0.1.0\n",
"evil.example/admin/felhom-controller:0.301.0\n",
"gitea.dooplex.hu/admin/felhom-controller:0.301.0\nalpine\n",
"gitea.dooplex.hu/admin/felhom-controller:0.301.0 x\n",
"", "\n"):
h, rc = self.go(ref)
self.assertEqual(rc, 3, f"{ref!r} was accepted")
self.assertFalse(h.guest_writes, f"{ref!r} wrote the guest file")
self.assertTrue(any("[I1]" in l for l in h.logs), h.logs)
def test_oversize_stdin_is_refused(self):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0" + " " * 300)
self.assertEqual(rc, 3)
self.assertFalse(h.guest_writes)
def test_vmid_must_be_numeric(self):
for argv in (("controller-image", "9201;id"), ("controller-image", "-1"), ("controller-image",),
("controller-image", "9201", "9202")):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0\n", *argv)
self.assertIn(rc, (2, 3), argv)
self.assertFalse(h.guest_writes, argv)
def test_self_check_names_the_verb(self):
import io, contextlib
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
pa.main(["--self-check"])
self.assertIn("controller-image", buf.getvalue())
if __name__ == "__main__":
unittest.main(verbosity=2)
@@ -241,3 +241,23 @@ func TestNewestArchiveTime_DistinctPhantomsEachAnnounced(t *testing.T) {
t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String())
}
}
// R-99 (`09` §3 decision 140): the WARN for a PBS phantom ends with the cleanup runbook, so whoever sees it knows the
// one sanctioned way to remove it; a tiny archive on a dir storage is not a PBS phantom and gets no pointer.
// RED-PROOF: drop the `msg += phantomCleanupPointer` line → "the PBS phantom WARN does not end with the runbook pointer".
func TestRejectedArchiveWarnNamesTheCleanupRunbook(t *testing.T) {
var buf bytes.Buffer
r := runnerWithContent(t, &buf, []proxmox.StorageContent{phantomEntry(), goodPBSEntry()})
if _, _, err := r.NewestArchiveTime(context.Background(), 9201); err != nil {
t.Fatal(err)
}
const want = "INCOMPLETE archive when computing tier freshness — it is not a successful backup — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
if !strings.Contains(buf.String(), want) {
t.Errorf("the PBS phantom WARN does not end with the runbook pointer:\n%s", buf.String())
}
local := phantomEntry()
local.Format, local.VolID = "tar.zst", "local:backup/vzdump-lxc-9201-2026_07_28-05_31_14.tar.zst"
if got := rejectedArchiveMessage(local); strings.Contains(got, "pbs-phantom-cleanup") {
t.Errorf("a dir-storage archive got the PBS runbook pointer: %s", got)
}
}
+131
View File
@@ -0,0 +1,131 @@
package backup
import (
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// BackupSuccessState persists the newest SUCCESSFUL whole-guest backup per tier and guest (R-894).
//
// Why it exists. The due-check (`localapi` handleBackupDue) asks the tier's storage when a backup last
// landed (R-84) and falls back to the in-memory record when the storage cannot be read. The in-memory
// record is empty after an agent restart (Store, R-348), so "storage unreadable" right after a restart
// read as "no record — DUE". Measured 2026-10-05 on demo-hp: the agent restarted at 04:57, the off-site
// storage answered "Can't connect" at 06:25, the 7-day tier — last copy 2026-10-01 — read DUE, the
// controller asked, and vzdump failed. This file is the last known copy the fallback reads instead.
//
// It is read ONLY when the storage cannot be read. A storage that answers is the ground truth and wins,
// in both directions: an archive found there counts, and an archive absent there is absent even when
// this file remembers a success (a pruned or deleted archive must make the tier due — the same reason
// R-84 chose the storage over a persisted record). Pinned by
// TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers.
//
// Only SUCCESSES are written (the RestoreTestState rule): a failure must stay due and be retried, so a
// record of a failure has no reader.
type BackupSuccessState struct {
path string
mu sync.Mutex
last map[string]savedSuccess // key(target, vmid) → the newest success
}
type savedSuccess struct {
target string
vmid int
at time.Time
}
// backupSuccessJSON is one entry on disk.
type backupSuccessJSON struct {
Target string `json:"target"`
VMID int `json:"vmid"`
StartedAt string `json:"started_at"`
}
func backupStateKey(target string, vmid int) string { return target + "/" + strconv.Itoa(vmid) }
// NewBackupSuccessState opens (or creates) the state at path. A missing or unreadable file degrades to
// "nothing known" — the pre-R-894 behaviour, which is DUE — and never wedges the daemon.
func NewBackupSuccessState(path string) *BackupSuccessState {
s := &BackupSuccessState{path: path, last: map[string]savedSuccess{}}
data, err := os.ReadFile(path)
if err != nil {
return s
}
var entries []backupSuccessJSON
if json.Unmarshal(data, &entries) != nil {
return s
}
for _, e := range entries {
t, perr := time.Parse(time.RFC3339, e.StartedAt)
if perr != nil {
continue // one unreadable entry must not lose the others
}
s.last[backupStateKey(e.Target, e.VMID)] = savedSuccess{target: e.Target, vmid: e.VMID, at: t.UTC()}
}
return s
}
// RecordBackupSuccess saves b when it is a success newer than the one on file. target is the tier the
// job ran on (the due-check's key); a failure or an unparseable time is ignored.
func (s *BackupSuccessState) RecordBackupSuccess(target string, b hub.Backup) error {
if s == nil || !b.Success {
return nil
}
t, err := time.Parse(time.RFC3339, b.StartedAt)
if err != nil {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
k := backupStateKey(target, b.VMID)
if old, ok := s.last[k]; ok && !t.After(old.at) {
return nil
}
s.last[k] = savedSuccess{target: target, vmid: b.VMID, at: t.UTC()}
return s.saveLocked()
}
// LastKnownSuccess returns the newest saved success for this tier and guest (ok=false = none on file).
func (s *BackupSuccessState) LastKnownSuccess(target string, vmid int) (time.Time, bool) {
if s == nil {
return time.Time{}, false
}
s.mu.Lock()
defer s.mu.Unlock()
e, ok := s.last[backupStateKey(target, vmid)]
return e.at, ok
}
func (s *BackupSuccessState) saveLocked() error {
entries := make([]backupSuccessJSON, 0, len(s.last))
for _, e := range s.last {
entries = append(entries, backupSuccessJSON{Target: e.target, VMID: e.vmid, StartedAt: e.at.Format(time.RFC3339)})
}
// Deterministic file content (Go's map order is random).
sort.Slice(entries, func(i, j int) bool {
if entries[i].Target != entries[j].Target {
return entries[i].Target < entries[j].Target
}
return entries[i].VMID < entries[j].VMID
})
data, err := json.MarshalIndent(entries, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(s.path), 0o755); err != nil {
return err
}
tmp := s.path + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, s.path)
}
+59
View File
@@ -0,0 +1,59 @@
package backup
import (
"os"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-894: the on-disk newest success per tier survives a restart (a new state from the same file).
func TestBackupSuccessState_SurvivesRestart(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
s := NewBackupSuccessState(path)
at := time.Date(2026, 10, 1, 20, 15, 0, 0, time.UTC)
if err := s.RecordBackupSuccess("felhom-pbs", hub.Backup{VMID: 9201, Success: true, StartedAt: at.Format(time.RFC3339)}); err != nil {
t.Fatal(err)
}
got, ok := NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 9201)
if !ok || !got.Equal(at) {
t.Fatalf("after a restart the saved copy must read back; got %v ok=%v", got, ok)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("local", 9201); ok {
t.Fatal("another tier must not borrow this tier's copy")
}
}
// Only a NEWER success replaces the saved one; failures and unparseable times are ignored.
func TestBackupSuccessState_KeepsNewestSuccessOnly(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
s := NewBackupSuccessState(path)
newer := time.Date(2026, 10, 5, 0, 0, 0, 0, time.UTC)
older := newer.Add(-48 * time.Hour)
for _, b := range []hub.Backup{
{VMID: 1, Success: true, StartedAt: newer.Format(time.RFC3339)},
{VMID: 1, Success: true, StartedAt: older.Format(time.RFC3339)}, // older: ignored
{VMID: 1, Success: false, StartedAt: newer.Add(time.Hour).Format(time.RFC3339)}, // failure: ignored
{VMID: 1, Success: true, StartedAt: "not-a-time"}, // unparseable: ignored
} {
if err := s.RecordBackupSuccess("t", b); err != nil {
t.Fatal(err)
}
}
if got, _ := NewBackupSuccessState(path).LastKnownSuccess("t", 1); !got.Equal(newer) {
t.Fatalf("want the newest success %v, got %v", newer, got)
}
}
// A corrupt file degrades to "nothing known" (the pre-R-894 DUE answer), never a crash.
func TestBackupSuccessState_CorruptFileIsEmpty(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
if err := os.WriteFile(path, []byte("{not json"), 0o600); err != nil {
t.Fatal(err)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("t", 1); ok {
t.Fatal("a corrupt file must read as nothing known")
}
}
+60
View File
@@ -0,0 +1,60 @@
package backup
import (
"context"
"sort"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// ForeignKeyLedger (R-366 slice 2, `09` §3 decision 168) holds, per backup tier, the whole-guest archives the
// restore-test pick skipped because another key wrote them (R-727 — an earlier install of this box). Since R-727 the
// skip was one INFO log line per archive and nothing else, so after a reinstall the operator was never told that the
// box's older whole-guest copies are unreadable to it. The host report carries this ledger; the hub raises ONE
// operator event when it changes.
//
// It reports nil until a tier has been evaluated since the agent started, so a restart does not read as "the set
// changed to empty" (the hub keeps its last state for an absent field).
type ForeignKeyLedger struct {
mu sync.Mutex
byTarget map[string]hub.ForeignKeyArchives
}
// NewForeignKeyLedger builds an empty ledger.
func NewForeignKeyLedger() *ForeignKeyLedger {
return &ForeignKeyLedger{}
}
func (l *ForeignKeyLedger) set(target string, n int, oldest, newest int64) {
l.mu.Lock()
defer l.mu.Unlock()
if l.byTarget == nil {
l.byTarget = map[string]hub.ForeignKeyArchives{}
}
e := hub.ForeignKeyArchives{Target: target, Count: n}
if n > 0 {
e.Oldest = time.Unix(oldest, 0).UTC().Format(time.RFC3339)
e.Newest = time.Unix(newest, 0).UTC().Format(time.RFC3339)
}
l.byTarget[target] = e
}
// ForeignKeyArchives implements hub.ForeignKeyArchiveReporter: nil before any evaluation; otherwise the tiers that
// hold such archives (`Tiers` empty, never nil, when none do), sorted by tier.
func (l *ForeignKeyLedger) ForeignKeyArchives(context.Context) *hub.ForeignKeyArchivesStanza {
l.mu.Lock()
defer l.mu.Unlock()
if l.byTarget == nil {
return nil
}
out := []hub.ForeignKeyArchives{}
for _, e := range l.byTarget {
if e.Count > 0 {
out = append(out, e)
}
}
sort.Slice(out, func(i, j int) bool { return out[i].Target < out[j].Target })
return &hub.ForeignKeyArchivesStanza{Tiers: out}
}
@@ -0,0 +1,69 @@
package backup
import (
"context"
"encoding/json"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-366 slice 2 (`09` §3 decision 168) — the restore-test's skip of an archive written with another key stops being
// silent: the pick records, per tier, how many it skipped and their time range, and the host report carries it.
//
// COMPANION RED-PROOF (observed): remove the `r.foreign.set(...)` call from PickSettledRestoreCandidateOn → this fails
// with "after one evaluation the ledger must report felhom-pbs: 2 archives …; got []". Restored.
func TestR366_PickRecordsArchivesWrittenWithAnotherKey(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T17:27:32Z", Content: "backup", VMID: 9201, Size: 4774114206, CTime: 1789579652, Encrypted: earlierBox2},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z", Content: "backup", VMID: 9201, Size: 20811501236, CTime: 1789595994, Encrypted: earlierBox1},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey},
},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
l := NewForeignKeyLedger()
r.SetForeignKeyLedger(l)
if got := l.ForeignKeyArchives(context.Background()); got != nil {
t.Fatalf("before any evaluation the ledger must be nil (the hub keeps its state); got %v", got)
}
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err != nil {
t.Fatal(err)
}
st := l.ForeignKeyArchives(context.Background())
want := hub.ForeignKeyArchives{Target: "felhom-pbs", Count: 2, Oldest: "2026-09-16T17:27:32Z", Newest: "2026-09-16T21:59:54Z"}
var got []hub.ForeignKeyArchives
if st != nil {
got = st.Tiers
}
if len(got) != 1 || got[0] != want {
t.Fatalf("after one evaluation the ledger must report felhom-pbs: 2 archives 2026-09-16T17:27:32Z…21:59:54Z; got %v", got)
}
}
// Evaluated and none found → the stanza with `tiers: []`; not evaluated → no stanza at all. Never a null on the wire.
func TestR366_EvaluatedWithNoneIsAnEmptyList(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey}},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
l := NewForeignKeyLedger()
r.SetForeignKeyLedger(l)
b, _ := json.Marshal(hub.HostReport{ForeignKeyArchives: l.ForeignKeyArchives(context.Background())})
if strings.Contains(string(b), "foreign_key_archives") {
t.Fatalf("not evaluated must omit the stanza; got %s", b)
}
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err != nil {
t.Fatal(err)
}
b, _ = json.Marshal(hub.HostReport{ForeignKeyArchives: l.ForeignKeyArchives(context.Background())})
if !strings.Contains(string(b), `"foreign_key_archives":{"tiers":[]}`) {
t.Fatalf("evaluated with none must report tiers: []; got %s", b)
}
}
+33 -1
View File
@@ -66,8 +66,13 @@ type BackupRunner struct {
// due-check is served from the local-API handler goroutines.
rejectedMu sync.Mutex
rejected map[string]struct{}
// foreign (R-366 slice 2) records, per tier, the archives the pick skipped as another key's. nil = not wired.
foreign *ForeignKeyLedger
}
// SetForeignKeyLedger wires the R-366 slice-2 ledger the host report reads.
func (r *BackupRunner) SetForeignKeyLedger(l *ForeignKeyLedger) { r.foreign = l }
// NewBackupRunner builds a runner. mode defaults to snapshot (works for a stopped guest and
// for lvm-thin); the caller may pass ModeStop for storages without snapshot support. retention is the
// per-run prune spec ("keep-last=N", or "" to never prune) — only the periodic local backup sets it.
@@ -370,6 +375,8 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
var best string
var bestCTime int64 = -1
known := map[int]bool{} // vmid → the guest exists on this node (asked once per vmid per pick)
var foreignN int // R-366 slice 2: archives skipped as another key's, and their time range
var foreignMin, foreignMax int64
for _, e := range contents {
if e.Content != "backup" {
continue
@@ -384,6 +391,13 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
}
if ownKey != "" && !strings.EqualFold(e.Encrypted, ownKey) {
r.noteNotAGuestBackupOnce(e, fmt.Sprintf("written by another box (key %s, this box's key %s) — not this box's proof", shortFP(e.Encrypted), shortFP(ownKey)))
foreignN++
if foreignMin == 0 || e.CTime < foreignMin {
foreignMin = e.CTime
}
if e.CTime > foreignMax {
foreignMax = e.CTime
}
continue
}
// R-689 (v0.136.0): … OF A GUEST THAT STILL EXISTS here. Measured on demo-hp 2026-09-27 right after
@@ -420,6 +434,9 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
bestCTime, best = e.CTime, e.VolID
}
}
if r.foreign != nil {
r.foreign.set(target, foreignN, foreignMin, foreignMax)
}
if best == "" {
return "", time.Time{}, nil
}
@@ -518,10 +535,25 @@ func (r *BackupRunner) warnRejectedArchiveOnce(e proxmox.StorageContent, why str
if seen {
return
}
r.logger.Warn("backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup",
r.logger.Warn(rejectedArchiveMessage(e),
"target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
}
// phantomCleanupPointer names the runbook that removes a PBS phantom (R-99, `09` §3 decision 140: a leftover of an
// aborted upload is deleted on the backup server, by a runbook, when one is seen — never automatically).
const phantomCleanupPointer = " — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
// rejectedArchiveMessage is the WARN text for a rejected archive. Only a PBS entry (format pbs-ct / pbs-vm) gets the
// cleanup pointer: the runbook deletes on a PBS datastore, and a tiny archive on a dir storage is not a PBS phantom.
// Pinned by TestRejectedArchiveWarnNamesTheCleanupRunbook.
func rejectedArchiveMessage(e proxmox.StorageContent) string {
msg := "backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup"
if strings.HasPrefix(e.Format, "pbs-") {
msg += phantomCleanupPointer
}
return msg
}
// demo-felhom in a single afternoon of deploys (2026-07-26).
//
// Asking the STORAGE rather than persisting the store is deliberate:
+3
View File
@@ -30,6 +30,9 @@ import (
// 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What
// is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/
// deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84).
// The due-check's fallback for an UNREADABLE storage no longer reads this store alone (R-894): the newest
// success per tier is also on disk (BackupSuccessState), so a restart followed by an unreachable storage
// reads the last known copy, not "never".
type Store struct {
mu sync.Mutex
byTarget map[string]hub.Backup // latest backup per target id
+6 -1
View File
@@ -145,12 +145,17 @@ var manifest = []Capability{
{"controllerswap-image-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "image", "inspect", "gitea.dooplex.hu/admin/felhom-controller:0.0.0"}, true, ""},
{"controllerswap-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "inspect", "-f", "{{.State.Running}}", "felhom-controller"}, true, ""},
{"controllerswap-restart", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "systemctl", "restart", "felhom-controller-bootstrap.service"}, true, ""},
{"controllerswap-write", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "tee", "/etc/felhom-controller-image"}, true, ""},
// R-861 (a) A1 (decision 165): the write goes through the ROOT verb that checks the ref; the agent has no `tee` grant.
{"controllerswap-write", "controller-swap / managed auto-update", "/usr/local/sbin/felhom-priv-apply", []string{"controller-image", "9201"}, true, ""},
// ---- Stale-lock recovery (FELHOM_STALELOCK, v0.49.0; Critical: a guest stuck behind a stale
// reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ----
{"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""},
// ---- Weekly guest disk trim (FELHOM_FSTRIM, R-444). NON-critical: a missing grant means the thin pool is not
// reclaimed this week (the trim job WARNs per guest and the report shows the failure), not a serving outage. ----
{"guest-fstrim", "weekly guest disk trim (thin-pool reclaim, R-444)", "/usr/sbin/pct", []string{"fstrim", "9201"}, false, ""},
// ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite
// backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the
// conf install, unit enable/restart and the handshake read gate the backup path. apt-install
+4 -3
View File
@@ -200,7 +200,8 @@ func TestRedProof_DroppedGrantFailsCheck(t *testing.T) {
}
// TestRedProof_DroppedControllerSwapTeeFailsCheck is the companion red-proof for the v0.45.0
// FELHOM_CONTROLLERSWAP grants: with the `tee /etc/felhom-controller-image` line removed, the
// FELHOM_CONTROLLERSWAP grants: with the write grant removed (since R-861 (a) A1 the `felhom-priv-apply controller-image`
// line; before it, an agent `tee /etc/felhom-controller-image`), the
// controllerswap-write capability MUST be reported uncovered. Proves the build gate watches the new
// swap write grant (so dropping it can't ship a non-root agent that silently can't auto-update).
func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
@@ -210,7 +211,7 @@ func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
}
var kept []string
for _, ln := range strings.Split(string(data), "\n") {
if strings.Contains(ln, "tee /etc/felhom-controller-image") {
if strings.Contains(ln, "felhom-priv-apply ^controller-image") { // R-861 (a) A1: the write's grant
continue
}
kept = append(kept, ln)
@@ -229,7 +230,7 @@ func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
}
cmdline := write.Binary + " " + strings.Join(write.ReprArgs, " ")
if matchesAny(cmdline, entries) {
t.Errorf("red-proof FAILED: controllerswap-write still matches after dropping the tee grant")
t.Errorf("red-proof FAILED: controllerswap-write still matches after dropping its grant")
}
if full := parseSudoersEntries(t, string(data)); !matchesAny(cmdline, full) {
t.Errorf("controllerswap-write should be covered by the real sudoers")
@@ -49,6 +49,10 @@ var r861Injections = []string{
"/usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount",
"/usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf",
"/usr/local/sbin/felhom-priv-apply wg /etc/shadow",
// R-861 (a) A1 (decision 165): the agent wrote ANY image ref into the guest by `tee` — now only the root verb may
"/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image",
"/usr/local/sbin/felhom-priv-apply controller-image 9201 9202",
"/usr/local/sbin/felhom-priv-apply controller-image 9201;id",
}
func TestSudoersRefusesTheR861Injections(t *testing.T) {
@@ -63,3 +67,72 @@ func TestSudoersRefusesTheR861Injections(t *testing.T) {
}
}
}
// R-444: the weekly trim's grant is ONE exact shape — `pct fstrim <vmid>` — and nothing smuggled after it.
// The manifest entry (guest-fstrim) proves the real call is still allowed (TestManifestCoveredBySudoers); this
// pins the other direction. RED-PROOF: write the rule as the glob `/usr/sbin/pct fstrim [0-9]*` → every decoy
// below with a trailing argument matches (the glob's `*` eats spaces).
func TestSudoersFstrimRuleIsExact(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
if !matchesAny("/usr/sbin/pct fstrim 9201", entries) {
t.Fatal("the sudoers does not allow `pct fstrim 9201` — the weekly trim cannot run")
}
for _, c := range []string{
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints",
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints 1",
"/usr/sbin/pct fstrim 9201; x",
"/usr/sbin/pct fstrim 9201 9202",
"/usr/sbin/pct fstrim 92a1",
"/usr/sbin/pct fstrim ",
"/usr/sbin/pct fstrim -- 9201",
"/usr/sbin/pct destroy 9201",
"/usr/sbin/pct destroy 9201 --purge",
} {
if matchesAny(c, entries) {
t.Errorf("the sudoers allows a command the trim rule must not: %q", c)
}
}
}
// R-861 (a) A1: the managed controller update still has its route — the root verb, one numeric vmid.
func TestSudoersAllowsTheControllerImageVerb(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
if !matchesAny("/usr/local/sbin/felhom-priv-apply controller-image 9201", parseSudoersEntries(t, string(data))) {
t.Fatal("the sudoers does not allow `felhom-priv-apply controller-image 9201` — a managed controller update cannot write its image")
}
}
// R-861 (b) B2 (decision 165, hygiene): felhom-op's `pct start|stop|unlock` grants are ONE numeric vmid each. The old
// glob `[0-9]*` eats spaces, so `pct stop 9201 --skiplock 1` and two vmids matched.
// RED-PROOF: on the pre-B2 felhom-op.sudoers (`/usr/sbin/pct stop [0-9]*`) the decoys match.
func TestFelhomOpSudoersPctIsExact(t *testing.T) {
data, err := os.ReadFile("../../configs/felhom-op.sudoers")
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
for _, ok := range []string{"/usr/sbin/pct start 9201", "/usr/sbin/pct stop 9201", "/usr/sbin/pct unlock 9201", "/usr/sbin/pct list"} {
if !matchesAny(ok, entries) {
t.Errorf("felhom-op lost a repair verb: %s", ok)
}
}
for _, bad := range []string{
"/usr/sbin/pct stop 9201 --skiplock 1",
"/usr/sbin/pct start 9201 9202",
"/usr/sbin/pct unlock 9201 --whatever",
"/usr/sbin/pct start 92a1",
"/usr/sbin/pct stop ",
"/usr/sbin/pct destroy 9201",
} {
if matchesAny(bad, entries) {
t.Errorf("felhom-op's sudoers allows %q", bad)
}
}
}
+11
View File
@@ -33,6 +33,7 @@ type Config struct {
LANResolver LANResolverConfig `json:"lan_resolver"`
WGTunnel WGTunnelConfig `json:"wg_tunnel"`
GuestNet GuestNetConfig `json:"guest_net"`
DiskTrim DiskTrimConfig `json:"disk_trim"`
OOB OOBConfig `json:"oob"`
SelfUpdate SelfUpdateConfig `json:"selfupdate"`
LogLevel string `json:"log_level"` // debug|info|warn|error (default info)
@@ -139,6 +140,16 @@ func (w WGTunnelConfig) WithDefaults() WGTunnelConfig {
return w
}
// DiskTrimConfig configures the R-444 weekly guest disk trim (internal/fstrim). DEFAULT-ON, like GuestNetConfig and for
// the same reason: it only acts on guests the agent already owns, and the operator ruled every box trims (`09` §3
// decision 139). Opting out is the explicit act: `"disk_trim": {"disable": true}`.
type DiskTrimConfig struct {
Disable bool `json:"disable"`
}
// Enabled reports whether the weekly trim should run.
func (d DiskTrimConfig) Enabled() bool { return !d.Disable }
// GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet).
//
// **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other
+346
View File
@@ -0,0 +1,346 @@
// Package fstrim is the weekly guest disk trim (R-444, operator ruling `09` §3 decision 139).
//
// Why: a thin pool only ever grows from blocks a guest has already FREED — `fstrim` inside the unprivileged container
// is refused (FITRIM: Operation not permitted), and nothing else on the box gives the blocks back. A full thin pool
// takes every guest on the host read-only, so the pool can reach 100 % from deleted data alone. Measured on demo-hp
// 2026-10-06 09:14Z: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %, 18/18 app probes 200, max 1.1 s
// (audits/ten-answers-2026-10-06/r444-measure.txt).
//
// The rule, each part pinned by a test in fstrim_test.go:
// - Weekly: a guest is DUE from Wednesday 10:00 local until it has been trimmed once since then (a box that was off
// on Wednesday catches up at its next eligible hour).
// - Daytime only: a trim starts only between 10:00 and 20:59 local — never in the night window (01:00–06:59) where
// the backups and the restore-tests run (TestEligibleHourNeverInTheNight).
// - Never beside a backup, a restore-test or another heavy operation: the pass holds the host-wide one-heavy-op gate
// (backup.InFlight) for its whole run; a busy gate DEFERS the pass to the next hourly tick.
// - A failed trim is retried at the next eligible hour, at most MaxAttemptsPerWeek times in one week.
// - The last result per guest (time, bytes, ok/fail) is persisted, so a restart neither loses it nor re-trims.
//
// The command is the ONE exact sudoers shape `pct fstrim <vmid>` (FELHOM_FSTRIM). Only guests from the pool-verified
// source (ListLXC ∩ the felhom pool, audit A1) and only RUNNING ones are trimmed.
package fstrim
import (
"context"
"encoding/json"
"fmt"
"log/slog"
"os"
"path/filepath"
"regexp"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// Schedule. The weekday/hours are fixed on purpose (one sentence the operator can read on the System page).
const (
Weekday = time.Wednesday
StartHour = 10 // first eligible local hour (inclusive)
EndHour = 21 // first NOT-eligible local hour (exclusive): last start is 20:59
MaxAttemptsPerWeek = 3
// TickInterval is how often the job looks; a deferred or failed pass is therefore retried the next hour.
TickInterval = time.Hour
// FirstTickDelay lets the agent settle after a start before the first look.
FirstTickDelay = 5 * time.Minute
// PerGuestTimeout bounds one `pct fstrim` (measured 24.4 s for 84 GiB).
PerGuestTimeout = 30 * time.Minute
)
// ScheduleText is the human description carried on the host report.
const ScheduleText = "weekly, due Wednesday from 10:00 host-local time; starts only 10:00-20:59; never beside a backup or restore-test"
// Runner runs a host command (proxmox.ExecRunner in production, through `sudo -n`).
type Runner interface {
Run(ctx context.Context, name string, args ...string) (stdout, stderr []byte, err error)
}
// GuestSource yields the guests this agent OWNS (the pool-verified source, never a bare ListLXC).
type GuestSource interface {
Guests(ctx context.Context) ([]proxmox.Guest, error)
}
// Gate is the host-wide one-heavy-operation gate (*backup.InFlight).
type Gate interface {
TryAcquire(what string) (release func(), busy string, ok bool)
}
// GateName is what the gate reports as busy while a trim runs.
const GateName = "guest-fstrim"
// Record is one guest's last trim attempt, as persisted.
type Record struct {
LastAttemptAt time.Time `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt time.Time `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
// Attempts counts the attempts since the current week's due time (reset by the first attempt of a new week).
Attempts int `json:"attempts"`
}
// Trimmer is the weekly job.
type Trimmer struct {
runner Runner
guests GuestSource
gate Gate
statePath string
logger *slog.Logger
loc *time.Location
now func() time.Time
mu sync.Mutex
records map[int]Record
}
// New builds the job and loads the persisted state. A missing state file is an empty state; a corrupt one is logged
// and treated as empty (the cost is one extra trim, never a missed one).
func New(runner Runner, guests GuestSource, gate Gate, statePath string, logger *slog.Logger) *Trimmer {
if logger == nil {
logger = slog.Default()
}
t := &Trimmer{runner: runner, guests: guests, gate: gate, statePath: statePath, logger: logger,
loc: time.Local, now: time.Now, records: map[int]Record{}}
t.load()
return t
}
func (t *Trimmer) load() {
data, err := os.ReadFile(t.statePath)
if err != nil {
if !os.IsNotExist(err) {
t.logger.Warn("fstrim: state read failed — starting empty", "path", t.statePath, "err", err)
}
return
}
var raw map[string]Record
if err := json.Unmarshal(data, &raw); err != nil {
t.logger.Warn("fstrim: state file corrupt — starting empty", "path", t.statePath, "err", err)
return
}
for k, r := range raw {
if id, err := strconv.Atoi(k); err == nil && id > 0 {
t.records[id] = r
}
}
}
func (t *Trimmer) saveLocked() error {
raw := make(map[string]Record, len(t.records))
for id, r := range t.records {
raw[strconv.Itoa(id)] = r
}
data, err := json.MarshalIndent(raw, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(t.statePath), 0o755); err != nil {
return err
}
tmp := t.statePath + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, t.statePath)
}
// EligibleHour reports whether a trim may START at local time lt.
func EligibleHour(lt time.Time) bool {
h := lt.Hour()
return h >= StartHour && h < EndHour
}
// weekAnchor is the most recent Wednesday StartHour:00 at or before lt (same location as lt).
func weekAnchor(lt time.Time) time.Time {
daysBack := (int(lt.Weekday()) - int(Weekday) + 7) % 7
d := lt.AddDate(0, 0, -daysBack)
a := time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
if a.After(lt) {
d = d.AddDate(0, 0, -7)
a = time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
}
return a
}
// due reports whether a guest with record r (ok=false: none) is due at local time lt.
func due(r Record, has bool, lt time.Time) bool {
if !has {
return true
}
anchor := weekAnchor(lt)
if r.LastAttemptAt.Before(anchor) {
return true // not tried this week
}
return !r.OK && r.Attempts < MaxAttemptsPerWeek
}
// Run looks every TickInterval until ctx ends. It never returns an error: a failed trim is a reported fact.
func (t *Trimmer) Run(ctx context.Context) {
t.logger.Info("fstrim: weekly guest disk trim starting", "schedule", ScheduleText)
timer := time.NewTimer(FirstTickDelay)
defer timer.Stop()
for {
select {
case <-ctx.Done():
return
case <-timer.C:
t.Pass(ctx)
timer.Reset(TickInterval)
}
}
}
// Pass is one look: outside the daytime window it does nothing; otherwise it trims every due, running, owned guest
// while holding the heavy-op gate.
func (t *Trimmer) Pass(ctx context.Context) {
lt := t.now().In(t.loc)
if !EligibleHour(lt) {
t.logger.Debug("fstrim: outside the daytime window — not looking", "local", lt.Format("Mon 15:04"))
return
}
guests, err := t.guests.Guests(ctx)
if err != nil {
t.logger.Warn("fstrim: owned-guest list unavailable — skipping this pass", "err", err)
return
}
owned := make(map[int]bool, len(guests))
var todo []int
t.mu.Lock()
for _, g := range guests {
owned[g.VMID] = true
r, has := t.records[g.VMID]
if !due(r, has, lt) {
continue
}
if g.Status != "running" {
t.logger.Info("fstrim: guest not running — trimmed when it runs", "vmid", g.VMID, "status", g.Status)
continue
}
todo = append(todo, g.VMID)
}
// A guest the agent no longer owns has no result to report.
pruned := false
for id := range t.records {
if !owned[id] {
delete(t.records, id)
pruned = true
}
}
if pruned {
if err := t.saveLocked(); err != nil {
t.logger.Warn("fstrim: state save failed", "err", err)
}
}
t.mu.Unlock()
if len(todo) == 0 {
return
}
sort.Ints(todo)
release, busy, ok := t.gate.TryAcquire(GateName)
if !ok {
t.logger.Info("fstrim: deferred — a heavy operation is in flight; retrying next hour", "busy", busy, "due_guests", len(todo))
return
}
defer release()
for _, vmid := range todo {
if ctx.Err() != nil {
return
}
t.trimOne(ctx, vmid, lt)
}
}
var trimmedLine = regexp.MustCompile(`\((\d+) bytes\) trimmed`)
// ParseTrimmed sums the "(N bytes) trimmed" lines of `pct fstrim` output and counts them (one per mount point), e.g.
// `/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed`.
func ParseTrimmed(out string) (bytes int64, mounts int) {
for _, m := range trimmedLine.FindAllStringSubmatch(out, -1) {
n, err := strconv.ParseInt(m[1], 10, 64)
if err != nil {
continue
}
bytes += n
mounts++
}
return bytes, mounts
}
// GiB renders bytes as "30.1 GiB".
func GiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
func (t *Trimmer) trimOne(ctx context.Context, vmid int, lt time.Time) {
start := t.now()
cctx, cancel := context.WithTimeout(ctx, PerGuestTimeout)
stdout, stderr, err := t.runner.Run(cctx, "pct", "fstrim", strconv.Itoa(vmid))
cancel()
dur := t.now().Sub(start)
bytes, mounts := ParseTrimmed(string(stdout) + "\n" + string(stderr))
t.mu.Lock()
prev, has := t.records[vmid]
r := Record{LastAttemptAt: start.UTC(), OK: err == nil, BytesTrimmed: bytes, Mounts: mounts,
DurationSeconds: float64(dur.Round(100*time.Millisecond)) / float64(time.Second), LastOKAt: prev.LastOKAt}
if has && !prev.LastAttemptAt.Before(weekAnchor(lt)) {
r.Attempts = prev.Attempts + 1
} else {
r.Attempts = 1
}
if err == nil {
r.LastOKAt = start.UTC()
} else {
msg := strings.TrimSpace(err.Error() + ": " + strings.TrimSpace(string(stderr)))
if len(msg) > 300 {
msg = msg[:300]
}
r.Error = msg
}
t.records[vmid] = r
saveErr := t.saveLocked()
t.mu.Unlock()
if err == nil {
t.logger.Info(fmt.Sprintf("fstrim: guest %d trimmed %s in %.1fs", vmid, GiB(bytes), r.DurationSeconds),
"vmid", vmid, "bytes_trimmed", bytes, "mounts", mounts, "duration_s", r.DurationSeconds)
if mounts == 0 {
t.logger.Warn("fstrim: pct fstrim succeeded but reported no trimmed mount — output not understood",
"vmid", vmid, "stdout", strings.TrimSpace(string(stdout)))
}
} else {
t.logger.Warn(fmt.Sprintf("fstrim: guest %d trim FAILED after %.1fs", vmid, r.DurationSeconds),
"vmid", vmid, "attempt", r.Attempts, "max_attempts_per_week", MaxAttemptsPerWeek, "err", r.Error)
}
if saveErr != nil {
t.logger.Warn("fstrim: state save failed — the result will not survive a restart", "path", t.statePath, "err", saveErr)
}
}
// GuestDiskTrimStatus implements hub.GuestDiskTrimReporter: a pure read of the persisted results (never runs pct).
func (t *Trimmer) GuestDiskTrimStatus(context.Context) *hub.GuestDiskTrimStatus {
t.mu.Lock()
defer t.mu.Unlock()
out := &hub.GuestDiskTrimStatus{Schedule: ScheduleText}
ids := make([]int, 0, len(t.records))
for id := range t.records {
ids = append(ids, id)
}
sort.Ints(ids)
for _, id := range ids {
r := t.records[id]
g := hub.GuestDiskTrim{VMID: id, LastAttemptAt: r.LastAttemptAt.UTC().Format(time.RFC3339), OK: r.OK,
BytesTrimmed: r.BytesTrimmed, Mounts: r.Mounts, DurationSeconds: r.DurationSeconds, Error: r.Error}
if !r.LastOKAt.IsZero() {
g.LastOKAt = r.LastOKAt.UTC().Format(time.RFC3339)
}
out.Guests = append(out.Guests, g)
}
return out
}
+267
View File
@@ -0,0 +1,267 @@
package fstrim
import (
"bytes"
"context"
"encoding/json"
"errors"
"log/slog"
"path/filepath"
"reflect"
"strings"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// The real `pct fstrim 9201` output measured on demo-hp 2026-10-06 (audits/ten-answers-2026-10-06/r444-measure.txt).
const measuredOut = "/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed\n" +
"/var/lib/lxc/9201/rootfs/var/lib/felhom: 53.9 GiB (57865633792 bytes) trimmed\n"
const measuredBytes = int64(32277680128 + 57865633792)
type fakeRunner struct {
mu sync.Mutex
calls [][]string
out string
err error
onRun func()
}
func (f *fakeRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
f.mu.Lock()
f.calls = append(f.calls, append([]string{name}, args...))
f.mu.Unlock()
if f.onRun != nil {
f.onRun()
}
if f.err != nil {
return nil, []byte("mount busy"), f.err
}
return []byte(f.out), nil, nil
}
type fakeGuests struct {
g []proxmox.Guest
err error
}
func (f fakeGuests) Guests(context.Context) ([]proxmox.Guest, error) { return f.g, f.err }
// A Wednesday 10:30 in a fixed zone (CEST-like), so the tests do not depend on the machine's zone.
var zone = time.FixedZone("CEST", 2*3600)
func at(day, hour, min int) time.Time { return time.Date(2026, 10, day, hour, min, 0, 0, zone) } // 2026-10-07 = Wednesday
func newT(t *testing.T, r Runner, g GuestSource, gate Gate, now *time.Time) (*Trimmer, *bytes.Buffer, string) {
t.Helper()
var logs bytes.Buffer
path := filepath.Join(t.TempDir(), "guest-disk-trim.json")
tr := New(r, g, gate, path, slog.New(slog.NewTextHandler(&logs, &slog.HandlerOptions{Level: slog.LevelDebug})))
tr.loc = zone
tr.now = func() time.Time { return *now }
return tr, &logs, path
}
func running(ids ...int) fakeGuests {
var g []proxmox.Guest
for _, id := range ids {
g = append(g, proxmox.Guest{VMID: id, Status: "running", Type: "lxc"})
}
return fakeGuests{g: g}
}
func TestParseTrimmedTheMeasuredOutput(t *testing.T) {
b, m := ParseTrimmed(measuredOut)
if b != measuredBytes || m != 2 {
t.Fatalf("ParseTrimmed = %d bytes over %d mounts, want %d over 2", b, m, measuredBytes)
}
if b, m := ParseTrimmed("something else\n"); b != 0 || m != 0 {
t.Fatalf("unrelated output parsed as %d/%d", b, m)
}
if got := GiB(measuredBytes); got != "84.0 GiB" {
t.Fatalf("GiB = %q", got)
}
}
// The night window (01:00–06:59) must never be eligible, and the daytime window is exactly 10:00–20:59.
func TestEligibleHourNeverInTheNight(t *testing.T) {
for h := 0; h < 24; h++ {
lt := time.Date(2026, 10, 7, h, 30, 0, 0, zone)
got := EligibleHour(lt)
if h >= 1 && h <= 6 && got {
t.Errorf("hour %02d is in the night window and must not be eligible", h)
}
if want := h >= 10 && h <= 20; got != want {
t.Errorf("EligibleHour(%02d:30) = %v, want %v", h, got, want)
}
}
}
func TestWeekAnchorIsTheLastWednesdayTen(t *testing.T) {
cases := map[time.Time]time.Time{
at(7, 10, 0): at(7, 10, 0), // Wednesday 10:00 itself
at(7, 9, 59): time.Date(2026, 9, 30, 10, 0, 0, 0, zone), // before 10:00 Wednesday → the previous week
at(8, 15, 0): at(7, 10, 0), // Thursday
at(13, 20, 0): at(7, 10, 0), // next Tuesday
at(14, 11, 0): at(14, 10, 0), // next Wednesday
}
for in, want := range cases {
if got := weekAnchor(in); !got.Equal(want) {
t.Errorf("weekAnchor(%s) = %s, want %s", in.Format("Mon 01-02 15:04"), got.Format("Mon 01-02 15:04"), want.Format("Mon 01-02 15:04"))
}
}
}
// The consequence: on Wednesday 10:30 a running owned guest is trimmed with the ONE exact argv, the bytes are parsed,
// the positive log line is written, the result is persisted, and the host report carries it.
func TestPassTrimsADueGuestAndReportsIt(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
tr, logs, path := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9201"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("runner calls = %q, want %q", r.calls, want)
}
if !strings.Contains(logs.String(), "fstrim: guest 9201 trimmed 84.0 GiB in ") {
t.Fatalf("no positive per-guest log line:\n%s", logs.String())
}
st := tr.GuestDiskTrimStatus(context.Background())
if st == nil || st.Schedule != ScheduleText || len(st.Guests) != 1 {
t.Fatalf("report stanza = %+v", st)
}
g := st.Guests[0]
if g.VMID != 9201 || !g.OK || g.BytesTrimmed != measuredBytes || g.Mounts != 2 || g.LastOKAt == "" || g.LastAttemptAt == "" {
t.Fatalf("report guest = %+v", g)
}
// Persisted: a NEW Trimmer over the same file (an agent restart) still has it and does not trim again this week.
now = at(8, 11, 0)
r2 := &fakeRunner{out: measuredOut}
tr2 := New(r2, running(9201), &backup.InFlight{}, path, slog.New(slog.NewTextHandler(&bytes.Buffer{}, nil)))
tr2.loc, tr2.now = zone, func() time.Time { return now }
if st2 := tr2.GuestDiskTrimStatus(context.Background()); len(st2.Guests) != 1 || st2.Guests[0].BytesTrimmed != measuredBytes {
t.Fatalf("result lost over a restart: %+v", st2)
}
tr2.Pass(context.Background())
if len(r2.calls) != 0 {
t.Fatalf("trimmed again in the same week after a restart: %q", r2.calls)
}
// Next week it is due again.
now = at(14, 10, 5)
tr2.Pass(context.Background())
if len(r2.calls) != 1 {
t.Fatalf("not trimmed in the next week: %q", r2.calls)
}
}
func TestPassNeverRunsInTheNight(t *testing.T) {
for _, h := range []int{1, 3, 6, 9, 21, 23} {
now := at(7, h, 15)
r := &fakeRunner{out: measuredOut}
tr, _, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Errorf("trimmed at %02d:15: %q", h, r.calls)
}
}
}
// A backup (or restore-test) holding the heavy-op gate DEFERS the trim; the next hour, gate free, it runs. And while
// a trim runs, the gate is held, so a backup cannot start beside it.
func TestPassDefersToAHeavyOperationAndRetriesNextHour(t *testing.T) {
now := at(7, 10, 30)
gate := &backup.InFlight{}
release, _, _ := gate.TryAcquire("backup:9201")
var busyDuringTrim string
r := &fakeRunner{out: measuredOut}
r.onRun = func() { busyDuringTrim = gate.Busy() }
tr, logs, _ := newT(t, r, running(9201), gate, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Fatalf("trimmed beside a running backup: %q", r.calls)
}
if !strings.Contains(logs.String(), "fstrim: deferred") || !strings.Contains(logs.String(), "backup:9201") {
t.Fatalf("the deferral is not logged with what holds the gate:\n%s", logs.String())
}
release()
now = now.Add(time.Hour)
tr.Pass(context.Background())
if len(r.calls) != 1 {
t.Fatalf("not retried the next hour: %q", r.calls)
}
if busyDuringTrim != GateName {
t.Fatalf("the heavy-op gate was %q during the trim, want %q", busyDuringTrim, GateName)
}
if gate.Busy() != "" {
t.Fatalf("the gate was not released after the pass: %q", gate.Busy())
}
}
func TestFailedTrimWarnsIsRecordedAndRetriedAtMostThreeTimes(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{err: errors.New("exit status 255")}
tr, logs, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
for i := 0; i < 6; i++ {
tr.Pass(context.Background())
now = now.Add(time.Hour)
}
if len(r.calls) != MaxAttemptsPerWeek {
t.Fatalf("attempts in one week = %d, want %d", len(r.calls), MaxAttemptsPerWeek)
}
if !strings.Contains(logs.String(), "level=WARN") || !strings.Contains(logs.String(), "fstrim: guest 9201 trim FAILED") {
t.Fatalf("no WARN for the failure:\n%s", logs.String())
}
g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]
if g.OK || g.LastOKAt != "" || !strings.Contains(g.Error, "exit status 255") || !strings.Contains(g.Error, "mount busy") {
t.Fatalf("failed result not recorded as a failure: %+v", g)
}
// A success later keeps a clean record.
r.err = nil
r.out = measuredOut
now = at(14, 10, 10)
tr.Pass(context.Background())
if g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]; !g.OK || g.Error != "" || g.BytesTrimmed != measuredBytes {
t.Fatalf("success after failure: %+v", g)
}
}
func TestOnlyRunningOwnedGuestsAndAFailedListActsOnNothing(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
g := fakeGuests{g: []proxmox.Guest{{VMID: 9201, Status: "stopped"}, {VMID: 9202, Status: "running"}}}
tr, _, _ := newT(t, r, g, &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9202"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("calls = %q, want only the running guest", r.calls)
}
r2 := &fakeRunner{out: measuredOut}
tr2, logs, _ := newT(t, r2, fakeGuests{err: errors.New("pool read 403")}, &backup.InFlight{}, &now)
tr2.Pass(context.Background())
if len(r2.calls) != 0 || !strings.Contains(logs.String(), "owned-guest list unavailable") {
t.Fatalf("a failed ownership read must act on nothing: calls %q", r2.calls)
}
}
func TestReportJSONShape(t *testing.T) {
now := at(7, 10, 30)
tr, _, _ := newT(t, &fakeRunner{out: measuredOut}, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
b, err := json.Marshal(tr.GuestDiskTrimStatus(context.Background()))
if err != nil {
t.Fatal(err)
}
for _, k := range []string{`"schedule":`, `"guests":[{"vmid":9201`, `"last_attempt_at":"2026-10-07T08:30:00Z"`, `"ok":true`,
`"bytes_trimmed":90143313920`, `"mounts":2`, `"duration_seconds":`, `"last_ok_at":"2026-10-07T08:30:00Z"`} {
if !strings.Contains(string(b), k) {
t.Errorf("report JSON lacks %s: %s", k, b)
}
}
if strings.Contains(string(b), `"error"`) {
t.Errorf("an ok result must omit error: %s", b)
}
}
+34
View File
@@ -90,6 +90,12 @@ type GuestNetReporter interface {
GuestNetStatus(ctx context.Context) *GuestNetStatus
}
// GuestDiskTrimReporter is the R-444 seam the weekly trim job plugs into (same consumer-side pattern — hub does not
// import fstrim). nil (feature not wired) → no guest_disk_trim stanza.
type GuestDiskTrimReporter interface {
GuestDiskTrimStatus(ctx context.Context) *GuestDiskTrimStatus
}
// Collector builds a HostReport from read-only sources. All deps are behind narrow
// interfaces for unit testing.
type Collector struct {
@@ -108,6 +114,8 @@ type Collector struct {
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
diskTrim GuestDiskTrimReporter // R-444: weekly guest disk trim (nil → stanza omitted)
foreignKey ForeignKeyArchiveReporter // R-366 slice 2 (nil → omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
oob OOBReporter // H1: operator-access health (nil → stanza omitted)
@@ -221,6 +229,24 @@ func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
return c
}
// ForeignKeyArchiveReporter is the R-366 slice-2 seam: the restore-test's ledger of archives written with another
// key. nil = not evaluated yet (the stanza is omitted and the hub keeps its state).
type ForeignKeyArchiveReporter interface {
ForeignKeyArchives(ctx context.Context) *ForeignKeyArchivesStanza
}
// SetForeignKeyArchiveReporter wires the restore-test's foreign-key ledger (R-366 slice 2; nil-safe → omitted).
func (c *Collector) SetForeignKeyArchiveReporter(r ForeignKeyArchiveReporter) *Collector {
c.foreignKey = r
return c
}
// SetGuestDiskTrimReporter wires the R-444 weekly trim job as a report source (nil-safe → stanza omitted).
func (c *Collector) SetGuestDiskTrimReporter(r GuestDiskTrimReporter) *Collector {
c.diskTrim = r
return c
}
// SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side
// pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report.
type SelfUpdateReporter interface {
@@ -377,6 +403,14 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
}
// R-444: the last weekly trim result per guest (nil reporter = not wired → stanza omitted).
if c.diskTrim != nil {
report.GuestDiskTrim = c.diskTrim.GuestDiskTrimStatus(ctx)
}
// R-366 slice 2: archives the restore-test skipped as another key's (nil → not evaluated yet → omitted).
if c.foreignKey != nil {
report.ForeignKeyArchives = c.foreignKey.ForeignKeyArchives(ctx)
}
// D1: agent self-update pending status (nil reporter → pending=false, the steady state).
if c.selfUpdate != nil {
report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending()
+54
View File
@@ -0,0 +1,54 @@
package hub
import (
"context"
"encoding/json"
"testing"
)
// R-444: the guest_disk_trim stanza must reach a report built through the PRODUCTION collect path, be absent from
// the wire when the job is not wired, and carry the keys the hub's System page reads.
type fakeDiskTrim struct{ st *GuestDiskTrimStatus }
func (f fakeDiskTrim) GuestDiskTrimStatus(context.Context) *GuestDiskTrimStatus { return f.st }
func TestCollect_GuestDiskTrim(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.150.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ := json.Marshal(r)
var m map[string]any
_ = json.Unmarshal(b, &m)
if _, ok := m["guest_disk_trim"]; ok {
t.Fatalf("guest_disk_trim on the wire with no reporter wired: %s", b)
}
c.SetGuestDiskTrimReporter(fakeDiskTrim{st: &GuestDiskTrimStatus{Schedule: "weekly", Guests: []GuestDiskTrim{{
VMID: 9201, LastAttemptAt: "2026-10-07T08:30:00Z", OK: true, BytesTrimmed: 90143313920, Mounts: 2,
DurationSeconds: 24.4, LastOKAt: "2026-10-07T08:30:00Z",
}}}})
r, err = c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ = json.Marshal(r)
m = nil
_ = json.Unmarshal(b, &m)
dt, ok := m["guest_disk_trim"].(map[string]any)
if !ok || dt["schedule"] != "weekly" {
t.Fatalf("guest_disk_trim missing or wrong on the wire: %s", b)
}
g := dt["guests"].([]any)[0].(map[string]any)
for _, k := range []string{"vmid", "last_attempt_at", "ok", "bytes_trimmed", "mounts", "duration_seconds", "last_ok_at"} {
if _, ok := g[k]; !ok {
t.Fatalf("guest_disk_trim.guests[0] lacks %q: %v", k, g)
}
}
if g["bytes_trimmed"] != float64(90143313920) || g["ok"] != true {
t.Fatalf("values did not survive the round trip: %v", g)
}
}
+58
View File
@@ -124,6 +124,17 @@ type HostReport struct {
// on HostReport would have been the only report block named against that convention.
GuestNet *GuestNetStatus `json:"guest_net,omitempty"`
// GuestDiskTrim is the weekly guest disk trim stanza (R-444, `09` §3 decision 139): the schedule and, per owned
// guest, the LAST trim result as persisted by the agent (it survives a restart). Present only when the trim job is
// wired; an empty `guests` list means the job runs and no guest has been trimmed yet. No secret.
GuestDiskTrim *GuestDiskTrimStatus `json:"guest_disk_trim,omitempty"`
// ForeignKeyArchives (R-366 slice 2, `09` §3 decision 168): per backup tier, the whole-guest archives the
// restore-test SKIPPED because they were written with another key (an earlier install of this box). This box
// cannot open them; the hub turns a CHANGE of this list into one operator event. Absent = not evaluated yet since
// the agent started (the hub keeps its last state); `tiers: []` = evaluated, none found.
ForeignKeyArchives *ForeignKeyArchivesStanza `json:"foreign_key_archives,omitempty"`
// LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent
// mirror of the controller's report log_tails channel. Present ONLY on the heartbeat
// right after the control envelope requested it (log_tail_requested); consume-once on
@@ -208,6 +219,27 @@ type GuestNetGuest struct {
Message string `json:"message,omitempty"`
}
// GuestDiskTrimStatus is the R-444 weekly trim stanza. `schedule` is a plain description of when the job runs (local
// time of the host); `guests` holds one entry per owned guest that has had at least one trim attempt.
type GuestDiskTrimStatus struct {
Schedule string `json:"schedule"`
Guests []GuestDiskTrim `json:"guests,omitempty"`
}
// GuestDiskTrim is one guest's LAST trim attempt. `ok` with `last_attempt_at` is the verdict of that attempt — never
// read the time alone as success; `last_ok_at` is the last attempt that succeeded ("" = never). `bytes_trimmed` is
// the sum of the "(N bytes) trimmed" lines `pct fstrim` printed, over `mounts` mount points.
type GuestDiskTrim struct {
VMID int `json:"vmid"`
LastAttemptAt string `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt string `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
}
type PBSDRStatus struct {
State string `json:"state"`
StorageID string `json:"storage_id,omitempty"`
@@ -428,6 +460,17 @@ type SmartSummary struct {
ReallocatedSectors *int `json:"reallocated_sectors"`
PendingSectors *int `json:"pending_sectors"`
OfflineUncorrectable *int `json:"offline_uncorrectable"`
// R-330 (disk health Phase 2): three more SATA raw counters. omitempty + pointer: absent (an
// older agent, an NVMe/USB device, or a drive that does not report the attribute) is OMITTED —
// unknown, never a zero (S-39). Wire only: no verdict reads them yet.
// 187 Reported_Uncorrect — the failing drive's most telling counter (normalized 1 vs thresh 0,
// raw 1001) while SMART still said PASSED.
// 188 Command_Timeout — some vendors PACK several counters into the 48-bit raw value, so the
// number is carried as reported and must not be compared across vendors.
// 199 UDMA_CRC_Error_Count — cabling / link errors, not the medium.
ReportedUncorrect *int64 `json:"reported_uncorrect,omitempty"`
CommandTimeout *int64 `json:"command_timeout,omitempty"`
UDMACRCErrors *int64 `json:"udma_crc_errors,omitempty"`
// NVMe attributes.
CriticalWarning *int `json:"critical_warning"`
@@ -689,3 +732,18 @@ type WireRestoreDirective struct {
Archive string `json:"archive,omitempty"` // source archive/snapshot to restore from
VMID int `json:"vmid,omitempty"`
}
// ForeignKeyArchivesStanza wraps the per-tier list so "evaluated, none" (`tiers: []`) differs from "not evaluated"
// (the stanza absent) without a null on the wire.
type ForeignKeyArchivesStanza struct {
Tiers []ForeignKeyArchives `json:"tiers"`
}
// ForeignKeyArchives is one tier's count of archives written with another key (R-366 slice 2): the count and the
// newest/oldest archive time (RFC3339, UTC). No key material — the fingerprints stay on the box.
type ForeignKeyArchives struct {
Target string `json:"target"`
Count int `json:"count"`
Oldest string `json:"oldest"`
Newest string `json:"newest"`
}
@@ -4,7 +4,6 @@ import (
"context"
"encoding/json"
"errors"
"io"
"log/slog"
"os"
"path/filepath"
@@ -54,8 +53,8 @@ func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string
}
return "", errors.New("supExec: unexpected args")
}
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) {
return "", errors.New("supExec: no stdin exec expected")
func (f *supExec) WriteControllerImage(context.Context, int, string) error {
return errors.New("supExec: no image write expected")
}
func (f *supExec) count(vmid int) int {
f.mu.Lock()
+9 -12
View File
@@ -4,7 +4,6 @@ import (
"context"
"encoding/json"
"fmt"
"io"
"log/slog"
"net/http"
"os"
@@ -41,9 +40,10 @@ func ValidControllerImage(ref string) bool { return controllerImageRe.MatchStrin
// faked in tests. The single seam the swap composes over (no hand-rolled pct).
type GuestExecutor interface {
GuestExec(ctx context.Context, vmid int, args ...string) (string, error)
// GuestExecStdin is GuestExec with the command's stdin fed from stdin — the swap write pipes the
// image ref into an in-guest `tee` (no shell vector).
GuestExecStdin(ctx context.Context, vmid int, stdin io.Reader, args ...string) (string, error)
// WriteControllerImage writes the image ref into the guest's /etc/felhom-controller-image through the ROOT
// verb `felhom-priv-apply controller-image <vmid>` (R-861 (a) A1, `09` §3 decision 165), which re-checks the
// ref against our registry + repository + x.y.z. The agent no longer holds a `tee` grant into the guest.
WriteControllerImage(ctx context.Context, vmid int, image string) error
}
// ControllerSwapState is the durable record of a swap (crash-safety + status). Written before the swap
@@ -141,14 +141,11 @@ func (c *ControllerSwapper) imagePresent(ctx context.Context, vmid int, image st
}
func (c *ControllerSwapper) writeImage(ctx context.Context, vmid int, image string) error {
// Non-root path: pipe the image ref into an in-guest `tee` over stdin — no shell, no
// interpolation, no `bash -c` (the only swap vector that would have needed an arbitrary-exec
// grant). The trailing "\n" makes the on-disk bytes byte-identical to the golden's
// `printf '%s\n'`; the bootstrap reads `IMAGE=$(cat …)` so the newline is stripped on read
// (spike SPIKE-controllerswap-narrow-grants-2026-06-29). image is strict-validated
// (controllerImageRe) upstream in Swap; defence-in-depth, the stdin path can't smuggle anyway.
_, err := c.exec.GuestExecStdin(ctx, vmid, strings.NewReader(image+"\n"), "tee", controllerImageFile)
return err
// R-861 (a) A1: the ROOT verb writes the file (it re-checks the ref — a compromised agent cannot hand the guest's
// bootstrap another image). The bytes are `image\n`, byte-identical to the golden's `printf '%s\n'`; the bootstrap
// reads `IMAGE=$(cat …)` so the newline is stripped on read. image is also strict-validated (controllerImageRe)
// upstream in Swap.
return c.exec.WriteControllerImage(ctx, vmid, image)
}
func (c *ControllerSwapper) restartBootstrap(ctx context.Context, vmid int) error {
+14 -28
View File
@@ -23,7 +23,7 @@ type fakeGuestExec struct {
present map[string]bool // images pulled into the guest
good map[string]bool // images that report healthy when running
containerImg string // image the running container currently has
teeStdin []string // raw bytes piped into each `tee` write (the swap's write vector)
teeStdin []string // image refs handed to WriteControllerImage (the root verb, R-861 (a) A1)
failRestart bool
noHealthBlock bool // if set, .State.Health is absent ("none")
restartCount int // F1: .RestartCount reported by docker inspect (a crash-looper has >0)
@@ -75,20 +75,15 @@ func (f *fakeGuestExec) GuestExec(_ context.Context, _ int, args ...string) (str
return "", fmt.Errorf("fake: unexpected exec %v", args)
}
// GuestExecStdin models the swap's write vector: `tee /etc/felhom-controller-image` with the image
// piped on stdin. It records the raw stdin bytes and sets the modeled file content (newline-stripped,
// as the bootstrap's `IMAGE=$(cat …)` read would see it).
func (f *fakeGuestExec) GuestExecStdin(_ context.Context, _ int, stdin io.Reader, args ...string) (string, error) {
// WriteControllerImage models the swap's write: the root verb `felhom-priv-apply controller-image <vmid>` (R-861
// (a) A1). It records the ref and sets the modeled file content, as the bootstrap's `IMAGE=$(cat …)` would read it.
func (f *fakeGuestExec) WriteControllerImage(_ context.Context, _ int, image string) error {
f.mu.Lock()
defer f.mu.Unlock()
f.calls = append(f.calls, args)
b, _ := io.ReadAll(stdin)
if len(args) >= 2 && args[0] == "tee" && args[1] == controllerImageFile {
f.teeStdin = append(f.teeStdin, string(b))
f.imageFile = strings.TrimSpace(string(b))
return string(b), nil // tee echoes stdin to stdout
}
return "", fmt.Errorf("fake: unexpected exec-stdin args=%v stdin=%q", args, string(b))
f.calls = append(f.calls, []string{"felhom-priv-apply", "controller-image", image})
f.teeStdin = append(f.teeStdin, image+"\n")
f.imageFile = image
return nil
}
// wrote reports whether the image was written via the stdin `tee` vector with the exact `image\n`
@@ -162,7 +157,7 @@ func TestControllerSwap_Happy(t *testing.T) {
// The write vector must be the stdin `tee` with byte-identical `image\n` and NO shell — the
// controllerswap.go writeImage rewrite. This would FAIL on the pre-change `bash -c "printf … >"` impl.
func TestControllerSwap_WriteViaStdinTee_NoShell(t *testing.T) {
func TestControllerSwap_WriteViaRootVerb_NoShell(t *testing.T) {
fe := &fakeGuestExec{
imageFile: prevImg,
present: map[string]bool{newImg: true},
@@ -173,22 +168,15 @@ func TestControllerSwap_WriteViaStdinTee_NoShell(t *testing.T) {
t.Fatalf("state = %q, want done", st.State)
}
if !fe.wrote(newImg) {
t.Errorf("expected a tee write of %q+\\n; teeStdin=%q", newImg, fe.teeStdin)
t.Errorf("expected the root verb to write %q; writes=%q", newImg, fe.teeStdin)
}
sawTee := false
for _, c := range fe.calls {
if len(c) >= 2 && c[0] == "tee" {
sawTee = true
if c[1] != controllerImageFile {
t.Errorf("tee target = %q, want fixed %q", c[1], controllerImageFile)
}
if len(c) >= 1 && c[0] == "tee" {
t.Errorf("the swap still uses an in-guest tee (R-861 (a) A1 removed that grant): %v", c)
}
}
if !sawTee {
t.Error("no tee call recorded — writeImage did not use the stdin tee vector")
}
if fe.usedShell() {
t.Errorf("swap used a shell vector (bash/-c/printf) — must be stdin tee only; calls=%v", fe.calls)
t.Errorf("swap used a shell vector (bash/-c/printf); calls=%v", fe.calls)
}
}
@@ -300,9 +288,7 @@ func (s *inspectScript) GuestExec(_ context.Context, _ int, args ...string) (str
}
return "", nil
}
func (s *inspectScript) GuestExecStdin(_ context.Context, _ int, _ io.Reader, _ ...string) (string, error) {
return "", nil
}
func (s *inspectScript) WriteControllerImage(context.Context, int, string) error { return nil }
func fastSwapper(exec GuestExecutor) *ControllerSwapper {
s := NewControllerSwapper(exec, "", discardLogger())
+102
View File
@@ -0,0 +1,102 @@
package localapi
import (
"encoding/json"
"errors"
"io"
"io/fs"
"net/http"
"os"
"time"
)
// GET /host/crash-guard (R-856, `09` §3 decision 143): what the host's crash guard
// (configs/felhom-crash-guard, `11` §5.9) recorded about the most recent HOST boot. The controller
// reads it once after it starts: when the host's last boot followed an UNCLEAN stop, its app mails
// wait ~15 minutes instead of the normal 90 s boot grace.
//
// Read-only and Proxmox-free: the agent reads the guard's state file (root-owned, 0644 — the
// non-root agent can read it) and passes four fields through. Host-wide, token-authed (any valid
// per-guest token sees the host's view, as GET /host/metrics does).
//
// NEVER an error page. A missing file (no guard installed, or no boot recorded yet), an unreadable
// one, or one that does not parse answers 200 with present:false — the controller reads that as
// UNKNOWN and keeps its normal boot grace. Pinned by TestR856_CrashGuard*.
// defaultCrashGuardStatePath is where configs/felhom-crash-guard writes its state (STATE_DIR there).
const defaultCrashGuardStatePath = "/var/lib/felhom-crash-guard/state.json"
// crashGuardStateMax bounds the read; the real file is well under 4 KiB.
const crashGuardStateMax = 1 << 20
// CrashGuardResponse is the data block of GET /host/crash-guard. Field names are the controller's
// agentapi.CrashGuardState (felhom-controller internal/agentapi/crashguard.go) — a wire contract,
// pinned by TestR856_CrashGuardWireMatchesControllerClient.
type CrashGuardResponse struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"` // RFC3339 UTC ("2006-01-02T15:04:05Z")
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// crashGuardFile is the subset of the guard's state.json the route passes through. Every other key
// (armed, boot_id, config, unclean_boots, last_trip, ...) is ignored.
type crashGuardFile struct {
LastBootAt string `json:"last_boot_at"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// readCrashGuardState reads and parses the guard's state file. ok=false on ANY failure (missing,
// unreadable, oversized, not a JSON object, a field of the wrong type); reason says which, for the log.
func readCrashGuardState(path string) (resp CrashGuardResponse, ok bool, reason string) {
f, err := os.Open(path)
if err != nil {
if errors.Is(err, fs.ErrNotExist) {
return resp, false, "no state file"
}
return resp, false, "unreadable: " + err.Error()
}
defer f.Close()
raw, err := io.ReadAll(io.LimitReader(f, crashGuardStateMax+1))
if err != nil {
return resp, false, "read: " + err.Error()
}
if len(raw) > crashGuardStateMax {
return resp, false, "state file too large"
}
var st crashGuardFile
// Unmarshal into a struct fails on a non-object top level (null decodes, so reject it below).
if err := json.Unmarshal(raw, &st); err != nil {
return resp, false, "unparseable: " + err.Error()
}
var probe map[string]json.RawMessage
if err := json.Unmarshal(raw, &probe); err != nil || probe == nil {
return resp, false, "unparseable: not a JSON object"
}
resp = CrashGuardResponse{Present: true, LastBootUnclean: st.LastBootUnclean, Tripped: st.Tripped}
// Normalise to RFC3339 UTC; an unparseable time passes through as-is (the controller reads an
// unparseable boot time as "not this start's boot" → its normal grace).
if t, perr := time.Parse(time.RFC3339, st.LastBootAt); perr == nil {
resp.LastBootAt = t.UTC().Format(time.RFC3339)
} else {
resp.LastBootAt = st.LastBootAt
}
return resp, true, ""
}
func (s *Server) handleCrashGuard(w http.ResponseWriter, r *http.Request, vmid int) {
path := s.crashGuardStatePath
if path == "" {
path = defaultCrashGuardStatePath
}
resp, ok, reason := readCrashGuardState(path)
if !ok {
s.logger.Debug("local-api: /host/crash-guard not present", "vmid", vmid, "reason", reason)
writeOK(w, CrashGuardResponse{Present: false})
return
}
s.logger.Debug("local-api: /host/crash-guard served", "vmid", vmid,
"last_boot_at", resp.LastBootAt, "last_boot_unclean", resp.LastBootUnclean, "tripped", resp.Tripped)
writeOK(w, resp)
}
+205
View File
@@ -0,0 +1,205 @@
package localapi
import (
"encoding/json"
"io"
"log/slog"
"net/http"
"os"
"path/filepath"
"testing"
)
// The shape of /var/lib/felhom-crash-guard/state.json as read on demo-hp on 2026-10-06 (values from
// that read where they matter; lists/objects kept to the same key set).
const crashGuardFixture = `{
"armed": true,
"boot_id": "3f1c0f1e-6a0b-4d7e-9b7a-0c2d4e6f8a1b",
"config": {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10},
"kernel_panic": 10,
"last_boot_at": "2026-10-05T07:56:41Z",
"last_boot_unclean": true,
"last_trip": {},
"rearmed_at": "2026-10-04T14:02:11Z",
"rearmed_by": "operator",
"tripped": false,
"unclean_boots": ["2026-10-05T07:56:41Z"],
"unclean_boots_24h": 1,
"unclean_boots_in_window": 1,
"updated_at": "2026-10-05T07:57:02Z",
"version": 1
}`
// controllerCrashGuardState is a COPY of the controller's wire type, felhom-controller
// controller/internal/agentapi/crashguard.go `CrashGuardState` (commit 8b13a5e) — same field names,
// same tags. If either side renames a key, the contract test below fails.
type controllerCrashGuardState struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
func newCrashGuardServer(t *testing.T, statePath string) http.Handler {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0",
Guests: &fakeGuests{},
Backups: &fakeBackups{},
Store: &fakeStore{},
Storage: fakeStorage{},
Tokens: staticTokens{"A": 8200, "B": 9300},
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatalf("new server: %v", err)
}
srv.crashGuardStatePath = statePath
return srv.Handler()
}
func writeCrashGuardFixture(t *testing.T, body string) string {
t.Helper()
p := filepath.Join(t.TempDir(), "state.json")
if err := os.WriteFile(p, []byte(body), 0o644); err != nil {
t.Fatal(err)
}
return p
}
// getCrashGuard calls the route and decodes the envelope with the CONTROLLER's type.
func getCrashGuard(t *testing.T, h http.Handler, token string) (int, controllerCrashGuardState, string) {
t.Helper()
w := do(t, h, "GET", "/host/crash-guard", token, "")
var env struct {
OK bool `json:"ok"`
Data controllerCrashGuardState `json:"data"`
}
if w.Code == http.StatusOK {
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil {
t.Fatalf("decode %q: %v", w.Body.String(), err)
}
if !env.OK {
t.Fatalf("ok=false: %s", w.Body.String())
}
}
return w.Code, env.Data, w.Body.String()
}
// A present state file (demo-hp's shape) passes the three facts through.
func TestR856_CrashGuardPresentFile(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, body)
}
want := controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: true, Tripped: false}
if st != want {
t.Fatalf("state = %+v, want %+v", st, want)
}
// A tripped, clean boot reads back as such (both bools are carried, not defaulted).
h = newCrashGuardServer(t, writeCrashGuardFixture(t,
`{"last_boot_at":"2026-10-05T09:56:41+02:00","last_boot_unclean":false,"tripped":true,"version":1}`))
_, st, _ = getCrashGuard(t, h, "B")
want = controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: false, Tripped: true}
if st != want {
t.Fatalf("offset time / tripped: state = %+v, want %+v (time normalised to UTC Z)", st, want)
}
}
// No state file (no guard on this host, or no boot recorded yet) → 200 present:false.
func TestR856_CrashGuardMissingFile(t *testing.T) {
h := newCrashGuardServer(t, filepath.Join(t.TempDir(), "absent", "state.json"))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("missing file: got %d, want 200 (%s)", code, body)
}
if st.Present || st.LastBootUnclean || st.Tripped || st.LastBootAt != "" {
t.Fatalf("missing file: state = %+v, want present:false and nothing else", st)
}
}
// A garbled file → 200 present:false, never a 5xx — every shape of garbage.
func TestR856_CrashGuardGarbageFile(t *testing.T) {
for name, body := range map[string]string{
"truncated": crashGuardFixture[:40],
"not json": "this is not json\n",
"empty": "",
"null": "null",
"array": `[{"last_boot_unclean":true}]`,
"wrong type": `{"last_boot_at":"2026-10-05T07:56:41Z","last_boot_unclean":"yes","tripped":false}`,
"lone brace": "{",
} {
t.Run(name, func(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, body))
code, st, raw := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, raw)
}
if st.Present || st.LastBootUnclean {
t.Fatalf("garbage %q read as %+v, want present:false", name, st)
}
})
}
// The path is a directory, not a file: unreadable → present:false, 200.
h := newCrashGuardServer(t, t.TempDir())
if code, st, raw := getCrashGuard(t, h, "A"); code != http.StatusOK || st.Present {
t.Fatalf("directory path: got %d %+v (%s), want 200 present:false", code, st, raw)
}
}
// No / unknown token → 401, like every sibling route; a cross-guest ?vmid= → 403.
func TestR856_CrashGuardRequiresGuestToken(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
for _, tok := range []string{"", "bogus"} {
w := do(t, h, "GET", "/host/crash-guard", tok, "")
if w.Code != http.StatusUnauthorized {
t.Fatalf("token %q: got %d, want 401", tok, w.Code)
}
if json.Valid(w.Body.Bytes()) {
var env struct {
Data controllerCrashGuardState `json:"data"`
}
_ = json.Unmarshal(w.Body.Bytes(), &env)
if env.Data.Present || env.Data.LastBootUnclean {
t.Fatalf("token %q: the refusal leaked the state: %s", tok, w.Body.String())
}
}
}
if w := do(t, h, "GET", "/host/crash-guard?vmid=9300", "A", ""); w.Code != http.StatusForbidden {
t.Fatalf("cross-guest query: got %d, want 403", w.Code)
}
}
// Wire contract: every key the controller's CrashGuardState decodes is emitted under exactly that
// name, and the agent emits no key the controller does not know.
func TestR856_CrashGuardWireMatchesControllerClient(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
w := do(t, h, "GET", "/host/crash-guard", "A", "")
var env struct {
OK bool `json:"ok"`
Data map[string]json.RawMessage `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil || !env.OK {
t.Fatalf("envelope: %v %s", err, w.Body.String())
}
want := []string{"present", "last_boot_at", "last_boot_unclean", "tripped"}
for _, k := range want {
if _, ok := env.Data[k]; !ok {
t.Errorf("agent does not emit %q, which the controller decodes (%s)", k, w.Body.String())
}
}
if len(env.Data) != len(want) {
t.Errorf("agent emits %d keys, controller knows %d: %s", len(env.Data), len(want), w.Body.String())
}
// And the agent's own type agrees with the controller's copy, field for field.
var mine CrashGuardResponse
var theirs controllerCrashGuardState
raw, _ := json.Marshal(env.Data)
_ = json.Unmarshal(raw, &mine)
_ = json.Unmarshal(raw, &theirs)
if (controllerCrashGuardState{mine.Present, mine.LastBootAt, mine.LastBootUnclean, mine.Tripped}) != theirs {
t.Errorf("agent %+v vs controller %+v", mine, theirs)
}
}
+10 -9
View File
@@ -3,7 +3,6 @@ package localapi
import (
"context"
"fmt"
"io"
"log/slog"
"strconv"
"strings"
@@ -126,14 +125,16 @@ func (b *GuestBinder) GuestExec(ctx context.Context, vmid int, args ...string) (
return string(out), nil
}
// GuestExecStdin is GuestExec with the in-guest command's stdin fed from stdin. The controller-swap
// write uses it to pipe the image ref into an in-guest `tee` (no shell vector, no interpolation),
// through the same fenced runner so the `sudo -n` prefix stays in one place.
func (b *GuestBinder) GuestExecStdin(ctx context.Context, vmid int, stdin io.Reader, args ...string) (string, error) {
pctArgs := append([]string{"exec", strconv.Itoa(vmid), "--"}, args...)
out, stderr, err := b.runner.RunStdin(ctx, stdin, "pct", pctArgs...)
// privApplyBin is the root content checker (R-861); its `controller-image` verb writes the guest's image file.
const privApplyBin = "/usr/local/sbin/felhom-priv-apply"
// WriteControllerImage pipes `image\n` to `felhom-priv-apply controller-image <vmid>` through the same fenced runner
// (the `sudo -n` prefix stays in one place). The verb checks the ref as root and writes the guest file itself
// (R-861 (a) A1, `09` §3 decision 165). Pinned by TestR861_WriteControllerImageUsesTheRootVerb.
func (b *GuestBinder) WriteControllerImage(ctx context.Context, vmid int, image string) error {
_, stderr, err := b.runner.RunStdin(ctx, strings.NewReader(image+"\n"), privApplyBin, "controller-image", strconv.Itoa(vmid))
if err != nil {
return string(out), fmt.Errorf("pct exec %d %v: %w: %s", vmid, args, err, strings.TrimSpace(string(stderr)))
return fmt.Errorf("felhom-priv-apply controller-image %d: %w: %s", vmid, err, strings.TrimSpace(string(stderr)))
}
return string(out), nil
return nil
}
@@ -0,0 +1,48 @@
package localapi
import (
"context"
"io"
"testing"
)
// stdinRecorder is a proxmox.Runner that records each call and the stdin it was handed.
type stdinRecorder struct {
name string
args []string
stdin string
}
func (r *stdinRecorder) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
r.name, r.args = name, args
return nil, nil, nil
}
func (r *stdinRecorder) RunStdin(_ context.Context, stdin io.Reader, name string, args ...string) ([]byte, []byte, error) {
b, _ := io.ReadAll(stdin)
r.name, r.args, r.stdin = name, args, string(b)
return nil, nil, nil
}
// R-861 (a) A1 (`09` §3 decision 165): the managed controller update writes the guest's image file through the ROOT
// verb, never through an in-guest `tee` the agent could feed any image.
//
// COMPANION RED-PROOF (observed): restore the pre-A1 body (`b.runner.RunStdin(ctx, …, "pct", "exec", vmid, "--",
// "tee", controllerImageFile)`) → this fails with "the image write ran pct …, want felhom-priv-apply". Restored.
func TestR861_WriteControllerImageUsesTheRootVerb(t *testing.T) {
rec := &stdinRecorder{}
b := NewGuestBinder(rec, discardLogger())
const img = "gitea.dooplex.hu/admin/felhom-controller:0.302.0"
if err := b.WriteControllerImage(context.Background(), 9201, img); err != nil {
t.Fatal(err)
}
if rec.name != privApplyBin {
t.Fatalf("the image write ran %s %v, want felhom-priv-apply", rec.name, rec.args)
}
if len(rec.args) != 2 || rec.args[0] != "controller-image" || rec.args[1] != "9201" {
t.Fatalf("argv = %v, want [controller-image 9201] (the sudoers line `^controller-image [0-9]+$`)", rec.args)
}
if rec.stdin != img+"\n" {
t.Fatalf("stdin = %q, want the ref plus one newline", rec.stdin)
}
}
+144
View File
@@ -0,0 +1,144 @@
package localapi
import (
"context"
"io"
"log/slog"
"net/http"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-894 — after an agent restart, an UNREADABLE storage must fall back to the last success saved on
// disk, not to "never". Measured 2026-10-05 on demo-hp: a restart at 04:57, the off-site storage
// unreachable at 06:25, the 7-day tier (last copy 4 days old) read DUE, vzdump failed.
//
// Every test here builds a NEW server and a NEW BackupSuccessState from the same file — that is the
// restart. The in-memory store (fakeStore) is always fresh, as after a real restart.
// r894Server builds a two-tier server whose off-site tier answers the storage listing with lister.
func r894Server(t *testing.T, path string, pbsSvc BackupService) *Server {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: &fakeStore{},
Storage: fakeStorage{targets: []hub.StorageTarget{{Name: "local"}, {Name: "felhom-pbs"}}},
Tokens: staticTokens{"A": 8200},
BackupTiers: []BackupTier{
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: pbsSvc},
},
LastKnownBackups: backup.NewBackupSuccessState(path),
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatal(err)
}
srv.baseCtx = context.Background()
srv.now = func() time.Time { return testNow }
return srv
}
// unreadable is the off-site storage as demo-hp saw it: "Can't connect to 10.77.0.1:8007".
func unreadable() archiveLister {
return archiveLister{fakeBackups: &fakeBackups{}, err: errStorageRead}
}
// THE R-894 CASE, end to end. Agent 1 takes an off-site backup through POST /backup (the fake
// runner's success is 12 h before testNow). The agent restarts. The storage cannot be read. The tier
// must read NOT due, from the copy saved on disk.
//
// COMPANION RED-PROOF (observed): delete the `lookup == archiveUnknown && s.lastKnown != nil` block in
// handleBackupDue → this fails with "after a restart an unreadable storage must fall back to the saved
// copy (12 h old, 7-day tier) — NOT due; got {… Due:true … AgeState:unknown …}". Restored.
func TestBackupDue_R894_RestartThenUnreadableStorage_FreshSavedCopyIsNotDue(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
// Agent 1: a real backup job through the endpoint the controller calls.
first := r894Server(t, path, &fakeBackups{})
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
}
waitFor(t, func() bool {
_, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200)
return ok
})
// Agent 2: a restart (new server, new state from the same file), and the storage is unreachable.
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if got.Due {
t.Fatalf("after a restart an unreadable storage must fall back to the saved copy (12 h old, 7-day tier) — NOT due; got %+v", got)
}
if got.AgeState != AgeStateKnown || got.AgeSecs == nil || *got.AgeSecs != int64((12*time.Hour).Seconds()) {
t.Fatalf("the age must come from the saved copy (12 h, known); got %+v", got)
}
}
// The deliberate rule stays: an unreadable storage must not suppress a backup that IS due. A saved
// copy older than the cadence reads DUE.
//
// COMPANION RED-PROOF (observed): make the fallback answer not-due whenever a saved copy exists
// (`if fromDisk { …Due:false… }` before the cadence check) → this fails with "a saved copy 9 days old
// under a 7-day cadence MUST read due". Restored.
func TestBackupDue_R894_RestartThenUnreadableStorage_OldSavedCopyIsDue(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
st := backup.NewBackupSuccessState(path)
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, 9*24*time.Hour, true)); err != nil {
t.Fatal(err)
}
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("a saved copy 9 days old under a 7-day cadence MUST read due; got %+v", got)
}
if got.AgeState != AgeStateKnown {
t.Fatalf("the age is known (from disk); got %+v", got)
}
}
// No saved copy → the pre-R-894 answer, byte for byte: DUE, age UNKNOWN (never ABSENT — the controller
// fires its window-gate valve only on absent, R-88).
func TestBackupDue_R894_RestartThenUnreadableStorage_NoSavedCopyIsDueUnknown(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due || got.AgeState != AgeStateUnknown || got.AgeSecs != nil {
t.Fatalf("no saved copy + unreadable storage must stay DUE with age unknown; got %+v", got)
}
}
// A storage that ANSWERS is the ground truth: an archive absent there makes the tier due even when the
// file remembers a fresh success (a pruned or deleted copy must be made again).
//
// COMPANION RED-PROOF (observed): drop `lookup == archiveUnknown &&` from the fallback condition → this
// fails with "the storage answered 'no archive' — the saved copy must NOT stand in for it". Restored.
func TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
st := backup.NewBackupSuccessState(path)
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, time.Hour, true)); err != nil {
t.Fatal(err)
}
absent := archiveLister{fakeBackups: &fakeBackups{}, found: false}
got := dueFor(t, r894Server(t, path, absent).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("the storage answered 'no archive' — the saved copy must NOT stand in for it; got %+v", got)
}
}
// A FAILED backup is never saved: it must not make a tier look fresh after a restart.
func TestBackupDue_R894_FailedBackupIsNotSaved(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
first := r894Server(t, path, &fakeBackups{failErr: "could not activate storage 'felhom-pbs'"})
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
}
waitFor(t, func() bool { return len(first.store.Backups(context.Background())) == 1 })
if _, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200); ok {
t.Fatal("a failed backup must never be saved as a success")
}
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("after a failed backup and a restart the tier must still be due; got %+v", got)
}
}
+67 -22
View File
@@ -85,6 +85,13 @@ type BackupStore interface {
RestoreTests(ctx context.Context) []hub.RestoreTest
}
// LastKnownBackupStore (R-894) is the on-disk newest-success-per-tier record. Satisfied by
// *backup.BackupSuccessState.
type LastKnownBackupStore interface {
RecordBackupSuccess(target string, b hub.Backup) error
LastKnownSuccess(target string, vmid int) (time.Time, bool)
}
// StorageView yields the host's observed storage targets (for mapping a mount's storage id →
// fast/slow class). Satisfied by *storage.Observer.
type StorageView interface {
@@ -164,6 +171,10 @@ type Options struct {
// PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg
// that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs.
AfterPrimaryBackup func(ctx context.Context, vmid int)
// LastKnownBackups (R-894) keeps the newest successful backup per tier ON DISK, so the due-check's
// fallback for an UNREADABLE storage after an agent restart is the last known copy, not "never".
// nil = the pre-R-894 behaviour (in-memory record only).
LastKnownBackups LastKnownBackupStore
// Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when
// nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner.
Privileged PrivilegedRunner
@@ -229,7 +240,6 @@ type Options struct {
// POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured"
// (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer.
EscrowRecovery EscrowRecoverer
}
// defaultBackupCadence is the fallback /backup/due window when none is configured.
@@ -279,30 +289,34 @@ type Server struct {
// is the pre-R-82 shape.
tiers []BackupTier
// inFlight (R-85) is shared with the restore-test scheduler so the two never run together.
inFlight *backup.InFlight
inFlight *backup.InFlight
afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none
logger *slog.Logger
now func() time.Time
lastKnown LastKnownBackupStore // R-894: on-disk newest success per tier; nil = none
logger *slog.Logger
now func() time.Time
disks DiskOps // slice 8C (optional)
diskGate StorageGate // slice 8C (optional)
guestList GuestLister // slice 8C (optional)
guestAttach GuestAttacher // slice 10 P2 (optional)
mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional)
memMu sync.Mutex // single-flight around a resize apply (one customer per host)
netStorage NetworkStorageOps // Part A1: NAS network mounts (optional)
netMountRoot string // the user-data namespace root for the network-mount role gate
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
disks DiskOps // slice 8C (optional)
diskGate StorageGate // slice 8C (optional)
guestList GuestLister // slice 8C (optional)
guestAttach GuestAttacher // slice 10 P2 (optional)
mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional)
memMu sync.Mutex // single-flight around a resize apply (one customer per host)
netStorage NetworkStorageOps // Part A1: NAS network mounts (optional)
netMountRoot string // the user-data namespace root for the network-mount role gate
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
// crashGuardStatePath (R-856) is the host crash guard's state file read by GET /host/crash-guard;
// empty = defaultCrashGuardStatePath. A seam: tests point it at a fixture.
crashGuardStatePath string
// escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
// offsite repository password. OPTIONAL — nil (no hub client configured) makes
// POST /escrow/recover-offsite-password answer 503 rather than pretending.
escrowRecovery EscrowRecoverer
intent IntentRecorder // slice 10 P3 (optional)
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
staleLock StaleLockController // F2-b startup stale-lock recovery (optional)
escrowRecovery EscrowRecoverer
intent IntentRecorder // slice 10 P3 (optional)
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
staleLock StaleLockController // F2-b startup stale-lock recovery (optional)
// guestPower (F-REBOOT) is per-guest start-attempt state for the guest-power watchdog.
// Guarded by guestPowerMu in guestpower.go; in-memory on purpose (see guestPowerState).
guestPower map[int]guestPowerState
@@ -475,6 +489,7 @@ func NewServer(o Options) (*Server, error) {
s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence)
s.inFlight = o.InFlight
s.afterPrimaryBackup = o.AfterPrimaryBackup
s.lastKnown = o.LastKnownBackups
if s.backups == nil && len(s.tiers) > 0 {
s.backups = s.tiers[0].Service
}
@@ -520,6 +535,9 @@ func (s *Server) Handler() http.Handler {
// Host metrics (slice 9): host-wide health + per-storage capacity for the customer's monitoring
// view. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot).
mux.HandleFunc("GET /host/metrics", s.withGuest(s.handleHostMetrics))
// R-856 (`09` §3 decision 143): the host crash guard's record of the last HOST boot — the controller
// waits ~15 min with app mails after an unclean one. Read-only; a missing/garbled file = present:false.
mux.HandleFunc("GET /host/crash-guard", s.withGuest(s.handleCrashGuard))
// Disk management (slice 8C) — self-scoped; format routes through the data-bearing classifier+gate.
mux.HandleFunc("GET /disks", s.withGuest(s.handleDisks))
mux.HandleFunc("GET /disks/candidates", s.withGuest(s.handleDiskCandidates))
@@ -899,6 +917,13 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
s.logger.Info("local-api: backup job complete", "vmid", vmid, "target", tier.TargetID, "job", jobID, "archive", b.Archive)
}
s.store.RecordBackup(b)
if b.Success && s.lastKnown != nil {
if err := s.lastKnown.RecordBackupSuccess(tier.TargetID, b); err != nil {
// Not fatal: the backup exists. Only the fallback after a restart loses this copy.
s.logger.Warn("local-api: could not save the backup on disk for the due-check fallback (R-894)",
"vmid", vmid, "target", tier.TargetID, "err", err)
}
}
s.finishJob(key, jobID, b)
// OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate.
if b.Success && tier.Primary && s.afterPrimaryBackup != nil {
@@ -1061,6 +1086,20 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
newest, haveNewest = t, true
unparseable = false // ground truth supersedes an unreadable in-memory timestamp
}
// R-894: the storage could not be read → the last success saved ON DISK stands in for the in-memory
// record a restart emptied. ONLY on archiveUnknown: a storage that answers is the ground truth, and an
// archive absent there must make the tier due even when the file remembers one (a pruned copy).
// A saved copy older than the cadence still reads DUE below — an unreadable storage never suppresses
// a backup that is due.
fromDisk := false
if lookup == archiveUnknown && s.lastKnown != nil {
if saved, ok := s.lastKnown.LastKnownSuccess(tier.TargetID, vmid); ok && (!haveNewest || saved.After(newest)) {
newest, haveNewest, fromDisk = saved, true, true
unparseable = false
s.logger.Info("local-api: backup storage unreadable — due-check uses the last success saved on disk (R-894)",
"vmid", vmid, "target", tier.TargetID, "saved", saved.UTC().Format(time.RFC3339))
}
}
if !haveNewest {
// R-88 Part 2: THREE distinct reasons for a nil age, each with its own state. Only ABSENT is a
// positive claim of "never backed up"; only that one may license the controller to bypass its
@@ -1083,13 +1122,17 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
}
age := s.now().Sub(newest)
ageSecs := int64(age.Seconds())
suffix := ""
if fromDisk {
suffix = " (storage unreadable — age from the last success saved on disk)"
}
if age >= tier.Cadence {
writeOK(w, BackupDueResponse{VMID: vmid, Due: true, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
Reason: "older than cadence", Target: echo})
Reason: "older than cadence" + suffix, Target: echo})
return
}
writeOK(w, BackupDueResponse{VMID: vmid, Due: false, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
Reason: "within cadence window", Target: echo})
Reason: "within cadence window" + suffix, Target: echo})
}
// BackupTiersResponse is GET /backup/tiers (R-82): the tiers this agent serves, primary first.
@@ -1477,4 +1520,6 @@ func writeStatus(w http.ResponseWriter, code int, ok bool, data any, errMsg stri
}
// SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0).
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) { s.afterPrimaryBackup = fn }
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) {
s.afterPrimaryBackup = fn
}
+146 -6
View File
@@ -20,6 +20,7 @@ import (
"log/slog"
"os"
"path/filepath"
"regexp"
"sort"
"strings"
"sync"
@@ -27,6 +28,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
)
// WrapperPath is the pinned sudoers vector (configs/felhom-agent.sudoers FELHOM_OSAPPLY).
@@ -40,8 +42,21 @@ const (
LayerGuest = "guest"
LayerHost = "host"
LayerDocker = "docker" // the guest's Docker engine set — slow lane (`11` §5.8)
// LayerPVE is the HOST's Proxmox userspace packages — slow lane (R-812 option A, `09` §3 decision 163, `11` §5.10):
// ring 0 in the night leg after a healthy host step, ring 1 only inside a signed os_pve_step. Never a kernel.
LayerPVE = "pve"
)
// PVEOrigin is the origin apt prints for download.proxmox.com (the wrapper's PVE_ORIGIN); InstalledPVEOrigin is how
// the wrapper's inventory names the same source.
const (
PVEOrigin = "Proxmox Debian Repository"
InstalledPVEOrigin = "Proxmox"
)
// hostSlowRE mirrors the wrapper's HOST_SLOW_RE: kernel, boot, firmware and microcode names never ride the pve lane.
var hostSlowRE = regexp.MustCompile(`^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr)`)
// DockerNames are the six packages of the Docker engine set (the wrapper's DOCKER_NAMES).
var DockerNames = map[string]bool{"containerd.io": true, "docker-buildx-plugin": true, "docker-ce": true,
"docker-ce-cli": true, "docker-ce-rootless-extras": true, "docker-compose-plugin": true}
@@ -106,6 +121,11 @@ type WrapperReport struct {
LiveRestore json.RawMessage `json:"live_restore"`
Facts json.RawMessage `json:"facts"`
Bundle json.RawMessage `json:"bundle"` // the config bundle's result (R-840, mode "bundle")
// OOMCheck (R-528, `09` decision 157): the docker layer's memory-kill check, {result, oom_killed, oom_event,
// exit_code, image, detail}. Carried to the hub UNCHANGED; the agent never reads it.
OOMCheck json.RawMessage `json:"oom_check"`
// PVEManager: the pve layer — pveversion's pve-manager version after the step ("unknown" = unreadable).
PVEManager string `json:"pve_manager"`
// R-868 (v0.144.0): the agent's ids, echoed from the plan, so a report kept on disk can be sent without the
// agent process that started the pass. ReleaseID / VMID were always in the report.
RunID string `json:"run_id"`
@@ -143,6 +163,11 @@ type Report struct {
DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step
Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed
Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade)
// OOMCheck: docker layer — the wrapper's oom_check object, byte-for-byte (R-528; the hub decides approval on it).
// Pinned by TestDocker_OOMCheckReachesTheHubUnchanged and TestR868_KeptCopyCarriesTheOOMCheck.
OOMCheck json.RawMessage `json:"oom_check,omitempty"`
// PVEManager: pve layer — pve-manager's version after the step (the hub's System page; R-812 option A).
PVEManager string `json:"pve_manager,omitempty"`
unsent string // R-868: the wrapper's kept copy of this pass's report — deleted once the hub has it
}
@@ -380,6 +405,36 @@ func EngineOf(pkgVersion string) string {
return v
}
// PVEHealthVerdict is THE Proxmox-package-step health rule (R-812 option A; pinned by TestPVEHealthVerdict): the host
// rule (every host service active, the guest running and healthy, the tunnel running), plus every container running at
// the start still runs as the SAME container (a Proxmox step must not restart the household's apps), plus pveversion
// now reports the pve-manager the step installed (wantPVE "" = pve-manager was not in the step).
func PVEHealthVerdict(before, after *Health, tunnel, wantPVE, gotPVE string) (bool, string) {
if ok, why := HostHealthVerdict(before, after, tunnel); !ok {
return false, why
}
if before != nil && before.Guest != nil && after.Guest != nil {
names := make([]string, 0, len(before.Guest.Containers))
for n := range before.Guest.Containers {
names = append(names, n)
}
sort.Strings(names)
for _, n := range names {
b := before.Guest.Containers[n]
if b.State != "running" || b.ID == "" {
continue
}
if a := after.Guest.Containers[n]; a.ID != b.ID {
return false, n + " is a new container (id changed) — the Proxmox step restarted the household's app"
}
}
}
if wantPVE != "" && gotPVE != wantPVE {
return false, "pveversion reads pve-manager " + gotPVE + ", not " + wantPVE
}
return true, ""
}
// DockerHealthVerdict is THE Docker-step health rule (`11` §5.8; pinned by TestDockerHealthVerdict): the guest rule,
// plus every container running at the start still runs as the SAME container (same id — a changed id means the
// household's apps restarted, which `live-restore` exists to prevent), plus the engine now reports the version the step
@@ -445,7 +500,7 @@ func (l *Leg) call(ctx context.Context, runID string, plan map[string]any) (Wrap
// Pass is one leg's reports; an empty Layer means the step did not run.
type Pass struct {
Guest, Host, Docker Report
Guest, Host, Docker, PVE Report
}
// Run is one pass: the guest layer, then (on an appliance, after a good guest step) the host layer, then (ring 0
@@ -473,13 +528,44 @@ func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Pass {
if err := l.EnsureLiveRestore(ctx, g.RunID, vmid); err != nil {
p.Docker = l.finish(ctx, lg, Report{RunID: g.RunID, Layer: LayerDocker, Trigger: trigger, Ring: 0, VMID: vmid,
Mode: "apply", Outcome: "failed", HealthReason: "live-restore could not be turned on: " + err.Error()})
return p
} else {
p.Docker = l.runLayer(ctx, g.RunID, LayerDocker, vmid, trigger, blk, dockerOpts{})
}
p.Docker = l.runLayer(ctx, g.RunID, LayerDocker, vmid, trigger, blk, dockerOpts{})
}
// R-812 option A: the Proxmox package step — ring 0, an appliance, after a HEALTHY host step (it is a host change).
// A Docker step's outcome does not gate it (the Docker set lives in the guest). Pinned by TestPVE_*.
switch {
case blk.Ring != 0 || !blk.Enabled:
lg.Info("osupdate: pve step skipped — ring 1 takes a Proxmox set only inside a signed operator job (`11` §5.10)", "ring", blk.Ring, "enabled", blk.Enabled)
case !l.Appliance || h.Layer == "" || !okStep(h):
lg.Info("osupdate: pve step skipped — no healthy host step this pass (an appliance only)", "appliance", l.Appliance, "host_outcome", h.Outcome)
default:
p.PVE = l.runPVE(ctx, g.RunID, vmid, trigger, blk, dockerOpts{})
}
return p
}
// pveDrainWait bounds how long a pve step waits for the agent's own /etc/pve writes in flight (pvegate).
var pveDrainWait = 2 * time.Minute
// runPVE runs the pve layer while holding pvegate: the agent's own /etc/pve writes wait until it ends.
func (l *Leg) runPVE(ctx context.Context, runID string, vmid int, trigger string, blk hub.WireOSUpdate, do dockerOpts) Report {
lg := l.log().With("run", runID, "layer", LayerPVE, "vmid", vmid, "trigger", trigger)
dctx, cancel := context.WithTimeout(ctx, pveDrainWait)
end, err := pvegate.Step(dctx)
cancel()
if err != nil {
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerPVE, Trigger: trigger, Ring: blk.Ring, VMID: vmid, Mode: "apply",
ReleaseID: do.releaseID, Outcome: "failed", HealthReason: "an agent write to /etc/pve did not finish in time (pvegate): " + err.Error()})
}
lg.Info("osupdate: pve step holds the /etc/pve write gate — the agent's own writes wait until it ends")
defer func() {
end()
lg.Info("osupdate: pve step released the /etc/pve write gate")
}()
return l.runLayer(ctx, runID, LayerPVE, vmid, trigger, blk, do)
}
// dockerOpts is a signed Docker step (DockerStepExecutor); the zero value is ring 0's unsigned "pending-docker".
type dockerOpts struct {
releaseID string
@@ -565,7 +651,7 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
wire = blk.HostRelease
}
lane := "fast"
if layer == LayerDocker {
if layer == LayerDocker || layer == LayerPVE {
lane = "slow"
if do.signed != nil {
rel = hub.WireOSRelease{ID: do.releaseID}
@@ -598,6 +684,13 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
}
case layer == LayerDocker:
plan["select"] = "pending-docker" // ring 0: the wrapper checks the box's ROOT-OWNED ring-0 mark itself
case layer == LayerPVE && do.signed != nil:
plan["packages"], plan["signed"] = do.packages, do.signed
for _, p := range do.packages {
planned[p.Name] = true
}
case layer == LayerPVE:
plan["select"] = "pending-pve" // ring 0: the same root-owned mark; the wrapper picks installed Proxmox userspace
case !blk.Enabled:
plan["mode"] = "inventory"
lg.Info("osupdate: switched OFF for this box — reporting only")
@@ -630,6 +723,8 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
}
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
rep.OOMCheck = rawOrNil(wr.OOMCheck)
rep.PVEManager = wr.PVEManager
if rep.Outcome == "" {
switch {
case rep.Mode == "inventory" && !blk.Enabled:
@@ -647,21 +742,27 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
}
// Health: compare with the start of the pass; give restarted services time (only after an install).
cur := wr.HealthAfter
wantEngine := ""
wantEngine, wantPVE := "", ""
for _, u := range wr.Upgraded {
if u.Name == "docker-ce" {
wantEngine = EngineOf(u.Version)
}
if u.Name == "pve-manager" {
wantPVE = u.Version
}
}
verdict := func(h *Health) (bool, string) {
if layer == LayerDocker {
return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine)
}
if layer == LayerHost {
if layer == LayerHost || layer == LayerPVE {
t := hub.TunnelUnknown
if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx)
}
if layer == LayerPVE {
return PVEHealthVerdict(wr.HealthBefore, h, t, wantPVE, wr.PVEManager)
}
return HostHealthVerdict(wr.HealthBefore, h, t)
}
return HealthVerdict(wr.HealthBefore, h)
@@ -701,12 +802,24 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
// the docker report carries the engine set only (the guest report already carries the Debian packages)
rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending)
rep.NotCovered = nil
} else if layer == LayerPVE {
// the pve report carries the Proxmox userspace set only — the hub's candidate is built from it
rep.Installed, rep.Pending = onlyPVE(wr.Installed), onlyPVEPending(wr.Pending)
rep.NotCovered = nil
} else {
rep.NotCovered = notCovered(wr.Pending, blk.Ring, planned)
}
return l.finish(ctx, lg, rep)
}
// rawOrNil: a wrapper field that is absent or JSON null stays out of the hub report (omitempty).
func rawOrNil(m json.RawMessage) json.RawMessage {
if len(m) == 0 || string(m) == "null" {
return nil
}
return m
}
func onlyDocker(in []Package) []Package {
var out []Package
for _, p := range in {
@@ -717,6 +830,33 @@ func onlyDocker(in []Package) []Package {
return out
}
// onlyPVE keeps the installed Proxmox-origin packages the pve lane may touch (never a kernel / boot / firmware name).
func onlyPVE(in []Package) []Package {
var out []Package
for _, p := range in {
if (p.Origin == InstalledPVEOrigin || p.Origin == PVEOrigin) && !hostSlowRE.MatchString(p.Name) && !DockerNames[p.Name] {
out = append(out, p)
}
}
return out
}
func onlyPVEPending(in []Pending) []Pending {
var out []Pending
for _, p := range in {
if p.From == "" || hostSlowRE.MatchString(p.Name) || DockerNames[p.Name] {
continue
}
for _, o := range p.Origin {
if o == PVEOrigin {
out = append(out, p)
break
}
}
}
return out
}
func onlyDockerPending(in []Pending) []Pending {
var out []Pending
for _, p := range in {
+66 -13
View File
@@ -13,16 +13,18 @@ import (
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
)
// fakeWrapper plays /usr/local/sbin/felhom-os-apply: it reads the plan the leg wrote and answers per layer and mode.
type fakeWrapper struct {
t *testing.T
pending []Pending
applyRep map[string]WrapperReport // per layer
healthSeq map[string][]*Health // per layer: answers to successive "health" calls
plans []map[string]any
keep bool // R-868: like the real wrapper, keep an apply report beside the plan
t *testing.T
pending []Pending
applyRep map[string]WrapperReport // per layer
healthSeq map[string][]*Health // per layer: answers to successive "health" calls
plans []map[string]any
keep bool // R-868: like the real wrapper, keep an apply report beside the plan
pveGateHeld bool
}
func yes() *bool { b := true; return &b }
@@ -50,9 +52,12 @@ func (f *fakeWrapper) Run(_ context.Context, name string, args ...string) ([]byt
f.plans = append(f.plans, plan)
layer := plan["layer"].(string)
ok := guestOK()
if layer == LayerHost {
if layer == LayerHost || layer == LayerPVE {
ok = hostOK()
}
if layer == LayerPVE && plan["mode"] == "apply" {
f.pveGateHeld = pvegate.Stepping() // R-812: the /etc/pve write gate must be held while the pve step runs
}
var rep WrapperReport
switch plan["mode"] {
case "inventory":
@@ -98,9 +103,13 @@ func (f *fakeWrapper) RunStdin(ctx context.Context, _ io.Reader, name string, ar
return f.Run(ctx, name, args...)
}
type fakeHub struct{ reports []Report }
type fakeHub struct {
reports []Report
bodies [][]byte // the exact bytes posted (R-528: the oom_check object must arrive unchanged)
}
func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error {
h.bodies = append(h.bodies, append([]byte(nil), body...))
var r Report
json.Unmarshal(body, &r)
h.reports = append(h.reports, r)
@@ -156,8 +165,8 @@ func TestRing0_OneCallPerLayer(t *testing.T) {
if g.Outcome != "applied" || !g.Healthy || ho.Outcome != "applied" || !ho.Healthy {
t.Fatalf("guest %+v\nhost %+v", g, ho)
}
if calls(w) != "guest:apply,host:apply,guest:live-restore-on,docker:apply" {
t.Fatalf("calls = %s, want one apply per layer, guest first, then live-restore and the ring-0 docker step", calls(w))
if calls(w) != "guest:apply,host:apply,guest:live-restore-on,docker:apply,pve:apply" {
t.Fatalf("calls = %s, want one apply per layer, guest first, then live-restore, the ring-0 docker step and the pve step", calls(w))
}
for _, p := range w.plans[:2] {
if p["select"] != "pending-fast" || p["snapshot"] != "" || len(p["packages"].([]any)) != 0 {
@@ -167,7 +176,7 @@ func TestRing0_OneCallPerLayer(t *testing.T) {
if len(g.NotCovered) != 1 || g.NotCovered[0] != "docker-ce" {
t.Fatalf("not covered = %v", g.NotCovered)
}
if len(h.reports) != 3 || h.reports[0].Layer != LayerGuest || h.reports[1].Layer != LayerHost || h.reports[2].Layer != LayerDocker {
if len(h.reports) != 4 || h.reports[0].Layer != LayerGuest || h.reports[1].Layer != LayerHost || h.reports[2].Layer != LayerDocker || h.reports[3].Layer != LayerPVE {
t.Fatalf("hub got %+v", h.reports)
}
}
@@ -385,7 +394,7 @@ func TestHostReport_CarriesRebootScanned(t *testing.T) {
}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
run2(l, "night")
if len(h.reports) != 3 || !h.reports[1].RebootScanned || !h.reports[1].RebootNeeded || h.reports[0].RebootScanned {
if len(h.reports) != 4 || !h.reports[1].RebootScanned || !h.reports[1].RebootNeeded || h.reports[0].RebootScanned {
t.Fatalf("hub got %+v", h.reports)
}
}
@@ -432,7 +441,12 @@ func TestDocker_Ring0PlanAndReport(t *testing.T) {
DockerEngine: "29.8.2", Authority: "ring0"}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
dp := w.plans[len(w.plans)-1]
var dp map[string]any
for _, x := range w.plans {
if x["layer"] == "docker" {
dp = x
}
}
if dp["layer"] != "docker" || dp["lane"] != "slow" || dp["select"] != "pending-docker" {
t.Fatalf("docker plan = %v", dp)
}
@@ -504,3 +518,42 @@ func TestHealthVerdict_ControllerBlindToDockerFails(t *testing.T) {
t.Fatal("an older wrapper (no field) must not fail")
}
}
// R-528 (`09` decision 157): the wrapper's oom_check object reaches the hub's docker report byte-for-byte; the guest
// and host reports carry none. COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in runLayer → "no
// oom_check in the docker report".
const oomCheckWire = `{"detail":"the engine reported the memory kill: OOMKilled=true and the oom event","exit_code":137,"image":"gitea.dooplex.hu/admin/felhom-controller:0.300.0","oom_event":true,"oom_killed":true,"result":"pass"}`
func TestDocker_OOMCheckReachesTheHubUnchanged(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}},
DockerEngine: "29.8.2", Authority: "ring0", OOMCheck: json.RawMessage(oomCheckWire)}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Run(context.Background(), 9201, "night")
found := false
for _, b := range h.bodies {
var m map[string]json.RawMessage
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
var layer string
json.Unmarshal(m["layer"], &layer)
oc, has := m["oom_check"]
if layer != LayerDocker {
if has {
t.Fatalf("the %s report carries an oom_check: %s", layer, oc)
}
continue
}
found = true
if !has {
t.Fatalf("no oom_check in the docker report: %s", b)
}
if string(oc) != oomCheckWire {
t.Fatalf("oom_check changed on the way:\n got %s\nwant %s", oc, oomCheckWire)
}
}
if !found {
t.Fatalf("no docker report posted: %s", calls(w))
}
}
+152
View File
@@ -0,0 +1,152 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// ---- the Proxmox package step (R-812 option A, `09` §3 decision 163, `11` §5.10) ----
// Ring 0: after a healthy host step the leg runs the pve layer — slow lane, select pending-pve — while holding the
// /etc/pve write gate; the report carries only Proxmox userspace packages and pve-manager's version.
//
// COMPANION RED-PROOF (observed): call runLayer instead of runPVE in Run → "the /etc/pve write gate was not held".
func TestPVE_Ring0PlanGateAndReport(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerPVE: {
Upgraded: []Package{{Name: "pve-manager", Version: "9.2.21"}},
Installed: []Package{{Name: "pve-manager", Version: "9.2.21", Origin: "Proxmox"},
{Name: "proxmox-kernel-helper", Version: "9.0.4", Origin: "Proxmox"}, {Name: "libc6", Version: "u4", Origin: "Debian"}},
Pending: []Pending{{Name: "qemu-server", From: "9.0.1", To: "9.0.9", Origin: []string{PVEOrigin}},
{Name: "proxmox-kernel-7.0", From: "7.0.2", To: "7.0.14", Origin: []string{PVEOrigin}},
{Name: "libc6", From: "u3", To: "u4", Origin: []string{"Debian"}}},
PVEManager: "9.2.21", Authority: "ring0"}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
var pp map[string]any
for _, x := range w.plans {
if x["layer"] == LayerPVE {
pp = x
}
}
if pp == nil || pp["lane"] != "slow" || pp["select"] != "pending-pve" {
t.Fatalf("pve plan = %v (calls %s)", pp, calls(w))
}
if !w.pveGateHeld {
t.Fatal("the /etc/pve write gate was not held while the pve step ran")
}
if pvegate.Stepping() {
t.Fatal("the gate must be released after the step")
}
r := p.PVE
if r.Outcome != "applied" || !r.Healthy || r.PVEManager != "9.2.21" {
t.Fatalf("pve report = %+v", r)
}
if len(r.Installed) != 1 || r.Installed[0].Name != "pve-manager" || len(r.Pending) != 1 || r.Pending[0].Name != "qemu-server" {
t.Fatalf("the pve report must carry Proxmox userspace only: installed=%v pending=%v", r.Installed, r.Pending)
}
if h.reports[len(h.reports)-1].Layer != LayerPVE {
t.Fatalf("the hub must get the pve report: %+v", h.reports)
}
}
// Ring 1 never takes a Proxmox step in the night leg.
func TestPVE_Ring1NightLegNeverSteps(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
if p := l.Run(context.Background(), 9201, "night"); p.PVE.Layer != "" || strings.Contains(calls(w), "pve") {
t.Fatalf("ring 1 took a pve step: %s", calls(w))
}
}
// No healthy host step (a BYO box, or an unhealthy host step) → no pve step.
func TestPVE_SkippedWithoutAHealthyHostStep(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Appliance = false
if p := l.Run(context.Background(), 9201, "night"); p.PVE.Layer != "" || strings.Contains(calls(w), "pve") {
t.Fatalf("a BYO box took a pve step: %s", calls(w))
}
w2 := &fakeWrapper{t: t}
l2, _ := newLeg(t, w2, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l2.Tunnel = fakeTunnel{"stopped"} // the host step reads unhealthy
w2.applyRep = map[string]WrapperReport{LayerHost: {Upgraded: []Package{{Name: "libc6", Version: "u4"}}}}
if p := l2.Run(context.Background(), 9201, "night"); p.PVE.Layer != "" || strings.Contains(calls(w2), "pve") {
t.Fatalf("a pve step ran after an unhealthy host step: %s", calls(w2))
}
}
// A write in flight that never finishes makes the pve step give up (failed), never run without the gate.
func TestPVE_GivesUpWhenAWriteDoesNotFinish(t *testing.T) {
old := pveDrainWait
pveDrainWait = 50_000_000 // 50 ms
defer func() { pveDrainWait = old }()
rel, _, _ := pvegate.Write(context.Background())
defer rel()
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
if p.PVE.Outcome != "failed" || strings.Contains(calls(w), "pve") {
t.Fatalf("the pve step must fail without a wrapper call: %+v calls=%s", p.PVE, calls(w))
}
}
// THE pve health rule. COMPANION RED-PROOF (observed): drop the container-id loop or the pve-manager check in
// PVEHealthVerdict → the matching case below fails.
func TestPVEHealthVerdict(t *testing.T) {
before, after := hostOK(), hostOK()
before.Guest.Containers["app"] = Container{State: "running", Health: "healthy", ID: "a1"}
after.Guest.Containers["app"] = Container{State: "running", Health: "healthy", ID: "a1"}
if ok, why := PVEHealthVerdict(before, after, hub.TunnelRunning, "9.2.21", "9.2.21"); !ok {
t.Fatalf("healthy step read unhealthy: %s", why)
}
if ok, _ := PVEHealthVerdict(before, after, hub.TunnelRunning, "9.2.21", "9.2.2"); ok {
t.Fatal("pveversion still on the old pve-manager must fail")
}
after.Guest.Containers["app"] = Container{State: "running", Health: "healthy", ID: "b2"}
if ok, why := PVEHealthVerdict(before, after, hub.TunnelRunning, "", "9.2.2"); ok || !strings.Contains(why, "id changed") {
t.Fatalf("an app restarted by the Proxmox step must fail, got ok=%v %q", ok, why)
}
}
// The signed executor hands the RAW envelope and the exact list to the wrapper's pve layer.
func TestPVEStepExecutor_PassesTheSignedEnvelope(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerPVE: {
Upgraded: []Package{{Name: "pve-manager", Version: "9.2.21"}}, PVEManager: "9.2.21", Authority: "signed"}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
e := PVEStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
params, _ := json.Marshal(PVEStepParams{ReleaseID: "os-pve-1", Packages: []Package{{Name: "pve-manager", Version: "9.2.21", Origin: PVEOrigin}}})
ctx := signedjobs.WithSignedOp(context.Background(), &reconcile.SignedOp{Blob: []byte(`{"op":"os_pve_step"}`), Sig: []byte("SIG")})
if err := e.Execute(ctx, OpPVEStep, params); err != nil {
t.Fatal(err)
}
pp := w.plans[len(w.plans)-1]
sg, _ := pp["signed"].(map[string]any)
if pp["layer"] != LayerPVE || pp["lane"] != "slow" || pp["release_id"] != "os-pve-1" || sg == nil ||
sg["blob_b64"] != base64.StdEncoding.EncodeToString([]byte(`{"op":"os_pve_step"}`)) || sg["sig"] != "SIG" {
t.Fatalf("pve plan = %v", pp)
}
if !w.pveGateHeld || calls(w) != "pve:apply" || len(h.reports) != 1 || h.reports[0].Trigger != "signed" {
t.Fatalf("gate=%v calls=%s reports=%+v", w.pveGateHeld, calls(w), h.reports)
}
if err := e.Execute(context.Background(), OpPVEStep, params); err == nil {
t.Fatal("no envelope must refuse")
}
if err := e.Execute(context.Background(), OpDockerStep, params); err != signedjobs.ErrNoExecutor {
t.Fatalf("another op must pass through the chain: %v", err)
}
}
// os_pve_step is never benign.
func TestPVEStep_IsDestructiveClass(t *testing.T) {
if reconcile.Classify(reconcile.ClassOSPVEStep, reconcile.Provenance{}) != reconcile.Destructive {
t.Fatal("os_pve_step must be destructive-class (signed, operational key)")
}
}
+85
View File
@@ -0,0 +1,85 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// OpPVEStep is the signed op class of a Proxmox package step (R-812 option A, `11` §5.10): a ring-1 box takes an
// approved Proxmox set only through it. No undo in this release. CC may sign it until the first paying customer.
const OpPVEStep = "os_pve_step"
// PVEStepParams are the signed params. The wrapper compares Packages with the plan byte-for-byte.
type PVEStepParams struct {
ReleaseID string `json:"release_id"`
Packages []Package `json:"packages"`
VMID int `json:"vmid,omitempty"`
}
// PVEStepExecutor runs a verified os_pve_step (signedjobs.Executor) under the host-wide heavy-op gate (Gate) and the
// /etc/pve write gate (inside runPVE).
type PVEStepExecutor struct {
Leg *Leg
Guest func(ctx context.Context) (int, error)
Gate func(ctx context.Context) (release func(), err error)
}
// Execute implements signedjobs.Executor.
func (e PVEStepExecutor) Execute(ctx context.Context, op string, params json.RawMessage) error {
if op != OpPVEStep {
return signedjobs.ErrNoExecutor
}
so, ok := signedjobs.SignedOpFrom(ctx)
if !ok {
return fmt.Errorf("os_pve_step: no signed envelope in the context — the wrapper could not verify it")
}
var p PVEStepParams
if err := json.Unmarshal(params, &p); err != nil || len(p.Packages) == 0 {
return fmt.Errorf("os_pve_step: params must name the Proxmox set: %v", err)
}
vmid := p.VMID
if vmid == 0 {
if e.Guest == nil {
return fmt.Errorf("os_pve_step: no vmid and no guest finder")
}
v, err := e.Guest(ctx)
if err != nil {
return fmt.Errorf("os_pve_step: find the customer guest: %w", err)
}
vmid = v
}
if e.Gate != nil {
release, err := e.Gate(ctx)
if err != nil {
return fmt.Errorf("os_pve_step: heavy-op gate busy (a backup or restore-test runs): %w", err)
}
defer release()
}
rep := e.Leg.RunPVESigned(ctx, vmid, p, so.Blob, string(so.Sig))
switch rep.Outcome {
case "applied", "nothing":
if rep.Healthy {
return nil
}
}
return fmt.Errorf("os_pve_step: %s (%s) %s", rep.Outcome, rep.HealthReason, string(rep.Refused))
}
// RunPVESigned is one signed Proxmox step (ring 1): the pve layer with the signed envelope, which the wrapper verifies
// itself, holding the /etc/pve write gate.
func (l *Leg) RunPVESigned(ctx context.Context, vmid int, p PVEStepParams, blob []byte, sig string) Report {
unlock := l.lockPass(true)
defer unlock()
l.sendUnsentLocked(ctx) // R-868
runID := l.now().UTC().Format("20060102T150405Z")
rid := p.ReleaseID
if rid == "" {
rid = "signed-" + runID
}
return l.runPVE(ctx, runID, vmid, "signed", l.Block(), dockerOpts{releaseID: rid, packages: p.Packages,
signed: map[string]string{"blob_b64": base64.StdEncoding.EncodeToString(blob), "sig": sig}})
}
+13 -3
View File
@@ -144,21 +144,29 @@ func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string)
}
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
wantEngine := ""
rep.OOMCheck = rawOrNil(wr.OOMCheck)
rep.PVEManager = wr.PVEManager
wantEngine, wantPVE := "", ""
for _, u := range wr.Upgraded {
if u.Name == "docker-ce" {
wantEngine = EngineOf(u.Version)
}
if u.Name == "pve-manager" {
wantPVE = u.Version
}
}
verdict := func(h *Health) (bool, string) {
switch wr.Layer {
case LayerDocker:
return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine)
case LayerHost:
case LayerHost, LayerPVE:
t := hub.TunnelUnknown
if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx)
}
if wr.Layer == LayerPVE {
return PVEHealthVerdict(wr.HealthBefore, h, t, wantPVE, wr.PVEManager)
}
return HostHealthVerdict(wr.HealthBefore, h, t)
}
return HealthVerdict(wr.HealthBefore, h)
@@ -166,7 +174,7 @@ func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string)
ok, why := verdict(wr.HealthAfter)
if !ok && len(wr.Upgraded) > 0 && wr.VMID > 0 {
lane := "fast"
if wr.Layer == LayerDocker {
if wr.Layer == LayerDocker || wr.Layer == LayerPVE {
lane = "slow"
}
if hr, err := l.call(ctx, "kept-"+runID, map[string]any{"release_id": "kept", "layer": wr.Layer, "lane": lane,
@@ -186,6 +194,8 @@ func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string)
rep.RebootScanned = wr.RebootScanned
if wr.Layer == LayerDocker {
rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending)
} else if wr.Layer == LayerPVE {
rep.Installed, rep.Pending = onlyPVE(wr.Installed), onlyPVEPending(wr.Pending)
} else {
planned := map[string]bool{}
for _, u := range wr.Upgraded {
+21
View File
@@ -150,3 +150,24 @@ func must(t *testing.T, err error) {
t.Fatal(err)
}
}
// R-528: a kept docker report (the agent was killed) still carries the oom_check object to the hub unchanged.
// COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in reportFromKept → "oom_check lost".
func TestR868_KeptCopyCarriesTheOOMCheck(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
ring := 0
kept := WrapperReport{Mode: "apply", Layer: LayerDocker, RunID: "20261007T020000Z", Trigger: "night", Ring: &ring, VMID: 9201,
ReleaseID: "ring0-20261007T020000Z", HealthBefore: guestOK(), HealthAfter: guestOK(), DockerEngine: "29.8.2",
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}}, OOMCheck: json.RawMessage(oomCheckWire)}
b, _ := json.Marshal(kept)
must(t, os.WriteFile(reportFile(l.PlanDir, kept.RunID, LayerDocker, "apply"), b, 0o600))
if n := l.SendUnsent(context.Background()); n != 1 || len(h.bodies) != 1 {
t.Fatalf("sent %d, bodies %d", n, len(h.bodies))
}
var m map[string]json.RawMessage
must(t, json.Unmarshal(h.bodies[0], &m))
if string(m["oom_check"]) != oomCheckWire {
t.Fatalf("oom_check lost or changed: %q", m["oom_check"])
}
}
+10
View File
@@ -5,6 +5,7 @@ import (
"context"
"encoding/json"
"fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
"io"
"net/http"
"net/url"
@@ -105,6 +106,15 @@ func (c *Client) do(ctx context.Context, method, path string, body io.Reader, ou
// doBody is the single HTTP chokepoint: builds the request, sets auth, executes,
// maps non-2xx to APIError, and decodes the data envelope.
func (c *Client) doBody(ctx context.Context, method, path string, body io.Reader, contentType string, out any) error {
// R-812 option A: every non-GET call may write /etc/pve — it waits while a Proxmox package step restarts pmxcfs
// (pvegate). Pinned by TestPVEGate_ClientWriteWaitsGetDoesNot.
if method != http.MethodGet {
release, _, gerr := pvegate.Write(ctx)
if gerr != nil {
return fmt.Errorf("proxmox: %s %s held back by a Proxmox package step: %w", method, path, gerr)
}
defer release()
}
req, err := http.NewRequestWithContext(ctx, method, c.base+path, body)
if err != nil {
return fmt.Errorf("proxmox: building request: %w", err)
+33
View File
@@ -4,8 +4,10 @@ import (
"context"
"encoding/json"
"fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
"io"
"os/exec"
"path/filepath"
"strconv"
)
@@ -59,6 +61,15 @@ func (r *ExecRunner) Run(ctx context.Context, name string, args ...string) ([]by
// RunStdin is Run with the process stdin fed from stdin (nil = no stdin). The sudo-prefix/mode
// handling is identical to Run — kept here so both paths share one place.
func (r *ExecRunner) RunStdin(ctx context.Context, stdin io.Reader, name string, args ...string) ([]byte, []byte, error) {
// R-812 option A: a root CLI that writes /etc/pve waits while a Proxmox package step runs (pvegate).
// Pinned by TestPVEGate_ExecRunnerPctSetWaits / TestWritesEtcPVE.
if WritesEtcPVE(name, args) {
release, _, gerr := pvegate.Write(ctx)
if gerr != nil {
return nil, nil, fmt.Errorf("proxmox: %s held back by a Proxmox package step: %w", name, gerr)
}
defer release()
}
var cmd *exec.Cmd
if r.Mode == RunnerSudo {
sudo := r.SudoPath
@@ -77,6 +88,28 @@ func (r *ExecRunner) RunStdin(ctx context.Context, stdin io.Reader, name string,
return stdout.b, stderr.b, err
}
// WritesEtcPVE reports whether a root command writes /etc/pve: `pct` with a config-changing verb, `pvesm`, `pveum`,
// and the PBS storage wrapper's create / reconcile verbs. `pct exec|status|list|config` and every other command do not
// (the os-update wrapper itself must never wait on the gate its own step holds). Pinned by TestWritesEtcPVE.
func WritesEtcPVE(name string, args []string) bool {
base := filepath.Base(name)
switch base {
case "pvesm", "pveum":
return true
case "pct":
if len(args) == 0 {
return false
}
switch args[0] {
case "set", "create", "destroy", "restore", "unlock", "resize", "snapshot", "delsnapshot", "rollback", "move-volume", "start", "stop", "reboot", "shutdown":
return true
}
case "felhom-pbs-apply":
return len(args) > 0 && (args[0] == "create" || args[0] == "reconcile")
}
return false
}
// Privileged is the root-CLI backend.
type Privileged struct {
runner Runner
+92
View File
@@ -0,0 +1,92 @@
package proxmox
import (
"context"
"net/http"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
)
// R-812 option A: while a Proxmox package step holds the gate, a non-GET API call waits and a GET does not.
//
// COMPANION RED-PROOF (observed): delete the pvegate.Write block in doBody → this fails with "a PUT reached the API
// while the Proxmox step held the gate". Restored. (audits/day-2026-10-07/B/red-pvegate-chokepoints.txt)
func TestPVEGate_ClientWriteWaitsGetDoesNot(t *testing.T) {
d := &mockDoer{fn: func(*http.Request) (*http.Response, error) { return jsonResp(200, `{"data":null}`), nil }}
c := newTestClient(d)
end, err := pvegate.Step(context.Background())
if err != nil {
t.Fatal(err)
}
if err := c.get(context.Background(), "/nodes", nil); err != nil {
t.Fatalf("a GET must not wait on the gate: %v", err)
}
if d.calls != 1 {
t.Fatalf("the GET must reach the API, calls=%d", d.calls)
}
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
err = c.postForm(ctx, http.MethodPut, "/nodes/x/lxc/9201/config", nil, nil)
if d.calls != 1 {
end()
t.Fatal("a PUT reached the API while the Proxmox step held the gate")
}
if err == nil {
end()
t.Fatal("a PUT held back past its deadline must fail")
}
end()
if err := c.postForm(context.Background(), http.MethodPut, "/nodes/x/lxc/9201/config", nil, nil); err != nil || d.calls != 2 {
t.Fatalf("after the step the PUT must go through (err=%v calls=%d)", err, d.calls)
}
}
// A root `pct set` waits on the gate; `pct exec` does not.
//
// COMPANION RED-PROOF (observed): delete the WritesEtcPVE block in RunStdin → this fails with "pct set ran while the
// Proxmox step held the gate". Restored.
func TestPVEGate_ExecRunnerPctSetWaits(t *testing.T) {
r := &ExecRunner{Mode: RunnerDirect}
end, err := pvegate.Step(context.Background())
if err != nil {
t.Fatal(err)
}
defer end()
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
start := time.Now()
_, _, err = r.Run(ctx, "/nonexistent/pct", "set", "9201", "-mp8", "/x")
if err == nil || time.Since(start) < 90*time.Millisecond {
t.Fatalf("pct set ran while the Proxmox step held the gate (err=%v after %s)", err, time.Since(start))
}
start = time.Now()
_, _, _ = r.Run(context.Background(), "/nonexistent/pct", "exec", "9201", "--", "true")
if time.Since(start) > 80*time.Millisecond {
t.Fatal("pct exec must not wait on the gate")
}
}
func TestWritesEtcPVE(t *testing.T) {
for _, c := range []struct {
name string
args []string
want bool
}{
{"pct", []string{"set", "9201", "-mp8", "x"}, true},
{"/usr/sbin/pct", []string{"create", "9201"}, true},
{"pct", []string{"exec", "9201", "--", "true"}, false},
{"pct", []string{"status", "9201"}, false},
{"pvesm", []string{"add", "dir", "x"}, true},
{"pveum", []string{"acl", "modify"}, true},
{"/usr/local/sbin/felhom-pbs-apply", []string{"reconcile"}, true},
{"/usr/local/sbin/felhom-pbs-apply", []string{"read"}, false},
{"/usr/local/sbin/felhom-os-apply", []string{"--plan", "x"}, false},
{"pct", nil, false},
} {
if got := WritesEtcPVE(c.name, c.args); got != c.want {
t.Errorf("WritesEtcPVE(%s %v) = %v, want %v", c.name, c.args, got, c.want)
}
}
}
+104
View File
@@ -0,0 +1,104 @@
// Package pvegate keeps the agent's own writes to /etc/pve out of the way of a Proxmox package step (R-812 option A,
// `09` §3 decision 163, `11` §5.10).
//
// WHY. A `pve` step upgrades pve-cluster / pve-manager / qemu-server / pve-container; their postinst scripts restart
// pmxcfs (the FUSE filesystem behind /etc/pve) and the API daemons. A write that lands while pmxcfs restarts fails or,
// worse, half-lands (design-R-812 §3 A, "can go wrong"). Backups and restore-tests are already kept out by the
// host-wide heavy-op gate; this gate covers everything else the agent writes: every non-GET Proxmox API call
// (proxmox.Client.doBody) and every root CLI that writes /etc/pve (proxmox.ExecRunner — `pct set|create|…`, `pvesm`,
// `pveum`, `felhom-pbs-apply create|reconcile`).
//
// THE RULE. Write waits while a step runs (bounded by its own context). Step marks the step and then waits until every
// write already in flight has finished; it never waits forever (its context bounds it, and the caller gives up and
// does not run the step). One step at a time. Pinned by pvegate_test.go and, at the two chokepoints, by
// proxmox TestPVEGate_*.
package pvegate
import (
"context"
"errors"
"sync"
"time"
)
var (
mu sync.Mutex
inFlight int
stepping bool
stepDone chan struct{}
)
// ErrStepRunning is returned by Step when another step already holds the gate.
var ErrStepRunning = errors.New("pvegate: a Proxmox package step is already running")
// Write marks one /etc/pve write in flight, first waiting while a Proxmox package step runs. The returned release must
// be called when the write has finished. waited reports how long the write was held back.
func Write(ctx context.Context) (release func(), waited time.Duration, err error) {
start := time.Now()
for {
mu.Lock()
if !stepping {
inFlight++
mu.Unlock()
var once sync.Once
return func() {
once.Do(func() {
mu.Lock()
inFlight--
mu.Unlock()
})
}, time.Since(start), nil
}
ch := stepDone
mu.Unlock()
select {
case <-ch:
case <-ctx.Done():
return nil, time.Since(start), ctx.Err()
}
}
}
// Step marks a Proxmox package step and waits until every /etc/pve write already in flight has finished. On error the
// gate is released again and the step must not run. end releases the gate and lets the held-back writes go.
func Step(ctx context.Context) (end func(), err error) {
mu.Lock()
if stepping {
mu.Unlock()
return nil, ErrStepRunning
}
stepping = true
done := make(chan struct{})
stepDone = done
mu.Unlock()
var once sync.Once
end = func() {
once.Do(func() {
mu.Lock()
stepping = false
close(done)
mu.Unlock()
})
}
for {
mu.Lock()
n := inFlight
mu.Unlock()
if n == 0 {
return end, nil
}
select {
case <-ctx.Done():
end()
return nil, ctx.Err()
case <-time.After(50 * time.Millisecond):
}
}
}
// Stepping reports whether a Proxmox package step holds the gate (for logs).
func Stepping() bool {
mu.Lock()
defer mu.Unlock()
return stepping
}
+104
View File
@@ -0,0 +1,104 @@
package pvegate
import (
"context"
"testing"
"time"
)
// A write that starts while a step runs waits until the step ends.
//
// COMPANION RED-PROOF (observed): make Write ignore `stepping` → this fails with "the write went through while the
// Proxmox step held the gate". Restored. (audits/day-2026-10-07/B/red-pvegate.txt)
func TestWrite_WaitsWhileAStepRuns(t *testing.T) {
end, err := Step(context.Background())
if err != nil {
t.Fatal(err)
}
got := make(chan time.Time, 1)
go func() {
rel, _, err := Write(context.Background())
if err == nil {
rel()
}
got <- time.Now()
}()
select {
case <-got:
end()
t.Fatal("the write went through while the Proxmox step held the gate")
case <-time.After(150 * time.Millisecond):
}
ended := time.Now()
end()
select {
case at := <-got:
if at.Before(ended) {
t.Fatal("the write finished before the step ended")
}
case <-time.After(2 * time.Second):
t.Fatal("the write never went through after the step ended")
}
}
// A step waits for a write already in flight before it starts.
func TestStep_WaitsForAWriteInFlight(t *testing.T) {
rel, _, err := Write(context.Background())
if err != nil {
t.Fatal(err)
}
started := make(chan struct{})
go func() {
end, err := Step(context.Background())
if err == nil {
close(started)
end()
}
}()
select {
case <-started:
rel()
t.Fatal("the step started while a write was in flight")
case <-time.After(150 * time.Millisecond):
}
rel()
select {
case <-started:
case <-time.After(2 * time.Second):
t.Fatal("the step never started after the write finished")
}
}
// A step that cannot drain the writes in time gives up and releases the gate (it never waits forever).
func TestStep_GivesUpAndReleases(t *testing.T) {
rel, _, _ := Write(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
if _, err := Step(ctx); err == nil {
t.Fatal("the step must give up while a write is in flight past its deadline")
}
if Stepping() {
t.Fatal("a step that gave up must release the gate")
}
rel()
}
// A write held back past its own deadline returns the context's error.
func TestWrite_HonoursItsContext(t *testing.T) {
end, _ := Step(context.Background())
defer end()
ctx, cancel := context.WithTimeout(context.Background(), 50*time.Millisecond)
defer cancel()
if _, _, err := Write(ctx); err == nil {
t.Fatal("a write held back past its deadline must fail")
}
}
// One step at a time.
func TestStep_OneAtATime(t *testing.T) {
end, _ := Step(context.Background())
defer end()
if _, err := Step(context.Background()); err != ErrStepRunning {
t.Fatalf("a second step must be refused, got %v", err)
}
}
+5 -1
View File
@@ -52,6 +52,10 @@ const (
// (signed, operational key) like agent_update; the root wrapper re-verifies the same signature itself.
ClassOSDockerStep OpClass = "os_docker_step"
// A Proxmox package step on the host (R-812 option A, `11` §5.10) — ring 1. Destructive-class (signed, operational
// key) like os_docker_step; the root wrapper re-verifies the same signature itself.
ClassOSPVEStep OpClass = "os_pve_step"
// The config bundle (agent v0.143.0, R-840, `11` §5.4.2): the box's ROOT-OWNED files (sudoers, wrappers, units).
// Destructive-class (signed, operational key) like agent_update; the root wrapper re-verifies the signature itself.
ClassAgentConfigUpdate OpClass = "agent_config_update"
@@ -123,7 +127,7 @@ func Classify(class OpClass, prov Provenance) Disposition {
return Destructive
case ClassKeyRotation:
return Destructive
case ClassAgentUpdate, ClassOSDockerStep, ClassAgentConfigUpdate:
case ClassAgentUpdate, ClassOSDockerStep, ClassOSPVEStep, ClassAgentConfigUpdate:
// Never benign — no agent-internal provenance can make replacing the agent binary
// unsigned-safe (a compromised process must not be able to self-bless an update).
return Destructive
+64
View File
@@ -0,0 +1,64 @@
package storage
import (
"encoding/json"
"strings"
"testing"
)
// R-330 (disk health Phase 2) — attributes 187, 188 and 199 ride the wire.
//
// The failing drive of 2026-08-14 carried 187 Reported_Uncorrect at raw 1001 while SMART said PASSED;
// none of the three reached the controller. The values are carried as RAW counters; an absent
// attribute is OMITTED from the JSON (unknown), never sent as 0 (S-39).
// The 2026-08-14 shape: PASSED, 187 raw 1001, plus 188/199 and the existing three.
const r330SATA = `{"smart_status":{"passed":true},"ata_smart_attributes":{"table":[
{"id":5,"raw":{"value":0}},
{"id":187,"raw":{"value":1001}},
{"id":188,"raw":{"value":4295032833}},
{"id":197,"raw":{"value":8}},
{"id":198,"raw":{"value":8}},
{"id":199,"raw":{"value":3}}]}}`
// COMPANION RED-PROOF (observed): delete the three R-330 cases in parseSMART → this fails with
// "187 Reported_Uncorrect must be carried (raw 1001); got <nil>". Restored.
func TestParseSMART_R330_CarriesTheThreeCounters(t *testing.T) {
s := parseSMART([]byte(r330SATA))
if s.ReportedUncorrect == nil || *s.ReportedUncorrect != 1001 {
t.Fatalf("187 Reported_Uncorrect must be carried (raw 1001); got %v", s.ReportedUncorrect)
}
// 188's raw value is vendor-packed on some drives (this one is 0x100010001): carried as reported, not truncated.
if s.CommandTimeout == nil || *s.CommandTimeout != 4295032833 {
t.Fatalf("188 Command_Timeout must be carried as the full raw value; got %v", s.CommandTimeout)
}
if s.UDMACRCErrors == nil || *s.UDMACRCErrors != 3 {
t.Fatalf("199 UDMA_CRC_Error_Count must be carried (raw 3); got %v", s.UDMACRCErrors)
}
if s.PendingSectors == nil || *s.PendingSectors != 8 {
t.Fatalf("the existing counters must be unchanged; pending=%v", s.PendingSectors)
}
}
// A drive (or an NVMe device) that does not report the attributes leaves them nil, and the JSON
// OMITS the keys — the receiver reads "unknown", never a measured zero.
//
// COMPANION RED-PROOF (observed): drop `,omitempty` from the three tags in hub.SmartSummary → this
// fails with "an unreported attribute must be omitted, not sent: … reported_uncorrect …". Restored.
func TestParseSMART_R330_AbsentIsOmittedNotZero(t *testing.T) {
for name, raw := range map[string]string{
"sata without the three": `{"smart_status":{"passed":true},"ata_smart_attributes":{"table":[{"id":5,"raw":{"value":0}}]}}`,
"nvme": `{"smart_status":{"passed":true},"nvme_smart_health_information_log":{"critical_warning":0,"media_errors":0,"percentage_used":3}}`,
} {
s := parseSMART([]byte(raw))
if s.ReportedUncorrect != nil || s.CommandTimeout != nil || s.UDMACRCErrors != nil {
t.Fatalf("%s: unreported attributes must stay nil; got %v %v %v", name, s.ReportedUncorrect, s.CommandTimeout, s.UDMACRCErrors)
}
b, _ := json.Marshal(s)
for _, k := range []string{"reported_uncorrect", "command_timeout", "udma_crc_errors"} {
if strings.Contains(string(b), `"`+k+`"`) {
t.Fatalf("%s: an unreported attribute must be omitted, not sent: %s", name, b)
}
}
}
}
+11
View File
@@ -44,6 +44,9 @@ const (
ataReallocatedSectorCt = 5
ataCurrentPending = 197
ataOfflineUncorrect = 198
ataReportedUncorrect = 187 // R-330
ataCommandTimeout = 188 // R-330
ataUDMACRCErrorCount = 199 // R-330
)
// parseSMART maps smartctl JSON to a hub.SmartSummary, handling SATA + NVMe and degrading
@@ -85,6 +88,12 @@ func parseSMART(raw []byte) hub.SmartSummary {
s.PendingSectors = intPtr(int(a.Raw.Value))
case ataOfflineUncorrect:
s.OfflineUncorrectable = intPtr(int(a.Raw.Value))
case ataReportedUncorrect:
s.ReportedUncorrect = int64Ptr(a.Raw.Value)
case ataCommandTimeout:
s.CommandTimeout = int64Ptr(a.Raw.Value)
case ataUDMACRCErrorCount:
s.UDMACRCErrors = int64Ptr(a.Raw.Value)
}
}
}
@@ -142,3 +151,5 @@ func parseThinPoolMetadata(raw []byte) (float64, bool) {
}
func intPtr(v int) *int { return &v }
func int64Ptr(v int64) *int64 { return &v }
+367
View File
@@ -0,0 +1,367 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""test_gate_decoys.py — can this repo's gates be fooled by a LABEL? (R-421, R-426)
The same instrument as `felhom.eu/scripts/test_gate_decoys.py`: a decoy is the LABEL without the
FACT, and a gate that passes on the label alone — or refuses the genuine article — is a live hole.
Every gate is asserted in BOTH directions: the decoy must be convicted, the genuine article passed.
Covered here (the `COVERS` literal is AST-read by `felhom.eu/scripts/decoy_coverage_gate.py`, which
never imports this file):
published check-published-versions.py, against a FAKE Gitea (see below).
release-complete check-release-complete.py, in a scratch clone whose `origin` is a scratch bare
repository, against the same fake Gitea.
reuse-refs, instructions, observations
the three SHARED felhom.eu scripts, run against a scratch clone of THIS repo —
so the decoy is planted in the agent's own REUSE.md / CLAUDE.md / REPORT.md and
coverage is per input, not per script.
NEVER THE REAL GITEA. Both network gates read `GITEA_BASE` from the environment (CI already sets it
to the in-cluster URL); here it points at an `http.server` bound to 127.0.0.1 inside this process,
and every proxy variable is removed from the child's environment so urllib cannot route around it.
A test that asked the real registry would pass or fail on whatever was published that day — the
constant-for-measurement shape — and would reach the network from a hook.
NEVER THE REAL TREE. Every planted file lives in a scratch directory: a workspace that holds a
clone of this repo beside symlinks to the sibling clones the shared scripts reach across to.
Run from the repo root: python3 scripts/test_gate_decoys.py
Exit 0 every decoy judged correctly · 1 a decoy passed or a genuine article was refused.
"""
import http.server
import io
import json
import os
import shutil
import socketserver
import subprocess
import sys
import tempfile
import threading
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
PARENT = os.path.dirname(ROOT)
# DECOY_SHARED_DIR exists for ONE purpose: the red-proof. It lets a mutated COPY of the shared scripts
# be judged without editing the felhom.eu clone. Unset, the suite judges the real shared scripts.
SHARED = os.environ.get("DECOY_SHARED_DIR") or os.path.join(PARENT, "felhom.eu", "scripts")
# ── WHAT THIS FILE COVERS ────────────────────────────────────────────────────────────────────────
# Read by felhom.eu/scripts/decoy_coverage_gate.py, which AST-parses this literal. A gate named here
# MUST have a decoy below that has been seen to fail.
COVERS = {
"published": "against a FAKE Gitea: a tag whose package 404s, a tag whose tree lacks the configs, a package one "
"version past the newest tag (published, never tagged), a patch-gap orphan, a missing package that "
"lexical sorting would drop out of the retention window (0.9.x vs 0.10.x); a tags api answering 500 "
"or a non-JSON 200 is INCONCLUSIVE, never a pass - vs a clean registry, a non-semver tag and a version "
"older than the retention window (not asserted, BY DESIGN) (R-426)",
"release-complete": "the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated commit, a tag "
"with no package; a registry 500 is INCONCLUSIVE - vs the genuine release, an `## Unreleased` "
"heading above it, a newer version named only in prose or under `###`, a tag that only "
"origin has (the shallow-CI shape); a LOCAL-only tag passes BY DESIGN (CI's fresh clone "
"and the published gate's converse probe are what see it) (R-426)",
"reuse-refs": "a cited .go and a cited .md path that do not exist, planted in THIS repo's REUSE.md - vs the "
"real file (R-426)",
"instructions": "a component version literal in THIS repo's CLAUDE.md effective text - vs the same sentence "
"inside an HTML comment (R-426)",
"observations": "R-419 in THIS repo's REPORT.md: an Observations note SAYING it carries no marker - vs the "
"two genuine markers (R-426)",
}
fails = []
ran = 0
def report(name, rc, out, expect_rc, must=()):
global ran
ran += 1
missing = [m for m in must if m not in out]
if rc == expect_rc and not missing:
print(" ok %-62s rc=%d (expected %d)" % (name, rc, expect_rc))
else:
hole = expect_rc != 0 and rc == 0
fails.append("%s: rc=%d expected %d%s; missing %s\n%s" % (
name, rc, expect_rc, " - LIVE HOLE" if hole else "", missing, out[-900:]))
# ── the fake Gitea ───────────────────────────────────────────────────────────────────────────────
class Fake(object):
"""What the fake registry serves. Reset per case."""
def reset(self):
self.tags = [] # tag names, as the tags api lists them
self.packages = set() # versions whose generic package downloads
self.raw = set() # versions whose tag tree serves the probe config
self.tags_status = 200
self.tags_body = None # override bytes for the tags api
self.pkg_status = None # override status for EVERY package request
self.hits = []
FAKE = Fake()
FAKE.reset()
PKG_PREFIX = "/api/packages/admin/generic/felhom-agent/"
RAW_PREFIX = "/admin/felhom-agent/raw/tag/v"
class Handler(http.server.BaseHTTPRequestHandler):
def log_message(self, *a):
pass
def _answer(self, status, body=b""):
self.send_response(status)
self.send_header("Content-Length", str(len(body)))
self.end_headers()
if self.command != "HEAD":
self.wfile.write(body)
def do_GET(self):
p = self.path
FAKE.hits.append(p)
if p.startswith("/api/v1/repos/admin/felhom-agent/tags"):
body = FAKE.tags_body if FAKE.tags_body is not None else \
json.dumps([{"name": t} for t in FAKE.tags]).encode()
return self._answer(FAKE.tags_status, body)
if p.startswith(PKG_PREFIX):
if FAKE.pkg_status is not None:
return self._answer(FAKE.pkg_status)
v = p[len(PKG_PREFIX):].split("/", 1)[0]
return self._answer(200, b"ELF") if v in FAKE.packages else self._answer(404)
if p.startswith(RAW_PREFIX):
v = p[len(RAW_PREFIX):].split("/", 1)[0]
ok = v in FAKE.raw and p.endswith("/configs/felhom-agent.service")
return self._answer(200, b"[Unit]\n") if ok else self._answer(404)
return self._answer(404)
do_HEAD = do_GET
class Server(socketserver.ThreadingMixIn, http.server.HTTPServer):
daemon_threads = True
def child_env(base):
env = {k: v for k, v in os.environ.items() if "proxy" not in k.lower()}
env["GITEA_BASE"] = base
env["NO_PROXY"] = env["no_proxy"] = "127.0.0.1,localhost"
return env
def run(argv, cwd, env=None):
# input="" — a child must never inherit (and block on) this process's stdin
p = subprocess.run(argv, cwd=cwd, env=env, capture_output=True, text=True, input="")
return p.returncode, p.stdout + p.stderr
def sh(argv, cwd):
rc, out = run(argv, cwd)
if rc != 0:
raise SystemExit("setup command failed (%s): %s" % (" ".join(argv), out))
return out.strip()
# ── published ────────────────────────────────────────────────────────────────────────────────────
def published_cases(base):
gate = os.path.join(ROOT, "scripts", "check-published-versions.py")
keep = json.load(io.open(os.path.join(ROOT, "scripts", "retention-policy.json"),
encoding="utf-8"))["generic_versions_kept"]
TAGS = ["v0.150.%d" % i for i in range(3)] # inside any retention window >= 3
VERS = [t[1:] for t in TAGS]
def case(name, setup, expect_rc, must=()):
FAKE.reset()
FAKE.tags = list(TAGS)
FAKE.packages = set(VERS)
FAKE.raw = set(VERS)
setup()
rc, out = run([sys.executable, gate], ROOT, child_env(base))
if not FAKE.hits:
fails.append("published/%s: the gate never asked the fake Gitea - the seam is not wired" % name)
report("published: " + name, rc, out, expect_rc, must)
case("GENUINE: every tag downloadable and serving its configs", lambda: None, 0,
("ALL RELEASED VERSIONS INSTALLABLE",))
case("GENUINE: a non-semver tag is not a release", lambda: FAKE.tags.append("v0.150.2-rc1"), 0,
("ALL RELEASED VERSIONS INSTALLABLE",))
case("FACT: a tag whose package 404s", lambda: FAKE.packages.discard("0.150.1"), 1,
("FAIL v0.150.1", "binary NOT downloadable"))
case("FACT: a tag whose tree does not serve the configs", lambda: FAKE.raw.discard("0.150.2"), 1,
("FAIL v0.150.2", "does not serve"))
case("FACT: published one patch past the newest tag, never tagged",
lambda: FAKE.packages.add("0.150.3"), 1, ("PUBLISHED VERSION(S) WITH NO TAG", "v0.150.3"))
case("FACT: published in a patch GAP between two tags",
lambda: (FAKE.tags.remove("v0.150.1"),), 1, ("v0.150.1 is downloadable", "has no git tag"))
def lexical():
# keep+1 tags: 0.9.0 and 0.10.0..0.10.<keep-1>. By SEMVER the oldest is 0.9.0 (dropped); by
# STRING sort "0.10.0" is the smallest and would be the one dropped - so its missing package
# is convicted only if the window is cut by semver.
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.10.0"}
FAKE.raw = set(t[1:] for t in FAKE.tags)
case("FACT: a missing package lexical sorting would drop (0.10.0 vs 0.9.0)", lexical, 1,
("FAIL v0.10.0",))
def retired():
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.9.0"}
FAKE.raw = set(t[1:] for t in FAKE.tags)
case("BY DESIGN: a version older than the retention window is not asserted", retired, 0,
("NOT ASSERTED", "0.9.0"))
def five_hundred():
FAKE.tags_status = 500
case("INCONCLUSIVE: the tags api answers 500", five_hundred, 2, ("INCONCLUSIVE",))
def html():
FAKE.tags_body = b"<html>sign in</html>"
case("INCONCLUSIVE: the tags api answers a 200 that is not JSON", html, 2, ("INCONCLUSIVE",))
# an unreachable Gitea: a port nothing listens on
s = Server(("127.0.0.1", 0), Handler)
dead = "http://127.0.0.1:%d" % s.server_address[1]
s.server_close()
rc, out = run([sys.executable, gate], ROOT, child_env(dead))
report("published: INCONCLUSIVE: Gitea unreachable", rc, out, 2, ("INCONCLUSIVE", "URLs tried"))
# ── release-complete ─────────────────────────────────────────────────────────────────────────────
def release_cases(base, ws):
bare = os.path.join(ws, "origin.git")
work = os.path.join(ws, "rc-work")
sh(["git", "clone", "-q", "--bare", "--no-tags", "file://" + ROOT, bare], ws)
sh(["git", "clone", "-q", "--no-tags", "file://" + bare, work], ws)
sh(["git", "config", "user.email", "decoy@gate.invalid"], work)
sh(["git", "config", "user.name", "decoy"], work)
# the WORKING-TREE gate, so the file under test is the one being edited, not HEAD's
shutil.copy(os.path.join(ROOT, "scripts", "check-release-complete.py"),
os.path.join(work, "scripts", "check-release-complete.py"))
sh(["git", "add", "scripts/check-release-complete.py"], work)
sh(["git", "commit", "-q", "--allow-empty", "-m", "the gate under test"], work)
base_sha = sh(["git", "rev-parse", "HEAD"], work)
ch = os.path.join(work, "CHANGELOG.md")
original = io.open(ch, encoding="utf-8").read()
V = "9.9.9"
def case(name, top, expect_rc, must=(), tag=None, origin_tag=False, packaged=True, pkg_status=None):
FAKE.reset()
if packaged:
FAKE.packages = {V}
FAKE.pkg_status = pkg_status
try:
io.open(ch, "w", encoding="utf-8").write(top + original)
sh(["git", "commit", "-q", "-am", name], work)
if tag == "head":
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", "HEAD"], work)
elif tag == "unrelated":
empty = sh(["git", "mktree"], work) # stdin is "" — the empty tree, written to this repo
orphan = sh(["git", "commit-tree", "-m", "unrelated", empty], work)
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", orphan], work)
if origin_tag:
sh(["git", "push", "-q", "origin", "HEAD:refs/tags/v" + V], work)
rc, out = run([sys.executable, os.path.join(work, "scripts", "check-release-complete.py")],
work, child_env(base))
report("release-complete: " + name, rc, out, expect_rc, must)
finally:
run(["git", "tag", "-d", "v" + V], work)
run(["git", "push", "-q", "origin", ":refs/tags/v" + V], work)
sh(["git", "reset", "-q", "--hard", base_sha], work)
HEAD = "## v%s — 2026-10-06\n\n- decoy release\n\n" % V
case("GENUINE: tagged at HEAD and published", HEAD, 0,
("newest CHANGELOG version: v9.9.9", "is tagged, placed and published"), tag="head")
case("GENUINE: an `## Unreleased` heading above the release", "## Unreleased\n\n- wip\n\n" + HEAD, 0,
("newest CHANGELOG version: v9.9.9",), tag="head")
case("GENUINE: a newer version named only in prose and under ###",
"The `## v10.0.0` heading is not written yet.\n### v10.0.0 notes\n\n" + HEAD, 0,
("newest CHANGELOG version: v9.9.9",), tag="head")
case("GENUINE: the tag only on origin (the shallow-CI shape)", HEAD, 0,
("exists on origin",), origin_tag=True)
case("BY DESIGN: a LOCAL-only tag passes (CI's fresh clone sees only origin)", HEAD, 0,
("an ancestor of HEAD",), tag="head")
case("FACT: the newest heading has no tag anywhere", HEAD, 1, ("DOES NOT EXIST",))
case("FACT: a tag parked on an unrelated commit", HEAD, 1, ("NOT an ancestor",), tag="unrelated")
case("FACT: tagged, never published", HEAD, 1, ("IS NOT PUBLISHED",), tag="head", packaged=False)
case("INCONCLUSIVE: the registry answers 500", HEAD, 2, ("INCONCLUSIVE",), tag="head", pkg_status=500)
case("FACT beats INCONCLUSIVE: no tag AND the registry answers 500", HEAD, 1, ("DOES NOT EXIST",),
pkg_status=500)
# ── the shared felhom.eu scripts, against THIS repo's inputs ─────────────────────────────────────
def shared_cases(ws):
"""A scratch WORKSPACE: a clone of this repo beside symlinks to the siblings, because the shared
scripts reach across (REUSE.md cites hub paths; instructions_gate reads the workspace CLAUDE.md)."""
for g in ("reuse_refs_check.py", "instructions_gate.py", "observations_gate.py"):
if not os.path.isfile(os.path.join(SHARED, g)):
fails.append("shared gate %s is MISSING beside this clone (tried %s) - a failure, never a skip"
% (g, SHARED))
return
space = os.path.join(ws, "workspace")
os.makedirs(space)
for entry in sorted(os.listdir(PARENT)):
if entry in ("felhom.eu", "felhom-controller", "app-catalog-felhom.eu", "homelab-manifests",
"CLAUDE.md", ".claude-memory"):
os.symlink(os.path.join(PARENT, entry), os.path.join(space, entry))
repo = os.path.join(space, "felhom-agent")
sh(["git", "clone", "-q", "--no-tags", "file://" + ROOT, repo], ws)
# the WORKING-TREE inputs the plants go into, so a case judges today's file
for f in ("REUSE.md", "CLAUDE.md", "REPORT.md"):
shutil.copy(os.path.join(ROOT, f), os.path.join(repo, f))
def case(name, gate, relpath, extra, expect_rc, must=()):
p = os.path.join(repo, relpath)
backup = io.open(p, encoding="utf-8").read()
try:
if extra:
io.open(p, "w", encoding="utf-8").write(backup + extra)
rc, out = run([sys.executable, os.path.join(SHARED, gate), repo], repo)
report(name, rc, out, expect_rc, must)
finally:
io.open(p, "w", encoding="utf-8").write(backup)
case("reuse-refs: GENUINE: this repo's REUSE.md", "reuse_refs_check.py", "REUSE.md", "", 0, ("FAILED 0",))
case("reuse-refs: FACT: a cited .go path that does not exist", "reuse_refs_check.py", "REUSE.md",
u"\n- see `internal/localapi/does_not_exist.go`\n", 1, ("does_not_exist.go",))
case("reuse-refs: FACT: a cited .md path that does not exist", "reuse_refs_check.py", "REUSE.md",
u"\n- see `docs/99-does-not-exist.md`\n", 1, ("99-does-not-exist.md",))
case("instructions: GENUINE: this repo's CLAUDE.md", "instructions_gate.py", "CLAUDE.md", "", 0,
("instructions_gate: OK",))
case("instructions: FACT: a version literal in effective text", "instructions_gate.py", "CLAUDE.md",
u"\nThe agent runs v0.148.0 today.\n", 1, ("v0.148.0",))
case("instructions: GENUINE: the same sentence in an HTML comment", "instructions_gate.py", "CLAUDE.md",
u"\n<!--\nThe agent ran v0.148.0 on 2026-10-06.\n-->\n", 0, ("instructions_gate: OK",))
case("observations: FACT: R-419, prose SAYING it has no marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** It carries no `FILED:` marker and no "
u"`NOT-A-FINDING:` marker, deliberately.\n", 1)
case("observations: GENUINE: a FILED marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** Something broke. **FILED: R-419**\n", 0)
case("observations: GENUINE: a NOT-A-FINDING marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** Odd. **NOT-A-FINDING: my own typo, corrected in "
u"the same minute.**\n", 0)
def main():
srv = Server(("127.0.0.1", 0), Handler)
threading.Thread(target=srv.serve_forever, daemon=True).start()
base = "http://127.0.0.1:%d" % srv.server_address[1]
ws = tempfile.mkdtemp(prefix="agent-decoys-")
print("agent gate decoys — fake Gitea at %s, scratch %s" % (base, ws))
try:
published_cases(base)
release_cases(base, ws)
shared_cases(ws)
finally:
srv.shutdown()
srv.server_close()
shutil.rmtree(ws, ignore_errors=True)
if fails:
print()
for f in fails:
print("FAIL: %s" % f)
return 1
print("\nagent gate decoys OK — %d case(s), every label judged on its fact (R-421)" % ran)
return 0
if __name__ == "__main__":
sys.exit(main())