Compare commits

..

42 Commits

Author SHA1 Message Date
admin 86adbcaf8a rules: "The system poster stays true" (section 6, all five copies identical)
gates / gates (push) Successful in 57s
The shared rule file carries the same wording in felhom.eu, felhom-controller,
felhom-agent, app-catalog-felhom.eu and the workspace root on DooPlex --
"change all five or none" -- so this is this repo's copy of a rule added in
felhom.eu a08bd3cbd5.

A session that changes a fact listed in
architecture/felhom-system-poster.facts.md (a machine, a role, a traffic path,
a backup tier, a time, a retention, a key, a known gap) updates that file in
the same commit; fixes the poster's text if the change is text only; or adds
"System poster needs a refresh: <what changed>" to STATUS. The report names
which of the three it did.

The poster is drawn in Claude Design and nothing in the repo renders it, so
scripts/poster_facts_gate.py only WARNS when the facts outrun the drawing --
it never fails a push, because a refresh needs the operator and another tool.

Verified identical to the other four copies by diff, before and after.
2026-10-09 18:10:19 +02:00
admin 24ea960229 REPORT: v0.154.0 released and delivered (2026-10-09)
gates / gates (push) Successful in 1m7s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-09 08:36:51 +02:00
admin cff6fb5a45 v0.154.0: CHANGELOG release entry (sha, bundle, tag)
gates / gates (push) Successful in 1m11s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-09 06:58:00 +02:00
admin db25b469ba Host report keeps the last backup across a restart; Secure Boot meta-package leaves the Proxmox lane
gates / gates (push) Successful in 1m7s
Found in the 2026-10-09 kernel-night read-back:
- demo-felhom: the kernel step restarted the host 4 minutes after the night
  backup; the in-memory backup list was empty after the restart, so the hub
  alarmed "host tier: newest backup is 48h old". The report now adds R-894's
  saved newest success per tier (one shared instance).
- demo-hp: proxmox-secure-boot-support pulled shim-signed-common into the
  ring-0 Proxmox plan; the step was refused R6 and the kernel step skipped.
  It joins HOST_SLOW_RE with shim and GRUB.

Red-proved: TestCollectBackups_SavedSuccessSurvivesARestart,
TestR894_LastKnownBackupsIsWiredIntoTheDaemon,
test_ring0_pending_pve_leaves_the_secure_boot_meta_out.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-09 06:47:58 +02:00
admin b228b44f0e REPORT/CONTEXT: 2026-10-08 day (R-899, R-304 on main, unreleased)
gates / gates (push) Successful in 52s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-08 08:08:53 +02:00
admin 91b9405c81 R-304: no 'wrong code' when earlier sealed packages were not all checked (424 older_unchecked)
gates / gates (push) Successful in 55s
Unreleased; ships with tomorrow's release.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-08 07:53:01 +02:00
admin 4c69c25b48 R-899: no OS leg after a household press (trigger=manual); after-boot kernel reports carry the saved ring
gates / gates (push) Successful in 50s
Unreleased; ships with tomorrow's release.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-08 07:34:19 +02:00
admin c9013bb47d CHANGELOG/REPORT: v0.153.0 released and delivered
gates / gates (push) Successful in 1m33s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 19:06:04 +02:00
admin 2d1e5d0774 R-898: ring 0 stages exactly the told kernel (select listed, KernelSet(kver))
gates / gates (push) Successful in 1m6s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 18:45:28 +02:00
admin f09c53efc1 REPORT/CONTEXT: v0.152.0 the kernel lane
gates / gates (push) Successful in 1m9s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 18:20:11 +02:00
admin 92d647a6f6 CHANGELOG: v0.152.0 released (shas, step bundle); host_stale is 45 min (comment fix)
gates / gates (push) Successful in 1m0s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 15:42:26 +02:00
admin d03ab7f1f5 R-836: the kernel lane — one-shot boot through the ESP flag, boot good / one self-revert, night step on a told night
gates / gates (push) Successful in 44s
Wrapper layer kernel (stage / reboot / boot / good / revert / cancel / status;
R20-R23), the two GRUB generators in the bundle (option C on the one-shot
entry), the agent's night step and after-boot judge (host health rule + hub
reached, 20 min measured), the signed os_kernel_step (stage only).
Red-proofs: felhom.eu audits/kernel-lane-2026-10-07/A/redproof.txt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 15:18:24 +02:00
admin b17d1c597d REPORT/CHANGELOG: 2026-10-07 day
gates / gates (push) Successful in 59s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 14:21:01 +02:00
admin 0ae01dfb52 CHANGELOG: v0.151.0 released (R-861, R-812 A, R-366, R-105)
gates / gates (push) Successful in 58s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 13:14:47 +02:00
admin dd7cdc09e7 R-366 slice 2 (decision 168): the host report carries the archives the restore-test skipped as another key's
gates / gates (push) Successful in 1m6s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin 7b0a8b234b R-105 (decision 169): remove the escrow-create -directive flag and the upload's directive field
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin ce1a4b4758 R-812 option A: the Proxmox package lane (layer pve) + the /etc/pve write gate
The wrapper gains layer "pve" (slow lane): the host's Proxmox userspace
packages only — origin "Proxmox Debian Repository", never a kernel / boot /
firmware / microcode name (R14), no removal, no undo, a new package only from
an allow-list; authority = a signed os_pve_step or the root-owned ring-0 mark.
The night leg runs it in ring 0 after a healthy host step; ring 1 only by a
signed job (PVEStepExecutor). While it runs, the agent's own /etc/pve writes
(every non-GET API call, pct config verbs, pvesm, pveum, felhom-pbs-apply)
wait on internal/pvegate. Health = the host rule + unchanged container ids +
pveversion reads the installed pve-manager.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:38 +02:00
admin ac90169a5d R-861 (a) A1 + (b) B2: the image ref goes to a root verb that checks it; the agent's in-guest tee grant is gone; felhom-op's pct lines are exact (09 §3 decision 165)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 10:19:28 +02:00
admin 154d6dcaa9 Shared rule file: no hub image build or deploy in a session the operator does not attend (09 §3 decision 162)
gates / gates (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 09:55:02 +02:00
admin 2f7072050f REPORT: released and delivered (2026-10-07)
gates / gates (push) Successful in 1m4s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 09:31:58 +02:00
admin 85e799f360 CHANGELOG: v0.150.0 released (R-528, R-894, R-330)
gates / gates (push) Successful in 46s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 08:59:08 +02:00
admin 3a72a4811b Shared rule file: rule 11 — every helper prompt carries the brief's fences in full (09 §3 decision 160)
gates / gates (push) Successful in 54s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 08:52:36 +02:00
admin a165d53c88 REPORT: the second burn-down night
gates / gates (push) Successful in 49s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 00:14:50 +02:00
admin adaf86ad57 R-330: the agent sends SMART 187/188/199 (raw; omitted when unknown)
gates / gates (push) Successful in 53s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 22:13:04 +02:00
admin 74b5eae5b0 R-894: after a restart the agent remembers the last backup per tier
gates / gates (push) Successful in 35s
An unreadable storage right after an agent restart read the off-site tier
DUE (the in-memory record was empty). The newest success per tier is now
kept on disk and read ONLY when the storage cannot be read: fresh -> not
due, older than the cadence -> due, none -> due (unknown) as before. A
storage that answers stays the ground truth.

Ships with v0.150.0 after the 2026-10-07 read-back; nothing delivered.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 20:33:27 +02:00
admin 7e82f325b8 CHANGELOG: the memory-kill check (unreleased, to ship as v0.150.0 after the night read-back)
gates / gates (push) Successful in 50s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:36:09 +02:00
admin acccb66bd3 R-528 (09 decision 157): after a Docker engine step the wrapper proves the engine reports a memory kill
felhom-os-apply: a docker-layer apply runs oom_check() after health_after and reports
"oom_check": {result pass|fail|error, oom_killed, oom_event, exit_code, image, detail}.
One throwaway container (the controller's image, --pull never, --network none, 64m cap,
label felhom.oomcheck=1) runs dd bs=200M; pass only with OOMKilled=true AND the oom event
(read after a 2 s settle, --until = guest epoch + 1: measured on demo-hp, an --until taken
right after the run missed the event). docker rm -f always runs in a finally; every call
is bounded (<= 90 s). It never changes the step's outcome or health. New wrapper-only mode
"oom-check" (docker layer) runs the check alone; check_guest etc. still apply.

Agent: WrapperReport/Report gain OOMCheck (json:"oom_check"), copied unchanged in runLayer
and in the kept-copy path.

Tests: 9 wrapper tests + 2 Go tests, each red-proofed (audits/readback-2026-10-07/F/red-*.txt).
Also: test_felhom_os_apply.py's `if __name__` sat mid-file, so 11 tests (UnsentReport,
SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran as a script or from
TestWrapperSuite; moved to the end (they pass).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:34:55 +02:00
admin de812bc027 The shared rule file (09 §3 decision 152), identical to the other copies; no code change
gates / gates (push) Successful in 44s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 16:06:46 +02:00
admin 3e8ebeb96c Instruction files kept true (09 §3 decision 150): stale gate lists, paths and facts corrected; no code change
gates / gates (push) Successful in 46s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 13:43:57 +02:00
admin cefdc731a4 REPORT: the operator's ten answers (2026-10-06)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 12:28:58 +02:00
admin e56dcb8a4c CHANGELOG/v0.149.0 released (shas)
gates / gates (push) Successful in 37s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:50:21 +02:00
admin f277e619e2 CHANGELOG: unreleased — R-856 GET /host/crash-guard
gates / gates (push) Successful in 38s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:40 +02:00
admin 386f51edc6 R-856: GET /host/crash-guard — the host crash guard's last-boot record for the controller (09 decision 143)
The controller waits ~15 minutes with app mails after a crash boot of the host; it learns of the
crash boot from this route. Reads /var/lib/felhom-crash-guard/state.json (read-only, no Proxmox call)
and passes present/last_boot_at/last_boot_unclean/tripped through; a missing, unreadable or garbled
file answers 200 present:false. Guest-token authed like every sibling route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:01 +02:00
admin 130e3ed882 CHANGELOG: unreleased — R-444 weekly guest disk trim, R-99 runbook pointer
gates / gates (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:25:22 +02:00
admin be398f92e8 R-99: the PBS phantom WARN names the cleanup runbook (09 §3 decision 140)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin ee71abd1d4 R-444: weekly guest disk trim (pct fstrim) outside the night, under the heavy-op gate
Operator ruling 09 §3 decision 139. One exact sudoers rule FELHOM_FSTRIM
(`/usr/sbin/pct ^fstrim [0-9]+$`) + manifest entry guest-fstrim; new
internal/fstrim job: due Wednesday from 10:00 host-local, starts only
10:00-20:59, holds backup.InFlight (busy -> deferred to the next hourly
tick), failed trim retried at most 3x per week, bytes parsed from the
"(N bytes) trimmed" lines, last result per guest persisted in
<state_dir>/guest-disk-trim.json and reported as guest_disk_trim.
Config opt-out: "disk_trim": {"disable": true}.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin 37e98f452b REPORT: the burn-down night (2026-10-06)
gates / gates (push) Successful in 45s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 02:13:10 +02:00
admin b2b82ae828 bundle test: the ISO first-boot files the installer names under KEPT are not installer-written (go test red since felhom.eu 85de3f9b); CHANGELOG unreleased (R-426 decoys)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:25 +02:00
admin b78a0ff3ac R-426: decoys for the shared reuse-refs, instructions and observations gates
COVERS "reuse-refs", "instructions", "observations": the three shared
felhom.eu scripts run against a scratch clone of THIS repo (in a scratch
workspace symlinking the sibling clones they reach across to), so the
plant is in the agent's own REUSE.md / CLAUDE.md / REPORT.md. Convicted:
a missing cited .go and .md path, a version literal in CLAUDE.md's
effective text, R-419's prose-only Observations note. Passed: the real
files, the version inside an HTML comment, both genuine markers.
DECOY_SHARED_DIR lets a red-proof judge a mutated copy of the shared
scripts without editing the felhom.eu clone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 96047453cb R-426: decoys for the release-complete gate
COVERS "release-complete": the working-tree gate runs in a scratch clone
whose origin is a scratch bare repo, against the fake Gitea. Convicted:
the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated
commit, a tag with no package, and no-tag wins over a registry 500.
Inconclusive: a registry 500. Passed: the genuine release, an
`## Unreleased` heading above it, a newer version named only in prose or
under `###` (the withdrawn sweep decoy, now asserted the right way round),
a tag only origin has (the shallow-CI shape), and a LOCAL-only tag BY
DESIGN (CI's fresh clone and the published gate's converse probe see it).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 4bf5db6875 R-426: decoy suite for the published gate, against a fake Gitea
scripts/test_gate_decoys.py (new; COVERS "published"): an http.server on
127.0.0.1 stands in for Gitea through the gate's existing GITEA_BASE
seam, proxies stripped, so no case reaches the real registry. Facts
convicted: a tag whose package 404s, a tag tree without the configs, a
package one patch past the newest tag (never tagged), a patch-gap orphan,
a missing package that lexical sorting would drop out of the retention
window. Inconclusive, never a pass: tags api 500, a non-JSON 200, Gitea
unreachable. Passed: a clean registry, a non-semver tag, a version older
than the retention window (BY DESIGN).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 87977ff40a CHANGELOG/v0.148.0 released (shas)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 00:37:07 +02:00
71 changed files with 6667 additions and 250 deletions
+3 -2
View File
@@ -21,6 +21,7 @@ fast, and wrong.
This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that
felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before
path-scoped rules existed. The single source is now felhom.eu/CLAUDE.md "Code quality rules"; this path-scoped rules existed. Deliberate scoped copies now live in felhom.eu/.claude/rules/hub.md and
file is the scoped copy that loads exactly where health checks are written. (2026-08-06) felhom-controller/.claude/rules/gates.md (hub.md's comment names them); none is the single source. This
file is the copy that loads exactly where agent health checks are written. (2026-08-06; corrected 2026-10-06)
--> -->
+91
View File
@@ -0,0 +1,91 @@
---
unconditional: true
---
# Unprompted work — rules for any session without a task file
> Goal sessions, nightly sessions, "work the register" sessions. **A session that starts from
> `/goal` or a standing brief inherits these rules exactly as it inherits the gates.** They are the
> part of `PROMPT-TEMPLATE.md` that a task file used to carry and a goal does not. Same wording lives
> in `felhom.eu`, `felhom-controller`, `felhom-agent` and `app-catalog-felhom.eu` `.claude/rules/`, and in the workspace
> root's unversioned `.claude/rules/`; change all five or none.
## 1. What you may pick up on your own
- A register row **you or another CC session filed**, with owner CC, at P3 or a bounded P2, that
needs **no operator decision**, touches **no customer data by design**, and introduces **no
mechanism nobody has measured**. Smallest first.
- A defect you find while exercising the product, filed as a row **before** you fix it — **unless it is small**:
a small finding is fixed in the session and never filed (the size rule, `OPEN-ITEMS.md` „How a row is filed").
- Hygiene: register compression, stale citations, rows with no owner, documents that contradict
live source.
**Not yours, ever, without a task file or an operator word:** money; anything that changes risk to
customer data; anything that changes a promise the product makes to a customer; anything that
reverses a documented design decision (`documentation/architecture/` — a design decision is not a
defect, R-370); anything on DooPlex or ep0; baking or vouching a golden; promoting a
catalog version; a new external dependency; **a hub image build or hub deploy in a session the operator does not
attend** (operator ruling 2026-10-07, `09` §3 decision 162).
## 2. When you may decide instead of ask (operator grant, 2026-09-14)
You may take a decision yourself when **all** of these hold: the architecture folder and the register
give a clear direction; your choice follows that direction; it is reversible without customer-data
risk; and you can write it in the `09-update-architecture.md` §3 shape — one answerable sentence, the
options, what each costs, why this one. **Then record it** as a dated decision in `CONTEXT.md` and
the owning architecture document, tagged *decided by CC unattended — operator may reverse*, and put
it **first** in the morning note. A decision you cannot write in that shape is one you do not take.
## 3. The discipline a task file used to carry
1. **Baselines first.** Read each repo's `main` hash and version from live source before touching it.
2. **Read the architecture document for the area, and name it** in the report, before any claim.
3. **Red-proof every correctness fix.** A test never seen failing has not been shown to test anything.
4. **Live-validate on a Tier-0 box** through the endpoints the UI invokes. `demo-hp` is `ssh hp`.
Throwaway apps only; the standing apps and `bentopdf` stay.
5. **Evidence off the machine at the end of each phase**, before any revert (R-320).
6. **One release per repo per session**, with a CHANGELOG entry (controller: with its `MinAgent`
line), REPORT overwritten, floor raised to deliver it. **No golden unless a drill or fresh install
needs one** (the waiver, R-468). **No `--no-verify`.**
7. **An enumerated gap becomes a row in the same session — or, if it is small, is fixed in it** (the size rule).
Prose is not a record.
8. **Hungarian text is searched with ASCII fragments**, with a positive and a negative control.
9. **Never leave a half-state.** If time runs out, revert to clean and say what was reverted.
10. **Teardown, three layers, stated** — machine, host, hub — or "provisioned nothing".
11. **Every helper prompt carries the brief's fences in full** (operator ruling 2026-10-07). A helper session (a
subagent, a fork, a workflow agent) gets the brief's fence list word for word — every protected machine, every
„no", every delivery and Docker limit — not a summary and not „the usual fences". Earned on 2026-10-06 night: two
helpers whose prompts carried only part of the fences ran `docker volume prune` on the bench and a Docker-using
gate on DooPlex.
## 4. The morning note
One screen, plain language, in this order: **decisions you took** (§2) first; what you exercised;
what broke and whether you fixed it; rows opened and closed with the register size before and after;
what needs the operator, each with what happens if they do nothing. No file paths, no function
names, no row numbers as the subject of a sentence.
## 5. Instruction files
**Instruction files (`CLAUDE.md`, `.claude/rules/*`) are kept true by the session that finds them wrong**
(operator ruling 2026-10-06, `09` §3 decision 150). A session MAY, without asking: correct a stale fact (a command, a
count, a version, a path, a description of what a gate does), add a fact it proved, and remove a reference to something
that no longer exists. Each edit is named in the report (file, line, before, after, why). A session MAY NOT, without the
operator's word: loosen a safety rule, a fence, a „never", a protected machine, a secret rule, or a review step; or
remove a rule. When in doubt, it is a rule change, and it goes to the operator. If Claude Code's own permission check
asks before such an edit, wait for the operator's click; if it refuses, record that and file the exact line.
## 6. The system poster stays true
**A session that changes a fact listed in `architecture/felhom-system-poster.facts.md` (a machine, a
role, a traffic path, a backup tier, a time, a retention, a key, a known gap) updates that file in
the same commit.** If the change is text only, it also edits the matching text in
`felhom-system-poster.html`. If it needs a new drawing, it adds the line
**"System poster needs a refresh: <what changed>"** to `STATUS.md`'s "waiting on the operator" list.
**The session report names which of the three it did.**
Why the three-way split: the poster is a drawing made in Claude Design, and **nothing in this
repository renders it** — unlike `where-felhom-stands.html`, which has `render_stands.py`. So the
facts file is the source of truth and the drawing trails it. `scripts/poster_facts_gate.py` warns
when the facts file has a newer commit than the poster; it **never fails a push**, because a refresh
needs the operator and another tool, and a gate nobody can clear is a gate people learn to route
around.
+208
View File
@@ -1,5 +1,213 @@
## v0.154.0 — no OS leg after a household press; the after-boot kernel report carries the real ring (R-899); no „wrong code" when older packages were not checked (R-304); the last backup survives a restart in the host report; the Secure Boot meta-package leaves the Proxmox lane (2026-10-09)
Released by `scripts/release-agent.sh`: binary sha256 `36a54ad20e896cd61d7b9944e2a521f4209c2b76940d8f1827af05b645463226`, bundle
`93487989f50d4a3039a05cec0f91ec478d62489a414705e49403e9b1340d20e0` (tag `v0.154.0` = `db25b46`).
**Delivery: the agent binary AND the config bundle** — the wrapper `felhom-os-apply` changed (2026-10-09).
- **The hub's backup evidence survives a restart (2026-10-09, found in the kernel-night read-back).** The host report's
`backups` list came only from memory, which is empty after an agent restart. On demo-felhom the night backup landed
02:40 UTC and the kernel step restarted the host at 02:44 — no report fell in between, two nights running — so the hub
alarmed „host tier: newest backup is 48h old" at 03:00 after a good backup. The report now adds R-894's saved newest
success per tier and guest (`backup-success-state.json`, ONE instance shared with the local API) where memory holds no
success at or after it. Only successes are saved, so a failure is never hidden. `hub.KnownBackupReporter`,
`BackupSuccessState.KnownBackupSuccesses`; tests `TestCollectBackups_SavedSuccessSurvivesARestart` (red-proved: the
restart case reported nothing), `TestBackupSuccessState_KnownBackupSuccessesSurviveAReopen`, and
`TestR894_LastKnownBackupsIsWiredIntoTheDaemon` now also asserts the collector gets the same instance (red-proved).
- **`proxmox-secure-boot-support` is a boot-chain package (2026-10-09).** On demo-hp (Secure Boot on) its upgrade pulled
`shim-signed-common`; in the ring-0 Proxmox plan that refused the whole step (R6) and skipped the night's kernel step,
after the household had been told „tonight". It joins `HOST_SLOW_RE` with shim and GRUB (they stay out of every lane,
as before). `test_ring0_pending_pve_leaves_the_secure_boot_meta_out` (red-proved).
- **R-899 (operator ruling 2026-10-08, option A):** `POST /backup?trigger=manual` (a „Mentés most" press, sent by the
controller from its next release) runs no OS leg after it: the leg belongs to the night, after the night's own copy.
Before, a press ran the leg at once (in the day; on 2026-10-07 08:51 demo-hp skipped it only by the 20-hour rule). A
request without the parameter (an older controller, the scheduled path) behaves as before.
`internal/localapi/server.go`; `TestAfterPrimaryBackup` gained two sub-cases (red-proved: ignoring `trigger`, the leg
ran once after a press).
- **R-304 — „wrong code" only when every earlier package was tried.** `POST /escrow/recover-offsite-password` answered
400 („the recovery code did not open the sealed bundle") whenever the current package and the TRIED earlier packages
refused the code — even when the hub withheld earlier packages (rows with no key material, rows over its serve cap), a
served package was malformed, the 6-attempt cap stopped the loop, or the retained list could not be read at all. Now
those cases answer **424** with `older_unchecked` (the count, -1 = unknown) and a sentence that does not call the code
wrong. `escrow.ErrRetainedUnchecked` / `RetainedUncheckedError`; the retained fetcher's second value is now every
withheld package (unopenable + truncated + malformed). Tests `TestR304_*` (real age crypto; red-proved: without the
check, four cases returned the wrong-code error) and a 424 row in `TestRecoverOffsitePassword_EachSituationGetsItsOwnStatus`.
An older controller maps the unknown 424 to its neutral „we do not know why" sentence.
- **The „ring 1" label after a boot:** in the first second after a reboot the agent has not fetched the hub's block, and
the after-boot kernel reports (`judging`, `good`, `revert`, `fell_back`…) said ring 1 on a ring-0 box (demo-felhom,
2026-10-08 night). Now they read the fetched block, else the block the daemon saved on disk before the reboot (R-866),
else ring 1 as before. A label only — the hub's approval reads its own ring list. `kernelReportRing`,
`TestKernelReportRing_BeforeFirstFetch` (red-proved: ring 1, want 0).
## v0.153.0 — ring 0 stages exactly the told kernel (R-898; `09` §3 decision 176) (2026-10-07)
Released by `scripts/release-agent.sh`: binary sha256 `b204ebe6d65944f43b70dcfa4ae0dfb90da38c92e2ba6a38a1826e488eeb6630`, bundle
`6db216275eb0dd39188d93481a2045998a69e8cb7838ad908a58d65d9a6a9b57` (tag `v0.153.0` = `2d1e5d0`). Delivered (binary only — the
bundle files are unchanged from 0.152.0) to demo-hp, demo-felhom, Tester 1 on 2026-10-07 19:03.
**Delivery: the agent binary only** — no root file changed (the wrapper is unchanged; its tests gained two cases).
- `internal/osupdate/kernel.go`: ring 0's night kernel step stages EXACTLY the kernel the household was told about
(select `listed`, `KernelSet(kver)` = the series meta-package and the signed image at the kernel's own version) instead
of "whatever is pending tonight". Seen 2026-10-07: demo-felhom was told about 7.0.14-20 while its sources offered
7.0.14-22 by night — the old code staged `pending-kernel` and the wrapper refused it (R23), losing the night. A told
version that is no longer installable is refused by the wrapper before any change (R7) and the hub tells the household
again for the newer kernel (hub v0.143.1). Tests `TestKernel_Ring0ToldNightStagesThenReboots` (red-proved against the
old select), `TestKernelSet`; wrapper `test_ring0_listed_installs_the_told_kernel_not_the_newest`,
`test_ring0_told_kernel_gone_is_refused_before_any_change`.
## v0.152.0 — the kernel lane (R-836; `09` §3 decisions 164, 172; `11` §5.11) (2026-10-07)
Released by `scripts/release-agent.sh`: binary sha256 `95ff42208e36ba49b6e2b09a97e81a6fa11562ecc8042f1ed18d378b8b1f88b1`,
config bundle sha256 `f0c2cec374b711b3c131c955012b33fd0ad495337d049b7a743c3eef9e85c20b` (tag `v0.152.0` = `d03ab7f`).
Step bundle `0.152.0-step1` sha256 `0b71d32b054cf3b7ade0234ffcbb0df159901f542cde540adaee411db466f48e` (the 0.151.0 bundle
with only `felhom-os-apply` replaced; `scripts/build-step-bundle.py`), published as package version `0.152.0-step1`.
**Delivery: agent binary, then the STEP bundle `0.152.0-step1`, then the bundle `0.152.0`** — the bundle ADDS two paths
(the GRUB generators), and an installed `felhom-os-apply` refuses a path its own table lacks (R16, R-880).
A new kernel boots ONCE; if it crashes the box comes back on the old kernel by itself; it becomes the default only after
a healthy boot; a booted-but-unhealthy kernel is reverted ONCE by the agent with no person. Built on the spike's
candidate 2 (`audits/kernel-spike-2026-10-07/`), with option C on the one-shot entry.
- `configs/felhom-grub-oneshot.sh` → `/etc/grub.d/01_felhom_oneshot` (bundle): reads `felhom_next` from a GRUB env block
on the ESP (`EFI/felhom/oneshot.env`), clears and saves it BEFORE the menu, and sets the default to that kernel's
one-shot entry only when the name is an installed kernel. No vfat ESP → prints nothing.
- `configs/felhom-grub-oneshot-entries.sh` → `/etc/grub.d/42_felhom_oneshot` (bundle): one entry per installed kernel,
id `felhom-oneshot-<ver>`, the normal entry plus `softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10`
(option C). Sorted after `10_linux`: never entry 0, never the default.
- `configs/felhom-os-apply`: layer `kernel` (lane slow; an appliance; authority = a signed `os_kernel_step` or the
root-owned ring-0 mark). Modes: `apply` STAGES (select `pending-kernel` or a signed `listed` set; `expect_kver` = the
kernel the household was told about): pins the GRUB default to the RUNNING kernel in
`/etc/default/grub.d/zz-felhom-kernel-default.cfg` and proves it from grub.cfg, installs, proves the default did not
move and the one-shot entry exists, writes the flag; never reboots. `kernel-reboot` (a staged step only),
`kernel-boot` (judging | fell_back | self_reverted | revert_failed), `kernel-good` (the new kernel becomes the
default, proved), `kernel-revert` (ONE per step; refused when the default is not the old kernel), `kernel-cancel`,
`kernel-status`. New refusals: R20 (the box cannot do a one-shot: not UEFI, no vfat ESP, a separate /boot, GRUB
without fat/loadenv, the generators missing, a hand pin), R21 (the crash guard tripped or an unclean boot in its
window), R22 (the phase does not allow the mode; never two steps within 20 h), R23 (not exactly one newer kernel, or
not the one signed / told). State `/var/lib/felhom-kernel/state.json`. Facts carry `kernel_lane`; the next-boot
kernel reads the flag and the grub.cfg default. Tests: `KernelLane` (27), red-proof
`felhom.eu/documentation/audits/kernel-lane-2026-10-07/A/redproof.txt`.
- `configs/felhom-crash-guard` unchanged; `KernelStepCannotLeaveTheBoxOff` (3 tests) pins that a step's planned reboot,
one crash and one self-revert add ONE unclean boot (a panic before userspace adds none), so the box cannot stay off.
- `internal/osupdate/kernel.go`: the night leg ends with the kernel step — after a healthy host step (and a healthy
Proxmox step when one ran), trigger `night` only, on a night the hub's `os_update.kernel` block marks `tonight` (the
household was mailed the day before — no mail, no step). Ring 0 stages + reboots; ring 1 reboots only a kernel a
signed `os_kernel_step` staged (`KernelStepExecutor`: stage only, under the heavy-op gate). The hub hears `staged`
BEFORE the reboot. At every start `KernelAfterBoot`: on the new kernel it JUDGES the boot — `KernelVerdict` = the
host health rule (`11` §8.2) AND the box reached the hub (the `judging` report itself) — for 20 minutes (measured:
everything healthy 68 s after the reboot on demo-felhom, 272 s on demo-hp; under the hub's 45-minute `host_stale` (`alerting.stale_threshold`)).
Healthy → `kernel-good`, outcome `applied`; not healthy → outcome `health_failed`, then ONE `kernel-revert`.
Tests: `TestKernel*` (13); red-proofs in the same file.
- `internal/hub`: `WireOSUpdate.Kernel` {kver, tonight, notified_at}. `internal/reconcile`: `os_kernel_step` is
destructive-class. `cmd/felhom-opsign`: the op is listed.
## v0.151.0 — the agent can no longer hand the guest any image; the Proxmox package lane; the other-key archives reported; the DR directive retired (R-861, R-812 A, R-366, R-105; `09` §3 163, 165, 168, 169) (2026-10-07)
Released by `scripts/release-agent.sh`: binary sha256 `0464354f2cdf452a7c5d2a74d9191fe91415fcfa244154480d26d5b30e10b194`
config bundle sha256 `bacd1d175a9392bc1755341d01a106abbd72aba48ab575e7a32dd819d8f5da4c` (tag `v0.151.0` = `dd7cdc0`).
**Deliver the binary FIRST, then the bundle:** the bundle's sudoers removes the `tee` grant the 0.150.0 binary still uses.
No path added (26 → 26), so no step bundle.
### Part of v0.151.0 — the agent can no longer hand the guest any image; felhom-op's pct lines are exact (R-861 (a) A1, (b) B2; `09` §3 decision 165)
**Delivery order: agent binary FIRST, then the config bundle.** The new sudoers drops the agent's in-guest `tee`
grant; an older binary still calls `tee`, so a bundle that lands before the binary would stop managed controller
updates (and the old binary's capability probe would read `controllerswap-write` degraded). No bundle path is added
(`felhom-priv-apply` and `/etc/sudoers.d/felhom-op` are already bundle files), so no step bundle.
- `configs/felhom-priv-apply`: new verb `controller-image <vmid>` — reads the ref on stdin (≤ 256 bytes, ASCII, one
optional trailing newline), requires `^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$` (the agent's
own `controllerImageRe`), then runs `pct exec <vmid> -- tee /etc/felhom-controller-image` AS ROOT; refusal rule `I1`
(rc 3), a bad vmid `A1` (rc 2); listed in `--self-check`.
- `configs/felhom-agent.sudoers` `FELHOM_CONTROLLERSWAP`: `pct ^exec [0-9]+ -- tee /etc/felhom-controller-image$`
REMOVED; `/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$` added.
- `internal/localapi`: `GuestExecutor.GuestExecStdin` replaced by `WriteControllerImage`; `GuestBinder.WriteControllerImage`
pipes `ref\n` to the verb through the fenced runner; the swap's `writeImage` calls it. Capability `controllerswap-write`
now probes the verb.
- `configs/felhom-op.sudoers` (B2, hygiene): `pct start|stop|unlock [0-9]*` → `pct ^start [0-9]+$` etc. (the glob's `*`
matched spaces: `pct stop 9201 --skiplock 1` passed).
- Tests: `ControllerImage` (5, `configs/test_felhom_priv_apply.py`), `TestSudoersRefusesTheR861Injections` (+3 lines),
`TestSudoersAllowsTheControllerImageVerb`, `TestFelhomOpSudoersPctIsExact`, `TestR861_WriteControllerImageUsesTheRootVerb`,
`TestControllerSwap_WriteViaRootVerb_NoShell`. Red-proofs: `felhom.eu/documentation/audits/day-2026-10-07/C/`.
- `README.md`: the controller-swap paragraph described the removed `tee` path — corrected.
### Part of v0.151.0 — the Proxmox package lane (R-812 option A, `09` §3 decision 163)
**MinAgent impact: none** (a new layer; an older hub ignores the pve report). **The bundle carries the new
`felhom-os-apply` — deliver it with the binary** (signed `agent_update`, then signed `agent_config_update`).
- `configs/felhom-os-apply`: new layer `pve`, lane `slow` only — the host's Proxmox USERSPACE packages: origin `Proxmox Debian Repository` only (R2), never a kernel / boot / firmware / microcode name (R14, `HOST_SLOW_RE` — the kernel is R-836's lane), no removal (R4), no undo (R5), a new package only from `PVE_NEW_ALLOW` (`proxmox-firewall-data`, measured on demo-felhom; R6 otherwise), an appliance only (R12), authority = a signed `os_pve_step` or the root-owned ring-0 mark (R3). Select `pending-pve` (ring 0): installed Proxmox-origin packages with a pending upgrade. The report carries `pve_manager` (pveversion after the step).
- `internal/pvegate` (new): the agent's own writes to /etc/pve wait while a pve step runs (pmxcfs restarts); the step waits for writes in flight (bounded, 2 min — then it fails and does not run). Wired at `proxmox.Client.doBody` (every non-GET) and `ExecRunner.RunStdin` (`WritesEtcPVE`: pct config verbs, pvesm, pveum, felhom-pbs-apply create/reconcile).
- `internal/osupdate`: `LayerPVE`; the night leg runs the pve step in ring 0 after a healthy host step (an appliance; ring 1 never in the night leg); `PVEHealthVerdict` = the host rule + every running container keeps its id + pveversion reads the installed pve-manager; the pve report carries Proxmox userspace only (the hub's candidate set). `PVEStepExecutor` (signed `os_pve_step`, ring 1, under the heavy-op gate and the /etc/pve gate); `reconcile.ClassOSPVEStep` (destructive-class); `felhom-opsign -op os_pve_step` (params by `-params`).
- Tests: wrapper `PVELane` (17; red first — the `pve-manager` plan was refused R12 on the old code), `pvegate` (5), `TestPVEGate_*` + `TestWritesEtcPVE`, `TestPVE_*`, `TestPVEHealthVerdict`, `TestPVEStepExecutor_*`. Red-proofs: `felhom.eu/documentation/audits/day-2026-10-07/B/`.
### Part of v0.151.0
- R-366 slice 2 (`09` §3 decision 168): the restore-test pick records, per tier, the archives it skipped as written with another key (count, oldest, newest — no key material) in a `ForeignKeyLedger`; the host report carries it as `foreign_key_archives.tiers` (the stanza absent until a tier was evaluated since start, `tiers: []` when none — no null on the wire, the report contract forbids it). The hub turns a change into one operator line. Tests `TestR366_PickRecordsArchivesWrittenWithAnotherKey`, `TestR366_EvaluatedWithNoneIsAnEmptyList` (red-proved, `felhom.eu/documentation/audits/day-2026-10-07/E/`).
- R-105 option A (`09` §3 decision 169): the `--selftest=escrow-create -directive <file>` flag and the escrow upload's `directive` field are removed — nothing read the directive; the DR path reads the recipe, tenantsync and the escrow blob. The hub ignores a `directive` from an older agent.
## v0.150.0 — the Docker step proves the engine reports a memory kill; after a restart the agent remembers the last backup per tier; three more SMART counters on the wire (R-528, R-894, R-330; `09` §3 decisions 157, 161) (2026-10-07)
Released by `scripts/release-agent.sh`: binary sha256 `a23d1c9085bc7fd4fc48fe0327f6504aa83e6331510fb4a3d23e042dddb26f9c`
config bundle sha256 `88456b386d9b1027bd22861cac8c23df004bf9fd9f67644d6595bfca8c94498e` (tag `v0.150.0` = `3a72a48`).
**The bundle carries the new `felhom-os-apply` (the memory-kill check) — deliver it with the binary:** signed
`agent_update`, then signed `agent_config_update`. No path added (26 → 26), so no step bundle.
### Part of v0.150.0 (2026-10-06 night, later) — after a restart the agent remembers the last backup per tier (R-894); three more SMART counters on the wire (R-330)
Ships with the memory-kill check below as v0.150.0, AFTER the 2026-10-07 night read-back. Nothing delivered tonight.
- **The defect (measured 2026-10-05 on demo-hp):** the agent restarted at 04:57; at 06:25 the off-site storage answered *Can't connect*; the per-tier backup record is in memory only, so the due-check fell back to an EMPTY record and the 7-day tier (last copy 4 days old) read DUE; the controller asked and vzdump failed.
- New `internal/backup/backup_state.go` `BackupSuccessState`: the newest SUCCESSFUL backup per tier and guest, on disk (`<oob state dir>/backup-success-state.json`, atomic tmp+rename, 0600). Only successes are written; a corrupt file reads as nothing known.
- `internal/localapi` `handleBackupDue`: when the tier's storage CANNOT be read, the saved copy stands in for the in-memory record. A fresh copy → not due („… (storage unreadable — age from the last success saved on disk)"); a copy older than the cadence → DUE; no copy → the old answer (DUE, age unknown). A storage that answers stays the ground truth: an archive absent there is due even when the file remembers one.
- Wired in `buildLocalAPIServer` (`LastKnownBackups`); the local API's backup job saves each success.
- Tests: `TestBackupDue_R894_*` (restart = a new server and a new state from the same file; fresh / old / none / storage answers / failed backup not saved), `TestBackupSuccessState_*`, `TestR894_LastKnownBackupsIsWiredIntoTheDaemon` (AST). Four red-proofs observed (`felhom.eu/documentation/audits/night-burndown-2026-10-06/s4/`).
- **R-330 (disk health Phase 2, the wire only):** the SMART summary carries three more SATA raw counters — `reported_uncorrect` (187), `command_timeout` (188, carried as the vendor reports it; some pack several counters), `udma_crc_errors` (199). Pointer + omitempty: an attribute the drive does not report is OMITTED (unknown), never 0. No verdict reads them yet. Tests `TestParseSMART_R330_*` (two red-proofs, `felhom.eu/documentation/audits/night-burndown-2026-10-06/r330/`).
### Part of v0.150.0 (2026-10-06 night) — the Docker step proves the engine reports a memory kill (`09` §3 decision 157, R-528)
To be released as v0.150.0 with its config bundle AFTER the 2026-10-07 night read-back (the night of 2026-10-06 runs v0.149.0 on purpose).
- `configs/felhom-os-apply`: after a docker-layer APPLY (after `health_after`) the wrapper runs `oom_check()`: a throwaway container from the image the running controller uses (`--pull never`, `--network none`, no volume, label `felhom.oomcheck=1`, 64 MB cap) asks for one 200 MB block; „pass" only when `OOMKilled=true` AND the `oom` event; it waits 2 s and reads the events window to the guest's epoch + 1 (measured: a window closed in the same second missed the event); the container is always removed. Reported as `oom_check`; it never changes the step's outcome or health. A wrapper-only mode `oom-check` runs the check alone (no apt, no engine change), by hand as root.
- `internal/osupdate`: `WrapperReport` and `Report` carry `oom_check` verbatim, on the normal pass and on the kept-copy path (R-868).
- `configs/test_felhom_os_apply.py`: its `unittest.main()` sat in the middle of the file, so 11 tests (UnsentReport, SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran — moved to the end; all pass.
- Tests: the OOMCheck class (pass, OOMKilled=false, no event, unreadable image, removal on an inspect error, not on other layers or in health mode, the events window after the settle wait, mode oom-check alone and its refusals); TestDocker_OOMCheckReachesTheHubUnchanged, TestR868_KeptCopyCarriesTheOOMCheck. 12 red-proofs in `felhom.eu/documentation/audits/readback-2026-10-07/F/`.
### Part of v0.150.0 (2026-10-06 evening) — the shared rule file (`09` §3 decision 152); no code change
- `.claude/rules/unprompted-work.md` added, byte-identical to the copies in felhom.eu, felhom-controller, app-catalog-felhom.eu and the workspace root (checked with `diff` against the controller's copy and one md5 across all five). Its copies line names five copies.
### Part of v0.150.0 (2026-10-06 afternoon) — instruction files kept true (`09` §3 decision 150); no code change
- `CLAUDE.md` „Gates — ONE entry point": the runner runs every gate in its `GATES` table (five: three shared, `published`, `release-complete`); `--fast` skips `published` (network). It said two gates and „all of them".
- `CLAUDE.md`: the decoy gate and its audit are named with their `felhom.eu/` prefix (they do not exist in this repo).
- `.claude/rules/health-checks.md` (comment): the health-check rule's copies live in felhom.eu `hub.md` and the controller's `gates.md`; it named felhom.eu `CLAUDE.md` „Code quality rules", which holds no such rule.
## v0.149.0 — a weekly disk trim of each customer guest, the crash-boot fact for the controller, the phantom WARN names its runbook (R-444, R-856, R-99; operator rulings `09` §3 139, 143, 140) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `6bcae9c2eb5d97e8285316583870059835793893299e291891a53a4ce505585f`
config bundle sha256 `e182c82dcf4a67faa3bcb74dbe4ffa7b06e0b27dc8451cb7574d6339ce91ad66` (tag `v0.149.0` = `f277e61`).
**The bundle carries the new sudoers rule for the trim (`FELHOM_FSTRIM`) — deliver it with the binary:** signed
`agent_update`, then signed `agent_config_update`.
- R-856 (`09` §3 decision 143): new local-API route `GET /host/crash-guard` — passes the host crash guard's last-boot record (present, last_boot_at, last_boot_unclean, tripped) from /var/lib/felhom-crash-guard/state.json to the controller, which waits ~15 min with app mails after a crash boot. Read-only, no Proxmox call, guest-token authed; a missing/unreadable/garbled file answers 200 present:false (never an error page). An older agent answers 404, which the controller reads as unknown (normal 90 s grace) — no controller MinAgent raise needed.
- R-444 (`09` §3 decision 139): weekly guest disk trim. New sudoers alias FELHOM_FSTRIM with ONE exact rule `/usr/sbin/pct ^fstrim [0-9]+$` (rides the signed config bundle; decoys pinned by TestSudoersFstrimRuleIsExact) and capability guest-fstrim (non-critical). New internal/fstrim job: each owned RUNNING guest gets `pct fstrim <vmid>` once a week - due Wednesday from 10:00 host-local, starts only 10:00-20:59 (never the 01:00-06:59 night), holds the one-heavy-op gate so it never runs beside a backup or restore-test (busy -> deferred to the next hourly tick; a box that was off catches up at its next daytime hour); a failed trim WARNs and is retried at most 3 times that week; bytes parsed from `pct fstrim`'s "(N bytes) trimmed" lines; positive log `fstrim: guest N trimmed X GiB in Ys`; last result per guest persisted in <state_dir>/guest-disk-trim.json and reported as the new omitempty host-report stanza `guest_disk_trim`. Opt-out: agent.json "disk_trim": {"disable": true}.
- R-99 (`09` §3 decision 140): the agent's WARN for a PBS archive below the 1 MiB plausibility floor now ends with the pointer to the sanctioned cleanup (`documentation/runbooks/pbs-phantom-cleanup.md`); detection only — nothing is deleted automatically. Dir-storage archives keep the old text.
## unreleased ## unreleased
- R-426: scripts/test_gate_decoys.py (new) — the published gate judged against a fake Gitea (127.0.0.1, via GITEA_BASE; never the real registry): 11 cases; COVERS published — `felhom-agent/published` leaves the decoy-coverage EXEMPT list.
- R-426: release-complete gate — 10 decoy cases (scratch clone + scratch bare origin + fake Gitea); COVERS release-complete — `felhom-agent/release-complete` leaves the decoy-coverage EXEMPT list.
- R-426: the shared reuse-refs/instructions/observations gates get agent-side decoys (9 cases on a scratch clone of this repo); COVERS reuse-refs, instructions, observations — three `felhom-agent/*` entries leave the decoy-coverage EXEMPT list.
- **Fixed without a row:** `configs/test_felhom_config_bundle.py` read the two ISO first-boot files that installer 1.32.0 now NAMES under KEPT (R-275) as files the installer writes — `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 (felhom.eu `85de3f9b`); they are listed with why. Test only; agent v0.148.0's code is unaffected.
## v0.148.0 — the host report names the running binary's sha; the format answer carries the new filesystem's UUID (burn-down night: R-349, R-25 agent halves) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `3e68a0870e0e2ce262cb4819294edddeb0a73e8c558a31611a20329a6d9ee283`
config bundle sha256 `a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de` (tag `v0.148.0` = `861d32a`).
Delivery order as for v0.147.0: signed `agent_update`, then signed `agent_config_update`.
- **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from - **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from
`/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub `/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub
comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved. comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved.
+7 -7
View File
@@ -52,11 +52,11 @@ This is in the core because breaching it is how this component stops being audit
## Gates — ONE entry point ## Gates — ONE entry point
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs this **Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs every
repo's gates — `reuse_refs_check` and `instructions_gate`, both the **shared** copies in gate in its `GATES` table (that table is the list); the shared ones — `reuse_refs_check`,
`felhom.eu/scripts/`, never copied into this repo (a copy recreates the drift they detect; an absent `instructions_gate`, `observations_gate` — are the copies in `felhom.eu/scripts/`, never copied into
sibling clone FAILS). `--fast` selects the gates touching no network and no container runtime; today this repo (a copy recreates the drift they detect; an absent sibling clone FAILS). `--fast` selects the
that is all of them. **A missing gate is a FAILURE, never a skip.** gates touching no network and no container runtime, and skips `published` (network), naming it. **A missing gate is a FAILURE, never a skip.**
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is **The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is
**per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS **per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS
@@ -104,8 +104,8 @@ the mechanism are exempt.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without **A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named lacks. `felhom.eu/scripts/decoy_coverage_gate.py` (run by felhom.eu's `repo_gates.py`, for all four repos) refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and decoys withdrawn as illegitimate: `felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over `felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list. `os.listdir`, and a glob over a hand-maintained list.
+7
View File
@@ -1,5 +1,12 @@
# CONTEXT — felhom-agent working state # CONTEXT — felhom-agent working state
> **2026-10-08 (day) — UNRELEASED on main, ships with tomorrow's release (decision 178).** R-899: `POST /backup?trigger=manual`
> runs no OS leg (`localapi/server.go`); after-boot kernel reports carry the saved block's ring (`kernelReportRing`). R-304:
> `escrow.ErrRetainedUnchecked` → HTTP 424 `older_unchecked` when not every earlier package was tried (withheld, caps,
> malformed, unreadable list). Binary-only delivery; no wrapper or root file changed.
> **2026-10-07 (evening) — v0.152.0, the kernel lane (R-836, decision 172, `11` §5.11).** Wrapper layer `kernel` + two GRUB generators in the bundle (delivered with step bundle `0.152.0-step1` — the bundle adds paths, R-880); `osupdate/kernel.go`: night step on told nights only, after-boot judge (host rule + hub reached, 20 min), ONE self-revert, `os_kernel_step` stages only. Proven on Tester 1 (panic → fell_back, held guest → self_reverted, healthy → 7.0.14-22 default). Open: R-898, R-897; the ring-0 night run.
> **2026-10-04 night — v0.143.0 RELEASED + vouched (R-840, decision 96): the config bundle.** `felhom-os-apply` mode > **2026-10-04 night — v0.143.0 RELEASED + vouched (R-840, decision 96): the config bundle.** `felhom-os-apply` mode
> `bundle` (signed `agent_config_update`, verified by the wrapper itself; trust files never bundle paths) + > `bundle` (signed `agent_config_update`, verified by the wrapper itself; trust files never bundle paths) +
> `--install-bundle` (installer 1.31.0); `BUNDLE_FILES` is the one table; `scripts/build-config-bundle.py`; > `--install-bundle` (installer 1.31.0); `BUNDLE_FILES` is the one table; `scripts/build-config-bundle.py`;
+8 -7
View File
@@ -44,13 +44,14 @@ unnoticed until a user hit them. `internal/capability` makes that loud:
(`HostCapabilityChecker`) alerts the operator on a Critical capability going degraded. Serve-degraded (`HostCapabilityChecker`) alerts the operator on a Critical capability going degraded. Serve-degraded
— the probe never blocks startup. (Next self-health slice: the controller↔agent channel check.) — the probe never blocks startup. (Next self-health slice: the controller↔agent channel check.)
**Controller-swap under non-root (v0.45.0).** The agent-owned controller image swap **Controller-swap under non-root (v0.45.0; the write since R-861 (a) A1).** The agent-owned controller image swap
(`internal/localapi/controllerswap.go`) no longer shells out: `writeImage` pipes the image ref on (`internal/localapi/controllerswap.go`) no longer shells out. The write goes on **stdin** to the ROOT verb
**stdin** into an in-guest `tee /etc/felhom-controller-image` (via `GuestExecStdin` → `felhom-priv-apply controller-image <vmid>` (`GuestBinder.WriteControllerImage` → `Runner.RunStdin`, the same fenced
`Runner.RunStdin`, the same fenced `sudo -n` runner) — no `bash -c`, no interpolation. Its 5 narrow `sudo -n` runner), which re-checks the ref against our registry + repository + an x.y.z tag and writes
grants live in the `FELHOM_CONTROLLERSWAP` sudoers alias (all read-only or fixed-target; the `tee` `/etc/felhom-controller-image` inside the guest itself; the agent has no in-guest `tee` grant any more (before, a
target is the FIXED image path, content stdin-fed) and in the capability manifest (Critical), so a compromised agent could feed any image — sudo cannot see stdin). Its grants live in the `FELHOM_CONTROLLERSWAP` sudoers
dropped grant is a build failure + a live degraded signal. No general `pct exec` is granted. alias (read-only or fixed-target) and in the capability manifest (Critical), so a dropped grant is a build failure + a
live degraded signal. No general `pct exec` is granted.
## The `storage` package — observe + watchdog (slice 5) ## The `storage` package — observe + watchdog (slice 5)
+13 -8
View File
@@ -1,10 +1,15 @@
# REPORT — agent v0.147.0 (2026-10-05, burn-down round 2) # REPORT — v0.154.0 released and delivered (2026-10-09)
Full session report: `felhom.eu/REPORT-burndown2-2026-10-05.md`. Baseline `d833163` (v0.146.1). Code commit `f1b9b41` Built from the 2026-10-09 kernel-night read-back: (1) the host report now adds the saved newest backup success per tier
(CI job 1365 success), tag `v0.147.0`, binary sha256 `642c4d19…`, bundle sha256 `326527d0…` (verified by download). (R-894's file, one shared instance), so a restart right after a night backup no longer empties the hub's evidence — on
demo-felhom it had caused a false „newest backup is 48h old" alarm; (2) `proxmox-secure-boot-support` joins the
boot-chain lane, so it no longer pulls `shim-signed-common` into the ring-0 Proxmox plan (demo-hp's step was refused R6
and its kernel step skipped). Red-proofs: `TestCollectBackups_SavedSuccessSurvivesARestart`,
`TestR894_LastKnownBackupsIsWiredIntoTheDaemon`, `test_ring0_pending_pve_leaves_the_secure_boot_meta_out`.
Rows: R-124 (recipe root namespace = ""), R-118 (no root size for an absent drive), R-269 (rotated-out token rejected at `scripts/release-agent.sh 0.154.0`: sha `36a54ad2…`, bundle `93487989…`, tag `v0.154.0` = `db25b46`, verified by
once), R-317 (dnsmasq install probed by its unit). Tests + red-proofs: `felhom.eu/documentation/audits/burndown2-2026-10-05/`. download. CI 1576 (code) success; 1579 (CHANGELOG) failed in its fetch step (Gitea answered 500), re-run → success.
`go build/vet/test ./...` green; `agent_gates.py --fast` green after the release (release-complete needs the tag). Signed `agent_update` → demo-hp, demo-felhom, Tester 1 (07:16–07:18 local); signed `agent_config_update` → `BUNDLE DONE
written=1 same=26 self-check=ok`, capability probe 68/68 on all three. NOT vouched (the vouch needs a golden at or above
Delivery: see the session report (vouch, signed jobs per box, the hub System page afterwards). the fleet's controller; no golden today). Live: demo-felhom's first report after the update carries the 02:40Z backup.
Evidence: `felhom.eu/documentation/audits/release-2026-10-09/delivery/`.
+7 -3
View File
@@ -15,7 +15,7 @@
| `SudoHostOps.run` | internal/storage/hostops.go | `run(ctx, name, args...) error` | allowlisted exec with stderr-wrapped error | Every arg pre-validated via validate.go before this is called | | `SudoHostOps.run` | internal/storage/hostops.go | `run(ctx, name, args...) error` | allowlisted exec with stderr-wrapped error | Every arg pre-validated via validate.go before this is called |
| `Prober.Probe` | internal/capability/probe.go | `Probe(ctx) []Status` | live sudo-policy capability check (`sudo -n -l --`) | Needs a DIRECT runner (never the sudo-prefixing one — double-sudo); never executes probed cmds. v0.86.0: config-gated caps (`Capability.GatedBy` + `Prober.GateActive`) report `inactive`/"disabled by configuration" ONLY when healthy — broken plumbing stays degraded; the pbsdr-* gate answers from `pbsdr.Manager.DRConfigured` (marker-backed across restarts) | | `Prober.Probe` | internal/capability/probe.go | `Probe(ctx) []Status` | live sudo-policy capability check (`sudo -n -l --`) | Needs a DIRECT runner (never the sudo-prefixing one — double-sudo); never executes probed cmds. v0.86.0: config-gated caps (`Capability.GatedBy` + `Prober.GateActive`) report `inactive`/"disabled by configuration" ONLY when healthy — broken plumbing stays degraded; the pbsdr-* gate answers from `pbsdr.Manager.DRConfigured` (marker-backed across restarts) |
| ~~`stageTemp`~~ (REMOVED v0.146.0, R-861) | — | — | — | Nothing the agent writes is `install`ed where root reads it any more: use `felhom-priv-apply` (below) or ship a fixed file in the bundle | | ~~`stageTemp`~~ (REMOVED v0.146.0, R-861) | — | — | — | Nothing the agent writes is `install`ed where root reads it any more: use `felhom-priv-apply` (below) or ship a fixed file in the bundle |
| `felhom-priv-apply` (v0.146.0, R-861) | configs/felhom-priv-apply | `felhom-priv-apply unit <name> \| dnsmasq <tmp> <name> \| wg \| sshd-config \| sshd-key` | ANY agent-rendered file a root program reads (systemd unit, dnsmasq drop-in, wg-quick conf, OOB sshd) — fixed source + destination, CONTENT checked against the agent's own renderers | A new renderer needs a verb + a contract test (`internal/privapplytest.Check`) feeding its REAL output; never a new `install` sudoers line | | `felhom-priv-apply` (v0.146.0, R-861) | configs/felhom-priv-apply | `felhom-priv-apply unit <name> \| dnsmasq <tmp> <name> \| wg \| sshd-config \| sshd-key \| controller-image <vmid>` (the last reads the ref on stdin, R-861 (a) A1) | ANY agent-rendered file a root program reads (systemd unit, dnsmasq drop-in, wg-quick conf, OOB sshd) — fixed source + destination, CONTENT checked against the agent's own renderers | A new renderer needs a verb + a contract test (`internal/privapplytest.Check`) feeding its REAL output; never a new `install` sudoers line |
| `privapplytest.Check` | internal/privapplytest/check.go | `Check(t, verb, name, content) string` | the Go↔root-checker contract: a renderer's real output must read `OK` | Skips without python3; one call per rendered shape + one refused control | | `privapplytest.Check` | internal/privapplytest/check.go | `Check(t, verb, name, content) string` | the Go↔root-checker contract: a renderer's real output must read `OK` | Skips without python3; one call per rendered shape + one refused control |
| `BUNDLE_FILES` + `Bundle` (mode `bundle`, `--install-bundle`; mode `agent_update` v0.146.0, R-861) | configs/felhom-os-apply | the ONE table of root-owned paths + the installer of them | ANY new root-owned file the installer writes (sudoers line, wrapper, unit) — add it to the table, never a new installer fetch (R-840) | The builder (`scripts/build-config-bundle.py`) and the installer read the same table; `test_every_root_file_the_installer_writes_is_in_the_bundle` fails on a path the bundle lacks. Trust files (`/etc/felhom/os-trust.json`, `operator-signers`) are NEVER bundle paths (R17) | | `BUNDLE_FILES` + `Bundle` (mode `bundle`, `--install-bundle`; mode `agent_update` v0.146.0, R-861) | configs/felhom-os-apply | the ONE table of root-owned paths + the installer of them | ANY new root-owned file the installer writes (sudoers line, wrapper, unit) — add it to the table, never a new installer fetch (R-840) | The builder (`scripts/build-config-bundle.py`) and the installer read the same table; `test_every_root_file_the_installer_writes_is_in_the_bundle` fails on a path the bundle lacks. Trust files (`/etc/felhom/os-trust.json`, `operator-signers`) are NEVER bundle paths (R17) |
| `osupdate.ConfigUpdateExecutor` | internal/osupdate/bundle.go | signed op `agent_config_update` {agent_version, bundle_sha256} | delivering the bundle to an installed box | a courier only: the root wrapper re-verifies signature, host, nonce and sha itself | | `osupdate.ConfigUpdateExecutor` | internal/osupdate/bundle.go | signed op `agent_config_update` {agent_version, bundle_sha256} | delivering the bundle to an installed box | a courier only: the root wrapper re-verifies signature, host, nonce and sha itself |
@@ -88,6 +88,8 @@
| Symbol | File | Short signature | Use for | Gotchas | | Symbol | File | Short signature | Use for | Gotchas |
|---|---|---|---|---| |---|---|---|---|---|
| `pvegate.Write` / `pvegate.Step` | internal/pvegate/pvegate.go | `Write(ctx) (release, waited, err)` / `Step(ctx) (end, err)` | R-812 option A: keep the agent's own /etc/pve writes out of a Proxmox package step (pmxcfs restarts) | Already wired at the two chokepoints — `Client.doBody` (every non-GET) and `ExecRunner.RunStdin` (`WritesEtcPVE`: pct config verbs, pvesm, pveum, felhom-pbs-apply create/reconcile). A new root CLI that writes /etc/pve goes into `WritesEtcPVE`, never its own lock. Never take `Step` around anything but the wrapper call (`Leg.runPVE`) — a `Write` inside a `Step` deadlocks until its context ends. |
| `osupdate.KernelVerdict` / `Leg.KernelAfterBoot` | internal/osupdate/kernel.go | `KernelVerdict(before, after, tunnel, hubReached) (ok, why)` | R-836: THE one-shot-boot rule — the host health rule (`HostHealthVerdict`) AND the box reached the hub since the boot | The only judge of a new kernel. Never reboot the host from Go: every reboot is the wrapper's (`kernel-reboot`, `kernel-revert`), and only for a staged step. A new kernel-lane state lives in the wrapper's `/var/lib/felhom-kernel/state.json`, never in the agent's own files (the agent can write those). |
| `Client.WaitTask` | internal/proxmox/task.go | `WaitTask(ctx, upid, opts) (TaskStatus, error)` | asserting EVERY mutating op | POST 200 ≠ success; authz can fail at task exec; `AllowWarnings` opt-in | | `Client.WaitTask` | internal/proxmox/task.go | `WaitTask(ctx, upid, opts) (TaskStatus, error)` | asserting EVERY mutating op | POST 200 ≠ success; authz can fail at task exec; `AllowWarnings` opt-in |
| `Client.Pool` | internal/proxmox/query.go | `Pool(ctx, name) (PoolInfo, error)` | felhom-pool membership (the ownership registry, A1) | Needs `Pool.Audit` at `/pool/<name>` (host-install v1.9.0+); `Pool.Allocate` does NOT satisfy the read; members can be storages (type `storage`, vmid 0) — filter them | | `Client.Pool` | internal/proxmox/query.go | `Pool(ctx, name) (PoolInfo, error)` | felhom-pool membership (the ownership registry, A1) | Needs `Pool.Audit` at `/pool/<name>` (host-install v1.9.0+); `Pool.Allocate` does NOT satisfy the read; members can be storages (type `storage`, vmid 0) — filter them |
| `Client` mutate wrappers (`RestoreLXC/Vzdump/DestroyLXC/Snapshot/Rollback/SetConfig/ResizeLXC/Start/Stop`) | internal/proxmox/mutate.go | return `(upid, error)` | all API mutations | Async → always pair with WaitTask; route via gate/queue, not ad-hoc | | `Client` mutate wrappers (`RestoreLXC/Vzdump/DestroyLXC/Snapshot/Rollback/SetConfig/ResizeLXC/Start/Stop`) | internal/proxmox/mutate.go | return `(upid, error)` | all API mutations | Async → always pair with WaitTask; route via gate/queue, not ad-hoc |
@@ -121,7 +123,7 @@
| Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only | | Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only |
| Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set | | Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set |
| Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction | | Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction |
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did | | Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`; R-856 `crashGuardStatePath` — GET /host/crash-guard's state file, internal/localapi/crashguard.go) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
| `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) | | `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) |
| `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs | | `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs |
| `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s | | `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s |
@@ -160,13 +162,15 @@
| `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring | | `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring |
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go | | `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). | | `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. | | `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. | | `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
| `backup.BackupSuccessState` | internal/backup/backup_state.go | `RecordBackupSuccess(target, b)` / `LastKnownSuccess(target, vmid)` | Newest SUCCESSFUL backup per tier+guest, persisted (atomic tmp+rename) — the due-check's fallback when the tier's storage cannot be read after a restart (R-894) | **Read ONLY when the storage cannot be read** — a storage that answers is the ground truth (R-84), and an archive absent there must make the tier due even when this file remembers one. A saved copy older than the cadence still reads due. Only successes are written. |
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. | | `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. | | `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). | | `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
| `localapi.StaleLockController` | internal/localapi/stalelock.go | `*staleLockController` (Client + Runner + pool) | `fakeStaleLock` (Server-level) stalelock_test.go; `fakeStaleLockAPI` (controller-level, tests the A1 pool intersect) stalelock_pool_test.go | | `localapi.StaleLockController` | internal/localapi/stalelock.go | `*staleLockController` (Client + Runner + pool) | `fakeStaleLock` (Server-level) stalelock_test.go; `fakeStaleLockAPI` (controller-level, tests the A1 pool intersect) stalelock_pool_test.go |
| `localapi.GuestExecutor` | internal/localapi/controllerswap.go | `*GuestBinder` (pct exec) | `fakeGuestExec` internal/localapi/controllerswap_test.go | | `localapi.GuestExecutor` | internal/localapi/controllerswap.go | `*GuestBinder` (pct exec; the image write via `felhom-priv-apply controller-image`) | `fakeGuestExec` internal/localapi/controllerswap_test.go |
| `guestnet.Runner` / `guestnet.GuestSource` (R-54, v0.92.0) | internal/guestnet/{probe,watchdog}.go | `*proxmox.ExecRunner`; the POOL-VERIFIED `localapi.StaleLockController.Guests` (ListLXC ∩ felhom pool, audit A1) | `scriptedRunner` + `fakeGuests` internal/guestnet/watchdog_test.go. **Never wire a bare `ListLXC` here** — under a broad token that would run dhclient inside a co-tenant's container. Every assertion is an exec COUNT, and the load-bearing ones are the negatives: a static guest, an unprobeable guest, a boot-race guest and an unproven guest list must record **zero** heal calls | | `guestnet.Runner` / `guestnet.GuestSource` (R-54, v0.92.0) | internal/guestnet/{probe,watchdog}.go | `*proxmox.ExecRunner`; the POOL-VERIFIED `localapi.StaleLockController.Guests` (ListLXC ∩ felhom pool, audit A1) | `scriptedRunner` + `fakeGuests` internal/guestnet/watchdog_test.go. **Never wire a bare `ListLXC` here** — under a broad token that would run dhclient inside a co-tenant's container. Every assertion is an exec COUNT, and the load-bearing ones are the negatives: a static guest, an unprobeable guest, a boot-race guest and an unproven guest list must record **zero** heal calls |
| `guestnet.Watchdog.SetDampers` / `now` (clock seam) | internal/guestnet/watchdog.go | config `guest_net.*`; `now` defaults to `time.Now` | tests advance a manual clock (the storage-watchdog pattern) and assert the heal ceilings EXACTLY — ≥10 min apart, ≤3/hour, and ≤30 over a scripted 10 hours of permanent failure. A damper with no test is a comment | | `guestnet.Watchdog.SetDampers` / `now` (clock seam) | internal/guestnet/watchdog.go | config `guest_net.*`; `now` defaults to `time.Now` | tests advance a manual clock (the storage-watchdog pattern) and assert the heal ceilings EXACTLY — ≥10 min apart, ≤3/hour, and ≤30 over a scripted 10 hours of permanent failure. A damper with no test is a comment |
| `hub.GuestNetReporter` (R-54) | internal/hub/collect.go | `*guestnet.Watchdog` (`GuestNetStatus`) | internal/hub/collect_guestnet_test.go asserts the stanza through the PRODUCTION `Collect` path AND that the `guest_net` key is ABSENT from the wire when no reporter is wired — an always-present empty stanza would make "not wired" and "found nothing" the same signal, which is the shape v0.91.0 hid behind | | `hub.GuestNetReporter` (R-54) | internal/hub/collect.go | `*guestnet.Watchdog` (`GuestNetStatus`) | internal/hub/collect_guestnet_test.go asserts the stanza through the PRODUCTION `Collect` path AND that the `guest_net` key is ABSENT from the wire when no reporter is wired — an always-present empty stanza would make "not wired" and "found nothing" the same signal, which is the shape v0.91.0 hid behind |
+53
View File
@@ -0,0 +1,53 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"testing"
)
// R-444: the weekly trim has the guestnet shape (component + reporter seam + goroutine), so its wiring is asserted
// from the AST like TestMainWiresGuestNetWatchdog — a unit-green trim job that main.go never starts is the inert-seam
// defect. It must also share the ONE heavy-op gate (heavyOps), or it could run beside a backup.
func TestMainWiresGuestDiskTrim(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, 0)
if err != nil {
t.Fatalf("parse main.go: %v", err)
}
var constructedWithGate, reporterWired, started bool
ast.Inspect(f, func(n ast.Node) bool {
switch node := n.(type) {
case *ast.CallExpr:
if fn, ok := node.Fun.(*ast.SelectorExpr); ok {
switch fn.Sel.Name {
case "New":
if pkg, ok := fn.X.(*ast.Ident); ok && pkg.Name == "fstrim" && len(node.Args) >= 3 {
if id, ok := node.Args[2].(*ast.Ident); ok && id.Name == "heavyOps" {
constructedWithGate = true
}
}
case "SetGuestDiskTrimReporter":
reporterWired = true
}
}
case *ast.GoStmt:
if sel, ok := node.Call.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "Run" {
if id, ok := sel.X.(*ast.Ident); ok && id.Name == "diskTrim" {
started = true
}
}
}
return true
})
if !constructedWithGate {
t.Error("main.go never calls fstrim.New(..., heavyOps, ...) — no trim job, or one outside the heavy-op gate")
}
if !reporterWired {
t.Error("main.go never calls collector.SetGuestDiskTrimReporter — the guest_disk_trim stanza never reaches the hub")
}
if !started {
t.Error("main.go never starts the trim job with `go diskTrim.Run(ctx)`")
}
}
+78 -26
View File
@@ -38,6 +38,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow" "gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
"gitea.dooplex.hu/admin/felhom-agent/internal/fasttick" "gitea.dooplex.hu/admin/felhom-agent/internal/fasttick"
"gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd" "gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd"
"gitea.dooplex.hu/admin/felhom-agent/internal/fstrim"
"gitea.dooplex.hu/admin/felhom-agent/internal/guesthook" "gitea.dooplex.hu/admin/felhom-agent/internal/guesthook"
"gitea.dooplex.hu/admin/felhom-agent/internal/guestnet" "gitea.dooplex.hu/admin/felhom-agent/internal/guestnet"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub" "gitea.dooplex.hu/admin/felhom-agent/internal/hub"
@@ -163,7 +164,6 @@ func main() {
keyDest string keyDest string
installWGKey bool installWGKey bool
idBundlePath string idBundlePath string
directivePath string
swapImage string swapImage string
outputMode string outputMode string
showVersion bool showVersion bool
@@ -197,7 +197,6 @@ func main() {
flag.StringVar(&keyDest, "keydest", "", "for --selftest=escrow-consume: where to install the recovered key (0600)") flag.StringVar(&keyDest, "keydest", "", "for --selftest=escrow-consume: where to install the recovered key (0600)")
flag.BoolVar(&installWGKey, "install-wg-key", false, "for --selftest=identity-consume: ALSO install the recovered wg_private_key into wgtunnel's key file (S5 DR; create-only, refuses to overwrite)") flag.BoolVar(&installWGKey, "install-wg-key", false, "for --selftest=identity-consume: ALSO install the recovered wg_private_key into wgtunnel's key file (S5 DR; create-only, refuses to overwrite)")
flag.StringVar(&idBundlePath, "identity-bundle", "", "for --selftest=escrow-create: a 0600 JSON file {tunnel_token,pbs_token} to ALSO escrow under R (10D)") flag.StringVar(&idBundlePath, "identity-bundle", "", "for --selftest=escrow-create: a 0600 JSON file {tunnel_token,pbs_token} to ALSO escrow under R (10D)")
flag.StringVar(&directivePath, "directive", "", "for --selftest=escrow-create: a JSON file with the non-secret DR directive (pbs repo/ns, expected fingerprint, tunnel id)")
flag.StringVar(&custID, "customer-id", "", "for --selftest=provision: the customer id — the hub config-pull target, baked into the guest's bootstrap") flag.StringVar(&custID, "customer-id", "", "for --selftest=provision: the customer id — the hub config-pull target, baked into the guest's bootstrap")
flag.StringVar(&hubPassword, "hub-password", "", "for --selftest=provision: the customer's hub RETRIEVAL PASSPHRASE (SECRET) — baked into bootstrap.json so the controller pulls its config (and the customer-scoped hub key) from the hub. The customer must already exist in the hub.") flag.StringVar(&hubPassword, "hub-password", "", "for --selftest=provision: the customer's hub RETRIEVAL PASSPHRASE (SECRET) — baked into bootstrap.json so the controller pulls its config (and the customer-scoped hub key) from the hub. The customer must already exist in the hub.")
flag.StringVar(&swapImage, "image", "", "for --selftest=controller-swap: the target controller image ref (gitea.dooplex.hu/admin/felhom-controller:<semver>) — must already be pulled in the guest") flag.StringVar(&swapImage, "image", "", "for --selftest=controller-swap: the target controller image ref (gitea.dooplex.hu/admin/felhom-controller:<semver>) — must already be pulled in the guest")
@@ -268,7 +267,7 @@ func main() {
SysDataGrowGB: sysDataGrow, SysDataMount: sysDataMount, Cores: cores, MemoryMB: memoryMB}, SysDataGrowGB: sysDataGrow, SysDataMount: sysDataMount, Cores: cores, MemoryMB: memoryMB},
})) }))
case "escrow-create": case "escrow-create":
os.Exit(runSelftestEscrowCreate(context.Background(), cfg, logger, pbsStorage, paperkey, offline, upload, idBundlePath, directivePath, outputMode)) os.Exit(runSelftestEscrowCreate(context.Background(), cfg, logger, pbsStorage, paperkey, offline, upload, idBundlePath, outputMode))
case "escrow-consume": case "escrow-consume":
os.Exit(runSelftestEscrowConsume(context.Background(), logger, blobPath, expectedFP, keyDest)) os.Exit(runSelftestEscrowConsume(context.Background(), logger, blobPath, expectedFP, keyDest))
case "identity-consume": case "identity-consume":
@@ -791,6 +790,9 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
pbsTargets := pbsTargetsFromPVE(cfg, px, logger) pbsTargets := pbsTargetsFromPVE(cfg, px, logger)
pbsReporter := pbs.NewLiveSnapshotReporter(pbsTargets, pbsStore, pbs.DefaultLiveSnapshotTimeout, logger) pbsReporter := pbs.NewLiveSnapshotReporter(pbsTargets, pbsStore, pbs.DefaultLiveSnapshotTimeout, logger)
collector := hub.NewCollector(px, newTunnelProber(cfg, px), observer, backupStore, backupStore, pbsReporter, cfg.Hub.HostID, version, logger) collector := hub.NewCollector(px, newTunnelProber(cfg, px), observer, backupStore, backupStore, pbsReporter, cfg.Hub.HostID, version, logger)
// R-366 slice 2: the restore-test's ledger of archives written with another key → the host report.
foreignKeys := backup.NewForeignKeyLedger()
collector.SetForeignKeyArchiveReporter(foreignKeys)
collector.SetBackupTargetResolver(primaryBackupTargetOf(cfg)) // R-109: the recipe names the live target collector.SetBackupTargetResolver(primaryBackupTargetOf(cfg)) // R-109: the recipe names the live target
// Privileged-capability self-check (v0.44.0): probe the sudoers grants the non-root agent // Privileged-capability self-check (v0.44.0): probe the sudoers grants the non-root agent
// depends on. The probe runs `sudo -n -l` LITERALLY (a policy LIST, never executing the // depends on. The probe runs `sudo -n -l` LITERALLY (a policy LIST, never executing the
@@ -875,6 +877,13 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
go osLeg.SendUnsentLoop(ctx, 5*time.Minute, func(n int) { go osLeg.SendUnsentLoop(ctx, 5*time.Minute, func(n int) {
logger.Info("osupdate: sent kept report(s)", "count", n) logger.Info("osupdate: sent kept report(s)", "count", n)
}) })
// R-836 (`09` §3 decision 172): what became of a kernel step across this boot; on a one-shot boot of a new kernel,
// judge it (the host health rule + the hub reached) for KernelJudgeWait, then make it the default or revert ONCE.
// The wait: measured 2026-10-07 (`audits/kernel-lane-2026-10-07/B/`) — every container healthy 68 s after the
// reboot on demo-felhom and 272 s on demo-hp (the hub reached at 63 s / 189 s); 20 minutes leaves room for a slow
// network and stays under the hub's 45-minute host_stale (its
// alerting.stale_threshold; host_down at 90). On a box without the kernel lane (an older wrapper, a BYO host) the check is refused and logged.
go osLeg.KernelAfterBoot(ctx, 0, osupdate.KernelJudge{Wait: osupdate.DefaultKernelJudgeWait})
// Reconcile (slice 4) runs alongside the hub loop, sharing the per-guest queue // Reconcile (slice 4) runs alongside the hub loop, sharing the per-guest queue
// (doc 03 §10). At slice 4 the desired-state provider is empty (no hub serving // (doc 03 §10). At slice 4 the desired-state provider is empty (no hub serving
@@ -1016,7 +1025,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// with the local API so a backup and a restore-test can never run together. // with the local API so a backup and a restore-test can never run together.
rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json")) rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json"))
heavyOps := &backup.InFlight{} heavyOps := &backup.InFlight{}
scheduler := buildRestoreTestScheduler(cfg, px, engine, backupStore, rtState, heavyOps, logger) scheduler := buildRestoreTestScheduler(cfg, px, engine, backupStore, rtState, heavyOps, foreignKeys, logger)
// R-189: the host report's restore_tests[] must survive an agent restart. The in-memory store // R-189: the host report's restore_tests[] must survive an agent restart. The in-memory store
// holds only this process's latest run, and under per-archive due-ness the agent will not // holds only this process's latest run, and under per-archive due-ness the agent will not
// re-test an archive it has already proven — so without this the hub can report a tier unproven // re-test an archive it has already proven — so without this the hub can report a tier unproven
@@ -1108,6 +1117,25 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
} }
return release, nil return release, nil
}} }}
// R-812 option A: a signed Proxmox package step (`11` §5.10) — ring 1; under the heavy-op gate and the /etc/pve gate.
pveExec := osupdate.PVEStepExecutor{Leg: osLeg, Guest: firstGuest(px),
Gate: func(ctx context.Context) (func(), error) {
release, busy, ok := heavyOps.TryAcquire("os-pve-step")
if !ok {
return nil, fmt.Errorf("busy: %s", busy)
}
return release, nil
}}
// R-836 (`09` §3 decision 172): a signed kernel step STAGES a kernel on a ring-1 box (install + the one-shot flag,
// never a reboot — the night leg reboots it on a night the household was told about); under the heavy-op gate.
kernelExec := osupdate.KernelStepExecutor{Leg: osLeg, Guest: firstGuest(px),
Gate: func(ctx context.Context) (func(), error) {
release, busy, ok := heavyOps.TryAcquire("os-kernel-step")
if !ok {
return nil, fmt.Errorf("busy: %s", busy)
}
return release, nil
}}
// Agent v0.143.0 (R-840): the config bundle — the box's root-owned files by a signed job; the wrapper verifies it. // Agent v0.143.0 (R-840): the config bundle — the box's root-owned files by a signed job; the wrapper verifies it.
bundleExec := osupdate.ConfigUpdateExecutor{Leg: osLeg, URLTemplate: suCfg.URLTemplate, Username: suCfg.Username, Token: suCfg.Token, bundleExec := osupdate.ConfigUpdateExecutor{Leg: osLeg, URLTemplate: suCfg.URLTemplate, Username: suCfg.Username, Token: suCfg.Token,
// The capability probe confirms from the agent's side: `sudo -l` lists every command the new sudoers grants. // The capability probe confirms from the agent's side: `sudo -l` lists every command the new sudoers grants.
@@ -1119,7 +1147,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
} }
logger.Warn("osupdate: capability probe after the config bundle", "ok", ok, "total", total, "degraded", strings.Join(names, ",")) logger.Warn("osupdate: capability probe after the config bundle", "ok", ok, "total", total, "degraded", strings.Join(names, ","))
}} }}
jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec, dockerExec, bundleExec}, cfg.Hub.HostID, logger) jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec, dockerExec, pveExec, kernelExec, bundleExec}, cfg.Hub.HostID, logger)
loop.SetEnvelopeObserver(hub.MultiObserver(desiredSyncer, jobsRunner)) loop.SetEnvelopeObserver(hub.MultiObserver(desiredSyncer, jobsRunner))
// Controller-driven escrow ceremony (v0.88.0): static config facts + the LATE-BOUND DR gate — // Controller-driven escrow ceremony (v0.88.0): static config facts + the LATE-BOUND DR gate —
@@ -1479,6 +1507,25 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
} }
go runJanitor(ctx, jd) go runJanitor(ctx, jd)
} }
// R-444 (`09` §3 decision 139): the weekly guest disk trim — `pct fstrim <vmid>` of every owned, running guest,
// Wednesday from 10:00 local, daytime only, under the one-heavy-op gate. Not part of the errc fan-out: a trim job
// must never be able to bring the agent down.
if cfg.DiskTrim.Enabled() {
dtMode := proxmox.RunnerMode(cfg.Privileged.Mode)
if dtMode == "" {
dtMode = proxmox.RunnerSudo
}
dtRunner := &proxmox.ExecRunner{Mode: dtMode, SudoPath: cfg.Privileged.SudoPath}
dtGuests := localapi.NewStaleLockController(px, dtRunner, reconcile.DefaultPool, logger)
if dtGuests != nil {
diskTrim := fstrim.New(dtRunner, dtGuests, heavyOps,
filepath.Join(cfg.OOB.WithDefaults().StateDir, "guest-disk-trim.json"), logger)
collector.SetGuestDiskTrimReporter(diskTrim)
go diskTrim.Run(ctx)
}
} else {
logger.Info("fstrim: weekly guest disk trim disabled by config (disk_trim.disable)")
}
if lanLoop != nil { if lanLoop != nil {
lanServers = 1 lanServers = 1
go func() { errc <- lanLoop.Run(ctx) }() go func() { errc <- lanLoop.Run(ctx) }()
@@ -1680,7 +1727,7 @@ func primaryBackupTargetOf(cfg config.Config) func() hub.ConfiguredBackupTarget
// disables the cadence (returns a scheduler that just waits) when the cadence is off or the // disables the cadence (returns a scheduler that just waits) when the cadence is off or the
// scratch band / restore storage is invalid — a misconfig must not crash the daemon, and the // scratch band / restore storage is invalid — a misconfig must not crash the daemon, and the
// machinery still works on-demand via --selftest=restore-test. // machinery still works on-demand via --selftest=restore-test.
func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *reconcile.Engine, store *backup.Store, rtState *backup.RestoreTestState, inFlight *backup.InFlight, logger *slog.Logger) *backup.Scheduler { func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *reconcile.Engine, store *backup.Store, rtState *backup.RestoreTestState, inFlight *backup.InFlight, foreign *backup.ForeignKeyLedger, logger *slog.Logger) *backup.Scheduler {
// R-86: this is the EVALUATION interval, not the trigger. What decides a test happens is the // R-86: this is the EVALUATION interval, not the trigger. What decides a test happens is the
// per-archive due-check in internal/backup/restoretest_due.go. // per-archive due-check in internal/backup/restoretest_due.go.
cadence := cfg.Backup.RestoreTestEvalInterval() cadence := cfg.Backup.RestoreTestEvalInterval()
@@ -1699,6 +1746,9 @@ func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *re
min, max := cfg.Backup.ScratchBand() min, max := cfg.Backup.ScratchBand()
target := cfg.Backup.BackupTarget() target := cfg.Backup.BackupTarget()
runner := backup.NewBackupRunner(px, target, "", "felhom restore-test", "", logger) runner := backup.NewBackupRunner(px, target, "", "felhom restore-test", "", logger)
if foreign != nil {
runner.SetForeignKeyLedger(foreign) // R-366 slice 2: the pick records archives written with another key
}
// Every configured tier is a rotation candidate, not just the primary. // Every configured tier is a rotation candidate, not just the primary.
cfgTiers, _ := cfg.Backup.BackupTiers() // warnings already logged where the tiers are armed cfgTiers, _ := cfg.Backup.BackupTiers() // warnings already logged where the tiers are armed
tierIDs := make([]string, 0, len(cfgTiers)) tierIDs := make([]string, 0, len(cfgTiers))
@@ -1854,11 +1904,15 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
return nil, 0, ferr return nil, 0, ferr
} }
out := make([]escrow.RetainedBlob, 0, len(resp.Packages)) out := make([]escrow.RetainedBlob, 0, len(resp.Packages))
// R-304: every package the hub holds and the code will NOT be tried against — no key material,
// over the hub's cap, or malformed here. A refusal may call the code wrong only when this is 0.
withheld := resp.UnopenableCount + resp.TruncatedCount
for _, p := range resp.Packages { for _, p := range resp.Packages {
blob, derr := base64.StdEncoding.DecodeString(p.IdentityEscrowB64) blob, derr := base64.StdEncoding.DecodeString(p.IdentityEscrowB64)
if derr != nil || len(blob) == 0 { if derr != nil || len(blob) == 0 {
// One malformed package must not sink the rest — the customer's code may open a // One malformed package must not sink the rest — the customer's code may open a
// later one, and a skipped entry is strictly better than a refusal we cannot justify. // later one, and a skipped entry is strictly better than a refusal we cannot justify.
withheld++
continue continue
} }
out = append(out, escrow.RetainedBlob{ out = append(out, escrow.RetainedBlob{
@@ -1868,9 +1922,13 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
Index: p.Index, Index: p.Index,
}) })
} }
return out, resp.UnopenableCount, nil return out, withheld, nil
}, },
} }
// R-894: ONE instance — the local API writes it, the host report reads it (2026-10-09: a restart right
// after the night backup erased the hub's evidence of it).
lastKnownBackups := backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json"))
collector.SetKnownBackupReporter(lastKnownBackups)
srv, err := localapi.NewServer(localapi.Options{ srv, err := localapi.NewServer(localapi.Options{
EscrowRecovery: escrowRecoverer, EscrowRecovery: escrowRecoverer,
ListenAddr: cfg.LocalAPI.ListenAddr, ListenAddr: cfg.LocalAPI.ListenAddr,
@@ -1881,6 +1939,9 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F) InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
Store: store, Store: store,
// R-894: the newest success per tier on disk — the due-check's fallback when the storage cannot be
// read right after a restart. Same state dir as restore-test-state.json.
LastKnownBackups: lastKnownBackups,
Storage: observer, Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages) DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
@@ -2241,7 +2302,7 @@ func runSelftestRestoreTestDue(ctx context.Context, cfg config.Config, logger *s
return 1 return 1
} }
rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json")) rtState := backup.NewRestoreTestState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "restore-test-state.json"))
sched := buildRestoreTestScheduler(cfg, px, nil, backup.NewStore(), rtState, &backup.InFlight{}, logger) sched := buildRestoreTestScheduler(cfg, px, nil, backup.NewStore(), rtState, &backup.InFlight{}, nil, logger)
fmt.Printf("eval_interval=%s settle=%s\n", cfg.Backup.RestoreTestEvalInterval(), cfg.Backup.RestoreTestSettle()) fmt.Printf("eval_interval=%s settle=%s\n", cfg.Backup.RestoreTestEvalInterval(), cfg.Backup.RestoreTestSettle())
start := time.Now() start := time.Now()
@@ -2667,7 +2728,6 @@ type escrowCeremonyOpts struct {
offline bool offline bool
upload bool upload bool
identityBundlePath string identityBundlePath string
directivePath string
} }
// escrowCeremonyOutcome is the shared core's result. R is the ONLY secret; Sum mirrors the // escrowCeremonyOutcome is the shared core's result. R is the ONLY secret; Sum mirrors the
@@ -2726,11 +2786,10 @@ func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger,
return out, &escrowCeremonyErr{kind: "setup", err: fmt.Errorf("PBS key for %q not found (%s): %v", storage, keyPath, err)} return out, &escrowCeremonyErr{kind: "setup", err: fmt.Errorf("PBS key for %q not found (%s): %v", storage, keyPath, err)}
} }
// Slice 10D.1: optionally ALSO wrap the identity bundle under the same R, and carry the non-secret // Slice 10D.1: optionally ALSO wrap the identity bundle under the same R. The bundle file is a 0600 secret
// directive for the hub. The bundle file is a 0600 secret (tunnel/pbs tokens); the directive is // (tunnel/pbs tokens). The non-secret "directive" that used to ride along is retired (R-105, `09` §3 decision 169:
// non-secret (pbs repo/ns, expected fingerprint, tunnel id). // nothing read it; the DR path reads the recipe, tenantsync and the escrow blob).
var identity *escrow.IdentityBundle var identity *escrow.IdentityBundle
var directive json.RawMessage
if opts.identityBundlePath != "" { if opts.identityBundlePath != "" {
raw, err := os.ReadFile(opts.identityBundlePath) raw, err := os.ReadFile(opts.identityBundlePath)
if err != nil { if err != nil {
@@ -2741,11 +2800,6 @@ func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger,
return out, &escrowCeremonyErr{kind: "setup", err: fmt.Errorf("identity bundle is not valid JSON {tunnel_token,pbs_token}: %v", err)} return out, &escrowCeremonyErr{kind: "setup", err: fmt.Errorf("identity bundle is not valid JSON {tunnel_token,pbs_token}: %v", err)}
} }
identity = &b identity = &b
if opts.directivePath != "" {
if d, err := os.ReadFile(opts.directivePath); err == nil && json.Valid(d) {
directive = d
}
}
} }
// S3: auto-inject the offsite WG private key into the escrowed identity when the key file // S3: auto-inject the offsite WG private key into the escrowed identity when the key file
// exists — a NEW escrow run should always capture the live tunnel identity. Field NAME only // exists — a NEW escrow run should always capture the live tunnel identity. Field NAME only
@@ -2824,7 +2878,7 @@ func escrowCeremony(ctx context.Context, cfg config.Config, logger *slog.Logger,
ResticPwSealed: resticStaged, ResticPwSealed: resticStaged,
} }
if opts.upload { if opts.upload {
if err := uploadEscrowBlob(ctx, cfg, res, directive, resticPwSHA256); err != nil { if err := uploadEscrowBlob(ctx, cfg, res, resticPwSHA256); err != nil {
// R is minted and the blob self-verified — only the hub leg failed. kind "upload" lets // R is minted and the blob self-verified — only the hub leg failed. kind "upload" lets
// the text shell keep the pre-extraction order (R surfaced, THEN the failure). // the text shell keep the pre-extraction order (R surfaced, THEN the failure).
return out, &escrowCeremonyErr{kind: "upload", err: err} return out, &escrowCeremonyErr{kind: "upload", err: err}
@@ -2881,7 +2935,7 @@ func printEscrowTextRBlock(out *escrowCeremonyOutcome) {
// NOTHING else there; every human/info line goes to stderr; failures exit non-zero with no // NOTHING else there; every human/info line goes to stderr; failures exit non-zero with no
// partial JSON. This is the controller-driven ceremony's parse surface (spike §2.3: the text // partial JSON. This is the controller-driven ceremony's parse surface (spike §2.3: the text
// banner is positionally brittle). // banner is positionally brittle).
func runSelftestEscrowCreate(ctx context.Context, cfg config.Config, logger *slog.Logger, storage string, paperkey, offline, upload bool, identityBundlePath, directivePath, outputMode string) int { func runSelftestEscrowCreate(ctx context.Context, cfg config.Config, logger *slog.Logger, storage string, paperkey, offline, upload bool, identityBundlePath, outputMode string) int {
switch outputMode { switch outputMode {
case "", "text", "json": case "", "text", "json":
default: default:
@@ -2899,7 +2953,7 @@ func runSelftestEscrowCreate(ctx context.Context, cfg config.Config, logger *slo
out, cerr := escrowCeremony(ctx, cfg, logger, escrowCeremonyOpts{ out, cerr := escrowCeremony(ctx, cfg, logger, escrowCeremonyOpts{
storage: storage, paperkey: paperkey, offline: offline, upload: upload, storage: storage, paperkey: paperkey, offline: offline, upload: upload,
identityBundlePath: identityBundlePath, directivePath: directivePath, identityBundlePath: identityBundlePath,
}) })
if cerr != nil { if cerr != nil {
switch cerr.kind { switch cerr.kind {
@@ -3082,9 +3136,8 @@ type escrowUploadRequest struct {
BlobB64 string `json:"blob_b64"` // base64 of the opaque R-wrapped blob (ciphertext) BlobB64 string `json:"blob_b64"` // base64 of the opaque R-wrapped blob (ciphertext)
KeyFingerprint string `json:"key_fingerprint"` // for operator display only KeyFingerprint string `json:"key_fingerprint"` // for operator display only
Posture string `json:"posture"` // e.g. "zero_knowledge" Posture string `json:"posture"` // e.g. "zero_knowledge"
// Slice 10D.1 — optional DR bundle (identity escrow + non-secret directive). Omitted in slice-7. // Slice 10D.1 — optional identity escrow. Omitted in slice-7. (The `directive` is retired — R-105.)
IdentityBlobB64 string `json:"identity_blob_b64,omitempty"` IdentityBlobB64 string `json:"identity_blob_b64,omitempty"`
DirectiveJSON json.RawMessage `json:"directive,omitempty"`
CreatedAt string `json:"created_at"` // RFC3339 CreatedAt string `json:"created_at"` // RFC3339
// SLICE 3 — sha256 hex of the offsite restic repo password sealed in the identity blob (present only // SLICE 3 — sha256 hex of the offsite restic repo password sealed in the identity blob (present only
// when a staged password was folded in). Non-reversible hash of a 256-bit random secret — safe to // when a staged password was folded in). Non-reversible hash of a 256-bit random secret — safe to
@@ -3092,10 +3145,10 @@ type escrowUploadRequest struct {
ResticPwSHA256 string `json:"restic_pw_sha256,omitempty"` ResticPwSHA256 string `json:"restic_pw_sha256,omitempty"`
} }
// uploadEscrowBlob PUTs the opaque blob (and, for 10D, the identity blob + non-secret directive) to // uploadEscrowBlob PUTs the opaque blob (and, for 10D, the identity blob) to
// the hub, authed with the per-host key. The hub stores ciphertext + non-secret fields; no usable // the hub, authed with the per-host key. The hub stores ciphertext + non-secret fields; no usable
// secret leaves the agent. // secret leaves the agent.
func uploadEscrowBlob(ctx context.Context, cfg config.Config, res escrow.CreateResult, directive json.RawMessage, resticPwSHA256 string) error { func uploadEscrowBlob(ctx context.Context, cfg config.Config, res escrow.CreateResult, resticPwSHA256 string) error {
if cfg.Hub.URL == "" || cfg.Hub.HostID == "" || cfg.Hub.APIKey == "" { if cfg.Hub.URL == "" || cfg.Hub.HostID == "" || cfg.Hub.APIKey == "" {
return fmt.Errorf("hub not configured (url/host_id/api_key)") return fmt.Errorf("hub not configured (url/host_id/api_key)")
} }
@@ -3108,7 +3161,6 @@ func uploadEscrowBlob(ctx context.Context, cfg config.Config, res escrow.CreateR
} }
if len(res.IdentityBlob) > 0 { if len(res.IdentityBlob) > 0 {
upReq.IdentityBlobB64 = base64.StdEncoding.EncodeToString(res.IdentityBlob) upReq.IdentityBlobB64 = base64.StdEncoding.EncodeToString(res.IdentityBlob)
upReq.DirectiveJSON = directive
} }
body, _ := json.Marshal(upReq) body, _ := json.Marshal(upReq)
url := strings.TrimRight(cfg.Hub.URL, "/") + "/api/v1/hosts/" + cfg.Hub.HostID + "/escrow" url := strings.TrimRight(cfg.Hub.URL, "/") + "/api/v1/hosts/" + cfg.Hub.HostID + "/escrow"
+105
View File
@@ -0,0 +1,105 @@
package main
import (
"go/ast"
"testing"
)
// R-894 — the on-disk backup record is WIRED on the daemon path (the built-but-never-wired class).
// main → runDaemon → buildLocalAPIServer, and inside it the localapi.Options literal carries
// LastKnownBackups built by backup.NewBackupSuccessState. An AST walk, not a string match, for the
// reasons in escrow_recover_wiring_test.go.
//
// COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this
// fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored.
//
// 2026-10-09: the SAME instance also feeds the host report (collector.SetKnownBackupReporter), so a
// restart right after a backup no longer erases the hub's evidence of it. The field may name a local
// variable; the variable must be built by backup.NewBackupSuccessState and be the one handed to the
// collector. RED-PROOF (observed): drop the SetKnownBackupReporter call → "not handed to the collector".
func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
_, f := parseMain(t)
if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one")
}
var field, built bool
var fieldVar, handed string
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
cl, ok := n.(*ast.CompositeLit)
if !ok {
return true
}
sel, ok := cl.Type.(*ast.SelectorExpr)
if !ok {
return true
}
if pkg, _ := sel.X.(*ast.Ident); pkg == nil || pkg.Name+"."+sel.Sel.Name != "localapi.Options" {
return true
}
for _, el := range cl.Elts {
kv, ok := el.(*ast.KeyValueExpr)
if !ok {
continue
}
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "LastKnownBackups" {
field = true
if callsIn(kv.Value)["backup.NewBackupSuccessState"] {
built = true
}
if id, ok := kv.Value.(*ast.Ident); ok {
fieldVar = id.Name
}
}
}
return true
})
// a local variable: built by NewBackupSuccessState, and handed to the collector
ast.Inspect(fd.Body, func(n ast.Node) bool {
switch x := n.(type) {
case *ast.AssignStmt:
for i, l := range x.Lhs {
if id, ok := l.(*ast.Ident); ok && fieldVar != "" && id.Name == fieldVar && i < len(x.Rhs) &&
callsIn(x.Rhs[i])["backup.NewBackupSuccessState"] {
built = true
}
}
case *ast.CallExpr:
if fn, ok := x.Fun.(*ast.SelectorExpr); ok && fn.Sel.Name == "SetKnownBackupReporter" && len(x.Args) == 1 {
if id, ok := x.Args[0].(*ast.Ident); ok {
handed = id.Name
}
}
}
return true
})
}
if !field {
t.Fatal("localapi.Options in buildLocalAPIServer has no LastKnownBackups field")
}
if !built {
t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState")
}
if handed == "" || handed != fieldVar {
t.Fatalf("the saved backups (%q) are not handed to the collector (SetKnownBackupReporter got %q)", fieldVar, handed)
}
}
func callsIn(n ast.Node) map[string]bool {
out := map[string]bool{}
ast.Inspect(n, func(n ast.Node) bool {
if ce, ok := n.(*ast.CallExpr); ok {
if fn, ok := ce.Fun.(*ast.SelectorExpr); ok {
if x, ok := fn.X.(*ast.Ident); ok {
out[x.Name+"."+fn.Sel.Name] = true
}
}
}
return true
})
return out
}
+1 -1
View File
@@ -43,7 +43,7 @@ func main() {
func run() error { func run() error {
var ( var (
op = flag.String("op", "", "op class to sign, e.g. storage_wipe | guest_destroy | decommission | agent_update | os_docker_step | agent_config_update") op = flag.String("op", "", "op class to sign, e.g. storage_wipe | guest_destroy | decommission | agent_update | os_docker_step | os_pve_step | os_kernel_step | agent_config_update")
host = flag.String("host", "", "target host_id (anti-retarget — the op runs ONLY on this host)") host = flag.String("host", "", "target host_id (anti-retarget — the op runs ONLY on this host)")
guest = flag.String("guest", "", "target guest_id (\"\" = host-scoped op)") guest = flag.String("guest", "", "target guest_id (\"\" = host-scoped op)")
keyID = flag.String("key-id", "", "key id of the signing key (must match a pinned agent signer)") keyID = flag.String("key-id", "", "key id of the signing key (must match a pinned agent signer)")
+14 -4
View File
@@ -111,15 +111,17 @@ Cmnd_Alias FELHOM_INTERMEDIARY = \
# docker inspect -f * — container running/health/image (read-only; `*` spans the -f template # docker inspect -f * — container running/health/image (read-only; `*` spans the -f template
# + container across spaces, spike-confirmed) # + container across spaces, spike-confirmed)
# systemctl restart <fixed unit> — re-run the golden's bootstrap (the only state change) # systemctl restart <fixed unit> — re-run the golden's bootstrap (the only state change)
# tee <FIXED image file> — WRITE the ref; content is fed on STDIN (no shell, no interpolation), # felhom-priv-apply controller-image <vmid> — WRITE the ref (R-861 (a) A1, `09` §3 decision 165): the ref goes on
# the agent strict-validates the ref (controllerImageRe) before the write. # STDIN to the ROOT wrapper, which requires our registry + repository + an x.y.z tag
# and writes the guest file itself. The agent's own `tee` grant is GONE: before, a
# compromised agent could hand the guest's bootstrap ANY image (sudo cannot see stdin).
# Validated GO: felhom.eu/documentation/audits/SPIKE-controllerswap-narrow-grants-2026-06-29.md. # Validated GO: felhom.eu/documentation/audits/SPIKE-controllerswap-narrow-grants-2026-06-29.md.
Cmnd_Alias FELHOM_CONTROLLERSWAP = \ Cmnd_Alias FELHOM_CONTROLLERSWAP = \
/usr/sbin/pct ^exec [0-9]+ -- cat /etc/felhom-controller-image$, \ /usr/sbin/pct ^exec [0-9]+ -- cat /etc/felhom-controller-image$, \
/usr/sbin/pct ^exec [0-9]+ -- docker image inspect gitea\.dooplex\.hu/admin/felhom-controller\:[0-9]+\.[0-9]+\.[0-9]+$, \ /usr/sbin/pct ^exec [0-9]+ -- docker image inspect gitea\.dooplex\.hu/admin/felhom-controller\:[0-9]+\.[0-9]+\.[0-9]+$, \
/usr/sbin/pct ^exec [0-9]+ -- docker inspect -f .+ (felhom-controller|cloudflared)$, \ /usr/sbin/pct ^exec [0-9]+ -- docker inspect -f .+ (felhom-controller|cloudflared)$, \
/usr/sbin/pct ^exec [0-9]+ -- systemctl restart felhom-controller-bootstrap\.service$, \ /usr/sbin/pct ^exec [0-9]+ -- systemctl restart felhom-controller-bootstrap\.service$, \
/usr/sbin/pct ^exec [0-9]+ -- tee /etc/felhom-controller-image$ /usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$
# Stale-lock recovery (F2-b, v0.49.0). A host reboot DURING a vzdump backup leaves the guest with a # Stale-lock recovery (F2-b, v0.49.0). A host reboot DURING a vzdump backup leaves the guest with a
# `snapshot-delete`/`backup` lock + `onboot:1` then can't start it → the customer box stays DOWN. The # `snapshot-delete`/`backup` lock + `onboot:1` then can't start it → the customer box stays DOWN. The
@@ -129,6 +131,14 @@ Cmnd_Alias FELHOM_CONTROLLERSWAP = \
Cmnd_Alias FELHOM_STALELOCK = \ Cmnd_Alias FELHOM_STALELOCK = \
/usr/sbin/pct ^unlock [0-9]+$ /usr/sbin/pct ^unlock [0-9]+$
# Weekly guest disk trim (R-444, operator ruling `09` §3 decision 139). A thin pool only ever grows from blocks the
# guest has FREED: `fstrim` inside the unprivileged container is refused (FITRIM: Operation not permitted), so the host
# trims the guest's mounts. Measured on demo-hp 2026-10-06: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %,
# apps kept answering. ONE exact pattern: a vmid and nothing else — no `--ignore-mountpoints`, no second argument
# (pinned: TestSudoersFstrimRuleIsExact). The agent runs it on a weekly daytime timer under the heavy-op gate.
Cmnd_Alias FELHOM_FSTRIM = \
/usr/sbin/pct ^fstrim [0-9]+$
# Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS # Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS
# leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and # leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and
# a guest joins that pool only when its restore COMPLETES, so a failed restore leaves a pool-less guest # a guest joins that pool only when its restore COMPLETES, so a failed restore leaves a pool-less guest
@@ -304,4 +314,4 @@ Cmnd_Alias FELHOM_GUESTNET = \
/usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \ /usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \
/usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$ /usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_FSTRIM, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
+39
View File
@@ -0,0 +1,39 @@
#!/bin/sh
# /etc/grub.d/42_felhom_oneshot — the kernel lane's one-shot ENTRIES (R-836, `09` §3 decision 172, `11` §5.11).
# Installed by the config bundle (felhom-os-apply BUNDLE_FILES), 0755 root. update-grub runs it.
#
# One menu entry per installed Proxmox kernel, id `felhom-oneshot-<version>`, booted ONLY when 01_felhom_oneshot found
# the flag naming it. It is the normal entry plus option C (decision 172): softlockup_panic=1 hardlockup_panic=1
# hung_task_panic=1 panic=10 — a lockup the kernel can detect becomes a panic, and a panic restarts the box in 10 s into
# the default (the old kernel). A true dead freeze still needs a person (spike candidate 3 failed on all three boxes).
# It sorts AFTER 10_linux, so it is never entry 0 and never the default.
#
# No vfat ESP at /boot/efi → prints nothing (no flag can name these entries).
set -e
prefix="/usr"
exec_prefix="/usr"
datarootdir="/usr/share"
. "$datarootdir/grub/grub-mkconfig_lib"
esp_uuid=$(findmnt -n -o UUID,FSTYPE /boot/efi 2>/dev/null | awk '$2 == "vfat" { print $1 }')
[ -n "$esp_uuid" ] || exit 0
kernels=$(ls /boot/vmlinuz-*-pve 2>/dev/null | sed 's#^/boot/vmlinuz-##' | grep -E '^[0-9]+\.[0-9]+\.[0-9]+-[0-9]+-pve$' || true)
[ -n "$kernels" ] || exit 0
case "${GRUB_DEVICE}" in
/dev/mapper/*|/dev/dm-*|"") root_arg="root=${GRUB_DEVICE}" ;;
*) if [ -n "${GRUB_DEVICE_UUID}" ]; then root_arg="root=UUID=${GRUB_DEVICE_UUID}"; else root_arg="root=${GRUB_DEVICE}"; fi ;;
esac
[ -n "${GRUB_DEVICE}" ] || root_arg="root=$(findmnt -n -o SOURCE /)"
rel=$(make_system_path_relative_to_its_root /boot)
prep=$(prepare_grub_to_access_device "$(${grub_probe:-grub-probe} --target=device /boot)" | sed 's/^/ /')
for k in $kernels; do
[ -f "/boot/initrd.img-$k" ] || continue
cat <<EOF
menuentry 'Felhom one-shot: $k' --class proxmox --id felhom-oneshot-$k {
insmod gzio
$prep
echo 'Loading Linux $k (felhom one-shot) ...'
linux $rel/vmlinuz-$k $root_arg ro ${GRUB_CMDLINE_LINUX} ${GRUB_CMDLINE_LINUX_DEFAULT} softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10
initrd $rel/initrd.img-$k
}
EOF
done
+36
View File
@@ -0,0 +1,36 @@
#!/bin/sh
# /etc/grub.d/01_felhom_oneshot — the kernel lane's ONE-SHOT boot (R-836, `09` §3 decisions 164 + 172, `11` §5.11).
# Installed by the config bundle (felhom-os-apply BUNDLE_FILES), 0755 root. update-grub runs it; it prints GRUB script.
#
# At boot, GRUB reads `felhom_next` from an environment block on the ESP (vfat — GRUB can rewrite a file there; it
# cannot on the LVM /boot, R-836), CLEARS it, and — only if it names an installed kernel — boots that kernel's one-shot
# entry (42_felhom_oneshot) instead of the default. The next boot uses the default again whatever happens: a new kernel
# that panics comes back on the old one by itself. felhom-os-apply writes the flag (mode apply) and never the default
# for a new kernel. Measured on the Tester 1 VM, demo-felhom and demo-hp (Secure Boot on):
# `audits/kernel-spike-2026-10-07/` (candidate 2).
#
# No vfat ESP at /boot/efi, or no Proxmox kernel → prints nothing (the box boots exactly as before).
set -e
esp_uuid=$(findmnt -n -o UUID,FSTYPE /boot/efi 2>/dev/null | awk '$2 == "vfat" { print $1 }')
[ -n "$esp_uuid" ] || exit 0
kernels=$(ls /boot/vmlinuz-*-pve 2>/dev/null | sed 's#^/boot/vmlinuz-##' | grep -E '^[0-9]+\.[0-9]+\.[0-9]+-[0-9]+-pve$' || true)
[ -n "$kernels" ] || exit 0
cat <<EOF
# felhom kernel lane: a one-shot kernel named on the ESP, read and cleared before the menu
insmod part_gpt
insmod fat
search --no-floppy --fs-uuid --set=felhom_esp $esp_uuid
if [ -f (\$felhom_esp)/EFI/felhom/oneshot.env ]; then
load_env -f (\$felhom_esp)/EFI/felhom/oneshot.env felhom_next
if [ "\${felhom_next}" ]; then
set felhom_boot="\${felhom_next}"
set felhom_next=
save_env -f (\$felhom_esp)/EFI/felhom/oneshot.env felhom_next
EOF
for k in $kernels; do
printf ' if [ "${felhom_boot}" = "%s" ]; then set default="felhom-oneshot-%s"; fi\n' "$k" "$k"
done
cat <<EOF
fi
fi
EOF
+5 -3
View File
@@ -16,8 +16,10 @@ Cmnd_Alias FELHOM_OP_REPAIR = \
/usr/bin/systemctl reset-failed felhom-sshd, \ /usr/bin/systemctl reset-failed felhom-sshd, \
/usr/bin/systemctl restart felhom-sshd, \ /usr/bin/systemctl restart felhom-sshd, \
/usr/sbin/pct list, \ /usr/sbin/pct list, \
/usr/sbin/pct start [0-9]*, \ /usr/sbin/pct ^start [0-9]+$, \
/usr/sbin/pct stop [0-9]*, \ /usr/sbin/pct ^stop [0-9]+$, \
/usr/sbin/pct unlock [0-9]* /usr/sbin/pct ^unlock [0-9]+$
# R-861 (b) B2 (`09` §3 decision 165, hygiene): one numeric vmid per pct verb, anchored — the old glob `[0-9]*` also
# matched spaces, so `pct stop 9201 --skiplock 1` passed. Pinned by TestFelhomOpSudoersPctIsExact.
felhom-op ALL=(root) NOPASSWD: FELHOM_OP_REPAIR felhom-op ALL=(root) NOPASSWD: FELHOM_OP_REPAIR
+657 -19
View File
@@ -21,6 +21,11 @@
# or, for an unsigned ring-0 step, the root-owned TRUST_FILE saying `"ring0_slow_lane": true` (set by hand on the demo # or, for an unsigned ring-0 step, the root-owned TRUST_FILE saying `"ring0_slow_lane": true` (set by hand on the demo
# boxes only). The agent's own config is NOT trusted for either: the agent can write it. A Docker step also needs # boxes only). The agent's own config is NOT trusted for either: the agent can write it. A Docker step also needs
# `live-restore` ON (R15) — without it every container restarts. # `live-restore` ON (R15) — without it every container restarts.
# Agent v0.151.0 (R-812 option A, `09` §3 decision 163) adds the layer "pve": the HOST's Proxmox USERSPACE packages,
# lane "slow" only, origin "Proxmox Debian Repository" only, never a kernel / boot / firmware / microcode name (R14 —
# the kernel is R-836's lane), no removal, no undo, a NEW package only from PVE_NEW_ALLOW, an appliance only (R12), and
# the same authority as the Docker step (R3: a signed `os_pve_step`, or the root-owned ring-0 mark). Select
# "pending-pve" (ring 0): every installed Proxmox-origin package with a pending upgrade, minus HOST_SLOW_RE.
# #
# Modes (plan field "mode"): # Modes (plan field "mode"):
# inventory `apt-get update`, then report what is installed (with origin), what is pending, and health. # inventory `apt-get update`, then report what is installed (with origin), what is pending, and health.
@@ -32,6 +37,14 @@
# held packages, kernel taint, the crash guard; guest Debian, Docker engine, containerd, live-restore. # held packages, kernel taint, the crash guard; guest Debian, Docker engine, containerd, live-restore.
# live-restore-on (v0.142.0, layer guest) the ONE-TIME step of `09` decision 87: merge `"live-restore": true` # live-restore-on (v0.142.0, layer guest) the ONE-TIME step of `09` decision 87: merge `"live-restore": true`
# into the guest's /etc/docker/daemon.json and `systemctl reload docker`. NEVER a restart (R-835). # into the guest's /etc/docker/daemon.json and `systemctl reload docker`. NEVER a restart (R-835).
# oom-check (R-528, `09` decision 157; layer docker, lane slow) ONLY the memory-kill check below: no apt, no engine
# change, no authority needed. Wrapper-only — the agent never writes this plan; run it by hand as root.
# R-528 (`09` decision 157): a docker-layer APPLY also runs `oom_check()` after health_after and reports it as
# "oom_check": {"result": "pass"|"fail"|"error", "oom_killed": bool, "oom_event": bool, "exit_code": int|null,
# "image": str|null, "detail": str}
# — one throwaway container (the controller's own image, no network / volume / port, 64 MB cap) is made to exceed its
# memory; "pass" only when the engine says OOMKilled=true AND emits the `oom` event. It never changes the step's
# outcome or health: the hub decides whether the engine set can be approved. Pinned by the OOMCheck tests.
# Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is # Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is
# OSAPPLY-REPORT <one JSON object> # OSAPPLY-REPORT <one JSON object>
# which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install. # which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install.
@@ -68,10 +81,25 @@ JOURNAL_MARK = "@@FELHOM-DPKG-JOURNAL@@"
DPKG_STATE_SCRIPT = "dpkg --audit; echo " + JOURNAL_MARK + "; ls -A /var/lib/dpkg/updates 2>/dev/null; true" DPKG_STATE_SCRIPT = "dpkg --audit; echo " + JOURNAL_MARK + "; ls -A /var/lib/dpkg/updates 2>/dev/null; true"
# The installer's ROOT-OWNED record (felhom-host-install.sh `state_set mode`); the agent cannot write it. # The installer's ROOT-OWNED record (felhom-host-install.sh `state_set mode`); the agent cannot write it.
INSTALL_STATE = "/var/lib/felhom-install/state.json" INSTALL_STATE = "/var/lib/felhom-install/state.json"
# R-528: the memory-kill check (oom_check). One 200 MB block under a 64 MB cap: measured on Docker 29.8.2 to be
# OOM-killed with OOMKilled=true and an `oom` event. Every call is bounded: timeouts (the clock read twice) + the
# settle wait stay within 90 s.
OOMCHECK_PREFIX = "felhom-oomcheck-"
OOMCHECK_SCRIPT = "dd if=/dev/zero of=/dev/null bs=200M count=1"
OOMCHECK_TIMEOUTS = {"image": 10, "clock": 5, "run": 30, "inspect": 10, "events": 10, "rm": 15}
# Measured on demo-hp 9201 (Docker 29.8.2, 2026-10-06, audits/readback-2026-10-07/F/F1, F2): an `--until` taken right
# after the run MISSED the oom event although OOMKilled=true; after a 2 s wait and `--until` = guest epoch + 1 it is
# seen. Pinned by test_events_window_ends_after_the_settle_wait.
OOMCHECK_SETTLE = 2
# Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host # Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host
# reboot is needed for them to take effect, and a bad one can stop the box from booting. # reboot is needed for them to take effect, and a bad one can stop the box from booting.
# proxmox-secure-boot-support is the Secure Boot meta-package: its only job is to pull shim-signed and the signed GRUB,
# so it belongs with them. Measured 2026-10-09 on demo-hp (Secure Boot on): left in the pve lane, its upgrade pulled
# shim-signed-common, the whole pve step was refused R6, and the night's kernel step was skipped with it
# (`audits/kernel-night-2026-10-08/`). Pinned by test_ring0_pending_pve_leaves_the_secure_boot_meta_out.
HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|" HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|"
r"pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr)") r"pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr|"
r"proxmox-secure-boot-support)")
# restart_needed() leaves out processes whose cgroup line matches (grep basic regex). Host: the LXC guests' own # restart_needed() leaves out processes whose cgroup line matches (grep basic regex). Host: the LXC guests' own
# processes (`0::/lxc/<vmid>/...`) -- NOT lxc-start itself, whose cgroup is `0::/lxc.monitor/<vmid>` (measured # processes (`0::/lxc/<vmid>/...`) -- NOT lxc-start itself, whose cgroup is `0::/lxc.monitor/<vmid>` (measured
# 2026-10-04 on demo-felhom: the old pattern "lxc" hid lxc-start with 20 deleted maps, so "reboot needed" stayed false # 2026-10-04 on demo-felhom: the old pattern "lxc" hid lxc-start with 20 deleted maps, so "reboot needed" stayed false
@@ -82,6 +110,12 @@ HOST_SERVICES = ["pveproxy", "pvedaemon", "pvestatd", "pve-cluster", "felhom-age
DOCKER_NAMES = ("containerd.io", "docker-buildx-plugin", "docker-ce", "docker-ce-cli", "docker-ce-rootless-extras", DOCKER_NAMES = ("containerd.io", "docker-buildx-plugin", "docker-ce", "docker-ce-cli", "docker-ce-rootless-extras",
"docker-compose-plugin") "docker-compose-plugin")
DOCKER_ORIGIN = "Docker CE" DOCKER_ORIGIN = "Docker CE"
# The Proxmox package lane (R-812 option A): the origin apt prints for download.proxmox.com, and the ONLY new packages
# a pve step may add (measured on demo-felhom 2026-10-07: a full upgrade adds proxmox-firewall-data and the kernel;
# the kernel is refused by R14 whatever this list says).
PVE_ORIGIN = "Proxmox Debian Repository"
PVE_NEW_ALLOW = ("proxmox-firewall-data",)
PVE_SIGNED_OP = "os_pve_step"
# ROOT-OWNED trust anchors (the installer writes them; the demo boxes got them by hand, R-840). Never the agent's config. # ROOT-OWNED trust anchors (the installer writes them; the demo boxes got them by hand, R-840). Never the agent's config.
TRUST_FILE = "/etc/felhom/os-trust.json" # {"host_id": "...", "ring0_slow_lane": false} TRUST_FILE = "/etc/felhom/os-trust.json" # {"host_id": "...", "ring0_slow_lane": false}
TRUST_SIGNERS = "/etc/felhom/operator-signers" # ssh allowed_signers: <key_id> namespaces="felhom-op-v1" <key> TRUST_SIGNERS = "/etc/felhom/operator-signers" # ssh allowed_signers: <key_id> namespaces="felhom-op-v1" <key>
@@ -95,6 +129,32 @@ DAEMON_JSON = "/etc/docker/daemon.json"
# wrapper restarts exactly the containers that mount one of these paths (never the apps, never the engine). # wrapper restarts exactly the containers that mount one of these paths (never the apps, never the engine).
DOCKER_SOCKETS = ("/var/run/docker.sock", "/run/docker.sock") DOCKER_SOCKETS = ("/var/run/docker.sock", "/run/docker.sock")
CRASH_GUARD_STATE = "/var/lib/felhom-crash-guard/state.json" CRASH_GUARD_STATE = "/var/lib/felhom-crash-guard/state.json"
# ---------- the kernel lane (R-836, `09` §3 decisions 164 + 172, `11` §5.11) ----------
# A new kernel boots ONCE through a flag in a GRUB environment block on the ESP (vfat — GRUB can rewrite it there; on
# the LVM /boot it cannot, R-836). The bundle's two GRUB generators read and clear the flag (01_felhom_oneshot) and
# give each installed kernel a one-shot entry with the lockup-to-panic options (42_felhom_oneshot, option C). The GRUB
# default is the kernel the box RUNS, pinned in KERNEL_DEFAULT_CFG; only "kernel-good" (after a healthy one-shot boot)
# moves it to the new kernel. Measured in the spike on the Tester 1 VM, demo-felhom and demo-hp (Secure Boot on):
# `audits/kernel-spike-2026-10-07/`.
KERNEL_OP = "os_kernel_step"
KERNEL_STATE = "/var/lib/felhom-kernel/state.json" # 0644 root: the step's phase (the agent reads it)
KERNEL_DEFAULT_CFG = "/etc/default/grub.d/zz-felhom-kernel-default.cfg" # sourced last: GRUB_DEFAULT = this kernel
ONESHOT_SNIPPET = "/etc/grub.d/01_felhom_oneshot" # bundle-owned: read + clear the flag, pick the entry
ONESHOT_ENTRIES = "/etc/grub.d/42_felhom_oneshot" # bundle-owned: the one-shot entries (option C options)
GRUB_CFG = "/boot/grub/grub.cfg"
ESP_MOUNT = "/boot/efi"
ONESHOT_ENV = ESP_MOUNT + "/EFI/felhom/oneshot.env"
ONESHOT_ARGS = "softlockup_panic=1 hardlockup_panic=1 hung_task_panic=1 panic=10"
KERNEL_PIN_FILE = "/etc/kernel/proxmox-boot-pin" # an operator's `proxmox-boot-tool kernel pin` — never fought
KVER_RE = re.compile(r"^[0-9]+\.[0-9]+\.[0-9]+-[0-9]+-pve$")
KERNEL_IMAGE_RE = re.compile(r"^proxmox-kernel-([0-9]+\.[0-9]+\.[0-9]+-[0-9]+-pve)(-signed)?$")
# the packages a kernel step may UPGRADE (never add): the kernel series meta-package, the default-kernel meta, the boot
# helper and the firmware the kernel loads. Everything else in HOST_SLOW_RE (grub, shim, microcode, efibootmgr) stays out.
KERNEL_UPGRADE_RE = re.compile(r"^(proxmox-default-kernel|proxmox-kernel-[0-9]+\.[0-9]+|proxmox-kernel-helper|pve-firmware)$")
KERNEL_MIN_GAP = 20 * 3600 # never two kernel steps in one night
KERNEL_MODES = ("kernel-status", "kernel-reboot", "kernel-boot", "kernel-good", "kernel-revert", "kernel-cancel")
# phases: staged → oneshot → judging → good | reverting → self_reverted | revert_failed; oneshot → fell_back; staged → cancelled
KERNEL_ACTIVE = ("staged", "oneshot", "judging", "reverting")
# ---------- the config bundle (R-840, agent v0.143.0, `11` §5.4.2) ---------- # ---------- the config bundle (R-840, agent v0.143.0, `11` §5.4.2) ----------
# A signed `agent_config_update` job carries {agent_version, bundle_sha256}; the bundle is ONE JSON file built from this # A signed `agent_config_update` job carries {agent_version, bundle_sha256}; the bundle is ONE JSON file built from this
@@ -141,6 +201,10 @@ BUNDLE_FILES = [
("/etc/systemd/system/felhom-crash-guard-check.service", "felhom-crash-guard-check.service", 0o644, "unit", "replace"), ("/etc/systemd/system/felhom-crash-guard-check.service", "felhom-crash-guard-check.service", 0o644, "unit", "replace"),
("/etc/systemd/system/felhom-crash-guard-check.timer", "felhom-crash-guard-check.timer", 0o644, "unit", "replace"), ("/etc/systemd/system/felhom-crash-guard-check.timer", "felhom-crash-guard-check.timer", 0o644, "unit", "replace"),
("/etc/felhom/crash-guard.conf", "crash-guard.conf", 0o644, "plain", "if-absent"), ("/etc/felhom/crash-guard.conf", "crash-guard.conf", 0o644, "plain", "if-absent"),
# the kernel lane's two GRUB generators (R-836, `11` §5.11): they take effect at the next update-grub, which the
# first kernel step runs itself; with no flag on the ESP they change nothing about how the box boots.
("/etc/grub.d/01_felhom_oneshot", "felhom-grub-oneshot.sh", 0o755, "sh", "replace"),
("/etc/grub.d/42_felhom_oneshot", "felhom-grub-oneshot-entries.sh", 0o755, "sh", "replace"),
("/etc/systemd/system/felhom-agent.service", "felhom-agent.service", 0o644, "agent-unit", "replace"), ("/etc/systemd/system/felhom-agent.service", "felhom-agent.service", 0o644, "agent-unit", "replace"),
("/etc/systemd/system/felhom-agent-rollback.service", "felhom-agent-rollback.service", 0o644, "unit", "replace"), ("/etc/systemd/system/felhom-agent-rollback.service", "felhom-agent-rollback.service", 0o644, "unit", "replace"),
("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", "felhom-agent-limits.conf", 0o644, "dropin", "replace"), ("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", "felhom-agent-limits.conf", 0o644, "dropin", "replace"),
@@ -405,7 +469,8 @@ class Apply:
def check_plan(self, plan): def check_plan(self, plan):
mode = plan.get("mode", "apply") mode = plan.get("mode", "apply")
if mode not in ("apply", "inventory", "health", "facts", "live-restore-on", "bundle", "agent_update"): if mode not in ("apply", "inventory", "health", "facts", "live-restore-on", "bundle", "agent_update", "oom-check") \
+ KERNEL_MODES:
raise Refused("R11", f"unknown mode {mode!r}") raise Refused("R11", f"unknown mode {mode!r}")
if mode == "agent_update": if mode == "agent_update":
if plan.get("layer") != "host": if plan.get("layer") != "host":
@@ -416,21 +481,31 @@ class Apply:
raise Refused("R11", "bundle is a host-layer mode") raise Refused("R11", "bundle is a host-layer mode")
return mode, "host", 0, "bundle" return mode, "host", 0, "bundle"
layer = plan.get("layer") layer = plan.get("layer")
if layer not in ("guest", "host", "docker"): if layer not in ("guest", "host", "docker", "pve", "kernel"):
raise Refused("R12", f"layer {layer!r} is not guest, host or docker") raise Refused("R12", f"layer {layer!r} is not guest, host, docker, pve or kernel")
if (mode in KERNEL_MODES) != (layer == "kernel" and mode not in ("apply", "health")):
raise Refused("R11", f"mode {mode!r} does not fit layer {layer!r}")
lane = plan.get("lane", "fast") lane = plan.get("lane", "fast")
if layer == "docker" and lane != "slow": if layer == "docker" and lane != "slow":
raise Refused("R3", "the Docker engine is the slow lane (`11` §5.8); a fast-lane Docker plan is refused") raise Refused("R3", "the Docker engine is the slow lane (`11` §5.8); a fast-lane Docker plan is refused")
if layer != "docker" and lane != "fast": if layer == "pve" and lane != "slow":
raise Refused("R3", "the Proxmox packages are the slow lane (`11` §5.10); a fast-lane pve plan is refused")
if layer == "kernel" and lane != "slow":
raise Refused("R3", "the kernel is the slow lane (`11` §5.11); a fast-lane kernel plan is refused")
if layer not in ("docker", "pve", "kernel") and lane != "fast":
raise Refused("R3", f"the {layer} layer has no slow lane in this release (kernel, Proxmox: `11` §8 step 6)") raise Refused("R3", f"the {layer} layer has no slow lane in this release (kernel, Proxmox: `11` §8 step 6)")
if mode == "facts" and layer != "host": if mode == "facts" and layer != "host":
raise Refused("R11", "facts is a host-layer mode (it reads the host and the guest)") raise Refused("R11", "facts is a host-layer mode (it reads the host and the guest)")
if mode == "live-restore-on" and layer != "guest": if mode == "live-restore-on" and layer != "guest":
raise Refused("R11", "live-restore-on is a guest-layer mode") raise Refused("R11", "live-restore-on is a guest-layer mode")
if plan.get("undo") and layer != "docker": if mode == "oom-check" and layer != "docker":
raise Refused("R11", "oom-check is a docker-layer mode (it checks the guest's Docker engine)")
if plan.get("undo") and layer != "docker": # the pve layer has no undo in this release (R-812 option A)
raise Refused("R5", "an undo (downgrade) exists only for the Docker layer, inside a signed job") raise Refused("R5", "an undo (downgrade) exists only for the Docker layer, inside a signed job")
vmid = plan.get("vmid") vmid = plan.get("vmid")
if not isinstance(vmid, int) or isinstance(vmid, bool) or vmid <= 0: if layer == "kernel" and mode in KERNEL_MODES and mode != "kernel-reboot" and vmid == 0:
pass # after a boot the guest may not run (that is what is judged); these modes never touch it
elif not isinstance(vmid, int) or isinstance(vmid, bool) or vmid <= 0:
raise Refused("R11", f"vmid must be a positive integer, got {vmid!r}") raise Refused("R11", f"vmid must be a positive integer, got {vmid!r}")
rid = plan.get("release_id", "") rid = plan.get("release_id", "")
if not isinstance(rid, str) or not re.match(r"^[A-Za-z0-9._:-]{1,80}$", rid): if not isinstance(rid, str) or not re.match(r"^[A-Za-z0-9._:-]{1,80}$", rid):
@@ -438,17 +513,24 @@ class Apply:
if plan.get("allow_new"): if plan.get("allow_new"):
raise Refused("R6", "allow_new is a slow-lane field; the fast lane never adds a package") raise Refused("R6", "allow_new is a slow-lane field; the fast lane never adds a package")
select = plan.get("select", "listed") select = plan.get("select", "listed")
if select not in ("listed", "pending-fast", "pending-docker"): if select not in ("listed", "pending-fast", "pending-docker", "pending-pve", "pending-kernel"):
raise Refused("R11", f"unknown select {select!r}") raise Refused("R11", f"unknown select {select!r}")
if (select == "pending-docker") != (layer == "docker" and select != "listed"): if (select == "pending-docker") != (layer == "docker" and select != "listed"):
if select == "pending-docker" or layer == "docker": if select == "pending-docker" or layer == "docker":
raise Refused("R11", f"select {select!r} does not fit layer {layer!r}") raise Refused("R11", f"select {select!r} does not fit layer {layer!r}")
if (select == "pending-pve") != (layer == "pve" and select != "listed"):
raise Refused("R11", f"select {select!r} does not fit layer {layer!r}")
if (select == "pending-kernel") != (layer == "kernel" and select != "listed"):
raise Refused("R11", f"select {select!r} does not fit layer {layer!r}")
ek = plan.get("expect_kver")
if ek is not None and (layer != "kernel" or not isinstance(ek, str) or not KVER_RE.match(ek)):
raise Refused("R11", f"expect_kver {ek!r} is not a kernel version of the kernel layer")
pk = plan.get("packages", []) pk = plan.get("packages", [])
if not isinstance(pk, list): if not isinstance(pk, list):
raise Refused("R11", "packages must be a list") raise Refused("R11", "packages must be a list")
if mode == "apply" and select == "listed" and not pk: if mode == "apply" and select == "listed" and not pk:
raise Refused("R11", "packages must be a non-empty list in apply mode (select listed)") raise Refused("R11", "packages must be a non-empty list in apply mode (select listed)")
if select in ("pending-fast", "pending-docker") and pk: if select in ("pending-fast", "pending-docker", "pending-pve", "pending-kernel") and pk:
raise Refused("R11", f"select {select} takes no package list") raise Refused("R11", f"select {select} takes no package list")
seen = set() seen = set()
for e in pk: for e in pk:
@@ -468,6 +550,19 @@ class Apply:
continue continue
if n in DOCKER_NAMES: if n in DOCKER_NAMES:
raise Refused("R2", f"{n} is a Docker package — the slow lane (`11` §5.8), never in a {layer} plan") raise Refused("R2", f"{n} is a Docker package — the slow lane (`11` §5.8), never in a {layer} plan")
if layer == "kernel":
if o != PVE_ORIGIN:
raise Refused("R2", f"{n}: origin {o!r} is not {PVE_ORIGIN!r} (the kernel layer)")
if not (KERNEL_UPGRADE_RE.match(n) or KERNEL_IMAGE_RE.match(n)):
raise Refused("R23", f"{n} is not a kernel package (a kernel image, the kernel meta-packages, "
f"proxmox-kernel-helper or pve-firmware)")
continue
if layer == "pve":
if o != PVE_ORIGIN:
raise Refused("R2", f"{n}: origin {o!r} is not {PVE_ORIGIN!r} (the pve layer)")
if HOST_SLOW_RE.match(n):
raise Refused("R14", f"{n} is a kernel / boot / firmware package — never the pve lane (R-836)")
continue
if o not in FAST_ORIGINS: if o not in FAST_ORIGINS:
raise Refused("R2", f"{n}: origin {o!r} is not Debian / Debian-Security (the fast lane, `11` C3)") raise Refused("R2", f"{n}: origin {o!r} is not Debian / Debian-Security (the fast lane, `11` C3)")
if layer == "host" and HOST_SLOW_RE.match(n): if layer == "host" and HOST_SLOW_RE.match(n):
@@ -610,12 +705,13 @@ class Apply:
now = self.r.now() now = self.r.now()
self.r.write_nonces({k: v for k, v in seen.items() if v > now}) self.r.write_nonces({k: v for k, v in seen.items() if v > now})
def docker_authority(self, plan): def docker_authority(self, plan, op_name=SIGNED_OP):
"""R3 for the docker layer: returns (who, undo). A signed job binds the EXACT package list and the undo flag.""" """R3 for the docker (and, op_name os_pve_step, the pve) layer: returns (who, undo). A signed job binds the EXACT
package list and the undo flag."""
trust = self.load_trust() trust = self.load_trust()
signed = plan.get("signed") signed = plan.get("signed")
if signed: if signed:
params = self.verify_signed(signed, trust) params = self.verify_signed(signed, trust, op_name=op_name)
want = sorted(f"{e.get('name')}={e.get('version')}" for e in params.get("packages") or []) want = sorted(f"{e.get('name')}={e.get('version')}" for e in params.get("packages") or [])
got = sorted(f"{e['name']}={e['version']}" for e in plan.get("packages", [])) got = sorted(f"{e['name']}={e['version']}" for e in plan.get("packages", []))
if not want or want != got: if not want or want != got:
@@ -629,7 +725,7 @@ class Apply:
raise Refused("R3", "an undo (downgrade) needs a signed operator job") raise Refused("R3", "an undo (downgrade) needs a signed operator job")
if trust.get("ring0_slow_lane") is True: if trust.get("ring0_slow_lane") is True:
return "ring0", False return "ring0", False
raise Refused("R3", "a Docker step needs a signed operator job (ring 1) or this box's root-owned ring-0 mark") raise Refused("R3", f"a {self.layer} slow-lane step needs a signed operator job (ring 1) or this box's root-owned ring-0 mark")
def live_restore(self): def live_restore(self):
rc, out, _ = self.g(["docker", "info", "--format", "{{.LiveRestoreEnabled}}"], timeout=60) rc, out, _ = self.g(["docker", "info", "--format", "{{.LiveRestoreEnabled}}"], timeout=60)
@@ -717,6 +813,17 @@ class Apply:
return v or "unknown" return v or "unknown"
h = {"debian": first(["cat", "/etc/debian_version"]), "kernel_running": first(["uname", "-r"])} h = {"debian": first(["cat", "/etc/debian_version"]), "kernel_running": first(["uname", "-r"])}
h["kernel_next_boot"], h["kernel_next_boot_source"] = self.kernel_next_boot() h["kernel_next_boot"], h["kernel_next_boot_source"] = self.kernel_next_boot()
try:
# the kernel lane (R-836): the default and the one-shot flag, read from grub.cfg and the ESP themselves
kv = Kernel(self, {}).view()
kv["setup_problems"] = Kernel(self, {}).setup_problems()
h["kernel_lane"] = kv
if kv.get("flag"):
h["kernel_next_boot"], h["kernel_next_boot_source"] = kv["flag"], "felhom one-shot flag (once; then the default)"
elif kv.get("default") not in (None, "unknown"):
h["kernel_next_boot"], h["kernel_next_boot_source"] = kv["default"], "grub.cfg default"
except Exception as e: # never cost the System page its other facts
h["kernel_lane"] = {"error": str(e)[:200]}
rc, out, _ = self.r.host(["apt-mark", "showhold"], 60) rc, out, _ = self.r.host(["apt-mark", "showhold"], 60)
h["held"] = sorted(out.split()) if rc == 0 else None h["held"] = sorted(out.split()) if rc == 0 else None
try: try:
@@ -775,8 +882,8 @@ class Apply:
# ---------- target helpers ---------- # ---------- target helpers ----------
def x(self, argv, timeout=1800): def x(self, argv, timeout=1800):
"""Run in the TARGET layer: the guest via pct exec, or the host directly.""" """Run in the TARGET layer: the guest via pct exec, or the host directly (host, pve and kernel)."""
if self.layer == "host": if self.layer in ("host", "pve", "kernel"):
return self.r.host(argv, timeout) return self.r.host(argv, timeout)
return self.r.guest(self.vmid, argv, timeout) # guest and docker both live in the customer guest return self.r.guest(self.vmid, argv, timeout) # guest and docker both live in the customer guest
@@ -871,7 +978,7 @@ class Apply:
def restart_needed(self): def restart_needed(self):
"""Processes still mapping deleted files, OUTSIDE containers (C11). Guest: outside docker; host: outside the """Processes still mapping deleted files, OUTSIDE containers (C11). Guest: outside docker; host: outside the
LXC guests (the host's /proc shows guest processes too).""" LXC guests (the host's /proc shows guest processes too)."""
skip = RESTART_SKIP_CGROUP["host" if self.layer == "host" else "guest"] skip = RESTART_SKIP_CGROUP["host" if self.layer in ("host", "pve", "kernel") else "guest"]
script = ('for p in /proc/[0-9]*; do grep -q "(deleted)" $p/maps 2>/dev/null || continue; ' script = ('for p in /proc/[0-9]*; do grep -q "(deleted)" $p/maps 2>/dev/null || continue; '
'grep -q "%s" $p/cgroup 2>/dev/null && continue; echo "${p#/proc/} $(cat $p/comm 2>/dev/null)"; done' % skip) 'grep -q "%s" $p/cgroup 2>/dev/null && continue; echo "${p#/proc/} $(cat $p/comm 2>/dev/null)"; done' % skip)
rc, out, _ = self.x(["sh", "-c", script], timeout=120) rc, out, _ = self.x(["sh", "-c", script], timeout=120)
@@ -945,13 +1052,24 @@ class Apply:
return Bundle(self).from_plan(plan) return Bundle(self).from_plan(plan)
if self.mode == "agent_update": if self.mode == "agent_update":
return self.agent_update(plan) return self.agent_update(plan)
if self.layer == "host": if self.layer in ("host", "pve", "kernel"):
self.check_appliance() self.check_appliance()
if self.layer == "kernel" and self.mode != "health":
# the kernel lane's own modes (R-836): most of them run while the guest is still starting after a boot, so
# the guest check is the stage's and the reboot's own (Kernel.run), not every mode's
return Kernel(self, plan).run()
self.check_guest(self.vmid) self.check_guest(self.vmid)
log = self.r.log log = self.r.log
if self.mode == "live-restore-on": if self.mode == "live-restore-on":
return self.live_restore_on() return self.live_restore_on()
if self.mode == "oom-check":
# R-528: the check alone — no apt, no engine change; check_guest above still applies.
self.report["oom_check"] = self.oom_check()
return 0
self.who, self.allow_downgrade = ("fast", False) self.who, self.allow_downgrade = ("fast", False)
if self.layer == "pve" and self.mode == "apply":
self.who, self.allow_downgrade = self.docker_authority(plan, op_name=PVE_SIGNED_OP)
self.report["authority"] = self.who
if self.layer == "docker" and self.mode == "apply": if self.layer == "docker" and self.mode == "apply":
self.who, self.allow_downgrade = self.docker_authority(plan) self.who, self.allow_downgrade = self.docker_authority(plan)
if self.live_restore() != "true": if self.live_restore() != "true":
@@ -964,7 +1082,7 @@ class Apply:
log(f"os-apply: START release={plan.get('release_id')} layer={self.layer}" + log(f"os-apply: START release={plan.get('release_id')} layer={self.layer}" +
(f":{self.vmid}" if self.layer != "host" else "") + (f":{self.vmid}" if self.layer != "host" else "") +
f" lane={plan.get('lane', 'fast')} mode={self.mode} select={self.select} packages={len(plan.get('packages', []))}" + f" lane={plan.get('lane', 'fast')} mode={self.mode} select={self.select} packages={len(plan.get('packages', []))}" +
(f" authority={self.who}{' UNDO' if self.allow_downgrade else ''}" if self.layer == "docker" else "")) (f" authority={self.who}{' UNDO' if self.allow_downgrade else ''}" if self.layer in ("docker", "pve") else ""))
if self.apt_lock_held(): if self.apt_lock_held():
raise Refused("R9", f"another apt/dpkg holds the lock on the {self.layer}") raise Refused("R9", f"another apt/dpkg holds the lock on the {self.layer}")
self.report["health_before"] = self.health() self.report["health_before"] = self.health()
@@ -988,10 +1106,91 @@ class Apply:
if self.layer == "docker": if self.layer == "docker":
rc_v, out_v, _ = self.g(["docker", "version", "--format", "{{.Server.Version}}"], timeout=60) rc_v, out_v, _ = self.g(["docker", "version", "--format", "{{.Server.Version}}"], timeout=60)
self.report["docker_engine"] = out_v.strip() if rc_v == 0 and out_v.strip() else "unknown" self.report["docker_engine"] = out_v.strip() if rc_v == 0 and out_v.strip() else "unknown"
if self.layer == "pve":
self.report["pve_manager"] = self.pve_manager()
self.report["reboot_scanned"] = "reboot_needed" in self.report self.report["reboot_scanned"] = "reboot_needed" in self.report
self.report["health_after"] = self.health() self.report["health_after"] = self.health()
if self.layer == "docker" and self.mode == "apply":
# R-528 (`09` decision 157): does the engine report a memory kill? Reported only — never the outcome.
self.report["oom_check"] = self.oom_check()
return 0 return 0
# ---------- R-528: the memory-kill check ----------
def oom_check(self):
"""Run one throwaway container over its memory cap in the guest and read what the engine says about it.
Never raises: any failure becomes result "error". The container is ALWAYS removed (finally), and a failed
removal is named in the detail."""
t = OOMCHECK_TIMEOUTS
res = {"result": "error", "oom_killed": False, "oom_event": False, "exit_code": None, "image": None, "detail": ""}
log = self.r.log
try:
rc, out, err = self.g(["docker", "inspect", "-f", "{{.Config.Image}}", "felhom-controller"], timeout=t["image"])
except Exception as e:
rc, out, err = -1, "", str(e)
img = out.strip() if rc == 0 else ""
if not img or any(c.isspace() for c in img):
res["detail"] = f"the controller's image could not be read (rc={rc}): {(err or out).strip()[:200]}"
log(f"os-apply: OOM-CHECK error — {res['detail']}")
return res
res["image"] = img
name = f"{OOMCHECK_PREFIX}{os.getpid()}-{os.urandom(4).hex()}"
notes = []
try:
t0 = self.guest_epoch()
if t0 is None:
raise RuntimeError("the guest clock could not be read")
rrc, rout, rerr = self.g(["docker", "run", "--name", name, "--pull", "never", "--network", "none",
"--memory", "64m", "--memory-swap", "64m", "--label", "felhom.oomcheck=1",
"--entrypoint", "sh", img, "-c", OOMCHECK_SCRIPT], timeout=t["run"])
self.r.sleep(OOMCHECK_SETTLE) # the engine publishes the oom event a moment after the run returns (F1/F2)
t1 = self.guest_epoch()
if t1 is None:
t1 = t0 + t["run"] + OOMCHECK_SETTLE + 1
irc, iout, ierr = self.g(["docker", "inspect", "-f", "{{.State.OOMKilled}} {{.State.ExitCode}}", name], timeout=t["inspect"])
if irc != 0:
raise RuntimeError(f"the check container could not be inspected (run rc={rrc}: {(rerr or rout).strip()[:120]}; "
f"inspect rc={irc}: {(ierr or iout).strip()[:120]})")
f = iout.split()
res["oom_killed"] = bool(f) and f[0] == "true"
try:
res["exit_code"] = int(f[1]) if len(f) > 1 else None
except ValueError:
res["exit_code"] = None
erc, eout, eerr = self.g(["docker", "events", "--since", str(t0 - 1), "--until", str(t1 + 1),
"--filter", f"container={name}", "--filter", "event=oom",
"--format", "{{.Action}}"], timeout=t["events"])
if erc != 0:
notes.append(f"the event read failed (rc={erc}): {(eerr or eout).strip()[:120]}")
res["oom_event"] = erc == 0 and any(l.strip() == "oom" for l in eout.splitlines())
if res["oom_killed"] and res["oom_event"]:
res["result"] = "pass"
notes.insert(0, "the engine reported the memory kill: OOMKilled=true and the oom event")
else:
res["result"] = "fail"
miss = [w for w, ok in (("OOMKilled=true", res["oom_killed"]), ("the oom event", res["oom_event"])) if not ok]
notes.insert(0, f"the engine did not report the memory kill: missing {' and '.join(miss)} (exit code {res['exit_code']})")
except Exception as e:
res["result"] = "error"
notes.insert(0, f"the check could not finish: {type(e).__name__}: {str(e)[:200]}")
finally:
try:
mrc, mout, merr = self.g(["docker", "rm", "-f", name], timeout=t["rm"])
if mrc != 0 and "no such container" not in (merr + mout).lower():
notes.append(f"the check container {name} could not be removed (rc={mrc}): {(merr or mout).strip()[:120]}")
except Exception as e:
notes.append(f"the check container {name} could not be removed: {type(e).__name__}: {str(e)[:120]}")
res["detail"] = "; ".join(notes)
log(f"os-apply: OOM-CHECK result={res['result']} oom_killed={res['oom_killed']} oom_event={res['oom_event']} "
f"exit={res['exit_code']} image={img} — {res['detail']}")
return res
def guest_epoch(self):
rc, out, _ = self.g(["date", "+%s"], timeout=OOMCHECK_TIMEOUTS["clock"])
try:
return int(out.strip()) if rc == 0 else None
except ValueError:
return None
def dpkg_state(self): def dpkg_state(self):
"""`dpkg --audit` AND dpkg's update journal, in ONE call (R-876, agent v0.145.0). A crash in the middle of an """`dpkg --audit` AND dpkg's update journal, in ONE call (R-876, agent v0.145.0). A crash in the middle of an
install can leave `/var/lib/dpkg/updates/` non-empty while `--audit` reads clean — measured on demo-hp install can leave `/var/lib/dpkg/updates/` non-empty while `--audit` reads clean — measured on demo-hp
@@ -1058,10 +1257,26 @@ class Apply:
return [{"name": p["name"], "version": p["to"], "origin": DOCKER_ORIGIN} for p in pend return [{"name": p["name"], "version": p["to"], "origin": DOCKER_ORIGIN} for p in pend
if p["from"] is not None and p["name"] in DOCKER_NAMES and self.origin_name(p["origin"]) == {DOCKER_ORIGIN}] if p["from"] is not None and p["name"] in DOCKER_NAMES and self.origin_name(p["origin"]) == {DOCKER_ORIGIN}]
def pending_pve(self):
"""Ring 0 (select pending-pve): the newest pending version of each INSTALLED Proxmox-origin package, never a
kernel / boot / firmware / microcode name (HOST_SLOW_RE, R-836's lane), never another origin."""
rc, pend, remv, _ = self.simulate(["dist-upgrade"])
return [{"name": p["name"], "version": p["to"], "origin": PVE_ORIGIN} for p in pend
if p["from"] is not None and not HOST_SLOW_RE.match(p["name"]) and p["name"] not in DOCKER_NAMES
and self.origin_name(p["origin"]) == {PVE_ORIGIN}]
def pve_manager(self):
"""pveversion's pve-manager version ("unknown" when it cannot be read)."""
rc, out, _ = self.r.host(["pveversion"], 60)
m = re.match(r"^pve-manager/([^/\s]+)", out.strip()) if rc == 0 else None
return m.group(1) if m else "unknown"
def origin_ok(self, origin): def origin_ok(self, origin):
o = self.origin_name(origin) o = self.origin_name(origin)
if self.layer == "docker": if self.layer == "docker":
return o == {DOCKER_ORIGIN} return o == {DOCKER_ORIGIN}
if self.layer == "pve":
return o == {PVE_ORIGIN}
return bool(o & set(FAST_ORIGINS)) return bool(o & set(FAST_ORIGINS))
def apply(self, plan): def apply(self, plan):
@@ -1070,6 +1285,8 @@ class Apply:
packages = plan["packages"] packages = plan["packages"]
elif self.select == "pending-docker": elif self.select == "pending-docker":
packages = self.pending_docker() packages = self.pending_docker()
elif self.select == "pending-pve":
packages = self.pending_pve()
else: else:
packages = self.pending_fast() packages = self.pending_fast()
cmp_op = "ne" if self.allow_downgrade else "gt" cmp_op = "ne" if self.allow_downgrade else "gt"
@@ -1118,6 +1335,11 @@ class Apply:
raise Refused("R4", f"the plan would remove {', '.join(remv[:5])}") raise Refused("R4", f"the plan would remove {', '.join(remv[:5])}")
want = dict(upgrade) want = dict(upgrade)
for p in sim: for p in sim:
if p["from"] is None and self.layer == "pve" and p["name"] in PVE_NEW_ALLOW and self.origin_ok(p["origin"]) \
and not HOST_SLOW_RE.match(p["name"]):
self.r.log(f"os-apply: NEW {p['name']}={p['to']} (on the pve lane's allow-list)")
self.report.setdefault("added", []).append({"name": p["name"], "version": p["to"]})
continue
if p["from"] is None: if p["from"] is None:
raise Refused("R6", f"the plan would add a package that is not installed: {p['name']}") raise Refused("R6", f"the plan would add a package that is not installed: {p['name']}")
if p["name"] not in want: if p["name"] not in want:
@@ -1128,7 +1350,7 @@ class Apply:
raise Refused("R5", f"{p['name']} would be downgraded {p['from']} -> {p['to']}") raise Refused("R5", f"{p['name']} would be downgraded {p['from']} -> {p['to']}")
if not self.origin_ok(p["origin"]): if not self.origin_ok(p["origin"]):
raise Refused("R2", f"{p['name']} would come from {p['origin']}, not the {self.layer} layer's origin") raise Refused("R2", f"{p['name']} would come from {p['origin']}, not the {self.layer} layer's origin")
if self.layer == "host" and HOST_SLOW_RE.match(p["name"]): if self.layer in ("host", "pve") and HOST_SLOW_RE.match(p["name"]):
raise Refused("R14", f"{p['name']} is a kernel / boot / firmware package — the host's slow lane") raise Refused("R14", f"{p['name']} is a kernel / boot / firmware package — the host's slow lane")
need = self.download_bytes(args) need = self.download_bytes(args)
free = self.free_bytes() free = self.free_bytes()
@@ -1211,6 +1433,422 @@ class Apply:
self.x(APT_ENV + ["apt-get", "-q", "update"], timeout=600) self.x(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
def _iso(t):
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(t))
def _parse_iso(s):
try:
return calendar.timegm(time.strptime(s, "%Y-%m-%dT%H:%M:%SZ"))
except (TypeError, ValueError):
return None
class Kernel:
"""The kernel lane (R-836, `09` §3 decisions 164 + 172, `11` §5.11). Refusal codes: R3 (authority), R4 (removal),
R6 (a package outside the step), R8 (space), R9 (locks), R12 (appliance), R20 (the box's boot setup cannot do a
one-shot), R21 (the crash guard is tripped or saw an unclean boot within its window), R22 (the step's phase does not
allow this mode), R23 (the kernel set is not one exact new kernel).
Modes (plan "mode", layer "kernel", lane "slow"):
apply STAGE: install the kernel set, keep the GRUB default on the kernel the box runs, write the flag.
Never reboots. select "pending-kernel" (ring 0, the root-owned mark) or "listed" (a signed
os_kernel_step). expect_kver: the kernel the hub told the household about — any other is R23.
kernel-reboot a STAGED step's reboot (the night leg, after the household was told): phase oneshot, then reboot.
kernel-boot after a boot: what became of the step (judging | fell_back | self_reverted | revert_failed).
kernel-good the one-shot boot was healthy: the new kernel becomes the GRUB default.
kernel-revert the one-shot boot was NOT healthy: reboot ONCE into the old kernel (still the default).
kernel-cancel drop a staged step: clear the flag (the package stays installed, the default never moved).
kernel-status read only."""
def __init__(self, apply, plan):
self.a, self.r, self.plan = apply, apply.r, plan
self.report = apply.report
self.log = apply.r.log
# ---------- reading the box ----------
def running(self):
rc, out, _ = self.r.host(["uname", "-r"], 30)
v = out.strip() if rc == 0 else ""
return v if KVER_RE.match(v) else ""
def state(self):
try:
s = json.loads(self.r.read_file(KERNEL_STATE))
return s if isinstance(s, dict) else {}
except (OSError, ValueError):
return {}
def save_state(self, s):
s["updated_at"] = _iso(self.r.now())
self.r.put_file(KERNEL_STATE, (json.dumps(s, indent=2, sort_keys=True) + "\n").encode(), 0o644)
def flag(self):
"""The one-shot flag: the kernel named, "" when the env block holds none, None when there is no env block."""
rc, out, _ = self.r.host(["grub-editenv", ONESHOT_ENV, "list"], 30)
if rc != 0:
return None
for l in out.splitlines():
if l.startswith("felhom_next="):
return l.split("=", 1)[1].strip()
return ""
def grub_cfg(self):
try:
return self.r.read_file(GRUB_CFG)
except OSError:
return ""
@staticmethod
def default_kver(cfg):
"""The kernel grub.cfg boots by default (00_header's `set default=`), or "unknown"."""
m = re.search(r'^\s*set default="(?:gnulinux-advanced-[^>"]*>)?gnulinux-([0-9][^"]*?-pve)-advanced-[^"]*"', cfg, re.M)
return m.group(1) if m else "unknown"
@staticmethod
def entry_id(cfg, kver):
"""The GRUB_DEFAULT value that names kver's normal entry, read from grub.cfg itself (10_linux's ids)."""
sub = re.search(r"\$menuentry_id_option '(gnulinux-advanced-[^']+)'", cfg)
m = re.search(r"\$menuentry_id_option '(gnulinux-" + re.escape(kver) + r"-advanced-[^']+)'", cfg)
if not m:
return None
return f"{sub.group(1)}>{m.group(1)}" if sub else m.group(1)
def setup_problems(self):
"""R20: why this box cannot do a one-shot boot (empty = it can). Measured shape: UEFI, a vfat ESP at /boot/efi,
/boot on the root filesystem, GRUB with fat + loadenv, the bundle's two generators, no hand pin."""
why = []
if not self.r.lexists("/sys/firmware/efi"):
why.append("the box does not boot UEFI")
rc, out, _ = self.r.host(["findmnt", "-n", "-o", "FSTYPE", ESP_MOUNT], 30)
if rc != 0 or out.strip() != "vfat":
why.append(f"{ESP_MOUNT} is not a mounted vfat ESP ({out.strip() or 'not mounted'})")
rc, out, _ = self.r.host(["findmnt", "-n", "-o", "TARGET", "/boot"], 30)
if rc == 0 and out.strip():
why.append("/boot is a separate filesystem (the one-shot entries assume /boot on the root filesystem)")
for m in ("fat", "loadenv"):
if not self.r.lexists(f"/usr/lib/grub/x86_64-efi/{m}.mod"):
why.append(f"GRUB has no {m} module")
for p in (ONESHOT_SNIPPET, ONESHOT_ENTRIES):
if not self.r.lexists(p):
why.append(f"{p} is missing (the config bundle installs it)")
if self.r.lexists(KERNEL_PIN_FILE):
why.append("a kernel is pinned by hand (proxmox-boot-tool kernel pin) — the lane never fights it")
return why
def guard(self):
try:
return json.loads(self.r.read_file(CRASH_GUARD_STATE))
except (OSError, ValueError):
return None
def check_guard(self):
"""R21: a kernel step only on a box whose crash guard is armed and saw no unclean boot within its window — so
the step's own reboots (clean), one crash and one self-revert (clean) can never reach the 3rd unclean boot that
leaves the box off (`11` §5.9). Pinned by test_felhom_crash_guard KernelStepCannotLeaveTheBoxOff."""
g = self.guard()
if not isinstance(g, dict):
raise Refused("R21", f"no crash guard state ({CRASH_GUARD_STATE}) — a kernel step needs the guard")
if g.get("tripped") or not g.get("armed"):
raise Refused("R21", "the crash guard is tripped — no kernel step until it re-arms")
if (g.get("unclean_boots_in_window") or 0) > 0:
raise Refused("R21", f"{g.get('unclean_boots_in_window')} unclean boot(s) within the guard's window — wait")
def view(self, st=None, cfg=None):
st = self.state() if st is None else st
cfg = self.grub_cfg() if cfg is None else cfg
return {"running": self.running() or "unknown", "default": self.default_kver(cfg), "flag": self.flag(),
"phase": st.get("phase", "none"), "from": st.get("from"), "to": st.get("to"),
"step_id": st.get("step_id"), "self_revert_used": bool(st.get("self_revert_used")), "vmid": st.get("vmid"),
"staged_at": st.get("staged_at"), "rebooted_at": st.get("rebooted_at"),
"result_at": st.get("result_at"), "reason": st.get("reason")}
# ---------- writing the box ----------
def write_default(self, kver):
"""Pin the GRUB default to kver's normal entry, regenerate grub.cfg, and PROVE it (the default read back)."""
cfg = self.grub_cfg()
eid = self.entry_id(cfg, kver)
if not eid:
raise Refused("R20", f"grub.cfg has no normal entry for {kver}")
body = ("# felhom kernel lane (R-836, `11` §5.11) — written by felhom-os-apply; the kernel that booted healthily\n"
f'GRUB_DEFAULT="{eid}"\n')
self.r.put_file(KERNEL_DEFAULT_CFG, body.encode(), 0o644)
rc, out, err = self.r.host(["update-grub"], 300)
got = self.default_kver(self.grub_cfg())
if rc != 0 or got != kver:
raise Refused("R20", f"update-grub rc={rc}: the default reads {got}, not {kver}: {(out + err).strip()[-200:]}")
self.log(f"os-apply: KERNEL default = {kver} (proved from grub.cfg)")
def set_flag(self, kver):
self.r.host(["mkdir", "-p", os.path.dirname(ONESHOT_ENV)], 30)
if self.flag() is None:
rc, out, err = self.r.host(["grub-editenv", ONESHOT_ENV, "create"], 30)
if rc != 0:
raise Refused("R20", f"cannot create the one-shot env block on the ESP: {(out + err).strip()[-200:]}")
rc, out, err = self.r.host(["grub-editenv", ONESHOT_ENV, "set", f"felhom_next={kver}"], 30)
if rc != 0 or self.flag() != kver:
raise Refused("R20", f"the one-shot flag did not read back as {kver}: {(out + err).strip()[-200:]}")
def clear_flag(self):
if self.flag():
self.r.host(["grub-editenv", ONESHOT_ENV, "unset", "felhom_next"], 30)
def reboot(self, why):
self.log(f"os-apply: KERNEL REBOOT — {why}")
rc, out, err = self.r.host(["systemctl", "reboot"], 60)
self.report["reboot_rc"] = rc
if rc != 0:
self.report["failed"] = {"rc": 3, "step": "reboot", "reason": (out + err).strip()[-200:]}
return 3
return 0
# ---------- the modes ----------
def run(self):
mode = self.a.mode
self.report["kernel_mode"] = mode
if mode == "kernel-status":
v = self.view()
v["setup_problems"] = self.setup_problems()
self.report["kernel"] = v
return 0
fn = {"apply": self.stage, "kernel-reboot": self.reboot_staged, "kernel-boot": self.after_boot,
"kernel-good": self.good, "kernel-revert": self.revert, "kernel-cancel": self.cancel}[mode]
rc = fn()
self.report["kernel"] = self.view()
return rc
def phase_is(self, st, *phases):
if st.get("phase") not in phases:
raise Refused("R22", f"the kernel step is {st.get('phase', 'none')!r}, not {' or '.join(phases)} — "
f"{self.a.mode} does not apply")
def stage(self):
a = self.a
st = self.state()
if st.get("phase") in KERNEL_ACTIVE:
raise Refused("R22", f"a kernel step is already {st['phase']} ({st.get('from')} -> {st.get('to')})")
last = _parse_iso(st.get("staged_at"))
if last is not None and self.r.now() - last < KERNEL_MIN_GAP:
raise Refused("R22", "a kernel step was staged within the last 20 hours — never two in one night")
why = self.setup_problems()
if why:
raise Refused("R20", "; ".join(why))
self.check_guard()
a.check_guest(a.vmid)
who, _ = a.docker_authority(self.plan, op_name=KERNEL_OP)
self.report["authority"] = who
old = self.running()
if not old:
raise Refused("R20", "the running kernel is not a Proxmox kernel version")
self.log(f"os-apply: START release={self.plan.get('release_id')} layer=kernel lane=slow mode=apply "
f"select={a.select} authority={who} running={old}")
if a.apt_lock_held():
raise Refused("R9", "another apt/dpkg holds the lock on the host")
self.report["health_before"] = a.health()
a.repair()
rc, out, err = a.x(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
if rc != 0:
raise Refused("R7", f"apt-get update failed on the host: {(out + err).strip().splitlines()[-1:]}")
inst = a.installed()
if a.select == "pending-kernel":
_, pend, _, _ = a.simulate(["dist-upgrade"])
want = [(p["name"], p["to"]) for p in pend if p["from"] is not None and KERNEL_UPGRADE_RE.match(p["name"])
and a.origin_name(p["origin"]) == {PVE_ORIGIN}]
else:
want = [(e["name"], e["version"]) for e in self.plan["packages"]]
# upgrades of installed names (never a downgrade), and at most the listed new kernel image
args, upg = [], {}
for n, v in want:
if n in inst:
if a.dpkg_cmp(v, "gt", inst[n]):
upg[n] = v
args.append(f"{n}={v}")
elif KERNEL_IMAGE_RE.match(n):
args.append(f"{n}={v}")
else:
raise Refused("R6", f"{n} is not installed and is not a kernel image")
if not args:
self.report["upgraded"], self.report["outcome_hint"] = [], "nothing"
self.log("os-apply: DONE rc=0 upgraded=0 (no pending kernel)")
return 0
rc, sim, remv, text = a.simulate(["install", "--no-install-recommends"] + args)
if rc != 0:
raise Refused("R7", "the simulation failed: " + (text.strip().splitlines()[-1] if text.strip() else ""))
if remv:
raise Refused("R4", f"the kernel step would remove {', '.join(remv[:5])}")
images = []
listed = dict(want)
for p in sim:
if a.origin_name(p["origin"]) != {PVE_ORIGIN}:
raise Refused("R2", f"{p['name']} would come from {p['origin']}, not {PVE_ORIGIN!r}")
m = KERNEL_IMAGE_RE.match(p["name"])
if p["from"] is None:
if not m:
raise Refused("R6", f"the kernel step would add {p['name']}, which is not a kernel image")
if a.select == "listed" and listed.get(p["name"]) != p["to"]:
raise Refused("R23", f"the kernel step would add {p['name']}={p['to']}, not the signed set")
images.append((m.group(1), p["to"]))
continue
if p["name"] not in upg or p["to"] != upg[p["name"]]:
raise Refused("R6", f"the kernel step would touch {p['name']} ({p['to']}), which is not in the step")
if not a.dpkg_cmp(p["to"], "gt", p["from"]):
raise Refused("R5", f"{p['name']} would be downgraded {p['from']} -> {p['to']}")
if len(images) > 1:
raise Refused("R23", f"the kernel step would add {len(images)} kernels — one at a time")
# the target kernel: the new image, else the series meta-package's version (its image is already installed)
if images:
new = images[0][0]
if images[0][1] + "-pve" != new:
raise Refused("R23", f"the image version {images[0][1]} does not name the kernel {new}")
else:
metas = [(n, v) for n, v in upg.items() if re.match(r"^proxmox-kernel-[0-9]+\.[0-9]+$", n)]
if len(metas) != 1:
self.report["upgraded"], self.report["outcome_hint"] = [], "nothing"
self.log("os-apply: DONE rc=0 upgraded=0 (the pending set names no kernel to boot)")
return 0
new = metas[0][1] + "-pve"
if not KVER_RE.match(new) or not a.dpkg_cmp(new[:-4], "gt", old[:-4]):
raise Refused("R23", f"the kernel {new} is not newer than the running {old}")
ek = self.plan.get("expect_kver")
if ek and ek != new:
raise Refused("R23", f"the step would boot {new}, but the household was told about {ek}")
need = a.download_bytes(["install", "--no-install-recommends"] + args)
free = a.free_bytes()
if free >= 0 and free < max(MIN_FREE, 3 * need):
raise Refused("R8", f"free space {free} B is below max(500 MB, 3 x download {need} B)")
# 1. the default = the kernel the box RUNS (it booted healthily), proved from grub.cfg BEFORE the install
default_before = self.default_kver(self.grub_cfg())
self.write_default(old)
# 2. install (the kernel's own postinst runs update-grub; our default file keeps the default on `old`)
t0 = time.time()
rc, out, err = a.x(APT_ENV + ["apt-get", "-y", "-q"] + DPKG_OPTS + ["install", "--no-install-recommends"] + args)
a.x(["apt-get", "clean"])
if rc != 0:
_, aud, _ = a.x(["dpkg", "--audit"])
self.report["failed"] = {"rc": rc, "step": "install", "dpkg_audit": (aud.strip().splitlines() or ["clean"])[0],
"tail": (out + err).strip().splitlines()[-3:]}
self.log(f"os-apply: FAILED rc={rc} step=install (the default stays {old}; no flag written)")
return 3
self.report["upgraded"] = [{"name": x.split("=", 1)[0], "version": x.split("=", 1)[1]} for x in args]
self.report["seconds"] = round(time.time() - t0, 1)
# 3. prove: the image is there, the default is still `old`, the one-shot entry for `new` exists
cfg = self.grub_cfg()
if f"felhom-oneshot-{new}" not in cfg or self.default_kver(cfg) != old:
self.r.host(["update-grub"], 300)
cfg = self.grub_cfg()
for f in (f"/boot/vmlinuz-{new}", f"/boot/initrd.img-{new}"):
if not self.r.lexists(f):
self.report["failed"] = {"rc": 3, "step": "verify", "reason": f"{f} is missing after the install"}
return 3
if f"felhom-oneshot-{new}" not in cfg or "felhom_next" not in cfg:
self.report["failed"] = {"rc": 3, "step": "verify", "reason": f"grub.cfg has no one-shot entry for {new}"}
return 3
if self.default_kver(cfg) != old:
self.report["failed"] = {"rc": 3, "step": "verify",
"reason": f"the install moved the default to {self.default_kver(cfg)} — no flag written"}
return 3
# 4. the flag — the ONLY thing that makes the next boot use `new`, and only once
self.set_flag(new)
self.save_state({"phase": "staged", "step_id": self.plan.get("release_id"), "from": old, "to": new, "vmid": a.vmid,
"staged_at": _iso(self.r.now()), "authority": who, "default_before": default_before,
"packages": self.report["upgraded"], "self_revert_used": False})
self.report["reboot_needed"] = True
self.log(f"os-apply: KERNEL STAGED {old} -> {new} (default {old}, one-shot flag {new}); upgraded={len(args)} "
f"seconds={self.report['seconds']}")
return 0
def reboot_staged(self):
st = self.state()
self.phase_is(st, "staged")
if self.flag() != st.get("to"):
raise Refused("R22", f"the one-shot flag reads {self.flag()!r}, not {st.get('to')!r}")
if self.running() != st.get("from"):
raise Refused("R22", f"the box runs {self.running()!r}, not the step's old kernel {st.get('from')!r}")
cfg = self.grub_cfg()
if self.default_kver(cfg) != st["from"] or f"felhom-oneshot-{st['to']}" not in cfg:
raise Refused("R20", "grub.cfg no longer keeps the old default with a one-shot entry for the new kernel")
if self.setup_problems():
raise Refused("R20", "; ".join(self.setup_problems()))
self.check_guard()
self.a.check_guest(self.a.vmid)
st["health_before"] = self.a.health()
st["phase"], st["rebooted_at"] = "oneshot", _iso(self.r.now())
self.save_state(st)
return self.reboot(f"one-shot boot of {st['to']} (the default stays {st['from']})")
def after_boot(self):
st = self.state()
ph, run = st.get("phase"), self.running()
old, new = st.get("from"), st.get("to")
event = "none"
if ph in ("oneshot", "staged", "judging") and run == new and new:
if ph != "judging":
st["phase"], st["judging_since"], event = "judging", _iso(self.r.now()), "judging"
else:
event = "judging"
elif ph in ("oneshot", "judging") and run == old:
# the new kernel did not come up, or crashed: GRUB already booted the default (the old kernel)
self.clear_flag()
st["phase"], st["result_at"], event = "fell_back", _iso(self.r.now()), "fell_back"
st["reason"] = "the box came back on the old kernel by itself (the new one did not boot, or crashed)"
elif ph == "reverting" and run == old:
st["phase"], st["result_at"], event = "self_reverted", _iso(self.r.now()), "self_reverted"
elif ph == "reverting" and run == new:
st["phase"], st["result_at"], event = "revert_failed", _iso(self.r.now()), "revert_failed"
st["reason"] = "the self-revert came back on the NEW kernel — never retried (one self-revert per step)"
if event not in ("none",) and st.get("phase") != ph:
self.save_state(st)
self.log(f"os-apply: KERNEL AFTER-BOOT {ph} -> {st['phase']} running={run} ({old} -> {new})")
self.report["kernel_event"] = event
if event == "judging":
self.report["health_before"] = st.get("health_before") # the agent judges the boot against it
return 0
def good(self):
st = self.state()
self.phase_is(st, "judging")
if self.running() != st.get("to"):
raise Refused("R22", f"the box runs {self.running()!r}, not the new kernel {st.get('to')!r}")
try:
self.write_default(st["to"])
except Refused:
# put the old default back — the box must never be left without a proved default
self.write_default(st["from"])
raise
self.clear_flag()
st["phase"], st["result_at"] = "good", _iso(self.r.now())
self.save_state(st)
self.log(f"os-apply: KERNEL GOOD {st['to']} is the default now (was {st['from']})")
return 0
def revert(self):
st = self.state()
self.phase_is(st, "judging")
if st.get("self_revert_used"):
raise Refused("R22", "this kernel step already used its one self-revert")
if self.running() != st.get("to"):
raise Refused("R22", f"the box runs {self.running()!r}, not the new kernel {st.get('to')!r}")
self.clear_flag()
cfg = self.grub_cfg()
if self.default_kver(cfg) != st.get("from"):
raise Refused("R20", f"the GRUB default reads {self.default_kver(cfg)}, not the old {st.get('from')} — "
f"no self-revert into an unknown kernel")
reason = self.plan.get("reason")
st["reason"] = reason[:300] if isinstance(reason, str) else "the one-shot boot was not healthy"
st["self_revert_used"], st["phase"], st["reverted_at"] = True, "reverting", _iso(self.r.now())
self.save_state(st)
return self.reboot(f"self-revert to {st['from']}: {st['reason']}")
def cancel(self):
st = self.state()
self.phase_is(st, "staged")
self.clear_flag()
st["phase"], st["result_at"], st["reason"] = "cancelled", _iso(self.r.now()), "cancelled before the reboot"
self.save_state(st)
self.log(f"os-apply: KERNEL CANCELLED {st.get('to')} (installed, never the default; flag cleared)")
return 0
class Bundle: class Bundle:
"""The config bundle (R-840, `11` §5.4.2): every root-owned file the installer's step 5 writes, installed as ONE """The config bundle (R-840, `11` §5.4.2): every root-owned file the installer's step 5 writes, installed as ONE
signed unit. Every check runs before the first write; a failed write or a failed self-check puts every previous signed unit. Every check runs before the first write; a failed write or a failed self-check puts every previous
+50 -1
View File
@@ -17,6 +17,8 @@ Verbs (each one sudoers line, exact-match pattern):
wg /var/lib/felhom-agent/wg/wg-felhom.conf -> /etc/wireguard/wg-felhom.conf (0600) wg /var/lib/felhom-agent/wg/wg-felhom.conf -> /etc/wireguard/wg-felhom.conf (0600)
sshd-config /var/lib/felhom-agent/felhom-sshd/sshd_config -> /etc/felhom-sshd/sshd_config sshd-config /var/lib/felhom-agent/felhom-sshd/sshd_config -> /etc/felhom-sshd/sshd_config
sshd-key /var/lib/felhom-agent/felhom-sshd/authorized_keys.felhom-op -> /etc/felhom-sshd/authorized_keys/felhom-op sshd-key /var/lib/felhom-agent/felhom-sshd/authorized_keys.felhom-op -> /etc/felhom-sshd/authorized_keys/felhom-op
controller-image <vmid> the ref on STDIN -> /etc/felhom-controller-image INSIDE guest <vmid> (R-861 (a) A1): only
our registry + our repository + an x.y.z tag; the agent no longer has a `tee` grant
--self-check prints "felhom-priv-apply ok verbs=..." (the bundle's self-check) --self-check prints "felhom-priv-apply ok verbs=..." (the bundle's self-check)
Exit codes: 0 installed (or already identical), 2 usage, 3 refused (content or source), 4 install failed. Exit codes: 0 installed (or already identical), 2 usage, 3 refused (content or source), 4 install failed.
@@ -39,7 +41,14 @@ WG_SRC, WG_DEST = STATE + "/wg/wg-felhom.conf", "/etc/wireguard/wg-felhom.conf"
SSHD_SRC, SSHD_DEST = STATE + "/felhom-sshd/sshd_config", "/etc/felhom-sshd/sshd_config" SSHD_SRC, SSHD_DEST = STATE + "/felhom-sshd/sshd_config", "/etc/felhom-sshd/sshd_config"
KEY_SRC, KEY_DEST = STATE + "/felhom-sshd/authorized_keys.felhom-op", "/etc/felhom-sshd/authorized_keys/felhom-op" KEY_SRC, KEY_DEST = STATE + "/felhom-sshd/authorized_keys.felhom-op", "/etc/felhom-sshd/authorized_keys/felhom-op"
MAX_BYTES = 64 * 1024 MAX_BYTES = 64 * 1024
VERBS = ("unit", "dnsmasq", "wg", "sshd-config", "sshd-key") VERBS = ("unit", "dnsmasq", "wg", "sshd-config", "sshd-key", "controller-image")
# R-861 (a) A1 (`09` §3 decision 165): the SAME pattern as the agent's controllerImageRe (internal/localapi/
# controllerswap.go) — a compromised agent cannot hand the guest's bootstrap any other image. Pinned by
# configs/test_felhom_priv_apply.py ControllerImage.
CONTROLLER_IMAGE_RE = re.compile(r"^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$")
CONTROLLER_IMAGE_FILE = "/etc/felhom-controller-image"
CONTROLLER_IMAGE_MAX = 256
VMID_RE = re.compile(r"^[0-9]{1,9}$")
UNIT_NAME_RE = re.compile(r"^mnt-[A-Za-z0-9_.\\-]+\.(mount|automount)$") UNIT_NAME_RE = re.compile(r"^mnt-[A-Za-z0-9_.\\-]+\.(mount|automount)$")
DNSMASQ_TMP_RE = re.compile(r"^/tmp/felhom-resolver-[0-9]+\.conf$") DNSMASQ_TMP_RE = re.compile(r"^/tmp/felhom-resolver-[0-9]+\.conf$")
@@ -123,6 +132,16 @@ class Host:
pass pass
raise raise
def read_stdin(self, limit):
return sys.stdin.buffer.read(limit + 1)
def write_guest_image(self, vmid, data):
"""As root: `pct exec <vmid> -- tee <the fixed file>` with the checked ref on stdin (no shell)."""
r = subprocess.run(["/usr/sbin/pct", "exec", str(vmid), "--", "tee", CONTROLLER_IMAGE_FILE],
input=data, stdout=subprocess.DEVNULL, stderr=subprocess.PIPE, timeout=60)
if r.returncode != 0:
raise OSError(f"pct exec {vmid} tee exited {r.returncode}: {r.stderr.decode(errors='replace').strip()[:200]}")
def log(self, line): def log(self, line):
print(line, file=sys.stderr) print(line, file=sys.stderr)
try: try:
@@ -377,6 +396,34 @@ def plan(argv):
raise Refused("A1", f"wrong arguments for {v}") raise Refused("A1", f"wrong arguments for {v}")
def controller_image(rest, host):
"""R-861 (a) A1: read the ref on stdin, check it, write it INSIDE the guest as root."""
try:
if len(rest) != 1 or not VMID_RE.match(rest[0]):
raise Refused("A1", "usage: felhom-priv-apply controller-image <vmid> (the ref on stdin)")
raw = host.read_stdin(CONTROLLER_IMAGE_MAX)
if len(raw) > CONTROLLER_IMAGE_MAX:
raise Refused("I1", f"the image ref is longer than {CONTROLLER_IMAGE_MAX} bytes")
try:
text = raw.decode("ascii")
except UnicodeDecodeError:
raise Refused("I1", "the image ref is not ASCII")
ref = text[:-1] if text.endswith("\n") else text
if not CONTROLLER_IMAGE_RE.match(ref) or "\n" in ref:
raise Refused("I1", "the image ref is not gitea.dooplex.hu/admin/felhom-controller:<x.y.z>")
except Refused as e:
host.log(f"felhom-priv-apply: REFUSED [{e.rule}] controller-image {' '.join(rest)[:40]}: {e.reason}")
return 2 if e.rule == "A1" else 3
vmid = int(rest[0])
try:
host.write_guest_image(vmid, (ref + "\n").encode())
except (OSError, subprocess.SubprocessError) as e:
host.log(f"felhom-priv-apply: FAILED controller-image {vmid}: {e}")
return 4
host.log(f"felhom-priv-apply: WROTE controller-image {vmid} {ref}")
return 0
def main(argv, host=None): def main(argv, host=None):
host = host or Host() host = host or Host()
if argv == ["--self-check"]: if argv == ["--self-check"]:
@@ -396,6 +443,8 @@ def main(argv, host=None):
return 3 return 3
print("OK") print("OK")
return 0 return 0
if argv and argv[0] == "controller-image":
return controller_image(argv[1:], host)
try: try:
verb, src, dest, mode, checker = plan(argv) verb, src, dest, mode, checker = plan(argv)
data = host.read_source(src) data = host.read_source(src)
+4 -1
View File
@@ -546,7 +546,10 @@ class Builder(unittest.TestCase):
r"/etc/felhom/[a-z.-]+)", text)) r"/etc/felhom/[a-z.-]+)", text))
agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"} agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"}
trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD} trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust) # Written by the appliance ISO's first boot (felhom.eu scripts/iso/felhom-bootstrap.sh), never by the installer:
# since installer 1.32.0 (R-275) the uninstall only NAMES them under KEPT.
iso_writes = {"/etc/felhom/.bootstrap-done", "/etc/felhom/appliance-pairing-code"}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust | iso_writes)
self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them") self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them")
# the limits drop-in is named through $AGENT_UNIT in the installer # the limits drop-in is named through $AGENT_UNIT in the installer
self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS) self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS)
+57
View File
@@ -139,5 +139,62 @@ class Guard(unittest.TestCase):
self.assertEqual(mode, 0o644) self.assertEqual(mode, 0o644)
class KernelStepCannotLeaveTheBoxOff(unittest.TestCase):
"""R-836 / `11` §5.11 Part B 4: a kernel step's planned reboot, one crash and one self-revert cannot add up to the
box staying off. The wrapper starts a step only when the guard is armed with NO unclean boot in its window
(felhom-os-apply Kernel.check_guard, R21); the planned reboot and the self-revert are orderly (`systemctl reboot` —
the clean-stop marker); so the step adds at most ONE unclean boot, and the box stays off only after the LIMIT-th
(3rd) within the hour. Red-proof: audits/kernel-lane-2026-10-07/A/redproof.txt."""
def setUp(self):
self.d = tempfile.TemporaryDirectory()
self.e = FakeEnv(self.d.name)
cg.main(["x", "boot"], self.e)
s = self.e.state()
self.assertTrue(s["armed"])
self.assertEqual(s["unclean_boots_in_window"], 0, "the wrapper's precondition (R21)")
def tearDown(self):
self.d.cleanup()
def planned(self, minutes):
cg.main(["x", "clean-stop"], self.e)
self.e.t += minutes * 60
cg.main(["x", "boot"], self.e)
def crash(self, minutes):
self.e.t += minutes * 60
cg.main(["x", "boot"], self.e)
def test_planned_reboot_one_crash_one_self_revert(self):
self.planned(2) # the step's one-shot reboot (orderly)
self.crash(3) # the new kernel crashes after the guard ran; the box restarts (panic=10)
self.planned(2) # the self-revert (orderly)
s = self.e.state()
self.assertFalse(s["tripped"], s)
self.assertTrue(s["armed"])
self.assertEqual(self.e.panic(), 10, "the box still restarts after a crash")
self.assertEqual(s["unclean_boots_in_window"], 1, "the step added exactly one unclean boot")
def test_a_panic_before_userspace_is_not_even_counted(self):
# the one-shot kernel panics before the guard's unit runs (measured: rdinit= and init= missing): the planned
# reboot's clean-stop marker is still there when the old kernel boots, so this boot counts as clean.
cg.main(["x", "clean-stop"], self.e)
self.e.t += 120 # the panicking boot: no userspace, the guard never ran
cg.main(["x", "boot"], self.e)
s = self.e.state()
self.assertEqual(s["unclean_boots_in_window"], 0, s)
self.assertEqual(self.e.panic(), 10)
def test_the_box_stays_off_only_after_two_more_crashes_than_the_step_makes(self):
self.planned(2)
self.crash(3) # the step's one crash
self.planned(2) # the self-revert
self.crash(5) # an UNRELATED crash within the hour: the guard trips (the 3rd would leave it off)
s = self.e.state()
self.assertTrue(s["tripped"])
self.assertEqual(s["unclean_boots_in_window"], 2, "two unclean boots: one from the step, one not")
if __name__ == "__main__": if __name__ == "__main__":
unittest.main() unittest.main()
+832 -4
View File
@@ -83,7 +83,7 @@ class Fake:
return self.clock return self.clock
def sleep(self, s): def sleep(self, s):
pass self.sleeps = getattr(self, "sleeps", []) + [(len(self.calls), s)] # (calls made before it, seconds)
def verify_sig(self, signers, key_id, ns, blob, sig): def verify_sig(self, signers, key_id, ns, blob, sig):
self.verified = (signers, key_id, ns, blob, sig) self.verified = (signers, key_id, ns, blob, sig)
@@ -196,6 +196,27 @@ class Fake:
return 0, self.engine + "\n", "" return 0, self.engine + "\n", ""
if cmd == "docker" and a[1:3] == ["ps", "-q"]: if cmd == "docker" and a[1:3] == ["ps", "-q"]:
return 0, "".join(i + "\n" for i in self.ids), "" return 0, "".join(i + "\n" for i in self.ids), ""
# R-528: the memory-kill check's engine (oom_image None = unreadable; oom_state / oom_event the engine's answer)
if cmd == "date" and a[1:] == ["+%s"]:
# each read is 3 s later than the last, so the order of the reads is visible in the values
self.oom_epochs = getattr(self, "oom_epochs", []) + [int(self.clock) + 3 * len(getattr(self, "oom_epochs", []))]
return 0, f"{self.oom_epochs[-1]}\n", ""
if cmd == "docker" and a[1:4] == ["inspect", "-f", "{{.Config.Image}}"]:
img = getattr(self, "oom_image", "gitea.dooplex.hu/admin/felhom-controller:0.300.0")
return (0, img + "\n", "") if img is not None else (1, "", "Error: No such object: felhom-controller")
if cmd == "docker" and a[1] == "run":
self.oom_runs = getattr(self, "oom_runs", []) + [a]
return 137, "", ""
if cmd == "docker" and a[1:4] == ["inspect", "-f", "{{.State.OOMKilled}} {{.State.ExitCode}}"]:
if getattr(self, "oom_inspect_raises", False):
raise subprocess.TimeoutExpired(a, 10)
return 0, getattr(self, "oom_state", "true 137") + "\n", ""
if cmd == "docker" and a[1] == "events":
self.oom_events_argv = a
return 0, ("oom\n" if getattr(self, "oom_event", True) else ""), ""
if cmd == "docker" and a[1:3] == ["rm", "-f"]:
self.oom_removed = getattr(self, "oom_removed", []) + a[3:]
return 0, a[3] + "\n", ""
if cmd == "docker" and a[1] == "inspect": if cmd == "docker" and a[1] == "inspect":
mounts = {"aaa111": "/felhom-controller|/var/run/docker.sock;/app/data;", "bbb222": "/app|/data;"} mounts = {"aaa111": "/felhom-controller|/var/run/docker.sock;/app/data;", "bbb222": "/app|/data;"}
return 0, mounts.get(a[-1], "/other|;") + "\n", "" return 0, mounts.get(a[-1], "/other|;") + "\n", ""
@@ -211,6 +232,8 @@ class Fake:
return (0, self.daemon_json, "") if self.daemon_json is not None else (1, "", "No such file") return (0, self.daemon_json, "") if self.daemon_json is not None else (1, "", "No such file")
if cmd == "cat" and a[1] == "/etc/debian_version": if cmd == "cat" and a[1] == "/etc/debian_version":
return 0, "13.7\n", "" return 0, "13.7\n", ""
if cmd == "pveversion":
return 0, f"pve-manager/{self.installed.get('pve-manager', '9.2.2')}/abcdef (running kernel: 7.0.14-20-pve)\n", ""
if cmd == "uname": if cmd == "uname":
return 0, "7.0.14-20-pve\n", "" return 0, "7.0.14-20-pve\n", ""
if cmd == "apt-mark": if cmd == "apt-mark":
@@ -263,7 +286,7 @@ class Fake:
n, v = x.split("=", 1) n, v = x.split("=", 1)
if v not in self.avail(n): if v not in self.avail(n):
return 100, "", f"E: Version '{v}' for '{n}' was not found" return 100, "", f"E: Version '{v}' for '{n}' was not found"
origin = "Docker CE:trixie" if n in osapply.DOCKER_NAMES else SEC if n == "openssl" else DEB origin = getattr(self, "origins", {}).get(n) or ("Docker CE:trixie" if n in osapply.DOCKER_NAMES else SEC if n == "openssl" else DEB)
out += f"Inst {n} [{self.installed[n]}] ({v} {origin} [amd64])\n" out += f"Inst {n} [{self.installed[n]}] ({v} {origin} [amd64])\n"
out += "".join(l + "\n" for l in self.extra_sim) out += "".join(l + "\n" for l in self.extra_sim)
return 0, out, "" return 0, out, ""
@@ -997,8 +1020,6 @@ class RealSignatureCheck(unittest.TestCase):
self.assertNotEqual(r.verify_sig(self.signers, "someone-else", "felhom-op-v1", blob, sig), 0) self.assertNotEqual(r.verify_sig(self.signers, "someone-else", "felhom-op-v1", blob, sig), 0)
self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob, ns="other-ns")), 0) self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob, ns="other-ns")), 0)
if __name__ == "__main__":
unittest.main()
class UnsentReport(unittest.TestCase): class UnsentReport(unittest.TestCase):
@@ -1179,3 +1200,810 @@ class CrashLeftTheJournal(unittest.TestCase):
self.assertEqual(rc, 0, rep) self.assertEqual(rc, 0, rep)
self.assertTrue(any("INTERRUPTED" in l for l in f.logs), f.logs) self.assertTrue(any("INTERRUPTED" in l for l in f.logs), f.logs)
self.assertTrue(any(l.startswith("os-apply: REPAIR ") and l.endswith("forced") for l in f.logs), f.logs) self.assertTrue(any(l.startswith("os-apply: REPAIR ") and l.endswith("forced") for l in f.logs), f.logs)
class OOMCheck(unittest.TestCase):
"""R-528 (`09` decision 157): after a Docker engine step the wrapper proves the engine reports a memory kill
(OOMKilled=true AND the `oom` event). Reported only; the hub decides. Red-proofs: audits/readback-2026-10-07/F/."""
def apply(self, **kw):
f = docker_fake(signed=signed_job())
for k, v in kw.items():
setattr(f, k, v)
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
return f, rep
def test_pass_when_oomkilled_and_the_event(self):
f, rep = self.apply()
oc = rep["oom_check"]
self.assertEqual(oc["result"], "pass", oc)
self.assertEqual((oc["oom_killed"], oc["oom_event"], oc["exit_code"]), (True, True, 137))
self.assertEqual(oc["image"], "gitea.dooplex.hu/admin/felhom-controller:0.300.0")
self.assertEqual(sorted(oc), ["detail", "exit_code", "image", "oom_event", "oom_killed", "result"])
run_argv = f.oom_runs[0]
name = run_argv[run_argv.index("--name") + 1]
self.assertTrue(re.match(r"^felhom-oomcheck-[0-9]+-[0-9a-f]{8}$", name), name)
for flag, val in (("--pull", "never"), ("--network", "none"), ("--memory", "64m"), ("--memory-swap", "64m"),
("--label", "felhom.oomcheck=1"), ("--entrypoint", "sh")):
self.assertEqual(run_argv[run_argv.index(flag) + 1], val, flag)
self.assertNotIn("-v", run_argv)
self.assertNotIn("-p", run_argv)
self.assertEqual(run_argv[-3:], ["gitea.dooplex.hu/admin/felhom-controller:0.300.0", "-c", osapply.OOMCHECK_SCRIPT])
self.assertIn(f"container={name}", f.oom_events_argv)
self.assertIn("event=oom", f.oom_events_argv)
self.assertEqual(f.oom_removed, [name], "the check container must be removed")
self.assertTrue(rep["health_after"], "health is read before the check")
# bounded: every call has a timeout and the clock is read twice — the worst case stays within 90 s
self.assertLessEqual(sum(osapply.OOMCHECK_TIMEOUTS.values()) + osapply.OOMCHECK_TIMEOUTS["clock"]
+ osapply.OOMCHECK_SETTLE, 90)
def test_events_window_ends_after_the_settle_wait(self):
# Measured (F1/F2): an --until taken right after the run missed the oom event. The window must end after a
# wait of >= 2 s that comes AFTER the run, at the guest epoch read after that wait, + 1.
f, rep = self.apply()
run_i = next(i for i, c in enumerate(f.calls) if c[-1][:2] == ["docker", "run"])
date_i = [i for i, c in enumerate(f.calls) if c[-1] == ["date", "+%s"]]
waits = [(i, s) for i, s in getattr(f, "sleeps", []) if i > run_i]
self.assertTrue(waits and waits[0][1] >= 2, f"no settle wait after the run: {getattr(f, 'sleeps', None)}")
self.assertTrue(date_i[-1] >= waits[0][0], "the end epoch must be read after the wait")
ev = f.oom_events_argv
since, until = int(ev[ev.index("--since") + 1]), int(ev[ev.index("--until") + 1])
self.assertEqual(until, f.oom_epochs[-1] + 1, "until = the guest epoch read after the wait, + 1")
self.assertGreater(until, f.oom_epochs[0] + 1)
self.assertLess(since, f.oom_epochs[0] + 1)
def test_oomkilled_false_is_fail(self):
f, rep = self.apply(oom_state="false 0")
oc = rep["oom_check"]
self.assertEqual(oc["result"], "fail", oc)
self.assertIn("OOMKilled=true", oc["detail"])
self.assertTrue(oc["oom_event"])
self.assertEqual(len(f.oom_removed), 1)
def test_no_event_is_fail(self):
f, rep = self.apply(oom_event=False)
oc = rep["oom_check"]
self.assertEqual(oc["result"], "fail", oc)
self.assertIn("the oom event", oc["detail"])
self.assertTrue(oc["oom_killed"])
def test_image_unreadable_is_error(self):
f, rep = self.apply(oom_image=None)
oc = rep["oom_check"]
self.assertEqual(oc["result"], "error", oc)
self.assertIsNone(oc["image"])
self.assertIn("image could not be read", oc["detail"])
self.assertFalse(hasattr(f, "oom_runs"), "no container is started without an image")
def test_container_removed_even_when_inspect_raises(self):
f, rep = self.apply(oom_inspect_raises=True)
oc = rep["oom_check"]
self.assertEqual(oc["result"], "error", oc)
self.assertEqual(len(f.oom_removed), 1, "the container must be removed even when inspect raised")
self.assertTrue(f.oom_removed[0].startswith(osapply.OOMCHECK_PREFIX))
self.assertNotIn("failed", rep, "the check never turns the step into a failure")
def test_never_on_guest_or_host_or_in_health_mode(self):
g = Fake()
rc, rep = run(g)
self.assertEqual(rc, 0, rep)
h = Fake()
h.plan["layer"] = "host"
h.plan["packages"] = [{"name": "bash", "version": "5.2.37-2+b10", "origin": "Debian"}]
h.installed["bash"] = "5.2.37-2+b9"
h.live["bash"] = {"5.2.37-2+b10"}
rc_h, rep_h = run(h)
self.assertEqual(rc_h, 0, rep_h)
d = docker_fake(signed=signed_job())
d.plan["mode"] = "health"
rc_d, rep_d = run(d)
self.assertEqual(rc_d, 0, rep_d)
for f, r in ((g, rep), (h, rep_h), (d, rep_d)):
self.assertNotIn("oom_check", r)
self.assertFalse(hasattr(f, "oom_runs"), r.get("layer"))
self.assertFalse(any(c[-1][:2] == ["docker", "run"] for c in f.calls))
def test_mode_oom_check_runs_only_the_check(self):
f = docker_fake() # no authority: the check changes nothing, so it needs none
f.plan["mode"], f.plan["packages"] = "oom-check", []
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["oom_check"]["result"], "pass", rep)
self.assertFalse(any("apt-get" in c[-1] or "dpkg-query" in c[-1] for c in f.calls), f.calls)
self.assertEqual(len(f.oom_removed), 1)
def test_mode_oom_check_keeps_the_refusals(self):
f = docker_fake()
f.plan["mode"], f.plan["packages"] = "oom-check", []
f.files["/etc/pve/lxc/9201.conf"] = "arch: amd64\n"
rc, rep = run(f)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R10"), rep)
self.assertFalse(hasattr(f, "oom_runs"))
g = Fake()
g.plan["mode"] = "oom-check"
rc, rep = run(g)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R11"), rep)
PVE = "Proxmox Debian Repository:stable"
PVE_SET = [{"name": "pve-manager", "version": "9.2.21", "origin": "Proxmox Debian Repository"},
{"name": "libpve-common-perl", "version": "9.1.9", "origin": "Proxmox Debian Repository"}]
def pve_fake(signed=None, ring0=False):
f = Fake()
f.installed.update({"pve-manager": "9.2.2", "libpve-common-perl": "9.1.1", "proxmox-kernel-helper": "9.0.4"})
f.live["pve-manager"] = {"9.2.21", "9.2.2"}
f.live["libpve-common-perl"] = {"9.1.9", "9.1.1"}
f.origins = {"pve-manager": PVE, "libpve-common-perl": PVE, "proxmox-kernel-helper": PVE, "shim-signed": PVE,
"proxmox-firewall-data": PVE}
f.plan = {"release_id": "os-pve-t1", "layer": "pve", "lane": "slow", "vmid": 9201, "mode": "apply",
"packages": [dict(p) for p in PVE_SET]}
if ring0:
f.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": True})
if signed is not None:
f.plan["signed"] = signed
return f
class PVELane(unittest.TestCase):
"""R-812 option A (`09` §3 decision 163): the host's Proxmox USERSPACE packages, slow lane, no kernel / boot /
firmware / microcode (R14), no removal, a new package only from PVE_NEW_ALLOW. Red-proof: audits/day-2026-10-07/B/."""
def refused(self, f, code):
rc, rep = run(f)
self.assertEqual(rc, 2, rep)
self.assertEqual(rep["refused"]["code"], code, rep)
self.assertEqual(f.installed["pve-manager"], "9.2.2", "nothing may be installed on a refusal")
return rep
# THE RED TEST (design-R-812 §5): before the pve layer existed this plan was refused R12.
def test_pve_manager_plan_is_installed_on_the_host(self):
f = pve_fake(signed=signed_job(packages=PVE_SET, op="os_pve_step"))
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(f.installed["pve-manager"], "9.2.21")
self.assertEqual(rep["authority"], "signed")
self.assertEqual(rep["pve_manager"], "9.2.21", "pveversion after the step is reported")
inst = [c for c in f.calls if c[0] == "host" and "install" in c[1] and "-s" not in c[1] and "-f" not in c[1]
and "--print-uris" not in c[1]]
self.assertTrue(inst, "the pve layer installs on the HOST")
self.assertFalse([c for c in f.calls if c[0] == "guest" and "install" in c[2]], "nothing installed in the guest")
self.assertEqual(sorted(rep["health_after"]["host_services"]), sorted(osapply.HOST_SERVICES))
def test_kernel_in_a_pve_plan_is_refused(self):
f = pve_fake(ring0=True)
f.plan["packages"].append({"name": "proxmox-kernel-7.0", "version": "7.0.14-20", "origin": "Proxmox Debian Repository"})
self.refused(f, "R14")
def test_shim_in_a_pve_plan_is_refused(self):
f = pve_fake(ring0=True)
f.plan["packages"].append({"name": "shim-signed", "version": "1.47+pmx1", "origin": "Proxmox Debian Repository"})
self.refused(f, "R14")
def test_kernel_pulled_in_by_the_simulation_is_refused(self):
f = pve_fake(ring0=True)
f.installed["proxmox-kernel-helper"] = "9.0.4"
f.extra_sim = ["Inst proxmox-kernel-helper [9.0.4] (9.0.6 Proxmox Debian Repository:stable [all])"]
self.refused(f, "R6") # not in the plan — refused before the name check; R14 below when it IS in the plan
g = pve_fake(ring0=True)
g.live["proxmox-kernel-helper"] = {"9.0.6"}
g.plan["packages"] = [dict(PVE_SET[0])]
g.extra_sim = ["Inst proxmox-kernel-helper [9.0.4] (9.0.6 Proxmox Debian Repository:stable [all])"]
# a kernel-helper the plan did not name is R6; the R14 name check covers a listed one (test above)
self.refused(g, "R6")
def test_debian_package_in_a_pve_plan_is_refused(self):
f = pve_fake(ring0=True)
f.plan["packages"].append({"name": "libc6", "version": "2.41-12+deb13u4", "origin": "Debian"})
self.refused(f, "R2")
def test_debian_origin_in_the_pve_simulation_is_refused(self):
f = pve_fake(ring0=True)
f.origins["libpve-common-perl"] = DEB
self.refused(f, "R2")
def test_docker_package_in_a_pve_plan_is_refused(self):
f = pve_fake(ring0=True)
f.plan["packages"].append({"name": "docker-ce", "version": "5:29.8.2-1~debian.13~trixie", "origin": "Proxmox Debian Repository"})
self.refused(f, "R2")
def test_pve_in_the_fast_lane_is_refused(self):
f = pve_fake(ring0=True)
f.plan["lane"] = "fast"
self.refused(f, "R3")
def test_no_authority_is_refused(self):
self.refused(pve_fake(), "R3")
def test_docker_signed_op_does_not_authorize_a_pve_step(self):
self.refused(pve_fake(signed=signed_job(packages=PVE_SET, op="os_docker_step")), "R3")
def test_pve_on_a_byo_box_is_refused(self):
f = pve_fake(ring0=True)
f.files[osapply.INSTALL_STATE] = json.dumps({"mode": "byo"})
self.refused(f, "R12")
def test_unlisted_new_package_is_refused(self):
f = pve_fake(ring0=True)
f.extra_sim = ["Inst proxmox-new-thing (1.0 Proxmox Debian Repository:stable [all])"]
self.refused(f, "R6")
def test_allow_listed_new_package_is_accepted(self):
f = pve_fake(ring0=True)
f.extra_sim = ["Inst proxmox-firewall-data (0.1 Proxmox Debian Repository:stable [all])"]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(f.installed["pve-manager"], "9.2.21")
def test_allow_new_is_never_an_open_door_in_the_fast_lane(self):
f = Fake()
f.plan["layer"] = "host"
f.extra_sim = ["Inst proxmox-firewall-data (0.1 Proxmox Debian Repository:stable [all])"]
rc, rep = run(f)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R6"), rep)
def test_undo_is_refused(self):
f = pve_fake(signed=signed_job(packages=PVE_SET, op="os_pve_step"))
f.plan["undo"] = True
self.refused(f, "R5")
def test_ring0_pending_pve_takes_only_installed_proxmox_userspace(self):
f = pve_fake(ring0=True)
f.plan["select"], f.plan["packages"] = "pending-pve", []
f.installed.update({"linux-image-amd64": "6.12.1", "tailscale": "1.102.2"})
f.pending_sim = [
"Inst pve-manager [9.2.2] (9.2.21 Proxmox Debian Repository:stable [amd64])",
"Inst libpve-common-perl [9.1.1] (9.1.9 Proxmox Debian Repository:stable [all])",
"Inst proxmox-kernel-helper [9.0.4] (9.0.6 Proxmox Debian Repository:stable [all])",
"Inst proxmox-kernel-7.0.14-20-pve-signed (7.0.14-20 Proxmox Debian Repository:stable [amd64])",
"Inst proxmox-firewall-data (0.1 Proxmox Debian Repository:stable [all])",
"Inst libc6 [2.41-12+deb13u3] (2.41-12+deb13u4 Debian:13.7/stable [amd64])",
"Inst tailscale [1.102.2] (1.102.5 Tailscale:pkgs.tailscale.com [amd64])",
]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["authority"], "ring0")
self.assertEqual(sorted(u["name"] for u in rep["upgraded"]), ["libpve-common-perl", "pve-manager"],
"pending-pve: installed, Proxmox-origin, never kernel/boot/firmware, never Debian or other origins")
self.assertEqual(f.installed["libc6"], "2.41-12+deb13u3")
# 2026-10-09, demo-hp (Secure Boot on): proxmox-secure-boot-support's upgrade pulls shim-signed-common; in the pve
# plan that refused the whole step (R6) and skipped the night's kernel step. It is a boot-chain package: left out.
# RED-PROOF: drop proxmox-secure-boot-support from HOST_SLOW_RE -> it is in the plan -> this test fails.
def test_ring0_pending_pve_leaves_the_secure_boot_meta_out(self):
f = pve_fake(ring0=True)
f.plan["select"], f.plan["packages"] = "pending-pve", []
f.installed.update({"proxmox-secure-boot-support": "9.0.1"})
f.pending_sim = [
"Inst pve-manager [9.2.2] (9.2.21 Proxmox Debian Repository:stable [amd64])",
"Inst proxmox-secure-boot-support [9.0.1] (9.0.2 Proxmox Debian Repository:stable [all])",
"Inst shim-signed-common [1.48+pmx1+16.1-1+pmx1] (1.51+pmx1+16.1-2+pmx1 Proxmox Debian Repository:stable [all])",
]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(sorted(u["name"] for u in rep["upgraded"]), ["pve-manager"])
self.assertEqual(f.installed["proxmox-secure-boot-support"], "9.0.1")
def test_select_pending_pve_needs_the_pve_layer(self):
f = Fake()
f.plan["layer"], f.plan["select"], f.plan["packages"] = "host", "pending-pve", []
rc, rep = run(f)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R11"), rep)
# ---------- the kernel lane (R-836, `09` §3 decision 172, `11` §5.11) ----------
U = "1af1fcc6-639c-416b-a7e5-c4470d41a502"
OLD, NEW = "7.0.2-6-pve", "7.0.14-22-pve"
KERNEL_SET = [{"name": "proxmox-kernel-7.0", "version": "7.0.14-22", "origin": "Proxmox Debian Repository"},
{"name": "proxmox-kernel-7.0.14-22-pve-signed", "version": "7.0.14-22", "origin": "Proxmox Debian Repository"}]
class KFake(Fake):
"""A Proxmox host with GRUB on UEFI: /boot on the root LV, a vfat ESP, the bundle's two generators. update-grub is
emulated like the real 10_linux: WITHOUT the felhom default file the NEWEST kernel becomes the default (measured,
R-836 — that is the defect the lane exists for); with it, the kernel it names."""
def __init__(self):
super().__init__()
self.running = OLD
self.boot = {OLD}
self.efi, self.esp_fs, self.boot_mount, self.mods, self.snippets, self.pin = True, "vfat", "", True, True, False
self.env = None # None = no env block on the ESP; else {"felhom_next": ...}
self.tree = {} # files put_file wrote
self.reboots = []
self.update_grubs = 0
self.grub_runs_in_postinst = True
self.installed.update({"proxmox-kernel-7.0": "7.0.2-6", "proxmox-default-kernel": "2.1.0",
"pve-firmware": "3.18-3", "proxmox-kernel-7.0.2-6-pve-signed": "7.0.2-6",
"pve-manager": "9.2.21"})
self.live.update({"proxmox-kernel-7.0": {"7.0.14-22", "7.0.2-6"},
"proxmox-kernel-7.0.14-22-pve-signed": {"7.0.14-22"},
"proxmox-kernel-7.0.14-23-pve-signed": {"7.0.14-23"},
"pve-firmware": {"3.18-7", "3.18-3"}})
self.origins = {}
self.kernel_pending = ["Inst pve-firmware [3.18-3] (3.18-7 Proxmox Debian Repository:stable [all])",
"Inst proxmox-kernel-7.0.14-22-pve-signed (7.0.14-22 Proxmox Debian Repository:stable [amd64])",
"Inst proxmox-kernel-7.0 [7.0.2-6] (7.0.14-22 Proxmox Debian Repository:stable [amd64])",
"Inst pve-manager [9.2.21] (9.2.22 Proxmox Debian Repository:stable [amd64])",
"Inst libc6 [2.41-12+deb13u3] (2.41-12+deb13u4 Debian:13.7/stable [amd64])"]
self.plan = {"release_id": "kernel-t1", "layer": "kernel", "lane": "slow", "vmid": 9201, "mode": "apply",
"select": "pending-kernel", "packages": []}
self.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": True})
self.files[osapply.CRASH_GUARD_STATE] = json.dumps({"armed": True, "tripped": False, "unclean_boots_in_window": 0})
self.cfg = self.render()
# -- GRUB --
def newest(self):
best = None
for k in self.boot:
if best is None or dpkg_cmp(k[:-4], "gt", best[:-4]):
best = k
return best
def render(self):
d = self.tree.get(osapply.KERNEL_DEFAULT_CFG)
if d:
dflt = re.search(r'GRUB_DEFAULT="([^"]+)"', d).group(1)
else:
dflt = f"gnulinux-advanced-{U}>gnulinux-{self.newest()}-advanced-{U}"
cfg = ('if [ "${next_entry}" ] ; then\n set default="${next_entry}"\nelse\n'
f' set default="{dflt}"\nfi\n'
f"submenu 'Advanced options' $menuentry_id_option 'gnulinux-advanced-{U}' {{\n")
for k in sorted(self.boot):
cfg += f" menuentry 'Proxmox VE, with Linux {k}' $menuentry_id_option 'gnulinux-{k}-advanced-{U}' {{ }}\n"
cfg += "}\n"
if self.snippets:
cfg += "### BEGIN /etc/grub.d/01_felhom_oneshot ###\nload_env felhom_next\n"
cfg += "".join(f"menuentry 'Felhom one-shot: {k}' --id felhom-oneshot-{k} {{ }}\n" for k in sorted(self.boot))
return cfg
def lexists(self, p):
if p == "/sys/firmware/efi":
return self.efi
if p.startswith("/usr/lib/grub/x86_64-efi/"):
return self.mods
if p in (osapply.ONESHOT_SNIPPET, osapply.ONESHOT_ENTRIES):
return self.snippets
if p == osapply.KERNEL_PIN_FILE:
return self.pin
m = re.match(r"^/boot/(vmlinuz|initrd\.img)-(.+)$", p)
if m:
return m.group(2) in self.boot
return p in self.tree
def put_file(self, path, data, mode):
self.tree[path] = data.decode()
def read_file(self, p):
if p == osapply.GRUB_CFG:
return self.cfg
if p in self.tree:
return self.tree[p]
return super().read_file(p)
def host(self, argv, timeout=600, stdin=None):
if argv[0] == "uname":
self.calls.append(("host", argv))
return 0, self.running + "\n", ""
if argv[0] == "findmnt":
self.calls.append(("host", argv))
if argv[-1] == osapply.ESP_MOUNT:
return (0, self.esp_fs + "\n", "") if self.esp_fs else (1, "", "")
return (0, self.boot_mount + "\n", "") if self.boot_mount else (1, "", "")
if argv[0] == "grub-editenv":
self.calls.append(("host", argv))
if argv[2] == "list":
return (1, "", "no such file") if self.env is None else \
(0, "".join(f"{k}={v}\n" for k, v in self.env.items()), "")
if argv[2] == "create":
self.env = {}
return 0, "", ""
if argv[2] == "set":
k, v = argv[3].split("=", 1)
self.env[k] = v
return 0, "", ""
if argv[2] == "unset":
self.env.pop(argv[3], None)
return 0, "", ""
if argv[0] == "update-grub":
self.calls.append(("host", argv))
self.update_grubs += 1
self.cfg = self.render()
return 0, "", ""
if argv[0] == "mkdir":
self.calls.append(("host", argv))
return 0, "", ""
if argv[:2] == ["systemctl", "reboot"]:
self.calls.append(("host", argv))
self.reboots.append(self.running)
return 0, "", ""
return super().host(argv, timeout, stdin)
def emulate(self, argv):
a = [x for x in argv if not re.match(r"^[A-Z_]+=", x) and x != "env"]
if a[0] == "apt-get" and "install" in a and "-s" not in a and "--print-uris" not in a and "-f" not in a:
if self.install_rc:
return self.install_rc, "", "E: boom"
for x in a:
if "=" in x and not x.startswith("-") and "::" not in x:
n, v = x.split("=", 1)
self.installed[n] = v
m = re.match(r"^proxmox-kernel-[0-9]+\.[0-9]+$", n)
if m:
self.installed[f"proxmox-kernel-{v}-pve-signed"] = v
if m or osapply.KERNEL_IMAGE_RE.match(n):
self.boot.add(v + "-pve")
if self.grub_runs_in_postinst:
self.cfg = self.render() # the kernel's postinst runs update-grub (zz-update-grub)
return 0, "Setting up proxmox-kernel ...\n", ""
return super().emulate(argv)
def sim(self, a):
if "--print-uris" in a:
return 0, "", ""
if "dist-upgrade" in a:
return 0, "\n".join(self.kernel_pending) + "\n", ""
out, named = "", set()
for x in a:
if "=" in x and not x.startswith("-") and "::" not in x:
named.add(x.split("=", 1)[0])
for x in a:
if "=" in x and not x.startswith("-") and "::" not in x:
n, v = x.split("=", 1)
if v not in self.avail(n):
return 100, "", f"E: Version '{v}' for '{n}' was not found"
o = self.origins.get(n, PVE)
out += f"Inst {n} [{self.installed[n]}] ({v} {o} [amd64])\n" if n in self.installed else f"Inst {n} ({v} {o} [amd64])\n"
if re.match(r"^proxmox-kernel-[0-9]+\.[0-9]+$", n):
img = f"proxmox-kernel-{v}-pve-signed"
if img not in self.installed and img not in named:
out += f"Inst {img} ({v} {PVE} [amd64])\n"
out += "".join(l + "\n" for l in self.extra_sim)
return 0, out, ""
def kfake(**kw):
f = KFake()
for k, v in kw.items():
setattr(f, k, v)
return f
def kmode(f, mode, **extra):
f.plan = {"release_id": "kernel-t1", "layer": "kernel", "lane": "slow", "vmid": 9201, "mode": mode, "packages": []}
f.plan.update(extra)
return run(f)
def staged(f=None):
f = f or kfake()
rc, rep = run(f)
assert rc == 0, rep
return f
class KernelLane(unittest.TestCase):
"""R-836 / `09` §3 decision 172: stage a kernel once through the ESP flag; the default moves only after a healthy
one-shot boot. Red-proof: audits/kernel-lane-2026-10-07/A/redproof.txt."""
def refused(self, f, code, mode=None, **extra):
rc, rep = kmode(f, mode, **extra) if mode else run(f)
self.assertEqual(rc, 2, rep)
self.assertEqual(rep["refused"]["code"], code, rep)
return rep
# --- the four red tests the brief names (Part A 3) ---
def test_installing_a_kernel_does_not_change_the_grub_default(self):
f = kfake()
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertIn(NEW, f.boot, "the new kernel was installed")
self.assertEqual(osapply.Kernel.default_kver(f.cfg), OLD,
"the default must stay the kernel the box runs (R-836: an install made the new one default)")
self.assertIn(f"gnulinux-{OLD}-advanced-{U}", f.tree[osapply.KERNEL_DEFAULT_CFG])
self.assertEqual(f.env, {"felhom_next": NEW}, "the ONLY way to the new kernel is the one-shot flag")
self.assertEqual(f.reboots, [], "staging never reboots")
st = json.loads(f.tree[osapply.KERNEL_STATE])
self.assertEqual((st["phase"], st["from"], st["to"]), ("staged", OLD, NEW))
self.assertEqual(rep["kernel"]["default"], OLD)
def test_a_box_without_the_esp_flag_is_refused(self):
for attr, val, frag in (("esp_fs", "ext4", "vfat"), ("efi", False, "UEFI"), ("snippets", False, "bundle"),
("mods", False, "module"), ("boot_mount", "/boot", "separate"), ("pin", True, "pinned")):
f = kfake(**{attr: val})
rep = self.refused(f, "R20")
self.assertIn(frag, rep["refused"]["reason"], attr)
self.assertNotIn(NEW, f.boot, f"{attr}: nothing may be installed on a refusal")
self.assertIsNone(f.env, f"{attr}: no flag on a refusal")
def test_a_reboot_without_the_flag_is_refused(self):
f = staged()
f.env = {"felhom_next": ""}
self.refused(f, "R22", mode="kernel-reboot")
self.assertEqual(f.reboots, [])
def test_a_kernel_outside_the_approved_set_is_refused(self):
# a signed set that names only the meta-package: the image the sources pull in was never signed for
f = kfake()
f.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": False})
f.plan.update(select="listed", packages=[dict(KERNEL_SET[0])],
signed=signed_job(packages=[KERNEL_SET[0]], op="os_kernel_step"))
self.refused(f, "R23")
self.assertNotIn(NEW, f.boot)
# a signed set names 7.0.14-22; the sources would ALSO bring 7.0.14-23
f2 = kfake()
f2.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": False})
f2.plan.update(select="listed", packages=[dict(p) for p in KERNEL_SET],
signed=signed_job(packages=KERNEL_SET, op="os_kernel_step"))
f2.extra_sim = ["Inst proxmox-kernel-7.0.14-23-pve-signed (7.0.14-23 Proxmox Debian Repository:stable [amd64])"]
self.refused(f2, "R23")
# a name that is not a kernel package at all
g = kfake()
g.plan.update(select="listed", packages=[{"name": "pve-manager", "version": "9.2.22", "origin": "Proxmox Debian Repository"}])
self.refused(g, "R23")
# the household was told about another kernel
h = kfake()
h.plan["expect_kver"] = "7.0.14-23-pve"
self.refused(h, "R23")
self.assertNotIn(NEW, h.boot)
def test_the_snippet_clears_the_flag_on_use(self):
"""01_felhom_oneshot: the flag is copied, CLEARED and SAVED before any `set default`, all inside the one branch
that runs only when the flag is set; a default is set only for an installed kernel's own entry."""
text = (HERE / "felhom-grub-oneshot.sh").read_text()
body = text[text.index("cat <<EOF"):]
i_copy = body.index('set felhom_boot="\\${felhom_next}"')
i_clear = body.index("set felhom_next=\n")
i_save = body.index("save_env -f (\\$felhom_esp)/EFI/felhom/oneshot.env felhom_next")
i_default = body.index('set default="felhom-oneshot-%s"')
self.assertLess(i_copy, i_clear)
self.assertLess(i_clear, i_save)
self.assertLess(i_save, i_default, "the flag must be cleared on disk BEFORE GRUB boots anything")
self.assertIn('if [ "${felhom_boot}" = "%s" ]', body, "a default only for a kernel that is installed")
self.assertIn(osapply.ONESHOT_ENV[len(osapply.ESP_MOUNT):], text, "the wrapper and GRUB name the same env file")
# --- the rest of the stage ---
def test_option_c_options_are_on_the_one_shot_entry_only(self):
text = (HERE / "felhom-grub-oneshot-entries.sh").read_text()
self.assertIn(osapply.ONESHOT_ARGS, text)
self.assertIn("--id felhom-oneshot-$k", text)
self.assertNotIn("set default", text, "the entries file never sets the default")
def test_both_generators_ride_the_bundle(self):
self.assertEqual(osapply.BUNDLE_DESTS[osapply.ONESHOT_SNIPPET][1:4], ("felhom-grub-oneshot.sh", 0o755, "sh"))
self.assertEqual(osapply.BUNDLE_DESTS[osapply.ONESHOT_ENTRIES][1:4], ("felhom-grub-oneshot-entries.sh", 0o755, "sh"))
def test_ring0_pending_kernel_takes_only_the_kernel_set(self):
f = kfake()
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["authority"], "ring0")
self.assertEqual(sorted(u["name"] for u in rep["upgraded"]), ["proxmox-kernel-7.0", "pve-firmware"])
self.assertEqual(f.installed["pve-manager"], "9.2.21", "a Proxmox userspace package is the pve lane's")
self.assertEqual(f.installed["libc6"], "2.41-12+deb13u3", "a Debian package is the fast lane's")
def test_signed_listed_set_is_installed(self):
f = kfake()
f.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": False})
f.plan.update(select="listed", packages=[dict(p) for p in KERNEL_SET],
signed=signed_job(packages=KERNEL_SET, op="os_kernel_step"))
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["authority"], "signed")
self.assertEqual(f.env, {"felhom_next": NEW})
def test_no_authority_and_the_wrong_op_are_refused(self):
f = kfake()
f.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": False})
self.refused(f, "R3")
g = kfake()
g.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": False})
g.plan.update(select="listed", packages=[dict(p) for p in KERNEL_SET],
signed=signed_job(packages=KERNEL_SET, op="os_pve_step"))
self.refused(g, "R3")
def test_fast_lane_byo_and_wrong_select_are_refused(self):
f = kfake()
f.plan["lane"] = "fast"
self.refused(f, "R3")
g = kfake()
g.files[osapply.INSTALL_STATE] = json.dumps({"mode": "byo"})
self.refused(g, "R12")
h = kfake()
h.plan["select"] = "pending-pve"
self.refused(h, "R11")
k = kfake()
k.plan.update(layer="host", lane="fast", mode="kernel-status")
self.refused(k, "R11")
def test_the_crash_guard_must_be_armed_and_quiet(self):
f = kfake()
f.files[osapply.CRASH_GUARD_STATE] = json.dumps({"armed": False, "tripped": True, "unclean_boots_in_window": 2})
self.refused(f, "R21")
g = kfake()
g.files[osapply.CRASH_GUARD_STATE] = json.dumps({"armed": True, "tripped": False, "unclean_boots_in_window": 1})
self.refused(g, "R21")
h = kfake()
del h.files[osapply.CRASH_GUARD_STATE]
self.refused(h, "R21")
self.assertNotIn(NEW, h.boot)
def test_never_two_steps_in_one_night(self):
f = staged()
self.refused(f, "R22") # still staged
g = staged()
st = json.loads(g.tree[osapply.KERNEL_STATE])
st["phase"] = "good"
g.tree[osapply.KERNEL_STATE] = json.dumps(st)
g.clock += 3600
self.refused(g, "R22") # done, but within 20 h
def test_removal_new_non_kernel_two_kernels_and_older_are_refused(self):
f = kfake(extra_sim=["Remv pve-firmware [3.18-3]"])
self.refused(f, "R4")
g = kfake(extra_sim=["Inst proxmox-new-thing (1.0 Proxmox Debian Repository:stable [all])"])
self.refused(g, "R6")
h = kfake(extra_sim=["Inst proxmox-kernel-7.0.14-23-pve-signed (7.0.14-23 Proxmox Debian Repository:stable [amd64])"])
self.refused(h, "R23")
k = kfake(running="7.0.14-23-pve")
k.boot = {"7.0.14-23-pve"}
k.cfg = k.render()
self.refused(k, "R23")
m = kfake()
m.origins = {"proxmox-kernel-7.0": DEB}
self.refused(m, "R2")
# R-898: ring 0 stages EXACTLY the told kernel (select listed, the set derived from it) — even when the sources
# offer a newer one by night; a told version that can no longer be installed is refused BEFORE any change (R7).
def test_ring0_listed_installs_the_told_kernel_not_the_newest(self):
f = kfake()
f.live["proxmox-kernel-7.0"] = {"7.0.14-20", "7.0.14-22", "7.0.2-6"}
f.live["proxmox-kernel-7.0.14-20-pve-signed"] = {"7.0.14-20"}
told = [{"name": "proxmox-kernel-7.0", "version": "7.0.14-20", "origin": "Proxmox Debian Repository"},
{"name": "proxmox-kernel-7.0.14-20-pve-signed", "version": "7.0.14-20", "origin": "Proxmox Debian Repository"}]
f.plan.update(select="listed", packages=told, expect_kver="7.0.14-20-pve")
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["authority"], "ring0")
self.assertEqual(f.env, {"felhom_next": "7.0.14-20-pve"})
self.assertEqual(f.installed["proxmox-kernel-7.0"], "7.0.14-20")
self.assertNotIn("7.0.14-22-pve", f.boot)
def test_ring0_told_kernel_gone_is_refused_before_any_change(self):
f = kfake()
f.live["proxmox-kernel-7.0"] = {"7.0.14-22"} # 7.0.14-20 is no longer in the archive
told = [{"name": "proxmox-kernel-7.0", "version": "7.0.14-20", "origin": "Proxmox Debian Repository"},
{"name": "proxmox-kernel-7.0.14-20-pve-signed", "version": "7.0.14-20", "origin": "Proxmox Debian Repository"}]
f.plan.update(select="listed", packages=told, expect_kver="7.0.14-20-pve")
rep = self.refused(f, "R7")
self.assertIsNone(f.env)
self.assertNotIn(osapply.KERNEL_DEFAULT_CFG, f.tree, "refused before the default was even pinned")
self.assertEqual(f.installed["proxmox-kernel-7.0"], "7.0.2-6")
def test_nothing_pending_changes_nothing(self):
f = kfake(kernel_pending=[])
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["upgraded"], [])
self.assertIsNone(f.env)
self.assertNotIn(osapply.KERNEL_STATE, f.tree)
def test_a_failed_install_writes_no_flag(self):
f = kfake(install_rc=100)
rc, rep = run(f)
self.assertEqual(rc, 3, rep)
self.assertIsNone(f.env)
self.assertNotIn(osapply.KERNEL_STATE, f.tree)
self.assertEqual(osapply.Kernel.default_kver(f.cfg), OLD)
def test_a_postinst_without_update_grub_is_regenerated_and_proved(self):
f = kfake(grub_runs_in_postinst=False)
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertIn(f"felhom-oneshot-{NEW}", f.cfg)
self.assertEqual(osapply.Kernel.default_kver(f.cfg), OLD)
# --- the reboot, the boot, good / revert ---
def test_reboot_of_a_staged_step(self):
f = staged()
rc, rep = kmode(f, "kernel-reboot")
self.assertEqual(rc, 0, rep)
self.assertEqual(f.reboots, [OLD])
st = json.loads(f.tree[osapply.KERNEL_STATE])
self.assertEqual(st["phase"], "oneshot")
self.assertIn("host_services", st["health_before"])
def test_reboot_needs_a_staged_step(self):
self.refused(kfake(), "R22", mode="kernel-reboot")
def boot_into(self, f, kver):
f.running = kver
if f.env and f.env.get("felhom_next"):
f.env["felhom_next"] = "" # GRUB cleared it (01_felhom_oneshot)
return kmode(f, "kernel-boot")
def test_healthy_one_shot_becomes_the_default(self):
f = staged()
kmode(f, "kernel-reboot")
rc, rep = self.boot_into(f, NEW)
self.assertEqual((rc, rep["kernel_event"], rep["kernel"]["phase"]), (0, "judging", "judging"), rep)
self.assertEqual(osapply.Kernel.default_kver(f.cfg), OLD, "judging does not move the default")
rc, rep = kmode(f, "kernel-good")
self.assertEqual(rc, 0, rep)
self.assertEqual(osapply.Kernel.default_kver(f.cfg), NEW)
self.assertEqual(rep["kernel"]["phase"], "good")
def test_a_panic_falls_back_and_is_recorded(self):
f = staged()
kmode(f, "kernel-reboot")
rc, rep = self.boot_into(f, OLD)
self.assertEqual((rc, rep["kernel_event"]), (0, "fell_back"), rep)
self.assertEqual(osapply.Kernel.default_kver(f.cfg), OLD)
self.refused(f, "R22", mode="kernel-good")
def test_one_self_revert_then_never_again(self):
f = staged()
kmode(f, "kernel-reboot")
self.boot_into(f, NEW)
rc, rep = kmode(f, "kernel-revert", reason="the customer guest is not running")
self.assertEqual(rc, 0, rep)
self.assertEqual(f.reboots, [OLD, NEW])
self.assertEqual(osapply.Kernel.default_kver(f.cfg), OLD, "the self-revert boots the default, the old kernel")
rc, rep = self.boot_into(f, OLD)
self.assertEqual(rep["kernel_event"], "self_reverted")
# the same step can never revert again
st = json.loads(f.tree[osapply.KERNEL_STATE])
st["phase"] = "judging"
f.tree[osapply.KERNEL_STATE] = json.dumps(st)
f.running = NEW
self.refused(f, "R22", mode="kernel-revert")
self.assertEqual(len(f.reboots), 2)
def test_a_revert_that_comes_back_new_is_never_retried(self):
f = staged()
kmode(f, "kernel-reboot")
self.boot_into(f, NEW)
kmode(f, "kernel-revert")
rc, rep = self.boot_into(f, NEW)
self.assertEqual(rep["kernel_event"], "revert_failed")
self.refused(f, "R22", mode="kernel-revert")
def test_revert_refuses_an_unknown_default(self):
f = staged()
kmode(f, "kernel-reboot")
self.boot_into(f, NEW)
f.tree[osapply.KERNEL_DEFAULT_CFG] = f'GRUB_DEFAULT="gnulinux-advanced-{U}>gnulinux-9.9.9-1-pve-advanced-{U}"\n'
f.cfg = f.render()
self.refused(f, "R20", mode="kernel-revert")
self.assertEqual(len(f.reboots), 1)
def test_cancel_clears_the_flag(self):
f = staged()
rc, rep = kmode(f, "kernel-cancel")
self.assertEqual(rc, 0, rep)
self.assertFalse(f.env.get("felhom_next"), "the flag is gone")
self.assertEqual(rep["kernel"]["phase"], "cancelled")
def test_status_is_read_only_and_names_setup_problems(self):
f = kfake(esp_fs="ext4")
rc, rep = kmode(f, "kernel-status")
self.assertEqual(rc, 0, rep)
self.assertTrue(rep["kernel"]["setup_problems"])
self.assertEqual(f.tree, {})
self.assertFalse([c for c in f.calls if c[1][0] in ("update-grub", "systemctl", "apt-get")])
def test_facts_carry_the_kernel_lane(self):
f = staged()
f.plan = {"release_id": "facts", "layer": "host", "lane": "fast", "vmid": 9201, "mode": "facts", "packages": []}
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
h = rep["facts"]["host"]
self.assertEqual(h["kernel_lane"]["phase"], "staged")
self.assertEqual(h["kernel_next_boot"], NEW)
self.assertIn("one-shot", h["kernel_next_boot_source"])
if __name__ == "__main__":
unittest.main()
+60
View File
@@ -56,6 +56,18 @@ class FakeHost:
def log(self, line): def log(self, line):
self.logs.append(line) self.logs.append(line)
# controller-image (R-861 (a) A1): the ref arrives on stdin and is written INSIDE the guest by root.
stdin = b""
guest_writes = None
def read_stdin(self, limit):
return self.stdin[:limit + 1]
def write_guest_image(self, vmid, data):
if self.guest_writes is None:
self.guest_writes = []
self.guest_writes.append((vmid, data))
LOCAL_UNIT = """# Managed by felhom-agent — do not edit by hand. LOCAL_UNIT = """# Managed by felhom-agent — do not edit by hand.
[Unit] [Unit]
@@ -316,5 +328,53 @@ class Refuses(unittest.TestCase):
self.refused(h, ["wg", "/etc/shadow"], "A1") self.refused(h, ["wg", "/etc/shadow"], "A1")
class ControllerImage(unittest.TestCase):
"""R-861 (a) A1 (`09` §3 decision 165): the agent can no longer `tee` any image ref into the guest. The root verb
reads the ref on stdin, requires our registry + our repository + an x.y.z tag, and writes the guest file itself.
RED-PROOF: on the pre-A1 wrapper `controller-image` is not a verb (A1 usage, rc 2) — the accepted case fails."""
def go(self, ref, *argv):
h = FakeHost()
h.stdin = ref.encode() if isinstance(ref, str) else ref
return h, run(h, *(argv or ("controller-image", "9201")))
def test_our_controller_ref_is_written_in_the_guest(self):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0\n")
self.assertEqual(rc, 0, h.logs)
self.assertEqual(h.guest_writes, [(9201, b"gitea.dooplex.hu/admin/felhom-controller:0.301.0\n")])
def test_a_foreign_image_is_refused(self):
for ref in ("docker.io/library/alpine:latest\n", "alpine\n",
"gitea.dooplex.hu/admin/felhom-controller:latest\n",
"gitea.dooplex.hu/admin/other:0.1.0\n",
"evil.example/admin/felhom-controller:0.301.0\n",
"gitea.dooplex.hu/admin/felhom-controller:0.301.0\nalpine\n",
"gitea.dooplex.hu/admin/felhom-controller:0.301.0 x\n",
"", "\n"):
h, rc = self.go(ref)
self.assertEqual(rc, 3, f"{ref!r} was accepted")
self.assertFalse(h.guest_writes, f"{ref!r} wrote the guest file")
self.assertTrue(any("[I1]" in l for l in h.logs), h.logs)
def test_oversize_stdin_is_refused(self):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0" + " " * 300)
self.assertEqual(rc, 3)
self.assertFalse(h.guest_writes)
def test_vmid_must_be_numeric(self):
for argv in (("controller-image", "9201;id"), ("controller-image", "-1"), ("controller-image",),
("controller-image", "9201", "9202")):
h, rc = self.go("gitea.dooplex.hu/admin/felhom-controller:0.301.0\n", *argv)
self.assertIn(rc, (2, 3), argv)
self.assertFalse(h.guest_writes, argv)
def test_self_check_names_the_verb(self):
import io, contextlib
buf = io.StringIO()
with contextlib.redirect_stdout(buf):
pa.main(["--self-check"])
self.assertIn("controller-image", buf.getvalue())
if __name__ == "__main__": if __name__ == "__main__":
unittest.main(verbosity=2) unittest.main(verbosity=2)
@@ -241,3 +241,23 @@ func TestNewestArchiveTime_DistinctPhantomsEachAnnounced(t *testing.T) {
t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String()) t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String())
} }
} }
// R-99 (`09` §3 decision 140): the WARN for a PBS phantom ends with the cleanup runbook, so whoever sees it knows the
// one sanctioned way to remove it; a tiny archive on a dir storage is not a PBS phantom and gets no pointer.
// RED-PROOF: drop the `msg += phantomCleanupPointer` line → "the PBS phantom WARN does not end with the runbook pointer".
func TestRejectedArchiveWarnNamesTheCleanupRunbook(t *testing.T) {
var buf bytes.Buffer
r := runnerWithContent(t, &buf, []proxmox.StorageContent{phantomEntry(), goodPBSEntry()})
if _, _, err := r.NewestArchiveTime(context.Background(), 9201); err != nil {
t.Fatal(err)
}
const want = "INCOMPLETE archive when computing tier freshness — it is not a successful backup — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
if !strings.Contains(buf.String(), want) {
t.Errorf("the PBS phantom WARN does not end with the runbook pointer:\n%s", buf.String())
}
local := phantomEntry()
local.Format, local.VolID = "tar.zst", "local:backup/vzdump-lxc-9201-2026_07_28-05_31_14.tar.zst"
if got := rejectedArchiveMessage(local); strings.Contains(got, "pbs-phantom-cleanup") {
t.Errorf("a dir-storage archive got the PBS runbook pointer: %s", got)
}
}
+154
View File
@@ -0,0 +1,154 @@
package backup
import (
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// BackupSuccessState persists the newest SUCCESSFUL whole-guest backup per tier and guest (R-894).
//
// Why it exists. The due-check (`localapi` handleBackupDue) asks the tier's storage when a backup last
// landed (R-84) and falls back to the in-memory record when the storage cannot be read. The in-memory
// record is empty after an agent restart (Store, R-348), so "storage unreadable" right after a restart
// read as "no record — DUE". Measured 2026-10-05 on demo-hp: the agent restarted at 04:57, the off-site
// storage answered "Can't connect" at 06:25, the 7-day tier — last copy 2026-10-01 — read DUE, the
// controller asked, and vzdump failed. This file is the last known copy the fallback reads instead.
//
// It is read ONLY when the storage cannot be read. A storage that answers is the ground truth and wins,
// in both directions: an archive found there counts, and an archive absent there is absent even when
// this file remembers a success (a pruned or deleted archive must make the tier due — the same reason
// R-84 chose the storage over a persisted record). Pinned by
// TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers.
//
// Only SUCCESSES are written (the RestoreTestState rule): a failure must stay due and be retried, so a
// record of a failure has no reader.
type BackupSuccessState struct {
path string
mu sync.Mutex
last map[string]savedSuccess // key(target, vmid) → the newest success
}
type savedSuccess struct {
target string
vmid int
at time.Time
}
// backupSuccessJSON is one entry on disk.
type backupSuccessJSON struct {
Target string `json:"target"`
VMID int `json:"vmid"`
StartedAt string `json:"started_at"`
}
func backupStateKey(target string, vmid int) string { return target + "/" + strconv.Itoa(vmid) }
// NewBackupSuccessState opens (or creates) the state at path. A missing or unreadable file degrades to
// "nothing known" — the pre-R-894 behaviour, which is DUE — and never wedges the daemon.
func NewBackupSuccessState(path string) *BackupSuccessState {
s := &BackupSuccessState{path: path, last: map[string]savedSuccess{}}
data, err := os.ReadFile(path)
if err != nil {
return s
}
var entries []backupSuccessJSON
if json.Unmarshal(data, &entries) != nil {
return s
}
for _, e := range entries {
t, perr := time.Parse(time.RFC3339, e.StartedAt)
if perr != nil {
continue // one unreadable entry must not lose the others
}
s.last[backupStateKey(e.Target, e.VMID)] = savedSuccess{target: e.Target, vmid: e.VMID, at: t.UTC()}
}
return s
}
// RecordBackupSuccess saves b when it is a success newer than the one on file. target is the tier the
// job ran on (the due-check's key); a failure or an unparseable time is ignored.
func (s *BackupSuccessState) RecordBackupSuccess(target string, b hub.Backup) error {
if s == nil || !b.Success {
return nil
}
t, err := time.Parse(time.RFC3339, b.StartedAt)
if err != nil {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
k := backupStateKey(target, b.VMID)
if old, ok := s.last[k]; ok && !t.After(old.at) {
return nil
}
s.last[k] = savedSuccess{target: target, vmid: b.VMID, at: t.UTC()}
return s.saveLocked()
}
// LastKnownSuccess returns the newest saved success for this tier and guest (ok=false = none on file).
func (s *BackupSuccessState) LastKnownSuccess(target string, vmid int) (time.Time, bool) {
if s == nil {
return time.Time{}, false
}
s.mu.Lock()
defer s.mu.Unlock()
e, ok := s.last[backupStateKey(target, vmid)]
return e.at, ok
}
// KnownBackupSuccesses returns every saved success as a host-report record (the collector's
// KnownBackupReporter, so the hub keeps its evidence across a restart). Only the tier, guest and start
// time are known; the archive name is not saved and stays empty. Sorted for a stable report.
func (s *BackupSuccessState) KnownBackupSuccesses() []hub.Backup {
if s == nil {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
out := make([]hub.Backup, 0, len(s.last))
for _, e := range s.last {
out = append(out, hub.Backup{TargetID: e.target, VMID: e.vmid, Success: true, CrashConsistent: true,
StartedAt: e.at.Format(time.RFC3339), UncoveredVolumes: []string{}})
}
sort.Slice(out, func(i, j int) bool {
if out[i].TargetID != out[j].TargetID {
return out[i].TargetID < out[j].TargetID
}
return out[i].VMID < out[j].VMID
})
return out
}
func (s *BackupSuccessState) saveLocked() error {
entries := make([]backupSuccessJSON, 0, len(s.last))
for _, e := range s.last {
entries = append(entries, backupSuccessJSON{Target: e.target, VMID: e.vmid, StartedAt: e.at.Format(time.RFC3339)})
}
// Deterministic file content (Go's map order is random).
sort.Slice(entries, func(i, j int) bool {
if entries[i].Target != entries[j].Target {
return entries[i].Target < entries[j].Target
}
return entries[i].VMID < entries[j].VMID
})
data, err := json.MarshalIndent(entries, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(s.path), 0o755); err != nil {
return err
}
tmp := s.path + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, s.path)
}
+76
View File
@@ -0,0 +1,76 @@
package backup
import (
"os"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-894: the on-disk newest success per tier survives a restart (a new state from the same file).
func TestBackupSuccessState_SurvivesRestart(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
s := NewBackupSuccessState(path)
at := time.Date(2026, 10, 1, 20, 15, 0, 0, time.UTC)
if err := s.RecordBackupSuccess("felhom-pbs", hub.Backup{VMID: 9201, Success: true, StartedAt: at.Format(time.RFC3339)}); err != nil {
t.Fatal(err)
}
got, ok := NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 9201)
if !ok || !got.Equal(at) {
t.Fatalf("after a restart the saved copy must read back; got %v ok=%v", got, ok)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("local", 9201); ok {
t.Fatal("another tier must not borrow this tier's copy")
}
}
// Only a NEWER success replaces the saved one; failures and unparseable times are ignored.
func TestBackupSuccessState_KeepsNewestSuccessOnly(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
s := NewBackupSuccessState(path)
newer := time.Date(2026, 10, 5, 0, 0, 0, 0, time.UTC)
older := newer.Add(-48 * time.Hour)
for _, b := range []hub.Backup{
{VMID: 1, Success: true, StartedAt: newer.Format(time.RFC3339)},
{VMID: 1, Success: true, StartedAt: older.Format(time.RFC3339)}, // older: ignored
{VMID: 1, Success: false, StartedAt: newer.Add(time.Hour).Format(time.RFC3339)}, // failure: ignored
{VMID: 1, Success: true, StartedAt: "not-a-time"}, // unparseable: ignored
} {
if err := s.RecordBackupSuccess("t", b); err != nil {
t.Fatal(err)
}
}
if got, _ := NewBackupSuccessState(path).LastKnownSuccess("t", 1); !got.Equal(newer) {
t.Fatalf("want the newest success %v, got %v", newer, got)
}
}
// A corrupt file degrades to "nothing known" (the pre-R-894 DUE answer), never a crash.
func TestBackupSuccessState_CorruptFileIsEmpty(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
if err := os.WriteFile(path, []byte("{not json"), 0o600); err != nil {
t.Fatal(err)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("t", 1); ok {
t.Fatal("a corrupt file must read as nothing known")
}
}
// The saved successes come back as host-report records: success, tier, guest, start time (R-894 → the
// host report, 2026-10-09). A new state read from the same file gives the same records (the restart case).
func TestBackupSuccessState_KnownBackupSuccessesSurviveAReopen(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
s := NewBackupSuccessState(path)
if err := s.RecordBackupSuccess("felhom-backup", hub.Backup{VMID: 9201, Success: true, StartedAt: "2026-10-09T02:40:16Z"}); err != nil {
t.Fatal(err)
}
got := NewBackupSuccessState(path).KnownBackupSuccesses()
if len(got) != 1 || !got[0].Success || got[0].TargetID != "felhom-backup" || got[0].VMID != 9201 || got[0].StartedAt != "2026-10-09T02:40:16Z" {
t.Fatalf("got %+v", got)
}
if got[0].UncoveredVolumes == nil {
t.Fatal("uncovered_volumes must marshal as [], not null")
}
}
+60
View File
@@ -0,0 +1,60 @@
package backup
import (
"context"
"sort"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// ForeignKeyLedger (R-366 slice 2, `09` §3 decision 168) holds, per backup tier, the whole-guest archives the
// restore-test pick skipped because another key wrote them (R-727 — an earlier install of this box). Since R-727 the
// skip was one INFO log line per archive and nothing else, so after a reinstall the operator was never told that the
// box's older whole-guest copies are unreadable to it. The host report carries this ledger; the hub raises ONE
// operator event when it changes.
//
// It reports nil until a tier has been evaluated since the agent started, so a restart does not read as "the set
// changed to empty" (the hub keeps its last state for an absent field).
type ForeignKeyLedger struct {
mu sync.Mutex
byTarget map[string]hub.ForeignKeyArchives
}
// NewForeignKeyLedger builds an empty ledger.
func NewForeignKeyLedger() *ForeignKeyLedger {
return &ForeignKeyLedger{}
}
func (l *ForeignKeyLedger) set(target string, n int, oldest, newest int64) {
l.mu.Lock()
defer l.mu.Unlock()
if l.byTarget == nil {
l.byTarget = map[string]hub.ForeignKeyArchives{}
}
e := hub.ForeignKeyArchives{Target: target, Count: n}
if n > 0 {
e.Oldest = time.Unix(oldest, 0).UTC().Format(time.RFC3339)
e.Newest = time.Unix(newest, 0).UTC().Format(time.RFC3339)
}
l.byTarget[target] = e
}
// ForeignKeyArchives implements hub.ForeignKeyArchiveReporter: nil before any evaluation; otherwise the tiers that
// hold such archives (`Tiers` empty, never nil, when none do), sorted by tier.
func (l *ForeignKeyLedger) ForeignKeyArchives(context.Context) *hub.ForeignKeyArchivesStanza {
l.mu.Lock()
defer l.mu.Unlock()
if l.byTarget == nil {
return nil
}
out := []hub.ForeignKeyArchives{}
for _, e := range l.byTarget {
if e.Count > 0 {
out = append(out, e)
}
}
sort.Slice(out, func(i, j int) bool { return out[i].Target < out[j].Target })
return &hub.ForeignKeyArchivesStanza{Tiers: out}
}
@@ -0,0 +1,69 @@
package backup
import (
"context"
"encoding/json"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-366 slice 2 (`09` §3 decision 168) — the restore-test's skip of an archive written with another key stops being
// silent: the pick records, per tier, how many it skipped and their time range, and the host report carries it.
//
// COMPANION RED-PROOF (observed): remove the `r.foreign.set(...)` call from PickSettledRestoreCandidateOn → this fails
// with "after one evaluation the ledger must report felhom-pbs: 2 archives …; got []". Restored.
func TestR366_PickRecordsArchivesWrittenWithAnotherKey(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T17:27:32Z", Content: "backup", VMID: 9201, Size: 4774114206, CTime: 1789579652, Encrypted: earlierBox2},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z", Content: "backup", VMID: 9201, Size: 20811501236, CTime: 1789595994, Encrypted: earlierBox1},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey},
},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
l := NewForeignKeyLedger()
r.SetForeignKeyLedger(l)
if got := l.ForeignKeyArchives(context.Background()); got != nil {
t.Fatalf("before any evaluation the ledger must be nil (the hub keeps its state); got %v", got)
}
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err != nil {
t.Fatal(err)
}
st := l.ForeignKeyArchives(context.Background())
want := hub.ForeignKeyArchives{Target: "felhom-pbs", Count: 2, Oldest: "2026-09-16T17:27:32Z", Newest: "2026-09-16T21:59:54Z"}
var got []hub.ForeignKeyArchives
if st != nil {
got = st.Tiers
}
if len(got) != 1 || got[0] != want {
t.Fatalf("after one evaluation the ledger must report felhom-pbs: 2 archives 2026-09-16T17:27:32Z…21:59:54Z; got %v", got)
}
}
// Evaluated and none found → the stanza with `tiers: []`; not evaluated → no stanza at all. Never a null on the wire.
func TestR366_EvaluatedWithNoneIsAnEmptyList(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey}},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
l := NewForeignKeyLedger()
r.SetForeignKeyLedger(l)
b, _ := json.Marshal(hub.HostReport{ForeignKeyArchives: l.ForeignKeyArchives(context.Background())})
if strings.Contains(string(b), "foreign_key_archives") {
t.Fatalf("not evaluated must omit the stanza; got %s", b)
}
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err != nil {
t.Fatal(err)
}
b, _ = json.Marshal(hub.HostReport{ForeignKeyArchives: l.ForeignKeyArchives(context.Background())})
if !strings.Contains(string(b), `"foreign_key_archives":{"tiers":[]}`) {
t.Fatalf("evaluated with none must report tiers: []; got %s", b)
}
}
+33 -1
View File
@@ -66,8 +66,13 @@ type BackupRunner struct {
// due-check is served from the local-API handler goroutines. // due-check is served from the local-API handler goroutines.
rejectedMu sync.Mutex rejectedMu sync.Mutex
rejected map[string]struct{} rejected map[string]struct{}
// foreign (R-366 slice 2) records, per tier, the archives the pick skipped as another key's. nil = not wired.
foreign *ForeignKeyLedger
} }
// SetForeignKeyLedger wires the R-366 slice-2 ledger the host report reads.
func (r *BackupRunner) SetForeignKeyLedger(l *ForeignKeyLedger) { r.foreign = l }
// NewBackupRunner builds a runner. mode defaults to snapshot (works for a stopped guest and // NewBackupRunner builds a runner. mode defaults to snapshot (works for a stopped guest and
// for lvm-thin); the caller may pass ModeStop for storages without snapshot support. retention is the // for lvm-thin); the caller may pass ModeStop for storages without snapshot support. retention is the
// per-run prune spec ("keep-last=N", or "" to never prune) — only the periodic local backup sets it. // per-run prune spec ("keep-last=N", or "" to never prune) — only the periodic local backup sets it.
@@ -370,6 +375,8 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
var best string var best string
var bestCTime int64 = -1 var bestCTime int64 = -1
known := map[int]bool{} // vmid → the guest exists on this node (asked once per vmid per pick) known := map[int]bool{} // vmid → the guest exists on this node (asked once per vmid per pick)
var foreignN int // R-366 slice 2: archives skipped as another key's, and their time range
var foreignMin, foreignMax int64
for _, e := range contents { for _, e := range contents {
if e.Content != "backup" { if e.Content != "backup" {
continue continue
@@ -384,6 +391,13 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
} }
if ownKey != "" && !strings.EqualFold(e.Encrypted, ownKey) { if ownKey != "" && !strings.EqualFold(e.Encrypted, ownKey) {
r.noteNotAGuestBackupOnce(e, fmt.Sprintf("written by another box (key %s, this box's key %s) — not this box's proof", shortFP(e.Encrypted), shortFP(ownKey))) r.noteNotAGuestBackupOnce(e, fmt.Sprintf("written by another box (key %s, this box's key %s) — not this box's proof", shortFP(e.Encrypted), shortFP(ownKey)))
foreignN++
if foreignMin == 0 || e.CTime < foreignMin {
foreignMin = e.CTime
}
if e.CTime > foreignMax {
foreignMax = e.CTime
}
continue continue
} }
// R-689 (v0.136.0): … OF A GUEST THAT STILL EXISTS here. Measured on demo-hp 2026-09-27 right after // R-689 (v0.136.0): … OF A GUEST THAT STILL EXISTS here. Measured on demo-hp 2026-09-27 right after
@@ -420,6 +434,9 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
bestCTime, best = e.CTime, e.VolID bestCTime, best = e.CTime, e.VolID
} }
} }
if r.foreign != nil {
r.foreign.set(target, foreignN, foreignMin, foreignMax)
}
if best == "" { if best == "" {
return "", time.Time{}, nil return "", time.Time{}, nil
} }
@@ -518,10 +535,25 @@ func (r *BackupRunner) warnRejectedArchiveOnce(e proxmox.StorageContent, why str
if seen { if seen {
return return
} }
r.logger.Warn("backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup", r.logger.Warn(rejectedArchiveMessage(e),
"target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why) "target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
} }
// phantomCleanupPointer names the runbook that removes a PBS phantom (R-99, `09` §3 decision 140: a leftover of an
// aborted upload is deleted on the backup server, by a runbook, when one is seen — never automatically).
const phantomCleanupPointer = " — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
// rejectedArchiveMessage is the WARN text for a rejected archive. Only a PBS entry (format pbs-ct / pbs-vm) gets the
// cleanup pointer: the runbook deletes on a PBS datastore, and a tiny archive on a dir storage is not a PBS phantom.
// Pinned by TestRejectedArchiveWarnNamesTheCleanupRunbook.
func rejectedArchiveMessage(e proxmox.StorageContent) string {
msg := "backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup"
if strings.HasPrefix(e.Format, "pbs-") {
msg += phantomCleanupPointer
}
return msg
}
// demo-felhom in a single afternoon of deploys (2026-07-26). // demo-felhom in a single afternoon of deploys (2026-07-26).
// //
// Asking the STORAGE rather than persisting the store is deliberate: // Asking the STORAGE rather than persisting the store is deliberate:
+3
View File
@@ -30,6 +30,9 @@ import (
// 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What // 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What
// is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/ // is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/
// deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84). // deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84).
// The due-check's fallback for an UNREADABLE storage no longer reads this store alone (R-894): the newest
// success per tier is also on disk (BackupSuccessState), so a restart followed by an unreachable storage
// reads the last known copy, not "never".
type Store struct { type Store struct {
mu sync.Mutex mu sync.Mutex
byTarget map[string]hub.Backup // latest backup per target id byTarget map[string]hub.Backup // latest backup per target id
+6 -1
View File
@@ -145,12 +145,17 @@ var manifest = []Capability{
{"controllerswap-image-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "image", "inspect", "gitea.dooplex.hu/admin/felhom-controller:0.0.0"}, true, ""}, {"controllerswap-image-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "image", "inspect", "gitea.dooplex.hu/admin/felhom-controller:0.0.0"}, true, ""},
{"controllerswap-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "inspect", "-f", "{{.State.Running}}", "felhom-controller"}, true, ""}, {"controllerswap-inspect", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "docker", "inspect", "-f", "{{.State.Running}}", "felhom-controller"}, true, ""},
{"controllerswap-restart", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "systemctl", "restart", "felhom-controller-bootstrap.service"}, true, ""}, {"controllerswap-restart", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "systemctl", "restart", "felhom-controller-bootstrap.service"}, true, ""},
{"controllerswap-write", "controller-swap / managed auto-update", "/usr/sbin/pct", []string{"exec", "9201", "--", "tee", "/etc/felhom-controller-image"}, true, ""}, // R-861 (a) A1 (decision 165): the write goes through the ROOT verb that checks the ref; the agent has no `tee` grant.
{"controllerswap-write", "controller-swap / managed auto-update", "/usr/local/sbin/felhom-priv-apply", []string{"controller-image", "9201"}, true, ""},
// ---- Stale-lock recovery (FELHOM_STALELOCK, v0.49.0; Critical: a guest stuck behind a stale // ---- Stale-lock recovery (FELHOM_STALELOCK, v0.49.0; Critical: a guest stuck behind a stale
// reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ---- // reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ----
{"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""}, {"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""},
// ---- Weekly guest disk trim (FELHOM_FSTRIM, R-444). NON-critical: a missing grant means the thin pool is not
// reclaimed this week (the trim job WARNs per guest and the report shows the failure), not a serving outage. ----
{"guest-fstrim", "weekly guest disk trim (thin-pool reclaim, R-444)", "/usr/sbin/pct", []string{"fstrim", "9201"}, false, ""},
// ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite // ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite
// backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the // backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the
// conf install, unit enable/restart and the handshake read gate the backup path. apt-install // conf install, unit enable/restart and the handshake read gate the backup path. apt-install
+4 -3
View File
@@ -200,7 +200,8 @@ func TestRedProof_DroppedGrantFailsCheck(t *testing.T) {
} }
// TestRedProof_DroppedControllerSwapTeeFailsCheck is the companion red-proof for the v0.45.0 // TestRedProof_DroppedControllerSwapTeeFailsCheck is the companion red-proof for the v0.45.0
// FELHOM_CONTROLLERSWAP grants: with the `tee /etc/felhom-controller-image` line removed, the // FELHOM_CONTROLLERSWAP grants: with the write grant removed (since R-861 (a) A1 the `felhom-priv-apply controller-image`
// line; before it, an agent `tee /etc/felhom-controller-image`), the
// controllerswap-write capability MUST be reported uncovered. Proves the build gate watches the new // controllerswap-write capability MUST be reported uncovered. Proves the build gate watches the new
// swap write grant (so dropping it can't ship a non-root agent that silently can't auto-update). // swap write grant (so dropping it can't ship a non-root agent that silently can't auto-update).
func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) { func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
@@ -210,7 +211,7 @@ func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
} }
var kept []string var kept []string
for _, ln := range strings.Split(string(data), "\n") { for _, ln := range strings.Split(string(data), "\n") {
if strings.Contains(ln, "tee /etc/felhom-controller-image") { if strings.Contains(ln, "felhom-priv-apply ^controller-image") { // R-861 (a) A1: the write's grant
continue continue
} }
kept = append(kept, ln) kept = append(kept, ln)
@@ -229,7 +230,7 @@ func TestRedProof_DroppedControllerSwapTeeFailsCheck(t *testing.T) {
} }
cmdline := write.Binary + " " + strings.Join(write.ReprArgs, " ") cmdline := write.Binary + " " + strings.Join(write.ReprArgs, " ")
if matchesAny(cmdline, entries) { if matchesAny(cmdline, entries) {
t.Errorf("red-proof FAILED: controllerswap-write still matches after dropping the tee grant") t.Errorf("red-proof FAILED: controllerswap-write still matches after dropping its grant")
} }
if full := parseSudoersEntries(t, string(data)); !matchesAny(cmdline, full) { if full := parseSudoersEntries(t, string(data)); !matchesAny(cmdline, full) {
t.Errorf("controllerswap-write should be covered by the real sudoers") t.Errorf("controllerswap-write should be covered by the real sudoers")
@@ -49,6 +49,10 @@ var r861Injections = []string{
"/usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount", "/usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount",
"/usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf", "/usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf",
"/usr/local/sbin/felhom-priv-apply wg /etc/shadow", "/usr/local/sbin/felhom-priv-apply wg /etc/shadow",
// R-861 (a) A1 (decision 165): the agent wrote ANY image ref into the guest by `tee` — now only the root verb may
"/usr/sbin/pct exec 9201 -- tee /etc/felhom-controller-image",
"/usr/local/sbin/felhom-priv-apply controller-image 9201 9202",
"/usr/local/sbin/felhom-priv-apply controller-image 9201;id",
} }
func TestSudoersRefusesTheR861Injections(t *testing.T) { func TestSudoersRefusesTheR861Injections(t *testing.T) {
@@ -63,3 +67,72 @@ func TestSudoersRefusesTheR861Injections(t *testing.T) {
} }
} }
} }
// R-444: the weekly trim's grant is ONE exact shape — `pct fstrim <vmid>` — and nothing smuggled after it.
// The manifest entry (guest-fstrim) proves the real call is still allowed (TestManifestCoveredBySudoers); this
// pins the other direction. RED-PROOF: write the rule as the glob `/usr/sbin/pct fstrim [0-9]*` → every decoy
// below with a trailing argument matches (the glob's `*` eats spaces).
func TestSudoersFstrimRuleIsExact(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
if !matchesAny("/usr/sbin/pct fstrim 9201", entries) {
t.Fatal("the sudoers does not allow `pct fstrim 9201` — the weekly trim cannot run")
}
for _, c := range []string{
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints",
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints 1",
"/usr/sbin/pct fstrim 9201; x",
"/usr/sbin/pct fstrim 9201 9202",
"/usr/sbin/pct fstrim 92a1",
"/usr/sbin/pct fstrim ",
"/usr/sbin/pct fstrim -- 9201",
"/usr/sbin/pct destroy 9201",
"/usr/sbin/pct destroy 9201 --purge",
} {
if matchesAny(c, entries) {
t.Errorf("the sudoers allows a command the trim rule must not: %q", c)
}
}
}
// R-861 (a) A1: the managed controller update still has its route — the root verb, one numeric vmid.
func TestSudoersAllowsTheControllerImageVerb(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
if !matchesAny("/usr/local/sbin/felhom-priv-apply controller-image 9201", parseSudoersEntries(t, string(data))) {
t.Fatal("the sudoers does not allow `felhom-priv-apply controller-image 9201` — a managed controller update cannot write its image")
}
}
// R-861 (b) B2 (decision 165, hygiene): felhom-op's `pct start|stop|unlock` grants are ONE numeric vmid each. The old
// glob `[0-9]*` eats spaces, so `pct stop 9201 --skiplock 1` and two vmids matched.
// RED-PROOF: on the pre-B2 felhom-op.sudoers (`/usr/sbin/pct stop [0-9]*`) the decoys match.
func TestFelhomOpSudoersPctIsExact(t *testing.T) {
data, err := os.ReadFile("../../configs/felhom-op.sudoers")
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
for _, ok := range []string{"/usr/sbin/pct start 9201", "/usr/sbin/pct stop 9201", "/usr/sbin/pct unlock 9201", "/usr/sbin/pct list"} {
if !matchesAny(ok, entries) {
t.Errorf("felhom-op lost a repair verb: %s", ok)
}
}
for _, bad := range []string{
"/usr/sbin/pct stop 9201 --skiplock 1",
"/usr/sbin/pct start 9201 9202",
"/usr/sbin/pct unlock 9201 --whatever",
"/usr/sbin/pct start 92a1",
"/usr/sbin/pct stop ",
"/usr/sbin/pct destroy 9201",
} {
if matchesAny(bad, entries) {
t.Errorf("felhom-op's sudoers allows %q", bad)
}
}
}
+11
View File
@@ -33,6 +33,7 @@ type Config struct {
LANResolver LANResolverConfig `json:"lan_resolver"` LANResolver LANResolverConfig `json:"lan_resolver"`
WGTunnel WGTunnelConfig `json:"wg_tunnel"` WGTunnel WGTunnelConfig `json:"wg_tunnel"`
GuestNet GuestNetConfig `json:"guest_net"` GuestNet GuestNetConfig `json:"guest_net"`
DiskTrim DiskTrimConfig `json:"disk_trim"`
OOB OOBConfig `json:"oob"` OOB OOBConfig `json:"oob"`
SelfUpdate SelfUpdateConfig `json:"selfupdate"` SelfUpdate SelfUpdateConfig `json:"selfupdate"`
LogLevel string `json:"log_level"` // debug|info|warn|error (default info) LogLevel string `json:"log_level"` // debug|info|warn|error (default info)
@@ -139,6 +140,16 @@ func (w WGTunnelConfig) WithDefaults() WGTunnelConfig {
return w return w
} }
// DiskTrimConfig configures the R-444 weekly guest disk trim (internal/fstrim). DEFAULT-ON, like GuestNetConfig and for
// the same reason: it only acts on guests the agent already owns, and the operator ruled every box trims (`09` §3
// decision 139). Opting out is the explicit act: `"disk_trim": {"disable": true}`.
type DiskTrimConfig struct {
Disable bool `json:"disable"`
}
// Enabled reports whether the weekly trim should run.
func (d DiskTrimConfig) Enabled() bool { return !d.Disable }
// GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet). // GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet).
// //
// **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other // **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other
+50 -10
View File
@@ -59,8 +59,26 @@ var (
// It carries no material and no code: only WHICH earlier package opened, by its supersession date, // It carries no material and no code: only WHICH earlier package opened, by its supersession date,
// which is the one fact the customer needs to recognise it. // which is the one fact the customer needs to recognise it.
ErrCodeOpensRetained = errors.New("escrow: the recovery code did not open the CURRENT sealed package, but it DID open a retained earlier one") ErrCodeOpensRetained = errors.New("escrow: the recovery code did not open the CURRENT sealed package, but it DID open a retained earlier one")
// ErrRetainedUnchecked — the code did NOT open the current package, and NOT every earlier package the hub
// holds for this host was tried (R-304, 2026-10-08): the retained list could not be fetched, the hub
// withheld rows (over its serve cap, or rows with no key material), a package was malformed, or the
// attempt cap stopped the loop. So „the code is wrong" is NOT known — it may be right for a package
// nobody tried. Distinct from a mistype for exactly the R-224 reason: never accuse the customer of
// something we did not check.
ErrRetainedUnchecked = errors.New("escrow: the recovery code did not open the current sealed package, and some earlier packages were NOT checked")
) )
// RetainedUncheckedError wraps ErrRetainedUnchecked with how many earlier packages went unchecked
// (-1 = unknown: the retained list itself could not be read). No secret.
type RetainedUncheckedError struct {
Unchecked int
}
func (e *RetainedUncheckedError) Error() string {
return fmt.Sprintf("%s (unchecked=%d)", ErrRetainedUnchecked.Error(), e.Unchecked)
}
func (e *RetainedUncheckedError) Unwrap() error { return ErrRetainedUnchecked }
// RetainedMatch says which retained package a code opened. Returned inside RetainedOpenedError; it // RetainedMatch says which retained package a code opened. Returned inside RetainedOpenedError; it
// carries no secret — not the code, not the bundle, not the repository password. // carries no secret — not the code, not the bundle, not the repository password.
type RetainedMatch struct { type RetainedMatch struct {
@@ -102,8 +120,10 @@ type RetainedBlob struct {
} }
// RetainedFetcher yields this host's RETAINED sealed packages, newest-superseded first. An empty // RetainedFetcher yields this host's RETAINED sealed packages, newest-superseded first. An empty
// slice is a clean "none". R-311. // slice is a clean "none". R-311. `withheld` counts the earlier packages the hub holds that are NOT in
type RetainedFetcher func(ctx context.Context) (blobs []RetainedBlob, unopenable int, err error) // blobs — rows with no key material, rows over the hub's serve cap, and packages dropped as malformed
// (R-304): each is a package the code was never tried against.
type RetainedFetcher func(ctx context.Context) (blobs []RetainedBlob, withheld int, err error)
// OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R. // OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R.
type OffsiteKeyRecoverer struct { type OffsiteKeyRecoverer struct {
@@ -154,9 +174,14 @@ func (r OffsiteKeyRecoverer) RecoverOffsiteRepoPassword(ctx context.Context, rec
// from the unwrap alone; the only way to tell is to try. Until this existed nobody tried, and // from the unwrap alone; the only way to tell is to try. Until this existed nobody tried, and
// the screen said so out loud ("innen nem tudjuk megkülönböztetni őket") — a true sentence // the screen said so out loud ("innen nem tudjuk megkülönböztetni őket") — a true sentence
// about our own incuriosity, read by the customer as a statement about their code. // about our own incuriosity, read by the customer as a statement about their code.
if m, ok := r.tryRetained(ctx, recoveryCode); ok { m, ok, unchecked := r.tryRetained(ctx, recoveryCode)
if ok {
return "", &RetainedOpenedError{Match: m} return "", &RetainedOpenedError{Match: m}
} }
// R-304: only when EVERY earlier package the hub holds was tried may this stay a wrong code.
if unchecked != 0 {
return "", &RetainedUncheckedError{Unchecked: unchecked}
}
return "", err // the fail-closed "the recovery code did not unwrap…" message; no secret in it return "", err // the fail-closed "the recovery code did not unwrap…" message; no secret in it
} }
if bundle.ResticRepoPassword == "" { if bundle.ResticRepoPassword == "" {
@@ -173,23 +198,38 @@ func (r OffsiteKeyRecoverer) RecoverOffsiteRepoPassword(ctx context.Context, rec
// function breaking is the behaviour we had before it existed. // function breaking is the behaviour we had before it existed.
// //
// NOTHING IS LOGGED HERE and no return value carries the code, a bundle or a password. // NOTHING IS LOGGED HERE and no return value carries the code, a bundle or a password.
func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode string) (RetainedMatch, bool) { //
// R-304 (2026-10-08): it also returns how many earlier packages were NOT tried — `withheld` from the hub, plus
// the ones past the attempt cap — or -1 when the retained list could not be read at all. A nil FetchRetained
// (an agent wired without the lookup) reports 0: the pre-R-311 refusal, unchanged.
func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode string) (RetainedMatch, bool, int) {
if r.FetchRetained == nil { if r.FetchRetained == nil {
return RetainedMatch{}, false return RetainedMatch{}, false, 0
} }
blobs, _, err := r.FetchRetained(ctx) blobs, withheld, err := r.FetchRetained(ctx)
if err != nil || len(blobs) == 0 { if err != nil {
return RetainedMatch{}, false return RetainedMatch{}, false, -1
}
if withheld < 0 {
withheld = 0
}
if len(blobs) == 0 {
return RetainedMatch{}, false, withheld
} }
limit := r.MaxRetainedTried limit := r.MaxRetainedTried
if limit <= 0 { if limit <= 0 {
limit = defaultMaxRetainedTried limit = defaultMaxRetainedTried
} }
unchecked := withheld
if len(blobs) > limit {
unchecked += len(blobs) - limit
}
for i, rb := range blobs { for i, rb := range blobs {
if i >= limit { if i >= limit {
break break
} }
if len(rb.Blob) == 0 { if len(rb.Blob) == 0 {
unchecked++
continue continue
} }
bundle, uerr := UnwrapIdentityBundle(ctx, rb.Blob, recoveryCode) bundle, uerr := UnwrapIdentityBundle(ctx, rb.Blob, recoveryCode)
@@ -204,7 +244,7 @@ func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode strin
// correct and must be told so — but the history behind it still cannot be reopened, and // correct and must be told so — but the history behind it still cannot be reopened, and
// saying otherwise would be a promise this path cannot keep. // saying otherwise would be a promise this path cannot keep.
HasResticPassword: bundle.ResticRepoPassword != "", HasResticPassword: bundle.ResticRepoPassword != "",
}, true }, true, 0
} }
return RetainedMatch{}, false return RetainedMatch{}, false, unchecked
} }
+76 -2
View File
@@ -115,8 +115,9 @@ func TestRecover_WrongCode_StaysAPlainRefusal(t *testing.T) {
} }
} }
// FAIL-SAFE — if the retained lookup itself fails, the original refusal must stand UNCHANGED. The // FAIL-SAFE — if the retained lookup itself fails, the lookup's error never reaches the customer and no
// worst outcome of this feature breaking is the behaviour we had before it. // retained package is claimed. R-304 (2026-10-08) changed what stands instead: not the wrong-code refusal (the
// earlier packages were never tried, so „wrong" is not known) but ErrRetainedUnchecked with Unchecked = -1.
// //
// RED-PROOF: make tryRetained propagate the fetch error instead of returning false → the customer // RED-PROOF: make tryRetained propagate the fetch error instead of returning false → the customer
// gets a new, unexplained failure mode → this FAILS. // gets a new, unexplained failure mode → this FAILS.
@@ -140,6 +141,10 @@ func TestRecover_RetainedFetchFails_OriginalRefusalStands(t *testing.T) {
if containsStr(err.Error(), "hub exploded") { if containsStr(err.Error(), "hub exploded") {
t.Error("the retained-lookup failure leaked into the customer-facing refusal — it must be silent") t.Error("the retained-lookup failure leaked into the customer-facing refusal — it must be silent")
} }
var ue *RetainedUncheckedError
if !errors.As(err, &ue) || ue.Unchecked != -1 {
t.Errorf("err = %v, want RetainedUncheckedError{-1} — the earlier packages were never tried (R-304)", err)
}
} }
// A nil FetchRetained keeps the pre-R-311 behaviour EXACTLY. An agent wired without it must be // A nil FetchRetained keeps the pre-R-311 behaviour EXACTLY. An agent wired without it must be
@@ -228,3 +233,72 @@ func containsStr(hay, needle string) bool {
return false return false
})() })()
} }
// ── R-304 (2026-10-08) — „wrong code" only when every earlier package was tried ──────────────────────
//
// The consequence asserted: is the refusal a WRONG CODE (the only error the screen may answer with „check your
// typing")? It may be only when the hub withheld nothing and every served package was tried.
//
// RED-PROOF: make tryRetained return 0 for `unchecked` → the three „unchecked" cases return the plain refusal
// → they FAIL; the all-tried case keeps passing.
func TestR304_WrongCodeOnlyWhenEveryEarlierPackageWasTried(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "4444567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
other := sealBundle(t, IdentityBundle{ResticRepoPassword: "5555567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
const code = "a code that opens nothing whatsoever in this test"
cases := []struct {
name string
blobs int
withheld int
limit int
unchecked int // 0 = must be a plain wrong code
}{
{"all tried, nothing withheld", 2, 0, 6, 0},
{"the hub withheld rows (no key material / over its cap)", 1, 3, 6, 3},
{"only withheld rows, nothing served", 0, 2, 6, 2},
{"the attempt cap stopped the loop", 5, 0, 2, 3},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
blobs := make([]RetainedBlob, 0, c.blobs)
for i := 0; i < c.blobs; i++ {
blobs = append(blobs, RetainedBlob{Blob: other, SupersededAt: "2026-08-01 00:00:00", Index: i})
}
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
return blobs, c.withheld, nil
},
MaxRetainedTried: c.limit,
}.RecoverOffsiteRepoPassword(context.Background(), code)
if err == nil {
t.Fatal("a code that opens nothing succeeded")
}
var ue *RetainedUncheckedError
isUnchecked := errors.As(err, &ue)
if c.unchecked == 0 {
if isUnchecked {
t.Fatalf("every earlier package was tried — this IS a wrong code, got %v", err)
}
return
}
if !isUnchecked || ue.Unchecked != c.unchecked || !errors.Is(err, ErrRetainedUnchecked) {
t.Fatalf("err = %v, want RetainedUncheckedError{%d} — never call the code wrong when packages were not tried", err, c.unchecked)
}
})
}
}
// A malformed (empty) served package counts as not tried.
func TestR304_EmptyServedPackageCountsAsUnchecked(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "6666567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: nil, SupersededAt: "2026-08-01 00:00:00"}),
}.RecoverOffsiteRepoPassword(context.Background(), testR2)
var ue *RetainedUncheckedError
if !errors.As(err, &ue) || ue.Unchecked != 1 {
t.Fatalf("err = %v, want RetainedUncheckedError{1}", err)
}
}
+346
View File
@@ -0,0 +1,346 @@
// Package fstrim is the weekly guest disk trim (R-444, operator ruling `09` §3 decision 139).
//
// Why: a thin pool only ever grows from blocks a guest has already FREED — `fstrim` inside the unprivileged container
// is refused (FITRIM: Operation not permitted), and nothing else on the box gives the blocks back. A full thin pool
// takes every guest on the host read-only, so the pool can reach 100 % from deleted data alone. Measured on demo-hp
// 2026-10-06 09:14Z: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %, 18/18 app probes 200, max 1.1 s
// (audits/ten-answers-2026-10-06/r444-measure.txt).
//
// The rule, each part pinned by a test in fstrim_test.go:
// - Weekly: a guest is DUE from Wednesday 10:00 local until it has been trimmed once since then (a box that was off
// on Wednesday catches up at its next eligible hour).
// - Daytime only: a trim starts only between 10:00 and 20:59 local — never in the night window (01:00–06:59) where
// the backups and the restore-tests run (TestEligibleHourNeverInTheNight).
// - Never beside a backup, a restore-test or another heavy operation: the pass holds the host-wide one-heavy-op gate
// (backup.InFlight) for its whole run; a busy gate DEFERS the pass to the next hourly tick.
// - A failed trim is retried at the next eligible hour, at most MaxAttemptsPerWeek times in one week.
// - The last result per guest (time, bytes, ok/fail) is persisted, so a restart neither loses it nor re-trims.
//
// The command is the ONE exact sudoers shape `pct fstrim <vmid>` (FELHOM_FSTRIM). Only guests from the pool-verified
// source (ListLXC ∩ the felhom pool, audit A1) and only RUNNING ones are trimmed.
package fstrim
import (
"context"
"encoding/json"
"fmt"
"log/slog"
"os"
"path/filepath"
"regexp"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// Schedule. The weekday/hours are fixed on purpose (one sentence the operator can read on the System page).
const (
Weekday = time.Wednesday
StartHour = 10 // first eligible local hour (inclusive)
EndHour = 21 // first NOT-eligible local hour (exclusive): last start is 20:59
MaxAttemptsPerWeek = 3
// TickInterval is how often the job looks; a deferred or failed pass is therefore retried the next hour.
TickInterval = time.Hour
// FirstTickDelay lets the agent settle after a start before the first look.
FirstTickDelay = 5 * time.Minute
// PerGuestTimeout bounds one `pct fstrim` (measured 24.4 s for 84 GiB).
PerGuestTimeout = 30 * time.Minute
)
// ScheduleText is the human description carried on the host report.
const ScheduleText = "weekly, due Wednesday from 10:00 host-local time; starts only 10:00-20:59; never beside a backup or restore-test"
// Runner runs a host command (proxmox.ExecRunner in production, through `sudo -n`).
type Runner interface {
Run(ctx context.Context, name string, args ...string) (stdout, stderr []byte, err error)
}
// GuestSource yields the guests this agent OWNS (the pool-verified source, never a bare ListLXC).
type GuestSource interface {
Guests(ctx context.Context) ([]proxmox.Guest, error)
}
// Gate is the host-wide one-heavy-operation gate (*backup.InFlight).
type Gate interface {
TryAcquire(what string) (release func(), busy string, ok bool)
}
// GateName is what the gate reports as busy while a trim runs.
const GateName = "guest-fstrim"
// Record is one guest's last trim attempt, as persisted.
type Record struct {
LastAttemptAt time.Time `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt time.Time `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
// Attempts counts the attempts since the current week's due time (reset by the first attempt of a new week).
Attempts int `json:"attempts"`
}
// Trimmer is the weekly job.
type Trimmer struct {
runner Runner
guests GuestSource
gate Gate
statePath string
logger *slog.Logger
loc *time.Location
now func() time.Time
mu sync.Mutex
records map[int]Record
}
// New builds the job and loads the persisted state. A missing state file is an empty state; a corrupt one is logged
// and treated as empty (the cost is one extra trim, never a missed one).
func New(runner Runner, guests GuestSource, gate Gate, statePath string, logger *slog.Logger) *Trimmer {
if logger == nil {
logger = slog.Default()
}
t := &Trimmer{runner: runner, guests: guests, gate: gate, statePath: statePath, logger: logger,
loc: time.Local, now: time.Now, records: map[int]Record{}}
t.load()
return t
}
func (t *Trimmer) load() {
data, err := os.ReadFile(t.statePath)
if err != nil {
if !os.IsNotExist(err) {
t.logger.Warn("fstrim: state read failed — starting empty", "path", t.statePath, "err", err)
}
return
}
var raw map[string]Record
if err := json.Unmarshal(data, &raw); err != nil {
t.logger.Warn("fstrim: state file corrupt — starting empty", "path", t.statePath, "err", err)
return
}
for k, r := range raw {
if id, err := strconv.Atoi(k); err == nil && id > 0 {
t.records[id] = r
}
}
}
func (t *Trimmer) saveLocked() error {
raw := make(map[string]Record, len(t.records))
for id, r := range t.records {
raw[strconv.Itoa(id)] = r
}
data, err := json.MarshalIndent(raw, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(t.statePath), 0o755); err != nil {
return err
}
tmp := t.statePath + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, t.statePath)
}
// EligibleHour reports whether a trim may START at local time lt.
func EligibleHour(lt time.Time) bool {
h := lt.Hour()
return h >= StartHour && h < EndHour
}
// weekAnchor is the most recent Wednesday StartHour:00 at or before lt (same location as lt).
func weekAnchor(lt time.Time) time.Time {
daysBack := (int(lt.Weekday()) - int(Weekday) + 7) % 7
d := lt.AddDate(0, 0, -daysBack)
a := time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
if a.After(lt) {
d = d.AddDate(0, 0, -7)
a = time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
}
return a
}
// due reports whether a guest with record r (ok=false: none) is due at local time lt.
func due(r Record, has bool, lt time.Time) bool {
if !has {
return true
}
anchor := weekAnchor(lt)
if r.LastAttemptAt.Before(anchor) {
return true // not tried this week
}
return !r.OK && r.Attempts < MaxAttemptsPerWeek
}
// Run looks every TickInterval until ctx ends. It never returns an error: a failed trim is a reported fact.
func (t *Trimmer) Run(ctx context.Context) {
t.logger.Info("fstrim: weekly guest disk trim starting", "schedule", ScheduleText)
timer := time.NewTimer(FirstTickDelay)
defer timer.Stop()
for {
select {
case <-ctx.Done():
return
case <-timer.C:
t.Pass(ctx)
timer.Reset(TickInterval)
}
}
}
// Pass is one look: outside the daytime window it does nothing; otherwise it trims every due, running, owned guest
// while holding the heavy-op gate.
func (t *Trimmer) Pass(ctx context.Context) {
lt := t.now().In(t.loc)
if !EligibleHour(lt) {
t.logger.Debug("fstrim: outside the daytime window — not looking", "local", lt.Format("Mon 15:04"))
return
}
guests, err := t.guests.Guests(ctx)
if err != nil {
t.logger.Warn("fstrim: owned-guest list unavailable — skipping this pass", "err", err)
return
}
owned := make(map[int]bool, len(guests))
var todo []int
t.mu.Lock()
for _, g := range guests {
owned[g.VMID] = true
r, has := t.records[g.VMID]
if !due(r, has, lt) {
continue
}
if g.Status != "running" {
t.logger.Info("fstrim: guest not running — trimmed when it runs", "vmid", g.VMID, "status", g.Status)
continue
}
todo = append(todo, g.VMID)
}
// A guest the agent no longer owns has no result to report.
pruned := false
for id := range t.records {
if !owned[id] {
delete(t.records, id)
pruned = true
}
}
if pruned {
if err := t.saveLocked(); err != nil {
t.logger.Warn("fstrim: state save failed", "err", err)
}
}
t.mu.Unlock()
if len(todo) == 0 {
return
}
sort.Ints(todo)
release, busy, ok := t.gate.TryAcquire(GateName)
if !ok {
t.logger.Info("fstrim: deferred — a heavy operation is in flight; retrying next hour", "busy", busy, "due_guests", len(todo))
return
}
defer release()
for _, vmid := range todo {
if ctx.Err() != nil {
return
}
t.trimOne(ctx, vmid, lt)
}
}
var trimmedLine = regexp.MustCompile(`\((\d+) bytes\) trimmed`)
// ParseTrimmed sums the "(N bytes) trimmed" lines of `pct fstrim` output and counts them (one per mount point), e.g.
// `/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed`.
func ParseTrimmed(out string) (bytes int64, mounts int) {
for _, m := range trimmedLine.FindAllStringSubmatch(out, -1) {
n, err := strconv.ParseInt(m[1], 10, 64)
if err != nil {
continue
}
bytes += n
mounts++
}
return bytes, mounts
}
// GiB renders bytes as "30.1 GiB".
func GiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
func (t *Trimmer) trimOne(ctx context.Context, vmid int, lt time.Time) {
start := t.now()
cctx, cancel := context.WithTimeout(ctx, PerGuestTimeout)
stdout, stderr, err := t.runner.Run(cctx, "pct", "fstrim", strconv.Itoa(vmid))
cancel()
dur := t.now().Sub(start)
bytes, mounts := ParseTrimmed(string(stdout) + "\n" + string(stderr))
t.mu.Lock()
prev, has := t.records[vmid]
r := Record{LastAttemptAt: start.UTC(), OK: err == nil, BytesTrimmed: bytes, Mounts: mounts,
DurationSeconds: float64(dur.Round(100*time.Millisecond)) / float64(time.Second), LastOKAt: prev.LastOKAt}
if has && !prev.LastAttemptAt.Before(weekAnchor(lt)) {
r.Attempts = prev.Attempts + 1
} else {
r.Attempts = 1
}
if err == nil {
r.LastOKAt = start.UTC()
} else {
msg := strings.TrimSpace(err.Error() + ": " + strings.TrimSpace(string(stderr)))
if len(msg) > 300 {
msg = msg[:300]
}
r.Error = msg
}
t.records[vmid] = r
saveErr := t.saveLocked()
t.mu.Unlock()
if err == nil {
t.logger.Info(fmt.Sprintf("fstrim: guest %d trimmed %s in %.1fs", vmid, GiB(bytes), r.DurationSeconds),
"vmid", vmid, "bytes_trimmed", bytes, "mounts", mounts, "duration_s", r.DurationSeconds)
if mounts == 0 {
t.logger.Warn("fstrim: pct fstrim succeeded but reported no trimmed mount — output not understood",
"vmid", vmid, "stdout", strings.TrimSpace(string(stdout)))
}
} else {
t.logger.Warn(fmt.Sprintf("fstrim: guest %d trim FAILED after %.1fs", vmid, r.DurationSeconds),
"vmid", vmid, "attempt", r.Attempts, "max_attempts_per_week", MaxAttemptsPerWeek, "err", r.Error)
}
if saveErr != nil {
t.logger.Warn("fstrim: state save failed — the result will not survive a restart", "path", t.statePath, "err", saveErr)
}
}
// GuestDiskTrimStatus implements hub.GuestDiskTrimReporter: a pure read of the persisted results (never runs pct).
func (t *Trimmer) GuestDiskTrimStatus(context.Context) *hub.GuestDiskTrimStatus {
t.mu.Lock()
defer t.mu.Unlock()
out := &hub.GuestDiskTrimStatus{Schedule: ScheduleText}
ids := make([]int, 0, len(t.records))
for id := range t.records {
ids = append(ids, id)
}
sort.Ints(ids)
for _, id := range ids {
r := t.records[id]
g := hub.GuestDiskTrim{VMID: id, LastAttemptAt: r.LastAttemptAt.UTC().Format(time.RFC3339), OK: r.OK,
BytesTrimmed: r.BytesTrimmed, Mounts: r.Mounts, DurationSeconds: r.DurationSeconds, Error: r.Error}
if !r.LastOKAt.IsZero() {
g.LastOKAt = r.LastOKAt.UTC().Format(time.RFC3339)
}
out.Guests = append(out.Guests, g)
}
return out
}
+267
View File
@@ -0,0 +1,267 @@
package fstrim
import (
"bytes"
"context"
"encoding/json"
"errors"
"log/slog"
"path/filepath"
"reflect"
"strings"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// The real `pct fstrim 9201` output measured on demo-hp 2026-10-06 (audits/ten-answers-2026-10-06/r444-measure.txt).
const measuredOut = "/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed\n" +
"/var/lib/lxc/9201/rootfs/var/lib/felhom: 53.9 GiB (57865633792 bytes) trimmed\n"
const measuredBytes = int64(32277680128 + 57865633792)
type fakeRunner struct {
mu sync.Mutex
calls [][]string
out string
err error
onRun func()
}
func (f *fakeRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
f.mu.Lock()
f.calls = append(f.calls, append([]string{name}, args...))
f.mu.Unlock()
if f.onRun != nil {
f.onRun()
}
if f.err != nil {
return nil, []byte("mount busy"), f.err
}
return []byte(f.out), nil, nil
}
type fakeGuests struct {
g []proxmox.Guest
err error
}
func (f fakeGuests) Guests(context.Context) ([]proxmox.Guest, error) { return f.g, f.err }
// A Wednesday 10:30 in a fixed zone (CEST-like), so the tests do not depend on the machine's zone.
var zone = time.FixedZone("CEST", 2*3600)
func at(day, hour, min int) time.Time { return time.Date(2026, 10, day, hour, min, 0, 0, zone) } // 2026-10-07 = Wednesday
func newT(t *testing.T, r Runner, g GuestSource, gate Gate, now *time.Time) (*Trimmer, *bytes.Buffer, string) {
t.Helper()
var logs bytes.Buffer
path := filepath.Join(t.TempDir(), "guest-disk-trim.json")
tr := New(r, g, gate, path, slog.New(slog.NewTextHandler(&logs, &slog.HandlerOptions{Level: slog.LevelDebug})))
tr.loc = zone
tr.now = func() time.Time { return *now }
return tr, &logs, path
}
func running(ids ...int) fakeGuests {
var g []proxmox.Guest
for _, id := range ids {
g = append(g, proxmox.Guest{VMID: id, Status: "running", Type: "lxc"})
}
return fakeGuests{g: g}
}
func TestParseTrimmedTheMeasuredOutput(t *testing.T) {
b, m := ParseTrimmed(measuredOut)
if b != measuredBytes || m != 2 {
t.Fatalf("ParseTrimmed = %d bytes over %d mounts, want %d over 2", b, m, measuredBytes)
}
if b, m := ParseTrimmed("something else\n"); b != 0 || m != 0 {
t.Fatalf("unrelated output parsed as %d/%d", b, m)
}
if got := GiB(measuredBytes); got != "84.0 GiB" {
t.Fatalf("GiB = %q", got)
}
}
// The night window (01:00–06:59) must never be eligible, and the daytime window is exactly 10:00–20:59.
func TestEligibleHourNeverInTheNight(t *testing.T) {
for h := 0; h < 24; h++ {
lt := time.Date(2026, 10, 7, h, 30, 0, 0, zone)
got := EligibleHour(lt)
if h >= 1 && h <= 6 && got {
t.Errorf("hour %02d is in the night window and must not be eligible", h)
}
if want := h >= 10 && h <= 20; got != want {
t.Errorf("EligibleHour(%02d:30) = %v, want %v", h, got, want)
}
}
}
func TestWeekAnchorIsTheLastWednesdayTen(t *testing.T) {
cases := map[time.Time]time.Time{
at(7, 10, 0): at(7, 10, 0), // Wednesday 10:00 itself
at(7, 9, 59): time.Date(2026, 9, 30, 10, 0, 0, 0, zone), // before 10:00 Wednesday → the previous week
at(8, 15, 0): at(7, 10, 0), // Thursday
at(13, 20, 0): at(7, 10, 0), // next Tuesday
at(14, 11, 0): at(14, 10, 0), // next Wednesday
}
for in, want := range cases {
if got := weekAnchor(in); !got.Equal(want) {
t.Errorf("weekAnchor(%s) = %s, want %s", in.Format("Mon 01-02 15:04"), got.Format("Mon 01-02 15:04"), want.Format("Mon 01-02 15:04"))
}
}
}
// The consequence: on Wednesday 10:30 a running owned guest is trimmed with the ONE exact argv, the bytes are parsed,
// the positive log line is written, the result is persisted, and the host report carries it.
func TestPassTrimsADueGuestAndReportsIt(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
tr, logs, path := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9201"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("runner calls = %q, want %q", r.calls, want)
}
if !strings.Contains(logs.String(), "fstrim: guest 9201 trimmed 84.0 GiB in ") {
t.Fatalf("no positive per-guest log line:\n%s", logs.String())
}
st := tr.GuestDiskTrimStatus(context.Background())
if st == nil || st.Schedule != ScheduleText || len(st.Guests) != 1 {
t.Fatalf("report stanza = %+v", st)
}
g := st.Guests[0]
if g.VMID != 9201 || !g.OK || g.BytesTrimmed != measuredBytes || g.Mounts != 2 || g.LastOKAt == "" || g.LastAttemptAt == "" {
t.Fatalf("report guest = %+v", g)
}
// Persisted: a NEW Trimmer over the same file (an agent restart) still has it and does not trim again this week.
now = at(8, 11, 0)
r2 := &fakeRunner{out: measuredOut}
tr2 := New(r2, running(9201), &backup.InFlight{}, path, slog.New(slog.NewTextHandler(&bytes.Buffer{}, nil)))
tr2.loc, tr2.now = zone, func() time.Time { return now }
if st2 := tr2.GuestDiskTrimStatus(context.Background()); len(st2.Guests) != 1 || st2.Guests[0].BytesTrimmed != measuredBytes {
t.Fatalf("result lost over a restart: %+v", st2)
}
tr2.Pass(context.Background())
if len(r2.calls) != 0 {
t.Fatalf("trimmed again in the same week after a restart: %q", r2.calls)
}
// Next week it is due again.
now = at(14, 10, 5)
tr2.Pass(context.Background())
if len(r2.calls) != 1 {
t.Fatalf("not trimmed in the next week: %q", r2.calls)
}
}
func TestPassNeverRunsInTheNight(t *testing.T) {
for _, h := range []int{1, 3, 6, 9, 21, 23} {
now := at(7, h, 15)
r := &fakeRunner{out: measuredOut}
tr, _, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Errorf("trimmed at %02d:15: %q", h, r.calls)
}
}
}
// A backup (or restore-test) holding the heavy-op gate DEFERS the trim; the next hour, gate free, it runs. And while
// a trim runs, the gate is held, so a backup cannot start beside it.
func TestPassDefersToAHeavyOperationAndRetriesNextHour(t *testing.T) {
now := at(7, 10, 30)
gate := &backup.InFlight{}
release, _, _ := gate.TryAcquire("backup:9201")
var busyDuringTrim string
r := &fakeRunner{out: measuredOut}
r.onRun = func() { busyDuringTrim = gate.Busy() }
tr, logs, _ := newT(t, r, running(9201), gate, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Fatalf("trimmed beside a running backup: %q", r.calls)
}
if !strings.Contains(logs.String(), "fstrim: deferred") || !strings.Contains(logs.String(), "backup:9201") {
t.Fatalf("the deferral is not logged with what holds the gate:\n%s", logs.String())
}
release()
now = now.Add(time.Hour)
tr.Pass(context.Background())
if len(r.calls) != 1 {
t.Fatalf("not retried the next hour: %q", r.calls)
}
if busyDuringTrim != GateName {
t.Fatalf("the heavy-op gate was %q during the trim, want %q", busyDuringTrim, GateName)
}
if gate.Busy() != "" {
t.Fatalf("the gate was not released after the pass: %q", gate.Busy())
}
}
func TestFailedTrimWarnsIsRecordedAndRetriedAtMostThreeTimes(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{err: errors.New("exit status 255")}
tr, logs, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
for i := 0; i < 6; i++ {
tr.Pass(context.Background())
now = now.Add(time.Hour)
}
if len(r.calls) != MaxAttemptsPerWeek {
t.Fatalf("attempts in one week = %d, want %d", len(r.calls), MaxAttemptsPerWeek)
}
if !strings.Contains(logs.String(), "level=WARN") || !strings.Contains(logs.String(), "fstrim: guest 9201 trim FAILED") {
t.Fatalf("no WARN for the failure:\n%s", logs.String())
}
g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]
if g.OK || g.LastOKAt != "" || !strings.Contains(g.Error, "exit status 255") || !strings.Contains(g.Error, "mount busy") {
t.Fatalf("failed result not recorded as a failure: %+v", g)
}
// A success later keeps a clean record.
r.err = nil
r.out = measuredOut
now = at(14, 10, 10)
tr.Pass(context.Background())
if g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]; !g.OK || g.Error != "" || g.BytesTrimmed != measuredBytes {
t.Fatalf("success after failure: %+v", g)
}
}
func TestOnlyRunningOwnedGuestsAndAFailedListActsOnNothing(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
g := fakeGuests{g: []proxmox.Guest{{VMID: 9201, Status: "stopped"}, {VMID: 9202, Status: "running"}}}
tr, _, _ := newT(t, r, g, &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9202"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("calls = %q, want only the running guest", r.calls)
}
r2 := &fakeRunner{out: measuredOut}
tr2, logs, _ := newT(t, r2, fakeGuests{err: errors.New("pool read 403")}, &backup.InFlight{}, &now)
tr2.Pass(context.Background())
if len(r2.calls) != 0 || !strings.Contains(logs.String(), "owned-guest list unavailable") {
t.Fatalf("a failed ownership read must act on nothing: calls %q", r2.calls)
}
}
func TestReportJSONShape(t *testing.T) {
now := at(7, 10, 30)
tr, _, _ := newT(t, &fakeRunner{out: measuredOut}, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
b, err := json.Marshal(tr.GuestDiskTrimStatus(context.Background()))
if err != nil {
t.Fatal(err)
}
for _, k := range []string{`"schedule":`, `"guests":[{"vmid":9201`, `"last_attempt_at":"2026-10-07T08:30:00Z"`, `"ok":true`,
`"bytes_trimmed":90143313920`, `"mounts":2`, `"duration_seconds":`, `"last_ok_at":"2026-10-07T08:30:00Z"`} {
if !strings.Contains(string(b), k) {
t.Errorf("report JSON lacks %s: %s", k, b)
}
}
if strings.Contains(string(b), `"error"`) {
t.Errorf("an ok result must omit error: %s", b)
}
}
+83 -5
View File
@@ -66,6 +66,13 @@ type ProvenRestoreTestReporter interface {
ProvenRestoreTests(ctx context.Context) []RestoreTest ProvenRestoreTests(ctx context.Context) []RestoreTest
} }
// KnownBackupReporter is the newest SUCCESSFUL backup per tier and guest kept on disk (R-894's
// backup-success-state.json). collectBackups folds it into the report so a restart does not erase the
// hub's evidence of a backup that ran minutes before it (the kernel-night shape, 2026-10-09).
type KnownBackupReporter interface {
KnownBackupSuccesses() []Backup
}
// PBSReporter is the slice-6-Phase-B seam the pbs verify loop plugs into (same pattern). // PBSReporter is the slice-6-Phase-B seam the pbs verify loop plugs into (same pattern).
// Returns the agent's latest-known PBS snapshot inventory + verify-state. nil → empty. // Returns the agent's latest-known PBS snapshot inventory + verify-state. nil → empty.
type PBSReporter interface { type PBSReporter interface {
@@ -90,6 +97,12 @@ type GuestNetReporter interface {
GuestNetStatus(ctx context.Context) *GuestNetStatus GuestNetStatus(ctx context.Context) *GuestNetStatus
} }
// GuestDiskTrimReporter is the R-444 seam the weekly trim job plugs into (same consumer-side pattern — hub does not
// import fstrim). nil (feature not wired) → no guest_disk_trim stanza.
type GuestDiskTrimReporter interface {
GuestDiskTrimStatus(ctx context.Context) *GuestDiskTrimStatus
}
// Collector builds a HostReport from read-only sources. All deps are behind narrow // Collector builds a HostReport from read-only sources. All deps are behind narrow
// interfaces for unit testing. // interfaces for unit testing.
type Collector struct { type Collector struct {
@@ -99,6 +112,7 @@ type Collector struct {
backups BackupReporter backups BackupReporter
restoreTests RestoreTestReporter restoreTests RestoreTestReporter
provenTests ProvenRestoreTestReporter provenTests ProvenRestoreTestReporter
knownBackups KnownBackupReporter // R-894 file → host report after a restart (nil → in-memory only)
pbs PBSReporter pbs PBSReporter
temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp) temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp)
capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty) capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty)
@@ -108,6 +122,8 @@ type Collector struct {
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted) pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted) ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted) guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
diskTrim GuestDiskTrimReporter // R-444: weekly guest disk trim (nil → stanza omitted)
foreignKey ForeignKeyArchiveReporter // R-366 slice 2 (nil → omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false) selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted) mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
oob OOBReporter // H1: operator-access health (nil → stanza omitted) oob OOBReporter // H1: operator-access health (nil → stanza omitted)
@@ -221,6 +237,24 @@ func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
return c return c
} }
// ForeignKeyArchiveReporter is the R-366 slice-2 seam: the restore-test's ledger of archives written with another
// key. nil = not evaluated yet (the stanza is omitted and the hub keeps its state).
type ForeignKeyArchiveReporter interface {
ForeignKeyArchives(ctx context.Context) *ForeignKeyArchivesStanza
}
// SetForeignKeyArchiveReporter wires the restore-test's foreign-key ledger (R-366 slice 2; nil-safe → omitted).
func (c *Collector) SetForeignKeyArchiveReporter(r ForeignKeyArchiveReporter) *Collector {
c.foreignKey = r
return c
}
// SetGuestDiskTrimReporter wires the R-444 weekly trim job as a report source (nil-safe → stanza omitted).
func (c *Collector) SetGuestDiskTrimReporter(r GuestDiskTrimReporter) *Collector {
c.diskTrim = r
return c
}
// SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side // SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side
// pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report. // pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report.
type SelfUpdateReporter interface { type SelfUpdateReporter interface {
@@ -377,6 +411,14 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.guestNet != nil { if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx) report.GuestNet = c.guestNet.GuestNetStatus(ctx)
} }
// R-444: the last weekly trim result per guest (nil reporter = not wired → stanza omitted).
if c.diskTrim != nil {
report.GuestDiskTrim = c.diskTrim.GuestDiskTrimStatus(ctx)
}
// R-366 slice 2: archives the restore-test skipped as another key's (nil → not evaluated yet → omitted).
if c.foreignKey != nil {
report.ForeignKeyArchives = c.foreignKey.ForeignKeyArchives(ctx)
}
// D1: agent self-update pending status (nil reporter → pending=false, the steady state). // D1: agent self-update pending status (nil reporter → pending=false, the steady state).
if c.selfUpdate != nil { if c.selfUpdate != nil {
report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending() report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending()
@@ -542,14 +584,50 @@ func (c *Collector) collectStorage(ctx context.Context) []StorageTarget {
// collectBackups / collectRestoreTests read the agent's latest backup + restore-test state // collectBackups / collectRestoreTests read the agent's latest backup + restore-test state
// via the seams. Best-effort: a nil reporter or nil slice degrades to an empty (non-nil) // via the seams. Best-effort: a nil reporter or nil slice degrades to an empty (non-nil)
// list so the collection always marshals as []. // list so the collection always marshals as [].
//
// The saved successes (SetKnownBackupReporter) are added for each tier and guest the in-memory list has
// no success for at or after the saved time. Why: the in-memory list is empty after an agent restart,
// and the hub's backup-freshness check reads only what the reports carried. Measured 2026-10-09 on
// demo-felhom: the night backup landed 02:40 UTC, the kernel step restarted the host at 02:44, no report
// fell in between, so two kernel nights in a row left no trace and the hub alarmed „newest backup is 48h
// old" at 03:00. Only SUCCESSES are saved, so a failure is never hidden and never invented.
// Pinned by TestCollectBackups_SavedSuccessSurvivesARestart.
func (c *Collector) collectBackups(ctx context.Context) []Backup { func (c *Collector) collectBackups(ctx context.Context) []Backup {
if c.backups == nil { out := []Backup{}
return []Backup{} if c.backups != nil {
}
if b := c.backups.Backups(ctx); b != nil { if b := c.backups.Backups(ctx); b != nil {
return b out = append(out, b...)
} }
return []Backup{} }
if c.knownBackups == nil {
return out
}
for _, k := range c.knownBackups.KnownBackupSuccesses() {
kt, err := time.Parse(time.RFC3339, k.StartedAt)
if err != nil {
continue
}
covered := false
for _, b := range out {
if !b.Success || b.TargetID != k.TargetID || b.VMID != k.VMID {
continue
}
if bt, err := time.Parse(time.RFC3339, b.StartedAt); err == nil && !bt.Before(kt) {
covered = true
break
}
}
if !covered {
out = append(out, k)
}
}
return out
}
// SetKnownBackupReporter wires the R-894 saved successes into the host report (nil-safe → in-memory only).
func (c *Collector) SetKnownBackupReporter(r KnownBackupReporter) *Collector {
c.knownBackups = r
return c
} }
// collectRestoreTests merges the in-memory result with the PERSISTED per-tier proofs (R-189). // collectRestoreTests merges the in-memory result with the PERSISTED per-tier proofs (R-189).
+54
View File
@@ -0,0 +1,54 @@
package hub
import (
"context"
"encoding/json"
"testing"
)
// R-444: the guest_disk_trim stanza must reach a report built through the PRODUCTION collect path, be absent from
// the wire when the job is not wired, and carry the keys the hub's System page reads.
type fakeDiskTrim struct{ st *GuestDiskTrimStatus }
func (f fakeDiskTrim) GuestDiskTrimStatus(context.Context) *GuestDiskTrimStatus { return f.st }
func TestCollect_GuestDiskTrim(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.150.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ := json.Marshal(r)
var m map[string]any
_ = json.Unmarshal(b, &m)
if _, ok := m["guest_disk_trim"]; ok {
t.Fatalf("guest_disk_trim on the wire with no reporter wired: %s", b)
}
c.SetGuestDiskTrimReporter(fakeDiskTrim{st: &GuestDiskTrimStatus{Schedule: "weekly", Guests: []GuestDiskTrim{{
VMID: 9201, LastAttemptAt: "2026-10-07T08:30:00Z", OK: true, BytesTrimmed: 90143313920, Mounts: 2,
DurationSeconds: 24.4, LastOKAt: "2026-10-07T08:30:00Z",
}}}})
r, err = c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ = json.Marshal(r)
m = nil
_ = json.Unmarshal(b, &m)
dt, ok := m["guest_disk_trim"].(map[string]any)
if !ok || dt["schedule"] != "weekly" {
t.Fatalf("guest_disk_trim missing or wrong on the wire: %s", b)
}
g := dt["guests"].([]any)[0].(map[string]any)
for _, k := range []string{"vmid", "last_attempt_at", "ok", "bytes_trimmed", "mounts", "duration_seconds", "last_ok_at"} {
if _, ok := g[k]; !ok {
t.Fatalf("guest_disk_trim.guests[0] lacks %q: %v", k, g)
}
}
if g["bytes_trimmed"] != float64(90143313920) || g["ok"] != true {
t.Fatalf("values did not survive the round trip: %v", g)
}
}
+62
View File
@@ -0,0 +1,62 @@
package hub
import (
"context"
"testing"
)
type fakeMemBackups []Backup
func (f fakeMemBackups) Backups(context.Context) []Backup { return f }
type fakeKnownBackups []Backup
func (f fakeKnownBackups) KnownBackupSuccesses() []Backup { return f }
// The consequence, not the mechanism: after a restart (empty in-memory list) the host report still carries
// the night's backup, so the hub's freshness check sees it. Measured 2026-10-09 on demo-felhom: without this
// the report carried backups: [] and the hub alarmed „newest backup is 48h old" an hour after a good backup.
// RED-PROOF: return before the saved-success loop in collectBackups → the first case reports nothing → fails.
func TestCollectBackups_SavedSuccessSurvivesARestart(t *testing.T) {
saved := Backup{TargetID: "felhom-backup", VMID: 9201, Success: true, StartedAt: "2026-10-09T02:40:16Z"}
t.Run("restart: memory empty, the saved success is reported", func(t *testing.T) {
c := &Collector{backups: fakeMemBackups(nil), knownBackups: fakeKnownBackups{saved}}
got := c.collectBackups(context.Background())
if len(got) != 1 || got[0].StartedAt != saved.StartedAt || !got[0].Success || got[0].TargetID != "felhom-backup" {
t.Fatalf("want the saved success, got %+v", got)
}
})
t.Run("memory holds the same or a newer success: no second entry", func(t *testing.T) {
mem := Backup{TargetID: "felhom-backup", VMID: 9201, Success: true, StartedAt: "2026-10-09T02:40:16Z", Archive: "a"}
c := &Collector{backups: fakeMemBackups{mem}, knownBackups: fakeKnownBackups{saved}}
if got := c.collectBackups(context.Background()); len(got) != 1 || got[0].Archive != "a" {
t.Fatalf("want only the in-memory record, got %+v", got)
}
})
t.Run("a newer FAILURE in memory never hides the saved success, and is kept", func(t *testing.T) {
fail := Backup{TargetID: "felhom-backup", VMID: 9201, Success: false, StartedAt: "2026-10-10T02:40:00Z", Error: "x"}
c := &Collector{backups: fakeMemBackups{fail}, knownBackups: fakeKnownBackups{saved}}
got := c.collectBackups(context.Background())
if len(got) != 2 || got[0].Success || !got[1].Success {
t.Fatalf("want the failure and the saved success, got %+v", got)
}
})
t.Run("another tier's success does not cover this tier", func(t *testing.T) {
other := Backup{TargetID: "felhom-pbs", VMID: 9201, Success: true, StartedAt: "2026-10-10T00:00:00Z"}
c := &Collector{backups: fakeMemBackups{other}, knownBackups: fakeKnownBackups{saved}}
if got := c.collectBackups(context.Background()); len(got) != 2 {
t.Fatalf("want both tiers, got %+v", got)
}
})
t.Run("no saved state wired: the in-memory list as before, never nil", func(t *testing.T) {
c := &Collector{}
if got := c.collectBackups(context.Background()); got == nil || len(got) != 0 {
t.Fatalf("want an empty non-nil list, got %#v", got)
}
})
}
+69
View File
@@ -124,6 +124,17 @@ type HostReport struct {
// on HostReport would have been the only report block named against that convention. // on HostReport would have been the only report block named against that convention.
GuestNet *GuestNetStatus `json:"guest_net,omitempty"` GuestNet *GuestNetStatus `json:"guest_net,omitempty"`
// GuestDiskTrim is the weekly guest disk trim stanza (R-444, `09` §3 decision 139): the schedule and, per owned
// guest, the LAST trim result as persisted by the agent (it survives a restart). Present only when the trim job is
// wired; an empty `guests` list means the job runs and no guest has been trimmed yet. No secret.
GuestDiskTrim *GuestDiskTrimStatus `json:"guest_disk_trim,omitempty"`
// ForeignKeyArchives (R-366 slice 2, `09` §3 decision 168): per backup tier, the whole-guest archives the
// restore-test SKIPPED because they were written with another key (an earlier install of this box). This box
// cannot open them; the hub turns a CHANGE of this list into one operator event. Absent = not evaluated yet since
// the agent started (the hub keeps its last state); `tiers: []` = evaluated, none found.
ForeignKeyArchives *ForeignKeyArchivesStanza `json:"foreign_key_archives,omitempty"`
// LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent // LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent
// mirror of the controller's report log_tails channel. Present ONLY on the heartbeat // mirror of the controller's report log_tails channel. Present ONLY on the heartbeat
// right after the control envelope requested it (log_tail_requested); consume-once on // right after the control envelope requested it (log_tail_requested); consume-once on
@@ -208,6 +219,27 @@ type GuestNetGuest struct {
Message string `json:"message,omitempty"` Message string `json:"message,omitempty"`
} }
// GuestDiskTrimStatus is the R-444 weekly trim stanza. `schedule` is a plain description of when the job runs (local
// time of the host); `guests` holds one entry per owned guest that has had at least one trim attempt.
type GuestDiskTrimStatus struct {
Schedule string `json:"schedule"`
Guests []GuestDiskTrim `json:"guests,omitempty"`
}
// GuestDiskTrim is one guest's LAST trim attempt. `ok` with `last_attempt_at` is the verdict of that attempt — never
// read the time alone as success; `last_ok_at` is the last attempt that succeeded ("" = never). `bytes_trimmed` is
// the sum of the "(N bytes) trimmed" lines `pct fstrim` printed, over `mounts` mount points.
type GuestDiskTrim struct {
VMID int `json:"vmid"`
LastAttemptAt string `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt string `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
}
type PBSDRStatus struct { type PBSDRStatus struct {
State string `json:"state"` State string `json:"state"`
StorageID string `json:"storage_id,omitempty"` StorageID string `json:"storage_id,omitempty"`
@@ -428,6 +460,17 @@ type SmartSummary struct {
ReallocatedSectors *int `json:"reallocated_sectors"` ReallocatedSectors *int `json:"reallocated_sectors"`
PendingSectors *int `json:"pending_sectors"` PendingSectors *int `json:"pending_sectors"`
OfflineUncorrectable *int `json:"offline_uncorrectable"` OfflineUncorrectable *int `json:"offline_uncorrectable"`
// R-330 (disk health Phase 2): three more SATA raw counters. omitempty + pointer: absent (an
// older agent, an NVMe/USB device, or a drive that does not report the attribute) is OMITTED —
// unknown, never a zero (S-39). Wire only: no verdict reads them yet.
// 187 Reported_Uncorrect — the failing drive's most telling counter (normalized 1 vs thresh 0,
// raw 1001) while SMART still said PASSED.
// 188 Command_Timeout — some vendors PACK several counters into the 48-bit raw value, so the
// number is carried as reported and must not be compared across vendors.
// 199 UDMA_CRC_Error_Count — cabling / link errors, not the medium.
ReportedUncorrect *int64 `json:"reported_uncorrect,omitempty"`
CommandTimeout *int64 `json:"command_timeout,omitempty"`
UDMACRCErrors *int64 `json:"udma_crc_errors,omitempty"`
// NVMe attributes. // NVMe attributes.
CriticalWarning *int `json:"critical_warning"` CriticalWarning *int `json:"critical_warning"`
@@ -593,6 +636,17 @@ type WireOSUpdate struct {
// HostRelease is the newest approved HOST release (hub v0.131.0, `11` §8 step 3) — a separate set: a version // HostRelease is the newest approved HOST release (hub v0.131.0, `11` §8 step 3) — a separate set: a version
// approved for the guest is not approved for the host by that fact alone. // approved for the guest is not approved for the host by that fact alone.
HostRelease *WireOSRelease `json:"host_release,omitempty"` HostRelease *WireOSRelease `json:"host_release,omitempty"`
// Kernel is the kernel lane's instruction for THIS box (R-836, `09` §3 decision 172, `11` §5.11): the kernel the
// household was told about, and whether tonight is a told night. Nil (an older hub, or no kernel due) = no kernel
// step. The hub sets Tonight only after the household's mail the day before went out — no mail, no step.
Kernel *WireKernelStep `json:"kernel,omitempty"`
}
// WireKernelStep is the hub's kernel-lane instruction (hub osupdates.KernelBlock — field-exact, cross-repo).
type WireKernelStep struct {
Kver string `json:"kver"` // e.g. "7.0.14-22-pve" — the wrapper refuses any other (R23)
Tonight bool `json:"tonight"` // the household was mailed the day before: tonight's leg may reboot
NotifiedAt string `json:"notified_at,omitempty"` // when that mail went out (RFC 3339), for the log
} }
// WireOSRelease is an approved version set; Snapshot is the approval time (YYYYMMDDTHHMMSSZ) the wrapper uses // WireOSRelease is an approved version set; Snapshot is the approval time (YYYYMMDDTHHMMSSZ) the wrapper uses
@@ -689,3 +743,18 @@ type WireRestoreDirective struct {
Archive string `json:"archive,omitempty"` // source archive/snapshot to restore from Archive string `json:"archive,omitempty"` // source archive/snapshot to restore from
VMID int `json:"vmid,omitempty"` VMID int `json:"vmid,omitempty"`
} }
// ForeignKeyArchivesStanza wraps the per-tier list so "evaluated, none" (`tiers: []`) differs from "not evaluated"
// (the stanza absent) without a null on the wire.
type ForeignKeyArchivesStanza struct {
Tiers []ForeignKeyArchives `json:"tiers"`
}
// ForeignKeyArchives is one tier's count of archives written with another key (R-366 slice 2): the count and the
// newest/oldest archive time (RFC3339, UTC). No key material — the fingerprints stay on the box.
type ForeignKeyArchives struct {
Target string `json:"target"`
Count int `json:"count"`
Oldest string `json:"oldest"`
Newest string `json:"newest"`
}
+17 -4
View File
@@ -15,7 +15,7 @@ import (
// Red-proof: drop the `b.Success &&` guard and the failed-backup sub-case fails; move the call after release() // Red-proof: drop the `b.Success &&` guard and the failed-backup sub-case fails; move the call after release()
// and the gate sub-case fails. // and the gate sub-case fails.
func TestAfterPrimaryBackup(t *testing.T) { func TestAfterPrimaryBackup(t *testing.T) {
run := func(t *testing.T, failErr string) (calls []int, gateHeld bool) { run := func(t *testing.T, failErr, path string) (calls []int, gateHeld bool) {
gate := &backup.InFlight{} gate := &backup.InFlight{}
b := &fakeBackups{failErr: failErr} b := &fakeBackups{failErr: failErr}
srv := newTestServerS(t, &fakeGuests{}, b, &fakeStore{}, nil) srv := newTestServerS(t, &fakeGuests{}, b, &fakeStore{}, nil)
@@ -34,7 +34,7 @@ func TestAfterPrimaryBackup(t *testing.T) {
done <- struct{}{} done <- struct{}{}
}) })
h := srv.Handler() h := srv.Handler()
if do(t, h, "POST", "/backup", "A", "").Code != http.StatusAccepted { if do(t, h, "POST", path, "A", "").Code != http.StatusAccepted {
t.Fatal("POST /backup not accepted") t.Fatal("POST /backup not accepted")
} }
select { select {
@@ -47,7 +47,7 @@ func TestAfterPrimaryBackup(t *testing.T) {
return calls, gateHeld return calls, gateHeld
} }
t.Run("success runs the leg under the gate", func(t *testing.T) { t.Run("success runs the leg under the gate", func(t *testing.T) {
calls, held := run(t, "") calls, held := run(t, "", "/backup")
if len(calls) != 1 { if len(calls) != 1 {
t.Fatalf("the leg ran %d time(s), want 1", len(calls)) t.Fatalf("the leg ran %d time(s), want 1", len(calls))
} }
@@ -56,8 +56,21 @@ func TestAfterPrimaryBackup(t *testing.T) {
} }
}) })
t.Run("a failed backup runs nothing", func(t *testing.T) { t.Run("a failed backup runs nothing", func(t *testing.T) {
if calls, _ := run(t, "vzdump exploded"); len(calls) != 0 { if calls, _ := run(t, "vzdump exploded", "/backup"); len(calls) != 0 {
t.Fatalf("the leg ran after a FAILED backup: %v", calls) t.Fatalf("the leg ran after a FAILED backup: %v", calls)
} }
}) })
// R-899: a household press is not the night's backup — no OS leg after it. Same fake, same successful backup as
// the first sub-case; only the query differs. Red-proof: make handleBackup ignore `trigger` and this sub-case
// fails with the leg run once.
t.Run("a manual press runs nothing", func(t *testing.T) {
if calls, _ := run(t, "", "/backup?trigger=manual"); len(calls) != 0 {
t.Fatalf("the OS leg ran after a manual press: %v", calls)
}
})
t.Run("the scheduled path with the new query still runs the leg", func(t *testing.T) {
if calls, _ := run(t, "", "/backup?trigger=night"); len(calls) != 1 {
t.Fatalf("the leg ran %d time(s) after a non-manual backup, want 1", len(calls))
}
})
} }
@@ -4,7 +4,6 @@ import (
"context" "context"
"encoding/json" "encoding/json"
"errors" "errors"
"io"
"log/slog" "log/slog"
"os" "os"
"path/filepath" "path/filepath"
@@ -54,8 +53,8 @@ func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string
} }
return "", errors.New("supExec: unexpected args") return "", errors.New("supExec: unexpected args")
} }
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) { func (f *supExec) WriteControllerImage(context.Context, int, string) error {
return "", errors.New("supExec: no stdin exec expected") return errors.New("supExec: no image write expected")
} }
func (f *supExec) count(vmid int) int { func (f *supExec) count(vmid int) int {
f.mu.Lock() f.mu.Lock()
+9 -12
View File
@@ -4,7 +4,6 @@ import (
"context" "context"
"encoding/json" "encoding/json"
"fmt" "fmt"
"io"
"log/slog" "log/slog"
"net/http" "net/http"
"os" "os"
@@ -41,9 +40,10 @@ func ValidControllerImage(ref string) bool { return controllerImageRe.MatchStrin
// faked in tests. The single seam the swap composes over (no hand-rolled pct). // faked in tests. The single seam the swap composes over (no hand-rolled pct).
type GuestExecutor interface { type GuestExecutor interface {
GuestExec(ctx context.Context, vmid int, args ...string) (string, error) GuestExec(ctx context.Context, vmid int, args ...string) (string, error)
// GuestExecStdin is GuestExec with the command's stdin fed from stdin — the swap write pipes the // WriteControllerImage writes the image ref into the guest's /etc/felhom-controller-image through the ROOT
// image ref into an in-guest `tee` (no shell vector). // verb `felhom-priv-apply controller-image <vmid>` (R-861 (a) A1, `09` §3 decision 165), which re-checks the
GuestExecStdin(ctx context.Context, vmid int, stdin io.Reader, args ...string) (string, error) // ref against our registry + repository + x.y.z. The agent no longer holds a `tee` grant into the guest.
WriteControllerImage(ctx context.Context, vmid int, image string) error
} }
// ControllerSwapState is the durable record of a swap (crash-safety + status). Written before the swap // ControllerSwapState is the durable record of a swap (crash-safety + status). Written before the swap
@@ -141,14 +141,11 @@ func (c *ControllerSwapper) imagePresent(ctx context.Context, vmid int, image st
} }
func (c *ControllerSwapper) writeImage(ctx context.Context, vmid int, image string) error { func (c *ControllerSwapper) writeImage(ctx context.Context, vmid int, image string) error {
// Non-root path: pipe the image ref into an in-guest `tee` over stdin — no shell, no // R-861 (a) A1: the ROOT verb writes the file (it re-checks the ref — a compromised agent cannot hand the guest's
// interpolation, no `bash -c` (the only swap vector that would have needed an arbitrary-exec // bootstrap another image). The bytes are `image\n`, byte-identical to the golden's `printf '%s\n'`; the bootstrap
// grant). The trailing "\n" makes the on-disk bytes byte-identical to the golden's // reads `IMAGE=$(cat …)` so the newline is stripped on read. image is also strict-validated (controllerImageRe)
// `printf '%s\n'`; the bootstrap reads `IMAGE=$(cat …)` so the newline is stripped on read // upstream in Swap.
// (spike SPIKE-controllerswap-narrow-grants-2026-06-29). image is strict-validated return c.exec.WriteControllerImage(ctx, vmid, image)
// (controllerImageRe) upstream in Swap; defence-in-depth, the stdin path can't smuggle anyway.
_, err := c.exec.GuestExecStdin(ctx, vmid, strings.NewReader(image+"\n"), "tee", controllerImageFile)
return err
} }
func (c *ControllerSwapper) restartBootstrap(ctx context.Context, vmid int) error { func (c *ControllerSwapper) restartBootstrap(ctx context.Context, vmid int) error {
+14 -28
View File
@@ -23,7 +23,7 @@ type fakeGuestExec struct {
present map[string]bool // images pulled into the guest present map[string]bool // images pulled into the guest
good map[string]bool // images that report healthy when running good map[string]bool // images that report healthy when running
containerImg string // image the running container currently has containerImg string // image the running container currently has
teeStdin []string // raw bytes piped into each `tee` write (the swap's write vector) teeStdin []string // image refs handed to WriteControllerImage (the root verb, R-861 (a) A1)
failRestart bool failRestart bool
noHealthBlock bool // if set, .State.Health is absent ("none") noHealthBlock bool // if set, .State.Health is absent ("none")
restartCount int // F1: .RestartCount reported by docker inspect (a crash-looper has >0) restartCount int // F1: .RestartCount reported by docker inspect (a crash-looper has >0)
@@ -75,20 +75,15 @@ func (f *fakeGuestExec) GuestExec(_ context.Context, _ int, args ...string) (str
return "", fmt.Errorf("fake: unexpected exec %v", args) return "", fmt.Errorf("fake: unexpected exec %v", args)
} }
// GuestExecStdin models the swap's write vector: `tee /etc/felhom-controller-image` with the image // WriteControllerImage models the swap's write: the root verb `felhom-priv-apply controller-image <vmid>` (R-861
// piped on stdin. It records the raw stdin bytes and sets the modeled file content (newline-stripped, // (a) A1). It records the ref and sets the modeled file content, as the bootstrap's `IMAGE=$(cat …)` would read it.
// as the bootstrap's `IMAGE=$(cat …)` read would see it). func (f *fakeGuestExec) WriteControllerImage(_ context.Context, _ int, image string) error {
func (f *fakeGuestExec) GuestExecStdin(_ context.Context, _ int, stdin io.Reader, args ...string) (string, error) {
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
f.calls = append(f.calls, args) f.calls = append(f.calls, []string{"felhom-priv-apply", "controller-image", image})
b, _ := io.ReadAll(stdin) f.teeStdin = append(f.teeStdin, image+"\n")
if len(args) >= 2 && args[0] == "tee" && args[1] == controllerImageFile { f.imageFile = image
f.teeStdin = append(f.teeStdin, string(b)) return nil
f.imageFile = strings.TrimSpace(string(b))
return string(b), nil // tee echoes stdin to stdout
}
return "", fmt.Errorf("fake: unexpected exec-stdin args=%v stdin=%q", args, string(b))
} }
// wrote reports whether the image was written via the stdin `tee` vector with the exact `image\n` // wrote reports whether the image was written via the stdin `tee` vector with the exact `image\n`
@@ -162,7 +157,7 @@ func TestControllerSwap_Happy(t *testing.T) {
// The write vector must be the stdin `tee` with byte-identical `image\n` and NO shell — the // The write vector must be the stdin `tee` with byte-identical `image\n` and NO shell — the
// controllerswap.go writeImage rewrite. This would FAIL on the pre-change `bash -c "printf … >"` impl. // controllerswap.go writeImage rewrite. This would FAIL on the pre-change `bash -c "printf … >"` impl.
func TestControllerSwap_WriteViaStdinTee_NoShell(t *testing.T) { func TestControllerSwap_WriteViaRootVerb_NoShell(t *testing.T) {
fe := &fakeGuestExec{ fe := &fakeGuestExec{
imageFile: prevImg, imageFile: prevImg,
present: map[string]bool{newImg: true}, present: map[string]bool{newImg: true},
@@ -173,22 +168,15 @@ func TestControllerSwap_WriteViaStdinTee_NoShell(t *testing.T) {
t.Fatalf("state = %q, want done", st.State) t.Fatalf("state = %q, want done", st.State)
} }
if !fe.wrote(newImg) { if !fe.wrote(newImg) {
t.Errorf("expected a tee write of %q+\\n; teeStdin=%q", newImg, fe.teeStdin) t.Errorf("expected the root verb to write %q; writes=%q", newImg, fe.teeStdin)
} }
sawTee := false
for _, c := range fe.calls { for _, c := range fe.calls {
if len(c) >= 2 && c[0] == "tee" { if len(c) >= 1 && c[0] == "tee" {
sawTee = true t.Errorf("the swap still uses an in-guest tee (R-861 (a) A1 removed that grant): %v", c)
if c[1] != controllerImageFile {
t.Errorf("tee target = %q, want fixed %q", c[1], controllerImageFile)
} }
} }
}
if !sawTee {
t.Error("no tee call recorded — writeImage did not use the stdin tee vector")
}
if fe.usedShell() { if fe.usedShell() {
t.Errorf("swap used a shell vector (bash/-c/printf) — must be stdin tee only; calls=%v", fe.calls) t.Errorf("swap used a shell vector (bash/-c/printf); calls=%v", fe.calls)
} }
} }
@@ -300,9 +288,7 @@ func (s *inspectScript) GuestExec(_ context.Context, _ int, args ...string) (str
} }
return "", nil return "", nil
} }
func (s *inspectScript) GuestExecStdin(_ context.Context, _ int, _ io.Reader, _ ...string) (string, error) { func (s *inspectScript) WriteControllerImage(context.Context, int, string) error { return nil }
return "", nil
}
func fastSwapper(exec GuestExecutor) *ControllerSwapper { func fastSwapper(exec GuestExecutor) *ControllerSwapper {
s := NewControllerSwapper(exec, "", discardLogger()) s := NewControllerSwapper(exec, "", discardLogger())
+102
View File
@@ -0,0 +1,102 @@
package localapi
import (
"encoding/json"
"errors"
"io"
"io/fs"
"net/http"
"os"
"time"
)
// GET /host/crash-guard (R-856, `09` §3 decision 143): what the host's crash guard
// (configs/felhom-crash-guard, `11` §5.9) recorded about the most recent HOST boot. The controller
// reads it once after it starts: when the host's last boot followed an UNCLEAN stop, its app mails
// wait ~15 minutes instead of the normal 90 s boot grace.
//
// Read-only and Proxmox-free: the agent reads the guard's state file (root-owned, 0644 — the
// non-root agent can read it) and passes four fields through. Host-wide, token-authed (any valid
// per-guest token sees the host's view, as GET /host/metrics does).
//
// NEVER an error page. A missing file (no guard installed, or no boot recorded yet), an unreadable
// one, or one that does not parse answers 200 with present:false — the controller reads that as
// UNKNOWN and keeps its normal boot grace. Pinned by TestR856_CrashGuard*.
// defaultCrashGuardStatePath is where configs/felhom-crash-guard writes its state (STATE_DIR there).
const defaultCrashGuardStatePath = "/var/lib/felhom-crash-guard/state.json"
// crashGuardStateMax bounds the read; the real file is well under 4 KiB.
const crashGuardStateMax = 1 << 20
// CrashGuardResponse is the data block of GET /host/crash-guard. Field names are the controller's
// agentapi.CrashGuardState (felhom-controller internal/agentapi/crashguard.go) — a wire contract,
// pinned by TestR856_CrashGuardWireMatchesControllerClient.
type CrashGuardResponse struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"` // RFC3339 UTC ("2006-01-02T15:04:05Z")
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// crashGuardFile is the subset of the guard's state.json the route passes through. Every other key
// (armed, boot_id, config, unclean_boots, last_trip, ...) is ignored.
type crashGuardFile struct {
LastBootAt string `json:"last_boot_at"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// readCrashGuardState reads and parses the guard's state file. ok=false on ANY failure (missing,
// unreadable, oversized, not a JSON object, a field of the wrong type); reason says which, for the log.
func readCrashGuardState(path string) (resp CrashGuardResponse, ok bool, reason string) {
f, err := os.Open(path)
if err != nil {
if errors.Is(err, fs.ErrNotExist) {
return resp, false, "no state file"
}
return resp, false, "unreadable: " + err.Error()
}
defer f.Close()
raw, err := io.ReadAll(io.LimitReader(f, crashGuardStateMax+1))
if err != nil {
return resp, false, "read: " + err.Error()
}
if len(raw) > crashGuardStateMax {
return resp, false, "state file too large"
}
var st crashGuardFile
// Unmarshal into a struct fails on a non-object top level (null decodes, so reject it below).
if err := json.Unmarshal(raw, &st); err != nil {
return resp, false, "unparseable: " + err.Error()
}
var probe map[string]json.RawMessage
if err := json.Unmarshal(raw, &probe); err != nil || probe == nil {
return resp, false, "unparseable: not a JSON object"
}
resp = CrashGuardResponse{Present: true, LastBootUnclean: st.LastBootUnclean, Tripped: st.Tripped}
// Normalise to RFC3339 UTC; an unparseable time passes through as-is (the controller reads an
// unparseable boot time as "not this start's boot" → its normal grace).
if t, perr := time.Parse(time.RFC3339, st.LastBootAt); perr == nil {
resp.LastBootAt = t.UTC().Format(time.RFC3339)
} else {
resp.LastBootAt = st.LastBootAt
}
return resp, true, ""
}
func (s *Server) handleCrashGuard(w http.ResponseWriter, r *http.Request, vmid int) {
path := s.crashGuardStatePath
if path == "" {
path = defaultCrashGuardStatePath
}
resp, ok, reason := readCrashGuardState(path)
if !ok {
s.logger.Debug("local-api: /host/crash-guard not present", "vmid", vmid, "reason", reason)
writeOK(w, CrashGuardResponse{Present: false})
return
}
s.logger.Debug("local-api: /host/crash-guard served", "vmid", vmid,
"last_boot_at", resp.LastBootAt, "last_boot_unclean", resp.LastBootUnclean, "tripped", resp.Tripped)
writeOK(w, resp)
}
+205
View File
@@ -0,0 +1,205 @@
package localapi
import (
"encoding/json"
"io"
"log/slog"
"net/http"
"os"
"path/filepath"
"testing"
)
// The shape of /var/lib/felhom-crash-guard/state.json as read on demo-hp on 2026-10-06 (values from
// that read where they matter; lists/objects kept to the same key set).
const crashGuardFixture = `{
"armed": true,
"boot_id": "3f1c0f1e-6a0b-4d7e-9b7a-0c2d4e6f8a1b",
"config": {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10},
"kernel_panic": 10,
"last_boot_at": "2026-10-05T07:56:41Z",
"last_boot_unclean": true,
"last_trip": {},
"rearmed_at": "2026-10-04T14:02:11Z",
"rearmed_by": "operator",
"tripped": false,
"unclean_boots": ["2026-10-05T07:56:41Z"],
"unclean_boots_24h": 1,
"unclean_boots_in_window": 1,
"updated_at": "2026-10-05T07:57:02Z",
"version": 1
}`
// controllerCrashGuardState is a COPY of the controller's wire type, felhom-controller
// controller/internal/agentapi/crashguard.go `CrashGuardState` (commit 8b13a5e) — same field names,
// same tags. If either side renames a key, the contract test below fails.
type controllerCrashGuardState struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
func newCrashGuardServer(t *testing.T, statePath string) http.Handler {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0",
Guests: &fakeGuests{},
Backups: &fakeBackups{},
Store: &fakeStore{},
Storage: fakeStorage{},
Tokens: staticTokens{"A": 8200, "B": 9300},
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatalf("new server: %v", err)
}
srv.crashGuardStatePath = statePath
return srv.Handler()
}
func writeCrashGuardFixture(t *testing.T, body string) string {
t.Helper()
p := filepath.Join(t.TempDir(), "state.json")
if err := os.WriteFile(p, []byte(body), 0o644); err != nil {
t.Fatal(err)
}
return p
}
// getCrashGuard calls the route and decodes the envelope with the CONTROLLER's type.
func getCrashGuard(t *testing.T, h http.Handler, token string) (int, controllerCrashGuardState, string) {
t.Helper()
w := do(t, h, "GET", "/host/crash-guard", token, "")
var env struct {
OK bool `json:"ok"`
Data controllerCrashGuardState `json:"data"`
}
if w.Code == http.StatusOK {
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil {
t.Fatalf("decode %q: %v", w.Body.String(), err)
}
if !env.OK {
t.Fatalf("ok=false: %s", w.Body.String())
}
}
return w.Code, env.Data, w.Body.String()
}
// A present state file (demo-hp's shape) passes the three facts through.
func TestR856_CrashGuardPresentFile(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, body)
}
want := controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: true, Tripped: false}
if st != want {
t.Fatalf("state = %+v, want %+v", st, want)
}
// A tripped, clean boot reads back as such (both bools are carried, not defaulted).
h = newCrashGuardServer(t, writeCrashGuardFixture(t,
`{"last_boot_at":"2026-10-05T09:56:41+02:00","last_boot_unclean":false,"tripped":true,"version":1}`))
_, st, _ = getCrashGuard(t, h, "B")
want = controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: false, Tripped: true}
if st != want {
t.Fatalf("offset time / tripped: state = %+v, want %+v (time normalised to UTC Z)", st, want)
}
}
// No state file (no guard on this host, or no boot recorded yet) → 200 present:false.
func TestR856_CrashGuardMissingFile(t *testing.T) {
h := newCrashGuardServer(t, filepath.Join(t.TempDir(), "absent", "state.json"))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("missing file: got %d, want 200 (%s)", code, body)
}
if st.Present || st.LastBootUnclean || st.Tripped || st.LastBootAt != "" {
t.Fatalf("missing file: state = %+v, want present:false and nothing else", st)
}
}
// A garbled file → 200 present:false, never a 5xx — every shape of garbage.
func TestR856_CrashGuardGarbageFile(t *testing.T) {
for name, body := range map[string]string{
"truncated": crashGuardFixture[:40],
"not json": "this is not json\n",
"empty": "",
"null": "null",
"array": `[{"last_boot_unclean":true}]`,
"wrong type": `{"last_boot_at":"2026-10-05T07:56:41Z","last_boot_unclean":"yes","tripped":false}`,
"lone brace": "{",
} {
t.Run(name, func(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, body))
code, st, raw := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, raw)
}
if st.Present || st.LastBootUnclean {
t.Fatalf("garbage %q read as %+v, want present:false", name, st)
}
})
}
// The path is a directory, not a file: unreadable → present:false, 200.
h := newCrashGuardServer(t, t.TempDir())
if code, st, raw := getCrashGuard(t, h, "A"); code != http.StatusOK || st.Present {
t.Fatalf("directory path: got %d %+v (%s), want 200 present:false", code, st, raw)
}
}
// No / unknown token → 401, like every sibling route; a cross-guest ?vmid= → 403.
func TestR856_CrashGuardRequiresGuestToken(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
for _, tok := range []string{"", "bogus"} {
w := do(t, h, "GET", "/host/crash-guard", tok, "")
if w.Code != http.StatusUnauthorized {
t.Fatalf("token %q: got %d, want 401", tok, w.Code)
}
if json.Valid(w.Body.Bytes()) {
var env struct {
Data controllerCrashGuardState `json:"data"`
}
_ = json.Unmarshal(w.Body.Bytes(), &env)
if env.Data.Present || env.Data.LastBootUnclean {
t.Fatalf("token %q: the refusal leaked the state: %s", tok, w.Body.String())
}
}
}
if w := do(t, h, "GET", "/host/crash-guard?vmid=9300", "A", ""); w.Code != http.StatusForbidden {
t.Fatalf("cross-guest query: got %d, want 403", w.Code)
}
}
// Wire contract: every key the controller's CrashGuardState decodes is emitted under exactly that
// name, and the agent emits no key the controller does not know.
func TestR856_CrashGuardWireMatchesControllerClient(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
w := do(t, h, "GET", "/host/crash-guard", "A", "")
var env struct {
OK bool `json:"ok"`
Data map[string]json.RawMessage `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil || !env.OK {
t.Fatalf("envelope: %v %s", err, w.Body.String())
}
want := []string{"present", "last_boot_at", "last_boot_unclean", "tripped"}
for _, k := range want {
if _, ok := env.Data[k]; !ok {
t.Errorf("agent does not emit %q, which the controller decodes (%s)", k, w.Body.String())
}
}
if len(env.Data) != len(want) {
t.Errorf("agent emits %d keys, controller knows %d: %s", len(env.Data), len(want), w.Body.String())
}
// And the agent's own type agrees with the controller's copy, field for field.
var mine CrashGuardResponse
var theirs controllerCrashGuardState
raw, _ := json.Marshal(env.Data)
_ = json.Unmarshal(raw, &mine)
_ = json.Unmarshal(raw, &theirs)
if (controllerCrashGuardState{mine.Present, mine.LastBootAt, mine.LastBootUnclean, mine.Tripped}) != theirs {
t.Errorf("agent %+v vs controller %+v", mine, theirs)
}
}
+16
View File
@@ -127,6 +127,22 @@ func (s *Server) handleRecoverOffsitePassword(w http.ResponseWriter, r *http.Req
"retained_has_restic_pw": match.HasResticPassword, "retained_has_restic_pw": match.HasResticPassword,
}, },
"the recovery code is correct, but it belongs to an EARLIER sealed package (superseded "+match.SupersededAt+"), not the one currently held") "the recovery code is correct, but it belongs to an EARLIER sealed package (superseded "+match.SupersededAt+"), not the one currently held")
// ── R-304 (2026-10-08) — NOT EVERY EARLIER PACKAGE WAS CHECKED. ─────────────────────────
//
// The current package refused the code, no retained package opened it — and at least one earlier
// package the hub holds was never tried (or the list could not be read). Saying „the code is
// wrong" here would claim a check that did not happen. 424 (Failed Dependency): the verdict
// depends on packages we could not try. The controller classifies on the status, never on this
// sentence; an older controller maps an unknown status to its neutral „we do not know why".
case errors.Is(err, escrow.ErrRetainedUnchecked):
n := -1
var ue *escrow.RetainedUncheckedError
if errors.As(err, &ue) {
n = ue.Unchecked
}
s.logger.Warn("local-api: offsite key recovery: the code did not open the current package and earlier packages were NOT all checked — not reported as a wrong code (R-304)", "vmid", vmid, "unchecked", n)
writeStatus(w, http.StatusFailedDependency, false, map[string]any{"older_unchecked": n},
"the recovery code did not open the current sealed package, and earlier packages the hub holds were not all checked — the code may belong to one of them; nothing was written")
case errors.Is(err, escrow.ErrNoResticPassword): case errors.Is(err, escrow.ErrNoResticPassword):
s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid) s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid)
writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)") writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)")
@@ -54,6 +54,13 @@ func TestRecoverOffsitePassword_EachSituationGetsItsOwnStatus(t *testing.T) {
wantStatus: 404, wantStatus: 404,
mustNotSay: []string{"did not open"}, mustNotSay: []string{"did not open"},
}, },
{
// R-304: earlier packages were not all tried — never „did not open the sealed bundle" (the wrong-code words).
name: "earlier packages not all checked — not a wrong code",
err: &escrow.RetainedUncheckedError{Unchecked: 2},
wantStatus: 424,
mustNotSay: []string{"did not open the sealed bundle", "could not be fetched"},
},
{ {
name: "the bundle predates the repository-password field", name: "the bundle predates the repository-password field",
err: escrow.ErrNoResticPassword, err: escrow.ErrNoResticPassword,
+10 -9
View File
@@ -3,7 +3,6 @@ package localapi
import ( import (
"context" "context"
"fmt" "fmt"
"io"
"log/slog" "log/slog"
"strconv" "strconv"
"strings" "strings"
@@ -126,14 +125,16 @@ func (b *GuestBinder) GuestExec(ctx context.Context, vmid int, args ...string) (
return string(out), nil return string(out), nil
} }
// GuestExecStdin is GuestExec with the in-guest command's stdin fed from stdin. The controller-swap // privApplyBin is the root content checker (R-861); its `controller-image` verb writes the guest's image file.
// write uses it to pipe the image ref into an in-guest `tee` (no shell vector, no interpolation), const privApplyBin = "/usr/local/sbin/felhom-priv-apply"
// through the same fenced runner so the `sudo -n` prefix stays in one place.
func (b *GuestBinder) GuestExecStdin(ctx context.Context, vmid int, stdin io.Reader, args ...string) (string, error) { // WriteControllerImage pipes `image\n` to `felhom-priv-apply controller-image <vmid>` through the same fenced runner
pctArgs := append([]string{"exec", strconv.Itoa(vmid), "--"}, args...) // (the `sudo -n` prefix stays in one place). The verb checks the ref as root and writes the guest file itself
out, stderr, err := b.runner.RunStdin(ctx, stdin, "pct", pctArgs...) // (R-861 (a) A1, `09` §3 decision 165). Pinned by TestR861_WriteControllerImageUsesTheRootVerb.
func (b *GuestBinder) WriteControllerImage(ctx context.Context, vmid int, image string) error {
_, stderr, err := b.runner.RunStdin(ctx, strings.NewReader(image+"\n"), privApplyBin, "controller-image", strconv.Itoa(vmid))
if err != nil { if err != nil {
return string(out), fmt.Errorf("pct exec %d %v: %w: %s", vmid, args, err, strings.TrimSpace(string(stderr))) return fmt.Errorf("felhom-priv-apply controller-image %d: %w: %s", vmid, err, strings.TrimSpace(string(stderr)))
} }
return string(out), nil return nil
} }
@@ -0,0 +1,48 @@
package localapi
import (
"context"
"io"
"testing"
)
// stdinRecorder is a proxmox.Runner that records each call and the stdin it was handed.
type stdinRecorder struct {
name string
args []string
stdin string
}
func (r *stdinRecorder) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
r.name, r.args = name, args
return nil, nil, nil
}
func (r *stdinRecorder) RunStdin(_ context.Context, stdin io.Reader, name string, args ...string) ([]byte, []byte, error) {
b, _ := io.ReadAll(stdin)
r.name, r.args, r.stdin = name, args, string(b)
return nil, nil, nil
}
// R-861 (a) A1 (`09` §3 decision 165): the managed controller update writes the guest's image file through the ROOT
// verb, never through an in-guest `tee` the agent could feed any image.
//
// COMPANION RED-PROOF (observed): restore the pre-A1 body (`b.runner.RunStdin(ctx, …, "pct", "exec", vmid, "--",
// "tee", controllerImageFile)`) → this fails with "the image write ran pct …, want felhom-priv-apply". Restored.
func TestR861_WriteControllerImageUsesTheRootVerb(t *testing.T) {
rec := &stdinRecorder{}
b := NewGuestBinder(rec, discardLogger())
const img = "gitea.dooplex.hu/admin/felhom-controller:0.302.0"
if err := b.WriteControllerImage(context.Background(), 9201, img); err != nil {
t.Fatal(err)
}
if rec.name != privApplyBin {
t.Fatalf("the image write ran %s %v, want felhom-priv-apply", rec.name, rec.args)
}
if len(rec.args) != 2 || rec.args[0] != "controller-image" || rec.args[1] != "9201" {
t.Fatalf("argv = %v, want [controller-image 9201] (the sudoers line `^controller-image [0-9]+$`)", rec.args)
}
if rec.stdin != img+"\n" {
t.Fatalf("stdin = %q, want the ref plus one newline", rec.stdin)
}
}
+144
View File
@@ -0,0 +1,144 @@
package localapi
import (
"context"
"io"
"log/slog"
"net/http"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-894 — after an agent restart, an UNREADABLE storage must fall back to the last success saved on
// disk, not to "never". Measured 2026-10-05 on demo-hp: a restart at 04:57, the off-site storage
// unreachable at 06:25, the 7-day tier (last copy 4 days old) read DUE, vzdump failed.
//
// Every test here builds a NEW server and a NEW BackupSuccessState from the same file — that is the
// restart. The in-memory store (fakeStore) is always fresh, as after a real restart.
// r894Server builds a two-tier server whose off-site tier answers the storage listing with lister.
func r894Server(t *testing.T, path string, pbsSvc BackupService) *Server {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: &fakeStore{},
Storage: fakeStorage{targets: []hub.StorageTarget{{Name: "local"}, {Name: "felhom-pbs"}}},
Tokens: staticTokens{"A": 8200},
BackupTiers: []BackupTier{
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: pbsSvc},
},
LastKnownBackups: backup.NewBackupSuccessState(path),
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatal(err)
}
srv.baseCtx = context.Background()
srv.now = func() time.Time { return testNow }
return srv
}
// unreadable is the off-site storage as demo-hp saw it: "Can't connect to 10.77.0.1:8007".
func unreadable() archiveLister {
return archiveLister{fakeBackups: &fakeBackups{}, err: errStorageRead}
}
// THE R-894 CASE, end to end. Agent 1 takes an off-site backup through POST /backup (the fake
// runner's success is 12 h before testNow). The agent restarts. The storage cannot be read. The tier
// must read NOT due, from the copy saved on disk.
//
// COMPANION RED-PROOF (observed): delete the `lookup == archiveUnknown && s.lastKnown != nil` block in
// handleBackupDue → this fails with "after a restart an unreadable storage must fall back to the saved
// copy (12 h old, 7-day tier) — NOT due; got {… Due:true … AgeState:unknown …}". Restored.
func TestBackupDue_R894_RestartThenUnreadableStorage_FreshSavedCopyIsNotDue(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
// Agent 1: a real backup job through the endpoint the controller calls.
first := r894Server(t, path, &fakeBackups{})
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
}
waitFor(t, func() bool {
_, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200)
return ok
})
// Agent 2: a restart (new server, new state from the same file), and the storage is unreachable.
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if got.Due {
t.Fatalf("after a restart an unreadable storage must fall back to the saved copy (12 h old, 7-day tier) — NOT due; got %+v", got)
}
if got.AgeState != AgeStateKnown || got.AgeSecs == nil || *got.AgeSecs != int64((12*time.Hour).Seconds()) {
t.Fatalf("the age must come from the saved copy (12 h, known); got %+v", got)
}
}
// The deliberate rule stays: an unreadable storage must not suppress a backup that IS due. A saved
// copy older than the cadence reads DUE.
//
// COMPANION RED-PROOF (observed): make the fallback answer not-due whenever a saved copy exists
// (`if fromDisk { …Due:false… }` before the cadence check) → this fails with "a saved copy 9 days old
// under a 7-day cadence MUST read due". Restored.
func TestBackupDue_R894_RestartThenUnreadableStorage_OldSavedCopyIsDue(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
st := backup.NewBackupSuccessState(path)
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, 9*24*time.Hour, true)); err != nil {
t.Fatal(err)
}
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("a saved copy 9 days old under a 7-day cadence MUST read due; got %+v", got)
}
if got.AgeState != AgeStateKnown {
t.Fatalf("the age is known (from disk); got %+v", got)
}
}
// No saved copy → the pre-R-894 answer, byte for byte: DUE, age UNKNOWN (never ABSENT — the controller
// fires its window-gate valve only on absent, R-88).
func TestBackupDue_R894_RestartThenUnreadableStorage_NoSavedCopyIsDueUnknown(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due || got.AgeState != AgeStateUnknown || got.AgeSecs != nil {
t.Fatalf("no saved copy + unreadable storage must stay DUE with age unknown; got %+v", got)
}
}
// A storage that ANSWERS is the ground truth: an archive absent there makes the tier due even when the
// file remembers a fresh success (a pruned or deleted copy must be made again).
//
// COMPANION RED-PROOF (observed): drop `lookup == archiveUnknown &&` from the fallback condition → this
// fails with "the storage answered 'no archive' — the saved copy must NOT stand in for it". Restored.
func TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
st := backup.NewBackupSuccessState(path)
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, time.Hour, true)); err != nil {
t.Fatal(err)
}
absent := archiveLister{fakeBackups: &fakeBackups{}, found: false}
got := dueFor(t, r894Server(t, path, absent).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("the storage answered 'no archive' — the saved copy must NOT stand in for it; got %+v", got)
}
}
// A FAILED backup is never saved: it must not make a tier look fresh after a restart.
func TestBackupDue_R894_FailedBackupIsNotSaved(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
first := r894Server(t, path, &fakeBackups{failErr: "could not activate storage 'felhom-pbs'"})
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
}
waitFor(t, func() bool { return len(first.store.Backups(context.Background())) == 1 })
if _, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200); ok {
t.Fatal("a failed backup must never be saved as a success")
}
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("after a failed backup and a restart the tier must still be due; got %+v", got)
}
}
+56 -5
View File
@@ -85,6 +85,13 @@ type BackupStore interface {
RestoreTests(ctx context.Context) []hub.RestoreTest RestoreTests(ctx context.Context) []hub.RestoreTest
} }
// LastKnownBackupStore (R-894) is the on-disk newest-success-per-tier record. Satisfied by
// *backup.BackupSuccessState.
type LastKnownBackupStore interface {
RecordBackupSuccess(target string, b hub.Backup) error
LastKnownSuccess(target string, vmid int) (time.Time, bool)
}
// StorageView yields the host's observed storage targets (for mapping a mount's storage id → // StorageView yields the host's observed storage targets (for mapping a mount's storage id →
// fast/slow class). Satisfied by *storage.Observer. // fast/slow class). Satisfied by *storage.Observer.
type StorageView interface { type StorageView interface {
@@ -164,6 +171,10 @@ type Options struct {
// PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg // PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg
// that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs. // that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs.
AfterPrimaryBackup func(ctx context.Context, vmid int) AfterPrimaryBackup func(ctx context.Context, vmid int)
// LastKnownBackups (R-894) keeps the newest successful backup per tier ON DISK, so the due-check's
// fallback for an UNREADABLE storage after an agent restart is the last known copy, not "never".
// nil = the pre-R-894 behaviour (in-memory record only).
LastKnownBackups LastKnownBackupStore
// Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when // Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when
// nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner. // nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner.
Privileged PrivilegedRunner Privileged PrivilegedRunner
@@ -229,7 +240,6 @@ type Options struct {
// POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured" // POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured"
// (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer. // (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer.
EscrowRecovery EscrowRecoverer EscrowRecovery EscrowRecoverer
} }
// defaultBackupCadence is the fallback /backup/due window when none is configured. // defaultBackupCadence is the fallback /backup/due window when none is configured.
@@ -281,6 +291,7 @@ type Server struct {
// inFlight (R-85) is shared with the restore-test scheduler so the two never run together. // inFlight (R-85) is shared with the restore-test scheduler so the two never run together.
inFlight *backup.InFlight inFlight *backup.InFlight
afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none
lastKnown LastKnownBackupStore // R-894: on-disk newest success per tier; nil = none
logger *slog.Logger logger *slog.Logger
now func() time.Time now func() time.Time
@@ -294,6 +305,9 @@ type Server struct {
netMountRoot string // the user-data namespace root for the network-mount role gate netMountRoot string // the user-data namespace root for the network-mount role gate
smbCredsDir string // where SMB creds files are written (out-of-band, 0600) smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
// crashGuardStatePath (R-856) is the host crash guard's state file read by GET /host/crash-guard;
// empty = defaultCrashGuardStatePath. A seam: tests point it at a fixture.
crashGuardStatePath string
// escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed // escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the // identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
// offsite repository password. OPTIONAL — nil (no hub client configured) makes // offsite repository password. OPTIONAL — nil (no hub client configured) makes
@@ -475,6 +489,7 @@ func NewServer(o Options) (*Server, error) {
s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence) s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence)
s.inFlight = o.InFlight s.inFlight = o.InFlight
s.afterPrimaryBackup = o.AfterPrimaryBackup s.afterPrimaryBackup = o.AfterPrimaryBackup
s.lastKnown = o.LastKnownBackups
if s.backups == nil && len(s.tiers) > 0 { if s.backups == nil && len(s.tiers) > 0 {
s.backups = s.tiers[0].Service s.backups = s.tiers[0].Service
} }
@@ -520,6 +535,9 @@ func (s *Server) Handler() http.Handler {
// Host metrics (slice 9): host-wide health + per-storage capacity for the customer's monitoring // Host metrics (slice 9): host-wide health + per-storage capacity for the customer's monitoring
// view. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot). // view. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot).
mux.HandleFunc("GET /host/metrics", s.withGuest(s.handleHostMetrics)) mux.HandleFunc("GET /host/metrics", s.withGuest(s.handleHostMetrics))
// R-856 (`09` §3 decision 143): the host crash guard's record of the last HOST boot — the controller
// waits ~15 min with app mails after an unclean one. Read-only; a missing/garbled file = present:false.
mux.HandleFunc("GET /host/crash-guard", s.withGuest(s.handleCrashGuard))
// Disk management (slice 8C) — self-scoped; format routes through the data-bearing classifier+gate. // Disk management (slice 8C) — self-scoped; format routes through the data-bearing classifier+gate.
mux.HandleFunc("GET /disks", s.withGuest(s.handleDisks)) mux.HandleFunc("GET /disks", s.withGuest(s.handleDisks))
mux.HandleFunc("GET /disks/candidates", s.withGuest(s.handleDiskCandidates)) mux.HandleFunc("GET /disks/candidates", s.withGuest(s.handleDiskCandidates))
@@ -807,6 +825,10 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
return return
} }
key := backupJobKey{vmid: vmid, target: tier.TargetID} key := backupJobKey{vmid: vmid, target: tier.TargetID}
// R-899 (operator ruling 2026-10-08): a household press („Mentés most") is not the night's backup. A controller
// that knows sends `trigger=manual`; then the OS leg does not follow (it belongs to the night, after the night's
// own copy). An older controller sends nothing and keeps the old behaviour.
manual := r.URL.Query().Get("trigger") == "manual"
// ONE BACKUP AT A TIME PER GUEST, ACROSS ALL TIERS (operator ruling 2026-07-26: "other backup // ONE BACKUP AT A TIME PER GUEST, ACROSS ALL TIERS (operator ruling 2026-07-26: "other backup
// shouldn't start until finished"). vzdump takes a guest lock, so a concurrent second backup // shouldn't start until finished"). vzdump takes a guest lock, so a concurrent second backup
@@ -899,9 +921,18 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
s.logger.Info("local-api: backup job complete", "vmid", vmid, "target", tier.TargetID, "job", jobID, "archive", b.Archive) s.logger.Info("local-api: backup job complete", "vmid", vmid, "target", tier.TargetID, "job", jobID, "archive", b.Archive)
} }
s.store.RecordBackup(b) s.store.RecordBackup(b)
if b.Success && s.lastKnown != nil {
if err := s.lastKnown.RecordBackupSuccess(tier.TargetID, b); err != nil {
// Not fatal: the backup exists. Only the fallback after a restart loses this copy.
s.logger.Warn("local-api: could not save the backup on disk for the due-check fallback (R-894)",
"vmid", vmid, "target", tier.TargetID, "err", err)
}
}
s.finishJob(key, jobID, b) s.finishJob(key, jobID, b)
// OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate. // OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate.
if b.Success && tier.Primary && s.afterPrimaryBackup != nil { if b.Success && tier.Primary && s.afterPrimaryBackup != nil && manual {
s.logger.Info("local-api: no OS leg after a manual backup — it follows the night's own backup (R-899)", "vmid", vmid, "job", jobID)
} else if b.Success && tier.Primary && s.afterPrimaryBackup != nil {
s.afterPrimaryBackup(base, vmid) s.afterPrimaryBackup(base, vmid)
} }
}() }()
@@ -1061,6 +1092,20 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
newest, haveNewest = t, true newest, haveNewest = t, true
unparseable = false // ground truth supersedes an unreadable in-memory timestamp unparseable = false // ground truth supersedes an unreadable in-memory timestamp
} }
// R-894: the storage could not be read → the last success saved ON DISK stands in for the in-memory
// record a restart emptied. ONLY on archiveUnknown: a storage that answers is the ground truth, and an
// archive absent there must make the tier due even when the file remembers one (a pruned copy).
// A saved copy older than the cadence still reads DUE below — an unreadable storage never suppresses
// a backup that is due.
fromDisk := false
if lookup == archiveUnknown && s.lastKnown != nil {
if saved, ok := s.lastKnown.LastKnownSuccess(tier.TargetID, vmid); ok && (!haveNewest || saved.After(newest)) {
newest, haveNewest, fromDisk = saved, true, true
unparseable = false
s.logger.Info("local-api: backup storage unreadable — due-check uses the last success saved on disk (R-894)",
"vmid", vmid, "target", tier.TargetID, "saved", saved.UTC().Format(time.RFC3339))
}
}
if !haveNewest { if !haveNewest {
// R-88 Part 2: THREE distinct reasons for a nil age, each with its own state. Only ABSENT is a // R-88 Part 2: THREE distinct reasons for a nil age, each with its own state. Only ABSENT is a
// positive claim of "never backed up"; only that one may license the controller to bypass its // positive claim of "never backed up"; only that one may license the controller to bypass its
@@ -1083,13 +1128,17 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
} }
age := s.now().Sub(newest) age := s.now().Sub(newest)
ageSecs := int64(age.Seconds()) ageSecs := int64(age.Seconds())
suffix := ""
if fromDisk {
suffix = " (storage unreadable — age from the last success saved on disk)"
}
if age >= tier.Cadence { if age >= tier.Cadence {
writeOK(w, BackupDueResponse{VMID: vmid, Due: true, AgeSecs: &ageSecs, AgeState: AgeStateKnown, writeOK(w, BackupDueResponse{VMID: vmid, Due: true, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
Reason: "older than cadence", Target: echo}) Reason: "older than cadence" + suffix, Target: echo})
return return
} }
writeOK(w, BackupDueResponse{VMID: vmid, Due: false, AgeSecs: &ageSecs, AgeState: AgeStateKnown, writeOK(w, BackupDueResponse{VMID: vmid, Due: false, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
Reason: "within cadence window", Target: echo}) Reason: "within cadence window" + suffix, Target: echo})
} }
// BackupTiersResponse is GET /backup/tiers (R-82): the tiers this agent serves, primary first. // BackupTiersResponse is GET /backup/tiers (R-82): the tiers this agent serves, primary first.
@@ -1477,4 +1526,6 @@ func writeStatus(w http.ResponseWriter, code int, ok bool, data any, errMsg stri
} }
// SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0). // SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0).
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) { s.afterPrimaryBackup = fn } func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) {
s.afterPrimaryBackup = fn
}
+446
View File
@@ -0,0 +1,446 @@
package osupdate
// The kernel lane (R-836, `09` §3 decisions 164 + 172, `11` §5.11). The root half is felhom-os-apply's layer "kernel"
// (configs/, its own tests); this file decides WHEN and judges the boot:
//
// - the night leg (Run → runKernel): after a healthy host step, on a night the hub marks as told (the household was
// mailed the day before — no mail, no step): ring 0 STAGES the pending kernel (select pending-kernel, the root-owned
// ring-0 mark) and reboots; ring 1 reboots only a step a signed os_kernel_step staged earlier (KernelStepExecutor).
// - after a boot (KernelAfterBoot, at daemon start): the wrapper says what became of the step. On the new kernel the
// agent JUDGES the boot — the host health rule (`11` §8.2: the Proxmox daemons, the guest running and healthy, the
// tunnel) AND the box reaching the hub — for KernelJudgeWait. Healthy → kernel-good (the new kernel becomes the
// default). Not healthy by the deadline → ONE self-revert (kernel-revert: a reboot into the old kernel, still the
// default). A crash on the new kernel needs nothing from the agent: GRUB already boots the old default.
//
// The host is rebooted by this file only through the wrapper (kernel-reboot, kernel-revert), and only for a staged step.
import (
"context"
"encoding/base64"
"encoding/json"
"fmt"
"log/slog"
"regexp"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// OpKernelStep is the signed op class that STAGES a kernel set on a ring-1 box (it never reboots: the night leg does,
// once the household was told). CC may sign it until the first paying customer (R-530 ruling).
const OpKernelStep = "os_kernel_step"
// DefaultKernelJudgeWait is how long a one-shot boot may take to come back healthy before the agent reverts it ONCE.
// Measured 2026-10-07 (`audits/kernel-lane-2026-10-07/B/`): every container healthy 68 s after a reboot on demo-felhom,
// 272 s on demo-hp; 20 minutes stays under the hub's 45-minute host_stale (a box that never comes back alarms after it).
const DefaultKernelJudgeWait = 20 * time.Minute
var kverRE = regexp.MustCompile(`^[0-9]+\.[0-9]+\.[0-9]+-[0-9]+-pve$`)
// KernelView is the wrapper's kernel object (felhom-os-apply Kernel.view).
type KernelView struct {
Running string `json:"running"`
Default string `json:"default"`
Flag *string `json:"flag"`
Phase string `json:"phase"`
From string `json:"from"`
To string `json:"to"`
SelfRevertUsed bool `json:"self_revert_used"`
Reason string `json:"reason"`
VMID int `json:"vmid"` // the customer guest the step was staged for (the health rule's guest)
}
// kernelReportRing is the ring an after-boot report carries. In the first second after a boot the agent has not
// fetched the hub's block yet, and Block() then answers ring 1 — so a ring-0 box's „judging" report said ring 1
// (seen on demo-felhom, 2026-10-08 night, `audits/kernel-night-2026-10-07/readback/`). Order: the fetched block; the
// block the daemon saved on disk before the reboot (R-866); ring 1 as before. A label only: the hub's approval reads its
// own ring list. Pinned by TestKernelReportRing_BeforeFirstFetch.
func (l *Leg) kernelReportRing() int {
l.mu.Lock()
fetched := l.block
l.mu.Unlock()
if fetched != nil {
return fetched.Ring
}
if b, _, ok := LoadSavedBlock(l.planDir()); ok && b != nil {
return b.Ring
}
return 1
}
func parseKernel(raw json.RawMessage) KernelView {
var v KernelView
_ = json.Unmarshal(raw, &v)
return v
}
// kernelPlan is one kernel-layer wrapper call.
func kernelPlan(mode string, vmid int, extra map[string]any) map[string]any {
p := map[string]any{"release_id": "kernel", "layer": LayerKernel, "lane": "slow", "vmid": vmid, "mode": mode,
"packages": []Package{}}
for k, v := range extra {
p[k] = v
}
return p
}
// KernelStatus reads the kernel lane's state (read only).
func (l *Leg) KernelStatus(ctx context.Context, vmid int) (KernelView, error) {
wr, err := l.call(ctx, "kstatus"+l.now().UTC().Format("150405"), kernelPlan("kernel-status", vmid, nil))
if err != nil {
return KernelView{}, err
}
if wr.refused() {
return KernelView{}, fmt.Errorf("kernel-status refused: %s", wr.Refused)
}
return parseKernel(wr.Kernel), nil
}
// kernelDue reports whether tonight's leg may take a kernel step, and why not.
func kernelDue(blk hub.WireOSUpdate, trigger string) (bool, string) {
switch {
case trigger != "night":
return false, "a kernel step runs only in the night leg (never a debug pass)"
case blk.Kernel == nil:
return false, "the hub names no kernel step for this box"
case !blk.Enabled:
return false, "OS updates are switched off for this box"
case !kverRE.MatchString(blk.Kernel.Kver):
return false, "the hub's kernel " + blk.Kernel.Kver + " is not a kernel version"
case !blk.Kernel.Tonight:
return false, "the household has not been told about tonight (no mail, no step — `09` §3 decision 172)"
}
return true, ""
}
// runKernel is the night leg's last step. It returns the stage report (Layer "" when nothing ran). On success the box
// is rebooting when it returns.
func (l *Leg) runKernel(ctx context.Context, runID string, vmid int, trigger string, blk hub.WireOSUpdate) Report {
lg := l.log().With("run", runID, "layer", LayerKernel, "vmid", vmid, "trigger", trigger, "ring", blk.Ring)
if ok, why := kernelDue(blk, trigger); !ok {
lg.Info("osupdate: kernel step skipped — " + why)
return Report{}
}
want := blk.Kernel.Kver
st, err := l.KernelStatus(ctx, vmid)
if err != nil {
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: trigger, Ring: blk.Ring, VMID: vmid,
Mode: "apply", Outcome: "failed", HealthReason: "kernel status unreadable: " + err.Error()})
}
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: trigger, Ring: blk.Ring, VMID: vmid, Mode: "apply", ReleaseID: want}
switch {
case st.Phase == "staged" && st.To == want:
lg.Info("osupdate: kernel step — a staged kernel waits for tonight", "from", st.From, "to", st.To)
rep.Outcome, rep.Healthy = "staged", true
case blk.Ring != 0:
lg.Info("osupdate: kernel step skipped — ring 1 boots only a kernel a signed os_kernel_step staged", "phase", st.Phase, "staged", st.To, "want", want)
return Report{}
default:
// R-898: EXACTLY the kernel the household was told about — never "whatever is pending tonight" (the sources can
// offer a newer one by night; the step then refused, R23, and the night was lost). A version no longer
// installable is refused by the wrapper before any change (R7) and the hub tells the household again.
wr, cerr := l.call(ctx, runID, kernelPlan("apply", vmid, map[string]any{"release_id": "ring0-" + runID,
"select": "listed", "packages": KernelSet(want), "expect_kver": want, "run_id": runID, "trigger": trigger,
"ring": blk.Ring}))
rep.unsent = reportFile(l.planDir(), runID, LayerKernel, "apply")
rep.Kernel = rawOrNil(wr.Kernel)
switch {
case cerr != nil:
rep.Outcome, rep.HealthReason = "failed", cerr.Error()
return l.finish(ctx, lg, rep)
case wr.refused():
rep.Outcome, rep.Refused = "refused", wr.Refused
return l.finish(ctx, lg, rep)
case wr.failed():
rep.Outcome, rep.Refused = "failed", wr.Failed
return l.finish(ctx, lg, rep)
case wr.OutcomeHint == "nothing" || len(wr.Upgraded) == 0:
rep.Outcome, rep.Healthy = "nothing", true
return l.finish(ctx, lg, rep)
}
rep.Outcome, rep.Healthy, rep.Upgraded, rep.Authority, rep.PassSeconds = "staged", true, wr.Upgraded, wr.Authority, wr.PassSeconds
rep.RebootNeeded = true
}
rep = l.finish(ctx, lg, rep) // the hub hears "staged" BEFORE the box goes down
wr, err := l.call(ctx, runID, kernelPlan("kernel-reboot", vmid, nil))
switch {
case err != nil:
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: trigger, Ring: blk.Ring, VMID: vmid,
Mode: "kernel-reboot", ReleaseID: want, Outcome: "failed", HealthReason: "kernel-reboot: " + err.Error()})
case wr.refused() || wr.failed():
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: trigger, Ring: blk.Ring, VMID: vmid,
Mode: "kernel-reboot", ReleaseID: want, Outcome: "refused", Refused: firstRaw(wr.Refused, wr.Failed),
Kernel: rawOrNil(wr.Kernel)})
}
lg.Warn("osupdate: kernel step — the box restarts now for its one-shot boot", "to", want)
return rep
}
func firstRaw(a, b json.RawMessage) json.RawMessage {
if r := rawOrNil(a); r != nil {
return r
}
return rawOrNil(b)
}
// KernelJudge is what KernelAfterBoot needs besides the leg: the hub reachability probe is the "judging" report itself.
type KernelJudge struct {
Wait time.Duration // default DefaultKernelJudgeWait
Poll time.Duration // default 30 s
}
// KernelAfterBoot runs once at daemon start: what became of a kernel step across the boot. On the new kernel it judges
// the boot (blocking up to the wait — run it in a goroutine). vmid 0 = the guest the step recorded (it may not run yet).
func (l *Leg) KernelAfterBoot(ctx context.Context, vmid int, j KernelJudge) Report {
runID := "boot-" + l.now().UTC().Format("20060102T150405Z")
lg := l.log().With("run", runID, "layer", LayerKernel, "vmid", vmid)
wr, err := l.call(ctx, runID, kernelPlan("kernel-boot", vmid, nil))
if err != nil {
lg.Warn("osupdate: kernel after-boot check failed", "err", err)
return Report{}
}
if wr.refused() {
lg.Info("osupdate: kernel after-boot check refused (an older wrapper, or a BYO host)", "refused", string(wr.Refused))
return Report{}
}
v := parseKernel(wr.Kernel)
if vmid <= 0 {
vmid = v.VMID // after a boot the guest may not run yet — the step's own record names it
}
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid, Mode: "kernel-boot",
ReleaseID: v.To, Kernel: rawOrNil(wr.Kernel)}
switch wr.KernelEvent {
case "fell_back":
rep.Outcome, rep.HealthReason = "fell_back", v.Reason
lg.Warn("osupdate: kernel step FELL BACK — the new kernel did not come up; the box runs the old one", "from", v.From, "to", v.To)
return l.finish(ctx, lg, rep)
case "self_reverted":
rep.Outcome, rep.HealthReason = "self_reverted", v.Reason
lg.Warn("osupdate: kernel step SELF-REVERTED — back on the old kernel", "from", v.From, "to", v.To, "reason", v.Reason)
return l.finish(ctx, lg, rep)
case "revert_failed":
rep.Outcome, rep.HealthReason = "revert_failed", v.Reason
lg.Error("osupdate: kernel self-revert came back on the NEW kernel — no second revert; the operator decides", "to", v.To)
return l.finish(ctx, lg, rep)
case "judging":
return l.judgeKernel(ctx, runID, vmid, v, wr.HealthBefore, j, lg)
}
return Report{}
}
// KernelVerdict is THE one-shot boot rule (R-836; pinned by TestKernelVerdict): the host health rule (`11` §8.2 —
// the Proxmox daemons and the agent active, the customer guest running and its own rule passing, the tunnel running)
// AND the box reached the hub since this boot.
func KernelVerdict(before, after *Health, tunnel string, hubReached bool) (bool, string) {
if ok, why := HostHealthVerdict(before, after, tunnel); !ok {
return false, why
}
if !hubReached {
return false, "the box has not reached the hub since the boot"
}
return true, ""
}
func (l *Leg) judgeKernel(ctx context.Context, runID string, vmid int, v KernelView, before *Health, j KernelJudge, lg *slog.Logger) Report {
wait, poll := j.Wait, j.Poll
if wait <= 0 {
wait = DefaultKernelJudgeWait
}
if poll <= 0 {
poll = 30 * time.Second
}
lg.Info("osupdate: kernel step — judging the one-shot boot", "from", v.From, "to", v.To, "wait", wait.String())
start := l.now()
deadline := start.Add(wait)
hubReached := false
var why string
for {
if !hubReached && l.Hub != nil {
// the hub's reachability IS this report reaching it (and the operator sees the box is back on the new kernel)
body, _ := json.Marshal(Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid,
Mode: "kernel-boot", ReleaseID: v.To, Outcome: "judging", Kernel: mustRaw(v)})
rctx, cancel := context.WithTimeout(ctx, 30*time.Second)
if err := l.Hub.PostOSReport(rctx, body); err == nil {
hubReached = true
lg.Info("osupdate: kernel step — the box reached the hub on the new kernel", "after", l.now().Sub(start).Round(time.Second).String())
}
cancel()
}
var h *Health
hr, err := l.call(ctx, runID, kernelPlan("health", vmid, nil))
switch {
case err != nil:
why = "no health reading: " + err.Error()
case hr.refused():
why = "no health reading: " + string(hr.Refused)
default:
h = hr.Health
}
if h != nil {
t := hub.TunnelUnknown
if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx)
}
var ok bool
ok, why = KernelVerdict(before, h, t, hubReached)
if ok {
return l.kernelGood(ctx, runID, vmid, v, start, lg)
}
}
if !l.now().Before(deadline) || ctx.Err() != nil {
break
}
l.sleep(ctx, poll)
}
if ctx.Err() != nil {
lg.Warn("osupdate: kernel judging stopped (the agent is stopping) — the next start judges again", "reason", why)
return Report{}
}
// not healthy by the deadline: tell the hub (best effort), then ONE self-revert into the old kernel
rep := l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid,
Mode: "kernel-revert", ReleaseID: v.To, Outcome: "health_failed", HealthReason: why + " — reverting to " + v.From,
Kernel: mustRaw(v)})
lg.Error("osupdate: kernel step — the one-shot boot is NOT healthy; restarting ONCE into the old kernel", "reason", why,
"waited", wait.String(), "from", v.From, "to", v.To)
wr, err := l.call(ctx, runID, kernelPlan("kernel-revert", vmid, map[string]any{"reason": truncate(why, 280)}))
if err != nil || wr.refused() || wr.failed() {
lg.Error("osupdate: kernel self-revert did not start — the box stays on the new kernel; the operator decides",
"err", err, "refused", string(firstRaw(wr.Refused, wr.Failed)))
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid,
Mode: "kernel-revert", ReleaseID: v.To, Outcome: "revert_failed", Refused: firstRaw(wr.Refused, wr.Failed),
HealthReason: "the self-revert did not start"})
}
return rep
}
func (l *Leg) kernelGood(ctx context.Context, runID string, vmid int, v KernelView, start time.Time, lg *slog.Logger) Report {
wr, err := l.call(ctx, runID, kernelPlan("kernel-good", vmid, nil))
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid, Mode: "kernel-good",
ReleaseID: v.To}
switch {
case err != nil:
rep.Outcome, rep.HealthReason = "failed", "kernel-good: "+err.Error()
case wr.refused() || wr.failed():
rep.Outcome, rep.Refused, rep.HealthReason = "failed", firstRaw(wr.Refused, wr.Failed), "kernel-good did not move the default"
default:
rep.Outcome, rep.Healthy = "applied", true
rep.HealthReason = fmt.Sprintf("healthy %s after the agent started; the new kernel is the default", l.now().Sub(start).Round(time.Second))
}
rep.Kernel = rawOrNil(wr.Kernel)
lg.Info("osupdate: kernel step — "+rep.Outcome, "to", v.To, "reason", rep.HealthReason)
return l.finish(ctx, lg, rep)
}
func mustRaw(v any) json.RawMessage {
b, _ := json.Marshal(v)
return b
}
func truncate(s string, n int) string {
if len(s) <= n {
return s
}
return s[:n]
}
// KernelSet is the package set that installs exactly kver (R-898; the hub's kernelSet, field-exact): the series
// meta-package and the signed image, both at the kernel's own version. Proxmox keeps old kernel versions in its archive.
// nil for a string that is not a kernel version.
func KernelSet(kver string) []Package {
m := kverSeriesRE.FindStringSubmatch(kver)
if m == nil {
return nil
}
v := kver[:len(kver)-len("-pve")]
return []Package{{Name: "proxmox-kernel-" + m[1], Version: v, Origin: PVEOrigin},
{Name: "proxmox-kernel-" + kver + "-signed", Version: v, Origin: PVEOrigin}}
}
var kverSeriesRE = regexp.MustCompile(`^([0-9]+\.[0-9]+)\.[0-9]+-[0-9]+-pve$`)
// KernelStepParams are a signed os_kernel_step's params: the exact kernel set (the wrapper compares it with the plan).
type KernelStepParams struct {
ReleaseID string `json:"release_id"`
Packages []Package `json:"packages"`
Kver string `json:"kver"`
VMID int `json:"vmid,omitempty"`
}
// KernelStepExecutor STAGES a verified os_kernel_step (signedjobs.Executor) under the heavy-op gate. It never reboots:
// the night leg reboots a staged kernel on a night the household was told about.
type KernelStepExecutor struct {
Leg *Leg
Guest func(ctx context.Context) (int, error)
Gate func(ctx context.Context) (release func(), err error)
}
// Execute implements signedjobs.Executor.
func (e KernelStepExecutor) Execute(ctx context.Context, op string, params json.RawMessage) error {
if op != OpKernelStep {
return signedjobs.ErrNoExecutor
}
so, ok := signedjobs.SignedOpFrom(ctx)
if !ok {
return fmt.Errorf("os_kernel_step: no signed envelope in the context — the wrapper could not verify it")
}
var p KernelStepParams
if err := json.Unmarshal(params, &p); err != nil || len(p.Packages) == 0 || !kverRE.MatchString(p.Kver) {
return fmt.Errorf("os_kernel_step: params must name the kernel set and its kver: %v", err)
}
vmid := p.VMID
if vmid == 0 {
if e.Guest == nil {
return fmt.Errorf("os_kernel_step: no vmid and no guest finder")
}
v, err := e.Guest(ctx)
if err != nil {
return fmt.Errorf("os_kernel_step: find the customer guest: %w", err)
}
vmid = v
}
if e.Gate != nil {
release, err := e.Gate(ctx)
if err != nil {
return fmt.Errorf("os_kernel_step: heavy-op gate busy (a backup or restore-test runs): %w", err)
}
defer release()
}
rep := e.Leg.StageKernelSigned(ctx, vmid, p, so.Blob, string(so.Sig))
if rep.Outcome == "staged" || rep.Outcome == "nothing" {
return nil
}
return fmt.Errorf("os_kernel_step: %s (%s) %s", rep.Outcome, rep.HealthReason, string(rep.Refused))
}
// StageKernelSigned stages a signed kernel set (ring 1): install + flag, no reboot.
func (l *Leg) StageKernelSigned(ctx context.Context, vmid int, p KernelStepParams, blob []byte, sig string) Report {
unlock := l.lockPass(true)
defer unlock()
l.sendUnsentLocked(ctx) // R-868
runID := l.now().UTC().Format("20060102T150405Z")
rid := p.ReleaseID
if rid == "" {
rid = "signed-" + runID
}
lg := l.log().With("run", runID, "layer", LayerKernel, "vmid", vmid, "trigger", "signed", "release", rid)
wr, err := l.call(ctx, runID, kernelPlan("apply", vmid, map[string]any{"release_id": rid, "select": "listed",
"packages": p.Packages, "expect_kver": p.Kver, "run_id": runID, "trigger": "signed", "ring": l.Block().Ring,
"signed": map[string]string{"blob_b64": base64.StdEncoding.EncodeToString(blob), "sig": sig}}))
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "signed", Ring: l.Block().Ring, VMID: vmid, Mode: "apply",
ReleaseID: rid, Kernel: rawOrNil(wr.Kernel), unsent: reportFile(l.planDir(), runID, LayerKernel, "apply")}
switch {
case err != nil:
rep.Outcome, rep.HealthReason = "failed", err.Error()
case wr.refused():
rep.Outcome, rep.Refused = "refused", wr.Refused
case wr.failed():
rep.Outcome, rep.Refused = "failed", wr.Failed
case wr.OutcomeHint == "nothing" || len(wr.Upgraded) == 0:
rep.Outcome, rep.Healthy = "nothing", true
default:
rep.Outcome, rep.Healthy, rep.Upgraded, rep.Authority, rep.PassSeconds = "staged", true, wr.Upgraded, wr.Authority, wr.PassSeconds
rep.RebootNeeded = true
}
return l.finish(ctx, lg, rep)
}
+328
View File
@@ -0,0 +1,328 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// ---- the kernel lane (R-836, `09` §3 decision 172, `11` §5.11) ----
const kOld, kNew = "7.0.2-6-pve", "7.0.14-22-pve"
func kview(phase string) json.RawMessage {
return mustRaw(KernelView{Running: kOld, Default: kOld, Phase: phase, From: kOld, To: kNew, VMID: 9201})
}
func tonight(ring int) *hub.WireOSUpdate {
return &hub.WireOSUpdate{Ring: ring, Enabled: true, Kernel: &hub.WireKernelStep{Kver: kNew, Tonight: true}}
}
func kernelCalls(w *fakeWrapper) []string {
var m []string
for _, p := range w.plans {
if p["layer"] == LayerKernel {
m = append(m, p["mode"].(string))
}
}
return m
}
// Ring 0, a told night: after the healthy host step the leg stages the pending kernel (select pending-kernel, the
// kernel the household was told about), tells the hub "staged", THEN reboots — the kernel step ends the night.
func TestKernel_Ring0ToldNightStagesThenReboots(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{
"kernel-status": {{Kernel: kview("none")}},
"apply": {{Upgraded: []Package{{Name: "proxmox-kernel-7.0", Version: "7.0.14-22"}}, Authority: "ring0", Kernel: kview("staged")}},
"kernel-reboot": {{Kernel: kview("oneshot")}},
}}
l, h := newLeg(t, w, tonight(0))
p := l.Run(context.Background(), 9201, "night")
if got := strings.Join(kernelCalls(w), ","); got != "kernel-status,apply,kernel-reboot" {
t.Fatalf("kernel calls = %s", got)
}
var ap map[string]any
for _, x := range w.plans {
if x["layer"] == LayerKernel && x["mode"] == "apply" {
ap = x
}
}
// R-898: EXACTLY the told kernel — the listed set derived from it, never "pending" (red before the fix: select was
// pending-kernel, so a newer kernel in the sources by night was refused R23 and the night was lost)
pk, _ := json.Marshal(ap["packages"])
if ap["select"] != "listed" || ap["expect_kver"] != kNew || ap["lane"] != "slow" ||
string(pk) != `[{"name":"proxmox-kernel-7.0","origin":"Proxmox Debian Repository","version":"7.0.14-22"},{"name":"proxmox-kernel-7.0.14-22-pve-signed","origin":"Proxmox Debian Repository","version":"7.0.14-22"}]` {
t.Fatalf("stage plan = %v (packages %s)", ap, pk)
}
if p.Kernel.Outcome != "staged" || !p.Kernel.Healthy {
t.Fatalf("kernel report = %+v", p.Kernel)
}
last := h.reports[len(h.reports)-1]
if last.Layer != LayerKernel || last.Outcome != "staged" {
t.Fatalf("the hub must hear 'staged' before the reboot: %+v", h.reports)
}
// the kernel step is the LAST wrapper call of the night
if lp := w.plans[len(w.plans)-1]; lp["layer"] != LayerKernel || lp["mode"] != "kernel-reboot" {
t.Fatalf("the reboot must end the night, last call = %v", lp)
}
}
// No mail, no step: a kernel the household was NOT told about never runs; nor in a debug pass; nor without a block.
// COMPANION RED-PROOF (observed): drop the `!blk.Kernel.Tonight` case in kernelDue → the first sub-case fails.
func TestKernel_NoMailNoStep(t *testing.T) {
cases := map[string]struct {
blk *hub.WireOSUpdate
trigger string
}{
"not told": {&hub.WireOSUpdate{Ring: 0, Enabled: true, Kernel: &hub.WireKernelStep{Kver: kNew, Tonight: false}}, "night"},
"debug pass": {tonight(0), "debug"},
"no block": {&hub.WireOSUpdate{Ring: 0, Enabled: true}, "night"},
"switch off": {&hub.WireOSUpdate{Ring: 0, Enabled: false, Kernel: &hub.WireKernelStep{Kver: kNew, Tonight: true}}, "night"},
"bad kver": {&hub.WireOSUpdate{Ring: 0, Enabled: true, Kernel: &hub.WireKernelStep{Kver: "7.0; reboot", Tonight: true}}, "night"},
}
for name, c := range cases {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, c.blk)
p := l.Run(context.Background(), 9201, c.trigger)
if len(kernelCalls(w)) != 0 || p.Kernel.Layer != "" {
t.Fatalf("%s: a kernel step ran: %v", name, kernelCalls(w))
}
}
}
// The kernel step needs a healthy host step on an appliance, and a healthy Proxmox step when one ran.
func TestKernel_SkippedWithoutHealthyEarlierSteps(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, tonight(0))
l.Appliance = false
l.Run(context.Background(), 9201, "night")
if len(kernelCalls(w)) != 0 {
t.Fatalf("a BYO box took a kernel step: %v", kernelCalls(w))
}
w2 := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerPVE: {Upgraded: []Package{{Name: "pve-manager", Version: "9.2.21"}},
PVEManager: "9.2.2"}}} // pveversion still old → the pve step is unhealthy
l2, _ := newLeg(t, w2, tonight(0))
l2.Run(context.Background(), 9201, "night")
if len(kernelCalls(w2)) != 0 {
t.Fatalf("a kernel step ran after an unhealthy Proxmox step: %v", kernelCalls(w2))
}
}
// Ring 1 reboots only a kernel a signed job staged — never stages one itself in the night leg.
func TestKernel_Ring1RebootsOnlyASignedStage(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-status": {{Kernel: kview("none")}}}}
l, _ := newLeg(t, w, tonight(1))
l.Run(context.Background(), 9201, "night")
if got := strings.Join(kernelCalls(w), ","); got != "kernel-status" {
t.Fatalf("ring 1 without a staged kernel: calls = %s", got)
}
w2 := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-status": {{Kernel: kview("staged")}},
"kernel-reboot": {{Kernel: kview("oneshot")}}}}
l2, _ := newLeg(t, w2, tonight(1))
p := l2.Run(context.Background(), 9201, "night")
if got := strings.Join(kernelCalls(w2), ","); got != "kernel-status,kernel-reboot" || p.Kernel.Outcome != "staged" {
t.Fatalf("ring 1 with a staged kernel: calls = %s report = %+v", got, p.Kernel)
}
}
// A refused stage never reboots.
func TestKernel_RefusedStageNeverReboots(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-status": {{Kernel: kview("none")}},
"apply": {{Refused: json.RawMessage(`{"code":"R20","reason":"/boot/efi is not a mounted vfat ESP"}`)}}}}
l, h := newLeg(t, w, tonight(0))
p := l.Run(context.Background(), 9201, "night")
if got := strings.Join(kernelCalls(w), ","); got != "kernel-status,apply" || p.Kernel.Outcome != "refused" {
t.Fatalf("calls = %s report = %+v", got, p.Kernel)
}
if last := h.reports[len(h.reports)-1]; last.Layer != LayerKernel || last.Outcome != "refused" {
t.Fatalf("the hub must hear the refusal: %+v", last)
}
}
// THE one-shot boot rule: the host rule AND the hub reached. COMPANION RED-PROOF (observed): drop the hubReached
// check in KernelVerdict → the second case fails.
func TestKernelVerdict(t *testing.T) {
if ok, why := KernelVerdict(hostOK(), hostOK(), hub.TunnelRunning, true); !ok {
t.Fatalf("a healthy boot read unhealthy: %s", why)
}
if ok, why := KernelVerdict(hostOK(), hostOK(), hub.TunnelRunning, false); ok || !strings.Contains(why, "hub") {
t.Fatalf("a box that has not reached the hub must not pass: ok=%v %q", ok, why)
}
down := hostOK()
down.GuestRunning = new(bool)
if ok, _ := KernelVerdict(hostOK(), down, hub.TunnelRunning, true); ok {
t.Fatal("a guest that does not run must fail")
}
if ok, _ := KernelVerdict(hostOK(), hostOK(), hub.TunnelUnknown, true); ok {
t.Fatal("an unknown tunnel must fail (the host rule)")
}
}
func judgingLeg(t *testing.T, w *fakeWrapper) (*Leg, *fakeHub) {
if w.kernelRep == nil {
w.kernelRep = map[string][]WrapperReport{}
}
if _, ok := w.kernelRep["kernel-boot"]; !ok {
w.kernelRep["kernel-boot"] = []WrapperReport{{KernelEvent: "judging", Kernel: mustRaw(KernelView{Running: kNew,
Default: kOld, Phase: "judging", From: kOld, To: kNew, VMID: 9201}), HealthBefore: hostOK()}}
}
return newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
}
// A healthy one-shot boot: the hub hears "judging", then kernel-good, then "applied".
func TestKernelAfterBoot_HealthyBecomesTheDefault(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-good": {{Kernel: kview("good")}}}}
l, h := judgingLeg(t, w)
r := l.KernelAfterBoot(context.Background(), 0, KernelJudge{Wait: 10 * time.Minute, Poll: 30 * time.Second})
if got := strings.Join(kernelCalls(w), ","); got != "kernel-boot,health,kernel-good" {
t.Fatalf("calls = %s", got)
}
if r.Outcome != "applied" || !r.Healthy {
t.Fatalf("report = %+v", r)
}
if len(h.reports) != 2 || h.reports[0].Outcome != "judging" || h.reports[1].Outcome != "applied" {
t.Fatalf("hub reports = %+v", h.reports)
}
if w.plans[1]["vmid"] != float64(9201) {
t.Fatalf("the health reading must use the step's own guest, got %v", w.plans[1]["vmid"])
}
}
// An unhealthy one-shot boot: wait the full judge time, tell the hub, then ONE kernel-revert.
// COMPANION RED-PROOF (observed): return before the kernel-revert call in judgeKernel → "calls" fails.
func TestKernelAfterBoot_UnhealthyRevertsOnceAfterTheWait(t *testing.T) {
down := hostOK()
down.GuestRunning = new(bool)
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"health": {{Health: down}},
"kernel-revert": {{Kernel: kview("reverting")}}}}
l, h := judgingLeg(t, w)
start := l.now()
r := l.KernelAfterBoot(context.Background(), 0, KernelJudge{Wait: 10 * time.Minute, Poll: time.Minute})
calls := kernelCalls(w)
if calls[len(calls)-1] != "kernel-revert" || strings.Count(strings.Join(calls, ","), "kernel-revert") != 1 {
t.Fatalf("calls = %v", calls)
}
if waited := l.now().Sub(start); waited < 10*time.Minute {
t.Fatalf("reverted after %s — before the judge wait", waited)
}
if r.Outcome != "health_failed" || !strings.Contains(r.HealthReason, "not running") {
t.Fatalf("report = %+v", r)
}
if last := h.reports[len(h.reports)-1]; last.Outcome != "health_failed" {
t.Fatalf("the hub must hear health_failed before the revert reboot: %+v", h.reports)
}
for _, p := range w.plans {
if p["mode"] == "kernel-revert" && !strings.Contains(p["reason"].(string), "not running") {
t.Fatalf("the revert must carry the reason: %v", p)
}
}
}
// A box that never reaches the hub is not "healthy" — it reverts too.
func TestKernelAfterBoot_NoHubMeansRevert(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-revert": {{Kernel: kview("reverting")}}}}
l, _ := judgingLeg(t, w)
l.Hub = unreachableHub{}
l.KernelAfterBoot(context.Background(), 0, KernelJudge{Wait: 5 * time.Minute, Poll: time.Minute})
if c := kernelCalls(w); c[len(c)-1] != "kernel-revert" {
t.Fatalf("calls = %v", c)
}
}
type unreachableHub struct{}
func (unreachableHub) PostOSReport(context.Context, []byte) error { return context.DeadlineExceeded }
// What kernel-boot found becomes the hub's outcome, with no judging and no reboot.
func TestKernelAfterBoot_FallBackAndRevertResultsAreReported(t *testing.T) {
for ev, want := range map[string]string{"fell_back": "fell_back", "self_reverted": "self_reverted", "revert_failed": "revert_failed"} {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-boot": {{KernelEvent: ev,
Kernel: mustRaw(KernelView{Running: kOld, Default: kOld, Phase: ev, From: kOld, To: kNew, Reason: "r", VMID: 9201})}}}}
l, h := judgingLeg(t, w)
r := l.KernelAfterBoot(context.Background(), 0, KernelJudge{})
if r.Outcome != want || len(h.reports) != 1 || h.reports[0].Outcome != want {
t.Fatalf("%s: report %+v hub %+v", ev, r, h.reports)
}
if got := strings.Join(kernelCalls(w), ","); got != "kernel-boot" {
t.Fatalf("%s: calls = %s", ev, got)
}
}
// nothing to do → nothing reported
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"kernel-boot": {{KernelEvent: "none", Kernel: kview("good")}}}}
l, h := judgingLeg(t, w)
if r := l.KernelAfterBoot(context.Background(), 0, KernelJudge{}); r.Layer != "" || len(h.reports) != 0 {
t.Fatalf("an ordinary boot must report nothing: %+v %+v", r, h.reports)
}
}
// The signed executor STAGES (listed + the raw envelope + the kver) and never reboots.
func TestKernelStepExecutor_StagesNeverReboots(t *testing.T) {
w := &fakeWrapper{t: t, kernelRep: map[string][]WrapperReport{"apply": {{Upgraded: []Package{{Name: "proxmox-kernel-7.0",
Version: "7.0.14-22"}}, Authority: "signed", Kernel: kview("staged")}}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
e := KernelStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
params, _ := json.Marshal(KernelStepParams{ReleaseID: "os-kernel-1", Kver: kNew,
Packages: []Package{{Name: "proxmox-kernel-7.0", Version: "7.0.14-22", Origin: PVEOrigin}}})
ctx := signedjobs.WithSignedOp(context.Background(), &reconcile.SignedOp{Blob: []byte(`{"op":"os_kernel_step"}`), Sig: []byte("SIG")})
if err := e.Execute(ctx, OpKernelStep, params); err != nil {
t.Fatal(err)
}
pp := w.plans[len(w.plans)-1]
sg, _ := pp["signed"].(map[string]any)
if pp["mode"] != "apply" || pp["select"] != "listed" || pp["expect_kver"] != kNew || sg == nil ||
sg["blob_b64"] != base64.StdEncoding.EncodeToString([]byte(`{"op":"os_kernel_step"}`)) {
t.Fatalf("plan = %v", pp)
}
if got := strings.Join(kernelCalls(w), ","); got != "apply" {
t.Fatalf("a signed stage must never reboot: %s", got)
}
if len(h.reports) != 1 || h.reports[0].Outcome != "staged" {
t.Fatalf("hub = %+v", h.reports)
}
if err := e.Execute(context.Background(), OpKernelStep, params); err == nil {
t.Fatal("no envelope must refuse")
}
bad, _ := json.Marshal(KernelStepParams{Kver: "x", Packages: []Package{{Name: "a"}}})
if err := e.Execute(ctx, OpKernelStep, bad); err == nil {
t.Fatal("a bad kver must refuse")
}
if err := e.Execute(ctx, OpPVEStep, params); err != signedjobs.ErrNoExecutor {
t.Fatalf("another op must pass through the chain: %v", err)
}
}
// os_kernel_step is never benign.
func TestKernelStep_IsDestructiveClass(t *testing.T) {
if reconcile.Classify(reconcile.ClassOSKernelStep, reconcile.Provenance{}) != reconcile.Destructive {
t.Fatal("os_kernel_step must be destructive-class (signed, operational key)")
}
}
// A kept stage report (the agent was killed mid-stage) reaches the hub as "staged" with its kernel view.
func TestKernel_KeptStageReportIsStaged(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, tonight(0))
ring := 0
rep := l.reportFromKept(context.Background(), WrapperReport{Layer: LayerKernel, Mode: "apply", Ring: &ring,
Upgraded: []Package{{Name: "proxmox-kernel-7.0", Version: "7.0.14-22"}}, Kernel: kview("staged")}, "/x/report-r-kernel-apply.json")
if rep.Outcome != "staged" || !rep.Healthy || len(rep.Kernel) == 0 {
t.Fatalf("kept = %+v", rep)
}
}
func TestKernelSet(t *testing.T) {
if got := KernelSet("7.0.14-20-pve"); len(got) != 2 || got[0].Name != "proxmox-kernel-7.0" || got[0].Version != "7.0.14-20" ||
got[1].Name != "proxmox-kernel-7.0.14-20-pve-signed" {
t.Fatalf("%+v", got)
}
if KernelSet("7.0; reboot") != nil || KernelSet("") != nil {
t.Fatal("a non-kernel string must give no set")
}
}
+168 -6
View File
@@ -20,6 +20,7 @@ import (
"log/slog" "log/slog"
"os" "os"
"path/filepath" "path/filepath"
"regexp"
"sort" "sort"
"strings" "strings"
"sync" "sync"
@@ -27,6 +28,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/hub" "gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox" "gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
) )
// WrapperPath is the pinned sudoers vector (configs/felhom-agent.sudoers FELHOM_OSAPPLY). // WrapperPath is the pinned sudoers vector (configs/felhom-agent.sudoers FELHOM_OSAPPLY).
@@ -40,8 +42,24 @@ const (
LayerGuest = "guest" LayerGuest = "guest"
LayerHost = "host" LayerHost = "host"
LayerDocker = "docker" // the guest's Docker engine set — slow lane (`11` §5.8) LayerDocker = "docker" // the guest's Docker engine set — slow lane (`11` §5.8)
// LayerPVE is the HOST's Proxmox userspace packages — slow lane (R-812 option A, `09` §3 decision 163, `11` §5.10):
// ring 0 in the night leg after a healthy host step, ring 1 only inside a signed os_pve_step. Never a kernel.
LayerPVE = "pve"
// LayerKernel is the HOST's kernel — the kernel lane (R-836, `09` §3 decision 172, `11` §5.11): a one-shot boot of
// the new kernel through the ESP flag, the default moved only after a healthy boot (kernel.go).
LayerKernel = "kernel"
) )
// PVEOrigin is the origin apt prints for download.proxmox.com (the wrapper's PVE_ORIGIN); InstalledPVEOrigin is how
// the wrapper's inventory names the same source.
const (
PVEOrigin = "Proxmox Debian Repository"
InstalledPVEOrigin = "Proxmox"
)
// hostSlowRE mirrors the wrapper's HOST_SLOW_RE: kernel, boot, firmware and microcode names never ride the pve lane.
var hostSlowRE = regexp.MustCompile(`^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr)`)
// DockerNames are the six packages of the Docker engine set (the wrapper's DOCKER_NAMES). // DockerNames are the six packages of the Docker engine set (the wrapper's DOCKER_NAMES).
var DockerNames = map[string]bool{"containerd.io": true, "docker-buildx-plugin": true, "docker-ce": true, var DockerNames = map[string]bool{"containerd.io": true, "docker-buildx-plugin": true, "docker-ce": true,
"docker-ce-cli": true, "docker-ce-rootless-extras": true, "docker-compose-plugin": true} "docker-ce-cli": true, "docker-ce-rootless-extras": true, "docker-compose-plugin": true}
@@ -106,6 +124,16 @@ type WrapperReport struct {
LiveRestore json.RawMessage `json:"live_restore"` LiveRestore json.RawMessage `json:"live_restore"`
Facts json.RawMessage `json:"facts"` Facts json.RawMessage `json:"facts"`
Bundle json.RawMessage `json:"bundle"` // the config bundle's result (R-840, mode "bundle") Bundle json.RawMessage `json:"bundle"` // the config bundle's result (R-840, mode "bundle")
// OOMCheck (R-528, `09` decision 157): the docker layer's memory-kill check, {result, oom_killed, oom_event,
// exit_code, image, detail}. Carried to the hub UNCHANGED; the agent never reads it.
OOMCheck json.RawMessage `json:"oom_check"`
// PVEManager: the pve layer — pveversion's pve-manager version after the step ("unknown" = unreadable).
PVEManager string `json:"pve_manager"`
// Kernel (R-836): the kernel layer's view {running, default, flag, phase, from, to, …}; KernelEvent what kernel-boot
// found after a boot; OutcomeHint "nothing" when no kernel was pending.
Kernel json.RawMessage `json:"kernel"`
KernelEvent string `json:"kernel_event"`
OutcomeHint string `json:"outcome_hint"`
// R-868 (v0.144.0): the agent's ids, echoed from the plan, so a report kept on disk can be sent without the // R-868 (v0.144.0): the agent's ids, echoed from the plan, so a report kept on disk can be sent without the
// agent process that started the pass. ReleaseID / VMID were always in the report. // agent process that started the pass. ReleaseID / VMID were always in the report.
RunID string `json:"run_id"` RunID string `json:"run_id"`
@@ -143,6 +171,15 @@ type Report struct {
DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step
Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed
Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade) Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade)
// OOMCheck: docker layer — the wrapper's oom_check object, byte-for-byte (R-528; the hub decides approval on it).
// Pinned by TestDocker_OOMCheckReachesTheHubUnchanged and TestR868_KeptCopyCarriesTheOOMCheck.
OOMCheck json.RawMessage `json:"oom_check,omitempty"`
// PVEManager: pve layer — pve-manager's version after the step (the hub's System page; R-812 option A).
PVEManager string `json:"pve_manager,omitempty"`
// Kernel: kernel layer — the wrapper's kernel view, byte for byte (running, default, flag, phase, from, to). The
// outcomes of this layer: staged | applied (the new kernel is the default) | fell_back | health_failed (self-revert
// started) | self_reverted | revert_failed | judging | nothing | refused | failed.
Kernel json.RawMessage `json:"kernel,omitempty"`
unsent string // R-868: the wrapper's kept copy of this pass's report — deleted once the hub has it unsent string // R-868: the wrapper's kept copy of this pass's report — deleted once the hub has it
} }
@@ -380,6 +417,36 @@ func EngineOf(pkgVersion string) string {
return v return v
} }
// PVEHealthVerdict is THE Proxmox-package-step health rule (R-812 option A; pinned by TestPVEHealthVerdict): the host
// rule (every host service active, the guest running and healthy, the tunnel running), plus every container running at
// the start still runs as the SAME container (a Proxmox step must not restart the household's apps), plus pveversion
// now reports the pve-manager the step installed (wantPVE "" = pve-manager was not in the step).
func PVEHealthVerdict(before, after *Health, tunnel, wantPVE, gotPVE string) (bool, string) {
if ok, why := HostHealthVerdict(before, after, tunnel); !ok {
return false, why
}
if before != nil && before.Guest != nil && after.Guest != nil {
names := make([]string, 0, len(before.Guest.Containers))
for n := range before.Guest.Containers {
names = append(names, n)
}
sort.Strings(names)
for _, n := range names {
b := before.Guest.Containers[n]
if b.State != "running" || b.ID == "" {
continue
}
if a := after.Guest.Containers[n]; a.ID != b.ID {
return false, n + " is a new container (id changed) — the Proxmox step restarted the household's app"
}
}
}
if wantPVE != "" && gotPVE != wantPVE {
return false, "pveversion reads pve-manager " + gotPVE + ", not " + wantPVE
}
return true, ""
}
// DockerHealthVerdict is THE Docker-step health rule (`11` §5.8; pinned by TestDockerHealthVerdict): the guest rule, // DockerHealthVerdict is THE Docker-step health rule (`11` §5.8; pinned by TestDockerHealthVerdict): the guest rule,
// plus every container running at the start still runs as the SAME container (same id — a changed id means the // plus every container running at the start still runs as the SAME container (same id — a changed id means the
// household's apps restarted, which `live-restore` exists to prevent), plus the engine now reports the version the step // household's apps restarted, which `live-restore` exists to prevent), plus the engine now reports the version the step
@@ -445,7 +512,7 @@ func (l *Leg) call(ctx context.Context, runID string, plan map[string]any) (Wrap
// Pass is one leg's reports; an empty Layer means the step did not run. // Pass is one leg's reports; an empty Layer means the step did not run.
type Pass struct { type Pass struct {
Guest, Host, Docker Report Guest, Host, Docker, PVE, Kernel Report
} }
// Run is one pass: the guest layer, then (on an appliance, after a good guest step) the host layer, then (ring 0 // Run is one pass: the guest layer, then (on an appliance, after a good guest step) the host layer, then (ring 0
@@ -473,13 +540,54 @@ func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Pass {
if err := l.EnsureLiveRestore(ctx, g.RunID, vmid); err != nil { if err := l.EnsureLiveRestore(ctx, g.RunID, vmid); err != nil {
p.Docker = l.finish(ctx, lg, Report{RunID: g.RunID, Layer: LayerDocker, Trigger: trigger, Ring: 0, VMID: vmid, p.Docker = l.finish(ctx, lg, Report{RunID: g.RunID, Layer: LayerDocker, Trigger: trigger, Ring: 0, VMID: vmid,
Mode: "apply", Outcome: "failed", HealthReason: "live-restore could not be turned on: " + err.Error()}) Mode: "apply", Outcome: "failed", HealthReason: "live-restore could not be turned on: " + err.Error()})
return p } else {
}
p.Docker = l.runLayer(ctx, g.RunID, LayerDocker, vmid, trigger, blk, dockerOpts{}) p.Docker = l.runLayer(ctx, g.RunID, LayerDocker, vmid, trigger, blk, dockerOpts{})
} }
}
// R-812 option A: the Proxmox package step — ring 0, an appliance, after a HEALTHY host step (it is a host change).
// A Docker step's outcome does not gate it (the Docker set lives in the guest). Pinned by TestPVE_*.
switch {
case blk.Ring != 0 || !blk.Enabled:
lg.Info("osupdate: pve step skipped — ring 1 takes a Proxmox set only inside a signed operator job (`11` §5.10)", "ring", blk.Ring, "enabled", blk.Enabled)
case !l.Appliance || h.Layer == "" || !okStep(h):
lg.Info("osupdate: pve step skipped — no healthy host step this pass (an appliance only)", "appliance", l.Appliance, "host_outcome", h.Outcome)
default:
p.PVE = l.runPVE(ctx, g.RunID, vmid, trigger, blk, dockerOpts{})
}
// R-836 (`09` §3 decision 172): the kernel step ENDS the night — an appliance, after a healthy host step (and a
// healthy Proxmox step when one ran), only on a night the hub marks as told. It reboots the box. Pinned by TestKernel_*.
switch {
case !l.Appliance || h.Layer == "" || !okStep(h):
lg.Info("osupdate: kernel step skipped — no healthy host step this pass (an appliance only)", "appliance", l.Appliance, "host_outcome", h.Outcome)
case p.PVE.Layer != "" && !okStep(p.PVE):
lg.Warn("osupdate: kernel step skipped — the Proxmox step did not end healthy", "pve_outcome", p.PVE.Outcome)
default:
p.Kernel = l.runKernel(ctx, g.RunID, vmid, trigger, blk)
}
return p return p
} }
// pveDrainWait bounds how long a pve step waits for the agent's own /etc/pve writes in flight (pvegate).
var pveDrainWait = 2 * time.Minute
// runPVE runs the pve layer while holding pvegate: the agent's own /etc/pve writes wait until it ends.
func (l *Leg) runPVE(ctx context.Context, runID string, vmid int, trigger string, blk hub.WireOSUpdate, do dockerOpts) Report {
lg := l.log().With("run", runID, "layer", LayerPVE, "vmid", vmid, "trigger", trigger)
dctx, cancel := context.WithTimeout(ctx, pveDrainWait)
end, err := pvegate.Step(dctx)
cancel()
if err != nil {
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerPVE, Trigger: trigger, Ring: blk.Ring, VMID: vmid, Mode: "apply",
ReleaseID: do.releaseID, Outcome: "failed", HealthReason: "an agent write to /etc/pve did not finish in time (pvegate): " + err.Error()})
}
lg.Info("osupdate: pve step holds the /etc/pve write gate — the agent's own writes wait until it ends")
defer func() {
end()
lg.Info("osupdate: pve step released the /etc/pve write gate")
}()
return l.runLayer(ctx, runID, LayerPVE, vmid, trigger, blk, do)
}
// dockerOpts is a signed Docker step (DockerStepExecutor); the zero value is ring 0's unsigned "pending-docker". // dockerOpts is a signed Docker step (DockerStepExecutor); the zero value is ring 0's unsigned "pending-docker".
type dockerOpts struct { type dockerOpts struct {
releaseID string releaseID string
@@ -565,7 +673,7 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
wire = blk.HostRelease wire = blk.HostRelease
} }
lane := "fast" lane := "fast"
if layer == LayerDocker { if layer == LayerDocker || layer == LayerPVE {
lane = "slow" lane = "slow"
if do.signed != nil { if do.signed != nil {
rel = hub.WireOSRelease{ID: do.releaseID} rel = hub.WireOSRelease{ID: do.releaseID}
@@ -598,6 +706,13 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
} }
case layer == LayerDocker: case layer == LayerDocker:
plan["select"] = "pending-docker" // ring 0: the wrapper checks the box's ROOT-OWNED ring-0 mark itself plan["select"] = "pending-docker" // ring 0: the wrapper checks the box's ROOT-OWNED ring-0 mark itself
case layer == LayerPVE && do.signed != nil:
plan["packages"], plan["signed"] = do.packages, do.signed
for _, p := range do.packages {
planned[p.Name] = true
}
case layer == LayerPVE:
plan["select"] = "pending-pve" // ring 0: the same root-owned mark; the wrapper picks installed Proxmox userspace
case !blk.Enabled: case !blk.Enabled:
plan["mode"] = "inventory" plan["mode"] = "inventory"
lg.Info("osupdate: switched OFF for this box — reporting only") lg.Info("osupdate: switched OFF for this box — reporting only")
@@ -630,6 +745,8 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
} }
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
rep.OOMCheck = rawOrNil(wr.OOMCheck)
rep.PVEManager = wr.PVEManager
if rep.Outcome == "" { if rep.Outcome == "" {
switch { switch {
case rep.Mode == "inventory" && !blk.Enabled: case rep.Mode == "inventory" && !blk.Enabled:
@@ -647,21 +764,27 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
} }
// Health: compare with the start of the pass; give restarted services time (only after an install). // Health: compare with the start of the pass; give restarted services time (only after an install).
cur := wr.HealthAfter cur := wr.HealthAfter
wantEngine := "" wantEngine, wantPVE := "", ""
for _, u := range wr.Upgraded { for _, u := range wr.Upgraded {
if u.Name == "docker-ce" { if u.Name == "docker-ce" {
wantEngine = EngineOf(u.Version) wantEngine = EngineOf(u.Version)
} }
if u.Name == "pve-manager" {
wantPVE = u.Version
}
} }
verdict := func(h *Health) (bool, string) { verdict := func(h *Health) (bool, string) {
if layer == LayerDocker { if layer == LayerDocker {
return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine) return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine)
} }
if layer == LayerHost { if layer == LayerHost || layer == LayerPVE {
t := hub.TunnelUnknown t := hub.TunnelUnknown
if l.Tunnel != nil { if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx) t, _ = l.Tunnel.Status(ctx)
} }
if layer == LayerPVE {
return PVEHealthVerdict(wr.HealthBefore, h, t, wantPVE, wr.PVEManager)
}
return HostHealthVerdict(wr.HealthBefore, h, t) return HostHealthVerdict(wr.HealthBefore, h, t)
} }
return HealthVerdict(wr.HealthBefore, h) return HealthVerdict(wr.HealthBefore, h)
@@ -701,12 +824,24 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
// the docker report carries the engine set only (the guest report already carries the Debian packages) // the docker report carries the engine set only (the guest report already carries the Debian packages)
rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending) rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending)
rep.NotCovered = nil rep.NotCovered = nil
} else if layer == LayerPVE {
// the pve report carries the Proxmox userspace set only — the hub's candidate is built from it
rep.Installed, rep.Pending = onlyPVE(wr.Installed), onlyPVEPending(wr.Pending)
rep.NotCovered = nil
} else { } else {
rep.NotCovered = notCovered(wr.Pending, blk.Ring, planned) rep.NotCovered = notCovered(wr.Pending, blk.Ring, planned)
} }
return l.finish(ctx, lg, rep) return l.finish(ctx, lg, rep)
} }
// rawOrNil: a wrapper field that is absent or JSON null stays out of the hub report (omitempty).
func rawOrNil(m json.RawMessage) json.RawMessage {
if len(m) == 0 || string(m) == "null" {
return nil
}
return m
}
func onlyDocker(in []Package) []Package { func onlyDocker(in []Package) []Package {
var out []Package var out []Package
for _, p := range in { for _, p := range in {
@@ -717,6 +852,33 @@ func onlyDocker(in []Package) []Package {
return out return out
} }
// onlyPVE keeps the installed Proxmox-origin packages the pve lane may touch (never a kernel / boot / firmware name).
func onlyPVE(in []Package) []Package {
var out []Package
for _, p := range in {
if (p.Origin == InstalledPVEOrigin || p.Origin == PVEOrigin) && !hostSlowRE.MatchString(p.Name) && !DockerNames[p.Name] {
out = append(out, p)
}
}
return out
}
func onlyPVEPending(in []Pending) []Pending {
var out []Pending
for _, p := range in {
if p.From == "" || hostSlowRE.MatchString(p.Name) || DockerNames[p.Name] {
continue
}
for _, o := range p.Origin {
if o == PVEOrigin {
out = append(out, p)
break
}
}
}
return out
}
func onlyDockerPending(in []Pending) []Pending { func onlyDockerPending(in []Pending) []Pending {
var out []Pending var out []Pending
for _, p := range in { for _, p := range in {
+71 -7
View File
@@ -13,6 +13,7 @@ import (
"time" "time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub" "gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
) )
// fakeWrapper plays /usr/local/sbin/felhom-os-apply: it reads the plan the leg wrote and answers per layer and mode. // fakeWrapper plays /usr/local/sbin/felhom-os-apply: it reads the plan the leg wrote and answers per layer and mode.
@@ -23,6 +24,8 @@ type fakeWrapper struct {
healthSeq map[string][]*Health // per layer: answers to successive "health" calls healthSeq map[string][]*Health // per layer: answers to successive "health" calls
plans []map[string]any plans []map[string]any
keep bool // R-868: like the real wrapper, keep an apply report beside the plan keep bool // R-868: like the real wrapper, keep an apply report beside the plan
pveGateHeld bool
kernelRep map[string][]WrapperReport // R-836: per kernel-layer mode, successive answers (the last one repeats)
} }
func yes() *bool { b := true; return &b } func yes() *bool { b := true; return &b }
@@ -50,9 +53,22 @@ func (f *fakeWrapper) Run(_ context.Context, name string, args ...string) ([]byt
f.plans = append(f.plans, plan) f.plans = append(f.plans, plan)
layer := plan["layer"].(string) layer := plan["layer"].(string)
ok := guestOK() ok := guestOK()
if layer == LayerHost { if layer == LayerHost || layer == LayerPVE || layer == LayerKernel {
ok = hostOK() ok = hostOK()
} }
if layer == LayerKernel {
if seq := f.kernelRep[plan["mode"].(string)]; len(seq) > 0 {
rep := seq[0]
if len(seq) > 1 {
f.kernelRep[plan["mode"].(string)] = seq[1:]
}
out, _ := json.Marshal(rep)
return []byte("OSAPPLY-REPORT " + string(out) + "\n"), []byte("os-apply: DONE rc=0\n"), nil
}
}
if layer == LayerPVE && plan["mode"] == "apply" {
f.pveGateHeld = pvegate.Stepping() // R-812: the /etc/pve write gate must be held while the pve step runs
}
var rep WrapperReport var rep WrapperReport
switch plan["mode"] { switch plan["mode"] {
case "inventory": case "inventory":
@@ -98,9 +114,13 @@ func (f *fakeWrapper) RunStdin(ctx context.Context, _ io.Reader, name string, ar
return f.Run(ctx, name, args...) return f.Run(ctx, name, args...)
} }
type fakeHub struct{ reports []Report } type fakeHub struct {
reports []Report
bodies [][]byte // the exact bytes posted (R-528: the oom_check object must arrive unchanged)
}
func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error { func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error {
h.bodies = append(h.bodies, append([]byte(nil), body...))
var r Report var r Report
json.Unmarshal(body, &r) json.Unmarshal(body, &r)
h.reports = append(h.reports, r) h.reports = append(h.reports, r)
@@ -156,8 +176,8 @@ func TestRing0_OneCallPerLayer(t *testing.T) {
if g.Outcome != "applied" || !g.Healthy || ho.Outcome != "applied" || !ho.Healthy { if g.Outcome != "applied" || !g.Healthy || ho.Outcome != "applied" || !ho.Healthy {
t.Fatalf("guest %+v\nhost %+v", g, ho) t.Fatalf("guest %+v\nhost %+v", g, ho)
} }
if calls(w) != "guest:apply,host:apply,guest:live-restore-on,docker:apply" { if calls(w) != "guest:apply,host:apply,guest:live-restore-on,docker:apply,pve:apply" {
t.Fatalf("calls = %s, want one apply per layer, guest first, then live-restore and the ring-0 docker step", calls(w)) t.Fatalf("calls = %s, want one apply per layer, guest first, then live-restore, the ring-0 docker step and the pve step", calls(w))
} }
for _, p := range w.plans[:2] { for _, p := range w.plans[:2] {
if p["select"] != "pending-fast" || p["snapshot"] != "" || len(p["packages"].([]any)) != 0 { if p["select"] != "pending-fast" || p["snapshot"] != "" || len(p["packages"].([]any)) != 0 {
@@ -167,7 +187,7 @@ func TestRing0_OneCallPerLayer(t *testing.T) {
if len(g.NotCovered) != 1 || g.NotCovered[0] != "docker-ce" { if len(g.NotCovered) != 1 || g.NotCovered[0] != "docker-ce" {
t.Fatalf("not covered = %v", g.NotCovered) t.Fatalf("not covered = %v", g.NotCovered)
} }
if len(h.reports) != 3 || h.reports[0].Layer != LayerGuest || h.reports[1].Layer != LayerHost || h.reports[2].Layer != LayerDocker { if len(h.reports) != 4 || h.reports[0].Layer != LayerGuest || h.reports[1].Layer != LayerHost || h.reports[2].Layer != LayerDocker || h.reports[3].Layer != LayerPVE {
t.Fatalf("hub got %+v", h.reports) t.Fatalf("hub got %+v", h.reports)
} }
} }
@@ -385,7 +405,7 @@ func TestHostReport_CarriesRebootScanned(t *testing.T) {
}} }}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true}) l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
run2(l, "night") run2(l, "night")
if len(h.reports) != 3 || !h.reports[1].RebootScanned || !h.reports[1].RebootNeeded || h.reports[0].RebootScanned { if len(h.reports) != 4 || !h.reports[1].RebootScanned || !h.reports[1].RebootNeeded || h.reports[0].RebootScanned {
t.Fatalf("hub got %+v", h.reports) t.Fatalf("hub got %+v", h.reports)
} }
} }
@@ -432,7 +452,12 @@ func TestDocker_Ring0PlanAndReport(t *testing.T) {
DockerEngine: "29.8.2", Authority: "ring0"}}} DockerEngine: "29.8.2", Authority: "ring0"}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true}) l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night") p := l.Run(context.Background(), 9201, "night")
dp := w.plans[len(w.plans)-1] var dp map[string]any
for _, x := range w.plans {
if x["layer"] == "docker" {
dp = x
}
}
if dp["layer"] != "docker" || dp["lane"] != "slow" || dp["select"] != "pending-docker" { if dp["layer"] != "docker" || dp["lane"] != "slow" || dp["select"] != "pending-docker" {
t.Fatalf("docker plan = %v", dp) t.Fatalf("docker plan = %v", dp)
} }
@@ -504,3 +529,42 @@ func TestHealthVerdict_ControllerBlindToDockerFails(t *testing.T) {
t.Fatal("an older wrapper (no field) must not fail") t.Fatal("an older wrapper (no field) must not fail")
} }
} }
// R-528 (`09` decision 157): the wrapper's oom_check object reaches the hub's docker report byte-for-byte; the guest
// and host reports carry none. COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in runLayer → "no
// oom_check in the docker report".
const oomCheckWire = `{"detail":"the engine reported the memory kill: OOMKilled=true and the oom event","exit_code":137,"image":"gitea.dooplex.hu/admin/felhom-controller:0.300.0","oom_event":true,"oom_killed":true,"result":"pass"}`
func TestDocker_OOMCheckReachesTheHubUnchanged(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}},
DockerEngine: "29.8.2", Authority: "ring0", OOMCheck: json.RawMessage(oomCheckWire)}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Run(context.Background(), 9201, "night")
found := false
for _, b := range h.bodies {
var m map[string]json.RawMessage
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
var layer string
json.Unmarshal(m["layer"], &layer)
oc, has := m["oom_check"]
if layer != LayerDocker {
if has {
t.Fatalf("the %s report carries an oom_check: %s", layer, oc)
}
continue
}
found = true
if !has {
t.Fatalf("no oom_check in the docker report: %s", b)
}
if string(oc) != oomCheckWire {
t.Fatalf("oom_check changed on the way:\n got %s\nwant %s", oc, oomCheckWire)
}
}
if !found {
t.Fatalf("no docker report posted: %s", calls(w))
}
}
+152
View File
@@ -0,0 +1,152 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// ---- the Proxmox package step (R-812 option A, `09` §3 decision 163, `11` §5.10) ----
// Ring 0: after a healthy host step the leg runs the pve layer — slow lane, select pending-pve — while holding the
// /etc/pve write gate; the report carries only Proxmox userspace packages and pve-manager's version.
//
// COMPANION RED-PROOF (observed): call runLayer instead of runPVE in Run → "the /etc/pve write gate was not held".
func TestPVE_Ring0PlanGateAndReport(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerPVE: {
Upgraded: []Package{{Name: "pve-manager", Version: "9.2.21"}},
Installed: []Package{{Name: "pve-manager", Version: "9.2.21", Origin: "Proxmox"},
{Name: "proxmox-kernel-helper", Version: "9.0.4", Origin: "Proxmox"}, {Name: "libc6", Version: "u4", Origin: "Debian"}},
Pending: []Pending{{Name: "qemu-server", From: "9.0.1", To: "9.0.9", Origin: []string{PVEOrigin}},
{Name: "proxmox-kernel-7.0", From: "7.0.2", To: "7.0.14", Origin: []string{PVEOrigin}},
{Name: "libc6", From: "u3", To: "u4", Origin: []string{"Debian"}}},
PVEManager: "9.2.21", Authority: "ring0"}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
var pp map[string]any
for _, x := range w.plans {
if x["layer"] == LayerPVE {
pp = x
}
}
if pp == nil || pp["lane"] != "slow" || pp["select"] != "pending-pve" {
t.Fatalf("pve plan = %v (calls %s)", pp, calls(w))
}
if !w.pveGateHeld {
t.Fatal("the /etc/pve write gate was not held while the pve step ran")
}
if pvegate.Stepping() {
t.Fatal("the gate must be released after the step")
}
r := p.PVE
if r.Outcome != "applied" || !r.Healthy || r.PVEManager != "9.2.21" {
t.Fatalf("pve report = %+v", r)
}
if len(r.Installed) != 1 || r.Installed[0].Name != "pve-manager" || len(r.Pending) != 1 || r.Pending[0].Name != "qemu-server" {
t.Fatalf("the pve report must carry Proxmox userspace only: installed=%v pending=%v", r.Installed, r.Pending)
}
if h.reports[len(h.reports)-1].Layer != LayerPVE {
t.Fatalf("the hub must get the pve report: %+v", h.reports)
}
}
// Ring 1 never takes a Proxmox step in the night leg.
func TestPVE_Ring1NightLegNeverSteps(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
if p := l.Run(context.Background(), 9201, "night"); p.PVE.Layer != "" || strings.Contains(calls(w), "pve") {
t.Fatalf("ring 1 took a pve step: %s", calls(w))
}
}
// No healthy host step (a BYO box, or an unhealthy host step) → no pve step.
func TestPVE_SkippedWithoutAHealthyHostStep(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Appliance = false
if p := l.Run(context.Background(), 9201, "night"); p.PVE.Layer != "" || strings.Contains(calls(w), "pve") {
t.Fatalf("a BYO box took a pve step: %s", calls(w))
}
w2 := &fakeWrapper{t: t}
l2, _ := newLeg(t, w2, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l2.Tunnel = fakeTunnel{"stopped"} // the host step reads unhealthy
w2.applyRep = map[string]WrapperReport{LayerHost: {Upgraded: []Package{{Name: "libc6", Version: "u4"}}}}
if p := l2.Run(context.Background(), 9201, "night"); p.PVE.Layer != "" || strings.Contains(calls(w2), "pve") {
t.Fatalf("a pve step ran after an unhealthy host step: %s", calls(w2))
}
}
// A write in flight that never finishes makes the pve step give up (failed), never run without the gate.
func TestPVE_GivesUpWhenAWriteDoesNotFinish(t *testing.T) {
old := pveDrainWait
pveDrainWait = 50_000_000 // 50 ms
defer func() { pveDrainWait = old }()
rel, _, _ := pvegate.Write(context.Background())
defer rel()
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
if p.PVE.Outcome != "failed" || strings.Contains(calls(w), "pve") {
t.Fatalf("the pve step must fail without a wrapper call: %+v calls=%s", p.PVE, calls(w))
}
}
// THE pve health rule. COMPANION RED-PROOF (observed): drop the container-id loop or the pve-manager check in
// PVEHealthVerdict → the matching case below fails.
func TestPVEHealthVerdict(t *testing.T) {
before, after := hostOK(), hostOK()
before.Guest.Containers["app"] = Container{State: "running", Health: "healthy", ID: "a1"}
after.Guest.Containers["app"] = Container{State: "running", Health: "healthy", ID: "a1"}
if ok, why := PVEHealthVerdict(before, after, hub.TunnelRunning, "9.2.21", "9.2.21"); !ok {
t.Fatalf("healthy step read unhealthy: %s", why)
}
if ok, _ := PVEHealthVerdict(before, after, hub.TunnelRunning, "9.2.21", "9.2.2"); ok {
t.Fatal("pveversion still on the old pve-manager must fail")
}
after.Guest.Containers["app"] = Container{State: "running", Health: "healthy", ID: "b2"}
if ok, why := PVEHealthVerdict(before, after, hub.TunnelRunning, "", "9.2.2"); ok || !strings.Contains(why, "id changed") {
t.Fatalf("an app restarted by the Proxmox step must fail, got ok=%v %q", ok, why)
}
}
// The signed executor hands the RAW envelope and the exact list to the wrapper's pve layer.
func TestPVEStepExecutor_PassesTheSignedEnvelope(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerPVE: {
Upgraded: []Package{{Name: "pve-manager", Version: "9.2.21"}}, PVEManager: "9.2.21", Authority: "signed"}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
e := PVEStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
params, _ := json.Marshal(PVEStepParams{ReleaseID: "os-pve-1", Packages: []Package{{Name: "pve-manager", Version: "9.2.21", Origin: PVEOrigin}}})
ctx := signedjobs.WithSignedOp(context.Background(), &reconcile.SignedOp{Blob: []byte(`{"op":"os_pve_step"}`), Sig: []byte("SIG")})
if err := e.Execute(ctx, OpPVEStep, params); err != nil {
t.Fatal(err)
}
pp := w.plans[len(w.plans)-1]
sg, _ := pp["signed"].(map[string]any)
if pp["layer"] != LayerPVE || pp["lane"] != "slow" || pp["release_id"] != "os-pve-1" || sg == nil ||
sg["blob_b64"] != base64.StdEncoding.EncodeToString([]byte(`{"op":"os_pve_step"}`)) || sg["sig"] != "SIG" {
t.Fatalf("pve plan = %v", pp)
}
if !w.pveGateHeld || calls(w) != "pve:apply" || len(h.reports) != 1 || h.reports[0].Trigger != "signed" {
t.Fatalf("gate=%v calls=%s reports=%+v", w.pveGateHeld, calls(w), h.reports)
}
if err := e.Execute(context.Background(), OpPVEStep, params); err == nil {
t.Fatal("no envelope must refuse")
}
if err := e.Execute(context.Background(), OpDockerStep, params); err != signedjobs.ErrNoExecutor {
t.Fatalf("another op must pass through the chain: %v", err)
}
}
// os_pve_step is never benign.
func TestPVEStep_IsDestructiveClass(t *testing.T) {
if reconcile.Classify(reconcile.ClassOSPVEStep, reconcile.Provenance{}) != reconcile.Destructive {
t.Fatal("os_pve_step must be destructive-class (signed, operational key)")
}
}
+85
View File
@@ -0,0 +1,85 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// OpPVEStep is the signed op class of a Proxmox package step (R-812 option A, `11` §5.10): a ring-1 box takes an
// approved Proxmox set only through it. No undo in this release. CC may sign it until the first paying customer.
const OpPVEStep = "os_pve_step"
// PVEStepParams are the signed params. The wrapper compares Packages with the plan byte-for-byte.
type PVEStepParams struct {
ReleaseID string `json:"release_id"`
Packages []Package `json:"packages"`
VMID int `json:"vmid,omitempty"`
}
// PVEStepExecutor runs a verified os_pve_step (signedjobs.Executor) under the host-wide heavy-op gate (Gate) and the
// /etc/pve write gate (inside runPVE).
type PVEStepExecutor struct {
Leg *Leg
Guest func(ctx context.Context) (int, error)
Gate func(ctx context.Context) (release func(), err error)
}
// Execute implements signedjobs.Executor.
func (e PVEStepExecutor) Execute(ctx context.Context, op string, params json.RawMessage) error {
if op != OpPVEStep {
return signedjobs.ErrNoExecutor
}
so, ok := signedjobs.SignedOpFrom(ctx)
if !ok {
return fmt.Errorf("os_pve_step: no signed envelope in the context — the wrapper could not verify it")
}
var p PVEStepParams
if err := json.Unmarshal(params, &p); err != nil || len(p.Packages) == 0 {
return fmt.Errorf("os_pve_step: params must name the Proxmox set: %v", err)
}
vmid := p.VMID
if vmid == 0 {
if e.Guest == nil {
return fmt.Errorf("os_pve_step: no vmid and no guest finder")
}
v, err := e.Guest(ctx)
if err != nil {
return fmt.Errorf("os_pve_step: find the customer guest: %w", err)
}
vmid = v
}
if e.Gate != nil {
release, err := e.Gate(ctx)
if err != nil {
return fmt.Errorf("os_pve_step: heavy-op gate busy (a backup or restore-test runs): %w", err)
}
defer release()
}
rep := e.Leg.RunPVESigned(ctx, vmid, p, so.Blob, string(so.Sig))
switch rep.Outcome {
case "applied", "nothing":
if rep.Healthy {
return nil
}
}
return fmt.Errorf("os_pve_step: %s (%s) %s", rep.Outcome, rep.HealthReason, string(rep.Refused))
}
// RunPVESigned is one signed Proxmox step (ring 1): the pve layer with the signed envelope, which the wrapper verifies
// itself, holding the /etc/pve write gate.
func (l *Leg) RunPVESigned(ctx context.Context, vmid int, p PVEStepParams, blob []byte, sig string) Report {
unlock := l.lockPass(true)
defer unlock()
l.sendUnsentLocked(ctx) // R-868
runID := l.now().UTC().Format("20060102T150405Z")
rid := p.ReleaseID
if rid == "" {
rid = "signed-" + runID
}
return l.runPVE(ctx, runID, vmid, "signed", l.Block(), dockerOpts{releaseID: rid, packages: p.Packages,
signed: map[string]string{"blob_b64": base64.StdEncoding.EncodeToString(blob), "sig": sig}})
}
+32
View File
@@ -0,0 +1,32 @@
package osupdate
import (
"os"
"path/filepath"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// After a boot the agent has not fetched the hub's block yet; the after-boot kernel report must still carry the
// box's real ring (seen 2026-10-08: a ring-0 box's „judging" report said ring 1). Red-proof: return Block().Ring
// from kernelReportRing and the saved-block case fails with 1.
func TestKernelReportRing_BeforeFirstFetch(t *testing.T) {
dir := t.TempDir()
l := &Leg{PlanDir: dir}
if got := l.kernelReportRing(); got != 1 {
t.Fatalf("nothing fetched, nothing saved: ring %d, want 1 (the old default)", got)
}
l.saveBlock(&hub.WireOSUpdate{Ring: 0, Enabled: true})
l2 := &Leg{PlanDir: dir} // a fresh daemon after the reboot: no block fetched yet
if got := l2.kernelReportRing(); got != 0 {
t.Fatalf("ring-0 block saved before the reboot: ring %d, want 0", got)
}
l2.SetBlock(&hub.WireOSUpdate{Ring: 1, Enabled: true})
if got := l2.kernelReportRing(); got != 1 {
t.Fatalf("a fetched block wins over the saved one: ring %d, want 1", got)
}
if _, err := os.Stat(filepath.Join(dir, SavedBlockFile)); err != nil {
t.Fatalf("saved block missing: %v", err)
}
}
+24 -3
View File
@@ -142,23 +142,42 @@ func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string)
default: default:
rep.Outcome = "applied" rep.Outcome = "applied"
} }
if wr.Layer == LayerKernel {
// a kernel STAGE changes nothing the box runs (the new kernel only boots once, at the night's reboot), so its
// kept copy needs no fresh health reading (R-836)
rep.Kernel = rawOrNil(wr.Kernel)
rep.Upgraded, rep.PassSeconds, rep.Authority = wr.Upgraded, wr.PassSeconds, wr.Authority
if rep.Outcome == "applied" {
rep.Outcome = "staged"
}
rep.Healthy, rep.HealthReason = rep.Outcome == "staged" || rep.Outcome == "nothing", prefix
return rep
}
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
wantEngine := "" rep.OOMCheck = rawOrNil(wr.OOMCheck)
rep.PVEManager = wr.PVEManager
wantEngine, wantPVE := "", ""
for _, u := range wr.Upgraded { for _, u := range wr.Upgraded {
if u.Name == "docker-ce" { if u.Name == "docker-ce" {
wantEngine = EngineOf(u.Version) wantEngine = EngineOf(u.Version)
} }
if u.Name == "pve-manager" {
wantPVE = u.Version
}
} }
verdict := func(h *Health) (bool, string) { verdict := func(h *Health) (bool, string) {
switch wr.Layer { switch wr.Layer {
case LayerDocker: case LayerDocker:
return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine) return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine)
case LayerHost: case LayerHost, LayerPVE:
t := hub.TunnelUnknown t := hub.TunnelUnknown
if l.Tunnel != nil { if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx) t, _ = l.Tunnel.Status(ctx)
} }
if wr.Layer == LayerPVE {
return PVEHealthVerdict(wr.HealthBefore, h, t, wantPVE, wr.PVEManager)
}
return HostHealthVerdict(wr.HealthBefore, h, t) return HostHealthVerdict(wr.HealthBefore, h, t)
} }
return HealthVerdict(wr.HealthBefore, h) return HealthVerdict(wr.HealthBefore, h)
@@ -166,7 +185,7 @@ func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string)
ok, why := verdict(wr.HealthAfter) ok, why := verdict(wr.HealthAfter)
if !ok && len(wr.Upgraded) > 0 && wr.VMID > 0 { if !ok && len(wr.Upgraded) > 0 && wr.VMID > 0 {
lane := "fast" lane := "fast"
if wr.Layer == LayerDocker { if wr.Layer == LayerDocker || wr.Layer == LayerPVE {
lane = "slow" lane = "slow"
} }
if hr, err := l.call(ctx, "kept-"+runID, map[string]any{"release_id": "kept", "layer": wr.Layer, "lane": lane, if hr, err := l.call(ctx, "kept-"+runID, map[string]any{"release_id": "kept", "layer": wr.Layer, "lane": lane,
@@ -186,6 +205,8 @@ func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string)
rep.RebootScanned = wr.RebootScanned rep.RebootScanned = wr.RebootScanned
if wr.Layer == LayerDocker { if wr.Layer == LayerDocker {
rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending) rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending)
} else if wr.Layer == LayerPVE {
rep.Installed, rep.Pending = onlyPVE(wr.Installed), onlyPVEPending(wr.Pending)
} else { } else {
planned := map[string]bool{} planned := map[string]bool{}
for _, u := range wr.Upgraded { for _, u := range wr.Upgraded {
+21
View File
@@ -150,3 +150,24 @@ func must(t *testing.T, err error) {
t.Fatal(err) t.Fatal(err)
} }
} }
// R-528: a kept docker report (the agent was killed) still carries the oom_check object to the hub unchanged.
// COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in reportFromKept → "oom_check lost".
func TestR868_KeptCopyCarriesTheOOMCheck(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
ring := 0
kept := WrapperReport{Mode: "apply", Layer: LayerDocker, RunID: "20261007T020000Z", Trigger: "night", Ring: &ring, VMID: 9201,
ReleaseID: "ring0-20261007T020000Z", HealthBefore: guestOK(), HealthAfter: guestOK(), DockerEngine: "29.8.2",
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}}, OOMCheck: json.RawMessage(oomCheckWire)}
b, _ := json.Marshal(kept)
must(t, os.WriteFile(reportFile(l.PlanDir, kept.RunID, LayerDocker, "apply"), b, 0o600))
if n := l.SendUnsent(context.Background()); n != 1 || len(h.bodies) != 1 {
t.Fatalf("sent %d, bodies %d", n, len(h.bodies))
}
var m map[string]json.RawMessage
must(t, json.Unmarshal(h.bodies[0], &m))
if string(m["oom_check"]) != oomCheckWire {
t.Fatalf("oom_check lost or changed: %q", m["oom_check"])
}
}
+10
View File
@@ -5,6 +5,7 @@ import (
"context" "context"
"encoding/json" "encoding/json"
"fmt" "fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
"io" "io"
"net/http" "net/http"
"net/url" "net/url"
@@ -105,6 +106,15 @@ func (c *Client) do(ctx context.Context, method, path string, body io.Reader, ou
// doBody is the single HTTP chokepoint: builds the request, sets auth, executes, // doBody is the single HTTP chokepoint: builds the request, sets auth, executes,
// maps non-2xx to APIError, and decodes the data envelope. // maps non-2xx to APIError, and decodes the data envelope.
func (c *Client) doBody(ctx context.Context, method, path string, body io.Reader, contentType string, out any) error { func (c *Client) doBody(ctx context.Context, method, path string, body io.Reader, contentType string, out any) error {
// R-812 option A: every non-GET call may write /etc/pve — it waits while a Proxmox package step restarts pmxcfs
// (pvegate). Pinned by TestPVEGate_ClientWriteWaitsGetDoesNot.
if method != http.MethodGet {
release, _, gerr := pvegate.Write(ctx)
if gerr != nil {
return fmt.Errorf("proxmox: %s %s held back by a Proxmox package step: %w", method, path, gerr)
}
defer release()
}
req, err := http.NewRequestWithContext(ctx, method, c.base+path, body) req, err := http.NewRequestWithContext(ctx, method, c.base+path, body)
if err != nil { if err != nil {
return fmt.Errorf("proxmox: building request: %w", err) return fmt.Errorf("proxmox: building request: %w", err)
+33
View File
@@ -4,8 +4,10 @@ import (
"context" "context"
"encoding/json" "encoding/json"
"fmt" "fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
"io" "io"
"os/exec" "os/exec"
"path/filepath"
"strconv" "strconv"
) )
@@ -59,6 +61,15 @@ func (r *ExecRunner) Run(ctx context.Context, name string, args ...string) ([]by
// RunStdin is Run with the process stdin fed from stdin (nil = no stdin). The sudo-prefix/mode // RunStdin is Run with the process stdin fed from stdin (nil = no stdin). The sudo-prefix/mode
// handling is identical to Run — kept here so both paths share one place. // handling is identical to Run — kept here so both paths share one place.
func (r *ExecRunner) RunStdin(ctx context.Context, stdin io.Reader, name string, args ...string) ([]byte, []byte, error) { func (r *ExecRunner) RunStdin(ctx context.Context, stdin io.Reader, name string, args ...string) ([]byte, []byte, error) {
// R-812 option A: a root CLI that writes /etc/pve waits while a Proxmox package step runs (pvegate).
// Pinned by TestPVEGate_ExecRunnerPctSetWaits / TestWritesEtcPVE.
if WritesEtcPVE(name, args) {
release, _, gerr := pvegate.Write(ctx)
if gerr != nil {
return nil, nil, fmt.Errorf("proxmox: %s held back by a Proxmox package step: %w", name, gerr)
}
defer release()
}
var cmd *exec.Cmd var cmd *exec.Cmd
if r.Mode == RunnerSudo { if r.Mode == RunnerSudo {
sudo := r.SudoPath sudo := r.SudoPath
@@ -77,6 +88,28 @@ func (r *ExecRunner) RunStdin(ctx context.Context, stdin io.Reader, name string,
return stdout.b, stderr.b, err return stdout.b, stderr.b, err
} }
// WritesEtcPVE reports whether a root command writes /etc/pve: `pct` with a config-changing verb, `pvesm`, `pveum`,
// and the PBS storage wrapper's create / reconcile verbs. `pct exec|status|list|config` and every other command do not
// (the os-update wrapper itself must never wait on the gate its own step holds). Pinned by TestWritesEtcPVE.
func WritesEtcPVE(name string, args []string) bool {
base := filepath.Base(name)
switch base {
case "pvesm", "pveum":
return true
case "pct":
if len(args) == 0 {
return false
}
switch args[0] {
case "set", "create", "destroy", "restore", "unlock", "resize", "snapshot", "delsnapshot", "rollback", "move-volume", "start", "stop", "reboot", "shutdown":
return true
}
case "felhom-pbs-apply":
return len(args) > 0 && (args[0] == "create" || args[0] == "reconcile")
}
return false
}
// Privileged is the root-CLI backend. // Privileged is the root-CLI backend.
type Privileged struct { type Privileged struct {
runner Runner runner Runner
+92
View File
@@ -0,0 +1,92 @@
package proxmox
import (
"context"
"net/http"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/pvegate"
)
// R-812 option A: while a Proxmox package step holds the gate, a non-GET API call waits and a GET does not.
//
// COMPANION RED-PROOF (observed): delete the pvegate.Write block in doBody → this fails with "a PUT reached the API
// while the Proxmox step held the gate". Restored. (audits/day-2026-10-07/B/red-pvegate-chokepoints.txt)
func TestPVEGate_ClientWriteWaitsGetDoesNot(t *testing.T) {
d := &mockDoer{fn: func(*http.Request) (*http.Response, error) { return jsonResp(200, `{"data":null}`), nil }}
c := newTestClient(d)
end, err := pvegate.Step(context.Background())
if err != nil {
t.Fatal(err)
}
if err := c.get(context.Background(), "/nodes", nil); err != nil {
t.Fatalf("a GET must not wait on the gate: %v", err)
}
if d.calls != 1 {
t.Fatalf("the GET must reach the API, calls=%d", d.calls)
}
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
err = c.postForm(ctx, http.MethodPut, "/nodes/x/lxc/9201/config", nil, nil)
if d.calls != 1 {
end()
t.Fatal("a PUT reached the API while the Proxmox step held the gate")
}
if err == nil {
end()
t.Fatal("a PUT held back past its deadline must fail")
}
end()
if err := c.postForm(context.Background(), http.MethodPut, "/nodes/x/lxc/9201/config", nil, nil); err != nil || d.calls != 2 {
t.Fatalf("after the step the PUT must go through (err=%v calls=%d)", err, d.calls)
}
}
// A root `pct set` waits on the gate; `pct exec` does not.
//
// COMPANION RED-PROOF (observed): delete the WritesEtcPVE block in RunStdin → this fails with "pct set ran while the
// Proxmox step held the gate". Restored.
func TestPVEGate_ExecRunnerPctSetWaits(t *testing.T) {
r := &ExecRunner{Mode: RunnerDirect}
end, err := pvegate.Step(context.Background())
if err != nil {
t.Fatal(err)
}
defer end()
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
start := time.Now()
_, _, err = r.Run(ctx, "/nonexistent/pct", "set", "9201", "-mp8", "/x")
if err == nil || time.Since(start) < 90*time.Millisecond {
t.Fatalf("pct set ran while the Proxmox step held the gate (err=%v after %s)", err, time.Since(start))
}
start = time.Now()
_, _, _ = r.Run(context.Background(), "/nonexistent/pct", "exec", "9201", "--", "true")
if time.Since(start) > 80*time.Millisecond {
t.Fatal("pct exec must not wait on the gate")
}
}
func TestWritesEtcPVE(t *testing.T) {
for _, c := range []struct {
name string
args []string
want bool
}{
{"pct", []string{"set", "9201", "-mp8", "x"}, true},
{"/usr/sbin/pct", []string{"create", "9201"}, true},
{"pct", []string{"exec", "9201", "--", "true"}, false},
{"pct", []string{"status", "9201"}, false},
{"pvesm", []string{"add", "dir", "x"}, true},
{"pveum", []string{"acl", "modify"}, true},
{"/usr/local/sbin/felhom-pbs-apply", []string{"reconcile"}, true},
{"/usr/local/sbin/felhom-pbs-apply", []string{"read"}, false},
{"/usr/local/sbin/felhom-os-apply", []string{"--plan", "x"}, false},
{"pct", nil, false},
} {
if got := WritesEtcPVE(c.name, c.args); got != c.want {
t.Errorf("WritesEtcPVE(%s %v) = %v, want %v", c.name, c.args, got, c.want)
}
}
}
+104
View File
@@ -0,0 +1,104 @@
// Package pvegate keeps the agent's own writes to /etc/pve out of the way of a Proxmox package step (R-812 option A,
// `09` §3 decision 163, `11` §5.10).
//
// WHY. A `pve` step upgrades pve-cluster / pve-manager / qemu-server / pve-container; their postinst scripts restart
// pmxcfs (the FUSE filesystem behind /etc/pve) and the API daemons. A write that lands while pmxcfs restarts fails or,
// worse, half-lands (design-R-812 §3 A, "can go wrong"). Backups and restore-tests are already kept out by the
// host-wide heavy-op gate; this gate covers everything else the agent writes: every non-GET Proxmox API call
// (proxmox.Client.doBody) and every root CLI that writes /etc/pve (proxmox.ExecRunner — `pct set|create|…`, `pvesm`,
// `pveum`, `felhom-pbs-apply create|reconcile`).
//
// THE RULE. Write waits while a step runs (bounded by its own context). Step marks the step and then waits until every
// write already in flight has finished; it never waits forever (its context bounds it, and the caller gives up and
// does not run the step). One step at a time. Pinned by pvegate_test.go and, at the two chokepoints, by
// proxmox TestPVEGate_*.
package pvegate
import (
"context"
"errors"
"sync"
"time"
)
var (
mu sync.Mutex
inFlight int
stepping bool
stepDone chan struct{}
)
// ErrStepRunning is returned by Step when another step already holds the gate.
var ErrStepRunning = errors.New("pvegate: a Proxmox package step is already running")
// Write marks one /etc/pve write in flight, first waiting while a Proxmox package step runs. The returned release must
// be called when the write has finished. waited reports how long the write was held back.
func Write(ctx context.Context) (release func(), waited time.Duration, err error) {
start := time.Now()
for {
mu.Lock()
if !stepping {
inFlight++
mu.Unlock()
var once sync.Once
return func() {
once.Do(func() {
mu.Lock()
inFlight--
mu.Unlock()
})
}, time.Since(start), nil
}
ch := stepDone
mu.Unlock()
select {
case <-ch:
case <-ctx.Done():
return nil, time.Since(start), ctx.Err()
}
}
}
// Step marks a Proxmox package step and waits until every /etc/pve write already in flight has finished. On error the
// gate is released again and the step must not run. end releases the gate and lets the held-back writes go.
func Step(ctx context.Context) (end func(), err error) {
mu.Lock()
if stepping {
mu.Unlock()
return nil, ErrStepRunning
}
stepping = true
done := make(chan struct{})
stepDone = done
mu.Unlock()
var once sync.Once
end = func() {
once.Do(func() {
mu.Lock()
stepping = false
close(done)
mu.Unlock()
})
}
for {
mu.Lock()
n := inFlight
mu.Unlock()
if n == 0 {
return end, nil
}
select {
case <-ctx.Done():
end()
return nil, ctx.Err()
case <-time.After(50 * time.Millisecond):
}
}
}
// Stepping reports whether a Proxmox package step holds the gate (for logs).
func Stepping() bool {
mu.Lock()
defer mu.Unlock()
return stepping
}
+104
View File
@@ -0,0 +1,104 @@
package pvegate
import (
"context"
"testing"
"time"
)
// A write that starts while a step runs waits until the step ends.
//
// COMPANION RED-PROOF (observed): make Write ignore `stepping` → this fails with "the write went through while the
// Proxmox step held the gate". Restored. (audits/day-2026-10-07/B/red-pvegate.txt)
func TestWrite_WaitsWhileAStepRuns(t *testing.T) {
end, err := Step(context.Background())
if err != nil {
t.Fatal(err)
}
got := make(chan time.Time, 1)
go func() {
rel, _, err := Write(context.Background())
if err == nil {
rel()
}
got <- time.Now()
}()
select {
case <-got:
end()
t.Fatal("the write went through while the Proxmox step held the gate")
case <-time.After(150 * time.Millisecond):
}
ended := time.Now()
end()
select {
case at := <-got:
if at.Before(ended) {
t.Fatal("the write finished before the step ended")
}
case <-time.After(2 * time.Second):
t.Fatal("the write never went through after the step ended")
}
}
// A step waits for a write already in flight before it starts.
func TestStep_WaitsForAWriteInFlight(t *testing.T) {
rel, _, err := Write(context.Background())
if err != nil {
t.Fatal(err)
}
started := make(chan struct{})
go func() {
end, err := Step(context.Background())
if err == nil {
close(started)
end()
}
}()
select {
case <-started:
rel()
t.Fatal("the step started while a write was in flight")
case <-time.After(150 * time.Millisecond):
}
rel()
select {
case <-started:
case <-time.After(2 * time.Second):
t.Fatal("the step never started after the write finished")
}
}
// A step that cannot drain the writes in time gives up and releases the gate (it never waits forever).
func TestStep_GivesUpAndReleases(t *testing.T) {
rel, _, _ := Write(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
if _, err := Step(ctx); err == nil {
t.Fatal("the step must give up while a write is in flight past its deadline")
}
if Stepping() {
t.Fatal("a step that gave up must release the gate")
}
rel()
}
// A write held back past its own deadline returns the context's error.
func TestWrite_HonoursItsContext(t *testing.T) {
end, _ := Step(context.Background())
defer end()
ctx, cancel := context.WithTimeout(context.Background(), 50*time.Millisecond)
defer cancel()
if _, _, err := Write(ctx); err == nil {
t.Fatal("a write held back past its deadline must fail")
}
}
// One step at a time.
func TestStep_OneAtATime(t *testing.T) {
end, _ := Step(context.Background())
defer end()
if _, err := Step(context.Background()); err != ErrStepRunning {
t.Fatalf("a second step must be refused, got %v", err)
}
}
+10 -1
View File
@@ -52,6 +52,15 @@ const (
// (signed, operational key) like agent_update; the root wrapper re-verifies the same signature itself. // (signed, operational key) like agent_update; the root wrapper re-verifies the same signature itself.
ClassOSDockerStep OpClass = "os_docker_step" ClassOSDockerStep OpClass = "os_docker_step"
// A Proxmox package step on the host (R-812 option A, `11` §5.10) — ring 1. Destructive-class (signed, operational
// key) like os_docker_step; the root wrapper re-verifies the same signature itself.
ClassOSPVEStep OpClass = "os_pve_step"
// A kernel step on the host (R-836, `09` §3 decision 172, `11` §5.11) — ring 1: it STAGES a kernel (install + the
// one-shot flag; the night leg reboots it). Destructive-class (signed, operational key) like os_pve_step; the root
// wrapper re-verifies the same signature itself.
ClassOSKernelStep OpClass = "os_kernel_step"
// The config bundle (agent v0.143.0, R-840, `11` §5.4.2): the box's ROOT-OWNED files (sudoers, wrappers, units). // The config bundle (agent v0.143.0, R-840, `11` §5.4.2): the box's ROOT-OWNED files (sudoers, wrappers, units).
// Destructive-class (signed, operational key) like agent_update; the root wrapper re-verifies the signature itself. // Destructive-class (signed, operational key) like agent_update; the root wrapper re-verifies the signature itself.
ClassAgentConfigUpdate OpClass = "agent_config_update" ClassAgentConfigUpdate OpClass = "agent_config_update"
@@ -123,7 +132,7 @@ func Classify(class OpClass, prov Provenance) Disposition {
return Destructive return Destructive
case ClassKeyRotation: case ClassKeyRotation:
return Destructive return Destructive
case ClassAgentUpdate, ClassOSDockerStep, ClassAgentConfigUpdate: case ClassAgentUpdate, ClassOSDockerStep, ClassOSPVEStep, ClassOSKernelStep, ClassAgentConfigUpdate:
// Never benign — no agent-internal provenance can make replacing the agent binary // Never benign — no agent-internal provenance can make replacing the agent binary
// unsigned-safe (a compromised process must not be able to self-bless an update). // unsigned-safe (a compromised process must not be able to self-bless an update).
return Destructive return Destructive
+64
View File
@@ -0,0 +1,64 @@
package storage
import (
"encoding/json"
"strings"
"testing"
)
// R-330 (disk health Phase 2) — attributes 187, 188 and 199 ride the wire.
//
// The failing drive of 2026-08-14 carried 187 Reported_Uncorrect at raw 1001 while SMART said PASSED;
// none of the three reached the controller. The values are carried as RAW counters; an absent
// attribute is OMITTED from the JSON (unknown), never sent as 0 (S-39).
// The 2026-08-14 shape: PASSED, 187 raw 1001, plus 188/199 and the existing three.
const r330SATA = `{"smart_status":{"passed":true},"ata_smart_attributes":{"table":[
{"id":5,"raw":{"value":0}},
{"id":187,"raw":{"value":1001}},
{"id":188,"raw":{"value":4295032833}},
{"id":197,"raw":{"value":8}},
{"id":198,"raw":{"value":8}},
{"id":199,"raw":{"value":3}}]}}`
// COMPANION RED-PROOF (observed): delete the three R-330 cases in parseSMART → this fails with
// "187 Reported_Uncorrect must be carried (raw 1001); got <nil>". Restored.
func TestParseSMART_R330_CarriesTheThreeCounters(t *testing.T) {
s := parseSMART([]byte(r330SATA))
if s.ReportedUncorrect == nil || *s.ReportedUncorrect != 1001 {
t.Fatalf("187 Reported_Uncorrect must be carried (raw 1001); got %v", s.ReportedUncorrect)
}
// 188's raw value is vendor-packed on some drives (this one is 0x100010001): carried as reported, not truncated.
if s.CommandTimeout == nil || *s.CommandTimeout != 4295032833 {
t.Fatalf("188 Command_Timeout must be carried as the full raw value; got %v", s.CommandTimeout)
}
if s.UDMACRCErrors == nil || *s.UDMACRCErrors != 3 {
t.Fatalf("199 UDMA_CRC_Error_Count must be carried (raw 3); got %v", s.UDMACRCErrors)
}
if s.PendingSectors == nil || *s.PendingSectors != 8 {
t.Fatalf("the existing counters must be unchanged; pending=%v", s.PendingSectors)
}
}
// A drive (or an NVMe device) that does not report the attributes leaves them nil, and the JSON
// OMITS the keys — the receiver reads "unknown", never a measured zero.
//
// COMPANION RED-PROOF (observed): drop `,omitempty` from the three tags in hub.SmartSummary → this
// fails with "an unreported attribute must be omitted, not sent: … reported_uncorrect …". Restored.
func TestParseSMART_R330_AbsentIsOmittedNotZero(t *testing.T) {
for name, raw := range map[string]string{
"sata without the three": `{"smart_status":{"passed":true},"ata_smart_attributes":{"table":[{"id":5,"raw":{"value":0}}]}}`,
"nvme": `{"smart_status":{"passed":true},"nvme_smart_health_information_log":{"critical_warning":0,"media_errors":0,"percentage_used":3}}`,
} {
s := parseSMART([]byte(raw))
if s.ReportedUncorrect != nil || s.CommandTimeout != nil || s.UDMACRCErrors != nil {
t.Fatalf("%s: unreported attributes must stay nil; got %v %v %v", name, s.ReportedUncorrect, s.CommandTimeout, s.UDMACRCErrors)
}
b, _ := json.Marshal(s)
for _, k := range []string{"reported_uncorrect", "command_timeout", "udma_crc_errors"} {
if strings.Contains(string(b), `"`+k+`"`) {
t.Fatalf("%s: an unreported attribute must be omitted, not sent: %s", name, b)
}
}
}
}
+11
View File
@@ -44,6 +44,9 @@ const (
ataReallocatedSectorCt = 5 ataReallocatedSectorCt = 5
ataCurrentPending = 197 ataCurrentPending = 197
ataOfflineUncorrect = 198 ataOfflineUncorrect = 198
ataReportedUncorrect = 187 // R-330
ataCommandTimeout = 188 // R-330
ataUDMACRCErrorCount = 199 // R-330
) )
// parseSMART maps smartctl JSON to a hub.SmartSummary, handling SATA + NVMe and degrading // parseSMART maps smartctl JSON to a hub.SmartSummary, handling SATA + NVMe and degrading
@@ -85,6 +88,12 @@ func parseSMART(raw []byte) hub.SmartSummary {
s.PendingSectors = intPtr(int(a.Raw.Value)) s.PendingSectors = intPtr(int(a.Raw.Value))
case ataOfflineUncorrect: case ataOfflineUncorrect:
s.OfflineUncorrectable = intPtr(int(a.Raw.Value)) s.OfflineUncorrectable = intPtr(int(a.Raw.Value))
case ataReportedUncorrect:
s.ReportedUncorrect = int64Ptr(a.Raw.Value)
case ataCommandTimeout:
s.CommandTimeout = int64Ptr(a.Raw.Value)
case ataUDMACRCErrorCount:
s.UDMACRCErrors = int64Ptr(a.Raw.Value)
} }
} }
} }
@@ -142,3 +151,5 @@ func parseThinPoolMetadata(raw []byte) (float64, bool) {
} }
func intPtr(v int) *int { return &v } func intPtr(v int) *int { return &v }
func int64Ptr(v int64) *int64 { return &v }
+367
View File
@@ -0,0 +1,367 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""test_gate_decoys.py — can this repo's gates be fooled by a LABEL? (R-421, R-426)
The same instrument as `felhom.eu/scripts/test_gate_decoys.py`: a decoy is the LABEL without the
FACT, and a gate that passes on the label alone — or refuses the genuine article — is a live hole.
Every gate is asserted in BOTH directions: the decoy must be convicted, the genuine article passed.
Covered here (the `COVERS` literal is AST-read by `felhom.eu/scripts/decoy_coverage_gate.py`, which
never imports this file):
published check-published-versions.py, against a FAKE Gitea (see below).
release-complete check-release-complete.py, in a scratch clone whose `origin` is a scratch bare
repository, against the same fake Gitea.
reuse-refs, instructions, observations
the three SHARED felhom.eu scripts, run against a scratch clone of THIS repo —
so the decoy is planted in the agent's own REUSE.md / CLAUDE.md / REPORT.md and
coverage is per input, not per script.
NEVER THE REAL GITEA. Both network gates read `GITEA_BASE` from the environment (CI already sets it
to the in-cluster URL); here it points at an `http.server` bound to 127.0.0.1 inside this process,
and every proxy variable is removed from the child's environment so urllib cannot route around it.
A test that asked the real registry would pass or fail on whatever was published that day — the
constant-for-measurement shape — and would reach the network from a hook.
NEVER THE REAL TREE. Every planted file lives in a scratch directory: a workspace that holds a
clone of this repo beside symlinks to the sibling clones the shared scripts reach across to.
Run from the repo root: python3 scripts/test_gate_decoys.py
Exit 0 every decoy judged correctly · 1 a decoy passed or a genuine article was refused.
"""
import http.server
import io
import json
import os
import shutil
import socketserver
import subprocess
import sys
import tempfile
import threading
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
PARENT = os.path.dirname(ROOT)
# DECOY_SHARED_DIR exists for ONE purpose: the red-proof. It lets a mutated COPY of the shared scripts
# be judged without editing the felhom.eu clone. Unset, the suite judges the real shared scripts.
SHARED = os.environ.get("DECOY_SHARED_DIR") or os.path.join(PARENT, "felhom.eu", "scripts")
# ── WHAT THIS FILE COVERS ────────────────────────────────────────────────────────────────────────
# Read by felhom.eu/scripts/decoy_coverage_gate.py, which AST-parses this literal. A gate named here
# MUST have a decoy below that has been seen to fail.
COVERS = {
"published": "against a FAKE Gitea: a tag whose package 404s, a tag whose tree lacks the configs, a package one "
"version past the newest tag (published, never tagged), a patch-gap orphan, a missing package that "
"lexical sorting would drop out of the retention window (0.9.x vs 0.10.x); a tags api answering 500 "
"or a non-JSON 200 is INCONCLUSIVE, never a pass - vs a clean registry, a non-semver tag and a version "
"older than the retention window (not asserted, BY DESIGN) (R-426)",
"release-complete": "the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated commit, a tag "
"with no package; a registry 500 is INCONCLUSIVE - vs the genuine release, an `## Unreleased` "
"heading above it, a newer version named only in prose or under `###`, a tag that only "
"origin has (the shallow-CI shape); a LOCAL-only tag passes BY DESIGN (CI's fresh clone "
"and the published gate's converse probe are what see it) (R-426)",
"reuse-refs": "a cited .go and a cited .md path that do not exist, planted in THIS repo's REUSE.md - vs the "
"real file (R-426)",
"instructions": "a component version literal in THIS repo's CLAUDE.md effective text - vs the same sentence "
"inside an HTML comment (R-426)",
"observations": "R-419 in THIS repo's REPORT.md: an Observations note SAYING it carries no marker - vs the "
"two genuine markers (R-426)",
}
fails = []
ran = 0
def report(name, rc, out, expect_rc, must=()):
global ran
ran += 1
missing = [m for m in must if m not in out]
if rc == expect_rc and not missing:
print(" ok %-62s rc=%d (expected %d)" % (name, rc, expect_rc))
else:
hole = expect_rc != 0 and rc == 0
fails.append("%s: rc=%d expected %d%s; missing %s\n%s" % (
name, rc, expect_rc, " - LIVE HOLE" if hole else "", missing, out[-900:]))
# ── the fake Gitea ───────────────────────────────────────────────────────────────────────────────
class Fake(object):
"""What the fake registry serves. Reset per case."""
def reset(self):
self.tags = [] # tag names, as the tags api lists them
self.packages = set() # versions whose generic package downloads
self.raw = set() # versions whose tag tree serves the probe config
self.tags_status = 200
self.tags_body = None # override bytes for the tags api
self.pkg_status = None # override status for EVERY package request
self.hits = []
FAKE = Fake()
FAKE.reset()
PKG_PREFIX = "/api/packages/admin/generic/felhom-agent/"
RAW_PREFIX = "/admin/felhom-agent/raw/tag/v"
class Handler(http.server.BaseHTTPRequestHandler):
def log_message(self, *a):
pass
def _answer(self, status, body=b""):
self.send_response(status)
self.send_header("Content-Length", str(len(body)))
self.end_headers()
if self.command != "HEAD":
self.wfile.write(body)
def do_GET(self):
p = self.path
FAKE.hits.append(p)
if p.startswith("/api/v1/repos/admin/felhom-agent/tags"):
body = FAKE.tags_body if FAKE.tags_body is not None else \
json.dumps([{"name": t} for t in FAKE.tags]).encode()
return self._answer(FAKE.tags_status, body)
if p.startswith(PKG_PREFIX):
if FAKE.pkg_status is not None:
return self._answer(FAKE.pkg_status)
v = p[len(PKG_PREFIX):].split("/", 1)[0]
return self._answer(200, b"ELF") if v in FAKE.packages else self._answer(404)
if p.startswith(RAW_PREFIX):
v = p[len(RAW_PREFIX):].split("/", 1)[0]
ok = v in FAKE.raw and p.endswith("/configs/felhom-agent.service")
return self._answer(200, b"[Unit]\n") if ok else self._answer(404)
return self._answer(404)
do_HEAD = do_GET
class Server(socketserver.ThreadingMixIn, http.server.HTTPServer):
daemon_threads = True
def child_env(base):
env = {k: v for k, v in os.environ.items() if "proxy" not in k.lower()}
env["GITEA_BASE"] = base
env["NO_PROXY"] = env["no_proxy"] = "127.0.0.1,localhost"
return env
def run(argv, cwd, env=None):
# input="" — a child must never inherit (and block on) this process's stdin
p = subprocess.run(argv, cwd=cwd, env=env, capture_output=True, text=True, input="")
return p.returncode, p.stdout + p.stderr
def sh(argv, cwd):
rc, out = run(argv, cwd)
if rc != 0:
raise SystemExit("setup command failed (%s): %s" % (" ".join(argv), out))
return out.strip()
# ── published ────────────────────────────────────────────────────────────────────────────────────
def published_cases(base):
gate = os.path.join(ROOT, "scripts", "check-published-versions.py")
keep = json.load(io.open(os.path.join(ROOT, "scripts", "retention-policy.json"),
encoding="utf-8"))["generic_versions_kept"]
TAGS = ["v0.150.%d" % i for i in range(3)] # inside any retention window >= 3
VERS = [t[1:] for t in TAGS]
def case(name, setup, expect_rc, must=()):
FAKE.reset()
FAKE.tags = list(TAGS)
FAKE.packages = set(VERS)
FAKE.raw = set(VERS)
setup()
rc, out = run([sys.executable, gate], ROOT, child_env(base))
if not FAKE.hits:
fails.append("published/%s: the gate never asked the fake Gitea - the seam is not wired" % name)
report("published: " + name, rc, out, expect_rc, must)
case("GENUINE: every tag downloadable and serving its configs", lambda: None, 0,
("ALL RELEASED VERSIONS INSTALLABLE",))
case("GENUINE: a non-semver tag is not a release", lambda: FAKE.tags.append("v0.150.2-rc1"), 0,
("ALL RELEASED VERSIONS INSTALLABLE",))
case("FACT: a tag whose package 404s", lambda: FAKE.packages.discard("0.150.1"), 1,
("FAIL v0.150.1", "binary NOT downloadable"))
case("FACT: a tag whose tree does not serve the configs", lambda: FAKE.raw.discard("0.150.2"), 1,
("FAIL v0.150.2", "does not serve"))
case("FACT: published one patch past the newest tag, never tagged",
lambda: FAKE.packages.add("0.150.3"), 1, ("PUBLISHED VERSION(S) WITH NO TAG", "v0.150.3"))
case("FACT: published in a patch GAP between two tags",
lambda: (FAKE.tags.remove("v0.150.1"),), 1, ("v0.150.1 is downloadable", "has no git tag"))
def lexical():
# keep+1 tags: 0.9.0 and 0.10.0..0.10.<keep-1>. By SEMVER the oldest is 0.9.0 (dropped); by
# STRING sort "0.10.0" is the smallest and would be the one dropped - so its missing package
# is convicted only if the window is cut by semver.
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.10.0"}
FAKE.raw = set(t[1:] for t in FAKE.tags)
case("FACT: a missing package lexical sorting would drop (0.10.0 vs 0.9.0)", lexical, 1,
("FAIL v0.10.0",))
def retired():
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.9.0"}
FAKE.raw = set(t[1:] for t in FAKE.tags)
case("BY DESIGN: a version older than the retention window is not asserted", retired, 0,
("NOT ASSERTED", "0.9.0"))
def five_hundred():
FAKE.tags_status = 500
case("INCONCLUSIVE: the tags api answers 500", five_hundred, 2, ("INCONCLUSIVE",))
def html():
FAKE.tags_body = b"<html>sign in</html>"
case("INCONCLUSIVE: the tags api answers a 200 that is not JSON", html, 2, ("INCONCLUSIVE",))
# an unreachable Gitea: a port nothing listens on
s = Server(("127.0.0.1", 0), Handler)
dead = "http://127.0.0.1:%d" % s.server_address[1]
s.server_close()
rc, out = run([sys.executable, gate], ROOT, child_env(dead))
report("published: INCONCLUSIVE: Gitea unreachable", rc, out, 2, ("INCONCLUSIVE", "URLs tried"))
# ── release-complete ─────────────────────────────────────────────────────────────────────────────
def release_cases(base, ws):
bare = os.path.join(ws, "origin.git")
work = os.path.join(ws, "rc-work")
sh(["git", "clone", "-q", "--bare", "--no-tags", "file://" + ROOT, bare], ws)
sh(["git", "clone", "-q", "--no-tags", "file://" + bare, work], ws)
sh(["git", "config", "user.email", "decoy@gate.invalid"], work)
sh(["git", "config", "user.name", "decoy"], work)
# the WORKING-TREE gate, so the file under test is the one being edited, not HEAD's
shutil.copy(os.path.join(ROOT, "scripts", "check-release-complete.py"),
os.path.join(work, "scripts", "check-release-complete.py"))
sh(["git", "add", "scripts/check-release-complete.py"], work)
sh(["git", "commit", "-q", "--allow-empty", "-m", "the gate under test"], work)
base_sha = sh(["git", "rev-parse", "HEAD"], work)
ch = os.path.join(work, "CHANGELOG.md")
original = io.open(ch, encoding="utf-8").read()
V = "9.9.9"
def case(name, top, expect_rc, must=(), tag=None, origin_tag=False, packaged=True, pkg_status=None):
FAKE.reset()
if packaged:
FAKE.packages = {V}
FAKE.pkg_status = pkg_status
try:
io.open(ch, "w", encoding="utf-8").write(top + original)
sh(["git", "commit", "-q", "-am", name], work)
if tag == "head":
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", "HEAD"], work)
elif tag == "unrelated":
empty = sh(["git", "mktree"], work) # stdin is "" — the empty tree, written to this repo
orphan = sh(["git", "commit-tree", "-m", "unrelated", empty], work)
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", orphan], work)
if origin_tag:
sh(["git", "push", "-q", "origin", "HEAD:refs/tags/v" + V], work)
rc, out = run([sys.executable, os.path.join(work, "scripts", "check-release-complete.py")],
work, child_env(base))
report("release-complete: " + name, rc, out, expect_rc, must)
finally:
run(["git", "tag", "-d", "v" + V], work)
run(["git", "push", "-q", "origin", ":refs/tags/v" + V], work)
sh(["git", "reset", "-q", "--hard", base_sha], work)
HEAD = "## v%s — 2026-10-06\n\n- decoy release\n\n" % V
case("GENUINE: tagged at HEAD and published", HEAD, 0,
("newest CHANGELOG version: v9.9.9", "is tagged, placed and published"), tag="head")
case("GENUINE: an `## Unreleased` heading above the release", "## Unreleased\n\n- wip\n\n" + HEAD, 0,
("newest CHANGELOG version: v9.9.9",), tag="head")
case("GENUINE: a newer version named only in prose and under ###",
"The `## v10.0.0` heading is not written yet.\n### v10.0.0 notes\n\n" + HEAD, 0,
("newest CHANGELOG version: v9.9.9",), tag="head")
case("GENUINE: the tag only on origin (the shallow-CI shape)", HEAD, 0,
("exists on origin",), origin_tag=True)
case("BY DESIGN: a LOCAL-only tag passes (CI's fresh clone sees only origin)", HEAD, 0,
("an ancestor of HEAD",), tag="head")
case("FACT: the newest heading has no tag anywhere", HEAD, 1, ("DOES NOT EXIST",))
case("FACT: a tag parked on an unrelated commit", HEAD, 1, ("NOT an ancestor",), tag="unrelated")
case("FACT: tagged, never published", HEAD, 1, ("IS NOT PUBLISHED",), tag="head", packaged=False)
case("INCONCLUSIVE: the registry answers 500", HEAD, 2, ("INCONCLUSIVE",), tag="head", pkg_status=500)
case("FACT beats INCONCLUSIVE: no tag AND the registry answers 500", HEAD, 1, ("DOES NOT EXIST",),
pkg_status=500)
# ── the shared felhom.eu scripts, against THIS repo's inputs ─────────────────────────────────────
def shared_cases(ws):
"""A scratch WORKSPACE: a clone of this repo beside symlinks to the siblings, because the shared
scripts reach across (REUSE.md cites hub paths; instructions_gate reads the workspace CLAUDE.md)."""
for g in ("reuse_refs_check.py", "instructions_gate.py", "observations_gate.py"):
if not os.path.isfile(os.path.join(SHARED, g)):
fails.append("shared gate %s is MISSING beside this clone (tried %s) - a failure, never a skip"
% (g, SHARED))
return
space = os.path.join(ws, "workspace")
os.makedirs(space)
for entry in sorted(os.listdir(PARENT)):
if entry in ("felhom.eu", "felhom-controller", "app-catalog-felhom.eu", "homelab-manifests",
"CLAUDE.md", ".claude-memory"):
os.symlink(os.path.join(PARENT, entry), os.path.join(space, entry))
repo = os.path.join(space, "felhom-agent")
sh(["git", "clone", "-q", "--no-tags", "file://" + ROOT, repo], ws)
# the WORKING-TREE inputs the plants go into, so a case judges today's file
for f in ("REUSE.md", "CLAUDE.md", "REPORT.md"):
shutil.copy(os.path.join(ROOT, f), os.path.join(repo, f))
def case(name, gate, relpath, extra, expect_rc, must=()):
p = os.path.join(repo, relpath)
backup = io.open(p, encoding="utf-8").read()
try:
if extra:
io.open(p, "w", encoding="utf-8").write(backup + extra)
rc, out = run([sys.executable, os.path.join(SHARED, gate), repo], repo)
report(name, rc, out, expect_rc, must)
finally:
io.open(p, "w", encoding="utf-8").write(backup)
case("reuse-refs: GENUINE: this repo's REUSE.md", "reuse_refs_check.py", "REUSE.md", "", 0, ("FAILED 0",))
case("reuse-refs: FACT: a cited .go path that does not exist", "reuse_refs_check.py", "REUSE.md",
u"\n- see `internal/localapi/does_not_exist.go`\n", 1, ("does_not_exist.go",))
case("reuse-refs: FACT: a cited .md path that does not exist", "reuse_refs_check.py", "REUSE.md",
u"\n- see `docs/99-does-not-exist.md`\n", 1, ("99-does-not-exist.md",))
case("instructions: GENUINE: this repo's CLAUDE.md", "instructions_gate.py", "CLAUDE.md", "", 0,
("instructions_gate: OK",))
case("instructions: FACT: a version literal in effective text", "instructions_gate.py", "CLAUDE.md",
u"\nThe agent runs v0.148.0 today.\n", 1, ("v0.148.0",))
case("instructions: GENUINE: the same sentence in an HTML comment", "instructions_gate.py", "CLAUDE.md",
u"\n<!--\nThe agent ran v0.148.0 on 2026-10-06.\n-->\n", 0, ("instructions_gate: OK",))
case("observations: FACT: R-419, prose SAYING it has no marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** It carries no `FILED:` marker and no "
u"`NOT-A-FINDING:` marker, deliberately.\n", 1)
case("observations: GENUINE: a FILED marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** Something broke. **FILED: R-419**\n", 0)
case("observations: GENUINE: a NOT-A-FINDING marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** Odd. **NOT-A-FINDING: my own typo, corrected in "
u"the same minute.**\n", 0)
def main():
srv = Server(("127.0.0.1", 0), Handler)
threading.Thread(target=srv.serve_forever, daemon=True).start()
base = "http://127.0.0.1:%d" % srv.server_address[1]
ws = tempfile.mkdtemp(prefix="agent-decoys-")
print("agent gate decoys — fake Gitea at %s, scratch %s" % (base, ws))
try:
published_cases(base)
release_cases(base, ws)
shared_cases(ws)
finally:
srv.shutdown()
srv.server_close()
shutil.rmtree(ws, ignore_errors=True)
if fails:
print()
for f in fails:
print("FAIL: %s" % f)
return 1
print("\nagent gate decoys OK — %d case(s), every label judged on its fact (R-421)" % ran)
return 0
if __name__ == "__main__":
sys.exit(main())