Compare commits

..

19 Commits

Author SHA1 Message Date
admin adaf86ad57 R-330: the agent sends SMART 187/188/199 (raw; omitted when unknown)
gates / gates (push) Successful in 53s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 22:13:04 +02:00
admin 74b5eae5b0 R-894: after a restart the agent remembers the last backup per tier
gates / gates (push) Successful in 35s
An unreadable storage right after an agent restart read the off-site tier
DUE (the in-memory record was empty). The newest success per tier is now
kept on disk and read ONLY when the storage cannot be read: fresh -> not
due, older than the cadence -> due, none -> due (unknown) as before. A
storage that answers stays the ground truth.

Ships with v0.150.0 after the 2026-10-07 read-back; nothing delivered.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 20:33:27 +02:00
admin 7e82f325b8 CHANGELOG: the memory-kill check (unreleased, to ship as v0.150.0 after the night read-back)
gates / gates (push) Successful in 50s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:36:09 +02:00
admin acccb66bd3 R-528 (09 decision 157): after a Docker engine step the wrapper proves the engine reports a memory kill
felhom-os-apply: a docker-layer apply runs oom_check() after health_after and reports
"oom_check": {result pass|fail|error, oom_killed, oom_event, exit_code, image, detail}.
One throwaway container (the controller's image, --pull never, --network none, 64m cap,
label felhom.oomcheck=1) runs dd bs=200M; pass only with OOMKilled=true AND the oom event
(read after a 2 s settle, --until = guest epoch + 1: measured on demo-hp, an --until taken
right after the run missed the event). docker rm -f always runs in a finally; every call
is bounded (<= 90 s). It never changes the step's outcome or health. New wrapper-only mode
"oom-check" (docker layer) runs the check alone; check_guest etc. still apply.

Agent: WrapperReport/Report gain OOMCheck (json:"oom_check"), copied unchanged in runLayer
and in the kept-copy path.

Tests: 9 wrapper tests + 2 Go tests, each red-proofed (audits/readback-2026-10-07/F/red-*.txt).
Also: test_felhom_os_apply.py's `if __name__` sat mid-file, so 11 tests (UnsentReport,
SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran as a script or from
TestWrapperSuite; moved to the end (they pass).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 19:34:55 +02:00
admin de812bc027 The shared rule file (09 §3 decision 152), identical to the other copies; no code change
gates / gates (push) Successful in 44s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 16:06:46 +02:00
admin 3e8ebeb96c Instruction files kept true (09 §3 decision 150): stale gate lists, paths and facts corrected; no code change
gates / gates (push) Successful in 46s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 13:43:57 +02:00
admin cefdc731a4 REPORT: the operator's ten answers (2026-10-06)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 12:28:58 +02:00
admin e56dcb8a4c CHANGELOG/v0.149.0 released (shas)
gates / gates (push) Successful in 37s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:50:21 +02:00
admin f277e619e2 CHANGELOG: unreleased — R-856 GET /host/crash-guard
gates / gates (push) Successful in 38s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:40 +02:00
admin 386f51edc6 R-856: GET /host/crash-guard — the host crash guard's last-boot record for the controller (09 decision 143)
The controller waits ~15 minutes with app mails after a crash boot of the host; it learns of the
crash boot from this route. Reads /var/lib/felhom-crash-guard/state.json (read-only, no Proxmox call)
and passes present/last_boot_at/last_boot_unclean/tripped through; a missing, unreadable or garbled
file answers 200 present:false. Guest-token authed like every sibling route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:48:01 +02:00
admin 130e3ed882 CHANGELOG: unreleased — R-444 weekly guest disk trim, R-99 runbook pointer
gates / gates (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:25:22 +02:00
admin be398f92e8 R-99: the PBS phantom WARN names the cleanup runbook (09 §3 decision 140)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin ee71abd1d4 R-444: weekly guest disk trim (pct fstrim) outside the night, under the heavy-op gate
Operator ruling 09 §3 decision 139. One exact sudoers rule FELHOM_FSTRIM
(`/usr/sbin/pct ^fstrim [0-9]+$`) + manifest entry guest-fstrim; new
internal/fstrim job: due Wednesday from 10:00 host-local, starts only
10:00-20:59, holds backup.InFlight (busy -> deferred to the next hourly
tick), failed trim retried at most 3x per week, bytes parsed from the
"(N bytes) trimmed" lines, last result per guest persisted in
<state_dir>/guest-disk-trim.json and reported as guest_disk_trim.
Config opt-out: "disk_trim": {"disable": true}.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 11:24:45 +02:00
admin 37e98f452b REPORT: the burn-down night (2026-10-06)
gates / gates (push) Successful in 45s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 02:13:10 +02:00
admin b2b82ae828 bundle test: the ISO first-boot files the installer names under KEPT are not installer-written (go test red since felhom.eu 85de3f9b); CHANGELOG unreleased (R-426 decoys)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:25 +02:00
admin b78a0ff3ac R-426: decoys for the shared reuse-refs, instructions and observations gates
COVERS "reuse-refs", "instructions", "observations": the three shared
felhom.eu scripts run against a scratch clone of THIS repo (in a scratch
workspace symlinking the sibling clones they reach across to), so the
plant is in the agent's own REUSE.md / CLAUDE.md / REPORT.md. Convicted:
a missing cited .go and .md path, a version literal in CLAUDE.md's
effective text, R-419's prose-only Observations note. Passed: the real
files, the version inside an HTML comment, both genuine markers.
DECOY_SHARED_DIR lets a red-proof judge a mutated copy of the shared
scripts without editing the felhom.eu clone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 96047453cb R-426: decoys for the release-complete gate
COVERS "release-complete": the working-tree gate runs in a scratch clone
whose origin is a scratch bare repo, against the fake Gitea. Convicted:
the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated
commit, a tag with no package, and no-tag wins over a registry 500.
Inconclusive: a registry 500. Passed: the genuine release, an
`## Unreleased` heading above it, a newer version named only in prose or
under `###` (the withdrawn sweep decoy, now asserted the right way round),
a tag only origin has (the shallow-CI shape), and a LOCAL-only tag BY
DESIGN (CI's fresh clone and the published gate's converse probe see it).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 4bf5db6875 R-426: decoy suite for the published gate, against a fake Gitea
scripts/test_gate_decoys.py (new; COVERS "published"): an http.server on
127.0.0.1 stands in for Gitea through the gate's existing GITEA_BASE
seam, proxies stripped, so no case reaches the real registry. Facts
convicted: a tag whose package 404s, a tag tree without the configs, a
package one patch past the newest tag (never tagged), a patch-gap orphan,
a missing package that lexical sorting would drop out of the retention
window. Inconclusive, never a pass: tags api 500, a non-JSON 200, Gitea
unreachable. Passed: a clean registry, a non-semver tag, a version older
than the retention window (BY DESIGN).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 01:42:03 +02:00
admin 87977ff40a CHANGELOG/v0.148.0 released (shas)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 00:37:07 +02:00
37 changed files with 2599 additions and 55 deletions
+3 -2
View File
@@ -21,6 +21,7 @@ fast, and wrong.
This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that
felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before
path-scoped rules existed. The single source is now felhom.eu/CLAUDE.md "Code quality rules"; this path-scoped rules existed. Deliberate scoped copies now live in felhom.eu/.claude/rules/hub.md and
file is the scoped copy that loads exactly where health checks are written. (2026-08-06) felhom-controller/.claude/rules/gates.md (hub.md's comment names them); none is the single source. This
file is the copy that loads exactly where agent health checks are written. (2026-08-06; corrected 2026-10-06)
--> -->
+69
View File
@@ -0,0 +1,69 @@
---
unconditional: true
---
# Unprompted work — rules for any session without a task file
> Goal sessions, nightly sessions, "work the register" sessions. **A session that starts from
> `/goal` or a standing brief inherits these rules exactly as it inherits the gates.** They are the
> part of `PROMPT-TEMPLATE.md` that a task file used to carry and a goal does not. Same wording lives
> in `felhom.eu`, `felhom-controller`, `felhom-agent` and `app-catalog-felhom.eu` `.claude/rules/`, and in the workspace
> root's unversioned `.claude/rules/`; change all five or none.
## 1. What you may pick up on your own
- A register row **you or another CC session filed**, with owner CC, at P3 or a bounded P2, that
needs **no operator decision**, touches **no customer data by design**, and introduces **no
mechanism nobody has measured**. Smallest first.
- A defect you find while exercising the product, filed as a row **before** you fix it — **unless it is small**:
a small finding is fixed in the session and never filed (the size rule, `OPEN-ITEMS.md` „How a row is filed").
- Hygiene: register compression, stale citations, rows with no owner, documents that contradict
live source.
**Not yours, ever, without a task file or an operator word:** money; anything that changes risk to
customer data; anything that changes a promise the product makes to a customer; anything that
reverses a documented design decision (`documentation/architecture/` — a design decision is not a
defect, R-370); anything on DooPlex or ep0; baking or vouching a golden; promoting a
catalog version; a new external dependency.
## 2. When you may decide instead of ask (operator grant, 2026-09-14)
You may take a decision yourself when **all** of these hold: the architecture folder and the register
give a clear direction; your choice follows that direction; it is reversible without customer-data
risk; and you can write it in the `09-update-architecture.md` §3 shape — one answerable sentence, the
options, what each costs, why this one. **Then record it** as a dated decision in `CONTEXT.md` and
the owning architecture document, tagged *decided by CC unattended — operator may reverse*, and put
it **first** in the morning note. A decision you cannot write in that shape is one you do not take.
## 3. The discipline a task file used to carry
1. **Baselines first.** Read each repo's `main` hash and version from live source before touching it.
2. **Read the architecture document for the area, and name it** in the report, before any claim.
3. **Red-proof every correctness fix.** A test never seen failing has not been shown to test anything.
4. **Live-validate on a Tier-0 box** through the endpoints the UI invokes. `demo-hp` is `ssh hp`.
Throwaway apps only; the standing apps and `bentopdf` stay.
5. **Evidence off the machine at the end of each phase**, before any revert (R-320).
6. **One release per repo per session**, with a CHANGELOG entry (controller: with its `MinAgent`
line), REPORT overwritten, floor raised to deliver it. **No golden unless a drill or fresh install
needs one** (the waiver, R-468). **No `--no-verify`.**
7. **An enumerated gap becomes a row in the same session — or, if it is small, is fixed in it** (the size rule).
Prose is not a record.
8. **Hungarian text is searched with ASCII fragments**, with a positive and a negative control.
9. **Never leave a half-state.** If time runs out, revert to clean and say what was reverted.
10. **Teardown, three layers, stated** — machine, host, hub — or "provisioned nothing".
## 4. The morning note
One screen, plain language, in this order: **decisions you took** (§2) first; what you exercised;
what broke and whether you fixed it; rows opened and closed with the register size before and after;
what needs the operator, each with what happens if they do nothing. No file paths, no function
names, no row numbers as the subject of a sentence.
## 5. Instruction files
**Instruction files (`CLAUDE.md`, `.claude/rules/*`) are kept true by the session that finds them wrong**
(operator ruling 2026-10-06, `09` §3 decision 150). A session MAY, without asking: correct a stale fact (a command, a
count, a version, a path, a description of what a gate does), add a fact it proved, and remove a reference to something
that no longer exists. Each edit is named in the report (file, line, before, after, why). A session MAY NOT, without the
operator's word: loosen a safety rule, a fence, a „never", a protected machine, a secret rule, or a review step; or
remove a rule. When in doubt, it is a rule change, and it goes to the operator. If Claude Code's own permission check
asks before such an edit, wait for the operator's click; if it refuses, record that and file the exact line.
+52
View File
@@ -1,5 +1,57 @@
## Unreleased (2026-10-06 night, later) — after a restart the agent remembers the last backup per tier (R-894); three more SMART counters on the wire (R-330)
Ships with the memory-kill check below as v0.150.0, AFTER the 2026-10-07 night read-back. Nothing delivered tonight.
- **The defect (measured 2026-10-05 on demo-hp):** the agent restarted at 04:57; at 06:25 the off-site storage answered *Can't connect*; the per-tier backup record is in memory only, so the due-check fell back to an EMPTY record and the 7-day tier (last copy 4 days old) read DUE; the controller asked and vzdump failed.
- New `internal/backup/backup_state.go` `BackupSuccessState`: the newest SUCCESSFUL backup per tier and guest, on disk (`<oob state dir>/backup-success-state.json`, atomic tmp+rename, 0600). Only successes are written; a corrupt file reads as nothing known.
- `internal/localapi` `handleBackupDue`: when the tier's storage CANNOT be read, the saved copy stands in for the in-memory record. A fresh copy → not due („… (storage unreadable — age from the last success saved on disk)"); a copy older than the cadence → DUE; no copy → the old answer (DUE, age unknown). A storage that answers stays the ground truth: an archive absent there is due even when the file remembers one.
- Wired in `buildLocalAPIServer` (`LastKnownBackups`); the local API's backup job saves each success.
- Tests: `TestBackupDue_R894_*` (restart = a new server and a new state from the same file; fresh / old / none / storage answers / failed backup not saved), `TestBackupSuccessState_*`, `TestR894_LastKnownBackupsIsWiredIntoTheDaemon` (AST). Four red-proofs observed (`felhom.eu/documentation/audits/night-burndown-2026-10-06/s4/`).
- **R-330 (disk health Phase 2, the wire only):** the SMART summary carries three more SATA raw counters — `reported_uncorrect` (187), `command_timeout` (188, carried as the vendor reports it; some pack several counters), `udma_crc_errors` (199). Pointer + omitempty: an attribute the drive does not report is OMITTED (unknown), never 0. No verdict reads them yet. Tests `TestParseSMART_R330_*` (two red-proofs, `felhom.eu/documentation/audits/night-burndown-2026-10-06/r330/`).
## Unreleased (2026-10-06 night) — the Docker step proves the engine reports a memory kill (`09` §3 decision 157, R-528)
To be released as v0.150.0 with its config bundle AFTER the 2026-10-07 night read-back (the night of 2026-10-06 runs v0.149.0 on purpose).
- `configs/felhom-os-apply`: after a docker-layer APPLY (after `health_after`) the wrapper runs `oom_check()`: a throwaway container from the image the running controller uses (`--pull never`, `--network none`, no volume, label `felhom.oomcheck=1`, 64 MB cap) asks for one 200 MB block; „pass" only when `OOMKilled=true` AND the `oom` event; it waits 2 s and reads the events window to the guest's epoch + 1 (measured: a window closed in the same second missed the event); the container is always removed. Reported as `oom_check`; it never changes the step's outcome or health. A wrapper-only mode `oom-check` runs the check alone (no apt, no engine change), by hand as root.
- `internal/osupdate`: `WrapperReport` and `Report` carry `oom_check` verbatim, on the normal pass and on the kept-copy path (R-868).
- `configs/test_felhom_os_apply.py`: its `unittest.main()` sat in the middle of the file, so 11 tests (UnsentReport, SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran — moved to the end; all pass.
- Tests: the OOMCheck class (pass, OOMKilled=false, no event, unreadable image, removal on an inspect error, not on other layers or in health mode, the events window after the settle wait, mode oom-check alone and its refusals); TestDocker_OOMCheckReachesTheHubUnchanged, TestR868_KeptCopyCarriesTheOOMCheck. 12 red-proofs in `felhom.eu/documentation/audits/readback-2026-10-07/F/`.
## Unreleased (2026-10-06 evening) — the shared rule file (`09` §3 decision 152); no code change
- `.claude/rules/unprompted-work.md` added, byte-identical to the copies in felhom.eu, felhom-controller, app-catalog-felhom.eu and the workspace root (checked with `diff` against the controller's copy and one md5 across all five). Its copies line names five copies.
## Unreleased (2026-10-06 afternoon) — instruction files kept true (`09` §3 decision 150); no code change
- `CLAUDE.md` „Gates — ONE entry point": the runner runs every gate in its `GATES` table (five: three shared, `published`, `release-complete`); `--fast` skips `published` (network). It said two gates and „all of them".
- `CLAUDE.md`: the decoy gate and its audit are named with their `felhom.eu/` prefix (they do not exist in this repo).
- `.claude/rules/health-checks.md` (comment): the health-check rule's copies live in felhom.eu `hub.md` and the controller's `gates.md`; it named felhom.eu `CLAUDE.md` „Code quality rules", which holds no such rule.
## v0.149.0 — a weekly disk trim of each customer guest, the crash-boot fact for the controller, the phantom WARN names its runbook (R-444, R-856, R-99; operator rulings `09` §3 139, 143, 140) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `6bcae9c2eb5d97e8285316583870059835793893299e291891a53a4ce505585f`
config bundle sha256 `e182c82dcf4a67faa3bcb74dbe4ffa7b06e0b27dc8451cb7574d6339ce91ad66` (tag `v0.149.0` = `f277e61`).
**The bundle carries the new sudoers rule for the trim (`FELHOM_FSTRIM`) — deliver it with the binary:** signed
`agent_update`, then signed `agent_config_update`.
- R-856 (`09` §3 decision 143): new local-API route `GET /host/crash-guard` — passes the host crash guard's last-boot record (present, last_boot_at, last_boot_unclean, tripped) from /var/lib/felhom-crash-guard/state.json to the controller, which waits ~15 min with app mails after a crash boot. Read-only, no Proxmox call, guest-token authed; a missing/unreadable/garbled file answers 200 present:false (never an error page). An older agent answers 404, which the controller reads as unknown (normal 90 s grace) — no controller MinAgent raise needed.
- R-444 (`09` §3 decision 139): weekly guest disk trim. New sudoers alias FELHOM_FSTRIM with ONE exact rule `/usr/sbin/pct ^fstrim [0-9]+$` (rides the signed config bundle; decoys pinned by TestSudoersFstrimRuleIsExact) and capability guest-fstrim (non-critical). New internal/fstrim job: each owned RUNNING guest gets `pct fstrim <vmid>` once a week - due Wednesday from 10:00 host-local, starts only 10:00-20:59 (never the 01:00-06:59 night), holds the one-heavy-op gate so it never runs beside a backup or restore-test (busy -> deferred to the next hourly tick; a box that was off catches up at its next daytime hour); a failed trim WARNs and is retried at most 3 times that week; bytes parsed from `pct fstrim`'s "(N bytes) trimmed" lines; positive log `fstrim: guest N trimmed X GiB in Ys`; last result per guest persisted in <state_dir>/guest-disk-trim.json and reported as the new omitempty host-report stanza `guest_disk_trim`. Opt-out: agent.json "disk_trim": {"disable": true}.
- R-99 (`09` §3 decision 140): the agent's WARN for a PBS archive below the 1 MiB plausibility floor now ends with the pointer to the sanctioned cleanup (`documentation/runbooks/pbs-phantom-cleanup.md`); detection only — nothing is deleted automatically. Dir-storage archives keep the old text.
## unreleased ## unreleased
- R-426: scripts/test_gate_decoys.py (new) — the published gate judged against a fake Gitea (127.0.0.1, via GITEA_BASE; never the real registry): 11 cases; COVERS published — `felhom-agent/published` leaves the decoy-coverage EXEMPT list.
- R-426: release-complete gate — 10 decoy cases (scratch clone + scratch bare origin + fake Gitea); COVERS release-complete — `felhom-agent/release-complete` leaves the decoy-coverage EXEMPT list.
- R-426: the shared reuse-refs/instructions/observations gates get agent-side decoys (9 cases on a scratch clone of this repo); COVERS reuse-refs, instructions, observations — three `felhom-agent/*` entries leave the decoy-coverage EXEMPT list.
- **Fixed without a row:** `configs/test_felhom_config_bundle.py` read the two ISO first-boot files that installer 1.32.0 now NAMES under KEPT (R-275) as files the installer writes — `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 (felhom.eu `85de3f9b`); they are listed with why. Test only; agent v0.148.0's code is unaffected.
## v0.148.0 — the host report names the running binary's sha; the format answer carries the new filesystem's UUID (burn-down night: R-349, R-25 agent halves) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `3e68a0870e0e2ce262cb4819294edddeb0a73e8c558a31611a20329a6d9ee283`
config bundle sha256 `a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de` (tag `v0.148.0` = `861d32a`).
Delivery order as for v0.147.0: signed `agent_update`, then signed `agent_config_update`.
- **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from - **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from
`/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub `/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub
comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved. comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved.
+7 -7
View File
@@ -52,11 +52,11 @@ This is in the core because breaching it is how this component stops being audit
## Gates — ONE entry point ## Gates — ONE entry point
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs this **Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs every
repo's gates — `reuse_refs_check` and `instructions_gate`, both the **shared** copies in gate in its `GATES` table (that table is the list); the shared ones — `reuse_refs_check`,
`felhom.eu/scripts/`, never copied into this repo (a copy recreates the drift they detect; an absent `instructions_gate`, `observations_gate` — are the copies in `felhom.eu/scripts/`, never copied into
sibling clone FAILS). `--fast` selects the gates touching no network and no container runtime; today this repo (a copy recreates the drift they detect; an absent sibling clone FAILS). `--fast` selects the
that is all of them. **A missing gate is a FAILURE, never a skip.** gates touching no network and no container runtime, and skips `published` (network), naming it. **A missing gate is a FAILURE, never a skip.**
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is **The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is
**per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS **per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS
@@ -104,8 +104,8 @@ the mechanism are exempt.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without **A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named lacks. `felhom.eu/scripts/decoy_coverage_gate.py` (run by felhom.eu's `repo_gates.py`, for all four repos) refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and decoys withdrawn as illegitimate: `felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over `felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list. `os.listdir`, and a glob over a hand-maintained list.
+6 -9
View File
@@ -1,10 +1,7 @@
# REPORT — agent v0.147.0 (2026-10-05, burn-down round 2) # REPORT — the shared rule file (2026-10-06 evening)
Full session report: `felhom.eu/REPORT-burndown2-2026-10-05.md`. Baseline `d833163` (v0.146.1). Code commit `f1b9b41` Operator ruling 2026-10-06 14:24 (`09` §3 decision 152): the agent repo gets its copy of the shared rule file. Added
(CI job 1365 success), tag `v0.147.0`, binary sha256 `642c4d19…`, bundle sha256 `326527d0…` (verified by download). `.claude/rules/unprompted-work.md`, byte-identical to the other copies (`diff` against felhom-controller's copy: no
output; md5 `c1e6c881…` across all five before the copies line changed, one md5 after). No code changed; no release.
Rows: R-124 (recipe root namespace = ""), R-118 (no root size for an absent drive), R-269 (rotated-out token rejected at `agent_gates.py --fast`: reuse-refs, instructions, release-complete, observations OK. The session report is
once), R-317 (dnsmasq install probed by its unit). Tests + red-proofs: `felhom.eu/documentation/audits/burndown2-2026-10-05/`. `felhom.eu/REPORT.md`.
`go build/vet/test ./...` green; `agent_gates.py --fast` green after the release (release-complete needs the tag).
Delivery: see the session report (vouch, signed jobs per box, the hub System page afterwards).
+3 -1
View File
@@ -121,7 +121,7 @@
| Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only | | Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only |
| Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set | | Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set |
| Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction | | Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction |
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did | | Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`; R-856 `crashGuardStatePath` — GET /host/crash-guard's state file, internal/localapi/crashguard.go) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
| `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) | | `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) |
| `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs | | `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs |
| `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s | | `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s |
@@ -160,8 +160,10 @@
| `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring | | `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring |
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go | | `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). | | `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. | | `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. | | `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
| `backup.BackupSuccessState` | internal/backup/backup_state.go | `RecordBackupSuccess(target, b)` / `LastKnownSuccess(target, vmid)` | Newest SUCCESSFUL backup per tier+guest, persisted (atomic tmp+rename) — the due-check's fallback when the tier's storage cannot be read after a restart (R-894) | **Read ONLY when the storage cannot be read** — a storage that answers is the ground truth (R-84), and an archive absent there must make the tier due even when this file remembers one. A saved copy older than the cadence still reads due. Only successes are written. |
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. | | `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. | | `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). | | `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
+53
View File
@@ -0,0 +1,53 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"testing"
)
// R-444: the weekly trim has the guestnet shape (component + reporter seam + goroutine), so its wiring is asserted
// from the AST like TestMainWiresGuestNetWatchdog — a unit-green trim job that main.go never starts is the inert-seam
// defect. It must also share the ONE heavy-op gate (heavyOps), or it could run beside a backup.
func TestMainWiresGuestDiskTrim(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, 0)
if err != nil {
t.Fatalf("parse main.go: %v", err)
}
var constructedWithGate, reporterWired, started bool
ast.Inspect(f, func(n ast.Node) bool {
switch node := n.(type) {
case *ast.CallExpr:
if fn, ok := node.Fun.(*ast.SelectorExpr); ok {
switch fn.Sel.Name {
case "New":
if pkg, ok := fn.X.(*ast.Ident); ok && pkg.Name == "fstrim" && len(node.Args) >= 3 {
if id, ok := node.Args[2].(*ast.Ident); ok && id.Name == "heavyOps" {
constructedWithGate = true
}
}
case "SetGuestDiskTrimReporter":
reporterWired = true
}
}
case *ast.GoStmt:
if sel, ok := node.Call.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "Run" {
if id, ok := sel.X.(*ast.Ident); ok && id.Name == "diskTrim" {
started = true
}
}
}
return true
})
if !constructedWithGate {
t.Error("main.go never calls fstrim.New(..., heavyOps, ...) — no trim job, or one outside the heavy-op gate")
}
if !reporterWired {
t.Error("main.go never calls collector.SetGuestDiskTrimReporter — the guest_disk_trim stanza never reaches the hub")
}
if !started {
t.Error("main.go never starts the trim job with `go diskTrim.Run(ctx)`")
}
}
+29 -6
View File
@@ -38,6 +38,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow" "gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
"gitea.dooplex.hu/admin/felhom-agent/internal/fasttick" "gitea.dooplex.hu/admin/felhom-agent/internal/fasttick"
"gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd" "gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd"
"gitea.dooplex.hu/admin/felhom-agent/internal/fstrim"
"gitea.dooplex.hu/admin/felhom-agent/internal/guesthook" "gitea.dooplex.hu/admin/felhom-agent/internal/guesthook"
"gitea.dooplex.hu/admin/felhom-agent/internal/guestnet" "gitea.dooplex.hu/admin/felhom-agent/internal/guestnet"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub" "gitea.dooplex.hu/admin/felhom-agent/internal/hub"
@@ -1479,6 +1480,25 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
} }
go runJanitor(ctx, jd) go runJanitor(ctx, jd)
} }
// R-444 (`09` §3 decision 139): the weekly guest disk trim — `pct fstrim <vmid>` of every owned, running guest,
// Wednesday from 10:00 local, daytime only, under the one-heavy-op gate. Not part of the errc fan-out: a trim job
// must never be able to bring the agent down.
if cfg.DiskTrim.Enabled() {
dtMode := proxmox.RunnerMode(cfg.Privileged.Mode)
if dtMode == "" {
dtMode = proxmox.RunnerSudo
}
dtRunner := &proxmox.ExecRunner{Mode: dtMode, SudoPath: cfg.Privileged.SudoPath}
dtGuests := localapi.NewStaleLockController(px, dtRunner, reconcile.DefaultPool, logger)
if dtGuests != nil {
diskTrim := fstrim.New(dtRunner, dtGuests, heavyOps,
filepath.Join(cfg.OOB.WithDefaults().StateDir, "guest-disk-trim.json"), logger)
collector.SetGuestDiskTrimReporter(diskTrim)
go diskTrim.Run(ctx)
}
} else {
logger.Info("fstrim: weekly guest disk trim disabled by config (disk_trim.disable)")
}
if lanLoop != nil { if lanLoop != nil {
lanServers = 1 lanServers = 1
go func() { errc <- lanLoop.Run(ctx) }() go func() { errc <- lanLoop.Run(ctx) }()
@@ -1881,12 +1901,15 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F) InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
Store: store, Store: store,
Storage: observer, // R-894: the newest success per tier on disk — the due-check's fallback when the storage cannot be
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages) // read right after a restart. Same state dir as restore-test-state.json.
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives LastKnownBackups: backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json")),
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate Storage: observer,
Tokens: tokens, DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
BackupCadence: cfg.Backup.BackupCadence(), Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
Tokens: tokens,
BackupCadence: cfg.Backup.BackupCadence(),
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate. // Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
Disks: hostOps, Disks: hostOps,
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID}, DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
+74
View File
@@ -0,0 +1,74 @@
package main
import (
"go/ast"
"testing"
)
// R-894 — the on-disk backup record is WIRED on the daemon path (the built-but-never-wired class).
// main → runDaemon → buildLocalAPIServer, and inside it the localapi.Options literal carries
// LastKnownBackups built by backup.NewBackupSuccessState. An AST walk, not a string match, for the
// reasons in escrow_recover_wiring_test.go.
//
// COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this
// fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored.
func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
_, f := parseMain(t)
if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one")
}
var field, built bool
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
cl, ok := n.(*ast.CompositeLit)
if !ok {
return true
}
sel, ok := cl.Type.(*ast.SelectorExpr)
if !ok {
return true
}
if pkg, _ := sel.X.(*ast.Ident); pkg == nil || pkg.Name+"."+sel.Sel.Name != "localapi.Options" {
return true
}
for _, el := range cl.Elts {
kv, ok := el.(*ast.KeyValueExpr)
if !ok {
continue
}
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "LastKnownBackups" {
field = true
if callsIn(kv.Value)["backup.NewBackupSuccessState"] {
built = true
}
}
}
return true
})
}
if !field {
t.Fatal("localapi.Options in buildLocalAPIServer has no LastKnownBackups field")
}
if !built {
t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState")
}
}
func callsIn(n ast.Node) map[string]bool {
out := map[string]bool{}
ast.Inspect(n, func(n ast.Node) bool {
if ce, ok := n.(*ast.CallExpr); ok {
if fn, ok := ce.Fun.(*ast.SelectorExpr); ok {
if x, ok := fn.X.(*ast.Ident); ok {
out[x.Name+"."+fn.Sel.Name] = true
}
}
}
return true
})
return out
}
+9 -1
View File
@@ -129,6 +129,14 @@ Cmnd_Alias FELHOM_CONTROLLERSWAP = \
Cmnd_Alias FELHOM_STALELOCK = \ Cmnd_Alias FELHOM_STALELOCK = \
/usr/sbin/pct ^unlock [0-9]+$ /usr/sbin/pct ^unlock [0-9]+$
# Weekly guest disk trim (R-444, operator ruling `09` §3 decision 139). A thin pool only ever grows from blocks the
# guest has FREED: `fstrim` inside the unprivileged container is refused (FITRIM: Operation not permitted), so the host
# trims the guest's mounts. Measured on demo-hp 2026-10-06: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %,
# apps kept answering. ONE exact pattern: a vmid and nothing else — no `--ignore-mountpoints`, no second argument
# (pinned: TestSudoersFstrimRuleIsExact). The agent runs it on a weekly daytime timer under the heavy-op gate.
Cmnd_Alias FELHOM_FSTRIM = \
/usr/sbin/pct ^fstrim [0-9]+$
# Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS # Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS
# leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and # leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and
# a guest joins that pool only when its restore COMPLETES, so a failed restore leaves a pool-less guest # a guest joins that pool only when its restore COMPLETES, so a failed restore leaves a pool-less guest
@@ -304,4 +312,4 @@ Cmnd_Alias FELHOM_GUESTNET = \
/usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \ /usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \
/usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$ /usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_FSTRIM, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
+104 -1
View File
@@ -32,6 +32,14 @@
# held packages, kernel taint, the crash guard; guest Debian, Docker engine, containerd, live-restore. # held packages, kernel taint, the crash guard; guest Debian, Docker engine, containerd, live-restore.
# live-restore-on (v0.142.0, layer guest) the ONE-TIME step of `09` decision 87: merge `"live-restore": true` # live-restore-on (v0.142.0, layer guest) the ONE-TIME step of `09` decision 87: merge `"live-restore": true`
# into the guest's /etc/docker/daemon.json and `systemctl reload docker`. NEVER a restart (R-835). # into the guest's /etc/docker/daemon.json and `systemctl reload docker`. NEVER a restart (R-835).
# oom-check (R-528, `09` decision 157; layer docker, lane slow) ONLY the memory-kill check below: no apt, no engine
# change, no authority needed. Wrapper-only — the agent never writes this plan; run it by hand as root.
# R-528 (`09` decision 157): a docker-layer APPLY also runs `oom_check()` after health_after and reports it as
# "oom_check": {"result": "pass"|"fail"|"error", "oom_killed": bool, "oom_event": bool, "exit_code": int|null,
# "image": str|null, "detail": str}
# — one throwaway container (the controller's own image, no network / volume / port, 64 MB cap) is made to exceed its
# memory; "pass" only when the engine says OOMKilled=true AND emits the `oom` event. It never changes the step's
# outcome or health: the hub decides whether the engine set can be approved. Pinned by the OOMCheck tests.
# Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is # Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is
# OSAPPLY-REPORT <one JSON object> # OSAPPLY-REPORT <one JSON object>
# which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install. # which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install.
@@ -68,6 +76,16 @@ JOURNAL_MARK = "@@FELHOM-DPKG-JOURNAL@@"
DPKG_STATE_SCRIPT = "dpkg --audit; echo " + JOURNAL_MARK + "; ls -A /var/lib/dpkg/updates 2>/dev/null; true" DPKG_STATE_SCRIPT = "dpkg --audit; echo " + JOURNAL_MARK + "; ls -A /var/lib/dpkg/updates 2>/dev/null; true"
# The installer's ROOT-OWNED record (felhom-host-install.sh `state_set mode`); the agent cannot write it. # The installer's ROOT-OWNED record (felhom-host-install.sh `state_set mode`); the agent cannot write it.
INSTALL_STATE = "/var/lib/felhom-install/state.json" INSTALL_STATE = "/var/lib/felhom-install/state.json"
# R-528: the memory-kill check (oom_check). One 200 MB block under a 64 MB cap: measured on Docker 29.8.2 to be
# OOM-killed with OOMKilled=true and an `oom` event. Every call is bounded: timeouts (the clock read twice) + the
# settle wait stay within 90 s.
OOMCHECK_PREFIX = "felhom-oomcheck-"
OOMCHECK_SCRIPT = "dd if=/dev/zero of=/dev/null bs=200M count=1"
OOMCHECK_TIMEOUTS = {"image": 10, "clock": 5, "run": 30, "inspect": 10, "events": 10, "rm": 15}
# Measured on demo-hp 9201 (Docker 29.8.2, 2026-10-06, audits/readback-2026-10-07/F/F1, F2): an `--until` taken right
# after the run MISSED the oom event although OOMKilled=true; after a 2 s wait and `--until` = guest epoch + 1 it is
# seen. Pinned by test_events_window_ends_after_the_settle_wait.
OOMCHECK_SETTLE = 2
# Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host # Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host
# reboot is needed for them to take effect, and a bad one can stop the box from booting. # reboot is needed for them to take effect, and a bad one can stop the box from booting.
HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|" HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|"
@@ -405,7 +423,7 @@ class Apply:
def check_plan(self, plan): def check_plan(self, plan):
mode = plan.get("mode", "apply") mode = plan.get("mode", "apply")
if mode not in ("apply", "inventory", "health", "facts", "live-restore-on", "bundle", "agent_update"): if mode not in ("apply", "inventory", "health", "facts", "live-restore-on", "bundle", "agent_update", "oom-check"):
raise Refused("R11", f"unknown mode {mode!r}") raise Refused("R11", f"unknown mode {mode!r}")
if mode == "agent_update": if mode == "agent_update":
if plan.get("layer") != "host": if plan.get("layer") != "host":
@@ -427,6 +445,8 @@ class Apply:
raise Refused("R11", "facts is a host-layer mode (it reads the host and the guest)") raise Refused("R11", "facts is a host-layer mode (it reads the host and the guest)")
if mode == "live-restore-on" and layer != "guest": if mode == "live-restore-on" and layer != "guest":
raise Refused("R11", "live-restore-on is a guest-layer mode") raise Refused("R11", "live-restore-on is a guest-layer mode")
if mode == "oom-check" and layer != "docker":
raise Refused("R11", "oom-check is a docker-layer mode (it checks the guest's Docker engine)")
if plan.get("undo") and layer != "docker": if plan.get("undo") and layer != "docker":
raise Refused("R5", "an undo (downgrade) exists only for the Docker layer, inside a signed job") raise Refused("R5", "an undo (downgrade) exists only for the Docker layer, inside a signed job")
vmid = plan.get("vmid") vmid = plan.get("vmid")
@@ -951,6 +971,10 @@ class Apply:
log = self.r.log log = self.r.log
if self.mode == "live-restore-on": if self.mode == "live-restore-on":
return self.live_restore_on() return self.live_restore_on()
if self.mode == "oom-check":
# R-528: the check alone — no apt, no engine change; check_guest above still applies.
self.report["oom_check"] = self.oom_check()
return 0
self.who, self.allow_downgrade = ("fast", False) self.who, self.allow_downgrade = ("fast", False)
if self.layer == "docker" and self.mode == "apply": if self.layer == "docker" and self.mode == "apply":
self.who, self.allow_downgrade = self.docker_authority(plan) self.who, self.allow_downgrade = self.docker_authority(plan)
@@ -990,8 +1014,87 @@ class Apply:
self.report["docker_engine"] = out_v.strip() if rc_v == 0 and out_v.strip() else "unknown" self.report["docker_engine"] = out_v.strip() if rc_v == 0 and out_v.strip() else "unknown"
self.report["reboot_scanned"] = "reboot_needed" in self.report self.report["reboot_scanned"] = "reboot_needed" in self.report
self.report["health_after"] = self.health() self.report["health_after"] = self.health()
if self.layer == "docker" and self.mode == "apply":
# R-528 (`09` decision 157): does the engine report a memory kill? Reported only — never the outcome.
self.report["oom_check"] = self.oom_check()
return 0 return 0
# ---------- R-528: the memory-kill check ----------
def oom_check(self):
"""Run one throwaway container over its memory cap in the guest and read what the engine says about it.
Never raises: any failure becomes result "error". The container is ALWAYS removed (finally), and a failed
removal is named in the detail."""
t = OOMCHECK_TIMEOUTS
res = {"result": "error", "oom_killed": False, "oom_event": False, "exit_code": None, "image": None, "detail": ""}
log = self.r.log
try:
rc, out, err = self.g(["docker", "inspect", "-f", "{{.Config.Image}}", "felhom-controller"], timeout=t["image"])
except Exception as e:
rc, out, err = -1, "", str(e)
img = out.strip() if rc == 0 else ""
if not img or any(c.isspace() for c in img):
res["detail"] = f"the controller's image could not be read (rc={rc}): {(err or out).strip()[:200]}"
log(f"os-apply: OOM-CHECK error — {res['detail']}")
return res
res["image"] = img
name = f"{OOMCHECK_PREFIX}{os.getpid()}-{os.urandom(4).hex()}"
notes = []
try:
t0 = self.guest_epoch()
if t0 is None:
raise RuntimeError("the guest clock could not be read")
rrc, rout, rerr = self.g(["docker", "run", "--name", name, "--pull", "never", "--network", "none",
"--memory", "64m", "--memory-swap", "64m", "--label", "felhom.oomcheck=1",
"--entrypoint", "sh", img, "-c", OOMCHECK_SCRIPT], timeout=t["run"])
self.r.sleep(OOMCHECK_SETTLE) # the engine publishes the oom event a moment after the run returns (F1/F2)
t1 = self.guest_epoch()
if t1 is None:
t1 = t0 + t["run"] + OOMCHECK_SETTLE + 1
irc, iout, ierr = self.g(["docker", "inspect", "-f", "{{.State.OOMKilled}} {{.State.ExitCode}}", name], timeout=t["inspect"])
if irc != 0:
raise RuntimeError(f"the check container could not be inspected (run rc={rrc}: {(rerr or rout).strip()[:120]}; "
f"inspect rc={irc}: {(ierr or iout).strip()[:120]})")
f = iout.split()
res["oom_killed"] = bool(f) and f[0] == "true"
try:
res["exit_code"] = int(f[1]) if len(f) > 1 else None
except ValueError:
res["exit_code"] = None
erc, eout, eerr = self.g(["docker", "events", "--since", str(t0 - 1), "--until", str(t1 + 1),
"--filter", f"container={name}", "--filter", "event=oom",
"--format", "{{.Action}}"], timeout=t["events"])
if erc != 0:
notes.append(f"the event read failed (rc={erc}): {(eerr or eout).strip()[:120]}")
res["oom_event"] = erc == 0 and any(l.strip() == "oom" for l in eout.splitlines())
if res["oom_killed"] and res["oom_event"]:
res["result"] = "pass"
notes.insert(0, "the engine reported the memory kill: OOMKilled=true and the oom event")
else:
res["result"] = "fail"
miss = [w for w, ok in (("OOMKilled=true", res["oom_killed"]), ("the oom event", res["oom_event"])) if not ok]
notes.insert(0, f"the engine did not report the memory kill: missing {' and '.join(miss)} (exit code {res['exit_code']})")
except Exception as e:
res["result"] = "error"
notes.insert(0, f"the check could not finish: {type(e).__name__}: {str(e)[:200]}")
finally:
try:
mrc, mout, merr = self.g(["docker", "rm", "-f", name], timeout=t["rm"])
if mrc != 0 and "no such container" not in (merr + mout).lower():
notes.append(f"the check container {name} could not be removed (rc={mrc}): {(merr or mout).strip()[:120]}")
except Exception as e:
notes.append(f"the check container {name} could not be removed: {type(e).__name__}: {str(e)[:120]}")
res["detail"] = "; ".join(notes)
log(f"os-apply: OOM-CHECK result={res['result']} oom_killed={res['oom_killed']} oom_event={res['oom_event']} "
f"exit={res['exit_code']} image={img} — {res['detail']}")
return res
def guest_epoch(self):
rc, out, _ = self.g(["date", "+%s"], timeout=OOMCHECK_TIMEOUTS["clock"])
try:
return int(out.strip()) if rc == 0 else None
except ValueError:
return None
def dpkg_state(self): def dpkg_state(self):
"""`dpkg --audit` AND dpkg's update journal, in ONE call (R-876, agent v0.145.0). A crash in the middle of an """`dpkg --audit` AND dpkg's update journal, in ONE call (R-876, agent v0.145.0). A crash in the middle of an
install can leave `/var/lib/dpkg/updates/` non-empty while `--audit` reads clean — measured on demo-hp install can leave `/var/lib/dpkg/updates/` non-empty while `--audit` reads clean — measured on demo-hp
+4 -1
View File
@@ -546,7 +546,10 @@ class Builder(unittest.TestCase):
r"/etc/felhom/[a-z.-]+)", text)) r"/etc/felhom/[a-z.-]+)", text))
agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"} agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"}
trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD} trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust) # Written by the appliance ISO's first boot (felhom.eu scripts/iso/felhom-bootstrap.sh), never by the installer:
# since installer 1.32.0 (R-275) the uninstall only NAMES them under KEPT.
iso_writes = {"/etc/felhom/.bootstrap-done", "/etc/felhom/appliance-pairing-code"}
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust | iso_writes)
self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them") self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them")
# the limits drop-in is named through $AGENT_UNIT in the installer # the limits drop-in is named through $AGENT_UNIT in the installer
self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS) self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS)
+150 -3
View File
@@ -83,7 +83,7 @@ class Fake:
return self.clock return self.clock
def sleep(self, s): def sleep(self, s):
pass self.sleeps = getattr(self, "sleeps", []) + [(len(self.calls), s)] # (calls made before it, seconds)
def verify_sig(self, signers, key_id, ns, blob, sig): def verify_sig(self, signers, key_id, ns, blob, sig):
self.verified = (signers, key_id, ns, blob, sig) self.verified = (signers, key_id, ns, blob, sig)
@@ -196,6 +196,27 @@ class Fake:
return 0, self.engine + "\n", "" return 0, self.engine + "\n", ""
if cmd == "docker" and a[1:3] == ["ps", "-q"]: if cmd == "docker" and a[1:3] == ["ps", "-q"]:
return 0, "".join(i + "\n" for i in self.ids), "" return 0, "".join(i + "\n" for i in self.ids), ""
# R-528: the memory-kill check's engine (oom_image None = unreadable; oom_state / oom_event the engine's answer)
if cmd == "date" and a[1:] == ["+%s"]:
# each read is 3 s later than the last, so the order of the reads is visible in the values
self.oom_epochs = getattr(self, "oom_epochs", []) + [int(self.clock) + 3 * len(getattr(self, "oom_epochs", []))]
return 0, f"{self.oom_epochs[-1]}\n", ""
if cmd == "docker" and a[1:4] == ["inspect", "-f", "{{.Config.Image}}"]:
img = getattr(self, "oom_image", "gitea.dooplex.hu/admin/felhom-controller:0.300.0")
return (0, img + "\n", "") if img is not None else (1, "", "Error: No such object: felhom-controller")
if cmd == "docker" and a[1] == "run":
self.oom_runs = getattr(self, "oom_runs", []) + [a]
return 137, "", ""
if cmd == "docker" and a[1:4] == ["inspect", "-f", "{{.State.OOMKilled}} {{.State.ExitCode}}"]:
if getattr(self, "oom_inspect_raises", False):
raise subprocess.TimeoutExpired(a, 10)
return 0, getattr(self, "oom_state", "true 137") + "\n", ""
if cmd == "docker" and a[1] == "events":
self.oom_events_argv = a
return 0, ("oom\n" if getattr(self, "oom_event", True) else ""), ""
if cmd == "docker" and a[1:3] == ["rm", "-f"]:
self.oom_removed = getattr(self, "oom_removed", []) + a[3:]
return 0, a[3] + "\n", ""
if cmd == "docker" and a[1] == "inspect": if cmd == "docker" and a[1] == "inspect":
mounts = {"aaa111": "/felhom-controller|/var/run/docker.sock;/app/data;", "bbb222": "/app|/data;"} mounts = {"aaa111": "/felhom-controller|/var/run/docker.sock;/app/data;", "bbb222": "/app|/data;"}
return 0, mounts.get(a[-1], "/other|;") + "\n", "" return 0, mounts.get(a[-1], "/other|;") + "\n", ""
@@ -997,8 +1018,6 @@ class RealSignatureCheck(unittest.TestCase):
self.assertNotEqual(r.verify_sig(self.signers, "someone-else", "felhom-op-v1", blob, sig), 0) self.assertNotEqual(r.verify_sig(self.signers, "someone-else", "felhom-op-v1", blob, sig), 0)
self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob, ns="other-ns")), 0) self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob, ns="other-ns")), 0)
if __name__ == "__main__":
unittest.main()
class UnsentReport(unittest.TestCase): class UnsentReport(unittest.TestCase):
@@ -1179,3 +1198,131 @@ class CrashLeftTheJournal(unittest.TestCase):
self.assertEqual(rc, 0, rep) self.assertEqual(rc, 0, rep)
self.assertTrue(any("INTERRUPTED" in l for l in f.logs), f.logs) self.assertTrue(any("INTERRUPTED" in l for l in f.logs), f.logs)
self.assertTrue(any(l.startswith("os-apply: REPAIR ") and l.endswith("forced") for l in f.logs), f.logs) self.assertTrue(any(l.startswith("os-apply: REPAIR ") and l.endswith("forced") for l in f.logs), f.logs)
class OOMCheck(unittest.TestCase):
"""R-528 (`09` decision 157): after a Docker engine step the wrapper proves the engine reports a memory kill
(OOMKilled=true AND the `oom` event). Reported only; the hub decides. Red-proofs: audits/readback-2026-10-07/F/."""
def apply(self, **kw):
f = docker_fake(signed=signed_job())
for k, v in kw.items():
setattr(f, k, v)
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
return f, rep
def test_pass_when_oomkilled_and_the_event(self):
f, rep = self.apply()
oc = rep["oom_check"]
self.assertEqual(oc["result"], "pass", oc)
self.assertEqual((oc["oom_killed"], oc["oom_event"], oc["exit_code"]), (True, True, 137))
self.assertEqual(oc["image"], "gitea.dooplex.hu/admin/felhom-controller:0.300.0")
self.assertEqual(sorted(oc), ["detail", "exit_code", "image", "oom_event", "oom_killed", "result"])
run_argv = f.oom_runs[0]
name = run_argv[run_argv.index("--name") + 1]
self.assertTrue(re.match(r"^felhom-oomcheck-[0-9]+-[0-9a-f]{8}$", name), name)
for flag, val in (("--pull", "never"), ("--network", "none"), ("--memory", "64m"), ("--memory-swap", "64m"),
("--label", "felhom.oomcheck=1"), ("--entrypoint", "sh")):
self.assertEqual(run_argv[run_argv.index(flag) + 1], val, flag)
self.assertNotIn("-v", run_argv)
self.assertNotIn("-p", run_argv)
self.assertEqual(run_argv[-3:], ["gitea.dooplex.hu/admin/felhom-controller:0.300.0", "-c", osapply.OOMCHECK_SCRIPT])
self.assertIn(f"container={name}", f.oom_events_argv)
self.assertIn("event=oom", f.oom_events_argv)
self.assertEqual(f.oom_removed, [name], "the check container must be removed")
self.assertTrue(rep["health_after"], "health is read before the check")
# bounded: every call has a timeout and the clock is read twice — the worst case stays within 90 s
self.assertLessEqual(sum(osapply.OOMCHECK_TIMEOUTS.values()) + osapply.OOMCHECK_TIMEOUTS["clock"]
+ osapply.OOMCHECK_SETTLE, 90)
def test_events_window_ends_after_the_settle_wait(self):
# Measured (F1/F2): an --until taken right after the run missed the oom event. The window must end after a
# wait of >= 2 s that comes AFTER the run, at the guest epoch read after that wait, + 1.
f, rep = self.apply()
run_i = next(i for i, c in enumerate(f.calls) if c[-1][:2] == ["docker", "run"])
date_i = [i for i, c in enumerate(f.calls) if c[-1] == ["date", "+%s"]]
waits = [(i, s) for i, s in getattr(f, "sleeps", []) if i > run_i]
self.assertTrue(waits and waits[0][1] >= 2, f"no settle wait after the run: {getattr(f, 'sleeps', None)}")
self.assertTrue(date_i[-1] >= waits[0][0], "the end epoch must be read after the wait")
ev = f.oom_events_argv
since, until = int(ev[ev.index("--since") + 1]), int(ev[ev.index("--until") + 1])
self.assertEqual(until, f.oom_epochs[-1] + 1, "until = the guest epoch read after the wait, + 1")
self.assertGreater(until, f.oom_epochs[0] + 1)
self.assertLess(since, f.oom_epochs[0] + 1)
def test_oomkilled_false_is_fail(self):
f, rep = self.apply(oom_state="false 0")
oc = rep["oom_check"]
self.assertEqual(oc["result"], "fail", oc)
self.assertIn("OOMKilled=true", oc["detail"])
self.assertTrue(oc["oom_event"])
self.assertEqual(len(f.oom_removed), 1)
def test_no_event_is_fail(self):
f, rep = self.apply(oom_event=False)
oc = rep["oom_check"]
self.assertEqual(oc["result"], "fail", oc)
self.assertIn("the oom event", oc["detail"])
self.assertTrue(oc["oom_killed"])
def test_image_unreadable_is_error(self):
f, rep = self.apply(oom_image=None)
oc = rep["oom_check"]
self.assertEqual(oc["result"], "error", oc)
self.assertIsNone(oc["image"])
self.assertIn("image could not be read", oc["detail"])
self.assertFalse(hasattr(f, "oom_runs"), "no container is started without an image")
def test_container_removed_even_when_inspect_raises(self):
f, rep = self.apply(oom_inspect_raises=True)
oc = rep["oom_check"]
self.assertEqual(oc["result"], "error", oc)
self.assertEqual(len(f.oom_removed), 1, "the container must be removed even when inspect raised")
self.assertTrue(f.oom_removed[0].startswith(osapply.OOMCHECK_PREFIX))
self.assertNotIn("failed", rep, "the check never turns the step into a failure")
def test_never_on_guest_or_host_or_in_health_mode(self):
g = Fake()
rc, rep = run(g)
self.assertEqual(rc, 0, rep)
h = Fake()
h.plan["layer"] = "host"
h.plan["packages"] = [{"name": "bash", "version": "5.2.37-2+b10", "origin": "Debian"}]
h.installed["bash"] = "5.2.37-2+b9"
h.live["bash"] = {"5.2.37-2+b10"}
rc_h, rep_h = run(h)
self.assertEqual(rc_h, 0, rep_h)
d = docker_fake(signed=signed_job())
d.plan["mode"] = "health"
rc_d, rep_d = run(d)
self.assertEqual(rc_d, 0, rep_d)
for f, r in ((g, rep), (h, rep_h), (d, rep_d)):
self.assertNotIn("oom_check", r)
self.assertFalse(hasattr(f, "oom_runs"), r.get("layer"))
self.assertFalse(any(c[-1][:2] == ["docker", "run"] for c in f.calls))
def test_mode_oom_check_runs_only_the_check(self):
f = docker_fake() # no authority: the check changes nothing, so it needs none
f.plan["mode"], f.plan["packages"] = "oom-check", []
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["oom_check"]["result"], "pass", rep)
self.assertFalse(any("apt-get" in c[-1] or "dpkg-query" in c[-1] for c in f.calls), f.calls)
self.assertEqual(len(f.oom_removed), 1)
def test_mode_oom_check_keeps_the_refusals(self):
f = docker_fake()
f.plan["mode"], f.plan["packages"] = "oom-check", []
f.files["/etc/pve/lxc/9201.conf"] = "arch: amd64\n"
rc, rep = run(f)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R10"), rep)
self.assertFalse(hasattr(f, "oom_runs"))
g = Fake()
g.plan["mode"] = "oom-check"
rc, rep = run(g)
self.assertEqual((rc, rep["refused"]["code"]), (2, "R11"), rep)
if __name__ == "__main__":
unittest.main()
@@ -241,3 +241,23 @@ func TestNewestArchiveTime_DistinctPhantomsEachAnnounced(t *testing.T) {
t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String()) t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String())
} }
} }
// R-99 (`09` §3 decision 140): the WARN for a PBS phantom ends with the cleanup runbook, so whoever sees it knows the
// one sanctioned way to remove it; a tiny archive on a dir storage is not a PBS phantom and gets no pointer.
// RED-PROOF: drop the `msg += phantomCleanupPointer` line → "the PBS phantom WARN does not end with the runbook pointer".
func TestRejectedArchiveWarnNamesTheCleanupRunbook(t *testing.T) {
var buf bytes.Buffer
r := runnerWithContent(t, &buf, []proxmox.StorageContent{phantomEntry(), goodPBSEntry()})
if _, _, err := r.NewestArchiveTime(context.Background(), 9201); err != nil {
t.Fatal(err)
}
const want = "INCOMPLETE archive when computing tier freshness — it is not a successful backup — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
if !strings.Contains(buf.String(), want) {
t.Errorf("the PBS phantom WARN does not end with the runbook pointer:\n%s", buf.String())
}
local := phantomEntry()
local.Format, local.VolID = "tar.zst", "local:backup/vzdump-lxc-9201-2026_07_28-05_31_14.tar.zst"
if got := rejectedArchiveMessage(local); strings.Contains(got, "pbs-phantom-cleanup") {
t.Errorf("a dir-storage archive got the PBS runbook pointer: %s", got)
}
}
+131
View File
@@ -0,0 +1,131 @@
package backup
import (
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// BackupSuccessState persists the newest SUCCESSFUL whole-guest backup per tier and guest (R-894).
//
// Why it exists. The due-check (`localapi` handleBackupDue) asks the tier's storage when a backup last
// landed (R-84) and falls back to the in-memory record when the storage cannot be read. The in-memory
// record is empty after an agent restart (Store, R-348), so "storage unreadable" right after a restart
// read as "no record — DUE". Measured 2026-10-05 on demo-hp: the agent restarted at 04:57, the off-site
// storage answered "Can't connect" at 06:25, the 7-day tier — last copy 2026-10-01 — read DUE, the
// controller asked, and vzdump failed. This file is the last known copy the fallback reads instead.
//
// It is read ONLY when the storage cannot be read. A storage that answers is the ground truth and wins,
// in both directions: an archive found there counts, and an archive absent there is absent even when
// this file remembers a success (a pruned or deleted archive must make the tier due — the same reason
// R-84 chose the storage over a persisted record). Pinned by
// TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers.
//
// Only SUCCESSES are written (the RestoreTestState rule): a failure must stay due and be retried, so a
// record of a failure has no reader.
type BackupSuccessState struct {
path string
mu sync.Mutex
last map[string]savedSuccess // key(target, vmid) → the newest success
}
type savedSuccess struct {
target string
vmid int
at time.Time
}
// backupSuccessJSON is one entry on disk.
type backupSuccessJSON struct {
Target string `json:"target"`
VMID int `json:"vmid"`
StartedAt string `json:"started_at"`
}
func backupStateKey(target string, vmid int) string { return target + "/" + strconv.Itoa(vmid) }
// NewBackupSuccessState opens (or creates) the state at path. A missing or unreadable file degrades to
// "nothing known" — the pre-R-894 behaviour, which is DUE — and never wedges the daemon.
func NewBackupSuccessState(path string) *BackupSuccessState {
s := &BackupSuccessState{path: path, last: map[string]savedSuccess{}}
data, err := os.ReadFile(path)
if err != nil {
return s
}
var entries []backupSuccessJSON
if json.Unmarshal(data, &entries) != nil {
return s
}
for _, e := range entries {
t, perr := time.Parse(time.RFC3339, e.StartedAt)
if perr != nil {
continue // one unreadable entry must not lose the others
}
s.last[backupStateKey(e.Target, e.VMID)] = savedSuccess{target: e.Target, vmid: e.VMID, at: t.UTC()}
}
return s
}
// RecordBackupSuccess saves b when it is a success newer than the one on file. target is the tier the
// job ran on (the due-check's key); a failure or an unparseable time is ignored.
func (s *BackupSuccessState) RecordBackupSuccess(target string, b hub.Backup) error {
if s == nil || !b.Success {
return nil
}
t, err := time.Parse(time.RFC3339, b.StartedAt)
if err != nil {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
k := backupStateKey(target, b.VMID)
if old, ok := s.last[k]; ok && !t.After(old.at) {
return nil
}
s.last[k] = savedSuccess{target: target, vmid: b.VMID, at: t.UTC()}
return s.saveLocked()
}
// LastKnownSuccess returns the newest saved success for this tier and guest (ok=false = none on file).
func (s *BackupSuccessState) LastKnownSuccess(target string, vmid int) (time.Time, bool) {
if s == nil {
return time.Time{}, false
}
s.mu.Lock()
defer s.mu.Unlock()
e, ok := s.last[backupStateKey(target, vmid)]
return e.at, ok
}
func (s *BackupSuccessState) saveLocked() error {
entries := make([]backupSuccessJSON, 0, len(s.last))
for _, e := range s.last {
entries = append(entries, backupSuccessJSON{Target: e.target, VMID: e.vmid, StartedAt: e.at.Format(time.RFC3339)})
}
// Deterministic file content (Go's map order is random).
sort.Slice(entries, func(i, j int) bool {
if entries[i].Target != entries[j].Target {
return entries[i].Target < entries[j].Target
}
return entries[i].VMID < entries[j].VMID
})
data, err := json.MarshalIndent(entries, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(s.path), 0o755); err != nil {
return err
}
tmp := s.path + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, s.path)
}
+59
View File
@@ -0,0 +1,59 @@
package backup
import (
"os"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-894: the on-disk newest success per tier survives a restart (a new state from the same file).
func TestBackupSuccessState_SurvivesRestart(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
s := NewBackupSuccessState(path)
at := time.Date(2026, 10, 1, 20, 15, 0, 0, time.UTC)
if err := s.RecordBackupSuccess("felhom-pbs", hub.Backup{VMID: 9201, Success: true, StartedAt: at.Format(time.RFC3339)}); err != nil {
t.Fatal(err)
}
got, ok := NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 9201)
if !ok || !got.Equal(at) {
t.Fatalf("after a restart the saved copy must read back; got %v ok=%v", got, ok)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("local", 9201); ok {
t.Fatal("another tier must not borrow this tier's copy")
}
}
// Only a NEWER success replaces the saved one; failures and unparseable times are ignored.
func TestBackupSuccessState_KeepsNewestSuccessOnly(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
s := NewBackupSuccessState(path)
newer := time.Date(2026, 10, 5, 0, 0, 0, 0, time.UTC)
older := newer.Add(-48 * time.Hour)
for _, b := range []hub.Backup{
{VMID: 1, Success: true, StartedAt: newer.Format(time.RFC3339)},
{VMID: 1, Success: true, StartedAt: older.Format(time.RFC3339)}, // older: ignored
{VMID: 1, Success: false, StartedAt: newer.Add(time.Hour).Format(time.RFC3339)}, // failure: ignored
{VMID: 1, Success: true, StartedAt: "not-a-time"}, // unparseable: ignored
} {
if err := s.RecordBackupSuccess("t", b); err != nil {
t.Fatal(err)
}
}
if got, _ := NewBackupSuccessState(path).LastKnownSuccess("t", 1); !got.Equal(newer) {
t.Fatalf("want the newest success %v, got %v", newer, got)
}
}
// A corrupt file degrades to "nothing known" (the pre-R-894 DUE answer), never a crash.
func TestBackupSuccessState_CorruptFileIsEmpty(t *testing.T) {
path := filepath.Join(t.TempDir(), "s.json")
if err := os.WriteFile(path, []byte("{not json"), 0o600); err != nil {
t.Fatal(err)
}
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("t", 1); ok {
t.Fatal("a corrupt file must read as nothing known")
}
}
+16 -1
View File
@@ -518,10 +518,25 @@ func (r *BackupRunner) warnRejectedArchiveOnce(e proxmox.StorageContent, why str
if seen { if seen {
return return
} }
r.logger.Warn("backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup", r.logger.Warn(rejectedArchiveMessage(e),
"target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why) "target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
} }
// phantomCleanupPointer names the runbook that removes a PBS phantom (R-99, `09` §3 decision 140: a leftover of an
// aborted upload is deleted on the backup server, by a runbook, when one is seen — never automatically).
const phantomCleanupPointer = " — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
// rejectedArchiveMessage is the WARN text for a rejected archive. Only a PBS entry (format pbs-ct / pbs-vm) gets the
// cleanup pointer: the runbook deletes on a PBS datastore, and a tiny archive on a dir storage is not a PBS phantom.
// Pinned by TestRejectedArchiveWarnNamesTheCleanupRunbook.
func rejectedArchiveMessage(e proxmox.StorageContent) string {
msg := "backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup"
if strings.HasPrefix(e.Format, "pbs-") {
msg += phantomCleanupPointer
}
return msg
}
// demo-felhom in a single afternoon of deploys (2026-07-26). // demo-felhom in a single afternoon of deploys (2026-07-26).
// //
// Asking the STORAGE rather than persisting the store is deliberate: // Asking the STORAGE rather than persisting the store is deliberate:
+3
View File
@@ -30,6 +30,9 @@ import (
// 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What // 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What
// is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/ // is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/
// deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84). // deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84).
// The due-check's fallback for an UNREADABLE storage no longer reads this store alone (R-894): the newest
// success per tier is also on disk (BackupSuccessState), so a restart followed by an unreachable storage
// reads the last known copy, not "never".
type Store struct { type Store struct {
mu sync.Mutex mu sync.Mutex
byTarget map[string]hub.Backup // latest backup per target id byTarget map[string]hub.Backup // latest backup per target id
+4
View File
@@ -151,6 +151,10 @@ var manifest = []Capability{
// reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ---- // reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ----
{"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""}, {"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""},
// ---- Weekly guest disk trim (FELHOM_FSTRIM, R-444). NON-critical: a missing grant means the thin pool is not
// reclaimed this week (the trim job WARNs per guest and the report shows the failure), not a serving outage. ----
{"guest-fstrim", "weekly guest disk trim (thin-pool reclaim, R-444)", "/usr/sbin/pct", []string{"fstrim", "9201"}, false, ""},
// ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite // ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite
// backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the // backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the
// conf install, unit enable/restart and the handshake read gate the backup path. apt-install // conf install, unit enable/restart and the handshake read gate the backup path. apt-install
@@ -63,3 +63,33 @@ func TestSudoersRefusesTheR861Injections(t *testing.T) {
} }
} }
} }
// R-444: the weekly trim's grant is ONE exact shape — `pct fstrim <vmid>` — and nothing smuggled after it.
// The manifest entry (guest-fstrim) proves the real call is still allowed (TestManifestCoveredBySudoers); this
// pins the other direction. RED-PROOF: write the rule as the glob `/usr/sbin/pct fstrim [0-9]*` → every decoy
// below with a trailing argument matches (the glob's `*` eats spaces).
func TestSudoersFstrimRuleIsExact(t *testing.T) {
data, err := os.ReadFile(sudoersPath)
if err != nil {
t.Fatal(err)
}
entries := parseSudoersEntries(t, string(data))
if !matchesAny("/usr/sbin/pct fstrim 9201", entries) {
t.Fatal("the sudoers does not allow `pct fstrim 9201` — the weekly trim cannot run")
}
for _, c := range []string{
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints",
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints 1",
"/usr/sbin/pct fstrim 9201; x",
"/usr/sbin/pct fstrim 9201 9202",
"/usr/sbin/pct fstrim 92a1",
"/usr/sbin/pct fstrim ",
"/usr/sbin/pct fstrim -- 9201",
"/usr/sbin/pct destroy 9201",
"/usr/sbin/pct destroy 9201 --purge",
} {
if matchesAny(c, entries) {
t.Errorf("the sudoers allows a command the trim rule must not: %q", c)
}
}
}
+11
View File
@@ -33,6 +33,7 @@ type Config struct {
LANResolver LANResolverConfig `json:"lan_resolver"` LANResolver LANResolverConfig `json:"lan_resolver"`
WGTunnel WGTunnelConfig `json:"wg_tunnel"` WGTunnel WGTunnelConfig `json:"wg_tunnel"`
GuestNet GuestNetConfig `json:"guest_net"` GuestNet GuestNetConfig `json:"guest_net"`
DiskTrim DiskTrimConfig `json:"disk_trim"`
OOB OOBConfig `json:"oob"` OOB OOBConfig `json:"oob"`
SelfUpdate SelfUpdateConfig `json:"selfupdate"` SelfUpdate SelfUpdateConfig `json:"selfupdate"`
LogLevel string `json:"log_level"` // debug|info|warn|error (default info) LogLevel string `json:"log_level"` // debug|info|warn|error (default info)
@@ -139,6 +140,16 @@ func (w WGTunnelConfig) WithDefaults() WGTunnelConfig {
return w return w
} }
// DiskTrimConfig configures the R-444 weekly guest disk trim (internal/fstrim). DEFAULT-ON, like GuestNetConfig and for
// the same reason: it only acts on guests the agent already owns, and the operator ruled every box trims (`09` §3
// decision 139). Opting out is the explicit act: `"disk_trim": {"disable": true}`.
type DiskTrimConfig struct {
Disable bool `json:"disable"`
}
// Enabled reports whether the weekly trim should run.
func (d DiskTrimConfig) Enabled() bool { return !d.Disable }
// GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet). // GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet).
// //
// **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other // **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other
+346
View File
@@ -0,0 +1,346 @@
// Package fstrim is the weekly guest disk trim (R-444, operator ruling `09` §3 decision 139).
//
// Why: a thin pool only ever grows from blocks a guest has already FREED — `fstrim` inside the unprivileged container
// is refused (FITRIM: Operation not permitted), and nothing else on the box gives the blocks back. A full thin pool
// takes every guest on the host read-only, so the pool can reach 100 % from deleted data alone. Measured on demo-hp
// 2026-10-06 09:14Z: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %, 18/18 app probes 200, max 1.1 s
// (audits/ten-answers-2026-10-06/r444-measure.txt).
//
// The rule, each part pinned by a test in fstrim_test.go:
// - Weekly: a guest is DUE from Wednesday 10:00 local until it has been trimmed once since then (a box that was off
// on Wednesday catches up at its next eligible hour).
// - Daytime only: a trim starts only between 10:00 and 20:59 local — never in the night window (01:00–06:59) where
// the backups and the restore-tests run (TestEligibleHourNeverInTheNight).
// - Never beside a backup, a restore-test or another heavy operation: the pass holds the host-wide one-heavy-op gate
// (backup.InFlight) for its whole run; a busy gate DEFERS the pass to the next hourly tick.
// - A failed trim is retried at the next eligible hour, at most MaxAttemptsPerWeek times in one week.
// - The last result per guest (time, bytes, ok/fail) is persisted, so a restart neither loses it nor re-trims.
//
// The command is the ONE exact sudoers shape `pct fstrim <vmid>` (FELHOM_FSTRIM). Only guests from the pool-verified
// source (ListLXC ∩ the felhom pool, audit A1) and only RUNNING ones are trimmed.
package fstrim
import (
"context"
"encoding/json"
"fmt"
"log/slog"
"os"
"path/filepath"
"regexp"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// Schedule. The weekday/hours are fixed on purpose (one sentence the operator can read on the System page).
const (
Weekday = time.Wednesday
StartHour = 10 // first eligible local hour (inclusive)
EndHour = 21 // first NOT-eligible local hour (exclusive): last start is 20:59
MaxAttemptsPerWeek = 3
// TickInterval is how often the job looks; a deferred or failed pass is therefore retried the next hour.
TickInterval = time.Hour
// FirstTickDelay lets the agent settle after a start before the first look.
FirstTickDelay = 5 * time.Minute
// PerGuestTimeout bounds one `pct fstrim` (measured 24.4 s for 84 GiB).
PerGuestTimeout = 30 * time.Minute
)
// ScheduleText is the human description carried on the host report.
const ScheduleText = "weekly, due Wednesday from 10:00 host-local time; starts only 10:00-20:59; never beside a backup or restore-test"
// Runner runs a host command (proxmox.ExecRunner in production, through `sudo -n`).
type Runner interface {
Run(ctx context.Context, name string, args ...string) (stdout, stderr []byte, err error)
}
// GuestSource yields the guests this agent OWNS (the pool-verified source, never a bare ListLXC).
type GuestSource interface {
Guests(ctx context.Context) ([]proxmox.Guest, error)
}
// Gate is the host-wide one-heavy-operation gate (*backup.InFlight).
type Gate interface {
TryAcquire(what string) (release func(), busy string, ok bool)
}
// GateName is what the gate reports as busy while a trim runs.
const GateName = "guest-fstrim"
// Record is one guest's last trim attempt, as persisted.
type Record struct {
LastAttemptAt time.Time `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt time.Time `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
// Attempts counts the attempts since the current week's due time (reset by the first attempt of a new week).
Attempts int `json:"attempts"`
}
// Trimmer is the weekly job.
type Trimmer struct {
runner Runner
guests GuestSource
gate Gate
statePath string
logger *slog.Logger
loc *time.Location
now func() time.Time
mu sync.Mutex
records map[int]Record
}
// New builds the job and loads the persisted state. A missing state file is an empty state; a corrupt one is logged
// and treated as empty (the cost is one extra trim, never a missed one).
func New(runner Runner, guests GuestSource, gate Gate, statePath string, logger *slog.Logger) *Trimmer {
if logger == nil {
logger = slog.Default()
}
t := &Trimmer{runner: runner, guests: guests, gate: gate, statePath: statePath, logger: logger,
loc: time.Local, now: time.Now, records: map[int]Record{}}
t.load()
return t
}
func (t *Trimmer) load() {
data, err := os.ReadFile(t.statePath)
if err != nil {
if !os.IsNotExist(err) {
t.logger.Warn("fstrim: state read failed — starting empty", "path", t.statePath, "err", err)
}
return
}
var raw map[string]Record
if err := json.Unmarshal(data, &raw); err != nil {
t.logger.Warn("fstrim: state file corrupt — starting empty", "path", t.statePath, "err", err)
return
}
for k, r := range raw {
if id, err := strconv.Atoi(k); err == nil && id > 0 {
t.records[id] = r
}
}
}
func (t *Trimmer) saveLocked() error {
raw := make(map[string]Record, len(t.records))
for id, r := range t.records {
raw[strconv.Itoa(id)] = r
}
data, err := json.MarshalIndent(raw, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(t.statePath), 0o755); err != nil {
return err
}
tmp := t.statePath + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, t.statePath)
}
// EligibleHour reports whether a trim may START at local time lt.
func EligibleHour(lt time.Time) bool {
h := lt.Hour()
return h >= StartHour && h < EndHour
}
// weekAnchor is the most recent Wednesday StartHour:00 at or before lt (same location as lt).
func weekAnchor(lt time.Time) time.Time {
daysBack := (int(lt.Weekday()) - int(Weekday) + 7) % 7
d := lt.AddDate(0, 0, -daysBack)
a := time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
if a.After(lt) {
d = d.AddDate(0, 0, -7)
a = time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
}
return a
}
// due reports whether a guest with record r (ok=false: none) is due at local time lt.
func due(r Record, has bool, lt time.Time) bool {
if !has {
return true
}
anchor := weekAnchor(lt)
if r.LastAttemptAt.Before(anchor) {
return true // not tried this week
}
return !r.OK && r.Attempts < MaxAttemptsPerWeek
}
// Run looks every TickInterval until ctx ends. It never returns an error: a failed trim is a reported fact.
func (t *Trimmer) Run(ctx context.Context) {
t.logger.Info("fstrim: weekly guest disk trim starting", "schedule", ScheduleText)
timer := time.NewTimer(FirstTickDelay)
defer timer.Stop()
for {
select {
case <-ctx.Done():
return
case <-timer.C:
t.Pass(ctx)
timer.Reset(TickInterval)
}
}
}
// Pass is one look: outside the daytime window it does nothing; otherwise it trims every due, running, owned guest
// while holding the heavy-op gate.
func (t *Trimmer) Pass(ctx context.Context) {
lt := t.now().In(t.loc)
if !EligibleHour(lt) {
t.logger.Debug("fstrim: outside the daytime window — not looking", "local", lt.Format("Mon 15:04"))
return
}
guests, err := t.guests.Guests(ctx)
if err != nil {
t.logger.Warn("fstrim: owned-guest list unavailable — skipping this pass", "err", err)
return
}
owned := make(map[int]bool, len(guests))
var todo []int
t.mu.Lock()
for _, g := range guests {
owned[g.VMID] = true
r, has := t.records[g.VMID]
if !due(r, has, lt) {
continue
}
if g.Status != "running" {
t.logger.Info("fstrim: guest not running — trimmed when it runs", "vmid", g.VMID, "status", g.Status)
continue
}
todo = append(todo, g.VMID)
}
// A guest the agent no longer owns has no result to report.
pruned := false
for id := range t.records {
if !owned[id] {
delete(t.records, id)
pruned = true
}
}
if pruned {
if err := t.saveLocked(); err != nil {
t.logger.Warn("fstrim: state save failed", "err", err)
}
}
t.mu.Unlock()
if len(todo) == 0 {
return
}
sort.Ints(todo)
release, busy, ok := t.gate.TryAcquire(GateName)
if !ok {
t.logger.Info("fstrim: deferred — a heavy operation is in flight; retrying next hour", "busy", busy, "due_guests", len(todo))
return
}
defer release()
for _, vmid := range todo {
if ctx.Err() != nil {
return
}
t.trimOne(ctx, vmid, lt)
}
}
var trimmedLine = regexp.MustCompile(`\((\d+) bytes\) trimmed`)
// ParseTrimmed sums the "(N bytes) trimmed" lines of `pct fstrim` output and counts them (one per mount point), e.g.
// `/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed`.
func ParseTrimmed(out string) (bytes int64, mounts int) {
for _, m := range trimmedLine.FindAllStringSubmatch(out, -1) {
n, err := strconv.ParseInt(m[1], 10, 64)
if err != nil {
continue
}
bytes += n
mounts++
}
return bytes, mounts
}
// GiB renders bytes as "30.1 GiB".
func GiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
func (t *Trimmer) trimOne(ctx context.Context, vmid int, lt time.Time) {
start := t.now()
cctx, cancel := context.WithTimeout(ctx, PerGuestTimeout)
stdout, stderr, err := t.runner.Run(cctx, "pct", "fstrim", strconv.Itoa(vmid))
cancel()
dur := t.now().Sub(start)
bytes, mounts := ParseTrimmed(string(stdout) + "\n" + string(stderr))
t.mu.Lock()
prev, has := t.records[vmid]
r := Record{LastAttemptAt: start.UTC(), OK: err == nil, BytesTrimmed: bytes, Mounts: mounts,
DurationSeconds: float64(dur.Round(100*time.Millisecond)) / float64(time.Second), LastOKAt: prev.LastOKAt}
if has && !prev.LastAttemptAt.Before(weekAnchor(lt)) {
r.Attempts = prev.Attempts + 1
} else {
r.Attempts = 1
}
if err == nil {
r.LastOKAt = start.UTC()
} else {
msg := strings.TrimSpace(err.Error() + ": " + strings.TrimSpace(string(stderr)))
if len(msg) > 300 {
msg = msg[:300]
}
r.Error = msg
}
t.records[vmid] = r
saveErr := t.saveLocked()
t.mu.Unlock()
if err == nil {
t.logger.Info(fmt.Sprintf("fstrim: guest %d trimmed %s in %.1fs", vmid, GiB(bytes), r.DurationSeconds),
"vmid", vmid, "bytes_trimmed", bytes, "mounts", mounts, "duration_s", r.DurationSeconds)
if mounts == 0 {
t.logger.Warn("fstrim: pct fstrim succeeded but reported no trimmed mount — output not understood",
"vmid", vmid, "stdout", strings.TrimSpace(string(stdout)))
}
} else {
t.logger.Warn(fmt.Sprintf("fstrim: guest %d trim FAILED after %.1fs", vmid, r.DurationSeconds),
"vmid", vmid, "attempt", r.Attempts, "max_attempts_per_week", MaxAttemptsPerWeek, "err", r.Error)
}
if saveErr != nil {
t.logger.Warn("fstrim: state save failed — the result will not survive a restart", "path", t.statePath, "err", saveErr)
}
}
// GuestDiskTrimStatus implements hub.GuestDiskTrimReporter: a pure read of the persisted results (never runs pct).
func (t *Trimmer) GuestDiskTrimStatus(context.Context) *hub.GuestDiskTrimStatus {
t.mu.Lock()
defer t.mu.Unlock()
out := &hub.GuestDiskTrimStatus{Schedule: ScheduleText}
ids := make([]int, 0, len(t.records))
for id := range t.records {
ids = append(ids, id)
}
sort.Ints(ids)
for _, id := range ids {
r := t.records[id]
g := hub.GuestDiskTrim{VMID: id, LastAttemptAt: r.LastAttemptAt.UTC().Format(time.RFC3339), OK: r.OK,
BytesTrimmed: r.BytesTrimmed, Mounts: r.Mounts, DurationSeconds: r.DurationSeconds, Error: r.Error}
if !r.LastOKAt.IsZero() {
g.LastOKAt = r.LastOKAt.UTC().Format(time.RFC3339)
}
out.Guests = append(out.Guests, g)
}
return out
}
+267
View File
@@ -0,0 +1,267 @@
package fstrim
import (
"bytes"
"context"
"encoding/json"
"errors"
"log/slog"
"path/filepath"
"reflect"
"strings"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// The real `pct fstrim 9201` output measured on demo-hp 2026-10-06 (audits/ten-answers-2026-10-06/r444-measure.txt).
const measuredOut = "/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed\n" +
"/var/lib/lxc/9201/rootfs/var/lib/felhom: 53.9 GiB (57865633792 bytes) trimmed\n"
const measuredBytes = int64(32277680128 + 57865633792)
type fakeRunner struct {
mu sync.Mutex
calls [][]string
out string
err error
onRun func()
}
func (f *fakeRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
f.mu.Lock()
f.calls = append(f.calls, append([]string{name}, args...))
f.mu.Unlock()
if f.onRun != nil {
f.onRun()
}
if f.err != nil {
return nil, []byte("mount busy"), f.err
}
return []byte(f.out), nil, nil
}
type fakeGuests struct {
g []proxmox.Guest
err error
}
func (f fakeGuests) Guests(context.Context) ([]proxmox.Guest, error) { return f.g, f.err }
// A Wednesday 10:30 in a fixed zone (CEST-like), so the tests do not depend on the machine's zone.
var zone = time.FixedZone("CEST", 2*3600)
func at(day, hour, min int) time.Time { return time.Date(2026, 10, day, hour, min, 0, 0, zone) } // 2026-10-07 = Wednesday
func newT(t *testing.T, r Runner, g GuestSource, gate Gate, now *time.Time) (*Trimmer, *bytes.Buffer, string) {
t.Helper()
var logs bytes.Buffer
path := filepath.Join(t.TempDir(), "guest-disk-trim.json")
tr := New(r, g, gate, path, slog.New(slog.NewTextHandler(&logs, &slog.HandlerOptions{Level: slog.LevelDebug})))
tr.loc = zone
tr.now = func() time.Time { return *now }
return tr, &logs, path
}
func running(ids ...int) fakeGuests {
var g []proxmox.Guest
for _, id := range ids {
g = append(g, proxmox.Guest{VMID: id, Status: "running", Type: "lxc"})
}
return fakeGuests{g: g}
}
func TestParseTrimmedTheMeasuredOutput(t *testing.T) {
b, m := ParseTrimmed(measuredOut)
if b != measuredBytes || m != 2 {
t.Fatalf("ParseTrimmed = %d bytes over %d mounts, want %d over 2", b, m, measuredBytes)
}
if b, m := ParseTrimmed("something else\n"); b != 0 || m != 0 {
t.Fatalf("unrelated output parsed as %d/%d", b, m)
}
if got := GiB(measuredBytes); got != "84.0 GiB" {
t.Fatalf("GiB = %q", got)
}
}
// The night window (01:00–06:59) must never be eligible, and the daytime window is exactly 10:00–20:59.
func TestEligibleHourNeverInTheNight(t *testing.T) {
for h := 0; h < 24; h++ {
lt := time.Date(2026, 10, 7, h, 30, 0, 0, zone)
got := EligibleHour(lt)
if h >= 1 && h <= 6 && got {
t.Errorf("hour %02d is in the night window and must not be eligible", h)
}
if want := h >= 10 && h <= 20; got != want {
t.Errorf("EligibleHour(%02d:30) = %v, want %v", h, got, want)
}
}
}
func TestWeekAnchorIsTheLastWednesdayTen(t *testing.T) {
cases := map[time.Time]time.Time{
at(7, 10, 0): at(7, 10, 0), // Wednesday 10:00 itself
at(7, 9, 59): time.Date(2026, 9, 30, 10, 0, 0, 0, zone), // before 10:00 Wednesday → the previous week
at(8, 15, 0): at(7, 10, 0), // Thursday
at(13, 20, 0): at(7, 10, 0), // next Tuesday
at(14, 11, 0): at(14, 10, 0), // next Wednesday
}
for in, want := range cases {
if got := weekAnchor(in); !got.Equal(want) {
t.Errorf("weekAnchor(%s) = %s, want %s", in.Format("Mon 01-02 15:04"), got.Format("Mon 01-02 15:04"), want.Format("Mon 01-02 15:04"))
}
}
}
// The consequence: on Wednesday 10:30 a running owned guest is trimmed with the ONE exact argv, the bytes are parsed,
// the positive log line is written, the result is persisted, and the host report carries it.
func TestPassTrimsADueGuestAndReportsIt(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
tr, logs, path := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9201"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("runner calls = %q, want %q", r.calls, want)
}
if !strings.Contains(logs.String(), "fstrim: guest 9201 trimmed 84.0 GiB in ") {
t.Fatalf("no positive per-guest log line:\n%s", logs.String())
}
st := tr.GuestDiskTrimStatus(context.Background())
if st == nil || st.Schedule != ScheduleText || len(st.Guests) != 1 {
t.Fatalf("report stanza = %+v", st)
}
g := st.Guests[0]
if g.VMID != 9201 || !g.OK || g.BytesTrimmed != measuredBytes || g.Mounts != 2 || g.LastOKAt == "" || g.LastAttemptAt == "" {
t.Fatalf("report guest = %+v", g)
}
// Persisted: a NEW Trimmer over the same file (an agent restart) still has it and does not trim again this week.
now = at(8, 11, 0)
r2 := &fakeRunner{out: measuredOut}
tr2 := New(r2, running(9201), &backup.InFlight{}, path, slog.New(slog.NewTextHandler(&bytes.Buffer{}, nil)))
tr2.loc, tr2.now = zone, func() time.Time { return now }
if st2 := tr2.GuestDiskTrimStatus(context.Background()); len(st2.Guests) != 1 || st2.Guests[0].BytesTrimmed != measuredBytes {
t.Fatalf("result lost over a restart: %+v", st2)
}
tr2.Pass(context.Background())
if len(r2.calls) != 0 {
t.Fatalf("trimmed again in the same week after a restart: %q", r2.calls)
}
// Next week it is due again.
now = at(14, 10, 5)
tr2.Pass(context.Background())
if len(r2.calls) != 1 {
t.Fatalf("not trimmed in the next week: %q", r2.calls)
}
}
func TestPassNeverRunsInTheNight(t *testing.T) {
for _, h := range []int{1, 3, 6, 9, 21, 23} {
now := at(7, h, 15)
r := &fakeRunner{out: measuredOut}
tr, _, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Errorf("trimmed at %02d:15: %q", h, r.calls)
}
}
}
// A backup (or restore-test) holding the heavy-op gate DEFERS the trim; the next hour, gate free, it runs. And while
// a trim runs, the gate is held, so a backup cannot start beside it.
func TestPassDefersToAHeavyOperationAndRetriesNextHour(t *testing.T) {
now := at(7, 10, 30)
gate := &backup.InFlight{}
release, _, _ := gate.TryAcquire("backup:9201")
var busyDuringTrim string
r := &fakeRunner{out: measuredOut}
r.onRun = func() { busyDuringTrim = gate.Busy() }
tr, logs, _ := newT(t, r, running(9201), gate, &now)
tr.Pass(context.Background())
if len(r.calls) != 0 {
t.Fatalf("trimmed beside a running backup: %q", r.calls)
}
if !strings.Contains(logs.String(), "fstrim: deferred") || !strings.Contains(logs.String(), "backup:9201") {
t.Fatalf("the deferral is not logged with what holds the gate:\n%s", logs.String())
}
release()
now = now.Add(time.Hour)
tr.Pass(context.Background())
if len(r.calls) != 1 {
t.Fatalf("not retried the next hour: %q", r.calls)
}
if busyDuringTrim != GateName {
t.Fatalf("the heavy-op gate was %q during the trim, want %q", busyDuringTrim, GateName)
}
if gate.Busy() != "" {
t.Fatalf("the gate was not released after the pass: %q", gate.Busy())
}
}
func TestFailedTrimWarnsIsRecordedAndRetriedAtMostThreeTimes(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{err: errors.New("exit status 255")}
tr, logs, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
for i := 0; i < 6; i++ {
tr.Pass(context.Background())
now = now.Add(time.Hour)
}
if len(r.calls) != MaxAttemptsPerWeek {
t.Fatalf("attempts in one week = %d, want %d", len(r.calls), MaxAttemptsPerWeek)
}
if !strings.Contains(logs.String(), "level=WARN") || !strings.Contains(logs.String(), "fstrim: guest 9201 trim FAILED") {
t.Fatalf("no WARN for the failure:\n%s", logs.String())
}
g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]
if g.OK || g.LastOKAt != "" || !strings.Contains(g.Error, "exit status 255") || !strings.Contains(g.Error, "mount busy") {
t.Fatalf("failed result not recorded as a failure: %+v", g)
}
// A success later keeps a clean record.
r.err = nil
r.out = measuredOut
now = at(14, 10, 10)
tr.Pass(context.Background())
if g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]; !g.OK || g.Error != "" || g.BytesTrimmed != measuredBytes {
t.Fatalf("success after failure: %+v", g)
}
}
func TestOnlyRunningOwnedGuestsAndAFailedListActsOnNothing(t *testing.T) {
now := at(7, 10, 30)
r := &fakeRunner{out: measuredOut}
g := fakeGuests{g: []proxmox.Guest{{VMID: 9201, Status: "stopped"}, {VMID: 9202, Status: "running"}}}
tr, _, _ := newT(t, r, g, &backup.InFlight{}, &now)
tr.Pass(context.Background())
if want := [][]string{{"pct", "fstrim", "9202"}}; !reflect.DeepEqual(r.calls, want) {
t.Fatalf("calls = %q, want only the running guest", r.calls)
}
r2 := &fakeRunner{out: measuredOut}
tr2, logs, _ := newT(t, r2, fakeGuests{err: errors.New("pool read 403")}, &backup.InFlight{}, &now)
tr2.Pass(context.Background())
if len(r2.calls) != 0 || !strings.Contains(logs.String(), "owned-guest list unavailable") {
t.Fatalf("a failed ownership read must act on nothing: calls %q", r2.calls)
}
}
func TestReportJSONShape(t *testing.T) {
now := at(7, 10, 30)
tr, _, _ := newT(t, &fakeRunner{out: measuredOut}, running(9201), &backup.InFlight{}, &now)
tr.Pass(context.Background())
b, err := json.Marshal(tr.GuestDiskTrimStatus(context.Background()))
if err != nil {
t.Fatal(err)
}
for _, k := range []string{`"schedule":`, `"guests":[{"vmid":9201`, `"last_attempt_at":"2026-10-07T08:30:00Z"`, `"ok":true`,
`"bytes_trimmed":90143313920`, `"mounts":2`, `"duration_seconds":`, `"last_ok_at":"2026-10-07T08:30:00Z"`} {
if !strings.Contains(string(b), k) {
t.Errorf("report JSON lacks %s: %s", k, b)
}
}
if strings.Contains(string(b), `"error"`) {
t.Errorf("an ok result must omit error: %s", b)
}
}
+17
View File
@@ -90,6 +90,12 @@ type GuestNetReporter interface {
GuestNetStatus(ctx context.Context) *GuestNetStatus GuestNetStatus(ctx context.Context) *GuestNetStatus
} }
// GuestDiskTrimReporter is the R-444 seam the weekly trim job plugs into (same consumer-side pattern — hub does not
// import fstrim). nil (feature not wired) → no guest_disk_trim stanza.
type GuestDiskTrimReporter interface {
GuestDiskTrimStatus(ctx context.Context) *GuestDiskTrimStatus
}
// Collector builds a HostReport from read-only sources. All deps are behind narrow // Collector builds a HostReport from read-only sources. All deps are behind narrow
// interfaces for unit testing. // interfaces for unit testing.
type Collector struct { type Collector struct {
@@ -108,6 +114,7 @@ type Collector struct {
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted) pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted) ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted) guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
diskTrim GuestDiskTrimReporter // R-444: weekly guest disk trim (nil → stanza omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false) selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted) mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
oob OOBReporter // H1: operator-access health (nil → stanza omitted) oob OOBReporter // H1: operator-access health (nil → stanza omitted)
@@ -221,6 +228,12 @@ func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
return c return c
} }
// SetGuestDiskTrimReporter wires the R-444 weekly trim job as a report source (nil-safe → stanza omitted).
func (c *Collector) SetGuestDiskTrimReporter(r GuestDiskTrimReporter) *Collector {
c.diskTrim = r
return c
}
// SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side // SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side
// pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report. // pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report.
type SelfUpdateReporter interface { type SelfUpdateReporter interface {
@@ -377,6 +390,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.guestNet != nil { if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx) report.GuestNet = c.guestNet.GuestNetStatus(ctx)
} }
// R-444: the last weekly trim result per guest (nil reporter = not wired → stanza omitted).
if c.diskTrim != nil {
report.GuestDiskTrim = c.diskTrim.GuestDiskTrimStatus(ctx)
}
// D1: agent self-update pending status (nil reporter → pending=false, the steady state). // D1: agent self-update pending status (nil reporter → pending=false, the steady state).
if c.selfUpdate != nil { if c.selfUpdate != nil {
report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending() report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending()
+54
View File
@@ -0,0 +1,54 @@
package hub
import (
"context"
"encoding/json"
"testing"
)
// R-444: the guest_disk_trim stanza must reach a report built through the PRODUCTION collect path, be absent from
// the wire when the job is not wired, and carry the keys the hub's System page reads.
type fakeDiskTrim struct{ st *GuestDiskTrimStatus }
func (f fakeDiskTrim) GuestDiskTrimStatus(context.Context) *GuestDiskTrimStatus { return f.st }
func TestCollect_GuestDiskTrim(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.150.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ := json.Marshal(r)
var m map[string]any
_ = json.Unmarshal(b, &m)
if _, ok := m["guest_disk_trim"]; ok {
t.Fatalf("guest_disk_trim on the wire with no reporter wired: %s", b)
}
c.SetGuestDiskTrimReporter(fakeDiskTrim{st: &GuestDiskTrimStatus{Schedule: "weekly", Guests: []GuestDiskTrim{{
VMID: 9201, LastAttemptAt: "2026-10-07T08:30:00Z", OK: true, BytesTrimmed: 90143313920, Mounts: 2,
DurationSeconds: 24.4, LastOKAt: "2026-10-07T08:30:00Z",
}}}})
r, err = c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
b, _ = json.Marshal(r)
m = nil
_ = json.Unmarshal(b, &m)
dt, ok := m["guest_disk_trim"].(map[string]any)
if !ok || dt["schedule"] != "weekly" {
t.Fatalf("guest_disk_trim missing or wrong on the wire: %s", b)
}
g := dt["guests"].([]any)[0].(map[string]any)
for _, k := range []string{"vmid", "last_attempt_at", "ok", "bytes_trimmed", "mounts", "duration_seconds", "last_ok_at"} {
if _, ok := g[k]; !ok {
t.Fatalf("guest_disk_trim.guests[0] lacks %q: %v", k, g)
}
}
if g["bytes_trimmed"] != float64(90143313920) || g["ok"] != true {
t.Fatalf("values did not survive the round trip: %v", g)
}
}
+37
View File
@@ -124,6 +124,11 @@ type HostReport struct {
// on HostReport would have been the only report block named against that convention. // on HostReport would have been the only report block named against that convention.
GuestNet *GuestNetStatus `json:"guest_net,omitempty"` GuestNet *GuestNetStatus `json:"guest_net,omitempty"`
// GuestDiskTrim is the weekly guest disk trim stanza (R-444, `09` §3 decision 139): the schedule and, per owned
// guest, the LAST trim result as persisted by the agent (it survives a restart). Present only when the trim job is
// wired; an empty `guests` list means the job runs and no guest has been trimmed yet. No secret.
GuestDiskTrim *GuestDiskTrimStatus `json:"guest_disk_trim,omitempty"`
// LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent // LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent
// mirror of the controller's report log_tails channel. Present ONLY on the heartbeat // mirror of the controller's report log_tails channel. Present ONLY on the heartbeat
// right after the control envelope requested it (log_tail_requested); consume-once on // right after the control envelope requested it (log_tail_requested); consume-once on
@@ -208,6 +213,27 @@ type GuestNetGuest struct {
Message string `json:"message,omitempty"` Message string `json:"message,omitempty"`
} }
// GuestDiskTrimStatus is the R-444 weekly trim stanza. `schedule` is a plain description of when the job runs (local
// time of the host); `guests` holds one entry per owned guest that has had at least one trim attempt.
type GuestDiskTrimStatus struct {
Schedule string `json:"schedule"`
Guests []GuestDiskTrim `json:"guests,omitempty"`
}
// GuestDiskTrim is one guest's LAST trim attempt. `ok` with `last_attempt_at` is the verdict of that attempt — never
// read the time alone as success; `last_ok_at` is the last attempt that succeeded ("" = never). `bytes_trimmed` is
// the sum of the "(N bytes) trimmed" lines `pct fstrim` printed, over `mounts` mount points.
type GuestDiskTrim struct {
VMID int `json:"vmid"`
LastAttemptAt string `json:"last_attempt_at"`
OK bool `json:"ok"`
BytesTrimmed int64 `json:"bytes_trimmed"`
Mounts int `json:"mounts"`
DurationSeconds float64 `json:"duration_seconds"`
LastOKAt string `json:"last_ok_at,omitempty"`
Error string `json:"error,omitempty"`
}
type PBSDRStatus struct { type PBSDRStatus struct {
State string `json:"state"` State string `json:"state"`
StorageID string `json:"storage_id,omitempty"` StorageID string `json:"storage_id,omitempty"`
@@ -428,6 +454,17 @@ type SmartSummary struct {
ReallocatedSectors *int `json:"reallocated_sectors"` ReallocatedSectors *int `json:"reallocated_sectors"`
PendingSectors *int `json:"pending_sectors"` PendingSectors *int `json:"pending_sectors"`
OfflineUncorrectable *int `json:"offline_uncorrectable"` OfflineUncorrectable *int `json:"offline_uncorrectable"`
// R-330 (disk health Phase 2): three more SATA raw counters. omitempty + pointer: absent (an
// older agent, an NVMe/USB device, or a drive that does not report the attribute) is OMITTED —
// unknown, never a zero (S-39). Wire only: no verdict reads them yet.
// 187 Reported_Uncorrect — the failing drive's most telling counter (normalized 1 vs thresh 0,
// raw 1001) while SMART still said PASSED.
// 188 Command_Timeout — some vendors PACK several counters into the 48-bit raw value, so the
// number is carried as reported and must not be compared across vendors.
// 199 UDMA_CRC_Error_Count — cabling / link errors, not the medium.
ReportedUncorrect *int64 `json:"reported_uncorrect,omitempty"`
CommandTimeout *int64 `json:"command_timeout,omitempty"`
UDMACRCErrors *int64 `json:"udma_crc_errors,omitempty"`
// NVMe attributes. // NVMe attributes.
CriticalWarning *int `json:"critical_warning"` CriticalWarning *int `json:"critical_warning"`
+102
View File
@@ -0,0 +1,102 @@
package localapi
import (
"encoding/json"
"errors"
"io"
"io/fs"
"net/http"
"os"
"time"
)
// GET /host/crash-guard (R-856, `09` §3 decision 143): what the host's crash guard
// (configs/felhom-crash-guard, `11` §5.9) recorded about the most recent HOST boot. The controller
// reads it once after it starts: when the host's last boot followed an UNCLEAN stop, its app mails
// wait ~15 minutes instead of the normal 90 s boot grace.
//
// Read-only and Proxmox-free: the agent reads the guard's state file (root-owned, 0644 — the
// non-root agent can read it) and passes four fields through. Host-wide, token-authed (any valid
// per-guest token sees the host's view, as GET /host/metrics does).
//
// NEVER an error page. A missing file (no guard installed, or no boot recorded yet), an unreadable
// one, or one that does not parse answers 200 with present:false — the controller reads that as
// UNKNOWN and keeps its normal boot grace. Pinned by TestR856_CrashGuard*.
// defaultCrashGuardStatePath is where configs/felhom-crash-guard writes its state (STATE_DIR there).
const defaultCrashGuardStatePath = "/var/lib/felhom-crash-guard/state.json"
// crashGuardStateMax bounds the read; the real file is well under 4 KiB.
const crashGuardStateMax = 1 << 20
// CrashGuardResponse is the data block of GET /host/crash-guard. Field names are the controller's
// agentapi.CrashGuardState (felhom-controller internal/agentapi/crashguard.go) — a wire contract,
// pinned by TestR856_CrashGuardWireMatchesControllerClient.
type CrashGuardResponse struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"` // RFC3339 UTC ("2006-01-02T15:04:05Z")
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// crashGuardFile is the subset of the guard's state.json the route passes through. Every other key
// (armed, boot_id, config, unclean_boots, last_trip, ...) is ignored.
type crashGuardFile struct {
LastBootAt string `json:"last_boot_at"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
// readCrashGuardState reads and parses the guard's state file. ok=false on ANY failure (missing,
// unreadable, oversized, not a JSON object, a field of the wrong type); reason says which, for the log.
func readCrashGuardState(path string) (resp CrashGuardResponse, ok bool, reason string) {
f, err := os.Open(path)
if err != nil {
if errors.Is(err, fs.ErrNotExist) {
return resp, false, "no state file"
}
return resp, false, "unreadable: " + err.Error()
}
defer f.Close()
raw, err := io.ReadAll(io.LimitReader(f, crashGuardStateMax+1))
if err != nil {
return resp, false, "read: " + err.Error()
}
if len(raw) > crashGuardStateMax {
return resp, false, "state file too large"
}
var st crashGuardFile
// Unmarshal into a struct fails on a non-object top level (null decodes, so reject it below).
if err := json.Unmarshal(raw, &st); err != nil {
return resp, false, "unparseable: " + err.Error()
}
var probe map[string]json.RawMessage
if err := json.Unmarshal(raw, &probe); err != nil || probe == nil {
return resp, false, "unparseable: not a JSON object"
}
resp = CrashGuardResponse{Present: true, LastBootUnclean: st.LastBootUnclean, Tripped: st.Tripped}
// Normalise to RFC3339 UTC; an unparseable time passes through as-is (the controller reads an
// unparseable boot time as "not this start's boot" → its normal grace).
if t, perr := time.Parse(time.RFC3339, st.LastBootAt); perr == nil {
resp.LastBootAt = t.UTC().Format(time.RFC3339)
} else {
resp.LastBootAt = st.LastBootAt
}
return resp, true, ""
}
func (s *Server) handleCrashGuard(w http.ResponseWriter, r *http.Request, vmid int) {
path := s.crashGuardStatePath
if path == "" {
path = defaultCrashGuardStatePath
}
resp, ok, reason := readCrashGuardState(path)
if !ok {
s.logger.Debug("local-api: /host/crash-guard not present", "vmid", vmid, "reason", reason)
writeOK(w, CrashGuardResponse{Present: false})
return
}
s.logger.Debug("local-api: /host/crash-guard served", "vmid", vmid,
"last_boot_at", resp.LastBootAt, "last_boot_unclean", resp.LastBootUnclean, "tripped", resp.Tripped)
writeOK(w, resp)
}
+205
View File
@@ -0,0 +1,205 @@
package localapi
import (
"encoding/json"
"io"
"log/slog"
"net/http"
"os"
"path/filepath"
"testing"
)
// The shape of /var/lib/felhom-crash-guard/state.json as read on demo-hp on 2026-10-06 (values from
// that read where they matter; lists/objects kept to the same key set).
const crashGuardFixture = `{
"armed": true,
"boot_id": "3f1c0f1e-6a0b-4d7e-9b7a-0c2d4e6f8a1b",
"config": {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10},
"kernel_panic": 10,
"last_boot_at": "2026-10-05T07:56:41Z",
"last_boot_unclean": true,
"last_trip": {},
"rearmed_at": "2026-10-04T14:02:11Z",
"rearmed_by": "operator",
"tripped": false,
"unclean_boots": ["2026-10-05T07:56:41Z"],
"unclean_boots_24h": 1,
"unclean_boots_in_window": 1,
"updated_at": "2026-10-05T07:57:02Z",
"version": 1
}`
// controllerCrashGuardState is a COPY of the controller's wire type, felhom-controller
// controller/internal/agentapi/crashguard.go `CrashGuardState` (commit 8b13a5e) — same field names,
// same tags. If either side renames a key, the contract test below fails.
type controllerCrashGuardState struct {
Present bool `json:"present"`
LastBootAt string `json:"last_boot_at,omitempty"`
LastBootUnclean bool `json:"last_boot_unclean"`
Tripped bool `json:"tripped"`
}
func newCrashGuardServer(t *testing.T, statePath string) http.Handler {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0",
Guests: &fakeGuests{},
Backups: &fakeBackups{},
Store: &fakeStore{},
Storage: fakeStorage{},
Tokens: staticTokens{"A": 8200, "B": 9300},
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatalf("new server: %v", err)
}
srv.crashGuardStatePath = statePath
return srv.Handler()
}
func writeCrashGuardFixture(t *testing.T, body string) string {
t.Helper()
p := filepath.Join(t.TempDir(), "state.json")
if err := os.WriteFile(p, []byte(body), 0o644); err != nil {
t.Fatal(err)
}
return p
}
// getCrashGuard calls the route and decodes the envelope with the CONTROLLER's type.
func getCrashGuard(t *testing.T, h http.Handler, token string) (int, controllerCrashGuardState, string) {
t.Helper()
w := do(t, h, "GET", "/host/crash-guard", token, "")
var env struct {
OK bool `json:"ok"`
Data controllerCrashGuardState `json:"data"`
}
if w.Code == http.StatusOK {
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil {
t.Fatalf("decode %q: %v", w.Body.String(), err)
}
if !env.OK {
t.Fatalf("ok=false: %s", w.Body.String())
}
}
return w.Code, env.Data, w.Body.String()
}
// A present state file (demo-hp's shape) passes the three facts through.
func TestR856_CrashGuardPresentFile(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, body)
}
want := controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: true, Tripped: false}
if st != want {
t.Fatalf("state = %+v, want %+v", st, want)
}
// A tripped, clean boot reads back as such (both bools are carried, not defaulted).
h = newCrashGuardServer(t, writeCrashGuardFixture(t,
`{"last_boot_at":"2026-10-05T09:56:41+02:00","last_boot_unclean":false,"tripped":true,"version":1}`))
_, st, _ = getCrashGuard(t, h, "B")
want = controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: false, Tripped: true}
if st != want {
t.Fatalf("offset time / tripped: state = %+v, want %+v (time normalised to UTC Z)", st, want)
}
}
// No state file (no guard on this host, or no boot recorded yet) → 200 present:false.
func TestR856_CrashGuardMissingFile(t *testing.T) {
h := newCrashGuardServer(t, filepath.Join(t.TempDir(), "absent", "state.json"))
code, st, body := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("missing file: got %d, want 200 (%s)", code, body)
}
if st.Present || st.LastBootUnclean || st.Tripped || st.LastBootAt != "" {
t.Fatalf("missing file: state = %+v, want present:false and nothing else", st)
}
}
// A garbled file → 200 present:false, never a 5xx — every shape of garbage.
func TestR856_CrashGuardGarbageFile(t *testing.T) {
for name, body := range map[string]string{
"truncated": crashGuardFixture[:40],
"not json": "this is not json\n",
"empty": "",
"null": "null",
"array": `[{"last_boot_unclean":true}]`,
"wrong type": `{"last_boot_at":"2026-10-05T07:56:41Z","last_boot_unclean":"yes","tripped":false}`,
"lone brace": "{",
} {
t.Run(name, func(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, body))
code, st, raw := getCrashGuard(t, h, "A")
if code != http.StatusOK {
t.Fatalf("got %d, want 200 (%s)", code, raw)
}
if st.Present || st.LastBootUnclean {
t.Fatalf("garbage %q read as %+v, want present:false", name, st)
}
})
}
// The path is a directory, not a file: unreadable → present:false, 200.
h := newCrashGuardServer(t, t.TempDir())
if code, st, raw := getCrashGuard(t, h, "A"); code != http.StatusOK || st.Present {
t.Fatalf("directory path: got %d %+v (%s), want 200 present:false", code, st, raw)
}
}
// No / unknown token → 401, like every sibling route; a cross-guest ?vmid= → 403.
func TestR856_CrashGuardRequiresGuestToken(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
for _, tok := range []string{"", "bogus"} {
w := do(t, h, "GET", "/host/crash-guard", tok, "")
if w.Code != http.StatusUnauthorized {
t.Fatalf("token %q: got %d, want 401", tok, w.Code)
}
if json.Valid(w.Body.Bytes()) {
var env struct {
Data controllerCrashGuardState `json:"data"`
}
_ = json.Unmarshal(w.Body.Bytes(), &env)
if env.Data.Present || env.Data.LastBootUnclean {
t.Fatalf("token %q: the refusal leaked the state: %s", tok, w.Body.String())
}
}
}
if w := do(t, h, "GET", "/host/crash-guard?vmid=9300", "A", ""); w.Code != http.StatusForbidden {
t.Fatalf("cross-guest query: got %d, want 403", w.Code)
}
}
// Wire contract: every key the controller's CrashGuardState decodes is emitted under exactly that
// name, and the agent emits no key the controller does not know.
func TestR856_CrashGuardWireMatchesControllerClient(t *testing.T) {
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
w := do(t, h, "GET", "/host/crash-guard", "A", "")
var env struct {
OK bool `json:"ok"`
Data map[string]json.RawMessage `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil || !env.OK {
t.Fatalf("envelope: %v %s", err, w.Body.String())
}
want := []string{"present", "last_boot_at", "last_boot_unclean", "tripped"}
for _, k := range want {
if _, ok := env.Data[k]; !ok {
t.Errorf("agent does not emit %q, which the controller decodes (%s)", k, w.Body.String())
}
}
if len(env.Data) != len(want) {
t.Errorf("agent emits %d keys, controller knows %d: %s", len(env.Data), len(want), w.Body.String())
}
// And the agent's own type agrees with the controller's copy, field for field.
var mine CrashGuardResponse
var theirs controllerCrashGuardState
raw, _ := json.Marshal(env.Data)
_ = json.Unmarshal(raw, &mine)
_ = json.Unmarshal(raw, &theirs)
if (controllerCrashGuardState{mine.Present, mine.LastBootAt, mine.LastBootUnclean, mine.Tripped}) != theirs {
t.Errorf("agent %+v vs controller %+v", mine, theirs)
}
}
+144
View File
@@ -0,0 +1,144 @@
package localapi
import (
"context"
"io"
"log/slog"
"net/http"
"path/filepath"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-894 — after an agent restart, an UNREADABLE storage must fall back to the last success saved on
// disk, not to "never". Measured 2026-10-05 on demo-hp: a restart at 04:57, the off-site storage
// unreachable at 06:25, the 7-day tier (last copy 4 days old) read DUE, vzdump failed.
//
// Every test here builds a NEW server and a NEW BackupSuccessState from the same file — that is the
// restart. The in-memory store (fakeStore) is always fresh, as after a real restart.
// r894Server builds a two-tier server whose off-site tier answers the storage listing with lister.
func r894Server(t *testing.T, path string, pbsSvc BackupService) *Server {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: &fakeStore{},
Storage: fakeStorage{targets: []hub.StorageTarget{{Name: "local"}, {Name: "felhom-pbs"}}},
Tokens: staticTokens{"A": 8200},
BackupTiers: []BackupTier{
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: pbsSvc},
},
LastKnownBackups: backup.NewBackupSuccessState(path),
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatal(err)
}
srv.baseCtx = context.Background()
srv.now = func() time.Time { return testNow }
return srv
}
// unreadable is the off-site storage as demo-hp saw it: "Can't connect to 10.77.0.1:8007".
func unreadable() archiveLister {
return archiveLister{fakeBackups: &fakeBackups{}, err: errStorageRead}
}
// THE R-894 CASE, end to end. Agent 1 takes an off-site backup through POST /backup (the fake
// runner's success is 12 h before testNow). The agent restarts. The storage cannot be read. The tier
// must read NOT due, from the copy saved on disk.
//
// COMPANION RED-PROOF (observed): delete the `lookup == archiveUnknown && s.lastKnown != nil` block in
// handleBackupDue → this fails with "after a restart an unreadable storage must fall back to the saved
// copy (12 h old, 7-day tier) — NOT due; got {… Due:true … AgeState:unknown …}". Restored.
func TestBackupDue_R894_RestartThenUnreadableStorage_FreshSavedCopyIsNotDue(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
// Agent 1: a real backup job through the endpoint the controller calls.
first := r894Server(t, path, &fakeBackups{})
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
}
waitFor(t, func() bool {
_, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200)
return ok
})
// Agent 2: a restart (new server, new state from the same file), and the storage is unreachable.
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if got.Due {
t.Fatalf("after a restart an unreadable storage must fall back to the saved copy (12 h old, 7-day tier) — NOT due; got %+v", got)
}
if got.AgeState != AgeStateKnown || got.AgeSecs == nil || *got.AgeSecs != int64((12*time.Hour).Seconds()) {
t.Fatalf("the age must come from the saved copy (12 h, known); got %+v", got)
}
}
// The deliberate rule stays: an unreadable storage must not suppress a backup that IS due. A saved
// copy older than the cadence reads DUE.
//
// COMPANION RED-PROOF (observed): make the fallback answer not-due whenever a saved copy exists
// (`if fromDisk { …Due:false… }` before the cadence check) → this fails with "a saved copy 9 days old
// under a 7-day cadence MUST read due". Restored.
func TestBackupDue_R894_RestartThenUnreadableStorage_OldSavedCopyIsDue(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
st := backup.NewBackupSuccessState(path)
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, 9*24*time.Hour, true)); err != nil {
t.Fatal(err)
}
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("a saved copy 9 days old under a 7-day cadence MUST read due; got %+v", got)
}
if got.AgeState != AgeStateKnown {
t.Fatalf("the age is known (from disk); got %+v", got)
}
}
// No saved copy → the pre-R-894 answer, byte for byte: DUE, age UNKNOWN (never ABSENT — the controller
// fires its window-gate valve only on absent, R-88).
func TestBackupDue_R894_RestartThenUnreadableStorage_NoSavedCopyIsDueUnknown(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due || got.AgeState != AgeStateUnknown || got.AgeSecs != nil {
t.Fatalf("no saved copy + unreadable storage must stay DUE with age unknown; got %+v", got)
}
}
// A storage that ANSWERS is the ground truth: an archive absent there makes the tier due even when the
// file remembers a fresh success (a pruned or deleted copy must be made again).
//
// COMPANION RED-PROOF (observed): drop `lookup == archiveUnknown &&` from the fallback condition → this
// fails with "the storage answered 'no archive' — the saved copy must NOT stand in for it". Restored.
func TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
st := backup.NewBackupSuccessState(path)
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, time.Hour, true)); err != nil {
t.Fatal(err)
}
absent := archiveLister{fakeBackups: &fakeBackups{}, found: false}
got := dueFor(t, r894Server(t, path, absent).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("the storage answered 'no archive' — the saved copy must NOT stand in for it; got %+v", got)
}
}
// A FAILED backup is never saved: it must not make a tier look fresh after a restart.
func TestBackupDue_R894_FailedBackupIsNotSaved(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
first := r894Server(t, path, &fakeBackups{failErr: "could not activate storage 'felhom-pbs'"})
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
}
waitFor(t, func() bool { return len(first.store.Backups(context.Background())) == 1 })
if _, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200); ok {
t.Fatal("a failed backup must never be saved as a success")
}
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
if !got.Due {
t.Fatalf("after a failed backup and a restart the tier must still be due; got %+v", got)
}
}
+67 -22
View File
@@ -85,6 +85,13 @@ type BackupStore interface {
RestoreTests(ctx context.Context) []hub.RestoreTest RestoreTests(ctx context.Context) []hub.RestoreTest
} }
// LastKnownBackupStore (R-894) is the on-disk newest-success-per-tier record. Satisfied by
// *backup.BackupSuccessState.
type LastKnownBackupStore interface {
RecordBackupSuccess(target string, b hub.Backup) error
LastKnownSuccess(target string, vmid int) (time.Time, bool)
}
// StorageView yields the host's observed storage targets (for mapping a mount's storage id → // StorageView yields the host's observed storage targets (for mapping a mount's storage id →
// fast/slow class). Satisfied by *storage.Observer. // fast/slow class). Satisfied by *storage.Observer.
type StorageView interface { type StorageView interface {
@@ -164,6 +171,10 @@ type Options struct {
// PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg // PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg
// that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs. // that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs.
AfterPrimaryBackup func(ctx context.Context, vmid int) AfterPrimaryBackup func(ctx context.Context, vmid int)
// LastKnownBackups (R-894) keeps the newest successful backup per tier ON DISK, so the due-check's
// fallback for an UNREADABLE storage after an agent restart is the last known copy, not "never".
// nil = the pre-R-894 behaviour (in-memory record only).
LastKnownBackups LastKnownBackupStore
// Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when // Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when
// nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner. // nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner.
Privileged PrivilegedRunner Privileged PrivilegedRunner
@@ -229,7 +240,6 @@ type Options struct {
// POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured" // POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured"
// (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer. // (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer.
EscrowRecovery EscrowRecoverer EscrowRecovery EscrowRecoverer
} }
// defaultBackupCadence is the fallback /backup/due window when none is configured. // defaultBackupCadence is the fallback /backup/due window when none is configured.
@@ -279,30 +289,34 @@ type Server struct {
// is the pre-R-82 shape. // is the pre-R-82 shape.
tiers []BackupTier tiers []BackupTier
// inFlight (R-85) is shared with the restore-test scheduler so the two never run together. // inFlight (R-85) is shared with the restore-test scheduler so the two never run together.
inFlight *backup.InFlight inFlight *backup.InFlight
afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none
logger *slog.Logger lastKnown LastKnownBackupStore // R-894: on-disk newest success per tier; nil = none
now func() time.Time logger *slog.Logger
now func() time.Time
disks DiskOps // slice 8C (optional) disks DiskOps // slice 8C (optional)
diskGate StorageGate // slice 8C (optional) diskGate StorageGate // slice 8C (optional)
guestList GuestLister // slice 8C (optional) guestList GuestLister // slice 8C (optional)
guestAttach GuestAttacher // slice 10 P2 (optional) guestAttach GuestAttacher // slice 10 P2 (optional)
mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional) mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional)
memMu sync.Mutex // single-flight around a resize apply (one customer per host) memMu sync.Mutex // single-flight around a resize apply (one customer per host)
netStorage NetworkStorageOps // Part A1: NAS network mounts (optional) netStorage NetworkStorageOps // Part A1: NAS network mounts (optional)
netMountRoot string // the user-data namespace root for the network-mount role gate netMountRoot string // the user-data namespace root for the network-mount role gate
smbCredsDir string // where SMB creds files are written (out-of-band, 0600) smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
// crashGuardStatePath (R-856) is the host crash guard's state file read by GET /host/crash-guard;
// empty = defaultCrashGuardStatePath. A seam: tests point it at a fixture.
crashGuardStatePath string
// escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed // escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the // identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
// offsite repository password. OPTIONAL — nil (no hub client configured) makes // offsite repository password. OPTIONAL — nil (no hub client configured) makes
// POST /escrow/recover-offsite-password answer 503 rather than pretending. // POST /escrow/recover-offsite-password answer 503 rather than pretending.
escrowRecovery EscrowRecoverer escrowRecovery EscrowRecoverer
intent IntentRecorder // slice 10 P3 (optional) intent IntentRecorder // slice 10 P3 (optional)
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional) guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional) formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
staleLock StaleLockController // F2-b startup stale-lock recovery (optional) staleLock StaleLockController // F2-b startup stale-lock recovery (optional)
// guestPower (F-REBOOT) is per-guest start-attempt state for the guest-power watchdog. // guestPower (F-REBOOT) is per-guest start-attempt state for the guest-power watchdog.
// Guarded by guestPowerMu in guestpower.go; in-memory on purpose (see guestPowerState). // Guarded by guestPowerMu in guestpower.go; in-memory on purpose (see guestPowerState).
guestPower map[int]guestPowerState guestPower map[int]guestPowerState
@@ -475,6 +489,7 @@ func NewServer(o Options) (*Server, error) {
s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence) s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence)
s.inFlight = o.InFlight s.inFlight = o.InFlight
s.afterPrimaryBackup = o.AfterPrimaryBackup s.afterPrimaryBackup = o.AfterPrimaryBackup
s.lastKnown = o.LastKnownBackups
if s.backups == nil && len(s.tiers) > 0 { if s.backups == nil && len(s.tiers) > 0 {
s.backups = s.tiers[0].Service s.backups = s.tiers[0].Service
} }
@@ -520,6 +535,9 @@ func (s *Server) Handler() http.Handler {
// Host metrics (slice 9): host-wide health + per-storage capacity for the customer's monitoring // Host metrics (slice 9): host-wide health + per-storage capacity for the customer's monitoring
// view. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot). // view. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot).
mux.HandleFunc("GET /host/metrics", s.withGuest(s.handleHostMetrics)) mux.HandleFunc("GET /host/metrics", s.withGuest(s.handleHostMetrics))
// R-856 (`09` §3 decision 143): the host crash guard's record of the last HOST boot — the controller
// waits ~15 min with app mails after an unclean one. Read-only; a missing/garbled file = present:false.
mux.HandleFunc("GET /host/crash-guard", s.withGuest(s.handleCrashGuard))
// Disk management (slice 8C) — self-scoped; format routes through the data-bearing classifier+gate. // Disk management (slice 8C) — self-scoped; format routes through the data-bearing classifier+gate.
mux.HandleFunc("GET /disks", s.withGuest(s.handleDisks)) mux.HandleFunc("GET /disks", s.withGuest(s.handleDisks))
mux.HandleFunc("GET /disks/candidates", s.withGuest(s.handleDiskCandidates)) mux.HandleFunc("GET /disks/candidates", s.withGuest(s.handleDiskCandidates))
@@ -899,6 +917,13 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
s.logger.Info("local-api: backup job complete", "vmid", vmid, "target", tier.TargetID, "job", jobID, "archive", b.Archive) s.logger.Info("local-api: backup job complete", "vmid", vmid, "target", tier.TargetID, "job", jobID, "archive", b.Archive)
} }
s.store.RecordBackup(b) s.store.RecordBackup(b)
if b.Success && s.lastKnown != nil {
if err := s.lastKnown.RecordBackupSuccess(tier.TargetID, b); err != nil {
// Not fatal: the backup exists. Only the fallback after a restart loses this copy.
s.logger.Warn("local-api: could not save the backup on disk for the due-check fallback (R-894)",
"vmid", vmid, "target", tier.TargetID, "err", err)
}
}
s.finishJob(key, jobID, b) s.finishJob(key, jobID, b)
// OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate. // OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate.
if b.Success && tier.Primary && s.afterPrimaryBackup != nil { if b.Success && tier.Primary && s.afterPrimaryBackup != nil {
@@ -1061,6 +1086,20 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
newest, haveNewest = t, true newest, haveNewest = t, true
unparseable = false // ground truth supersedes an unreadable in-memory timestamp unparseable = false // ground truth supersedes an unreadable in-memory timestamp
} }
// R-894: the storage could not be read → the last success saved ON DISK stands in for the in-memory
// record a restart emptied. ONLY on archiveUnknown: a storage that answers is the ground truth, and an
// archive absent there must make the tier due even when the file remembers one (a pruned copy).
// A saved copy older than the cadence still reads DUE below — an unreadable storage never suppresses
// a backup that is due.
fromDisk := false
if lookup == archiveUnknown && s.lastKnown != nil {
if saved, ok := s.lastKnown.LastKnownSuccess(tier.TargetID, vmid); ok && (!haveNewest || saved.After(newest)) {
newest, haveNewest, fromDisk = saved, true, true
unparseable = false
s.logger.Info("local-api: backup storage unreadable — due-check uses the last success saved on disk (R-894)",
"vmid", vmid, "target", tier.TargetID, "saved", saved.UTC().Format(time.RFC3339))
}
}
if !haveNewest { if !haveNewest {
// R-88 Part 2: THREE distinct reasons for a nil age, each with its own state. Only ABSENT is a // R-88 Part 2: THREE distinct reasons for a nil age, each with its own state. Only ABSENT is a
// positive claim of "never backed up"; only that one may license the controller to bypass its // positive claim of "never backed up"; only that one may license the controller to bypass its
@@ -1083,13 +1122,17 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
} }
age := s.now().Sub(newest) age := s.now().Sub(newest)
ageSecs := int64(age.Seconds()) ageSecs := int64(age.Seconds())
suffix := ""
if fromDisk {
suffix = " (storage unreadable — age from the last success saved on disk)"
}
if age >= tier.Cadence { if age >= tier.Cadence {
writeOK(w, BackupDueResponse{VMID: vmid, Due: true, AgeSecs: &ageSecs, AgeState: AgeStateKnown, writeOK(w, BackupDueResponse{VMID: vmid, Due: true, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
Reason: "older than cadence", Target: echo}) Reason: "older than cadence" + suffix, Target: echo})
return return
} }
writeOK(w, BackupDueResponse{VMID: vmid, Due: false, AgeSecs: &ageSecs, AgeState: AgeStateKnown, writeOK(w, BackupDueResponse{VMID: vmid, Due: false, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
Reason: "within cadence window", Target: echo}) Reason: "within cadence window" + suffix, Target: echo})
} }
// BackupTiersResponse is GET /backup/tiers (R-82): the tiers this agent serves, primary first. // BackupTiersResponse is GET /backup/tiers (R-82): the tiers this agent serves, primary first.
@@ -1477,4 +1520,6 @@ func writeStatus(w http.ResponseWriter, code int, ok bool, data any, errMsg stri
} }
// SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0). // SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0).
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) { s.afterPrimaryBackup = fn } func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) {
s.afterPrimaryBackup = fn
}
+15
View File
@@ -106,6 +106,9 @@ type WrapperReport struct {
LiveRestore json.RawMessage `json:"live_restore"` LiveRestore json.RawMessage `json:"live_restore"`
Facts json.RawMessage `json:"facts"` Facts json.RawMessage `json:"facts"`
Bundle json.RawMessage `json:"bundle"` // the config bundle's result (R-840, mode "bundle") Bundle json.RawMessage `json:"bundle"` // the config bundle's result (R-840, mode "bundle")
// OOMCheck (R-528, `09` decision 157): the docker layer's memory-kill check, {result, oom_killed, oom_event,
// exit_code, image, detail}. Carried to the hub UNCHANGED; the agent never reads it.
OOMCheck json.RawMessage `json:"oom_check"`
// R-868 (v0.144.0): the agent's ids, echoed from the plan, so a report kept on disk can be sent without the // R-868 (v0.144.0): the agent's ids, echoed from the plan, so a report kept on disk can be sent without the
// agent process that started the pass. ReleaseID / VMID were always in the report. // agent process that started the pass. ReleaseID / VMID were always in the report.
RunID string `json:"run_id"` RunID string `json:"run_id"`
@@ -143,6 +146,9 @@ type Report struct {
DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step
Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed
Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade) Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade)
// OOMCheck: docker layer — the wrapper's oom_check object, byte-for-byte (R-528; the hub decides approval on it).
// Pinned by TestDocker_OOMCheckReachesTheHubUnchanged and TestR868_KeptCopyCarriesTheOOMCheck.
OOMCheck json.RawMessage `json:"oom_check,omitempty"`
unsent string // R-868: the wrapper's kept copy of this pass's report — deleted once the hub has it unsent string // R-868: the wrapper's kept copy of this pass's report — deleted once the hub has it
} }
@@ -630,6 +636,7 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
} }
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
rep.OOMCheck = rawOrNil(wr.OOMCheck)
if rep.Outcome == "" { if rep.Outcome == "" {
switch { switch {
case rep.Mode == "inventory" && !blk.Enabled: case rep.Mode == "inventory" && !blk.Enabled:
@@ -707,6 +714,14 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
return l.finish(ctx, lg, rep) return l.finish(ctx, lg, rep)
} }
// rawOrNil: a wrapper field that is absent or JSON null stays out of the hub report (omitempty).
func rawOrNil(m json.RawMessage) json.RawMessage {
if len(m) == 0 || string(m) == "null" {
return nil
}
return m
}
func onlyDocker(in []Package) []Package { func onlyDocker(in []Package) []Package {
var out []Package var out []Package
for _, p := range in { for _, p := range in {
+44 -1
View File
@@ -98,9 +98,13 @@ func (f *fakeWrapper) RunStdin(ctx context.Context, _ io.Reader, name string, ar
return f.Run(ctx, name, args...) return f.Run(ctx, name, args...)
} }
type fakeHub struct{ reports []Report } type fakeHub struct {
reports []Report
bodies [][]byte // the exact bytes posted (R-528: the oom_check object must arrive unchanged)
}
func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error { func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error {
h.bodies = append(h.bodies, append([]byte(nil), body...))
var r Report var r Report
json.Unmarshal(body, &r) json.Unmarshal(body, &r)
h.reports = append(h.reports, r) h.reports = append(h.reports, r)
@@ -504,3 +508,42 @@ func TestHealthVerdict_ControllerBlindToDockerFails(t *testing.T) {
t.Fatal("an older wrapper (no field) must not fail") t.Fatal("an older wrapper (no field) must not fail")
} }
} }
// R-528 (`09` decision 157): the wrapper's oom_check object reaches the hub's docker report byte-for-byte; the guest
// and host reports carry none. COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in runLayer → "no
// oom_check in the docker report".
const oomCheckWire = `{"detail":"the engine reported the memory kill: OOMKilled=true and the oom event","exit_code":137,"image":"gitea.dooplex.hu/admin/felhom-controller:0.300.0","oom_event":true,"oom_killed":true,"result":"pass"}`
func TestDocker_OOMCheckReachesTheHubUnchanged(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}},
DockerEngine: "29.8.2", Authority: "ring0", OOMCheck: json.RawMessage(oomCheckWire)}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Run(context.Background(), 9201, "night")
found := false
for _, b := range h.bodies {
var m map[string]json.RawMessage
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
var layer string
json.Unmarshal(m["layer"], &layer)
oc, has := m["oom_check"]
if layer != LayerDocker {
if has {
t.Fatalf("the %s report carries an oom_check: %s", layer, oc)
}
continue
}
found = true
if !has {
t.Fatalf("no oom_check in the docker report: %s", b)
}
if string(oc) != oomCheckWire {
t.Fatalf("oom_check changed on the way:\n got %s\nwant %s", oc, oomCheckWire)
}
}
if !found {
t.Fatalf("no docker report posted: %s", calls(w))
}
}
+1
View File
@@ -144,6 +144,7 @@ func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string)
} }
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
rep.OOMCheck = rawOrNil(wr.OOMCheck)
wantEngine := "" wantEngine := ""
for _, u := range wr.Upgraded { for _, u := range wr.Upgraded {
if u.Name == "docker-ce" { if u.Name == "docker-ce" {
+21
View File
@@ -150,3 +150,24 @@ func must(t *testing.T, err error) {
t.Fatal(err) t.Fatal(err)
} }
} }
// R-528: a kept docker report (the agent was killed) still carries the oom_check object to the hub unchanged.
// COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in reportFromKept → "oom_check lost".
func TestR868_KeptCopyCarriesTheOOMCheck(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
ring := 0
kept := WrapperReport{Mode: "apply", Layer: LayerDocker, RunID: "20261007T020000Z", Trigger: "night", Ring: &ring, VMID: 9201,
ReleaseID: "ring0-20261007T020000Z", HealthBefore: guestOK(), HealthAfter: guestOK(), DockerEngine: "29.8.2",
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}}, OOMCheck: json.RawMessage(oomCheckWire)}
b, _ := json.Marshal(kept)
must(t, os.WriteFile(reportFile(l.PlanDir, kept.RunID, LayerDocker, "apply"), b, 0o600))
if n := l.SendUnsent(context.Background()); n != 1 || len(h.bodies) != 1 {
t.Fatalf("sent %d, bodies %d", n, len(h.bodies))
}
var m map[string]json.RawMessage
must(t, json.Unmarshal(h.bodies[0], &m))
if string(m["oom_check"]) != oomCheckWire {
t.Fatalf("oom_check lost or changed: %q", m["oom_check"])
}
}
+64
View File
@@ -0,0 +1,64 @@
package storage
import (
"encoding/json"
"strings"
"testing"
)
// R-330 (disk health Phase 2) — attributes 187, 188 and 199 ride the wire.
//
// The failing drive of 2026-08-14 carried 187 Reported_Uncorrect at raw 1001 while SMART said PASSED;
// none of the three reached the controller. The values are carried as RAW counters; an absent
// attribute is OMITTED from the JSON (unknown), never sent as 0 (S-39).
// The 2026-08-14 shape: PASSED, 187 raw 1001, plus 188/199 and the existing three.
const r330SATA = `{"smart_status":{"passed":true},"ata_smart_attributes":{"table":[
{"id":5,"raw":{"value":0}},
{"id":187,"raw":{"value":1001}},
{"id":188,"raw":{"value":4295032833}},
{"id":197,"raw":{"value":8}},
{"id":198,"raw":{"value":8}},
{"id":199,"raw":{"value":3}}]}}`
// COMPANION RED-PROOF (observed): delete the three R-330 cases in parseSMART → this fails with
// "187 Reported_Uncorrect must be carried (raw 1001); got <nil>". Restored.
func TestParseSMART_R330_CarriesTheThreeCounters(t *testing.T) {
s := parseSMART([]byte(r330SATA))
if s.ReportedUncorrect == nil || *s.ReportedUncorrect != 1001 {
t.Fatalf("187 Reported_Uncorrect must be carried (raw 1001); got %v", s.ReportedUncorrect)
}
// 188's raw value is vendor-packed on some drives (this one is 0x100010001): carried as reported, not truncated.
if s.CommandTimeout == nil || *s.CommandTimeout != 4295032833 {
t.Fatalf("188 Command_Timeout must be carried as the full raw value; got %v", s.CommandTimeout)
}
if s.UDMACRCErrors == nil || *s.UDMACRCErrors != 3 {
t.Fatalf("199 UDMA_CRC_Error_Count must be carried (raw 3); got %v", s.UDMACRCErrors)
}
if s.PendingSectors == nil || *s.PendingSectors != 8 {
t.Fatalf("the existing counters must be unchanged; pending=%v", s.PendingSectors)
}
}
// A drive (or an NVMe device) that does not report the attributes leaves them nil, and the JSON
// OMITS the keys — the receiver reads "unknown", never a measured zero.
//
// COMPANION RED-PROOF (observed): drop `,omitempty` from the three tags in hub.SmartSummary → this
// fails with "an unreported attribute must be omitted, not sent: … reported_uncorrect …". Restored.
func TestParseSMART_R330_AbsentIsOmittedNotZero(t *testing.T) {
for name, raw := range map[string]string{
"sata without the three": `{"smart_status":{"passed":true},"ata_smart_attributes":{"table":[{"id":5,"raw":{"value":0}}]}}`,
"nvme": `{"smart_status":{"passed":true},"nvme_smart_health_information_log":{"critical_warning":0,"media_errors":0,"percentage_used":3}}`,
} {
s := parseSMART([]byte(raw))
if s.ReportedUncorrect != nil || s.CommandTimeout != nil || s.UDMACRCErrors != nil {
t.Fatalf("%s: unreported attributes must stay nil; got %v %v %v", name, s.ReportedUncorrect, s.CommandTimeout, s.UDMACRCErrors)
}
b, _ := json.Marshal(s)
for _, k := range []string{"reported_uncorrect", "command_timeout", "udma_crc_errors"} {
if strings.Contains(string(b), `"`+k+`"`) {
t.Fatalf("%s: an unreported attribute must be omitted, not sent: %s", name, b)
}
}
}
}
+11
View File
@@ -44,6 +44,9 @@ const (
ataReallocatedSectorCt = 5 ataReallocatedSectorCt = 5
ataCurrentPending = 197 ataCurrentPending = 197
ataOfflineUncorrect = 198 ataOfflineUncorrect = 198
ataReportedUncorrect = 187 // R-330
ataCommandTimeout = 188 // R-330
ataUDMACRCErrorCount = 199 // R-330
) )
// parseSMART maps smartctl JSON to a hub.SmartSummary, handling SATA + NVMe and degrading // parseSMART maps smartctl JSON to a hub.SmartSummary, handling SATA + NVMe and degrading
@@ -85,6 +88,12 @@ func parseSMART(raw []byte) hub.SmartSummary {
s.PendingSectors = intPtr(int(a.Raw.Value)) s.PendingSectors = intPtr(int(a.Raw.Value))
case ataOfflineUncorrect: case ataOfflineUncorrect:
s.OfflineUncorrectable = intPtr(int(a.Raw.Value)) s.OfflineUncorrectable = intPtr(int(a.Raw.Value))
case ataReportedUncorrect:
s.ReportedUncorrect = int64Ptr(a.Raw.Value)
case ataCommandTimeout:
s.CommandTimeout = int64Ptr(a.Raw.Value)
case ataUDMACRCErrorCount:
s.UDMACRCErrors = int64Ptr(a.Raw.Value)
} }
} }
} }
@@ -142,3 +151,5 @@ func parseThinPoolMetadata(raw []byte) (float64, bool) {
} }
func intPtr(v int) *int { return &v } func intPtr(v int) *int { return &v }
func int64Ptr(v int64) *int64 { return &v }
+367
View File
@@ -0,0 +1,367 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""test_gate_decoys.py — can this repo's gates be fooled by a LABEL? (R-421, R-426)
The same instrument as `felhom.eu/scripts/test_gate_decoys.py`: a decoy is the LABEL without the
FACT, and a gate that passes on the label alone — or refuses the genuine article — is a live hole.
Every gate is asserted in BOTH directions: the decoy must be convicted, the genuine article passed.
Covered here (the `COVERS` literal is AST-read by `felhom.eu/scripts/decoy_coverage_gate.py`, which
never imports this file):
published check-published-versions.py, against a FAKE Gitea (see below).
release-complete check-release-complete.py, in a scratch clone whose `origin` is a scratch bare
repository, against the same fake Gitea.
reuse-refs, instructions, observations
the three SHARED felhom.eu scripts, run against a scratch clone of THIS repo —
so the decoy is planted in the agent's own REUSE.md / CLAUDE.md / REPORT.md and
coverage is per input, not per script.
NEVER THE REAL GITEA. Both network gates read `GITEA_BASE` from the environment (CI already sets it
to the in-cluster URL); here it points at an `http.server` bound to 127.0.0.1 inside this process,
and every proxy variable is removed from the child's environment so urllib cannot route around it.
A test that asked the real registry would pass or fail on whatever was published that day — the
constant-for-measurement shape — and would reach the network from a hook.
NEVER THE REAL TREE. Every planted file lives in a scratch directory: a workspace that holds a
clone of this repo beside symlinks to the sibling clones the shared scripts reach across to.
Run from the repo root: python3 scripts/test_gate_decoys.py
Exit 0 every decoy judged correctly · 1 a decoy passed or a genuine article was refused.
"""
import http.server
import io
import json
import os
import shutil
import socketserver
import subprocess
import sys
import tempfile
import threading
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
PARENT = os.path.dirname(ROOT)
# DECOY_SHARED_DIR exists for ONE purpose: the red-proof. It lets a mutated COPY of the shared scripts
# be judged without editing the felhom.eu clone. Unset, the suite judges the real shared scripts.
SHARED = os.environ.get("DECOY_SHARED_DIR") or os.path.join(PARENT, "felhom.eu", "scripts")
# ── WHAT THIS FILE COVERS ────────────────────────────────────────────────────────────────────────
# Read by felhom.eu/scripts/decoy_coverage_gate.py, which AST-parses this literal. A gate named here
# MUST have a decoy below that has been seen to fail.
COVERS = {
"published": "against a FAKE Gitea: a tag whose package 404s, a tag whose tree lacks the configs, a package one "
"version past the newest tag (published, never tagged), a patch-gap orphan, a missing package that "
"lexical sorting would drop out of the retention window (0.9.x vs 0.10.x); a tags api answering 500 "
"or a non-JSON 200 is INCONCLUSIVE, never a pass - vs a clean registry, a non-semver tag and a version "
"older than the retention window (not asserted, BY DESIGN) (R-426)",
"release-complete": "the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated commit, a tag "
"with no package; a registry 500 is INCONCLUSIVE - vs the genuine release, an `## Unreleased` "
"heading above it, a newer version named only in prose or under `###`, a tag that only "
"origin has (the shallow-CI shape); a LOCAL-only tag passes BY DESIGN (CI's fresh clone "
"and the published gate's converse probe are what see it) (R-426)",
"reuse-refs": "a cited .go and a cited .md path that do not exist, planted in THIS repo's REUSE.md - vs the "
"real file (R-426)",
"instructions": "a component version literal in THIS repo's CLAUDE.md effective text - vs the same sentence "
"inside an HTML comment (R-426)",
"observations": "R-419 in THIS repo's REPORT.md: an Observations note SAYING it carries no marker - vs the "
"two genuine markers (R-426)",
}
fails = []
ran = 0
def report(name, rc, out, expect_rc, must=()):
global ran
ran += 1
missing = [m for m in must if m not in out]
if rc == expect_rc and not missing:
print(" ok %-62s rc=%d (expected %d)" % (name, rc, expect_rc))
else:
hole = expect_rc != 0 and rc == 0
fails.append("%s: rc=%d expected %d%s; missing %s\n%s" % (
name, rc, expect_rc, " - LIVE HOLE" if hole else "", missing, out[-900:]))
# ── the fake Gitea ───────────────────────────────────────────────────────────────────────────────
class Fake(object):
"""What the fake registry serves. Reset per case."""
def reset(self):
self.tags = [] # tag names, as the tags api lists them
self.packages = set() # versions whose generic package downloads
self.raw = set() # versions whose tag tree serves the probe config
self.tags_status = 200
self.tags_body = None # override bytes for the tags api
self.pkg_status = None # override status for EVERY package request
self.hits = []
FAKE = Fake()
FAKE.reset()
PKG_PREFIX = "/api/packages/admin/generic/felhom-agent/"
RAW_PREFIX = "/admin/felhom-agent/raw/tag/v"
class Handler(http.server.BaseHTTPRequestHandler):
def log_message(self, *a):
pass
def _answer(self, status, body=b""):
self.send_response(status)
self.send_header("Content-Length", str(len(body)))
self.end_headers()
if self.command != "HEAD":
self.wfile.write(body)
def do_GET(self):
p = self.path
FAKE.hits.append(p)
if p.startswith("/api/v1/repos/admin/felhom-agent/tags"):
body = FAKE.tags_body if FAKE.tags_body is not None else \
json.dumps([{"name": t} for t in FAKE.tags]).encode()
return self._answer(FAKE.tags_status, body)
if p.startswith(PKG_PREFIX):
if FAKE.pkg_status is not None:
return self._answer(FAKE.pkg_status)
v = p[len(PKG_PREFIX):].split("/", 1)[0]
return self._answer(200, b"ELF") if v in FAKE.packages else self._answer(404)
if p.startswith(RAW_PREFIX):
v = p[len(RAW_PREFIX):].split("/", 1)[0]
ok = v in FAKE.raw and p.endswith("/configs/felhom-agent.service")
return self._answer(200, b"[Unit]\n") if ok else self._answer(404)
return self._answer(404)
do_HEAD = do_GET
class Server(socketserver.ThreadingMixIn, http.server.HTTPServer):
daemon_threads = True
def child_env(base):
env = {k: v for k, v in os.environ.items() if "proxy" not in k.lower()}
env["GITEA_BASE"] = base
env["NO_PROXY"] = env["no_proxy"] = "127.0.0.1,localhost"
return env
def run(argv, cwd, env=None):
# input="" — a child must never inherit (and block on) this process's stdin
p = subprocess.run(argv, cwd=cwd, env=env, capture_output=True, text=True, input="")
return p.returncode, p.stdout + p.stderr
def sh(argv, cwd):
rc, out = run(argv, cwd)
if rc != 0:
raise SystemExit("setup command failed (%s): %s" % (" ".join(argv), out))
return out.strip()
# ── published ────────────────────────────────────────────────────────────────────────────────────
def published_cases(base):
gate = os.path.join(ROOT, "scripts", "check-published-versions.py")
keep = json.load(io.open(os.path.join(ROOT, "scripts", "retention-policy.json"),
encoding="utf-8"))["generic_versions_kept"]
TAGS = ["v0.150.%d" % i for i in range(3)] # inside any retention window >= 3
VERS = [t[1:] for t in TAGS]
def case(name, setup, expect_rc, must=()):
FAKE.reset()
FAKE.tags = list(TAGS)
FAKE.packages = set(VERS)
FAKE.raw = set(VERS)
setup()
rc, out = run([sys.executable, gate], ROOT, child_env(base))
if not FAKE.hits:
fails.append("published/%s: the gate never asked the fake Gitea - the seam is not wired" % name)
report("published: " + name, rc, out, expect_rc, must)
case("GENUINE: every tag downloadable and serving its configs", lambda: None, 0,
("ALL RELEASED VERSIONS INSTALLABLE",))
case("GENUINE: a non-semver tag is not a release", lambda: FAKE.tags.append("v0.150.2-rc1"), 0,
("ALL RELEASED VERSIONS INSTALLABLE",))
case("FACT: a tag whose package 404s", lambda: FAKE.packages.discard("0.150.1"), 1,
("FAIL v0.150.1", "binary NOT downloadable"))
case("FACT: a tag whose tree does not serve the configs", lambda: FAKE.raw.discard("0.150.2"), 1,
("FAIL v0.150.2", "does not serve"))
case("FACT: published one patch past the newest tag, never tagged",
lambda: FAKE.packages.add("0.150.3"), 1, ("PUBLISHED VERSION(S) WITH NO TAG", "v0.150.3"))
case("FACT: published in a patch GAP between two tags",
lambda: (FAKE.tags.remove("v0.150.1"),), 1, ("v0.150.1 is downloadable", "has no git tag"))
def lexical():
# keep+1 tags: 0.9.0 and 0.10.0..0.10.<keep-1>. By SEMVER the oldest is 0.9.0 (dropped); by
# STRING sort "0.10.0" is the smallest and would be the one dropped - so its missing package
# is convicted only if the window is cut by semver.
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.10.0"}
FAKE.raw = set(t[1:] for t in FAKE.tags)
case("FACT: a missing package lexical sorting would drop (0.10.0 vs 0.9.0)", lexical, 1,
("FAIL v0.10.0",))
def retired():
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.9.0"}
FAKE.raw = set(t[1:] for t in FAKE.tags)
case("BY DESIGN: a version older than the retention window is not asserted", retired, 0,
("NOT ASSERTED", "0.9.0"))
def five_hundred():
FAKE.tags_status = 500
case("INCONCLUSIVE: the tags api answers 500", five_hundred, 2, ("INCONCLUSIVE",))
def html():
FAKE.tags_body = b"<html>sign in</html>"
case("INCONCLUSIVE: the tags api answers a 200 that is not JSON", html, 2, ("INCONCLUSIVE",))
# an unreachable Gitea: a port nothing listens on
s = Server(("127.0.0.1", 0), Handler)
dead = "http://127.0.0.1:%d" % s.server_address[1]
s.server_close()
rc, out = run([sys.executable, gate], ROOT, child_env(dead))
report("published: INCONCLUSIVE: Gitea unreachable", rc, out, 2, ("INCONCLUSIVE", "URLs tried"))
# ── release-complete ─────────────────────────────────────────────────────────────────────────────
def release_cases(base, ws):
bare = os.path.join(ws, "origin.git")
work = os.path.join(ws, "rc-work")
sh(["git", "clone", "-q", "--bare", "--no-tags", "file://" + ROOT, bare], ws)
sh(["git", "clone", "-q", "--no-tags", "file://" + bare, work], ws)
sh(["git", "config", "user.email", "decoy@gate.invalid"], work)
sh(["git", "config", "user.name", "decoy"], work)
# the WORKING-TREE gate, so the file under test is the one being edited, not HEAD's
shutil.copy(os.path.join(ROOT, "scripts", "check-release-complete.py"),
os.path.join(work, "scripts", "check-release-complete.py"))
sh(["git", "add", "scripts/check-release-complete.py"], work)
sh(["git", "commit", "-q", "--allow-empty", "-m", "the gate under test"], work)
base_sha = sh(["git", "rev-parse", "HEAD"], work)
ch = os.path.join(work, "CHANGELOG.md")
original = io.open(ch, encoding="utf-8").read()
V = "9.9.9"
def case(name, top, expect_rc, must=(), tag=None, origin_tag=False, packaged=True, pkg_status=None):
FAKE.reset()
if packaged:
FAKE.packages = {V}
FAKE.pkg_status = pkg_status
try:
io.open(ch, "w", encoding="utf-8").write(top + original)
sh(["git", "commit", "-q", "-am", name], work)
if tag == "head":
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", "HEAD"], work)
elif tag == "unrelated":
empty = sh(["git", "mktree"], work) # stdin is "" — the empty tree, written to this repo
orphan = sh(["git", "commit-tree", "-m", "unrelated", empty], work)
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", orphan], work)
if origin_tag:
sh(["git", "push", "-q", "origin", "HEAD:refs/tags/v" + V], work)
rc, out = run([sys.executable, os.path.join(work, "scripts", "check-release-complete.py")],
work, child_env(base))
report("release-complete: " + name, rc, out, expect_rc, must)
finally:
run(["git", "tag", "-d", "v" + V], work)
run(["git", "push", "-q", "origin", ":refs/tags/v" + V], work)
sh(["git", "reset", "-q", "--hard", base_sha], work)
HEAD = "## v%s — 2026-10-06\n\n- decoy release\n\n" % V
case("GENUINE: tagged at HEAD and published", HEAD, 0,
("newest CHANGELOG version: v9.9.9", "is tagged, placed and published"), tag="head")
case("GENUINE: an `## Unreleased` heading above the release", "## Unreleased\n\n- wip\n\n" + HEAD, 0,
("newest CHANGELOG version: v9.9.9",), tag="head")
case("GENUINE: a newer version named only in prose and under ###",
"The `## v10.0.0` heading is not written yet.\n### v10.0.0 notes\n\n" + HEAD, 0,
("newest CHANGELOG version: v9.9.9",), tag="head")
case("GENUINE: the tag only on origin (the shallow-CI shape)", HEAD, 0,
("exists on origin",), origin_tag=True)
case("BY DESIGN: a LOCAL-only tag passes (CI's fresh clone sees only origin)", HEAD, 0,
("an ancestor of HEAD",), tag="head")
case("FACT: the newest heading has no tag anywhere", HEAD, 1, ("DOES NOT EXIST",))
case("FACT: a tag parked on an unrelated commit", HEAD, 1, ("NOT an ancestor",), tag="unrelated")
case("FACT: tagged, never published", HEAD, 1, ("IS NOT PUBLISHED",), tag="head", packaged=False)
case("INCONCLUSIVE: the registry answers 500", HEAD, 2, ("INCONCLUSIVE",), tag="head", pkg_status=500)
case("FACT beats INCONCLUSIVE: no tag AND the registry answers 500", HEAD, 1, ("DOES NOT EXIST",),
pkg_status=500)
# ── the shared felhom.eu scripts, against THIS repo's inputs ─────────────────────────────────────
def shared_cases(ws):
"""A scratch WORKSPACE: a clone of this repo beside symlinks to the siblings, because the shared
scripts reach across (REUSE.md cites hub paths; instructions_gate reads the workspace CLAUDE.md)."""
for g in ("reuse_refs_check.py", "instructions_gate.py", "observations_gate.py"):
if not os.path.isfile(os.path.join(SHARED, g)):
fails.append("shared gate %s is MISSING beside this clone (tried %s) - a failure, never a skip"
% (g, SHARED))
return
space = os.path.join(ws, "workspace")
os.makedirs(space)
for entry in sorted(os.listdir(PARENT)):
if entry in ("felhom.eu", "felhom-controller", "app-catalog-felhom.eu", "homelab-manifests",
"CLAUDE.md", ".claude-memory"):
os.symlink(os.path.join(PARENT, entry), os.path.join(space, entry))
repo = os.path.join(space, "felhom-agent")
sh(["git", "clone", "-q", "--no-tags", "file://" + ROOT, repo], ws)
# the WORKING-TREE inputs the plants go into, so a case judges today's file
for f in ("REUSE.md", "CLAUDE.md", "REPORT.md"):
shutil.copy(os.path.join(ROOT, f), os.path.join(repo, f))
def case(name, gate, relpath, extra, expect_rc, must=()):
p = os.path.join(repo, relpath)
backup = io.open(p, encoding="utf-8").read()
try:
if extra:
io.open(p, "w", encoding="utf-8").write(backup + extra)
rc, out = run([sys.executable, os.path.join(SHARED, gate), repo], repo)
report(name, rc, out, expect_rc, must)
finally:
io.open(p, "w", encoding="utf-8").write(backup)
case("reuse-refs: GENUINE: this repo's REUSE.md", "reuse_refs_check.py", "REUSE.md", "", 0, ("FAILED 0",))
case("reuse-refs: FACT: a cited .go path that does not exist", "reuse_refs_check.py", "REUSE.md",
u"\n- see `internal/localapi/does_not_exist.go`\n", 1, ("does_not_exist.go",))
case("reuse-refs: FACT: a cited .md path that does not exist", "reuse_refs_check.py", "REUSE.md",
u"\n- see `docs/99-does-not-exist.md`\n", 1, ("99-does-not-exist.md",))
case("instructions: GENUINE: this repo's CLAUDE.md", "instructions_gate.py", "CLAUDE.md", "", 0,
("instructions_gate: OK",))
case("instructions: FACT: a version literal in effective text", "instructions_gate.py", "CLAUDE.md",
u"\nThe agent runs v0.148.0 today.\n", 1, ("v0.148.0",))
case("instructions: GENUINE: the same sentence in an HTML comment", "instructions_gate.py", "CLAUDE.md",
u"\n<!--\nThe agent ran v0.148.0 on 2026-10-06.\n-->\n", 0, ("instructions_gate: OK",))
case("observations: FACT: R-419, prose SAYING it has no marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** It carries no `FILED:` marker and no "
u"`NOT-A-FINDING:` marker, deliberately.\n", 1)
case("observations: GENUINE: a FILED marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** Something broke. **FILED: R-419**\n", 0)
case("observations: GENUINE: a NOT-A-FINDING marker", "observations_gate.py", "REPORT.md",
u"\n## Observations\n\n1. **A real finding.** Odd. **NOT-A-FINDING: my own typo, corrected in "
u"the same minute.**\n", 0)
def main():
srv = Server(("127.0.0.1", 0), Handler)
threading.Thread(target=srv.serve_forever, daemon=True).start()
base = "http://127.0.0.1:%d" % srv.server_address[1]
ws = tempfile.mkdtemp(prefix="agent-decoys-")
print("agent gate decoys — fake Gitea at %s, scratch %s" % (base, ws))
try:
published_cases(base)
release_cases(base, ws)
shared_cases(ws)
finally:
srv.shutdown()
srv.server_close()
shutil.rmtree(ws, ignore_errors=True)
if fails:
print()
for f in fails:
print("FAIL: %s" % f)
return 1
print("\nagent gate decoys OK — %d case(s), every label judged on its fact (R-421)" % ran)
return 0
if __name__ == "__main__":
sys.exit(main())