Compare commits
15 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| f277e619e2 | |||
| 386f51edc6 | |||
| 130e3ed882 | |||
| be398f92e8 | |||
| ee71abd1d4 | |||
| 37e98f452b | |||
| b2b82ae828 | |||
| b78a0ff3ac | |||
| 96047453cb | |||
| 4bf5db6875 | |||
| 87977ff40a | |||
| 861d32a4b4 | |||
| 769c4c3cf2 | |||
| 64f704d0f7 | |||
| 208fac8027 |
+32
-1
@@ -1,4 +1,35 @@
|
||||
## unreleased (to be released as v0.147.0) — the recovery recipe spells the root namespace the way PBS does; a removed drive no longer shows the root disk's size; a rotated-out token stops at once; the dnsmasq check looks at the right package (burn-down round 2: R-124, R-118, R-269, R-317) (2026-10-05)
|
||||
## unreleased
|
||||
|
||||
- R-856 (`09` §3 decision 143): new local-API route `GET /host/crash-guard` — passes the host crash guard's last-boot record (present, last_boot_at, last_boot_unclean, tripped) from /var/lib/felhom-crash-guard/state.json to the controller, which waits ~15 min with app mails after a crash boot. Read-only, no Proxmox call, guest-token authed; a missing/unreadable/garbled file answers 200 present:false (never an error page). An older agent answers 404, which the controller reads as unknown (normal 90 s grace) — no controller MinAgent raise needed.
|
||||
- R-444 (`09` §3 decision 139): weekly guest disk trim. New sudoers alias FELHOM_FSTRIM with ONE exact rule `/usr/sbin/pct ^fstrim [0-9]+$` (rides the signed config bundle; decoys pinned by TestSudoersFstrimRuleIsExact) and capability guest-fstrim (non-critical). New internal/fstrim job: each owned RUNNING guest gets `pct fstrim <vmid>` once a week - due Wednesday from 10:00 host-local, starts only 10:00-20:59 (never the 01:00-06:59 night), holds the one-heavy-op gate so it never runs beside a backup or restore-test (busy -> deferred to the next hourly tick; a box that was off catches up at its next daytime hour); a failed trim WARNs and is retried at most 3 times that week; bytes parsed from `pct fstrim`'s "(N bytes) trimmed" lines; positive log `fstrim: guest N trimmed X GiB in Ys`; last result per guest persisted in <state_dir>/guest-disk-trim.json and reported as the new omitempty host-report stanza `guest_disk_trim`. Opt-out: agent.json "disk_trim": {"disable": true}.
|
||||
- R-99 (`09` §3 decision 140): the agent's WARN for a PBS archive below the 1 MiB plausibility floor now ends with the pointer to the sanctioned cleanup (`documentation/runbooks/pbs-phantom-cleanup.md`); detection only — nothing is deleted automatically. Dir-storage archives keep the old text.
|
||||
|
||||
## unreleased
|
||||
|
||||
- R-426: scripts/test_gate_decoys.py (new) — the published gate judged against a fake Gitea (127.0.0.1, via GITEA_BASE; never the real registry): 11 cases; COVERS published — `felhom-agent/published` leaves the decoy-coverage EXEMPT list.
|
||||
- R-426: release-complete gate — 10 decoy cases (scratch clone + scratch bare origin + fake Gitea); COVERS release-complete — `felhom-agent/release-complete` leaves the decoy-coverage EXEMPT list.
|
||||
- R-426: the shared reuse-refs/instructions/observations gates get agent-side decoys (9 cases on a scratch clone of this repo); COVERS reuse-refs, instructions, observations — three `felhom-agent/*` entries leave the decoy-coverage EXEMPT list.
|
||||
- **Fixed without a row:** `configs/test_felhom_config_bundle.py` read the two ISO first-boot files that installer 1.32.0 now NAMES under KEPT (R-275) as files the installer writes — `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 (felhom.eu `85de3f9b`); they are listed with why. Test only; agent v0.148.0's code is unaffected.
|
||||
|
||||
## v0.148.0 — the host report names the running binary's sha; the format answer carries the new filesystem's UUID (burn-down night: R-349, R-25 agent halves) (2026-10-06)
|
||||
|
||||
Released by `scripts/release-agent.sh`: binary sha256 `3e68a0870e0e2ce262cb4819294edddeb0a73e8c558a31611a20329a6d9ee283`
|
||||
config bundle sha256 `a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de` (tag `v0.148.0` = `861d32a`).
|
||||
Delivery order as for v0.147.0: signed `agent_update`, then signed `agent_config_update`.
|
||||
|
||||
- **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from
|
||||
`/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub
|
||||
comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved.
|
||||
- **R-25 (agent half):** `POST /disks/format` and `GET /disks/format/status` return `fs_uuid`, the new filesystem's UUID
|
||||
read back after mkfs only when the bound durable id still resolves to the formatted device and the superblock is the
|
||||
requested type (empty = not verified); `DeviceProbe` gains `FSUUID` from blkid. Tests `TestFormat_*FSUUID*`; three
|
||||
red-proofs. The controller half (mount that UUID) is a controller change.
|
||||
|
||||
## v0.147.0 — the recovery recipe spells the root namespace the way PBS does; a removed drive no longer shows the root disk's size; a rotated-out token stops at once; the dnsmasq check looks at the right package (burn-down round 2: R-124, R-118, R-269, R-317) (2026-10-05)
|
||||
|
||||
Released by `scripts/release-agent.sh`: binary sha256 `642c4d196c48671c14ff653118303c5903af1670b7abb550edeaf73e701cd5b8`,
|
||||
config bundle sha256 `326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d` (tag `v0.147.0` = `f1b9b41`).
|
||||
Delivery order as for v0.146.1: signed `agent_update`, then signed `agent_config_update`.
|
||||
|
||||
MinAgent impact: none (the controller needs nothing new from this agent). Config bundle content unchanged from v0.146.1.
|
||||
|
||||
|
||||
@@ -1,7 +1,16 @@
|
||||
# REPORT — 2026-10-05 burn-down: two record corrections (no release)
|
||||
# REPORT — agent v0.148.0 (2026-10-06, the burn-down night)
|
||||
|
||||
Full session report: `felhom.eu/REPORT-burndown-2026-10-05.md`. Baseline `e06ed97` (v0.146.1). No binary, bundle or
|
||||
version change; `go build`, `go vet ./internal/backup/`, `agent_gates.py --fast` green.
|
||||
Full session report: `felhom.eu/REPORT-burndown3-2026-10-06.md`. Baseline `208fac8` (v0.147.0).
|
||||
|
||||
- R-291 — `scripts/retention-policy.json`: the source of the number (R-267 newest-10 prune, R-287) and readers corrected.
|
||||
- R-348 — `internal/backup/store.go`: the restart comment says what a restart really blanks.
|
||||
**Released:** v0.148.0 (tag = `861d32a`; binary sha256 `3e68a087…`, bundle `a6fa4f58…`, verified by download). R-349
|
||||
(the host report carries `agent_sha256` — the hub's host pages read „matches vouched” for all three boxes) and R-25's
|
||||
agent half (the format answer carries the verified `fs_uuid`). **Delivered** by signed `agent_update` (all three on
|
||||
0.148.0 by 22:50Z) and `agent_config_update` (root files 0.148.0 by 22:58Z) to demo-hp, demo-felhom, Tester 1. Tester 2:
|
||||
nothing sent (off).
|
||||
|
||||
**On main, unreleased:** R-426 decoys (new `scripts/test_gate_decoys.py`: published against a fake Gitea,
|
||||
release-complete, the shared reuse-refs/instructions/observations).
|
||||
|
||||
**Said plainly:** `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 — `configs/test_felhom_config_bundle.py`
|
||||
read the installer 1.32.0 KEPT names as installer-written files. The v0.148.0 binary was released inside that window;
|
||||
its code is unaffected (a test-only sibling coupling; CI has no Go). Fixed in `b2b82ae`.
|
||||
|
||||
@@ -66,7 +66,7 @@
|
||||
|---|---|---|---|---|
|
||||
| `IntentStore` (`Get/SetEnrolled/SetEjected/SetDecommissioned/OnAbsent`) | internal/storage/intent.go | `OpenIntentStore(path)` | drive intent (4-state self-heal) | Keyed by durable-id only; `OnAbsent` is the ONLY ejected→enrolled path; refuses empty ids |
|
||||
| `GuestBindStore` (`Record/Remove/Guests`) | internal/localapi/guestbindstore.go | `OpenGuestBindStore(path)` | per-guest enrolled binds (F9 re-assert) | Same tmp+rename 0600 pattern as IntentStore |
|
||||
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) <-chan error` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank |
|
||||
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) (*formatJob, <-chan error)` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank; on success `job.FSUUID` = the new filesystem's UUID, read back only when the durable id still resolves to the formatted device (R-25) — read it only after `done` delivers |
|
||||
| `TokenStore.Mint` / `Lookup` | internal/localapi/tokenstore.go | `Mint(vmid) (plaintext, error)` | per-guest local-API tokens | Only the SHA-256 hash persists (fsync'd append log); constant-time compare on lookup; plaintext returned exactly once. Lookup stats the file on EVERY call and reloads BEFORE answering when the append-only log grew (R-269; was reload-on-miss only, v0.63.0 B3, which let a token rotated out by another process keep authorizing as a map hit): the one-shot provisioner mints into the same file the daemon indexes — cross-process coherence both ways without a restart; unchanged size = no re-read. Pinned by `TestTokenStore_RotatedOutTokenRejectedFirst` |
|
||||
| `FileNonceStore.SeenOrRecord` | internal/authz/noncestore.go | `SeenOrRecord(nonce, exp) bool` | durable anti-replay | fsync'd before returning false; prune only after exp |
|
||||
| `Journal` (`Append/Latest/InFlight/AlreadyApplied`) | internal/reconcile/journal.go | `OpenJournal(path)` | op journal + idempotency + crash recovery | `Recover` consumes `InFlight()`; scratch entries special-cased |
|
||||
@@ -121,7 +121,7 @@
|
||||
| Anti-retarget durable-id binding | internal/localapi/wipe_reresolve.go | resolve id → re-derive + exact match → re-inspect expected state → act on RE-RESOLVED device only |
|
||||
| Atomic single-file JSON store | internal/storage/intent.go | `Open*` loads (missing=empty, corrupt=fail-loud), mutex, tmp+rename 0600, idempotent set |
|
||||
| Durable append-only log + index | internal/authz/noncestore.go (`FileNonceStore`) | fsync before returning "new"; replay into index on open; expiry-only compaction |
|
||||
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
|
||||
| Injectable seam funcs on Server | internal/localapi/server.go (`reresolveWipe`, `deviceDurableID`, `boundCheck`, `deviceCheck`, `livenessCheck`, net-verify: `netTrigger`/`netMounted`/`netJournal`/`netReachable`; R-856 `crashGuardStatePath` — GET /host/crash-guard's state file, internal/localapi/crashguard.go) | prod default wired in `NewServer`; tests override — no real /dev, /proc/mounts, journalctl or TCP in tests. **For mount-table predicates prefer the DATA seams `procSelfMountinfo` / `procGuestMountinfo` (internal/localapi/intermediary.go) over `boundCheck`/`livenessCheck`**: pointing them at a captured fixture runs the real parser, the real predicate and the real handler, so the test cannot go hollow the way R-116's did |
|
||||
| `Server.devicePresent` (R-113, v0.114.0) | internal/localapi/disks.go | `devicePresent(rawMountPath) bool`; seam `deviceCheck`, default `isHostMountpoint` | the agent's DEVICE-presence signal — asks whether the drive's RAW mount is still mounted | **Use this, never the bind, to answer "is the drive there".** The raw mount is a device-bound systemd unit and dies with its device; the agent's own bind under the shared parent is NOT device-bound and outlives it as a stale shell. `BoundUnderParent` is now `boundUnderParent(...) && devicePresent(...)` at BOTH /disks construction sites — dropping either half is a regression with its own red-proof. Empty path ⇒ **true** (unknown is never absent: absent stops a customer's apps) |
|
||||
| `bindLiveness` + `BindLiveness` (R-117, v0.117.0) | internal/localapi/intermediary.go | `bindLiveness(stable, raw) BindLiveness`; seam `livenessCheck`; read verdicts ONLY via `.Usable()` | the agent's bind-LIVENESS signal — the third term of `BoundUnderParent` | **`devicePresent` and `boundUnderParent` are both PATH-PRESENCE tests and neither is liveness.** They compare only mountinfo field 5, so both stay true over a bind that names the drive that went away while the raw mount healed onto the returning one (measured: raw 8:32 /dev/sdc, bind 8:16 /dev/sdb `shutdown`, EIO both ways, payload healthy). Two dead states, and a fix needs BOTH checks: devno mismatch (the detach/return case) AND the ext4 abort tokens `shutdown`/`emergency_ro` (the steady-state case, where the devnos AGREE because the device never left). **THREE states, never a bool** — `BindUnknown` must exist and `Usable()` treats it as PRESENT (absent stops a customer's apps). **Order matters:** compare devices first and read the abort flag off the RAW mount in the stale case — abort-first classifies the real return state as aborted and refuses the re-bind that repairs it. **NO BLOCK I/O, ever** (CLAUDE.md rule; a probe on a wedged device survives SIGKILL). 6 red-proofs |
|
||||
| `AttachDrive` repair ruling (R-117, v0.117.0) | internal/localapi/intermediary.go | the `switch bindLiveness(...)` inside the `n == 1 && GuestSeesMount` arm | decides whether the existing self-heal runs | `BindStaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, no guest restart). `BindAborted` ⇒ **quiet no-op** — a re-bind lands on the SAME dead superblock and this runs every 20 s, so re-binding is an infinite silent retry that also masks the state; it must surface via `BoundUnderParent=false`. `BindLive`/`BindUnknown` ⇒ no-op, unchanged. **Do not return an error for the aborted case** — the reconcile loop would log a failure every 20 s |
|
||||
@@ -160,6 +160,7 @@
|
||||
| `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring |
|
||||
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
|
||||
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
|
||||
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
|
||||
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
|
||||
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
|
||||
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
|
||||
|
||||
@@ -0,0 +1,53 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"go/ast"
|
||||
"go/parser"
|
||||
"go/token"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-444: the weekly trim has the guestnet shape (component + reporter seam + goroutine), so its wiring is asserted
|
||||
// from the AST like TestMainWiresGuestNetWatchdog — a unit-green trim job that main.go never starts is the inert-seam
|
||||
// defect. It must also share the ONE heavy-op gate (heavyOps), or it could run beside a backup.
|
||||
func TestMainWiresGuestDiskTrim(t *testing.T) {
|
||||
fset := token.NewFileSet()
|
||||
f, err := parser.ParseFile(fset, "main.go", nil, 0)
|
||||
if err != nil {
|
||||
t.Fatalf("parse main.go: %v", err)
|
||||
}
|
||||
var constructedWithGate, reporterWired, started bool
|
||||
ast.Inspect(f, func(n ast.Node) bool {
|
||||
switch node := n.(type) {
|
||||
case *ast.CallExpr:
|
||||
if fn, ok := node.Fun.(*ast.SelectorExpr); ok {
|
||||
switch fn.Sel.Name {
|
||||
case "New":
|
||||
if pkg, ok := fn.X.(*ast.Ident); ok && pkg.Name == "fstrim" && len(node.Args) >= 3 {
|
||||
if id, ok := node.Args[2].(*ast.Ident); ok && id.Name == "heavyOps" {
|
||||
constructedWithGate = true
|
||||
}
|
||||
}
|
||||
case "SetGuestDiskTrimReporter":
|
||||
reporterWired = true
|
||||
}
|
||||
}
|
||||
case *ast.GoStmt:
|
||||
if sel, ok := node.Call.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "Run" {
|
||||
if id, ok := sel.X.(*ast.Ident); ok && id.Name == "diskTrim" {
|
||||
started = true
|
||||
}
|
||||
}
|
||||
}
|
||||
return true
|
||||
})
|
||||
if !constructedWithGate {
|
||||
t.Error("main.go never calls fstrim.New(..., heavyOps, ...) — no trim job, or one outside the heavy-op gate")
|
||||
}
|
||||
if !reporterWired {
|
||||
t.Error("main.go never calls collector.SetGuestDiskTrimReporter — the guest_disk_trim stanza never reaches the hub")
|
||||
}
|
||||
if !started {
|
||||
t.Error("main.go never starts the trim job with `go diskTrim.Run(ctx)`")
|
||||
}
|
||||
}
|
||||
@@ -38,6 +38,7 @@ import (
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/fasttick"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/fstrim"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/guesthook"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/guestnet"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
@@ -1479,6 +1480,25 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
}
|
||||
go runJanitor(ctx, jd)
|
||||
}
|
||||
// R-444 (`09` §3 decision 139): the weekly guest disk trim — `pct fstrim <vmid>` of every owned, running guest,
|
||||
// Wednesday from 10:00 local, daytime only, under the one-heavy-op gate. Not part of the errc fan-out: a trim job
|
||||
// must never be able to bring the agent down.
|
||||
if cfg.DiskTrim.Enabled() {
|
||||
dtMode := proxmox.RunnerMode(cfg.Privileged.Mode)
|
||||
if dtMode == "" {
|
||||
dtMode = proxmox.RunnerSudo
|
||||
}
|
||||
dtRunner := &proxmox.ExecRunner{Mode: dtMode, SudoPath: cfg.Privileged.SudoPath}
|
||||
dtGuests := localapi.NewStaleLockController(px, dtRunner, reconcile.DefaultPool, logger)
|
||||
if dtGuests != nil {
|
||||
diskTrim := fstrim.New(dtRunner, dtGuests, heavyOps,
|
||||
filepath.Join(cfg.OOB.WithDefaults().StateDir, "guest-disk-trim.json"), logger)
|
||||
collector.SetGuestDiskTrimReporter(diskTrim)
|
||||
go diskTrim.Run(ctx)
|
||||
}
|
||||
} else {
|
||||
logger.Info("fstrim: weekly guest disk trim disabled by config (disk_trim.disable)")
|
||||
}
|
||||
if lanLoop != nil {
|
||||
lanServers = 1
|
||||
go func() { errc <- lanLoop.Run(ctx) }()
|
||||
|
||||
@@ -129,6 +129,14 @@ Cmnd_Alias FELHOM_CONTROLLERSWAP = \
|
||||
Cmnd_Alias FELHOM_STALELOCK = \
|
||||
/usr/sbin/pct ^unlock [0-9]+$
|
||||
|
||||
# Weekly guest disk trim (R-444, operator ruling `09` §3 decision 139). A thin pool only ever grows from blocks the
|
||||
# guest has FREED: `fstrim` inside the unprivileged container is refused (FITRIM: Operation not permitted), so the host
|
||||
# trims the guest's mounts. Measured on demo-hp 2026-10-06: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %,
|
||||
# apps kept answering. ONE exact pattern: a vmid and nothing else — no `--ignore-mountpoints`, no second argument
|
||||
# (pinned: TestSudoersFstrimRuleIsExact). The agent runs it on a weekly daytime timer under the heavy-op gate.
|
||||
Cmnd_Alias FELHOM_FSTRIM = \
|
||||
/usr/sbin/pct ^fstrim [0-9]+$
|
||||
|
||||
# Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS
|
||||
# leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and
|
||||
# a guest joins that pool only when its restore COMPLETES, so a failed restore leaves a pool-less guest
|
||||
@@ -304,4 +312,4 @@ Cmnd_Alias FELHOM_GUESTNET = \
|
||||
/usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \
|
||||
/usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$
|
||||
|
||||
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
|
||||
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_FSTRIM, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
|
||||
|
||||
@@ -546,7 +546,10 @@ class Builder(unittest.TestCase):
|
||||
r"/etc/felhom/[a-z.-]+)", text))
|
||||
agent_writes = {"/usr/local/sbin/felhom-shared-parent", "/etc/systemd/system/felhom-shared-parent.service"}
|
||||
trust = {osapply.TRUST_FILE, osapply.TRUST_SIGNERS, osapply.TRUST_SIGNERS + ".tmp", osapply.BUNDLE_RECORD}
|
||||
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust)
|
||||
# Written by the appliance ISO's first boot (felhom.eu scripts/iso/felhom-bootstrap.sh), never by the installer:
|
||||
# since installer 1.32.0 (R-275) the uninstall only NAMES them under KEPT.
|
||||
iso_writes = {"/etc/felhom/.bootstrap-done", "/etc/felhom/appliance-pairing-code"}
|
||||
missing = sorted(p for p in found if p not in osapply.BUNDLE_DESTS and p not in agent_writes | trust | iso_writes)
|
||||
self.assertEqual(missing, [], "the installer writes these root files, but the bundle does not carry them")
|
||||
# the limits drop-in is named through $AGENT_UNIT in the installer
|
||||
self.assertIn("/etc/systemd/system/felhom-agent.service.d/felhom-agent-limits.conf", osapply.BUNDLE_DESTS)
|
||||
|
||||
@@ -241,3 +241,23 @@ func TestNewestArchiveTime_DistinctPhantomsEachAnnounced(t *testing.T) {
|
||||
t.Errorf("got %d rejection lines for 2 distinct phantoms across 3 polls, want 2:\n%s", n, buf.String())
|
||||
}
|
||||
}
|
||||
|
||||
// R-99 (`09` §3 decision 140): the WARN for a PBS phantom ends with the cleanup runbook, so whoever sees it knows the
|
||||
// one sanctioned way to remove it; a tiny archive on a dir storage is not a PBS phantom and gets no pointer.
|
||||
// RED-PROOF: drop the `msg += phantomCleanupPointer` line → "the PBS phantom WARN does not end with the runbook pointer".
|
||||
func TestRejectedArchiveWarnNamesTheCleanupRunbook(t *testing.T) {
|
||||
var buf bytes.Buffer
|
||||
r := runnerWithContent(t, &buf, []proxmox.StorageContent{phantomEntry(), goodPBSEntry()})
|
||||
if _, _, err := r.NewestArchiveTime(context.Background(), 9201); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
const want = "INCOMPLETE archive when computing tier freshness — it is not a successful backup — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
|
||||
if !strings.Contains(buf.String(), want) {
|
||||
t.Errorf("the PBS phantom WARN does not end with the runbook pointer:\n%s", buf.String())
|
||||
}
|
||||
local := phantomEntry()
|
||||
local.Format, local.VolID = "tar.zst", "local:backup/vzdump-lxc-9201-2026_07_28-05_31_14.tar.zst"
|
||||
if got := rejectedArchiveMessage(local); strings.Contains(got, "pbs-phantom-cleanup") {
|
||||
t.Errorf("a dir-storage archive got the PBS runbook pointer: %s", got)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -518,10 +518,25 @@ func (r *BackupRunner) warnRejectedArchiveOnce(e proxmox.StorageContent, why str
|
||||
if seen {
|
||||
return
|
||||
}
|
||||
r.logger.Warn("backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup",
|
||||
r.logger.Warn(rejectedArchiveMessage(e),
|
||||
"target", r.target, "vmid", e.VMID, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
|
||||
}
|
||||
|
||||
// phantomCleanupPointer names the runbook that removes a PBS phantom (R-99, `09` §3 decision 140: a leftover of an
|
||||
// aborted upload is deleted on the backup server, by a runbook, when one is seen — never automatically).
|
||||
const phantomCleanupPointer = " — a phantom leftover; delete it by felhom.eu documentation/runbooks/pbs-phantom-cleanup.md (09 §3 decision 140)"
|
||||
|
||||
// rejectedArchiveMessage is the WARN text for a rejected archive. Only a PBS entry (format pbs-ct / pbs-vm) gets the
|
||||
// cleanup pointer: the runbook deletes on a PBS datastore, and a tiny archive on a dir storage is not a PBS phantom.
|
||||
// Pinned by TestRejectedArchiveWarnNamesTheCleanupRunbook.
|
||||
func rejectedArchiveMessage(e proxmox.StorageContent) string {
|
||||
msg := "backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup"
|
||||
if strings.HasPrefix(e.Format, "pbs-") {
|
||||
msg += phantomCleanupPointer
|
||||
}
|
||||
return msg
|
||||
}
|
||||
|
||||
// demo-felhom in a single afternoon of deploys (2026-07-26).
|
||||
//
|
||||
// Asking the STORAGE rather than persisting the store is deliberate:
|
||||
|
||||
@@ -151,6 +151,10 @@ var manifest = []Capability{
|
||||
// reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ----
|
||||
{"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""},
|
||||
|
||||
// ---- Weekly guest disk trim (FELHOM_FSTRIM, R-444). NON-critical: a missing grant means the thin pool is not
|
||||
// reclaimed this week (the trim job WARNs per guest and the report shows the failure), not a serving outage. ----
|
||||
{"guest-fstrim", "weekly guest disk trim (thin-pool reclaim, R-444)", "/usr/sbin/pct", []string{"fstrim", "9201"}, false, ""},
|
||||
|
||||
// ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite
|
||||
// backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the
|
||||
// conf install, unit enable/restart and the handshake read gate the backup path. apt-install
|
||||
|
||||
@@ -63,3 +63,33 @@ func TestSudoersRefusesTheR861Injections(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// R-444: the weekly trim's grant is ONE exact shape — `pct fstrim <vmid>` — and nothing smuggled after it.
|
||||
// The manifest entry (guest-fstrim) proves the real call is still allowed (TestManifestCoveredBySudoers); this
|
||||
// pins the other direction. RED-PROOF: write the rule as the glob `/usr/sbin/pct fstrim [0-9]*` → every decoy
|
||||
// below with a trailing argument matches (the glob's `*` eats spaces).
|
||||
func TestSudoersFstrimRuleIsExact(t *testing.T) {
|
||||
data, err := os.ReadFile(sudoersPath)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
entries := parseSudoersEntries(t, string(data))
|
||||
if !matchesAny("/usr/sbin/pct fstrim 9201", entries) {
|
||||
t.Fatal("the sudoers does not allow `pct fstrim 9201` — the weekly trim cannot run")
|
||||
}
|
||||
for _, c := range []string{
|
||||
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints",
|
||||
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints 1",
|
||||
"/usr/sbin/pct fstrim 9201; x",
|
||||
"/usr/sbin/pct fstrim 9201 9202",
|
||||
"/usr/sbin/pct fstrim 92a1",
|
||||
"/usr/sbin/pct fstrim ",
|
||||
"/usr/sbin/pct fstrim -- 9201",
|
||||
"/usr/sbin/pct destroy 9201",
|
||||
"/usr/sbin/pct destroy 9201 --purge",
|
||||
} {
|
||||
if matchesAny(c, entries) {
|
||||
t.Errorf("the sudoers allows a command the trim rule must not: %q", c)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -33,6 +33,7 @@ type Config struct {
|
||||
LANResolver LANResolverConfig `json:"lan_resolver"`
|
||||
WGTunnel WGTunnelConfig `json:"wg_tunnel"`
|
||||
GuestNet GuestNetConfig `json:"guest_net"`
|
||||
DiskTrim DiskTrimConfig `json:"disk_trim"`
|
||||
OOB OOBConfig `json:"oob"`
|
||||
SelfUpdate SelfUpdateConfig `json:"selfupdate"`
|
||||
LogLevel string `json:"log_level"` // debug|info|warn|error (default info)
|
||||
@@ -139,6 +140,16 @@ func (w WGTunnelConfig) WithDefaults() WGTunnelConfig {
|
||||
return w
|
||||
}
|
||||
|
||||
// DiskTrimConfig configures the R-444 weekly guest disk trim (internal/fstrim). DEFAULT-ON, like GuestNetConfig and for
|
||||
// the same reason: it only acts on guests the agent already owns, and the operator ruled every box trims (`09` §3
|
||||
// decision 139). Opting out is the explicit act: `"disk_trim": {"disable": true}`.
|
||||
type DiskTrimConfig struct {
|
||||
Disable bool `json:"disable"`
|
||||
}
|
||||
|
||||
// Enabled reports whether the weekly trim should run.
|
||||
func (d DiskTrimConfig) Enabled() bool { return !d.Disable }
|
||||
|
||||
// GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet).
|
||||
//
|
||||
// **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other
|
||||
|
||||
@@ -0,0 +1,346 @@
|
||||
// Package fstrim is the weekly guest disk trim (R-444, operator ruling `09` §3 decision 139).
|
||||
//
|
||||
// Why: a thin pool only ever grows from blocks a guest has already FREED — `fstrim` inside the unprivileged container
|
||||
// is refused (FITRIM: Operation not permitted), and nothing else on the box gives the blocks back. A full thin pool
|
||||
// takes every guest on the host read-only, so the pool can reach 100 % from deleted data alone. Measured on demo-hp
|
||||
// 2026-10-06 09:14Z: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %, 18/18 app probes 200, max 1.1 s
|
||||
// (audits/ten-answers-2026-10-06/r444-measure.txt).
|
||||
//
|
||||
// The rule, each part pinned by a test in fstrim_test.go:
|
||||
// - Weekly: a guest is DUE from Wednesday 10:00 local until it has been trimmed once since then (a box that was off
|
||||
// on Wednesday catches up at its next eligible hour).
|
||||
// - Daytime only: a trim starts only between 10:00 and 20:59 local — never in the night window (01:00–06:59) where
|
||||
// the backups and the restore-tests run (TestEligibleHourNeverInTheNight).
|
||||
// - Never beside a backup, a restore-test or another heavy operation: the pass holds the host-wide one-heavy-op gate
|
||||
// (backup.InFlight) for its whole run; a busy gate DEFERS the pass to the next hourly tick.
|
||||
// - A failed trim is retried at the next eligible hour, at most MaxAttemptsPerWeek times in one week.
|
||||
// - The last result per guest (time, bytes, ok/fail) is persisted, so a restart neither loses it nor re-trims.
|
||||
//
|
||||
// The command is the ONE exact sudoers shape `pct fstrim <vmid>` (FELHOM_FSTRIM). Only guests from the pool-verified
|
||||
// source (ListLXC ∩ the felhom pool, audit A1) and only RUNNING ones are trimmed.
|
||||
package fstrim
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// Schedule. The weekday/hours are fixed on purpose (one sentence the operator can read on the System page).
|
||||
const (
|
||||
Weekday = time.Wednesday
|
||||
StartHour = 10 // first eligible local hour (inclusive)
|
||||
EndHour = 21 // first NOT-eligible local hour (exclusive): last start is 20:59
|
||||
MaxAttemptsPerWeek = 3
|
||||
// TickInterval is how often the job looks; a deferred or failed pass is therefore retried the next hour.
|
||||
TickInterval = time.Hour
|
||||
// FirstTickDelay lets the agent settle after a start before the first look.
|
||||
FirstTickDelay = 5 * time.Minute
|
||||
// PerGuestTimeout bounds one `pct fstrim` (measured 24.4 s for 84 GiB).
|
||||
PerGuestTimeout = 30 * time.Minute
|
||||
)
|
||||
|
||||
// ScheduleText is the human description carried on the host report.
|
||||
const ScheduleText = "weekly, due Wednesday from 10:00 host-local time; starts only 10:00-20:59; never beside a backup or restore-test"
|
||||
|
||||
// Runner runs a host command (proxmox.ExecRunner in production, through `sudo -n`).
|
||||
type Runner interface {
|
||||
Run(ctx context.Context, name string, args ...string) (stdout, stderr []byte, err error)
|
||||
}
|
||||
|
||||
// GuestSource yields the guests this agent OWNS (the pool-verified source, never a bare ListLXC).
|
||||
type GuestSource interface {
|
||||
Guests(ctx context.Context) ([]proxmox.Guest, error)
|
||||
}
|
||||
|
||||
// Gate is the host-wide one-heavy-operation gate (*backup.InFlight).
|
||||
type Gate interface {
|
||||
TryAcquire(what string) (release func(), busy string, ok bool)
|
||||
}
|
||||
|
||||
// GateName is what the gate reports as busy while a trim runs.
|
||||
const GateName = "guest-fstrim"
|
||||
|
||||
// Record is one guest's last trim attempt, as persisted.
|
||||
type Record struct {
|
||||
LastAttemptAt time.Time `json:"last_attempt_at"`
|
||||
OK bool `json:"ok"`
|
||||
BytesTrimmed int64 `json:"bytes_trimmed"`
|
||||
Mounts int `json:"mounts"`
|
||||
DurationSeconds float64 `json:"duration_seconds"`
|
||||
LastOKAt time.Time `json:"last_ok_at,omitempty"`
|
||||
Error string `json:"error,omitempty"`
|
||||
// Attempts counts the attempts since the current week's due time (reset by the first attempt of a new week).
|
||||
Attempts int `json:"attempts"`
|
||||
}
|
||||
|
||||
// Trimmer is the weekly job.
|
||||
type Trimmer struct {
|
||||
runner Runner
|
||||
guests GuestSource
|
||||
gate Gate
|
||||
statePath string
|
||||
logger *slog.Logger
|
||||
loc *time.Location
|
||||
now func() time.Time
|
||||
|
||||
mu sync.Mutex
|
||||
records map[int]Record
|
||||
}
|
||||
|
||||
// New builds the job and loads the persisted state. A missing state file is an empty state; a corrupt one is logged
|
||||
// and treated as empty (the cost is one extra trim, never a missed one).
|
||||
func New(runner Runner, guests GuestSource, gate Gate, statePath string, logger *slog.Logger) *Trimmer {
|
||||
if logger == nil {
|
||||
logger = slog.Default()
|
||||
}
|
||||
t := &Trimmer{runner: runner, guests: guests, gate: gate, statePath: statePath, logger: logger,
|
||||
loc: time.Local, now: time.Now, records: map[int]Record{}}
|
||||
t.load()
|
||||
return t
|
||||
}
|
||||
|
||||
func (t *Trimmer) load() {
|
||||
data, err := os.ReadFile(t.statePath)
|
||||
if err != nil {
|
||||
if !os.IsNotExist(err) {
|
||||
t.logger.Warn("fstrim: state read failed — starting empty", "path", t.statePath, "err", err)
|
||||
}
|
||||
return
|
||||
}
|
||||
var raw map[string]Record
|
||||
if err := json.Unmarshal(data, &raw); err != nil {
|
||||
t.logger.Warn("fstrim: state file corrupt — starting empty", "path", t.statePath, "err", err)
|
||||
return
|
||||
}
|
||||
for k, r := range raw {
|
||||
if id, err := strconv.Atoi(k); err == nil && id > 0 {
|
||||
t.records[id] = r
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func (t *Trimmer) saveLocked() error {
|
||||
raw := make(map[string]Record, len(t.records))
|
||||
for id, r := range t.records {
|
||||
raw[strconv.Itoa(id)] = r
|
||||
}
|
||||
data, err := json.MarshalIndent(raw, "", " ")
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := os.MkdirAll(filepath.Dir(t.statePath), 0o755); err != nil {
|
||||
return err
|
||||
}
|
||||
tmp := t.statePath + ".tmp"
|
||||
if err := os.WriteFile(tmp, data, 0o600); err != nil {
|
||||
os.Remove(tmp)
|
||||
return err
|
||||
}
|
||||
return os.Rename(tmp, t.statePath)
|
||||
}
|
||||
|
||||
// EligibleHour reports whether a trim may START at local time lt.
|
||||
func EligibleHour(lt time.Time) bool {
|
||||
h := lt.Hour()
|
||||
return h >= StartHour && h < EndHour
|
||||
}
|
||||
|
||||
// weekAnchor is the most recent Wednesday StartHour:00 at or before lt (same location as lt).
|
||||
func weekAnchor(lt time.Time) time.Time {
|
||||
daysBack := (int(lt.Weekday()) - int(Weekday) + 7) % 7
|
||||
d := lt.AddDate(0, 0, -daysBack)
|
||||
a := time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
|
||||
if a.After(lt) {
|
||||
d = d.AddDate(0, 0, -7)
|
||||
a = time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
|
||||
}
|
||||
return a
|
||||
}
|
||||
|
||||
// due reports whether a guest with record r (ok=false: none) is due at local time lt.
|
||||
func due(r Record, has bool, lt time.Time) bool {
|
||||
if !has {
|
||||
return true
|
||||
}
|
||||
anchor := weekAnchor(lt)
|
||||
if r.LastAttemptAt.Before(anchor) {
|
||||
return true // not tried this week
|
||||
}
|
||||
return !r.OK && r.Attempts < MaxAttemptsPerWeek
|
||||
}
|
||||
|
||||
// Run looks every TickInterval until ctx ends. It never returns an error: a failed trim is a reported fact.
|
||||
func (t *Trimmer) Run(ctx context.Context) {
|
||||
t.logger.Info("fstrim: weekly guest disk trim starting", "schedule", ScheduleText)
|
||||
timer := time.NewTimer(FirstTickDelay)
|
||||
defer timer.Stop()
|
||||
for {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-timer.C:
|
||||
t.Pass(ctx)
|
||||
timer.Reset(TickInterval)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Pass is one look: outside the daytime window it does nothing; otherwise it trims every due, running, owned guest
|
||||
// while holding the heavy-op gate.
|
||||
func (t *Trimmer) Pass(ctx context.Context) {
|
||||
lt := t.now().In(t.loc)
|
||||
if !EligibleHour(lt) {
|
||||
t.logger.Debug("fstrim: outside the daytime window — not looking", "local", lt.Format("Mon 15:04"))
|
||||
return
|
||||
}
|
||||
guests, err := t.guests.Guests(ctx)
|
||||
if err != nil {
|
||||
t.logger.Warn("fstrim: owned-guest list unavailable — skipping this pass", "err", err)
|
||||
return
|
||||
}
|
||||
owned := make(map[int]bool, len(guests))
|
||||
var todo []int
|
||||
t.mu.Lock()
|
||||
for _, g := range guests {
|
||||
owned[g.VMID] = true
|
||||
r, has := t.records[g.VMID]
|
||||
if !due(r, has, lt) {
|
||||
continue
|
||||
}
|
||||
if g.Status != "running" {
|
||||
t.logger.Info("fstrim: guest not running — trimmed when it runs", "vmid", g.VMID, "status", g.Status)
|
||||
continue
|
||||
}
|
||||
todo = append(todo, g.VMID)
|
||||
}
|
||||
// A guest the agent no longer owns has no result to report.
|
||||
pruned := false
|
||||
for id := range t.records {
|
||||
if !owned[id] {
|
||||
delete(t.records, id)
|
||||
pruned = true
|
||||
}
|
||||
}
|
||||
if pruned {
|
||||
if err := t.saveLocked(); err != nil {
|
||||
t.logger.Warn("fstrim: state save failed", "err", err)
|
||||
}
|
||||
}
|
||||
t.mu.Unlock()
|
||||
if len(todo) == 0 {
|
||||
return
|
||||
}
|
||||
sort.Ints(todo)
|
||||
release, busy, ok := t.gate.TryAcquire(GateName)
|
||||
if !ok {
|
||||
t.logger.Info("fstrim: deferred — a heavy operation is in flight; retrying next hour", "busy", busy, "due_guests", len(todo))
|
||||
return
|
||||
}
|
||||
defer release()
|
||||
for _, vmid := range todo {
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
t.trimOne(ctx, vmid, lt)
|
||||
}
|
||||
}
|
||||
|
||||
var trimmedLine = regexp.MustCompile(`\((\d+) bytes\) trimmed`)
|
||||
|
||||
// ParseTrimmed sums the "(N bytes) trimmed" lines of `pct fstrim` output and counts them (one per mount point), e.g.
|
||||
// `/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed`.
|
||||
func ParseTrimmed(out string) (bytes int64, mounts int) {
|
||||
for _, m := range trimmedLine.FindAllStringSubmatch(out, -1) {
|
||||
n, err := strconv.ParseInt(m[1], 10, 64)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
bytes += n
|
||||
mounts++
|
||||
}
|
||||
return bytes, mounts
|
||||
}
|
||||
|
||||
// GiB renders bytes as "30.1 GiB".
|
||||
func GiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
|
||||
|
||||
func (t *Trimmer) trimOne(ctx context.Context, vmid int, lt time.Time) {
|
||||
start := t.now()
|
||||
cctx, cancel := context.WithTimeout(ctx, PerGuestTimeout)
|
||||
stdout, stderr, err := t.runner.Run(cctx, "pct", "fstrim", strconv.Itoa(vmid))
|
||||
cancel()
|
||||
dur := t.now().Sub(start)
|
||||
bytes, mounts := ParseTrimmed(string(stdout) + "\n" + string(stderr))
|
||||
|
||||
t.mu.Lock()
|
||||
prev, has := t.records[vmid]
|
||||
r := Record{LastAttemptAt: start.UTC(), OK: err == nil, BytesTrimmed: bytes, Mounts: mounts,
|
||||
DurationSeconds: float64(dur.Round(100*time.Millisecond)) / float64(time.Second), LastOKAt: prev.LastOKAt}
|
||||
if has && !prev.LastAttemptAt.Before(weekAnchor(lt)) {
|
||||
r.Attempts = prev.Attempts + 1
|
||||
} else {
|
||||
r.Attempts = 1
|
||||
}
|
||||
if err == nil {
|
||||
r.LastOKAt = start.UTC()
|
||||
} else {
|
||||
msg := strings.TrimSpace(err.Error() + ": " + strings.TrimSpace(string(stderr)))
|
||||
if len(msg) > 300 {
|
||||
msg = msg[:300]
|
||||
}
|
||||
r.Error = msg
|
||||
}
|
||||
t.records[vmid] = r
|
||||
saveErr := t.saveLocked()
|
||||
t.mu.Unlock()
|
||||
|
||||
if err == nil {
|
||||
t.logger.Info(fmt.Sprintf("fstrim: guest %d trimmed %s in %.1fs", vmid, GiB(bytes), r.DurationSeconds),
|
||||
"vmid", vmid, "bytes_trimmed", bytes, "mounts", mounts, "duration_s", r.DurationSeconds)
|
||||
if mounts == 0 {
|
||||
t.logger.Warn("fstrim: pct fstrim succeeded but reported no trimmed mount — output not understood",
|
||||
"vmid", vmid, "stdout", strings.TrimSpace(string(stdout)))
|
||||
}
|
||||
} else {
|
||||
t.logger.Warn(fmt.Sprintf("fstrim: guest %d trim FAILED after %.1fs", vmid, r.DurationSeconds),
|
||||
"vmid", vmid, "attempt", r.Attempts, "max_attempts_per_week", MaxAttemptsPerWeek, "err", r.Error)
|
||||
}
|
||||
if saveErr != nil {
|
||||
t.logger.Warn("fstrim: state save failed — the result will not survive a restart", "path", t.statePath, "err", saveErr)
|
||||
}
|
||||
}
|
||||
|
||||
// GuestDiskTrimStatus implements hub.GuestDiskTrimReporter: a pure read of the persisted results (never runs pct).
|
||||
func (t *Trimmer) GuestDiskTrimStatus(context.Context) *hub.GuestDiskTrimStatus {
|
||||
t.mu.Lock()
|
||||
defer t.mu.Unlock()
|
||||
out := &hub.GuestDiskTrimStatus{Schedule: ScheduleText}
|
||||
ids := make([]int, 0, len(t.records))
|
||||
for id := range t.records {
|
||||
ids = append(ids, id)
|
||||
}
|
||||
sort.Ints(ids)
|
||||
for _, id := range ids {
|
||||
r := t.records[id]
|
||||
g := hub.GuestDiskTrim{VMID: id, LastAttemptAt: r.LastAttemptAt.UTC().Format(time.RFC3339), OK: r.OK,
|
||||
BytesTrimmed: r.BytesTrimmed, Mounts: r.Mounts, DurationSeconds: r.DurationSeconds, Error: r.Error}
|
||||
if !r.LastOKAt.IsZero() {
|
||||
g.LastOKAt = r.LastOKAt.UTC().Format(time.RFC3339)
|
||||
}
|
||||
out.Guests = append(out.Guests, g)
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,267 @@
|
||||
package fstrim
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"log/slog"
|
||||
"path/filepath"
|
||||
"reflect"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// The real `pct fstrim 9201` output measured on demo-hp 2026-10-06 (audits/ten-answers-2026-10-06/r444-measure.txt).
|
||||
const measuredOut = "/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed\n" +
|
||||
"/var/lib/lxc/9201/rootfs/var/lib/felhom: 53.9 GiB (57865633792 bytes) trimmed\n"
|
||||
|
||||
const measuredBytes = int64(32277680128 + 57865633792)
|
||||
|
||||
type fakeRunner struct {
|
||||
mu sync.Mutex
|
||||
calls [][]string
|
||||
out string
|
||||
err error
|
||||
onRun func()
|
||||
}
|
||||
|
||||
func (f *fakeRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
|
||||
f.mu.Lock()
|
||||
f.calls = append(f.calls, append([]string{name}, args...))
|
||||
f.mu.Unlock()
|
||||
if f.onRun != nil {
|
||||
f.onRun()
|
||||
}
|
||||
if f.err != nil {
|
||||
return nil, []byte("mount busy"), f.err
|
||||
}
|
||||
return []byte(f.out), nil, nil
|
||||
}
|
||||
|
||||
type fakeGuests struct {
|
||||
g []proxmox.Guest
|
||||
err error
|
||||
}
|
||||
|
||||
func (f fakeGuests) Guests(context.Context) ([]proxmox.Guest, error) { return f.g, f.err }
|
||||
|
||||
// A Wednesday 10:30 in a fixed zone (CEST-like), so the tests do not depend on the machine's zone.
|
||||
var zone = time.FixedZone("CEST", 2*3600)
|
||||
|
||||
func at(day, hour, min int) time.Time { return time.Date(2026, 10, day, hour, min, 0, 0, zone) } // 2026-10-07 = Wednesday
|
||||
|
||||
func newT(t *testing.T, r Runner, g GuestSource, gate Gate, now *time.Time) (*Trimmer, *bytes.Buffer, string) {
|
||||
t.Helper()
|
||||
var logs bytes.Buffer
|
||||
path := filepath.Join(t.TempDir(), "guest-disk-trim.json")
|
||||
tr := New(r, g, gate, path, slog.New(slog.NewTextHandler(&logs, &slog.HandlerOptions{Level: slog.LevelDebug})))
|
||||
tr.loc = zone
|
||||
tr.now = func() time.Time { return *now }
|
||||
return tr, &logs, path
|
||||
}
|
||||
|
||||
func running(ids ...int) fakeGuests {
|
||||
var g []proxmox.Guest
|
||||
for _, id := range ids {
|
||||
g = append(g, proxmox.Guest{VMID: id, Status: "running", Type: "lxc"})
|
||||
}
|
||||
return fakeGuests{g: g}
|
||||
}
|
||||
|
||||
func TestParseTrimmedTheMeasuredOutput(t *testing.T) {
|
||||
b, m := ParseTrimmed(measuredOut)
|
||||
if b != measuredBytes || m != 2 {
|
||||
t.Fatalf("ParseTrimmed = %d bytes over %d mounts, want %d over 2", b, m, measuredBytes)
|
||||
}
|
||||
if b, m := ParseTrimmed("something else\n"); b != 0 || m != 0 {
|
||||
t.Fatalf("unrelated output parsed as %d/%d", b, m)
|
||||
}
|
||||
if got := GiB(measuredBytes); got != "84.0 GiB" {
|
||||
t.Fatalf("GiB = %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// The night window (01:00–06:59) must never be eligible, and the daytime window is exactly 10:00–20:59.
|
||||
func TestEligibleHourNeverInTheNight(t *testing.T) {
|
||||
for h := 0; h < 24; h++ {
|
||||
lt := time.Date(2026, 10, 7, h, 30, 0, 0, zone)
|
||||
got := EligibleHour(lt)
|
||||
if h >= 1 && h <= 6 && got {
|
||||
t.Errorf("hour %02d is in the night window and must not be eligible", h)
|
||||
}
|
||||
if want := h >= 10 && h <= 20; got != want {
|
||||
t.Errorf("EligibleHour(%02d:30) = %v, want %v", h, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestWeekAnchorIsTheLastWednesdayTen(t *testing.T) {
|
||||
cases := map[time.Time]time.Time{
|
||||
at(7, 10, 0): at(7, 10, 0), // Wednesday 10:00 itself
|
||||
at(7, 9, 59): time.Date(2026, 9, 30, 10, 0, 0, 0, zone), // before 10:00 Wednesday → the previous week
|
||||
at(8, 15, 0): at(7, 10, 0), // Thursday
|
||||
at(13, 20, 0): at(7, 10, 0), // next Tuesday
|
||||
at(14, 11, 0): at(14, 10, 0), // next Wednesday
|
||||
}
|
||||
for in, want := range cases {
|
||||
if got := weekAnchor(in); !got.Equal(want) {
|
||||
t.Errorf("weekAnchor(%s) = %s, want %s", in.Format("Mon 01-02 15:04"), got.Format("Mon 01-02 15:04"), want.Format("Mon 01-02 15:04"))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The consequence: on Wednesday 10:30 a running owned guest is trimmed with the ONE exact argv, the bytes are parsed,
|
||||
// the positive log line is written, the result is persisted, and the host report carries it.
|
||||
func TestPassTrimsADueGuestAndReportsIt(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
r := &fakeRunner{out: measuredOut}
|
||||
tr, logs, path := newT(t, r, running(9201), &backup.InFlight{}, &now)
|
||||
tr.Pass(context.Background())
|
||||
|
||||
if want := [][]string{{"pct", "fstrim", "9201"}}; !reflect.DeepEqual(r.calls, want) {
|
||||
t.Fatalf("runner calls = %q, want %q", r.calls, want)
|
||||
}
|
||||
if !strings.Contains(logs.String(), "fstrim: guest 9201 trimmed 84.0 GiB in ") {
|
||||
t.Fatalf("no positive per-guest log line:\n%s", logs.String())
|
||||
}
|
||||
st := tr.GuestDiskTrimStatus(context.Background())
|
||||
if st == nil || st.Schedule != ScheduleText || len(st.Guests) != 1 {
|
||||
t.Fatalf("report stanza = %+v", st)
|
||||
}
|
||||
g := st.Guests[0]
|
||||
if g.VMID != 9201 || !g.OK || g.BytesTrimmed != measuredBytes || g.Mounts != 2 || g.LastOKAt == "" || g.LastAttemptAt == "" {
|
||||
t.Fatalf("report guest = %+v", g)
|
||||
}
|
||||
// Persisted: a NEW Trimmer over the same file (an agent restart) still has it and does not trim again this week.
|
||||
now = at(8, 11, 0)
|
||||
r2 := &fakeRunner{out: measuredOut}
|
||||
tr2 := New(r2, running(9201), &backup.InFlight{}, path, slog.New(slog.NewTextHandler(&bytes.Buffer{}, nil)))
|
||||
tr2.loc, tr2.now = zone, func() time.Time { return now }
|
||||
if st2 := tr2.GuestDiskTrimStatus(context.Background()); len(st2.Guests) != 1 || st2.Guests[0].BytesTrimmed != measuredBytes {
|
||||
t.Fatalf("result lost over a restart: %+v", st2)
|
||||
}
|
||||
tr2.Pass(context.Background())
|
||||
if len(r2.calls) != 0 {
|
||||
t.Fatalf("trimmed again in the same week after a restart: %q", r2.calls)
|
||||
}
|
||||
// Next week it is due again.
|
||||
now = at(14, 10, 5)
|
||||
tr2.Pass(context.Background())
|
||||
if len(r2.calls) != 1 {
|
||||
t.Fatalf("not trimmed in the next week: %q", r2.calls)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPassNeverRunsInTheNight(t *testing.T) {
|
||||
for _, h := range []int{1, 3, 6, 9, 21, 23} {
|
||||
now := at(7, h, 15)
|
||||
r := &fakeRunner{out: measuredOut}
|
||||
tr, _, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
|
||||
tr.Pass(context.Background())
|
||||
if len(r.calls) != 0 {
|
||||
t.Errorf("trimmed at %02d:15: %q", h, r.calls)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A backup (or restore-test) holding the heavy-op gate DEFERS the trim; the next hour, gate free, it runs. And while
|
||||
// a trim runs, the gate is held, so a backup cannot start beside it.
|
||||
func TestPassDefersToAHeavyOperationAndRetriesNextHour(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
gate := &backup.InFlight{}
|
||||
release, _, _ := gate.TryAcquire("backup:9201")
|
||||
var busyDuringTrim string
|
||||
r := &fakeRunner{out: measuredOut}
|
||||
r.onRun = func() { busyDuringTrim = gate.Busy() }
|
||||
tr, logs, _ := newT(t, r, running(9201), gate, &now)
|
||||
|
||||
tr.Pass(context.Background())
|
||||
if len(r.calls) != 0 {
|
||||
t.Fatalf("trimmed beside a running backup: %q", r.calls)
|
||||
}
|
||||
if !strings.Contains(logs.String(), "fstrim: deferred") || !strings.Contains(logs.String(), "backup:9201") {
|
||||
t.Fatalf("the deferral is not logged with what holds the gate:\n%s", logs.String())
|
||||
}
|
||||
release()
|
||||
now = now.Add(time.Hour)
|
||||
tr.Pass(context.Background())
|
||||
if len(r.calls) != 1 {
|
||||
t.Fatalf("not retried the next hour: %q", r.calls)
|
||||
}
|
||||
if busyDuringTrim != GateName {
|
||||
t.Fatalf("the heavy-op gate was %q during the trim, want %q", busyDuringTrim, GateName)
|
||||
}
|
||||
if gate.Busy() != "" {
|
||||
t.Fatalf("the gate was not released after the pass: %q", gate.Busy())
|
||||
}
|
||||
}
|
||||
|
||||
func TestFailedTrimWarnsIsRecordedAndRetriedAtMostThreeTimes(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
r := &fakeRunner{err: errors.New("exit status 255")}
|
||||
tr, logs, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
|
||||
for i := 0; i < 6; i++ {
|
||||
tr.Pass(context.Background())
|
||||
now = now.Add(time.Hour)
|
||||
}
|
||||
if len(r.calls) != MaxAttemptsPerWeek {
|
||||
t.Fatalf("attempts in one week = %d, want %d", len(r.calls), MaxAttemptsPerWeek)
|
||||
}
|
||||
if !strings.Contains(logs.String(), "level=WARN") || !strings.Contains(logs.String(), "fstrim: guest 9201 trim FAILED") {
|
||||
t.Fatalf("no WARN for the failure:\n%s", logs.String())
|
||||
}
|
||||
g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]
|
||||
if g.OK || g.LastOKAt != "" || !strings.Contains(g.Error, "exit status 255") || !strings.Contains(g.Error, "mount busy") {
|
||||
t.Fatalf("failed result not recorded as a failure: %+v", g)
|
||||
}
|
||||
// A success later keeps a clean record.
|
||||
r.err = nil
|
||||
r.out = measuredOut
|
||||
now = at(14, 10, 10)
|
||||
tr.Pass(context.Background())
|
||||
if g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]; !g.OK || g.Error != "" || g.BytesTrimmed != measuredBytes {
|
||||
t.Fatalf("success after failure: %+v", g)
|
||||
}
|
||||
}
|
||||
|
||||
func TestOnlyRunningOwnedGuestsAndAFailedListActsOnNothing(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
r := &fakeRunner{out: measuredOut}
|
||||
g := fakeGuests{g: []proxmox.Guest{{VMID: 9201, Status: "stopped"}, {VMID: 9202, Status: "running"}}}
|
||||
tr, _, _ := newT(t, r, g, &backup.InFlight{}, &now)
|
||||
tr.Pass(context.Background())
|
||||
if want := [][]string{{"pct", "fstrim", "9202"}}; !reflect.DeepEqual(r.calls, want) {
|
||||
t.Fatalf("calls = %q, want only the running guest", r.calls)
|
||||
}
|
||||
r2 := &fakeRunner{out: measuredOut}
|
||||
tr2, logs, _ := newT(t, r2, fakeGuests{err: errors.New("pool read 403")}, &backup.InFlight{}, &now)
|
||||
tr2.Pass(context.Background())
|
||||
if len(r2.calls) != 0 || !strings.Contains(logs.String(), "owned-guest list unavailable") {
|
||||
t.Fatalf("a failed ownership read must act on nothing: calls %q", r2.calls)
|
||||
}
|
||||
}
|
||||
|
||||
func TestReportJSONShape(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
tr, _, _ := newT(t, &fakeRunner{out: measuredOut}, running(9201), &backup.InFlight{}, &now)
|
||||
tr.Pass(context.Background())
|
||||
b, err := json.Marshal(tr.GuestDiskTrimStatus(context.Background()))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, k := range []string{`"schedule":`, `"guests":[{"vmid":9201`, `"last_attempt_at":"2026-10-07T08:30:00Z"`, `"ok":true`,
|
||||
`"bytes_trimmed":90143313920`, `"mounts":2`, `"duration_seconds":`, `"last_ok_at":"2026-10-07T08:30:00Z"`} {
|
||||
if !strings.Contains(string(b), k) {
|
||||
t.Errorf("report JSON lacks %s: %s", k, b)
|
||||
}
|
||||
}
|
||||
if strings.Contains(string(b), `"error"`) {
|
||||
t.Errorf("an ok result must omit error: %s", b)
|
||||
}
|
||||
}
|
||||
+43
-2
@@ -10,6 +10,7 @@ import (
|
||||
"log/slog"
|
||||
"os"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
|
||||
@@ -89,6 +90,12 @@ type GuestNetReporter interface {
|
||||
GuestNetStatus(ctx context.Context) *GuestNetStatus
|
||||
}
|
||||
|
||||
// GuestDiskTrimReporter is the R-444 seam the weekly trim job plugs into (same consumer-side pattern — hub does not
|
||||
// import fstrim). nil (feature not wired) → no guest_disk_trim stanza.
|
||||
type GuestDiskTrimReporter interface {
|
||||
GuestDiskTrimStatus(ctx context.Context) *GuestDiskTrimStatus
|
||||
}
|
||||
|
||||
// Collector builds a HostReport from read-only sources. All deps are behind narrow
|
||||
// interfaces for unit testing.
|
||||
type Collector struct {
|
||||
@@ -107,6 +114,7 @@ type Collector struct {
|
||||
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
|
||||
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
|
||||
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
|
||||
diskTrim GuestDiskTrimReporter // R-444: weekly guest disk trim (nil → stanza omitted)
|
||||
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
|
||||
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
|
||||
oob OOBReporter // H1: operator-access health (nil → stanza omitted)
|
||||
@@ -115,6 +123,7 @@ type Collector struct {
|
||||
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
|
||||
hostID string
|
||||
agentVersion string
|
||||
selfSHA func() string // R-349: sha256 of the running binary; default runningBinarySHA256
|
||||
logger *slog.Logger
|
||||
now func() time.Time
|
||||
}
|
||||
@@ -135,6 +144,7 @@ func NewCollector(px proxmoxReader, cf CloudflaredProber, storage StorageObserve
|
||||
temp: SysfsTempReader{}, // slice 9: real sysfs reader by default; tests inject a fake
|
||||
hostID: hostID,
|
||||
agentVersion: agentVersion,
|
||||
selfSHA: runningBinarySHA256,
|
||||
logger: logger,
|
||||
now: func() time.Time { return time.Now().UTC() },
|
||||
}
|
||||
@@ -218,6 +228,12 @@ func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
|
||||
return c
|
||||
}
|
||||
|
||||
// SetGuestDiskTrimReporter wires the R-444 weekly trim job as a report source (nil-safe → stanza omitted).
|
||||
func (c *Collector) SetGuestDiskTrimReporter(r GuestDiskTrimReporter) *Collector {
|
||||
c.diskTrim = r
|
||||
return c
|
||||
}
|
||||
|
||||
// SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side
|
||||
// pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report.
|
||||
type SelfUpdateReporter interface {
|
||||
@@ -337,6 +353,7 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
|
||||
HostID: c.hostID,
|
||||
ReportedAt: c.now().Format(time.RFC3339),
|
||||
AgentVersion: c.agentVersion,
|
||||
AgentSHA256: c.agentSHA256(),
|
||||
Host: host,
|
||||
Guests: c.collectGuests(ctx),
|
||||
// storage_targets populated this slice (slice 5) via the observer; the rest stay
|
||||
@@ -373,6 +390,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
|
||||
if c.guestNet != nil {
|
||||
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
|
||||
}
|
||||
// R-444: the last weekly trim result per guest (nil reporter = not wired → stanza omitted).
|
||||
if c.diskTrim != nil {
|
||||
report.GuestDiskTrim = c.diskTrim.GuestDiskTrimStatus(ctx)
|
||||
}
|
||||
// D1: agent self-update pending status (nil reporter → pending=false, the steady state).
|
||||
if c.selfUpdate != nil {
|
||||
report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending()
|
||||
@@ -431,8 +452,28 @@ const pbsWrapperPath = "/usr/local/sbin/felhom-pbs-apply"
|
||||
// unreadable file yields "", which the hub reads as UNKNOWN rather than as drift — a host that
|
||||
// legitimately has no DR wrapper must not light up amber. The file is 0755, so no privilege is
|
||||
// needed to read it.
|
||||
func pbsWrapperSHA256() string {
|
||||
f, err := os.Open(pbsWrapperPath)
|
||||
func pbsWrapperSHA256() string { return fileSHA256(pbsWrapperPath) }
|
||||
|
||||
// selfExePath is the running binary as the kernel holds it. /proc/self/exe, not the installed path:
|
||||
// after an A/B flip the file at /usr/local/bin/felhom-agent may already be the NEXT binary while this
|
||||
// process still runs the old one, and the report must describe what runs (R-349). Test seam.
|
||||
var selfExePath = "/proc/self/exe"
|
||||
|
||||
// runningBinarySHA256 hashes the running binary ONCE per process — the bytes cannot change under a
|
||||
// running process, and re-hashing ~20 MB every report cycle buys nothing. A failed read is cached as
|
||||
// "" (UNKNOWN); it never fails the report.
|
||||
var runningBinarySHA256 = sync.OnceValue(func() string { return fileSHA256(selfExePath) })
|
||||
|
||||
func (c *Collector) agentSHA256() string {
|
||||
if c.selfSHA == nil {
|
||||
return ""
|
||||
}
|
||||
return c.selfSHA()
|
||||
}
|
||||
|
||||
// fileSHA256 is the hex sha256 of a file's bytes, or "" when it cannot be read.
|
||||
func fileSHA256(path string) string {
|
||||
f, err := os.Open(path)
|
||||
if err != nil {
|
||||
return ""
|
||||
}
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
package hub
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-349: the report carries the sha256 of the binary that is RUNNING, so the hub can tell a
|
||||
// hand-built proof binary from the vouched artifact of the same version string. The consequence
|
||||
// asserted: the wire field equals the hash of this very test binary's bytes (read independently via
|
||||
// os.Executable, a different channel from /proc/self/exe), and it is on the wire as agent_sha256.
|
||||
func TestCollect_AgentSHA256IsTheRunningBinary(t *testing.T) {
|
||||
exe, err := os.Executable()
|
||||
if err != nil {
|
||||
t.Skipf("os.Executable: %v", err)
|
||||
}
|
||||
raw, err := os.ReadFile(exe)
|
||||
if err != nil {
|
||||
t.Fatalf("read own binary: %v", err)
|
||||
}
|
||||
sum := sha256.Sum256(raw)
|
||||
want := hex.EncodeToString(sum[:])
|
||||
|
||||
px := &fakePx{node: "n", ns: newTestNodeStatus()}
|
||||
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
|
||||
r, err := c.Collect(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("Collect: %v", err)
|
||||
}
|
||||
if r.AgentSHA256 != want {
|
||||
t.Fatalf("agent_sha256 = %q, want the running binary's %q", r.AgentSHA256, want)
|
||||
}
|
||||
b, err := json.Marshal(r)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(string(b), `"agent_sha256":"`+want+`"`) {
|
||||
t.Fatalf("agent_sha256 not on the wire: %s", b)
|
||||
}
|
||||
}
|
||||
|
||||
// An unreadable binary is UNKNOWN (empty, omitted) — never a made-up hash, never a failed report.
|
||||
func TestFileSHA256_UnreadableIsEmpty(t *testing.T) {
|
||||
if got := fileSHA256(filepath.Join(t.TempDir(), "absent")); got != "" {
|
||||
t.Fatalf("absent file hashed to %q, want empty", got)
|
||||
}
|
||||
px := &fakePx{node: "n", ns: newTestNodeStatus()}
|
||||
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
|
||||
c.selfSHA = func() string { return "" }
|
||||
r, err := c.Collect(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("Collect must not fail on an unreadable binary: %v", err)
|
||||
}
|
||||
b, _ := json.Marshal(r)
|
||||
if strings.Contains(string(b), "agent_sha256") {
|
||||
t.Fatalf("empty agent_sha256 must be omitted: %s", b)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
package hub
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-444: the guest_disk_trim stanza must reach a report built through the PRODUCTION collect path, be absent from
|
||||
// the wire when the job is not wired, and carry the keys the hub's System page reads.
|
||||
|
||||
type fakeDiskTrim struct{ st *GuestDiskTrimStatus }
|
||||
|
||||
func (f fakeDiskTrim) GuestDiskTrimStatus(context.Context) *GuestDiskTrimStatus { return f.st }
|
||||
|
||||
func TestCollect_GuestDiskTrim(t *testing.T) {
|
||||
px := &fakePx{node: "n", ns: newTestNodeStatus()}
|
||||
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.150.0", quietLogger())
|
||||
r, err := c.Collect(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("Collect: %v", err)
|
||||
}
|
||||
b, _ := json.Marshal(r)
|
||||
var m map[string]any
|
||||
_ = json.Unmarshal(b, &m)
|
||||
if _, ok := m["guest_disk_trim"]; ok {
|
||||
t.Fatalf("guest_disk_trim on the wire with no reporter wired: %s", b)
|
||||
}
|
||||
|
||||
c.SetGuestDiskTrimReporter(fakeDiskTrim{st: &GuestDiskTrimStatus{Schedule: "weekly", Guests: []GuestDiskTrim{{
|
||||
VMID: 9201, LastAttemptAt: "2026-10-07T08:30:00Z", OK: true, BytesTrimmed: 90143313920, Mounts: 2,
|
||||
DurationSeconds: 24.4, LastOKAt: "2026-10-07T08:30:00Z",
|
||||
}}}})
|
||||
r, err = c.Collect(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("Collect: %v", err)
|
||||
}
|
||||
b, _ = json.Marshal(r)
|
||||
m = nil
|
||||
_ = json.Unmarshal(b, &m)
|
||||
dt, ok := m["guest_disk_trim"].(map[string]any)
|
||||
if !ok || dt["schedule"] != "weekly" {
|
||||
t.Fatalf("guest_disk_trim missing or wrong on the wire: %s", b)
|
||||
}
|
||||
g := dt["guests"].([]any)[0].(map[string]any)
|
||||
for _, k := range []string{"vmid", "last_attempt_at", "ok", "bytes_trimmed", "mounts", "duration_seconds", "last_ok_at"} {
|
||||
if _, ok := g[k]; !ok {
|
||||
t.Fatalf("guest_disk_trim.guests[0] lacks %q: %v", k, g)
|
||||
}
|
||||
}
|
||||
if g["bytes_trimmed"] != float64(90143313920) || g["ok"] != true {
|
||||
t.Fatalf("values did not survive the round trip: %v", g)
|
||||
}
|
||||
}
|
||||
@@ -18,6 +18,16 @@ type HostReport struct {
|
||||
HostID string `json:"host_id"` // echoes config.Hub.HostID
|
||||
ReportedAt string `json:"reported_at"` // RFC3339, agent clock
|
||||
AgentVersion string `json:"agent_version"`
|
||||
// AgentSHA256 is the sha256 of the binary this process is RUNNING (read through /proc/self/exe,
|
||||
// once per process), R-349. The version string cannot tell a hand-built proof binary from the
|
||||
// published, vouched artifact of the same version — same source, different bytes (`-trimpath
|
||||
// -buildvcs=false` in release-agent.sh) — so self-update sees "already installed" and never
|
||||
// corrects it. Reporting the bytes lets the hub compare against the vouched agent_sha256, the
|
||||
// same mechanism host.wrapper_sha256 is for the PBS wrapper (R-50b(a)).
|
||||
//
|
||||
// Empty = unreadable, which the hub must treat as UNKNOWN, never as drift. Pinned by
|
||||
// TestCollect_AgentSHA256IsTheRunningBinary.
|
||||
AgentSHA256 string `json:"agent_sha256,omitempty"`
|
||||
|
||||
Host HostMetrics `json:"host"`
|
||||
Guests []Guest `json:"guests"`
|
||||
@@ -114,6 +124,11 @@ type HostReport struct {
|
||||
// on HostReport would have been the only report block named against that convention.
|
||||
GuestNet *GuestNetStatus `json:"guest_net,omitempty"`
|
||||
|
||||
// GuestDiskTrim is the weekly guest disk trim stanza (R-444, `09` §3 decision 139): the schedule and, per owned
|
||||
// guest, the LAST trim result as persisted by the agent (it survives a restart). Present only when the trim job is
|
||||
// wired; an empty `guests` list means the job runs and no guest has been trimmed yet. No secret.
|
||||
GuestDiskTrim *GuestDiskTrimStatus `json:"guest_disk_trim,omitempty"`
|
||||
|
||||
// LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent
|
||||
// mirror of the controller's report log_tails channel. Present ONLY on the heartbeat
|
||||
// right after the control envelope requested it (log_tail_requested); consume-once on
|
||||
@@ -198,6 +213,27 @@ type GuestNetGuest struct {
|
||||
Message string `json:"message,omitempty"`
|
||||
}
|
||||
|
||||
// GuestDiskTrimStatus is the R-444 weekly trim stanza. `schedule` is a plain description of when the job runs (local
|
||||
// time of the host); `guests` holds one entry per owned guest that has had at least one trim attempt.
|
||||
type GuestDiskTrimStatus struct {
|
||||
Schedule string `json:"schedule"`
|
||||
Guests []GuestDiskTrim `json:"guests,omitempty"`
|
||||
}
|
||||
|
||||
// GuestDiskTrim is one guest's LAST trim attempt. `ok` with `last_attempt_at` is the verdict of that attempt — never
|
||||
// read the time alone as success; `last_ok_at` is the last attempt that succeeded ("" = never). `bytes_trimmed` is
|
||||
// the sum of the "(N bytes) trimmed" lines `pct fstrim` printed, over `mounts` mount points.
|
||||
type GuestDiskTrim struct {
|
||||
VMID int `json:"vmid"`
|
||||
LastAttemptAt string `json:"last_attempt_at"`
|
||||
OK bool `json:"ok"`
|
||||
BytesTrimmed int64 `json:"bytes_trimmed"`
|
||||
Mounts int `json:"mounts"`
|
||||
DurationSeconds float64 `json:"duration_seconds"`
|
||||
LastOKAt string `json:"last_ok_at,omitempty"`
|
||||
Error string `json:"error,omitempty"`
|
||||
}
|
||||
|
||||
type PBSDRStatus struct {
|
||||
State string `json:"state"`
|
||||
StorageID string `json:"storage_id,omitempty"`
|
||||
|
||||
@@ -0,0 +1,102 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"io"
|
||||
"io/fs"
|
||||
"net/http"
|
||||
"os"
|
||||
"time"
|
||||
)
|
||||
|
||||
// GET /host/crash-guard (R-856, `09` §3 decision 143): what the host's crash guard
|
||||
// (configs/felhom-crash-guard, `11` §5.9) recorded about the most recent HOST boot. The controller
|
||||
// reads it once after it starts: when the host's last boot followed an UNCLEAN stop, its app mails
|
||||
// wait ~15 minutes instead of the normal 90 s boot grace.
|
||||
//
|
||||
// Read-only and Proxmox-free: the agent reads the guard's state file (root-owned, 0644 — the
|
||||
// non-root agent can read it) and passes four fields through. Host-wide, token-authed (any valid
|
||||
// per-guest token sees the host's view, as GET /host/metrics does).
|
||||
//
|
||||
// NEVER an error page. A missing file (no guard installed, or no boot recorded yet), an unreadable
|
||||
// one, or one that does not parse answers 200 with present:false — the controller reads that as
|
||||
// UNKNOWN and keeps its normal boot grace. Pinned by TestR856_CrashGuard*.
|
||||
|
||||
// defaultCrashGuardStatePath is where configs/felhom-crash-guard writes its state (STATE_DIR there).
|
||||
const defaultCrashGuardStatePath = "/var/lib/felhom-crash-guard/state.json"
|
||||
|
||||
// crashGuardStateMax bounds the read; the real file is well under 4 KiB.
|
||||
const crashGuardStateMax = 1 << 20
|
||||
|
||||
// CrashGuardResponse is the data block of GET /host/crash-guard. Field names are the controller's
|
||||
// agentapi.CrashGuardState (felhom-controller internal/agentapi/crashguard.go) — a wire contract,
|
||||
// pinned by TestR856_CrashGuardWireMatchesControllerClient.
|
||||
type CrashGuardResponse struct {
|
||||
Present bool `json:"present"`
|
||||
LastBootAt string `json:"last_boot_at,omitempty"` // RFC3339 UTC ("2006-01-02T15:04:05Z")
|
||||
LastBootUnclean bool `json:"last_boot_unclean"`
|
||||
Tripped bool `json:"tripped"`
|
||||
}
|
||||
|
||||
// crashGuardFile is the subset of the guard's state.json the route passes through. Every other key
|
||||
// (armed, boot_id, config, unclean_boots, last_trip, ...) is ignored.
|
||||
type crashGuardFile struct {
|
||||
LastBootAt string `json:"last_boot_at"`
|
||||
LastBootUnclean bool `json:"last_boot_unclean"`
|
||||
Tripped bool `json:"tripped"`
|
||||
}
|
||||
|
||||
// readCrashGuardState reads and parses the guard's state file. ok=false on ANY failure (missing,
|
||||
// unreadable, oversized, not a JSON object, a field of the wrong type); reason says which, for the log.
|
||||
func readCrashGuardState(path string) (resp CrashGuardResponse, ok bool, reason string) {
|
||||
f, err := os.Open(path)
|
||||
if err != nil {
|
||||
if errors.Is(err, fs.ErrNotExist) {
|
||||
return resp, false, "no state file"
|
||||
}
|
||||
return resp, false, "unreadable: " + err.Error()
|
||||
}
|
||||
defer f.Close()
|
||||
raw, err := io.ReadAll(io.LimitReader(f, crashGuardStateMax+1))
|
||||
if err != nil {
|
||||
return resp, false, "read: " + err.Error()
|
||||
}
|
||||
if len(raw) > crashGuardStateMax {
|
||||
return resp, false, "state file too large"
|
||||
}
|
||||
var st crashGuardFile
|
||||
// Unmarshal into a struct fails on a non-object top level (null decodes, so reject it below).
|
||||
if err := json.Unmarshal(raw, &st); err != nil {
|
||||
return resp, false, "unparseable: " + err.Error()
|
||||
}
|
||||
var probe map[string]json.RawMessage
|
||||
if err := json.Unmarshal(raw, &probe); err != nil || probe == nil {
|
||||
return resp, false, "unparseable: not a JSON object"
|
||||
}
|
||||
resp = CrashGuardResponse{Present: true, LastBootUnclean: st.LastBootUnclean, Tripped: st.Tripped}
|
||||
// Normalise to RFC3339 UTC; an unparseable time passes through as-is (the controller reads an
|
||||
// unparseable boot time as "not this start's boot" → its normal grace).
|
||||
if t, perr := time.Parse(time.RFC3339, st.LastBootAt); perr == nil {
|
||||
resp.LastBootAt = t.UTC().Format(time.RFC3339)
|
||||
} else {
|
||||
resp.LastBootAt = st.LastBootAt
|
||||
}
|
||||
return resp, true, ""
|
||||
}
|
||||
|
||||
func (s *Server) handleCrashGuard(w http.ResponseWriter, r *http.Request, vmid int) {
|
||||
path := s.crashGuardStatePath
|
||||
if path == "" {
|
||||
path = defaultCrashGuardStatePath
|
||||
}
|
||||
resp, ok, reason := readCrashGuardState(path)
|
||||
if !ok {
|
||||
s.logger.Debug("local-api: /host/crash-guard not present", "vmid", vmid, "reason", reason)
|
||||
writeOK(w, CrashGuardResponse{Present: false})
|
||||
return
|
||||
}
|
||||
s.logger.Debug("local-api: /host/crash-guard served", "vmid", vmid,
|
||||
"last_boot_at", resp.LastBootAt, "last_boot_unclean", resp.LastBootUnclean, "tripped", resp.Tripped)
|
||||
writeOK(w, resp)
|
||||
}
|
||||
@@ -0,0 +1,205 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"io"
|
||||
"log/slog"
|
||||
"net/http"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// The shape of /var/lib/felhom-crash-guard/state.json as read on demo-hp on 2026-10-06 (values from
|
||||
// that read where they matter; lists/objects kept to the same key set).
|
||||
const crashGuardFixture = `{
|
||||
"armed": true,
|
||||
"boot_id": "3f1c0f1e-6a0b-4d7e-9b7a-0c2d4e6f8a1b",
|
||||
"config": {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10},
|
||||
"kernel_panic": 10,
|
||||
"last_boot_at": "2026-10-05T07:56:41Z",
|
||||
"last_boot_unclean": true,
|
||||
"last_trip": {},
|
||||
"rearmed_at": "2026-10-04T14:02:11Z",
|
||||
"rearmed_by": "operator",
|
||||
"tripped": false,
|
||||
"unclean_boots": ["2026-10-05T07:56:41Z"],
|
||||
"unclean_boots_24h": 1,
|
||||
"unclean_boots_in_window": 1,
|
||||
"updated_at": "2026-10-05T07:57:02Z",
|
||||
"version": 1
|
||||
}`
|
||||
|
||||
// controllerCrashGuardState is a COPY of the controller's wire type, felhom-controller
|
||||
// controller/internal/agentapi/crashguard.go `CrashGuardState` (commit 8b13a5e) — same field names,
|
||||
// same tags. If either side renames a key, the contract test below fails.
|
||||
type controllerCrashGuardState struct {
|
||||
Present bool `json:"present"`
|
||||
LastBootAt string `json:"last_boot_at,omitempty"`
|
||||
LastBootUnclean bool `json:"last_boot_unclean"`
|
||||
Tripped bool `json:"tripped"`
|
||||
}
|
||||
|
||||
func newCrashGuardServer(t *testing.T, statePath string) http.Handler {
|
||||
t.Helper()
|
||||
srv, err := NewServer(Options{
|
||||
ListenAddr: "127.0.0.1:0",
|
||||
Guests: &fakeGuests{},
|
||||
Backups: &fakeBackups{},
|
||||
Store: &fakeStore{},
|
||||
Storage: fakeStorage{},
|
||||
Tokens: staticTokens{"A": 8200, "B": 9300},
|
||||
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("new server: %v", err)
|
||||
}
|
||||
srv.crashGuardStatePath = statePath
|
||||
return srv.Handler()
|
||||
}
|
||||
|
||||
func writeCrashGuardFixture(t *testing.T, body string) string {
|
||||
t.Helper()
|
||||
p := filepath.Join(t.TempDir(), "state.json")
|
||||
if err := os.WriteFile(p, []byte(body), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
// getCrashGuard calls the route and decodes the envelope with the CONTROLLER's type.
|
||||
func getCrashGuard(t *testing.T, h http.Handler, token string) (int, controllerCrashGuardState, string) {
|
||||
t.Helper()
|
||||
w := do(t, h, "GET", "/host/crash-guard", token, "")
|
||||
var env struct {
|
||||
OK bool `json:"ok"`
|
||||
Data controllerCrashGuardState `json:"data"`
|
||||
}
|
||||
if w.Code == http.StatusOK {
|
||||
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil {
|
||||
t.Fatalf("decode %q: %v", w.Body.String(), err)
|
||||
}
|
||||
if !env.OK {
|
||||
t.Fatalf("ok=false: %s", w.Body.String())
|
||||
}
|
||||
}
|
||||
return w.Code, env.Data, w.Body.String()
|
||||
}
|
||||
|
||||
// A present state file (demo-hp's shape) passes the three facts through.
|
||||
func TestR856_CrashGuardPresentFile(t *testing.T) {
|
||||
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
|
||||
code, st, body := getCrashGuard(t, h, "A")
|
||||
if code != http.StatusOK {
|
||||
t.Fatalf("got %d, want 200 (%s)", code, body)
|
||||
}
|
||||
want := controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: true, Tripped: false}
|
||||
if st != want {
|
||||
t.Fatalf("state = %+v, want %+v", st, want)
|
||||
}
|
||||
|
||||
// A tripped, clean boot reads back as such (both bools are carried, not defaulted).
|
||||
h = newCrashGuardServer(t, writeCrashGuardFixture(t,
|
||||
`{"last_boot_at":"2026-10-05T09:56:41+02:00","last_boot_unclean":false,"tripped":true,"version":1}`))
|
||||
_, st, _ = getCrashGuard(t, h, "B")
|
||||
want = controllerCrashGuardState{Present: true, LastBootAt: "2026-10-05T07:56:41Z", LastBootUnclean: false, Tripped: true}
|
||||
if st != want {
|
||||
t.Fatalf("offset time / tripped: state = %+v, want %+v (time normalised to UTC Z)", st, want)
|
||||
}
|
||||
}
|
||||
|
||||
// No state file (no guard on this host, or no boot recorded yet) → 200 present:false.
|
||||
func TestR856_CrashGuardMissingFile(t *testing.T) {
|
||||
h := newCrashGuardServer(t, filepath.Join(t.TempDir(), "absent", "state.json"))
|
||||
code, st, body := getCrashGuard(t, h, "A")
|
||||
if code != http.StatusOK {
|
||||
t.Fatalf("missing file: got %d, want 200 (%s)", code, body)
|
||||
}
|
||||
if st.Present || st.LastBootUnclean || st.Tripped || st.LastBootAt != "" {
|
||||
t.Fatalf("missing file: state = %+v, want present:false and nothing else", st)
|
||||
}
|
||||
}
|
||||
|
||||
// A garbled file → 200 present:false, never a 5xx — every shape of garbage.
|
||||
func TestR856_CrashGuardGarbageFile(t *testing.T) {
|
||||
for name, body := range map[string]string{
|
||||
"truncated": crashGuardFixture[:40],
|
||||
"not json": "this is not json\n",
|
||||
"empty": "",
|
||||
"null": "null",
|
||||
"array": `[{"last_boot_unclean":true}]`,
|
||||
"wrong type": `{"last_boot_at":"2026-10-05T07:56:41Z","last_boot_unclean":"yes","tripped":false}`,
|
||||
"lone brace": "{",
|
||||
} {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
h := newCrashGuardServer(t, writeCrashGuardFixture(t, body))
|
||||
code, st, raw := getCrashGuard(t, h, "A")
|
||||
if code != http.StatusOK {
|
||||
t.Fatalf("got %d, want 200 (%s)", code, raw)
|
||||
}
|
||||
if st.Present || st.LastBootUnclean {
|
||||
t.Fatalf("garbage %q read as %+v, want present:false", name, st)
|
||||
}
|
||||
})
|
||||
}
|
||||
// The path is a directory, not a file: unreadable → present:false, 200.
|
||||
h := newCrashGuardServer(t, t.TempDir())
|
||||
if code, st, raw := getCrashGuard(t, h, "A"); code != http.StatusOK || st.Present {
|
||||
t.Fatalf("directory path: got %d %+v (%s), want 200 present:false", code, st, raw)
|
||||
}
|
||||
}
|
||||
|
||||
// No / unknown token → 401, like every sibling route; a cross-guest ?vmid= → 403.
|
||||
func TestR856_CrashGuardRequiresGuestToken(t *testing.T) {
|
||||
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
|
||||
for _, tok := range []string{"", "bogus"} {
|
||||
w := do(t, h, "GET", "/host/crash-guard", tok, "")
|
||||
if w.Code != http.StatusUnauthorized {
|
||||
t.Fatalf("token %q: got %d, want 401", tok, w.Code)
|
||||
}
|
||||
if json.Valid(w.Body.Bytes()) {
|
||||
var env struct {
|
||||
Data controllerCrashGuardState `json:"data"`
|
||||
}
|
||||
_ = json.Unmarshal(w.Body.Bytes(), &env)
|
||||
if env.Data.Present || env.Data.LastBootUnclean {
|
||||
t.Fatalf("token %q: the refusal leaked the state: %s", tok, w.Body.String())
|
||||
}
|
||||
}
|
||||
}
|
||||
if w := do(t, h, "GET", "/host/crash-guard?vmid=9300", "A", ""); w.Code != http.StatusForbidden {
|
||||
t.Fatalf("cross-guest query: got %d, want 403", w.Code)
|
||||
}
|
||||
}
|
||||
|
||||
// Wire contract: every key the controller's CrashGuardState decodes is emitted under exactly that
|
||||
// name, and the agent emits no key the controller does not know.
|
||||
func TestR856_CrashGuardWireMatchesControllerClient(t *testing.T) {
|
||||
h := newCrashGuardServer(t, writeCrashGuardFixture(t, crashGuardFixture))
|
||||
w := do(t, h, "GET", "/host/crash-guard", "A", "")
|
||||
var env struct {
|
||||
OK bool `json:"ok"`
|
||||
Data map[string]json.RawMessage `json:"data"`
|
||||
}
|
||||
if err := json.Unmarshal(w.Body.Bytes(), &env); err != nil || !env.OK {
|
||||
t.Fatalf("envelope: %v %s", err, w.Body.String())
|
||||
}
|
||||
want := []string{"present", "last_boot_at", "last_boot_unclean", "tripped"}
|
||||
for _, k := range want {
|
||||
if _, ok := env.Data[k]; !ok {
|
||||
t.Errorf("agent does not emit %q, which the controller decodes (%s)", k, w.Body.String())
|
||||
}
|
||||
}
|
||||
if len(env.Data) != len(want) {
|
||||
t.Errorf("agent emits %d keys, controller knows %d: %s", len(env.Data), len(want), w.Body.String())
|
||||
}
|
||||
// And the agent's own type agrees with the controller's copy, field for field.
|
||||
var mine CrashGuardResponse
|
||||
var theirs controllerCrashGuardState
|
||||
raw, _ := json.Marshal(env.Data)
|
||||
_ = json.Unmarshal(raw, &mine)
|
||||
_ = json.Unmarshal(raw, &theirs)
|
||||
if (controllerCrashGuardState{mine.Present, mine.LastBootAt, mine.LastBootUnclean, mine.Tripped}) != theirs {
|
||||
t.Errorf("agent %+v vs controller %+v", mine, theirs)
|
||||
}
|
||||
}
|
||||
@@ -768,6 +768,11 @@ type FormatResponse struct {
|
||||
// signature — the customer authorizes the wipe of their own data drive.
|
||||
NeedsConfirmation bool `json:"needs_confirmation,omitempty"`
|
||||
DurableID string `json:"durable_id,omitempty"` // the durable id to confirm against (user-data)
|
||||
// FSUUID (on Formatted) is the UUID of the filesystem the agent just made, verified against the
|
||||
// bound durable id after mkfs (R-25). The caller mounts THIS — re-resolving a UUID from the /dev
|
||||
// path later can name another disk if /dev re-enumerated. "" = not verified: the caller must not
|
||||
// substitute a path-resolved guess silently.
|
||||
FSUUID string `json:"fs_uuid,omitempty"`
|
||||
// PendingOp is set on a SYSTEM/BACKUP data-bearing refusal — the exact op the operator must sign.
|
||||
PendingOp *PendingOp `json:"pending_op,omitempty"`
|
||||
}
|
||||
@@ -815,7 +820,7 @@ func (s *Server) handleDiskFormatStatus(w http.ResponseWriter, r *http.Request,
|
||||
writeOK(w, map[string]any{
|
||||
"vmid": vmid, "phase": job.Phase, "device": job.Device, "fstype": job.FSType,
|
||||
"durable_id": job.DurableID, "error": job.Error, "started_at": job.StartedAt, "updated_at": job.UpdatedAt,
|
||||
"job_id": job.JobID,
|
||||
"job_id": job.JobID, "fs_uuid": job.FSUUID, // R-25: "" until done + verified
|
||||
})
|
||||
}
|
||||
|
||||
@@ -878,7 +883,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
|
||||
"format refused (device may have changed since inspection): "+rerr.Error())
|
||||
return
|
||||
}
|
||||
done := s.startFormatDetached(device, blankDurable, req.FSType, true)
|
||||
job, done := s.startFormatDetached(device, blankDurable, req.FSType, true)
|
||||
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
|
||||
if err == errFormatClientGone {
|
||||
return // client gone; mkfs continues detached + the job record records the outcome
|
||||
@@ -887,7 +892,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
|
||||
writeErr(w, http.StatusBadGateway, "format failed: "+err.Error())
|
||||
return
|
||||
}
|
||||
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, Reason: "blank device formatted " + req.FSType})
|
||||
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, FSUUID: job.FSUUID, Reason: "blank device formatted " + req.FSType})
|
||||
return
|
||||
}
|
||||
|
||||
@@ -924,7 +929,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
|
||||
// F20-BUG3: run the destructive mkfs DETACHED off s.baseCtx (bound durable id recorded for
|
||||
// restart-recovery), so a request/client deadline can never SIGKILL it mid-write and corrupt the
|
||||
// disk. We still wait to return the synchronous result (backward-compatible with the controller).
|
||||
done := s.startFormatDetached(device, deviceDurable, req.FSType, false)
|
||||
job, done := s.startFormatDetached(device, deviceDurable, req.FSType, false)
|
||||
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
|
||||
if err == errFormatClientGone {
|
||||
return // client gone; the wipe continues detached + survives a restart via the job record
|
||||
@@ -936,7 +941,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
|
||||
s.logger.Warn("local-api: USER-DATA data-bearing format — CUSTOMER CONFIRMED (no operator signature)",
|
||||
"vmid", vmid, "device", device, "durable_id", deviceDurable, "fstype", req.FSType, "why", probe.Reason())
|
||||
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: true,
|
||||
Role: string(role), DurableID: deviceDurable, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"})
|
||||
Role: string(role), DurableID: deviceDurable, FSUUID: job.FSUUID, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"})
|
||||
return
|
||||
}
|
||||
|
||||
|
||||
@@ -28,6 +28,7 @@ type fakeDiskOps struct {
|
||||
unmountCalls []string
|
||||
candidates []storage.CandidateDisk // returned by ListCandidateDisks
|
||||
candErr error
|
||||
afterFormat *storage.DeviceProbe // R-25: when set, InspectDevice returns it once a format ran
|
||||
}
|
||||
|
||||
func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.CandidateDisk, error) {
|
||||
@@ -35,7 +36,12 @@ func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.Candidate
|
||||
}
|
||||
|
||||
func (f *fakeDiskOps) InspectDevice(_ context.Context, device string) (storage.DeviceProbe, error) {
|
||||
f.mu.Lock()
|
||||
p := f.probe
|
||||
if f.afterFormat != nil && len(f.formatCalls) > 0 {
|
||||
p = *f.afterFormat
|
||||
}
|
||||
f.mu.Unlock()
|
||||
p.Device = device
|
||||
return p, f.inspectErr
|
||||
}
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"net/http"
|
||||
"sync"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/storage"
|
||||
)
|
||||
|
||||
// R-25: the format answer carries the UUID of the filesystem the agent JUST made, verified against the
|
||||
// bound durable id after mkfs, so the controller mounts that filesystem rather than whatever the /dev
|
||||
// path resolves to a few requests later. The consequence asserted: the UUID on the wire (and in the
|
||||
// polled job record) is the new superblock's — and is EMPTY whenever the binding cannot be re-proved.
|
||||
|
||||
const newFSUUID = "0fc63daf-8483-4772-8e79-3d69d8477de4"
|
||||
|
||||
func confirmedFormat(t *testing.T, d *fakeDiskOps, srv *Server, fj *FormatJobStore) (string, *formatJob) {
|
||||
t.Helper()
|
||||
w := do(t, srv.Handler(), "POST", "/disks/format", "A", `{"device":"/dev/sdb1","fstype":"ext4","confirmed":true,"durable_id":"byid:wwn-/dev/sdb1"}`)
|
||||
if w.Code != http.StatusOK {
|
||||
t.Fatalf("confirmed format: %d (%s)", w.Code, w.Body.String())
|
||||
}
|
||||
var resp struct {
|
||||
Data FormatResponse `json:"data"`
|
||||
}
|
||||
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
|
||||
t.Fatalf("decode: %v (%s)", err, w.Body.String())
|
||||
}
|
||||
if !resp.Data.Formatted {
|
||||
t.Fatalf("not formatted: %s", w.Body.String())
|
||||
}
|
||||
return resp.Data.FSUUID, waitFormatPhase(t, fj, formatPhaseDone)
|
||||
}
|
||||
|
||||
func confirmedGate() *fakeGate {
|
||||
return &fakeGate{decision: WipeDecision{Allowed: true, Tier: "customer_confirmable", Reason: "customer_confirmed"}}
|
||||
}
|
||||
|
||||
func TestFormat_ReportsNewFilesystemUUID(t *testing.T) {
|
||||
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
|
||||
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
|
||||
fj := tempFormatStore(t)
|
||||
srv := formatServer(t, d, confirmedGate(), fj)
|
||||
|
||||
got, job := confirmedFormat(t, d, srv, fj)
|
||||
if got != newFSUUID {
|
||||
t.Fatalf("response fs_uuid = %q, want the new filesystem's %q", got, newFSUUID)
|
||||
}
|
||||
if job.FSUUID != newFSUUID {
|
||||
t.Fatalf("job record fs_uuid = %q, want %q (the polled status path)", job.FSUUID, newFSUUID)
|
||||
}
|
||||
}
|
||||
|
||||
// The node moved between mkfs and the read-back: the bound durable id now resolves elsewhere. The
|
||||
// UUID must NOT be reported — reading it would name the other disk's filesystem.
|
||||
func TestFormat_FSUUIDWithheldWhenDurableIDMoved(t *testing.T) {
|
||||
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
|
||||
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
|
||||
fj := tempFormatStore(t)
|
||||
srv := formatServer(t, d, confirmedGate(), fj)
|
||||
var mu sync.Mutex
|
||||
calls := 0
|
||||
srv.reresolveWipe = func(_ context.Context, _ string) (string, error) {
|
||||
mu.Lock()
|
||||
defer mu.Unlock()
|
||||
calls++
|
||||
if calls == 1 {
|
||||
return "/dev/sdb", nil // the pre-mkfs anti-retarget re-resolve
|
||||
}
|
||||
return "/dev/sdc", nil // after mkfs: the durable id now names another node
|
||||
}
|
||||
|
||||
got, job := confirmedFormat(t, d, srv, fj)
|
||||
if got != "" || job.FSUUID != "" {
|
||||
t.Fatalf("fs_uuid reported after the durable id moved (response %q, job %q) — must be empty", got, job.FSUUID)
|
||||
}
|
||||
if calls < 2 {
|
||||
t.Fatalf("the post-mkfs re-resolve never ran (calls=%d)", calls)
|
||||
}
|
||||
}
|
||||
|
||||
// The superblock did not read back as the requested filesystem → not verified → empty.
|
||||
func TestFormat_FSUUIDWithheldOnFSTypeMismatch(t *testing.T) {
|
||||
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
|
||||
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "xfs", FSUUID: newFSUUID}}
|
||||
fj := tempFormatStore(t)
|
||||
srv := formatServer(t, d, confirmedGate(), fj)
|
||||
|
||||
got, job := confirmedFormat(t, d, srv, fj)
|
||||
if got != "" || job.FSUUID != "" {
|
||||
t.Fatalf("fs_uuid reported for a superblock of the wrong type (response %q, job %q)", got, job.FSUUID)
|
||||
}
|
||||
}
|
||||
@@ -22,6 +22,10 @@ type formatJob struct {
|
||||
Blank bool `json:"blank,omitempty"` // audit D3: blank (benign) format — recovery re-checks STILL-blank, not data-bearing
|
||||
Phase string `json:"phase"` // running | done | failed
|
||||
Error string `json:"error,omitempty"`
|
||||
// FSUUID is the filesystem UUID of the NEW filesystem, read by the agent right after mkfs on the
|
||||
// device the bound durable id still resolves to (R-25). "" = not verified (the controller must not
|
||||
// read that as a UUID). Set only on phase done.
|
||||
FSUUID string `json:"fs_uuid,omitempty"`
|
||||
StartedAt string `json:"started_at"`
|
||||
UpdatedAt string `json:"updated_at"`
|
||||
}
|
||||
@@ -96,7 +100,10 @@ func (s *FormatJobStore) save(j *formatJob) error {
|
||||
// runs to completion and records the outcome. device is the ALREADY anti-retarget-resolved device; the
|
||||
// record carries durableID so a restart can re-resolve + re-run. blank marks a benign (blank-device)
|
||||
// format, so restart recovery re-checks STILL-blank rather than data-bearing (audit D3).
|
||||
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) <-chan error {
|
||||
//
|
||||
// The returned job may be read (FSUUID) only AFTER a value arrives on done — the goroutine writes it
|
||||
// before the send, which is the happens-before edge.
|
||||
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) (*formatJob, <-chan error) {
|
||||
base := s.baseCtx
|
||||
if base == nil {
|
||||
base = context.Background()
|
||||
@@ -116,10 +123,39 @@ func (s *Server) startFormatDetached(device, durableID, fstype string, blank boo
|
||||
ctx, cancel := context.WithTimeout(base, 60*time.Minute)
|
||||
defer cancel()
|
||||
err := s.disks.Format(ctx, device, fstype)
|
||||
if err == nil {
|
||||
job.FSUUID = s.formattedFSUUID(ctx, device, durableID, fstype)
|
||||
}
|
||||
s.finishFormatJob(job, err)
|
||||
done <- err
|
||||
}()
|
||||
return done
|
||||
return job, done
|
||||
}
|
||||
|
||||
// formattedFSUUID reads the UUID of the filesystem the agent has JUST made (R-25). The caller used to
|
||||
// re-resolve the UUID from the mutable /dev path afterwards, over separate requests — a re-enumeration
|
||||
// in that window could hand it ANOTHER disk's filesystem to mount. Here the bound durable id must still
|
||||
// resolve to the very device that was formatted (and re-derive to the same id), and the superblock must
|
||||
// carry the fstype that was asked for; anything else returns "" (not verified), never a guess.
|
||||
func (s *Server) formattedFSUUID(ctx context.Context, device, durableID, fstype string) string {
|
||||
if durableID == "" || s.reresolveWipe == nil {
|
||||
return ""
|
||||
}
|
||||
// The device now holds a filesystem, so the data-bearing anti-retarget re-resolve is the right one.
|
||||
now, err := s.reresolveWipe(ctx, durableID)
|
||||
if err != nil || now != device {
|
||||
s.logger.Warn("format: new filesystem UUID NOT reported — bound durable id no longer resolves to the formatted device",
|
||||
"device", device, "durable_id", durableID, "resolves_to", now, "err", err)
|
||||
return ""
|
||||
}
|
||||
probe, err := s.disks.InspectDevice(ctx, device)
|
||||
if err != nil || !probe.Probed || probe.FSType != fstype || probe.FSUUID == "" {
|
||||
s.logger.Warn("format: new filesystem UUID NOT reported — superblock did not read back as the requested filesystem",
|
||||
"device", device, "want_fstype", fstype, "got_fstype", probe.FSType, "has_uuid", probe.FSUUID != "", "err", err)
|
||||
return ""
|
||||
}
|
||||
s.logger.Info("format: new filesystem bound to its durable id", "device", device, "durable_id", durableID, "fs_uuid", probe.FSUUID)
|
||||
return probe.FSUUID
|
||||
}
|
||||
|
||||
// finishFormatJob updates the persisted record to done/failed.
|
||||
@@ -172,7 +208,7 @@ func (s *Server) RecoverFormatJob(ctx context.Context) {
|
||||
return
|
||||
}
|
||||
s.logger.Warn("format-job recover: re-running interrupted format detached", "durable_id", job.DurableID, "device", device, "fstype", job.FSType, "blank", job.Blank)
|
||||
_ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion
|
||||
_, _ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion
|
||||
}
|
||||
|
||||
// nowFn returns the server clock (testable), defaulting to time.Now.
|
||||
|
||||
@@ -294,6 +294,9 @@ type Server struct {
|
||||
netMountRoot string // the user-data namespace root for the network-mount role gate
|
||||
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
|
||||
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
|
||||
// crashGuardStatePath (R-856) is the host crash guard's state file read by GET /host/crash-guard;
|
||||
// empty = defaultCrashGuardStatePath. A seam: tests point it at a fixture.
|
||||
crashGuardStatePath string
|
||||
// escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed
|
||||
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
|
||||
// offsite repository password. OPTIONAL — nil (no hub client configured) makes
|
||||
@@ -520,6 +523,9 @@ func (s *Server) Handler() http.Handler {
|
||||
// Host metrics (slice 9): host-wide health + per-storage capacity for the customer's monitoring
|
||||
// view. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot).
|
||||
mux.HandleFunc("GET /host/metrics", s.withGuest(s.handleHostMetrics))
|
||||
// R-856 (`09` §3 decision 143): the host crash guard's record of the last HOST boot — the controller
|
||||
// waits ~15 min with app mails after an unclean one. Read-only; a missing/garbled file = present:false.
|
||||
mux.HandleFunc("GET /host/crash-guard", s.withGuest(s.handleCrashGuard))
|
||||
// Disk management (slice 8C) — self-scoped; format routes through the data-bearing classifier+gate.
|
||||
mux.HandleFunc("GET /disks", s.withGuest(s.handleDisks))
|
||||
mux.HandleFunc("GET /disks/candidates", s.withGuest(s.handleDiskCandidates))
|
||||
|
||||
@@ -57,6 +57,10 @@ type DeviceProbe struct {
|
||||
HasPartitions bool `json:"has_partitions"` // child partitions present (lsblk)
|
||||
Mounted bool `json:"mounted"` // currently mounted somewhere
|
||||
FSType string `json:"fstype,omitempty"`
|
||||
// FSUUID is the filesystem UUID blkid read from the on-disk superblock (`blkid -p`, no cache),
|
||||
// "" when there is none. R-25: the format path reports it so the caller mounts the filesystem the
|
||||
// agent just made, not whatever a /dev path resolves to later.
|
||||
FSUUID string `json:"fs_uuid,omitempty"`
|
||||
}
|
||||
|
||||
// DataBearing is the conservative verdict: any signature / partition table / partition / mount —
|
||||
@@ -423,6 +427,8 @@ func (h *SudoHostOps) InspectDevice(ctx context.Context, device string) (DeviceP
|
||||
probe.FSType = v
|
||||
case "PTTYPE":
|
||||
probe.HasPartitionTable = true
|
||||
case "UUID":
|
||||
probe.FSUUID = v
|
||||
case "USAGE":
|
||||
if v != "" {
|
||||
probe.HasFilesystem = true // filesystem/raid/crypto member = data-bearing
|
||||
|
||||
@@ -208,3 +208,17 @@ func TestFormat_RejectsBadArgs(t *testing.T) {
|
||||
t.Fatalf("mkfs ran despite invalid input: %v", r.calls)
|
||||
}
|
||||
}
|
||||
|
||||
// R-25: the probe carries the superblock's filesystem UUID so the format path can report the new one.
|
||||
func TestInspect_ReadsFilesystemUUID(t *testing.T) {
|
||||
r := &scriptedRunner{
|
||||
outputs: map[string][]byte{
|
||||
"blkid": []byte("DEVNAME=/dev/sdb\nUUID=0fc63daf-8483-4772-8e79-3d69d8477de4\nTYPE=ext4\nUSAGE=filesystem\n"),
|
||||
"lsblk": []byte(`{"blockdevices":[{"name":"sdb","fstype":"ext4","pttype":null,"mountpoint":null}]}`),
|
||||
},
|
||||
}
|
||||
p, _ := newSudo(r).InspectDevice(context.Background(), "/dev/sdb")
|
||||
if p.FSUUID != "0fc63daf-8483-4772-8e79-3d69d8477de4" {
|
||||
t.Fatalf("FSUUID = %q, want the blkid UUID", p.FSUUID)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,367 @@
|
||||
#!/usr/bin/env python3
|
||||
# -*- coding: utf-8 -*-
|
||||
"""test_gate_decoys.py — can this repo's gates be fooled by a LABEL? (R-421, R-426)
|
||||
|
||||
The same instrument as `felhom.eu/scripts/test_gate_decoys.py`: a decoy is the LABEL without the
|
||||
FACT, and a gate that passes on the label alone — or refuses the genuine article — is a live hole.
|
||||
Every gate is asserted in BOTH directions: the decoy must be convicted, the genuine article passed.
|
||||
|
||||
Covered here (the `COVERS` literal is AST-read by `felhom.eu/scripts/decoy_coverage_gate.py`, which
|
||||
never imports this file):
|
||||
|
||||
published check-published-versions.py, against a FAKE Gitea (see below).
|
||||
release-complete check-release-complete.py, in a scratch clone whose `origin` is a scratch bare
|
||||
repository, against the same fake Gitea.
|
||||
reuse-refs, instructions, observations
|
||||
the three SHARED felhom.eu scripts, run against a scratch clone of THIS repo —
|
||||
so the decoy is planted in the agent's own REUSE.md / CLAUDE.md / REPORT.md and
|
||||
coverage is per input, not per script.
|
||||
|
||||
NEVER THE REAL GITEA. Both network gates read `GITEA_BASE` from the environment (CI already sets it
|
||||
to the in-cluster URL); here it points at an `http.server` bound to 127.0.0.1 inside this process,
|
||||
and every proxy variable is removed from the child's environment so urllib cannot route around it.
|
||||
A test that asked the real registry would pass or fail on whatever was published that day — the
|
||||
constant-for-measurement shape — and would reach the network from a hook.
|
||||
|
||||
NEVER THE REAL TREE. Every planted file lives in a scratch directory: a workspace that holds a
|
||||
clone of this repo beside symlinks to the sibling clones the shared scripts reach across to.
|
||||
|
||||
Run from the repo root: python3 scripts/test_gate_decoys.py
|
||||
Exit 0 every decoy judged correctly · 1 a decoy passed or a genuine article was refused.
|
||||
"""
|
||||
import http.server
|
||||
import io
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import socketserver
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import threading
|
||||
|
||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
PARENT = os.path.dirname(ROOT)
|
||||
# DECOY_SHARED_DIR exists for ONE purpose: the red-proof. It lets a mutated COPY of the shared scripts
|
||||
# be judged without editing the felhom.eu clone. Unset, the suite judges the real shared scripts.
|
||||
SHARED = os.environ.get("DECOY_SHARED_DIR") or os.path.join(PARENT, "felhom.eu", "scripts")
|
||||
|
||||
# ── WHAT THIS FILE COVERS ────────────────────────────────────────────────────────────────────────
|
||||
# Read by felhom.eu/scripts/decoy_coverage_gate.py, which AST-parses this literal. A gate named here
|
||||
# MUST have a decoy below that has been seen to fail.
|
||||
COVERS = {
|
||||
"published": "against a FAKE Gitea: a tag whose package 404s, a tag whose tree lacks the configs, a package one "
|
||||
"version past the newest tag (published, never tagged), a patch-gap orphan, a missing package that "
|
||||
"lexical sorting would drop out of the retention window (0.9.x vs 0.10.x); a tags api answering 500 "
|
||||
"or a non-JSON 200 is INCONCLUSIVE, never a pass - vs a clean registry, a non-semver tag and a version "
|
||||
"older than the retention window (not asserted, BY DESIGN) (R-426)",
|
||||
"release-complete": "the newest `## vX.Y.Z` with no tag anywhere, a tag parked on an unrelated commit, a tag "
|
||||
"with no package; a registry 500 is INCONCLUSIVE - vs the genuine release, an `## Unreleased` "
|
||||
"heading above it, a newer version named only in prose or under `###`, a tag that only "
|
||||
"origin has (the shallow-CI shape); a LOCAL-only tag passes BY DESIGN (CI's fresh clone "
|
||||
"and the published gate's converse probe are what see it) (R-426)",
|
||||
"reuse-refs": "a cited .go and a cited .md path that do not exist, planted in THIS repo's REUSE.md - vs the "
|
||||
"real file (R-426)",
|
||||
"instructions": "a component version literal in THIS repo's CLAUDE.md effective text - vs the same sentence "
|
||||
"inside an HTML comment (R-426)",
|
||||
"observations": "R-419 in THIS repo's REPORT.md: an Observations note SAYING it carries no marker - vs the "
|
||||
"two genuine markers (R-426)",
|
||||
}
|
||||
|
||||
fails = []
|
||||
ran = 0
|
||||
|
||||
|
||||
def report(name, rc, out, expect_rc, must=()):
|
||||
global ran
|
||||
ran += 1
|
||||
missing = [m for m in must if m not in out]
|
||||
if rc == expect_rc and not missing:
|
||||
print(" ok %-62s rc=%d (expected %d)" % (name, rc, expect_rc))
|
||||
else:
|
||||
hole = expect_rc != 0 and rc == 0
|
||||
fails.append("%s: rc=%d expected %d%s; missing %s\n%s" % (
|
||||
name, rc, expect_rc, " - LIVE HOLE" if hole else "", missing, out[-900:]))
|
||||
|
||||
|
||||
# ── the fake Gitea ───────────────────────────────────────────────────────────────────────────────
|
||||
class Fake(object):
|
||||
"""What the fake registry serves. Reset per case."""
|
||||
|
||||
def reset(self):
|
||||
self.tags = [] # tag names, as the tags api lists them
|
||||
self.packages = set() # versions whose generic package downloads
|
||||
self.raw = set() # versions whose tag tree serves the probe config
|
||||
self.tags_status = 200
|
||||
self.tags_body = None # override bytes for the tags api
|
||||
self.pkg_status = None # override status for EVERY package request
|
||||
self.hits = []
|
||||
|
||||
|
||||
FAKE = Fake()
|
||||
FAKE.reset()
|
||||
PKG_PREFIX = "/api/packages/admin/generic/felhom-agent/"
|
||||
RAW_PREFIX = "/admin/felhom-agent/raw/tag/v"
|
||||
|
||||
|
||||
class Handler(http.server.BaseHTTPRequestHandler):
|
||||
def log_message(self, *a):
|
||||
pass
|
||||
|
||||
def _answer(self, status, body=b""):
|
||||
self.send_response(status)
|
||||
self.send_header("Content-Length", str(len(body)))
|
||||
self.end_headers()
|
||||
if self.command != "HEAD":
|
||||
self.wfile.write(body)
|
||||
|
||||
def do_GET(self):
|
||||
p = self.path
|
||||
FAKE.hits.append(p)
|
||||
if p.startswith("/api/v1/repos/admin/felhom-agent/tags"):
|
||||
body = FAKE.tags_body if FAKE.tags_body is not None else \
|
||||
json.dumps([{"name": t} for t in FAKE.tags]).encode()
|
||||
return self._answer(FAKE.tags_status, body)
|
||||
if p.startswith(PKG_PREFIX):
|
||||
if FAKE.pkg_status is not None:
|
||||
return self._answer(FAKE.pkg_status)
|
||||
v = p[len(PKG_PREFIX):].split("/", 1)[0]
|
||||
return self._answer(200, b"ELF") if v in FAKE.packages else self._answer(404)
|
||||
if p.startswith(RAW_PREFIX):
|
||||
v = p[len(RAW_PREFIX):].split("/", 1)[0]
|
||||
ok = v in FAKE.raw and p.endswith("/configs/felhom-agent.service")
|
||||
return self._answer(200, b"[Unit]\n") if ok else self._answer(404)
|
||||
return self._answer(404)
|
||||
|
||||
do_HEAD = do_GET
|
||||
|
||||
|
||||
class Server(socketserver.ThreadingMixIn, http.server.HTTPServer):
|
||||
daemon_threads = True
|
||||
|
||||
|
||||
def child_env(base):
|
||||
env = {k: v for k, v in os.environ.items() if "proxy" not in k.lower()}
|
||||
env["GITEA_BASE"] = base
|
||||
env["NO_PROXY"] = env["no_proxy"] = "127.0.0.1,localhost"
|
||||
return env
|
||||
|
||||
|
||||
def run(argv, cwd, env=None):
|
||||
# input="" — a child must never inherit (and block on) this process's stdin
|
||||
p = subprocess.run(argv, cwd=cwd, env=env, capture_output=True, text=True, input="")
|
||||
return p.returncode, p.stdout + p.stderr
|
||||
|
||||
|
||||
def sh(argv, cwd):
|
||||
rc, out = run(argv, cwd)
|
||||
if rc != 0:
|
||||
raise SystemExit("setup command failed (%s): %s" % (" ".join(argv), out))
|
||||
return out.strip()
|
||||
|
||||
|
||||
# ── published ────────────────────────────────────────────────────────────────────────────────────
|
||||
def published_cases(base):
|
||||
gate = os.path.join(ROOT, "scripts", "check-published-versions.py")
|
||||
keep = json.load(io.open(os.path.join(ROOT, "scripts", "retention-policy.json"),
|
||||
encoding="utf-8"))["generic_versions_kept"]
|
||||
TAGS = ["v0.150.%d" % i for i in range(3)] # inside any retention window >= 3
|
||||
VERS = [t[1:] for t in TAGS]
|
||||
|
||||
def case(name, setup, expect_rc, must=()):
|
||||
FAKE.reset()
|
||||
FAKE.tags = list(TAGS)
|
||||
FAKE.packages = set(VERS)
|
||||
FAKE.raw = set(VERS)
|
||||
setup()
|
||||
rc, out = run([sys.executable, gate], ROOT, child_env(base))
|
||||
if not FAKE.hits:
|
||||
fails.append("published/%s: the gate never asked the fake Gitea - the seam is not wired" % name)
|
||||
report("published: " + name, rc, out, expect_rc, must)
|
||||
|
||||
case("GENUINE: every tag downloadable and serving its configs", lambda: None, 0,
|
||||
("ALL RELEASED VERSIONS INSTALLABLE",))
|
||||
case("GENUINE: a non-semver tag is not a release", lambda: FAKE.tags.append("v0.150.2-rc1"), 0,
|
||||
("ALL RELEASED VERSIONS INSTALLABLE",))
|
||||
case("FACT: a tag whose package 404s", lambda: FAKE.packages.discard("0.150.1"), 1,
|
||||
("FAIL v0.150.1", "binary NOT downloadable"))
|
||||
case("FACT: a tag whose tree does not serve the configs", lambda: FAKE.raw.discard("0.150.2"), 1,
|
||||
("FAIL v0.150.2", "does not serve"))
|
||||
case("FACT: published one patch past the newest tag, never tagged",
|
||||
lambda: FAKE.packages.add("0.150.3"), 1, ("PUBLISHED VERSION(S) WITH NO TAG", "v0.150.3"))
|
||||
case("FACT: published in a patch GAP between two tags",
|
||||
lambda: (FAKE.tags.remove("v0.150.1"),), 1, ("v0.150.1 is downloadable", "has no git tag"))
|
||||
|
||||
def lexical():
|
||||
# keep+1 tags: 0.9.0 and 0.10.0..0.10.<keep-1>. By SEMVER the oldest is 0.9.0 (dropped); by
|
||||
# STRING sort "0.10.0" is the smallest and would be the one dropped - so its missing package
|
||||
# is convicted only if the window is cut by semver.
|
||||
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
|
||||
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.10.0"}
|
||||
FAKE.raw = set(t[1:] for t in FAKE.tags)
|
||||
case("FACT: a missing package lexical sorting would drop (0.10.0 vs 0.9.0)", lexical, 1,
|
||||
("FAIL v0.10.0",))
|
||||
|
||||
def retired():
|
||||
FAKE.tags = ["v0.9.0"] + ["v0.10.%d" % i for i in range(keep)]
|
||||
FAKE.packages = set(t[1:] for t in FAKE.tags) - {"0.9.0"}
|
||||
FAKE.raw = set(t[1:] for t in FAKE.tags)
|
||||
case("BY DESIGN: a version older than the retention window is not asserted", retired, 0,
|
||||
("NOT ASSERTED", "0.9.0"))
|
||||
|
||||
def five_hundred():
|
||||
FAKE.tags_status = 500
|
||||
case("INCONCLUSIVE: the tags api answers 500", five_hundred, 2, ("INCONCLUSIVE",))
|
||||
|
||||
def html():
|
||||
FAKE.tags_body = b"<html>sign in</html>"
|
||||
case("INCONCLUSIVE: the tags api answers a 200 that is not JSON", html, 2, ("INCONCLUSIVE",))
|
||||
|
||||
# an unreachable Gitea: a port nothing listens on
|
||||
s = Server(("127.0.0.1", 0), Handler)
|
||||
dead = "http://127.0.0.1:%d" % s.server_address[1]
|
||||
s.server_close()
|
||||
rc, out = run([sys.executable, gate], ROOT, child_env(dead))
|
||||
report("published: INCONCLUSIVE: Gitea unreachable", rc, out, 2, ("INCONCLUSIVE", "URLs tried"))
|
||||
|
||||
|
||||
# ── release-complete ─────────────────────────────────────────────────────────────────────────────
|
||||
def release_cases(base, ws):
|
||||
bare = os.path.join(ws, "origin.git")
|
||||
work = os.path.join(ws, "rc-work")
|
||||
sh(["git", "clone", "-q", "--bare", "--no-tags", "file://" + ROOT, bare], ws)
|
||||
sh(["git", "clone", "-q", "--no-tags", "file://" + bare, work], ws)
|
||||
sh(["git", "config", "user.email", "decoy@gate.invalid"], work)
|
||||
sh(["git", "config", "user.name", "decoy"], work)
|
||||
# the WORKING-TREE gate, so the file under test is the one being edited, not HEAD's
|
||||
shutil.copy(os.path.join(ROOT, "scripts", "check-release-complete.py"),
|
||||
os.path.join(work, "scripts", "check-release-complete.py"))
|
||||
sh(["git", "add", "scripts/check-release-complete.py"], work)
|
||||
sh(["git", "commit", "-q", "--allow-empty", "-m", "the gate under test"], work)
|
||||
base_sha = sh(["git", "rev-parse", "HEAD"], work)
|
||||
ch = os.path.join(work, "CHANGELOG.md")
|
||||
original = io.open(ch, encoding="utf-8").read()
|
||||
V = "9.9.9"
|
||||
|
||||
def case(name, top, expect_rc, must=(), tag=None, origin_tag=False, packaged=True, pkg_status=None):
|
||||
FAKE.reset()
|
||||
if packaged:
|
||||
FAKE.packages = {V}
|
||||
FAKE.pkg_status = pkg_status
|
||||
try:
|
||||
io.open(ch, "w", encoding="utf-8").write(top + original)
|
||||
sh(["git", "commit", "-q", "-am", name], work)
|
||||
if tag == "head":
|
||||
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", "HEAD"], work)
|
||||
elif tag == "unrelated":
|
||||
empty = sh(["git", "mktree"], work) # stdin is "" — the empty tree, written to this repo
|
||||
orphan = sh(["git", "commit-tree", "-m", "unrelated", empty], work)
|
||||
sh(["git", "tag", "-a", "v" + V, "-m", "decoy", orphan], work)
|
||||
if origin_tag:
|
||||
sh(["git", "push", "-q", "origin", "HEAD:refs/tags/v" + V], work)
|
||||
rc, out = run([sys.executable, os.path.join(work, "scripts", "check-release-complete.py")],
|
||||
work, child_env(base))
|
||||
report("release-complete: " + name, rc, out, expect_rc, must)
|
||||
finally:
|
||||
run(["git", "tag", "-d", "v" + V], work)
|
||||
run(["git", "push", "-q", "origin", ":refs/tags/v" + V], work)
|
||||
sh(["git", "reset", "-q", "--hard", base_sha], work)
|
||||
|
||||
HEAD = "## v%s — 2026-10-06\n\n- decoy release\n\n" % V
|
||||
case("GENUINE: tagged at HEAD and published", HEAD, 0,
|
||||
("newest CHANGELOG version: v9.9.9", "is tagged, placed and published"), tag="head")
|
||||
case("GENUINE: an `## Unreleased` heading above the release", "## Unreleased\n\n- wip\n\n" + HEAD, 0,
|
||||
("newest CHANGELOG version: v9.9.9",), tag="head")
|
||||
case("GENUINE: a newer version named only in prose and under ###",
|
||||
"The `## v10.0.0` heading is not written yet.\n### v10.0.0 notes\n\n" + HEAD, 0,
|
||||
("newest CHANGELOG version: v9.9.9",), tag="head")
|
||||
case("GENUINE: the tag only on origin (the shallow-CI shape)", HEAD, 0,
|
||||
("exists on origin",), origin_tag=True)
|
||||
case("BY DESIGN: a LOCAL-only tag passes (CI's fresh clone sees only origin)", HEAD, 0,
|
||||
("an ancestor of HEAD",), tag="head")
|
||||
case("FACT: the newest heading has no tag anywhere", HEAD, 1, ("DOES NOT EXIST",))
|
||||
case("FACT: a tag parked on an unrelated commit", HEAD, 1, ("NOT an ancestor",), tag="unrelated")
|
||||
case("FACT: tagged, never published", HEAD, 1, ("IS NOT PUBLISHED",), tag="head", packaged=False)
|
||||
case("INCONCLUSIVE: the registry answers 500", HEAD, 2, ("INCONCLUSIVE",), tag="head", pkg_status=500)
|
||||
case("FACT beats INCONCLUSIVE: no tag AND the registry answers 500", HEAD, 1, ("DOES NOT EXIST",),
|
||||
pkg_status=500)
|
||||
|
||||
|
||||
# ── the shared felhom.eu scripts, against THIS repo's inputs ─────────────────────────────────────
|
||||
def shared_cases(ws):
|
||||
"""A scratch WORKSPACE: a clone of this repo beside symlinks to the siblings, because the shared
|
||||
scripts reach across (REUSE.md cites hub paths; instructions_gate reads the workspace CLAUDE.md)."""
|
||||
for g in ("reuse_refs_check.py", "instructions_gate.py", "observations_gate.py"):
|
||||
if not os.path.isfile(os.path.join(SHARED, g)):
|
||||
fails.append("shared gate %s is MISSING beside this clone (tried %s) - a failure, never a skip"
|
||||
% (g, SHARED))
|
||||
return
|
||||
space = os.path.join(ws, "workspace")
|
||||
os.makedirs(space)
|
||||
for entry in sorted(os.listdir(PARENT)):
|
||||
if entry in ("felhom.eu", "felhom-controller", "app-catalog-felhom.eu", "homelab-manifests",
|
||||
"CLAUDE.md", ".claude-memory"):
|
||||
os.symlink(os.path.join(PARENT, entry), os.path.join(space, entry))
|
||||
repo = os.path.join(space, "felhom-agent")
|
||||
sh(["git", "clone", "-q", "--no-tags", "file://" + ROOT, repo], ws)
|
||||
# the WORKING-TREE inputs the plants go into, so a case judges today's file
|
||||
for f in ("REUSE.md", "CLAUDE.md", "REPORT.md"):
|
||||
shutil.copy(os.path.join(ROOT, f), os.path.join(repo, f))
|
||||
|
||||
def case(name, gate, relpath, extra, expect_rc, must=()):
|
||||
p = os.path.join(repo, relpath)
|
||||
backup = io.open(p, encoding="utf-8").read()
|
||||
try:
|
||||
if extra:
|
||||
io.open(p, "w", encoding="utf-8").write(backup + extra)
|
||||
rc, out = run([sys.executable, os.path.join(SHARED, gate), repo], repo)
|
||||
report(name, rc, out, expect_rc, must)
|
||||
finally:
|
||||
io.open(p, "w", encoding="utf-8").write(backup)
|
||||
|
||||
case("reuse-refs: GENUINE: this repo's REUSE.md", "reuse_refs_check.py", "REUSE.md", "", 0, ("FAILED 0",))
|
||||
case("reuse-refs: FACT: a cited .go path that does not exist", "reuse_refs_check.py", "REUSE.md",
|
||||
u"\n- see `internal/localapi/does_not_exist.go`\n", 1, ("does_not_exist.go",))
|
||||
case("reuse-refs: FACT: a cited .md path that does not exist", "reuse_refs_check.py", "REUSE.md",
|
||||
u"\n- see `docs/99-does-not-exist.md`\n", 1, ("99-does-not-exist.md",))
|
||||
case("instructions: GENUINE: this repo's CLAUDE.md", "instructions_gate.py", "CLAUDE.md", "", 0,
|
||||
("instructions_gate: OK",))
|
||||
case("instructions: FACT: a version literal in effective text", "instructions_gate.py", "CLAUDE.md",
|
||||
u"\nThe agent runs v0.148.0 today.\n", 1, ("v0.148.0",))
|
||||
case("instructions: GENUINE: the same sentence in an HTML comment", "instructions_gate.py", "CLAUDE.md",
|
||||
u"\n<!--\nThe agent ran v0.148.0 on 2026-10-06.\n-->\n", 0, ("instructions_gate: OK",))
|
||||
case("observations: FACT: R-419, prose SAYING it has no marker", "observations_gate.py", "REPORT.md",
|
||||
u"\n## Observations\n\n1. **A real finding.** It carries no `FILED:` marker and no "
|
||||
u"`NOT-A-FINDING:` marker, deliberately.\n", 1)
|
||||
case("observations: GENUINE: a FILED marker", "observations_gate.py", "REPORT.md",
|
||||
u"\n## Observations\n\n1. **A real finding.** Something broke. **FILED: R-419**\n", 0)
|
||||
case("observations: GENUINE: a NOT-A-FINDING marker", "observations_gate.py", "REPORT.md",
|
||||
u"\n## Observations\n\n1. **A real finding.** Odd. **NOT-A-FINDING: my own typo, corrected in "
|
||||
u"the same minute.**\n", 0)
|
||||
|
||||
|
||||
def main():
|
||||
srv = Server(("127.0.0.1", 0), Handler)
|
||||
threading.Thread(target=srv.serve_forever, daemon=True).start()
|
||||
base = "http://127.0.0.1:%d" % srv.server_address[1]
|
||||
ws = tempfile.mkdtemp(prefix="agent-decoys-")
|
||||
print("agent gate decoys — fake Gitea at %s, scratch %s" % (base, ws))
|
||||
try:
|
||||
published_cases(base)
|
||||
release_cases(base, ws)
|
||||
shared_cases(ws)
|
||||
finally:
|
||||
srv.shutdown()
|
||||
srv.server_close()
|
||||
shutil.rmtree(ws, ignore_errors=True)
|
||||
if fails:
|
||||
print()
|
||||
for f in fails:
|
||||
print("FAIL: %s" % f)
|
||||
return 1
|
||||
print("\nagent gate decoys OK — %d case(s), every label judged on its fact (R-421)" % ran)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Reference in New Issue
Block a user