R-444: weekly guest disk trim (pct fstrim) outside the night, under the heavy-op gate
Operator ruling 09 §3 decision 139. One exact sudoers rule FELHOM_FSTRIM
(`/usr/sbin/pct ^fstrim [0-9]+$`) + manifest entry guest-fstrim; new
internal/fstrim job: due Wednesday from 10:00 host-local, starts only
10:00-20:59, holds backup.InFlight (busy -> deferred to the next hourly
tick), failed trim retried at most 3x per week, bytes parsed from the
"(N bytes) trimmed" lines, last result per guest persisted in
<state_dir>/guest-disk-trim.json and reported as guest_disk_trim.
Config opt-out: "disk_trim": {"disable": true}.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -160,6 +160,7 @@
|
||||
| `lanresolver.hostRoot` + `dnsmasqUnitPaths` (data seam, R-317) | internal/lanresolver/lanresolver.go | prod `hostRoot = "/"`; probe = the `dnsmasq` package's systemd UNIT, never `/usr/sbin/dnsmasq` (owned by `dnsmasq-base`) | internal/lanresolver/ensure_dnsmasq_test.go — fixture root tree + recording `proxmox.Runner`; the REAL `os.Stat` probe and `EnsureDnsmasq` run. `TestEnsureDnsmasq_ProductionProbeIsTheUnit` pins the production wiring |
|
||||
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
|
||||
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
|
||||
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
|
||||
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
|
||||
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
|
||||
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
|
||||
|
||||
@@ -0,0 +1,53 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"go/ast"
|
||||
"go/parser"
|
||||
"go/token"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-444: the weekly trim has the guestnet shape (component + reporter seam + goroutine), so its wiring is asserted
|
||||
// from the AST like TestMainWiresGuestNetWatchdog — a unit-green trim job that main.go never starts is the inert-seam
|
||||
// defect. It must also share the ONE heavy-op gate (heavyOps), or it could run beside a backup.
|
||||
func TestMainWiresGuestDiskTrim(t *testing.T) {
|
||||
fset := token.NewFileSet()
|
||||
f, err := parser.ParseFile(fset, "main.go", nil, 0)
|
||||
if err != nil {
|
||||
t.Fatalf("parse main.go: %v", err)
|
||||
}
|
||||
var constructedWithGate, reporterWired, started bool
|
||||
ast.Inspect(f, func(n ast.Node) bool {
|
||||
switch node := n.(type) {
|
||||
case *ast.CallExpr:
|
||||
if fn, ok := node.Fun.(*ast.SelectorExpr); ok {
|
||||
switch fn.Sel.Name {
|
||||
case "New":
|
||||
if pkg, ok := fn.X.(*ast.Ident); ok && pkg.Name == "fstrim" && len(node.Args) >= 3 {
|
||||
if id, ok := node.Args[2].(*ast.Ident); ok && id.Name == "heavyOps" {
|
||||
constructedWithGate = true
|
||||
}
|
||||
}
|
||||
case "SetGuestDiskTrimReporter":
|
||||
reporterWired = true
|
||||
}
|
||||
}
|
||||
case *ast.GoStmt:
|
||||
if sel, ok := node.Call.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "Run" {
|
||||
if id, ok := sel.X.(*ast.Ident); ok && id.Name == "diskTrim" {
|
||||
started = true
|
||||
}
|
||||
}
|
||||
}
|
||||
return true
|
||||
})
|
||||
if !constructedWithGate {
|
||||
t.Error("main.go never calls fstrim.New(..., heavyOps, ...) — no trim job, or one outside the heavy-op gate")
|
||||
}
|
||||
if !reporterWired {
|
||||
t.Error("main.go never calls collector.SetGuestDiskTrimReporter — the guest_disk_trim stanza never reaches the hub")
|
||||
}
|
||||
if !started {
|
||||
t.Error("main.go never starts the trim job with `go diskTrim.Run(ctx)`")
|
||||
}
|
||||
}
|
||||
@@ -38,6 +38,7 @@ import (
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/fasttick"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/fstrim"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/guesthook"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/guestnet"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
@@ -1479,6 +1480,25 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
}
|
||||
go runJanitor(ctx, jd)
|
||||
}
|
||||
// R-444 (`09` §3 decision 139): the weekly guest disk trim — `pct fstrim <vmid>` of every owned, running guest,
|
||||
// Wednesday from 10:00 local, daytime only, under the one-heavy-op gate. Not part of the errc fan-out: a trim job
|
||||
// must never be able to bring the agent down.
|
||||
if cfg.DiskTrim.Enabled() {
|
||||
dtMode := proxmox.RunnerMode(cfg.Privileged.Mode)
|
||||
if dtMode == "" {
|
||||
dtMode = proxmox.RunnerSudo
|
||||
}
|
||||
dtRunner := &proxmox.ExecRunner{Mode: dtMode, SudoPath: cfg.Privileged.SudoPath}
|
||||
dtGuests := localapi.NewStaleLockController(px, dtRunner, reconcile.DefaultPool, logger)
|
||||
if dtGuests != nil {
|
||||
diskTrim := fstrim.New(dtRunner, dtGuests, heavyOps,
|
||||
filepath.Join(cfg.OOB.WithDefaults().StateDir, "guest-disk-trim.json"), logger)
|
||||
collector.SetGuestDiskTrimReporter(diskTrim)
|
||||
go diskTrim.Run(ctx)
|
||||
}
|
||||
} else {
|
||||
logger.Info("fstrim: weekly guest disk trim disabled by config (disk_trim.disable)")
|
||||
}
|
||||
if lanLoop != nil {
|
||||
lanServers = 1
|
||||
go func() { errc <- lanLoop.Run(ctx) }()
|
||||
|
||||
@@ -129,6 +129,14 @@ Cmnd_Alias FELHOM_CONTROLLERSWAP = \
|
||||
Cmnd_Alias FELHOM_STALELOCK = \
|
||||
/usr/sbin/pct ^unlock [0-9]+$
|
||||
|
||||
# Weekly guest disk trim (R-444, operator ruling `09` §3 decision 139). A thin pool only ever grows from blocks the
|
||||
# guest has FREED: `fstrim` inside the unprivileged container is refused (FITRIM: Operation not permitted), so the host
|
||||
# trims the guest's mounts. Measured on demo-hp 2026-10-06: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %,
|
||||
# apps kept answering. ONE exact pattern: a vmid and nothing else — no `--ignore-mountpoints`, no second argument
|
||||
# (pinned: TestSudoersFstrimRuleIsExact). The agent runs it on a weekly daytime timer under the heavy-op gate.
|
||||
Cmnd_Alias FELHOM_FSTRIM = \
|
||||
/usr/sbin/pct ^fstrim [0-9]+$
|
||||
|
||||
# Restore-test scratch teardown (F-LEAK, Campaign 8, v0.110.0). A restore-test whose restore FAILS
|
||||
# leaves a scratch guest the API token CANNOT destroy: `FelhomAgentGuest` is granted at /pool/felhom and
|
||||
# a guest joins that pool only when its restore COMPLETES, so a failed restore leaves a pool-less guest
|
||||
@@ -304,4 +312,4 @@ Cmnd_Alias FELHOM_GUESTNET = \
|
||||
/usr/sbin/pct ^exec [0-9]+ -- pgrep -x dhclient$, \
|
||||
/usr/sbin/pct ^exec [0-9]+ -- dhclient -pf /run/dhclient\.eth0\.pid -lf /var/lib/dhcp/dhclient\.eth0\.leases eth0$
|
||||
|
||||
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
|
||||
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_FSTRIM, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
|
||||
|
||||
@@ -151,6 +151,10 @@ var manifest = []Capability{
|
||||
// reboot-during-backup lock can't start → the customer box stays DOWN until this clears it) ----
|
||||
{"stalelock-unlock", "reboot-during-backup stale-lock recovery", "/usr/sbin/pct", []string{"unlock", "9201"}, true, ""},
|
||||
|
||||
// ---- Weekly guest disk trim (FELHOM_FSTRIM, R-444). NON-critical: a missing grant means the thin pool is not
|
||||
// reclaimed this week (the trim job WARNs per guest and the report shows the failure), not a serving outage. ----
|
||||
{"guest-fstrim", "weekly guest disk trim (thin-pool reclaim, R-444)", "/usr/sbin/pct", []string{"fstrim", "9201"}, false, ""},
|
||||
|
||||
// ---- Offsite WG tunnel (FELHOM_WG, S3/v0.64.0; Critical FLIPPED in S4/v0.66.0 — offsite
|
||||
// backups now RIDE the tunnel, so a degraded tunnel capability is operator-alert-worthy: the
|
||||
// conf install, unit enable/restart and the handshake read gate the backup path. apt-install
|
||||
|
||||
@@ -63,3 +63,33 @@ func TestSudoersRefusesTheR861Injections(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// R-444: the weekly trim's grant is ONE exact shape — `pct fstrim <vmid>` — and nothing smuggled after it.
|
||||
// The manifest entry (guest-fstrim) proves the real call is still allowed (TestManifestCoveredBySudoers); this
|
||||
// pins the other direction. RED-PROOF: write the rule as the glob `/usr/sbin/pct fstrim [0-9]*` → every decoy
|
||||
// below with a trailing argument matches (the glob's `*` eats spaces).
|
||||
func TestSudoersFstrimRuleIsExact(t *testing.T) {
|
||||
data, err := os.ReadFile(sudoersPath)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
entries := parseSudoersEntries(t, string(data))
|
||||
if !matchesAny("/usr/sbin/pct fstrim 9201", entries) {
|
||||
t.Fatal("the sudoers does not allow `pct fstrim 9201` — the weekly trim cannot run")
|
||||
}
|
||||
for _, c := range []string{
|
||||
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints",
|
||||
"/usr/sbin/pct fstrim 9201 --ignore-mountpoints 1",
|
||||
"/usr/sbin/pct fstrim 9201; x",
|
||||
"/usr/sbin/pct fstrim 9201 9202",
|
||||
"/usr/sbin/pct fstrim 92a1",
|
||||
"/usr/sbin/pct fstrim ",
|
||||
"/usr/sbin/pct fstrim -- 9201",
|
||||
"/usr/sbin/pct destroy 9201",
|
||||
"/usr/sbin/pct destroy 9201 --purge",
|
||||
} {
|
||||
if matchesAny(c, entries) {
|
||||
t.Errorf("the sudoers allows a command the trim rule must not: %q", c)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -33,6 +33,7 @@ type Config struct {
|
||||
LANResolver LANResolverConfig `json:"lan_resolver"`
|
||||
WGTunnel WGTunnelConfig `json:"wg_tunnel"`
|
||||
GuestNet GuestNetConfig `json:"guest_net"`
|
||||
DiskTrim DiskTrimConfig `json:"disk_trim"`
|
||||
OOB OOBConfig `json:"oob"`
|
||||
SelfUpdate SelfUpdateConfig `json:"selfupdate"`
|
||||
LogLevel string `json:"log_level"` // debug|info|warn|error (default info)
|
||||
@@ -139,6 +140,16 @@ func (w WGTunnelConfig) WithDefaults() WGTunnelConfig {
|
||||
return w
|
||||
}
|
||||
|
||||
// DiskTrimConfig configures the R-444 weekly guest disk trim (internal/fstrim). DEFAULT-ON, like GuestNetConfig and for
|
||||
// the same reason: it only acts on guests the agent already owns, and the operator ruled every box trims (`09` §3
|
||||
// decision 139). Opting out is the explicit act: `"disk_trim": {"disable": true}`.
|
||||
type DiskTrimConfig struct {
|
||||
Disable bool `json:"disable"`
|
||||
}
|
||||
|
||||
// Enabled reports whether the weekly trim should run.
|
||||
func (d DiskTrimConfig) Enabled() bool { return !d.Disable }
|
||||
|
||||
// GuestNetConfig configures the R-54 guest-network watchdog (internal/guestnet).
|
||||
//
|
||||
// **This is the repo's first DEFAULT-ON feature gate, and the inversion is deliberate.** Every other
|
||||
|
||||
@@ -0,0 +1,346 @@
|
||||
// Package fstrim is the weekly guest disk trim (R-444, operator ruling `09` §3 decision 139).
|
||||
//
|
||||
// Why: a thin pool only ever grows from blocks a guest has already FREED — `fstrim` inside the unprivileged container
|
||||
// is refused (FITRIM: Operation not permitted), and nothing else on the box gives the blocks back. A full thin pool
|
||||
// takes every guest on the host read-only, so the pool can reach 100 % from deleted data alone. Measured on demo-hp
|
||||
// 2026-10-06 09:14Z: `pct fstrim 9201` rc 0 in 24.4 s, pool 65.53 % -> 33.40 %, 18/18 app probes 200, max 1.1 s
|
||||
// (audits/ten-answers-2026-10-06/r444-measure.txt).
|
||||
//
|
||||
// The rule, each part pinned by a test in fstrim_test.go:
|
||||
// - Weekly: a guest is DUE from Wednesday 10:00 local until it has been trimmed once since then (a box that was off
|
||||
// on Wednesday catches up at its next eligible hour).
|
||||
// - Daytime only: a trim starts only between 10:00 and 20:59 local — never in the night window (01:00–06:59) where
|
||||
// the backups and the restore-tests run (TestEligibleHourNeverInTheNight).
|
||||
// - Never beside a backup, a restore-test or another heavy operation: the pass holds the host-wide one-heavy-op gate
|
||||
// (backup.InFlight) for its whole run; a busy gate DEFERS the pass to the next hourly tick.
|
||||
// - A failed trim is retried at the next eligible hour, at most MaxAttemptsPerWeek times in one week.
|
||||
// - The last result per guest (time, bytes, ok/fail) is persisted, so a restart neither loses it nor re-trims.
|
||||
//
|
||||
// The command is the ONE exact sudoers shape `pct fstrim <vmid>` (FELHOM_FSTRIM). Only guests from the pool-verified
|
||||
// source (ListLXC ∩ the felhom pool, audit A1) and only RUNNING ones are trimmed.
|
||||
package fstrim
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// Schedule. The weekday/hours are fixed on purpose (one sentence the operator can read on the System page).
|
||||
const (
|
||||
Weekday = time.Wednesday
|
||||
StartHour = 10 // first eligible local hour (inclusive)
|
||||
EndHour = 21 // first NOT-eligible local hour (exclusive): last start is 20:59
|
||||
MaxAttemptsPerWeek = 3
|
||||
// TickInterval is how often the job looks; a deferred or failed pass is therefore retried the next hour.
|
||||
TickInterval = time.Hour
|
||||
// FirstTickDelay lets the agent settle after a start before the first look.
|
||||
FirstTickDelay = 5 * time.Minute
|
||||
// PerGuestTimeout bounds one `pct fstrim` (measured 24.4 s for 84 GiB).
|
||||
PerGuestTimeout = 30 * time.Minute
|
||||
)
|
||||
|
||||
// ScheduleText is the human description carried on the host report.
|
||||
const ScheduleText = "weekly, due Wednesday from 10:00 host-local time; starts only 10:00-20:59; never beside a backup or restore-test"
|
||||
|
||||
// Runner runs a host command (proxmox.ExecRunner in production, through `sudo -n`).
|
||||
type Runner interface {
|
||||
Run(ctx context.Context, name string, args ...string) (stdout, stderr []byte, err error)
|
||||
}
|
||||
|
||||
// GuestSource yields the guests this agent OWNS (the pool-verified source, never a bare ListLXC).
|
||||
type GuestSource interface {
|
||||
Guests(ctx context.Context) ([]proxmox.Guest, error)
|
||||
}
|
||||
|
||||
// Gate is the host-wide one-heavy-operation gate (*backup.InFlight).
|
||||
type Gate interface {
|
||||
TryAcquire(what string) (release func(), busy string, ok bool)
|
||||
}
|
||||
|
||||
// GateName is what the gate reports as busy while a trim runs.
|
||||
const GateName = "guest-fstrim"
|
||||
|
||||
// Record is one guest's last trim attempt, as persisted.
|
||||
type Record struct {
|
||||
LastAttemptAt time.Time `json:"last_attempt_at"`
|
||||
OK bool `json:"ok"`
|
||||
BytesTrimmed int64 `json:"bytes_trimmed"`
|
||||
Mounts int `json:"mounts"`
|
||||
DurationSeconds float64 `json:"duration_seconds"`
|
||||
LastOKAt time.Time `json:"last_ok_at,omitempty"`
|
||||
Error string `json:"error,omitempty"`
|
||||
// Attempts counts the attempts since the current week's due time (reset by the first attempt of a new week).
|
||||
Attempts int `json:"attempts"`
|
||||
}
|
||||
|
||||
// Trimmer is the weekly job.
|
||||
type Trimmer struct {
|
||||
runner Runner
|
||||
guests GuestSource
|
||||
gate Gate
|
||||
statePath string
|
||||
logger *slog.Logger
|
||||
loc *time.Location
|
||||
now func() time.Time
|
||||
|
||||
mu sync.Mutex
|
||||
records map[int]Record
|
||||
}
|
||||
|
||||
// New builds the job and loads the persisted state. A missing state file is an empty state; a corrupt one is logged
|
||||
// and treated as empty (the cost is one extra trim, never a missed one).
|
||||
func New(runner Runner, guests GuestSource, gate Gate, statePath string, logger *slog.Logger) *Trimmer {
|
||||
if logger == nil {
|
||||
logger = slog.Default()
|
||||
}
|
||||
t := &Trimmer{runner: runner, guests: guests, gate: gate, statePath: statePath, logger: logger,
|
||||
loc: time.Local, now: time.Now, records: map[int]Record{}}
|
||||
t.load()
|
||||
return t
|
||||
}
|
||||
|
||||
func (t *Trimmer) load() {
|
||||
data, err := os.ReadFile(t.statePath)
|
||||
if err != nil {
|
||||
if !os.IsNotExist(err) {
|
||||
t.logger.Warn("fstrim: state read failed — starting empty", "path", t.statePath, "err", err)
|
||||
}
|
||||
return
|
||||
}
|
||||
var raw map[string]Record
|
||||
if err := json.Unmarshal(data, &raw); err != nil {
|
||||
t.logger.Warn("fstrim: state file corrupt — starting empty", "path", t.statePath, "err", err)
|
||||
return
|
||||
}
|
||||
for k, r := range raw {
|
||||
if id, err := strconv.Atoi(k); err == nil && id > 0 {
|
||||
t.records[id] = r
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func (t *Trimmer) saveLocked() error {
|
||||
raw := make(map[string]Record, len(t.records))
|
||||
for id, r := range t.records {
|
||||
raw[strconv.Itoa(id)] = r
|
||||
}
|
||||
data, err := json.MarshalIndent(raw, "", " ")
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := os.MkdirAll(filepath.Dir(t.statePath), 0o755); err != nil {
|
||||
return err
|
||||
}
|
||||
tmp := t.statePath + ".tmp"
|
||||
if err := os.WriteFile(tmp, data, 0o600); err != nil {
|
||||
os.Remove(tmp)
|
||||
return err
|
||||
}
|
||||
return os.Rename(tmp, t.statePath)
|
||||
}
|
||||
|
||||
// EligibleHour reports whether a trim may START at local time lt.
|
||||
func EligibleHour(lt time.Time) bool {
|
||||
h := lt.Hour()
|
||||
return h >= StartHour && h < EndHour
|
||||
}
|
||||
|
||||
// weekAnchor is the most recent Wednesday StartHour:00 at or before lt (same location as lt).
|
||||
func weekAnchor(lt time.Time) time.Time {
|
||||
daysBack := (int(lt.Weekday()) - int(Weekday) + 7) % 7
|
||||
d := lt.AddDate(0, 0, -daysBack)
|
||||
a := time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
|
||||
if a.After(lt) {
|
||||
d = d.AddDate(0, 0, -7)
|
||||
a = time.Date(d.Year(), d.Month(), d.Day(), StartHour, 0, 0, 0, lt.Location())
|
||||
}
|
||||
return a
|
||||
}
|
||||
|
||||
// due reports whether a guest with record r (ok=false: none) is due at local time lt.
|
||||
func due(r Record, has bool, lt time.Time) bool {
|
||||
if !has {
|
||||
return true
|
||||
}
|
||||
anchor := weekAnchor(lt)
|
||||
if r.LastAttemptAt.Before(anchor) {
|
||||
return true // not tried this week
|
||||
}
|
||||
return !r.OK && r.Attempts < MaxAttemptsPerWeek
|
||||
}
|
||||
|
||||
// Run looks every TickInterval until ctx ends. It never returns an error: a failed trim is a reported fact.
|
||||
func (t *Trimmer) Run(ctx context.Context) {
|
||||
t.logger.Info("fstrim: weekly guest disk trim starting", "schedule", ScheduleText)
|
||||
timer := time.NewTimer(FirstTickDelay)
|
||||
defer timer.Stop()
|
||||
for {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-timer.C:
|
||||
t.Pass(ctx)
|
||||
timer.Reset(TickInterval)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Pass is one look: outside the daytime window it does nothing; otherwise it trims every due, running, owned guest
|
||||
// while holding the heavy-op gate.
|
||||
func (t *Trimmer) Pass(ctx context.Context) {
|
||||
lt := t.now().In(t.loc)
|
||||
if !EligibleHour(lt) {
|
||||
t.logger.Debug("fstrim: outside the daytime window — not looking", "local", lt.Format("Mon 15:04"))
|
||||
return
|
||||
}
|
||||
guests, err := t.guests.Guests(ctx)
|
||||
if err != nil {
|
||||
t.logger.Warn("fstrim: owned-guest list unavailable — skipping this pass", "err", err)
|
||||
return
|
||||
}
|
||||
owned := make(map[int]bool, len(guests))
|
||||
var todo []int
|
||||
t.mu.Lock()
|
||||
for _, g := range guests {
|
||||
owned[g.VMID] = true
|
||||
r, has := t.records[g.VMID]
|
||||
if !due(r, has, lt) {
|
||||
continue
|
||||
}
|
||||
if g.Status != "running" {
|
||||
t.logger.Info("fstrim: guest not running — trimmed when it runs", "vmid", g.VMID, "status", g.Status)
|
||||
continue
|
||||
}
|
||||
todo = append(todo, g.VMID)
|
||||
}
|
||||
// A guest the agent no longer owns has no result to report.
|
||||
pruned := false
|
||||
for id := range t.records {
|
||||
if !owned[id] {
|
||||
delete(t.records, id)
|
||||
pruned = true
|
||||
}
|
||||
}
|
||||
if pruned {
|
||||
if err := t.saveLocked(); err != nil {
|
||||
t.logger.Warn("fstrim: state save failed", "err", err)
|
||||
}
|
||||
}
|
||||
t.mu.Unlock()
|
||||
if len(todo) == 0 {
|
||||
return
|
||||
}
|
||||
sort.Ints(todo)
|
||||
release, busy, ok := t.gate.TryAcquire(GateName)
|
||||
if !ok {
|
||||
t.logger.Info("fstrim: deferred — a heavy operation is in flight; retrying next hour", "busy", busy, "due_guests", len(todo))
|
||||
return
|
||||
}
|
||||
defer release()
|
||||
for _, vmid := range todo {
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
t.trimOne(ctx, vmid, lt)
|
||||
}
|
||||
}
|
||||
|
||||
var trimmedLine = regexp.MustCompile(`\((\d+) bytes\) trimmed`)
|
||||
|
||||
// ParseTrimmed sums the "(N bytes) trimmed" lines of `pct fstrim` output and counts them (one per mount point), e.g.
|
||||
// `/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed`.
|
||||
func ParseTrimmed(out string) (bytes int64, mounts int) {
|
||||
for _, m := range trimmedLine.FindAllStringSubmatch(out, -1) {
|
||||
n, err := strconv.ParseInt(m[1], 10, 64)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
bytes += n
|
||||
mounts++
|
||||
}
|
||||
return bytes, mounts
|
||||
}
|
||||
|
||||
// GiB renders bytes as "30.1 GiB".
|
||||
func GiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
|
||||
|
||||
func (t *Trimmer) trimOne(ctx context.Context, vmid int, lt time.Time) {
|
||||
start := t.now()
|
||||
cctx, cancel := context.WithTimeout(ctx, PerGuestTimeout)
|
||||
stdout, stderr, err := t.runner.Run(cctx, "pct", "fstrim", strconv.Itoa(vmid))
|
||||
cancel()
|
||||
dur := t.now().Sub(start)
|
||||
bytes, mounts := ParseTrimmed(string(stdout) + "\n" + string(stderr))
|
||||
|
||||
t.mu.Lock()
|
||||
prev, has := t.records[vmid]
|
||||
r := Record{LastAttemptAt: start.UTC(), OK: err == nil, BytesTrimmed: bytes, Mounts: mounts,
|
||||
DurationSeconds: float64(dur.Round(100*time.Millisecond)) / float64(time.Second), LastOKAt: prev.LastOKAt}
|
||||
if has && !prev.LastAttemptAt.Before(weekAnchor(lt)) {
|
||||
r.Attempts = prev.Attempts + 1
|
||||
} else {
|
||||
r.Attempts = 1
|
||||
}
|
||||
if err == nil {
|
||||
r.LastOKAt = start.UTC()
|
||||
} else {
|
||||
msg := strings.TrimSpace(err.Error() + ": " + strings.TrimSpace(string(stderr)))
|
||||
if len(msg) > 300 {
|
||||
msg = msg[:300]
|
||||
}
|
||||
r.Error = msg
|
||||
}
|
||||
t.records[vmid] = r
|
||||
saveErr := t.saveLocked()
|
||||
t.mu.Unlock()
|
||||
|
||||
if err == nil {
|
||||
t.logger.Info(fmt.Sprintf("fstrim: guest %d trimmed %s in %.1fs", vmid, GiB(bytes), r.DurationSeconds),
|
||||
"vmid", vmid, "bytes_trimmed", bytes, "mounts", mounts, "duration_s", r.DurationSeconds)
|
||||
if mounts == 0 {
|
||||
t.logger.Warn("fstrim: pct fstrim succeeded but reported no trimmed mount — output not understood",
|
||||
"vmid", vmid, "stdout", strings.TrimSpace(string(stdout)))
|
||||
}
|
||||
} else {
|
||||
t.logger.Warn(fmt.Sprintf("fstrim: guest %d trim FAILED after %.1fs", vmid, r.DurationSeconds),
|
||||
"vmid", vmid, "attempt", r.Attempts, "max_attempts_per_week", MaxAttemptsPerWeek, "err", r.Error)
|
||||
}
|
||||
if saveErr != nil {
|
||||
t.logger.Warn("fstrim: state save failed — the result will not survive a restart", "path", t.statePath, "err", saveErr)
|
||||
}
|
||||
}
|
||||
|
||||
// GuestDiskTrimStatus implements hub.GuestDiskTrimReporter: a pure read of the persisted results (never runs pct).
|
||||
func (t *Trimmer) GuestDiskTrimStatus(context.Context) *hub.GuestDiskTrimStatus {
|
||||
t.mu.Lock()
|
||||
defer t.mu.Unlock()
|
||||
out := &hub.GuestDiskTrimStatus{Schedule: ScheduleText}
|
||||
ids := make([]int, 0, len(t.records))
|
||||
for id := range t.records {
|
||||
ids = append(ids, id)
|
||||
}
|
||||
sort.Ints(ids)
|
||||
for _, id := range ids {
|
||||
r := t.records[id]
|
||||
g := hub.GuestDiskTrim{VMID: id, LastAttemptAt: r.LastAttemptAt.UTC().Format(time.RFC3339), OK: r.OK,
|
||||
BytesTrimmed: r.BytesTrimmed, Mounts: r.Mounts, DurationSeconds: r.DurationSeconds, Error: r.Error}
|
||||
if !r.LastOKAt.IsZero() {
|
||||
g.LastOKAt = r.LastOKAt.UTC().Format(time.RFC3339)
|
||||
}
|
||||
out.Guests = append(out.Guests, g)
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,267 @@
|
||||
package fstrim
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"log/slog"
|
||||
"path/filepath"
|
||||
"reflect"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// The real `pct fstrim 9201` output measured on demo-hp 2026-10-06 (audits/ten-answers-2026-10-06/r444-measure.txt).
|
||||
const measuredOut = "/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed\n" +
|
||||
"/var/lib/lxc/9201/rootfs/var/lib/felhom: 53.9 GiB (57865633792 bytes) trimmed\n"
|
||||
|
||||
const measuredBytes = int64(32277680128 + 57865633792)
|
||||
|
||||
type fakeRunner struct {
|
||||
mu sync.Mutex
|
||||
calls [][]string
|
||||
out string
|
||||
err error
|
||||
onRun func()
|
||||
}
|
||||
|
||||
func (f *fakeRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
|
||||
f.mu.Lock()
|
||||
f.calls = append(f.calls, append([]string{name}, args...))
|
||||
f.mu.Unlock()
|
||||
if f.onRun != nil {
|
||||
f.onRun()
|
||||
}
|
||||
if f.err != nil {
|
||||
return nil, []byte("mount busy"), f.err
|
||||
}
|
||||
return []byte(f.out), nil, nil
|
||||
}
|
||||
|
||||
type fakeGuests struct {
|
||||
g []proxmox.Guest
|
||||
err error
|
||||
}
|
||||
|
||||
func (f fakeGuests) Guests(context.Context) ([]proxmox.Guest, error) { return f.g, f.err }
|
||||
|
||||
// A Wednesday 10:30 in a fixed zone (CEST-like), so the tests do not depend on the machine's zone.
|
||||
var zone = time.FixedZone("CEST", 2*3600)
|
||||
|
||||
func at(day, hour, min int) time.Time { return time.Date(2026, 10, day, hour, min, 0, 0, zone) } // 2026-10-07 = Wednesday
|
||||
|
||||
func newT(t *testing.T, r Runner, g GuestSource, gate Gate, now *time.Time) (*Trimmer, *bytes.Buffer, string) {
|
||||
t.Helper()
|
||||
var logs bytes.Buffer
|
||||
path := filepath.Join(t.TempDir(), "guest-disk-trim.json")
|
||||
tr := New(r, g, gate, path, slog.New(slog.NewTextHandler(&logs, &slog.HandlerOptions{Level: slog.LevelDebug})))
|
||||
tr.loc = zone
|
||||
tr.now = func() time.Time { return *now }
|
||||
return tr, &logs, path
|
||||
}
|
||||
|
||||
func running(ids ...int) fakeGuests {
|
||||
var g []proxmox.Guest
|
||||
for _, id := range ids {
|
||||
g = append(g, proxmox.Guest{VMID: id, Status: "running", Type: "lxc"})
|
||||
}
|
||||
return fakeGuests{g: g}
|
||||
}
|
||||
|
||||
func TestParseTrimmedTheMeasuredOutput(t *testing.T) {
|
||||
b, m := ParseTrimmed(measuredOut)
|
||||
if b != measuredBytes || m != 2 {
|
||||
t.Fatalf("ParseTrimmed = %d bytes over %d mounts, want %d over 2", b, m, measuredBytes)
|
||||
}
|
||||
if b, m := ParseTrimmed("something else\n"); b != 0 || m != 0 {
|
||||
t.Fatalf("unrelated output parsed as %d/%d", b, m)
|
||||
}
|
||||
if got := GiB(measuredBytes); got != "84.0 GiB" {
|
||||
t.Fatalf("GiB = %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// The night window (01:00–06:59) must never be eligible, and the daytime window is exactly 10:00–20:59.
|
||||
func TestEligibleHourNeverInTheNight(t *testing.T) {
|
||||
for h := 0; h < 24; h++ {
|
||||
lt := time.Date(2026, 10, 7, h, 30, 0, 0, zone)
|
||||
got := EligibleHour(lt)
|
||||
if h >= 1 && h <= 6 && got {
|
||||
t.Errorf("hour %02d is in the night window and must not be eligible", h)
|
||||
}
|
||||
if want := h >= 10 && h <= 20; got != want {
|
||||
t.Errorf("EligibleHour(%02d:30) = %v, want %v", h, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestWeekAnchorIsTheLastWednesdayTen(t *testing.T) {
|
||||
cases := map[time.Time]time.Time{
|
||||
at(7, 10, 0): at(7, 10, 0), // Wednesday 10:00 itself
|
||||
at(7, 9, 59): time.Date(2026, 9, 30, 10, 0, 0, 0, zone), // before 10:00 Wednesday → the previous week
|
||||
at(8, 15, 0): at(7, 10, 0), // Thursday
|
||||
at(13, 20, 0): at(7, 10, 0), // next Tuesday
|
||||
at(14, 11, 0): at(14, 10, 0), // next Wednesday
|
||||
}
|
||||
for in, want := range cases {
|
||||
if got := weekAnchor(in); !got.Equal(want) {
|
||||
t.Errorf("weekAnchor(%s) = %s, want %s", in.Format("Mon 01-02 15:04"), got.Format("Mon 01-02 15:04"), want.Format("Mon 01-02 15:04"))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The consequence: on Wednesday 10:30 a running owned guest is trimmed with the ONE exact argv, the bytes are parsed,
|
||||
// the positive log line is written, the result is persisted, and the host report carries it.
|
||||
func TestPassTrimsADueGuestAndReportsIt(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
r := &fakeRunner{out: measuredOut}
|
||||
tr, logs, path := newT(t, r, running(9201), &backup.InFlight{}, &now)
|
||||
tr.Pass(context.Background())
|
||||
|
||||
if want := [][]string{{"pct", "fstrim", "9201"}}; !reflect.DeepEqual(r.calls, want) {
|
||||
t.Fatalf("runner calls = %q, want %q", r.calls, want)
|
||||
}
|
||||
if !strings.Contains(logs.String(), "fstrim: guest 9201 trimmed 84.0 GiB in ") {
|
||||
t.Fatalf("no positive per-guest log line:\n%s", logs.String())
|
||||
}
|
||||
st := tr.GuestDiskTrimStatus(context.Background())
|
||||
if st == nil || st.Schedule != ScheduleText || len(st.Guests) != 1 {
|
||||
t.Fatalf("report stanza = %+v", st)
|
||||
}
|
||||
g := st.Guests[0]
|
||||
if g.VMID != 9201 || !g.OK || g.BytesTrimmed != measuredBytes || g.Mounts != 2 || g.LastOKAt == "" || g.LastAttemptAt == "" {
|
||||
t.Fatalf("report guest = %+v", g)
|
||||
}
|
||||
// Persisted: a NEW Trimmer over the same file (an agent restart) still has it and does not trim again this week.
|
||||
now = at(8, 11, 0)
|
||||
r2 := &fakeRunner{out: measuredOut}
|
||||
tr2 := New(r2, running(9201), &backup.InFlight{}, path, slog.New(slog.NewTextHandler(&bytes.Buffer{}, nil)))
|
||||
tr2.loc, tr2.now = zone, func() time.Time { return now }
|
||||
if st2 := tr2.GuestDiskTrimStatus(context.Background()); len(st2.Guests) != 1 || st2.Guests[0].BytesTrimmed != measuredBytes {
|
||||
t.Fatalf("result lost over a restart: %+v", st2)
|
||||
}
|
||||
tr2.Pass(context.Background())
|
||||
if len(r2.calls) != 0 {
|
||||
t.Fatalf("trimmed again in the same week after a restart: %q", r2.calls)
|
||||
}
|
||||
// Next week it is due again.
|
||||
now = at(14, 10, 5)
|
||||
tr2.Pass(context.Background())
|
||||
if len(r2.calls) != 1 {
|
||||
t.Fatalf("not trimmed in the next week: %q", r2.calls)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPassNeverRunsInTheNight(t *testing.T) {
|
||||
for _, h := range []int{1, 3, 6, 9, 21, 23} {
|
||||
now := at(7, h, 15)
|
||||
r := &fakeRunner{out: measuredOut}
|
||||
tr, _, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
|
||||
tr.Pass(context.Background())
|
||||
if len(r.calls) != 0 {
|
||||
t.Errorf("trimmed at %02d:15: %q", h, r.calls)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A backup (or restore-test) holding the heavy-op gate DEFERS the trim; the next hour, gate free, it runs. And while
|
||||
// a trim runs, the gate is held, so a backup cannot start beside it.
|
||||
func TestPassDefersToAHeavyOperationAndRetriesNextHour(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
gate := &backup.InFlight{}
|
||||
release, _, _ := gate.TryAcquire("backup:9201")
|
||||
var busyDuringTrim string
|
||||
r := &fakeRunner{out: measuredOut}
|
||||
r.onRun = func() { busyDuringTrim = gate.Busy() }
|
||||
tr, logs, _ := newT(t, r, running(9201), gate, &now)
|
||||
|
||||
tr.Pass(context.Background())
|
||||
if len(r.calls) != 0 {
|
||||
t.Fatalf("trimmed beside a running backup: %q", r.calls)
|
||||
}
|
||||
if !strings.Contains(logs.String(), "fstrim: deferred") || !strings.Contains(logs.String(), "backup:9201") {
|
||||
t.Fatalf("the deferral is not logged with what holds the gate:\n%s", logs.String())
|
||||
}
|
||||
release()
|
||||
now = now.Add(time.Hour)
|
||||
tr.Pass(context.Background())
|
||||
if len(r.calls) != 1 {
|
||||
t.Fatalf("not retried the next hour: %q", r.calls)
|
||||
}
|
||||
if busyDuringTrim != GateName {
|
||||
t.Fatalf("the heavy-op gate was %q during the trim, want %q", busyDuringTrim, GateName)
|
||||
}
|
||||
if gate.Busy() != "" {
|
||||
t.Fatalf("the gate was not released after the pass: %q", gate.Busy())
|
||||
}
|
||||
}
|
||||
|
||||
func TestFailedTrimWarnsIsRecordedAndRetriedAtMostThreeTimes(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
r := &fakeRunner{err: errors.New("exit status 255")}
|
||||
tr, logs, _ := newT(t, r, running(9201), &backup.InFlight{}, &now)
|
||||
for i := 0; i < 6; i++ {
|
||||
tr.Pass(context.Background())
|
||||
now = now.Add(time.Hour)
|
||||
}
|
||||
if len(r.calls) != MaxAttemptsPerWeek {
|
||||
t.Fatalf("attempts in one week = %d, want %d", len(r.calls), MaxAttemptsPerWeek)
|
||||
}
|
||||
if !strings.Contains(logs.String(), "level=WARN") || !strings.Contains(logs.String(), "fstrim: guest 9201 trim FAILED") {
|
||||
t.Fatalf("no WARN for the failure:\n%s", logs.String())
|
||||
}
|
||||
g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]
|
||||
if g.OK || g.LastOKAt != "" || !strings.Contains(g.Error, "exit status 255") || !strings.Contains(g.Error, "mount busy") {
|
||||
t.Fatalf("failed result not recorded as a failure: %+v", g)
|
||||
}
|
||||
// A success later keeps a clean record.
|
||||
r.err = nil
|
||||
r.out = measuredOut
|
||||
now = at(14, 10, 10)
|
||||
tr.Pass(context.Background())
|
||||
if g := tr.GuestDiskTrimStatus(context.Background()).Guests[0]; !g.OK || g.Error != "" || g.BytesTrimmed != measuredBytes {
|
||||
t.Fatalf("success after failure: %+v", g)
|
||||
}
|
||||
}
|
||||
|
||||
func TestOnlyRunningOwnedGuestsAndAFailedListActsOnNothing(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
r := &fakeRunner{out: measuredOut}
|
||||
g := fakeGuests{g: []proxmox.Guest{{VMID: 9201, Status: "stopped"}, {VMID: 9202, Status: "running"}}}
|
||||
tr, _, _ := newT(t, r, g, &backup.InFlight{}, &now)
|
||||
tr.Pass(context.Background())
|
||||
if want := [][]string{{"pct", "fstrim", "9202"}}; !reflect.DeepEqual(r.calls, want) {
|
||||
t.Fatalf("calls = %q, want only the running guest", r.calls)
|
||||
}
|
||||
r2 := &fakeRunner{out: measuredOut}
|
||||
tr2, logs, _ := newT(t, r2, fakeGuests{err: errors.New("pool read 403")}, &backup.InFlight{}, &now)
|
||||
tr2.Pass(context.Background())
|
||||
if len(r2.calls) != 0 || !strings.Contains(logs.String(), "owned-guest list unavailable") {
|
||||
t.Fatalf("a failed ownership read must act on nothing: calls %q", r2.calls)
|
||||
}
|
||||
}
|
||||
|
||||
func TestReportJSONShape(t *testing.T) {
|
||||
now := at(7, 10, 30)
|
||||
tr, _, _ := newT(t, &fakeRunner{out: measuredOut}, running(9201), &backup.InFlight{}, &now)
|
||||
tr.Pass(context.Background())
|
||||
b, err := json.Marshal(tr.GuestDiskTrimStatus(context.Background()))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, k := range []string{`"schedule":`, `"guests":[{"vmid":9201`, `"last_attempt_at":"2026-10-07T08:30:00Z"`, `"ok":true`,
|
||||
`"bytes_trimmed":90143313920`, `"mounts":2`, `"duration_seconds":`, `"last_ok_at":"2026-10-07T08:30:00Z"`} {
|
||||
if !strings.Contains(string(b), k) {
|
||||
t.Errorf("report JSON lacks %s: %s", k, b)
|
||||
}
|
||||
}
|
||||
if strings.Contains(string(b), `"error"`) {
|
||||
t.Errorf("an ok result must omit error: %s", b)
|
||||
}
|
||||
}
|
||||
@@ -90,6 +90,12 @@ type GuestNetReporter interface {
|
||||
GuestNetStatus(ctx context.Context) *GuestNetStatus
|
||||
}
|
||||
|
||||
// GuestDiskTrimReporter is the R-444 seam the weekly trim job plugs into (same consumer-side pattern — hub does not
|
||||
// import fstrim). nil (feature not wired) → no guest_disk_trim stanza.
|
||||
type GuestDiskTrimReporter interface {
|
||||
GuestDiskTrimStatus(ctx context.Context) *GuestDiskTrimStatus
|
||||
}
|
||||
|
||||
// Collector builds a HostReport from read-only sources. All deps are behind narrow
|
||||
// interfaces for unit testing.
|
||||
type Collector struct {
|
||||
@@ -108,6 +114,7 @@ type Collector struct {
|
||||
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
|
||||
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
|
||||
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
|
||||
diskTrim GuestDiskTrimReporter // R-444: weekly guest disk trim (nil → stanza omitted)
|
||||
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
|
||||
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
|
||||
oob OOBReporter // H1: operator-access health (nil → stanza omitted)
|
||||
@@ -221,6 +228,12 @@ func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
|
||||
return c
|
||||
}
|
||||
|
||||
// SetGuestDiskTrimReporter wires the R-444 weekly trim job as a report source (nil-safe → stanza omitted).
|
||||
func (c *Collector) SetGuestDiskTrimReporter(r GuestDiskTrimReporter) *Collector {
|
||||
c.diskTrim = r
|
||||
return c
|
||||
}
|
||||
|
||||
// SelfUpdateReporter is the D1 seam the selfupdate commit-manager plugs into (same consumer-side
|
||||
// pattern — hub does not import selfupdate). nil (feature not wired) → pending=false on the report.
|
||||
type SelfUpdateReporter interface {
|
||||
@@ -377,6 +390,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
|
||||
if c.guestNet != nil {
|
||||
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
|
||||
}
|
||||
// R-444: the last weekly trim result per guest (nil reporter = not wired → stanza omitted).
|
||||
if c.diskTrim != nil {
|
||||
report.GuestDiskTrim = c.diskTrim.GuestDiskTrimStatus(ctx)
|
||||
}
|
||||
// D1: agent self-update pending status (nil reporter → pending=false, the steady state).
|
||||
if c.selfUpdate != nil {
|
||||
report.SelfUpdatePending, report.SelfUpdatePendingVersion = c.selfUpdate.SelfUpdatePending()
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
package hub
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-444: the guest_disk_trim stanza must reach a report built through the PRODUCTION collect path, be absent from
|
||||
// the wire when the job is not wired, and carry the keys the hub's System page reads.
|
||||
|
||||
type fakeDiskTrim struct{ st *GuestDiskTrimStatus }
|
||||
|
||||
func (f fakeDiskTrim) GuestDiskTrimStatus(context.Context) *GuestDiskTrimStatus { return f.st }
|
||||
|
||||
func TestCollect_GuestDiskTrim(t *testing.T) {
|
||||
px := &fakePx{node: "n", ns: newTestNodeStatus()}
|
||||
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.150.0", quietLogger())
|
||||
r, err := c.Collect(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("Collect: %v", err)
|
||||
}
|
||||
b, _ := json.Marshal(r)
|
||||
var m map[string]any
|
||||
_ = json.Unmarshal(b, &m)
|
||||
if _, ok := m["guest_disk_trim"]; ok {
|
||||
t.Fatalf("guest_disk_trim on the wire with no reporter wired: %s", b)
|
||||
}
|
||||
|
||||
c.SetGuestDiskTrimReporter(fakeDiskTrim{st: &GuestDiskTrimStatus{Schedule: "weekly", Guests: []GuestDiskTrim{{
|
||||
VMID: 9201, LastAttemptAt: "2026-10-07T08:30:00Z", OK: true, BytesTrimmed: 90143313920, Mounts: 2,
|
||||
DurationSeconds: 24.4, LastOKAt: "2026-10-07T08:30:00Z",
|
||||
}}}})
|
||||
r, err = c.Collect(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("Collect: %v", err)
|
||||
}
|
||||
b, _ = json.Marshal(r)
|
||||
m = nil
|
||||
_ = json.Unmarshal(b, &m)
|
||||
dt, ok := m["guest_disk_trim"].(map[string]any)
|
||||
if !ok || dt["schedule"] != "weekly" {
|
||||
t.Fatalf("guest_disk_trim missing or wrong on the wire: %s", b)
|
||||
}
|
||||
g := dt["guests"].([]any)[0].(map[string]any)
|
||||
for _, k := range []string{"vmid", "last_attempt_at", "ok", "bytes_trimmed", "mounts", "duration_seconds", "last_ok_at"} {
|
||||
if _, ok := g[k]; !ok {
|
||||
t.Fatalf("guest_disk_trim.guests[0] lacks %q: %v", k, g)
|
||||
}
|
||||
}
|
||||
if g["bytes_trimmed"] != float64(90143313920) || g["ok"] != true {
|
||||
t.Fatalf("values did not survive the round trip: %v", g)
|
||||
}
|
||||
}
|
||||
@@ -124,6 +124,11 @@ type HostReport struct {
|
||||
// on HostReport would have been the only report block named against that convention.
|
||||
GuestNet *GuestNetStatus `json:"guest_net,omitempty"`
|
||||
|
||||
// GuestDiskTrim is the weekly guest disk trim stanza (R-444, `09` §3 decision 139): the schedule and, per owned
|
||||
// guest, the LAST trim result as persisted by the agent (it survives a restart). Present only when the trim job is
|
||||
// wired; an empty `guests` list means the job runs and no guest has been trimmed yet. No secret.
|
||||
GuestDiskTrim *GuestDiskTrimStatus `json:"guest_disk_trim,omitempty"`
|
||||
|
||||
// LogTail is the agent's on-demand debug-ring tail (v0.83.0 observability) — the agent
|
||||
// mirror of the controller's report log_tails channel. Present ONLY on the heartbeat
|
||||
// right after the control envelope requested it (log_tail_requested); consume-once on
|
||||
@@ -208,6 +213,27 @@ type GuestNetGuest struct {
|
||||
Message string `json:"message,omitempty"`
|
||||
}
|
||||
|
||||
// GuestDiskTrimStatus is the R-444 weekly trim stanza. `schedule` is a plain description of when the job runs (local
|
||||
// time of the host); `guests` holds one entry per owned guest that has had at least one trim attempt.
|
||||
type GuestDiskTrimStatus struct {
|
||||
Schedule string `json:"schedule"`
|
||||
Guests []GuestDiskTrim `json:"guests,omitempty"`
|
||||
}
|
||||
|
||||
// GuestDiskTrim is one guest's LAST trim attempt. `ok` with `last_attempt_at` is the verdict of that attempt — never
|
||||
// read the time alone as success; `last_ok_at` is the last attempt that succeeded ("" = never). `bytes_trimmed` is
|
||||
// the sum of the "(N bytes) trimmed" lines `pct fstrim` printed, over `mounts` mount points.
|
||||
type GuestDiskTrim struct {
|
||||
VMID int `json:"vmid"`
|
||||
LastAttemptAt string `json:"last_attempt_at"`
|
||||
OK bool `json:"ok"`
|
||||
BytesTrimmed int64 `json:"bytes_trimmed"`
|
||||
Mounts int `json:"mounts"`
|
||||
DurationSeconds float64 `json:"duration_seconds"`
|
||||
LastOKAt string `json:"last_ok_at,omitempty"`
|
||||
Error string `json:"error,omitempty"`
|
||||
}
|
||||
|
||||
type PBSDRStatus struct {
|
||||
State string `json:"state"`
|
||||
StorageID string `json:"storage_id,omitempty"`
|
||||
|
||||
Reference in New Issue
Block a user