R-894: after a restart the agent remembers the last backup per tier
gates / gates (push) Successful in 35s
gates / gates (push) Successful in 35s
An unreadable storage right after an agent restart read the off-site tier DUE (the in-memory record was empty). The newest success per tier is now kept on disk and read ONLY when the storage cannot be read: fresh -> not due, older than the cadence -> due, none -> due (unknown) as before. A storage that answers stays the ground truth. Ships with v0.150.0 after the 2026-10-07 read-back; nothing delivered. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,3 +1,13 @@
|
||||
## Unreleased (2026-10-06 night, later) — after a restart the agent remembers the last backup per tier (R-894)
|
||||
|
||||
Ships with the memory-kill check below as v0.150.0, AFTER the 2026-10-07 night read-back. Nothing delivered tonight.
|
||||
|
||||
- **The defect (measured 2026-10-05 on demo-hp):** the agent restarted at 04:57; at 06:25 the off-site storage answered *Can't connect*; the per-tier backup record is in memory only, so the due-check fell back to an EMPTY record and the 7-day tier (last copy 4 days old) read DUE; the controller asked and vzdump failed.
|
||||
- New `internal/backup/backup_state.go` `BackupSuccessState`: the newest SUCCESSFUL backup per tier and guest, on disk (`<oob state dir>/backup-success-state.json`, atomic tmp+rename, 0600). Only successes are written; a corrupt file reads as nothing known.
|
||||
- `internal/localapi` `handleBackupDue`: when the tier's storage CANNOT be read, the saved copy stands in for the in-memory record. A fresh copy → not due („… (storage unreadable — age from the last success saved on disk)"); a copy older than the cadence → DUE; no copy → the old answer (DUE, age unknown). A storage that answers stays the ground truth: an archive absent there is due even when the file remembers one.
|
||||
- Wired in `buildLocalAPIServer` (`LastKnownBackups`); the local API's backup job saves each success.
|
||||
- Tests: `TestBackupDue_R894_*` (restart = a new server and a new state from the same file; fresh / old / none / storage answers / failed backup not saved), `TestBackupSuccessState_*`, `TestR894_LastKnownBackupsIsWiredIntoTheDaemon` (AST). Four red-proofs observed (`felhom.eu/documentation/audits/night-burndown-2026-10-06/s4/`).
|
||||
|
||||
## Unreleased (2026-10-06 night) — the Docker step proves the engine reports a memory kill (`09` §3 decision 157, R-528)
|
||||
|
||||
To be released as v0.150.0 with its config bundle AFTER the 2026-10-07 night read-back (the night of 2026-10-06 runs v0.149.0 on purpose).
|
||||
|
||||
@@ -163,6 +163,7 @@
|
||||
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
|
||||
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
|
||||
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
|
||||
| `backup.BackupSuccessState` | internal/backup/backup_state.go | `RecordBackupSuccess(target, b)` / `LastKnownSuccess(target, vmid)` | Newest SUCCESSFUL backup per tier+guest, persisted (atomic tmp+rename) — the due-check's fallback when the tier's storage cannot be read after a restart (R-894) | **Read ONLY when the storage cannot be read** — a storage that answers is the ground truth (R-84), and an archive absent there must make the tier due even when this file remembers one. A saved copy older than the cadence still reads due. Only successes are written. |
|
||||
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
|
||||
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
|
||||
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
|
||||
|
||||
@@ -1901,12 +1901,15 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
|
||||
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
|
||||
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
|
||||
Store: store,
|
||||
Storage: observer,
|
||||
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
|
||||
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
|
||||
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
|
||||
Tokens: tokens,
|
||||
BackupCadence: cfg.Backup.BackupCadence(),
|
||||
// R-894: the newest success per tier on disk — the due-check's fallback when the storage cannot be
|
||||
// read right after a restart. Same state dir as restore-test-state.json.
|
||||
LastKnownBackups: backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json")),
|
||||
Storage: observer,
|
||||
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
|
||||
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
|
||||
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
|
||||
Tokens: tokens,
|
||||
BackupCadence: cfg.Backup.BackupCadence(),
|
||||
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
|
||||
Disks: hostOps,
|
||||
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
|
||||
|
||||
@@ -0,0 +1,74 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"go/ast"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-894 — the on-disk backup record is WIRED on the daemon path (the built-but-never-wired class).
|
||||
// main → runDaemon → buildLocalAPIServer, and inside it the localapi.Options literal carries
|
||||
// LastKnownBackups built by backup.NewBackupSuccessState. An AST walk, not a string match, for the
|
||||
// reasons in escrow_recover_wiring_test.go.
|
||||
//
|
||||
// COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this
|
||||
// fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored.
|
||||
func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
|
||||
_, f := parseMain(t)
|
||||
if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
|
||||
t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one")
|
||||
}
|
||||
var field, built bool
|
||||
for _, d := range f.Decls {
|
||||
fd, ok := d.(*ast.FuncDecl)
|
||||
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
|
||||
continue
|
||||
}
|
||||
ast.Inspect(fd.Body, func(n ast.Node) bool {
|
||||
cl, ok := n.(*ast.CompositeLit)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
sel, ok := cl.Type.(*ast.SelectorExpr)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
if pkg, _ := sel.X.(*ast.Ident); pkg == nil || pkg.Name+"."+sel.Sel.Name != "localapi.Options" {
|
||||
return true
|
||||
}
|
||||
for _, el := range cl.Elts {
|
||||
kv, ok := el.(*ast.KeyValueExpr)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "LastKnownBackups" {
|
||||
field = true
|
||||
if callsIn(kv.Value)["backup.NewBackupSuccessState"] {
|
||||
built = true
|
||||
}
|
||||
}
|
||||
}
|
||||
return true
|
||||
})
|
||||
}
|
||||
if !field {
|
||||
t.Fatal("localapi.Options in buildLocalAPIServer has no LastKnownBackups field")
|
||||
}
|
||||
if !built {
|
||||
t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState")
|
||||
}
|
||||
}
|
||||
|
||||
func callsIn(n ast.Node) map[string]bool {
|
||||
out := map[string]bool{}
|
||||
ast.Inspect(n, func(n ast.Node) bool {
|
||||
if ce, ok := n.(*ast.CallExpr); ok {
|
||||
if fn, ok := ce.Fun.(*ast.SelectorExpr); ok {
|
||||
if x, ok := fn.X.(*ast.Ident); ok {
|
||||
out[x.Name+"."+fn.Sel.Name] = true
|
||||
}
|
||||
}
|
||||
}
|
||||
return true
|
||||
})
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,131 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strconv"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
)
|
||||
|
||||
// BackupSuccessState persists the newest SUCCESSFUL whole-guest backup per tier and guest (R-894).
|
||||
//
|
||||
// Why it exists. The due-check (`localapi` handleBackupDue) asks the tier's storage when a backup last
|
||||
// landed (R-84) and falls back to the in-memory record when the storage cannot be read. The in-memory
|
||||
// record is empty after an agent restart (Store, R-348), so "storage unreadable" right after a restart
|
||||
// read as "no record — DUE". Measured 2026-10-05 on demo-hp: the agent restarted at 04:57, the off-site
|
||||
// storage answered "Can't connect" at 06:25, the 7-day tier — last copy 2026-10-01 — read DUE, the
|
||||
// controller asked, and vzdump failed. This file is the last known copy the fallback reads instead.
|
||||
//
|
||||
// It is read ONLY when the storage cannot be read. A storage that answers is the ground truth and wins,
|
||||
// in both directions: an archive found there counts, and an archive absent there is absent even when
|
||||
// this file remembers a success (a pruned or deleted archive must make the tier due — the same reason
|
||||
// R-84 chose the storage over a persisted record). Pinned by
|
||||
// TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers.
|
||||
//
|
||||
// Only SUCCESSES are written (the RestoreTestState rule): a failure must stay due and be retried, so a
|
||||
// record of a failure has no reader.
|
||||
type BackupSuccessState struct {
|
||||
path string
|
||||
mu sync.Mutex
|
||||
last map[string]savedSuccess // key(target, vmid) → the newest success
|
||||
}
|
||||
|
||||
type savedSuccess struct {
|
||||
target string
|
||||
vmid int
|
||||
at time.Time
|
||||
}
|
||||
|
||||
// backupSuccessJSON is one entry on disk.
|
||||
type backupSuccessJSON struct {
|
||||
Target string `json:"target"`
|
||||
VMID int `json:"vmid"`
|
||||
StartedAt string `json:"started_at"`
|
||||
}
|
||||
|
||||
func backupStateKey(target string, vmid int) string { return target + "/" + strconv.Itoa(vmid) }
|
||||
|
||||
// NewBackupSuccessState opens (or creates) the state at path. A missing or unreadable file degrades to
|
||||
// "nothing known" — the pre-R-894 behaviour, which is DUE — and never wedges the daemon.
|
||||
func NewBackupSuccessState(path string) *BackupSuccessState {
|
||||
s := &BackupSuccessState{path: path, last: map[string]savedSuccess{}}
|
||||
data, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return s
|
||||
}
|
||||
var entries []backupSuccessJSON
|
||||
if json.Unmarshal(data, &entries) != nil {
|
||||
return s
|
||||
}
|
||||
for _, e := range entries {
|
||||
t, perr := time.Parse(time.RFC3339, e.StartedAt)
|
||||
if perr != nil {
|
||||
continue // one unreadable entry must not lose the others
|
||||
}
|
||||
s.last[backupStateKey(e.Target, e.VMID)] = savedSuccess{target: e.Target, vmid: e.VMID, at: t.UTC()}
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// RecordBackupSuccess saves b when it is a success newer than the one on file. target is the tier the
|
||||
// job ran on (the due-check's key); a failure or an unparseable time is ignored.
|
||||
func (s *BackupSuccessState) RecordBackupSuccess(target string, b hub.Backup) error {
|
||||
if s == nil || !b.Success {
|
||||
return nil
|
||||
}
|
||||
t, err := time.Parse(time.RFC3339, b.StartedAt)
|
||||
if err != nil {
|
||||
return nil
|
||||
}
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
k := backupStateKey(target, b.VMID)
|
||||
if old, ok := s.last[k]; ok && !t.After(old.at) {
|
||||
return nil
|
||||
}
|
||||
s.last[k] = savedSuccess{target: target, vmid: b.VMID, at: t.UTC()}
|
||||
return s.saveLocked()
|
||||
}
|
||||
|
||||
// LastKnownSuccess returns the newest saved success for this tier and guest (ok=false = none on file).
|
||||
func (s *BackupSuccessState) LastKnownSuccess(target string, vmid int) (time.Time, bool) {
|
||||
if s == nil {
|
||||
return time.Time{}, false
|
||||
}
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
e, ok := s.last[backupStateKey(target, vmid)]
|
||||
return e.at, ok
|
||||
}
|
||||
|
||||
func (s *BackupSuccessState) saveLocked() error {
|
||||
entries := make([]backupSuccessJSON, 0, len(s.last))
|
||||
for _, e := range s.last {
|
||||
entries = append(entries, backupSuccessJSON{Target: e.target, VMID: e.vmid, StartedAt: e.at.Format(time.RFC3339)})
|
||||
}
|
||||
// Deterministic file content (Go's map order is random).
|
||||
sort.Slice(entries, func(i, j int) bool {
|
||||
if entries[i].Target != entries[j].Target {
|
||||
return entries[i].Target < entries[j].Target
|
||||
}
|
||||
return entries[i].VMID < entries[j].VMID
|
||||
})
|
||||
data, err := json.MarshalIndent(entries, "", " ")
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := os.MkdirAll(filepath.Dir(s.path), 0o755); err != nil {
|
||||
return err
|
||||
}
|
||||
tmp := s.path + ".tmp"
|
||||
if err := os.WriteFile(tmp, data, 0o600); err != nil {
|
||||
os.Remove(tmp)
|
||||
return err
|
||||
}
|
||||
return os.Rename(tmp, s.path)
|
||||
}
|
||||
@@ -0,0 +1,59 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
)
|
||||
|
||||
// R-894: the on-disk newest success per tier survives a restart (a new state from the same file).
|
||||
func TestBackupSuccessState_SurvivesRestart(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||
s := NewBackupSuccessState(path)
|
||||
at := time.Date(2026, 10, 1, 20, 15, 0, 0, time.UTC)
|
||||
if err := s.RecordBackupSuccess("felhom-pbs", hub.Backup{VMID: 9201, Success: true, StartedAt: at.Format(time.RFC3339)}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
got, ok := NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 9201)
|
||||
if !ok || !got.Equal(at) {
|
||||
t.Fatalf("after a restart the saved copy must read back; got %v ok=%v", got, ok)
|
||||
}
|
||||
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("local", 9201); ok {
|
||||
t.Fatal("another tier must not borrow this tier's copy")
|
||||
}
|
||||
}
|
||||
|
||||
// Only a NEWER success replaces the saved one; failures and unparseable times are ignored.
|
||||
func TestBackupSuccessState_KeepsNewestSuccessOnly(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "s.json")
|
||||
s := NewBackupSuccessState(path)
|
||||
newer := time.Date(2026, 10, 5, 0, 0, 0, 0, time.UTC)
|
||||
older := newer.Add(-48 * time.Hour)
|
||||
for _, b := range []hub.Backup{
|
||||
{VMID: 1, Success: true, StartedAt: newer.Format(time.RFC3339)},
|
||||
{VMID: 1, Success: true, StartedAt: older.Format(time.RFC3339)}, // older: ignored
|
||||
{VMID: 1, Success: false, StartedAt: newer.Add(time.Hour).Format(time.RFC3339)}, // failure: ignored
|
||||
{VMID: 1, Success: true, StartedAt: "not-a-time"}, // unparseable: ignored
|
||||
} {
|
||||
if err := s.RecordBackupSuccess("t", b); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if got, _ := NewBackupSuccessState(path).LastKnownSuccess("t", 1); !got.Equal(newer) {
|
||||
t.Fatalf("want the newest success %v, got %v", newer, got)
|
||||
}
|
||||
}
|
||||
|
||||
// A corrupt file degrades to "nothing known" (the pre-R-894 DUE answer), never a crash.
|
||||
func TestBackupSuccessState_CorruptFileIsEmpty(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "s.json")
|
||||
if err := os.WriteFile(path, []byte("{not json"), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("t", 1); ok {
|
||||
t.Fatal("a corrupt file must read as nothing known")
|
||||
}
|
||||
}
|
||||
@@ -30,6 +30,9 @@ import (
|
||||
// 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What
|
||||
// is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/
|
||||
// deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84).
|
||||
// The due-check's fallback for an UNREADABLE storage no longer reads this store alone (R-894): the newest
|
||||
// success per tier is also on disk (BackupSuccessState), so a restart followed by an unreachable storage
|
||||
// reads the last known copy, not "never".
|
||||
type Store struct {
|
||||
mu sync.Mutex
|
||||
byTarget map[string]hub.Backup // latest backup per target id
|
||||
|
||||
@@ -0,0 +1,144 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"io"
|
||||
"log/slog"
|
||||
"net/http"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
)
|
||||
|
||||
// R-894 — after an agent restart, an UNREADABLE storage must fall back to the last success saved on
|
||||
// disk, not to "never". Measured 2026-10-05 on demo-hp: a restart at 04:57, the off-site storage
|
||||
// unreachable at 06:25, the 7-day tier (last copy 4 days old) read DUE, vzdump failed.
|
||||
//
|
||||
// Every test here builds a NEW server and a NEW BackupSuccessState from the same file — that is the
|
||||
// restart. The in-memory store (fakeStore) is always fresh, as after a real restart.
|
||||
|
||||
// r894Server builds a two-tier server whose off-site tier answers the storage listing with lister.
|
||||
func r894Server(t *testing.T, path string, pbsSvc BackupService) *Server {
|
||||
t.Helper()
|
||||
srv, err := NewServer(Options{
|
||||
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: &fakeStore{},
|
||||
Storage: fakeStorage{targets: []hub.StorageTarget{{Name: "local"}, {Name: "felhom-pbs"}}},
|
||||
Tokens: staticTokens{"A": 8200},
|
||||
BackupTiers: []BackupTier{
|
||||
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
|
||||
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: pbsSvc},
|
||||
},
|
||||
LastKnownBackups: backup.NewBackupSuccessState(path),
|
||||
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
srv.baseCtx = context.Background()
|
||||
srv.now = func() time.Time { return testNow }
|
||||
return srv
|
||||
}
|
||||
|
||||
// unreadable is the off-site storage as demo-hp saw it: "Can't connect to 10.77.0.1:8007".
|
||||
func unreadable() archiveLister {
|
||||
return archiveLister{fakeBackups: &fakeBackups{}, err: errStorageRead}
|
||||
}
|
||||
|
||||
// THE R-894 CASE, end to end. Agent 1 takes an off-site backup through POST /backup (the fake
|
||||
// runner's success is 12 h before testNow). The agent restarts. The storage cannot be read. The tier
|
||||
// must read NOT due, from the copy saved on disk.
|
||||
//
|
||||
// COMPANION RED-PROOF (observed): delete the `lookup == archiveUnknown && s.lastKnown != nil` block in
|
||||
// handleBackupDue → this fails with "after a restart an unreadable storage must fall back to the saved
|
||||
// copy (12 h old, 7-day tier) — NOT due; got {… Due:true … AgeState:unknown …}". Restored.
|
||||
func TestBackupDue_R894_RestartThenUnreadableStorage_FreshSavedCopyIsNotDue(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||
|
||||
// Agent 1: a real backup job through the endpoint the controller calls.
|
||||
first := r894Server(t, path, &fakeBackups{})
|
||||
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
|
||||
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
|
||||
}
|
||||
waitFor(t, func() bool {
|
||||
_, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200)
|
||||
return ok
|
||||
})
|
||||
|
||||
// Agent 2: a restart (new server, new state from the same file), and the storage is unreachable.
|
||||
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
|
||||
if got.Due {
|
||||
t.Fatalf("after a restart an unreadable storage must fall back to the saved copy (12 h old, 7-day tier) — NOT due; got %+v", got)
|
||||
}
|
||||
if got.AgeState != AgeStateKnown || got.AgeSecs == nil || *got.AgeSecs != int64((12*time.Hour).Seconds()) {
|
||||
t.Fatalf("the age must come from the saved copy (12 h, known); got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// The deliberate rule stays: an unreadable storage must not suppress a backup that IS due. A saved
|
||||
// copy older than the cadence reads DUE.
|
||||
//
|
||||
// COMPANION RED-PROOF (observed): make the fallback answer not-due whenever a saved copy exists
|
||||
// (`if fromDisk { …Due:false… }` before the cadence check) → this fails with "a saved copy 9 days old
|
||||
// under a 7-day cadence MUST read due". Restored.
|
||||
func TestBackupDue_R894_RestartThenUnreadableStorage_OldSavedCopyIsDue(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||
st := backup.NewBackupSuccessState(path)
|
||||
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, 9*24*time.Hour, true)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
|
||||
if !got.Due {
|
||||
t.Fatalf("a saved copy 9 days old under a 7-day cadence MUST read due; got %+v", got)
|
||||
}
|
||||
if got.AgeState != AgeStateKnown {
|
||||
t.Fatalf("the age is known (from disk); got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// No saved copy → the pre-R-894 answer, byte for byte: DUE, age UNKNOWN (never ABSENT — the controller
|
||||
// fires its window-gate valve only on absent, R-88).
|
||||
func TestBackupDue_R894_RestartThenUnreadableStorage_NoSavedCopyIsDueUnknown(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
|
||||
if !got.Due || got.AgeState != AgeStateUnknown || got.AgeSecs != nil {
|
||||
t.Fatalf("no saved copy + unreadable storage must stay DUE with age unknown; got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// A storage that ANSWERS is the ground truth: an archive absent there makes the tier due even when the
|
||||
// file remembers a fresh success (a pruned or deleted copy must be made again).
|
||||
//
|
||||
// COMPANION RED-PROOF (observed): drop `lookup == archiveUnknown &&` from the fallback condition → this
|
||||
// fails with "the storage answered 'no archive' — the saved copy must NOT stand in for it". Restored.
|
||||
func TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||
st := backup.NewBackupSuccessState(path)
|
||||
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, time.Hour, true)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
absent := archiveLister{fakeBackups: &fakeBackups{}, found: false}
|
||||
got := dueFor(t, r894Server(t, path, absent).Handler(), "felhom-pbs")
|
||||
if !got.Due {
|
||||
t.Fatalf("the storage answered 'no archive' — the saved copy must NOT stand in for it; got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// A FAILED backup is never saved: it must not make a tier look fresh after a restart.
|
||||
func TestBackupDue_R894_FailedBackupIsNotSaved(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||
first := r894Server(t, path, &fakeBackups{failErr: "could not activate storage 'felhom-pbs'"})
|
||||
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
|
||||
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
|
||||
}
|
||||
waitFor(t, func() bool { return len(first.store.Backups(context.Background())) == 1 })
|
||||
if _, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200); ok {
|
||||
t.Fatal("a failed backup must never be saved as a success")
|
||||
}
|
||||
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
|
||||
if !got.Due {
|
||||
t.Fatalf("after a failed backup and a restart the tier must still be due; got %+v", got)
|
||||
}
|
||||
}
|
||||
+61
-22
@@ -85,6 +85,13 @@ type BackupStore interface {
|
||||
RestoreTests(ctx context.Context) []hub.RestoreTest
|
||||
}
|
||||
|
||||
// LastKnownBackupStore (R-894) is the on-disk newest-success-per-tier record. Satisfied by
|
||||
// *backup.BackupSuccessState.
|
||||
type LastKnownBackupStore interface {
|
||||
RecordBackupSuccess(target string, b hub.Backup) error
|
||||
LastKnownSuccess(target string, vmid int) (time.Time, bool)
|
||||
}
|
||||
|
||||
// StorageView yields the host's observed storage targets (for mapping a mount's storage id →
|
||||
// fast/slow class). Satisfied by *storage.Observer.
|
||||
type StorageView interface {
|
||||
@@ -164,6 +171,10 @@ type Options struct {
|
||||
// PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg
|
||||
// that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs.
|
||||
AfterPrimaryBackup func(ctx context.Context, vmid int)
|
||||
// LastKnownBackups (R-894) keeps the newest successful backup per tier ON DISK, so the due-check's
|
||||
// fallback for an UNREADABLE storage after an agent restart is the last known copy, not "never".
|
||||
// nil = the pre-R-894 behaviour (in-memory record only).
|
||||
LastKnownBackups LastKnownBackupStore
|
||||
// Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when
|
||||
// nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner.
|
||||
Privileged PrivilegedRunner
|
||||
@@ -229,7 +240,6 @@ type Options struct {
|
||||
// POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured"
|
||||
// (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer.
|
||||
EscrowRecovery EscrowRecoverer
|
||||
|
||||
}
|
||||
|
||||
// defaultBackupCadence is the fallback /backup/due window when none is configured.
|
||||
@@ -279,21 +289,22 @@ type Server struct {
|
||||
// is the pre-R-82 shape.
|
||||
tiers []BackupTier
|
||||
// inFlight (R-85) is shared with the restore-test scheduler so the two never run together.
|
||||
inFlight *backup.InFlight
|
||||
inFlight *backup.InFlight
|
||||
afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none
|
||||
logger *slog.Logger
|
||||
now func() time.Time
|
||||
lastKnown LastKnownBackupStore // R-894: on-disk newest success per tier; nil = none
|
||||
logger *slog.Logger
|
||||
now func() time.Time
|
||||
|
||||
disks DiskOps // slice 8C (optional)
|
||||
diskGate StorageGate // slice 8C (optional)
|
||||
guestList GuestLister // slice 8C (optional)
|
||||
guestAttach GuestAttacher // slice 10 P2 (optional)
|
||||
mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional)
|
||||
memMu sync.Mutex // single-flight around a resize apply (one customer per host)
|
||||
netStorage NetworkStorageOps // Part A1: NAS network mounts (optional)
|
||||
netMountRoot string // the user-data namespace root for the network-mount role gate
|
||||
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
|
||||
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
|
||||
disks DiskOps // slice 8C (optional)
|
||||
diskGate StorageGate // slice 8C (optional)
|
||||
guestList GuestLister // slice 8C (optional)
|
||||
guestAttach GuestAttacher // slice 10 P2 (optional)
|
||||
mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional)
|
||||
memMu sync.Mutex // single-flight around a resize apply (one customer per host)
|
||||
netStorage NetworkStorageOps // Part A1: NAS network mounts (optional)
|
||||
netMountRoot string // the user-data namespace root for the network-mount role gate
|
||||
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
|
||||
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
|
||||
// crashGuardStatePath (R-856) is the host crash guard's state file read by GET /host/crash-guard;
|
||||
// empty = defaultCrashGuardStatePath. A seam: tests point it at a fixture.
|
||||
crashGuardStatePath string
|
||||
@@ -301,11 +312,11 @@ type Server struct {
|
||||
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
|
||||
// offsite repository password. OPTIONAL — nil (no hub client configured) makes
|
||||
// POST /escrow/recover-offsite-password answer 503 rather than pretending.
|
||||
escrowRecovery EscrowRecoverer
|
||||
intent IntentRecorder // slice 10 P3 (optional)
|
||||
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
|
||||
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
|
||||
staleLock StaleLockController // F2-b startup stale-lock recovery (optional)
|
||||
escrowRecovery EscrowRecoverer
|
||||
intent IntentRecorder // slice 10 P3 (optional)
|
||||
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
|
||||
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
|
||||
staleLock StaleLockController // F2-b startup stale-lock recovery (optional)
|
||||
// guestPower (F-REBOOT) is per-guest start-attempt state for the guest-power watchdog.
|
||||
// Guarded by guestPowerMu in guestpower.go; in-memory on purpose (see guestPowerState).
|
||||
guestPower map[int]guestPowerState
|
||||
@@ -478,6 +489,7 @@ func NewServer(o Options) (*Server, error) {
|
||||
s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence)
|
||||
s.inFlight = o.InFlight
|
||||
s.afterPrimaryBackup = o.AfterPrimaryBackup
|
||||
s.lastKnown = o.LastKnownBackups
|
||||
if s.backups == nil && len(s.tiers) > 0 {
|
||||
s.backups = s.tiers[0].Service
|
||||
}
|
||||
@@ -905,6 +917,13 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
|
||||
s.logger.Info("local-api: backup job complete", "vmid", vmid, "target", tier.TargetID, "job", jobID, "archive", b.Archive)
|
||||
}
|
||||
s.store.RecordBackup(b)
|
||||
if b.Success && s.lastKnown != nil {
|
||||
if err := s.lastKnown.RecordBackupSuccess(tier.TargetID, b); err != nil {
|
||||
// Not fatal: the backup exists. Only the fallback after a restart loses this copy.
|
||||
s.logger.Warn("local-api: could not save the backup on disk for the due-check fallback (R-894)",
|
||||
"vmid", vmid, "target", tier.TargetID, "err", err)
|
||||
}
|
||||
}
|
||||
s.finishJob(key, jobID, b)
|
||||
// OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate.
|
||||
if b.Success && tier.Primary && s.afterPrimaryBackup != nil {
|
||||
@@ -1067,6 +1086,20 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
|
||||
newest, haveNewest = t, true
|
||||
unparseable = false // ground truth supersedes an unreadable in-memory timestamp
|
||||
}
|
||||
// R-894: the storage could not be read → the last success saved ON DISK stands in for the in-memory
|
||||
// record a restart emptied. ONLY on archiveUnknown: a storage that answers is the ground truth, and an
|
||||
// archive absent there must make the tier due even when the file remembers one (a pruned copy).
|
||||
// A saved copy older than the cadence still reads DUE below — an unreadable storage never suppresses
|
||||
// a backup that is due.
|
||||
fromDisk := false
|
||||
if lookup == archiveUnknown && s.lastKnown != nil {
|
||||
if saved, ok := s.lastKnown.LastKnownSuccess(tier.TargetID, vmid); ok && (!haveNewest || saved.After(newest)) {
|
||||
newest, haveNewest, fromDisk = saved, true, true
|
||||
unparseable = false
|
||||
s.logger.Info("local-api: backup storage unreadable — due-check uses the last success saved on disk (R-894)",
|
||||
"vmid", vmid, "target", tier.TargetID, "saved", saved.UTC().Format(time.RFC3339))
|
||||
}
|
||||
}
|
||||
if !haveNewest {
|
||||
// R-88 Part 2: THREE distinct reasons for a nil age, each with its own state. Only ABSENT is a
|
||||
// positive claim of "never backed up"; only that one may license the controller to bypass its
|
||||
@@ -1089,13 +1122,17 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
|
||||
}
|
||||
age := s.now().Sub(newest)
|
||||
ageSecs := int64(age.Seconds())
|
||||
suffix := ""
|
||||
if fromDisk {
|
||||
suffix = " (storage unreadable — age from the last success saved on disk)"
|
||||
}
|
||||
if age >= tier.Cadence {
|
||||
writeOK(w, BackupDueResponse{VMID: vmid, Due: true, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
|
||||
Reason: "older than cadence", Target: echo})
|
||||
Reason: "older than cadence" + suffix, Target: echo})
|
||||
return
|
||||
}
|
||||
writeOK(w, BackupDueResponse{VMID: vmid, Due: false, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
|
||||
Reason: "within cadence window", Target: echo})
|
||||
Reason: "within cadence window" + suffix, Target: echo})
|
||||
}
|
||||
|
||||
// BackupTiersResponse is GET /backup/tiers (R-82): the tiers this agent serves, primary first.
|
||||
@@ -1483,4 +1520,6 @@ func writeStatus(w http.ResponseWriter, code int, ok bool, data any, errMsg stri
|
||||
}
|
||||
|
||||
// SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0).
|
||||
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) { s.afterPrimaryBackup = fn }
|
||||
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) {
|
||||
s.afterPrimaryBackup = fn
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user