Host report keeps the last backup across a restart; Secure Boot meta-package leaves the Proxmox lane
gates / gates (push) Successful in 1m7s

Found in the 2026-10-09 kernel-night read-back:
- demo-felhom: the kernel step restarted the host 4 minutes after the night
  backup; the in-memory backup list was empty after the restart, so the hub
  alarmed "host tier: newest backup is 48h old". The report now adds R-894's
  saved newest success per tier (one shared instance).
- demo-hp: proxmox-secure-boot-support pulled shim-signed-common into the
  ring-0 Proxmox plan; the step was refused R6 and the kernel step skipped.
  It joins HOST_SLOW_RE with shim and GRUB.

Red-proved: TestCollectBackups_SavedSuccessSurvivesARestart,
TestR894_LastKnownBackupsIsWiredIntoTheDaemon,
test_ring0_pending_pve_leaves_the_secure_boot_meta_out.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-09 06:47:58 +02:00
parent b228b44f0e
commit db25b469ba
9 changed files with 226 additions and 9 deletions
+16 -2
View File
@@ -1,6 +1,20 @@
## Unreleased (2026-10-08) — no OS leg after a household press; the after-boot kernel report carries the real ring (R-899); no „wrong code" when older packages were not checked (R-304) — ships with tomorrow's release
## Unreleased (2026-10-08) — no OS leg after a household press; the after-boot kernel report carries the real ring (R-899); no „wrong code" when older packages were not checked (R-304); the last backup survives a restart in the host report; the Secure Boot meta-package leaves the Proxmox lane
**Delivery: the agent binary only** — no root file changed.
**Delivery: the agent binary AND the config bundle** — the wrapper `felhom-os-apply` changed (2026-10-09).
- **The hub's backup evidence survives a restart (2026-10-09, found in the kernel-night read-back).** The host report's
`backups` list came only from memory, which is empty after an agent restart. On demo-felhom the night backup landed
02:40 UTC and the kernel step restarted the host at 02:44 — no report fell in between, two nights running — so the hub
alarmed „host tier: newest backup is 48h old" at 03:00 after a good backup. The report now adds R-894's saved newest
success per tier and guest (`backup-success-state.json`, ONE instance shared with the local API) where memory holds no
success at or after it. Only successes are saved, so a failure is never hidden. `hub.KnownBackupReporter`,
`BackupSuccessState.KnownBackupSuccesses`; tests `TestCollectBackups_SavedSuccessSurvivesARestart` (red-proved: the
restart case reported nothing), `TestBackupSuccessState_KnownBackupSuccessesSurviveAReopen`, and
`TestR894_LastKnownBackupsIsWiredIntoTheDaemon` now also asserts the collector gets the same instance (red-proved).
- **`proxmox-secure-boot-support` is a boot-chain package (2026-10-09).** On demo-hp (Secure Boot on) its upgrade pulled
`shim-signed-common`; in the ring-0 Proxmox plan that refused the whole step (R6) and skipped the night's kernel step,
after the household had been told „tonight". It joins `HOST_SLOW_RE` with shim and GRUB (they stay out of every lane,
as before). `test_ring0_pending_pve_leaves_the_secure_boot_meta_out` (red-proved).
- **R-899 (operator ruling 2026-10-08, option A):** `POST /backup?trigger=manual` (a „Mentés most" press, sent by the
controller from its next release) runs no OS leg after it: the leg belongs to the night, after the night's own copy.
+5 -1
View File
@@ -1925,6 +1925,10 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
return out, withheld, nil
},
}
// R-894: ONE instance — the local API writes it, the host report reads it (2026-10-09: a restart right
// after the night backup erased the hub's evidence of it).
lastKnownBackups := backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json"))
collector.SetKnownBackupReporter(lastKnownBackups)
srv, err := localapi.NewServer(localapi.Options{
EscrowRecovery: escrowRecoverer,
ListenAddr: cfg.LocalAPI.ListenAddr,
@@ -1937,7 +1941,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
Store: store,
// R-894: the newest success per tier on disk — the due-check's fallback when the storage cannot be
// read right after a restart. Same state dir as restore-test-state.json.
LastKnownBackups: backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json")),
LastKnownBackups: lastKnownBackups,
Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
+31
View File
@@ -12,12 +12,18 @@ import (
//
// COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this
// fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored.
//
// 2026-10-09: the SAME instance also feeds the host report (collector.SetKnownBackupReporter), so a
// restart right after a backup no longer erases the hub's evidence of it. The field may name a local
// variable; the variable must be built by backup.NewBackupSuccessState and be the one handed to the
// collector. RED-PROOF (observed): drop the SetKnownBackupReporter call → "not handed to the collector".
func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
_, f := parseMain(t)
if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one")
}
var field, built bool
var fieldVar, handed string
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
@@ -45,6 +51,28 @@ func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
if callsIn(kv.Value)["backup.NewBackupSuccessState"] {
built = true
}
if id, ok := kv.Value.(*ast.Ident); ok {
fieldVar = id.Name
}
}
}
return true
})
// a local variable: built by NewBackupSuccessState, and handed to the collector
ast.Inspect(fd.Body, func(n ast.Node) bool {
switch x := n.(type) {
case *ast.AssignStmt:
for i, l := range x.Lhs {
if id, ok := l.(*ast.Ident); ok && fieldVar != "" && id.Name == fieldVar && i < len(x.Rhs) &&
callsIn(x.Rhs[i])["backup.NewBackupSuccessState"] {
built = true
}
}
case *ast.CallExpr:
if fn, ok := x.Fun.(*ast.SelectorExpr); ok && fn.Sel.Name == "SetKnownBackupReporter" && len(x.Args) == 1 {
if id, ok := x.Args[0].(*ast.Ident); ok {
handed = id.Name
}
}
}
return true
@@ -56,6 +84,9 @@ func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
if !built {
t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState")
}
if handed == "" || handed != fieldVar {
t.Fatalf("the saved backups (%q) are not handed to the collector (SetKnownBackupReporter got %q)", fieldVar, handed)
}
}
func callsIn(n ast.Node) map[string]bool {
+6 -1
View File
@@ -93,8 +93,13 @@ OOMCHECK_TIMEOUTS = {"image": 10, "clock": 5, "run": 30, "inspect": 10, "events"
OOMCHECK_SETTLE = 2
# Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host
# reboot is needed for them to take effect, and a bad one can stop the box from booting.
# proxmox-secure-boot-support is the Secure Boot meta-package: its only job is to pull shim-signed and the signed GRUB,
# so it belongs with them. Measured 2026-10-09 on demo-hp (Secure Boot on): left in the pve lane, its upgrade pulled
# shim-signed-common, the whole pve step was refused R6, and the night's kernel step was skipped with it
# (`audits/kernel-night-2026-10-08/`). Pinned by test_ring0_pending_pve_leaves_the_secure_boot_meta_out.
HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|"
r"pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr)")
r"pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr|"
r"proxmox-secure-boot-support)")
# restart_needed() leaves out processes whose cgroup line matches (grep basic regex). Host: the LXC guests' own
# processes (`0::/lxc/<vmid>/...`) -- NOT lxc-start itself, whose cgroup is `0::/lxc.monitor/<vmid>` (measured
# 2026-10-04 on demo-felhom: the old pattern "lxc" hid lxc-start with 20 deleted maps, so "reboot needed" stayed false
+17
View File
@@ -1469,6 +1469,23 @@ class PVELane(unittest.TestCase):
"pending-pve: installed, Proxmox-origin, never kernel/boot/firmware, never Debian or other origins")
self.assertEqual(f.installed["libc6"], "2.41-12+deb13u3")
# 2026-10-09, demo-hp (Secure Boot on): proxmox-secure-boot-support's upgrade pulls shim-signed-common; in the pve
# plan that refused the whole step (R6) and skipped the night's kernel step. It is a boot-chain package: left out.
# RED-PROOF: drop proxmox-secure-boot-support from HOST_SLOW_RE -> it is in the plan -> this test fails.
def test_ring0_pending_pve_leaves_the_secure_boot_meta_out(self):
f = pve_fake(ring0=True)
f.plan["select"], f.plan["packages"] = "pending-pve", []
f.installed.update({"proxmox-secure-boot-support": "9.0.1"})
f.pending_sim = [
"Inst pve-manager [9.2.2] (9.2.21 Proxmox Debian Repository:stable [amd64])",
"Inst proxmox-secure-boot-support [9.0.1] (9.0.2 Proxmox Debian Repository:stable [all])",
"Inst shim-signed-common [1.48+pmx1+16.1-1+pmx1] (1.51+pmx1+16.1-2+pmx1 Proxmox Debian Repository:stable [all])",
]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(sorted(u["name"] for u in rep["upgraded"]), ["pve-manager"])
self.assertEqual(f.installed["proxmox-secure-boot-support"], "9.0.1")
def test_select_pending_pve_needs_the_pve_layer(self):
f = Fake()
f.plan["layer"], f.plan["select"], f.plan["packages"] = "host", "pending-pve", []
+23
View File
@@ -103,6 +103,29 @@ func (s *BackupSuccessState) LastKnownSuccess(target string, vmid int) (time.Tim
return e.at, ok
}
// KnownBackupSuccesses returns every saved success as a host-report record (the collector's
// KnownBackupReporter, so the hub keeps its evidence across a restart). Only the tier, guest and start
// time are known; the archive name is not saved and stays empty. Sorted for a stable report.
func (s *BackupSuccessState) KnownBackupSuccesses() []hub.Backup {
if s == nil {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
out := make([]hub.Backup, 0, len(s.last))
for _, e := range s.last {
out = append(out, hub.Backup{TargetID: e.target, VMID: e.vmid, Success: true, CrashConsistent: true,
StartedAt: e.at.Format(time.RFC3339), UncoveredVolumes: []string{}})
}
sort.Slice(out, func(i, j int) bool {
if out[i].TargetID != out[j].TargetID {
return out[i].TargetID < out[j].TargetID
}
return out[i].VMID < out[j].VMID
})
return out
}
func (s *BackupSuccessState) saveLocked() error {
entries := make([]backupSuccessJSON, 0, len(s.last))
for _, e := range s.last {
+17
View File
@@ -57,3 +57,20 @@ func TestBackupSuccessState_CorruptFileIsEmpty(t *testing.T) {
t.Fatal("a corrupt file must read as nothing known")
}
}
// The saved successes come back as host-report records: success, tier, guest, start time (R-894 → the
// host report, 2026-10-09). A new state read from the same file gives the same records (the restart case).
func TestBackupSuccessState_KnownBackupSuccessesSurviveAReopen(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
s := NewBackupSuccessState(path)
if err := s.RecordBackupSuccess("felhom-backup", hub.Backup{VMID: 9201, Success: true, StartedAt: "2026-10-09T02:40:16Z"}); err != nil {
t.Fatal(err)
}
got := NewBackupSuccessState(path).KnownBackupSuccesses()
if len(got) != 1 || !got[0].Success || got[0].TargetID != "felhom-backup" || got[0].VMID != 9201 || got[0].StartedAt != "2026-10-09T02:40:16Z" {
t.Fatalf("got %+v", got)
}
if got[0].UncoveredVolumes == nil {
t.Fatal("uncovered_volumes must marshal as [], not null")
}
}
+49 -5
View File
@@ -66,6 +66,13 @@ type ProvenRestoreTestReporter interface {
ProvenRestoreTests(ctx context.Context) []RestoreTest
}
// KnownBackupReporter is the newest SUCCESSFUL backup per tier and guest kept on disk (R-894's
// backup-success-state.json). collectBackups folds it into the report so a restart does not erase the
// hub's evidence of a backup that ran minutes before it (the kernel-night shape, 2026-10-09).
type KnownBackupReporter interface {
KnownBackupSuccesses() []Backup
}
// PBSReporter is the slice-6-Phase-B seam the pbs verify loop plugs into (same pattern).
// Returns the agent's latest-known PBS snapshot inventory + verify-state. nil → empty.
type PBSReporter interface {
@@ -105,6 +112,7 @@ type Collector struct {
backups BackupReporter
restoreTests RestoreTestReporter
provenTests ProvenRestoreTestReporter
knownBackups KnownBackupReporter // R-894 file → host report after a restart (nil → in-memory only)
pbs PBSReporter
temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp)
capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty)
@@ -576,14 +584,50 @@ func (c *Collector) collectStorage(ctx context.Context) []StorageTarget {
// collectBackups / collectRestoreTests read the agent's latest backup + restore-test state
// via the seams. Best-effort: a nil reporter or nil slice degrades to an empty (non-nil)
// list so the collection always marshals as [].
//
// The saved successes (SetKnownBackupReporter) are added for each tier and guest the in-memory list has
// no success for at or after the saved time. Why: the in-memory list is empty after an agent restart,
// and the hub's backup-freshness check reads only what the reports carried. Measured 2026-10-09 on
// demo-felhom: the night backup landed 02:40 UTC, the kernel step restarted the host at 02:44, no report
// fell in between, so two kernel nights in a row left no trace and the hub alarmed „newest backup is 48h
// old" at 03:00. Only SUCCESSES are saved, so a failure is never hidden and never invented.
// Pinned by TestCollectBackups_SavedSuccessSurvivesARestart.
func (c *Collector) collectBackups(ctx context.Context) []Backup {
if c.backups == nil {
return []Backup{}
out := []Backup{}
if c.backups != nil {
if b := c.backups.Backups(ctx); b != nil {
out = append(out, b...)
}
}
if b := c.backups.Backups(ctx); b != nil {
return b
if c.knownBackups == nil {
return out
}
return []Backup{}
for _, k := range c.knownBackups.KnownBackupSuccesses() {
kt, err := time.Parse(time.RFC3339, k.StartedAt)
if err != nil {
continue
}
covered := false
for _, b := range out {
if !b.Success || b.TargetID != k.TargetID || b.VMID != k.VMID {
continue
}
if bt, err := time.Parse(time.RFC3339, b.StartedAt); err == nil && !bt.Before(kt) {
covered = true
break
}
}
if !covered {
out = append(out, k)
}
}
return out
}
// SetKnownBackupReporter wires the R-894 saved successes into the host report (nil-safe → in-memory only).
func (c *Collector) SetKnownBackupReporter(r KnownBackupReporter) *Collector {
c.knownBackups = r
return c
}
// collectRestoreTests merges the in-memory result with the PERSISTED per-tier proofs (R-189).
+62
View File
@@ -0,0 +1,62 @@
package hub
import (
"context"
"testing"
)
type fakeMemBackups []Backup
func (f fakeMemBackups) Backups(context.Context) []Backup { return f }
type fakeKnownBackups []Backup
func (f fakeKnownBackups) KnownBackupSuccesses() []Backup { return f }
// The consequence, not the mechanism: after a restart (empty in-memory list) the host report still carries
// the night's backup, so the hub's freshness check sees it. Measured 2026-10-09 on demo-felhom: without this
// the report carried backups: [] and the hub alarmed „newest backup is 48h old" an hour after a good backup.
// RED-PROOF: return before the saved-success loop in collectBackups → the first case reports nothing → fails.
func TestCollectBackups_SavedSuccessSurvivesARestart(t *testing.T) {
saved := Backup{TargetID: "felhom-backup", VMID: 9201, Success: true, StartedAt: "2026-10-09T02:40:16Z"}
t.Run("restart: memory empty, the saved success is reported", func(t *testing.T) {
c := &Collector{backups: fakeMemBackups(nil), knownBackups: fakeKnownBackups{saved}}
got := c.collectBackups(context.Background())
if len(got) != 1 || got[0].StartedAt != saved.StartedAt || !got[0].Success || got[0].TargetID != "felhom-backup" {
t.Fatalf("want the saved success, got %+v", got)
}
})
t.Run("memory holds the same or a newer success: no second entry", func(t *testing.T) {
mem := Backup{TargetID: "felhom-backup", VMID: 9201, Success: true, StartedAt: "2026-10-09T02:40:16Z", Archive: "a"}
c := &Collector{backups: fakeMemBackups{mem}, knownBackups: fakeKnownBackups{saved}}
if got := c.collectBackups(context.Background()); len(got) != 1 || got[0].Archive != "a" {
t.Fatalf("want only the in-memory record, got %+v", got)
}
})
t.Run("a newer FAILURE in memory never hides the saved success, and is kept", func(t *testing.T) {
fail := Backup{TargetID: "felhom-backup", VMID: 9201, Success: false, StartedAt: "2026-10-10T02:40:00Z", Error: "x"}
c := &Collector{backups: fakeMemBackups{fail}, knownBackups: fakeKnownBackups{saved}}
got := c.collectBackups(context.Background())
if len(got) != 2 || got[0].Success || !got[1].Success {
t.Fatalf("want the failure and the saved success, got %+v", got)
}
})
t.Run("another tier's success does not cover this tier", func(t *testing.T) {
other := Backup{TargetID: "felhom-pbs", VMID: 9201, Success: true, StartedAt: "2026-10-10T00:00:00Z"}
c := &Collector{backups: fakeMemBackups{other}, knownBackups: fakeKnownBackups{saved}}
if got := c.collectBackups(context.Background()); len(got) != 2 {
t.Fatalf("want both tiers, got %+v", got)
}
})
t.Run("no saved state wired: the in-memory list as before, never nil", func(t *testing.T) {
c := &Collector{}
if got := c.collectBackups(context.Background()); got == nil || len(got) != 0 {
t.Fatalf("want an empty non-nil list, got %#v", got)
}
})
}