Compare commits

..

7 Commits

Author SHA1 Message Date
admin 24ea960229 REPORT: v0.154.0 released and delivered (2026-10-09)
gates / gates (push) Successful in 1m7s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-09 08:36:51 +02:00
admin cff6fb5a45 v0.154.0: CHANGELOG release entry (sha, bundle, tag)
gates / gates (push) Successful in 1m11s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-09 06:58:00 +02:00
admin db25b469ba Host report keeps the last backup across a restart; Secure Boot meta-package leaves the Proxmox lane
gates / gates (push) Successful in 1m7s
Found in the 2026-10-09 kernel-night read-back:
- demo-felhom: the kernel step restarted the host 4 minutes after the night
  backup; the in-memory backup list was empty after the restart, so the hub
  alarmed "host tier: newest backup is 48h old". The report now adds R-894's
  saved newest success per tier (one shared instance).
- demo-hp: proxmox-secure-boot-support pulled shim-signed-common into the
  ring-0 Proxmox plan; the step was refused R6 and the kernel step skipped.
  It joins HOST_SLOW_RE with shim and GRUB.

Red-proved: TestCollectBackups_SavedSuccessSurvivesARestart,
TestR894_LastKnownBackupsIsWiredIntoTheDaemon,
test_ring0_pending_pve_leaves_the_secure_boot_meta_out.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-09 06:47:58 +02:00
admin b228b44f0e REPORT/CONTEXT: 2026-10-08 day (R-899, R-304 on main, unreleased)
gates / gates (push) Successful in 52s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-08 08:08:53 +02:00
admin 91b9405c81 R-304: no 'wrong code' when earlier sealed packages were not all checked (424 older_unchecked)
gates / gates (push) Successful in 55s
Unreleased; ships with tomorrow's release.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-08 07:53:01 +02:00
admin 4c69c25b48 R-899: no OS leg after a household press (trigger=manual); after-boot kernel reports carry the saved ring
gates / gates (push) Successful in 50s
Unreleased; ships with tomorrow's release.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-08 07:34:19 +02:00
admin c9013bb47d CHANGELOG/REPORT: v0.153.0 released and delivered
gates / gates (push) Successful in 1m33s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-07 19:06:04 +02:00
19 changed files with 508 additions and 47 deletions
+47 -1
View File
@@ -1,4 +1,50 @@
## Unreleased — part of v0.153.0: ring 0 stages exactly the told kernel (R-898; `09` §3 decision 176) (2026-10-07) ## v0.154.0 — no OS leg after a household press; the after-boot kernel report carries the real ring (R-899); no „wrong code" when older packages were not checked (R-304); the last backup survives a restart in the host report; the Secure Boot meta-package leaves the Proxmox lane (2026-10-09)
Released by `scripts/release-agent.sh`: binary sha256 `36a54ad20e896cd61d7b9944e2a521f4209c2b76940d8f1827af05b645463226`, bundle
`93487989f50d4a3039a05cec0f91ec478d62489a414705e49403e9b1340d20e0` (tag `v0.154.0` = `db25b46`).
**Delivery: the agent binary AND the config bundle** — the wrapper `felhom-os-apply` changed (2026-10-09).
- **The hub's backup evidence survives a restart (2026-10-09, found in the kernel-night read-back).** The host report's
`backups` list came only from memory, which is empty after an agent restart. On demo-felhom the night backup landed
02:40 UTC and the kernel step restarted the host at 02:44 — no report fell in between, two nights running — so the hub
alarmed „host tier: newest backup is 48h old" at 03:00 after a good backup. The report now adds R-894's saved newest
success per tier and guest (`backup-success-state.json`, ONE instance shared with the local API) where memory holds no
success at or after it. Only successes are saved, so a failure is never hidden. `hub.KnownBackupReporter`,
`BackupSuccessState.KnownBackupSuccesses`; tests `TestCollectBackups_SavedSuccessSurvivesARestart` (red-proved: the
restart case reported nothing), `TestBackupSuccessState_KnownBackupSuccessesSurviveAReopen`, and
`TestR894_LastKnownBackupsIsWiredIntoTheDaemon` now also asserts the collector gets the same instance (red-proved).
- **`proxmox-secure-boot-support` is a boot-chain package (2026-10-09).** On demo-hp (Secure Boot on) its upgrade pulled
`shim-signed-common`; in the ring-0 Proxmox plan that refused the whole step (R6) and skipped the night's kernel step,
after the household had been told „tonight". It joins `HOST_SLOW_RE` with shim and GRUB (they stay out of every lane,
as before). `test_ring0_pending_pve_leaves_the_secure_boot_meta_out` (red-proved).
- **R-899 (operator ruling 2026-10-08, option A):** `POST /backup?trigger=manual` (a „Mentés most" press, sent by the
controller from its next release) runs no OS leg after it: the leg belongs to the night, after the night's own copy.
Before, a press ran the leg at once (in the day; on 2026-10-07 08:51 demo-hp skipped it only by the 20-hour rule). A
request without the parameter (an older controller, the scheduled path) behaves as before.
`internal/localapi/server.go`; `TestAfterPrimaryBackup` gained two sub-cases (red-proved: ignoring `trigger`, the leg
ran once after a press).
- **R-304 — „wrong code" only when every earlier package was tried.** `POST /escrow/recover-offsite-password` answered
400 („the recovery code did not open the sealed bundle") whenever the current package and the TRIED earlier packages
refused the code — even when the hub withheld earlier packages (rows with no key material, rows over its serve cap), a
served package was malformed, the 6-attempt cap stopped the loop, or the retained list could not be read at all. Now
those cases answer **424** with `older_unchecked` (the count, -1 = unknown) and a sentence that does not call the code
wrong. `escrow.ErrRetainedUnchecked` / `RetainedUncheckedError`; the retained fetcher's second value is now every
withheld package (unopenable + truncated + malformed). Tests `TestR304_*` (real age crypto; red-proved: without the
check, four cases returned the wrong-code error) and a 424 row in `TestRecoverOffsitePassword_EachSituationGetsItsOwnStatus`.
An older controller maps the unknown 424 to its neutral „we do not know why" sentence.
- **The „ring 1" label after a boot:** in the first second after a reboot the agent has not fetched the hub's block, and
the after-boot kernel reports (`judging`, `good`, `revert`, `fell_back`…) said ring 1 on a ring-0 box (demo-felhom,
2026-10-08 night). Now they read the fetched block, else the block the daemon saved on disk before the reboot (R-866),
else ring 1 as before. A label only — the hub's approval reads its own ring list. `kernelReportRing`,
`TestKernelReportRing_BeforeFirstFetch` (red-proved: ring 1, want 0).
## v0.153.0 — ring 0 stages exactly the told kernel (R-898; `09` §3 decision 176) (2026-10-07)
Released by `scripts/release-agent.sh`: binary sha256 `b204ebe6d65944f43b70dcfa4ae0dfb90da38c92e2ba6a38a1826e488eeb6630`, bundle
`6db216275eb0dd39188d93481a2045998a69e8cb7838ad908a58d65d9a6a9b57` (tag `v0.153.0` = `2d1e5d0`). Delivered (binary only — the
bundle files are unchanged from 0.152.0) to demo-hp, demo-felhom, Tester 1 on 2026-10-07 19:03.
**Delivery: the agent binary only** — no root file changed (the wrapper is unchanged; its tests gained two cases). **Delivery: the agent binary only** — no root file changed (the wrapper is unchanged; its tests gained two cases).
+5
View File
@@ -1,5 +1,10 @@
# CONTEXT — felhom-agent working state # CONTEXT — felhom-agent working state
> **2026-10-08 (day) — UNRELEASED on main, ships with tomorrow's release (decision 178).** R-899: `POST /backup?trigger=manual`
> runs no OS leg (`localapi/server.go`); after-boot kernel reports carry the saved block's ring (`kernelReportRing`). R-304:
> `escrow.ErrRetainedUnchecked` → HTTP 424 `older_unchecked` when not every earlier package was tried (withheld, caps,
> malformed, unreadable list). Binary-only delivery; no wrapper or root file changed.
> **2026-10-07 (evening) — v0.152.0, the kernel lane (R-836, decision 172, `11` §5.11).** Wrapper layer `kernel` + two GRUB generators in the bundle (delivered with step bundle `0.152.0-step1` — the bundle adds paths, R-880); `osupdate/kernel.go`: night step on told nights only, after-boot judge (host rule + hub reached, 20 min), ONE self-revert, `os_kernel_step` stages only. Proven on Tester 1 (panic → fell_back, held guest → self_reverted, healthy → 7.0.14-22 default). Open: R-898, R-897; the ring-0 night run. > **2026-10-07 (evening) — v0.152.0, the kernel lane (R-836, decision 172, `11` §5.11).** Wrapper layer `kernel` + two GRUB generators in the bundle (delivered with step bundle `0.152.0-step1` — the bundle adds paths, R-880); `osupdate/kernel.go`: night step on told nights only, after-boot judge (host rule + hub reached, 20 min), ONE self-revert, `os_kernel_step` stages only. Proven on Tester 1 (panic → fell_back, held guest → self_reverted, healthy → 7.0.14-22 default). Open: R-898, R-897; the ring-0 night run.
> **2026-10-04 night — v0.143.0 RELEASED + vouched (R-840, decision 96): the config bundle.** `felhom-os-apply` mode > **2026-10-04 night — v0.143.0 RELEASED + vouched (R-840, decision 96): the config bundle.** `felhom-os-apply` mode
+13 -16
View File
@@ -1,18 +1,15 @@
# REPORT — agent v0.152.0: the kernel lane (2026-10-07) # REPORT — v0.154.0 released and delivered (2026-10-09)
**What:** R-836, `09` §3 decision 172 — a new kernel boots ONCE through a flag on the ESP; a crash falls back to the old Built from the 2026-10-09 kernel-night read-back: (1) the host report now adds the saved newest backup success per tier
kernel by itself; a healthy boot (host health rule + the hub reached, 20 min) makes it the default; a booted-but-unhealthy (R-894's file, one shared instance), so a restart right after a night backup no longer empties the hub's evidence — on
one is reverted ONCE. Design: `felhom.eu/documentation/architecture/11-os-updates.md` §5.11. demo-felhom it had caused a false „newest backup is 48h old" alarm; (2) `proxmox-secure-boot-support` joins the
boot-chain lane, so it no longer pulls `shim-signed-common` into the ring-0 Proxmox plan (demo-hp's step was refused R6
and its kernel step skipped). Red-proofs: `TestCollectBackups_SavedSuccessSurvivesARestart`,
`TestR894_LastKnownBackupsIsWiredIntoTheDaemon`, `test_ring0_pending_pve_leaves_the_secure_boot_meta_out`.
- Wrapper `configs/felhom-os-apply`: layer `kernel` (stage, kernel-reboot, kernel-boot, kernel-good, kernel-revert, `scripts/release-agent.sh 0.154.0`: sha `36a54ad2…`, bundle `93487989…`, tag `v0.154.0` = `db25b46`, verified by
kernel-cancel, kernel-status; R20–R23). Bundle: `/etc/grub.d/01_felhom_oneshot`, `/etc/grub.d/42_felhom_oneshot`. download. CI 1576 (code) success; 1579 (CHANGELOG) failed in its fetch step (Gitea answered 500), re-run → success.
- Agent `internal/osupdate/kernel.go`: the night step (told nights only), `KernelAfterBoot`, `KernelStepExecutor`. Signed `agent_update` → demo-hp, demo-felhom, Tester 1 (07:16–07:18 local); signed `agent_config_update` → `BUNDLE DONE
- Tests: `KernelLane` (27), `KernelStepCannotLeaveTheBoxOff` (3), `TestKernel*` (13); red-proofs written=1 same=26 self-check=ok`, capability probe 68/68 on all three. NOT vouched (the vouch needs a golden at or above
`felhom.eu/documentation/audits/kernel-lane-2026-10-07/A/redproof.txt`. the fleet's controller; no golden today). Live: demo-felhom's first report after the update carries the 02:40Z backup.
- Released `v0.152.0` (`d03ab7f`; binary `95ff4220…`, bundle `f0c2cec3…`) + step bundle `0.152.0-step1` (`0b71d32b…`). Evidence: `felhom.eu/documentation/audits/release-2026-10-09/delivery/`.
Delivered by signed jobs to Tester 1, demo-hp, demo-felhom (binary → step bundle → bundle). Not vouched (the golden is
behind; waiver to 2026-10-13).
- **Proven on the Tester 1 box** (`felhom.eu/documentation/audits/kernel-lane-2026-10-07/E/RESULT.md`): forced panic →
fell_back by itself; held guest → one self-revert after 20 min; healthy → 7.0.14-22 the default.
- **Open:** the ring-0 night run (R-836 says where it stopped); R-898 (ring 0 stages the pending kernel, not exactly the
told one); R-897 (post-reboot drive re-bind races the first backup).
+10 -2
View File
@@ -1904,11 +1904,15 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
return nil, 0, ferr return nil, 0, ferr
} }
out := make([]escrow.RetainedBlob, 0, len(resp.Packages)) out := make([]escrow.RetainedBlob, 0, len(resp.Packages))
// R-304: every package the hub holds and the code will NOT be tried against — no key material,
// over the hub's cap, or malformed here. A refusal may call the code wrong only when this is 0.
withheld := resp.UnopenableCount + resp.TruncatedCount
for _, p := range resp.Packages { for _, p := range resp.Packages {
blob, derr := base64.StdEncoding.DecodeString(p.IdentityEscrowB64) blob, derr := base64.StdEncoding.DecodeString(p.IdentityEscrowB64)
if derr != nil || len(blob) == 0 { if derr != nil || len(blob) == 0 {
// One malformed package must not sink the rest — the customer's code may open a // One malformed package must not sink the rest — the customer's code may open a
// later one, and a skipped entry is strictly better than a refusal we cannot justify. // later one, and a skipped entry is strictly better than a refusal we cannot justify.
withheld++
continue continue
} }
out = append(out, escrow.RetainedBlob{ out = append(out, escrow.RetainedBlob{
@@ -1918,9 +1922,13 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
Index: p.Index, Index: p.Index,
}) })
} }
return out, resp.UnopenableCount, nil return out, withheld, nil
}, },
} }
// R-894: ONE instance — the local API writes it, the host report reads it (2026-10-09: a restart right
// after the night backup erased the hub's evidence of it).
lastKnownBackups := backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json"))
collector.SetKnownBackupReporter(lastKnownBackups)
srv, err := localapi.NewServer(localapi.Options{ srv, err := localapi.NewServer(localapi.Options{
EscrowRecovery: escrowRecoverer, EscrowRecovery: escrowRecoverer,
ListenAddr: cfg.LocalAPI.ListenAddr, ListenAddr: cfg.LocalAPI.ListenAddr,
@@ -1933,7 +1941,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
Store: store, Store: store,
// R-894: the newest success per tier on disk — the due-check's fallback when the storage cannot be // R-894: the newest success per tier on disk — the due-check's fallback when the storage cannot be
// read right after a restart. Same state dir as restore-test-state.json. // read right after a restart. Same state dir as restore-test-state.json.
LastKnownBackups: backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json")), LastKnownBackups: lastKnownBackups,
Storage: observer, Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages) DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
+31
View File
@@ -12,12 +12,18 @@ import (
// //
// COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this // COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this
// fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored. // fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored.
//
// 2026-10-09: the SAME instance also feeds the host report (collector.SetKnownBackupReporter), so a
// restart right after a backup no longer erases the hub's evidence of it. The field may name a local
// variable; the variable must be built by backup.NewBackupSuccessState and be the one handed to the
// collector. RED-PROOF (observed): drop the SetKnownBackupReporter call → "not handed to the collector".
func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) { func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
_, f := parseMain(t) _, f := parseMain(t)
if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] { if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one") t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one")
} }
var field, built bool var field, built bool
var fieldVar, handed string
for _, d := range f.Decls { for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl) fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil { if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
@@ -45,6 +51,28 @@ func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
if callsIn(kv.Value)["backup.NewBackupSuccessState"] { if callsIn(kv.Value)["backup.NewBackupSuccessState"] {
built = true built = true
} }
if id, ok := kv.Value.(*ast.Ident); ok {
fieldVar = id.Name
}
}
}
return true
})
// a local variable: built by NewBackupSuccessState, and handed to the collector
ast.Inspect(fd.Body, func(n ast.Node) bool {
switch x := n.(type) {
case *ast.AssignStmt:
for i, l := range x.Lhs {
if id, ok := l.(*ast.Ident); ok && fieldVar != "" && id.Name == fieldVar && i < len(x.Rhs) &&
callsIn(x.Rhs[i])["backup.NewBackupSuccessState"] {
built = true
}
}
case *ast.CallExpr:
if fn, ok := x.Fun.(*ast.SelectorExpr); ok && fn.Sel.Name == "SetKnownBackupReporter" && len(x.Args) == 1 {
if id, ok := x.Args[0].(*ast.Ident); ok {
handed = id.Name
}
} }
} }
return true return true
@@ -56,6 +84,9 @@ func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
if !built { if !built {
t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState") t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState")
} }
if handed == "" || handed != fieldVar {
t.Fatalf("the saved backups (%q) are not handed to the collector (SetKnownBackupReporter got %q)", fieldVar, handed)
}
} }
func callsIn(n ast.Node) map[string]bool { func callsIn(n ast.Node) map[string]bool {
+6 -1
View File
@@ -93,8 +93,13 @@ OOMCHECK_TIMEOUTS = {"image": 10, "clock": 5, "run": 30, "inspect": 10, "events"
OOMCHECK_SETTLE = 2 OOMCHECK_SETTLE = 2
# Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host # Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host
# reboot is needed for them to take effect, and a bad one can stop the box from booting. # reboot is needed for them to take effect, and a bad one can stop the box from booting.
# proxmox-secure-boot-support is the Secure Boot meta-package: its only job is to pull shim-signed and the signed GRUB,
# so it belongs with them. Measured 2026-10-09 on demo-hp (Secure Boot on): left in the pve lane, its upgrade pulled
# shim-signed-common, the whole pve step was refused R6, and the night's kernel step was skipped with it
# (`audits/kernel-night-2026-10-08/`). Pinned by test_ring0_pending_pve_leaves_the_secure_boot_meta_out.
HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|" HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|"
r"pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr)") r"pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr|"
r"proxmox-secure-boot-support)")
# restart_needed() leaves out processes whose cgroup line matches (grep basic regex). Host: the LXC guests' own # restart_needed() leaves out processes whose cgroup line matches (grep basic regex). Host: the LXC guests' own
# processes (`0::/lxc/<vmid>/...`) -- NOT lxc-start itself, whose cgroup is `0::/lxc.monitor/<vmid>` (measured # processes (`0::/lxc/<vmid>/...`) -- NOT lxc-start itself, whose cgroup is `0::/lxc.monitor/<vmid>` (measured
# 2026-10-04 on demo-felhom: the old pattern "lxc" hid lxc-start with 20 deleted maps, so "reboot needed" stayed false # 2026-10-04 on demo-felhom: the old pattern "lxc" hid lxc-start with 20 deleted maps, so "reboot needed" stayed false
+17
View File
@@ -1469,6 +1469,23 @@ class PVELane(unittest.TestCase):
"pending-pve: installed, Proxmox-origin, never kernel/boot/firmware, never Debian or other origins") "pending-pve: installed, Proxmox-origin, never kernel/boot/firmware, never Debian or other origins")
self.assertEqual(f.installed["libc6"], "2.41-12+deb13u3") self.assertEqual(f.installed["libc6"], "2.41-12+deb13u3")
# 2026-10-09, demo-hp (Secure Boot on): proxmox-secure-boot-support's upgrade pulls shim-signed-common; in the pve
# plan that refused the whole step (R6) and skipped the night's kernel step. It is a boot-chain package: left out.
# RED-PROOF: drop proxmox-secure-boot-support from HOST_SLOW_RE -> it is in the plan -> this test fails.
def test_ring0_pending_pve_leaves_the_secure_boot_meta_out(self):
f = pve_fake(ring0=True)
f.plan["select"], f.plan["packages"] = "pending-pve", []
f.installed.update({"proxmox-secure-boot-support": "9.0.1"})
f.pending_sim = [
"Inst pve-manager [9.2.2] (9.2.21 Proxmox Debian Repository:stable [amd64])",
"Inst proxmox-secure-boot-support [9.0.1] (9.0.2 Proxmox Debian Repository:stable [all])",
"Inst shim-signed-common [1.48+pmx1+16.1-1+pmx1] (1.51+pmx1+16.1-2+pmx1 Proxmox Debian Repository:stable [all])",
]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(sorted(u["name"] for u in rep["upgraded"]), ["pve-manager"])
self.assertEqual(f.installed["proxmox-secure-boot-support"], "9.0.1")
def test_select_pending_pve_needs_the_pve_layer(self): def test_select_pending_pve_needs_the_pve_layer(self):
f = Fake() f = Fake()
f.plan["layer"], f.plan["select"], f.plan["packages"] = "host", "pending-pve", [] f.plan["layer"], f.plan["select"], f.plan["packages"] = "host", "pending-pve", []
+23
View File
@@ -103,6 +103,29 @@ func (s *BackupSuccessState) LastKnownSuccess(target string, vmid int) (time.Tim
return e.at, ok return e.at, ok
} }
// KnownBackupSuccesses returns every saved success as a host-report record (the collector's
// KnownBackupReporter, so the hub keeps its evidence across a restart). Only the tier, guest and start
// time are known; the archive name is not saved and stays empty. Sorted for a stable report.
func (s *BackupSuccessState) KnownBackupSuccesses() []hub.Backup {
if s == nil {
return nil
}
s.mu.Lock()
defer s.mu.Unlock()
out := make([]hub.Backup, 0, len(s.last))
for _, e := range s.last {
out = append(out, hub.Backup{TargetID: e.target, VMID: e.vmid, Success: true, CrashConsistent: true,
StartedAt: e.at.Format(time.RFC3339), UncoveredVolumes: []string{}})
}
sort.Slice(out, func(i, j int) bool {
if out[i].TargetID != out[j].TargetID {
return out[i].TargetID < out[j].TargetID
}
return out[i].VMID < out[j].VMID
})
return out
}
func (s *BackupSuccessState) saveLocked() error { func (s *BackupSuccessState) saveLocked() error {
entries := make([]backupSuccessJSON, 0, len(s.last)) entries := make([]backupSuccessJSON, 0, len(s.last))
for _, e := range s.last { for _, e := range s.last {
+17
View File
@@ -57,3 +57,20 @@ func TestBackupSuccessState_CorruptFileIsEmpty(t *testing.T) {
t.Fatal("a corrupt file must read as nothing known") t.Fatal("a corrupt file must read as nothing known")
} }
} }
// The saved successes come back as host-report records: success, tier, guest, start time (R-894 → the
// host report, 2026-10-09). A new state read from the same file gives the same records (the restart case).
func TestBackupSuccessState_KnownBackupSuccessesSurviveAReopen(t *testing.T) {
path := filepath.Join(t.TempDir(), "backup-success-state.json")
s := NewBackupSuccessState(path)
if err := s.RecordBackupSuccess("felhom-backup", hub.Backup{VMID: 9201, Success: true, StartedAt: "2026-10-09T02:40:16Z"}); err != nil {
t.Fatal(err)
}
got := NewBackupSuccessState(path).KnownBackupSuccesses()
if len(got) != 1 || !got[0].Success || got[0].TargetID != "felhom-backup" || got[0].VMID != 9201 || got[0].StartedAt != "2026-10-09T02:40:16Z" {
t.Fatalf("got %+v", got)
}
if got[0].UncoveredVolumes == nil {
t.Fatal("uncovered_volumes must marshal as [], not null")
}
}
+50 -10
View File
@@ -59,8 +59,26 @@ var (
// It carries no material and no code: only WHICH earlier package opened, by its supersession date, // It carries no material and no code: only WHICH earlier package opened, by its supersession date,
// which is the one fact the customer needs to recognise it. // which is the one fact the customer needs to recognise it.
ErrCodeOpensRetained = errors.New("escrow: the recovery code did not open the CURRENT sealed package, but it DID open a retained earlier one") ErrCodeOpensRetained = errors.New("escrow: the recovery code did not open the CURRENT sealed package, but it DID open a retained earlier one")
// ErrRetainedUnchecked — the code did NOT open the current package, and NOT every earlier package the hub
// holds for this host was tried (R-304, 2026-10-08): the retained list could not be fetched, the hub
// withheld rows (over its serve cap, or rows with no key material), a package was malformed, or the
// attempt cap stopped the loop. So „the code is wrong" is NOT known — it may be right for a package
// nobody tried. Distinct from a mistype for exactly the R-224 reason: never accuse the customer of
// something we did not check.
ErrRetainedUnchecked = errors.New("escrow: the recovery code did not open the current sealed package, and some earlier packages were NOT checked")
) )
// RetainedUncheckedError wraps ErrRetainedUnchecked with how many earlier packages went unchecked
// (-1 = unknown: the retained list itself could not be read). No secret.
type RetainedUncheckedError struct {
Unchecked int
}
func (e *RetainedUncheckedError) Error() string {
return fmt.Sprintf("%s (unchecked=%d)", ErrRetainedUnchecked.Error(), e.Unchecked)
}
func (e *RetainedUncheckedError) Unwrap() error { return ErrRetainedUnchecked }
// RetainedMatch says which retained package a code opened. Returned inside RetainedOpenedError; it // RetainedMatch says which retained package a code opened. Returned inside RetainedOpenedError; it
// carries no secret — not the code, not the bundle, not the repository password. // carries no secret — not the code, not the bundle, not the repository password.
type RetainedMatch struct { type RetainedMatch struct {
@@ -102,8 +120,10 @@ type RetainedBlob struct {
} }
// RetainedFetcher yields this host's RETAINED sealed packages, newest-superseded first. An empty // RetainedFetcher yields this host's RETAINED sealed packages, newest-superseded first. An empty
// slice is a clean "none". R-311. // slice is a clean "none". R-311. `withheld` counts the earlier packages the hub holds that are NOT in
type RetainedFetcher func(ctx context.Context) (blobs []RetainedBlob, unopenable int, err error) // blobs — rows with no key material, rows over the hub's serve cap, and packages dropped as malformed
// (R-304): each is a package the code was never tried against.
type RetainedFetcher func(ctx context.Context) (blobs []RetainedBlob, withheld int, err error)
// OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R. // OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R.
type OffsiteKeyRecoverer struct { type OffsiteKeyRecoverer struct {
@@ -154,9 +174,14 @@ func (r OffsiteKeyRecoverer) RecoverOffsiteRepoPassword(ctx context.Context, rec
// from the unwrap alone; the only way to tell is to try. Until this existed nobody tried, and // from the unwrap alone; the only way to tell is to try. Until this existed nobody tried, and
// the screen said so out loud ("innen nem tudjuk megkülönböztetni őket") — a true sentence // the screen said so out loud ("innen nem tudjuk megkülönböztetni őket") — a true sentence
// about our own incuriosity, read by the customer as a statement about their code. // about our own incuriosity, read by the customer as a statement about their code.
if m, ok := r.tryRetained(ctx, recoveryCode); ok { m, ok, unchecked := r.tryRetained(ctx, recoveryCode)
if ok {
return "", &RetainedOpenedError{Match: m} return "", &RetainedOpenedError{Match: m}
} }
// R-304: only when EVERY earlier package the hub holds was tried may this stay a wrong code.
if unchecked != 0 {
return "", &RetainedUncheckedError{Unchecked: unchecked}
}
return "", err // the fail-closed "the recovery code did not unwrap…" message; no secret in it return "", err // the fail-closed "the recovery code did not unwrap…" message; no secret in it
} }
if bundle.ResticRepoPassword == "" { if bundle.ResticRepoPassword == "" {
@@ -173,23 +198,38 @@ func (r OffsiteKeyRecoverer) RecoverOffsiteRepoPassword(ctx context.Context, rec
// function breaking is the behaviour we had before it existed. // function breaking is the behaviour we had before it existed.
// //
// NOTHING IS LOGGED HERE and no return value carries the code, a bundle or a password. // NOTHING IS LOGGED HERE and no return value carries the code, a bundle or a password.
func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode string) (RetainedMatch, bool) { //
// R-304 (2026-10-08): it also returns how many earlier packages were NOT tried — `withheld` from the hub, plus
// the ones past the attempt cap — or -1 when the retained list could not be read at all. A nil FetchRetained
// (an agent wired without the lookup) reports 0: the pre-R-311 refusal, unchanged.
func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode string) (RetainedMatch, bool, int) {
if r.FetchRetained == nil { if r.FetchRetained == nil {
return RetainedMatch{}, false return RetainedMatch{}, false, 0
} }
blobs, _, err := r.FetchRetained(ctx) blobs, withheld, err := r.FetchRetained(ctx)
if err != nil || len(blobs) == 0 { if err != nil {
return RetainedMatch{}, false return RetainedMatch{}, false, -1
}
if withheld < 0 {
withheld = 0
}
if len(blobs) == 0 {
return RetainedMatch{}, false, withheld
} }
limit := r.MaxRetainedTried limit := r.MaxRetainedTried
if limit <= 0 { if limit <= 0 {
limit = defaultMaxRetainedTried limit = defaultMaxRetainedTried
} }
unchecked := withheld
if len(blobs) > limit {
unchecked += len(blobs) - limit
}
for i, rb := range blobs { for i, rb := range blobs {
if i >= limit { if i >= limit {
break break
} }
if len(rb.Blob) == 0 { if len(rb.Blob) == 0 {
unchecked++
continue continue
} }
bundle, uerr := UnwrapIdentityBundle(ctx, rb.Blob, recoveryCode) bundle, uerr := UnwrapIdentityBundle(ctx, rb.Blob, recoveryCode)
@@ -204,7 +244,7 @@ func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode strin
// correct and must be told so — but the history behind it still cannot be reopened, and // correct and must be told so — but the history behind it still cannot be reopened, and
// saying otherwise would be a promise this path cannot keep. // saying otherwise would be a promise this path cannot keep.
HasResticPassword: bundle.ResticRepoPassword != "", HasResticPassword: bundle.ResticRepoPassword != "",
}, true }, true, 0
} }
return RetainedMatch{}, false return RetainedMatch{}, false, unchecked
} }
+76 -2
View File
@@ -115,8 +115,9 @@ func TestRecover_WrongCode_StaysAPlainRefusal(t *testing.T) {
} }
} }
// FAIL-SAFE — if the retained lookup itself fails, the original refusal must stand UNCHANGED. The // FAIL-SAFE — if the retained lookup itself fails, the lookup's error never reaches the customer and no
// worst outcome of this feature breaking is the behaviour we had before it. // retained package is claimed. R-304 (2026-10-08) changed what stands instead: not the wrong-code refusal (the
// earlier packages were never tried, so „wrong" is not known) but ErrRetainedUnchecked with Unchecked = -1.
// //
// RED-PROOF: make tryRetained propagate the fetch error instead of returning false → the customer // RED-PROOF: make tryRetained propagate the fetch error instead of returning false → the customer
// gets a new, unexplained failure mode → this FAILS. // gets a new, unexplained failure mode → this FAILS.
@@ -140,6 +141,10 @@ func TestRecover_RetainedFetchFails_OriginalRefusalStands(t *testing.T) {
if containsStr(err.Error(), "hub exploded") { if containsStr(err.Error(), "hub exploded") {
t.Error("the retained-lookup failure leaked into the customer-facing refusal — it must be silent") t.Error("the retained-lookup failure leaked into the customer-facing refusal — it must be silent")
} }
var ue *RetainedUncheckedError
if !errors.As(err, &ue) || ue.Unchecked != -1 {
t.Errorf("err = %v, want RetainedUncheckedError{-1} — the earlier packages were never tried (R-304)", err)
}
} }
// A nil FetchRetained keeps the pre-R-311 behaviour EXACTLY. An agent wired without it must be // A nil FetchRetained keeps the pre-R-311 behaviour EXACTLY. An agent wired without it must be
@@ -228,3 +233,72 @@ func containsStr(hay, needle string) bool {
return false return false
})() })()
} }
// ── R-304 (2026-10-08) — „wrong code" only when every earlier package was tried ──────────────────────
//
// The consequence asserted: is the refusal a WRONG CODE (the only error the screen may answer with „check your
// typing")? It may be only when the hub withheld nothing and every served package was tried.
//
// RED-PROOF: make tryRetained return 0 for `unchecked` → the three „unchecked" cases return the plain refusal
// → they FAIL; the all-tried case keeps passing.
func TestR304_WrongCodeOnlyWhenEveryEarlierPackageWasTried(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "4444567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
other := sealBundle(t, IdentityBundle{ResticRepoPassword: "5555567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
const code = "a code that opens nothing whatsoever in this test"
cases := []struct {
name string
blobs int
withheld int
limit int
unchecked int // 0 = must be a plain wrong code
}{
{"all tried, nothing withheld", 2, 0, 6, 0},
{"the hub withheld rows (no key material / over its cap)", 1, 3, 6, 3},
{"only withheld rows, nothing served", 0, 2, 6, 2},
{"the attempt cap stopped the loop", 5, 0, 2, 3},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
blobs := make([]RetainedBlob, 0, c.blobs)
for i := 0; i < c.blobs; i++ {
blobs = append(blobs, RetainedBlob{Blob: other, SupersededAt: "2026-08-01 00:00:00", Index: i})
}
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
return blobs, c.withheld, nil
},
MaxRetainedTried: c.limit,
}.RecoverOffsiteRepoPassword(context.Background(), code)
if err == nil {
t.Fatal("a code that opens nothing succeeded")
}
var ue *RetainedUncheckedError
isUnchecked := errors.As(err, &ue)
if c.unchecked == 0 {
if isUnchecked {
t.Fatalf("every earlier package was tried — this IS a wrong code, got %v", err)
}
return
}
if !isUnchecked || ue.Unchecked != c.unchecked || !errors.Is(err, ErrRetainedUnchecked) {
t.Fatalf("err = %v, want RetainedUncheckedError{%d} — never call the code wrong when packages were not tried", err, c.unchecked)
}
})
}
}
// A malformed (empty) served package counts as not tried.
func TestR304_EmptyServedPackageCountsAsUnchecked(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "6666567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: nil, SupersededAt: "2026-08-01 00:00:00"}),
}.RecoverOffsiteRepoPassword(context.Background(), testR2)
var ue *RetainedUncheckedError
if !errors.As(err, &ue) || ue.Unchecked != 1 {
t.Fatalf("err = %v, want RetainedUncheckedError{1}", err)
}
}
+49 -5
View File
@@ -66,6 +66,13 @@ type ProvenRestoreTestReporter interface {
ProvenRestoreTests(ctx context.Context) []RestoreTest ProvenRestoreTests(ctx context.Context) []RestoreTest
} }
// KnownBackupReporter is the newest SUCCESSFUL backup per tier and guest kept on disk (R-894's
// backup-success-state.json). collectBackups folds it into the report so a restart does not erase the
// hub's evidence of a backup that ran minutes before it (the kernel-night shape, 2026-10-09).
type KnownBackupReporter interface {
KnownBackupSuccesses() []Backup
}
// PBSReporter is the slice-6-Phase-B seam the pbs verify loop plugs into (same pattern). // PBSReporter is the slice-6-Phase-B seam the pbs verify loop plugs into (same pattern).
// Returns the agent's latest-known PBS snapshot inventory + verify-state. nil → empty. // Returns the agent's latest-known PBS snapshot inventory + verify-state. nil → empty.
type PBSReporter interface { type PBSReporter interface {
@@ -105,6 +112,7 @@ type Collector struct {
backups BackupReporter backups BackupReporter
restoreTests RestoreTestReporter restoreTests RestoreTestReporter
provenTests ProvenRestoreTestReporter provenTests ProvenRestoreTestReporter
knownBackups KnownBackupReporter // R-894 file → host report after a restart (nil → in-memory only)
pbs PBSReporter pbs PBSReporter
temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp) temp TempReader // slice 9: host CPU/chassis temp (nil-safe → nil temp)
capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty) capProbe func(ctx context.Context) []capability.Status // v0.44.0: privileged-capability self-check (nil → empty)
@@ -576,14 +584,50 @@ func (c *Collector) collectStorage(ctx context.Context) []StorageTarget {
// collectBackups / collectRestoreTests read the agent's latest backup + restore-test state // collectBackups / collectRestoreTests read the agent's latest backup + restore-test state
// via the seams. Best-effort: a nil reporter or nil slice degrades to an empty (non-nil) // via the seams. Best-effort: a nil reporter or nil slice degrades to an empty (non-nil)
// list so the collection always marshals as []. // list so the collection always marshals as [].
//
// The saved successes (SetKnownBackupReporter) are added for each tier and guest the in-memory list has
// no success for at or after the saved time. Why: the in-memory list is empty after an agent restart,
// and the hub's backup-freshness check reads only what the reports carried. Measured 2026-10-09 on
// demo-felhom: the night backup landed 02:40 UTC, the kernel step restarted the host at 02:44, no report
// fell in between, so two kernel nights in a row left no trace and the hub alarmed „newest backup is 48h
// old" at 03:00. Only SUCCESSES are saved, so a failure is never hidden and never invented.
// Pinned by TestCollectBackups_SavedSuccessSurvivesARestart.
func (c *Collector) collectBackups(ctx context.Context) []Backup { func (c *Collector) collectBackups(ctx context.Context) []Backup {
if c.backups == nil { out := []Backup{}
return []Backup{} if c.backups != nil {
}
if b := c.backups.Backups(ctx); b != nil { if b := c.backups.Backups(ctx); b != nil {
return b out = append(out, b...)
} }
return []Backup{} }
if c.knownBackups == nil {
return out
}
for _, k := range c.knownBackups.KnownBackupSuccesses() {
kt, err := time.Parse(time.RFC3339, k.StartedAt)
if err != nil {
continue
}
covered := false
for _, b := range out {
if !b.Success || b.TargetID != k.TargetID || b.VMID != k.VMID {
continue
}
if bt, err := time.Parse(time.RFC3339, b.StartedAt); err == nil && !bt.Before(kt) {
covered = true
break
}
}
if !covered {
out = append(out, k)
}
}
return out
}
// SetKnownBackupReporter wires the R-894 saved successes into the host report (nil-safe → in-memory only).
func (c *Collector) SetKnownBackupReporter(r KnownBackupReporter) *Collector {
c.knownBackups = r
return c
} }
// collectRestoreTests merges the in-memory result with the PERSISTED per-tier proofs (R-189). // collectRestoreTests merges the in-memory result with the PERSISTED per-tier proofs (R-189).
+62
View File
@@ -0,0 +1,62 @@
package hub
import (
"context"
"testing"
)
type fakeMemBackups []Backup
func (f fakeMemBackups) Backups(context.Context) []Backup { return f }
type fakeKnownBackups []Backup
func (f fakeKnownBackups) KnownBackupSuccesses() []Backup { return f }
// The consequence, not the mechanism: after a restart (empty in-memory list) the host report still carries
// the night's backup, so the hub's freshness check sees it. Measured 2026-10-09 on demo-felhom: without this
// the report carried backups: [] and the hub alarmed „newest backup is 48h old" an hour after a good backup.
// RED-PROOF: return before the saved-success loop in collectBackups → the first case reports nothing → fails.
func TestCollectBackups_SavedSuccessSurvivesARestart(t *testing.T) {
saved := Backup{TargetID: "felhom-backup", VMID: 9201, Success: true, StartedAt: "2026-10-09T02:40:16Z"}
t.Run("restart: memory empty, the saved success is reported", func(t *testing.T) {
c := &Collector{backups: fakeMemBackups(nil), knownBackups: fakeKnownBackups{saved}}
got := c.collectBackups(context.Background())
if len(got) != 1 || got[0].StartedAt != saved.StartedAt || !got[0].Success || got[0].TargetID != "felhom-backup" {
t.Fatalf("want the saved success, got %+v", got)
}
})
t.Run("memory holds the same or a newer success: no second entry", func(t *testing.T) {
mem := Backup{TargetID: "felhom-backup", VMID: 9201, Success: true, StartedAt: "2026-10-09T02:40:16Z", Archive: "a"}
c := &Collector{backups: fakeMemBackups{mem}, knownBackups: fakeKnownBackups{saved}}
if got := c.collectBackups(context.Background()); len(got) != 1 || got[0].Archive != "a" {
t.Fatalf("want only the in-memory record, got %+v", got)
}
})
t.Run("a newer FAILURE in memory never hides the saved success, and is kept", func(t *testing.T) {
fail := Backup{TargetID: "felhom-backup", VMID: 9201, Success: false, StartedAt: "2026-10-10T02:40:00Z", Error: "x"}
c := &Collector{backups: fakeMemBackups{fail}, knownBackups: fakeKnownBackups{saved}}
got := c.collectBackups(context.Background())
if len(got) != 2 || got[0].Success || !got[1].Success {
t.Fatalf("want the failure and the saved success, got %+v", got)
}
})
t.Run("another tier's success does not cover this tier", func(t *testing.T) {
other := Backup{TargetID: "felhom-pbs", VMID: 9201, Success: true, StartedAt: "2026-10-10T00:00:00Z"}
c := &Collector{backups: fakeMemBackups{other}, knownBackups: fakeKnownBackups{saved}}
if got := c.collectBackups(context.Background()); len(got) != 2 {
t.Fatalf("want both tiers, got %+v", got)
}
})
t.Run("no saved state wired: the in-memory list as before, never nil", func(t *testing.T) {
c := &Collector{}
if got := c.collectBackups(context.Background()); got == nil || len(got) != 0 {
t.Fatalf("want an empty non-nil list, got %#v", got)
}
})
}
+17 -4
View File
@@ -15,7 +15,7 @@ import (
// Red-proof: drop the `b.Success &&` guard and the failed-backup sub-case fails; move the call after release() // Red-proof: drop the `b.Success &&` guard and the failed-backup sub-case fails; move the call after release()
// and the gate sub-case fails. // and the gate sub-case fails.
func TestAfterPrimaryBackup(t *testing.T) { func TestAfterPrimaryBackup(t *testing.T) {
run := func(t *testing.T, failErr string) (calls []int, gateHeld bool) { run := func(t *testing.T, failErr, path string) (calls []int, gateHeld bool) {
gate := &backup.InFlight{} gate := &backup.InFlight{}
b := &fakeBackups{failErr: failErr} b := &fakeBackups{failErr: failErr}
srv := newTestServerS(t, &fakeGuests{}, b, &fakeStore{}, nil) srv := newTestServerS(t, &fakeGuests{}, b, &fakeStore{}, nil)
@@ -34,7 +34,7 @@ func TestAfterPrimaryBackup(t *testing.T) {
done <- struct{}{} done <- struct{}{}
}) })
h := srv.Handler() h := srv.Handler()
if do(t, h, "POST", "/backup", "A", "").Code != http.StatusAccepted { if do(t, h, "POST", path, "A", "").Code != http.StatusAccepted {
t.Fatal("POST /backup not accepted") t.Fatal("POST /backup not accepted")
} }
select { select {
@@ -47,7 +47,7 @@ func TestAfterPrimaryBackup(t *testing.T) {
return calls, gateHeld return calls, gateHeld
} }
t.Run("success runs the leg under the gate", func(t *testing.T) { t.Run("success runs the leg under the gate", func(t *testing.T) {
calls, held := run(t, "") calls, held := run(t, "", "/backup")
if len(calls) != 1 { if len(calls) != 1 {
t.Fatalf("the leg ran %d time(s), want 1", len(calls)) t.Fatalf("the leg ran %d time(s), want 1", len(calls))
} }
@@ -56,8 +56,21 @@ func TestAfterPrimaryBackup(t *testing.T) {
} }
}) })
t.Run("a failed backup runs nothing", func(t *testing.T) { t.Run("a failed backup runs nothing", func(t *testing.T) {
if calls, _ := run(t, "vzdump exploded"); len(calls) != 0 { if calls, _ := run(t, "vzdump exploded", "/backup"); len(calls) != 0 {
t.Fatalf("the leg ran after a FAILED backup: %v", calls) t.Fatalf("the leg ran after a FAILED backup: %v", calls)
} }
}) })
// R-899: a household press is not the night's backup — no OS leg after it. Same fake, same successful backup as
// the first sub-case; only the query differs. Red-proof: make handleBackup ignore `trigger` and this sub-case
// fails with the leg run once.
t.Run("a manual press runs nothing", func(t *testing.T) {
if calls, _ := run(t, "", "/backup?trigger=manual"); len(calls) != 0 {
t.Fatalf("the OS leg ran after a manual press: %v", calls)
}
})
t.Run("the scheduled path with the new query still runs the leg", func(t *testing.T) {
if calls, _ := run(t, "", "/backup?trigger=night"); len(calls) != 1 {
t.Fatalf("the leg ran %d time(s) after a non-manual backup, want 1", len(calls))
}
})
} }
+16
View File
@@ -127,6 +127,22 @@ func (s *Server) handleRecoverOffsitePassword(w http.ResponseWriter, r *http.Req
"retained_has_restic_pw": match.HasResticPassword, "retained_has_restic_pw": match.HasResticPassword,
}, },
"the recovery code is correct, but it belongs to an EARLIER sealed package (superseded "+match.SupersededAt+"), not the one currently held") "the recovery code is correct, but it belongs to an EARLIER sealed package (superseded "+match.SupersededAt+"), not the one currently held")
// ── R-304 (2026-10-08) — NOT EVERY EARLIER PACKAGE WAS CHECKED. ─────────────────────────
//
// The current package refused the code, no retained package opened it — and at least one earlier
// package the hub holds was never tried (or the list could not be read). Saying „the code is
// wrong" here would claim a check that did not happen. 424 (Failed Dependency): the verdict
// depends on packages we could not try. The controller classifies on the status, never on this
// sentence; an older controller maps an unknown status to its neutral „we do not know why".
case errors.Is(err, escrow.ErrRetainedUnchecked):
n := -1
var ue *escrow.RetainedUncheckedError
if errors.As(err, &ue) {
n = ue.Unchecked
}
s.logger.Warn("local-api: offsite key recovery: the code did not open the current package and earlier packages were NOT all checked — not reported as a wrong code (R-304)", "vmid", vmid, "unchecked", n)
writeStatus(w, http.StatusFailedDependency, false, map[string]any{"older_unchecked": n},
"the recovery code did not open the current sealed package, and earlier packages the hub holds were not all checked — the code may belong to one of them; nothing was written")
case errors.Is(err, escrow.ErrNoResticPassword): case errors.Is(err, escrow.ErrNoResticPassword):
s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid) s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid)
writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)") writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)")
@@ -54,6 +54,13 @@ func TestRecoverOffsitePassword_EachSituationGetsItsOwnStatus(t *testing.T) {
wantStatus: 404, wantStatus: 404,
mustNotSay: []string{"did not open"}, mustNotSay: []string{"did not open"},
}, },
{
// R-304: earlier packages were not all tried — never „did not open the sealed bundle" (the wrong-code words).
name: "earlier packages not all checked — not a wrong code",
err: &escrow.RetainedUncheckedError{Unchecked: 2},
wantStatus: 424,
mustNotSay: []string{"did not open the sealed bundle", "could not be fetched"},
},
{ {
name: "the bundle predates the repository-password field", name: "the bundle predates the repository-password field",
err: escrow.ErrNoResticPassword, err: escrow.ErrNoResticPassword,
+7 -1
View File
@@ -825,6 +825,10 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
return return
} }
key := backupJobKey{vmid: vmid, target: tier.TargetID} key := backupJobKey{vmid: vmid, target: tier.TargetID}
// R-899 (operator ruling 2026-10-08): a household press („Mentés most") is not the night's backup. A controller
// that knows sends `trigger=manual`; then the OS leg does not follow (it belongs to the night, after the night's
// own copy). An older controller sends nothing and keeps the old behaviour.
manual := r.URL.Query().Get("trigger") == "manual"
// ONE BACKUP AT A TIME PER GUEST, ACROSS ALL TIERS (operator ruling 2026-07-26: "other backup // ONE BACKUP AT A TIME PER GUEST, ACROSS ALL TIERS (operator ruling 2026-07-26: "other backup
// shouldn't start until finished"). vzdump takes a guest lock, so a concurrent second backup // shouldn't start until finished"). vzdump takes a guest lock, so a concurrent second backup
@@ -926,7 +930,9 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
} }
s.finishJob(key, jobID, b) s.finishJob(key, jobID, b)
// OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate. // OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate.
if b.Success && tier.Primary && s.afterPrimaryBackup != nil { if b.Success && tier.Primary && s.afterPrimaryBackup != nil && manual {
s.logger.Info("local-api: no OS leg after a manual backup — it follows the night's own backup (R-899)", "vmid", vmid, "job", jobID)
} else if b.Success && tier.Primary && s.afterPrimaryBackup != nil {
s.afterPrimaryBackup(base, vmid) s.afterPrimaryBackup(base, vmid)
} }
}() }()
+23 -5
View File
@@ -51,6 +51,24 @@ type KernelView struct {
VMID int `json:"vmid"` // the customer guest the step was staged for (the health rule's guest) VMID int `json:"vmid"` // the customer guest the step was staged for (the health rule's guest)
} }
// kernelReportRing is the ring an after-boot report carries. In the first second after a boot the agent has not
// fetched the hub's block yet, and Block() then answers ring 1 — so a ring-0 box's „judging" report said ring 1
// (seen on demo-felhom, 2026-10-08 night, `audits/kernel-night-2026-10-07/readback/`). Order: the fetched block; the
// block the daemon saved on disk before the reboot (R-866); ring 1 as before. A label only: the hub's approval reads its
// own ring list. Pinned by TestKernelReportRing_BeforeFirstFetch.
func (l *Leg) kernelReportRing() int {
l.mu.Lock()
fetched := l.block
l.mu.Unlock()
if fetched != nil {
return fetched.Ring
}
if b, _, ok := LoadSavedBlock(l.planDir()); ok && b != nil {
return b.Ring
}
return 1
}
func parseKernel(raw json.RawMessage) KernelView { func parseKernel(raw json.RawMessage) KernelView {
var v KernelView var v KernelView
_ = json.Unmarshal(raw, &v) _ = json.Unmarshal(raw, &v)
@@ -190,7 +208,7 @@ func (l *Leg) KernelAfterBoot(ctx context.Context, vmid int, j KernelJudge) Repo
if vmid <= 0 { if vmid <= 0 {
vmid = v.VMID // after a boot the guest may not run yet — the step's own record names it vmid = v.VMID // after a boot the guest may not run yet — the step's own record names it
} }
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid, Mode: "kernel-boot", rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid, Mode: "kernel-boot",
ReleaseID: v.To, Kernel: rawOrNil(wr.Kernel)} ReleaseID: v.To, Kernel: rawOrNil(wr.Kernel)}
switch wr.KernelEvent { switch wr.KernelEvent {
case "fell_back": case "fell_back":
@@ -240,7 +258,7 @@ func (l *Leg) judgeKernel(ctx context.Context, runID string, vmid int, v KernelV
for { for {
if !hubReached && l.Hub != nil { if !hubReached && l.Hub != nil {
// the hub's reachability IS this report reaching it (and the operator sees the box is back on the new kernel) // the hub's reachability IS this report reaching it (and the operator sees the box is back on the new kernel)
body, _ := json.Marshal(Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid, body, _ := json.Marshal(Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid,
Mode: "kernel-boot", ReleaseID: v.To, Outcome: "judging", Kernel: mustRaw(v)}) Mode: "kernel-boot", ReleaseID: v.To, Outcome: "judging", Kernel: mustRaw(v)})
rctx, cancel := context.WithTimeout(ctx, 30*time.Second) rctx, cancel := context.WithTimeout(ctx, 30*time.Second)
if err := l.Hub.PostOSReport(rctx, body); err == nil { if err := l.Hub.PostOSReport(rctx, body); err == nil {
@@ -280,7 +298,7 @@ func (l *Leg) judgeKernel(ctx context.Context, runID string, vmid int, v KernelV
return Report{} return Report{}
} }
// not healthy by the deadline: tell the hub (best effort), then ONE self-revert into the old kernel // not healthy by the deadline: tell the hub (best effort), then ONE self-revert into the old kernel
rep := l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid, rep := l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid,
Mode: "kernel-revert", ReleaseID: v.To, Outcome: "health_failed", HealthReason: why + " — reverting to " + v.From, Mode: "kernel-revert", ReleaseID: v.To, Outcome: "health_failed", HealthReason: why + " — reverting to " + v.From,
Kernel: mustRaw(v)}) Kernel: mustRaw(v)})
lg.Error("osupdate: kernel step — the one-shot boot is NOT healthy; restarting ONCE into the old kernel", "reason", why, lg.Error("osupdate: kernel step — the one-shot boot is NOT healthy; restarting ONCE into the old kernel", "reason", why,
@@ -289,7 +307,7 @@ func (l *Leg) judgeKernel(ctx context.Context, runID string, vmid int, v KernelV
if err != nil || wr.refused() || wr.failed() { if err != nil || wr.refused() || wr.failed() {
lg.Error("osupdate: kernel self-revert did not start — the box stays on the new kernel; the operator decides", lg.Error("osupdate: kernel self-revert did not start — the box stays on the new kernel; the operator decides",
"err", err, "refused", string(firstRaw(wr.Refused, wr.Failed))) "err", err, "refused", string(firstRaw(wr.Refused, wr.Failed)))
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid, return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid,
Mode: "kernel-revert", ReleaseID: v.To, Outcome: "revert_failed", Refused: firstRaw(wr.Refused, wr.Failed), Mode: "kernel-revert", ReleaseID: v.To, Outcome: "revert_failed", Refused: firstRaw(wr.Refused, wr.Failed),
HealthReason: "the self-revert did not start"}) HealthReason: "the self-revert did not start"})
} }
@@ -298,7 +316,7 @@ func (l *Leg) judgeKernel(ctx context.Context, runID string, vmid int, v KernelV
func (l *Leg) kernelGood(ctx context.Context, runID string, vmid int, v KernelView, start time.Time, lg *slog.Logger) Report { func (l *Leg) kernelGood(ctx context.Context, runID string, vmid int, v KernelView, start time.Time, lg *slog.Logger) Report {
wr, err := l.call(ctx, runID, kernelPlan("kernel-good", vmid, nil)) wr, err := l.call(ctx, runID, kernelPlan("kernel-good", vmid, nil))
rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.Block().Ring, VMID: vmid, Mode: "kernel-good", rep := Report{RunID: runID, Layer: LayerKernel, Trigger: "boot", Ring: l.kernelReportRing(), VMID: vmid, Mode: "kernel-good",
ReleaseID: v.To} ReleaseID: v.To}
switch { switch {
case err != nil: case err != nil:
+32
View File
@@ -0,0 +1,32 @@
package osupdate
import (
"os"
"path/filepath"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// After a boot the agent has not fetched the hub's block yet; the after-boot kernel report must still carry the
// box's real ring (seen 2026-10-08: a ring-0 box's „judging" report said ring 1). Red-proof: return Block().Ring
// from kernelReportRing and the saved-block case fails with 1.
func TestKernelReportRing_BeforeFirstFetch(t *testing.T) {
dir := t.TempDir()
l := &Leg{PlanDir: dir}
if got := l.kernelReportRing(); got != 1 {
t.Fatalf("nothing fetched, nothing saved: ring %d, want 1 (the old default)", got)
}
l.saveBlock(&hub.WireOSUpdate{Ring: 0, Enabled: true})
l2 := &Leg{PlanDir: dir} // a fresh daemon after the reboot: no block fetched yet
if got := l2.kernelReportRing(); got != 0 {
t.Fatalf("ring-0 block saved before the reboot: ring %d, want 0", got)
}
l2.SetBlock(&hub.WireOSUpdate{Ring: 1, Enabled: true})
if got := l2.kernelReportRing(); got != 1 {
t.Fatalf("a fetched block wins over the saved one: ring %d, want 1", got)
}
if _, err := os.Stat(filepath.Join(dir, SavedBlockFile)); err != nil {
t.Fatalf("saved block missing: %v", err)
}
}