Compare commits

...

7 Commits

Author SHA1 Message Date
admin b1746c25af Docker engine slow lane (live-restore once by reload; ring-0 pending-docker under a root-owned ring-0 mark; ring 1 and undo only by a signed os_docker_step the wrapper re-verifies against a root-owned signers file; same-container-id health), the version report (facts mode -> host report system stanza, R-852), guest restart scan every pass (R-849), the crash guard (kernel.panic=10, the 3rd unclean stop in 60 min stays off, 24 h re-arm)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 16:14:36 +02:00
admin 2e2e8f56b8 CHANGELOG + REPORT: v0.141.1 released (host reboot-needed fixes)
gates / gates (push) Successful in 17s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 13:39:35 +02:00
admin a6bc3f1197 os-apply: the host restart scan no longer hides lxc-start (skip ':/lxc/' not 'lxc'), and the host scans on every pass so a reboot clears 'reboot needed' (reboot_scanned reaches the hub) — both found live on demo-felhom
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 13:39:14 +02:00
admin 3bf77c3423 CHANGELOG + REPORT: v0.141.0 released (host fast lane, true tunnel status, fast leg)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 13:19:02 +02:00
admin cfba0d022a contract: the desired-state golden gains host_release (byte-identical with hub v0.131.0); TestOSUpdateGolden_Decodes checks it
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 13:18:43 +02:00
admin c3d08b4821 OS updates host fast lane + true tunnel status + fast leg: wrapper host layer (R12 appliance proof from the root-owned install record, R14 kernel/boot/firmware refused), select pending-fast, one call per layer, host-side version checks, restart scan only after an install, reboot-needed for PID 1/lxc-start; the leg runs the host step after a healthy guest step; GuestTunnelProber reads the cloudflared container + its readiness check (R-841)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 12:54:02 +02:00
admin a55eedcf2c CHANGELOG + REPORT: v0.140.0 released (OS updates, guest fast lane)
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 11:30:14 +02:00
29 changed files with 2760 additions and 430 deletions
+79
View File
@@ -1,3 +1,82 @@
## v0.141.1 — "reboot needed" is true on the host (found live on demo-felhom, 2026-10-04)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.141.1` (`a6bc3f1`), sha256
> `b712f577099fd2d374f648df1e825874821302fe1b46e93c7b4b70044fbe84b5`, verified by download. Not vouched at release time.
**MinAgent impact:** none. **Pairs with hub v0.131.1** (reads `reboot_scanned`); an older hub ignores the field.
A patch release in the same session as v0.141.0 — a deliberate exception to "one release per repo": v0.141.0's host
"reboot needed" was wrong in two ways, and it feeds an operator alarm.
- **The scan hid `lxc-start`.** The host scan skipped every process whose cgroup line contains `lxc`, to leave out the
guests' own processes. `lxc-start` lives in `0::/lxc.monitor/<vmid>`, so it was skipped too. Measured: after a
108-package host pass (libc6 included) `lxc-start` mapped 20 deleted files and the report said `reboot_needed:
false`. The pattern is now `:/lxc/` (`RESTART_SKIP_CGROUP`), pinned by a test that runs `grep` against the measured
cgroup lines.
- **A reboot never cleared it.** v0.141.0 scanned only after an install (R-845). The host now scans on EVERY pass (it
is local, no `pct exec`); the guest still scans only after an install. The report carries `reboot_scanned`.
- Red-proofs: `felhom.eu/documentation/audits/os-host-lane-2026-10-04/partB/live-defects-redproofs.txt`.
## v0.141.0 — OS updates: the host fast lane (`11` §8 step 3); the tunnel status is true (R-841); the leg is fast (R-845)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.141.0` (`cfba0d0`), sha256
> `6eaad9809613c4fca64aefc30cc16415528499bf48efe3a1daaa08af8a611aff`, verified by download. Not vouched at release time.
**MinAgent impact:** none required by any controller. **Needs hub v0.131.0** (`host_release`, the tunnel's three
states, layer-tagged OS reports). With an older hub the host step finds no host release (ring 1 → nothing) and the
tunnel's `detail` is ignored. Reads the cloudflared health check controller v0.292.0 adds; with an older controller
the tunnel is judged on the container state alone and `detail` says so.
- **R-841 — the tunnel.** The agent used to run `systemctl is-active cloudflared` on the HOST — a unit that does not
exist (cloudflared is a container in the customer guest), so every box reported `inactive`. `GuestTunnelProber`
now reads the guest's `cloudflared` container through the EXISTING sudoers line (`pct exec N -- docker inspect -f
*`): state, exit code, and the Docker health status. Three states: `running` (healthy), `not_running` (stopped,
absent, or running but NOT connected), `unknown` (could not ask, or the check is still starting). `detail` says why.
- **The host step** (`11` §8 step 3). After a healthy guest step, under the same heavy-op gate, the leg runs the same
wrapper with `layer: host`. Debian origin only; never kernel, boot or firmware packages (new refusal **R14**);
**only on an appliance** (**R12** lifted for the host fast lane: proof is the ROOT-owned install record
`/var/lib/felhom-install/state.json` `mode: appliance` — the agent-writable `agent.json` is not trusted for this).
A failed or unhealthy guest step skips it. **Host health rule:** the agent, pveproxy, pvedaemon, pvestatd and
pve-cluster are active; the customer guest runs; the guest health rule passes; the tunnel is `running` (an
`unknown` tunnel does not fail the rule; `not_running` does). Never reboots: "reboot needed" is reported when PID 1
or `lxc-start` runs a replaced library. No automatic undo — the by-hand runbook is `os-updates-host-undo.md`.
- **R-845 — the leg is fast.** The wrapper asked about each package in its own `pct exec` (about 0.9 s each). It now
makes one call per layer for the version checks (`apt-cache madison` for all names at once, `dpkg --compare-versions`
on the host), scans for restart-needed only after an install, repairs only when `dpkg --audit` reports something,
and reports its own `pass_seconds`.
- **Wrapper tests:** `configs/test_felhom_os_apply.py` 46 tests (host layer, R12, R14, lxc-start, the speed rules).
Red-proofs: `felhom.eu/documentation/audits/os-host-lane-2026-10-04/partA/`, `partB/`, `partC/agent-golden-redproof.txt`.
- **Contract:** the desired-state golden gains `host_release` (byte-identical with the hub's); `TestOSUpdateGolden_Decodes`
checks it.
## v0.140.0 — OS updates, guest fast lane (`11-os-updates.md` §8 step 2; `09` §3 decisions 76, 79, 80)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.140.0` (`9cac346`), sha256
> `ae2d60b794869c51d6b063c8e31e2da98ecdbd36c6b75febd64e2fb4266c1250`, verified by download. Not vouched at release time.
**MinAgent impact:** none required by any controller. **Needs hub v0.130.0** (`os-report`, the `os_update` block); an
older hub serves no block and the leg then reports and installs nothing (ring 1, no release).
- **`configs/felhom-os-apply`** — the root wrapper (Python 3, stdlib). One sudoers entry, `FELHOM_OSAPPLY`:
`felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json`. Modes `inventory` / `apply` / `health`. Refuses (exit
2, nothing changed) on R1–R13: plan path/owner/JSON, a non-Debian origin, the slow lane, any removal, any downgrade,
a new or unlisted package, a version not downloadable even from the snapshot, low space, a lock (apt or a guest lock
such as a backup), a vmid that is not the box's own customer guest (it must bind `/mnt/felhom-drives`), malformed
names/versions, the host layer, dpkg still broken after the repair. Repairs first (`dpkg --configure -a`,
`apt-get -f install`). A version Debian already replaced comes from `snapshot.debian.org` at the approval time
(decision 79). Reports the full installed set with origins, pending, restart-needed (outside containers), health.
`configs/test_felhom_os_apply.py`: 35 tests; every refusal red-proved.
- **`internal/osupdate`** — the leg: after a SUCCESSFUL primary whole-guest backup, still holding the heavy-op gate
(never beside another backup or a restore-test), once per night, 90 s after the backup. Ring 0 installs every pending
Debian / Debian-Security fix; ring 1 exactly the hub's newest approved release; switched OFF → reports only. The
health rule: docker answers, the network resolves, the controller is healthy, every container running at the START
of the leg runs (and is healthy if it was) — a 5-minute wait. **No automatic undo:** a customer guest cannot be
snapshotted (R-837). A failure is `health_failed` → the hub mails the operator.
- **`--selftest=os-update -vmid N`** — the debug action (trigger `debug`, never throttled, not a night run).
- **The `--selftest` flag also accepts `wgtunnel`** — it was dispatched but refused since S3 (found by the new
`TestSelftestFlag_AcceptsEveryDispatchedMode`, which also caught `os-update` live).
- Proven live 2026-10-04 on both demo boxes (ring 0: 53 packages each; ring 1: exactly 3 approved versions; a failed
health check → `health_failed`, operator mailed): `felhom.eu/documentation/audits/os-guest-lane-2026-10-04/`.
## v0.139.0 — a DR restore never lands beside a live original (2026-10-04, R-834)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.139.0` (`475bdce`), sha256
+9 -8
View File
@@ -1,10 +1,11 @@
# REPORT — 2026-10-04: v0.139.0 (R-834)
# REPORT — 2026-10-04: v0.141.0 + v0.141.1, the host fast lane, the true tunnel status, the fast leg
Full session report: `felhom.eu/REPORT-backup-close-os-spike-2026-10-04.md`.
Full session report: `felhom.eu/REPORT-os-host-lane-2026-10-04.md`.
- **Measured** on demo-hp: the scheduled restore-test's scratch guest has `onboot: 0` and throwaway stand-ins for
mp8/mp9 on every config read until teardown — it was already safe. Evidence `felhom.eu/documentation/audits/backup-close-2026-10-04/partA/`.
- **Fixed:** the DR bring-up refuses beside a live original (source guest present, a drives bind on another guest,
or an unreadable config). On a replaced host it is unchanged.
- Tests `TestRunBringUp_DRRefusesBesideALiveOriginal`, `TestRestoreTest_NoHostPathBindBesideTheOriginal` and two
more; both rules red-proved. No sudoers change (said why in the CHANGELOG).
- Tunnel (R-841): the agent reads the guest's cloudflared container and its health check; three states.
- Host fast lane: the wrapper gains the host layer (R12 appliance proof from the root-owned install record, R14 no
kernel/boot/firmware); the leg runs the host step after a healthy guest step; host health rule.
- Speed (R-845): one call per layer instead of one per package; measured before/after in the session report.
- Tests green; red-proofs in the audit folder.
- v0.141.1 (same day): the host "reboot needed" scan hid `lxc-start` and was never cleared by a reboot — both fixed,
found live on demo-felhom, red-proved.
+184 -17
View File
@@ -44,11 +44,11 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/localapi"
applog "gitea.dooplex.hu/admin/felhom-agent/internal/log"
"gitea.dooplex.hu/admin/felhom-agent/internal/mgmtplane"
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
"gitea.dooplex.hu/admin/felhom-agent/internal/pbs"
"gitea.dooplex.hu/admin/felhom-agent/internal/pbsdr"
"gitea.dooplex.hu/admin/felhom-agent/internal/poke"
"gitea.dooplex.hu/admin/felhom-agent/internal/provision"
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/restorespace"
@@ -245,6 +245,10 @@ func main() {
os.Exit(runSelftestRestoreTestDue(context.Background(), cfg, logger))
case "os-update":
os.Exit(runSelftestOSUpdate(context.Background(), cfg, logger, vmid))
case "os-facts":
os.Exit(runSelftestFacts(context.Background(), cfg, logger, vmid))
case "live-restore":
os.Exit(runSelftestLiveRestore(context.Background(), cfg, logger, vmid))
case "pbs-verify":
os.Exit(runSelftestPBSVerify(context.Background(), cfg, logger))
case "lanresolver":
@@ -785,7 +789,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
pbsStore := pbs.NewSnapshotStore()
pbsTargets := pbsTargetsFromPVE(cfg, px, logger)
pbsReporter := pbs.NewLiveSnapshotReporter(pbsTargets, pbsStore, pbs.DefaultLiveSnapshotTimeout, logger)
collector := hub.NewCollector(px, hub.SystemctlProber{}, observer, backupStore, backupStore, pbsReporter, cfg.Hub.HostID, version, logger)
collector := hub.NewCollector(px, newTunnelProber(cfg, px), observer, backupStore, backupStore, pbsReporter, cfg.Hub.HostID, version, logger)
collector.SetBackupTargetResolver(primaryBackupTargetOf(cfg)) // R-109: the recipe names the live target
// Privileged-capability self-check (v0.44.0): probe the sudoers grants the non-root agent
// depends on. The probe runs `sudo -n -l` LITERALLY (a policy LIST, never executing the
@@ -844,7 +848,8 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
desiredSyncer := desired.NewSyncer(client, desiredProvider, logger)
// OS updates, guest fast lane (agent v0.140.0, `11-os-updates.md` §8 step 2): the leg consumes the hub's
// os_update block and runs after each successful primary whole-guest backup (wired on the local API below).
osLeg := newOSLeg(cfg, client, logger)
osLeg := newOSLeg(cfg, client, px, logger)
collector.SetSystemReporter(&factsReporter{leg: osLeg, guest: firstGuest(px)}) // R-852: the versions
desiredSyncer.AddConsumer(osLeg)
// S5: consume a host_loss restore_directive into an inspectable restore PLAN (derive + surface,
// execute nothing). The recipe is fetched on-demand (rare directive) via a fresh Collect.
@@ -1087,7 +1092,16 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// recurring clobber before it locks the box out. Port 22 for G1 (H1 passes the felhom-sshd port).
collector.SetMgmtPlaneReporter(mgmtplane.NewReporter(mgmtplane.DefaultPrivsepDir, mgmtplane.DefaultHealMarker, mgmtplane.DefaultSshdPort))
jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec}, cfg.Hub.HostID, logger)
// Agent v0.142.0: a signed Docker engine step (`11` §5.8) — ring 1 and every undo; under the heavy-op gate.
dockerExec := osupdate.DockerStepExecutor{Leg: osLeg, Guest: firstGuest(px),
Gate: func(ctx context.Context) (func(), error) {
release, busy, ok := heavyOps.TryAcquire("os-docker-step")
if !ok {
return nil, fmt.Errorf("busy: %s", busy)
}
return release, nil
}}
jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec, dockerExec}, cfg.Hub.HostID, logger)
loop.SetEnvelopeObserver(hub.MultiObserver(desiredSyncer, jobsRunner))
// Controller-driven escrow ceremony (v0.88.0): static config facts + the LATE-BOUND DR gate —
@@ -1127,7 +1141,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
return
case <-time.After(90 * time.Second):
}
osLeg.Run(ctx, vmid, "night")
_ = osLeg.Run(ctx, vmid, "night")
})
}
if localTokens != nil {
@@ -1869,7 +1883,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
ConfigPath: cfg.SourcePath,
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
SmbCredsDir: cfg.Privileged.SmbCredsDir,
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
@@ -2026,7 +2040,7 @@ func runSelftestHub(ctx context.Context, cfg config.Config, logger *slog.Logger)
// pbs coord. The live reporter lists snapshots directly (fresh store, last-known-good fallback) so
// the selftest reflects exactly what a freshly-restarted daemon's first collect emits.
pbsReporter := pbs.NewLiveSnapshotReporter(pbsTargetsFromPVE(cfg, px, logger), pbs.NewSnapshotStore(), pbs.DefaultLiveSnapshotTimeout, logger)
collector := hub.NewCollector(px, hub.SystemctlProber{}, observer, nil, nil, pbsReporter, cfg.Hub.HostID, version, logger)
collector := hub.NewCollector(px, newTunnelProber(cfg, px), observer, nil, nil, pbsReporter, cfg.Hub.HostID, version, logger)
// R-109: wire the backup-target resolver here TOO. Without it selftest=hub would print a recipe whose
// backup_target reads unknown/agent_backup_config_unavailable while the daemon's is resolved — and
// this one-shot exists precisely so "the report it would send" can be trusted to match.
@@ -3498,17 +3512,19 @@ func (f *selftestFlag) Set(v string) error {
f.mode = "controller-swap"
case "os-update":
f.mode = "os-update"
case "os-facts", "live-restore": // agent v0.142.0
f.mode = v
case "wgtunnel": // dispatched since S3 but refused here until 2026-10-04 (TestSelftestFlag_AcceptsEveryDispatchedMode)
f.mode = "wgtunnel"
default:
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|restore-test-due|pbs-verify|lanresolver|wgtunnel|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap|os-update)", v)
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|restore-test-due|pbs-verify|lanresolver|wgtunnel|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap|os-update|os-facts|live-restore)", v)
}
return nil
}
// newOSLeg builds the OS-update leg (agent v0.140.0). The wrapper runs through sudo (FELHOM_OSAPPLY); the plan and
// the once-per-night marker live in the agent's own os/ dir.
func newOSLeg(cfg config.Config, client *hub.Client, logger *slog.Logger) *osupdate.Leg {
func newOSLeg(cfg config.Config, client *hub.Client, px *proxmox.Client, logger *slog.Logger) *osupdate.Leg {
mode := proxmox.RunnerMode(cfg.Privileged.Mode)
if mode == "" {
mode = proxmox.RunnerSudo
@@ -3518,6 +3534,9 @@ func newOSLeg(cfg config.Config, client *hub.Client, logger *slog.Logger) *osupd
Logger: logger,
PlanDir: osupdate.DefaultPlanDir,
StatePath: filepath.Join(osupdate.DefaultPlanDir, "last-night-run"),
// The host step (agent v0.141.0) runs only on an appliance install; the wrapper re-checks the ROOT-owned record.
Appliance: cfg.IsAppliance(),
Tunnel: newTunnelProber(cfg, px),
}
if client != nil {
l.Hub = client
@@ -3539,7 +3558,12 @@ func runSelftestOSUpdate(ctx context.Context, cfg config.Config, logger *slog.Lo
fmt.Fprintln(os.Stderr, "selftest=os-update: hub client:", err)
return 1
}
leg := newOSLeg(cfg, client, logger)
px, perr := newProxmoxClient(cfg)
if perr != nil {
fmt.Fprintln(os.Stderr, "selftest=os-update: proxmox client:", perr)
return 1
}
leg := newOSLeg(cfg, client, px, logger)
resp, err := client.FetchDesiredState(ctx)
if err != nil {
fmt.Fprintln(os.Stderr, "selftest=os-update: desired state:", err)
@@ -3547,15 +3571,158 @@ func runSelftestOSUpdate(ctx context.Context, cfg config.Config, logger *slog.Lo
}
leg.SetBlock(resp.DesiredState.OSUpdate)
b := leg.Block()
fmt.Printf("=== felhom-agent %s selftest=os-update vmid=%d ring=%d enabled=%v release=%v ===\n", version, vmid, b.Ring, b.Enabled, b.Release != nil)
rep := leg.Run(ctx, vmid, "debug")
printJSON("os-update report", map[string]any{"run_id": rep.RunID, "ring": rep.Ring, "release_id": rep.ReleaseID,
"mode": rep.Mode, "outcome": rep.Outcome, "healthy": rep.Healthy, "health_reason": rep.HealthReason,
"upgraded": rep.Upgraded, "pending": len(rep.Pending), "not_covered": rep.NotCovered,
"restart_needed": rep.RestartNeeded, "docker_restart_needed": rep.DockerRestartNeeded, "refused": rep.Refused})
switch rep.Outcome {
fmt.Printf("=== felhom-agent %s selftest=os-update vmid=%d ring=%d enabled=%v guest-release=%v host-release=%v appliance=%v ===\n",
version, vmid, b.Ring, b.Enabled, b.Release != nil, b.HostRelease != nil, leg.Appliance)
start := time.Now()
pass := leg.Run(ctx, vmid, "debug")
worst := pass.Guest
for i, rep := range []osupdate.Report{pass.Guest, pass.Host, pass.Docker} {
if rep.Layer == "" {
fmt.Printf(" %s step: skipped (see the log line above)\n", []string{"guest", "host", "docker"}[i])
continue
}
printJSON("os-update report ("+rep.Layer+")", map[string]any{"run_id": rep.RunID, "ring": rep.Ring, "release_id": rep.ReleaseID,
"mode": rep.Mode, "outcome": rep.Outcome, "healthy": rep.Healthy, "health_reason": rep.HealthReason,
"upgraded": rep.Upgraded, "pending": len(rep.Pending), "not_covered": rep.NotCovered,
"restart_needed": rep.RestartNeeded, "reboot_needed": rep.RebootNeeded, "wrapper_seconds": rep.PassSeconds,
"refused": rep.Refused, "docker_engine": rep.DockerEngine, "authority": rep.Authority})
if !(rep.Outcome == "applied" || rep.Outcome == "nothing" || rep.Outcome == "inventory" || rep.Outcome == "skipped") {
worst = rep
}
}
fmt.Printf(" pass took %s\n", time.Since(start).Round(100*time.Millisecond))
switch worst.Outcome {
case "applied", "nothing", "inventory", "skipped":
return 0
}
return 1
}
// runSelftestFacts prints the versions the host report carries (R-852, agent v0.142.0) — read-only.
//
// sudo -u felhom-agent felhom-agent --config … --selftest=os-facts -vmid 9201
func runSelftestFacts(ctx context.Context, cfg config.Config, logger *slog.Logger, vmid int) int {
if vmid <= 0 {
fmt.Fprintln(os.Stderr, "selftest=os-facts: -vmid is required")
return 2
}
px, _ := newProxmoxClient(cfg)
leg := newOSLeg(cfg, nil, px, logger)
start := time.Now()
f, err := leg.Facts(ctx, vmid)
if err != nil {
fmt.Fprintln(os.Stderr, "selftest=os-facts:", err)
return 1
}
var v any
_ = json.Unmarshal(f, &v)
printJSON(fmt.Sprintf("facts (vmid %d, %s)", vmid, time.Since(start).Round(100*time.Millisecond)), v)
return 0
}
// runSelftestLiveRestore is the ONE-TIME live-restore step (`09` decision 87) as a debug action — the night leg does
// the same before a ring-0 Docker step. It prints the container ids before and after (they must not change).
//
// sudo -u felhom-agent felhom-agent --config … --selftest=live-restore -vmid 9202
func runSelftestLiveRestore(ctx context.Context, cfg config.Config, logger *slog.Logger, vmid int) int {
if vmid <= 0 {
fmt.Fprintln(os.Stderr, "selftest=live-restore: -vmid is required")
return 2
}
px, _ := newProxmoxClient(cfg)
leg := newOSLeg(cfg, nil, px, logger)
if err := leg.EnsureLiveRestore(ctx, time.Now().UTC().Format("20060102T150405Z"), vmid); err != nil {
fmt.Fprintln(os.Stderr, "selftest=live-restore:", err)
return 1
}
fmt.Println("live-restore: on (see the wrapper's LIVE-RESTORE line above for the container ids)")
return 0
}
// newTunnelProber reads the box's REAL tunnel (R-841, agent v0.141.0): the cloudflared container in each running
// customer guest — a guest that binds /mnt/felhom-drives, the same rule the OS wrapper's R10 uses — through the
// existing `pct exec [0-9]* -- docker inspect -f *` sudoers line.
func newTunnelProber(cfg config.Config, px *proxmox.Client) hub.CloudflaredProber {
mode := proxmox.RunnerMode(cfg.Privileged.Mode)
if mode == "" {
mode = proxmox.RunnerSudo
}
return hub.GuestTunnelProber{
Runner: &proxmox.ExecRunner{Mode: mode, SudoPath: cfg.Privileged.SudoPath},
Guests: customerGuests(px),
}
}
// customerGuests lists the running guests that bind /mnt/felhom-drives — the box's customer guest(s), the same rule the
// OS wrapper's R10 uses. Shared by the tunnel probe, the facts read and the signed Docker step.
func customerGuests(px *proxmox.Client) func(ctx context.Context) ([]int, error) {
return func(ctx context.Context) ([]int, error) {
if px == nil {
return nil, fmt.Errorf("no proxmox client")
}
gs, err := px.ListLXC(ctx)
if err != nil {
return nil, err
}
var out []int
for _, g := range gs {
if g.Status != "running" {
continue
}
gc, err := px.GuestConfig(ctx, g.VMID)
if err != nil {
continue
}
for _, v := range gc.MountPoints() {
if src, _, _ := strings.Cut(v, ","); src == "/mnt/felhom-drives" {
out = append(out, g.VMID)
break
}
}
}
return out, nil
}
}
// firstGuest is the single customer guest (an error when there is none).
func firstGuest(px *proxmox.Client) func(ctx context.Context) (int, error) {
f := customerGuests(px)
return func(ctx context.Context) (int, error) {
v, err := f(ctx)
if err != nil {
return 0, err
}
if len(v) == 0 {
return 0, fmt.Errorf("no running customer guest")
}
return v[0], nil
}
}
// factsReporter feeds the host report's `system` stanza (R-852, agent v0.142.0) from the wrapper's read-only facts
// mode, at most every 10 minutes (each read is ~2 s of pct exec; the host reports every 15 min).
type factsReporter struct {
leg *osupdate.Leg
guest func(ctx context.Context) (int, error)
mu sync.Mutex
at time.Time
vmid int
facts json.RawMessage
err error
}
func (f *factsReporter) SystemFacts(ctx context.Context) (int, json.RawMessage, error) {
f.mu.Lock()
defer f.mu.Unlock()
if !f.at.IsZero() && time.Since(f.at) < 10*time.Minute {
return f.vmid, f.facts, f.err
}
f.at = time.Now()
f.vmid, f.err = f.guest(ctx)
if f.err != nil {
f.facts = nil
return 0, nil, f.err
}
f.facts, f.err = f.leg.Facts(ctx, f.vmid)
return f.vmid, f.facts, f.err
}
+8
View File
@@ -0,0 +1,8 @@
# /etc/felhom/crash-guard.conf — read by /usr/local/sbin/felhom-crash-guard (`11` §5.9).
# The LIMIT-th unclean stop within WINDOW_MINUTES leaves the box off. Decided by CC unattended — operator may reverse.
LIMIT=3
WINDOW_MINUTES=60
# kernel.panic while armed: seconds after a crash before the kernel restarts the box.
PANIC_SECONDS=10
# A tripped guard re-arms after this many hours of normal running (or `felhom-crash-guard rearm`).
REARM_HOURS=24
+235
View File
@@ -0,0 +1,235 @@
#!/usr/bin/python3
# felhom-crash-guard — a crashed host restarts by itself, but not forever (`09` decision 88, R-851, `11` §5.9).
#
# Install as /usr/local/sbin/felhom-crash-guard (0755 root:root), with felhom-crash-guard.service (boot / clean-stop)
# and felhom-crash-guard-check.timer (hourly re-arm check). Python 3, standard library only.
# Tests: configs/test_felhom_crash_guard.py (temp dirs; nothing real is touched).
#
# WHAT IT DOES
# boot early at every boot. Was the previous boot ended CLEANLY? (the clean-stop marker exists). If not, this
# boot follows an UNCLEAN stop — a kernel crash, a power cut or a hard reset (they cannot be told apart
# on these boxes: measured 2026-10-04 on demo-hp, efi_pstore is on yet saved NOTHING for a real panic;
# the journal and `last` show only "no shutdown"). It records the unclean boot, counts those in the last
# WINDOW_MINUTES, and sets kernel.panic:
# - fewer than LIMIT-1 recent unclean boots → kernel.panic = PANIC_SECONDS (a crash restarts the box);
# - LIMIT-1 or more → the guard TRIPS: kernel.panic = 0, so the LIMIT-th crash within the window
# leaves the box OFF (operator's own words: "if it crashes 3 times within one hour, it stays off").
# A tripped guard stays tripped across further boots until it re-arms.
# clean-stop ExecStop of the service: writes the clean-stop marker during an orderly shutdown or reboot.
# check hourly: a tripped guard re-arms after REARM_HOURS of normal running (since the trip AND since boot).
# rearm the operator re-arms by hand (`felhom-crash-guard rearm`).
# status prints the state.
# The state is /var/lib/felhom-crash-guard/state.json (0644: the non-root agent reads it into its host report).
# Before the service runs (very early boot) the kernel default kernel.panic = 0 applies, so a crash THAT early leaves
# the box off — the safe side: a box that cannot reach userspace must not loop.
import json
import os
import sys
import time
CONF = "/etc/felhom/crash-guard.conf"
STATE_DIR = "/var/lib/felhom-crash-guard"
DEFAULTS = {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10, "REARM_HOURS": 24}
class Env:
"""Paths and clock; tests replace them."""
def __init__(self, conf=CONF, state_dir=STATE_DIR, panic_path="/proc/sys/kernel/panic",
uptime_path="/proc/uptime", boot_id_path="/proc/sys/kernel/random/boot_id"):
self.conf, self.state_dir = conf, state_dir
self.panic_path, self.uptime_path, self.boot_id_path = panic_path, uptime_path, boot_id_path
def now(self):
return time.time()
def log(self, line):
print(line, file=sys.stderr, flush=True)
try:
import subprocess
subprocess.run(["logger", "-t", "felhom-crash-guard", line], timeout=10)
except Exception:
pass
def iso(t):
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(t))
def parse_iso(s):
import calendar
return calendar.timegm(time.strptime(s, "%Y-%m-%dT%H:%M:%SZ"))
def load_conf(env):
c = dict(DEFAULTS)
try:
for line in open(env.conf):
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
k, v = (x.strip() for x in line.split("=", 1))
if k in c and v.isdigit() and int(v) >= (1 if k != "PANIC_SECONDS" else 1):
c[k] = int(v)
except OSError:
pass
return c
def state_path(env):
return os.path.join(env.state_dir, "state.json")
def marker_path(env):
return os.path.join(env.state_dir, "clean-stop")
def load_state(env):
try:
with open(state_path(env)) as f:
s = json.load(f)
return s if isinstance(s, dict) else None
except (OSError, ValueError):
return None
def save_state(env, s):
os.makedirs(env.state_dir, mode=0o755, exist_ok=True)
tmp = state_path(env) + ".tmp"
with open(tmp, "w") as f:
json.dump(s, f, indent=2, sort_keys=True)
f.write("\n")
os.chmod(tmp, 0o644)
os.replace(tmp, state_path(env))
def set_panic(env, seconds):
with open(env.panic_path, "w") as f:
f.write(f"{seconds}\n")
def read(path, default=""):
try:
with open(path) as f:
return f.read().strip()
except OSError:
return default
def summarize(s, c, now):
window = c["WINDOW_MINUTES"] * 60
times = [parse_iso(t) for t in s.get("unclean_boots", [])]
after = parse_iso(s["rearmed_at"]) if s.get("rearmed_at") else 0
# a re-arm starts a fresh window (or the next unclean boot would trip again at once); the history stays
s["unclean_boots_in_window"] = sum(1 for t in times if now - t <= window and t > after)
s["unclean_boots_24h"] = sum(1 for t in times if now - t <= 86400)
s["config"] = c
s["updated_at"] = iso(now)
def boot(env):
c = load_conf(env)
now = env.now()
try:
up = float(read(env.uptime_path, "0").split()[0])
except (ValueError, IndexError):
up = 0.0
boot_at = now - up
prev = load_state(env)
first = prev is None
s = prev or {"version": 1, "unclean_boots": [], "tripped": False}
clean = os.path.exists(marker_path(env))
unclean = (not first) and (not clean)
try:
os.remove(marker_path(env))
except OSError:
pass
# keep 7 days of history (the 24 h figure and the operator's view), drop older
s["unclean_boots"] = [t for t in s.get("unclean_boots", []) if now - parse_iso(t) <= 7 * 86400]
if unclean:
s["unclean_boots"].append(iso(boot_at))
s["last_boot_at"] = iso(boot_at)
s["last_boot_unclean"] = unclean
s["boot_id"] = read(env.boot_id_path, "unknown")
summarize(s, c, now)
if not s.get("tripped") and s["unclean_boots_in_window"] >= c["LIMIT"] - 1:
s["tripped"], s["tripped_at"] = True, iso(now)
s["tripped_reason"] = (f"{s['unclean_boots_in_window']} unclean boots within {c['WINDOW_MINUTES']} minutes — "
f"the next crash leaves the box off (limit {c['LIMIT']})")
env.log(f"crash-guard: TRIPPED: {s['tripped_reason']}")
panic = 0 if s.get("tripped") else c["PANIC_SECONDS"]
set_panic(env, panic)
s["kernel_panic"] = panic
s["armed"] = not s.get("tripped")
save_state(env, s)
env.log(f"crash-guard: boot first={first} unclean={unclean} in-window={s['unclean_boots_in_window']} "
f"tripped={s.get('tripped')} kernel.panic={panic}")
return 0
def clean_stop(env):
os.makedirs(env.state_dir, mode=0o755, exist_ok=True)
with open(marker_path(env), "w") as f:
f.write(iso(env.now()) + "\n")
env.log("crash-guard: clean stop recorded")
return 0
def rearm(env, by):
c = load_conf(env)
now = env.now()
s = load_state(env) or {"version": 1, "unclean_boots": []}
was = bool(s.get("tripped"))
s["tripped"] = False
s["armed"] = True
s["rearmed_at"], s["rearmed_by"] = iso(now), by
if was:
s["last_trip"] = {"at": s.get("tripped_at"), "reason": s.get("tripped_reason")}
s.pop("tripped_at", None)
s.pop("tripped_reason", None)
summarize(s, c, now)
set_panic(env, c["PANIC_SECONDS"])
s["kernel_panic"] = c["PANIC_SECONDS"]
save_state(env, s)
env.log(f"crash-guard: RE-ARMED by {by} (was tripped: {was}); kernel.panic={c['PANIC_SECONDS']}")
return 0
def check(env):
c = load_conf(env)
now = env.now()
s = load_state(env)
if not s:
return 0
if s.get("tripped"):
since = max(parse_iso(s["tripped_at"]), parse_iso(s.get("last_boot_at", s["tripped_at"])))
if now - since >= c["REARM_HOURS"] * 3600:
return rearm(env, f"timer ({c['REARM_HOURS']} h of normal running)")
summarize(s, c, now)
save_state(env, s)
return 0
def main(argv, env=None):
env = env or Env()
cmd = argv[1] if len(argv) == 2 else ""
if cmd == "boot":
return boot(env)
if cmd == "clean-stop":
return clean_stop(env)
if cmd == "check":
return check(env)
if cmd == "rearm":
return rearm(env, "operator")
if cmd == "status":
print(json.dumps(load_state(env), indent=2, sort_keys=True))
return 0
print("usage: felhom-crash-guard boot|clean-stop|check|rearm|status", file=sys.stderr)
return 2
if __name__ == "__main__":
if os.geteuid() != 0:
print("felhom-crash-guard: must run as root", file=sys.stderr)
sys.exit(2)
sys.exit(main(sys.argv))
+7
View File
@@ -0,0 +1,7 @@
# Hourly: a tripped crash guard re-arms after REARM_HOURS of normal running (felhom-crash-guard check).
[Unit]
Description=Felhom crash guard re-arm check
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/felhom-crash-guard check
+9
View File
@@ -0,0 +1,9 @@
[Unit]
Description=Felhom crash guard re-arm check (hourly)
[Timer]
OnBootSec=15min
OnUnitActiveSec=1h
[Install]
WantedBy=timers.target
+19
View File
@@ -0,0 +1,19 @@
# felhom-crash-guard — a crashed host restarts by itself, with a limit (`09` decision 88, R-851, `11` §5.9).
# Starts early at boot (sets kernel.panic for THIS boot); its ExecStop writes the clean-stop marker during an orderly
# shutdown or reboot. A boot that finds no marker followed a crash, a power cut or a hard reset.
[Unit]
Description=Felhom crash guard (restart after a kernel crash, with a limit)
DefaultDependencies=no
After=local-fs.target
Before=sysinit.target shutdown.target
Conflicts=shutdown.target
RequiresMountsFor=/var/lib
[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=/usr/local/sbin/felhom-crash-guard boot
ExecStop=/usr/local/sbin/felhom-crash-guard clean-stop
[Install]
WantedBy=sysinit.target
+520 -101
View File
@@ -1,5 +1,5 @@
#!/usr/bin/python3
# felhom-os-apply — the ROOT half of the agent's operating-system update leg (`11-os-updates.md` §5.4.1).
# felhom-os-apply — the ROOT half of the agent's operating-system update leg (`11-os-updates.md` §5.4.1, §8.1–8.2).
#
# Install as /usr/local/sbin/felhom-os-apply (0755 root:root). The non-root agent invokes it via `sudo -n`
# (FELHOM_OSAPPLY alias) with EXACTLY: felhom-os-apply --plan /var/lib/felhom-agent/os/plan-<id>.json
@@ -8,22 +8,39 @@
#
# THE TRUST MODEL. The plan is written by the agent, so a broken-into agent writes whatever plan it likes. The
# protection is therefore what this file REFUSES, not where the plan came from: no removal, no downgrade, no new
# package, no package outside the plan, only Debian origin in the fast lane, only the box's own customer guest.
# Package signatures stay Debian's: apt checks every Release file against the guest's keyring, including the
# snapshot.debian.org fallback (decision 79). Nothing here is overridable from the environment.
# package, no package outside the plan, only Debian origin in the fast lane, no kernel / boot package on the host,
# only the box's own customer guest, and the host layer only on a box whose ROOT-OWNED install record says
# "appliance" (a BYO host belongs to its owner, `11` §1). Package signatures stay Debian's: apt checks every Release
# file, including the snapshot.debian.org fallback (decision 79). Nothing here is overridable from the environment.
#
# THIS RELEASE: layer "guest", lane "fast" only. The host layer and the slow lane exist in the interface and are
# REFUSED (R3, R12) until `11` §8 steps 3 and 5 enable them.
# LAYERS (agent v0.141.0): "guest" (the customer LXC, entered with `pct exec`) and "host" (this Proxmox host, run
# directly). LANE: "fast" for those two. Agent v0.142.0 adds the layer "docker" (the guest's Docker engine set, `11`
# §5.8), which is the SLOW lane: lane "slow" only, the six Docker packages only, origin "Docker CE" only, and only
# with an authority this file checks ITSELF (R3): a signed operator job verified with `ssh-keygen -Y verify` against
# the ROOT-OWNED signers file (TRUST_SIGNERS), bound to this host (TRUST_FILE host_id), unexpired and never replayed;
# or, for an unsigned ring-0 step, the root-owned TRUST_FILE saying `"ring0_slow_lane": true` (set by hand on the demo
# boxes only). The agent's own config is NOT trusted for either: the agent can write it. A Docker step also needs
# `live-restore` ON (R15) — without it every container restarts.
#
# Modes (plan field "mode"):
# inventory read-only for packages: `apt-get update` in the guest, then report what is installed (with origin),
# what is pending, restart-needed and health. Installs nothing.
# apply repair first, check every refusal on an `apt-get -s` simulation of EXACTLY name=version, then
# install, clean, and report the same as inventory.
# inventory `apt-get update`, then report what is installed (with origin), what is pending, and health.
# apply repair first, pick the packages (select "listed": the plan's name=version list; "pending-fast": every
# pending Debian / Debian-Security upgrade, for ring 0), check every refusal on an `apt-get -s`
# simulation of EXACTLY name=version, install, clean, scan for restart-needed, report as inventory.
# health report health only (the agent polls it after a run).
# facts (v0.142.0) read-only versions for the hub's System page: host Debian, running and next-boot kernel,
# held packages, kernel taint, the crash guard; guest Debian, Docker engine, containerd, live-restore.
# live-restore-on (v0.142.0, layer guest) the ONE-TIME step of `09` decision 87: merge `"live-restore": true`
# into the guest's /etc/docker/daemon.json and `systemctl reload docker`. NEVER a restart (R-835).
# Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is
# OSAPPLY-REPORT <one JSON object>
# which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install.
#
# SPEED (R-845, agent v0.141.0). Every `pct exec` costs ~0.9 s (measured on demo-hp), and v0.140.0 made one per
# package for version comparisons — 272 packages ≈ 4 minutes. Versions are now compared with the HOST's dpkg (the same
# Debian algorithm), madison/policy run once per pass for all packages, the restart scan runs only after an install,
# and one wrapper call does the whole pass (no separate inventory call before an apply).
import calendar
import json
import os
import re
@@ -46,6 +63,30 @@ SNAPSHOT_LIST = "/etc/apt/sources.list.d/felhom-os-snapshot.list"
APT_ENV = ["env", "DEBIAN_FRONTEND=noninteractive", "APT_LISTCHANGES_FRONTEND=none", "NEEDRESTART_MODE=l", "LC_ALL=C"]
DPKG_OPTS = ["-o", "Dpkg::Options::=--force-confold", "-o", "Dpkg::Options::=--force-confdef"]
MIN_FREE = 500 * 1024 * 1024
# The installer's ROOT-OWNED record (felhom-host-install.sh `state_set mode`); the agent cannot write it.
INSTALL_STATE = "/var/lib/felhom-install/state.json"
# Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host
# reboot is needed for them to take effect, and a bad one can stop the box from booting.
HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|"
r"pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr)")
# restart_needed() leaves out processes whose cgroup line matches (grep basic regex). Host: the LXC guests' own
# processes (`0::/lxc/<vmid>/...`) -- NOT lxc-start itself, whose cgroup is `0::/lxc.monitor/<vmid>` (measured
# 2026-10-04 on demo-felhom: the old pattern "lxc" hid lxc-start with 20 deleted maps, so "reboot needed" stayed false
# after a libc6 update). Pinned by test_restart_skip_patterns_against_real_cgroups.
RESTART_SKIP_CGROUP = {"guest": "docker", "host": ":/lxc/"}
HOST_SERVICES = ["pveproxy", "pvedaemon", "pvestatd", "pve-cluster", "felhom-agent"]
# The Docker engine set (`11` §5.8): the only names the docker layer may touch, from the only origin it may use.
DOCKER_NAMES = ("containerd.io", "docker-buildx-plugin", "docker-ce", "docker-ce-cli", "docker-ce-rootless-extras",
"docker-compose-plugin")
DOCKER_ORIGIN = "Docker CE"
# ROOT-OWNED trust anchors (the installer writes them; the demo boxes got them by hand, R-840). Never the agent's config.
TRUST_FILE = "/etc/felhom/os-trust.json" # {"host_id": "...", "ring0_slow_lane": false}
TRUST_SIGNERS = "/etc/felhom/operator-signers" # ssh allowed_signers: <key_id> namespaces="felhom-op-v1" <key>
SIG_NAMESPACE = "felhom-op-v1"
SIGNED_OP = "os_docker_step"
NONCE_FILE = "/var/lib/felhom-os-apply/nonces.json"
DAEMON_JSON = "/etc/docker/daemon.json"
CRASH_GUARD_STATE = "/var/lib/felhom-crash-guard/state.json"
class Refused(Exception):
@@ -57,8 +98,8 @@ class Refused(Exception):
class Runner:
"""Runs commands for real. Tests replace it with a fake. `guest` runs inside the container via pct exec."""
def host(self, argv, timeout=600):
p = subprocess.run(argv, capture_output=True, text=True, timeout=timeout)
def host(self, argv, timeout=600, stdin=None):
p = subprocess.run(argv, capture_output=True, text=True, timeout=timeout, input=stdin)
return p.returncode, p.stdout, p.stderr
def guest(self, vmid, argv, timeout=1800):
@@ -75,6 +116,48 @@ class Runner:
import pwd
return pwd.getpwnam(AGENT_USER).pw_uid
def write_file(self, layer, vmid, path, body):
"""Write a small text file in the target layer — never via a shell string."""
if layer == "host":
with open(path, "w") as f:
f.write(body)
return
rc, _, _ = self.host(["/usr/sbin/pct", "exec", str(vmid), "--", "tee", path], 60, stdin=body)
if rc != 0:
raise Refused("R7", f"could not write {path} in the guest")
def now(self):
return time.time()
def sleep(self, s):
time.sleep(s)
def verify_sig(self, signers, key_id, namespace, blob, sig):
"""`ssh-keygen -Y verify` over the EXACT signed bytes. Files in a root-only temp dir; nothing via a shell."""
import tempfile
with tempfile.TemporaryDirectory(prefix="felhom-os-apply-") as d:
sp = os.path.join(d, "sig")
with open(sp, "w") as f:
f.write(sig)
p = subprocess.run(["ssh-keygen", "-Y", "verify", "-f", signers, "-I", key_id, "-n", namespace, "-s", sp],
input=blob, capture_output=True, timeout=30)
return p.returncode
def read_nonces(self):
try:
with open(NONCE_FILE) as f:
d = json.load(f)
return d if isinstance(d, dict) else {}
except (OSError, ValueError):
return {}
def write_nonces(self, d):
os.makedirs(os.path.dirname(NONCE_FILE), mode=0o700, exist_ok=True)
tmp = NONCE_FILE + ".tmp"
with open(tmp, "w") as f:
json.dump(d, f)
os.replace(tmp, NONCE_FILE)
def log(self, line):
print(line, file=sys.stderr, flush=True)
try:
@@ -115,12 +198,22 @@ class Apply:
def check_plan(self, plan):
mode = plan.get("mode", "apply")
if mode not in ("apply", "inventory", "health"):
if mode not in ("apply", "inventory", "health", "facts", "live-restore-on"):
raise Refused("R11", f"unknown mode {mode!r}")
if plan.get("layer") != "guest":
raise Refused("R12", f"layer {plan.get('layer')!r} is refused in this release (guest only)")
if plan.get("lane", "fast") != "fast":
raise Refused("R3", "the slow lane is refused in this release")
layer = plan.get("layer")
if layer not in ("guest", "host", "docker"):
raise Refused("R12", f"layer {layer!r} is not guest, host or docker")
lane = plan.get("lane", "fast")
if layer == "docker" and lane != "slow":
raise Refused("R3", "the Docker engine is the slow lane (`11` §5.8); a fast-lane Docker plan is refused")
if layer != "docker" and lane != "fast":
raise Refused("R3", f"the {layer} layer has no slow lane in this release (kernel, Proxmox: `11` §8 step 6)")
if mode == "facts" and layer != "host":
raise Refused("R11", "facts is a host-layer mode (it reads the host and the guest)")
if mode == "live-restore-on" and layer != "guest":
raise Refused("R11", "live-restore-on is a guest-layer mode")
if plan.get("undo") and layer != "docker":
raise Refused("R5", "an undo (downgrade) exists only for the Docker layer, inside a signed job")
vmid = plan.get("vmid")
if not isinstance(vmid, int) or isinstance(vmid, bool) or vmid <= 0:
raise Refused("R11", f"vmid must be a positive integer, got {vmid!r}")
@@ -129,9 +222,19 @@ class Apply:
raise Refused("R11", f"release_id {rid!r} is not a plain id")
if plan.get("allow_new"):
raise Refused("R6", "allow_new is a slow-lane field; the fast lane never adds a package")
select = plan.get("select", "listed")
if select not in ("listed", "pending-fast", "pending-docker"):
raise Refused("R11", f"unknown select {select!r}")
if (select == "pending-docker") != (layer == "docker" and select != "listed"):
if select == "pending-docker" or layer == "docker":
raise Refused("R11", f"select {select!r} does not fit layer {layer!r}")
pk = plan.get("packages", [])
if not isinstance(pk, list) or (mode == "apply" and not pk):
raise Refused("R11", "packages must be a non-empty list in apply mode")
if not isinstance(pk, list):
raise Refused("R11", "packages must be a list")
if mode == "apply" and select == "listed" and not pk:
raise Refused("R11", "packages must be a non-empty list in apply mode (select listed)")
if select in ("pending-fast", "pending-docker") and pk:
raise Refused("R11", f"select {select} takes no package list")
seen = set()
for e in pk:
if not isinstance(e, dict):
@@ -144,12 +247,237 @@ class Apply:
if n in seen:
raise Refused("R11", f"package {n} is named twice")
seen.add(n)
if layer == "docker":
if n not in DOCKER_NAMES or o != DOCKER_ORIGIN:
raise Refused("R2", f"{n} ({o!r}) is not one of the six Docker packages from {DOCKER_ORIGIN!r}")
continue
if n in DOCKER_NAMES:
raise Refused("R2", f"{n} is a Docker package — the slow lane (`11` §5.8), never in a {layer} plan")
if o not in FAST_ORIGINS:
raise Refused("R2", f"{n}: origin {o!r} is not Debian / Debian-Security (the fast lane, `11` C3)")
if layer == "host" and HOST_SLOW_RE.match(n):
raise Refused("R14", f"{n} is a kernel / boot / firmware package — the host's slow lane")
snap = plan.get("snapshot", "")
if snap and not SNAP_RE.match(snap):
raise Refused("R11", f"snapshot {snap!r} is not YYYYMMDDTHHMMSSZ")
return mode, vmid
return mode, layer, vmid, select
def check_appliance(self):
"""R12: the host layer only on a box whose ROOT-OWNED install record says appliance (`11` §1: never BYO)."""
try:
st = self.r.stat(INSTALL_STATE)
except OSError:
raise Refused("R12", f"no install record ({INSTALL_STATE}) — this box cannot prove it is an appliance")
if st.st_uid != 0 or (st.st_mode & 0o022):
raise Refused("R12", f"{INSTALL_STATE} is not root-owned and root-only-writable — it proves nothing")
try:
mode = json.loads(self.r.read_file(INSTALL_STATE)).get("mode")
except (OSError, ValueError, AttributeError):
raise Refused("R12", f"{INSTALL_STATE} is unreadable — this box cannot prove it is an appliance")
if mode != "appliance":
raise Refused("R12", f"this box was installed as {mode!r}, not appliance — its host belongs to its owner")
def load_trust(self):
"""The ROOT-OWNED trust record. Absent or agent-writable → no slow-lane authority at all (R3)."""
try:
st = self.r.stat(TRUST_FILE)
except OSError:
raise Refused("R3", f"no {TRUST_FILE} — this box has no slow-lane trust anchor")
if st.st_uid != 0 or (st.st_mode & 0o022):
raise Refused("R3", f"{TRUST_FILE} is not root-owned and root-only-writable — it proves nothing")
try:
t = json.loads(self.r.read_file(TRUST_FILE))
except (OSError, ValueError):
raise Refused("R3", f"{TRUST_FILE} is unreadable")
if not isinstance(t, dict) or not isinstance(t.get("host_id"), str) or not t["host_id"]:
raise Refused("R3", f"{TRUST_FILE} names no host_id")
return t
def verify_signed(self, signed, trust):
"""R3: an operator-signed os_docker_step, checked HERE (not by the agent): signature against the root-owned
signers file, op, host binding, time window, and a root-owned nonce record (no replay)."""
import base64
if not isinstance(signed, dict) or not isinstance(signed.get("blob_b64"), str) or not isinstance(signed.get("sig"), str):
raise Refused("R3", "the signed job is malformed")
try:
blob = base64.b64decode(signed["blob_b64"], validate=True)
op = json.loads(blob)
except (ValueError, TypeError):
raise Refused("R3", "the signed blob is not base64 JSON")
key_id = op.get("key_id", "")
if not isinstance(key_id, str) or not re.match(r"^[A-Za-z0-9._-]{1,64}$", key_id):
raise Refused("R3", "the signed blob names no plain key_id")
try:
st = self.r.stat(TRUST_SIGNERS)
except OSError:
raise Refused("R3", f"no {TRUST_SIGNERS} — no operator key to check a signed job against")
if st.st_uid != 0 or (st.st_mode & 0o022):
raise Refused("R3", f"{TRUST_SIGNERS} is not root-owned and root-only-writable")
rc = self.r.verify_sig(TRUST_SIGNERS, key_id, SIG_NAMESPACE, blob, signed["sig"])
if rc != 0:
raise Refused("R3", f"the operator signature does not verify (ssh-keygen rc={rc})")
if op.get("op") != SIGNED_OP:
raise Refused("R3", f"the signed op is {op.get('op')!r}, not {SIGNED_OP}")
if (op.get("target") or {}).get("host_id") != trust["host_id"]:
raise Refused("R3", "the signed job is for another host")
now = self.r.now()
try:
exp = calendar.timegm(time.strptime(op["expires_at"], "%Y-%m-%dT%H:%M:%SZ"))
iss = calendar.timegm(time.strptime(op["issued_at"], "%Y-%m-%dT%H:%M:%SZ"))
except (KeyError, ValueError, TypeError):
raise Refused("R3", "the signed job has no readable time window")
if now > exp or now < iss - 120:
raise Refused("R3", "the signed job is expired or not yet valid")
nonce = op.get("nonce")
if not isinstance(nonce, str) or not nonce:
raise Refused("R3", "the signed job has no nonce")
seen = self.r.read_nonces()
if nonce in seen:
raise Refused("R3", "the signed job was already used (replay)")
seen[nonce] = exp
self.r.write_nonces({k: v for k, v in seen.items() if v > now})
return op.get("params") or {}
def docker_authority(self, plan):
"""R3 for the docker layer: returns (who, undo). A signed job binds the EXACT package list and the undo flag."""
trust = self.load_trust()
signed = plan.get("signed")
if signed:
params = self.verify_signed(signed, trust)
want = sorted(f"{e.get('name')}={e.get('version')}" for e in params.get("packages") or [])
got = sorted(f"{e['name']}={e['version']}" for e in plan.get("packages", []))
if not want or want != got:
raise Refused("R3", "the plan's packages are not exactly the signed job's packages")
if bool(params.get("undo")) != bool(plan.get("undo")):
raise Refused("R3", "the plan's undo flag is not the signed job's")
if params.get("vmid") not in (None, self.vmid):
raise Refused("R3", "the signed job names another guest")
return "signed", bool(plan.get("undo"))
if plan.get("undo"):
raise Refused("R3", "an undo (downgrade) needs a signed operator job")
if trust.get("ring0_slow_lane") is True:
return "ring0", False
raise Refused("R3", "a Docker step needs a signed operator job (ring 1) or this box's root-owned ring-0 mark")
def live_restore(self):
rc, out, _ = self.g(["docker", "info", "--format", "{{.LiveRestoreEnabled}}"], timeout=60)
return out.strip() if rc == 0 and out.strip() in ("true", "false") else "unknown"
def container_ids(self):
rc, out, _ = self.g(["docker", "ps", "-q", "--no-trunc"], timeout=60)
return sorted(out.split()) if rc == 0 else None
def live_restore_on(self):
"""`09` decision 87: merge live-restore into daemon.json and RELOAD (C5: a reload turns it on, no restart)."""
log = self.r.log
if self.live_restore() == "true":
self.report["live_restore"] = {"result": "already on"}
log("os-apply: LIVE-RESTORE already on")
return 0
rc, cur, _ = self.g(["cat", DAEMON_JSON], timeout=30)
try:
conf = json.loads(cur) if rc == 0 and cur.strip() else {}
except ValueError:
raise Refused("R16", f"{DAEMON_JSON} in the guest is not valid JSON — not touched")
if not isinstance(conf, dict):
raise Refused("R16", f"{DAEMON_JSON} is not a JSON object — not touched")
before = self.container_ids()
conf["live-restore"] = True
self.r.write_file("guest", self.vmid, DAEMON_JSON, json.dumps(conf, indent=2, sort_keys=True) + "\n")
rrc, _, rerr = self.g(["systemctl", "reload", "docker"], timeout=120)
state = "unknown"
for _ in range(10):
state = self.live_restore()
if state == "true":
break
self.r.sleep(1)
after = self.container_ids()
same = before is not None and before == after
self.report["live_restore"] = {"result": "on" if state == "true" else "failed", "reload_rc": rrc,
"containers_before": len(before or []), "same_ids": same}
log(f"os-apply: LIVE-RESTORE reload_rc={rrc} state={state} containers={len(before or [])} same-ids={'yes' if same else 'NO'}")
if state != "true":
# put the old file back (and reload again) — still never a restart
self.r.write_file("guest", self.vmid, DAEMON_JSON, cur if rc == 0 else "{}\n")
self.g(["systemctl", "reload", "docker"], timeout=120)
self.report["failed"] = {"rc": 3, "step": "live-restore", "reason": (rerr or "").strip()[-200:]}
return 3
return 0
def kernel_next_boot(self):
"""Which kernel GRUB boots next, read without root-only files (grubenv + /etc/default/grub + /boot)."""
try:
dflt = re.search(r'^GRUB_DEFAULT=["\']?([^"\'\n]*)', self.r.read_file("/etc/default/grub"), re.M)
dflt = dflt.group(1) if dflt else "0"
except OSError:
dflt = "0"
env = {}
try:
for l in self.r.read_file("/boot/grub/grubenv").splitlines():
if "=" in l and not l.startswith("#"):
k, v = l.split("=", 1)
env[k] = v
except OSError:
pass
def ver(entry):
m = re.search(r"gnulinux-([0-9][^>\s]*?-pve)-(?:advanced|recovery)", entry)
return m.group(1) if m else "unknown"
if env.get("next_entry"):
return ver(env["next_entry"]), "next_entry (a one-shot GRUB cannot clear on LVM /boot)"
if dflt == "saved":
return (ver(env["saved_entry"]), "saved default") if env.get("saved_entry") else ("unknown", "saved default unset")
if dflt == "0":
rc, out, _ = self.r.host(["sh", "-c", "ls /boot/vmlinuz-* 2>/dev/null"], 30)
vers = [l.split("vmlinuz-", 1)[1] for l in out.split() if "vmlinuz-" in l]
best = None
for v in vers:
if best is None or self.dpkg_cmp(v, "gt", best):
best = v
return (best or "unknown"), "GRUB_DEFAULT=0 (the newest installed)"
return "unknown", f"GRUB_DEFAULT={dflt}"
def facts(self):
"""Read-only versions for the hub's System page (R-852). A value that cannot be read is "unknown"."""
def first(cmd, timeout=30):
rc, out, _ = self.r.host(cmd, timeout)
v = out.strip().splitlines()[0].strip() if rc == 0 and out.strip() else ""
return v or "unknown"
h = {"debian": first(["cat", "/etc/debian_version"]), "kernel_running": first(["uname", "-r"])}
h["kernel_next_boot"], h["kernel_next_boot_source"] = self.kernel_next_boot()
rc, out, _ = self.r.host(["apt-mark", "showhold"], 60)
h["held"] = sorted(out.split()) if rc == 0 else None
try:
t = int(self.r.read_file("/proc/sys/kernel/tainted").strip())
h["tainted"], h["oops_this_boot"], h["warn_this_boot"] = t, bool(t & 128), bool(t & 512)
except (OSError, ValueError):
h["tainted"], h["oops_this_boot"], h["warn_this_boot"] = None, None, None
try:
h["kernel_panic"] = int(self.r.read_file("/proc/sys/kernel/panic").strip())
except (OSError, ValueError):
h["kernel_panic"] = None
try:
h["crash_guard"] = json.loads(self.r.read_file(CRASH_GUARD_STATE))
except (OSError, ValueError):
h["crash_guard"] = None
g = {"debian": "unknown", "docker_engine": "unknown", "containerd": "unknown", "live_restore": "unknown"}
try:
self.check_guest(self.vmid)
running = True
except Refused as e:
running, g["unknown_reason"] = False, f"{e.code} {e.reason}"
if running:
script = ('echo "debian=$(cat /etc/debian_version 2>/dev/null)"; '
'echo "engine=$(docker version --format \'{{.Server.Version}}\' 2>/dev/null)"; '
'echo "containerd=$(dpkg-query -W -f \'${Version}\' containerd.io 2>/dev/null)"; '
'echo "live=$(docker info --format \'{{.LiveRestoreEnabled}}\' 2>/dev/null)"')
rc, out, _ = self.g(["sh", "-c", script], timeout=60)
kv = dict(l.split("=", 1) for l in out.splitlines() if "=" in l)
for k, src in (("debian", "debian"), ("docker_engine", "engine"), ("containerd", "containerd")):
g[k] = kv.get(src, "").strip() or "unknown"
g["live_restore"] = {"true": "on", "false": "off"}.get(kv.get("live", "").strip(), "unknown")
self.report["facts"] = {"host": h, "guest": g}
return 0
def check_guest(self, vmid):
if vmid in RESERVED_VMIDS:
@@ -169,12 +497,24 @@ class Apply:
if rc != 0 or "running" not in out:
raise Refused("R10", f"vmid {vmid} is not running")
# ---------- guest helpers ----------
# ---------- target helpers ----------
def x(self, argv, timeout=1800):
"""Run in the TARGET layer: the guest via pct exec, or the host directly."""
if self.layer == "host":
return self.r.host(argv, timeout)
return self.r.guest(self.vmid, argv, timeout) # guest and docker both live in the customer guest
def g(self, argv, timeout=1800):
"""Run in the customer GUEST whatever the layer (its health)."""
return self.r.guest(self.vmid, argv, timeout)
def dpkg_cmp(self, a, op, b):
# The HOST's dpkg: the same Debian version algorithm, and no `pct exec` (0.9 s) per comparison (R-845).
rc, _, _ = self.r.host(["dpkg", "--compare-versions", a, op, b], 30)
return rc == 0
def installed(self):
rc, out, _ = self.g(["dpkg-query", "-W", "-f", "${Package}\t${Version}\t${db:Status-Abbrev}\n"])
rc, out, _ = self.x(["dpkg-query", "-W", "-f", "${Package}\t${Version}\t${db:Status-Abbrev}\n"])
res = {}
for l in out.splitlines():
parts = l.split("\t")
@@ -182,21 +522,20 @@ class Apply:
res[parts[0]] = parts[1]
return res
def dpkg_cmp(self, a, op, b):
rc, _, _ = self.g(["dpkg", "--compare-versions", a, op, b])
return rc == 0
def madison(self, name):
rc, out, _ = self.g(["apt-cache", "madison", name])
vs = set()
def madison_all(self, names):
"""name -> set of downloadable versions, ONE call for all names."""
res = {n: set() for n in names}
if not names:
return res
rc, out, _ = self.x(["apt-cache", "madison"] + sorted(names))
for l in out.splitlines():
f = [x.strip() for x in l.split("|")]
if len(f) >= 3 and f[0] == name:
vs.add(f[1])
return vs
if len(f) >= 3 and f[0] in res:
res[f[0]].add(f[1])
return res
def simulate(self, args):
rc, out, err = self.g(APT_ENV + ["apt-get", "-s", "-q"] + args)
rc, out, err = self.x(APT_ENV + ["apt-get", "-s", "-q"] + args)
inst, remv = [], []
for l in out.splitlines():
m = re.match(r"^Inst (\S+) (?:\[([^]]*)\] )?\((\S+) (.*?) \[[a-z0-9]+\]\)", l)
@@ -213,46 +552,61 @@ class Apply:
return {o.strip().split(":")[0] for o in origin.split(",") if o.strip()}
def free_bytes(self):
rc, out, _ = self.g(["df", "-B1", "--output=avail", "/"])
rc, out, _ = self.x(["df", "-B1", "--output=avail", "/"])
try:
return int(out.strip().splitlines()[-1])
except (ValueError, IndexError):
return -1
def apt_lock_held(self):
rc, out, _ = self.g(["fuser", "/var/lib/dpkg/lock-frontend", "/var/lib/dpkg/lock"])
rc, out, _ = self.x(["fuser", "/var/lib/dpkg/lock-frontend", "/var/lib/dpkg/lock"])
return rc == 0 and out.strip() != ""
def health(self):
def guest_health(self):
"""The guest's signals: every container's state + health, the controller's own health, the network."""
rc, out, _ = self.g(["docker", "ps", "-a", "--format", "{{.Names}}\t{{.State}}\t{{.Status}}"], timeout=60)
rc, out, _ = self.g(["docker", "ps", "-a", "--no-trunc", "--format", "{{.Names}}\t{{.State}}\t{{.Status}}\t{{.ID}}"], timeout=60)
cont = {}
for l in out.splitlines():
p = l.split("\t")
if len(p) == 3:
if len(p) >= 3:
h = "healthy" if "(healthy)" in p[2] else "unhealthy" if "(unhealthy)" in p[2] else \
"starting" if "(health: starting)" in p[2] else "none"
cont[p[0]] = {"state": p[1], "health": h}
if len(p) >= 4 and p[3]:
cont[p[0]]["id"] = p[3]
nrc, _, _ = self.g(["getent", "hosts", "deb.debian.org"], timeout=30)
return {"docker_ok": rc == 0, "containers": cont,
"controller": cont.get("felhom-controller", {}).get("health", "absent"),
"network_ok": nrc == 0}
def restart_needed(self):
"""Processes still mapping deleted files, OUTSIDE docker containers (C11)."""
script = ('for p in /proc/[0-9]*; do grep -q "(deleted)" $p/maps 2>/dev/null || continue; '
'grep -q "docker" $p/cgroup 2>/dev/null && continue; echo "${p#/proc/} $(cat $p/comm 2>/dev/null)"; done')
rc, out, _ = self.g(["sh", "-c", script], timeout=120)
procs = sorted({l.split(" ", 1)[1] for l in out.splitlines() if " " in l})
pid1 = any(l.split(" ", 1)[0] == "1" for l in out.splitlines())
return procs, pid1
def health(self):
if self.layer in ("guest", "docker"):
return self.guest_health()
rc, out, _ = self.r.host(["systemctl", "is-active"] + HOST_SERVICES, 30)
states = out.split()
svc = {s: (states[i] if i < len(states) else "unknown") for i, s in enumerate(HOST_SERVICES)}
src, sout, _ = self.r.host(["/usr/sbin/pct", "status", str(self.vmid)], 30)
running = src == 0 and "running" in sout
return {"host_services": svc, "guest_running": running, "guest": self.guest_health() if running else None}
def inventory(self):
inst = self.installed()
def restart_needed(self):
"""Processes still mapping deleted files, OUTSIDE containers (C11). Guest: outside docker; host: outside the
LXC guests (the host's /proc shows guest processes too)."""
skip = RESTART_SKIP_CGROUP["host" if self.layer == "host" else "guest"]
script = ('for p in /proc/[0-9]*; do grep -q "(deleted)" $p/maps 2>/dev/null || continue; '
'grep -q "%s" $p/cgroup 2>/dev/null && continue; echo "${p#/proc/} $(cat $p/comm 2>/dev/null)"; done' % skip)
rc, out, _ = self.x(["sh", "-c", script], timeout=120)
lines = [l for l in out.splitlines() if " " in l]
procs = sorted({l.split(" ", 1)[1] for l in lines})
pid1 = any(l.split(" ", 1)[0] == "1" for l in lines)
return procs, pid1 or "lxc-start" in procs
def inventory(self, inst=None):
inst = inst if inst is not None else self.installed()
names = sorted(inst)
origins = {}
for i in range(0, len(names), 200):
rc, out, _ = self.g(["apt-cache", "policy"] + names[i:i + 200])
if names:
rc, out, _ = self.x(["apt-cache", "policy"] + names) # ONE call (R-845)
cur, star = None, False
for l in out.splitlines():
if not l.startswith(" "):
@@ -270,10 +624,12 @@ class Apply:
continue
if star and not re.match(r"^[0-9-]+ ", s):
star = False
# Map an index URL to an origin name the hub understands.
def oname(src):
if src in (None, "local"):
return "unknown"
if "proxmox" in src:
return "Proxmox"
if "security" in src and "debian" in src:
return "Debian-Security"
if "docker.com" in src:
@@ -282,6 +638,7 @@ class Apply:
return "Debian"
return "other"
rc, pend, remv, _ = self.simulate(["dist-upgrade"])
self._pending = pend
return {
"installed": [{"name": n, "version": inst[n], "origin": oname(origins.get(n))} for n in names],
"pending": [{"name": p["name"], "from": p["from"], "to": p["to"],
@@ -291,67 +648,131 @@ class Apply:
# ---------- the run ----------
def run(self):
plan = self.load_plan()
self.mode, self.vmid = self.check_plan(plan)
self.report.update(mode=self.mode, release_id=plan.get("release_id"), vmid=self.vmid)
self.mode, self.layer, self.vmid, self.select = self.check_plan(plan)
self.report.update(mode=self.mode, layer=self.layer, release_id=plan.get("release_id"), vmid=self.vmid)
if self.mode == "facts":
return self.facts()
if self.layer == "host":
self.check_appliance()
self.check_guest(self.vmid)
log = self.r.log
if self.mode == "live-restore-on":
return self.live_restore_on()
self.who, self.allow_downgrade = ("fast", False)
if self.layer == "docker" and self.mode == "apply":
self.who, self.allow_downgrade = self.docker_authority(plan)
if self.live_restore() != "true":
raise Refused("R15", "live-restore is not ON in the guest — a Docker step would restart every container")
self.report["authority"] = self.who
self.report["undo"] = self.allow_downgrade
if self.mode == "health":
self.report["health"] = self.health()
return 0
log(f"os-apply: START release={plan.get('release_id')} layer=guest:{self.vmid} lane=fast mode={self.mode} packages={len(plan.get('packages', []))}")
log(f"os-apply: START release={plan.get('release_id')} layer={self.layer}" +
(f":{self.vmid}" if self.layer != "host" else "") +
f" lane={plan.get('lane', 'fast')} mode={self.mode} select={self.select} packages={len(plan.get('packages', []))}" +
(f" authority={self.who}{' UNDO' if self.allow_downgrade else ''}" if self.layer == "docker" else ""))
if self.apt_lock_held():
raise Refused("R9", "another apt/dpkg holds the lock in the guest")
raise Refused("R9", f"another apt/dpkg holds the lock on the {self.layer}")
self.report["health_before"] = self.health()
if self.mode == "apply":
self.repair()
rc, out, err = self.g(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
rc, out, err = self.x(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
if rc != 0:
raise Refused("R7", f"apt-get update failed in the guest: {(out + err).strip().splitlines()[-1:]}")
raise Refused("R7", f"apt-get update failed on the {self.layer}: {(out + err).strip().splitlines()[-1:]}")
installed_after = None
if self.mode == "apply":
rc = self.apply(plan)
rc, installed_after = self.apply(plan)
if rc:
return rc
procs, pid1 = self.restart_needed()
self.report.update(self.inventory())
self.report["restart_needed"] = procs
self.report["docker_restart_needed"] = any(p in ("dockerd", "containerd") for p in procs)
self.report["reboot_needed"] = pid1
self.report.update(self.inventory(installed_after))
if "reboot_needed" not in self.report:
# EVERY layer is scanned on EVERY pass: a reboot (host) or a restart (guest) must CLEAR "restart needed",
# or the fleet view keeps a stale date (R-849, v0.142.0; the host since v0.141.1). One pct exec, ~1 s.
# Pinned by test_every_layer_scans_every_pass.
self.report["restart_needed"], self.report["reboot_needed"] = self.restart_needed()
self.report["docker_restart_needed"] = any(p in ("dockerd", "containerd") for p in self.report["restart_needed"])
if self.layer == "docker":
rc_v, out_v, _ = self.g(["docker", "version", "--format", "{{.Server.Version}}"], timeout=60)
self.report["docker_engine"] = out_v.strip() if rc_v == 0 and out_v.strip() else "unknown"
self.report["reboot_scanned"] = "reboot_needed" in self.report
self.report["health_after"] = self.health()
return 0
def repair(self):
rc, before, _ = self.g(["dpkg", "--audit"])
self.g(APT_ENV + ["dpkg", "--configure", "-a", "--force-confold"])
rc2, out, err = self.g(APT_ENV + ["apt-get", "-f", "install", "-y", "-q"] + DPKG_OPTS)
_, after, _ = self.g(["dpkg", "--audit"])
rc, before, _ = self.x(["dpkg", "--audit"])
configured = len([l for l in before.splitlines() if l.startswith(" ")])
fixed = len(re.findall(r"^Setting up ", out, re.M))
fixed = 0
after = ""
if before.strip(): # nothing half-done → nothing to run (R-845: two calls saved on every clean pass)
self.x(APT_ENV + ["dpkg", "--configure", "-a", "--force-confold"])
rc2, out, err = self.x(APT_ENV + ["apt-get", "-f", "install", "-y", "-q"] + DPKG_OPTS)
_, after, _ = self.x(["dpkg", "--audit"])
fixed = len(re.findall(r"^Setting up ", out, re.M))
self.report["repair"] = {"half_configured_before": configured, "fixed": fixed, "clean_after": after.strip() == ""}
self.r.log(f"os-apply: REPAIR configured={configured} fixed={fixed}")
if after.strip():
raise Refused("R13", "dpkg is still broken after the repair: " + after.strip().splitlines()[0])
def pending_fast(self):
"""Ring 0 (select pending-fast): every pending upgrade of an INSTALLED package whose every origin is Debian /
Debian-Security — and, on the host, not a kernel / boot / firmware package."""
rc, pend, remv, _ = self.simulate(["dist-upgrade"])
out = []
for p in pend:
o = self.origin_name(p["origin"])
if p["from"] is None or not o or not o <= set(FAST_ORIGINS):
continue
if self.layer == "host" and HOST_SLOW_RE.match(p["name"]):
continue
out.append({"name": p["name"], "version": p["to"], "origin": "Debian-Security" if "Debian-Security" in o else "Debian"})
return out
def pending_docker(self):
"""Ring 0 (select pending-docker): the newest pending version of each INSTALLED Docker package, Docker origin."""
rc, pend, remv, _ = self.simulate(["dist-upgrade"])
return [{"name": p["name"], "version": p["to"], "origin": DOCKER_ORIGIN} for p in pend
if p["from"] is not None and p["name"] in DOCKER_NAMES and self.origin_name(p["origin"]) == {DOCKER_ORIGIN}]
def origin_ok(self, origin):
o = self.origin_name(origin)
if self.layer == "docker":
return o == {DOCKER_ORIGIN}
return bool(o & set(FAST_ORIGINS))
def apply(self, plan):
log = self.r.log
if self.select == "listed":
packages = plan["packages"]
elif self.select == "pending-docker":
packages = self.pending_docker()
else:
packages = self.pending_fast()
cmp_op = "ne" if self.allow_downgrade else "gt"
inst = self.installed()
upgrade, already, notinst = [], 0, 0
for e in plan["packages"]:
for e in packages:
n, v = e["name"], e["version"]
if n not in inst:
notinst += 1
continue
if not self.dpkg_cmp(v, "gt", inst[n]):
if not self.dpkg_cmp(v, cmp_op, inst[n]):
already += 1
continue
upgrade.append((n, v))
from_snap = 0
missing = [(n, v) for n, v in upgrade if v not in self.madison(n)]
if upgrade:
avail = self.madison_all([n for n, _ in upgrade])
missing = [(n, v) for n, v in upgrade if v not in avail[n]]
else:
missing = []
if missing:
snap = plan.get("snapshot", "")
if not snap:
raise Refused("R7", f"{missing[0][0]}={missing[0][1]} is not downloadable and the plan names no snapshot")
self.add_snapshot_sources(snap)
still = [(n, v) for n, v in missing if v not in self.madison(n)]
avail = self.madison_all([n for n, _ in missing])
still = [(n, v) for n, v in missing if v not in avail[n]]
if still:
self.remove_snapshot_sources()
raise Refused("R7", f"{still[0][0]}={still[0][1]} is not downloadable, not even from snapshot {snap}")
@@ -362,8 +783,9 @@ class Apply:
if not upgrade:
self.report["upgraded"] = []
log("os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)")
return 0
args = ["install", "--only-upgrade", "--no-install-recommends"] + [f"{n}={v}" for n, v in upgrade]
return 0, inst
args = ["install", "--only-upgrade", "--no-install-recommends"] + \
(["--allow-downgrades"] if self.allow_downgrade else []) + [f"{n}={v}" for n, v in upgrade]
rc, sim, remv, text = self.simulate(args)
if rc != 0:
tail = text.strip().splitlines()[-1] if text.strip() else ""
@@ -378,16 +800,18 @@ class Apply:
raise Refused("R6", f"the plan would touch {p['name']}, which is not in the plan")
if p["to"] != want[p["name"]]:
raise Refused("R6", f"{p['name']} would go to {p['to']}, not the approved {want[p['name']]}")
if not self.dpkg_cmp(p["to"], "gt", p["from"]):
if not self.allow_downgrade and not self.dpkg_cmp(p["to"], "gt", p["from"]):
raise Refused("R5", f"{p['name']} would be downgraded {p['from']} -> {p['to']}")
if not self.origin_name(p["origin"]) & set(FAST_ORIGINS):
raise Refused("R2", f"{p['name']} would come from {p['origin']}, not Debian")
if not self.origin_ok(p["origin"]):
raise Refused("R2", f"{p['name']} would come from {p['origin']}, not the {self.layer} layer's origin")
if self.layer == "host" and HOST_SLOW_RE.match(p["name"]):
raise Refused("R14", f"{p['name']} is a kernel / boot / firmware package — the host's slow lane")
need = self.download_bytes(args)
free = self.free_bytes()
if free >= 0 and free < max(MIN_FREE, 3 * need):
raise Refused("R8", f"free space {free} B is below max(500 MB, 3 x download {need} B)")
t0 = time.time()
rc, out, err = self.g(APT_ENV + ["apt-get", "-y", "-q"] + DPKG_OPTS + args)
rc, out, err = self.x(APT_ENV + ["apt-get", "-y", "-q"] + DPKG_OPTS + args)
secs = time.time() - t0
# dpkg says "Installing new version of config file X" when X was NOT changed locally (the package's new
# version is taken), and "Configuration file 'X'" + "Keeping old config file" when it was (--force-confold
@@ -404,23 +828,27 @@ class Apply:
log(f"os-apply: CONFFILE kept {conflict} (changed locally; the package's version is {conflict}.dpkg-dist)")
self.report.setdefault("conffiles_kept", []).append(conflict)
conflict = None
self.g(["apt-get", "clean"])
self.x(["apt-get", "clean"])
if rc != 0:
_, aud, _ = self.g(["dpkg", "--audit"])
_, aud, _ = self.x(["dpkg", "--audit"])
first = aud.strip().splitlines()[0] if aud.strip() else "clean"
log(f"os-apply: FAILED rc={rc} step=install — dpkg state: {first}")
self.report["failed"] = {"rc": rc, "dpkg_audit": first, "tail": (out + err).strip().splitlines()[-3:]}
return 3
return 3, None
self.report["upgraded"] = [{"name": n, "version": v} for n, v in upgrade]
self.report["seconds"] = round(secs, 1)
log(f"os-apply: DONE rc=0 seconds={secs:.1f} upgraded={len(upgrade)}")
return 0
procs, reboot = self.restart_needed()
self.report["restart_needed"] = procs
self.report["docker_restart_needed"] = any(p in ("dockerd", "containerd") for p in procs)
self.report["reboot_needed"] = reboot
log(f"os-apply: DONE rc=0 seconds={secs:.1f} upgraded={len(upgrade)} restart-needed={','.join(procs) or '-'} reboot-needed={'yes' if reboot else 'no'}")
return 0, None
finally:
if from_snap:
self.remove_snapshot_sources()
def download_bytes(self, args):
rc, out, _ = self.g(APT_ENV + ["apt-get", "-s", "-o", "Debug::NoLocking=1", "--print-uris", "-q"] + args)
rc, out, _ = self.x(APT_ENV + ["apt-get", "-s", "-o", "Debug::NoLocking=1", "--print-uris", "-q"] + args)
total = 0
for l in out.splitlines():
m = re.match(r"^'[^']+' \S+ ([0-9]+) ", l)
@@ -429,33 +857,22 @@ class Apply:
return total
def add_snapshot_sources(self, snap):
rc, out, _ = self.g(["sh", "-c", ". /etc/os-release && echo $VERSION_CODENAME"])
rc, out, _ = self.x(["sh", "-c", ". /etc/os-release && echo $VERSION_CODENAME"])
code = out.strip()
if not re.match(r"^[a-z]+$", code):
raise Refused("R7", f"cannot read the guest's Debian codename ({code!r})")
raise Refused("R7", f"cannot read the {self.layer}'s Debian codename ({code!r})")
body = (f"deb [check-valid-until=no] http://snapshot.debian.org/archive/debian/{snap} {code} main\n"
f"deb [check-valid-until=no] http://snapshot.debian.org/archive/debian-security/{snap} {code}-security main\n")
self.r.guest_write(self.vmid, SNAPSHOT_LIST, body)
self.r.write_file(self.layer, self.vmid, SNAPSHOT_LIST, body)
self.r.log(f"os-apply: SNAPSHOT using snapshot.debian.org/{snap} for versions no longer published (decision 79)")
rc, out, err = self.g(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
rc, out, err = self.x(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
if rc != 0:
self.remove_snapshot_sources()
raise Refused("R7", "apt-get update against snapshot.debian.org failed")
def remove_snapshot_sources(self):
self.g(["rm", "-f", SNAPSHOT_LIST])
self.g(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
def guest_write(self, vmid, path, body):
"""Write a small text file inside the guest via `pct exec … tee` (stdin), never via a shell string."""
p = subprocess.run(["/usr/sbin/pct", "exec", str(vmid), "--", "tee", path], input=body,
capture_output=True, text=True, timeout=60)
if p.returncode != 0:
raise Refused("R7", f"could not write {path} in the guest")
Runner.guest_write = guest_write
self.x(["rm", "-f", SNAPSHOT_LIST])
self.x(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
def main(argv, runner=None):
@@ -465,6 +882,7 @@ def main(argv, runner=None):
print("OSAPPLY-REPORT " + json.dumps({"refused": {"code": "R1", "reason": "usage"}}))
return 2
a = Apply(r, argv[2])
t0 = time.time()
try:
rc = a.run()
except Refused as e:
@@ -475,6 +893,7 @@ def main(argv, runner=None):
r.log(f"os-apply: FAILED rc=124 step=timeout — {e.cmd}")
a.report["failed"] = {"rc": 124, "timeout": str(e.cmd)[:200]}
rc = 3
a.report["pass_seconds"] = round(time.time() - t0, 1)
print("OSAPPLY-REPORT " + json.dumps(a.report, sort_keys=True))
return rc
+143
View File
@@ -0,0 +1,143 @@
#!/usr/bin/python3
"""Tests for felhom-crash-guard (`11` §5.9). Temp dirs only; nothing real is touched. Red-proof seam: CRASHGUARD_UNDER_TEST."""
import importlib.machinery
import importlib.util
import json
import os
import pathlib
import tempfile
import unittest
HERE = pathlib.Path(__file__).resolve().parent
_loader = importlib.machinery.SourceFileLoader("crashguard", os.environ.get("CRASHGUARD_UNDER_TEST", str(HERE / "felhom-crash-guard")))
_spec = importlib.util.spec_from_loader("crashguard", _loader)
cg = importlib.util.module_from_spec(_spec)
_loader.exec_module(cg)
T0 = 1791115200.0 # 2026-10-04T12:00:00Z
class FakeEnv(cg.Env):
def __init__(self, d):
super().__init__(conf=os.path.join(d, "conf"), state_dir=os.path.join(d, "state"),
panic_path=os.path.join(d, "panic"), uptime_path=os.path.join(d, "uptime"),
boot_id_path=os.path.join(d, "bootid"))
self.t = T0
self.logs = []
open(self.panic_path, "w").write("0\n")
open(self.uptime_path, "w").write("20.00 10.00\n")
def now(self):
return self.t
def log(self, line):
self.logs.append(line)
def panic(self):
return int(open(self.panic_path).read())
def state(self):
return json.load(open(os.path.join(self.state_dir, "state.json")))
class Guard(unittest.TestCase):
def setUp(self):
self.d = tempfile.TemporaryDirectory()
self.e = FakeEnv(self.d.name)
def tearDown(self):
self.d.cleanup()
def crash_boot(self, minutes_later):
self.e.t += minutes_later * 60
cg.main(["x", "boot"], self.e) # no clean-stop before it: an unclean stop
def clean_reboot(self, minutes_later):
cg.main(["x", "clean-stop"], self.e)
self.e.t += minutes_later * 60
cg.main(["x", "boot"], self.e)
def test_first_boot_is_not_a_crash_and_arms(self):
cg.main(["x", "boot"], self.e)
s = self.e.state()
self.assertFalse(s["last_boot_unclean"])
self.assertEqual(self.e.panic(), 10)
self.assertTrue(s["armed"])
def test_clean_reboots_never_count(self):
cg.main(["x", "boot"], self.e)
for _ in range(5):
self.clean_reboot(1)
s = self.e.state()
self.assertEqual(s["unclean_boots_in_window"], 0)
self.assertEqual(self.e.panic(), 10)
def test_third_crash_in_an_hour_leaves_the_box_off(self):
# operator's words: "if it crashes 3 times within one hour, it stays off" — after crash 2 the guard trips,
# so crash 3 (kernel.panic = 0) does not restart the box.
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.assertEqual(self.e.panic(), 10, "one crash: still restarts")
self.crash_boot(5)
s = self.e.state()
self.assertTrue(s["tripped"], s)
self.assertEqual(self.e.panic(), 0, "after the 2nd crash boot the 3rd crash must leave the box off")
self.assertIn("2 unclean boots within 60 minutes", s["tripped_reason"])
def test_crashes_spread_over_more_than_the_window_do_not_trip(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(61)
self.assertFalse(self.e.state()["tripped"])
self.assertEqual(self.e.panic(), 10)
def test_tripped_stays_tripped_across_boots(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(5)
self.clean_reboot(30) # the operator switched it on; even a clean boot keeps the trip
self.assertTrue(self.e.state()["tripped"])
self.assertEqual(self.e.panic(), 0)
def test_rearms_after_24h_of_normal_running(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(5)
self.e.t += 23 * 3600
cg.main(["x", "check"], self.e)
self.assertTrue(self.e.state()["tripped"], "not before 24 h")
self.e.t += 3600
cg.main(["x", "check"], self.e)
s = self.e.state()
self.assertFalse(s["tripped"])
self.assertEqual(self.e.panic(), 10)
self.assertIn("timer", s["rearmed_by"])
def test_operator_rearm_starts_a_fresh_window(self):
cg.main(["x", "boot"], self.e)
self.crash_boot(5)
self.crash_boot(5)
self.e.t += 60
cg.main(["x", "rearm"], self.e)
s = self.e.state()
self.assertFalse(s["tripped"])
self.assertEqual(s["rearmed_by"], "operator")
self.assertEqual(s["unclean_boots_24h"], 2, "the history stays")
self.crash_boot(5)
self.assertFalse(self.e.state()["tripped"], "one crash after a re-arm must not trip at once")
def test_config_numbers_are_read(self):
open(self.e.conf, "w").write("LIMIT=2\nPANIC_SECONDS=30\n")
cg.main(["x", "boot"], self.e)
self.assertEqual(self.e.panic(), 30)
self.crash_boot(1)
self.assertTrue(self.e.state()["tripped"], "LIMIT=2: the first crash boot trips")
def test_state_is_world_readable_for_the_agent(self):
cg.main(["x", "boot"], self.e)
mode = os.stat(os.path.join(self.e.state_dir, "state.json")).st_mode & 0o777
self.assertEqual(mode, 0o644)
if __name__ == "__main__":
unittest.main()
+478 -9
View File
@@ -61,6 +61,38 @@ class Fake:
self.written = {}
self.status = "status: running"
self.lock_held = False
self.services = {}
self.files[osapply.INSTALL_STATE] = json.dumps({"mode": "appliance"})
self.stats[osapply.INSTALL_STATE] = St(mode=statmod.S_IFREG | 0o644, uid=0)
# Docker / trust / facts state (v0.142.0)
self.live_restore = "true"
self.ids = ["aaa111", "bbb222"]
self.engine = "29.7.2"
self.daemon_json = '{"log-driver": "json-file"}'
self.reload_enables = True
self.sig_rc = 0
self.nonces = {}
self.clock = 1791115200.0 # 2026-10-04T12:00:00Z
self.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": False})
self.stats[osapply.TRUST_FILE] = St(mode=statmod.S_IFREG | 0o644, uid=0)
self.files[osapply.TRUST_SIGNERS] = 'felhom-op-1 namespaces="felhom-op-v1" ssh-ed25519 AAAA\n'
self.stats[osapply.TRUST_SIGNERS] = St(mode=statmod.S_IFREG | 0o644, uid=0)
def now(self):
return self.clock
def sleep(self, s):
pass
def verify_sig(self, signers, key_id, ns, blob, sig):
self.verified = (signers, key_id, ns, blob, sig)
return self.sig_rc
def read_nonces(self):
return dict(self.nonces)
def write_nonces(self, d):
self.nonces = dict(d)
# Runner interface
def read_file(self, p):
@@ -81,14 +113,19 @@ class Fake:
def log(self, line):
self.logs.append(line)
def host(self, argv, timeout=600):
def host(self, argv, timeout=600, stdin=None):
self.calls.append(("host", argv))
if argv[1] == "status":
if argv[0] == "/usr/sbin/pct" and argv[1] == "status":
return 0, self.status + "\n", ""
return 1, "", "unexpected host call"
if argv[0] == "dpkg" and argv[1] == "--compare-versions":
return (0 if dpkg_cmp(argv[2], argv[3], argv[4]) else 1), "", ""
if argv[0] == "systemctl" and argv[1] == "is-active":
return 0, "\n".join(self.services.get(s, "active") for s in argv[2:]) + "\n", ""
return self.emulate(argv)
def guest_write(self, vmid, path, body):
def write_file(self, layer, vmid, path, body):
self.written[path] = body
self.write_layer = layer
if path == osapply.SNAPSHOT_LIST:
self.snap_active = True
@@ -100,6 +137,9 @@ class Fake:
def guest(self, vmid, argv, timeout=1800):
self.calls.append(("guest", vmid, argv))
return self.emulate(argv)
def emulate(self, argv):
a = [x for x in argv if not re.match(r"^[A-Z_]+=", x) and x != "env"]
cmd = a[0]
if cmd == "dpkg-query":
@@ -113,7 +153,8 @@ class Fake:
if cmd == "fuser":
return (0, " 123", "") if self.lock_held else (1, "", "")
if cmd == "apt-cache" and a[1] == "madison":
return 0, "".join(f" {a[2]} | {v} | http://deb.debian.org trixie/main amd64 Packages\n" for v in self.avail(a[2])), ""
self.madison_calls = getattr(self, "madison_calls", 0) + 1
return 0, "".join(f" {n} | {v} | http://deb.debian.org trixie/main amd64 Packages\n" for n in a[2:] for v in self.avail(n)), ""
if cmd == "apt-cache" and a[1] == "policy":
out = ""
for n in a[2:]:
@@ -139,13 +180,42 @@ class Fake:
return 0, getattr(self, "install_out", "Setting up libc6 ...\n"), ""
if cmd == "df":
return 0, f"Avail\n{self.free}\n", ""
if cmd == "docker" and a[1] == "info":
return 0, self.live_restore + "\n", ""
if cmd == "docker" and a[1] == "version":
return 0, self.engine + "\n", ""
if cmd == "docker" and a[1:3] == ["ps", "-q"]:
return 0, "".join(i + "\n" for i in self.ids), ""
if cmd == "docker":
return 0, "felhom-controller\trunning\tUp 1 hour (healthy)\napp\trunning\tUp 1 hour (healthy)\n", ""
ids = self.ids + ["x"] * 2
return 0, f"felhom-controller\trunning\tUp 1 hour (healthy)\t{ids[0]}\napp\trunning\tUp 1 hour (healthy)\t{ids[1]}\n", ""
if cmd == "cat" and a[1] == osapply.DAEMON_JSON:
return (0, self.daemon_json, "") if self.daemon_json is not None else (1, "", "No such file")
if cmd == "cat" and a[1] == "/etc/debian_version":
return 0, "13.7\n", ""
if cmd == "uname":
return 0, "7.0.14-20-pve\n", ""
if cmd == "apt-mark":
return 0, getattr(self, "held", ""), ""
if cmd == "systemctl" and a[1] == "reload":
self.reloads = getattr(self, "reloads", 0) + 1
if self.reload_enables and '"live-restore": true' in self.written.get(osapply.DAEMON_JSON, ""):
self.live_restore = "true"
return 0, "", ""
if cmd == "systemctl" and a[1] == "restart":
self.restarted = True
return 0, "", ""
if cmd == "getent":
return 0, "1.2.3.4 deb.debian.org\n", ""
if cmd == "sh":
if "vmlinuz" in a[2]:
return 0, "/boot/vmlinuz-7.0.2-6-pve\n/boot/vmlinuz-7.0.14-20-pve\n", ""
if "engine=" in a[2]:
return 0, f"debian=13.7\nengine={self.engine}\ncontainerd=2.3.3-1~debian.13~trixie\nlive={self.live_restore}\n", ""
if "os-release" in a[2]:
return 0, "trixie\n", ""
if "(deleted)" in a[2]:
return 0, getattr(self, "restart_out", ""), ""
return 0, "", ""
if cmd == "rm":
self.snap_active = False
@@ -156,6 +226,9 @@ class Fake:
if "--print-uris" in a:
return 0, "'http://x/libc6.deb' libc6.deb 4000000 SHA256:x\n", ""
if "dist-upgrade" in a:
if getattr(self, "pending_sim", None) is not None and not getattr(self, "_pending_used", False):
self._pending_used = True
return 0, "\n".join(self.pending_sim) + "\n", ""
return 0, "Inst bash [5.2.37-2+b9] (5.2.37-2+b10 Debian:13.7/stable [amd64])\n", ""
out = ""
for x in a:
@@ -163,7 +236,7 @@ class Fake:
n, v = x.split("=", 1)
if v not in self.avail(n):
return 100, "", f"E: Version '{v}' for '{n}' was not found"
origin = SEC if n == "openssl" else DEB
origin = "Docker CE:trixie" if n in osapply.DOCKER_NAMES else SEC if n == "openssl" else DEB
out += f"Inst {n} [{self.installed[n]}] ({v} {origin} [amd64])\n"
out += "".join(l + "\n" for l in self.extra_sim)
return 0, out, ""
@@ -209,7 +282,6 @@ class Happy(unittest.TestCase):
self.assertEqual(rc, 0, rep)
self.assertEqual(f.installed["libc6"], "2.41-12+deb13u3")
self.assertIn("installed", rep)
self.assertIn("restart_needed", rep)
def test_health_mode(self):
f = Fake()
@@ -419,11 +491,44 @@ class Refusals(unittest.TestCase):
f.plan["packages"][0]["name"] = "--purge"
self.refused(f, "R11")
def test_R12_host_layer(self):
def test_R12_unknown_layer(self):
f = Fake()
f.plan["layer"] = "vm"
self.refused(f, "R12")
def test_R12_host_on_a_byo_box(self):
f = Fake()
f.plan["layer"] = "host"
f.files[osapply.INSTALL_STATE] = json.dumps({"mode": "byo"})
self.refused(f, "R12")
def test_R12_host_without_an_install_record(self):
f = Fake()
f.plan["layer"] = "host"
del f.stats[osapply.INSTALL_STATE]
self.refused(f, "R12")
def test_R12_host_record_not_root_owned(self):
# the agent can write agent.json's deployment_mode; only a ROOT-owned record proves anything
f = Fake()
f.plan["layer"] = "host"
f.stats[osapply.INSTALL_STATE] = St(mode=statmod.S_IFREG | 0o644, uid=999)
self.refused(f, "R12")
def test_R14_kernel_package_in_a_host_plan(self):
f = Fake()
f.plan["layer"] = "host"
f.plan["packages"].append({"name": "linux-image-amd64", "version": "6.12.1-1", "origin": "Debian"})
self.refused(f, "R14")
def test_R14_kernel_package_pulled_by_the_simulation(self):
f = Fake()
f.plan["layer"] = "host"
f.extra_sim = ["Inst grub-common [2.12-9] (2.12-10 Debian:13.7/stable [amd64])"]
f.installed["grub-common"] = "2.12-9"
f.plan["packages"].append({"name": "grub-common", "version": "2.12-10", "origin": "Debian"})
self.refused(f, "R14")
def test_R13_repair_does_not_fix_it(self):
f = Fake()
f.dpkg_audit = "The following packages are broken\n perl\n"
@@ -453,6 +558,96 @@ class Conffiles(unittest.TestCase):
self.assertFalse(any("kept /etc/debian_version" in l for l in f.logs), "an updated file must not be reported as kept")
class HostLayer(unittest.TestCase):
def test_host_runs_on_the_host_not_in_the_guest(self):
f = Fake()
f.plan["layer"] = "host"
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
inst = [c for c in f.calls if c[0] == "host" and "install" in c[1] and "-s" not in c[1] and "-f" not in c[1]]
self.assertTrue(inst, "the host install must run on the host")
self.assertFalse([c for c in f.calls if c[0] == "guest" and "install" in c[2]], "nothing installed in the guest")
self.assertEqual(sorted(rep["health_after"]["host_services"]), sorted(osapply.HOST_SERVICES))
self.assertTrue(rep["health_after"]["guest_running"])
def test_pending_fast_skips_proxmox_docker_and_kernel(self):
f = Fake()
f.plan["layer"] = "host"
f.plan["select"] = "pending-fast"
f.plan["packages"] = []
f.installed.update({"pve-manager": "9.2.2", "linux-image-amd64": "6.12.1", "docker-ce": "29.7"})
f.pending_sim = [
"Inst libc6 [2.41-12+deb13u3] (2.41-12+deb13u4 Debian:13.7/stable [amd64])",
"Inst openssl [3.5.6-1~deb13u1] (3.5.7-1~deb13u3 Debian:13.7/stable, Debian-Security:13/stable-security [amd64])",
"Inst pve-manager [9.2.2] (9.2.21 Proxmox Debian Repository:stable [amd64])",
"Inst linux-image-amd64 [6.12.1] (6.12.9 Debian:13.7/stable [amd64])",
"Inst docker-ce [29.7] (29.8 Docker CE:trixie [amd64])",
"Inst brand-new (1.0 Debian:13.7/stable [amd64])",
]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
got = sorted(u["name"] for u in rep["upgraded"])
self.assertEqual(got, ["libc6", "openssl"], "pending-fast must take only installed, Debian-origin, non-kernel packages")
def test_reboot_needed_when_pid1_or_lxc_start(self):
f = Fake()
f.plan["layer"] = "host"
f.restart_out = "1 systemd\n2101 lxc-start\n530 sshd\n"
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertTrue(rep["reboot_needed"])
self.assertIn("lxc-start", rep["restart_needed"])
def test_reboot_needed_for_lxc_start_alone(self):
# lxc-start runs the guest; only a guest restart (or a host reboot) replaces it
f = Fake()
f.plan["layer"] = "host"
f.restart_out = "2101 lxc-start\n530 sshd\n"
rc, rep = run(f)
self.assertTrue(rep["reboot_needed"], rep)
def test_every_layer_scans_every_pass(self):
# R-849 (v0.142.0): host AND guest are scanned on every pass, so a reboot / restart clears the flag.
for layer in ("host", "guest"):
f = Fake()
f.plan["layer"] = layer
f.plan["mode"] = "inventory"
f.restart_out = "2101 lxc-start\n" if layer == "host" else "1 systemd\n"
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertTrue(rep["reboot_scanned"] and rep["reboot_needed"], (layer, rep))
f = Fake()
f.plan["layer"] = layer
f.plan["mode"] = "inventory"
f.restart_out = ""
rc, rep = run(f)
self.assertTrue(rep["reboot_scanned"] and rep["reboot_needed"] is False, (layer, rep))
def test_no_reboot_for_ordinary_daemons(self):
f = Fake()
f.plan["layer"] = "host"
f.restart_out = "530 sshd\n611 cron\n"
rc, rep = run(f)
self.assertFalse(rep["reboot_needed"], rep)
class Speed(unittest.TestCase):
# R-845: no `pct exec` per package — version checks on the host, madison once, the restart scan only after an install.
def test_no_per_package_guest_calls(self):
f = Fake()
for i in range(40):
f.installed[f"pkg{i}"] = "1.0-1"
f.live[f"pkg{i}"] = {"1.0-2"}
f.plan["packages"].append({"name": f"pkg{i}", "version": "1.0-2", "origin": "Debian"})
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
guest_cmp = [c for c in f.calls if c[0] == "guest" and "--compare-versions" in c[2]]
self.assertEqual(guest_cmp, [], "version comparisons must run on the host")
self.assertEqual(f.madison_calls, 1, "madison must run once for all packages")
guest_calls = len([c for c in f.calls if c[0] == "guest"])
self.assertLess(guest_calls, 30, f"{guest_calls} guest calls for 42 packages — something is per-package again")
class Failure(unittest.TestCase):
def test_install_failure_is_rc3_with_dpkg_state(self):
f = Fake()
@@ -463,5 +658,279 @@ class Failure(unittest.TestCase):
self.assertTrue(any(l.startswith("os-apply: FAILED rc=100 step=install") for l in f.logs))
class RestartSkipPattern(unittest.TestCase):
"""The cgroup filter runs as `grep -q PATTERN /proc/<pid>/cgroup`; check it with grep itself against the cgroup
lines measured on demo-felhom 2026-10-04."""
def grep(self, pattern, line):
return subprocess.run(["grep", "-q", pattern], input=line + "\n", text=True).returncode == 0
def test_restart_skip_patterns_against_real_cgroups(self):
host = osapply.RESTART_SKIP_CGROUP["host"]
self.assertTrue(self.grep(host, "0::/lxc/9201/ns/system.slice/docker.service"), "a guest process must be skipped")
self.assertFalse(self.grep(host, "0::/lxc.monitor/9201"), "lxc-start must NOT be skipped (it runs the guest)")
self.assertFalse(self.grep(host, "0::/system.slice/pve-cluster.service"), "a host daemon must NOT be skipped")
guest = osapply.RESTART_SKIP_CGROUP["guest"]
self.assertTrue(self.grep(guest, "0::/system.slice/docker-0123abcd.scope"))
self.assertFalse(self.grep(guest, "0::/system.slice/cron.service"))
DOCKER_SET = [{"name": "docker-ce", "version": "5:29.8.2-1~debian.13~trixie", "origin": "Docker CE"},
{"name": "containerd.io", "version": "2.3.6-1~debian.13~trixie", "origin": "Docker CE"}]
def docker_fake(signed=None, undo=False, ring0=False):
f = Fake()
f.installed.update({"docker-ce": "5:29.7.2-1~debian.13~trixie", "containerd.io": "2.3.3-1~debian.13~trixie"})
f.live["docker-ce"] = {"5:29.8.2-1~debian.13~trixie", "5:29.7.2-1~debian.13~trixie"}
f.live["containerd.io"] = {"2.3.6-1~debian.13~trixie", "2.3.3-1~debian.13~trixie"}
f.plan = {"release_id": "os-docker-t1", "layer": "docker", "lane": "slow", "vmid": 9201, "mode": "apply",
"packages": [dict(p) for p in DOCKER_SET]}
if undo:
f.plan["undo"] = True
if ring0:
f.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": True})
if signed is not None:
f.plan["signed"] = signed
return f
def signed_job(packages=DOCKER_SET, host="demo-hp-bb76ea", op="os_docker_step", undo=False, nonce="n1",
issued="2026-10-04T11:50:00Z", expires="2026-10-04T12:30:00Z"):
import base64
params = {"packages": packages, "undo": undo}
blob = json.dumps({"expires_at": expires, "issued_at": issued, "key_id": "felhom-op-1", "nonce": nonce, "op": op,
"params": params, "target": {"guest_id": "", "host_id": host}}, sort_keys=True).encode()
return {"blob_b64": base64.b64encode(blob).decode(), "sig": "-----BEGIN SSH SIGNATURE-----\nx\n-----END SSH SIGNATURE-----\n"}
class DockerLane(unittest.TestCase):
"""`11` §5.8, agent v0.142.0. Each test names the refusal it pins; the red-proof file mutates each one."""
def refused(self, f, code):
rc, rep = run(f)
self.assertEqual(rc, 2, rep)
self.assertEqual(rep["refused"]["code"], code, rep)
self.assertEqual(f.installed.get("docker-ce", "5:29.7.2-1~debian.13~trixie"), "5:29.7.2-1~debian.13~trixie")
return rep
def test_docker_package_in_a_fast_plan_is_refused(self):
f = Fake()
f.plan["packages"].append({"name": "docker-ce", "version": "5:29.8.2-1~debian.13~trixie", "origin": "Debian"})
self.refused(f, "R2")
def test_docker_layer_in_the_fast_lane_is_refused(self):
f = docker_fake(ring0=True)
f.plan["lane"] = "fast"
self.refused(f, "R3")
def test_no_authority_is_refused(self):
self.refused(docker_fake(), "R3")
def test_ring0_mark_allows_pending_docker(self):
f = docker_fake(ring0=True)
f.plan["select"], f.plan["packages"] = "pending-docker", []
f.pending_sim = ["Inst docker-ce [5:29.7.2-1~debian.13~trixie] (5:29.8.2-1~debian.13~trixie Docker CE:trixie [amd64])",
"Inst containerd.io [2.3.3-1~debian.13~trixie] (2.3.6-1~debian.13~trixie Docker CE:trixie [amd64])",
"Inst bash [5.2.37-2+b9] (5.2.37-2+b10 Debian:13.7/stable [amd64])"]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["authority"], "ring0")
self.assertEqual(f.installed["docker-ce"], "5:29.8.2-1~debian.13~trixie")
self.assertEqual(f.installed["bash"], "5.2.37-2+b9", "a Debian package must not ride a Docker step")
self.assertFalse(getattr(f, "restarted", False), "never a docker restart")
def test_signed_job_applies_exactly_its_packages(self):
f = docker_fake(signed=signed_job())
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["authority"], "signed")
self.assertEqual(f.installed["containerd.io"], "2.3.6-1~debian.13~trixie")
self.assertEqual(f.verified[0], osapply.TRUST_SIGNERS, "the ROOT-OWNED signers file, not the agent's config")
self.assertIn("n1", f.nonces)
def test_bad_signature_is_refused(self):
f = docker_fake(signed=signed_job())
f.sig_rc = 255
self.refused(f, "R3")
self.assertEqual(f.nonces, {}, "a bad signature must not burn a nonce")
def test_signed_job_for_another_host_is_refused(self):
self.refused(docker_fake(signed=signed_job(host="demo-felhom-8363b5")), "R3")
def test_signed_job_other_op_is_refused(self):
self.refused(docker_fake(signed=signed_job(op="agent_update")), "R3")
def test_expired_signed_job_is_refused(self):
self.refused(docker_fake(signed=signed_job(expires="2026-10-04T11:55:00Z")), "R3")
def test_replayed_signed_job_is_refused(self):
f = docker_fake(signed=signed_job())
f.nonces = {"n1": f.clock + 600}
self.refused(f, "R3")
def test_plan_must_equal_the_signed_packages(self):
f = docker_fake(signed=signed_job(packages=DOCKER_SET[:1]))
self.refused(f, "R3")
def test_agent_writable_trust_file_is_refused(self):
f = docker_fake(ring0=True)
f.stats[osapply.TRUST_FILE] = St(mode=statmod.S_IFREG | 0o644, uid=999)
self.refused(f, "R3")
def test_live_restore_off_is_refused(self):
f = docker_fake(signed=signed_job())
f.live_restore = "false"
self.refused(f, "R15")
def test_undo_needs_a_signed_job(self):
self.refused(docker_fake(undo=True, ring0=True), "R3")
def test_signed_undo_downgrades(self):
old = [{"name": "docker-ce", "version": "5:29.7.2-1~debian.13~trixie", "origin": "Docker CE"}]
f = docker_fake(signed=signed_job(packages=old, undo=True), undo=True)
f.plan["packages"] = [dict(p) for p in old]
f.installed["docker-ce"] = "5:29.8.2-1~debian.13~trixie"
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(f.installed["docker-ce"], "5:29.7.2-1~debian.13~trixie")
self.assertTrue(rep["undo"])
def test_unsigned_downgrade_is_refused(self):
f = docker_fake(ring0=True)
f.plan["packages"] = [{"name": "docker-ce", "version": "5:29.6.0-1~debian.13~trixie", "origin": "Docker CE"}]
f.live["docker-ce"].add("5:29.6.0-1~debian.13~trixie")
rc, rep = run(f)
self.assertEqual(rep["plan"]["upgrade"], 0, "an older version on an unsigned step is 'already', never installed")
def test_health_carries_container_ids(self):
f = docker_fake(signed=signed_job())
rc, rep = run(f)
self.assertEqual(rep["health_after"]["containers"]["app"]["id"], "bbb222")
self.assertEqual(rep["docker_engine"], "29.7.2")
class LiveRestore(unittest.TestCase):
def lr(self):
f = Fake()
f.plan = {"release_id": "lr", "layer": "guest", "lane": "fast", "vmid": 9201, "mode": "live-restore-on", "packages": []}
f.live_restore = "false"
return f
def test_turns_it_on_with_a_reload_never_a_restart(self):
f = self.lr()
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(json.loads(f.written[osapply.DAEMON_JSON]), {"log-driver": "json-file", "live-restore": True})
self.assertEqual(f.reloads, 1)
self.assertFalse(getattr(f, "restarted", False))
self.assertEqual(rep["live_restore"]["result"], "on")
self.assertTrue(rep["live_restore"]["same_ids"])
def test_already_on_writes_nothing(self):
f = self.lr()
f.live_restore = "true"
rc, rep = run(f)
self.assertEqual(rc, 0)
self.assertNotIn(osapply.DAEMON_JSON, f.written)
def test_invalid_daemon_json_is_left_alone(self):
f = self.lr()
f.daemon_json = "{not json"
rc, rep = run(f)
self.assertEqual(rep["refused"]["code"], "R16")
self.assertNotIn(osapply.DAEMON_JSON, f.written)
def test_reload_that_does_not_enable_puts_the_file_back(self):
f = self.lr()
f.reload_enables = False
rc, rep = run(f)
self.assertEqual(rc, 3, rep)
self.assertEqual(f.written[osapply.DAEMON_JSON], '{"log-driver": "json-file"}')
self.assertFalse(getattr(f, "restarted", False))
class Facts(unittest.TestCase):
def facts(self, f=None):
f = f or Fake()
f.plan = {"release_id": "facts", "layer": "host", "lane": "fast", "vmid": 9201, "mode": "facts", "packages": []}
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
return f, rep["facts"]
def test_reads_host_and_guest(self):
f = Fake()
f.files["/proc/sys/kernel/tainted"] = "4225\n" # 4096 + 128 (D: oops) + 1
f.files["/proc/sys/kernel/panic"] = "10\n"
f.held = "tzdata\n"
f.files["/etc/default/grub"] = "GRUB_DEFAULT=saved\n"
f.files["/boot/grub/grubenv"] = "# GRUB Environment Block\nsaved_entry=gnulinux-advanced-x>gnulinux-7.0.14-20-pve-advanced-x\n"
_, fa = self.facts(f)
h, g = fa["host"], fa["guest"]
self.assertEqual((h["debian"], h["kernel_running"], h["kernel_next_boot"]), ("13.7", "7.0.14-20-pve", "7.0.14-20-pve"))
self.assertEqual(h["held"], ["tzdata"])
self.assertTrue(h["oops_this_boot"])
self.assertEqual(h["kernel_panic"], 10)
self.assertEqual((g["debian"], g["docker_engine"], g["live_restore"]), ("13.7", "29.7.2", "on"))
self.assertEqual(g["containerd"], "2.3.3-1~debian.13~trixie")
def test_next_entry_wins_and_default_zero_is_the_newest(self):
f = Fake()
f.files["/etc/default/grub"] = "GRUB_DEFAULT=saved\n"
f.files["/boot/grub/grubenv"] = "saved_entry=gnulinux-advanced-x>gnulinux-7.0.2-6-pve-advanced-x\nnext_entry=gnulinux-advanced-x>gnulinux-7.0.14-20-pve-advanced-x\n"
_, fa = self.facts(f)
self.assertEqual(fa["host"]["kernel_next_boot"], "7.0.14-20-pve")
self.assertIn("next_entry", fa["host"]["kernel_next_boot_source"])
f = Fake()
f.files["/etc/default/grub"] = "GRUB_DEFAULT=0\n"
_, fa = self.facts(f)
self.assertEqual(fa["host"]["kernel_next_boot"], "7.0.14-20-pve", "dpkg order, not string order (7.0.2 < 7.0.14)")
def test_stopped_guest_is_unknown_never_guessed(self):
f = Fake()
f.status = "status: stopped"
_, fa = self.facts(f)
self.assertEqual(fa["guest"]["docker_engine"], "unknown")
self.assertIn("R10", fa["guest"]["unknown_reason"])
self.assertEqual(fa["host"]["tainted"], None, "unreadable -> None, not 0")
class RealSignatureCheck(unittest.TestCase):
"""The REAL Runner.verify_sig with the real ssh-keygen and a throwaway key, in the installer's allowed_signers form
(`<key_id> namespaces="felhom-op-v1" <type> <b64> <comment>`). Nothing leaves the temp dir."""
def setUp(self):
import shutil
if not shutil.which("ssh-keygen"):
self.skipTest("ssh-keygen not available")
import tempfile
self.d = tempfile.mkdtemp()
self.key = os.path.join(self.d, "k")
subprocess.run(["ssh-keygen", "-q", "-t", "ed25519", "-N", "", "-C", "felhom-op-1", "-f", self.key], check=True)
pub = open(self.key + ".pub").read().strip()
self.signers = os.path.join(self.d, "signers")
open(self.signers, "w").write(f'felhom-op-1 namespaces="felhom-op-v1" {pub}\n')
def sign(self, blob, ns="felhom-op-v1"):
bp = os.path.join(self.d, "blob")
open(bp, "wb").write(blob)
if os.path.exists(bp + ".sig"):
os.remove(bp + ".sig") # ssh-keygen -Y sign asks before overwriting (it would wait on stdin)
subprocess.run(["ssh-keygen", "-q", "-Y", "sign", "-f", self.key, "-n", ns, bp], check=True, stdin=subprocess.DEVNULL, timeout=30)
return open(bp + ".sig").read()
def test_good_signature_verifies(self):
blob = b'{"op":"os_docker_step"}'
self.assertEqual(osapply.Runner().verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob)), 0)
def test_changed_blob_wrong_namespace_or_wrong_principal_fail(self):
blob = b'{"op":"os_docker_step"}'
sig = self.sign(blob)
r = osapply.Runner()
self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob + b" ", sig), 0)
self.assertNotEqual(r.verify_sig(self.signers, "someone-else", "felhom-op-v1", blob, sig), 0)
self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob, ns="other-ns")), 0)
if __name__ == "__main__":
unittest.main()
+97 -31
View File
@@ -2,45 +2,111 @@ package hub
import (
"context"
"os/exec"
"fmt"
"strings"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// CloudflaredProber reports the cloudflared tunnel service health. It is a
// READ-ONLY probe: the agent does NOT manage or restart cloudflared in this slice
// (that is the tunnel-management slice — this is the seam for it). Injectable so
// tests use a fake and never exec.
// Tunnel states the agent reports (R-841, agent v0.141.0). THREE, never two: a probe that could not ask is
// `unknown`, which the hub never shows as up or down and never alarms on (R-96 rule 3).
const (
TunnelRunning = "running" // the cloudflared container runs AND its readiness check says CONNECTED
TunnelNotRunning = "not_running" // stopped / exited / absent, or running but NOT connected (Detail says which)
TunnelUnknown = "unknown" // the probe could not ask (guest down, pct/sudo error, health still starting)
)
// CloudflaredProber reports the box's tunnel. Injectable so tests use a fake and never exec.
type CloudflaredProber interface {
// Status returns one of: "active" | "inactive" | "failed" | "unknown".
Status(ctx context.Context) (string, error)
// Status returns one of the Tunnel* states and a short detail (why not_running / why unknown).
Status(ctx context.Context) (status, detail string)
}
// SystemctlProber runs `systemctl is-active cloudflared`. This is NOT a Privileged
// (root-CLI) op — `is-active` is non-root readable and is not one of the three
// proven root exceptions, so it does not go through internal/proxmox.Privileged.
type SystemctlProber struct {
Unit string // defaults to "cloudflared"
// GuestTunnelProber reads the REAL tunnel: the `cloudflared` container in the box's own customer guest.
//
// Before v0.141.0 the agent ran `systemctl is-active cloudflared` on the HOST — a unit that does not exist (cloudflared
// is a guest container, `11-os-updates.md` C8), so every box reported `inactive` (R-841).
//
// It uses ONLY the existing sudoers line `pct exec [0-9]* -- docker inspect -f *` (03 §3): the container's state, exit
// code and Docker health status. The health status comes from the compose health check controller v0.292.0 adds
// (`cloudflared tunnel --metrics localhost:20241 ready` → /ready: 200 only with ≥ 1 connection). Measured 2026-10-04:
// with a wrong token the container stays "running" while /ready answers 503 — so the container state alone would lie.
// A container with no health check (an older controller) is judged on its state alone, and Detail says so.
type GuestTunnelProber struct {
Runner proxmox.Runner
// Guests returns the box's customer guest vmids (running pool guests that bind /mnt/felhom-drives).
Guests func(ctx context.Context) ([]int, error)
}
// Status maps `systemctl is-active` output to the report vocabulary. systemctl
// exits non-zero for inactive/failed, so the output string is authoritative over
// the exit code; any exec error (binary missing, etc.) maps to "unknown".
func (p SystemctlProber) Status(ctx context.Context) (string, error) {
unit := p.Unit
if unit == "" {
unit = "cloudflared"
const tunnelInspect = `{{.State.Status}}|{{.State.ExitCode}}|{{if .State.Health}}{{.State.Health.Status}}{{else}}none{{end}}`
// Status probes every customer guest and reports the worst state (normally there is exactly one guest).
func (p GuestTunnelProber) Status(ctx context.Context) (string, string) {
if p.Runner == nil || p.Guests == nil {
return TunnelUnknown, "no probe wired"
}
out, _ := exec.CommandContext(ctx, "systemctl", "is-active", unit).Output()
switch strings.TrimSpace(string(out)) {
case "active":
return "active", nil
case "failed":
return "failed", nil
case "inactive", "deactivating", "activating":
return "inactive", nil
case "":
return "unknown", nil // no output → systemctl/exec problem
default:
return "unknown", nil
vmids, err := p.Guests(ctx)
if err != nil {
return TunnelUnknown, "could not list the customer guest: " + err.Error()
}
if len(vmids) == 0 {
return TunnelUnknown, "no running customer guest"
}
worst, wdetail := "", ""
rank := map[string]int{TunnelRunning: 0, TunnelUnknown: 1, TunnelNotRunning: 2}
for _, v := range vmids {
out, errOut, err := p.Runner.Run(ctx, "/usr/sbin/pct", "exec", fmt.Sprint(v), "--", "docker", "inspect", "-f", tunnelInspect, "cloudflared")
st, d := ClassifyTunnel(string(out), string(errOut), err)
if len(vmids) > 1 {
d = fmt.Sprintf("guest %d: %s", v, d)
}
if worst == "" || rank[st] > rank[worst] {
worst, wdetail = st, d
}
}
return worst, wdetail
}
// ClassifyTunnel maps one `docker inspect` answer to a state. Pure; pinned by TestClassifyTunnel.
func ClassifyTunnel(stdout, stderr string, err error) (string, string) {
out := strings.TrimSpace(stdout)
if err != nil || out == "" {
if strings.Contains(stderr, "No such object") || strings.Contains(stderr, "No such container") {
return TunnelNotRunning, "no cloudflared container in the guest"
}
return TunnelUnknown, "could not ask the guest: " + firstLine(stderr, err)
}
parts := strings.Split(out, "|")
if len(parts) != 3 {
return TunnelUnknown, "unreadable docker answer: " + out
}
state, code, health := parts[0], parts[1], parts[2]
if state != "running" {
return TunnelNotRunning, fmt.Sprintf("container %s, exit code %s", state, code)
}
switch health {
case "healthy":
return TunnelRunning, "connected"
case "unhealthy":
return TunnelNotRunning, "container running but the tunnel is NOT connected (cloudflared /ready fails)"
case "starting":
return TunnelUnknown, "container running, readiness check still starting"
case "none":
return TunnelRunning, "container running (no readiness check on this controller — connection not checked)"
}
return TunnelUnknown, "unknown health state " + health
}
func firstLine(stderr string, err error) string {
s := strings.TrimSpace(stderr)
if i := strings.IndexByte(s, '\n'); i >= 0 {
s = s[:i]
}
if s == "" && err != nil {
s = err.Error()
}
if len(s) > 160 {
s = s[:160]
}
return s
}
+67
View File
@@ -0,0 +1,67 @@
package hub
import (
"context"
"errors"
"io"
"strings"
"testing"
)
// R-841: the three states from one `docker inspect` answer. Red-proof: map "unhealthy" to running (the container
// state alone — what a plain "is it running" probe would say) and the "running but not connected" case fails.
func TestClassifyTunnel(t *testing.T) {
cases := []struct {
name, out, errOut string
err error
want string
detail string
}{
{"connected", "running|0|healthy\n", "", nil, TunnelRunning, "connected"},
{"running but not connected", "running|0|unhealthy\n", "", nil, TunnelNotRunning, "NOT connected"},
{"stopped", "exited|137|unhealthy\n", "", nil, TunnelNotRunning, "exit code 137"},
{"absent", "", "Error: No such object: cloudflared", errors.New("exit status 1"), TunnelNotRunning, "no cloudflared container"},
{"still starting", "running|0|starting\n", "", nil, TunnelUnknown, "starting"},
{"no health check (older controller)", "running|0|none\n", "", nil, TunnelRunning, "connection not checked"},
{"guest not running", "", "CT 9201 not running", errors.New("exit status 255"), TunnelUnknown, "could not ask"},
{"sudo refused", "", "sudo: a password is required", errors.New("exit status 1"), TunnelUnknown, "could not ask"},
}
for _, c := range cases {
st, d := ClassifyTunnel(c.out, c.errOut, c.err)
if st != c.want || !strings.Contains(d, c.detail) {
t.Errorf("%s: got %q (%s), want %q (…%s…)", c.name, st, d, c.want, c.detail)
}
}
}
type tunnelRunner struct {
calls []string
out map[string]string
}
func (r *tunnelRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
line := name + " " + strings.Join(args, " ")
r.calls = append(r.calls, line)
return []byte(r.out[args[1]]), nil, nil
}
func (r *tunnelRunner) RunStdin(ctx context.Context, _ io.Reader, name string, args ...string) ([]byte, []byte, error) {
return r.Run(ctx, name, args...)
}
// The probe uses EXACTLY the existing sudoers shape `pct exec <vmid> -- docker inspect -f <tmpl> cloudflared`, and
// with no customer guest it is unknown, never down.
func TestGuestTunnelProber(t *testing.T) {
r := &tunnelRunner{out: map[string]string{"9201": "running|0|healthy"}}
p := GuestTunnelProber{Runner: r, Guests: func(context.Context) ([]int, error) { return []int{9201}, nil }}
if st, d := p.Status(context.Background()); st != TunnelRunning || d != "connected" {
t.Fatalf("got %q %q", st, d)
}
want := "/usr/sbin/pct exec 9201 -- docker inspect -f " + tunnelInspect + " cloudflared"
if len(r.calls) != 1 || r.calls[0] != want {
t.Fatalf("command = %q, want %q", r.calls, want)
}
none := GuestTunnelProber{Runner: r, Guests: func(context.Context) ([]int, error) { return nil, nil }}
if st, _ := none.Status(context.Background()); st != TunnelUnknown {
t.Fatalf("no guest → %q, want unknown", st)
}
}
+49 -8
View File
@@ -4,10 +4,12 @@ import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"log/slog"
"os"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
@@ -108,6 +110,7 @@ type Collector struct {
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
oob OOBReporter // H1: operator-access health (nil → stanza omitted)
system SystemReporter // R-852: the box versions (nil → API fields only)
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
hostID string
agentVersion string
@@ -248,6 +251,41 @@ type OOBReporter interface {
}
// SetOOBReporter wires the operator-access health source (H1; nil-safe → stanza omitted).
// SystemReporter reads the box's versions (R-852): the customer guest's vmid and the wrapper's raw facts.
type SystemReporter interface {
SystemFacts(ctx context.Context) (vmid int, facts json.RawMessage, err error)
}
// SetSystemReporter wires the facts read (agent v0.142.0). Without it the stanza carries the Proxmox API fields only.
func (c *Collector) SetSystemReporter(r SystemReporter) *Collector {
c.system = r
return c
}
func unknownIfEmpty(s string) string {
if strings.TrimSpace(s) == "" {
return "unknown"
}
return s
}
// systemInfo builds the `system` stanza. Never fatal: a failed facts read is FactsError, the API fields stay.
func (c *Collector) systemInfo(ctx context.Context, ns proxmox.NodeStatus) *SystemInfo {
si := &SystemInfo{PVEVersion: unknownIfEmpty(ns.PVEVersion), KernelVersion: unknownIfEmpty(ns.KVersion),
ReadAt: c.now().Format(time.RFC3339)}
if c.system == nil {
si.FactsError = "no facts reader wired"
return si
}
vmid, f, err := c.system.SystemFacts(ctx)
si.VMID, si.Facts = vmid, f
if err != nil {
si.FactsError = err.Error()
c.logger.Debug("hub: system facts unavailable", "err", err)
}
return si
}
func (c *Collector) SetOOBReporter(o OOBReporter) *Collector {
c.oob = o
return c
@@ -280,10 +318,11 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
PBSSnapshots: c.collectPBSSnapshots(ctx),
AuditTail: []AuditEntry{},
Cloudflared: Cloudflared{Status: c.cloudflaredStatus(ctx)},
Cloudflared: c.cloudflared(ctx),
Capabilities: c.capabilities(ctx),
LeafFingerprint: c.leafFP,
Addresses: c.collectAddresses(),
System: c.systemInfo(ctx, ns),
}
// DR recipe host-half — derived from the just-collected guest/storage/PBS facts (no new reads).
// Secret-free by construction (identifiers/intents/sizes/coordinates only).
@@ -555,16 +594,18 @@ func (c *Collector) collectPBSSnapshots(ctx context.Context) []PBSSnapshot {
return []PBSSnapshot{}
}
func (c *Collector) cloudflaredStatus(ctx context.Context) string {
func (c *Collector) cloudflared(ctx context.Context) Cloudflared {
if c.cf == nil {
return "unknown"
return Cloudflared{Status: TunnelUnknown, Detail: "no probe wired"}
}
st, err := c.cf.Status(ctx)
if err != nil || st == "" {
c.logger.Warn("hub: cloudflared probe failed", "err", err)
return "unknown"
st, d := c.cf.Status(ctx)
if st == "" {
st = TunnelUnknown
}
return st
if st == TunnelUnknown {
c.logger.Debug("hub: tunnel probe could not decide", "detail", d)
}
return Cloudflared{Status: st, Detail: d}
}
func percent(used, total int64) float64 {
+2 -2
View File
@@ -16,7 +16,7 @@ func (f fakeGuestNet) GuestNetStatus(context.Context) *GuestNetStatus { return f
func TestCollect_GuestNetOmittedWhenReporterNil(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.92.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.92.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
@@ -38,7 +38,7 @@ func TestCollect_GuestNetOmittedWhenReporterNil(t *testing.T) {
func TestCollect_GuestNetPopulatedWhenWired(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.92.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.92.0", quietLogger())
c.SetGuestNetReporter(fakeGuestNet{st: &GuestNetStatus{
CheckedAt: "2026-07-21T10:00:00Z",
Guests: []GuestNetGuest{{
+4 -4
View File
@@ -12,7 +12,7 @@ func (f fakeMgmtPlane) MgmtPlaneStatus(context.Context) *MgmtPlaneStatus { retur
func TestCollect_MgmtPlaneOmittedWhenReporterNil(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.71.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.71.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
@@ -24,7 +24,7 @@ func TestCollect_MgmtPlaneOmittedWhenReporterNil(t *testing.T) {
func TestCollect_MgmtPlanePopulatedWhenWired(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.71.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.71.0", quietLogger())
c.SetMgmtPlaneReporter(fakeMgmtPlane{st: &MgmtPlaneStatus{
PrivsepDirOK: true, SshdReachable: true, HealedRecently: true, PrivsepHealedAt: "2026-07-05T16:42:17Z",
}})
@@ -47,7 +47,7 @@ func (f fakeOOB) OOBStatus(context.Context) *OOBStatus { return f.st }
func TestCollect_OOBOmittedWhenNil(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.72.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.72.0", quietLogger())
r, _ := c.Collect(context.Background())
if r.OOB != nil {
t.Fatalf("no reporter → oob omitted, got %+v", r.OOB)
@@ -56,7 +56,7 @@ func TestCollect_OOBOmittedWhenNil(t *testing.T) {
func TestCollect_OOBPopulatedWhenWired(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{}, nil, nil, nil, "h", "0.72.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{}, nil, nil, nil, "h", "0.72.0", quietLogger())
c.SetOOBReporter(fakeOOB{st: &OOBStatus{FelhomSshdActive: true, FelhomSshdPort: 8822, Reachable: true}})
r, _ := c.Collect(context.Background())
if r.OOB == nil || r.OOB.FelhomSshdPort != 8822 || !r.OOB.Reachable {
+9 -9
View File
@@ -33,7 +33,7 @@ func TestCollect_StorageTargetsFromObserver(t *testing.T) {
obs := fakeObserver{targets: []StorageTarget{
{Name: "local-lvm", Type: StorageTypeLVMThin, State: StorageStateAttached, Reachable: true},
}}
c := NewCollector(px, fakeProber{status: "active"}, obs, nil, nil, nil, "h", "0.5.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, obs, nil, nil, nil, "h", "0.5.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
@@ -45,7 +45,7 @@ func TestCollect_StorageTargetsFromObserver(t *testing.T) {
func TestCollect_StorageObserverErrorDegradesToEmpty(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{err: errors.New("proxmox down")}, nil, nil, nil, "h", "0.5.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{err: errors.New("proxmox down")}, nil, nil, nil, "h", "0.5.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("a storage observe error must not sink the heartbeat: %v", err)
@@ -64,7 +64,7 @@ func TestCollect_HostAndGuests(t *testing.T) {
},
cfg: map[int]proxmox.GuestConfig{100: {Cores: 2, Memory: 2048}},
}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "demo-host-01", "0.3.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "demo-host-01", "0.3.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
@@ -88,7 +88,7 @@ func TestCollect_HostAndGuests(t *testing.T) {
if g.Spec.Cores != 2 || g.Spec.MemoryBytes != 2147483648 || g.Spec.DiskBytes != 21474836480 {
t.Errorf("spec = %+v", g.Spec)
}
if r.Cloudflared.Status != "active" {
if r.Cloudflared.Status != "running" || r.Cloudflared.Detail != "connected" {
t.Errorf("cloudflared = %q", r.Cloudflared.Status)
}
}
@@ -104,7 +104,7 @@ func TestCollect_GuestConfigFailureKeepsStatusOmitsSpec(t *testing.T) {
cfg: map[int]proxmox.GuestConfig{100: {Cores: 2}},
cfgErr: map[int]error{200: errors.New("config read failed")},
}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.3.1", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.3.1", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("a per-guest failure must NOT fail the whole report: %v", err)
@@ -125,7 +125,7 @@ func TestCollect_GuestConfigFailureKeepsStatusOmitsSpec(t *testing.T) {
func TestCollect_NodeStatusFailureIsHardError(t *testing.T) {
px := &fakePx{node: "n", nsErr: errors.New("proxmox down")}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
if _, err := c.Collect(context.Background()); err == nil {
t.Fatal("NodeStatus failure must be a hard error (no useful report)")
}
@@ -133,7 +133,7 @@ func TestCollect_NodeStatusFailureIsHardError(t *testing.T) {
func TestCollect_CloudflaredProbeErrorIsUnknown(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{err: errors.New("no systemctl")}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
c := NewCollector(px, fakeProber{status: "", detail: "could not ask"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("cloudflared failure must not be fatal: %v", err)
@@ -153,7 +153,7 @@ func TestCollect_LeafFingerprint(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
const fp = "60b5974d586f5f3c8ec41eb998d0f07406178219c36bf6d3ff377570279d8245"
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.48.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.48.0", quietLogger())
c.SetLeafFingerprint(fp)
r, err := c.Collect(context.Background())
if err != nil {
@@ -164,7 +164,7 @@ func TestCollect_LeafFingerprint(t *testing.T) {
}
// Companion: no SetLeafFingerprint (local API disabled) → empty, never a fabricated value.
c2 := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.48.0", quietLogger())
c2 := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.48.0", quietLogger())
r2, _ := c2.Collect(context.Background())
if r2.LeafFingerprint != "" {
t.Fatalf("unset leaf_fingerprint = %q, want empty", r2.LeafFingerprint)
+2 -2
View File
@@ -384,7 +384,7 @@ func TestCollectDRRecipe_ProductionPath(t *testing.T) {
obs := fakeObserver{targets: capturedDemoFelhomTargets()}
pbsRep := fakePBSReporter{snaps: capturedDemoFelhomSnapshots()}
c := NewCollector(px, fakeProber{status: "active"}, obs, nil, nil, pbsRep, "h", "0.118.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, obs, nil, nil, pbsRep, "h", "0.118.0", quietLogger())
c.SetBackupTargetResolver(func() ConfiguredBackupTarget {
return ConfiguredBackupTarget{StorageID: "felhom-backup", Known: true}
})
@@ -409,7 +409,7 @@ func TestCollectDRRecipe_ProductionPath(t *testing.T) {
// test that would have caught shipping the seam without wiring it.
func TestCollectDRRecipe_UnwiredSeamReportsUnknown(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, fakeObserver{targets: capturedDemoFelhomTargets()},
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, fakeObserver{targets: capturedDemoFelhomTargets()},
nil, nil, nil, "h", "0.118.0", quietLogger())
r, err := c.Collect(context.Background())
+4 -4
View File
@@ -16,7 +16,7 @@ func intp(v int) *int { return &v }
// HostMetricsNow returns a fresh host block with cpu% from NodeStatus and the temp from the reader.
func TestHostMetricsNow_PopulatesTemp(t *testing.T) {
px := &fakePx{node: "demo-felhom", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
SetTempReader(fakeTemp{c: intp(46)})
h, err := c.HostMetricsNow(context.Background())
if err != nil {
@@ -36,7 +36,7 @@ func TestHostMetricsNow_PopulatesTemp(t *testing.T) {
// A missing temp sensor gracefully nulls cpu_temp_c without failing the host read.
func TestHostMetricsNow_GracefulNullTemp(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
SetTempReader(fakeTemp{c: nil})
h, err := c.HostMetricsNow(context.Background())
if err != nil {
@@ -50,7 +50,7 @@ func TestHostMetricsNow_GracefulNullTemp(t *testing.T) {
// A NodeStatus failure is a hard error (no useful host view).
func TestHostMetricsNow_NodeStatusErrorIsHard(t *testing.T) {
px := &fakePx{node: "n", nsErr: errors.New("proxmox down")}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger())
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger())
if _, err := c.HostMetricsNow(context.Background()); err == nil {
t.Fatal("NodeStatus failure must be a hard error")
}
@@ -59,7 +59,7 @@ func TestHostMetricsNow_NodeStatusErrorIsHard(t *testing.T) {
// Collect() (the hub report) also carries the temp now — the operator freebie.
func TestCollect_HostReportCarriesTemp(t *testing.T) {
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "active"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
c := NewCollector(px, fakeProber{status: "running", detail: "connected"}, nil, nil, nil, nil, "h", "0.14.0", quietLogger()).
SetTempReader(fakeTemp{c: intp(51)})
r, err := c.Collect(context.Background())
if err != nil {
+2 -2
View File
@@ -56,7 +56,7 @@ func (f *fakePx) GuestConfig(ctx context.Context, vmid int) (proxmox.GuestConfig
// fakeProber is a fake CloudflaredProber.
type fakeProber struct {
status string
err error
detail string
}
func (p fakeProber) Status(ctx context.Context) (string, error) { return p.status, p.err }
func (p fakeProber) Status(ctx context.Context) (string, string) { return p.status, p.detail }
+7 -1
View File
@@ -23,8 +23,14 @@ func TestOSUpdateGolden_Decodes(t *testing.T) {
t.Fatalf("os_update = %+v", o)
}
r := o.Release
if r.ID != "os-20261004-120000" || r.Snapshot != "20261004T120000Z" || len(r.Packages) != 2 ||
if r.ID != "os-guest-20261004-120000" || r.Snapshot != "20261004T120000Z" || len(r.Packages) != 2 ||
r.Packages[1].Name != "openssl" || r.Packages[1].Version != "3.5.7-1~deb13u3" || r.Packages[1].Origin != "Debian-Security" {
t.Fatalf("release = %+v", r)
}
// v0.141.0: the host layer's own approved set (`11` §8 step 3).
h := o.HostRelease
if h == nil || h.ID != "os-host-20261004-120000" || h.Snapshot != "20261004T120000Z" || len(h.Packages) != 1 ||
h.Packages[0].Name != "libssl3t64" || h.Packages[0].Origin != "Debian-Security" {
t.Fatalf("host_release = %+v", h)
}
}
+22 -2
View File
@@ -89,6 +89,12 @@ type HostReport struct {
// hub-schema change and are absent when the reporter is not wired.
MgmtPlane *MgmtPlaneStatus `json:"mgmt_plane,omitempty"`
// System is the box's versions for the hub's System page (agent v0.142.0, R-852, `09` decision 89): Proxmox and the
// running kernel from the Proxmox API, and the wrapper's read-only facts (host Debian, next-boot kernel, held
// packages, taint, the crash guard; guest Debian, Docker engine, containerd, live-restore). A value nobody could
// read is "unknown", never empty and never guessed. The hub v0.132.0 consumes it (hosts + System pages).
System *SystemInfo `json:"system,omitempty"`
// PBSDR is the PBS-DR-tier bridge status stanza (slice 2). Present only when the pbsdr
// consumer is wired. `consumed_failed` is the LOUD persistent state: the one-time token
// secret was consumed but the apply failed afterwards — the secret is burned, the bridge
@@ -241,6 +247,16 @@ type WireguardStatus struct {
AssignedIP string `json:"assigned_ip,omitempty"` // from the marker, e.g. "10.77.0.2/32"
}
// SystemInfo is the `system` stanza (see HostReport.System).
type SystemInfo struct {
PVEVersion string `json:"pve_version"` // GET /nodes/{node}/status pveversion
KernelVersion string `json:"kernel_version"` // GET /nodes/{node}/status kversion
VMID int `json:"vmid,omitempty"` // the customer guest the facts read
Facts json.RawMessage `json:"facts,omitempty"`
FactsError string `json:"facts_error,omitempty"`
ReadAt string `json:"read_at"`
}
// HostMetrics is the host block, sourced from proxmox NodeStatus.
type HostMetrics struct {
Node string `json:"node"`
@@ -289,9 +305,10 @@ type GuestSpec struct {
DiskBytes int64 `json:"disk_bytes"`
}
// Cloudflared is the tunnel service health (read-only probe this slice).
// Cloudflared is the box's tunnel (R-841, agent v0.141.0): the cloudflared container in the customer guest.
type Cloudflared struct {
Status string `json:"status"` // active | inactive | failed | unknown
Status string `json:"status"` // running | not_running | unknown (TunnelRunning …)
Detail string `json:"detail,omitempty"` // why not_running / unknown, or "connected"
}
// The following element types are declared now so the empty collections above are
@@ -559,6 +576,9 @@ type WireOSUpdate struct {
Ring int `json:"ring"`
Enabled bool `json:"enabled"`
Release *WireOSRelease `json:"release,omitempty"`
// HostRelease is the newest approved HOST release (hub v0.131.0, `11` §8 step 3) — a separate set: a version
// approved for the guest is not approved for the host by that fact alone.
HostRelease *WireOSRelease `json:"host_release,omitempty"`
}
// WireOSRelease is an approved version set; Snapshot is the approval time (YYYYMMDDTHHMMSSZ) the wrapper uses
+8 -1
View File
@@ -5,12 +5,19 @@
"ring": 1,
"enabled": true,
"release": {
"id": "os-20261004-120000",
"id": "os-guest-20261004-120000",
"snapshot": "20261004T120000Z",
"packages": [
{"name": "libc6", "version": "2.41-12+deb13u4", "origin": "Debian"},
{"name": "openssl", "version": "3.5.7-1~deb13u3", "origin": "Debian-Security"}
]
},
"host_release": {
"id": "os-host-20261004-120000",
"snapshot": "20261004T120000Z",
"packages": [
{"name": "libssl3t64", "version": "3.5.7-1~deb13u3", "origin": "Debian-Security"}
]
}
}
}
+88
View File
@@ -0,0 +1,88 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"fmt"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// OpDockerStep is the signed op class of a Docker engine step (`11` §5.8): a ring-1 box takes an approved engine set,
// and every box takes an UNDO, only through it. CC may sign it until the first paying customer (R-530 ruling).
const OpDockerStep = "os_docker_step"
// DockerStepParams are the signed params. The wrapper compares Packages and Undo with the plan byte-for-byte.
type DockerStepParams struct {
ReleaseID string `json:"release_id"`
Packages []Package `json:"packages"`
Undo bool `json:"undo"`
VMID int `json:"vmid,omitempty"`
}
// DockerStepExecutor runs a verified os_docker_step (signedjobs.Executor). Guest finds the box's customer guest when
// the params name none; Gate (optional) takes the host-wide heavy-op gate so a step never runs beside a backup.
type DockerStepExecutor struct {
Leg *Leg
Guest func(ctx context.Context) (int, error)
Gate func(ctx context.Context) (release func(), err error)
}
// Execute implements signedjobs.Executor.
func (e DockerStepExecutor) Execute(ctx context.Context, op string, params json.RawMessage) error {
if op != OpDockerStep {
return signedjobs.ErrNoExecutor
}
so, ok := signedjobs.SignedOpFrom(ctx)
if !ok {
return fmt.Errorf("os_docker_step: no signed envelope in the context — the wrapper could not verify it")
}
var p DockerStepParams
if err := json.Unmarshal(params, &p); err != nil || len(p.Packages) == 0 {
return fmt.Errorf("os_docker_step: params must name the engine set: %v", err)
}
vmid := p.VMID
if vmid == 0 {
if e.Guest == nil {
return fmt.Errorf("os_docker_step: no vmid and no guest finder")
}
v, err := e.Guest(ctx)
if err != nil {
return fmt.Errorf("os_docker_step: find the customer guest: %w", err)
}
vmid = v
}
if e.Gate != nil {
release, err := e.Gate(ctx)
if err != nil {
return fmt.Errorf("os_docker_step: heavy-op gate busy (a backup or restore-test runs): %w", err)
}
defer release()
}
rep := e.Leg.RunDockerSigned(ctx, vmid, p, so.Blob, string(so.Sig))
switch rep.Outcome {
case "applied", "nothing":
if rep.Healthy {
return nil
}
}
return fmt.Errorf("os_docker_step: %s (%s) %s", rep.Outcome, rep.HealthReason, string(rep.Refused))
}
// RunDockerSigned is one signed Docker step (ring 1 or an undo): live-restore first (decision 87, a no-op when on),
// then the docker layer with the signed envelope, which the wrapper verifies itself.
func (l *Leg) RunDockerSigned(ctx context.Context, vmid int, p DockerStepParams, blob []byte, sig string) Report {
runID := l.now().UTC().Format("20060102T150405Z")
lg := l.log().With("run", runID, "vmid", vmid, "trigger", "signed", "release", p.ReleaseID, "undo", p.Undo)
if err := l.EnsureLiveRestore(ctx, runID, vmid); err != nil {
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerDocker, Trigger: "signed", Ring: l.Block().Ring, VMID: vmid,
Mode: "apply", ReleaseID: p.ReleaseID, Outcome: "failed", HealthReason: "live-restore could not be turned on: " + err.Error()})
}
rid := p.ReleaseID
if rid == "" {
rid = "signed-" + runID
}
return l.runLayer(ctx, runID, LayerDocker, vmid, "signed", l.Block(), dockerOpts{releaseID: rid, packages: p.Packages,
undo: p.Undo, signed: map[string]string{"blob_b64": base64.StdEncoding.EncodeToString(blob), "sig": sig}})
}
+49
View File
@@ -0,0 +1,49 @@
package osupdate
import (
"context"
"encoding/base64"
"encoding/json"
"errors"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
)
// The executor hands the RAW signed bytes to the wrapper (which verifies them itself) and the exact signed package
// list. Red-proof: drop the `signed` field from the docker plan in runLayer and the plan check fails.
func TestDockerStepExecutor_PassesTheSignedEnvelope(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}}, DockerEngine: "29.8.2", Authority: "signed"}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
e := DockerStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
params, _ := json.Marshal(DockerStepParams{ReleaseID: "os-docker-1", Packages: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie", Origin: "Docker CE"}}})
ctx := signedjobs.WithSignedOp(context.Background(), &reconcile.SignedOp{Blob: []byte(`{"op":"os_docker_step"}`), Sig: []byte("SIG")})
if err := e.Execute(ctx, OpDockerStep, params); err != nil {
t.Fatal(err)
}
dp := w.plans[len(w.plans)-1]
sg, _ := dp["signed"].(map[string]any)
if dp["layer"] != "docker" || dp["lane"] != "slow" || dp["release_id"] != "os-docker-1" || sg == nil ||
sg["blob_b64"] != base64.StdEncoding.EncodeToString([]byte(`{"op":"os_docker_step"}`)) || sg["sig"] != "SIG" {
t.Fatalf("docker plan = %v", dp)
}
if calls(w) != "guest:live-restore-on,docker:apply" || len(h.reports) != 1 || h.reports[0].Trigger != "signed" {
t.Fatalf("calls=%s reports=%+v", calls(w), h.reports)
}
}
func TestDockerStepExecutor_RefusesWithoutEnvelopeAndPassesOtherOps(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
e := DockerStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
if err := e.Execute(context.Background(), "agent_update", nil); !errors.Is(err, signedjobs.ErrNoExecutor) {
t.Fatalf("another op must pass through the chain: %v", err)
}
params, _ := json.Marshal(DockerStepParams{Packages: []Package{{Name: "docker-ce", Version: "1"}}})
if err := e.Execute(context.Background(), OpDockerStep, params); err == nil || len(w.plans) != 0 {
t.Fatalf("no envelope must refuse before any wrapper call: err=%v calls=%s", err, calls(w))
}
}
+351 -126
View File
@@ -1,14 +1,16 @@
// Package osupdate is the agent's OS-update leg for the customer GUEST (`11-os-updates.md` §8 step 2, agent v0.140.0).
// Package osupdate is the agent's OS-update leg (`11-os-updates.md` §8 steps 2–3, §5.8): the customer GUEST's Debian
// fast lane (agent v0.140.0), after it in the same pass the HOST's (agent v0.141.0), and then — ring 0 only — the
// guest's DOCKER engine set, the slow lane (agent v0.142.0; a ring-1 box takes a Docker step only inside a signed
// operator job, DockerStepExecutor). It also reads the box's versions for the hub's System page (Facts, R-852).
//
// It runs right after the night's successful whole-guest backup, while the backup goroutine still holds the host-wide
// heavy-op gate (so it never overlaps a backup or a restore-test, `11` C10), at most once per night. All root work is
// the wrapper `felhom-os-apply` (configs/, its own tests); this package only builds plans, calls the wrapper through
// sudo, judges health and reports to the hub.
// sudo, judges health and reports to the hub — one report per layer.
//
// NO AUTOMATIC UNDO in this release (R-837, measured 2026-10-04): a customer guest cannot be snapshotted — PVE
// refuses any snapshot not named `vzdump` when the guest has host-path binds (mp8, mp9). A failed health check
// therefore stops, reports `health_failed` and the hub mails the operator; the whole-guest backup taken minutes
// earlier is the undo, by hand.
// NO AUTOMATIC UNDO (R-837 measured; `09` §3 decision 81): a failed health check stops, reports `health_failed` and the
// hub mails the operator; the whole-guest backup taken minutes earlier is the guest's undo, by hand; a host package is
// put back by hand from the previous release's snapshot (runbook). The host is NEVER rebooted by this package.
package osupdate
import (
@@ -33,6 +35,17 @@ const WrapperPath = "/usr/local/sbin/felhom-os-apply"
// DefaultPlanDir is where plans are written (the sudoers glob names it).
const DefaultPlanDir = "/var/lib/felhom-agent/os"
// Layers.
const (
LayerGuest = "guest"
LayerHost = "host"
LayerDocker = "docker" // the guest's Docker engine set — slow lane (`11` §5.8)
)
// DockerNames are the six packages of the Docker engine set (the wrapper's DOCKER_NAMES).
var DockerNames = map[string]bool{"containerd.io": true, "docker-buildx-plugin": true, "docker-ce": true,
"docker-ce-cli": true, "docker-ce-rootless-extras": true, "docker-compose-plugin": true}
// Package is one name=version with its origin.
type Package struct {
Name string `json:"name"`
@@ -40,7 +53,7 @@ type Package struct {
Origin string `json:"origin"`
}
// Pending is one update the guest's sources offer (origin as apt names it, possibly several).
// Pending is one update the sources offer (origin as apt names it, possibly several).
type Pending struct {
Name string `json:"name"`
From string `json:"from"`
@@ -52,19 +65,25 @@ type Pending struct {
type Container struct {
State string `json:"state"`
Health string `json:"health"` // healthy | unhealthy | starting | none
ID string `json:"id,omitempty"`
}
// Health is the guest's health snapshot (the wrapper's `health` object).
// Health is one health reading. Guest layer: DockerOK..Containers. Host layer: HostServices, GuestRunning and the
// guest's own reading in Guest.
type Health struct {
DockerOK bool `json:"docker_ok"`
NetworkOK bool `json:"network_ok"`
Controller string `json:"controller"`
Containers map[string]Container `json:"containers"`
DockerOK bool `json:"docker_ok"`
NetworkOK bool `json:"network_ok"`
Controller string `json:"controller"`
Containers map[string]Container `json:"containers"`
HostServices map[string]string `json:"host_services,omitempty"`
GuestRunning *bool `json:"guest_running,omitempty"`
Guest *Health `json:"guest,omitempty"`
}
// WrapperReport is the wrapper's OSAPPLY-REPORT object.
type WrapperReport struct {
Mode string `json:"mode"`
Layer string `json:"layer"`
Refused json.RawMessage `json:"refused"`
Failed json.RawMessage `json:"failed"`
Upgraded []Package `json:"upgraded"`
@@ -73,17 +92,25 @@ type WrapperReport struct {
RestartNeeded []string `json:"restart_needed"`
DockerRestartNeeded bool `json:"docker_restart_needed"`
RebootNeeded bool `json:"reboot_needed"`
RebootScanned bool `json:"reboot_scanned"`
HealthBefore *Health `json:"health_before"`
HealthAfter *Health `json:"health_after"`
Health *Health `json:"health"`
PassSeconds float64 `json:"pass_seconds"`
DockerEngine string `json:"docker_engine"`
Authority string `json:"authority"`
Undo bool `json:"undo"`
LiveRestore json.RawMessage `json:"live_restore"`
Facts json.RawMessage `json:"facts"`
}
func (w WrapperReport) refused() bool { return len(w.Refused) > 0 && string(w.Refused) != "null" }
func (w WrapperReport) failed() bool { return len(w.Failed) > 0 && string(w.Failed) != "null" }
// Report is what the hub receives (hub osupdates.Report — field-exact).
// Report is what the hub receives per layer (hub osupdates.Report — field-exact).
type Report struct {
RunID string `json:"run_id"`
Layer string `json:"layer"`
Trigger string `json:"trigger"`
Mode string `json:"mode"`
Ring int `json:"ring"`
@@ -99,7 +126,12 @@ type Report struct {
RestartNeeded []string `json:"restart_needed,omitempty"`
DockerRestartNeeded bool `json:"docker_restart_needed,omitempty"`
RebootNeeded bool `json:"reboot_needed,omitempty"`
RebootScanned bool `json:"reboot_scanned,omitempty"` // the pass looked (host: every pass) — a false RebootNeeded then means "not needed"
Refused json.RawMessage `json:"refused,omitempty"`
PassSeconds float64 `json:"pass_seconds,omitempty"`
DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step
Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed
Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade)
}
// Reporter posts a report to the hub (*hub.Client).
@@ -107,10 +139,12 @@ type Reporter interface {
PostOSReport(ctx context.Context, body []byte) error
}
// Leg runs one OS-update pass for the customer guest.
// Leg runs one OS-update pass for the customer guest and then the host.
type Leg struct {
Runner proxmox.Runner
Hub Reporter
Tunnel hub.CloudflaredProber // the host health rule needs the tunnel `running` (R-841)
Appliance bool // agent.json deployment_mode; the wrapper re-checks the ROOT-owned record (R12)
Logger *slog.Logger
PlanDir string
StatePath string // last night run (once per night)
@@ -122,7 +156,6 @@ type Leg struct {
mu sync.Mutex
block *hub.WireOSUpdate
have bool
}
// OnDesiredState stores the hub's os_update block (desired.RawConsumer — store only, never block).
@@ -132,7 +165,7 @@ func (l *Leg) OnDesiredState(_ context.Context, resp *hub.DesiredStateResponse)
}
l.mu.Lock()
defer l.mu.Unlock()
l.block, l.have = resp.DesiredState.OSUpdate, true
l.block = resp.DesiredState.OSUpdate
}
// Block returns the newest os_update block. No block (an older hub, or nothing fetched yet) = ring 1, ON, no
@@ -150,7 +183,7 @@ func (l *Leg) Block() hub.WireOSUpdate {
func (l *Leg) SetBlock(b *hub.WireOSUpdate) {
l.mu.Lock()
defer l.mu.Unlock()
l.block, l.have = b, true
l.block = b
}
func (l *Leg) now() time.Time {
@@ -191,18 +224,9 @@ func IsFast(origins []string) bool {
return true
}
func fastOrigin(origins []string) string {
for _, o := range origins {
if o == "Debian-Security" {
return o
}
}
return "Debian"
}
// HealthVerdict is THE health rule (`11` §5.4.1; pinned by TestHealthVerdict_*): after the run, docker answers, the
// guest's network resolves, the controller's own health check is `healthy`, and every container that was running
// before is running again — and healthy again if it was healthy before. "starting" is not yet healthy.
// HealthVerdict is THE guest health rule (`11` §8.1; pinned by TestHealthVerdict*): docker answers, the guest's
// network resolves, the controller's own health check is `healthy`, and every container that was running at the
// start of the pass runs again — and healthy again if it was. "starting" is not yet healthy.
func HealthVerdict(before, after *Health) (bool, string) {
if after == nil {
return false, "no health reading"
@@ -240,34 +264,87 @@ func HealthVerdict(before, after *Health) (bool, string) {
return true, ""
}
// MergeBaseline is the health baseline from two readings: a container counts as running (and healthy) if EITHER
// reading saw it so. nil-safe. Pinned by TestHealth_BaselineIsTheStartOfTheLeg.
func MergeBaseline(a, b *Health) *Health {
if a == nil {
return b
// HostHealthVerdict is THE host health rule (`11` §8.2; pinned by TestHostHealthVerdict): the Proxmox daemons and the
// agent are active, the customer guest still runs, the guest's own rule still passes against the start of the pass,
// and the tunnel is `running` (R-841).
func HostHealthVerdict(before, after *Health, tunnel string) (bool, string) {
if after == nil {
return false, "no health reading"
}
if b == nil {
return a
svcs := make([]string, 0, len(after.HostServices))
for s := range after.HostServices {
svcs = append(svcs, s)
}
out := &Health{DockerOK: a.DockerOK && b.DockerOK, NetworkOK: a.NetworkOK && b.NetworkOK, Controller: b.Controller,
Containers: map[string]Container{}}
for _, h := range []*Health{a, b} {
for n, c := range h.Containers {
cur, seen := out.Containers[n]
if !seen || (cur.State != "running" && c.State == "running") {
out.Containers[n] = c
sort.Strings(svcs)
if len(svcs) == 0 {
return false, "no host service reading"
}
for _, s := range svcs {
if after.HostServices[s] != "active" {
return false, s + " is " + after.HostServices[s]
}
}
if after.GuestRunning == nil || !*after.GuestRunning {
return false, "the customer guest is not running"
}
var gb *Health
if before != nil {
gb = before.Guest
}
if ok, why := HealthVerdict(gb, after.Guest); !ok {
return false, "guest: " + why
}
if tunnel != hub.TunnelRunning {
return false, "the tunnel is " + tunnel
}
return true, ""
}
// EngineOf is the engine version `docker version` prints for a docker-ce package version: "5:29.8.2-1~debian.13~trixie"
// → "29.8.2" (no epoch, no Debian revision).
func EngineOf(pkgVersion string) string {
v := pkgVersion
if i := strings.Index(v, ":"); i >= 0 {
v = v[i+1:]
}
if i := strings.Index(v, "-"); i >= 0 {
v = v[:i]
}
return v
}
// DockerHealthVerdict is THE Docker-step health rule (`11` §5.8; pinned by TestDockerHealthVerdict): the guest rule,
// plus every container running at the start still runs as the SAME container (same id — a changed id means the
// household's apps restarted, which `live-restore` exists to prevent), plus the engine now reports the version the step
// installed (wantEngine "" = no engine change expected).
func DockerHealthVerdict(before, after *Health, wantEngine, gotEngine string) (bool, string) {
if ok, why := HealthVerdict(before, after); !ok {
return false, why
}
if before != nil {
names := make([]string, 0, len(before.Containers))
for n := range before.Containers {
names = append(names, n)
}
sort.Strings(names)
for _, n := range names {
b := before.Containers[n]
if b.State != "running" || b.ID == "" {
continue
}
if c.State == "running" && c.Health == "healthy" {
out.Containers[n] = c
if a := after.Containers[n]; a.ID != b.ID {
return false, n + " is a new container (id changed) — the engine step restarted it"
}
}
}
return out
if wantEngine != "" && gotEngine != wantEngine {
return false, "the engine is " + gotEngine + ", not " + wantEngine
}
return true, ""
}
// call writes the plan and runs the wrapper once.
func (l *Leg) call(ctx context.Context, runID, mode string, vmid int, rel hub.WireOSRelease, pkgs []Package) (WrapperReport, error) {
func (l *Leg) call(ctx context.Context, runID string, plan map[string]any) (WrapperReport, error) {
dir := l.PlanDir
if dir == "" {
dir = DefaultPlanDir
@@ -275,13 +352,8 @@ func (l *Leg) call(ctx context.Context, runID, mode string, vmid int, rel hub.Wi
if err := os.MkdirAll(dir, 0o700); err != nil {
return WrapperReport{}, fmt.Errorf("osupdate: plan dir: %w", err)
}
plan := map[string]any{"release_id": rel.ID, "layer": "guest", "lane": "fast", "vmid": vmid, "mode": mode,
"snapshot": rel.Snapshot, "packages": pkgs}
if pkgs == nil {
plan["packages"] = []Package{}
}
b, _ := json.Marshal(plan)
path := filepath.Join(dir, "plan-"+runID+"-"+mode+".json")
path := filepath.Join(dir, fmt.Sprintf("plan-%s-%s-%s.json", runID, plan["layer"], plan["mode"]))
if err := os.WriteFile(path, b, 0o600); err != nil {
return WrapperReport{}, fmt.Errorf("osupdate: write plan: %w", err)
}
@@ -307,20 +379,87 @@ func (l *Leg) call(ctx context.Context, runID, mode string, vmid int, rel hub.Wi
return rep, nil // a refusal / failure is IN the report (exit 2 / 3), not an error here
}
// Run is one pass: inventory → (switch, ring, plan) → apply → health → report. trigger is "night" or "debug".
func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Report {
runID := l.now().UTC().Format("20060102T150405Z")
blk := l.Block()
rel := hub.WireOSRelease{ID: "ring0-" + runID}
if blk.Ring == 1 {
rel = hub.WireOSRelease{}
if blk.Release != nil {
rel = *blk.Release
}
}
rep := Report{RunID: runID, Trigger: trigger, Ring: blk.Ring, VMID: vmid, ReleaseID: rel.ID, Mode: "inventory"}
lg := l.log().With("run", runID, "vmid", vmid, "ring", blk.Ring, "trigger", trigger)
// Pass is one leg's reports; an empty Layer means the step did not run.
type Pass struct {
Guest, Host, Docker Report
}
// Run is one pass: the guest layer, then (on an appliance, after a good guest step) the host layer, then (ring 0
// only, after good earlier steps) the Docker engine set. trigger is "night" or "debug".
func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Pass {
g, h := l.runFast(ctx, vmid, trigger)
p := Pass{Guest: g, Host: h}
if g.Outcome == "skipped" {
return p
}
blk := l.Block()
okStep := func(r Report) bool {
return (r.Outcome == "applied" || r.Outcome == "nothing" || r.Outcome == "inventory") && r.Healthy
}
lg := l.log().With("run", g.RunID, "vmid", vmid, "trigger", trigger)
switch {
case blk.Ring != 0 || !blk.Enabled:
lg.Info("osupdate: docker step skipped — ring 1 takes an engine set only inside a signed operator job (`11` §5.8)", "ring", blk.Ring, "enabled", blk.Enabled)
case !okStep(g) || (h.Layer != "" && !okStep(h)):
lg.Warn("osupdate: docker step skipped — an earlier step did not end healthy")
default:
if err := l.EnsureLiveRestore(ctx, g.RunID, vmid); err != nil {
p.Docker = l.finish(ctx, lg, Report{RunID: g.RunID, Layer: LayerDocker, Trigger: trigger, Ring: 0, VMID: vmid,
Mode: "apply", Outcome: "failed", HealthReason: "live-restore could not be turned on: " + err.Error()})
return p
}
p.Docker = l.runLayer(ctx, g.RunID, LayerDocker, vmid, trigger, blk, dockerOpts{})
}
return p
}
// dockerOpts is a signed Docker step (DockerStepExecutor); the zero value is ring 0's unsigned "pending-docker".
type dockerOpts struct {
releaseID string
packages []Package
undo bool
signed map[string]string // blob_b64, sig — the wrapper verifies them ITSELF
}
// EnsureLiveRestore is the ONE-TIME step of `09` decision 87: the wrapper merges `"live-restore": true` into the guest's
// daemon.json and RELOADS docker (never a restart, R-835). A no-op when it is already on.
func (l *Leg) EnsureLiveRestore(ctx context.Context, runID string, vmid int) error {
wr, err := l.call(ctx, runID, map[string]any{"release_id": "live-restore", "layer": LayerGuest, "lane": "fast",
"vmid": vmid, "mode": "live-restore-on", "packages": []Package{}})
if err != nil {
return err
}
if wr.refused() {
return fmt.Errorf("refused: %s", wr.Refused)
}
if wr.failed() {
return fmt.Errorf("failed: %s", wr.Failed)
}
l.log().Info("osupdate: live-restore", "vmid", vmid, "result", string(wr.LiveRestore))
return nil
}
// Facts reads the box's versions through the wrapper's read-only facts mode (R-852): host Debian, kernels, held
// packages, taint, the crash guard; guest Debian, Docker engine, containerd, live-restore. Raw JSON, the wrapper's shape.
func (l *Leg) Facts(ctx context.Context, vmid int) (json.RawMessage, error) {
wr, err := l.call(ctx, "facts"+l.now().UTC().Format("150405"), map[string]any{"release_id": "facts", "layer": LayerHost,
"lane": "fast", "vmid": vmid, "mode": "facts", "packages": []Package{}})
if err != nil {
return nil, err
}
if wr.refused() {
return nil, fmt.Errorf("facts refused: %s", wr.Refused)
}
if len(wr.Facts) == 0 {
return nil, fmt.Errorf("facts: the wrapper returned none (an older wrapper?)")
}
return wr.Facts, nil
}
// runFast is the guest + host fast lane (agent v0.141.x behaviour).
func (l *Leg) runFast(ctx context.Context, vmid int, trigger string) (guest Report, host Report) {
runID := l.now().UTC().Format("20060102T150405Z")
lg := l.log().With("run", runID, "vmid", vmid, "trigger", trigger)
if trigger == "night" && l.StatePath != "" {
gap := l.MinGap
if gap == 0 {
@@ -329,68 +468,134 @@ func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Report {
if b, err := os.ReadFile(l.StatePath); err == nil {
if last, perr := time.Parse(time.RFC3339, strings.TrimSpace(string(b))); perr == nil && l.now().Sub(last) < gap {
lg.Info("osupdate: skipped — already ran tonight", "last", last.UTC().Format(time.RFC3339))
rep.Outcome = "skipped"
return rep
return Report{RunID: runID, Layer: LayerGuest, Outcome: "skipped"}, Report{}
}
}
}
blk := l.Block()
guest = l.runLayer(ctx, runID, LayerGuest, vmid, trigger, blk, dockerOpts{})
if trigger == "night" && l.StatePath != "" {
_ = os.WriteFile(l.StatePath, []byte(l.now().UTC().Format(time.RFC3339)), 0o600)
}
switch {
case !l.Appliance:
lg.Info("osupdate: host step skipped — not an appliance install (a BYO host belongs to its owner, `11` §1)")
case !(guest.Outcome == "applied" || guest.Outcome == "nothing" || guest.Outcome == "inventory") || !guest.Healthy:
lg.Warn("osupdate: host step skipped — the guest step did not end healthy", "guest_outcome", guest.Outcome, "reason", guest.HealthReason)
default:
host = l.runLayer(ctx, runID, LayerHost, vmid, trigger, blk, dockerOpts{})
}
return guest, host
}
func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigger string, blk hub.WireOSUpdate, do dockerOpts) Report {
rel := hub.WireOSRelease{ID: "ring0-" + runID}
var wire *hub.WireOSRelease
switch layer {
case LayerGuest:
wire = blk.Release
case LayerHost:
wire = blk.HostRelease
}
lane := "fast"
if layer == LayerDocker {
lane = "slow"
if do.signed != nil {
rel = hub.WireOSRelease{ID: do.releaseID}
}
} else if blk.Ring == 1 {
rel = hub.WireOSRelease{}
if wire != nil {
rel = *wire
}
}
rep := Report{RunID: runID, Layer: layer, Trigger: trigger, Ring: blk.Ring, VMID: vmid, ReleaseID: rel.ID}
lg := l.log().With("run", runID, "layer", layer, "vmid", vmid, "ring", blk.Ring, "trigger", trigger)
lg.Info("osupdate: START", "enabled", blk.Enabled, "release", rel.ID)
inv, err := l.call(ctx, runID, "inventory", vmid, rel, nil)
if err != nil || inv.refused() || inv.failed() {
rep.Outcome, rep.Refused = "refused", inv.Refused
if err != nil {
rep.Outcome, rep.HealthReason = "failed", err.Error()
}
return l.finish(ctx, lg, rep)
plan := map[string]any{"release_id": rel.ID, "layer": layer, "lane": lane, "vmid": vmid, "snapshot": rel.Snapshot,
"packages": []Package{}, "mode": "apply", "select": "listed"}
if rel.ID == "" {
plan["release_id"] = "none"
}
var plan []Package
planned := map[string]bool{}
switch {
case layer == LayerDocker && do.signed != nil:
plan["packages"], plan["signed"] = do.packages, do.signed
if do.undo {
plan["undo"] = true
}
for _, p := range do.packages {
planned[p.Name] = true
}
case layer == LayerDocker:
plan["select"] = "pending-docker" // ring 0: the wrapper checks the box's ROOT-OWNED ring-0 mark itself
case !blk.Enabled:
rep.Outcome = "inventory"
plan["mode"] = "inventory"
lg.Info("osupdate: switched OFF for this box — reporting only")
case blk.Ring == 0:
for _, p := range inv.Pending {
if IsFast(p.Origin) {
plan = append(plan, Package{Name: p.Name, Version: p.To, Origin: fastOrigin(p.Origin)})
}
}
plan["select"] = "pending-fast" // the wrapper picks every pending Debian / Debian-Security upgrade
case len(rel.Packages) == 0:
plan["mode"] = "inventory" // ring 1 with no approved release for this layer: nothing to install
default:
var pk []Package
for _, p := range rel.Packages {
plan = append(plan, Package{Name: p.Name, Version: p.Version, Origin: p.Origin})
pk = append(pk, Package{Name: p.Name, Version: p.Version, Origin: p.Origin})
planned[p.Name] = true
}
plan["packages"] = pk
}
pkgNames := map[string]bool{}
for _, p := range plan {
pkgNames[p.Name] = true
rep.Mode = plan["mode"].(string)
wr, err := l.call(ctx, runID, plan)
switch {
case err != nil:
rep.Outcome, rep.HealthReason = "failed", err.Error()
return l.finish(ctx, lg, rep)
case wr.refused():
rep.Outcome, rep.Refused = "refused", wr.Refused
return l.finish(ctx, lg, rep)
case wr.failed():
rep.Outcome, rep.Refused = "failed", wr.Failed
}
if blk.Enabled && len(plan) == 0 {
rep.Outcome = "nothing"
}
final := inv
if blk.Enabled && len(plan) > 0 {
rep.Mode = "apply"
ap, err := l.call(ctx, runID, "apply", vmid, rel, plan)
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
if rep.Outcome == "" {
switch {
case err != nil:
rep.Outcome, rep.HealthReason = "failed", err.Error()
return l.finish(ctx, lg, rep)
case ap.refused():
rep.Outcome, rep.Refused = "refused", ap.Refused
return l.finish(ctx, lg, rep)
case ap.failed():
rep.Outcome, rep.Refused = "failed", ap.Failed
case rep.Mode == "inventory" && !blk.Enabled:
rep.Outcome = "inventory"
case len(wr.Upgraded) == 0:
rep.Outcome = "nothing"
default:
rep.Outcome = "applied"
}
final = ap
rep.Upgraded = ap.Upgraded
if rep.Outcome == "" {
if len(ap.Upgraded) == 0 {
rep.Outcome = "nothing"
} else {
rep.Outcome = "applied"
}
if blk.Ring == 0 {
for _, u := range wr.Upgraded {
planned[u.Name] = true
}
}
// Health: compare with the start of the pass; give restarted services time (only after an install).
cur := wr.HealthAfter
wantEngine := ""
for _, u := range wr.Upgraded {
if u.Name == "docker-ce" {
wantEngine = EngineOf(u.Version)
}
}
verdict := func(h *Health) (bool, string) {
if layer == LayerDocker {
return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine)
}
if layer == LayerHost {
t := hub.TunnelUnknown
if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx)
}
return HostHealthVerdict(wr.HealthBefore, h, t)
}
// Health: compare with what the guest looked like BEFORE the run; give restarted services time.
return HealthVerdict(wr.HealthBefore, h)
}
if len(wr.Upgraded) > 0 {
wait, poll := l.HealthWait, l.HealthPoll
if wait == 0 {
wait = 5 * time.Minute
@@ -399,19 +604,15 @@ func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Report {
poll = 15 * time.Second
}
deadline := l.now().Add(wait)
cur := ap.HealthAfter
// The baseline is the guest as it was at the START of the leg (the inventory's reading) merged with the
// apply's own "before": an app that stops at any point during the run counts. Found live 2026-10-04: an app
// stopped between the inventory and the apply's own reading was taken as "stopped before" and ignored.
base := MergeBaseline(inv.HealthAfter, ap.HealthBefore)
for {
ok, why := HealthVerdict(base, cur)
ok, why := verdict(cur)
rep.Healthy, rep.HealthReason = ok, why
if ok || !l.now().Before(deadline) || ctx.Err() != nil {
break
}
l.sleep(ctx, poll)
hr, herr := l.call(ctx, runID, "health", vmid, rel, nil)
hp := map[string]any{"release_id": plan["release_id"], "layer": layer, "lane": lane, "vmid": vmid, "mode": "health", "packages": []Package{}}
hr, herr := l.call(ctx, runID, hp)
if herr == nil && hr.Health != nil {
cur = hr.Health
}
@@ -420,20 +621,43 @@ func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Report {
rep.Outcome = "health_failed"
}
} else {
ok, why := HealthVerdict(nil, inv.HealthAfter)
rep.Healthy, rep.HealthReason = ok, why
rep.Healthy, rep.HealthReason = verdict(cur)
}
rep.Installed, rep.Pending = final.Installed, final.Pending
rep.RestartNeeded, rep.DockerRestartNeeded, rep.RebootNeeded = final.RestartNeeded, final.DockerRestartNeeded, final.RebootNeeded
rep.NotCovered = notCovered(final.Pending, blk.Ring, pkgNames)
if trigger == "night" && l.StatePath != "" {
_ = os.WriteFile(l.StatePath, []byte(l.now().UTC().Format(time.RFC3339)), 0o600)
rep.Installed, rep.Pending = wr.Installed, wr.Pending
rep.RestartNeeded, rep.DockerRestartNeeded, rep.RebootNeeded = wr.RestartNeeded, wr.DockerRestartNeeded, wr.RebootNeeded
rep.RebootScanned = wr.RebootScanned
if layer == LayerDocker {
// the docker report carries the engine set only (the guest report already carries the Debian packages)
rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending)
rep.NotCovered = nil
} else {
rep.NotCovered = notCovered(wr.Pending, blk.Ring, planned)
}
return l.finish(ctx, lg, rep)
}
// notCovered lists pending updates no approved release covers: in ring 0 everything outside the fast lane; in
// ring 1 also every fast-lane update the release did not name.
func onlyDocker(in []Package) []Package {
var out []Package
for _, p := range in {
if DockerNames[p.Name] {
out = append(out, p)
}
}
return out
}
func onlyDockerPending(in []Pending) []Pending {
var out []Pending
for _, p := range in {
if DockerNames[p.Name] {
out = append(out, p)
}
}
return out
}
// notCovered lists pending updates no approved release covers: in ring 0 everything outside the fast lane; in ring 1
// also every fast-lane update the release did not name.
func notCovered(pending []Pending, ring int, planned map[string]bool) []string {
var out []string
for _, p := range pending {
@@ -446,7 +670,8 @@ func notCovered(pending []Pending, ring int, planned map[string]bool) []string {
func (l *Leg) finish(ctx context.Context, lg *slog.Logger, rep Report) Report {
lg.Info("osupdate: DONE", "outcome", rep.Outcome, "healthy", rep.Healthy, "reason", rep.HealthReason,
"upgraded", len(rep.Upgraded), "pending", len(rep.Pending), "not_covered", len(rep.NotCovered), "restart_needed", len(rep.RestartNeeded))
"upgraded", len(rep.Upgraded), "pending", len(rep.Pending), "not_covered", len(rep.NotCovered),
"restart_needed", len(rep.RestartNeeded), "reboot_needed", rep.RebootNeeded, "wrapper_seconds", rep.PassSeconds)
if l.Hub != nil {
body, _ := json.Marshal(rep)
rctx, cancel := context.WithTimeout(context.WithoutCancel(ctx), time.Minute)
+287 -101
View File
@@ -14,15 +14,27 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// fakeWrapper plays /usr/local/sbin/felhom-os-apply: it reads the plan the leg wrote and answers per mode.
// fakeWrapper plays /usr/local/sbin/felhom-os-apply: it reads the plan the leg wrote and answers per layer and mode.
type fakeWrapper struct {
t *testing.T
pending []Pending
applyRep WrapperReport
healthSeq []*Health // answers to successive "health" calls
applyRep map[string]WrapperReport // per layer
healthSeq map[string][]*Health // per layer: answers to successive "health" calls
plans []map[string]any
}
func yes() *bool { b := true; return &b }
func guestOK() *Health {
return &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{
"felhom-controller": {State: "running", Health: "healthy"}, "app": {State: "running", Health: "healthy"}}}
}
func hostOK() *Health {
return &Health{HostServices: map[string]string{"pveproxy": "active", "pvedaemon": "active", "pvestatd": "active",
"pve-cluster": "active", "felhom-agent": "active"}, GuestRunning: yes(), Guest: guestOK()}
}
func (f *fakeWrapper) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
if name != WrapperPath || len(args) != 2 || args[0] != "--plan" {
f.t.Fatalf("unexpected command %s %v", name, args)
@@ -34,27 +46,30 @@ func (f *fakeWrapper) Run(_ context.Context, name string, args ...string) ([]byt
var plan map[string]any
json.Unmarshal(b, &plan)
f.plans = append(f.plans, plan)
layer := plan["layer"].(string)
ok := guestOK()
if layer == LayerHost {
ok = hostOK()
}
var rep WrapperReport
healthy := &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{
"felhom-controller": {State: "running", Health: "healthy"}, "app": {State: "running", Health: "healthy"}}}
switch plan["mode"] {
case "inventory":
rep = WrapperReport{Mode: "inventory", Pending: f.pending, HealthAfter: healthy,
rep = WrapperReport{Mode: "inventory", Pending: f.pending, HealthBefore: ok, HealthAfter: ok,
Installed: []Package{{Name: "libc6", Version: "u3", Origin: "Debian"}}}
case "apply":
rep = f.applyRep
rep = f.applyRep[layer]
rep.Mode = "apply"
if rep.HealthBefore == nil {
rep.HealthBefore = healthy
rep.HealthBefore = ok
}
if rep.HealthAfter == nil {
rep.HealthAfter = healthy
rep.HealthAfter = ok
}
case "health":
if len(f.healthSeq) > 0 {
rep.Health, f.healthSeq = f.healthSeq[0], f.healthSeq[1:]
if seq := f.healthSeq[layer]; len(seq) > 0 {
rep.Health, f.healthSeq[layer] = seq[0], seq[1:]
} else {
rep.Health = healthy
rep.Health = ok
}
}
out, _ := json.Marshal(rep)
@@ -74,11 +89,21 @@ func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error {
return nil
}
type fakeTunnel struct{ st string }
func (t fakeTunnel) Status(context.Context) (string, string) { return t.st, "" }
func newLeg(t *testing.T, w *fakeWrapper, blk *hub.WireOSUpdate) (*Leg, *fakeHub) {
h := &fakeHub{}
now := time.Date(2026, 10, 4, 4, 0, 0, 0, time.UTC)
if w.applyRep == nil {
w.applyRep = map[string]WrapperReport{}
}
if w.healthSeq == nil {
w.healthSeq = map[string][]*Health{}
}
l := &Leg{Runner: w, Hub: h, PlanDir: t.TempDir(), StatePath: filepath.Join(t.TempDir(), "last"),
HealthWait: time.Minute, HealthPoll: 10 * time.Second,
HealthWait: time.Minute, HealthPoll: 10 * time.Second, Appliance: true, Tunnel: fakeTunnel{hub.TunnelRunning},
Now: func() time.Time { return now },
Sleep: func(_ context.Context, d time.Duration) { now = now.Add(d) }}
if blk != nil {
@@ -93,66 +118,73 @@ var pend = []Pending{
{Name: "docker-ce", From: "29.7", To: "29.8", Origin: []string{"Docker CE"}},
}
func modes(w *fakeWrapper) string {
func calls(w *fakeWrapper) string {
var m []string
for _, p := range w.plans {
m = append(m, p["mode"].(string))
m = append(m, p["layer"].(string)+":"+p["mode"].(string))
}
return strings.Join(m, ",")
}
// Ring 0 plans every pending FAST-LANE update (Docker excluded, `11` C3) and reports what it now runs.
func TestRing0_PlansTheFastLaneOnly(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6", Version: "u4"}, {Name: "openssl", Version: "u3"}}, Pending: pend[2:]}}
// Ring 0: ONE wrapper call per layer (R-845), select pending-fast (the wrapper picks every Debian / Debian-Security
// upgrade), from live sources; the guest step first, then the host step.
func TestRing0_OneCallPerLayer(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{
LayerGuest: {Upgraded: []Package{{Name: "libc6", Version: "u4"}, {Name: "openssl", Version: "u3"}}, Pending: pend[2:]},
LayerHost: {Upgraded: []Package{{Name: "openssl", Version: "u3"}}},
}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
rep := l.Run(context.Background(), 9201, "night")
if rep.Outcome != "applied" || !rep.Healthy {
t.Fatalf("rep = %+v", rep)
g, ho := run2(l, "night")
if g.Outcome != "applied" || !g.Healthy || ho.Outcome != "applied" || !ho.Healthy {
t.Fatalf("guest %+v\nhost %+v", g, ho)
}
pk := w.plans[1]["packages"].([]any)
if len(pk) != 2 {
t.Fatalf("ring-0 plan = %v, want libc6 + openssl only", pk)
if calls(w) != "guest:apply,host:apply,guest:live-restore-on,docker:apply" {
t.Fatalf("calls = %s, want one apply per layer, guest first, then live-restore and the ring-0 docker step", calls(w))
}
if o := pk[1].(map[string]any)["origin"]; o != "Debian-Security" {
t.Fatalf("openssl origin = %v", o)
for _, p := range w.plans[:2] {
if p["select"] != "pending-fast" || p["snapshot"] != "" || len(p["packages"].([]any)) != 0 {
t.Fatalf("ring-0 plan = %v", p)
}
}
if w.plans[1]["snapshot"] != "" {
t.Fatalf("ring 0 installs from live sources, snapshot = %v", w.plans[1]["snapshot"])
if len(g.NotCovered) != 1 || g.NotCovered[0] != "docker-ce" {
t.Fatalf("not covered = %v", g.NotCovered)
}
if len(rep.NotCovered) != 1 || rep.NotCovered[0] != "docker-ce" {
t.Fatalf("not covered = %v", rep.NotCovered)
}
if len(h.reports) != 1 || h.reports[0].Outcome != "applied" {
if len(h.reports) != 3 || h.reports[0].Layer != LayerGuest || h.reports[1].Layer != LayerHost || h.reports[2].Layer != LayerDocker {
t.Fatalf("hub got %+v", h.reports)
}
}
// Ring 1 installs EXACTLY the approved release (its versions, its snapshot), nothing it computed itself.
func TestRing1_InstallsExactlyTheRelease(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6", Version: "u4-approved"}}, Pending: pend[1:]}}
rel := &hub.WireOSRelease{ID: "os-1", Snapshot: "20261004T080000Z", Packages: []hub.WireOSPackage{{Name: "libc6", Version: "u4-approved", Origin: "Debian"}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true, Release: rel})
rep := l.Run(context.Background(), 9201, "night")
if rep.Outcome != "applied" || rep.ReleaseID != "os-1" {
t.Fatalf("rep = %+v", rep)
// Ring 1 installs EXACTLY each layer's own approved release (a guest release is not a host release).
func TestRing1_EachLayerItsOwnRelease(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{
LayerGuest: {Upgraded: []Package{{Name: "libc6", Version: "g-u4"}}, Pending: pend[1:]},
LayerHost: {Upgraded: []Package{{Name: "openssl", Version: "h-u3"}}},
}}
gr := &hub.WireOSRelease{ID: "os-g", Snapshot: "20261004T080000Z", Packages: []hub.WireOSPackage{{Name: "libc6", Version: "g-u4", Origin: "Debian"}}}
hr := &hub.WireOSRelease{ID: "os-h", Snapshot: "20261004T090000Z", Packages: []hub.WireOSPackage{{Name: "openssl", Version: "h-u3", Origin: "Debian-Security"}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true, Release: gr, HostRelease: hr})
g, ho := run2(l, "night")
if g.ReleaseID != "os-g" || ho.ReleaseID != "os-h" {
t.Fatalf("release ids %q %q", g.ReleaseID, ho.ReleaseID)
}
ap := w.plans[1]
pk := ap["packages"].([]any)
if len(pk) != 1 || pk[0].(map[string]any)["version"] != "u4-approved" || ap["snapshot"] != "20261004T080000Z" || ap["release_id"] != "os-1" {
t.Fatalf("ring-1 plan = %v", ap)
gp, hp := w.plans[0], w.plans[1]
if gp["snapshot"] != "20261004T080000Z" || gp["packages"].([]any)[0].(map[string]any)["version"] != "g-u4" {
t.Fatalf("guest plan %v", gp)
}
// openssl is pending and fast-lane but NOT in the release → not covered; docker-ce is never covered.
if strings.Join(rep.NotCovered, ",") != "openssl,docker-ce" {
t.Fatalf("not covered = %v", rep.NotCovered)
if hp["layer"] != LayerHost || hp["snapshot"] != "20261004T090000Z" || hp["packages"].([]any)[0].(map[string]any)["name"] != "openssl" {
t.Fatalf("host plan %v", hp)
}
if strings.Join(g.NotCovered, ",") != "openssl,docker-ce" {
t.Fatalf("guest not covered = %v", g.NotCovered)
}
}
func TestRing1_NoReleaseInstallsNothing(t *testing.T) {
func TestRing1_NoReleaseIsInventory(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
rep := l.Run(context.Background(), 9201, "night")
if rep.Outcome != "nothing" || modes(w) != "inventory" {
t.Fatalf("rep=%+v modes=%s", rep, modes(w))
g, ho := run2(l, "night")
if g.Outcome != "nothing" || ho.Outcome != "nothing" || calls(w) != "guest:inventory,host:inventory" {
t.Fatalf("g=%+v h=%+v calls=%s", g, ho, calls(w))
}
}
@@ -160,31 +192,54 @@ func TestRing1_NoReleaseInstallsNothing(t *testing.T) {
func TestNoBlock_IsRing1Nothing(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend}
l, _ := newLeg(t, w, nil)
if rep := l.Run(context.Background(), 9201, "night"); rep.Outcome != "nothing" || rep.Ring != 1 || modes(w) != "inventory" {
t.Fatalf("rep=%+v modes=%s", rep, modes(w))
if g, _ := run2(l, "night"); g.Outcome != "nothing" || g.Ring != 1 {
t.Fatalf("g=%+v calls=%s", g, calls(w))
}
}
// Switched OFF: the box reports but installs nothing (the brief, Part D 2).
// Switched OFF: the box reports but installs nothing, on both layers.
func TestSwitchOff_ReportsOnly(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: false})
rep := l.Run(context.Background(), 9201, "night")
if rep.Outcome != "inventory" || modes(w) != "inventory" || len(h.reports) != 1 || len(rep.Pending) != 3 {
t.Fatalf("rep=%+v modes=%s", rep, modes(w))
g, ho := run2(l, "night")
if g.Outcome != "inventory" || ho.Outcome != "inventory" || calls(w) != "guest:inventory,host:inventory" || len(h.reports) != 2 {
t.Fatalf("g=%+v h=%+v calls=%s", g, ho, calls(w))
}
}
// Unhealthy after the run, and still unhealthy at the end of the wait → health_failed (the hub mails the operator).
// Not an appliance (BYO, `11` §1): the host step never runs — no host plan at all.
func TestBYO_NoHostPlan(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Appliance = false
_, ho := run2(l, "night")
// the guest (and so its Docker engine) is ours on a BYO box too: only the HOST is the owner's
if ho.Outcome != "" || calls(w) != "guest:apply,guest:live-restore-on,docker:apply" || len(h.reports) != 2 {
t.Fatalf("a BYO box got a host step: host=%+v calls=%s", ho, calls(w))
}
}
// A failed guest step skips the host step that night.
func TestGuestFailure_SkipsTheHost(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{
LayerGuest: {Refused: json.RawMessage(`{"code":"R6","reason":"x"}`)}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
g, ho := run2(l, "night")
if g.Outcome != "refused" || ho.Outcome != "" || calls(w) != "guest:apply" {
t.Fatalf("g=%+v h=%+v calls=%s", g, ho, calls(w))
}
}
// Unhealthy after the run, and still unhealthy at the end of the wait → health_failed; the host step is skipped.
func TestHealth_FailsAfterTheWait(t *testing.T) {
bad := &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{
"felhom-controller": {State: "running", Health: "healthy"}, "app": {State: "exited"}}}
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6"}}, HealthAfter: bad},
healthSeq: []*Health{bad, bad, bad, bad, bad, bad, bad, bad}}
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}, HealthAfter: bad}},
healthSeq: map[string][]*Health{LayerGuest: {bad, bad, bad, bad, bad, bad, bad, bad}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
rep := l.Run(context.Background(), 9201, "night")
if rep.Outcome != "health_failed" || rep.Healthy || !strings.Contains(rep.HealthReason, "app was running") {
t.Fatalf("rep = %+v", rep)
g, ho := run2(l, "night")
if g.Outcome != "health_failed" || g.Healthy || !strings.Contains(g.HealthReason, "app was running") || ho.Outcome != "" {
t.Fatalf("g=%+v h=%+v", g, ho)
}
if h.reports[0].Outcome != "health_failed" {
t.Fatal("the hub was not told")
@@ -194,11 +249,22 @@ func TestHealth_FailsAfterTheWait(t *testing.T) {
// A service that takes a moment to come back is not a failure: the poll sees it recover inside the wait.
func TestHealth_RecoversInsideTheWait(t *testing.T) {
starting := &Health{DockerOK: true, NetworkOK: true, Controller: "starting"}
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6"}}, HealthAfter: starting},
healthSeq: []*Health{starting}}
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}, HealthAfter: starting}},
healthSeq: map[string][]*Health{LayerGuest: {starting}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
if rep := l.Run(context.Background(), 9201, "night"); rep.Outcome != "applied" || !rep.Healthy {
t.Fatalf("rep = %+v", rep)
if g, _ := run2(l, "night"); g.Outcome != "applied" || !g.Healthy {
t.Fatalf("g = %+v", g)
}
}
// The host step judged unhealthy when the tunnel is down after it.
func TestHost_TunnelDownFailsTheHostStep(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerHost: {Upgraded: []Package{{Name: "openssl"}}}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Tunnel = fakeTunnel{hub.TunnelNotRunning}
_, ho := run2(l, "night")
if ho.Outcome != "health_failed" || !strings.Contains(ho.HealthReason, "tunnel") {
t.Fatalf("host = %+v", ho)
}
}
@@ -227,24 +293,49 @@ func TestHealthVerdict(t *testing.T) {
}
}
func TestOncePerNight(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6"}}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Run(context.Background(), 9201, "night")
n := len(w.plans)
if rep := l.Run(context.Background(), 9201, "night"); rep.Outcome != "skipped" || len(w.plans) != n {
t.Fatalf("a second night run in the same night ran: %+v", rep)
// The host rule: every listed daemon active, the guest running and passing its own rule, the tunnel running.
// Red-proofs: drop any one check and its case fails.
func TestHostHealthVerdict(t *testing.T) {
no := false
svcDown := hostOK()
svcDown.HostServices["pveproxy"] = "failed"
guestDown := hostOK()
guestDown.GuestRunning = &no
guestApp := hostOK()
guestApp.Guest = &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{"felhom-controller": {State: "running", Health: "healthy"}}}
cases := []struct {
name string
after *Health
tunnel string
want bool
why string
}{
{"all good", hostOK(), hub.TunnelRunning, true, ""},
{"a daemon down", svcDown, hub.TunnelRunning, false, "pveproxy"},
{"the guest stopped", guestDown, hub.TunnelRunning, false, "guest is not running"},
{"an app in the guest gone", guestApp, hub.TunnelRunning, false, "app was running"},
{"the tunnel down", hostOK(), hub.TunnelNotRunning, false, "tunnel"},
{"the tunnel unknown", hostOK(), hub.TunnelUnknown, false, "tunnel"},
{"no services read", &Health{GuestRunning: yes(), Guest: guestOK()}, hub.TunnelRunning, false, "no host service"},
}
if rep := l.Run(context.Background(), 9201, "debug"); rep.Outcome == "skipped" {
t.Fatal("the debug action must not be throttled")
for _, c := range cases {
got, why := HostHealthVerdict(hostOK(), c.after, c.tunnel)
if got != c.want || !strings.Contains(why, c.why) {
t.Errorf("%s: got %v (%s), want %v (…%s…)", c.name, got, why, c.want, c.why)
}
}
}
func TestRefusedIsReported(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Refused: json.RawMessage(`{"code":"R6","reason":"x"}`)}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
if rep := l.Run(context.Background(), 9201, "night"); rep.Outcome != "refused" || h.reports[0].Outcome != "refused" {
t.Fatalf("rep = %+v", rep)
func TestOncePerNight(t *testing.T) {
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
run2(l, "night")
n := len(w.plans)
if g, _ := run2(l, "night"); g.Outcome != "skipped" || len(w.plans) != n {
t.Fatalf("a second night run in the same night ran: %+v", g)
}
if g, _ := run2(l, "debug"); g.Outcome == "skipped" {
t.Fatal("the debug action must not be throttled")
}
}
@@ -254,28 +345,123 @@ func TestWrapperSuite(t *testing.T) {
if err != nil {
t.Skip("python3 not available")
}
cmd := exec.Command(py, "../../configs/test_felhom_os_apply.py")
out, err := cmd.CombinedOutput()
if err != nil {
t.Fatalf("wrapper suite failed: %v\n%s", err, out)
}
if !strings.Contains(string(out), "OK") {
t.Fatalf("wrapper suite did not report OK:\n%s", out)
// the OS wrapper and (agent v0.142.0) the crash guard — both root programs in configs/ with their own suites
for _, suite := range []string{"../../configs/test_felhom_os_apply.py", "../../configs/test_felhom_crash_guard.py"} {
cmd := exec.Command(py, "-B", suite)
out, err := cmd.CombinedOutput()
if err != nil {
t.Fatalf("%s failed: %v\n%s", suite, err, out)
}
if !strings.Contains(string(out), "OK") {
t.Fatalf("%s did not report OK:\n%s", suite, out)
}
}
}
// An app that stops BETWEEN the start of the leg and the apply's own "before" reading still fails the run.
// Measured live 2026-10-04 on demo-hp (privatebin stopped 1 s after the apply plan was written: the old rule
// passed). Red-proof: use ap.HealthBefore alone as the baseline and this fails.
func TestHealth_BaselineIsTheStartOfTheLeg(t *testing.T) {
stoppedEarly := &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{
"felhom-controller": {State: "running", Health: "healthy"}, "app": {State: "exited"}}}
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6"}},
HealthBefore: stoppedEarly, HealthAfter: stoppedEarly},
healthSeq: []*Health{stoppedEarly, stoppedEarly, stoppedEarly, stoppedEarly, stoppedEarly, stoppedEarly, stoppedEarly}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
rep := l.Run(context.Background(), 9201, "night")
if rep.Outcome != "health_failed" || !strings.Contains(rep.HealthReason, "app was running") {
t.Fatalf("an app that stopped during the run passed: %+v", rep)
// The host's restart scan result reaches the hub with reboot_scanned, so the hub can tell "looked: not needed" (a
// reboot cleared it) from "did not look". Red-proof: drop the RebootScanned copy in runLayer and this fails.
func TestHostReport_CarriesRebootScanned(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{
LayerGuest: {},
LayerHost: {RebootScanned: true, RebootNeeded: true, RestartNeeded: []string{"lxc-start"}},
}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
run2(l, "night")
if len(h.reports) != 3 || !h.reports[1].RebootScanned || !h.reports[1].RebootNeeded || h.reports[0].RebootScanned {
t.Fatalf("hub got %+v", h.reports)
}
}
// run2 is the guest + host reports of one pass (the tests written before the docker step).
func run2(l *Leg, trigger string) (Report, Report) {
p := l.Run(context.Background(), 9201, trigger)
return p.Guest, p.Host
}
// ---- the Docker step (`11` §5.8, agent v0.142.0) ----
// Ring 1 never takes an engine step in the night leg — only inside a signed operator job. Red-proof: drop the
// `blk.Ring != 0` case in Run and the ring-1 pass makes a docker call.
func TestDocker_Ring1NightLegNeverSteps(t *testing.T) {
w := &fakeWrapper{t: t}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
if p.Docker.Layer != "" || strings.Contains(calls(w), "docker") || strings.Contains(calls(w), "live-restore") {
t.Fatalf("ring 1 took a docker step: %s", calls(w))
}
}
// An unhealthy earlier step skips the docker step.
func TestDocker_SkippedAfterAnUnhealthyStep(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}}},
healthSeq: map[string][]*Health{}}
bad := guestOK()
bad.Controller = "unhealthy"
w.applyRep[LayerGuest] = WrapperReport{Upgraded: []Package{{Name: "libc6"}}, HealthAfter: bad}
w.healthSeq[LayerGuest] = []*Health{bad, bad, bad, bad, bad, bad, bad, bad}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
if p.Docker.Layer != "" || strings.Contains(calls(w), "docker") {
t.Fatalf("docker step ran after an unhealthy guest step: %s", calls(w))
}
}
// The docker plan is the slow lane, pending-docker for ring 0; the report carries only the engine set.
func TestDocker_Ring0PlanAndReport(t *testing.T) {
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}},
Installed: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie", Origin: "Docker"}, {Name: "libc6", Version: "u4", Origin: "Debian"}},
DockerEngine: "29.8.2", Authority: "ring0"}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
dp := w.plans[len(w.plans)-1]
if dp["layer"] != "docker" || dp["lane"] != "slow" || dp["select"] != "pending-docker" {
t.Fatalf("docker plan = %v", dp)
}
d := p.Docker
if d.Outcome != "applied" || !d.Healthy || d.DockerEngine != "29.8.2" || len(d.Installed) != 1 || d.Installed[0].Name != "docker-ce" {
t.Fatalf("docker report = %+v", d)
}
}
// THE docker health rule. Red-proof: drop the id comparison (or the engine check) in DockerHealthVerdict and a case fails.
func TestDockerHealthVerdict(t *testing.T) {
before := guestOK()
before.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"},
"app": {State: "running", Health: "healthy", ID: "b"}}
same := guestOK()
same.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"},
"app": {State: "running", Health: "healthy", ID: "b"}}
moved := guestOK()
moved.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"},
"app": {State: "running", Health: "healthy", ID: "c"}}
if ok, why := DockerHealthVerdict(before, same, "29.8.2", "29.8.2"); !ok {
t.Fatalf("same ids, right engine: %s", why)
}
if ok, _ := DockerHealthVerdict(before, moved, "29.8.2", "29.8.2"); ok {
t.Fatal("a changed container id passed — live-restore failed and the apps restarted")
}
if ok, _ := DockerHealthVerdict(before, same, "29.8.2", "29.7.2"); ok {
t.Fatal("the engine did not move and the step passed")
}
if EngineOf("5:29.8.2-1~debian.13~trixie") != "29.8.2" {
t.Fatalf("EngineOf = %q", EngineOf("5:29.8.2-1~debian.13~trixie"))
}
}
// A changed id after the step → health_failed (the consequence, not only the verdict).
func TestDocker_ChangedIDIsHealthFailed(t *testing.T) {
before := guestOK()
before.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"}}
after := guestOK()
after.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "z"}}
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1"}}, DockerEngine: "29.8.2",
HealthBefore: before, HealthAfter: after}}, healthSeq: map[string][]*Health{}}
w.healthSeq[LayerDocker] = []*Health{after, after, after, after, after, after, after, after}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
p := l.Run(context.Background(), 9201, "night")
if p.Docker.Outcome != "health_failed" || !strings.Contains(p.Docker.HealthReason, "id changed") {
t.Fatalf("docker = %+v", p.Docker)
}
}
+5 -1
View File
@@ -47,6 +47,10 @@ const (
// Destructive for unknown classes — this named constant documents the class and keeps the
// signed-op vocabulary explicit, it does not (and must not) loosen anything.
ClassAgentUpdate OpClass = "agent_update"
// A Docker engine step in the customer guest (agent v0.142.0, `11` §5.8) — ring 1, and every undo. Destructive-class
// (signed, operational key) like agent_update; the root wrapper re-verifies the same signature itself.
ClassOSDockerStep OpClass = "os_docker_step"
)
// Disposition is the classifier verdict.
@@ -115,7 +119,7 @@ func Classify(class OpClass, prov Provenance) Disposition {
return Destructive
case ClassKeyRotation:
return Destructive
case ClassAgentUpdate:
case ClassAgentUpdate, ClassOSDockerStep:
// Never benign — no agent-internal provenance can make replacing the agent binary
// unsigned-safe (a compromised process must not be able to self-bless an update).
return Destructive
+16 -1
View File
@@ -49,6 +49,21 @@ type Executor interface {
Execute(ctx context.Context, op string, params json.RawMessage) error
}
type signedOpKey struct{}
// WithSignedOp / SignedOpFrom carry the RAW verified envelope (blob bytes + armored signature) to an executor whose
// ROOT half verifies it AGAIN against a root-owned key file (agent v0.142.0, the Docker slow lane: the agent's own
// config is agent-writable, so a root wrapper must not take the agent's word for a signature).
func WithSignedOp(ctx context.Context, s *reconcile.SignedOp) context.Context {
return context.WithValue(ctx, signedOpKey{}, s)
}
// SignedOpFrom returns the envelope set by WithSignedOp.
func SignedOpFrom(ctx context.Context) (*reconcile.SignedOp, bool) {
s, ok := ctx.Value(signedOpKey{}).(*reconcile.SignedOp)
return s, ok && s != nil
}
// ErrNoExecutor signals an op class with no executor wired in this build (don't clear the job).
var ErrNoExecutor = fmt.Errorf("signedjobs: no executor for this op class in this build")
@@ -158,7 +173,7 @@ func (r *Runner) processJob(ctx context.Context, j hub.JobWire) bool {
// Allowed: the nonce is already durably burned (Verify, before this point). Execute.
r.logger.Warn("signedjobs: AUTHORIZED signed op — executing",
"job", j.JobID, "op", ob.Op, "key_id", dec.Verified.KeyID, "nonce", dec.Verified.Nonce)
err := r.exec.Execute(ctx, ob.Op, ob.Params)
err := r.exec.Execute(WithSignedOp(ctx, signed), ob.Op, ob.Params)
switch {
case err == nil:
r.logger.Warn("signedjobs: signed op COMPLETED", "job", j.JobID, "op", ob.Op)