Compare commits
7 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 495003051b | |||
| 42af3ab9bc | |||
| f24dce5b95 | |||
| b1746c25af | |||
| 2e2e8f56b8 | |||
| a6bc3f1197 | |||
| 3bf77c3423 |
@@ -1,3 +1,96 @@
|
||||
## v0.142.0 — the Docker engine slow lane, the version report, the crash guard (`11` §5.8, §5.9; `09` decisions 87–89)
|
||||
|
||||
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.142.0` (`b1746c2`), sha256
|
||||
> `7beb32224d6495e9561acfd3ad8a48393a799520f196080011cb27fceb6d1de6`, verified by download. Not vouched at release time.
|
||||
|
||||
**MinAgent impact:** none. **Needs hub v0.132.0** for the System page, the Docker approval and the crash events; an older
|
||||
hub stores the new `system` stanza unread. **Needs the new root files** on an installed box (R-840): the wrapper,
|
||||
`/etc/felhom/os-trust.json`, `/etc/felhom/operator-signers` and the crash guard — the installer 1.30.0 writes them; the
|
||||
demo boxes got them by hand.
|
||||
|
||||
- **Docker `live-restore` ON** (decision 87). Wrapper mode `live-restore-on`: merge `"live-restore": true` into the
|
||||
guest's `/etc/docker/daemon.json` and `systemctl reload docker` — never a restart (R-835). An invalid daemon.json is
|
||||
left alone (R16); a reload that does not enable it puts the old file back. The leg runs it once before a Docker step;
|
||||
`--selftest=live-restore -vmid N` runs it by hand. Measured: demo-hp 24 containers, demo-felhom 5 — the same ids after.
|
||||
- **The Docker engine slow lane** (`11` §5.8). Wrapper layer `docker`, lane `slow` only, the six Docker packages only,
|
||||
origin `Docker CE` only (R2; a Docker package in a fast-lane plan is refused). **R3 — the wrapper checks the authority
|
||||
itself**, never the agent's config: a ring-1 step or ANY undo needs a signed `os_docker_step` it verifies with
|
||||
`ssh-keygen -Y verify` against the ROOT-owned `/etc/felhom/operator-signers` (namespace `felhom-op-v1`, the blob's
|
||||
`key_id`), bound to `/etc/felhom/os-trust.json` `host_id`, inside its time window, never replayed (a root-owned nonce
|
||||
file), with exactly the signed packages and undo flag; an unsigned ring-0 step needs that file's
|
||||
`"ring0_slow_lane": true` (the demo boxes only, set by hand). **R15:** live-restore must be on. An undo may downgrade
|
||||
(`--allow-downgrades`) only inside a signed job. The leg: ring 0 runs the Docker step at night after a healthy guest
|
||||
and host step (`select pending-docker`); ring 1 never does — only `DockerStepExecutor` (signed job, heavy-op gate).
|
||||
**Health:** the guest rule + every container running at the start has the SAME id after + the engine reports the
|
||||
installed version; a changed id is `health_failed`. Measured ring 0: 29.7.x → 29.8.2 on both demo boxes, every id kept.
|
||||
- **The version report** (R-852, decision 89). Wrapper mode `facts` (read-only): host Debian, running and next-boot
|
||||
kernel (`next_entry` > saved default > newest installed, by dpkg order), held packages (R-848), kernel taint (oops,
|
||||
warn), `kernel.panic`, the crash guard state; guest Debian, Docker engine, containerd, live-restore. The host report
|
||||
gains `system {pve_version, kernel_version, vmid, facts, facts_error}`, read at most every 10 min (~2 s);
|
||||
`--selftest=os-facts -vmid N`. A value nobody could read is `unknown`.
|
||||
- **R-849:** the guest is scanned for "restart needed" on every pass too, so a guest restart clears it.
|
||||
- **The crash guard** (decision 88, R-851): `configs/felhom-crash-guard` + `felhom-crash-guard.service` (early boot;
|
||||
its ExecStop writes a clean-stop marker) + an hourly re-arm timer + `/etc/felhom/crash-guard.conf`. A boot without
|
||||
the marker followed an unclean stop (a crash, a power cut or a hard reset — pstore saved nothing for a real panic on
|
||||
demo-hp, so they cannot be told apart). Armed: `kernel.panic = 10`. After the 2nd unclean boot within 60 min it
|
||||
TRIPS (`kernel.panic = 0`), so the 3rd crash within the hour leaves the box off; it re-arms after 24 h of normal
|
||||
running or `felhom-crash-guard rearm`. State in `/var/lib/felhom-crash-guard/state.json` (0644; read by facts).
|
||||
- **After the tag (main only, not shipped to boxes): `configs/build-golden.sh` 3.1.0** — the golden's `daemon.json`
|
||||
carries `"live-restore": true` with a fail-closed assertion, and `GOLDEN_DOCKER_PKGS` pins the approved Docker engine
|
||||
set (all six `name=version`); without it the bake log warns that the set is the newest, not an approved one.
|
||||
- Tests: wrapper 76 (DockerLane, LiveRestore, Facts, RealSignatureCheck with a throwaway key), crash guard 9, Go leg +
|
||||
executor; red-proofs `felhom.eu/documentation/audits/os-docker-crash-2026-10-04/partB/agent-redproofs.txt` (17 caught).
|
||||
|
||||
## v0.141.1 — "reboot needed" is true on the host (found live on demo-felhom, 2026-10-04)
|
||||
|
||||
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.141.1` (`a6bc3f1`), sha256
|
||||
> `b712f577099fd2d374f648df1e825874821302fe1b46e93c7b4b70044fbe84b5`, verified by download. Not vouched at release time.
|
||||
|
||||
**MinAgent impact:** none. **Pairs with hub v0.131.1** (reads `reboot_scanned`); an older hub ignores the field.
|
||||
|
||||
A patch release in the same session as v0.141.0 — a deliberate exception to "one release per repo": v0.141.0's host
|
||||
"reboot needed" was wrong in two ways, and it feeds an operator alarm.
|
||||
- **The scan hid `lxc-start`.** The host scan skipped every process whose cgroup line contains `lxc`, to leave out the
|
||||
guests' own processes. `lxc-start` lives in `0::/lxc.monitor/<vmid>`, so it was skipped too. Measured: after a
|
||||
108-package host pass (libc6 included) `lxc-start` mapped 20 deleted files and the report said `reboot_needed:
|
||||
false`. The pattern is now `:/lxc/` (`RESTART_SKIP_CGROUP`), pinned by a test that runs `grep` against the measured
|
||||
cgroup lines.
|
||||
- **A reboot never cleared it.** v0.141.0 scanned only after an install (R-845). The host now scans on EVERY pass (it
|
||||
is local, no `pct exec`); the guest still scans only after an install. The report carries `reboot_scanned`.
|
||||
- Red-proofs: `felhom.eu/documentation/audits/os-host-lane-2026-10-04/partB/live-defects-redproofs.txt`.
|
||||
|
||||
## v0.141.0 — OS updates: the host fast lane (`11` §8 step 3); the tunnel status is true (R-841); the leg is fast (R-845)
|
||||
|
||||
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.141.0` (`cfba0d0`), sha256
|
||||
> `6eaad9809613c4fca64aefc30cc16415528499bf48efe3a1daaa08af8a611aff`, verified by download. Not vouched at release time.
|
||||
|
||||
**MinAgent impact:** none required by any controller. **Needs hub v0.131.0** (`host_release`, the tunnel's three
|
||||
states, layer-tagged OS reports). With an older hub the host step finds no host release (ring 1 → nothing) and the
|
||||
tunnel's `detail` is ignored. Reads the cloudflared health check controller v0.292.0 adds; with an older controller
|
||||
the tunnel is judged on the container state alone and `detail` says so.
|
||||
|
||||
- **R-841 — the tunnel.** The agent used to run `systemctl is-active cloudflared` on the HOST — a unit that does not
|
||||
exist (cloudflared is a container in the customer guest), so every box reported `inactive`. `GuestTunnelProber`
|
||||
now reads the guest's `cloudflared` container through the EXISTING sudoers line (`pct exec N -- docker inspect -f
|
||||
*`): state, exit code, and the Docker health status. Three states: `running` (healthy), `not_running` (stopped,
|
||||
absent, or running but NOT connected), `unknown` (could not ask, or the check is still starting). `detail` says why.
|
||||
- **The host step** (`11` §8 step 3). After a healthy guest step, under the same heavy-op gate, the leg runs the same
|
||||
wrapper with `layer: host`. Debian origin only; never kernel, boot or firmware packages (new refusal **R14**);
|
||||
**only on an appliance** (**R12** lifted for the host fast lane: proof is the ROOT-owned install record
|
||||
`/var/lib/felhom-install/state.json` `mode: appliance` — the agent-writable `agent.json` is not trusted for this).
|
||||
A failed or unhealthy guest step skips it. **Host health rule:** the agent, pveproxy, pvedaemon, pvestatd and
|
||||
pve-cluster are active; the customer guest runs; the guest health rule passes; the tunnel is `running` (an
|
||||
`unknown` tunnel does not fail the rule; `not_running` does). Never reboots: "reboot needed" is reported when PID 1
|
||||
or `lxc-start` runs a replaced library. No automatic undo — the by-hand runbook is `os-updates-host-undo.md`.
|
||||
- **R-845 — the leg is fast.** The wrapper asked about each package in its own `pct exec` (about 0.9 s each). It now
|
||||
makes one call per layer for the version checks (`apt-cache madison` for all names at once, `dpkg --compare-versions`
|
||||
on the host), scans for restart-needed only after an install, repairs only when `dpkg --audit` reports something,
|
||||
and reports its own `pass_seconds`.
|
||||
- **Wrapper tests:** `configs/test_felhom_os_apply.py` 46 tests (host layer, R12, R14, lxc-start, the speed rules).
|
||||
Red-proofs: `felhom.eu/documentation/audits/os-host-lane-2026-10-04/partA/`, `partB/`, `partC/agent-golden-redproof.txt`.
|
||||
- **Contract:** the desired-state golden gains `host_release` (byte-identical with the hub's); `TestOSUpdateGolden_Decodes`
|
||||
checks it.
|
||||
|
||||
## v0.140.0 — OS updates, guest fast lane (`11-os-updates.md` §8 step 2; `09` §3 decisions 76, 79, 80)
|
||||
|
||||
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.140.0` (`9cac346`), sha256
|
||||
|
||||
@@ -1,11 +1,9 @@
|
||||
# REPORT — 2026-10-04: v0.140.0, OS updates (guest fast lane)
|
||||
# REPORT — 2026-10-04: v0.142.0, Docker slow lane + version report + crash guard
|
||||
|
||||
Full session report: `felhom.eu/REPORT-os-guest-lane-2026-10-04.md`.
|
||||
Full session report: `felhom.eu/REPORT-os-docker-crash-2026-10-04.md`.
|
||||
|
||||
- `felhom-os-apply` wrapper (R1–R13, repair first, snapshot.debian.org fallback), `FELHOM_OSAPPLY` sudoers, the OS leg
|
||||
after the primary backup, `--selftest=os-update`. Released `9cac346`, sha256 `ae2d60b7…1250`, verified by download.
|
||||
- Live on both demo boxes: ring 0 installed 53 packages each, healthy; ring 1 installed exactly the 3 approved versions;
|
||||
a deliberately failed health check reported `health_failed` and mailed the operator.
|
||||
- Found and fixed live: the `--selftest` flag refused `os-update` (and `wgtunnel`, since S3); the conffile log line
|
||||
called an updated file "kept"; an app stopped between the inventory and the apply escaped the health check.
|
||||
- No automatic undo (R-837: PVE refuses a snapshot of a guest with host-path binds).
|
||||
- Docker live-restore turned on by reload (never a restart); the Docker engine set as a slow lane whose authority the
|
||||
root wrapper checks itself (signed job against a root-owned key file, or the root-owned ring-0 mark); same-id health.
|
||||
- The box reports its versions (Proxmox, kernels, Debian, Docker, live-restore, held packages, taint, crash guard).
|
||||
- The crash guard: a crashed host restarts, the 3rd unclean stop within an hour leaves it off, 24 h re-arm.
|
||||
- Tests and 17 red-proofs; live on both demo boxes (see the session report).
|
||||
|
||||
+147
-40
@@ -44,11 +44,11 @@ import (
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/localapi"
|
||||
applog "gitea.dooplex.hu/admin/felhom-agent/internal/log"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/mgmtplane"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/pbs"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/pbsdr"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/poke"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/provision"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/restorespace"
|
||||
@@ -245,6 +245,10 @@ func main() {
|
||||
os.Exit(runSelftestRestoreTestDue(context.Background(), cfg, logger))
|
||||
case "os-update":
|
||||
os.Exit(runSelftestOSUpdate(context.Background(), cfg, logger, vmid))
|
||||
case "os-facts":
|
||||
os.Exit(runSelftestFacts(context.Background(), cfg, logger, vmid))
|
||||
case "live-restore":
|
||||
os.Exit(runSelftestLiveRestore(context.Background(), cfg, logger, vmid))
|
||||
case "pbs-verify":
|
||||
os.Exit(runSelftestPBSVerify(context.Background(), cfg, logger))
|
||||
case "lanresolver":
|
||||
@@ -845,6 +849,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
// OS updates, guest fast lane (agent v0.140.0, `11-os-updates.md` §8 step 2): the leg consumes the hub's
|
||||
// os_update block and runs after each successful primary whole-guest backup (wired on the local API below).
|
||||
osLeg := newOSLeg(cfg, client, px, logger)
|
||||
collector.SetSystemReporter(&factsReporter{leg: osLeg, guest: firstGuest(px)}) // R-852: the versions
|
||||
desiredSyncer.AddConsumer(osLeg)
|
||||
// S5: consume a host_loss restore_directive into an inspectable restore PLAN (derive + surface,
|
||||
// execute nothing). The recipe is fetched on-demand (rare directive) via a fresh Collect.
|
||||
@@ -1087,7 +1092,16 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
// recurring clobber before it locks the box out. Port 22 for G1 (H1 passes the felhom-sshd port).
|
||||
collector.SetMgmtPlaneReporter(mgmtplane.NewReporter(mgmtplane.DefaultPrivsepDir, mgmtplane.DefaultHealMarker, mgmtplane.DefaultSshdPort))
|
||||
|
||||
jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec}, cfg.Hub.HostID, logger)
|
||||
// Agent v0.142.0: a signed Docker engine step (`11` §5.8) — ring 1 and every undo; under the heavy-op gate.
|
||||
dockerExec := osupdate.DockerStepExecutor{Leg: osLeg, Guest: firstGuest(px),
|
||||
Gate: func(ctx context.Context) (func(), error) {
|
||||
release, busy, ok := heavyOps.TryAcquire("os-docker-step")
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("busy: %s", busy)
|
||||
}
|
||||
return release, nil
|
||||
}}
|
||||
jobsRunner := signedjobs.NewRunner(client, gate, signedjobs.ExecutorChain{wipeExec, decommExec, updateExec, dockerExec}, cfg.Hub.HostID, logger)
|
||||
loop.SetEnvelopeObserver(hub.MultiObserver(desiredSyncer, jobsRunner))
|
||||
|
||||
// Controller-driven escrow ceremony (v0.88.0): static config facts + the LATE-BOUND DR gate —
|
||||
@@ -1127,7 +1141,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
return
|
||||
case <-time.After(90 * time.Second):
|
||||
}
|
||||
_, _ = osLeg.Run(ctx, vmid, "night")
|
||||
_ = osLeg.Run(ctx, vmid, "night")
|
||||
})
|
||||
}
|
||||
if localTokens != nil {
|
||||
@@ -1869,7 +1883,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
|
||||
ConfigPath: cfg.SourcePath,
|
||||
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
|
||||
SmbCredsDir: cfg.Privileged.SmbCredsDir,
|
||||
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
|
||||
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
|
||||
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
|
||||
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
|
||||
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
|
||||
@@ -3498,10 +3512,12 @@ func (f *selftestFlag) Set(v string) error {
|
||||
f.mode = "controller-swap"
|
||||
case "os-update":
|
||||
f.mode = "os-update"
|
||||
case "os-facts", "live-restore": // agent v0.142.0
|
||||
f.mode = v
|
||||
case "wgtunnel": // dispatched since S3 but refused here until 2026-10-04 (TestSelftestFlag_AcceptsEveryDispatchedMode)
|
||||
f.mode = "wgtunnel"
|
||||
default:
|
||||
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|restore-test-due|pbs-verify|lanresolver|wgtunnel|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap|os-update)", v)
|
||||
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|restore-test-due|pbs-verify|lanresolver|wgtunnel|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap|os-update|os-facts|live-restore)", v)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
@@ -3558,29 +3574,71 @@ func runSelftestOSUpdate(ctx context.Context, cfg config.Config, logger *slog.Lo
|
||||
fmt.Printf("=== felhom-agent %s selftest=os-update vmid=%d ring=%d enabled=%v guest-release=%v host-release=%v appliance=%v ===\n",
|
||||
version, vmid, b.Ring, b.Enabled, b.Release != nil, b.HostRelease != nil, leg.Appliance)
|
||||
start := time.Now()
|
||||
g, h := leg.Run(ctx, vmid, "debug")
|
||||
for _, rep := range []osupdate.Report{g, h} {
|
||||
pass := leg.Run(ctx, vmid, "debug")
|
||||
worst := pass.Guest
|
||||
for i, rep := range []osupdate.Report{pass.Guest, pass.Host, pass.Docker} {
|
||||
if rep.Layer == "" {
|
||||
fmt.Println(" host step: skipped (see the log line above)")
|
||||
fmt.Printf(" %s step: skipped (see the log line above)\n", []string{"guest", "host", "docker"}[i])
|
||||
continue
|
||||
}
|
||||
printJSON("os-update report ("+rep.Layer+")", map[string]any{"run_id": rep.RunID, "ring": rep.Ring, "release_id": rep.ReleaseID,
|
||||
"mode": rep.Mode, "outcome": rep.Outcome, "healthy": rep.Healthy, "health_reason": rep.HealthReason,
|
||||
"upgraded": rep.Upgraded, "pending": len(rep.Pending), "not_covered": rep.NotCovered,
|
||||
"restart_needed": rep.RestartNeeded, "reboot_needed": rep.RebootNeeded, "wrapper_seconds": rep.PassSeconds, "refused": rep.Refused})
|
||||
"restart_needed": rep.RestartNeeded, "reboot_needed": rep.RebootNeeded, "wrapper_seconds": rep.PassSeconds,
|
||||
"refused": rep.Refused, "docker_engine": rep.DockerEngine, "authority": rep.Authority})
|
||||
if !(rep.Outcome == "applied" || rep.Outcome == "nothing" || rep.Outcome == "inventory" || rep.Outcome == "skipped") {
|
||||
worst = rep
|
||||
}
|
||||
}
|
||||
fmt.Printf(" pass took %s\n", time.Since(start).Round(100*time.Millisecond))
|
||||
rep := g
|
||||
if h.Layer != "" && !(h.Outcome == "applied" || h.Outcome == "nothing" || h.Outcome == "inventory") {
|
||||
rep = h
|
||||
}
|
||||
switch rep.Outcome {
|
||||
switch worst.Outcome {
|
||||
case "applied", "nothing", "inventory", "skipped":
|
||||
return 0
|
||||
}
|
||||
return 1
|
||||
}
|
||||
|
||||
// runSelftestFacts prints the versions the host report carries (R-852, agent v0.142.0) — read-only.
|
||||
//
|
||||
// sudo -u felhom-agent felhom-agent --config … --selftest=os-facts -vmid 9201
|
||||
func runSelftestFacts(ctx context.Context, cfg config.Config, logger *slog.Logger, vmid int) int {
|
||||
if vmid <= 0 {
|
||||
fmt.Fprintln(os.Stderr, "selftest=os-facts: -vmid is required")
|
||||
return 2
|
||||
}
|
||||
px, _ := newProxmoxClient(cfg)
|
||||
leg := newOSLeg(cfg, nil, px, logger)
|
||||
start := time.Now()
|
||||
f, err := leg.Facts(ctx, vmid)
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, "selftest=os-facts:", err)
|
||||
return 1
|
||||
}
|
||||
var v any
|
||||
_ = json.Unmarshal(f, &v)
|
||||
printJSON(fmt.Sprintf("facts (vmid %d, %s)", vmid, time.Since(start).Round(100*time.Millisecond)), v)
|
||||
return 0
|
||||
}
|
||||
|
||||
// runSelftestLiveRestore is the ONE-TIME live-restore step (`09` decision 87) as a debug action — the night leg does
|
||||
// the same before a ring-0 Docker step. It prints the container ids before and after (they must not change).
|
||||
//
|
||||
// sudo -u felhom-agent felhom-agent --config … --selftest=live-restore -vmid 9202
|
||||
func runSelftestLiveRestore(ctx context.Context, cfg config.Config, logger *slog.Logger, vmid int) int {
|
||||
if vmid <= 0 {
|
||||
fmt.Fprintln(os.Stderr, "selftest=live-restore: -vmid is required")
|
||||
return 2
|
||||
}
|
||||
px, _ := newProxmoxClient(cfg)
|
||||
leg := newOSLeg(cfg, nil, px, logger)
|
||||
if err := leg.EnsureLiveRestore(ctx, time.Now().UTC().Format("20060102T150405Z"), vmid); err != nil {
|
||||
fmt.Fprintln(os.Stderr, "selftest=live-restore:", err)
|
||||
return 1
|
||||
}
|
||||
fmt.Println("live-restore: on (see the wrapper's LIVE-RESTORE line above for the container ids)")
|
||||
return 0
|
||||
}
|
||||
|
||||
// newTunnelProber reads the box's REAL tunnel (R-841, agent v0.141.0): the cloudflared container in each running
|
||||
// customer guest — a guest that binds /mnt/felhom-drives, the same rule the OS wrapper's R10 uses — through the
|
||||
// existing `pct exec [0-9]* -- docker inspect -f *` sudoers line.
|
||||
@@ -3591,31 +3649,80 @@ func newTunnelProber(cfg config.Config, px *proxmox.Client) hub.CloudflaredProbe
|
||||
}
|
||||
return hub.GuestTunnelProber{
|
||||
Runner: &proxmox.ExecRunner{Mode: mode, SudoPath: cfg.Privileged.SudoPath},
|
||||
Guests: func(ctx context.Context) ([]int, error) {
|
||||
if px == nil {
|
||||
return nil, fmt.Errorf("no proxmox client")
|
||||
}
|
||||
gs, err := px.ListLXC(ctx)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
var out []int
|
||||
for _, g := range gs {
|
||||
if g.Status != "running" {
|
||||
continue
|
||||
}
|
||||
gc, err := px.GuestConfig(ctx, g.VMID)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
for _, v := range gc.MountPoints() {
|
||||
if src, _, _ := strings.Cut(v, ","); src == "/mnt/felhom-drives" {
|
||||
out = append(out, g.VMID)
|
||||
break
|
||||
}
|
||||
}
|
||||
}
|
||||
return out, nil
|
||||
},
|
||||
Guests: customerGuests(px),
|
||||
}
|
||||
}
|
||||
|
||||
// customerGuests lists the running guests that bind /mnt/felhom-drives — the box's customer guest(s), the same rule the
|
||||
// OS wrapper's R10 uses. Shared by the tunnel probe, the facts read and the signed Docker step.
|
||||
func customerGuests(px *proxmox.Client) func(ctx context.Context) ([]int, error) {
|
||||
return func(ctx context.Context) ([]int, error) {
|
||||
if px == nil {
|
||||
return nil, fmt.Errorf("no proxmox client")
|
||||
}
|
||||
gs, err := px.ListLXC(ctx)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
var out []int
|
||||
for _, g := range gs {
|
||||
if g.Status != "running" {
|
||||
continue
|
||||
}
|
||||
gc, err := px.GuestConfig(ctx, g.VMID)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
for _, v := range gc.MountPoints() {
|
||||
if src, _, _ := strings.Cut(v, ","); src == "/mnt/felhom-drives" {
|
||||
out = append(out, g.VMID)
|
||||
break
|
||||
}
|
||||
}
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
}
|
||||
|
||||
// firstGuest is the single customer guest (an error when there is none).
|
||||
func firstGuest(px *proxmox.Client) func(ctx context.Context) (int, error) {
|
||||
f := customerGuests(px)
|
||||
return func(ctx context.Context) (int, error) {
|
||||
v, err := f(ctx)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if len(v) == 0 {
|
||||
return 0, fmt.Errorf("no running customer guest")
|
||||
}
|
||||
return v[0], nil
|
||||
}
|
||||
}
|
||||
|
||||
// factsReporter feeds the host report's `system` stanza (R-852, agent v0.142.0) from the wrapper's read-only facts
|
||||
// mode, at most every 10 minutes (each read is ~2 s of pct exec; the host reports every 15 min).
|
||||
type factsReporter struct {
|
||||
leg *osupdate.Leg
|
||||
guest func(ctx context.Context) (int, error)
|
||||
mu sync.Mutex
|
||||
at time.Time
|
||||
vmid int
|
||||
facts json.RawMessage
|
||||
err error
|
||||
}
|
||||
|
||||
func (f *factsReporter) SystemFacts(ctx context.Context) (int, json.RawMessage, error) {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
if !f.at.IsZero() && time.Since(f.at) < 10*time.Minute {
|
||||
return f.vmid, f.facts, f.err
|
||||
}
|
||||
f.at = time.Now()
|
||||
f.vmid, f.err = f.guest(ctx)
|
||||
if f.err != nil {
|
||||
f.facts = nil
|
||||
return 0, nil, f.err
|
||||
}
|
||||
f.facts, f.err = f.leg.Facts(ctx, f.vmid)
|
||||
return f.vmid, f.facts, f.err
|
||||
}
|
||||
|
||||
+21
-3
@@ -61,7 +61,7 @@ set -euo pipefail
|
||||
|
||||
# Script provenance — logged into every bake transcript next to the baked controller tag, so an
|
||||
# archive can always be traced to the script that produced it. Bump on any behavior change.
|
||||
GOLDEN_SCRIPT_VERSION="3.0.0"
|
||||
GOLDEN_SCRIPT_VERSION="3.1.0"
|
||||
|
||||
VMID="${1:-9100}"
|
||||
TEMPLATE="${2:-local:vztmpl/debian-13-standard_13.1-2_amd64.tar.zst}"
|
||||
@@ -112,7 +112,15 @@ for i in $(seq 1 30); do
|
||||
if pct exec "$VMID" -- getent hosts download.docker.com >/dev/null 2>&1; then break; fi
|
||||
sleep 1
|
||||
done
|
||||
pct exec "$VMID" -- bash -c '
|
||||
# v3.1.0 (`11` §5.8): GOLDEN_DOCKER_PKGS pins the APPROVED Docker engine set (the hub's newest Docker release, all six
|
||||
# "name=version"); without it the newest stable set is installed and the bake log says so.
|
||||
GOLDEN_DOCKER_PKGS="${GOLDEN_DOCKER_PKGS:-}"
|
||||
if [[ -n "$GOLDEN_DOCKER_PKGS" ]]; then
|
||||
echo "[golden] Docker engine set PINNED to the approved release: $GOLDEN_DOCKER_PKGS"
|
||||
else
|
||||
echo "[golden] WARNING: GOLDEN_DOCKER_PKGS not set — installing the newest stable Docker set, not an approved one"
|
||||
fi
|
||||
pct exec "$VMID" -- env GOLDEN_DOCKER_PKGS="$GOLDEN_DOCKER_PKGS" bash -c '
|
||||
set -e
|
||||
export DEBIAN_FRONTEND=noninteractive
|
||||
apt-get update -qq
|
||||
@@ -122,7 +130,12 @@ pct exec "$VMID" -- bash -c '
|
||||
echo "deb [signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/debian trixie stable" \
|
||||
> /etc/apt/sources.list.d/docker.list
|
||||
apt-get update -qq
|
||||
apt-get install -y -qq docker-ce docker-ce-cli containerd.io >/dev/null
|
||||
if [ -n "$GOLDEN_DOCKER_PKGS" ]; then
|
||||
apt-get install -y -qq $GOLDEN_DOCKER_PKGS >/dev/null
|
||||
else
|
||||
apt-get install -y -qq docker-ce docker-ce-cli containerd.io >/dev/null
|
||||
fi
|
||||
dpkg-query -W containerd.io docker-buildx-plugin docker-ce docker-ce-cli docker-ce-rootless-extras docker-compose-plugin 2>/dev/null | sed "s/^/ installed: /"
|
||||
'
|
||||
echo "[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …"
|
||||
# containerd-snapshotter (Docker 28+/29 default) keeps the IMAGE content store under
|
||||
@@ -135,9 +148,12 @@ echo "[golden] baking daemon.json: classic overlay2 driver (containerd-snapshott
|
||||
# layout (a container's `df /` reports the single volume, phase-0 spike). Since v3.0.0 /var/lib/docker
|
||||
# is a BIND of <volume>/docker rather than the mp0 mount itself, wired immediately below; data-root
|
||||
# still needs no override because the path is unchanged. Log caps kill the most common runaway.
|
||||
# v3.1.0: "live-restore": true (`09` decision 87) — a Docker engine update then restarts no app (`11` C5). A box made
|
||||
# from this golden never needs the agent's one-time live-restore-on step. NEVER removed by a plain restart (R-835).
|
||||
pct exec "$VMID" -- bash -c 'mkdir -p /etc/docker; cat > /etc/docker/daemon.json <<JSON
|
||||
{
|
||||
"features": { "containerd-snapshotter": false },
|
||||
"live-restore": true,
|
||||
"log-driver": "json-file",
|
||||
"log-opts": { "max-size": "10m", "max-file": "3" }
|
||||
}
|
||||
@@ -175,6 +191,8 @@ pct exec "$VMID" -- bash -c 'systemctl restart docker; sleep 3; docker run --rm
|
||||
# Guard: the image store MUST be on the data volume now. /var/lib/containerd holding the images would
|
||||
# mean containerd-snapshotter is still on (the split would leave images on the rootfs).
|
||||
pct exec "$VMID" -- bash -c 'drv=$(docker info 2>/dev/null | sed -n "s/.*Storage Driver: //p"); [ "$drv" = "overlay2" ] || { echo "[golden] FATAL: storage driver is $drv, expected overlay2 — images would not land on the data volume"; exit 1; }'
|
||||
# v3.1.0 ASSERTION: live-restore is ON in the running daemon (decision 87), or the bake fails closed.
|
||||
pct exec "$VMID" -- bash -c 'lr=$(docker info --format "{{.LiveRestoreEnabled}}" 2>/dev/null); [ "$lr" = "true" ] && echo " live-restore: on" || { echo "[golden] FATAL: live-restore is $lr, expected true (decision 87)"; exit 1; }'
|
||||
# ASSERTION 1 (RETARGETED v3.0.0, not removed). /var/lib/docker must be a real mount — now the V-c
|
||||
# bind of <volume>/docker rather than the mp0 mount itself. Still fails closed on the same failure:
|
||||
# if the bind did not take, Docker's data-root silently sits on the OS rootfs and the golden ships
|
||||
|
||||
@@ -0,0 +1,8 @@
|
||||
# /etc/felhom/crash-guard.conf — read by /usr/local/sbin/felhom-crash-guard (`11` §5.9).
|
||||
# The LIMIT-th unclean stop within WINDOW_MINUTES leaves the box off. Decided by CC unattended — operator may reverse.
|
||||
LIMIT=3
|
||||
WINDOW_MINUTES=60
|
||||
# kernel.panic while armed: seconds after a crash before the kernel restarts the box.
|
||||
PANIC_SECONDS=10
|
||||
# A tripped guard re-arms after this many hours of normal running (or `felhom-crash-guard rearm`).
|
||||
REARM_HOURS=24
|
||||
@@ -0,0 +1,235 @@
|
||||
#!/usr/bin/python3
|
||||
# felhom-crash-guard — a crashed host restarts by itself, but not forever (`09` decision 88, R-851, `11` §5.9).
|
||||
#
|
||||
# Install as /usr/local/sbin/felhom-crash-guard (0755 root:root), with felhom-crash-guard.service (boot / clean-stop)
|
||||
# and felhom-crash-guard-check.timer (hourly re-arm check). Python 3, standard library only.
|
||||
# Tests: configs/test_felhom_crash_guard.py (temp dirs; nothing real is touched).
|
||||
#
|
||||
# WHAT IT DOES
|
||||
# boot early at every boot. Was the previous boot ended CLEANLY? (the clean-stop marker exists). If not, this
|
||||
# boot follows an UNCLEAN stop — a kernel crash, a power cut or a hard reset (they cannot be told apart
|
||||
# on these boxes: measured 2026-10-04 on demo-hp, efi_pstore is on yet saved NOTHING for a real panic;
|
||||
# the journal and `last` show only "no shutdown"). It records the unclean boot, counts those in the last
|
||||
# WINDOW_MINUTES, and sets kernel.panic:
|
||||
# - fewer than LIMIT-1 recent unclean boots → kernel.panic = PANIC_SECONDS (a crash restarts the box);
|
||||
# - LIMIT-1 or more → the guard TRIPS: kernel.panic = 0, so the LIMIT-th crash within the window
|
||||
# leaves the box OFF (operator's own words: "if it crashes 3 times within one hour, it stays off").
|
||||
# A tripped guard stays tripped across further boots until it re-arms.
|
||||
# clean-stop ExecStop of the service: writes the clean-stop marker during an orderly shutdown or reboot.
|
||||
# check hourly: a tripped guard re-arms after REARM_HOURS of normal running (since the trip AND since boot).
|
||||
# rearm the operator re-arms by hand (`felhom-crash-guard rearm`).
|
||||
# status prints the state.
|
||||
# The state is /var/lib/felhom-crash-guard/state.json (0644: the non-root agent reads it into its host report).
|
||||
# Before the service runs (very early boot) the kernel default kernel.panic = 0 applies, so a crash THAT early leaves
|
||||
# the box off — the safe side: a box that cannot reach userspace must not loop.
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
|
||||
CONF = "/etc/felhom/crash-guard.conf"
|
||||
STATE_DIR = "/var/lib/felhom-crash-guard"
|
||||
DEFAULTS = {"LIMIT": 3, "WINDOW_MINUTES": 60, "PANIC_SECONDS": 10, "REARM_HOURS": 24}
|
||||
|
||||
|
||||
class Env:
|
||||
"""Paths and clock; tests replace them."""
|
||||
|
||||
def __init__(self, conf=CONF, state_dir=STATE_DIR, panic_path="/proc/sys/kernel/panic",
|
||||
uptime_path="/proc/uptime", boot_id_path="/proc/sys/kernel/random/boot_id"):
|
||||
self.conf, self.state_dir = conf, state_dir
|
||||
self.panic_path, self.uptime_path, self.boot_id_path = panic_path, uptime_path, boot_id_path
|
||||
|
||||
def now(self):
|
||||
return time.time()
|
||||
|
||||
def log(self, line):
|
||||
print(line, file=sys.stderr, flush=True)
|
||||
try:
|
||||
import subprocess
|
||||
subprocess.run(["logger", "-t", "felhom-crash-guard", line], timeout=10)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
def iso(t):
|
||||
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(t))
|
||||
|
||||
|
||||
def parse_iso(s):
|
||||
import calendar
|
||||
return calendar.timegm(time.strptime(s, "%Y-%m-%dT%H:%M:%SZ"))
|
||||
|
||||
|
||||
def load_conf(env):
|
||||
c = dict(DEFAULTS)
|
||||
try:
|
||||
for line in open(env.conf):
|
||||
line = line.strip()
|
||||
if not line or line.startswith("#") or "=" not in line:
|
||||
continue
|
||||
k, v = (x.strip() for x in line.split("=", 1))
|
||||
if k in c and v.isdigit() and int(v) >= (1 if k != "PANIC_SECONDS" else 1):
|
||||
c[k] = int(v)
|
||||
except OSError:
|
||||
pass
|
||||
return c
|
||||
|
||||
|
||||
def state_path(env):
|
||||
return os.path.join(env.state_dir, "state.json")
|
||||
|
||||
|
||||
def marker_path(env):
|
||||
return os.path.join(env.state_dir, "clean-stop")
|
||||
|
||||
|
||||
def load_state(env):
|
||||
try:
|
||||
with open(state_path(env)) as f:
|
||||
s = json.load(f)
|
||||
return s if isinstance(s, dict) else None
|
||||
except (OSError, ValueError):
|
||||
return None
|
||||
|
||||
|
||||
def save_state(env, s):
|
||||
os.makedirs(env.state_dir, mode=0o755, exist_ok=True)
|
||||
tmp = state_path(env) + ".tmp"
|
||||
with open(tmp, "w") as f:
|
||||
json.dump(s, f, indent=2, sort_keys=True)
|
||||
f.write("\n")
|
||||
os.chmod(tmp, 0o644)
|
||||
os.replace(tmp, state_path(env))
|
||||
|
||||
|
||||
def set_panic(env, seconds):
|
||||
with open(env.panic_path, "w") as f:
|
||||
f.write(f"{seconds}\n")
|
||||
|
||||
|
||||
def read(path, default=""):
|
||||
try:
|
||||
with open(path) as f:
|
||||
return f.read().strip()
|
||||
except OSError:
|
||||
return default
|
||||
|
||||
|
||||
def summarize(s, c, now):
|
||||
window = c["WINDOW_MINUTES"] * 60
|
||||
times = [parse_iso(t) for t in s.get("unclean_boots", [])]
|
||||
after = parse_iso(s["rearmed_at"]) if s.get("rearmed_at") else 0
|
||||
# a re-arm starts a fresh window (or the next unclean boot would trip again at once); the history stays
|
||||
s["unclean_boots_in_window"] = sum(1 for t in times if now - t <= window and t > after)
|
||||
s["unclean_boots_24h"] = sum(1 for t in times if now - t <= 86400)
|
||||
s["config"] = c
|
||||
s["updated_at"] = iso(now)
|
||||
|
||||
|
||||
def boot(env):
|
||||
c = load_conf(env)
|
||||
now = env.now()
|
||||
try:
|
||||
up = float(read(env.uptime_path, "0").split()[0])
|
||||
except (ValueError, IndexError):
|
||||
up = 0.0
|
||||
boot_at = now - up
|
||||
prev = load_state(env)
|
||||
first = prev is None
|
||||
s = prev or {"version": 1, "unclean_boots": [], "tripped": False}
|
||||
clean = os.path.exists(marker_path(env))
|
||||
unclean = (not first) and (not clean)
|
||||
try:
|
||||
os.remove(marker_path(env))
|
||||
except OSError:
|
||||
pass
|
||||
# keep 7 days of history (the 24 h figure and the operator's view), drop older
|
||||
s["unclean_boots"] = [t for t in s.get("unclean_boots", []) if now - parse_iso(t) <= 7 * 86400]
|
||||
if unclean:
|
||||
s["unclean_boots"].append(iso(boot_at))
|
||||
s["last_boot_at"] = iso(boot_at)
|
||||
s["last_boot_unclean"] = unclean
|
||||
s["boot_id"] = read(env.boot_id_path, "unknown")
|
||||
summarize(s, c, now)
|
||||
if not s.get("tripped") and s["unclean_boots_in_window"] >= c["LIMIT"] - 1:
|
||||
s["tripped"], s["tripped_at"] = True, iso(now)
|
||||
s["tripped_reason"] = (f"{s['unclean_boots_in_window']} unclean boots within {c['WINDOW_MINUTES']} minutes — "
|
||||
f"the next crash leaves the box off (limit {c['LIMIT']})")
|
||||
env.log(f"crash-guard: TRIPPED: {s['tripped_reason']}")
|
||||
panic = 0 if s.get("tripped") else c["PANIC_SECONDS"]
|
||||
set_panic(env, panic)
|
||||
s["kernel_panic"] = panic
|
||||
s["armed"] = not s.get("tripped")
|
||||
save_state(env, s)
|
||||
env.log(f"crash-guard: boot first={first} unclean={unclean} in-window={s['unclean_boots_in_window']} "
|
||||
f"tripped={s.get('tripped')} kernel.panic={panic}")
|
||||
return 0
|
||||
|
||||
|
||||
def clean_stop(env):
|
||||
os.makedirs(env.state_dir, mode=0o755, exist_ok=True)
|
||||
with open(marker_path(env), "w") as f:
|
||||
f.write(iso(env.now()) + "\n")
|
||||
env.log("crash-guard: clean stop recorded")
|
||||
return 0
|
||||
|
||||
|
||||
def rearm(env, by):
|
||||
c = load_conf(env)
|
||||
now = env.now()
|
||||
s = load_state(env) or {"version": 1, "unclean_boots": []}
|
||||
was = bool(s.get("tripped"))
|
||||
s["tripped"] = False
|
||||
s["armed"] = True
|
||||
s["rearmed_at"], s["rearmed_by"] = iso(now), by
|
||||
if was:
|
||||
s["last_trip"] = {"at": s.get("tripped_at"), "reason": s.get("tripped_reason")}
|
||||
s.pop("tripped_at", None)
|
||||
s.pop("tripped_reason", None)
|
||||
summarize(s, c, now)
|
||||
set_panic(env, c["PANIC_SECONDS"])
|
||||
s["kernel_panic"] = c["PANIC_SECONDS"]
|
||||
save_state(env, s)
|
||||
env.log(f"crash-guard: RE-ARMED by {by} (was tripped: {was}); kernel.panic={c['PANIC_SECONDS']}")
|
||||
return 0
|
||||
|
||||
|
||||
def check(env):
|
||||
c = load_conf(env)
|
||||
now = env.now()
|
||||
s = load_state(env)
|
||||
if not s:
|
||||
return 0
|
||||
if s.get("tripped"):
|
||||
since = max(parse_iso(s["tripped_at"]), parse_iso(s.get("last_boot_at", s["tripped_at"])))
|
||||
if now - since >= c["REARM_HOURS"] * 3600:
|
||||
return rearm(env, f"timer ({c['REARM_HOURS']} h of normal running)")
|
||||
summarize(s, c, now)
|
||||
save_state(env, s)
|
||||
return 0
|
||||
|
||||
|
||||
def main(argv, env=None):
|
||||
env = env or Env()
|
||||
cmd = argv[1] if len(argv) == 2 else ""
|
||||
if cmd == "boot":
|
||||
return boot(env)
|
||||
if cmd == "clean-stop":
|
||||
return clean_stop(env)
|
||||
if cmd == "check":
|
||||
return check(env)
|
||||
if cmd == "rearm":
|
||||
return rearm(env, "operator")
|
||||
if cmd == "status":
|
||||
print(json.dumps(load_state(env), indent=2, sort_keys=True))
|
||||
return 0
|
||||
print("usage: felhom-crash-guard boot|clean-stop|check|rearm|status", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
if os.geteuid() != 0:
|
||||
print("felhom-crash-guard: must run as root", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
sys.exit(main(sys.argv))
|
||||
@@ -0,0 +1,7 @@
|
||||
# Hourly: a tripped crash guard re-arms after REARM_HOURS of normal running (felhom-crash-guard check).
|
||||
[Unit]
|
||||
Description=Felhom crash guard re-arm check
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=/usr/local/sbin/felhom-crash-guard check
|
||||
@@ -0,0 +1,9 @@
|
||||
[Unit]
|
||||
Description=Felhom crash guard re-arm check (hourly)
|
||||
|
||||
[Timer]
|
||||
OnBootSec=15min
|
||||
OnUnitActiveSec=1h
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
@@ -0,0 +1,19 @@
|
||||
# felhom-crash-guard — a crashed host restarts by itself, with a limit (`09` decision 88, R-851, `11` §5.9).
|
||||
# Starts early at boot (sets kernel.panic for THIS boot); its ExecStop writes the clean-stop marker during an orderly
|
||||
# shutdown or reboot. A boot that finds no marker followed a crash, a power cut or a hard reset.
|
||||
[Unit]
|
||||
Description=Felhom crash guard (restart after a kernel crash, with a limit)
|
||||
DefaultDependencies=no
|
||||
After=local-fs.target
|
||||
Before=sysinit.target shutdown.target
|
||||
Conflicts=shutdown.target
|
||||
RequiresMountsFor=/var/lib
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
RemainAfterExit=yes
|
||||
ExecStart=/usr/local/sbin/felhom-crash-guard boot
|
||||
ExecStop=/usr/local/sbin/felhom-crash-guard clean-stop
|
||||
|
||||
[Install]
|
||||
WantedBy=sysinit.target
|
||||
+374
-24
@@ -14,7 +14,13 @@
|
||||
# file, including the snapshot.debian.org fallback (decision 79). Nothing here is overridable from the environment.
|
||||
#
|
||||
# LAYERS (agent v0.141.0): "guest" (the customer LXC, entered with `pct exec`) and "host" (this Proxmox host, run
|
||||
# directly). LANE: "fast" only — the slow lane (kernel, Proxmox, Docker) is REFUSED (R3, R14) until `11` §8 steps 5–6.
|
||||
# directly). LANE: "fast" for those two. Agent v0.142.0 adds the layer "docker" (the guest's Docker engine set, `11`
|
||||
# §5.8), which is the SLOW lane: lane "slow" only, the six Docker packages only, origin "Docker CE" only, and only
|
||||
# with an authority this file checks ITSELF (R3): a signed operator job verified with `ssh-keygen -Y verify` against
|
||||
# the ROOT-OWNED signers file (TRUST_SIGNERS), bound to this host (TRUST_FILE host_id), unexpired and never replayed;
|
||||
# or, for an unsigned ring-0 step, the root-owned TRUST_FILE saying `"ring0_slow_lane": true` (set by hand on the demo
|
||||
# boxes only). The agent's own config is NOT trusted for either: the agent can write it. A Docker step also needs
|
||||
# `live-restore` ON (R15) — without it every container restarts.
|
||||
#
|
||||
# Modes (plan field "mode"):
|
||||
# inventory `apt-get update`, then report what is installed (with origin), what is pending, and health.
|
||||
@@ -22,6 +28,10 @@
|
||||
# pending Debian / Debian-Security upgrade, for ring 0), check every refusal on an `apt-get -s`
|
||||
# simulation of EXACTLY name=version, install, clean, scan for restart-needed, report as inventory.
|
||||
# health report health only (the agent polls it after a run).
|
||||
# facts (v0.142.0) read-only versions for the hub's System page: host Debian, running and next-boot kernel,
|
||||
# held packages, kernel taint, the crash guard; guest Debian, Docker engine, containerd, live-restore.
|
||||
# live-restore-on (v0.142.0, layer guest) the ONE-TIME step of `09` decision 87: merge `"live-restore": true`
|
||||
# into the guest's /etc/docker/daemon.json and `systemctl reload docker`. NEVER a restart (R-835).
|
||||
# Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is
|
||||
# OSAPPLY-REPORT <one JSON object>
|
||||
# which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install.
|
||||
@@ -30,6 +40,7 @@
|
||||
# package for version comparisons — 272 packages ≈ 4 minutes. Versions are now compared with the HOST's dpkg (the same
|
||||
# Debian algorithm), madison/policy run once per pass for all packages, the restart scan runs only after an install,
|
||||
# and one wrapper call does the whole pass (no separate inventory call before an apply).
|
||||
import calendar
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
@@ -58,7 +69,29 @@ INSTALL_STATE = "/var/lib/felhom-install/state.json"
|
||||
# reboot is needed for them to take effect, and a bad one can stop the box from booting.
|
||||
HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|"
|
||||
r"pve-firmware|firmware-|grub|shim|systemd-boot|intel-microcode|amd64-microcode|efibootmgr)")
|
||||
# restart_needed() leaves out processes whose cgroup line matches (grep basic regex). Host: the LXC guests' own
|
||||
# processes (`0::/lxc/<vmid>/...`) -- NOT lxc-start itself, whose cgroup is `0::/lxc.monitor/<vmid>` (measured
|
||||
# 2026-10-04 on demo-felhom: the old pattern "lxc" hid lxc-start with 20 deleted maps, so "reboot needed" stayed false
|
||||
# after a libc6 update). Pinned by test_restart_skip_patterns_against_real_cgroups.
|
||||
RESTART_SKIP_CGROUP = {"guest": "docker", "host": ":/lxc/"}
|
||||
HOST_SERVICES = ["pveproxy", "pvedaemon", "pvestatd", "pve-cluster", "felhom-agent"]
|
||||
# The Docker engine set (`11` §5.8): the only names the docker layer may touch, from the only origin it may use.
|
||||
DOCKER_NAMES = ("containerd.io", "docker-buildx-plugin", "docker-ce", "docker-ce-cli", "docker-ce-rootless-extras",
|
||||
"docker-compose-plugin")
|
||||
DOCKER_ORIGIN = "Docker CE"
|
||||
# ROOT-OWNED trust anchors (the installer writes them; the demo boxes got them by hand, R-840). Never the agent's config.
|
||||
TRUST_FILE = "/etc/felhom/os-trust.json" # {"host_id": "...", "ring0_slow_lane": false}
|
||||
TRUST_SIGNERS = "/etc/felhom/operator-signers" # ssh allowed_signers: <key_id> namespaces="felhom-op-v1" <key>
|
||||
SIG_NAMESPACE = "felhom-op-v1"
|
||||
SIGNED_OP = "os_docker_step"
|
||||
NONCE_FILE = "/var/lib/felhom-os-apply/nonces.json"
|
||||
DAEMON_JSON = "/etc/docker/daemon.json"
|
||||
# R-858 (v0.142.1): a Docker engine step restarts dockerd, which RECREATES the socket file. With live-restore the
|
||||
# containers keep running — and one that bind-mounts the socket FILE keeps the deleted inode: measured 2026-10-04 on
|
||||
# demo-felhom, the controller and traefik were blind to Docker for 1h44m. After a step that installed something, the
|
||||
# wrapper restarts exactly the containers that mount one of these paths (never the apps, never the engine).
|
||||
DOCKER_SOCKETS = ("/var/run/docker.sock", "/run/docker.sock")
|
||||
CRASH_GUARD_STATE = "/var/lib/felhom-crash-guard/state.json"
|
||||
|
||||
|
||||
class Refused(Exception):
|
||||
@@ -98,6 +131,38 @@ class Runner:
|
||||
if rc != 0:
|
||||
raise Refused("R7", f"could not write {path} in the guest")
|
||||
|
||||
def now(self):
|
||||
return time.time()
|
||||
|
||||
def sleep(self, s):
|
||||
time.sleep(s)
|
||||
|
||||
def verify_sig(self, signers, key_id, namespace, blob, sig):
|
||||
"""`ssh-keygen -Y verify` over the EXACT signed bytes. Files in a root-only temp dir; nothing via a shell."""
|
||||
import tempfile
|
||||
with tempfile.TemporaryDirectory(prefix="felhom-os-apply-") as d:
|
||||
sp = os.path.join(d, "sig")
|
||||
with open(sp, "w") as f:
|
||||
f.write(sig)
|
||||
p = subprocess.run(["ssh-keygen", "-Y", "verify", "-f", signers, "-I", key_id, "-n", namespace, "-s", sp],
|
||||
input=blob, capture_output=True, timeout=30)
|
||||
return p.returncode
|
||||
|
||||
def read_nonces(self):
|
||||
try:
|
||||
with open(NONCE_FILE) as f:
|
||||
d = json.load(f)
|
||||
return d if isinstance(d, dict) else {}
|
||||
except (OSError, ValueError):
|
||||
return {}
|
||||
|
||||
def write_nonces(self, d):
|
||||
os.makedirs(os.path.dirname(NONCE_FILE), mode=0o700, exist_ok=True)
|
||||
tmp = NONCE_FILE + ".tmp"
|
||||
with open(tmp, "w") as f:
|
||||
json.dump(d, f)
|
||||
os.replace(tmp, NONCE_FILE)
|
||||
|
||||
def log(self, line):
|
||||
print(line, file=sys.stderr, flush=True)
|
||||
try:
|
||||
@@ -138,13 +203,22 @@ class Apply:
|
||||
|
||||
def check_plan(self, plan):
|
||||
mode = plan.get("mode", "apply")
|
||||
if mode not in ("apply", "inventory", "health"):
|
||||
if mode not in ("apply", "inventory", "health", "facts", "live-restore-on"):
|
||||
raise Refused("R11", f"unknown mode {mode!r}")
|
||||
layer = plan.get("layer")
|
||||
if layer not in ("guest", "host"):
|
||||
raise Refused("R12", f"layer {layer!r} is not guest or host")
|
||||
if plan.get("lane", "fast") != "fast":
|
||||
raise Refused("R3", "the slow lane is refused in this release")
|
||||
if layer not in ("guest", "host", "docker"):
|
||||
raise Refused("R12", f"layer {layer!r} is not guest, host or docker")
|
||||
lane = plan.get("lane", "fast")
|
||||
if layer == "docker" and lane != "slow":
|
||||
raise Refused("R3", "the Docker engine is the slow lane (`11` §5.8); a fast-lane Docker plan is refused")
|
||||
if layer != "docker" and lane != "fast":
|
||||
raise Refused("R3", f"the {layer} layer has no slow lane in this release (kernel, Proxmox: `11` §8 step 6)")
|
||||
if mode == "facts" and layer != "host":
|
||||
raise Refused("R11", "facts is a host-layer mode (it reads the host and the guest)")
|
||||
if mode == "live-restore-on" and layer != "guest":
|
||||
raise Refused("R11", "live-restore-on is a guest-layer mode")
|
||||
if plan.get("undo") and layer != "docker":
|
||||
raise Refused("R5", "an undo (downgrade) exists only for the Docker layer, inside a signed job")
|
||||
vmid = plan.get("vmid")
|
||||
if not isinstance(vmid, int) or isinstance(vmid, bool) or vmid <= 0:
|
||||
raise Refused("R11", f"vmid must be a positive integer, got {vmid!r}")
|
||||
@@ -154,15 +228,18 @@ class Apply:
|
||||
if plan.get("allow_new"):
|
||||
raise Refused("R6", "allow_new is a slow-lane field; the fast lane never adds a package")
|
||||
select = plan.get("select", "listed")
|
||||
if select not in ("listed", "pending-fast"):
|
||||
if select not in ("listed", "pending-fast", "pending-docker"):
|
||||
raise Refused("R11", f"unknown select {select!r}")
|
||||
if (select == "pending-docker") != (layer == "docker" and select != "listed"):
|
||||
if select == "pending-docker" or layer == "docker":
|
||||
raise Refused("R11", f"select {select!r} does not fit layer {layer!r}")
|
||||
pk = plan.get("packages", [])
|
||||
if not isinstance(pk, list):
|
||||
raise Refused("R11", "packages must be a list")
|
||||
if mode == "apply" and select == "listed" and not pk:
|
||||
raise Refused("R11", "packages must be a non-empty list in apply mode (select listed)")
|
||||
if select == "pending-fast" and pk:
|
||||
raise Refused("R11", "select pending-fast takes no package list")
|
||||
if select in ("pending-fast", "pending-docker") and pk:
|
||||
raise Refused("R11", f"select {select} takes no package list")
|
||||
seen = set()
|
||||
for e in pk:
|
||||
if not isinstance(e, dict):
|
||||
@@ -175,6 +252,12 @@ class Apply:
|
||||
if n in seen:
|
||||
raise Refused("R11", f"package {n} is named twice")
|
||||
seen.add(n)
|
||||
if layer == "docker":
|
||||
if n not in DOCKER_NAMES or o != DOCKER_ORIGIN:
|
||||
raise Refused("R2", f"{n} ({o!r}) is not one of the six Docker packages from {DOCKER_ORIGIN!r}")
|
||||
continue
|
||||
if n in DOCKER_NAMES:
|
||||
raise Refused("R2", f"{n} is a Docker package — the slow lane (`11` §5.8), never in a {layer} plan")
|
||||
if o not in FAST_ORIGINS:
|
||||
raise Refused("R2", f"{n}: origin {o!r} is not Debian / Debian-Security (the fast lane, `11` C3)")
|
||||
if layer == "host" and HOST_SLOW_RE.match(n):
|
||||
@@ -199,6 +282,208 @@ class Apply:
|
||||
if mode != "appliance":
|
||||
raise Refused("R12", f"this box was installed as {mode!r}, not appliance — its host belongs to its owner")
|
||||
|
||||
def load_trust(self):
|
||||
"""The ROOT-OWNED trust record. Absent or agent-writable → no slow-lane authority at all (R3)."""
|
||||
try:
|
||||
st = self.r.stat(TRUST_FILE)
|
||||
except OSError:
|
||||
raise Refused("R3", f"no {TRUST_FILE} — this box has no slow-lane trust anchor")
|
||||
if st.st_uid != 0 or (st.st_mode & 0o022):
|
||||
raise Refused("R3", f"{TRUST_FILE} is not root-owned and root-only-writable — it proves nothing")
|
||||
try:
|
||||
t = json.loads(self.r.read_file(TRUST_FILE))
|
||||
except (OSError, ValueError):
|
||||
raise Refused("R3", f"{TRUST_FILE} is unreadable")
|
||||
if not isinstance(t, dict) or not isinstance(t.get("host_id"), str) or not t["host_id"]:
|
||||
raise Refused("R3", f"{TRUST_FILE} names no host_id")
|
||||
return t
|
||||
|
||||
def verify_signed(self, signed, trust):
|
||||
"""R3: an operator-signed os_docker_step, checked HERE (not by the agent): signature against the root-owned
|
||||
signers file, op, host binding, time window, and a root-owned nonce record (no replay)."""
|
||||
import base64
|
||||
if not isinstance(signed, dict) or not isinstance(signed.get("blob_b64"), str) or not isinstance(signed.get("sig"), str):
|
||||
raise Refused("R3", "the signed job is malformed")
|
||||
try:
|
||||
blob = base64.b64decode(signed["blob_b64"], validate=True)
|
||||
op = json.loads(blob)
|
||||
except (ValueError, TypeError):
|
||||
raise Refused("R3", "the signed blob is not base64 JSON")
|
||||
key_id = op.get("key_id", "")
|
||||
if not isinstance(key_id, str) or not re.match(r"^[A-Za-z0-9._-]{1,64}$", key_id):
|
||||
raise Refused("R3", "the signed blob names no plain key_id")
|
||||
try:
|
||||
st = self.r.stat(TRUST_SIGNERS)
|
||||
except OSError:
|
||||
raise Refused("R3", f"no {TRUST_SIGNERS} — no operator key to check a signed job against")
|
||||
if st.st_uid != 0 or (st.st_mode & 0o022):
|
||||
raise Refused("R3", f"{TRUST_SIGNERS} is not root-owned and root-only-writable")
|
||||
rc = self.r.verify_sig(TRUST_SIGNERS, key_id, SIG_NAMESPACE, blob, signed["sig"])
|
||||
if rc != 0:
|
||||
raise Refused("R3", f"the operator signature does not verify (ssh-keygen rc={rc})")
|
||||
if op.get("op") != SIGNED_OP:
|
||||
raise Refused("R3", f"the signed op is {op.get('op')!r}, not {SIGNED_OP}")
|
||||
if (op.get("target") or {}).get("host_id") != trust["host_id"]:
|
||||
raise Refused("R3", "the signed job is for another host")
|
||||
now = self.r.now()
|
||||
try:
|
||||
exp = calendar.timegm(time.strptime(op["expires_at"], "%Y-%m-%dT%H:%M:%SZ"))
|
||||
iss = calendar.timegm(time.strptime(op["issued_at"], "%Y-%m-%dT%H:%M:%SZ"))
|
||||
except (KeyError, ValueError, TypeError):
|
||||
raise Refused("R3", "the signed job has no readable time window")
|
||||
if now > exp or now < iss - 120:
|
||||
raise Refused("R3", "the signed job is expired or not yet valid")
|
||||
nonce = op.get("nonce")
|
||||
if not isinstance(nonce, str) or not nonce:
|
||||
raise Refused("R3", "the signed job has no nonce")
|
||||
seen = self.r.read_nonces()
|
||||
if nonce in seen:
|
||||
raise Refused("R3", "the signed job was already used (replay)")
|
||||
seen[nonce] = exp
|
||||
self.r.write_nonces({k: v for k, v in seen.items() if v > now})
|
||||
return op.get("params") or {}
|
||||
|
||||
def docker_authority(self, plan):
|
||||
"""R3 for the docker layer: returns (who, undo). A signed job binds the EXACT package list and the undo flag."""
|
||||
trust = self.load_trust()
|
||||
signed = plan.get("signed")
|
||||
if signed:
|
||||
params = self.verify_signed(signed, trust)
|
||||
want = sorted(f"{e.get('name')}={e.get('version')}" for e in params.get("packages") or [])
|
||||
got = sorted(f"{e['name']}={e['version']}" for e in plan.get("packages", []))
|
||||
if not want or want != got:
|
||||
raise Refused("R3", "the plan's packages are not exactly the signed job's packages")
|
||||
if bool(params.get("undo")) != bool(plan.get("undo")):
|
||||
raise Refused("R3", "the plan's undo flag is not the signed job's")
|
||||
if params.get("vmid") not in (None, self.vmid):
|
||||
raise Refused("R3", "the signed job names another guest")
|
||||
return "signed", bool(plan.get("undo"))
|
||||
if plan.get("undo"):
|
||||
raise Refused("R3", "an undo (downgrade) needs a signed operator job")
|
||||
if trust.get("ring0_slow_lane") is True:
|
||||
return "ring0", False
|
||||
raise Refused("R3", "a Docker step needs a signed operator job (ring 1) or this box's root-owned ring-0 mark")
|
||||
|
||||
def live_restore(self):
|
||||
rc, out, _ = self.g(["docker", "info", "--format", "{{.LiveRestoreEnabled}}"], timeout=60)
|
||||
return out.strip() if rc == 0 and out.strip() in ("true", "false") else "unknown"
|
||||
|
||||
def container_ids(self):
|
||||
rc, out, _ = self.g(["docker", "ps", "-q", "--no-trunc"], timeout=60)
|
||||
return sorted(out.split()) if rc == 0 else None
|
||||
|
||||
def live_restore_on(self):
|
||||
"""`09` decision 87: merge live-restore into daemon.json and RELOAD (C5: a reload turns it on, no restart)."""
|
||||
log = self.r.log
|
||||
if self.live_restore() == "true":
|
||||
self.report["live_restore"] = {"result": "already on"}
|
||||
log("os-apply: LIVE-RESTORE already on")
|
||||
return 0
|
||||
rc, cur, _ = self.g(["cat", DAEMON_JSON], timeout=30)
|
||||
try:
|
||||
conf = json.loads(cur) if rc == 0 and cur.strip() else {}
|
||||
except ValueError:
|
||||
raise Refused("R16", f"{DAEMON_JSON} in the guest is not valid JSON — not touched")
|
||||
if not isinstance(conf, dict):
|
||||
raise Refused("R16", f"{DAEMON_JSON} is not a JSON object — not touched")
|
||||
before = self.container_ids()
|
||||
conf["live-restore"] = True
|
||||
self.r.write_file("guest", self.vmid, DAEMON_JSON, json.dumps(conf, indent=2, sort_keys=True) + "\n")
|
||||
rrc, _, rerr = self.g(["systemctl", "reload", "docker"], timeout=120)
|
||||
state = "unknown"
|
||||
for _ in range(10):
|
||||
state = self.live_restore()
|
||||
if state == "true":
|
||||
break
|
||||
self.r.sleep(1)
|
||||
after = self.container_ids()
|
||||
same = before is not None and before == after
|
||||
self.report["live_restore"] = {"result": "on" if state == "true" else "failed", "reload_rc": rrc,
|
||||
"containers_before": len(before or []), "same_ids": same}
|
||||
log(f"os-apply: LIVE-RESTORE reload_rc={rrc} state={state} containers={len(before or [])} same-ids={'yes' if same else 'NO'}")
|
||||
if state != "true":
|
||||
# put the old file back (and reload again) — still never a restart
|
||||
self.r.write_file("guest", self.vmid, DAEMON_JSON, cur if rc == 0 else "{}\n")
|
||||
self.g(["systemctl", "reload", "docker"], timeout=120)
|
||||
self.report["failed"] = {"rc": 3, "step": "live-restore", "reason": (rerr or "").strip()[-200:]}
|
||||
return 3
|
||||
return 0
|
||||
|
||||
def kernel_next_boot(self):
|
||||
"""Which kernel GRUB boots next, read without root-only files (grubenv + /etc/default/grub + /boot)."""
|
||||
try:
|
||||
dflt = re.search(r'^GRUB_DEFAULT=["\']?([^"\'\n]*)', self.r.read_file("/etc/default/grub"), re.M)
|
||||
dflt = dflt.group(1) if dflt else "0"
|
||||
except OSError:
|
||||
dflt = "0"
|
||||
env = {}
|
||||
try:
|
||||
for l in self.r.read_file("/boot/grub/grubenv").splitlines():
|
||||
if "=" in l and not l.startswith("#"):
|
||||
k, v = l.split("=", 1)
|
||||
env[k] = v
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
def ver(entry):
|
||||
m = re.search(r"gnulinux-([0-9][^>\s]*?-pve)-(?:advanced|recovery)", entry)
|
||||
return m.group(1) if m else "unknown"
|
||||
if env.get("next_entry"):
|
||||
return ver(env["next_entry"]), "next_entry (a one-shot GRUB cannot clear on LVM /boot)"
|
||||
if dflt == "saved":
|
||||
return (ver(env["saved_entry"]), "saved default") if env.get("saved_entry") else ("unknown", "saved default unset")
|
||||
if dflt == "0":
|
||||
rc, out, _ = self.r.host(["sh", "-c", "ls /boot/vmlinuz-* 2>/dev/null"], 30)
|
||||
vers = [l.split("vmlinuz-", 1)[1] for l in out.split() if "vmlinuz-" in l]
|
||||
best = None
|
||||
for v in vers:
|
||||
if best is None or self.dpkg_cmp(v, "gt", best):
|
||||
best = v
|
||||
return (best or "unknown"), "GRUB_DEFAULT=0 (the newest installed)"
|
||||
return "unknown", f"GRUB_DEFAULT={dflt}"
|
||||
|
||||
def facts(self):
|
||||
"""Read-only versions for the hub's System page (R-852). A value that cannot be read is "unknown"."""
|
||||
def first(cmd, timeout=30):
|
||||
rc, out, _ = self.r.host(cmd, timeout)
|
||||
v = out.strip().splitlines()[0].strip() if rc == 0 and out.strip() else ""
|
||||
return v or "unknown"
|
||||
h = {"debian": first(["cat", "/etc/debian_version"]), "kernel_running": first(["uname", "-r"])}
|
||||
h["kernel_next_boot"], h["kernel_next_boot_source"] = self.kernel_next_boot()
|
||||
rc, out, _ = self.r.host(["apt-mark", "showhold"], 60)
|
||||
h["held"] = sorted(out.split()) if rc == 0 else None
|
||||
try:
|
||||
t = int(self.r.read_file("/proc/sys/kernel/tainted").strip())
|
||||
h["tainted"], h["oops_this_boot"], h["warn_this_boot"] = t, bool(t & 128), bool(t & 512)
|
||||
except (OSError, ValueError):
|
||||
h["tainted"], h["oops_this_boot"], h["warn_this_boot"] = None, None, None
|
||||
try:
|
||||
h["kernel_panic"] = int(self.r.read_file("/proc/sys/kernel/panic").strip())
|
||||
except (OSError, ValueError):
|
||||
h["kernel_panic"] = None
|
||||
try:
|
||||
h["crash_guard"] = json.loads(self.r.read_file(CRASH_GUARD_STATE))
|
||||
except (OSError, ValueError):
|
||||
h["crash_guard"] = None
|
||||
g = {"debian": "unknown", "docker_engine": "unknown", "containerd": "unknown", "live_restore": "unknown"}
|
||||
try:
|
||||
self.check_guest(self.vmid)
|
||||
running = True
|
||||
except Refused as e:
|
||||
running, g["unknown_reason"] = False, f"{e.code} {e.reason}"
|
||||
if running:
|
||||
script = ('echo "debian=$(cat /etc/debian_version 2>/dev/null)"; '
|
||||
'echo "engine=$(docker version --format \'{{.Server.Version}}\' 2>/dev/null)"; '
|
||||
'echo "containerd=$(dpkg-query -W -f \'${Version}\' containerd.io 2>/dev/null)"; '
|
||||
'echo "live=$(docker info --format \'{{.LiveRestoreEnabled}}\' 2>/dev/null)"')
|
||||
rc, out, _ = self.g(["sh", "-c", script], timeout=60)
|
||||
kv = dict(l.split("=", 1) for l in out.splitlines() if "=" in l)
|
||||
for k, src in (("debian", "debian"), ("docker_engine", "engine"), ("containerd", "containerd")):
|
||||
g[k] = kv.get(src, "").strip() or "unknown"
|
||||
g["live_restore"] = {"true": "on", "false": "off"}.get(kv.get("live", "").strip(), "unknown")
|
||||
self.report["facts"] = {"host": h, "guest": g}
|
||||
return 0
|
||||
|
||||
def check_guest(self, vmid):
|
||||
if vmid in RESERVED_VMIDS:
|
||||
raise Refused("R10", f"vmid {vmid} is a reserved scratch vmid")
|
||||
@@ -222,7 +507,7 @@ class Apply:
|
||||
"""Run in the TARGET layer: the guest via pct exec, or the host directly."""
|
||||
if self.layer == "host":
|
||||
return self.r.host(argv, timeout)
|
||||
return self.r.guest(self.vmid, argv, timeout)
|
||||
return self.r.guest(self.vmid, argv, timeout) # guest and docker both live in the customer guest
|
||||
|
||||
def g(self, argv, timeout=1800):
|
||||
"""Run in the customer GUEST whatever the layer (its health)."""
|
||||
@@ -284,21 +569,26 @@ class Apply:
|
||||
|
||||
def guest_health(self):
|
||||
"""The guest's signals: every container's state + health, the controller's own health, the network."""
|
||||
rc, out, _ = self.g(["docker", "ps", "-a", "--format", "{{.Names}}\t{{.State}}\t{{.Status}}"], timeout=60)
|
||||
rc, out, _ = self.g(["docker", "ps", "-a", "--no-trunc", "--format", "{{.Names}}\t{{.State}}\t{{.Status}}\t{{.ID}}"], timeout=60)
|
||||
cont = {}
|
||||
for l in out.splitlines():
|
||||
p = l.split("\t")
|
||||
if len(p) == 3:
|
||||
if len(p) >= 3:
|
||||
h = "healthy" if "(healthy)" in p[2] else "unhealthy" if "(unhealthy)" in p[2] else \
|
||||
"starting" if "(health: starting)" in p[2] else "none"
|
||||
cont[p[0]] = {"state": p[1], "health": h}
|
||||
if len(p) >= 4 and p[3]:
|
||||
cont[p[0]]["id"] = p[3]
|
||||
nrc, _, _ = self.g(["getent", "hosts", "deb.debian.org"], timeout=30)
|
||||
return {"docker_ok": rc == 0, "containers": cont,
|
||||
# R-858: the controller's own health check stayed "healthy" while it could not reach Docker at all — so ask
|
||||
# the consequence directly: can the controller talk to the engine from inside its container?
|
||||
crc, _, _ = self.g(["docker", "exec", "felhom-controller", "docker", "version", "--format", "{{.Server.Version}}"], timeout=60)
|
||||
return {"docker_ok": rc == 0, "containers": cont, "controller_docker_ok": crc == 0,
|
||||
"controller": cont.get("felhom-controller", {}).get("health", "absent"),
|
||||
"network_ok": nrc == 0}
|
||||
|
||||
def health(self):
|
||||
if self.layer == "guest":
|
||||
if self.layer in ("guest", "docker"):
|
||||
return self.guest_health()
|
||||
rc, out, _ = self.r.host(["systemctl", "is-active"] + HOST_SERVICES, 30)
|
||||
states = out.split()
|
||||
@@ -310,7 +600,7 @@ class Apply:
|
||||
def restart_needed(self):
|
||||
"""Processes still mapping deleted files, OUTSIDE containers (C11). Guest: outside docker; host: outside the
|
||||
LXC guests (the host's /proc shows guest processes too)."""
|
||||
skip = "docker" if self.layer == "guest" else "lxc"
|
||||
skip = RESTART_SKIP_CGROUP["host" if self.layer == "host" else "guest"]
|
||||
script = ('for p in /proc/[0-9]*; do grep -q "(deleted)" $p/maps 2>/dev/null || continue; '
|
||||
'grep -q "%s" $p/cgroup 2>/dev/null && continue; echo "${p#/proc/} $(cat $p/comm 2>/dev/null)"; done' % skip)
|
||||
rc, out, _ = self.x(["sh", "-c", script], timeout=120)
|
||||
@@ -368,16 +658,28 @@ class Apply:
|
||||
plan = self.load_plan()
|
||||
self.mode, self.layer, self.vmid, self.select = self.check_plan(plan)
|
||||
self.report.update(mode=self.mode, layer=self.layer, release_id=plan.get("release_id"), vmid=self.vmid)
|
||||
if self.mode == "facts":
|
||||
return self.facts()
|
||||
if self.layer == "host":
|
||||
self.check_appliance()
|
||||
self.check_guest(self.vmid)
|
||||
log = self.r.log
|
||||
if self.mode == "live-restore-on":
|
||||
return self.live_restore_on()
|
||||
self.who, self.allow_downgrade = ("fast", False)
|
||||
if self.layer == "docker" and self.mode == "apply":
|
||||
self.who, self.allow_downgrade = self.docker_authority(plan)
|
||||
if self.live_restore() != "true":
|
||||
raise Refused("R15", "live-restore is not ON in the guest — a Docker step would restart every container")
|
||||
self.report["authority"] = self.who
|
||||
self.report["undo"] = self.allow_downgrade
|
||||
if self.mode == "health":
|
||||
self.report["health"] = self.health()
|
||||
return 0
|
||||
log(f"os-apply: START release={plan.get('release_id')} layer={self.layer}" +
|
||||
(f":{self.vmid}" if self.layer == "guest" else "") +
|
||||
f" lane=fast mode={self.mode} select={self.select} packages={len(plan.get('packages', []))}")
|
||||
(f":{self.vmid}" if self.layer != "host" else "") +
|
||||
f" lane={plan.get('lane', 'fast')} mode={self.mode} select={self.select} packages={len(plan.get('packages', []))}" +
|
||||
(f" authority={self.who}{' UNDO' if self.allow_downgrade else ''}" if self.layer == "docker" else ""))
|
||||
if self.apt_lock_held():
|
||||
raise Refused("R9", f"another apt/dpkg holds the lock on the {self.layer}")
|
||||
self.report["health_before"] = self.health()
|
||||
@@ -392,6 +694,16 @@ class Apply:
|
||||
if rc:
|
||||
return rc
|
||||
self.report.update(self.inventory(installed_after))
|
||||
if "reboot_needed" not in self.report:
|
||||
# EVERY layer is scanned on EVERY pass: a reboot (host) or a restart (guest) must CLEAR "restart needed",
|
||||
# or the fleet view keeps a stale date (R-849, v0.142.0; the host since v0.141.1). One pct exec, ~1 s.
|
||||
# Pinned by test_every_layer_scans_every_pass.
|
||||
self.report["restart_needed"], self.report["reboot_needed"] = self.restart_needed()
|
||||
self.report["docker_restart_needed"] = any(p in ("dockerd", "containerd") for p in self.report["restart_needed"])
|
||||
if self.layer == "docker":
|
||||
rc_v, out_v, _ = self.g(["docker", "version", "--format", "{{.Server.Version}}"], timeout=60)
|
||||
self.report["docker_engine"] = out_v.strip() if rc_v == 0 and out_v.strip() else "unknown"
|
||||
self.report["reboot_scanned"] = "reboot_needed" in self.report
|
||||
self.report["health_after"] = self.health()
|
||||
return 0
|
||||
|
||||
@@ -424,9 +736,44 @@ class Apply:
|
||||
out.append({"name": p["name"], "version": p["to"], "origin": "Debian-Security" if "Debian-Security" in o else "Debian"})
|
||||
return out
|
||||
|
||||
def restart_socket_users(self):
|
||||
"""R-858: restart ONLY the containers that bind-mount the Docker socket, so they attach to the new one."""
|
||||
rc, out, _ = self.g(["docker", "ps", "-q", "--no-trunc"], timeout=60)
|
||||
users = []
|
||||
for cid in out.split():
|
||||
irc, iout, _ = self.g(["docker", "inspect", "-f", "{{.Name}}|{{range .Mounts}}{{.Destination}};{{end}}", cid], timeout=60)
|
||||
if irc != 0 or "|" not in iout:
|
||||
continue
|
||||
name, mounts = iout.strip().split("|", 1)
|
||||
if any(m in DOCKER_SOCKETS for m in mounts.split(";")):
|
||||
users.append(name.lstrip("/"))
|
||||
users.sort()
|
||||
if users:
|
||||
rrc, _, rerr = self.g(["docker", "restart"] + users, timeout=300)
|
||||
self.r.log(f"os-apply: SOCKET-USERS restarted={','.join(users)} rc={rrc} (R-858: they held the old docker socket)")
|
||||
return users
|
||||
|
||||
def pending_docker(self):
|
||||
"""Ring 0 (select pending-docker): the newest pending version of each INSTALLED Docker package, Docker origin."""
|
||||
rc, pend, remv, _ = self.simulate(["dist-upgrade"])
|
||||
return [{"name": p["name"], "version": p["to"], "origin": DOCKER_ORIGIN} for p in pend
|
||||
if p["from"] is not None and p["name"] in DOCKER_NAMES and self.origin_name(p["origin"]) == {DOCKER_ORIGIN}]
|
||||
|
||||
def origin_ok(self, origin):
|
||||
o = self.origin_name(origin)
|
||||
if self.layer == "docker":
|
||||
return o == {DOCKER_ORIGIN}
|
||||
return bool(o & set(FAST_ORIGINS))
|
||||
|
||||
def apply(self, plan):
|
||||
log = self.r.log
|
||||
packages = plan["packages"] if self.select == "listed" else self.pending_fast()
|
||||
if self.select == "listed":
|
||||
packages = plan["packages"]
|
||||
elif self.select == "pending-docker":
|
||||
packages = self.pending_docker()
|
||||
else:
|
||||
packages = self.pending_fast()
|
||||
cmp_op = "ne" if self.allow_downgrade else "gt"
|
||||
inst = self.installed()
|
||||
upgrade, already, notinst = [], 0, 0
|
||||
for e in packages:
|
||||
@@ -434,7 +781,7 @@ class Apply:
|
||||
if n not in inst:
|
||||
notinst += 1
|
||||
continue
|
||||
if not self.dpkg_cmp(v, "gt", inst[n]):
|
||||
if not self.dpkg_cmp(v, cmp_op, inst[n]):
|
||||
already += 1
|
||||
continue
|
||||
upgrade.append((n, v))
|
||||
@@ -462,7 +809,8 @@ class Apply:
|
||||
self.report["upgraded"] = []
|
||||
log("os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)")
|
||||
return 0, inst
|
||||
args = ["install", "--only-upgrade", "--no-install-recommends"] + [f"{n}={v}" for n, v in upgrade]
|
||||
args = ["install", "--only-upgrade", "--no-install-recommends"] + \
|
||||
(["--allow-downgrades"] if self.allow_downgrade else []) + [f"{n}={v}" for n, v in upgrade]
|
||||
rc, sim, remv, text = self.simulate(args)
|
||||
if rc != 0:
|
||||
tail = text.strip().splitlines()[-1] if text.strip() else ""
|
||||
@@ -477,10 +825,10 @@ class Apply:
|
||||
raise Refused("R6", f"the plan would touch {p['name']}, which is not in the plan")
|
||||
if p["to"] != want[p["name"]]:
|
||||
raise Refused("R6", f"{p['name']} would go to {p['to']}, not the approved {want[p['name']]}")
|
||||
if not self.dpkg_cmp(p["to"], "gt", p["from"]):
|
||||
if not self.allow_downgrade and not self.dpkg_cmp(p["to"], "gt", p["from"]):
|
||||
raise Refused("R5", f"{p['name']} would be downgraded {p['from']} -> {p['to']}")
|
||||
if not self.origin_name(p["origin"]) & set(FAST_ORIGINS):
|
||||
raise Refused("R2", f"{p['name']} would come from {p['origin']}, not Debian")
|
||||
if not self.origin_ok(p["origin"]):
|
||||
raise Refused("R2", f"{p['name']} would come from {p['origin']}, not the {self.layer} layer's origin")
|
||||
if self.layer == "host" and HOST_SLOW_RE.match(p["name"]):
|
||||
raise Refused("R14", f"{p['name']} is a kernel / boot / firmware package — the host's slow lane")
|
||||
need = self.download_bytes(args)
|
||||
@@ -514,7 +862,9 @@ class Apply:
|
||||
return 3, None
|
||||
self.report["upgraded"] = [{"name": n, "version": v} for n, v in upgrade]
|
||||
self.report["seconds"] = round(secs, 1)
|
||||
procs, reboot = self.restart_needed() # only after an install (R-845)
|
||||
if self.layer == "docker":
|
||||
self.report["socket_restarted"] = self.restart_socket_users()
|
||||
procs, reboot = self.restart_needed()
|
||||
self.report["restart_needed"] = procs
|
||||
self.report["docker_restart_needed"] = any(p in ("dockerd", "containerd") for p in procs)
|
||||
self.report["reboot_needed"] = reboot
|
||||
|
||||
@@ -0,0 +1,143 @@
|
||||
#!/usr/bin/python3
|
||||
"""Tests for felhom-crash-guard (`11` §5.9). Temp dirs only; nothing real is touched. Red-proof seam: CRASHGUARD_UNDER_TEST."""
|
||||
import importlib.machinery
|
||||
import importlib.util
|
||||
import json
|
||||
import os
|
||||
import pathlib
|
||||
import tempfile
|
||||
import unittest
|
||||
|
||||
HERE = pathlib.Path(__file__).resolve().parent
|
||||
_loader = importlib.machinery.SourceFileLoader("crashguard", os.environ.get("CRASHGUARD_UNDER_TEST", str(HERE / "felhom-crash-guard")))
|
||||
_spec = importlib.util.spec_from_loader("crashguard", _loader)
|
||||
cg = importlib.util.module_from_spec(_spec)
|
||||
_loader.exec_module(cg)
|
||||
|
||||
T0 = 1791115200.0 # 2026-10-04T12:00:00Z
|
||||
|
||||
|
||||
class FakeEnv(cg.Env):
|
||||
def __init__(self, d):
|
||||
super().__init__(conf=os.path.join(d, "conf"), state_dir=os.path.join(d, "state"),
|
||||
panic_path=os.path.join(d, "panic"), uptime_path=os.path.join(d, "uptime"),
|
||||
boot_id_path=os.path.join(d, "bootid"))
|
||||
self.t = T0
|
||||
self.logs = []
|
||||
open(self.panic_path, "w").write("0\n")
|
||||
open(self.uptime_path, "w").write("20.00 10.00\n")
|
||||
|
||||
def now(self):
|
||||
return self.t
|
||||
|
||||
def log(self, line):
|
||||
self.logs.append(line)
|
||||
|
||||
def panic(self):
|
||||
return int(open(self.panic_path).read())
|
||||
|
||||
def state(self):
|
||||
return json.load(open(os.path.join(self.state_dir, "state.json")))
|
||||
|
||||
|
||||
class Guard(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.d = tempfile.TemporaryDirectory()
|
||||
self.e = FakeEnv(self.d.name)
|
||||
|
||||
def tearDown(self):
|
||||
self.d.cleanup()
|
||||
|
||||
def crash_boot(self, minutes_later):
|
||||
self.e.t += minutes_later * 60
|
||||
cg.main(["x", "boot"], self.e) # no clean-stop before it: an unclean stop
|
||||
|
||||
def clean_reboot(self, minutes_later):
|
||||
cg.main(["x", "clean-stop"], self.e)
|
||||
self.e.t += minutes_later * 60
|
||||
cg.main(["x", "boot"], self.e)
|
||||
|
||||
def test_first_boot_is_not_a_crash_and_arms(self):
|
||||
cg.main(["x", "boot"], self.e)
|
||||
s = self.e.state()
|
||||
self.assertFalse(s["last_boot_unclean"])
|
||||
self.assertEqual(self.e.panic(), 10)
|
||||
self.assertTrue(s["armed"])
|
||||
|
||||
def test_clean_reboots_never_count(self):
|
||||
cg.main(["x", "boot"], self.e)
|
||||
for _ in range(5):
|
||||
self.clean_reboot(1)
|
||||
s = self.e.state()
|
||||
self.assertEqual(s["unclean_boots_in_window"], 0)
|
||||
self.assertEqual(self.e.panic(), 10)
|
||||
|
||||
def test_third_crash_in_an_hour_leaves_the_box_off(self):
|
||||
# operator's words: "if it crashes 3 times within one hour, it stays off" — after crash 2 the guard trips,
|
||||
# so crash 3 (kernel.panic = 0) does not restart the box.
|
||||
cg.main(["x", "boot"], self.e)
|
||||
self.crash_boot(5)
|
||||
self.assertEqual(self.e.panic(), 10, "one crash: still restarts")
|
||||
self.crash_boot(5)
|
||||
s = self.e.state()
|
||||
self.assertTrue(s["tripped"], s)
|
||||
self.assertEqual(self.e.panic(), 0, "after the 2nd crash boot the 3rd crash must leave the box off")
|
||||
self.assertIn("2 unclean boots within 60 minutes", s["tripped_reason"])
|
||||
|
||||
def test_crashes_spread_over_more_than_the_window_do_not_trip(self):
|
||||
cg.main(["x", "boot"], self.e)
|
||||
self.crash_boot(5)
|
||||
self.crash_boot(61)
|
||||
self.assertFalse(self.e.state()["tripped"])
|
||||
self.assertEqual(self.e.panic(), 10)
|
||||
|
||||
def test_tripped_stays_tripped_across_boots(self):
|
||||
cg.main(["x", "boot"], self.e)
|
||||
self.crash_boot(5)
|
||||
self.crash_boot(5)
|
||||
self.clean_reboot(30) # the operator switched it on; even a clean boot keeps the trip
|
||||
self.assertTrue(self.e.state()["tripped"])
|
||||
self.assertEqual(self.e.panic(), 0)
|
||||
|
||||
def test_rearms_after_24h_of_normal_running(self):
|
||||
cg.main(["x", "boot"], self.e)
|
||||
self.crash_boot(5)
|
||||
self.crash_boot(5)
|
||||
self.e.t += 23 * 3600
|
||||
cg.main(["x", "check"], self.e)
|
||||
self.assertTrue(self.e.state()["tripped"], "not before 24 h")
|
||||
self.e.t += 3600
|
||||
cg.main(["x", "check"], self.e)
|
||||
s = self.e.state()
|
||||
self.assertFalse(s["tripped"])
|
||||
self.assertEqual(self.e.panic(), 10)
|
||||
self.assertIn("timer", s["rearmed_by"])
|
||||
|
||||
def test_operator_rearm_starts_a_fresh_window(self):
|
||||
cg.main(["x", "boot"], self.e)
|
||||
self.crash_boot(5)
|
||||
self.crash_boot(5)
|
||||
self.e.t += 60
|
||||
cg.main(["x", "rearm"], self.e)
|
||||
s = self.e.state()
|
||||
self.assertFalse(s["tripped"])
|
||||
self.assertEqual(s["rearmed_by"], "operator")
|
||||
self.assertEqual(s["unclean_boots_24h"], 2, "the history stays")
|
||||
self.crash_boot(5)
|
||||
self.assertFalse(self.e.state()["tripped"], "one crash after a re-arm must not trip at once")
|
||||
|
||||
def test_config_numbers_are_read(self):
|
||||
open(self.e.conf, "w").write("LIMIT=2\nPANIC_SECONDS=30\n")
|
||||
cg.main(["x", "boot"], self.e)
|
||||
self.assertEqual(self.e.panic(), 30)
|
||||
self.crash_boot(1)
|
||||
self.assertTrue(self.e.state()["tripped"], "LIMIT=2: the first crash boot trips")
|
||||
|
||||
def test_state_is_world_readable_for_the_agent(self):
|
||||
cg.main(["x", "boot"], self.e)
|
||||
mode = os.stat(os.path.join(self.e.state_dir, "state.json")).st_mode & 0o777
|
||||
self.assertEqual(mode, 0o644)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -64,6 +64,35 @@ class Fake:
|
||||
self.services = {}
|
||||
self.files[osapply.INSTALL_STATE] = json.dumps({"mode": "appliance"})
|
||||
self.stats[osapply.INSTALL_STATE] = St(mode=statmod.S_IFREG | 0o644, uid=0)
|
||||
# Docker / trust / facts state (v0.142.0)
|
||||
self.live_restore = "true"
|
||||
self.ids = ["aaa111", "bbb222"]
|
||||
self.engine = "29.7.2"
|
||||
self.daemon_json = '{"log-driver": "json-file"}'
|
||||
self.reload_enables = True
|
||||
self.sig_rc = 0
|
||||
self.nonces = {}
|
||||
self.clock = 1791115200.0 # 2026-10-04T12:00:00Z
|
||||
self.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": False})
|
||||
self.stats[osapply.TRUST_FILE] = St(mode=statmod.S_IFREG | 0o644, uid=0)
|
||||
self.files[osapply.TRUST_SIGNERS] = 'felhom-op-1 namespaces="felhom-op-v1" ssh-ed25519 AAAA\n'
|
||||
self.stats[osapply.TRUST_SIGNERS] = St(mode=statmod.S_IFREG | 0o644, uid=0)
|
||||
|
||||
def now(self):
|
||||
return self.clock
|
||||
|
||||
def sleep(self, s):
|
||||
pass
|
||||
|
||||
def verify_sig(self, signers, key_id, ns, blob, sig):
|
||||
self.verified = (signers, key_id, ns, blob, sig)
|
||||
return self.sig_rc
|
||||
|
||||
def read_nonces(self):
|
||||
return dict(self.nonces)
|
||||
|
||||
def write_nonces(self, d):
|
||||
self.nonces = dict(d)
|
||||
|
||||
# Runner interface
|
||||
def read_file(self, p):
|
||||
@@ -151,11 +180,46 @@ class Fake:
|
||||
return 0, getattr(self, "install_out", "Setting up libc6 ...\n"), ""
|
||||
if cmd == "df":
|
||||
return 0, f"Avail\n{self.free}\n", ""
|
||||
if cmd == "docker" and a[1] == "info":
|
||||
return 0, self.live_restore + "\n", ""
|
||||
if cmd == "docker" and a[1] == "version":
|
||||
return 0, self.engine + "\n", ""
|
||||
if cmd == "docker" and a[1:3] == ["ps", "-q"]:
|
||||
return 0, "".join(i + "\n" for i in self.ids), ""
|
||||
if cmd == "docker" and a[1] == "inspect":
|
||||
mounts = {"aaa111": "/felhom-controller|/var/run/docker.sock;/app/data;", "bbb222": "/app|/data;"}
|
||||
return 0, mounts.get(a[-1], "/other|;") + "\n", ""
|
||||
if cmd == "docker" and a[1] == "restart":
|
||||
self.restarted_containers = a[2:]
|
||||
return 0, "", ""
|
||||
if cmd == "docker" and a[1] == "exec":
|
||||
return (1, "", "Cannot connect to the Docker daemon") if getattr(self, "controller_blind", False) else (0, self.engine + "\n", "")
|
||||
if cmd == "docker":
|
||||
return 0, "felhom-controller\trunning\tUp 1 hour (healthy)\napp\trunning\tUp 1 hour (healthy)\n", ""
|
||||
ids = self.ids + ["x"] * 2
|
||||
return 0, f"felhom-controller\trunning\tUp 1 hour (healthy)\t{ids[0]}\napp\trunning\tUp 1 hour (healthy)\t{ids[1]}\n", ""
|
||||
if cmd == "cat" and a[1] == osapply.DAEMON_JSON:
|
||||
return (0, self.daemon_json, "") if self.daemon_json is not None else (1, "", "No such file")
|
||||
if cmd == "cat" and a[1] == "/etc/debian_version":
|
||||
return 0, "13.7\n", ""
|
||||
if cmd == "uname":
|
||||
return 0, "7.0.14-20-pve\n", ""
|
||||
if cmd == "apt-mark":
|
||||
return 0, getattr(self, "held", ""), ""
|
||||
if cmd == "systemctl" and a[1] == "reload":
|
||||
self.reloads = getattr(self, "reloads", 0) + 1
|
||||
if self.reload_enables and '"live-restore": true' in self.written.get(osapply.DAEMON_JSON, ""):
|
||||
self.live_restore = "true"
|
||||
return 0, "", ""
|
||||
if cmd == "systemctl" and a[1] == "restart":
|
||||
self.restarted = True
|
||||
return 0, "", ""
|
||||
if cmd == "getent":
|
||||
return 0, "1.2.3.4 deb.debian.org\n", ""
|
||||
if cmd == "sh":
|
||||
if "vmlinuz" in a[2]:
|
||||
return 0, "/boot/vmlinuz-7.0.2-6-pve\n/boot/vmlinuz-7.0.14-20-pve\n", ""
|
||||
if "engine=" in a[2]:
|
||||
return 0, f"debian=13.7\nengine={self.engine}\ncontainerd=2.3.3-1~debian.13~trixie\nlive={self.live_restore}\n", ""
|
||||
if "os-release" in a[2]:
|
||||
return 0, "trixie\n", ""
|
||||
if "(deleted)" in a[2]:
|
||||
@@ -180,7 +244,7 @@ class Fake:
|
||||
n, v = x.split("=", 1)
|
||||
if v not in self.avail(n):
|
||||
return 100, "", f"E: Version '{v}' for '{n}' was not found"
|
||||
origin = SEC if n == "openssl" else DEB
|
||||
origin = "Docker CE:trixie" if n in osapply.DOCKER_NAMES else SEC if n == "openssl" else DEB
|
||||
out += f"Inst {n} [{self.installed[n]}] ({v} {origin} [amd64])\n"
|
||||
out += "".join(l + "\n" for l in self.extra_sim)
|
||||
return 0, out, ""
|
||||
@@ -226,7 +290,6 @@ class Happy(unittest.TestCase):
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(f.installed["libc6"], "2.41-12+deb13u3")
|
||||
self.assertIn("installed", rep)
|
||||
self.assertNotIn("restart_needed", rep, "the restart scan runs only after an install (R-845)")
|
||||
|
||||
def test_health_mode(self):
|
||||
f = Fake()
|
||||
@@ -551,6 +614,23 @@ class HostLayer(unittest.TestCase):
|
||||
rc, rep = run(f)
|
||||
self.assertTrue(rep["reboot_needed"], rep)
|
||||
|
||||
def test_every_layer_scans_every_pass(self):
|
||||
# R-849 (v0.142.0): host AND guest are scanned on every pass, so a reboot / restart clears the flag.
|
||||
for layer in ("host", "guest"):
|
||||
f = Fake()
|
||||
f.plan["layer"] = layer
|
||||
f.plan["mode"] = "inventory"
|
||||
f.restart_out = "2101 lxc-start\n" if layer == "host" else "1 systemd\n"
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertTrue(rep["reboot_scanned"] and rep["reboot_needed"], (layer, rep))
|
||||
f = Fake()
|
||||
f.plan["layer"] = layer
|
||||
f.plan["mode"] = "inventory"
|
||||
f.restart_out = ""
|
||||
rc, rep = run(f)
|
||||
self.assertTrue(rep["reboot_scanned"] and rep["reboot_needed"] is False, (layer, rep))
|
||||
|
||||
def test_no_reboot_for_ordinary_daemons(self):
|
||||
f = Fake()
|
||||
f.plan["layer"] = "host"
|
||||
@@ -586,5 +666,299 @@ class Failure(unittest.TestCase):
|
||||
self.assertTrue(any(l.startswith("os-apply: FAILED rc=100 step=install") for l in f.logs))
|
||||
|
||||
|
||||
|
||||
class RestartSkipPattern(unittest.TestCase):
|
||||
"""The cgroup filter runs as `grep -q PATTERN /proc/<pid>/cgroup`; check it with grep itself against the cgroup
|
||||
lines measured on demo-felhom 2026-10-04."""
|
||||
|
||||
def grep(self, pattern, line):
|
||||
return subprocess.run(["grep", "-q", pattern], input=line + "\n", text=True).returncode == 0
|
||||
|
||||
def test_restart_skip_patterns_against_real_cgroups(self):
|
||||
host = osapply.RESTART_SKIP_CGROUP["host"]
|
||||
self.assertTrue(self.grep(host, "0::/lxc/9201/ns/system.slice/docker.service"), "a guest process must be skipped")
|
||||
self.assertFalse(self.grep(host, "0::/lxc.monitor/9201"), "lxc-start must NOT be skipped (it runs the guest)")
|
||||
self.assertFalse(self.grep(host, "0::/system.slice/pve-cluster.service"), "a host daemon must NOT be skipped")
|
||||
guest = osapply.RESTART_SKIP_CGROUP["guest"]
|
||||
self.assertTrue(self.grep(guest, "0::/system.slice/docker-0123abcd.scope"))
|
||||
self.assertFalse(self.grep(guest, "0::/system.slice/cron.service"))
|
||||
|
||||
|
||||
DOCKER_SET = [{"name": "docker-ce", "version": "5:29.8.2-1~debian.13~trixie", "origin": "Docker CE"},
|
||||
{"name": "containerd.io", "version": "2.3.6-1~debian.13~trixie", "origin": "Docker CE"}]
|
||||
|
||||
|
||||
def docker_fake(signed=None, undo=False, ring0=False):
|
||||
f = Fake()
|
||||
f.installed.update({"docker-ce": "5:29.7.2-1~debian.13~trixie", "containerd.io": "2.3.3-1~debian.13~trixie"})
|
||||
f.live["docker-ce"] = {"5:29.8.2-1~debian.13~trixie", "5:29.7.2-1~debian.13~trixie"}
|
||||
f.live["containerd.io"] = {"2.3.6-1~debian.13~trixie", "2.3.3-1~debian.13~trixie"}
|
||||
f.plan = {"release_id": "os-docker-t1", "layer": "docker", "lane": "slow", "vmid": 9201, "mode": "apply",
|
||||
"packages": [dict(p) for p in DOCKER_SET]}
|
||||
if undo:
|
||||
f.plan["undo"] = True
|
||||
if ring0:
|
||||
f.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": True})
|
||||
if signed is not None:
|
||||
f.plan["signed"] = signed
|
||||
return f
|
||||
|
||||
|
||||
def signed_job(packages=DOCKER_SET, host="demo-hp-bb76ea", op="os_docker_step", undo=False, nonce="n1",
|
||||
issued="2026-10-04T11:50:00Z", expires="2026-10-04T12:30:00Z"):
|
||||
import base64
|
||||
params = {"packages": packages, "undo": undo}
|
||||
blob = json.dumps({"expires_at": expires, "issued_at": issued, "key_id": "felhom-op-1", "nonce": nonce, "op": op,
|
||||
"params": params, "target": {"guest_id": "", "host_id": host}}, sort_keys=True).encode()
|
||||
return {"blob_b64": base64.b64encode(blob).decode(), "sig": "-----BEGIN SSH SIGNATURE-----\nx\n-----END SSH SIGNATURE-----\n"}
|
||||
|
||||
|
||||
class DockerLane(unittest.TestCase):
|
||||
"""`11` §5.8, agent v0.142.0. Each test names the refusal it pins; the red-proof file mutates each one."""
|
||||
|
||||
def refused(self, f, code):
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 2, rep)
|
||||
self.assertEqual(rep["refused"]["code"], code, rep)
|
||||
self.assertEqual(f.installed.get("docker-ce", "5:29.7.2-1~debian.13~trixie"), "5:29.7.2-1~debian.13~trixie")
|
||||
return rep
|
||||
|
||||
def test_docker_package_in_a_fast_plan_is_refused(self):
|
||||
f = Fake()
|
||||
f.plan["packages"].append({"name": "docker-ce", "version": "5:29.8.2-1~debian.13~trixie", "origin": "Debian"})
|
||||
self.refused(f, "R2")
|
||||
|
||||
def test_docker_layer_in_the_fast_lane_is_refused(self):
|
||||
f = docker_fake(ring0=True)
|
||||
f.plan["lane"] = "fast"
|
||||
self.refused(f, "R3")
|
||||
|
||||
def test_no_authority_is_refused(self):
|
||||
self.refused(docker_fake(), "R3")
|
||||
|
||||
def test_ring0_mark_allows_pending_docker(self):
|
||||
f = docker_fake(ring0=True)
|
||||
f.plan["select"], f.plan["packages"] = "pending-docker", []
|
||||
f.pending_sim = ["Inst docker-ce [5:29.7.2-1~debian.13~trixie] (5:29.8.2-1~debian.13~trixie Docker CE:trixie [amd64])",
|
||||
"Inst containerd.io [2.3.3-1~debian.13~trixie] (2.3.6-1~debian.13~trixie Docker CE:trixie [amd64])",
|
||||
"Inst bash [5.2.37-2+b9] (5.2.37-2+b10 Debian:13.7/stable [amd64])"]
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(rep["authority"], "ring0")
|
||||
self.assertEqual(f.installed["docker-ce"], "5:29.8.2-1~debian.13~trixie")
|
||||
self.assertEqual(f.installed["bash"], "5.2.37-2+b9", "a Debian package must not ride a Docker step")
|
||||
self.assertFalse(getattr(f, "restarted", False), "never a docker restart")
|
||||
|
||||
def test_signed_job_applies_exactly_its_packages(self):
|
||||
f = docker_fake(signed=signed_job())
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(rep["authority"], "signed")
|
||||
self.assertEqual(f.installed["containerd.io"], "2.3.6-1~debian.13~trixie")
|
||||
self.assertEqual(f.verified[0], osapply.TRUST_SIGNERS, "the ROOT-OWNED signers file, not the agent's config")
|
||||
self.assertIn("n1", f.nonces)
|
||||
|
||||
def test_bad_signature_is_refused(self):
|
||||
f = docker_fake(signed=signed_job())
|
||||
f.sig_rc = 255
|
||||
self.refused(f, "R3")
|
||||
self.assertEqual(f.nonces, {}, "a bad signature must not burn a nonce")
|
||||
|
||||
def test_signed_job_for_another_host_is_refused(self):
|
||||
self.refused(docker_fake(signed=signed_job(host="demo-felhom-8363b5")), "R3")
|
||||
|
||||
def test_signed_job_other_op_is_refused(self):
|
||||
self.refused(docker_fake(signed=signed_job(op="agent_update")), "R3")
|
||||
|
||||
def test_expired_signed_job_is_refused(self):
|
||||
self.refused(docker_fake(signed=signed_job(expires="2026-10-04T11:55:00Z")), "R3")
|
||||
|
||||
def test_replayed_signed_job_is_refused(self):
|
||||
f = docker_fake(signed=signed_job())
|
||||
f.nonces = {"n1": f.clock + 600}
|
||||
self.refused(f, "R3")
|
||||
|
||||
def test_plan_must_equal_the_signed_packages(self):
|
||||
f = docker_fake(signed=signed_job(packages=DOCKER_SET[:1]))
|
||||
self.refused(f, "R3")
|
||||
|
||||
def test_agent_writable_trust_file_is_refused(self):
|
||||
f = docker_fake(ring0=True)
|
||||
f.stats[osapply.TRUST_FILE] = St(mode=statmod.S_IFREG | 0o644, uid=999)
|
||||
self.refused(f, "R3")
|
||||
|
||||
def test_live_restore_off_is_refused(self):
|
||||
f = docker_fake(signed=signed_job())
|
||||
f.live_restore = "false"
|
||||
self.refused(f, "R15")
|
||||
|
||||
def test_undo_needs_a_signed_job(self):
|
||||
self.refused(docker_fake(undo=True, ring0=True), "R3")
|
||||
|
||||
def test_signed_undo_downgrades(self):
|
||||
old = [{"name": "docker-ce", "version": "5:29.7.2-1~debian.13~trixie", "origin": "Docker CE"}]
|
||||
f = docker_fake(signed=signed_job(packages=old, undo=True), undo=True)
|
||||
f.plan["packages"] = [dict(p) for p in old]
|
||||
f.installed["docker-ce"] = "5:29.8.2-1~debian.13~trixie"
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(f.installed["docker-ce"], "5:29.7.2-1~debian.13~trixie")
|
||||
self.assertTrue(rep["undo"])
|
||||
|
||||
def test_unsigned_downgrade_is_refused(self):
|
||||
f = docker_fake(ring0=True)
|
||||
f.plan["packages"] = [{"name": "docker-ce", "version": "5:29.6.0-1~debian.13~trixie", "origin": "Docker CE"}]
|
||||
f.live["docker-ce"].add("5:29.6.0-1~debian.13~trixie")
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rep["plan"]["upgrade"], 0, "an older version on an unsigned step is 'already', never installed")
|
||||
|
||||
def test_step_restarts_only_the_socket_users(self):
|
||||
# R-858: after an engine step, ONLY the container that mounts the docker socket is restarted (here the controller)
|
||||
f = docker_fake(signed=signed_job())
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(f.restarted_containers, ["felhom-controller"])
|
||||
self.assertEqual(rep["socket_restarted"], ["felhom-controller"])
|
||||
|
||||
def test_no_install_restarts_nothing(self):
|
||||
f = docker_fake(signed=signed_job())
|
||||
f.installed.update({"docker-ce": "5:29.8.2-1~debian.13~trixie", "containerd.io": "2.3.6-1~debian.13~trixie"})
|
||||
rc, rep = run(f)
|
||||
self.assertFalse(hasattr(f, "restarted_containers"), "nothing installed -> no container restart")
|
||||
|
||||
def test_health_says_when_the_controller_cannot_reach_docker(self):
|
||||
f = docker_fake(signed=signed_job())
|
||||
f.controller_blind = True
|
||||
rc, rep = run(f)
|
||||
self.assertFalse(rep["health_after"]["controller_docker_ok"])
|
||||
|
||||
def test_health_carries_container_ids(self):
|
||||
f = docker_fake(signed=signed_job())
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rep["health_after"]["containers"]["app"]["id"], "bbb222")
|
||||
self.assertEqual(rep["docker_engine"], "29.7.2")
|
||||
|
||||
|
||||
class LiveRestore(unittest.TestCase):
|
||||
def lr(self):
|
||||
f = Fake()
|
||||
f.plan = {"release_id": "lr", "layer": "guest", "lane": "fast", "vmid": 9201, "mode": "live-restore-on", "packages": []}
|
||||
f.live_restore = "false"
|
||||
return f
|
||||
|
||||
def test_turns_it_on_with_a_reload_never_a_restart(self):
|
||||
f = self.lr()
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(json.loads(f.written[osapply.DAEMON_JSON]), {"log-driver": "json-file", "live-restore": True})
|
||||
self.assertEqual(f.reloads, 1)
|
||||
self.assertFalse(getattr(f, "restarted", False))
|
||||
self.assertEqual(rep["live_restore"]["result"], "on")
|
||||
self.assertTrue(rep["live_restore"]["same_ids"])
|
||||
|
||||
def test_already_on_writes_nothing(self):
|
||||
f = self.lr()
|
||||
f.live_restore = "true"
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0)
|
||||
self.assertNotIn(osapply.DAEMON_JSON, f.written)
|
||||
|
||||
def test_invalid_daemon_json_is_left_alone(self):
|
||||
f = self.lr()
|
||||
f.daemon_json = "{not json"
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rep["refused"]["code"], "R16")
|
||||
self.assertNotIn(osapply.DAEMON_JSON, f.written)
|
||||
|
||||
def test_reload_that_does_not_enable_puts_the_file_back(self):
|
||||
f = self.lr()
|
||||
f.reload_enables = False
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 3, rep)
|
||||
self.assertEqual(f.written[osapply.DAEMON_JSON], '{"log-driver": "json-file"}')
|
||||
self.assertFalse(getattr(f, "restarted", False))
|
||||
|
||||
|
||||
class Facts(unittest.TestCase):
|
||||
def facts(self, f=None):
|
||||
f = f or Fake()
|
||||
f.plan = {"release_id": "facts", "layer": "host", "lane": "fast", "vmid": 9201, "mode": "facts", "packages": []}
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
return f, rep["facts"]
|
||||
|
||||
def test_reads_host_and_guest(self):
|
||||
f = Fake()
|
||||
f.files["/proc/sys/kernel/tainted"] = "4225\n" # 4096 + 128 (D: oops) + 1
|
||||
f.files["/proc/sys/kernel/panic"] = "10\n"
|
||||
f.held = "tzdata\n"
|
||||
f.files["/etc/default/grub"] = "GRUB_DEFAULT=saved\n"
|
||||
f.files["/boot/grub/grubenv"] = "# GRUB Environment Block\nsaved_entry=gnulinux-advanced-x>gnulinux-7.0.14-20-pve-advanced-x\n"
|
||||
_, fa = self.facts(f)
|
||||
h, g = fa["host"], fa["guest"]
|
||||
self.assertEqual((h["debian"], h["kernel_running"], h["kernel_next_boot"]), ("13.7", "7.0.14-20-pve", "7.0.14-20-pve"))
|
||||
self.assertEqual(h["held"], ["tzdata"])
|
||||
self.assertTrue(h["oops_this_boot"])
|
||||
self.assertEqual(h["kernel_panic"], 10)
|
||||
self.assertEqual((g["debian"], g["docker_engine"], g["live_restore"]), ("13.7", "29.7.2", "on"))
|
||||
self.assertEqual(g["containerd"], "2.3.3-1~debian.13~trixie")
|
||||
|
||||
def test_next_entry_wins_and_default_zero_is_the_newest(self):
|
||||
f = Fake()
|
||||
f.files["/etc/default/grub"] = "GRUB_DEFAULT=saved\n"
|
||||
f.files["/boot/grub/grubenv"] = "saved_entry=gnulinux-advanced-x>gnulinux-7.0.2-6-pve-advanced-x\nnext_entry=gnulinux-advanced-x>gnulinux-7.0.14-20-pve-advanced-x\n"
|
||||
_, fa = self.facts(f)
|
||||
self.assertEqual(fa["host"]["kernel_next_boot"], "7.0.14-20-pve")
|
||||
self.assertIn("next_entry", fa["host"]["kernel_next_boot_source"])
|
||||
f = Fake()
|
||||
f.files["/etc/default/grub"] = "GRUB_DEFAULT=0\n"
|
||||
_, fa = self.facts(f)
|
||||
self.assertEqual(fa["host"]["kernel_next_boot"], "7.0.14-20-pve", "dpkg order, not string order (7.0.2 < 7.0.14)")
|
||||
|
||||
def test_stopped_guest_is_unknown_never_guessed(self):
|
||||
f = Fake()
|
||||
f.status = "status: stopped"
|
||||
_, fa = self.facts(f)
|
||||
self.assertEqual(fa["guest"]["docker_engine"], "unknown")
|
||||
self.assertIn("R10", fa["guest"]["unknown_reason"])
|
||||
self.assertEqual(fa["host"]["tainted"], None, "unreadable -> None, not 0")
|
||||
|
||||
|
||||
class RealSignatureCheck(unittest.TestCase):
|
||||
"""The REAL Runner.verify_sig with the real ssh-keygen and a throwaway key, in the installer's allowed_signers form
|
||||
(`<key_id> namespaces="felhom-op-v1" <type> <b64> <comment>`). Nothing leaves the temp dir."""
|
||||
|
||||
def setUp(self):
|
||||
import shutil
|
||||
if not shutil.which("ssh-keygen"):
|
||||
self.skipTest("ssh-keygen not available")
|
||||
import tempfile
|
||||
self.d = tempfile.mkdtemp()
|
||||
self.key = os.path.join(self.d, "k")
|
||||
subprocess.run(["ssh-keygen", "-q", "-t", "ed25519", "-N", "", "-C", "felhom-op-1", "-f", self.key], check=True)
|
||||
pub = open(self.key + ".pub").read().strip()
|
||||
self.signers = os.path.join(self.d, "signers")
|
||||
open(self.signers, "w").write(f'felhom-op-1 namespaces="felhom-op-v1" {pub}\n')
|
||||
|
||||
def sign(self, blob, ns="felhom-op-v1"):
|
||||
bp = os.path.join(self.d, "blob")
|
||||
open(bp, "wb").write(blob)
|
||||
if os.path.exists(bp + ".sig"):
|
||||
os.remove(bp + ".sig") # ssh-keygen -Y sign asks before overwriting (it would wait on stdin)
|
||||
subprocess.run(["ssh-keygen", "-q", "-Y", "sign", "-f", self.key, "-n", ns, bp], check=True, stdin=subprocess.DEVNULL, timeout=30)
|
||||
return open(bp + ".sig").read()
|
||||
|
||||
def test_good_signature_verifies(self):
|
||||
blob = b'{"op":"os_docker_step"}'
|
||||
self.assertEqual(osapply.Runner().verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob)), 0)
|
||||
|
||||
def test_changed_blob_wrong_namespace_or_wrong_principal_fail(self):
|
||||
blob = b'{"op":"os_docker_step"}'
|
||||
sig = self.sign(blob)
|
||||
r = osapply.Runner()
|
||||
self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob + b" ", sig), 0)
|
||||
self.assertNotEqual(r.verify_sig(self.signers, "someone-else", "felhom-op-v1", blob, sig), 0)
|
||||
self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob, ns="other-ns")), 0)
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
|
||||
@@ -4,10 +4,12 @@ import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"log/slog"
|
||||
"os"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
|
||||
@@ -108,6 +110,7 @@ type Collector struct {
|
||||
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
|
||||
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
|
||||
oob OOBReporter // H1: operator-access health (nil → stanza omitted)
|
||||
system SystemReporter // R-852: the box versions (nil → API fields only)
|
||||
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
|
||||
hostID string
|
||||
agentVersion string
|
||||
@@ -248,6 +251,41 @@ type OOBReporter interface {
|
||||
}
|
||||
|
||||
// SetOOBReporter wires the operator-access health source (H1; nil-safe → stanza omitted).
|
||||
// SystemReporter reads the box's versions (R-852): the customer guest's vmid and the wrapper's raw facts.
|
||||
type SystemReporter interface {
|
||||
SystemFacts(ctx context.Context) (vmid int, facts json.RawMessage, err error)
|
||||
}
|
||||
|
||||
// SetSystemReporter wires the facts read (agent v0.142.0). Without it the stanza carries the Proxmox API fields only.
|
||||
func (c *Collector) SetSystemReporter(r SystemReporter) *Collector {
|
||||
c.system = r
|
||||
return c
|
||||
}
|
||||
|
||||
func unknownIfEmpty(s string) string {
|
||||
if strings.TrimSpace(s) == "" {
|
||||
return "unknown"
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// systemInfo builds the `system` stanza. Never fatal: a failed facts read is FactsError, the API fields stay.
|
||||
func (c *Collector) systemInfo(ctx context.Context, ns proxmox.NodeStatus) *SystemInfo {
|
||||
si := &SystemInfo{PVEVersion: unknownIfEmpty(ns.PVEVersion), KernelVersion: unknownIfEmpty(ns.KVersion),
|
||||
ReadAt: c.now().Format(time.RFC3339)}
|
||||
if c.system == nil {
|
||||
si.FactsError = "no facts reader wired"
|
||||
return si
|
||||
}
|
||||
vmid, f, err := c.system.SystemFacts(ctx)
|
||||
si.VMID, si.Facts = vmid, f
|
||||
if err != nil {
|
||||
si.FactsError = err.Error()
|
||||
c.logger.Debug("hub: system facts unavailable", "err", err)
|
||||
}
|
||||
return si
|
||||
}
|
||||
|
||||
func (c *Collector) SetOOBReporter(o OOBReporter) *Collector {
|
||||
c.oob = o
|
||||
return c
|
||||
@@ -284,6 +322,7 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
|
||||
Capabilities: c.capabilities(ctx),
|
||||
LeafFingerprint: c.leafFP,
|
||||
Addresses: c.collectAddresses(),
|
||||
System: c.systemInfo(ctx, ns),
|
||||
}
|
||||
// DR recipe host-half — derived from the just-collected guest/storage/PBS facts (no new reads).
|
||||
// Secret-free by construction (identifiers/intents/sizes/coordinates only).
|
||||
|
||||
@@ -89,6 +89,12 @@ type HostReport struct {
|
||||
// hub-schema change and are absent when the reporter is not wired.
|
||||
MgmtPlane *MgmtPlaneStatus `json:"mgmt_plane,omitempty"`
|
||||
|
||||
// System is the box's versions for the hub's System page (agent v0.142.0, R-852, `09` decision 89): Proxmox and the
|
||||
// running kernel from the Proxmox API, and the wrapper's read-only facts (host Debian, next-boot kernel, held
|
||||
// packages, taint, the crash guard; guest Debian, Docker engine, containerd, live-restore). A value nobody could
|
||||
// read is "unknown", never empty and never guessed. The hub v0.132.0 consumes it (hosts + System pages).
|
||||
System *SystemInfo `json:"system,omitempty"`
|
||||
|
||||
// PBSDR is the PBS-DR-tier bridge status stanza (slice 2). Present only when the pbsdr
|
||||
// consumer is wired. `consumed_failed` is the LOUD persistent state: the one-time token
|
||||
// secret was consumed but the apply failed afterwards — the secret is burned, the bridge
|
||||
@@ -241,6 +247,16 @@ type WireguardStatus struct {
|
||||
AssignedIP string `json:"assigned_ip,omitempty"` // from the marker, e.g. "10.77.0.2/32"
|
||||
}
|
||||
|
||||
// SystemInfo is the `system` stanza (see HostReport.System).
|
||||
type SystemInfo struct {
|
||||
PVEVersion string `json:"pve_version"` // GET /nodes/{node}/status pveversion
|
||||
KernelVersion string `json:"kernel_version"` // GET /nodes/{node}/status kversion
|
||||
VMID int `json:"vmid,omitempty"` // the customer guest the facts read
|
||||
Facts json.RawMessage `json:"facts,omitempty"`
|
||||
FactsError string `json:"facts_error,omitempty"`
|
||||
ReadAt string `json:"read_at"`
|
||||
}
|
||||
|
||||
// HostMetrics is the host block, sourced from proxmox NodeStatus.
|
||||
type HostMetrics struct {
|
||||
Node string `json:"node"`
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
package osupdate
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/base64"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
|
||||
)
|
||||
|
||||
// OpDockerStep is the signed op class of a Docker engine step (`11` §5.8): a ring-1 box takes an approved engine set,
|
||||
// and every box takes an UNDO, only through it. CC may sign it until the first paying customer (R-530 ruling).
|
||||
const OpDockerStep = "os_docker_step"
|
||||
|
||||
// DockerStepParams are the signed params. The wrapper compares Packages and Undo with the plan byte-for-byte.
|
||||
type DockerStepParams struct {
|
||||
ReleaseID string `json:"release_id"`
|
||||
Packages []Package `json:"packages"`
|
||||
Undo bool `json:"undo"`
|
||||
VMID int `json:"vmid,omitempty"`
|
||||
}
|
||||
|
||||
// DockerStepExecutor runs a verified os_docker_step (signedjobs.Executor). Guest finds the box's customer guest when
|
||||
// the params name none; Gate (optional) takes the host-wide heavy-op gate so a step never runs beside a backup.
|
||||
type DockerStepExecutor struct {
|
||||
Leg *Leg
|
||||
Guest func(ctx context.Context) (int, error)
|
||||
Gate func(ctx context.Context) (release func(), err error)
|
||||
}
|
||||
|
||||
// Execute implements signedjobs.Executor.
|
||||
func (e DockerStepExecutor) Execute(ctx context.Context, op string, params json.RawMessage) error {
|
||||
if op != OpDockerStep {
|
||||
return signedjobs.ErrNoExecutor
|
||||
}
|
||||
so, ok := signedjobs.SignedOpFrom(ctx)
|
||||
if !ok {
|
||||
return fmt.Errorf("os_docker_step: no signed envelope in the context — the wrapper could not verify it")
|
||||
}
|
||||
var p DockerStepParams
|
||||
if err := json.Unmarshal(params, &p); err != nil || len(p.Packages) == 0 {
|
||||
return fmt.Errorf("os_docker_step: params must name the engine set: %v", err)
|
||||
}
|
||||
vmid := p.VMID
|
||||
if vmid == 0 {
|
||||
if e.Guest == nil {
|
||||
return fmt.Errorf("os_docker_step: no vmid and no guest finder")
|
||||
}
|
||||
v, err := e.Guest(ctx)
|
||||
if err != nil {
|
||||
return fmt.Errorf("os_docker_step: find the customer guest: %w", err)
|
||||
}
|
||||
vmid = v
|
||||
}
|
||||
if e.Gate != nil {
|
||||
release, err := e.Gate(ctx)
|
||||
if err != nil {
|
||||
return fmt.Errorf("os_docker_step: heavy-op gate busy (a backup or restore-test runs): %w", err)
|
||||
}
|
||||
defer release()
|
||||
}
|
||||
rep := e.Leg.RunDockerSigned(ctx, vmid, p, so.Blob, string(so.Sig))
|
||||
switch rep.Outcome {
|
||||
case "applied", "nothing":
|
||||
if rep.Healthy {
|
||||
return nil
|
||||
}
|
||||
}
|
||||
return fmt.Errorf("os_docker_step: %s (%s) %s", rep.Outcome, rep.HealthReason, string(rep.Refused))
|
||||
}
|
||||
|
||||
// RunDockerSigned is one signed Docker step (ring 1 or an undo): live-restore first (decision 87, a no-op when on),
|
||||
// then the docker layer with the signed envelope, which the wrapper verifies itself.
|
||||
func (l *Leg) RunDockerSigned(ctx context.Context, vmid int, p DockerStepParams, blob []byte, sig string) Report {
|
||||
runID := l.now().UTC().Format("20060102T150405Z")
|
||||
lg := l.log().With("run", runID, "vmid", vmid, "trigger", "signed", "release", p.ReleaseID, "undo", p.Undo)
|
||||
if err := l.EnsureLiveRestore(ctx, runID, vmid); err != nil {
|
||||
return l.finish(ctx, lg, Report{RunID: runID, Layer: LayerDocker, Trigger: "signed", Ring: l.Block().Ring, VMID: vmid,
|
||||
Mode: "apply", ReleaseID: p.ReleaseID, Outcome: "failed", HealthReason: "live-restore could not be turned on: " + err.Error()})
|
||||
}
|
||||
rid := p.ReleaseID
|
||||
if rid == "" {
|
||||
rid = "signed-" + runID
|
||||
}
|
||||
return l.runLayer(ctx, runID, LayerDocker, vmid, "signed", l.Block(), dockerOpts{releaseID: rid, packages: p.Packages,
|
||||
undo: p.Undo, signed: map[string]string{"blob_b64": base64.StdEncoding.EncodeToString(blob), "sig": sig}})
|
||||
}
|
||||
@@ -0,0 +1,49 @@
|
||||
package osupdate
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/base64"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
|
||||
)
|
||||
|
||||
// The executor hands the RAW signed bytes to the wrapper (which verifies them itself) and the exact signed package
|
||||
// list. Red-proof: drop the `signed` field from the docker plan in runLayer and the plan check fails.
|
||||
func TestDockerStepExecutor_PassesTheSignedEnvelope(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
|
||||
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}}, DockerEngine: "29.8.2", Authority: "signed"}}}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
|
||||
e := DockerStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
|
||||
params, _ := json.Marshal(DockerStepParams{ReleaseID: "os-docker-1", Packages: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie", Origin: "Docker CE"}}})
|
||||
ctx := signedjobs.WithSignedOp(context.Background(), &reconcile.SignedOp{Blob: []byte(`{"op":"os_docker_step"}`), Sig: []byte("SIG")})
|
||||
if err := e.Execute(ctx, OpDockerStep, params); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
dp := w.plans[len(w.plans)-1]
|
||||
sg, _ := dp["signed"].(map[string]any)
|
||||
if dp["layer"] != "docker" || dp["lane"] != "slow" || dp["release_id"] != "os-docker-1" || sg == nil ||
|
||||
sg["blob_b64"] != base64.StdEncoding.EncodeToString([]byte(`{"op":"os_docker_step"}`)) || sg["sig"] != "SIG" {
|
||||
t.Fatalf("docker plan = %v", dp)
|
||||
}
|
||||
if calls(w) != "guest:live-restore-on,docker:apply" || len(h.reports) != 1 || h.reports[0].Trigger != "signed" {
|
||||
t.Fatalf("calls=%s reports=%+v", calls(w), h.reports)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDockerStepExecutor_RefusesWithoutEnvelopeAndPassesOtherOps(t *testing.T) {
|
||||
w := &fakeWrapper{t: t}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
|
||||
e := DockerStepExecutor{Leg: l, Guest: func(context.Context) (int, error) { return 9201, nil }}
|
||||
if err := e.Execute(context.Background(), "agent_update", nil); !errors.Is(err, signedjobs.ErrNoExecutor) {
|
||||
t.Fatalf("another op must pass through the chain: %v", err)
|
||||
}
|
||||
params, _ := json.Marshal(DockerStepParams{Packages: []Package{{Name: "docker-ce", Version: "1"}}})
|
||||
if err := e.Execute(context.Background(), OpDockerStep, params); err == nil || len(w.plans) != 0 {
|
||||
t.Fatalf("no envelope must refuse before any wrapper call: err=%v calls=%s", err, calls(w))
|
||||
}
|
||||
}
|
||||
+213
-16
@@ -1,5 +1,7 @@
|
||||
// Package osupdate is the agent's OS-update leg (`11-os-updates.md` §8 steps 2–3): the customer GUEST's Debian fast
|
||||
// lane (agent v0.140.0) and, after it in the same pass, the HOST's (agent v0.141.0).
|
||||
// Package osupdate is the agent's OS-update leg (`11-os-updates.md` §8 steps 2–3, §5.8): the customer GUEST's Debian
|
||||
// fast lane (agent v0.140.0), after it in the same pass the HOST's (agent v0.141.0), and then — ring 0 only — the
|
||||
// guest's DOCKER engine set, the slow lane (agent v0.142.0; a ring-1 box takes a Docker step only inside a signed
|
||||
// operator job, DockerStepExecutor). It also reads the box's versions for the hub's System page (Facts, R-852).
|
||||
//
|
||||
// It runs right after the night's successful whole-guest backup, while the backup goroutine still holds the host-wide
|
||||
// heavy-op gate (so it never overlaps a backup or a restore-test, `11` C10), at most once per night. All root work is
|
||||
@@ -35,10 +37,15 @@ const DefaultPlanDir = "/var/lib/felhom-agent/os"
|
||||
|
||||
// Layers.
|
||||
const (
|
||||
LayerGuest = "guest"
|
||||
LayerHost = "host"
|
||||
LayerGuest = "guest"
|
||||
LayerHost = "host"
|
||||
LayerDocker = "docker" // the guest's Docker engine set — slow lane (`11` §5.8)
|
||||
)
|
||||
|
||||
// DockerNames are the six packages of the Docker engine set (the wrapper's DOCKER_NAMES).
|
||||
var DockerNames = map[string]bool{"containerd.io": true, "docker-buildx-plugin": true, "docker-ce": true,
|
||||
"docker-ce-cli": true, "docker-ce-rootless-extras": true, "docker-compose-plugin": true}
|
||||
|
||||
// Package is one name=version with its origin.
|
||||
type Package struct {
|
||||
Name string `json:"name"`
|
||||
@@ -58,6 +65,7 @@ type Pending struct {
|
||||
type Container struct {
|
||||
State string `json:"state"`
|
||||
Health string `json:"health"` // healthy | unhealthy | starting | none
|
||||
ID string `json:"id,omitempty"`
|
||||
}
|
||||
|
||||
// Health is one health reading. Guest layer: DockerOK..Containers. Host layer: HostServices, GuestRunning and the
|
||||
@@ -67,6 +75,9 @@ type Health struct {
|
||||
NetworkOK bool `json:"network_ok"`
|
||||
Controller string `json:"controller"`
|
||||
Containers map[string]Container `json:"containers"`
|
||||
// ControllerDockerOK: the controller reaches the engine from INSIDE its container (R-858, wrapper ≥ v0.142.1;
|
||||
// nil from an older wrapper = not checked). Its own health check stayed "healthy" while it was blind.
|
||||
ControllerDockerOK *bool `json:"controller_docker_ok,omitempty"`
|
||||
HostServices map[string]string `json:"host_services,omitempty"`
|
||||
GuestRunning *bool `json:"guest_running,omitempty"`
|
||||
Guest *Health `json:"guest,omitempty"`
|
||||
@@ -84,10 +95,16 @@ type WrapperReport struct {
|
||||
RestartNeeded []string `json:"restart_needed"`
|
||||
DockerRestartNeeded bool `json:"docker_restart_needed"`
|
||||
RebootNeeded bool `json:"reboot_needed"`
|
||||
RebootScanned bool `json:"reboot_scanned"`
|
||||
HealthBefore *Health `json:"health_before"`
|
||||
HealthAfter *Health `json:"health_after"`
|
||||
Health *Health `json:"health"`
|
||||
PassSeconds float64 `json:"pass_seconds"`
|
||||
DockerEngine string `json:"docker_engine"`
|
||||
Authority string `json:"authority"`
|
||||
Undo bool `json:"undo"`
|
||||
LiveRestore json.RawMessage `json:"live_restore"`
|
||||
Facts json.RawMessage `json:"facts"`
|
||||
}
|
||||
|
||||
func (w WrapperReport) refused() bool { return len(w.Refused) > 0 && string(w.Refused) != "null" }
|
||||
@@ -112,8 +129,12 @@ type Report struct {
|
||||
RestartNeeded []string `json:"restart_needed,omitempty"`
|
||||
DockerRestartNeeded bool `json:"docker_restart_needed,omitempty"`
|
||||
RebootNeeded bool `json:"reboot_needed,omitempty"`
|
||||
RebootScanned bool `json:"reboot_scanned,omitempty"` // the pass looked (host: every pass) — a false RebootNeeded then means "not needed"
|
||||
Refused json.RawMessage `json:"refused,omitempty"`
|
||||
PassSeconds float64 `json:"pass_seconds,omitempty"`
|
||||
DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step
|
||||
Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed
|
||||
Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade)
|
||||
}
|
||||
|
||||
// Reporter posts a report to the hub (*hub.Client).
|
||||
@@ -222,6 +243,9 @@ func HealthVerdict(before, after *Health) (bool, string) {
|
||||
if after.Controller != "healthy" {
|
||||
return false, "the controller is " + after.Controller
|
||||
}
|
||||
if after.ControllerDockerOK != nil && !*after.ControllerDockerOK {
|
||||
return false, "the controller cannot reach Docker (it holds an old socket — R-858)"
|
||||
}
|
||||
if before == nil {
|
||||
return true, ""
|
||||
}
|
||||
@@ -282,6 +306,49 @@ func HostHealthVerdict(before, after *Health, tunnel string) (bool, string) {
|
||||
return true, ""
|
||||
}
|
||||
|
||||
// EngineOf is the engine version `docker version` prints for a docker-ce package version: "5:29.8.2-1~debian.13~trixie"
|
||||
// → "29.8.2" (no epoch, no Debian revision).
|
||||
func EngineOf(pkgVersion string) string {
|
||||
v := pkgVersion
|
||||
if i := strings.Index(v, ":"); i >= 0 {
|
||||
v = v[i+1:]
|
||||
}
|
||||
if i := strings.Index(v, "-"); i >= 0 {
|
||||
v = v[:i]
|
||||
}
|
||||
return v
|
||||
}
|
||||
|
||||
// DockerHealthVerdict is THE Docker-step health rule (`11` §5.8; pinned by TestDockerHealthVerdict): the guest rule,
|
||||
// plus every container running at the start still runs as the SAME container (same id — a changed id means the
|
||||
// household's apps restarted, which `live-restore` exists to prevent), plus the engine now reports the version the step
|
||||
// installed (wantEngine "" = no engine change expected).
|
||||
func DockerHealthVerdict(before, after *Health, wantEngine, gotEngine string) (bool, string) {
|
||||
if ok, why := HealthVerdict(before, after); !ok {
|
||||
return false, why
|
||||
}
|
||||
if before != nil {
|
||||
names := make([]string, 0, len(before.Containers))
|
||||
for n := range before.Containers {
|
||||
names = append(names, n)
|
||||
}
|
||||
sort.Strings(names)
|
||||
for _, n := range names {
|
||||
b := before.Containers[n]
|
||||
if b.State != "running" || b.ID == "" {
|
||||
continue
|
||||
}
|
||||
if a := after.Containers[n]; a.ID != b.ID {
|
||||
return false, n + " is a new container (id changed) — the engine step restarted it"
|
||||
}
|
||||
}
|
||||
}
|
||||
if wantEngine != "" && gotEngine != wantEngine {
|
||||
return false, "the engine is " + gotEngine + ", not " + wantEngine
|
||||
}
|
||||
return true, ""
|
||||
}
|
||||
|
||||
// call writes the plan and runs the wrapper once.
|
||||
func (l *Leg) call(ctx context.Context, runID string, plan map[string]any) (WrapperReport, error) {
|
||||
dir := l.PlanDir
|
||||
@@ -318,9 +385,85 @@ func (l *Leg) call(ctx context.Context, runID string, plan map[string]any) (Wrap
|
||||
return rep, nil // a refusal / failure is IN the report (exit 2 / 3), not an error here
|
||||
}
|
||||
|
||||
// Run is one pass: the guest layer, then (on an appliance, after a good guest step) the host layer. Returns both
|
||||
// reports (host empty when skipped). trigger is "night" or "debug".
|
||||
func (l *Leg) Run(ctx context.Context, vmid int, trigger string) (guest Report, host Report) {
|
||||
// Pass is one leg's reports; an empty Layer means the step did not run.
|
||||
type Pass struct {
|
||||
Guest, Host, Docker Report
|
||||
}
|
||||
|
||||
// Run is one pass: the guest layer, then (on an appliance, after a good guest step) the host layer, then (ring 0
|
||||
// only, after good earlier steps) the Docker engine set. trigger is "night" or "debug".
|
||||
func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Pass {
|
||||
g, h := l.runFast(ctx, vmid, trigger)
|
||||
p := Pass{Guest: g, Host: h}
|
||||
if g.Outcome == "skipped" {
|
||||
return p
|
||||
}
|
||||
blk := l.Block()
|
||||
okStep := func(r Report) bool {
|
||||
return (r.Outcome == "applied" || r.Outcome == "nothing" || r.Outcome == "inventory") && r.Healthy
|
||||
}
|
||||
lg := l.log().With("run", g.RunID, "vmid", vmid, "trigger", trigger)
|
||||
switch {
|
||||
case blk.Ring != 0 || !blk.Enabled:
|
||||
lg.Info("osupdate: docker step skipped — ring 1 takes an engine set only inside a signed operator job (`11` §5.8)", "ring", blk.Ring, "enabled", blk.Enabled)
|
||||
case !okStep(g) || (h.Layer != "" && !okStep(h)):
|
||||
lg.Warn("osupdate: docker step skipped — an earlier step did not end healthy")
|
||||
default:
|
||||
if err := l.EnsureLiveRestore(ctx, g.RunID, vmid); err != nil {
|
||||
p.Docker = l.finish(ctx, lg, Report{RunID: g.RunID, Layer: LayerDocker, Trigger: trigger, Ring: 0, VMID: vmid,
|
||||
Mode: "apply", Outcome: "failed", HealthReason: "live-restore could not be turned on: " + err.Error()})
|
||||
return p
|
||||
}
|
||||
p.Docker = l.runLayer(ctx, g.RunID, LayerDocker, vmid, trigger, blk, dockerOpts{})
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
// dockerOpts is a signed Docker step (DockerStepExecutor); the zero value is ring 0's unsigned "pending-docker".
|
||||
type dockerOpts struct {
|
||||
releaseID string
|
||||
packages []Package
|
||||
undo bool
|
||||
signed map[string]string // blob_b64, sig — the wrapper verifies them ITSELF
|
||||
}
|
||||
|
||||
// EnsureLiveRestore is the ONE-TIME step of `09` decision 87: the wrapper merges `"live-restore": true` into the guest's
|
||||
// daemon.json and RELOADS docker (never a restart, R-835). A no-op when it is already on.
|
||||
func (l *Leg) EnsureLiveRestore(ctx context.Context, runID string, vmid int) error {
|
||||
wr, err := l.call(ctx, runID, map[string]any{"release_id": "live-restore", "layer": LayerGuest, "lane": "fast",
|
||||
"vmid": vmid, "mode": "live-restore-on", "packages": []Package{}})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if wr.refused() {
|
||||
return fmt.Errorf("refused: %s", wr.Refused)
|
||||
}
|
||||
if wr.failed() {
|
||||
return fmt.Errorf("failed: %s", wr.Failed)
|
||||
}
|
||||
l.log().Info("osupdate: live-restore", "vmid", vmid, "result", string(wr.LiveRestore))
|
||||
return nil
|
||||
}
|
||||
|
||||
// Facts reads the box's versions through the wrapper's read-only facts mode (R-852): host Debian, kernels, held
|
||||
// packages, taint, the crash guard; guest Debian, Docker engine, containerd, live-restore. Raw JSON, the wrapper's shape.
|
||||
func (l *Leg) Facts(ctx context.Context, vmid int) (json.RawMessage, error) {
|
||||
wr, err := l.call(ctx, "facts"+l.now().UTC().Format("150405"), map[string]any{"release_id": "facts", "layer": LayerHost,
|
||||
"lane": "fast", "vmid": vmid, "mode": "facts", "packages": []Package{}})
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if wr.refused() {
|
||||
return nil, fmt.Errorf("facts refused: %s", wr.Refused)
|
||||
}
|
||||
if len(wr.Facts) == 0 {
|
||||
return nil, fmt.Errorf("facts: the wrapper returned none (an older wrapper?)")
|
||||
}
|
||||
return wr.Facts, nil
|
||||
}
|
||||
|
||||
// runFast is the guest + host fast lane (agent v0.141.x behaviour).
|
||||
func (l *Leg) runFast(ctx context.Context, vmid int, trigger string) (guest Report, host Report) {
|
||||
runID := l.now().UTC().Format("20060102T150405Z")
|
||||
lg := l.log().With("run", runID, "vmid", vmid, "trigger", trigger)
|
||||
if trigger == "night" && l.StatePath != "" {
|
||||
@@ -336,7 +479,7 @@ func (l *Leg) Run(ctx context.Context, vmid int, trigger string) (guest Report,
|
||||
}
|
||||
}
|
||||
blk := l.Block()
|
||||
guest = l.runLayer(ctx, runID, LayerGuest, vmid, trigger, blk)
|
||||
guest = l.runLayer(ctx, runID, LayerGuest, vmid, trigger, blk, dockerOpts{})
|
||||
if trigger == "night" && l.StatePath != "" {
|
||||
_ = os.WriteFile(l.StatePath, []byte(l.now().UTC().Format(time.RFC3339)), 0o600)
|
||||
}
|
||||
@@ -346,20 +489,27 @@ func (l *Leg) Run(ctx context.Context, vmid int, trigger string) (guest Report,
|
||||
case !(guest.Outcome == "applied" || guest.Outcome == "nothing" || guest.Outcome == "inventory") || !guest.Healthy:
|
||||
lg.Warn("osupdate: host step skipped — the guest step did not end healthy", "guest_outcome", guest.Outcome, "reason", guest.HealthReason)
|
||||
default:
|
||||
host = l.runLayer(ctx, runID, LayerHost, vmid, trigger, blk)
|
||||
host = l.runLayer(ctx, runID, LayerHost, vmid, trigger, blk, dockerOpts{})
|
||||
}
|
||||
return guest, host
|
||||
}
|
||||
|
||||
func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigger string, blk hub.WireOSUpdate) Report {
|
||||
func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigger string, blk hub.WireOSUpdate, do dockerOpts) Report {
|
||||
rel := hub.WireOSRelease{ID: "ring0-" + runID}
|
||||
var wire *hub.WireOSRelease
|
||||
if layer == LayerGuest {
|
||||
switch layer {
|
||||
case LayerGuest:
|
||||
wire = blk.Release
|
||||
} else {
|
||||
case LayerHost:
|
||||
wire = blk.HostRelease
|
||||
}
|
||||
if blk.Ring == 1 {
|
||||
lane := "fast"
|
||||
if layer == LayerDocker {
|
||||
lane = "slow"
|
||||
if do.signed != nil {
|
||||
rel = hub.WireOSRelease{ID: do.releaseID}
|
||||
}
|
||||
} else if blk.Ring == 1 {
|
||||
rel = hub.WireOSRelease{}
|
||||
if wire != nil {
|
||||
rel = *wire
|
||||
@@ -369,13 +519,23 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
|
||||
lg := l.log().With("run", runID, "layer", layer, "vmid", vmid, "ring", blk.Ring, "trigger", trigger)
|
||||
lg.Info("osupdate: START", "enabled", blk.Enabled, "release", rel.ID)
|
||||
|
||||
plan := map[string]any{"release_id": rel.ID, "layer": layer, "lane": "fast", "vmid": vmid, "snapshot": rel.Snapshot,
|
||||
plan := map[string]any{"release_id": rel.ID, "layer": layer, "lane": lane, "vmid": vmid, "snapshot": rel.Snapshot,
|
||||
"packages": []Package{}, "mode": "apply", "select": "listed"}
|
||||
if rel.ID == "" {
|
||||
plan["release_id"] = "none"
|
||||
}
|
||||
planned := map[string]bool{}
|
||||
switch {
|
||||
case layer == LayerDocker && do.signed != nil:
|
||||
plan["packages"], plan["signed"] = do.packages, do.signed
|
||||
if do.undo {
|
||||
plan["undo"] = true
|
||||
}
|
||||
for _, p := range do.packages {
|
||||
planned[p.Name] = true
|
||||
}
|
||||
case layer == LayerDocker:
|
||||
plan["select"] = "pending-docker" // ring 0: the wrapper checks the box's ROOT-OWNED ring-0 mark itself
|
||||
case !blk.Enabled:
|
||||
plan["mode"] = "inventory"
|
||||
lg.Info("osupdate: switched OFF for this box — reporting only")
|
||||
@@ -404,6 +564,7 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
|
||||
rep.Outcome, rep.Refused = "failed", wr.Failed
|
||||
}
|
||||
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
|
||||
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
|
||||
if rep.Outcome == "" {
|
||||
switch {
|
||||
case rep.Mode == "inventory" && !blk.Enabled:
|
||||
@@ -421,7 +582,16 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
|
||||
}
|
||||
// Health: compare with the start of the pass; give restarted services time (only after an install).
|
||||
cur := wr.HealthAfter
|
||||
wantEngine := ""
|
||||
for _, u := range wr.Upgraded {
|
||||
if u.Name == "docker-ce" {
|
||||
wantEngine = EngineOf(u.Version)
|
||||
}
|
||||
}
|
||||
verdict := func(h *Health) (bool, string) {
|
||||
if layer == LayerDocker {
|
||||
return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine)
|
||||
}
|
||||
if layer == LayerHost {
|
||||
t := hub.TunnelUnknown
|
||||
if l.Tunnel != nil {
|
||||
@@ -447,7 +617,7 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
|
||||
break
|
||||
}
|
||||
l.sleep(ctx, poll)
|
||||
hp := map[string]any{"release_id": plan["release_id"], "layer": layer, "lane": "fast", "vmid": vmid, "mode": "health", "packages": []Package{}}
|
||||
hp := map[string]any{"release_id": plan["release_id"], "layer": layer, "lane": lane, "vmid": vmid, "mode": "health", "packages": []Package{}}
|
||||
hr, herr := l.call(ctx, runID, hp)
|
||||
if herr == nil && hr.Health != nil {
|
||||
cur = hr.Health
|
||||
@@ -461,10 +631,37 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
|
||||
}
|
||||
rep.Installed, rep.Pending = wr.Installed, wr.Pending
|
||||
rep.RestartNeeded, rep.DockerRestartNeeded, rep.RebootNeeded = wr.RestartNeeded, wr.DockerRestartNeeded, wr.RebootNeeded
|
||||
rep.NotCovered = notCovered(wr.Pending, blk.Ring, planned)
|
||||
rep.RebootScanned = wr.RebootScanned
|
||||
if layer == LayerDocker {
|
||||
// the docker report carries the engine set only (the guest report already carries the Debian packages)
|
||||
rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending)
|
||||
rep.NotCovered = nil
|
||||
} else {
|
||||
rep.NotCovered = notCovered(wr.Pending, blk.Ring, planned)
|
||||
}
|
||||
return l.finish(ctx, lg, rep)
|
||||
}
|
||||
|
||||
func onlyDocker(in []Package) []Package {
|
||||
var out []Package
|
||||
for _, p := range in {
|
||||
if DockerNames[p.Name] {
|
||||
out = append(out, p)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func onlyDockerPending(in []Pending) []Pending {
|
||||
var out []Pending
|
||||
for _, p := range in {
|
||||
if DockerNames[p.Name] {
|
||||
out = append(out, p)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// notCovered lists pending updates no approved release covers: in ring 0 everything outside the fast lane; in ring 1
|
||||
// also every fast-lane update the release did not name.
|
||||
func notCovered(pending []Pending, ring int, planned map[string]bool) []string {
|
||||
|
||||
+158
-25
@@ -134,14 +134,14 @@ func TestRing0_OneCallPerLayer(t *testing.T) {
|
||||
LayerHost: {Upgraded: []Package{{Name: "openssl", Version: "u3"}}},
|
||||
}}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
g, ho := l.Run(context.Background(), 9201, "night")
|
||||
g, ho := run2(l, "night")
|
||||
if g.Outcome != "applied" || !g.Healthy || ho.Outcome != "applied" || !ho.Healthy {
|
||||
t.Fatalf("guest %+v\nhost %+v", g, ho)
|
||||
}
|
||||
if calls(w) != "guest:apply,host:apply" {
|
||||
t.Fatalf("calls = %s, want one apply per layer, guest first", calls(w))
|
||||
if calls(w) != "guest:apply,host:apply,guest:live-restore-on,docker:apply" {
|
||||
t.Fatalf("calls = %s, want one apply per layer, guest first, then live-restore and the ring-0 docker step", calls(w))
|
||||
}
|
||||
for _, p := range w.plans {
|
||||
for _, p := range w.plans[:2] {
|
||||
if p["select"] != "pending-fast" || p["snapshot"] != "" || len(p["packages"].([]any)) != 0 {
|
||||
t.Fatalf("ring-0 plan = %v", p)
|
||||
}
|
||||
@@ -149,7 +149,7 @@ func TestRing0_OneCallPerLayer(t *testing.T) {
|
||||
if len(g.NotCovered) != 1 || g.NotCovered[0] != "docker-ce" {
|
||||
t.Fatalf("not covered = %v", g.NotCovered)
|
||||
}
|
||||
if len(h.reports) != 2 || h.reports[0].Layer != LayerGuest || h.reports[1].Layer != LayerHost {
|
||||
if len(h.reports) != 3 || h.reports[0].Layer != LayerGuest || h.reports[1].Layer != LayerHost || h.reports[2].Layer != LayerDocker {
|
||||
t.Fatalf("hub got %+v", h.reports)
|
||||
}
|
||||
}
|
||||
@@ -163,7 +163,7 @@ func TestRing1_EachLayerItsOwnRelease(t *testing.T) {
|
||||
gr := &hub.WireOSRelease{ID: "os-g", Snapshot: "20261004T080000Z", Packages: []hub.WireOSPackage{{Name: "libc6", Version: "g-u4", Origin: "Debian"}}}
|
||||
hr := &hub.WireOSRelease{ID: "os-h", Snapshot: "20261004T090000Z", Packages: []hub.WireOSPackage{{Name: "openssl", Version: "h-u3", Origin: "Debian-Security"}}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true, Release: gr, HostRelease: hr})
|
||||
g, ho := l.Run(context.Background(), 9201, "night")
|
||||
g, ho := run2(l, "night")
|
||||
if g.ReleaseID != "os-g" || ho.ReleaseID != "os-h" {
|
||||
t.Fatalf("release ids %q %q", g.ReleaseID, ho.ReleaseID)
|
||||
}
|
||||
@@ -182,7 +182,7 @@ func TestRing1_EachLayerItsOwnRelease(t *testing.T) {
|
||||
func TestRing1_NoReleaseIsInventory(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
|
||||
g, ho := l.Run(context.Background(), 9201, "night")
|
||||
g, ho := run2(l, "night")
|
||||
if g.Outcome != "nothing" || ho.Outcome != "nothing" || calls(w) != "guest:inventory,host:inventory" {
|
||||
t.Fatalf("g=%+v h=%+v calls=%s", g, ho, calls(w))
|
||||
}
|
||||
@@ -192,7 +192,7 @@ func TestRing1_NoReleaseIsInventory(t *testing.T) {
|
||||
func TestNoBlock_IsRing1Nothing(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend}
|
||||
l, _ := newLeg(t, w, nil)
|
||||
if g, _ := l.Run(context.Background(), 9201, "night"); g.Outcome != "nothing" || g.Ring != 1 {
|
||||
if g, _ := run2(l, "night"); g.Outcome != "nothing" || g.Ring != 1 {
|
||||
t.Fatalf("g=%+v calls=%s", g, calls(w))
|
||||
}
|
||||
}
|
||||
@@ -201,7 +201,7 @@ func TestNoBlock_IsRing1Nothing(t *testing.T) {
|
||||
func TestSwitchOff_ReportsOnly(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: false})
|
||||
g, ho := l.Run(context.Background(), 9201, "night")
|
||||
g, ho := run2(l, "night")
|
||||
if g.Outcome != "inventory" || ho.Outcome != "inventory" || calls(w) != "guest:inventory,host:inventory" || len(h.reports) != 2 {
|
||||
t.Fatalf("g=%+v h=%+v calls=%s", g, ho, calls(w))
|
||||
}
|
||||
@@ -212,8 +212,9 @@ func TestBYO_NoHostPlan(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}}}}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
l.Appliance = false
|
||||
_, ho := l.Run(context.Background(), 9201, "night")
|
||||
if ho.Outcome != "" || calls(w) != "guest:apply" || len(h.reports) != 1 {
|
||||
_, ho := run2(l, "night")
|
||||
// the guest (and so its Docker engine) is ours on a BYO box too: only the HOST is the owner's
|
||||
if ho.Outcome != "" || calls(w) != "guest:apply,guest:live-restore-on,docker:apply" || len(h.reports) != 2 {
|
||||
t.Fatalf("a BYO box got a host step: host=%+v calls=%s", ho, calls(w))
|
||||
}
|
||||
}
|
||||
@@ -223,7 +224,7 @@ func TestGuestFailure_SkipsTheHost(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{
|
||||
LayerGuest: {Refused: json.RawMessage(`{"code":"R6","reason":"x"}`)}}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
g, ho := l.Run(context.Background(), 9201, "night")
|
||||
g, ho := run2(l, "night")
|
||||
if g.Outcome != "refused" || ho.Outcome != "" || calls(w) != "guest:apply" {
|
||||
t.Fatalf("g=%+v h=%+v calls=%s", g, ho, calls(w))
|
||||
}
|
||||
@@ -236,7 +237,7 @@ func TestHealth_FailsAfterTheWait(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}, HealthAfter: bad}},
|
||||
healthSeq: map[string][]*Health{LayerGuest: {bad, bad, bad, bad, bad, bad, bad, bad}}}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
g, ho := l.Run(context.Background(), 9201, "night")
|
||||
g, ho := run2(l, "night")
|
||||
if g.Outcome != "health_failed" || g.Healthy || !strings.Contains(g.HealthReason, "app was running") || ho.Outcome != "" {
|
||||
t.Fatalf("g=%+v h=%+v", g, ho)
|
||||
}
|
||||
@@ -251,7 +252,7 @@ func TestHealth_RecoversInsideTheWait(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}, HealthAfter: starting}},
|
||||
healthSeq: map[string][]*Health{LayerGuest: {starting}}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
if g, _ := l.Run(context.Background(), 9201, "night"); g.Outcome != "applied" || !g.Healthy {
|
||||
if g, _ := run2(l, "night"); g.Outcome != "applied" || !g.Healthy {
|
||||
t.Fatalf("g = %+v", g)
|
||||
}
|
||||
}
|
||||
@@ -261,7 +262,7 @@ func TestHost_TunnelDownFailsTheHostStep(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerHost: {Upgraded: []Package{{Name: "openssl"}}}}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
l.Tunnel = fakeTunnel{hub.TunnelNotRunning}
|
||||
_, ho := l.Run(context.Background(), 9201, "night")
|
||||
_, ho := run2(l, "night")
|
||||
if ho.Outcome != "health_failed" || !strings.Contains(ho.HealthReason, "tunnel") {
|
||||
t.Fatalf("host = %+v", ho)
|
||||
}
|
||||
@@ -328,12 +329,12 @@ func TestHostHealthVerdict(t *testing.T) {
|
||||
func TestOncePerNight(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}}}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
l.Run(context.Background(), 9201, "night")
|
||||
run2(l, "night")
|
||||
n := len(w.plans)
|
||||
if g, _ := l.Run(context.Background(), 9201, "night"); g.Outcome != "skipped" || len(w.plans) != n {
|
||||
if g, _ := run2(l, "night"); g.Outcome != "skipped" || len(w.plans) != n {
|
||||
t.Fatalf("a second night run in the same night ran: %+v", g)
|
||||
}
|
||||
if g, _ := l.Run(context.Background(), 9201, "debug"); g.Outcome == "skipped" {
|
||||
if g, _ := run2(l, "debug"); g.Outcome == "skipped" {
|
||||
t.Fatal("the debug action must not be throttled")
|
||||
}
|
||||
}
|
||||
@@ -344,12 +345,144 @@ func TestWrapperSuite(t *testing.T) {
|
||||
if err != nil {
|
||||
t.Skip("python3 not available")
|
||||
}
|
||||
cmd := exec.Command(py, "-B", "../../configs/test_felhom_os_apply.py")
|
||||
out, err := cmd.CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("wrapper suite failed: %v\n%s", err, out)
|
||||
}
|
||||
if !strings.Contains(string(out), "OK") {
|
||||
t.Fatalf("wrapper suite did not report OK:\n%s", out)
|
||||
// the OS wrapper and (agent v0.142.0) the crash guard — both root programs in configs/ with their own suites
|
||||
for _, suite := range []string{"../../configs/test_felhom_os_apply.py", "../../configs/test_felhom_crash_guard.py"} {
|
||||
cmd := exec.Command(py, "-B", suite)
|
||||
out, err := cmd.CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("%s failed: %v\n%s", suite, err, out)
|
||||
}
|
||||
if !strings.Contains(string(out), "OK") {
|
||||
t.Fatalf("%s did not report OK:\n%s", suite, out)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The host's restart scan result reaches the hub with reboot_scanned, so the hub can tell "looked: not needed" (a
|
||||
// reboot cleared it) from "did not look". Red-proof: drop the RebootScanned copy in runLayer and this fails.
|
||||
func TestHostReport_CarriesRebootScanned(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{
|
||||
LayerGuest: {},
|
||||
LayerHost: {RebootScanned: true, RebootNeeded: true, RestartNeeded: []string{"lxc-start"}},
|
||||
}}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
run2(l, "night")
|
||||
if len(h.reports) != 3 || !h.reports[1].RebootScanned || !h.reports[1].RebootNeeded || h.reports[0].RebootScanned {
|
||||
t.Fatalf("hub got %+v", h.reports)
|
||||
}
|
||||
}
|
||||
|
||||
// run2 is the guest + host reports of one pass (the tests written before the docker step).
|
||||
func run2(l *Leg, trigger string) (Report, Report) {
|
||||
p := l.Run(context.Background(), 9201, trigger)
|
||||
return p.Guest, p.Host
|
||||
}
|
||||
|
||||
// ---- the Docker step (`11` §5.8, agent v0.142.0) ----
|
||||
|
||||
// Ring 1 never takes an engine step in the night leg — only inside a signed operator job. Red-proof: drop the
|
||||
// `blk.Ring != 0` case in Run and the ring-1 pass makes a docker call.
|
||||
func TestDocker_Ring1NightLegNeverSteps(t *testing.T) {
|
||||
w := &fakeWrapper{t: t}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
|
||||
p := l.Run(context.Background(), 9201, "night")
|
||||
if p.Docker.Layer != "" || strings.Contains(calls(w), "docker") || strings.Contains(calls(w), "live-restore") {
|
||||
t.Fatalf("ring 1 took a docker step: %s", calls(w))
|
||||
}
|
||||
}
|
||||
|
||||
// An unhealthy earlier step skips the docker step.
|
||||
func TestDocker_SkippedAfterAnUnhealthyStep(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6"}}}},
|
||||
healthSeq: map[string][]*Health{}}
|
||||
bad := guestOK()
|
||||
bad.Controller = "unhealthy"
|
||||
w.applyRep[LayerGuest] = WrapperReport{Upgraded: []Package{{Name: "libc6"}}, HealthAfter: bad}
|
||||
w.healthSeq[LayerGuest] = []*Health{bad, bad, bad, bad, bad, bad, bad, bad}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
p := l.Run(context.Background(), 9201, "night")
|
||||
if p.Docker.Layer != "" || strings.Contains(calls(w), "docker") {
|
||||
t.Fatalf("docker step ran after an unhealthy guest step: %s", calls(w))
|
||||
}
|
||||
}
|
||||
|
||||
// The docker plan is the slow lane, pending-docker for ring 0; the report carries only the engine set.
|
||||
func TestDocker_Ring0PlanAndReport(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
|
||||
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}},
|
||||
Installed: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie", Origin: "Docker"}, {Name: "libc6", Version: "u4", Origin: "Debian"}},
|
||||
DockerEngine: "29.8.2", Authority: "ring0"}}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
p := l.Run(context.Background(), 9201, "night")
|
||||
dp := w.plans[len(w.plans)-1]
|
||||
if dp["layer"] != "docker" || dp["lane"] != "slow" || dp["select"] != "pending-docker" {
|
||||
t.Fatalf("docker plan = %v", dp)
|
||||
}
|
||||
d := p.Docker
|
||||
if d.Outcome != "applied" || !d.Healthy || d.DockerEngine != "29.8.2" || len(d.Installed) != 1 || d.Installed[0].Name != "docker-ce" {
|
||||
t.Fatalf("docker report = %+v", d)
|
||||
}
|
||||
}
|
||||
|
||||
// THE docker health rule. Red-proof: drop the id comparison (or the engine check) in DockerHealthVerdict and a case fails.
|
||||
func TestDockerHealthVerdict(t *testing.T) {
|
||||
before := guestOK()
|
||||
before.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"},
|
||||
"app": {State: "running", Health: "healthy", ID: "b"}}
|
||||
same := guestOK()
|
||||
same.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"},
|
||||
"app": {State: "running", Health: "healthy", ID: "b"}}
|
||||
moved := guestOK()
|
||||
moved.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"},
|
||||
"app": {State: "running", Health: "healthy", ID: "c"}}
|
||||
if ok, why := DockerHealthVerdict(before, same, "29.8.2", "29.8.2"); !ok {
|
||||
t.Fatalf("same ids, right engine: %s", why)
|
||||
}
|
||||
if ok, _ := DockerHealthVerdict(before, moved, "29.8.2", "29.8.2"); ok {
|
||||
t.Fatal("a changed container id passed — live-restore failed and the apps restarted")
|
||||
}
|
||||
if ok, _ := DockerHealthVerdict(before, same, "29.8.2", "29.7.2"); ok {
|
||||
t.Fatal("the engine did not move and the step passed")
|
||||
}
|
||||
if EngineOf("5:29.8.2-1~debian.13~trixie") != "29.8.2" {
|
||||
t.Fatalf("EngineOf = %q", EngineOf("5:29.8.2-1~debian.13~trixie"))
|
||||
}
|
||||
}
|
||||
|
||||
// A changed id after the step → health_failed (the consequence, not only the verdict).
|
||||
func TestDocker_ChangedIDIsHealthFailed(t *testing.T) {
|
||||
before := guestOK()
|
||||
before.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "a"}}
|
||||
after := guestOK()
|
||||
after.Containers = map[string]Container{"felhom-controller": {State: "running", Health: "healthy", ID: "z"}}
|
||||
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
|
||||
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1"}}, DockerEngine: "29.8.2",
|
||||
HealthBefore: before, HealthAfter: after}}, healthSeq: map[string][]*Health{}}
|
||||
w.healthSeq[LayerDocker] = []*Health{after, after, after, after, after, after, after, after}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
p := l.Run(context.Background(), 9201, "night")
|
||||
if p.Docker.Outcome != "health_failed" || !strings.Contains(p.Docker.HealthReason, "id changed") {
|
||||
t.Fatalf("docker = %+v", p.Docker)
|
||||
}
|
||||
}
|
||||
|
||||
// R-858: a controller that cannot reach Docker fails the health rule even though its own check says healthy.
|
||||
// Red-proof: drop the ControllerDockerOK check in HealthVerdict and this fails.
|
||||
func TestHealthVerdict_ControllerBlindToDockerFails(t *testing.T) {
|
||||
no, yes2 := false, true
|
||||
after := guestOK()
|
||||
after.ControllerDockerOK = &no
|
||||
if ok, why := HealthVerdict(guestOK(), after); ok || !strings.Contains(why, "R-858") {
|
||||
t.Fatalf("a blind controller passed: %v %q", ok, why)
|
||||
}
|
||||
if ok, _ := DockerHealthVerdict(guestOK(), after, "", ""); ok {
|
||||
t.Fatal("the Docker rule passed a blind controller")
|
||||
}
|
||||
after.ControllerDockerOK = &yes2
|
||||
if ok, why := HealthVerdict(guestOK(), after); !ok {
|
||||
t.Fatalf("a seeing controller failed: %s", why)
|
||||
}
|
||||
if ok, _ := HealthVerdict(guestOK(), guestOK()); !ok {
|
||||
t.Fatal("an older wrapper (no field) must not fail")
|
||||
}
|
||||
}
|
||||
|
||||
@@ -47,6 +47,10 @@ const (
|
||||
// Destructive for unknown classes — this named constant documents the class and keeps the
|
||||
// signed-op vocabulary explicit, it does not (and must not) loosen anything.
|
||||
ClassAgentUpdate OpClass = "agent_update"
|
||||
|
||||
// A Docker engine step in the customer guest (agent v0.142.0, `11` §5.8) — ring 1, and every undo. Destructive-class
|
||||
// (signed, operational key) like agent_update; the root wrapper re-verifies the same signature itself.
|
||||
ClassOSDockerStep OpClass = "os_docker_step"
|
||||
)
|
||||
|
||||
// Disposition is the classifier verdict.
|
||||
@@ -115,7 +119,7 @@ func Classify(class OpClass, prov Provenance) Disposition {
|
||||
return Destructive
|
||||
case ClassKeyRotation:
|
||||
return Destructive
|
||||
case ClassAgentUpdate:
|
||||
case ClassAgentUpdate, ClassOSDockerStep:
|
||||
// Never benign — no agent-internal provenance can make replacing the agent binary
|
||||
// unsigned-safe (a compromised process must not be able to self-bless an update).
|
||||
return Destructive
|
||||
|
||||
@@ -49,6 +49,21 @@ type Executor interface {
|
||||
Execute(ctx context.Context, op string, params json.RawMessage) error
|
||||
}
|
||||
|
||||
type signedOpKey struct{}
|
||||
|
||||
// WithSignedOp / SignedOpFrom carry the RAW verified envelope (blob bytes + armored signature) to an executor whose
|
||||
// ROOT half verifies it AGAIN against a root-owned key file (agent v0.142.0, the Docker slow lane: the agent's own
|
||||
// config is agent-writable, so a root wrapper must not take the agent's word for a signature).
|
||||
func WithSignedOp(ctx context.Context, s *reconcile.SignedOp) context.Context {
|
||||
return context.WithValue(ctx, signedOpKey{}, s)
|
||||
}
|
||||
|
||||
// SignedOpFrom returns the envelope set by WithSignedOp.
|
||||
func SignedOpFrom(ctx context.Context) (*reconcile.SignedOp, bool) {
|
||||
s, ok := ctx.Value(signedOpKey{}).(*reconcile.SignedOp)
|
||||
return s, ok && s != nil
|
||||
}
|
||||
|
||||
// ErrNoExecutor signals an op class with no executor wired in this build (don't clear the job).
|
||||
var ErrNoExecutor = fmt.Errorf("signedjobs: no executor for this op class in this build")
|
||||
|
||||
@@ -158,7 +173,7 @@ func (r *Runner) processJob(ctx context.Context, j hub.JobWire) bool {
|
||||
// Allowed: the nonce is already durably burned (Verify, before this point). Execute.
|
||||
r.logger.Warn("signedjobs: AUTHORIZED signed op — executing",
|
||||
"job", j.JobID, "op", ob.Op, "key_id", dec.Verified.KeyID, "nonce", dec.Verified.Nonce)
|
||||
err := r.exec.Execute(ctx, ob.Op, ob.Params)
|
||||
err := r.exec.Execute(WithSignedOp(ctx, signed), ob.Op, ob.Params)
|
||||
switch {
|
||||
case err == nil:
|
||||
r.logger.Warn("signedjobs: signed op COMPLETED", "job", j.JobID, "op", ob.Op)
|
||||
|
||||
Reference in New Issue
Block a user