Compare commits

..

8 Commits

Author SHA1 Message Date
admin 56ef1d6655 v0.145.0 code: the OS wrapper repairs dpkg's update journal by itself after a power cut (R-876); restore-test first check 30 min after start (R-874); neutral "sent late" text (R-875)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 09:23:23 +02:00
admin 78c890e4bb REPORT: the 2026-10-05 night-fixes session
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 08:25:49 +02:00
admin d48f1bbb23 v0.144.1: CHANGELOG (released 6ccd521d…, bundle e89a9ddf…)
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:49:29 +02:00
admin 8401a30917 v0.144.1 code: the wrapper survives a dead reader (BrokenPipe) so a killed pass still keeps its report; the agent looks for kept copies every 5 min (R-868, measured live)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:49:04 +02:00
admin d1b6004458 v0.144.0: CHANGELOG (released f18093c3…, bundle 6acf42fe…)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:15:07 +02:00
admin ca78c17b29 v0.144.0 code: R8 measures the real download (R-865); an OS pass's report survives a killed agent (R-868); the debug pass runs from the saved block when the hub is away (R-866)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:14:32 +02:00
admin c8d12f1f2a REPORT + CONTEXT: 2026-10-04 night (R-840 / R-860)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 20:39:47 +02:00
admin dc9164c5af CHANGELOG v0.143.0 (R-840); build-golden.sh 3.2.0: GOLDEN_GUEST_PKGS — the approved guest release at bake time, first-night count
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 20:08:16 +02:00
16 changed files with 1089 additions and 46 deletions
+68
View File
@@ -1,3 +1,71 @@
## v0.144.1 — a killed pass really keeps its report: the wrapper survives a dead reader; the agent looks again every 5 minutes (R-868, measured live) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `6ccd521d47e64999017e8eb5bc613d724543cfdc5ef9b13bcae9e3ea8c53b8f3`,
config bundle sha256 `e89a9ddfb767e177f8874d56f3dcd3bfd47409d157ff830d362653333bf815e8`. The wrapper changed again:
a box needs the signed `agent_update` AND the signed `agent_config_update`.
- **Found live on demo-hp 2026-10-05 05:45 UTC with v0.144.0** (the night's A5 shape: kill -9 of the pass and the
daemon while apt-get ran): apt finished all 13 packages, but the wrapper's next log line went to a stderr pipe no
process read any more → `BrokenPipeError` → the wrapper died before it saved its report copy (journal: `PLAN
upgrade=13`, then nothing; no copy; the hub got nothing). v0.144.0's mechanism was right and never reached.
`Runner.log` and the final `OSAPPLY-REPORT` line now survive a dead reader (the journal still gets every line).
Test `AgentDiesMidPass` drives the REAL `log()` into a pipe that breaks while apt-get runs.
- **Also found live:** the restarted daemon looked for kept copies ~7 s before the orphaned wrapper wrote one. The
daemon now looks at start and every 5 minutes (`Leg.SendUnsentLoop`); `TestR868_ACopyWrittenAfterTheStartIsSentByTheLoop`.
- Red-proofs: `felhom.eu/documentation/audits/night-fixes-2026-10-05/partD/r868-brokenpipe-red-proof.txt`.
- A second agent release in one session, against "one release per repo": recorded as `09` decision 108 (operator may
reverse) — the alternative was to ship a fix proven not to work.
## v0.144.0 — R8 measures the real download; an OS pass reports even when its agent was killed; the debug pass runs with the hub away (R-865, R-868, R-866) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `f18093c3466749ec4cd47f83f97a401160a24bad1e704f1183051adad14db928`,
config bundle `felhom-config-bundle.json` sha256 `6acf42fe46df5223384d767801cb2bf73238ba4dab82debb7811ca2f790591d8`.
The wrapper `felhom-os-apply` changed, so a box needs BOTH the signed `agent_update` and the signed `agent_config_update`.
- **R-865.** `download_bytes` runs `apt-get --print-uris` WITHOUT `-s`: with `-s` apt prints the simulation and no URI
list, so R8 summed 0 B and only its 500 MB floor ever applied. `--print-uris` alone downloads nothing (measured on
9202: the archive cache and the versions unchanged). The test fake now answers like real apt (with `-s`: no URIs),
and `test_R8_counts_the_real_download` / `test_download_bytes_never_simulates` pin it.
- **R-868.** The wrapper writes every apply pass's report to `<plan dir>/report-<run>-<layer>-apply.json` before it
prints it (root writes into the agent's dir: the dir opened O_NOFOLLOW and checked to be the agent's own, the file
created O_EXCL|O_NOFOLLOW, 0600, handed to the agent). The plan now carries `run_id`, `trigger`, `ring`, echoed in
the report. The agent deletes the copy once the hub has the report; a copy left on disk (the agent was killed, or
the hub was away) is sent at the agent's start and before every pass (`Leg.SendUnsent`), then deleted. A pass lock
(flock on `pass.lock`, across the daemon and a selftest) keeps the sender off a pass that is still running.
- **R-866.** The daemon saves the hub's newest os_update block (`os-update-block.json`); `--selftest=os-update` uses it
when the hub cannot be reached and says so in its header (`block=SAVED(<time>; hub unreachable: …)`); with no hub
and nothing saved it does not run.
- Tests: `configs/test_felhom_os_apply.py` (UnsentReport, SaveReportOnDisk — real files, symlink cases), Go
`TestR868_*`, `TestR866_*`. Red-proofs: `felhom.eu/documentation/audits/night-fixes-2026-10-05/part{C,D}/`.
## v0.143.0 — the config bundle: a signed route for a box's root-owned files (R-840, decision 96) (2026-10-04)
Released by `scripts/release-agent.sh`: binary sha256 `41c0d3060013dfda795262147454248149bee0888c0935170fdb85de6e7a35da`,
config bundle `felhom-config-bundle.json` sha256 `8d7273cf5313ef62b867cb6f831c631923a436452d6f90b8ff7f0771170396ba`.
- **The bundle.** Every root-owned file the installer's step 5 writes (sudoers ×2, the five wrappers, the crash guard
and its units, the agent and rollback units, the start-limit drop-in, the mgmt watchdog, the OOB belt's files) as ONE
reproducible JSON file, built by `scripts/build-config-bundle.py` from `BUNDLE_FILES` in `configs/felhom-os-apply`
(one table) and published beside the binary. The installer (1.31.0) installs the same file.
- **The route.** A signed `agent_config_update` {agent_version, bundle_sha256} (`felhom-opsign -op agent_config_update
-bundle-sha256 …`). The agent is the courier (downloads, checks the sha, hands over); `felhom-os-apply` mode `bundle`
verifies it ITSELF: the operator signature against the root-owned `/etc/felhom/operator-signers` (or, when that file
is missing, ONLY the installer's pinned key, after which it creates the file with exactly that key), the host binding,
the window, its own nonce; the bundle sha; every path in `BUNDLE_FILES` (R16) and never a trust file (R17); every
content check before the first write (visudo, sh/bash -n, python, unit sections, the RuntimeDirectory guard, User=,
nft -c, and that the route itself survives). Atomic per file, previous copies kept under
`/var/lib/felhom-os-apply/bundle-prev/`; a self-check after (visudo -c, `sudo -l` lists the route, the new wrapper's
`--self-check`, the self-update wrapper's usage, the crash guard's status = kernel.panic); any failure puts every
previous copy back. A newly installed crash guard is started (`enable --now`: kernel.panic for this boot, no reboot).
Record `/etc/felhom/config-bundle.json`; the agent reports it as `system.config_bundle`, the facts mode adds drift.
- **Bootstrap.** A box whose `felhom-os-apply` predates 0.143.0 cannot take the first bundle by the route (nothing on it
can write a root file from a signed job): `felhom.eu/scripts/felhom-bundle-bootstrap.sh` is the one by-hand step.
- **`build-golden.sh` 3.2.0** (not part of the binary): `GOLDEN_GUEST_PKGS` brings the template to exactly the
approved guest release (only installed packages, never newer, never a removal or a new package) and prints the
first-night count.
- Tests: `configs/test_felhom_config_bundle.py` (43; 22 of 22 mutants red), `internal/osupdate/bundle_test.go`,
`internal/hub/bundle_record_test.go`. Live: `felhom.eu/documentation/audits/r840-config-bundle-2026-10-04/partB/`.
## v0.142.1 — a Docker step no longer leaves the controller and traefik blind (R-858, `09` decision 95)
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.142.1` (`4950030`), sha256
+8
View File
@@ -1,5 +1,13 @@
# CONTEXT — felhom-agent working state
> **2026-10-04 night — v0.143.0 RELEASED + vouched (R-840, decision 96): the config bundle.** `felhom-os-apply` mode
> `bundle` (signed `agent_config_update`, verified by the wrapper itself; trust files never bundle paths) +
> `--install-bundle` (installer 1.31.0); `BUNDLE_FILES` is the one table; `scripts/build-config-bundle.py`;
> `release-agent.sh` publishes it. A box whose `felhom-os-apply` predates 0.143.0 needs ONE by-hand bootstrap
> (`felhom.eu/scripts/felhom-bundle-bootstrap.sh`) — done on both demo boxes; Tester 2 waits for the operator (R-862).
> Both demo boxes: agent 0.143.0, bundle 0.143.0 (record `/etc/felhom/config-bundle.json`). R-861 found: the sudoers is
> root-equivalent. `build-golden.sh` 3.2.0 (`GOLDEN_GUEST_PKGS`). Runbook `felhom.eu/documentation/runbooks/config-bundle.md`.
> **2026-09-25 night — v0.133.0 AND v0.134.0 DELIVERED to both demo boxes (CC-signed `agent_update`, ruling 1);
> restore test back ON (the `-1` config kept as `agent.json.night-0925-off`). v0.134.0 = R-685:** `backup/runner.go`
+19 -2
View File
@@ -1,2 +1,19 @@
- v0.142.1 (same day, ruling 95): a Docker step restarts the containers that mount the docker socket, and the health
rule checks the controller reaches Docker (R-858, found live by the operator on demo-felhom).
# REPORT — agent v0.144.0 + v0.144.1 (2026-10-05): the night's fixes
Brief: the 2026-10-05 night-fixes brief (operator), Parts C, D, E. Full session report:
`felhom.eu/REPORT-night-fixes-2026-10-05.md`. Architecture read: `11-os-updates.md` (§5.4.1 R8, §8.1–8.3).
| Row | Fix | Proof |
|---|---|---|
| R-865 | R8: `--print-uris` WITHOUT `-s` (with `-s` apt lists no URIs → 0 B) | fake answers like real apt (verbatim 9202 output); red-proof; live: the installed wrapper read 12 802 456 B for 13 pending upgrades on demo-hp |
| R-868 | the wrapper keeps its apply report beside the plan; the agent sends kept copies (start + every 5 min) and deletes them; pass lock (flock) | v0.144.0 measured NOT to work live (the wrapper died on a broken stderr pipe); v0.144.1: survives a dead reader + the 5-min look; live A5 shape on demo-hp → ONE `applied` report (13 packages) at the hub |
| R-866 | the daemon saves the hub's block; the selftest uses it with the hub away and says so | live on demo-felhom with the hub blackholed: `block=SAVED(…)`, pass ran; its kept reports reached the hub at the next start |
Released by `scripts/release-agent.sh`: v0.144.0 (`f18093c3…`, bundle `6acf42fe…`) and v0.144.1 (`6ccd521d…`, bundle
`e89a9ddf…`), both verified by download. Delivered by signed `agent_update` + `agent_config_update` to demo-hp,
demo-felhom and tester-1 (71/71 capability probe after each bundle). Vouched: agent 0.144.1, golden 0.294.0,
min_agent 0.131.0. A second release in one session is `09` decision 108.
**Found, not fixed: R-876 (P2)** — after a power cut mid-update (Part E, demo-hp) `dpkg --audit` is clean but dpkg's
update journal is not; `repair()` skips, every pass fails until `dpkg --configure -a` by hand. Next agent release.
Also R-875 (P4): the kept-report reason text is wrong for the hub-away case.
+2
View File
@@ -17,6 +17,8 @@
| `stageTemp` | internal/localapi/intermediary.go | `stageTemp(pattern, content) (path, err)` | random-named temp before a root `install` (audit B1) | Fixed /tmp names are a TOCTOU — sudoers globs expect `/tmp/felhom-*-*.ext` |
| `BUNDLE_FILES` + `Bundle` (mode `bundle`, `--install-bundle`) | configs/felhom-os-apply | the ONE table of root-owned paths + the installer of them | ANY new root-owned file the installer writes (sudoers line, wrapper, unit) — add it to the table, never a new installer fetch (R-840) | The builder (`scripts/build-config-bundle.py`) and the installer read the same table; `test_every_root_file_the_installer_writes_is_in_the_bundle` fails on a path the bundle lacks. Trust files (`/etc/felhom/os-trust.json`, `operator-signers`) are NEVER bundle paths (R17) |
| `osupdate.ConfigUpdateExecutor` | internal/osupdate/bundle.go | signed op `agent_config_update` {agent_version, bundle_sha256} | delivering the bundle to an installed box | a courier only: the root wrapper re-verifies signature, host, nonce and sha itself |
| `osupdate.Leg.SendUnsent` / `lockPass` (v0.144.0, R-868) | internal/osupdate/unsent.go | `(ctx) int` | an OS-pass report the agent never sent (killed mid-pass): the wrapper keeps `report-<run>-<layer>-apply.json` beside the plan; the agent deletes it once the hub has it | any new caller that runs an apply pass must hold `lockPass` (flock, across processes) — the sender must never take a running pass's copy |
| `osupdate.LoadSavedBlock` (v0.144.0, R-866) | internal/osupdate/leg.go | `(planDir) (block, savedAt, ok)` | the hub's newest os_update block as the daemon last received it (`os-update-block.json`) | the debug pass uses it ONLY when the hub cannot be reached, and says so in its header; no saved block → no pass |
| `guesthook.InstallSnippet` / `Register` | internal/guesthook/install.go | `InstallSnippet(ctx, runner) error` | pre-start self-heal hook install (C1 net) | Same random-temp+install pattern; snippet delegates to the agent binary (no shell logic). Issues `mkdir -p /var/lib/vz/snippets` FIRST (v0.63.0, B2 — fresh boxes lack the dir; sudoers grants exactly that argv) |
### Disk / format safety (role gates, durable IDs, format guards)
+30 -6
View File
@@ -868,6 +868,12 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
logger.Info("felhom-agent daemon starting",
"version", version, "host_id", cfg.Hub.HostID, "hub_url", cfg.Hub.URL,
"interval_s", hcfg.PollSeconds) // hub key intentionally not logged
// R-868 (v0.144.0): an OS pass whose agent was killed kept its report on disk — send it now (a pass that starts
// first sends it itself; the pass lock keeps the two apart).
// v0.144.1: and again every 5 minutes — the wrapper of a killed pass can finish AFTER the restart (measured).
go osLeg.SendUnsentLoop(ctx, 5*time.Minute, func(n int) {
logger.Info("osupdate: sent kept report(s)", "count", n)
})
// Reconcile (slice 4) runs alongside the hub loop, sharing the per-guest queue
// (doc 03 §10). At slice 4 the desired-state provider is empty (no hub serving
@@ -3575,15 +3581,15 @@ func runSelftestOSUpdate(ctx context.Context, cfg config.Config, logger *slog.Lo
return 1
}
leg := newOSLeg(cfg, client, px, logger)
resp, err := client.FetchDesiredState(ctx)
if err != nil {
fmt.Fprintln(os.Stderr, "selftest=os-update: desired state:", err)
blk, source, ok := selftestOSBlock(ctx, client, leg.PlanDir)
if !ok {
fmt.Fprintln(os.Stderr, "selftest=os-update:", source)
return 1
}
leg.SetBlock(resp.DesiredState.OSUpdate)
leg.SetBlock(blk)
b := leg.Block()
fmt.Printf("=== felhom-agent %s selftest=os-update vmid=%d ring=%d enabled=%v guest-release=%v host-release=%v appliance=%v ===\n",
version, vmid, b.Ring, b.Enabled, b.Release != nil, b.HostRelease != nil, leg.Appliance)
fmt.Printf("=== felhom-agent %s selftest=os-update vmid=%d ring=%d enabled=%v guest-release=%v host-release=%v appliance=%v block=%s ===\n",
version, vmid, b.Ring, b.Enabled, b.Release != nil, b.HostRelease != nil, leg.Appliance, source)
start := time.Now()
pass := leg.Run(ctx, vmid, "debug")
worst := pass.Guest
@@ -3609,6 +3615,24 @@ func runSelftestOSUpdate(ctx context.Context, cfg config.Config, logger *slog.Lo
return 1
}
// selftestOSBlock is the debug pass's os_update block (R-866, v0.144.0): the hub's, fetched fresh; when the hub cannot
// be reached, the block the daemon last saved — named in the selftest's first line, so a pass with the hub away can
// be exercised by hand (the daemon's own leg already ran from its last block; the selftest stopped). No saved block
// and no hub → not run (ok=false), never a guessed block.
func selftestOSBlock(ctx context.Context, f interface {
FetchDesiredState(context.Context) (*hub.DesiredStateResponse, error)
}, planDir string) (*hub.WireOSUpdate, string, bool) {
resp, err := f.FetchDesiredState(ctx)
if err == nil {
return resp.DesiredState.OSUpdate, "hub", true
}
saved, at, ok := osupdate.LoadSavedBlock(planDir)
if !ok {
return nil, fmt.Sprintf("desired state: %v — and no saved block (the daemon saves one when the hub sends it)", err), false
}
return saved, fmt.Sprintf("SAVED(%s; hub unreachable: %v)", at.UTC().Format(time.RFC3339), err), true
}
// runSelftestFacts prints the versions the host report carries (R-852, agent v0.142.0) — read-only.
//
// sudo -u felhom-agent felhom-agent --config … --selftest=os-facts -vmid 9201
@@ -0,0 +1,55 @@
package main
import (
"context"
"errors"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
)
type r866Fetcher struct{ err error }
func (f r866Fetcher) FetchDesiredState(context.Context) (*hub.DesiredStateResponse, error) {
if f.err != nil {
return nil, f.err
}
r := &hub.DesiredStateResponse{}
r.DesiredState.OSUpdate = &hub.WireOSUpdate{Ring: 1, Enabled: true}
return r, nil
}
// R-866 (v0.144.0). THE NIGHT'S SHAPE (A3, Tester 1 box, hub blocked): `selftest=os-update: desired state: hub:
// transport error … connect: invalid argument` — the debug pass could not run at all. Now it runs from the block the
// daemon saved, and its header says so.
// COMPANION RED-PROOF: return at once on a fetch error in selftestOSBlock → "the pass did not run from the saved block".
func TestR866_DebugPassUsesTheSavedBlockWhenTheHubIsAway(t *testing.T) {
dir := t.TempDir()
leg := &osupdate.Leg{PlanDir: dir}
r := &hub.DesiredStateResponse{}
r.DesiredState.OSUpdate = &hub.WireOSUpdate{Ring: 0, Enabled: true}
leg.OnDesiredState(context.Background(), r) // the daemon received a block and saved it
away := r866Fetcher{err: errors.New("hub: transport error: connect: invalid argument")}
b, src, ok := selftestOSBlock(context.Background(), away, dir)
if !ok || b == nil || b.Ring != 0 {
t.Fatalf("the pass did not run from the saved block: ok=%v block=%+v src=%q", ok, b, src)
}
if !strings.HasPrefix(src, "SAVED(") || !strings.Contains(src, "hub unreachable") {
t.Fatalf("the header must say the block is the saved one: %q", src)
}
// the hub reachable: its block wins, and the header says "hub"
b, src, ok = selftestOSBlock(context.Background(), r866Fetcher{}, dir)
if !ok || b.Ring != 1 || src != "hub" {
t.Fatalf("hub block not used: %+v %q", b, src)
}
}
// No hub and nothing saved: the pass does not run on a guessed block.
func TestR866_NoHubNoSavedBlockDoesNotRun(t *testing.T) {
_, why, ok := selftestOSBlock(context.Background(), r866Fetcher{err: errors.New("down")}, t.TempDir())
if ok || !strings.Contains(why, "no saved block") {
t.Fatalf("ok=%v why=%q", ok, why)
}
}
+43 -1
View File
@@ -61,7 +61,7 @@ set -euo pipefail
# Script provenance — logged into every bake transcript next to the baked controller tag, so an
# archive can always be traced to the script that produced it. Bump on any behavior change.
GOLDEN_SCRIPT_VERSION="3.1.0"
GOLDEN_SCRIPT_VERSION="3.2.0"
VMID="${1:-9100}"
TEMPLATE="${2:-local:vztmpl/debian-13-standard_13.1-2_amd64.tar.zst}"
@@ -137,6 +137,48 @@ pct exec "$VMID" -- env GOLDEN_DOCKER_PKGS="$GOLDEN_DOCKER_PKGS" bash -c '
fi
dpkg-query -W containerd.io docker-buildx-plugin docker-ce docker-ce-cli docker-ce-rootless-extras docker-compose-plugin 2>/dev/null | sed "s/^/ installed: /"
'
# v3.2.0 (Part F of the R-840 brief, `11` §5.3): GOLDEN_GUEST_PKGS = the newest APPROVED guest release, as
# "name=version …" (the hub's os_releases row for layer guest, IN FORCE — never a cancelled test approval). The bake
# brings every package the template HAS to exactly that version, under felhom-os-apply's rules: never a package the
# template lacks (--only-upgrade), never newer than approved, never a removal or a new package (a simulation is checked
# first and the bake FAILS on either). Empty = no approved guest release in force: the template's versions stay, and
# the box's first night installs whatever release is approved then. Either way the bake PRINTS the first-night count:
# how many installed packages are older than the approved version (target 0).
GOLDEN_GUEST_PKGS="${GOLDEN_GUEST_PKGS:-}"
pct exec "$VMID" -- env GOLDEN_GUEST_PKGS="$GOLDEN_GUEST_PKGS" bash -c '
set -e
export DEBIAN_FRONTEND=noninteractive
if [ -z "$GOLDEN_GUEST_PKGS" ]; then
echo "[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a"
echo "[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): $(apt list --upgradable 2>/dev/null | grep -c /)"
exit 0
fi
want=""
for nv in $GOLDEN_GUEST_PKGS; do
n=${nv%%=*}; v=${nv#*=}
cur=$(dpkg-query -W -f="\${Version}" "$n" 2>/dev/null) || continue # not in the template: never added
dpkg --compare-versions "$cur" lt "$v" && want="$want $n=$v"
done
if [ -n "$want" ]; then
sim=$(apt-get -s install --only-upgrade -o Dpkg::Options::=--force-confold $want)
if echo "$sim" | grep -q "^Remv "; then echo "[golden] FATAL: the approved guest set would REMOVE a package"; echo "$sim" | grep "^Remv "; exit 1; fi
for p in $(echo "$sim" | awk "/^Inst /{print \$2}"); do
dpkg-query -W "$p" >/dev/null 2>&1 || { echo "[golden] FATAL: the approved guest set would ADD $p - not in the template"; exit 1; }
done
apt-get install -y -qq --only-upgrade -o Dpkg::Options::=--force-confold -o Dpkg::Options::=--force-confdef $want >/dev/null
echo "[golden] approved guest release installed: $(echo $want | wc -w) package(s) brought to the approved version"
else
echo "[golden] approved guest release: the template already runs every approved version"
fi
left=0
for nv in $GOLDEN_GUEST_PKGS; do
n=${nv%%=*}; v=${nv#*=}
cur=$(dpkg-query -W -f="\${Version}" "$n" 2>/dev/null) || continue
dpkg --compare-versions "$cur" lt "$v" && { left=$((left+1)); echo " still older: $n $cur < $v"; }
done
echo "[golden] first-night count vs the approved guest release: $left (target 0)"
[ "$left" -eq 0 ] || { echo "[golden] FATAL: $left package(s) stayed older than the approved release"; exit 1; }
'
echo "[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …"
# containerd-snapshotter (Docker 28+/29 default) keeps the IMAGE content store under
# /var/lib/containerd — which is NOT /var/lib/docker, so it would stay on the OS rootfs and the split
+91 -10
View File
@@ -63,6 +63,9 @@ SNAPSHOT_LIST = "/etc/apt/sources.list.d/felhom-os-snapshot.list"
APT_ENV = ["env", "DEBIAN_FRONTEND=noninteractive", "APT_LISTCHANGES_FRONTEND=none", "NEEDRESTART_MODE=l", "LC_ALL=C"]
DPKG_OPTS = ["-o", "Dpkg::Options::=--force-confold", "-o", "Dpkg::Options::=--force-confdef"]
MIN_FREE = 500 * 1024 * 1024
# R-876: dpkg's state in ONE call — `--audit`, a marker line, then the update journal's file names.
JOURNAL_MARK = "@@FELHOM-DPKG-JOURNAL@@"
DPKG_STATE_SCRIPT = "dpkg --audit; echo " + JOURNAL_MARK + "; ls -A /var/lib/dpkg/updates 2>/dev/null; true"
# The installer's ROOT-OWNED record (felhom-host-install.sh `state_set mode`); the agent cannot write it.
INSTALL_STATE = "/var/lib/felhom-install/state.json"
# Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host
@@ -200,6 +203,32 @@ class Runner:
if rc != 0:
raise Refused("R7", f"could not write {path} in the guest")
def save_report(self, plan_path, report):
"""R-868: keep an apply pass's report on disk until the agent has sent it (the agent deletes it). Written
INTO the agent's own plan dir as root, so: the dir is opened with O_NOFOLLOW and must be a real directory
owned by the agent (a symlink swapped in for it is refused); the file is created O_EXCL|O_NOFOLLOW after
removing an old one, then handed to the agent (0600). Any failure only loses the copy — never the run."""
base = os.path.basename(plan_path)
name = "report-" + base[len("plan-"):]
dfd = os.open(PLAN_DIR, os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW)
try:
st = os.fstat(dfd)
if not stat.S_ISDIR(st.st_mode) or st.st_uid != self.agent_uid():
raise OSError(f"{PLAN_DIR} is not the agent's own directory")
try:
os.unlink(name, dir_fd=dfd)
except FileNotFoundError:
pass
fd = os.open(name, os.O_WRONLY | os.O_CREAT | os.O_EXCL | os.O_NOFOLLOW, 0o600, dir_fd=dfd)
try:
os.write(fd, (json.dumps(report, sort_keys=True) + "\n").encode())
os.fchown(fd, self.agent_uid(), -1)
finally:
os.close(fd)
finally:
os.close(dfd)
return os.path.join(PLAN_DIR, name)
def now(self):
return time.time()
@@ -237,7 +266,14 @@ class Runner:
os.replace(tmp, NONCE_FILE)
def log(self, line):
print(line, file=sys.stderr, flush=True)
# R-868 (v0.144.1): the agent that reads stderr may be GONE (killed mid-pass, measured live on demo-hp
# 2026-10-05): the write then raises BrokenPipeError, and v0.144.0 died right there — after apt had installed
# everything, before the report copy was saved. A dead reader must never stop the pass; the journal still
# gets every line. Pinned by AgentDiesMidPass.
try:
print(line, file=sys.stderr, flush=True)
except OSError:
pass
try:
subprocess.run(["logger", "-t", "felhom-os-apply", line], timeout=10)
except Exception:
@@ -808,6 +844,16 @@ class Apply:
plan = self.load_plan()
self.mode, self.layer, self.vmid, self.select = self.check_plan(plan)
self.report.update(mode=self.mode, layer=self.layer, release_id=plan.get("release_id"), vmid=self.vmid)
# R-868 (v0.144.0): the agent's run id, trigger and ring travel in the report, so a report the agent never
# received (it was killed mid-pass) can be sent later from the saved copy. Plain ids only; anything else
# is dropped, never refused (the plan's other checks decide).
rid, trig, ring = plan.get("run_id"), plan.get("trigger"), plan.get("ring")
if isinstance(rid, str) and re.match(r"^[A-Za-z0-9._-]{1,80}$", rid):
self.report["run_id"] = rid
if isinstance(trig, str) and re.match(r"^[a-z0-9_-]{1,20}$", trig):
self.report["trigger"] = trig
if ring in (0, 1) and not isinstance(ring, bool):
self.report["ring"] = ring
if self.mode == "facts":
return self.facts()
if self.mode == "bundle":
@@ -859,20 +905,34 @@ class Apply:
self.report["health_after"] = self.health()
return 0
def repair(self):
rc, before, _ = self.x(["dpkg", "--audit"])
def dpkg_state(self):
"""`dpkg --audit` AND dpkg's update journal, in ONE call (R-876, agent v0.145.0). A crash in the middle of an
install can leave `/var/lib/dpkg/updates/` non-empty while `--audit` reads clean — measured on demo-hp
2026-10-05 — and that journal is exactly what apt refuses on ("dpkg was interrupted"). One `sh -c` with a
constant script keeps R-845's speed: a clean pass still costs one call here, as before."""
rc, out, _ = self.x(["sh", "-c", DPKG_STATE_SCRIPT])
audit, _, journal = out.partition(JOURNAL_MARK + "\n")
return audit, [l for l in journal.split() if l]
def repair(self, force=False):
before, journal = self.dpkg_state()
configured = len([l for l in before.splitlines() if l.startswith(" ")])
fixed = 0
after = ""
if before.strip(): # nothing half-done → nothing to run (R-845: two calls saved on every clean pass)
after, journal_after = "", []
# nothing half-done and no update journal → nothing to run (R-845: two calls saved on every clean pass).
# R-876: the JOURNAL counts too, and `force` (apt said "dpkg was interrupted") always repairs.
if before.strip() or journal or force:
self.x(APT_ENV + ["dpkg", "--configure", "-a", "--force-confold"])
rc2, out, err = self.x(APT_ENV + ["apt-get", "-f", "install", "-y", "-q"] + DPKG_OPTS)
_, after, _ = self.x(["dpkg", "--audit"])
after, journal_after = self.dpkg_state()
fixed = len(re.findall(r"^Setting up ", out, re.M))
self.report["repair"] = {"half_configured_before": configured, "fixed": fixed, "clean_after": after.strip() == ""}
self.r.log(f"os-apply: REPAIR configured={configured} fixed={fixed}")
self.report["repair"] = {"half_configured_before": configured, "journal_before": len(journal), "fixed": fixed,
"clean_after": after.strip() == "" and not journal_after}
self.r.log(f"os-apply: REPAIR configured={configured} journal={len(journal)} fixed={fixed}" + (" forced" if force else ""))
if after.strip():
raise Refused("R13", "dpkg is still broken after the repair: " + after.strip().splitlines()[0])
if journal_after:
raise Refused("R13", f"dpkg's update journal is still not empty after the repair ({len(journal_after)} file(s))")
def pending_fast(self):
"""Ring 0 (select pending-fast): every pending upgrade of an INSTALLED package whose every origin is Debian /
@@ -989,6 +1049,12 @@ class Apply:
raise Refused("R8", f"free space {free} B is below max(500 MB, 3 x download {need} B)")
t0 = time.time()
rc, out, err = self.x(APT_ENV + ["apt-get", "-y", "-q"] + DPKG_OPTS + args)
if rc != 0 and "dpkg was interrupted" in (out + err):
# R-876 (belt): apt says dpkg was interrupted although the repair found nothing — repair and
# retry ONCE. Never a loop.
self.r.log("os-apply: INTERRUPTED apt says dpkg was interrupted — repairing and retrying once")
self.repair(force=True)
rc, out, err = self.x(APT_ENV + ["apt-get", "-y", "-q"] + DPKG_OPTS + args)
secs = time.time() - t0
# dpkg says "Installing new version of config file X" when X was NOT changed locally (the package's new
# version is taken), and "Configuration file 'X'" + "Keeping old config file" when it was (--force-confold
@@ -1027,7 +1093,11 @@ class Apply:
self.remove_snapshot_sources()
def download_bytes(self, args):
rc, out, _ = self.x(APT_ENV + ["apt-get", "-s", "-o", "Debug::NoLocking=1", "--print-uris", "-q"] + args)
# R-865 (v0.144.0): NO `-s`. With `-s` apt prints the simulation ("Inst …") and no URI list, so this summed
# 0 B and R8 only ever applied its 500 MB floor (measured 2026-10-04: 0 URIs with -s, 3 URIs without).
# `--print-uris` alone downloads nothing — measured on 9202 2026-10-05: the archive cache and the versions
# unchanged. Pinned by test_R8_counts_the_real_download / test_download_bytes_never_simulates.
rc, out, _ = self.x(APT_ENV + ["apt-get", "-o", "Debug::NoLocking=1", "--print-uris", "-q"] + args)
total = 0
for l in out.splitlines():
m = re.match(r"^'[^']+' \S+ ([0-9]+) ", l)
@@ -1406,7 +1476,18 @@ def main(argv, runner=None, environ=None):
a.report["failed"] = {"rc": 124, "timeout": str(e.cmd)[:200]}
rc = 3
a.report["pass_seconds"] = round(time.time() - t0, 1)
print("OSAPPLY-REPORT " + json.dumps(a.report, sort_keys=True))
if a.report.get("mode") == "apply":
# R-868: BEFORE the stdout line — a killed agent never reads stdout, and this copy is how its report still
# reaches the hub (the agent sends an unsent copy when it starts, and deletes it once sent).
try:
r.save_report(argv[2], a.report)
except Exception as e:
r.log(f"os-apply: the report copy could not be saved (the run is unaffected): {e}")
try:
print("OSAPPLY-REPORT " + json.dumps(a.report, sort_keys=True), flush=True)
except OSError:
# R-868: nobody reads stdout any more (the agent was killed); the copy above carries the report.
sys.stdout = open(os.devnull, "w")
return rc
+222 -5
View File
@@ -73,6 +73,7 @@ class Fake:
self.sig_rc = 0
self.nonces = {}
self.clock = 1791115200.0 # 2026-10-04T12:00:00Z
self.saved_reports = [] # R-868: (plan path, report) the wrapper kept on disk
self.files[osapply.TRUST_FILE] = json.dumps({"host_id": "demo-hp-bb76ea", "ring0_slow_lane": False})
self.stats[osapply.TRUST_FILE] = St(mode=statmod.S_IFREG | 0o644, uid=0)
self.files[osapply.TRUST_SIGNERS] = 'felhom-op-1 namespaces="felhom-op-v1" ssh-ed25519 AAAA\n'
@@ -113,6 +114,10 @@ class Fake:
def log(self, line):
self.logs.append(line)
def save_report(self, plan_path, report):
self.saved_reports.append((plan_path, json.loads(json.dumps(report))))
return plan_path.replace("/plan-", "/report-")
def host(self, argv, timeout=600, stdin=None):
self.calls.append(("host", argv))
if argv[0] == "/usr/sbin/pct" and argv[1] == "status":
@@ -149,6 +154,8 @@ class Fake:
if cmd == "dpkg" and a[1] == "--audit":
return 0, self.dpkg_audit, ""
if cmd == "dpkg" and a[1] == "--configure":
self.configured_calls = getattr(self, "configured_calls", 0) + 1
self.dpkg_journal = [] # `dpkg --configure -a` replays and empties the update journal
return 0, "", ""
if cmd == "fuser":
return (0, " 123", "") if self.lock_held else (1, "", "")
@@ -168,9 +175,12 @@ class Fake:
if "-f" in a:
self.dpkg_audit = ""
return 0, "Setting up x (1) ...\n" if getattr(self, "repaired", False) else "", ""
if "-s" in a:
return self.sim(a)
if "--print-uris" in a or "-s" in a:
return self.sim(a) # --print-uris prints and installs nothing, with or without -s (9202, 2026-10-05)
if "install" in a:
if getattr(self, "dpkg_journal", []):
# real apt (demo-hp 2026-10-05): a non-empty update journal refuses every install
return 100, "", "E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a' to correct the problem.\n"
if self.install_rc:
return self.install_rc, "", "E: boom"
for x in a:
@@ -215,6 +225,8 @@ class Fake:
return 0, "", ""
if cmd == "getent":
return 0, "1.2.3.4 deb.debian.org\n", ""
if cmd == "sh" and a[2] == osapply.DPKG_STATE_SCRIPT:
return 0, self.dpkg_audit + osapply.JOURNAL_MARK + "\n" + "".join(j + "\n" for j in getattr(self, "dpkg_journal", [])), ""
if cmd == "sh":
if "vmlinuz" in a[2]:
return 0, "/boot/vmlinuz-7.0.2-6-pve\n/boot/vmlinuz-7.0.14-20-pve\n", ""
@@ -232,7 +244,14 @@ class Fake:
def sim(self, a):
if "--print-uris" in a:
return 0, "'http://x/libc6.deb' libc6.deb 4000000 SHA256:x\n", ""
if "-s" in a:
# real apt (measured 9202 2026-10-05): with -s it prints the SIMULATION, no URI list
return 0, "Inst libc6 [2.41-12+deb13u4] (2.41-12+deb13u4 Debian:13.7/stable [amd64])\n", ""
# real apt's line shape, verbatim from 9202 2026-10-05 (audits/night-fixes-2026-10-05/partC/)
return 0, ("Need to get 4347 kB of archives.\n"
"'http://deb.debian.org/debian/pool/main/b/bash/bash_5.2.37-2%2bb10_amd64.deb' bash_5.2.37-2+b10_amd64.deb 1500792 MD5Sum:27b11721fea83d73b96e0f7023863771\n"
"'http://deb.debian.org/debian/pool/main/g/glibc/libc6_2.41-12%2bdeb13u4_amd64.deb' libc6_2.41-12+deb13u4_amd64.deb 2846580 MD5Sum:5559581916477ef1f57ea9f82cecf22e\n"
+ getattr(self, "extra_uris", "")), ""
if "dist-upgrade" in a:
if getattr(self, "pending_sim", None) is not None and not getattr(self, "_pending_used", False):
self._pending_used = True
@@ -273,7 +292,7 @@ class Happy(unittest.TestCase):
self.assertEqual(rep["pending"][0]["name"], "bash")
self.assertTrue(any(l.startswith("os-apply: REPAIR ") for l in f.logs), "the repair line must always print")
self.assertTrue(any(l.startswith("os-apply: DONE rc=0") for l in f.logs))
inst = [c for c in f.calls if c[0] == "guest" and "install" in c[2] and "-s" not in c[2] and "-f" not in c[2]]
inst = [c for c in f.calls if c[0] == "guest" and "install" in c[2] and "-s" not in c[2] and "-f" not in c[2] and "--print-uris" not in c[2]]
self.assertTrue(inst and "Dpkg::Options::=--force-confold" in inst[0][2], "must keep existing config files")
def test_already_current_is_a_no_op(self):
@@ -343,7 +362,7 @@ class Refusals(unittest.TestCase):
self.assertEqual(rc, 2, rep)
self.assertEqual(rep["refused"]["code"], code, rep)
self.assertTrue(any(l.startswith(f"os-apply: REFUSED: {code} ") for l in f.logs), f.logs)
inst = [c for c in f.calls if c[0] == "guest" and "install" in c[2] and "-s" not in c[2] and "-f" not in c[2]]
inst = [c for c in f.calls if c[0] == "guest" and "install" in c[2] and "-s" not in c[2] and "-f" not in c[2] and "--print-uris" not in c[2]]
self.assertEqual(inst, [], "a refusal must install nothing")
return rep
@@ -454,6 +473,24 @@ class Refusals(unittest.TestCase):
f.free = 100 * 1024 * 1024
self.refused(f, "R8")
# R-865: the download is the real one. 2 GB of URIs, 5 GB free: 5 GB < 3 x 2 GB -> R8, with the size in the line.
# COMPANION RED-PROOF: put "-s" back into download_bytes -> 0 B -> no refusal -> this test fails.
def test_R8_counts_the_real_download(self):
f = Fake()
f.free = 5 * 1024 ** 3
f.extra_uris = "'http://deb.debian.org/debian/pool/main/b/big/big_1_amd64.deb' big_1_amd64.deb 2000000000 MD5Sum:x\n"
rep = self.refused(f, "R8")
self.assertIn("download 2004347372 B", str(rep))
def test_download_bytes_never_simulates(self):
f = Fake()
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
calls = [c[2] for c in f.calls if c[0] == "guest" and "--print-uris" in c[2]]
self.assertTrue(calls, "download_bytes was never called")
for c in calls:
self.assertNotIn("-s", c, f"--print-uris with -s prints no URIs: {c}")
def test_R9_guest_locked_by_a_backup(self):
f = Fake()
f.files["/etc/pve/lxc/9201.conf"] = CONF_OK + "lock: backup\n"
@@ -962,3 +999,183 @@ class RealSignatureCheck(unittest.TestCase):
if __name__ == "__main__":
unittest.main()
class UnsentReport(unittest.TestCase):
"""R-868 (v0.144.0): an apply pass keeps its report on disk until the agent has sent it.
COMPANION RED-PROOF: drop the r.save_report call in main() -> test_apply_keeps_a_copy fails."""
def test_apply_keeps_a_copy_with_the_agents_ids(self):
f = Fake()
f.plan.update(run_id="20261005T0257-ab12", trigger="debug", ring=0)
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(len(f.saved_reports), 1, "an apply pass must keep its report on disk")
path, saved = f.saved_reports[0]
self.assertEqual(path, PLAN)
self.assertEqual(saved, rep, "the copy is the report the agent would have read")
self.assertEqual((saved["run_id"], saved["trigger"], saved["ring"]), ("20261005T0257-ab12", "debug", 0))
def test_a_refusal_is_kept_too(self):
f = Fake()
f.free = 1
rc, rep = run(f)
self.assertEqual(rc, 2)
self.assertEqual(f.saved_reports[0][1]["refused"]["code"], "R8")
def test_other_modes_keep_nothing(self):
f = Fake()
f.plan["mode"] = "health"
run(f)
self.assertEqual(f.saved_reports, [])
def test_odd_ids_are_dropped_not_trusted(self):
f = Fake()
f.plan.update(run_id="../../etc/x", trigger="Night; rm", ring=True)
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
for k in ("run_id", "trigger", "ring"):
self.assertNotIn(k, rep)
class SaveReportOnDisk(unittest.TestCase):
"""The real Runner.save_report on a temp dir: root writes into the AGENT's directory, so a symlink must never
be followed — neither for the directory nor for the file name."""
def setUp(self):
import tempfile
self.tmp = tempfile.mkdtemp()
self.dir = os.path.join(self.tmp, "os")
os.mkdir(self.dir)
self.prev = osapply.PLAN_DIR
osapply.PLAN_DIR = self.dir
self.r = osapply.Runner()
self.r.agent_uid = lambda: os.getuid()
def tearDown(self):
import shutil
osapply.PLAN_DIR = self.prev
shutil.rmtree(self.tmp)
def test_writes_0600_next_to_the_plan(self):
p = self.r.save_report(os.path.join(self.dir, "plan-r1-guest-apply.json"), {"mode": "apply"})
self.assertEqual(p, os.path.join(self.dir, "report-r1-guest-apply.json"))
st = os.stat(p)
self.assertEqual(statmod.S_IMODE(st.st_mode), 0o600)
with open(p) as fh:
self.assertEqual(json.load(fh), {"mode": "apply"})
def test_a_symlink_at_the_name_is_replaced_not_followed(self):
victim = os.path.join(self.tmp, "victim")
with open(victim, "w") as fh:
fh.write("untouched")
os.symlink(victim, os.path.join(self.dir, "report-r2-guest-apply.json"))
self.r.save_report(os.path.join(self.dir, "plan-r2-guest-apply.json"), {"mode": "apply"})
with open(victim) as fh:
self.assertEqual(fh.read(), "untouched")
self.assertFalse(os.path.islink(os.path.join(self.dir, "report-r2-guest-apply.json")))
def test_a_symlinked_directory_is_refused(self):
real = os.path.join(self.tmp, "elsewhere")
os.mkdir(real)
link = os.path.join(self.tmp, "linked")
os.symlink(real, link)
osapply.PLAN_DIR = link
with self.assertRaises(OSError):
self.r.save_report(os.path.join(link, "plan-r3-guest-apply.json"), {"mode": "apply"})
self.assertEqual(os.listdir(real), [])
class AgentDiesMidPass(unittest.TestCase):
"""R-868, MEASURED LIVE 2026-10-05 05:45 UTC on demo-hp (agent v0.144.0): the agent was kill -9-ed while apt-get
ran; apt finished all 13 packages, but the wrapper's next log line went to a stderr pipe nobody reads any more ->
BrokenPipeError -> the wrapper died before save_report: no DONE in the journal, no kept copy, no report.
COMPANION RED-PROOF: let Runner.log write to stderr unguarded -> this test fails (BrokenPipeError)."""
def test_a_dead_reader_does_not_stop_the_report_copy(self):
import io, sys, contextlib
class DeadPipe(io.TextIOBase):
dead = False
def write(self, s):
if DeadPipe.dead:
raise BrokenPipeError(32, "Broken pipe")
return len(s)
def flush(self):
if DeadPipe.dead:
raise BrokenPipeError(32, "Broken pipe")
class DyingFake(Fake):
def emulate(self, argv):
a = [x for x in argv if not re.match(r"^[A-Z_]+=", x) and x != "env"]
if a and a[0] == "apt-get" and "install" in a and "-s" not in a and "--print-uris" not in a:
DeadPipe.dead = True # the agent is killed while apt-get runs
return super().emulate(argv)
f = DyingFake()
journal = []
f.log = lambda line: osapply.Runner.log(f, line) # the REAL log(): stderr, then the journal
prev_run = osapply.subprocess.run
osapply.subprocess.run = lambda argv, *a, **kw: (journal.append(argv[-1]) if argv[0] == "logger"
else prev_run(argv, *a, **kw)) # never the real `logger`
prev_err, prev_out = sys.stderr, sys.stdout
sys.stderr, sys.stdout = DeadPipe(), DeadPipe()
try:
rc = osapply.main(["felhom-os-apply", "--plan", PLAN], runner=f)
finally:
sys.stderr, sys.stdout = prev_err, prev_out
osapply.subprocess.run = prev_run
DeadPipe.dead = False
self.assertEqual(rc, 0)
self.assertEqual(len(f.saved_reports), 1, "the kept copy is the only way this report reaches the hub")
self.assertEqual(len(f.saved_reports[0][1]["upgraded"]), 2)
self.assertTrue(any(l.startswith("os-apply: DONE") for l in journal), "the journal must still get the DONE line")
class CrashLeftTheJournal(unittest.TestCase):
"""R-876 — THE MEASURED SHAPE (demo-hp 2026-10-05, a crash while dpkg unpacked): `dpkg --audit` CLEAN, but
/var/lib/dpkg/updates holds 3 files; apt refuses every install ("dpkg was interrupted"). v0.144.1 logged
`REPAIR configured=0 fixed=0`, then `FAILED rc=100`, every pass, until a person ran `dpkg --configure -a`.
COMPANION RED-PROOFS: drop `or journal` from repair()'s condition -> test 1 fails; drop the interrupted-retry
-> test 3 fails; make repair() always run -> test 2 fails (R-845's speed)."""
def test_the_next_pass_repairs_by_itself_and_finishes(self):
f = Fake()
f.dpkg_audit = ""
f.dpkg_journal = ["0000", "0001", "0002"]
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertEqual(rep["repair"]["journal_before"], 3)
self.assertTrue(rep["repair"]["clean_after"])
self.assertEqual(len(rep["upgraded"]), 2)
cfg = next(i for i, c in enumerate(f.calls) if c[0] == "guest" and "--configure" in c[2])
inst = next(i for i, c in enumerate(f.calls) if c[0] == "guest" and "install" in c[2] and "-s" not in c[2]
and "-f" not in c[2] and "--print-uris" not in c[2])
self.assertLess(cfg, inst, "the repair must run before the install")
self.assertFalse(any("INTERRUPTED" in l for l in f.logs), "the journal check must catch it BEFORE apt refuses")
def test_a_clean_pass_still_costs_one_state_call(self):
f = Fake()
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
state = [c for c in f.calls if c[0] == "guest" and c[2][:2] == ["sh", "-c"] and c[2][2] == osapply.DPKG_STATE_SCRIPT]
self.assertEqual(len(state), 1, "R-845: a clean pass reads dpkg's state ONCE and runs no repair")
self.assertEqual(getattr(f, "configured_calls", 0), 0)
def test_apt_interrupted_is_repaired_and_retried_once(self):
"""The belt: the journal probe saw nothing (a race, an odd layout), apt still says interrupted."""
f = Fake()
orig = f.emulate
state = {"first": True}
def emulate(argv):
a = [x for x in argv if not re.match(r"^[A-Z_]+=", x) and x != "env"]
if a[0] == "apt-get" and "install" in a and "-s" not in a and "-f" not in a and "--print-uris" not in a and state["first"]:
state["first"] = False
return 100, "", "E: dpkg was interrupted, you must manually run 'sudo dpkg --configure -a' to correct the problem.\n"
return orig(argv)
f.emulate = emulate
rc, rep = run(f)
self.assertEqual(rc, 0, rep)
self.assertTrue(any("INTERRUPTED" in l for l in f.logs), f.logs)
self.assertTrue(any(l.startswith("os-apply: REPAIR ") and l.endswith("forced") for l in f.logs), f.logs)
+63
View File
@@ -0,0 +1,63 @@
package backup
import (
"context"
"fmt"
"sync/atomic"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-874 (v0.145.0). THE MEASURED SHAPE (Part F spike, Tester 2): power-on sessions of ~1.5 h and ~5 min against a
// 6 h evaluation ticker that restarts at every start — no restore-test ever evaluated. Now the first evaluation runs
// FirstEval after start.
// COMPANION RED-PROOF: drop the first-evaluation timer in Run (back to the bare ticker) → "no evaluation within".
func TestR874_FirstEvaluationAfterStart(t *testing.T) {
var picks int32
s := NewScheduler(SchedulerOptions{
Runner: &fakeRTRunner{res: reconcile.RestoreTestResult{Pass: true, Verified: "boot+running"}},
Pick: func(context.Context) (string, error) {
atomic.AddInt32(&picks, 1)
return fmt.Sprintf("local:backup/vzdump-lxc-9201-%d.tar.zst", atomic.LoadInt32(&picks)), nil
},
Store: NewStore(), Spec: (&specSpy{}).build,
Cadence: 6 * time.Hour, FirstEval: 30 * time.Millisecond, Logger: quiet(),
})
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
go func() { _ = s.Run(ctx); close(done) }()
deadline := time.Now().Add(3 * time.Second)
for atomic.LoadInt32(&picks) == 0 && time.Now().Before(deadline) {
time.Sleep(10 * time.Millisecond)
}
cancel()
<-done
if atomic.LoadInt32(&picks) == 0 {
t.Fatal("no evaluation within 3 s of start (FirstEval 30 ms) — a box with short sessions never gets a restore-test")
}
}
// The earned restraint stays: an agent that restarts before FirstEval never evaluates (a crash loop does not
// hammer a failing tier).
func TestR874_CrashLoopNeverEvaluates(t *testing.T) {
var picks int32
for i := 0; i < 5; i++ { // five quick "restarts"
s := NewScheduler(SchedulerOptions{
Runner: &fakeRTRunner{res: reconcile.RestoreTestResult{Pass: true}},
Pick: func(context.Context) (string, error) { atomic.AddInt32(&picks, 1); return "x", nil },
Store: NewStore(), Spec: (&specSpy{}).build,
Cadence: 6 * time.Hour, FirstEval: 200 * time.Millisecond, Logger: quiet(),
})
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Millisecond)
_ = s.Run(ctx)
cancel()
}
if n := atomic.LoadInt32(&picks); n != 0 {
t.Fatalf("a restart before FirstEval evaluated %d time(s)", n)
}
if DefaultFirstEval != 30*time.Minute {
t.Fatalf("DefaultFirstEval = %s, the documented 30 min", DefaultFirstEval)
}
}
+34 -6
View File
@@ -60,12 +60,21 @@ type Scheduler struct {
// R-85 tier rotation. All optional: without them the scheduler behaves exactly as before
// (single tier via `pick`), which keeps every existing caller and test working untouched.
tiers []string // configured tier target ids, primary first
tierPick TierPicker // newest archive on a named tier
rtState *RestoreTestState // persisted last-successful-per-tier (drives oldest-first)
inFlight *InFlight // shared with the backup path — Scenario F
tiers []string // configured tier target ids, primary first
tierPick TierPicker // newest archive on a named tier
rtState *RestoreTestState // persisted last-successful-per-tier (drives oldest-first)
inFlight *InFlight // shared with the backup path — Scenario F
firstEval time.Duration // R-874: the first evaluation after start
}
// DefaultFirstEval (R-874): the first due-ness evaluation runs 30 minutes after the agent starts, then every
// cadence. MEASURED need (2026-10-05 Part F spike): a box whose power-on sessions are all shorter than the 6 h
// interval (Tester 2: ~1.5 h and ~5 min) NEVER evaluated, because the ticker restarts at each start. 30 minutes
// keeps the earned restraint below — a crash-looping agent restarts far more often than that and still never
// evaluates — while a box that stays on for half an hour gets its due test. Pinned by
// TestR874_FirstEvaluationAfterStart and TestR874_CrashLoopNeverEvaluates.
const DefaultFirstEval = 30 * time.Minute
// SchedulerOptions configures a Scheduler.
type SchedulerOptions struct {
Runner RestoreTestRunner
@@ -81,6 +90,8 @@ type SchedulerOptions struct {
// 0 → no settle requirement (any archive is a candidate).
Settle time.Duration
Logger *slog.Logger
// FirstEval (R-874, v0.145.0) is when the FIRST evaluation runs after start; 0 → DefaultFirstEval.
FirstEval time.Duration
// R-85 (all optional — omit for the pre-R-85 single-tier behaviour):
// Tiers are the configured tier target ids (primary first); TierPick resolves an archive on a
@@ -110,6 +121,12 @@ func NewScheduler(opts SchedulerOptions) *Scheduler {
tierPick: opts.TierPick,
rtState: opts.State,
inFlight: opts.InFlight,
firstEval: func() time.Duration {
if opts.FirstEval > 0 {
return opts.FirstEval
}
return DefaultFirstEval
}(),
}
}
@@ -120,8 +137,9 @@ func NewScheduler(opts SchedulerOptions) *Scheduler {
// trigger any more: its phase is the process's uptime, and agent deploys reset it, which is exactly
// the defect R-86 removes. What decides that a test happens is `EvaluateDue`.
//
// It still does NOT evaluate immediately on start — the first evaluation is one interval in. That
// is an EARNED restraint, kept deliberately: a restore is heavy, agent restarts are routine, and a
// It still does NOT evaluate immediately on start. v0.145.0 (R-874): the first evaluation is
// firstEval (30 min) in, then every interval — it was one full interval in, which a box with short
// power-on sessions never reached. The restraint itself is EARNED and kept: a restore is heavy, agent restarts are routine, and a
// crash-loop that evaluated at start would hammer a permanently-failing tier as fast as it could
// restart. Due-ness does not expire while we wait, so the only cost is up to one interval of
// latency on a tier that just became due. On-demand runs use `--selftest=restore-test`.
@@ -135,6 +153,16 @@ func (s *Scheduler) Run(ctx context.Context) error {
}
s.logger.Info("backup: restore-test scheduler starting (per-archive due-check)",
"eval_interval", s.cadence, "settle", s.settle)
first := time.NewTimer(s.firstEval)
defer first.Stop()
select {
case <-ctx.Done():
s.logger.Info("backup: restore-test scheduler shutting down", "reason", ctx.Err())
return nil
case <-first.C:
s.logger.Info("backup: restore-test first evaluation after start (R-874)", "after", s.firstEval)
s.tick(ctx)
}
t := time.NewTicker(s.cadence)
defer t.Stop()
for {
+3
View File
@@ -73,6 +73,9 @@ func (e DockerStepExecutor) Execute(ctx context.Context, op string, params json.
// RunDockerSigned is one signed Docker step (ring 1 or an undo): live-restore first (decision 87, a no-op when on),
// then the docker layer with the signed envelope, which the wrapper verifies itself.
func (l *Leg) RunDockerSigned(ctx context.Context, vmid int, p DockerStepParams, blob []byte, sig string) Report {
unlock := l.lockPass(true)
defer unlock()
l.sendUnsentLocked(ctx) // R-868
runID := l.now().UTC().Format("20060102T150405Z")
lg := l.log().With("run", runID, "vmid", vmid, "trigger", "signed", "release", p.ReleaseID, "undo", p.Undo)
if err := l.EnsureLiveRestore(ctx, runID, vmid); err != nil {
+84 -16
View File
@@ -71,16 +71,16 @@ type Container struct {
// Health is one health reading. Guest layer: DockerOK..Containers. Host layer: HostServices, GuestRunning and the
// guest's own reading in Guest.
type Health struct {
DockerOK bool `json:"docker_ok"`
NetworkOK bool `json:"network_ok"`
Controller string `json:"controller"`
Containers map[string]Container `json:"containers"`
DockerOK bool `json:"docker_ok"`
NetworkOK bool `json:"network_ok"`
Controller string `json:"controller"`
Containers map[string]Container `json:"containers"`
// ControllerDockerOK: the controller reaches the engine from INSIDE its container (R-858, wrapper ≥ v0.142.1;
// nil from an older wrapper = not checked). Its own health check stayed "healthy" while it was blind.
ControllerDockerOK *bool `json:"controller_docker_ok,omitempty"`
HostServices map[string]string `json:"host_services,omitempty"`
GuestRunning *bool `json:"guest_running,omitempty"`
Guest *Health `json:"guest,omitempty"`
ControllerDockerOK *bool `json:"controller_docker_ok,omitempty"`
HostServices map[string]string `json:"host_services,omitempty"`
GuestRunning *bool `json:"guest_running,omitempty"`
Guest *Health `json:"guest,omitempty"`
}
// WrapperReport is the wrapper's OSAPPLY-REPORT object.
@@ -106,6 +106,13 @@ type WrapperReport struct {
LiveRestore json.RawMessage `json:"live_restore"`
Facts json.RawMessage `json:"facts"`
Bundle json.RawMessage `json:"bundle"` // the config bundle's result (R-840, mode "bundle")
// R-868 (v0.144.0): the agent's ids, echoed from the plan, so a report kept on disk can be sent without the
// agent process that started the pass. ReleaseID / VMID were always in the report.
RunID string `json:"run_id"`
Trigger string `json:"trigger"`
Ring *int `json:"ring"`
ReleaseID string `json:"release_id"`
VMID int `json:"vmid"`
}
func (w WrapperReport) refused() bool { return len(w.Refused) > 0 && string(w.Refused) != "null" }
@@ -136,6 +143,8 @@ type Report struct {
DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step
Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed
Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade)
unsent string // R-868: the wrapper's kept copy of this pass's report — deleted once the hub has it
}
// Reporter posts a report to the hub (*hub.Client).
@@ -162,14 +171,65 @@ type Leg struct {
block *hub.WireOSUpdate
}
// planFile / reportFile: the plan the agent writes and the copy of the report the wrapper keeps beside it (R-868).
func planFile(dir, runID, layer, mode string) string {
return filepath.Join(dir, fmt.Sprintf("plan-%s-%s-%s.json", runID, layer, mode))
}
func reportFile(dir, runID, layer, mode string) string {
return filepath.Join(dir, fmt.Sprintf("report-%s-%s-%s.json", runID, layer, mode))
}
func (l *Leg) planDir() string {
if l.PlanDir == "" {
return DefaultPlanDir
}
return l.PlanDir
}
// OnDesiredState stores the hub's os_update block (desired.RawConsumer — store only, never block).
func (l *Leg) OnDesiredState(_ context.Context, resp *hub.DesiredStateResponse) {
if resp == nil {
return
}
l.mu.Lock()
defer l.mu.Unlock()
l.block = resp.DesiredState.OSUpdate
l.mu.Unlock()
l.saveBlock(resp.DesiredState.OSUpdate)
}
// SavedBlockFile is the hub's newest os_update block as the daemon last received it (R-866, v0.144.0): the debug
// pass falls back to it when the hub cannot be reached, and says so.
const SavedBlockFile = "os-update-block.json"
type savedBlock struct {
SavedAt time.Time `json:"saved_at"`
Block *hub.WireOSUpdate `json:"block"`
}
func (l *Leg) saveBlock(b *hub.WireOSUpdate) {
dir := l.planDir()
if err := os.MkdirAll(dir, 0o700); err != nil {
return
}
body, _ := json.Marshal(savedBlock{SavedAt: l.now().UTC(), Block: b})
tmp := filepath.Join(dir, SavedBlockFile+".tmp")
if err := os.WriteFile(tmp, body, 0o600); err == nil {
_ = os.Rename(tmp, filepath.Join(dir, SavedBlockFile))
}
}
// LoadSavedBlock reads the block the daemon saved (R-866). ok=false: none saved yet.
func LoadSavedBlock(dir string) (b *hub.WireOSUpdate, savedAt time.Time, ok bool) {
raw, err := os.ReadFile(filepath.Join(dir, SavedBlockFile))
if err != nil {
return nil, time.Time{}, false
}
var s savedBlock
if json.Unmarshal(raw, &s) != nil {
return nil, time.Time{}, false
}
return s.Block, s.SavedAt, true
}
// Block returns the newest os_update block. No block (an older hub, or nothing fetched yet) = ring 1, ON, no
@@ -352,15 +412,12 @@ func DockerHealthVerdict(before, after *Health, wantEngine, gotEngine string) (b
// call writes the plan and runs the wrapper once.
func (l *Leg) call(ctx context.Context, runID string, plan map[string]any) (WrapperReport, error) {
dir := l.PlanDir
if dir == "" {
dir = DefaultPlanDir
}
dir := l.planDir()
if err := os.MkdirAll(dir, 0o700); err != nil {
return WrapperReport{}, fmt.Errorf("osupdate: plan dir: %w", err)
}
b, _ := json.Marshal(plan)
path := filepath.Join(dir, fmt.Sprintf("plan-%s-%s-%s.json", runID, plan["layer"], plan["mode"]))
path := planFile(dir, runID, fmt.Sprint(plan["layer"]), fmt.Sprint(plan["mode"]))
if err := os.WriteFile(path, b, 0o600); err != nil {
return WrapperReport{}, fmt.Errorf("osupdate: write plan: %w", err)
}
@@ -394,6 +451,9 @@ type Pass struct {
// Run is one pass: the guest layer, then (on an appliance, after a good guest step) the host layer, then (ring 0
// only, after good earlier steps) the Docker engine set. trigger is "night" or "debug".
func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Pass {
unlock := l.lockPass(true)
defer unlock()
l.sendUnsentLocked(ctx) // R-868: a report a killed agent never sent goes first
g, h := l.runFast(ctx, vmid, trigger)
p := Pass{Guest: g, Host: h}
if g.Outcome == "skipped" {
@@ -521,7 +581,8 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
lg.Info("osupdate: START", "enabled", blk.Enabled, "release", rel.ID)
plan := map[string]any{"release_id": rel.ID, "layer": layer, "lane": lane, "vmid": vmid, "snapshot": rel.Snapshot,
"packages": []Package{}, "mode": "apply", "select": "listed"}
"packages": []Package{}, "mode": "apply", "select": "listed",
"run_id": runID, "trigger": trigger, "ring": blk.Ring} // R-868: echoed into the wrapper's kept copy
if rel.ID == "" {
plan["release_id"] = "none"
}
@@ -554,6 +615,9 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
}
rep.Mode = plan["mode"].(string)
wr, err := l.call(ctx, runID, plan)
if rep.Mode == "apply" {
rep.unsent = reportFile(l.planDir(), runID, layer, rep.Mode) // the wrapper kept a copy (R-868)
}
switch {
case err != nil:
rep.Outcome, rep.HealthReason = "failed", err.Error()
@@ -684,7 +748,11 @@ func (l *Leg) finish(ctx context.Context, lg *slog.Logger, rep Report) Report {
rctx, cancel := context.WithTimeout(context.WithoutCancel(ctx), time.Minute)
defer cancel()
if err := l.Hub.PostOSReport(rctx, body); err != nil {
lg.Warn("osupdate: reporting to the hub failed (the run itself is done)", "err", err)
lg.Warn("osupdate: reporting to the hub failed (the run itself is done; the kept copy is sent at the next start or pass)", "err", err)
} else if rep.unsent != "" {
if rerr := os.Remove(rep.unsent); rerr != nil && !os.IsNotExist(rerr) {
lg.Warn("osupdate: could not delete the sent report's kept copy (it may be sent twice)", "path", rep.unsent, "err", rerr)
}
}
}
return rep
+18
View File
@@ -3,6 +3,7 @@ package osupdate
import (
"context"
"encoding/json"
"fmt"
"io"
"os"
"os/exec"
@@ -21,6 +22,7 @@ type fakeWrapper struct {
applyRep map[string]WrapperReport // per layer
healthSeq map[string][]*Health // per layer: answers to successive "health" calls
plans []map[string]any
keep bool // R-868: like the real wrapper, keep an apply report beside the plan
}
func yes() *bool { b := true; return &b }
@@ -72,6 +74,22 @@ func (f *fakeWrapper) Run(_ context.Context, name string, args ...string) ([]byt
rep.Health = ok
}
}
if f.keep && plan["mode"] == "apply" {
kept := rep
kept.Layer, kept.RunID, _ = layer, fmt.Sprint(plan["run_id"]), 0
if tr, ok := plan["trigger"].(string); ok {
kept.Trigger = tr
}
if r, ok := plan["ring"].(float64); ok {
ri := int(r)
kept.Ring = &ri
}
kb, _ := json.Marshal(kept)
dst := filepath.Join(filepath.Dir(args[1]), "report-"+strings.TrimPrefix(filepath.Base(args[1]), "plan-"))
if err := os.WriteFile(dst, kb, 0o600); err != nil {
f.t.Fatal(err)
}
}
out, _ := json.Marshal(rep)
return []byte("OSAPPLY-REPORT " + string(out) + "\n"), []byte("os-apply: DONE rc=0\n"), nil
}
+197
View File
@@ -0,0 +1,197 @@
package osupdate
import (
"context"
"encoding/json"
"os"
"path/filepath"
"strings"
"syscall"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// ── R-868 (v0.144.0): a pass whose agent was killed still reports ─────────────────────────────────────
//
// MEASURED 2026-10-05 02:57 UTC on demo-hp (night drill A5): the debug pass and the agent daemon were kill -9-ed while
// apt-get ran. The root wrapper (its own process under sudo) finished all six packages, but the agent that would have
// read its stdout and posted the report was gone; the hub never learned what the pass installed.
//
// THE MECHANISM: the wrapper writes its apply report to <plan dir>/report-<run>-<layer>-apply.json BEFORE printing
// it (configs/felhom-os-apply save_report). The agent deletes that copy once the hub has the report (finish). A copy
// still on disk is a report nobody sent: SendUnsent posts it — at the agent's start, and before every pass — and then
// deletes it. A pass lock (flock on <plan dir>/pass.lock, released by the kernel when a process dies) keeps the
// sender from picking up the copy of a pass that is still running, also across the daemon and a selftest process.
// Pinned by TestR868_* (unsent_test.go).
// lockPass takes the pass lock. block=false returns ok=false at once when another pass holds it. A lock that cannot
// be opened at all (no plan dir yet) does not stop a pass: the unlock is then a no-op.
func (l *Leg) lockPass(block bool) (unlock func()) {
u, _ := l.tryLockPass(block)
return u
}
func (l *Leg) tryLockPass(block bool) (unlock func(), ok bool) {
dir := l.planDir()
_ = os.MkdirAll(dir, 0o700)
f, err := os.OpenFile(filepath.Join(dir, "pass.lock"), os.O_CREATE|os.O_RDWR, 0o600)
if err != nil {
l.log().Warn("osupdate: pass lock unavailable — continuing without it", "err", err)
return func() {}, true
}
how := syscall.LOCK_EX
if !block {
how |= syscall.LOCK_NB
}
if err := syscall.Flock(int(f.Fd()), how); err != nil {
f.Close()
return func() {}, false
}
return func() { _ = syscall.Flock(int(f.Fd()), syscall.LOCK_UN); f.Close() }, true
}
// SendUnsent posts every report a pass kept on disk and nobody sent (R-868). It skips when a pass runs now (that
// pass sends them first). Called at the agent's start.
func (l *Leg) SendUnsent(ctx context.Context) int {
unlock, ok := l.tryLockPass(false)
if !ok {
l.log().Info("osupdate: a pass is running — its start sends any kept report")
return 0
}
defer unlock()
return l.sendUnsentLocked(ctx)
}
// SendUnsentLoop runs SendUnsent now and then every `every` until ctx ends (v0.144.1). MEASURED live on demo-hp
// 2026-10-05: after a kill -9 the daemon restarted in ~5 s, while the orphaned root wrapper was still installing — its
// copy appeared ~7 s AFTER the start-time sender had looked. One look at start is therefore not enough. A glob of the
// plan dir every few minutes costs nothing; the pass lock keeps it off a running pass. Pinned by
// TestR868_ACopyWrittenAfterTheStartIsSentByTheLoop.
func (l *Leg) SendUnsentLoop(ctx context.Context, every time.Duration, onSent func(int)) {
for {
if n := l.SendUnsent(ctx); n > 0 && onSent != nil {
onSent(n)
}
select {
case <-ctx.Done():
return
case <-time.After(every):
}
}
}
func (l *Leg) sendUnsentLocked(ctx context.Context) int {
if l.Hub == nil {
return 0 // nobody to send to: keep the copies for a process that has the hub
}
files, _ := filepath.Glob(filepath.Join(l.planDir(), "report-*.json"))
sent := 0
for _, f := range files {
b, err := os.ReadFile(f)
if err != nil {
l.log().Warn("osupdate: a kept report cannot be read", "path", f, "err", err)
continue
}
var wr WrapperReport
if err := json.Unmarshal(b, &wr); err != nil || wr.Layer == "" {
l.log().Warn("osupdate: a kept report is not a report — moved aside", "path", f, "err", err)
_ = os.Rename(f, f+".bad")
continue
}
rep := l.reportFromKept(ctx, wr, f)
lg := l.log().With("run", rep.RunID, "layer", rep.Layer, "vmid", rep.VMID, "ring", rep.Ring, "trigger", rep.Trigger)
lg.Info("osupdate: sending a kept report late (the agent was stopped mid-pass or the hub was away, R-868)", "path", f)
before := rep.unsent
_ = l.finish(ctx, lg, rep)
if _, err := os.Stat(before); os.IsNotExist(err) {
sent++
// the pass's plan file is left behind too when the agent was killed inside call()
_ = os.Remove(filepath.Join(filepath.Dir(f), "plan-"+strings.TrimPrefix(filepath.Base(f), "report-")))
}
}
return sent
}
// reportFromKept builds the hub report from a kept wrapper report, as runLayer would have. Health: the copy's own
// before/after reading; when that is not healthy after an install, one fresh reading (services restart after an
// install, and the pass that would have waited for them is gone).
func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string) Report {
ring := 1
if wr.Ring != nil {
ring = *wr.Ring
}
runID, trigger := wr.RunID, wr.Trigger
if runID == "" {
runID = strings.TrimSuffix(strings.TrimPrefix(filepath.Base(path), "report-"), ".json")
}
if trigger == "" {
trigger = "unknown"
}
rep := Report{RunID: runID, Layer: wr.Layer, Trigger: trigger, Mode: wr.Mode, Ring: ring, ReleaseID: wr.ReleaseID,
VMID: wr.VMID, unsent: path}
prefix := "sent late — kept on the box until the hub could take it (R-868, R-875)"
switch {
case wr.refused():
rep.Outcome, rep.Refused, rep.HealthReason = "refused", wr.Refused, prefix
return rep
case wr.failed():
rep.Outcome, rep.Refused = "failed", wr.Failed
case len(wr.Upgraded) == 0:
rep.Outcome = "nothing"
default:
rep.Outcome = "applied"
}
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
wantEngine := ""
for _, u := range wr.Upgraded {
if u.Name == "docker-ce" {
wantEngine = EngineOf(u.Version)
}
}
verdict := func(h *Health) (bool, string) {
switch wr.Layer {
case LayerDocker:
return DockerHealthVerdict(wr.HealthBefore, h, wantEngine, wr.DockerEngine)
case LayerHost:
t := hub.TunnelUnknown
if l.Tunnel != nil {
t, _ = l.Tunnel.Status(ctx)
}
return HostHealthVerdict(wr.HealthBefore, h, t)
}
return HealthVerdict(wr.HealthBefore, h)
}
ok, why := verdict(wr.HealthAfter)
if !ok && len(wr.Upgraded) > 0 && wr.VMID > 0 {
lane := "fast"
if wr.Layer == LayerDocker {
lane = "slow"
}
if hr, err := l.call(ctx, "kept-"+runID, map[string]any{"release_id": "kept", "layer": wr.Layer, "lane": lane,
"vmid": wr.VMID, "mode": "health", "packages": []Package{}}); err == nil && hr.Health != nil {
ok, why = verdict(hr.Health)
}
}
rep.Healthy, rep.HealthReason = ok, prefix
if why != "" {
rep.HealthReason = prefix + ": " + why
}
if !ok && rep.Outcome == "applied" {
rep.Outcome = "health_failed"
}
rep.Installed, rep.Pending = wr.Installed, wr.Pending
rep.RestartNeeded, rep.DockerRestartNeeded, rep.RebootNeeded = wr.RestartNeeded, wr.DockerRestartNeeded, wr.RebootNeeded
rep.RebootScanned = wr.RebootScanned
if wr.Layer == LayerDocker {
rep.Installed, rep.Pending = onlyDocker(wr.Installed), onlyDockerPending(wr.Pending)
} else {
planned := map[string]bool{}
for _, u := range wr.Upgraded {
planned[u.Name] = true
}
rep.NotCovered = notCovered(wr.Pending, ring, planned)
}
return rep
}
+152
View File
@@ -0,0 +1,152 @@
package osupdate
import (
"context"
"encoding/json"
"errors"
"os"
"path/filepath"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-868 (v0.144.0). THE NIGHT'S SHAPE (A5, demo-hp 2026-10-05 02:57 UTC): the wrapper finished six packages, the agent
// was kill -9-ed before it read the report; the hub never got it. Here: the wrapper's kept copy and the plan file are
// on disk, a NEW agent process starts — it must send exactly one `applied` report and delete both files.
// COMPANION RED-PROOF: drop the SendUnsent body (return 0) → "the hub got no report".
func TestR868_KilledPassIsReportedAtStart(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
ring := 0
kept := WrapperReport{Mode: "apply", Layer: LayerGuest, RunID: "20261005T025700Z", Trigger: "debug", Ring: &ring, VMID: 9201,
ReleaseID: "ring0-20261005T025700Z", HealthBefore: guestOK(), HealthAfter: guestOK(),
Upgraded: []Package{{Name: "libc6", Version: "u4"}, {Name: "openssl", Version: "u3"}}}
b, _ := json.Marshal(kept)
rp := reportFile(l.PlanDir, kept.RunID, LayerGuest, "apply")
pp := planFile(l.PlanDir, kept.RunID, LayerGuest, "apply")
must(t, os.WriteFile(rp, b, 0o600))
must(t, os.WriteFile(pp, []byte("{}"), 0o600))
if n := l.SendUnsent(context.Background()); n != 1 {
t.Fatalf("sent %d, want 1", n)
}
if len(h.reports) != 1 {
t.Fatalf("the hub got no report (or several): %+v", h.reports)
}
r := h.reports[0]
if r.Outcome != "applied" || !r.Healthy || r.Trigger != "debug" || r.RunID != kept.RunID || r.Ring != 0 || len(r.Upgraded) != 2 || r.VMID != 9201 {
t.Fatalf("report = %+v", r)
}
// R-875 (v0.145.0): neutral — the copy cannot tell a killed agent from an absent hub.
if !strings.HasPrefix(r.HealthReason, "sent late") || strings.Contains(r.HealthReason, "stopped mid-pass") {
t.Fatalf("health_reason = %q, want the neutral \"sent late …\"", r.HealthReason)
}
for _, p := range []string{rp, pp} {
if _, err := os.Stat(p); !os.IsNotExist(err) {
t.Fatalf("%s still on disk after the hub got it", filepath.Base(p))
}
}
// a second start sends nothing again — no duplicate report
if n := l.SendUnsent(context.Background()); n != 0 || len(h.reports) != 1 {
t.Fatalf("sent again: %d, reports %d", n, len(h.reports))
}
}
// A normal pass: the hub gets ONE report per layer and the kept copies are gone after it (nothing resent later).
// COMPANION RED-PROOF: drop the os.Remove(rep.unsent) in finish → the next pass resends → "2 guest reports".
func TestR868_NormalPassLeavesNoCopyAndNoDuplicate(t *testing.T) {
w := &fakeWrapper{t: t, keep: true, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6", Version: "u4"}}}}}
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Run(context.Background(), 9201, "debug")
left, _ := filepath.Glob(filepath.Join(l.PlanDir, "report-*.json"))
if len(left) != 0 {
t.Fatalf("kept copies left after the hub got them: %v", left)
}
l.Run(context.Background(), 9201, "debug") // the next pass sends kept copies first
guest := 0
for _, r := range h.reports {
if r.Layer == LayerGuest && r.Outcome == "applied" {
guest++
}
}
if guest != 2 {
t.Fatalf("%d guest reports for 2 passes (a duplicate or a loss)", guest)
}
}
type failingHub struct{ n int }
func (h *failingHub) PostOSReport(context.Context, []byte) error {
h.n++
return errors.New("hub away")
}
// The hub away: the copy stays, and goes at the next chance.
func TestR868_HubAwayKeepsTheCopy(t *testing.T) {
w := &fakeWrapper{t: t, keep: true, applyRep: map[string]WrapperReport{LayerGuest: {Upgraded: []Package{{Name: "libc6", Version: "u4"}}}}}
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
l.Appliance = false
l.Hub = &failingHub{}
l.Run(context.Background(), 9201, "night")
left, _ := filepath.Glob(filepath.Join(l.PlanDir, "report-*.json"))
if len(left) != 2 { // ring 0: the guest step and the Docker step each kept one
t.Fatalf("the copies must stay while the hub is away: %v", left)
}
h := &fakeHub{}
l.Hub = h
if n := l.SendUnsent(context.Background()); n != 2 {
t.Fatalf("sent %d: %+v", n, h.reports)
}
for _, r := range h.reports {
if r.Trigger != "night" || r.Ring != 0 {
t.Fatalf("the kept report lost its ids: %+v", r)
}
}
}
// A pass in progress holds the lock: the sender at start must not take that pass's copy (it would be sent twice).
func TestR868_RunningPassKeepsTheSenderOff(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, nil)
must(t, os.WriteFile(reportFile(l.PlanDir, "r1", LayerGuest, "apply"), []byte(`{"mode":"apply","layer":"guest"}`), 0o600))
unlock := l.lockPass(true)
if n := l.SendUnsent(context.Background()); n != 0 || len(h.reports) != 0 {
t.Fatalf("sent while a pass ran: %d", n)
}
unlock()
if n := l.SendUnsent(context.Background()); n != 1 {
t.Fatalf("not sent after the pass: %d", n)
}
}
// v0.144.1 — THE LIVE SHAPE (demo-hp 2026-10-05 05:45 UTC): the restarted daemon looked at 05:45:09, the orphaned
// wrapper wrote its copy at ~05:45:16. The loop must still send it.
// COMPANION RED-PROOF: make SendUnsentLoop return after the first look → "the late copy was never sent".
func TestR868_ACopyWrittenAfterTheStartIsSentByTheLoop(t *testing.T) {
w := &fakeWrapper{t: t}
l, h := newLeg(t, w, nil)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
sent := make(chan int, 4)
go l.SendUnsentLoop(ctx, 20*time.Millisecond, func(n int) { sent <- n })
time.Sleep(50 * time.Millisecond) // the start-time look found nothing
must(t, os.WriteFile(reportFile(l.PlanDir, "late", LayerGuest, "apply"),
[]byte(`{"mode":"apply","layer":"guest","run_id":"late","trigger":"debug","ring":0,"vmid":9201,"upgraded":[{"name":"openssl","version":"u3"}]}`), 0o600))
select {
case n := <-sent:
if n != 1 || len(h.reports) != 1 || h.reports[0].RunID != "late" || h.reports[0].Outcome == "" {
t.Fatalf("sent %d: %+v", n, h.reports)
}
case <-time.After(3 * time.Second):
t.Fatal("the late copy was never sent")
}
}
func must(t *testing.T, err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}