Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
11 KiB
REPORT — the config bundle (R-840), test approvals (R-859), the Docker-socket self-heal (R-860), Tester 2, drill-r50
2026-10-04, evening–night. Brief: "a signed route that brings an installed box's root-owned parts up to date (R-840)…".
Architecture read first: 11-os-updates.md (§5.3, §5.4, §5.8, §5.9, §8), 03-host-agent.md (§3, §4, §11),
04-control-plane-authorization.md (§3), 07-backup-architecture.md (§6.4). Evidence: documentation/audits/r840-config-bundle-2026-10-04/.
The Part table
| Part | Result | Why / what changed |
|---|---|---|
| A — what Tester 2 has | done (read through the hub; no route to the box) | The brief's timestamp reasoning was wrong (UTC vs local). Tester 2 has installer 1.30.0, agent 0.142.0, the crash guard armed, live-restore on, Docker 29.8.2, the re-made golden. It lacks only the R-858 wrapper fix. |
A — the felhom-pbs warning |
done: not a fault | By design (R-723): a new box's first-hour tier skip is recorded, never mailed; the tier came 7 min later and backed up at 16:31 UTC. No row. |
| B — the bundle route | done, live on both demo boxes | Agent 0.143.0 (bundle mode in felhom-os-apply, signed agent_config_update), hub 0.133.0, installer 1.31.0. Changed: the installer is NOT the self-update wrapper (it cannot do it); a one-time bootstrap is needed on older boxes. |
| C — Tester 2 | act 1 sent, NOT delivered (box offline since 18:06 UTC); act 2 waits for the operator; act 3 not needed | Act 1: queued 18:35 UTC — see "Part C" below. Act 2 needs the one by-hand bootstrap (R-862, the operator's tunnel). Act 3: live-restore is already on. |
| D — test approvals end | done, live | Hub 0.133.0 cancelled the 4 test approvals of 2026-10-04 at start (18:20 UTC). Changed: the operator's own Docker button approval was not backfilled (CC decision). |
| E — self-heal after any Docker restart | done, live (9202 + demo-hp) | Controller 0.293.0. Changed: only a docker.socket restart breaks it; a dockerd crash or systemctl restart docker does not. |
| F — new installs start with the approved fixes | built; tonight's golden carries none | build-golden.sh 3.2.0 GOLDEN_GUEST_PKGS. No guest release is in force tonight (the test one was cancelled), so the golden keeps the template (49 Debian updates pending for the next real approval). |
| G — drill-r50 | done | Deleted through the product (operator's yes: it touched ep0). On ep0: its PBS namespace did not exist; its WireGuard peer left (5 → 4 peers pushed). |
| H — release, golden, records | done | Agent 0.143.0, controller 0.293.0, hub 0.133.0, installer 1.31.0, golden 0.293.0 baked + vouched. Floor 0.293.0 for the two demo customers only (Tester 2 untouched). |
Claims in the brief that turned out wrong
- The timestamp reasoning about Tester 2 (§1). Tester 2 was bound at 16:06 UTC = 18:06 local; the releases were
compared in local time. At 18:06 local, installer 1.30.0 (17:36), agent 0.142.0 (vouched ~17:49) and the re-made golden
(17:48) were all current. Measured (hub records): agent 0.142.0, crash guard armed (
kernel.panic10),live-restoreon, Docker 29.8.2; the wrapper reports a facts mode (so it is not "old"). Only agent 0.142.1's R-858 wrapper fix is missing — and the root files of 0.142.0 equal 0.143.0's in every file butfelhom-os-apply. - "The felhom-pbs warning is a fault." It is the designed first-hour behaviour (R-723): recorded, not mailed.
- "The self-update wrapper can install a bundle safely." It cannot: it is a fixed sh script that swaps one binary. The
bundle is installed by
felhom-os-apply(already reachable through the agent's sudoers line). And no existing root helper can install the first bundle on an older box — one by-hand bootstrap per older box (done on both demo boxes). - "A restart of an existing container picks up the new socket." True — measured twice (controller exit → restart policy;
docker restart traefik): both came back on the current inode, same container ids. - "dockerd restarts on its own … the box shows DOWN until a person acts." Only when the socket FILE is re-created
(
docker.socketrestart, i.e. a docker-ce upgrade or by hand). A dockerd crash orsystemctl restart dockerkeeps it (measured). - Act 3 for Tester 2 (
live-restorereload) is not needed: it is already on (from the golden).
Part A — Tester 2, read back (hub records; CC has no route to the box)
| Item | Brief expected | Measured | Source |
|---|---|---|---|
| bound / enrolled | 16:06 (read as local) | 16:06:27 UTC bind, 16:07:04 UTC host | appliance_registrations, hosts |
| installer | 1.29.0 | 1.30.0 (the crash guard is installed — only 1.30.0 does that) | host report system.facts.host.crash_guard |
| agent | 0.141.1 | 0.142.0 | /hosts |
| PBS wrapper sha | — | 104db0a4… = every release's |
host report wrapper_sha256 |
felhom-os-apply |
old (no facts, no Docker) | 0.142.0's (it answers the facts mode) | facts present |
| sudoers | — | not readable from the hub; 0.142.0's file equals 0.143.0's (tag diff); capability probe all ok | /hosts/Tester-2-be8404 |
crash guard / kernel.panic |
absent / 0 | armed / 10 | facts |
| operator-signers | absent | not readable from the hub; installer 1.30.0 writes it | inference |
live-restore / Docker |
off / unpinned | on / 29.8.2 | facts |
| golden | first 0.292.0 bake | the re-bake (live-restore + 29.8.2) | facts |
| OS releases installed | TEST-approved | guest os-guest-20261004-123933 (49 pkgs, 16:24) and host os-host-20261004-124133 (106, 16:25), both approved "auto" 1.5 h after first seen = the TEST wait |
os_reports, os_releases |
Part C — Tester 2, what happened
- Act 1 (signed
agent_updateto 0.143.0): queued 18:35 UTC, not delivered — Tester 2 stopped reporting at 18:06 UTC (host 18:05:48, controller 18:06:24). It was also silent 17:13–18:05 UTC, and its controller started at 18:06:20 UTC, so the box restarted or was switched on in between. CC sent Tester 2 nothing before 18:35 UTC; the cause is outside this work (power, network, or the tester). The job expires 19:20 UTC; if Tester 2 returns later, it is refused as expired (harmless) and must be re-signed. - Act 2 (bundle): waits for the operator's one-time bootstrap (R-862).
- Act 3 (
live-restore): not needed — already on. - No Docker step, no reboot, no app change was sent. Read-back of the System page line: still agent 0.142.0, Root files
unknown, guard armed,live-restoreon (last report 18:05 UTC).
Part B — the bundle
Shape, checks, trust rule, bootstrap: 11 §5.4.2; runbook runbooks/config-bundle.md. Red-proofs: 22 of 22 wrapper rules
(partB/redproof.txt), hub 6 of 6 (partB/hub-bundle-redproof.txt), Go executor tests. Live:
| Step | Box | Result |
|---|---|---|
| read the box's files vs the bundle | both | 21 of 22 identical; only felhom-os-apply differed (b1) |
| bootstrap, wrong sha | demo-hp | STOP, nothing changed (b2) |
| bootstrap | both | the new felhom-os-apply, self-check ok (b2, b3) |
| job A, wrong sha | demo-hp | refused by the agent, nothing changed (b4) |
| job B, the 0.143.0 bundle | both | written 0, same 21, kept 1, self-check ok, probe 71/71 (b4) |
| a bundle with one deliberate change | demo-hp | written 1 (that file), self-check ok (b6) |
| its undo (the 0.143.0 bundle again) | demo-hp | written 1, the release file back (sha b71d8698…), probe 71/71 (b6) |
| replay of job B | demo-hp | REJECTED … replay (nonce already seen) (b6) |
| the installer's new path | demo-felhom | written 0, record installer, services active (b5) |
Part D — test approvals
Marked, cancelled at start, amber, operator event, ring-1 bump; red-proofs 6 of 6 (partD/d-redproof.txt). Live at the
18:20 UTC start: os-guest-20261004-123933, os-20261004-091417, os-host-20261004-124133, os-host-20261004-124034
cancelled; one mail (three held by the cooldown); the System page lists them; the Docker release stays (partD/d2). The
2026-10-04 fact and why the risk was small: runbooks/os-updates-test-waits.md.
Part E — the measured heal
| Box | Break | Controller back | traefik back | Container ids |
|---|---|---|---|---|
| 9202 (controller 0.291.0, before the fix) | systemctl restart docker.socket |
never by itself (blind 2+ min, health "healthy") | never | same |
| 9202 (0.293.0) | same | +73 s | +104 s | same (e7) |
| demo-hp 9201 (0.293.0) | same | +88 s | +120 s | 21 of 21 same (e8) |
systemctl restart docker and kill -9 dockerd: socket inode unchanged, nothing to heal (e1, e2). Red-proof 9 of 9 (e6).
Part F — the first-night count
Golden 0.293.0 (documentation/tests/golden-0.293.0-2026-10-04/): no guest release in force → template versions kept;
49 Debian updates pending in the baked guest, which the next REAL approval (tomorrow, after 24 h + a night) will bring.
A ring-1 box from this golden installs 0 on its first night until then. The host: the agent's first leg already runs
right after the first whole-guest backup (Tester 2: 17 min after enrolment, 106 host packages in 44 s) — an installer pass
would cost ~45 s and run before any backup exists; not built.
Part G — drill-r50
What the product delete touched (read first, partG/g1): hub rows (host, guest, reports 185, telemetry 185, the recovery
credential, the DR recipe, the claim, an appliance registration, the WireGuard peer 10.77.0.4), ep0's PBS tenancy
(tenantsync.Deprovision, namespace drill-r50), ep0's WireGuard peer list. No Storage Box (no off-site tier), no mail
(no address). Asked the operator (ep0); yes. Result: deprovision ok (existed=false) — nothing destroyed on ep0; the peer
left ep0 at the next push (5 → 4); /hosts without drill-r50; the delete preview 404s (partG/).
Rows
Before 333, after 334. Closed: R-840 (built), R-859 (opened and closed), R-860 (opened and closed). Opened: R-861 (P2 — the agent's sudoers is root-equivalent; read, not exploited), R-862 (P3 — Tester 2's bootstrap, operator). R-857 not fixed: avoided by baking under a new controller version.
Decisions taken by CC (operator may reverse)
- The one-time backfill marks only the AUTOMATIC test approvals; your Docker button approval of 2026-10-04 stays in force.
- The bundle alarm fires after 7 days behind (
OS_ALARM_BUNDLE_BEHIND_AFTER). - The bundle lives inside
felhom-os-apply(no new sudoers line), not in a new tool. - The controller fixes Part E (it reaches every box by the floor), not the agent (that would need a sudoers change).
- The demo customers' floor is 0.293.0; the global floor stays 0.292.0 so Tester 2's controller does not move this session.
Teardown
- Machines: 9202 left on controller 0.293.0, working (all apps up). demo-hp and demo-felhom: agent 0.143.0, bundle 0.143.0,
controller 0.293.0; the test line removed again on demo-hp. Drill VM off and at
virgin. - Hosts: the bootstrap script, the installer harness and the test script removed from
/rooton both demo hosts. The bundles' previous copies stay in/var/lib/felhom-os-apply/bundle-prev/by design. - Hub: the registry's test version
0.143.0-r840testdeleted (204, then 404); drill-r50 removed; the scratchpad copy of the hub DB deleted at the end.