# REPORT — the config bundle (R-840), test approvals (R-859), the Docker-socket self-heal (R-860), Tester 2, drill-r50 2026-10-04, evening–night. Brief: "a signed route that brings an installed box's root-owned parts up to date (R-840)…". Architecture read first: `11-os-updates.md` (§5.3, §5.4, §5.8, §5.9, §8), `03-host-agent.md` (§3, §4, §11), `04-control-plane-authorization.md` (§3), `07-backup-architecture.md` (§6.4). Evidence: `documentation/audits/r840-config-bundle-2026-10-04/`. ## The Part table | Part | Result | Why / what changed | |---|---|---| | A — what Tester 2 has | **done (read through the hub; no route to the box)** | The brief's timestamp reasoning was wrong (UTC vs local). Tester 2 has installer 1.30.0, agent 0.142.0, the crash guard armed, `live-restore` on, Docker 29.8.2, the re-made golden. It lacks only the R-858 wrapper fix. | | A — the `felhom-pbs` warning | **done: not a fault** | By design (R-723): a new box's first-hour tier skip is recorded, never mailed; the tier came 7 min later and backed up at 16:31 UTC. No row. | | B — the bundle route | **done, live on both demo boxes** | Agent 0.143.0 (bundle mode in `felhom-os-apply`, signed `agent_config_update`), hub 0.133.0, installer 1.31.0. Changed: the installer is NOT the self-update wrapper (it cannot do it); a one-time bootstrap is needed on older boxes. | | C — Tester 2 | **act 1 sent, NOT delivered (box offline since 18:06 UTC); act 2 waits for the operator; act 3 not needed** | Act 1: queued 18:35 UTC — see "Part C" below. Act 2 needs the one by-hand bootstrap (R-862, the operator's tunnel). Act 3: `live-restore` is already on. | | D — test approvals end | **done, live** | Hub 0.133.0 cancelled the 4 test approvals of 2026-10-04 at start (18:20 UTC). Changed: the operator's own Docker button approval was not backfilled (CC decision). | | E — self-heal after any Docker restart | **done, live (9202 + demo-hp)** | Controller 0.293.0. Changed: only a `docker.socket` restart breaks it; a dockerd crash or `systemctl restart docker` does not. | | F — new installs start with the approved fixes | **built; tonight's golden carries none** | `build-golden.sh` 3.2.0 `GOLDEN_GUEST_PKGS`. No guest release is in force tonight (the test one was cancelled), so the golden keeps the template (49 Debian updates pending for the next real approval). | | G — drill-r50 | **done** | Deleted through the product (operator's yes: it touched ep0). On ep0: its PBS namespace did not exist; its WireGuard peer left (5 → 4 peers pushed). | | H — release, golden, records | **done** | Agent 0.143.0, controller 0.293.0, hub 0.133.0, installer 1.31.0, golden 0.293.0 baked + vouched. Floor 0.293.0 for the two demo customers only (Tester 2 untouched). | ## Claims in the brief that turned out wrong 1. **The timestamp reasoning about Tester 2 (§1).** Tester 2 was bound at **16:06 UTC = 18:06 local**; the releases were compared in local time. At 18:06 local, installer 1.30.0 (17:36), agent 0.142.0 (vouched ~17:49) and the re-made golden (17:48) were all current. Measured (hub records): agent **0.142.0**, crash guard **armed** (`kernel.panic` 10), `live-restore` **on**, Docker **29.8.2**; the wrapper reports a facts mode (so it is not "old"). Only agent 0.142.1's R-858 wrapper fix is missing — and the root files of 0.142.0 equal 0.143.0's in every file but `felhom-os-apply`. 2. **"The felhom-pbs warning is a fault."** It is the designed first-hour behaviour (R-723): recorded, not mailed. 3. **"The self-update wrapper can install a bundle safely."** It cannot: it is a fixed sh script that swaps one binary. The bundle is installed by `felhom-os-apply` (already reachable through the agent's sudoers line). And **no** existing root helper can install the first bundle on an older box — one by-hand bootstrap per older box (done on both demo boxes). 4. **"A restart of an existing container picks up the new socket."** True — measured twice (controller exit → restart policy; `docker restart traefik`): both came back on the current inode, same container ids. 5. **"dockerd restarts on its own … the box shows DOWN until a person acts."** Only when the socket FILE is re-created (`docker.socket` restart, i.e. a docker-ce upgrade or by hand). A dockerd crash or `systemctl restart docker` keeps it (measured). 6. **Act 3 for Tester 2 (`live-restore` reload)** is not needed: it is already on (from the golden). ## Part A — Tester 2, read back (hub records; CC has no route to the box) | Item | Brief expected | Measured | Source | |---|---|---|---| | bound / enrolled | 16:06 (read as local) | 16:06:27 UTC bind, 16:07:04 UTC host | `appliance_registrations`, `hosts` | | installer | 1.29.0 | **1.30.0** (the crash guard is installed — only 1.30.0 does that) | host report `system.facts.host.crash_guard` | | agent | 0.141.1 | **0.142.0** | `/hosts` | | PBS wrapper sha | — | `104db0a4…` = every release's | host report `wrapper_sha256` | | `felhom-os-apply` | old (no facts, no Docker) | 0.142.0's (it answers the facts mode) | facts present | | sudoers | — | not readable from the hub; 0.142.0's file equals 0.143.0's (tag diff); capability probe all ok | `/hosts/Tester-2-be8404` | | crash guard / `kernel.panic` | absent / 0 | **armed / 10** | facts | | operator-signers | absent | not readable from the hub; installer 1.30.0 writes it | inference | | `live-restore` / Docker | off / unpinned | **on / 29.8.2** | facts | | golden | first 0.292.0 bake | the re-bake (live-restore + 29.8.2) | facts | | OS releases installed | TEST-approved | guest `os-guest-20261004-123933` (49 pkgs, 16:24) and host `os-host-20261004-124133` (106, 16:25), both approved "auto" 1.5 h after first seen = the TEST wait | `os_reports`, `os_releases` | ## Part C — Tester 2, what happened - **Act 1** (signed `agent_update` to 0.143.0): queued 18:35 UTC, **not delivered** — Tester 2 stopped reporting at 18:06 UTC (host 18:05:48, controller 18:06:24). It was also silent 17:13–18:05 UTC, and its controller started at 18:06:20 UTC, so the box restarted or was switched on in between. CC sent Tester 2 nothing before 18:35 UTC; the cause is outside this work (power, network, or the tester). The job expires 19:20 UTC; if Tester 2 returns later, it is refused as expired (harmless) and must be re-signed. - **Act 2** (bundle): waits for the operator's one-time bootstrap (R-862). - **Act 3** (`live-restore`): not needed — already on. - No Docker step, no reboot, no app change was sent. Read-back of the System page line: still agent 0.142.0, Root files `unknown`, guard armed, `live-restore` on (last report 18:05 UTC). ## Part B — the bundle Shape, checks, trust rule, bootstrap: `11` §5.4.2; runbook `runbooks/config-bundle.md`. Red-proofs: 22 of 22 wrapper rules (`partB/redproof.txt`), hub 6 of 6 (`partB/hub-bundle-redproof.txt`), Go executor tests. Live: | Step | Box | Result | |---|---|---| | read the box's files vs the bundle | both | 21 of 22 identical; only `felhom-os-apply` differed (`b1`) | | bootstrap, wrong sha | demo-hp | `STOP`, nothing changed (`b2`) | | bootstrap | both | the new `felhom-os-apply`, self-check ok (`b2`, `b3`) | | job A, wrong sha | demo-hp | refused by the agent, nothing changed (`b4`) | | job B, the 0.143.0 bundle | both | written 0, same 21, kept 1, self-check ok, probe 71/71 (`b4`) | | a bundle with one deliberate change | demo-hp | written 1 (that file), self-check ok (`b6`) | | its undo (the 0.143.0 bundle again) | demo-hp | written 1, the release file back (sha `b71d8698…`), probe 71/71 (`b6`) | | replay of job B | demo-hp | `REJECTED … replay (nonce already seen)` (`b6`) | | the installer's new path | demo-felhom | written 0, record `installer`, services active (`b5`) | ## Part D — test approvals Marked, cancelled at start, amber, operator event, ring-1 bump; red-proofs 6 of 6 (`partD/d-redproof.txt`). Live at the 18:20 UTC start: `os-guest-20261004-123933`, `os-20261004-091417`, `os-host-20261004-124133`, `os-host-20261004-124034` cancelled; one mail (three held by the cooldown); the System page lists them; the Docker release stays (`partD/d2`). The 2026-10-04 fact and why the risk was small: `runbooks/os-updates-test-waits.md`. ## Part E — the measured heal | Box | Break | Controller back | traefik back | Container ids | |---|---|---|---|---| | 9202 (controller 0.291.0, before the fix) | `systemctl restart docker.socket` | never by itself (blind 2+ min, health "healthy") | never | same | | 9202 (0.293.0) | same | +73 s | +104 s | same (`e7`) | | demo-hp 9201 (0.293.0) | same | +88 s | +120 s | 21 of 21 same (`e8`) | `systemctl restart docker` and `kill -9 dockerd`: socket inode unchanged, nothing to heal (`e1`, `e2`). Red-proof 9 of 9 (`e6`). ## Part F — the first-night count Golden 0.293.0 (`documentation/tests/golden-0.293.0-2026-10-04/`): no guest release in force → template versions kept; **49** Debian updates pending in the baked guest, which the next REAL approval (tomorrow, after 24 h + a night) will bring. A ring-1 box from this golden installs **0** on its first night until then. The host: the agent's first leg already runs right after the first whole-guest backup (Tester 2: 17 min after enrolment, 106 host packages in 44 s) — an installer pass would cost ~45 s and run before any backup exists; not built. ## Part G — drill-r50 What the product delete touched (read first, `partG/g1`): hub rows (host, guest, reports 185, telemetry 185, the recovery credential, the DR recipe, the claim, an appliance registration, the WireGuard peer 10.77.0.4), ep0's PBS tenancy (`tenantsync.Deprovision`, namespace `drill-r50`), ep0's WireGuard peer list. No Storage Box (no off-site tier), no mail (no address). Asked the operator (ep0); yes. Result: `deprovision ok (existed=false)` — nothing destroyed on ep0; the peer left ep0 at the next push (5 → 4); `/hosts` without drill-r50; the delete preview 404s (`partG/`). ## Rows Before **333**, after **334**. Closed: **R-840** (built), **R-859** (opened and closed), **R-860** (opened and closed). Opened: **R-861** (P2 — the agent's sudoers is root-equivalent; read, not exploited), **R-862** (P3 — Tester 2's bootstrap, operator). R-857 not fixed: avoided by baking under a new controller version. ## Decisions taken by CC (operator may reverse) 1. The one-time backfill marks only the AUTOMATIC test approvals; your Docker button approval of 2026-10-04 stays in force. 2. The bundle alarm fires after 7 days behind (`OS_ALARM_BUNDLE_BEHIND_AFTER`). 3. The bundle lives inside `felhom-os-apply` (no new sudoers line), not in a new tool. 4. The controller fixes Part E (it reaches every box by the floor), not the agent (that would need a sudoers change). 5. The demo customers' floor is 0.293.0; the global floor stays 0.292.0 so Tester 2's controller does not move this session. ## Teardown - Machines: 9202 left on controller 0.293.0, working (all apps up). demo-hp and demo-felhom: agent 0.143.0, bundle 0.143.0, controller 0.293.0; the test line removed again on demo-hp. Drill VM off and at `virgin`. - Hosts: the bootstrap script, the installer harness and the test script removed from `/root` on both demo hosts. The bundles' previous copies stay in `/var/lib/felhom-os-apply/bundle-prev/` by design. - Hub: the registry's test version `0.143.0-r840test` deleted (204, then 404); drill-r50 removed; the scratchpad copy of the hub DB deleted at the end.