R-840/R-859/R-860 records: 11 §5.3.1 + §5.4.2, 03, 04, 00; runbooks config-bundle + os-updates-test-waits; register 333→334 (R-840/859/860 closed, R-861/862 opened); STATUS; report; evidence
gates / gates (push) Successful in 31s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-04 20:40:08 +02:00
parent c01d48f457
commit 79f07a7f10
18 changed files with 478 additions and 10 deletions
+125
View File
@@ -0,0 +1,125 @@
# REPORT — the config bundle (R-840), test approvals (R-859), the Docker-socket self-heal (R-860), Tester 2, drill-r50
2026-10-04, evening–night. Brief: "a signed route that brings an installed box's root-owned parts up to date (R-840)…".
Architecture read first: `11-os-updates.md` (§5.3, §5.4, §5.8, §5.9, §8), `03-host-agent.md` (§3, §4, §11),
`04-control-plane-authorization.md` (§3), `07-backup-architecture.md` (§6.4). Evidence: `documentation/audits/r840-config-bundle-2026-10-04/`.
## The Part table
| Part | Result | Why / what changed |
|---|---|---|
| A — what Tester 2 has | **done (read through the hub; no route to the box)** | The brief's timestamp reasoning was wrong (UTC vs local). Tester 2 has installer 1.30.0, agent 0.142.0, the crash guard armed, `live-restore` on, Docker 29.8.2, the re-made golden. It lacks only the R-858 wrapper fix. |
| A — the `felhom-pbs` warning | **done: not a fault** | By design (R-723): a new box's first-hour tier skip is recorded, never mailed; the tier came 7 min later and backed up at 16:31 UTC. No row. |
| B — the bundle route | **done, live on both demo boxes** | Agent 0.143.0 (bundle mode in `felhom-os-apply`, signed `agent_config_update`), hub 0.133.0, installer 1.31.0. Changed: the installer is NOT the self-update wrapper (it cannot do it); a one-time bootstrap is needed on older boxes. |
| C — Tester 2 | **act 1 sent; act 2 waits for the operator; act 3 not needed** | Act 1 (signed agent update to 0.143.0): queued 18:35 UTC — see "Tester 2" below. Act 2 needs the one by-hand bootstrap (R-862, the operator's tunnel). Act 3: `live-restore` is already on. |
| D — test approvals end | **done, live** | Hub 0.133.0 cancelled the 4 test approvals of 2026-10-04 at start (18:20 UTC). Changed: the operator's own Docker button approval was not backfilled (CC decision). |
| E — self-heal after any Docker restart | **done, live (9202 + demo-hp)** | Controller 0.293.0. Changed: only a `docker.socket` restart breaks it; a dockerd crash or `systemctl restart docker` does not. |
| F — new installs start with the approved fixes | **built; tonight's golden carries none** | `build-golden.sh` 3.2.0 `GOLDEN_GUEST_PKGS`. No guest release is in force tonight (the test one was cancelled), so the golden keeps the template (49 Debian updates pending for the next real approval). |
| G — drill-r50 | **done** | Deleted through the product (operator's yes: it touched ep0). On ep0: its PBS namespace did not exist; its WireGuard peer left (5 → 4 peers pushed). |
| H — release, golden, records | **done** | Agent 0.143.0, controller 0.293.0, hub 0.133.0, installer 1.31.0, golden 0.293.0 baked + vouched. Floor 0.293.0 for the two demo customers only (Tester 2 untouched). |
## Claims in the brief that turned out wrong
1. **The timestamp reasoning about Tester 2 (§1).** Tester 2 was bound at **16:06 UTC = 18:06 local**; the releases were
compared in local time. At 18:06 local, installer 1.30.0 (17:36), agent 0.142.0 (vouched ~17:49) and the re-made golden
(17:48) were all current. Measured (hub records): agent **0.142.0**, crash guard **armed** (`kernel.panic` 10), `live-restore`
**on**, Docker **29.8.2**; the wrapper reports a facts mode (so it is not "old"). Only agent 0.142.1's R-858 wrapper fix is
missing — and the root files of 0.142.0 equal 0.143.0's in every file but `felhom-os-apply`.
2. **"The felhom-pbs warning is a fault."** It is the designed first-hour behaviour (R-723): recorded, not mailed.
3. **"The self-update wrapper can install a bundle safely."** It cannot: it is a fixed sh script that swaps one binary. The
bundle is installed by `felhom-os-apply` (already reachable through the agent's sudoers line). And **no** existing root
helper can install the first bundle on an older box — one by-hand bootstrap per older box (done on both demo boxes).
4. **"A restart of an existing container picks up the new socket."** True — measured twice (controller exit → restart policy;
`docker restart traefik`): both came back on the current inode, same container ids.
5. **"dockerd restarts on its own … the box shows DOWN until a person acts."** Only when the socket FILE is re-created
(`docker.socket` restart, i.e. a docker-ce upgrade or by hand). A dockerd crash or `systemctl restart docker` keeps it (measured).
6. **Act 3 for Tester 2 (`live-restore` reload)** is not needed: it is already on (from the golden).
## Part A — Tester 2, read back (hub records; CC has no route to the box)
| Item | Brief expected | Measured | Source |
|---|---|---|---|
| bound / enrolled | 16:06 (read as local) | 16:06:27 UTC bind, 16:07:04 UTC host | `appliance_registrations`, `hosts` |
| installer | 1.29.0 | **1.30.0** (the crash guard is installed — only 1.30.0 does that) | host report `system.facts.host.crash_guard` |
| agent | 0.141.1 | **0.142.0** | `/hosts` |
| PBS wrapper sha | — | `104db0a4…` = every release's | host report `wrapper_sha256` |
| `felhom-os-apply` | old (no facts, no Docker) | 0.142.0's (it answers the facts mode) | facts present |
| sudoers | — | not readable from the hub; 0.142.0's file equals 0.143.0's (tag diff); capability probe all ok | `/hosts/Tester-2-be8404` |
| crash guard / `kernel.panic` | absent / 0 | **armed / 10** | facts |
| operator-signers | absent | not readable from the hub; installer 1.30.0 writes it | inference |
| `live-restore` / Docker | off / unpinned | **on / 29.8.2** | facts |
| golden | first 0.292.0 bake | the re-bake (live-restore + 29.8.2) | facts |
| OS releases installed | TEST-approved | guest `os-guest-20261004-123933` (49 pkgs, 16:24) and host `os-host-20261004-124133` (106, 16:25), both approved "auto" 1.5 h after first seen = the TEST wait | `os_reports`, `os_releases` |
## Part B — the bundle
Shape, checks, trust rule, bootstrap: `11` §5.4.2; runbook `runbooks/config-bundle.md`. Red-proofs: 22 of 22 wrapper rules
(`partB/redproof.txt`), hub 6 of 6 (`partB/hub-bundle-redproof.txt`), Go executor tests. Live:
| Step | Box | Result |
|---|---|---|
| read the box's files vs the bundle | both | 21 of 22 identical; only `felhom-os-apply` differed (`b1`) |
| bootstrap, wrong sha | demo-hp | `STOP`, nothing changed (`b2`) |
| bootstrap | both | the new `felhom-os-apply`, self-check ok (`b2`, `b3`) |
| job A, wrong sha | demo-hp | refused by the agent, nothing changed (`b4`) |
| job B, the 0.143.0 bundle | both | written 0, same 21, kept 1, self-check ok, probe 71/71 (`b4`) |
| a bundle with one deliberate change | demo-hp | written 1 (that file), self-check ok (`b6`) |
| its undo (the 0.143.0 bundle again) | demo-hp | written 1, the release file back (sha `b71d8698…`), probe 71/71 (`b6`) |
| replay of job B | demo-hp | `REJECTED … replay (nonce already seen)` (`b6`) |
| the installer's new path | demo-felhom | written 0, record `installer`, services active (`b5`) |
## Part D — test approvals
Marked, cancelled at start, amber, operator event, ring-1 bump; red-proofs 6 of 6 (`partD/d-redproof.txt`). Live at the
18:20 UTC start: `os-guest-20261004-123933`, `os-20261004-091417`, `os-host-20261004-124133`, `os-host-20261004-124034`
cancelled; one mail (three held by the cooldown); the System page lists them; the Docker release stays (`partD/d2`). The
2026-10-04 fact and why the risk was small: `runbooks/os-updates-test-waits.md`.
## Part E — the measured heal
| Box | Break | Controller back | traefik back | Container ids |
|---|---|---|---|---|
| 9202 (controller 0.291.0, before the fix) | `systemctl restart docker.socket` | never by itself (blind 2+ min, health "healthy") | never | same |
| 9202 (0.293.0) | same | +73 s | +104 s | same (`e7`) |
| demo-hp 9201 (0.293.0) | same | +88 s | +120 s | 21 of 21 same (`e8`) |
`systemctl restart docker` and `kill -9 dockerd`: socket inode unchanged, nothing to heal (`e1`, `e2`). Red-proof 9 of 9 (`e6`).
## Part F — the first-night count
Golden 0.293.0 (`documentation/tests/golden-0.293.0-2026-10-04/`): no guest release in force → template versions kept;
**49** Debian updates pending in the baked guest, which the next REAL approval (tomorrow, after 24 h + a night) will bring.
A ring-1 box from this golden installs **0** on its first night until then. The host: the agent's first leg already runs
right after the first whole-guest backup (Tester 2: 17 min after enrolment, 106 host packages in 44 s) — an installer pass
would cost ~45 s and run before any backup exists; not built.
## Part G — drill-r50
What the product delete touched (read first, `partG/g1`): hub rows (host, guest, reports 185, telemetry 185, the recovery
credential, the DR recipe, the claim, an appliance registration, the WireGuard peer 10.77.0.4), ep0's PBS tenancy
(`tenantsync.Deprovision`, namespace `drill-r50`), ep0's WireGuard peer list. No Storage Box (no off-site tier), no mail
(no address). Asked the operator (ep0); yes. Result: `deprovision ok (existed=false)` — nothing destroyed on ep0; the peer
left ep0 at the next push (5 → 4); `/hosts` without drill-r50; the delete preview 404s (`partG/`).
## Rows
Before **333**, after **334**. Closed: **R-840** (built), **R-859** (opened and closed), **R-860** (opened and closed).
Opened: **R-861** (P2 — the agent's sudoers is root-equivalent; read, not exploited), **R-862** (P3 — Tester 2's bootstrap,
operator). R-857 not fixed: avoided by baking under a new controller version.
## Decisions taken by CC (operator may reverse)
1. The one-time backfill marks only the AUTOMATIC test approvals; your Docker button approval of 2026-10-04 stays in force.
2. The bundle alarm fires after 7 days behind (`OS_ALARM_BUNDLE_BEHIND_AFTER`).
3. The bundle lives inside `felhom-os-apply` (no new sudoers line), not in a new tool.
4. The controller fixes Part E (it reaches every box by the floor), not the agent (that would need a sudoers change).
5. The demo customers' floor is 0.293.0; the global floor stays 0.292.0 so Tester 2's controller does not move this session.
## Teardown
- Machines: 9202 left on controller 0.293.0, working (all apps up). demo-hp and demo-felhom: agent 0.143.0, bundle 0.143.0,
controller 0.293.0; the test line removed again on demo-hp. Drill VM off and at `virgin`.
- Hosts: the bootstrap script, the installer harness and the test script removed from `/root` on both demo hosts. The
bundles' previous copies stay in `/var/lib/felhom-os-apply/bundle-prev/` by design.
- Hub: the registry's test version `0.143.0-r840test` deleted (204, then 404); drill-r50 removed; the scratchpad copy of
the hub DB deleted at the end.