Files
felhom.eu/REPORT-r840-config-bundle-2026-10-04.md
T
2026-10-04 20:52:18 +02:00

138 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — the config bundle (R-840), test approvals (R-859), the Docker-socket self-heal (R-860), Tester 2, drill-r50
2026-10-04, evening–night. Brief: "a signed route that brings an installed box's root-owned parts up to date (R-840)…".
Architecture read first: `11-os-updates.md` (§5.3, §5.4, §5.8, §5.9, §8), `03-host-agent.md` (§3, §4, §11),
`04-control-plane-authorization.md` (§3), `07-backup-architecture.md` (§6.4). Evidence: `documentation/audits/r840-config-bundle-2026-10-04/`.
## The Part table
| Part | Result | Why / what changed |
|---|---|---|
| A — what Tester 2 has | **done (read through the hub; no route to the box)** | The brief's timestamp reasoning was wrong (UTC vs local). Tester 2 has installer 1.30.0, agent 0.142.0, the crash guard armed, `live-restore` on, Docker 29.8.2, the re-made golden. It lacks only the R-858 wrapper fix. |
| A — the `felhom-pbs` warning | **done: not a fault** | By design (R-723): a new box's first-hour tier skip is recorded, never mailed; the tier came 7 min later and backed up at 16:31 UTC. No row. |
| B — the bundle route | **done, live on both demo boxes** | Agent 0.143.0 (bundle mode in `felhom-os-apply`, signed `agent_config_update`), hub 0.133.0, installer 1.31.0. Changed: the installer is NOT the self-update wrapper (it cannot do it); a one-time bootstrap is needed on older boxes. |
| C — Tester 2 | **act 1 sent, NOT delivered (box offline since 18:06 UTC); act 2 waits for the operator; act 3 not needed** | Act 1: queued 18:35 UTC — see "Part C" below. Act 2 needs the one by-hand bootstrap (R-862, the operator's tunnel). Act 3: `live-restore` is already on. |
| D — test approvals end | **done, live** | Hub 0.133.0 cancelled the 4 test approvals of 2026-10-04 at start (18:20 UTC). Changed: the operator's own Docker button approval was not backfilled (CC decision). |
| E — self-heal after any Docker restart | **done, live (9202 + demo-hp)** | Controller 0.293.0. Changed: only a `docker.socket` restart breaks it; a dockerd crash or `systemctl restart docker` does not. |
| F — new installs start with the approved fixes | **built; tonight's golden carries none** | `build-golden.sh` 3.2.0 `GOLDEN_GUEST_PKGS`. No guest release is in force tonight (the test one was cancelled), so the golden keeps the template (49 Debian updates pending for the next real approval). |
| G — drill-r50 | **done** | Deleted through the product (operator's yes: it touched ep0). On ep0: its PBS namespace did not exist; its WireGuard peer left (5 → 4 peers pushed). |
| H — release, golden, records | **done** | Agent 0.143.0, controller 0.293.0, hub 0.133.0, installer 1.31.0, golden 0.293.0 baked + vouched. Floor 0.293.0 for the two demo customers only (Tester 2 untouched). |
## Claims in the brief that turned out wrong
1. **The timestamp reasoning about Tester 2 (§1).** Tester 2 was bound at **16:06 UTC = 18:06 local**; the releases were
compared in local time. At 18:06 local, installer 1.30.0 (17:36), agent 0.142.0 (vouched ~17:49) and the re-made golden
(17:48) were all current. Measured (hub records): agent **0.142.0**, crash guard **armed** (`kernel.panic` 10), `live-restore`
**on**, Docker **29.8.2**; the wrapper reports a facts mode (so it is not "old"). Only agent 0.142.1's R-858 wrapper fix is
missing — and the root files of 0.142.0 equal 0.143.0's in every file but `felhom-os-apply`.
2. **"The felhom-pbs warning is a fault."** It is the designed first-hour behaviour (R-723): recorded, not mailed.
3. **"The self-update wrapper can install a bundle safely."** It cannot: it is a fixed sh script that swaps one binary. The
bundle is installed by `felhom-os-apply` (already reachable through the agent's sudoers line). And **no** existing root
helper can install the first bundle on an older box — one by-hand bootstrap per older box (done on both demo boxes).
4. **"A restart of an existing container picks up the new socket."** True — measured twice (controller exit → restart policy;
`docker restart traefik`): both came back on the current inode, same container ids.
5. **"dockerd restarts on its own … the box shows DOWN until a person acts."** Only when the socket FILE is re-created
(`docker.socket` restart, i.e. a docker-ce upgrade or by hand). A dockerd crash or `systemctl restart docker` keeps it (measured).
6. **Act 3 for Tester 2 (`live-restore` reload)** is not needed: it is already on (from the golden).
## Part A — Tester 2, read back (hub records; CC has no route to the box)
| Item | Brief expected | Measured | Source |
|---|---|---|---|
| bound / enrolled | 16:06 (read as local) | 16:06:27 UTC bind, 16:07:04 UTC host | `appliance_registrations`, `hosts` |
| installer | 1.29.0 | **1.30.0** (the crash guard is installed — only 1.30.0 does that) | host report `system.facts.host.crash_guard` |
| agent | 0.141.1 | **0.142.0** | `/hosts` |
| PBS wrapper sha | — | `104db0a4…` = every release's | host report `wrapper_sha256` |
| `felhom-os-apply` | old (no facts, no Docker) | 0.142.0's (it answers the facts mode) | facts present |
| sudoers | — | not readable from the hub; 0.142.0's file equals 0.143.0's (tag diff); capability probe all ok | `/hosts/Tester-2-be8404` |
| crash guard / `kernel.panic` | absent / 0 | **armed / 10** | facts |
| operator-signers | absent | not readable from the hub; installer 1.30.0 writes it | inference |
| `live-restore` / Docker | off / unpinned | **on / 29.8.2** | facts |
| golden | first 0.292.0 bake | the re-bake (live-restore + 29.8.2) | facts |
| OS releases installed | TEST-approved | guest `os-guest-20261004-123933` (49 pkgs, 16:24) and host `os-host-20261004-124133` (106, 16:25), both approved "auto" 1.5 h after first seen = the TEST wait | `os_reports`, `os_releases` |
## Part C — Tester 2, what happened
- **Act 1** (signed `agent_update` to 0.143.0): queued 18:35 UTC, **not delivered** — Tester 2 stopped reporting at
18:06 UTC (host 18:05:48, controller 18:06:24). It was also silent 17:13–18:05 UTC, and its controller started at
18:06:20 UTC, so the box restarted or was switched on in between. CC sent Tester 2 nothing before 18:35 UTC; the cause
is outside this work (power, network, or the tester). The job expires 19:20 UTC; if Tester 2 returns later, it is
refused as expired (harmless) and must be re-signed.
- **Act 2** (bundle): waits for the operator's one-time bootstrap (R-862).
- **Act 3** (`live-restore`): not needed — already on.
- No Docker step, no reboot, no app change was sent. Read-back of the System page line: still agent 0.142.0, Root files
`unknown`, guard armed, `live-restore` on (last report 18:05 UTC).
## Part B — the bundle
Shape, checks, trust rule, bootstrap: `11` §5.4.2; runbook `runbooks/config-bundle.md`. Red-proofs: 22 of 22 wrapper rules
(`partB/redproof.txt`), hub 6 of 6 (`partB/hub-bundle-redproof.txt`), Go executor tests. Live:
| Step | Box | Result |
|---|---|---|
| read the box's files vs the bundle | both | 21 of 22 identical; only `felhom-os-apply` differed (`b1`) |
| bootstrap, wrong sha | demo-hp | `STOP`, nothing changed (`b2`) |
| bootstrap | both | the new `felhom-os-apply`, self-check ok (`b2`, `b3`) |
| job A, wrong sha | demo-hp | refused by the agent, nothing changed (`b4`) |
| job B, the 0.143.0 bundle | both | written 0, same 21, kept 1, self-check ok, probe 71/71 (`b4`) |
| a bundle with one deliberate change | demo-hp | written 1 (that file), self-check ok (`b6`) |
| its undo (the 0.143.0 bundle again) | demo-hp | written 1, the release file back (sha `b71d8698…`), probe 71/71 (`b6`) |
| replay of job B | demo-hp | `REJECTED … replay (nonce already seen)` (`b6`) |
| the installer's new path | demo-felhom | written 0, record `installer`, services active (`b5`) |
## Part D — test approvals
Marked, cancelled at start, amber, operator event, ring-1 bump; red-proofs 6 of 6 (`partD/d-redproof.txt`). Live at the
18:20 UTC start: `os-guest-20261004-123933`, `os-20261004-091417`, `os-host-20261004-124133`, `os-host-20261004-124034`
cancelled; one mail (three held by the cooldown); the System page lists them; the Docker release stays (`partD/d2`). The
2026-10-04 fact and why the risk was small: `runbooks/os-updates-test-waits.md`.
## Part E — the measured heal
| Box | Break | Controller back | traefik back | Container ids |
|---|---|---|---|---|
| 9202 (controller 0.291.0, before the fix) | `systemctl restart docker.socket` | never by itself (blind 2+ min, health "healthy") | never | same |
| 9202 (0.293.0) | same | +73 s | +104 s | same (`e7`) |
| demo-hp 9201 (0.293.0) | same | +88 s | +120 s | 21 of 21 same (`e8`) |
`systemctl restart docker` and `kill -9 dockerd`: socket inode unchanged, nothing to heal (`e1`, `e2`). Red-proof 9 of 9 (`e6`).
## Part F — the first-night count
Golden 0.293.0 (`documentation/tests/golden-0.293.0-2026-10-04/`): no guest release in force → template versions kept;
**49** Debian updates pending in the baked guest, which the next REAL approval (tomorrow, after 24 h + a night) will bring.
A ring-1 box from this golden installs **0** on its first night until then. The host: the agent's first leg already runs
right after the first whole-guest backup (Tester 2: 17 min after enrolment, 106 host packages in 44 s) — an installer pass
would cost ~45 s and run before any backup exists; not built.
## Part G — drill-r50
What the product delete touched (read first, `partG/g1`): hub rows (host, guest, reports 185, telemetry 185, the recovery
credential, the DR recipe, the claim, an appliance registration, the WireGuard peer 10.77.0.4), ep0's PBS tenancy
(`tenantsync.Deprovision`, namespace `drill-r50`), ep0's WireGuard peer list. No Storage Box (no off-site tier), no mail
(no address). Asked the operator (ep0); yes. Result: `deprovision ok (existed=false)` — nothing destroyed on ep0; the peer
left ep0 at the next push (5 → 4); `/hosts` without drill-r50; the delete preview 404s (`partG/`).
## Rows
Before **333**, after **334**. Closed: **R-840** (built), **R-859** (opened and closed), **R-860** (opened and closed).
Opened: **R-861** (P2 — the agent's sudoers is root-equivalent; read, not exploited), **R-862** (P3 — Tester 2's bootstrap,
operator). R-857 not fixed: avoided by baking under a new controller version.
## Decisions taken by CC (operator may reverse)
1. The one-time backfill marks only the AUTOMATIC test approvals; your Docker button approval of 2026-10-04 stays in force.
2. The bundle alarm fires after 7 days behind (`OS_ALARM_BUNDLE_BEHIND_AFTER`).
3. The bundle lives inside `felhom-os-apply` (no new sudoers line), not in a new tool.
4. The controller fixes Part E (it reaches every box by the floor), not the agent (that would need a sudoers change).
5. The demo customers' floor is 0.293.0; the global floor stays 0.292.0 so Tester 2's controller does not move this session.
## Teardown
- Machines: 9202 left on controller 0.293.0, working (all apps up). demo-hp and demo-felhom: agent 0.143.0, bundle 0.143.0,
controller 0.293.0; the test line removed again on demo-hp. Drill VM off and at `virgin`.
- Hosts: the bootstrap script, the installer harness and the test script removed from `/root` on both demo hosts. The
bundles' previous copies stay in `/var/lib/felhom-os-apply/bundle-prev/` by design.
- Hub: the registry's test version `0.143.0-r840test` deleted (204, then 404); drill-r50 removed; the scratchpad copy of
the hub DB deleted at the end.