Files
felhom.eu/REPORT-r840-config-bundle-2026-10-04.md
T
2026-10-04 20:52:18 +02:00

11 KiB
Raw Blame History

REPORT — the config bundle (R-840), test approvals (R-859), the Docker-socket self-heal (R-860), Tester 2, drill-r50

2026-10-04, evening–night. Brief: "a signed route that brings an installed box's root-owned parts up to date (R-840)…". Architecture read first: 11-os-updates.md (§5.3, §5.4, §5.8, §5.9, §8), 03-host-agent.md (§3, §4, §11), 04-control-plane-authorization.md (§3), 07-backup-architecture.md (§6.4). Evidence: documentation/audits/r840-config-bundle-2026-10-04/.

The Part table

Part Result Why / what changed
A — what Tester 2 has done (read through the hub; no route to the box) The brief's timestamp reasoning was wrong (UTC vs local). Tester 2 has installer 1.30.0, agent 0.142.0, the crash guard armed, live-restore on, Docker 29.8.2, the re-made golden. It lacks only the R-858 wrapper fix.
A — the felhom-pbs warning done: not a fault By design (R-723): a new box's first-hour tier skip is recorded, never mailed; the tier came 7 min later and backed up at 16:31 UTC. No row.
B — the bundle route done, live on both demo boxes Agent 0.143.0 (bundle mode in felhom-os-apply, signed agent_config_update), hub 0.133.0, installer 1.31.0. Changed: the installer is NOT the self-update wrapper (it cannot do it); a one-time bootstrap is needed on older boxes.
C — Tester 2 act 1 sent, NOT delivered (box offline since 18:06 UTC); act 2 waits for the operator; act 3 not needed Act 1: queued 18:35 UTC — see "Part C" below. Act 2 needs the one by-hand bootstrap (R-862, the operator's tunnel). Act 3: live-restore is already on.
D — test approvals end done, live Hub 0.133.0 cancelled the 4 test approvals of 2026-10-04 at start (18:20 UTC). Changed: the operator's own Docker button approval was not backfilled (CC decision).
E — self-heal after any Docker restart done, live (9202 + demo-hp) Controller 0.293.0. Changed: only a docker.socket restart breaks it; a dockerd crash or systemctl restart docker does not.
F — new installs start with the approved fixes built; tonight's golden carries none build-golden.sh 3.2.0 GOLDEN_GUEST_PKGS. No guest release is in force tonight (the test one was cancelled), so the golden keeps the template (49 Debian updates pending for the next real approval).
G — drill-r50 done Deleted through the product (operator's yes: it touched ep0). On ep0: its PBS namespace did not exist; its WireGuard peer left (5 → 4 peers pushed).
H — release, golden, records done Agent 0.143.0, controller 0.293.0, hub 0.133.0, installer 1.31.0, golden 0.293.0 baked + vouched. Floor 0.293.0 for the two demo customers only (Tester 2 untouched).

Claims in the brief that turned out wrong

  1. The timestamp reasoning about Tester 2 (§1). Tester 2 was bound at 16:06 UTC = 18:06 local; the releases were compared in local time. At 18:06 local, installer 1.30.0 (17:36), agent 0.142.0 (vouched ~17:49) and the re-made golden (17:48) were all current. Measured (hub records): agent 0.142.0, crash guard armed (kernel.panic 10), live-restore on, Docker 29.8.2; the wrapper reports a facts mode (so it is not "old"). Only agent 0.142.1's R-858 wrapper fix is missing — and the root files of 0.142.0 equal 0.143.0's in every file but felhom-os-apply.
  2. "The felhom-pbs warning is a fault." It is the designed first-hour behaviour (R-723): recorded, not mailed.
  3. "The self-update wrapper can install a bundle safely." It cannot: it is a fixed sh script that swaps one binary. The bundle is installed by felhom-os-apply (already reachable through the agent's sudoers line). And no existing root helper can install the first bundle on an older box — one by-hand bootstrap per older box (done on both demo boxes).
  4. "A restart of an existing container picks up the new socket." True — measured twice (controller exit → restart policy; docker restart traefik): both came back on the current inode, same container ids.
  5. "dockerd restarts on its own … the box shows DOWN until a person acts." Only when the socket FILE is re-created (docker.socket restart, i.e. a docker-ce upgrade or by hand). A dockerd crash or systemctl restart docker keeps it (measured).
  6. Act 3 for Tester 2 (live-restore reload) is not needed: it is already on (from the golden).

Part A — Tester 2, read back (hub records; CC has no route to the box)

Item Brief expected Measured Source
bound / enrolled 16:06 (read as local) 16:06:27 UTC bind, 16:07:04 UTC host appliance_registrations, hosts
installer 1.29.0 1.30.0 (the crash guard is installed — only 1.30.0 does that) host report system.facts.host.crash_guard
agent 0.141.1 0.142.0 /hosts
PBS wrapper sha — 104db0a4… = every release's host report wrapper_sha256
felhom-os-apply old (no facts, no Docker) 0.142.0's (it answers the facts mode) facts present
sudoers — not readable from the hub; 0.142.0's file equals 0.143.0's (tag diff); capability probe all ok /hosts/Tester-2-be8404
crash guard / kernel.panic absent / 0 armed / 10 facts
operator-signers absent not readable from the hub; installer 1.30.0 writes it inference
live-restore / Docker off / unpinned on / 29.8.2 facts
golden first 0.292.0 bake the re-bake (live-restore + 29.8.2) facts
OS releases installed TEST-approved guest os-guest-20261004-123933 (49 pkgs, 16:24) and host os-host-20261004-124133 (106, 16:25), both approved "auto" 1.5 h after first seen = the TEST wait os_reports, os_releases

Part C — Tester 2, what happened

  • Act 1 (signed agent_update to 0.143.0): queued 18:35 UTC, not delivered — Tester 2 stopped reporting at 18:06 UTC (host 18:05:48, controller 18:06:24). It was also silent 17:13–18:05 UTC, and its controller started at 18:06:20 UTC, so the box restarted or was switched on in between. CC sent Tester 2 nothing before 18:35 UTC; the cause is outside this work (power, network, or the tester). The job expires 19:20 UTC; if Tester 2 returns later, it is refused as expired (harmless) and must be re-signed.
  • Act 2 (bundle): waits for the operator's one-time bootstrap (R-862).
  • Act 3 (live-restore): not needed — already on.
  • No Docker step, no reboot, no app change was sent. Read-back of the System page line: still agent 0.142.0, Root files unknown, guard armed, live-restore on (last report 18:05 UTC).

Part B — the bundle

Shape, checks, trust rule, bootstrap: 11 §5.4.2; runbook runbooks/config-bundle.md. Red-proofs: 22 of 22 wrapper rules (partB/redproof.txt), hub 6 of 6 (partB/hub-bundle-redproof.txt), Go executor tests. Live:

Step Box Result
read the box's files vs the bundle both 21 of 22 identical; only felhom-os-apply differed (b1)
bootstrap, wrong sha demo-hp STOP, nothing changed (b2)
bootstrap both the new felhom-os-apply, self-check ok (b2, b3)
job A, wrong sha demo-hp refused by the agent, nothing changed (b4)
job B, the 0.143.0 bundle both written 0, same 21, kept 1, self-check ok, probe 71/71 (b4)
a bundle with one deliberate change demo-hp written 1 (that file), self-check ok (b6)
its undo (the 0.143.0 bundle again) demo-hp written 1, the release file back (sha b71d8698…), probe 71/71 (b6)
replay of job B demo-hp REJECTED … replay (nonce already seen) (b6)
the installer's new path demo-felhom written 0, record installer, services active (b5)

Part D — test approvals

Marked, cancelled at start, amber, operator event, ring-1 bump; red-proofs 6 of 6 (partD/d-redproof.txt). Live at the 18:20 UTC start: os-guest-20261004-123933, os-20261004-091417, os-host-20261004-124133, os-host-20261004-124034 cancelled; one mail (three held by the cooldown); the System page lists them; the Docker release stays (partD/d2). The 2026-10-04 fact and why the risk was small: runbooks/os-updates-test-waits.md.

Part E — the measured heal

Box Break Controller back traefik back Container ids
9202 (controller 0.291.0, before the fix) systemctl restart docker.socket never by itself (blind 2+ min, health "healthy") never same
9202 (0.293.0) same +73 s +104 s same (e7)
demo-hp 9201 (0.293.0) same +88 s +120 s 21 of 21 same (e8)

systemctl restart docker and kill -9 dockerd: socket inode unchanged, nothing to heal (e1, e2). Red-proof 9 of 9 (e6).

Part F — the first-night count

Golden 0.293.0 (documentation/tests/golden-0.293.0-2026-10-04/): no guest release in force → template versions kept; 49 Debian updates pending in the baked guest, which the next REAL approval (tomorrow, after 24 h + a night) will bring. A ring-1 box from this golden installs 0 on its first night until then. The host: the agent's first leg already runs right after the first whole-guest backup (Tester 2: 17 min after enrolment, 106 host packages in 44 s) — an installer pass would cost ~45 s and run before any backup exists; not built.

Part G — drill-r50

What the product delete touched (read first, partG/g1): hub rows (host, guest, reports 185, telemetry 185, the recovery credential, the DR recipe, the claim, an appliance registration, the WireGuard peer 10.77.0.4), ep0's PBS tenancy (tenantsync.Deprovision, namespace drill-r50), ep0's WireGuard peer list. No Storage Box (no off-site tier), no mail (no address). Asked the operator (ep0); yes. Result: deprovision ok (existed=false) — nothing destroyed on ep0; the peer left ep0 at the next push (5 → 4); /hosts without drill-r50; the delete preview 404s (partG/).

Rows

Before 333, after 334. Closed: R-840 (built), R-859 (opened and closed), R-860 (opened and closed). Opened: R-861 (P2 — the agent's sudoers is root-equivalent; read, not exploited), R-862 (P3 — Tester 2's bootstrap, operator). R-857 not fixed: avoided by baking under a new controller version.

Decisions taken by CC (operator may reverse)

  1. The one-time backfill marks only the AUTOMATIC test approvals; your Docker button approval of 2026-10-04 stays in force.
  2. The bundle alarm fires after 7 days behind (OS_ALARM_BUNDLE_BEHIND_AFTER).
  3. The bundle lives inside felhom-os-apply (no new sudoers line), not in a new tool.
  4. The controller fixes Part E (it reaches every box by the floor), not the agent (that would need a sudoers change).
  5. The demo customers' floor is 0.293.0; the global floor stays 0.292.0 so Tester 2's controller does not move this session.

Teardown

  • Machines: 9202 left on controller 0.293.0, working (all apps up). demo-hp and demo-felhom: agent 0.143.0, bundle 0.143.0, controller 0.293.0; the test line removed again on demo-hp. Drill VM off and at virgin.
  • Hosts: the bootstrap script, the installer harness and the test script removed from /root on both demo hosts. The bundles' previous copies stay in /var/lib/felhom-os-apply/bundle-prev/ by design.
  • Hub: the registry's test version 0.143.0-r840test deleted (204, then 404); drill-r50 removed; the scratchpad copy of the hub DB deleted at the end.