golden 0.292.0 re-vouched (re-bake: live-restore + approved Docker set) with agent 0.142.0; R-857 filed; STATUS; REPORT-os-docker-crash-2026-10-04
gates / gates (push) Successful in 32s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-04 17:49:30 +02:00
parent 208d21d18f
commit 21986d03b5
5 changed files with 129 additions and 4 deletions
+83
View File
@@ -0,0 +1,83 @@
# REPORT — System page, Docker slow lane, crash restart (2026-10-04, late afternoon–evening)
Brief: "OS updates — a System page …; the Docker engine slow lane built (live-restore ON, decision); a crashed host
restarts by itself, with a limit (decision)". Architecture read first: `documentation/architecture/11-os-updates.md`
(owner; §5.6, §5.8, §8.1–8.3), `05-hub-architecture.md`, `03-host-agent.md`, `08-alarm-ladder.md`, `07` §6.1. Rulings
recorded before the work: `09` §3 decisions 87–89. Evidence for every claim: `documentation/audits/os-docker-crash-2026-10-04/`.
## Part table
| Part | What | Result | Evidence |
|---|---|---|---|
| A | System page + version report (R-852) | **DONE.** The box reports Proxmox + kernel (API) and the wrapper's read-only facts (host Debian, next-boot kernel, held packages, taint, `kernel.panic`, crash guard; guest Debian, Docker, containerd, live-restore); unreadable = `unknown`. Hub: **System** tab on every page — per box ring + switch with buttons, tunnel, host, guest, Docker, last leg; releases per layer; what ring 0 runs; "Approve now" (confirm) and "Approve Docker set" (only when allowed). Hosts gets a Proxmox / kernel column. R-849: guest scanned every pass. Rendered with real data from both demo boxes. | `partA/` (`live/system-page.html`, `hosts-page.html`, facts) |
| B | Docker engine slow lane (`11` §5.8) | **DONE.** live-restore on by RELOAD: same ids on 9202 (6/6, by hand — R10 refuses a scratch guest by design), demo-hp (24/24, 7.5 s), demo-felhom (5/5). Ring 0 on both: 29.7.x → **29.8.2**, every id kept (122 s / 99 s whole pass). Approval with the page button under a TEST 0-night wait (logged, reverted): `os-docker-20261004-142842`. **Signed undo** on demo-hp → 29.7.2 (wrapper `authority=signed UNDO`, 24/24 ids, 41 s) and back to 29.8.2 by its ring-0 pass. **demo-felhom as ring 1**: its night pass skips Docker; a signed job with the approved set verified by the wrapper (already current → nothing); the **replayed** job refused ("nonce already seen"); back to ring 0. | `partB/` |
| C | Crash restart with a limit (R-851) | **DONE.** Spike (your word before each crash): `kernel.panic=10` + `echo c` → back by itself in 54 s, same kernel; pstore saved nothing; no oops history. Guard built (decision 88, 90–92). Live: crash 1 → 54 s, crash 2 → 53 s and the guard **tripped**, crash 3 → **stayed off** (184 s watched) until you switched it on; the hub mailed `host_crash_guard_tripped` (+3 restart events, 3 household lines); System page red; re-armed by `felhom-crash-guard rearm`. | `partC/` |
| D | Releases, golden, records | Agent **v0.142.0** (signed to both demo boxes), hub **v0.132.0**, installer **1.30.0** (tagged, public, verified), `build-golden.sh` 3.1.0. Golden: see below. Records: `11` §5.7/§5.8/§5.9, `00`, `03`, `07`, `08`, decisions 90–94, two runbooks. | `partD/`, git |
## Claims in the brief that turned out wrong (or half right)
- **"No box reports its versions"** — half right: every box already sent `pveversion` and `kversion` inside the Proxmox
API answer the agent reads each report (`NodeStatus`); nothing stored or showed them. Debian, the next-boot kernel and
everything Docker were truly not reported.
- **"A reload turns live-restore on with no container restart, on the demo boxes too"** — TRUE, measured: 24/24 and 5/5
ids kept. But "prove on 9202 first" could not use the product path: 9202 binds a scratch folder, not the drives, so the
wrapper refuses it (R10, by design); proved there by hand with the same two steps.
- **"A crash boot can be told apart from a clean one"** — only from a CLEAN one. Not from a power cut or a hard reset:
`efi_pstore` is on, yet a real panic saved nothing on demo-hp. The guard counts every unclean stop (decision 90).
- **"`echo c` crashes the host and `kernel.panic` brings it back"** — TRUE (54 s, 53 s). But `sysctl -w` does not survive
the restart (back to 0), so it must be set at every boot — the guard does that.
- **The guard's count** — the brief said both "at 3 … the next crash leaves the box off" (the 4th) and "crashes 3 times
within one hour, it stays off" (the 3rd). Built per your words: the 3rd (decision 92).
## Decisions taken by CC unattended (operator may reverse) — `09` §3
90 crash signal = clean-stop marker · 91 `panic_on_oops` stays 0, an oops is mailed · 92 the 3rd unclean stop in 60 min
stays off; 10 s; 24 h · 93 the wrapper verifies Docker authority against root-owned files · 94 page colours = alarm
thresholds.
## Releases and what was copied by hand
- **Agent v0.142.0** (`b1746c2`, sha256 `7beb3222…de6`): signed `agent_update` to both demo boxes, both COMPLETED.
- **Hub v0.132.0** (`175ecfc`): image tag verified on the pod; ArgoCD Synced/Healthy.
- **Installer 1.30.0**: tag `installer-v1.30.0`, both git-sync refs; `https://felhom.eu/scripts/felhom-host-install.sh`
serves `SCRIPT_VERSION="1.30.0"`.
- **By hand on both demo hosts** (R-840; `partD/copied-by-hand-*.txt`, hashes equal to tag v0.142.0):
`/usr/local/sbin/felhom-os-apply`, `/usr/local/sbin/felhom-crash-guard`, `/etc/systemd/system/felhom-crash-guard.service`,
`…/felhom-crash-guard-check.service`, `…/felhom-crash-guard-check.timer`, `/etc/felhom/crash-guard.conf`,
`/etc/felhom/operator-signers`, `/etc/felhom/os-trust.json` (with `ring0_slow_lane: true` — the demo boxes only);
units enabled. Previous wrapper kept as `/root/felhom-os-apply.bak-0.141.1`. A test binary (`felhom-agent-0.142.0-rc1`)
ran the debug actions before the release and was removed after.
## Golden
- **Re-baked golden 0.292.0** with `build-golden.sh` 3.1.0 (by a helper agent, RUNBOOK-manual-build §4.0/§4.1; the
documented publish replaces the same version: pre-delete 204, upload 201). The log shows "Docker engine set PINNED",
the six approved versions (docker-ce 5:29.8.2, containerd.io 2.3.6, buildx 0.37.1, compose 5.6.0, …) and
"live-restore: on". New sha256 `79a1dce3…d43a`, re-hashed by download in the main session: match. Drill VM destroyed,
drill disk back to `virgin`. Evidence `documentation/tests/golden-0.292.0-2026-10-04-rebake/` (commit `208d21d`).
- **Re-vouched:** agent 0.142.0 + golden 0.292.0 (new sha), `min_agent` 0.131.0 → `artifacts_set` (17:48). Between the
re-bake upload and the re-vouch the hub vouched the old sha — a fresh install would have failed closed; none ran (R-857).
- The crash guard is NOT in the golden (it is a host program): the installer 1.30.0 installs it.
- The golden waiver stays deleted: the newest controller (0.292.0) has its golden.
## Register
Open rows **334 → 333** (333 at the start + R-852 filed first). Closed: **R-852, R-835, R-848, R-849, R-851**; filed and
closed: **R-854**. Opened: **R-853** (facts reach the hub ~15 min late after a boot), **R-855** (cosmetic TEST log line),
**R-856** (after a crash the household also gets app mails — your choice later), **R-857** (a same-version golden
re-bake: the gate shows the first sha; a window until the re-vouch). Narrowed: **R-812** (Docker lane built;
kernel left), **R-840** (by hand again). `unproven.py`: unchanged (35 of 55 not walked).
## Teardown — three layers
- **Machine:** demo-hp and demo-felhom customer guests running, live-restore on, Docker 29.8.2, all apps up. 9202
running, live-restore on (its old daemon.json kept as `/root/daemon.json.bak-2026-10-04` in the guest).
- **Host:** both hosts run agent 0.142.0, the crash guard ARMED (`kernel.panic = 10`), ring 0, switch ON; the test
binary and the id lists removed. demo-hp booted 3 extra times today (the crash test); kernel unchanged (7.0.14-20).
- **Hub:** the TEST Docker wait reverted (log: 2 nights); the approval `os-docker-20261004-142842` stays (a real
approval of what ring 0 runs). The crash and trip mails for demo-hp were sent on purpose. The hub password copy in
the scratchpad shredded at the end. The hub announced the re-arm (`host_crash_guard_rearmed`, 17:47).
## CI
CI_SECTION
+37 -3
View File
@@ -2,9 +2,43 @@
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
**Updated 2026-10-04 (evening): the HOST's security fixes install themselves too, the hub has a fleet view and four
OS alarms, and the tunnel status is true. Both demo boxes run controller 0.292.0 and host agent 0.141.1. Hub 0.131.1.
New installs get golden 0.292.0 with agent 0.141.1 (vouched).**
**Updated 2026-10-04 (late evening): the hub has a System page with every box's versions and the update buttons; Docker
updates are built; a crashed box restarts by itself, at most twice an hour. Both demo boxes run host agent 0.142.0 and
controller 0.292.0. Hub 0.132.0. Installer 1.30.0. New installs get golden 0.292.0 (re-made: live-restore on, Docker 29.8.2) with agent 0.142.0.**
## Today (2026-10-04, late evening): the System page, Docker updates, the crash restart
**Decisions I took myself (you may reverse each):**
- The box knows a crash only as "it did not shut down cleanly" (the crash memory chip saved nothing). So a power cut
also counts as a crash.
- An "oops" (a kernel error the box survives) does not restart the box; you get a mail instead.
- Your words win where the brief disagreed: the **3rd** crash within one hour leaves the box off. Restart after 10 s,
re-arm after 24 h. All settings.
- Only the box itself decides whether a Docker update is allowed: it checks your signature with a key file only root
can change (not the agent's own settings, which the agent could change).
- The System page uses the alarm limits for its colours (red = an alarm would fire).
**No decision needed from you today.**
**What I did:**
- **The System tab** in the hub: per box the Proxmox, kernel (now and next boot), Debian and Docker versions, what is
waiting, held packages, "restart needed", the crash guard and the last update run — with buttons for ring, on/off,
"Approve now" and "Approve Docker set". The Hosts page shows Proxmox and kernel too.
- **Docker updates:** "live-restore" is on in every box (no app restarted: 24 of 24 and 5 of 5 containers kept running).
Both demo boxes moved to Docker 29.8.2, every app kept running. I approved that set with the new button (a 0-night
test wait, then back to 2 nights). An undo signed by your key put demo-hp back one version and forward again, apps
running throughout. A copied (replayed) signed job was refused.
- **Crash restart (your 3 crashes on demo-hp):** crash 1 and 2 — back by itself in under a minute; crash 3 — it stayed
off until you switched it on. You got the "guard tripped" mail. I re-armed it.
- **Found:** the hub learns about a crash up to 15 minutes late (nothing lost). After a crash, the household can also get
an "app stopped" mail besides "restarted after a crash" — whether to calm that is a later choice for you.
- **New-install image re-made** with live-restore on and the approved Docker version, and approved in the hub with host
agent 0.142.0. New boxes also get the crash guard (installer 1.30.0).
- **Rows:** 5 closed, 1 opened-and-closed the same day, 4 opened. The list went from 334 to 333.
**Needs you later (nothing breaks if you wait):**
- Installed boxes still get new root files only by hand (the long-standing gap; there are no other boxes today).
- Kernel updates are still not built (a hung new kernel would stay — needs a fix first).
## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status
@@ -0,0 +1 @@
2026/10/04 17:47:51 [WARN] host demo-hp-bb76ea crash guard: host_crash_guard_rearmed
+2 -1
View File
@@ -398,7 +398,7 @@ stopping line that lies.
| **R-793** | Business & legal | P4 | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. **Watch:** never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the `-enterprise` image, and re-read on each major. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them.** | — | — | CC |
| **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC |
## Process & tooling — 86 rows (P3 4, P4 82)
## Process & tooling — 87 rows (P3 4, P4 83)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -489,6 +489,7 @@ stopping line that lies.
| **R-817** | Process & tooling | P4 | **`09` decision 56 and R-745 disagree about what the controller self-update rolls back to.** Decision 56 (`documentation/architecture/09-update-architecture.md:619`) says a box keeps the image before the running one as *"the self-update's roll-back target"*; R-745 (closed, `CLOSED-ITEMS.md`) measured that *"the agent rolls back to the RUNNING image (what `/etc/felhom-controller-image` named), never to a previous one the controller hands it"*. Found 2026-10-03 while carrying R-745's rule to `CONTEXT.md`. | **VERIFY — filed 2026-10-03 (triage); owner: CC.** Read the agent's rollback path and correct whichever text is wrong. | — | — | CC |
| **R-818** | Process & tooling | P4 | **Two changelogs cite register ids for other findings.** `hub/CHANGELOG.md:526-538` (hub v0.109.0, the Backup card) and its note at `:533` use **R-331** and **R-330** for a hub display fault and a nightly false app alarm; in the register R-330 and R-331 are the disk-health Phase 2 and Phase 3 rows. A reader following the id lands on the wrong finding. Found 2026-10-03 by the triage. | **READY — filed 2026-10-03 (triage); owner: CC.** Add a dated correction line under each changelog entry naming the right rows (the real ids are in `CLOSED-ITEMS.md`); do not renumber anything. | — | — | CC |
| **R-819** | Process & tooling | P4 | **`scripts/check_stands.py` is red and runs in no runner.** Measured 2026-10-03 on `9e2786c` (before the triage): it convicts `where-felhom-stands.yaml` for citing R-273 and R-356, which are in neither `OPEN-ITEMS.md` nor anywhere it reads. After the triage it also convicts R-281, R-198 and R-201, because its rule 3 reads only `OPEN-ITEMS.md` and those rows are closed. It is in neither `repo_gates.py` nor CI, so nobody saw it — the R-29 shape. | **READY — filed 2026-10-03 (triage); owner: CC.** Let rule 3 accept an id in `CLOSED-ITEMS.md` (and check the stand's status agrees), fix the two dangling ids, then register it in `repo_gates.py` with a decoy. | — | — | CC |
| **R-857** | Process & tooling | P4 | **Baking a golden twice under the SAME version leaves two stale facts.** 2026-10-04 (golden 0.292.0 re-baked with live-restore): (1) `golden_currency_gate.py` keeps reporting the FIRST bake directory's sha (`golden-0.292.0-2026-10-04`, d6cf8b33…) — it picks one of two directories with the same version, not the newest; (2) the publish replaces the package in place (pre-delete + upload), so until the operator re-vouches, the hub vouches a sha the registry no longer holds — a fresh install in that window fails its sha check (fail-closed; none ran; re-vouched 17:48). Fix direction: the gate prefers the newest bake directory (or refuses two for one version); the runbook says "re-vouch at once after a same-version re-bake". `documentation/tests/golden-0.292.0-2026-10-04-rebake/` | **READY — owner: CC** | — | — | CC |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
@@ -0,0 +1,6 @@
Re-vouch 2026-10-04 17:48 CEST by CC (main session) after the same-version re-bake, POST /configuration/artifacts (operator Basic auth, ClusterIP):
agent_version=0.142.0 agent_sha256=7beb32224d6495e9561acfd3ad8a48393a799520f196080011cb27fceb6d1de6
golden_version=0.292.0 golden_sha256=79a1dce3bd3c5a636a03c82be0f4ef969a1e0e9cbb139c5890b84c98fd69d43a (re-hashed by download in the main session: match)
min_agent=0.131.0 -> 303 artifacts_set
hub log: Artifact manifest set: agent=0.142.0 golden=0.292.0 min_agent="0.131.0" wrapper_sha=false
Between the re-bake upload and this re-vouch the hub vouched the OLD sha (d6cf8b33...) for a package that no longer had it: a fresh install in that window would have failed its sha check (fail-closed). None ran. R-857.