From 21986d03b5f350527ddef908f277cb7cf608ac6f Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 4 Oct 2026 17:49:30 +0200 Subject: [PATCH] golden 0.292.0 re-vouched (re-bake: live-restore + approved Docker set) with agent 0.142.0; R-857 filed; STATUS; REPORT-os-docker-crash-2026-10-04 Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT-os-docker-crash-2026-10-04.md | 83 +++++++++++++++++++ STATUS.md | 40 ++++++++- .../partC/c7-hub-rearmed.txt | 1 + documentation/backlog/OPEN-ITEMS.md | 3 +- .../05-vouch.txt | 6 ++ 5 files changed, 129 insertions(+), 4 deletions(-) create mode 100644 REPORT-os-docker-crash-2026-10-04.md create mode 100644 documentation/audits/os-docker-crash-2026-10-04/partC/c7-hub-rearmed.txt create mode 100644 documentation/tests/golden-0.292.0-2026-10-04-rebake/05-vouch.txt diff --git a/REPORT-os-docker-crash-2026-10-04.md b/REPORT-os-docker-crash-2026-10-04.md new file mode 100644 index 00000000..5be271a9 --- /dev/null +++ b/REPORT-os-docker-crash-2026-10-04.md @@ -0,0 +1,83 @@ +# REPORT — System page, Docker slow lane, crash restart (2026-10-04, late afternoon–evening) + +Brief: "OS updates — a System page …; the Docker engine slow lane built (live-restore ON, decision); a crashed host +restarts by itself, with a limit (decision)". Architecture read first: `documentation/architecture/11-os-updates.md` +(owner; §5.6, §5.8, §8.1–8.3), `05-hub-architecture.md`, `03-host-agent.md`, `08-alarm-ladder.md`, `07` §6.1. Rulings +recorded before the work: `09` §3 decisions 87–89. Evidence for every claim: `documentation/audits/os-docker-crash-2026-10-04/`. + +## Part table + +| Part | What | Result | Evidence | +|---|---|---|---| +| A | System page + version report (R-852) | **DONE.** The box reports Proxmox + kernel (API) and the wrapper's read-only facts (host Debian, next-boot kernel, held packages, taint, `kernel.panic`, crash guard; guest Debian, Docker, containerd, live-restore); unreadable = `unknown`. Hub: **System** tab on every page — per box ring + switch with buttons, tunnel, host, guest, Docker, last leg; releases per layer; what ring 0 runs; "Approve now" (confirm) and "Approve Docker set" (only when allowed). Hosts gets a Proxmox / kernel column. R-849: guest scanned every pass. Rendered with real data from both demo boxes. | `partA/` (`live/system-page.html`, `hosts-page.html`, facts) | +| B | Docker engine slow lane (`11` §5.8) | **DONE.** live-restore on by RELOAD: same ids on 9202 (6/6, by hand — R10 refuses a scratch guest by design), demo-hp (24/24, 7.5 s), demo-felhom (5/5). Ring 0 on both: 29.7.x → **29.8.2**, every id kept (122 s / 99 s whole pass). Approval with the page button under a TEST 0-night wait (logged, reverted): `os-docker-20261004-142842`. **Signed undo** on demo-hp → 29.7.2 (wrapper `authority=signed UNDO`, 24/24 ids, 41 s) and back to 29.8.2 by its ring-0 pass. **demo-felhom as ring 1**: its night pass skips Docker; a signed job with the approved set verified by the wrapper (already current → nothing); the **replayed** job refused ("nonce already seen"); back to ring 0. | `partB/` | +| C | Crash restart with a limit (R-851) | **DONE.** Spike (your word before each crash): `kernel.panic=10` + `echo c` → back by itself in 54 s, same kernel; pstore saved nothing; no oops history. Guard built (decision 88, 90–92). Live: crash 1 → 54 s, crash 2 → 53 s and the guard **tripped**, crash 3 → **stayed off** (184 s watched) until you switched it on; the hub mailed `host_crash_guard_tripped` (+3 restart events, 3 household lines); System page red; re-armed by `felhom-crash-guard rearm`. | `partC/` | +| D | Releases, golden, records | Agent **v0.142.0** (signed to both demo boxes), hub **v0.132.0**, installer **1.30.0** (tagged, public, verified), `build-golden.sh` 3.1.0. Golden: see below. Records: `11` §5.7/§5.8/§5.9, `00`, `03`, `07`, `08`, decisions 90–94, two runbooks. | `partD/`, git | + +## Claims in the brief that turned out wrong (or half right) + +- **"No box reports its versions"** — half right: every box already sent `pveversion` and `kversion` inside the Proxmox + API answer the agent reads each report (`NodeStatus`); nothing stored or showed them. Debian, the next-boot kernel and + everything Docker were truly not reported. +- **"A reload turns live-restore on with no container restart, on the demo boxes too"** — TRUE, measured: 24/24 and 5/5 + ids kept. But "prove on 9202 first" could not use the product path: 9202 binds a scratch folder, not the drives, so the + wrapper refuses it (R10, by design); proved there by hand with the same two steps. +- **"A crash boot can be told apart from a clean one"** — only from a CLEAN one. Not from a power cut or a hard reset: + `efi_pstore` is on, yet a real panic saved nothing on demo-hp. The guard counts every unclean stop (decision 90). +- **"`echo c` crashes the host and `kernel.panic` brings it back"** — TRUE (54 s, 53 s). But `sysctl -w` does not survive + the restart (back to 0), so it must be set at every boot — the guard does that. +- **The guard's count** — the brief said both "at 3 … the next crash leaves the box off" (the 4th) and "crashes 3 times + within one hour, it stays off" (the 3rd). Built per your words: the 3rd (decision 92). + +## Decisions taken by CC unattended (operator may reverse) — `09` §3 + +90 crash signal = clean-stop marker · 91 `panic_on_oops` stays 0, an oops is mailed · 92 the 3rd unclean stop in 60 min +stays off; 10 s; 24 h · 93 the wrapper verifies Docker authority against root-owned files · 94 page colours = alarm +thresholds. + +## Releases and what was copied by hand + +- **Agent v0.142.0** (`b1746c2`, sha256 `7beb3222…de6`): signed `agent_update` to both demo boxes, both COMPLETED. +- **Hub v0.132.0** (`175ecfc`): image tag verified on the pod; ArgoCD Synced/Healthy. +- **Installer 1.30.0**: tag `installer-v1.30.0`, both git-sync refs; `https://felhom.eu/scripts/felhom-host-install.sh` + serves `SCRIPT_VERSION="1.30.0"`. +- **By hand on both demo hosts** (R-840; `partD/copied-by-hand-*.txt`, hashes equal to tag v0.142.0): + `/usr/local/sbin/felhom-os-apply`, `/usr/local/sbin/felhom-crash-guard`, `/etc/systemd/system/felhom-crash-guard.service`, + `…/felhom-crash-guard-check.service`, `…/felhom-crash-guard-check.timer`, `/etc/felhom/crash-guard.conf`, + `/etc/felhom/operator-signers`, `/etc/felhom/os-trust.json` (with `ring0_slow_lane: true` — the demo boxes only); + units enabled. Previous wrapper kept as `/root/felhom-os-apply.bak-0.141.1`. A test binary (`felhom-agent-0.142.0-rc1`) + ran the debug actions before the release and was removed after. + +## Golden + +- **Re-baked golden 0.292.0** with `build-golden.sh` 3.1.0 (by a helper agent, RUNBOOK-manual-build §4.0/§4.1; the + documented publish replaces the same version: pre-delete 204, upload 201). The log shows "Docker engine set PINNED", + the six approved versions (docker-ce 5:29.8.2, containerd.io 2.3.6, buildx 0.37.1, compose 5.6.0, …) and + "live-restore: on". New sha256 `79a1dce3…d43a`, re-hashed by download in the main session: match. Drill VM destroyed, + drill disk back to `virgin`. Evidence `documentation/tests/golden-0.292.0-2026-10-04-rebake/` (commit `208d21d`). +- **Re-vouched:** agent 0.142.0 + golden 0.292.0 (new sha), `min_agent` 0.131.0 → `artifacts_set` (17:48). Between the + re-bake upload and the re-vouch the hub vouched the old sha — a fresh install would have failed closed; none ran (R-857). +- The crash guard is NOT in the golden (it is a host program): the installer 1.30.0 installs it. +- The golden waiver stays deleted: the newest controller (0.292.0) has its golden. + +## Register + +Open rows **334 → 333** (333 at the start + R-852 filed first). Closed: **R-852, R-835, R-848, R-849, R-851**; filed and +closed: **R-854**. Opened: **R-853** (facts reach the hub ~15 min late after a boot), **R-855** (cosmetic TEST log line), +**R-856** (after a crash the household also gets app mails — your choice later), **R-857** (a same-version golden +re-bake: the gate shows the first sha; a window until the re-vouch). Narrowed: **R-812** (Docker lane built; +kernel left), **R-840** (by hand again). `unproven.py`: unchanged (35 of 55 not walked). + +## Teardown — three layers + +- **Machine:** demo-hp and demo-felhom customer guests running, live-restore on, Docker 29.8.2, all apps up. 9202 + running, live-restore on (its old daemon.json kept as `/root/daemon.json.bak-2026-10-04` in the guest). +- **Host:** both hosts run agent 0.142.0, the crash guard ARMED (`kernel.panic = 10`), ring 0, switch ON; the test + binary and the id lists removed. demo-hp booted 3 extra times today (the crash test); kernel unchanged (7.0.14-20). +- **Hub:** the TEST Docker wait reverted (log: 2 nights); the approval `os-docker-20261004-142842` stays (a real + approval of what ring 0 runs). The crash and trip mails for demo-hp were sent on purpose. The hub password copy in + the scratchpad shredded at the end. The hub announced the re-arm (`host_crash_guard_rearmed`, 17:47). + +## CI + +CI_SECTION diff --git a/STATUS.md b/STATUS.md index ec814800..b86a9d13 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,9 +2,43 @@ **Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).** -**Updated 2026-10-04 (evening): the HOST's security fixes install themselves too, the hub has a fleet view and four -OS alarms, and the tunnel status is true. Both demo boxes run controller 0.292.0 and host agent 0.141.1. Hub 0.131.1. -New installs get golden 0.292.0 with agent 0.141.1 (vouched).** +**Updated 2026-10-04 (late evening): the hub has a System page with every box's versions and the update buttons; Docker +updates are built; a crashed box restarts by itself, at most twice an hour. Both demo boxes run host agent 0.142.0 and +controller 0.292.0. Hub 0.132.0. Installer 1.30.0. New installs get golden 0.292.0 (re-made: live-restore on, Docker 29.8.2) with agent 0.142.0.** + +## Today (2026-10-04, late evening): the System page, Docker updates, the crash restart + +**Decisions I took myself (you may reverse each):** +- The box knows a crash only as "it did not shut down cleanly" (the crash memory chip saved nothing). So a power cut + also counts as a crash. +- An "oops" (a kernel error the box survives) does not restart the box; you get a mail instead. +- Your words win where the brief disagreed: the **3rd** crash within one hour leaves the box off. Restart after 10 s, + re-arm after 24 h. All settings. +- Only the box itself decides whether a Docker update is allowed: it checks your signature with a key file only root + can change (not the agent's own settings, which the agent could change). +- The System page uses the alarm limits for its colours (red = an alarm would fire). + +**No decision needed from you today.** + +**What I did:** +- **The System tab** in the hub: per box the Proxmox, kernel (now and next boot), Debian and Docker versions, what is + waiting, held packages, "restart needed", the crash guard and the last update run — with buttons for ring, on/off, + "Approve now" and "Approve Docker set". The Hosts page shows Proxmox and kernel too. +- **Docker updates:** "live-restore" is on in every box (no app restarted: 24 of 24 and 5 of 5 containers kept running). + Both demo boxes moved to Docker 29.8.2, every app kept running. I approved that set with the new button (a 0-night + test wait, then back to 2 nights). An undo signed by your key put demo-hp back one version and forward again, apps + running throughout. A copied (replayed) signed job was refused. +- **Crash restart (your 3 crashes on demo-hp):** crash 1 and 2 — back by itself in under a minute; crash 3 — it stayed + off until you switched it on. You got the "guard tripped" mail. I re-armed it. +- **Found:** the hub learns about a crash up to 15 minutes late (nothing lost). After a crash, the household can also get + an "app stopped" mail besides "restarted after a crash" — whether to calm that is a later choice for you. +- **New-install image re-made** with live-restore on and the approved Docker version, and approved in the hub with host + agent 0.142.0. New boxes also get the crash guard (installer 1.30.0). +- **Rows:** 5 closed, 1 opened-and-closed the same day, 4 opened. The list went from 334 to 333. + +**Needs you later (nothing breaks if you wait):** +- Installed boxes still get new root files only by hand (the long-standing gap; there are no other boxes today). +- Kernel updates are still not built (a hung new kernel would stay — needs a fix first). ## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status diff --git a/documentation/audits/os-docker-crash-2026-10-04/partC/c7-hub-rearmed.txt b/documentation/audits/os-docker-crash-2026-10-04/partC/c7-hub-rearmed.txt new file mode 100644 index 00000000..d201bf4c --- /dev/null +++ b/documentation/audits/os-docker-crash-2026-10-04/partC/c7-hub-rearmed.txt @@ -0,0 +1 @@ +2026/10/04 17:47:51 [WARN] host demo-hp-bb76ea crash guard: host_crash_guard_rearmed diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 5c1c8cbb..37b2257e 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -398,7 +398,7 @@ stopping line that lies. | **R-793** | Business & legal | P4 | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. **Watch:** never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the `-enterprise` image, and re-read on each major. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them.** | — | — | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 86 rows (P3 4, P4 82) +## Process & tooling — 87 rows (P3 4, P4 83) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -489,6 +489,7 @@ stopping line that lies. | **R-817** | Process & tooling | P4 | **`09` decision 56 and R-745 disagree about what the controller self-update rolls back to.** Decision 56 (`documentation/architecture/09-update-architecture.md:619`) says a box keeps the image before the running one as *"the self-update's roll-back target"*; R-745 (closed, `CLOSED-ITEMS.md`) measured that *"the agent rolls back to the RUNNING image (what `/etc/felhom-controller-image` named), never to a previous one the controller hands it"*. Found 2026-10-03 while carrying R-745's rule to `CONTEXT.md`. | **VERIFY — filed 2026-10-03 (triage); owner: CC.** Read the agent's rollback path and correct whichever text is wrong. | — | — | CC | | **R-818** | Process & tooling | P4 | **Two changelogs cite register ids for other findings.** `hub/CHANGELOG.md:526-538` (hub v0.109.0, the Backup card) and its note at `:533` use **R-331** and **R-330** for a hub display fault and a nightly false app alarm; in the register R-330 and R-331 are the disk-health Phase 2 and Phase 3 rows. A reader following the id lands on the wrong finding. Found 2026-10-03 by the triage. | **READY — filed 2026-10-03 (triage); owner: CC.** Add a dated correction line under each changelog entry naming the right rows (the real ids are in `CLOSED-ITEMS.md`); do not renumber anything. | — | — | CC | | **R-819** | Process & tooling | P4 | **`scripts/check_stands.py` is red and runs in no runner.** Measured 2026-10-03 on `9e2786c` (before the triage): it convicts `where-felhom-stands.yaml` for citing R-273 and R-356, which are in neither `OPEN-ITEMS.md` nor anywhere it reads. After the triage it also convicts R-281, R-198 and R-201, because its rule 3 reads only `OPEN-ITEMS.md` and those rows are closed. It is in neither `repo_gates.py` nor CI, so nobody saw it — the R-29 shape. | **READY — filed 2026-10-03 (triage); owner: CC.** Let rule 3 accept an id in `CLOSED-ITEMS.md` (and check the stand's status agrees), fix the two dangling ids, then register it in `repo_gates.py` with a decoy. | — | — | CC | +| **R-857** | Process & tooling | P4 | **Baking a golden twice under the SAME version leaves two stale facts.** 2026-10-04 (golden 0.292.0 re-baked with live-restore): (1) `golden_currency_gate.py` keeps reporting the FIRST bake directory's sha (`golden-0.292.0-2026-10-04`, d6cf8b33…) — it picks one of two directories with the same version, not the newest; (2) the publish replaces the package in place (pre-delete + upload), so until the operator re-vouches, the hub vouches a sha the registry no longer holds — a fresh install in that window fails its sha check (fail-closed; none ran; re-vouched 17:48). Fix direction: the gate prefers the newest bake directory (or refuses two for one version); the runbook says "re-vouch at once after a same-version re-bake". `documentation/tests/golden-0.292.0-2026-10-04-rebake/` | **READY — owner: CC** | — | — | CC |