golden 0.292.0 re-vouched (re-bake: live-restore + approved Docker set) with agent 0.142.0; R-857 filed; STATUS; REPORT-os-docker-crash-2026-10-04
gates / gates (push) Successful in 32s
gates / gates (push) Successful in 32s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,83 @@
|
||||
# REPORT — System page, Docker slow lane, crash restart (2026-10-04, late afternoon–evening)
|
||||
|
||||
Brief: "OS updates — a System page …; the Docker engine slow lane built (live-restore ON, decision); a crashed host
|
||||
restarts by itself, with a limit (decision)". Architecture read first: `documentation/architecture/11-os-updates.md`
|
||||
(owner; §5.6, §5.8, §8.1–8.3), `05-hub-architecture.md`, `03-host-agent.md`, `08-alarm-ladder.md`, `07` §6.1. Rulings
|
||||
recorded before the work: `09` §3 decisions 87–89. Evidence for every claim: `documentation/audits/os-docker-crash-2026-10-04/`.
|
||||
|
||||
## Part table
|
||||
|
||||
| Part | What | Result | Evidence |
|
||||
|---|---|---|---|
|
||||
| A | System page + version report (R-852) | **DONE.** The box reports Proxmox + kernel (API) and the wrapper's read-only facts (host Debian, next-boot kernel, held packages, taint, `kernel.panic`, crash guard; guest Debian, Docker, containerd, live-restore); unreadable = `unknown`. Hub: **System** tab on every page — per box ring + switch with buttons, tunnel, host, guest, Docker, last leg; releases per layer; what ring 0 runs; "Approve now" (confirm) and "Approve Docker set" (only when allowed). Hosts gets a Proxmox / kernel column. R-849: guest scanned every pass. Rendered with real data from both demo boxes. | `partA/` (`live/system-page.html`, `hosts-page.html`, facts) |
|
||||
| B | Docker engine slow lane (`11` §5.8) | **DONE.** live-restore on by RELOAD: same ids on 9202 (6/6, by hand — R10 refuses a scratch guest by design), demo-hp (24/24, 7.5 s), demo-felhom (5/5). Ring 0 on both: 29.7.x → **29.8.2**, every id kept (122 s / 99 s whole pass). Approval with the page button under a TEST 0-night wait (logged, reverted): `os-docker-20261004-142842`. **Signed undo** on demo-hp → 29.7.2 (wrapper `authority=signed UNDO`, 24/24 ids, 41 s) and back to 29.8.2 by its ring-0 pass. **demo-felhom as ring 1**: its night pass skips Docker; a signed job with the approved set verified by the wrapper (already current → nothing); the **replayed** job refused ("nonce already seen"); back to ring 0. | `partB/` |
|
||||
| C | Crash restart with a limit (R-851) | **DONE.** Spike (your word before each crash): `kernel.panic=10` + `echo c` → back by itself in 54 s, same kernel; pstore saved nothing; no oops history. Guard built (decision 88, 90–92). Live: crash 1 → 54 s, crash 2 → 53 s and the guard **tripped**, crash 3 → **stayed off** (184 s watched) until you switched it on; the hub mailed `host_crash_guard_tripped` (+3 restart events, 3 household lines); System page red; re-armed by `felhom-crash-guard rearm`. | `partC/` |
|
||||
| D | Releases, golden, records | Agent **v0.142.0** (signed to both demo boxes), hub **v0.132.0**, installer **1.30.0** (tagged, public, verified), `build-golden.sh` 3.1.0. Golden: see below. Records: `11` §5.7/§5.8/§5.9, `00`, `03`, `07`, `08`, decisions 90–94, two runbooks. | `partD/`, git |
|
||||
|
||||
## Claims in the brief that turned out wrong (or half right)
|
||||
|
||||
- **"No box reports its versions"** — half right: every box already sent `pveversion` and `kversion` inside the Proxmox
|
||||
API answer the agent reads each report (`NodeStatus`); nothing stored or showed them. Debian, the next-boot kernel and
|
||||
everything Docker were truly not reported.
|
||||
- **"A reload turns live-restore on with no container restart, on the demo boxes too"** — TRUE, measured: 24/24 and 5/5
|
||||
ids kept. But "prove on 9202 first" could not use the product path: 9202 binds a scratch folder, not the drives, so the
|
||||
wrapper refuses it (R10, by design); proved there by hand with the same two steps.
|
||||
- **"A crash boot can be told apart from a clean one"** — only from a CLEAN one. Not from a power cut or a hard reset:
|
||||
`efi_pstore` is on, yet a real panic saved nothing on demo-hp. The guard counts every unclean stop (decision 90).
|
||||
- **"`echo c` crashes the host and `kernel.panic` brings it back"** — TRUE (54 s, 53 s). But `sysctl -w` does not survive
|
||||
the restart (back to 0), so it must be set at every boot — the guard does that.
|
||||
- **The guard's count** — the brief said both "at 3 … the next crash leaves the box off" (the 4th) and "crashes 3 times
|
||||
within one hour, it stays off" (the 3rd). Built per your words: the 3rd (decision 92).
|
||||
|
||||
## Decisions taken by CC unattended (operator may reverse) — `09` §3
|
||||
|
||||
90 crash signal = clean-stop marker · 91 `panic_on_oops` stays 0, an oops is mailed · 92 the 3rd unclean stop in 60 min
|
||||
stays off; 10 s; 24 h · 93 the wrapper verifies Docker authority against root-owned files · 94 page colours = alarm
|
||||
thresholds.
|
||||
|
||||
## Releases and what was copied by hand
|
||||
|
||||
- **Agent v0.142.0** (`b1746c2`, sha256 `7beb3222…de6`): signed `agent_update` to both demo boxes, both COMPLETED.
|
||||
- **Hub v0.132.0** (`175ecfc`): image tag verified on the pod; ArgoCD Synced/Healthy.
|
||||
- **Installer 1.30.0**: tag `installer-v1.30.0`, both git-sync refs; `https://felhom.eu/scripts/felhom-host-install.sh`
|
||||
serves `SCRIPT_VERSION="1.30.0"`.
|
||||
- **By hand on both demo hosts** (R-840; `partD/copied-by-hand-*.txt`, hashes equal to tag v0.142.0):
|
||||
`/usr/local/sbin/felhom-os-apply`, `/usr/local/sbin/felhom-crash-guard`, `/etc/systemd/system/felhom-crash-guard.service`,
|
||||
`…/felhom-crash-guard-check.service`, `…/felhom-crash-guard-check.timer`, `/etc/felhom/crash-guard.conf`,
|
||||
`/etc/felhom/operator-signers`, `/etc/felhom/os-trust.json` (with `ring0_slow_lane: true` — the demo boxes only);
|
||||
units enabled. Previous wrapper kept as `/root/felhom-os-apply.bak-0.141.1`. A test binary (`felhom-agent-0.142.0-rc1`)
|
||||
ran the debug actions before the release and was removed after.
|
||||
|
||||
## Golden
|
||||
|
||||
- **Re-baked golden 0.292.0** with `build-golden.sh` 3.1.0 (by a helper agent, RUNBOOK-manual-build §4.0/§4.1; the
|
||||
documented publish replaces the same version: pre-delete 204, upload 201). The log shows "Docker engine set PINNED",
|
||||
the six approved versions (docker-ce 5:29.8.2, containerd.io 2.3.6, buildx 0.37.1, compose 5.6.0, …) and
|
||||
"live-restore: on". New sha256 `79a1dce3…d43a`, re-hashed by download in the main session: match. Drill VM destroyed,
|
||||
drill disk back to `virgin`. Evidence `documentation/tests/golden-0.292.0-2026-10-04-rebake/` (commit `208d21d`).
|
||||
- **Re-vouched:** agent 0.142.0 + golden 0.292.0 (new sha), `min_agent` 0.131.0 → `artifacts_set` (17:48). Between the
|
||||
re-bake upload and the re-vouch the hub vouched the old sha — a fresh install would have failed closed; none ran (R-857).
|
||||
- The crash guard is NOT in the golden (it is a host program): the installer 1.30.0 installs it.
|
||||
- The golden waiver stays deleted: the newest controller (0.292.0) has its golden.
|
||||
|
||||
## Register
|
||||
|
||||
Open rows **334 → 333** (333 at the start + R-852 filed first). Closed: **R-852, R-835, R-848, R-849, R-851**; filed and
|
||||
closed: **R-854**. Opened: **R-853** (facts reach the hub ~15 min late after a boot), **R-855** (cosmetic TEST log line),
|
||||
**R-856** (after a crash the household also gets app mails — your choice later), **R-857** (a same-version golden
|
||||
re-bake: the gate shows the first sha; a window until the re-vouch). Narrowed: **R-812** (Docker lane built;
|
||||
kernel left), **R-840** (by hand again). `unproven.py`: unchanged (35 of 55 not walked).
|
||||
|
||||
## Teardown — three layers
|
||||
|
||||
- **Machine:** demo-hp and demo-felhom customer guests running, live-restore on, Docker 29.8.2, all apps up. 9202
|
||||
running, live-restore on (its old daemon.json kept as `/root/daemon.json.bak-2026-10-04` in the guest).
|
||||
- **Host:** both hosts run agent 0.142.0, the crash guard ARMED (`kernel.panic = 10`), ring 0, switch ON; the test
|
||||
binary and the id lists removed. demo-hp booted 3 extra times today (the crash test); kernel unchanged (7.0.14-20).
|
||||
- **Hub:** the TEST Docker wait reverted (log: 2 nights); the approval `os-docker-20261004-142842` stays (a real
|
||||
approval of what ring 0 runs). The crash and trip mails for demo-hp were sent on purpose. The hub password copy in
|
||||
the scratchpad shredded at the end. The hub announced the re-arm (`host_crash_guard_rearmed`, 17:47).
|
||||
|
||||
## CI
|
||||
|
||||
CI_SECTION
|
||||
@@ -2,9 +2,43 @@
|
||||
|
||||
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
|
||||
|
||||
**Updated 2026-10-04 (evening): the HOST's security fixes install themselves too, the hub has a fleet view and four
|
||||
OS alarms, and the tunnel status is true. Both demo boxes run controller 0.292.0 and host agent 0.141.1. Hub 0.131.1.
|
||||
New installs get golden 0.292.0 with agent 0.141.1 (vouched).**
|
||||
**Updated 2026-10-04 (late evening): the hub has a System page with every box's versions and the update buttons; Docker
|
||||
updates are built; a crashed box restarts by itself, at most twice an hour. Both demo boxes run host agent 0.142.0 and
|
||||
controller 0.292.0. Hub 0.132.0. Installer 1.30.0. New installs get golden 0.292.0 (re-made: live-restore on, Docker 29.8.2) with agent 0.142.0.**
|
||||
|
||||
## Today (2026-10-04, late evening): the System page, Docker updates, the crash restart
|
||||
|
||||
**Decisions I took myself (you may reverse each):**
|
||||
- The box knows a crash only as "it did not shut down cleanly" (the crash memory chip saved nothing). So a power cut
|
||||
also counts as a crash.
|
||||
- An "oops" (a kernel error the box survives) does not restart the box; you get a mail instead.
|
||||
- Your words win where the brief disagreed: the **3rd** crash within one hour leaves the box off. Restart after 10 s,
|
||||
re-arm after 24 h. All settings.
|
||||
- Only the box itself decides whether a Docker update is allowed: it checks your signature with a key file only root
|
||||
can change (not the agent's own settings, which the agent could change).
|
||||
- The System page uses the alarm limits for its colours (red = an alarm would fire).
|
||||
|
||||
**No decision needed from you today.**
|
||||
|
||||
**What I did:**
|
||||
- **The System tab** in the hub: per box the Proxmox, kernel (now and next boot), Debian and Docker versions, what is
|
||||
waiting, held packages, "restart needed", the crash guard and the last update run — with buttons for ring, on/off,
|
||||
"Approve now" and "Approve Docker set". The Hosts page shows Proxmox and kernel too.
|
||||
- **Docker updates:** "live-restore" is on in every box (no app restarted: 24 of 24 and 5 of 5 containers kept running).
|
||||
Both demo boxes moved to Docker 29.8.2, every app kept running. I approved that set with the new button (a 0-night
|
||||
test wait, then back to 2 nights). An undo signed by your key put demo-hp back one version and forward again, apps
|
||||
running throughout. A copied (replayed) signed job was refused.
|
||||
- **Crash restart (your 3 crashes on demo-hp):** crash 1 and 2 — back by itself in under a minute; crash 3 — it stayed
|
||||
off until you switched it on. You got the "guard tripped" mail. I re-armed it.
|
||||
- **Found:** the hub learns about a crash up to 15 minutes late (nothing lost). After a crash, the household can also get
|
||||
an "app stopped" mail besides "restarted after a crash" — whether to calm that is a later choice for you.
|
||||
- **New-install image re-made** with live-restore on and the approved Docker version, and approved in the hub with host
|
||||
agent 0.142.0. New boxes also get the crash guard (installer 1.30.0).
|
||||
- **Rows:** 5 closed, 1 opened-and-closed the same day, 4 opened. The list went from 334 to 333.
|
||||
|
||||
**Needs you later (nothing breaks if you wait):**
|
||||
- Installed boxes still get new root files only by hand (the long-standing gap; there are no other boxes today).
|
||||
- Kernel updates are still not built (a hung new kernel would stay — needs a fix first).
|
||||
|
||||
## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status
|
||||
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
2026/10/04 17:47:51 [WARN] host demo-hp-bb76ea crash guard: host_crash_guard_rearmed
|
||||
@@ -398,7 +398,7 @@ stopping line that lies.
|
||||
| **R-793** | Business & legal | P4 | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. **Watch:** never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the `-enterprise` image, and re-read on each major. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them.** | — | — | CC |
|
||||
| **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC |
|
||||
|
||||
## Process & tooling — 86 rows (P3 4, P4 82)
|
||||
## Process & tooling — 87 rows (P3 4, P4 83)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -489,6 +489,7 @@ stopping line that lies.
|
||||
| **R-817** | Process & tooling | P4 | **`09` decision 56 and R-745 disagree about what the controller self-update rolls back to.** Decision 56 (`documentation/architecture/09-update-architecture.md:619`) says a box keeps the image before the running one as *"the self-update's roll-back target"*; R-745 (closed, `CLOSED-ITEMS.md`) measured that *"the agent rolls back to the RUNNING image (what `/etc/felhom-controller-image` named), never to a previous one the controller hands it"*. Found 2026-10-03 while carrying R-745's rule to `CONTEXT.md`. | **VERIFY — filed 2026-10-03 (triage); owner: CC.** Read the agent's rollback path and correct whichever text is wrong. | — | — | CC |
|
||||
| **R-818** | Process & tooling | P4 | **Two changelogs cite register ids for other findings.** `hub/CHANGELOG.md:526-538` (hub v0.109.0, the Backup card) and its note at `:533` use **R-331** and **R-330** for a hub display fault and a nightly false app alarm; in the register R-330 and R-331 are the disk-health Phase 2 and Phase 3 rows. A reader following the id lands on the wrong finding. Found 2026-10-03 by the triage. | **READY — filed 2026-10-03 (triage); owner: CC.** Add a dated correction line under each changelog entry naming the right rows (the real ids are in `CLOSED-ITEMS.md`); do not renumber anything. | — | — | CC |
|
||||
| **R-819** | Process & tooling | P4 | **`scripts/check_stands.py` is red and runs in no runner.** Measured 2026-10-03 on `9e2786c` (before the triage): it convicts `where-felhom-stands.yaml` for citing R-273 and R-356, which are in neither `OPEN-ITEMS.md` nor anywhere it reads. After the triage it also convicts R-281, R-198 and R-201, because its rule 3 reads only `OPEN-ITEMS.md` and those rows are closed. It is in neither `repo_gates.py` nor CI, so nobody saw it — the R-29 shape. | **READY — filed 2026-10-03 (triage); owner: CC.** Let rule 3 accept an id in `CLOSED-ITEMS.md` (and check the stand's status agrees), fix the two dangling ids, then register it in `repo_gates.py` with a decoy. | — | — | CC |
|
||||
| **R-857** | Process & tooling | P4 | **Baking a golden twice under the SAME version leaves two stale facts.** 2026-10-04 (golden 0.292.0 re-baked with live-restore): (1) `golden_currency_gate.py` keeps reporting the FIRST bake directory's sha (`golden-0.292.0-2026-10-04`, d6cf8b33…) — it picks one of two directories with the same version, not the newest; (2) the publish replaces the package in place (pre-delete + upload), so until the operator re-vouches, the hub vouches a sha the registry no longer holds — a fresh install in that window fails its sha check (fail-closed; none ran; re-vouched 17:48). Fix direction: the gate prefers the newest bake directory (or refuses two for one version); the runbook says "re-vouch at once after a same-version re-bake". `documentation/tests/golden-0.292.0-2026-10-04-rebake/` | **READY — owner: CC** | — | — | CC |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
Re-vouch 2026-10-04 17:48 CEST by CC (main session) after the same-version re-bake, POST /configuration/artifacts (operator Basic auth, ClusterIP):
|
||||
agent_version=0.142.0 agent_sha256=7beb32224d6495e9561acfd3ad8a48393a799520f196080011cb27fceb6d1de6
|
||||
golden_version=0.292.0 golden_sha256=79a1dce3bd3c5a636a03c82be0f4ef969a1e0e9cbb139c5890b84c98fd69d43a (re-hashed by download in the main session: match)
|
||||
min_agent=0.131.0 -> 303 artifacts_set
|
||||
hub log: Artifact manifest set: agent=0.142.0 golden=0.292.0 min_agent="0.131.0" wrapper_sha=false
|
||||
Between the re-bake upload and this re-vouch the hub vouched the OLD sha (d6cf8b33...) for a package that no longer had it: a fresh install in that window would have failed its sha check (fail-closed). None ran. R-857.
|
||||
Reference in New Issue
Block a user