golden 0.292.0 re-vouched (re-bake: live-restore + approved Docker set) with agent 0.142.0; R-857 filed; STATUS; REPORT-os-docker-crash-2026-10-04
gates / gates (push) Successful in 32s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-04 17:49:30 +02:00
parent 208d21d18f
commit 21986d03b5
5 changed files with 129 additions and 4 deletions
+37 -3
View File
@@ -2,9 +2,43 @@
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
**Updated 2026-10-04 (evening): the HOST's security fixes install themselves too, the hub has a fleet view and four
OS alarms, and the tunnel status is true. Both demo boxes run controller 0.292.0 and host agent 0.141.1. Hub 0.131.1.
New installs get golden 0.292.0 with agent 0.141.1 (vouched).**
**Updated 2026-10-04 (late evening): the hub has a System page with every box's versions and the update buttons; Docker
updates are built; a crashed box restarts by itself, at most twice an hour. Both demo boxes run host agent 0.142.0 and
controller 0.292.0. Hub 0.132.0. Installer 1.30.0. New installs get golden 0.292.0 (re-made: live-restore on, Docker 29.8.2) with agent 0.142.0.**
## Today (2026-10-04, late evening): the System page, Docker updates, the crash restart
**Decisions I took myself (you may reverse each):**
- The box knows a crash only as "it did not shut down cleanly" (the crash memory chip saved nothing). So a power cut
also counts as a crash.
- An "oops" (a kernel error the box survives) does not restart the box; you get a mail instead.
- Your words win where the brief disagreed: the **3rd** crash within one hour leaves the box off. Restart after 10 s,
re-arm after 24 h. All settings.
- Only the box itself decides whether a Docker update is allowed: it checks your signature with a key file only root
can change (not the agent's own settings, which the agent could change).
- The System page uses the alarm limits for its colours (red = an alarm would fire).
**No decision needed from you today.**
**What I did:**
- **The System tab** in the hub: per box the Proxmox, kernel (now and next boot), Debian and Docker versions, what is
waiting, held packages, "restart needed", the crash guard and the last update run — with buttons for ring, on/off,
"Approve now" and "Approve Docker set". The Hosts page shows Proxmox and kernel too.
- **Docker updates:** "live-restore" is on in every box (no app restarted: 24 of 24 and 5 of 5 containers kept running).
Both demo boxes moved to Docker 29.8.2, every app kept running. I approved that set with the new button (a 0-night
test wait, then back to 2 nights). An undo signed by your key put demo-hp back one version and forward again, apps
running throughout. A copied (replayed) signed job was refused.
- **Crash restart (your 3 crashes on demo-hp):** crash 1 and 2 — back by itself in under a minute; crash 3 — it stayed
off until you switched it on. You got the "guard tripped" mail. I re-armed it.
- **Found:** the hub learns about a crash up to 15 minutes late (nothing lost). After a crash, the household can also get
an "app stopped" mail besides "restarted after a crash" — whether to calm that is a later choice for you.
- **New-install image re-made** with live-restore on and the approved Docker version, and approved in the hub with host
agent 0.142.0. New boxes also get the crash guard (installer 1.30.0).
- **Rows:** 5 closed, 1 opened-and-closed the same day, 4 opened. The list went from 334 to 333.
**Needs you later (nothing breaks if you wait):**
- Installed boxes still get new root files only by hand (the long-standing gap; there are no other boxes today).
- Kernel updates are still not built (a hung new kernel would stay — needs a fix first).
## Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status