docs: 11 §5.7 System page, §5.8 Docker slow lane BUILT, §5.9 crash restart; 00/03/07/08; decisions 90-94 (CC unattended); runbooks docker-undo + crash-guard; register R-852 R-835 R-848 R-849 R-851 R-854 closed, R-853 R-855 R-856 opened, R-812 R-840 narrowed (334 -> 332); live evidence
gates / gates (push) Successful in 33s
gates / gates (push) Successful in 33s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,19 @@
|
||||
# The crash guard — read it, re-arm it
|
||||
|
||||
> **Design:** `architecture/11-os-updates.md` §5.9 (decision 88). **When:** the hub mailed `host_crash_guard_tripped`,
|
||||
> or the System page shows **TRIPPED** for a box.
|
||||
|
||||
A tripped guard means: the box stopped uncleanly twice within an hour (a crash, a power cut or a hard reset), so it set
|
||||
`kernel.panic = 0` — **the next crash leaves it off** until someone switches it on. It re-arms by itself after 24 h of
|
||||
normal running.
|
||||
|
||||
1. **Read why** (as root on the box's host): `felhom-crash-guard status` — `unclean_boots`, `tripped_at`,
|
||||
`tripped_reason`. Then the end of each crashed boot: `journalctl --list-boots` and `journalctl -b -1 -n 50` (a crashed
|
||||
boot ends with no shutdown lines). `journalctl -k -b -1 | grep -iE "panic|oops|BUG:"` for a kernel message.
|
||||
2. **Fix the cause first** if you found one (a bad kernel → boot the previous one; a power problem → the PSU, the cable).
|
||||
3. **Re-arm** (operator's choice): `felhom-crash-guard rearm` → `kernel.panic = 10` now; the history stays; a fresh
|
||||
60-minute window starts. The hub announces `host_crash_guard_rearmed` with the next report that carries the facts
|
||||
(up to ~15 min, R-853).
|
||||
4. **Check:** `sysctl kernel.panic` = 10; the System page shows "armed".
|
||||
|
||||
The numbers live in `/etc/felhom/crash-guard.conf` (LIMIT, WINDOW_MINUTES, PANIC_SECONDS, REARM_HOURS).
|
||||
@@ -0,0 +1,54 @@
|
||||
# Put a box's Docker engine back one set (OS updates, Docker slow lane)
|
||||
|
||||
> **When:** the hub mailed `os_update_health_failed` for the **docker** layer, or the System page shows a box whose
|
||||
> apps broke after a Docker engine step.
|
||||
> **Who:** the operator, or CC with the operator's word (CC may sign until the first paying customer, R-530 ruling).
|
||||
> **Owner design:** `architecture/11-os-updates.md` §5.8. **Proved** on demo-hp 2026-10-04 —
|
||||
> evidence `audits/os-docker-crash-2026-10-04/partB/undo/`.
|
||||
|
||||
The undo is an ordinary **signed `os_docker_step` with `"undo": true`** that names the previous engine set. The root
|
||||
wrapper re-verifies the signature itself (against the root-owned `/etc/felhom/operator-signers`), allows the downgrade
|
||||
only because the signed job says `undo`, and refuses unless `live-restore` is on — so the undo, like the step, restarts
|
||||
no app. Nothing is done by hand on the box.
|
||||
|
||||
## 1. Find the previous set
|
||||
|
||||
- The box's previous Docker report (System page → Docker release column, or the hub's `os_reports` for the box, layer
|
||||
`docker`): its `installed` list before the bad step. Or, on the host: `pct exec <vmid> -- grep -A3 "Start-Date"
|
||||
/var/log/apt/history.log | tail` — each line `name:amd64 (old, new)`.
|
||||
- All **six** names, each with its old version. Docker's repository keeps old versions (`11` C2), so no snapshot is
|
||||
needed.
|
||||
|
||||
## 2. Sign and queue it
|
||||
|
||||
```bash
|
||||
cd /mnt/5_hdd/felhom.eu/git/felhom-agent && go build -o /tmp/felhom-opsign ./cmd/felhom-opsign
|
||||
P='{"release_id":"undo-to-<engine>","undo":true,"packages":[
|
||||
{"name":"containerd.io","version":"<old>","origin":"Docker CE"},
|
||||
{"name":"docker-buildx-plugin","version":"<old>","origin":"Docker CE"},
|
||||
{"name":"docker-ce","version":"<old>","origin":"Docker CE"},
|
||||
{"name":"docker-ce-cli","version":"<old>","origin":"Docker CE"},
|
||||
{"name":"docker-ce-rootless-extras","version":"<old>","origin":"Docker CE"},
|
||||
{"name":"docker-compose-plugin","version":"<old>","origin":"Docker CE"}]}'
|
||||
KEY=$(sudo kubectl -n felhom-system get secret report-api -o jsonpath='{.data.REPORT_API_KEY}' | base64 -d)
|
||||
/tmp/felhom-opsign -op os_docker_step -host <host_id> -key-id felhom-op-1 \
|
||||
-key /mnt/5_hdd/felhom.eu/felhom-op-operational -params "$P" -ttl 45m \
|
||||
-upload http://<hub ClusterIP>:8080 -hub-key "$KEY"
|
||||
unset KEY
|
||||
```
|
||||
|
||||
The box takes the job at its next poll (up to 15 minutes), under the heavy-op gate (never beside a backup).
|
||||
|
||||
## 3. Check
|
||||
|
||||
- Agent journal: `signedjobs: AUTHORIZED signed op` → `os-apply: START … layer=docker:… authority=signed UNDO` →
|
||||
`signedjobs: signed op COMPLETED`. A `REFUSED: R3` means the signature, host, time window, nonce or package list did
|
||||
not match; `R15` means live-restore is off.
|
||||
- The hub: a docker-layer report, `outcome applied`, `healthy`, `undo: true`, the old engine.
|
||||
- **The same container ids before and after** (the health rule checks it; a changed id is `health_failed`).
|
||||
|
||||
## Afterwards
|
||||
|
||||
On a **ring-0** box the next night step installs the newest set again — switch the box's updates OFF on the System page
|
||||
first if that set is the bad one, and record it in the incident's register row. A **ring-1** box takes a set only by a
|
||||
signed job, so it stays on the undone set until someone signs the next one.
|
||||
Reference in New Issue
Block a user