docs: 11 §5.7 System page, §5.8 Docker slow lane BUILT, §5.9 crash restart; 00/03/07/08; decisions 90-94 (CC unattended); runbooks docker-undo + crash-guard; register R-852 R-835 R-848 R-849 R-851 R-854 closed, R-853 R-855 R-856 opened, R-812 R-840 narrowed (334 -> 332); live evidence
gates / gates (push) Successful in 33s
gates / gates (push) Successful in 33s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.141.1, hub v0.131.1 | **PARTIAL — the GUEST and HOST Debian fast lanes are PROVEN-LIVE (2026-10-04), with the fleet view and four operator alarms; Docker and the kernel are MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/` — ring 0 on both demo hosts (demo-felhom 108 Debian host packages, healthy, 0 Proxmox-origin); a 605-package host release approved (TEST wait, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the one host version it lacked; the by-hand host undo proved (`runbooks/os-updates-host-undo.md`); `tunnel_down` / `tunnel_recovered` live. Design `architecture/11-os-updates.md` §8.1–§8.3 | **No automatic undo** (guest: last night's backup, decision 81; host: the by-hand runbook); existing boxes need the wrapper + sudoers by hand (R-840); Docker and kernel lanes not built (R-812, R-835, R-836); a host panic is not restarted (R-851) |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.142.0, hub v0.132.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes need the wrapper, the trust files and the guard by hand (R-840); the kernel lane (R-836); facts reach the hub late after a boot (R-853) |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
||||
|
||||
|
||||
@@ -47,7 +47,13 @@ Owns:
|
||||
1. **Proxmox lifecycle** — create/start/stop/destroy guests, snapshots, storage allocation. Via a scoped Proxmox API token (the **`FelhomAgent` operator role** — `proxmox-platform.md` §3.6, validated Phase 3 B3) for everything the API covers; raw host ops only where unavoidable.
|
||||
2. **Storage management** — attach/classify targets, reconcile the storage manifest, mount USB-by-UUID, present mounts into guests.
|
||||
3. **Backup/restore orchestration** — vzdump to the tiers, PBS, snapshot management, and the **self-restore-test**.
|
||||
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841. **[FACT, FIXED agent v0.141.0 / controller v0.292.0, R-841 CLOSED]** The agent now reads the guest's `cloudflared` container through the existing `pct exec [0-9]* -- docker inspect -f *` sudoers line: state, exit code and the Docker health status of a check the controller adds (`cloudflared tunnel --metrics localhost:20241 ready` → cloudflared's own `/ready`, 200 only with a connection). Three states: `running` (healthy), `not_running` (stopped, absent, or running but NOT connected), `unknown` (could not ask, or the check is still starting) — `unknown` never alarms. No new sudoers line.
|
||||
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841. **[FACT, FIXED agent v0.141.0 / controller v0.292.0, R-841 CLOSED]** The agent now reads the guest's `cloudflared` container through the existing `pct exec [0-9]* -- docker inspect -f *` sudoers line: state, exit code and the Docker health status of a check the controller adds (`cloudflared tunnel --metrics localhost:20241 ready` → cloudflared's own `/ready`, 200 only with a connection). Three states: `running` (healthy), `not_running` (stopped, absent, or running but NOT connected), `unknown` (could not ask, or the check is still starting) — `unknown` never alarms. No new sudoers line. **[FACT, agent v0.142.0, R-852]** The host report gains `system`: the Proxmox version and kernel from the
|
||||
Proxmox API, plus the OS wrapper's read-only `facts` (host Debian, next-boot kernel, held packages, taint, the crash
|
||||
guard; guest Debian, Docker engine, containerd, live-restore), at most every 10 min. **Still no new sudoers line** — the
|
||||
facts, `live-restore-on` and the Docker layer all ride `FELHOM_OSAPPLY`. New ROOT-owned files instead (installer 1.30.0;
|
||||
by hand on installed boxes, R-840): `/etc/felhom/os-trust.json`, `/etc/felhom/operator-signers` (the wrapper verifies
|
||||
Docker authority against them, never against the agent-writable config — `09` decision 93), and the crash guard
|
||||
(`/usr/local/sbin/felhom-crash-guard`, its units, `/etc/felhom/crash-guard.conf`).
|
||||
5. **Provisioning** — provision a guest **by restoring the golden base image** (§9), deploy the controller into it, hand it its bootstrap config; also **build and refresh the golden base image** itself.
|
||||
6. **Hub control loop** — poll for desired state + signed jobs, reconcile, execute, report, heartbeat.
|
||||
7. **Local API** — the per-guest authorization gate the controller calls.
|
||||
|
||||
@@ -369,7 +369,10 @@ backup or a restore-test; at most once per 20 h. The backup minutes old is the g
|
||||
R-837). A failed or missed backup → no OS leg that night. **[FACT, 2026-10-04 — agent v0.141.1, `11` §8.2]** On an
|
||||
appliance, the HOST step follows under the same gate: after a healthy guest step only (a failed or unhealthy guest
|
||||
step skips it), Debian-origin fixes only, never a kernel, boot or firmware package, never a reboot. Measured: both
|
||||
steps with nothing to install, 23–32 s; a 108-package host pass, 70 s.
|
||||
steps with nothing to install, 23–32 s; a 108-package host pass, 70 s. **[FACT, 2026-10-04 — agent v0.142.0, `11` §5.8]** On a RING-0 box a third step
|
||||
follows, still under the gate: the guest's Docker engine set (live-restore turned on first, once, by reload; every
|
||||
container must keep its id). Ring 1 never takes it in the night leg. Measured: the three steps with nothing to install,
|
||||
38–51 s; a six-package Docker step, ~50 s.
|
||||
|
||||
**[FACT]** The three nightly legs derive from **one** customer-settable window start W at fixed
|
||||
offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered
|
||||
|
||||
@@ -312,7 +312,8 @@ message, not a wider cooldown.
|
||||
|
||||
## 6.3 Box alarms outside the app ladder: the tunnel and OS updates [DESIGN, hub v0.131.0, 2026-10-04]
|
||||
|
||||
These are **operator-only** (the household can act on none of them; `operatorOnlyEvents`, pinned by
|
||||
These are **operator-only** (the household can act on none of them — except `host_restarted_after_crash`, the household's
|
||||
one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
|
||||
`TestOSUpdateEvents_OperatorOnlyExceptApplied`). They go through the same dispatcher and severity contract (§6.1):
|
||||
`info` is recorded and never mailed; `warning` and `error` are mailed.
|
||||
|
||||
@@ -323,6 +324,9 @@ These are **operator-only** (the household can act on none of them; `operatorOnl
|
||||
| `os_reboot_needed` | warning | the host has needed a reboot for **14 days** (from the FIRST scanned report that said so) | a scanned pass that finds nothing (agent ≥ 0.141.1 scans the host every pass) | `TestAlarm_RebootNeeded`, `TestRebootNeeded_ClearedByAScannedPass` |
|
||||
| `os_ring0_stalled` | error | ring 0 approved nothing for **7 days** in a layer while it has pending FAST-lane updates (a pending kernel does not count) | a new release | `TestAlarm_Ring0Stalled` |
|
||||
| `os_not_covered` | warning | a ring-1 box has had fast-lane packages no approved release names for **14 days** | the packages are covered or gone | `TestAlarm_NotCovered` |
|
||||
| `host_crash_restart` | warning | the box's crash guard reports a NEW unclean boot (a crash, a power cut or a hard reset; hub v0.132.0) | — (one per boot) | `api/crash_test.go` |
|
||||
| `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` |
|
||||
| `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` |
|
||||
|
||||
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
|
||||
breaks a `not_running` run.
|
||||
|
||||
@@ -778,6 +778,28 @@ its length, and both fixes cost something the household would notice — operato
|
||||
83. **Next: the tunnel status (R-841), the host fast lane (`11` §8 step 3), and other OS-update improvements.**
|
||||
*Operator ruling 2026-10-04 ~12:20.*
|
||||
|
||||
### 2026-10-04 (evening) — decided by CC unattended — operator may reverse (System page / Docker / crash-restart brief)
|
||||
|
||||
90. **How does a box tell a crash boot from a clean one?** Options: (a) `pstore` — measured on demo-hp: `efi_pstore` is on,
|
||||
yet a real panic saved NOTHING; (b) the previous boot's journal ends without shutdown lines — readable only after
|
||||
the journal is up, and slow; (c) a marker written by an ExecStop at every orderly shutdown. **Chosen (c):** cheap,
|
||||
early, measured to work; cost: a power cut and a hard reset count as a crash too (they cannot be told apart on these
|
||||
boxes). `11` §5.9.
|
||||
91. **`kernel.panic_on_oops`?** Options: set it (an oops becomes a restart, which may loop on a bad driver) or leave it
|
||||
0 and REPORT an oops (taint bit D) to the operator. **Chosen: leave 0, report** — a box with an oops usually keeps
|
||||
serving the household; the hub mails `host_kernel_oops`. `11` §5.9.
|
||||
92. **The guard numbers.** The brief said both "at 3 it sets panic=0 so the next crash leaves the box off" (= the 4th)
|
||||
and "if it crashes 3 times within one hour, it stays off" (= the 3rd, the operator's own page). **Chosen: the
|
||||
operator's words** — the 3rd unclean stop within 60 minutes leaves the box off (the guard trips at the 2nd crash
|
||||
boot); `kernel.panic` 10 s; re-arm after 24 h of normal running. All four are `/etc/felhom/crash-guard.conf`.
|
||||
93. **Who may authorize a Docker step on the box itself?** Options: the agent's word (its config is agent-writable);
|
||||
the wrapper verifying the signed job against ROOT-owned files. **Chosen: the wrapper verifies** — `ssh-keygen -Y
|
||||
verify` against `/etc/felhom/operator-signers`, bound to `/etc/felhom/os-trust.json` `host_id`, a root-owned nonce
|
||||
record; an unsigned ring-0 step only with that file's `ring0_slow_lane: true`, set by hand on the demo boxes only.
|
||||
Cost: two more root files per box (the installer writes them; R-840 for installed boxes). `11` §5.8.
|
||||
94. **The System page colours** use the alarm thresholds themselves (red = an `08` alarm would fire; amber = worth a
|
||||
look; `unknown` = amber with its reason). No second set of numbers. `11` §5.7.
|
||||
|
||||
### 2026-10-04 (~15:17) — three operator rulings (recorded before the work)
|
||||
|
||||
87. **Docker `live-restore` is ON for every box** (`11` §5.8, option A). It is turned on once — by the golden and by a
|
||||
|
||||
@@ -332,11 +332,31 @@ must never overlap a backup, a restore-test or a self-update.~~
|
||||
- **Operator:** a hub event for each update and each failure; a fleet view showing each box's OS release,
|
||||
how far behind it is, and whether it needs a reboot. A box that is more than N days behind the newest
|
||||
approved release raises an alarm (`08`).
|
||||
- **[FACT, hub v0.132.0] The System page** (`/system`, R-852, decision 89): one row per box — ring and switch with
|
||||
buttons, the tunnel, host Proxmox / running and next-boot kernel / Debian / release / pending / not covered / held /
|
||||
reboot needed / `kernel.panic` / oops / crash restarts / the guard, guest Debian / release / pending / restart needed,
|
||||
Docker engine / containerd / live-restore / release, the last leg — and above it the releases, what ring 0 runs, "Approve
|
||||
now" and "Approve Docker set". Colours are the alarm thresholds (decision 94). Hosts shows Proxmox / kernel too.
|
||||
- **Household:** one line on the timeline in both languages, informal voice: what was updated and
|
||||
whether the box restarted. Telling households in advance that the box may restart at night is a
|
||||
**promise to users**. That is the operator's decision when the slow lane is built.
|
||||
|
||||
### 5.8 The Docker engine slow lane — DESIGN (2026-10-04, nothing built) `[PROPOSAL]`
|
||||
### 5.8 The Docker engine slow lane — BUILT 2026-10-04 (agent v0.142.0, hub v0.132.0) `[FACT]`
|
||||
|
||||
**As built** (evidence `audits/os-docker-crash-2026-10-04/partB/`; decisions 87, 93):
|
||||
- **live-restore ON**: the golden bakes it (`build-golden.sh` 3.1.0, fail-closed assertion); an installed box gets it once
|
||||
by the wrapper's `live-restore-on` (merge into daemon.json + `systemctl reload docker`). Measured: the same container ids
|
||||
after (9202 6/6 by hand — R10 refuses a scratch guest by design —, demo-hp 24/24, demo-felhom 5/5).
|
||||
- **The step** runs in the night leg after a healthy guest and host step, **ring 0 only**, `select pending-docker`,
|
||||
allowed by the wrapper only with the root-owned `ring0_slow_lane` mark. **Ring 1 and every undo** only through a signed
|
||||
`os_docker_step` the wrapper re-verifies itself (decision 93). Measured: 29.7.x → 29.8.2 on both demo boxes, every id
|
||||
kept; a signed undo to 29.7.2 on demo-hp (41 s, ids kept) and back; demo-felhom as ring 1 by a signed job; a replayed
|
||||
job refused by the agent (nonce).
|
||||
- **Approval**: the hub never approves a Docker set automatically; the System page's button works after every ring-0 box
|
||||
ran the set in 2 healthy night Docker steps (`OS_DOCKER_APPROVE_NIGHTS` TEST override, logged). An approval nudges no
|
||||
box. Undo: `runbooks/os-updates-docker-undo.md`.
|
||||
|
||||
**The design as written before the build:**
|
||||
|
||||
Built from C5 (measured) and R-835. **The one decision it needs is in STATUS: `live-restore` on, fleet-wide.**
|
||||
|
||||
@@ -360,6 +380,25 @@ Built from C5 (measured) and R-835. **The one decision it needs is in STATUS: `l
|
||||
- **Not covered here:** the golden's own engine (baked weekly; a new golden carries the approved set), and BYO hosts
|
||||
(the guest is ours on both, so the lane applies there too).
|
||||
|
||||
### 5.9 A crashed host restarts, with a limit — BUILT 2026-10-04 (agent v0.142.0, installer 1.30.0) `[FACT]`
|
||||
|
||||
Decision 88 (R-851): *"Yes, but maybe not indefinitely."* Evidence `audits/os-docker-crash-2026-10-04/partC/`.
|
||||
- **Measured first (demo-hp, the operator's word before each crash):** `kernel.panic = 10` + `echo c >
|
||||
/proc/sysrq-trigger` → the box restarted by itself in 54 s, on the same kernel (the saved default), the agent up 18 s
|
||||
after the boot. `kernel.panic` set by `sysctl -w` is gone after the restart (back to 0) — it must be set at every boot.
|
||||
- **The crash signal** (decision 90): `efi_pstore` is on, yet a real panic saved NOTHING; the journal and `last` show
|
||||
only "no shutdown". The guard uses a **clean-stop marker** (the unit's ExecStop at every orderly shutdown); a boot
|
||||
without it followed a crash, a power cut or a hard reset — counted alike.
|
||||
- **The guard** (`felhom-crash-guard`, early boot unit + hourly re-arm timer): armed → `kernel.panic = 10`; the 2nd
|
||||
unclean boot within 60 minutes TRIPS it (`kernel.panic = 0`), so **the 3rd crash within the hour leaves the box off**
|
||||
(decision 92 — the operator's words); re-arms after 24 h of normal running or `felhom-crash-guard rearm`. A crash before
|
||||
the unit runs (very early boot) leaves the box off: the safe side. `panic_on_oops` stays 0; an oops is reported
|
||||
(decision 91).
|
||||
- **Telling people:** each unclean boot = an operator mail (`host_crash_restart`) + the household's line; the trip = an
|
||||
alarm (`host_crash_guard_tripped`); re-arm and oops announced once (`08` §6.3). The System page shows the guard.
|
||||
- **Measured live:** crash 1 → back in 54 s; crash 2 → back in 53 s, guard tripped; crash 3 → **stayed off** until the
|
||||
operator switched it on; the hub mailed the trip (15 min after the boot — R-853); re-armed by hand.
|
||||
|
||||
---
|
||||
|
||||
## 6. Risks and edge cases
|
||||
@@ -432,7 +471,7 @@ Each step returns to the operator for go or no-go.
|
||||
3. **Host Debian, fast lane** (no kernel, no Proxmox packages). **BUILT 2026-10-04** — agent v0.141.1, hub v0.131.1;
|
||||
§8.2.
|
||||
4. **Fleet view and alarms** (§5.7). **BUILT 2026-10-04** — hub v0.131.0/v0.131.1; §8.3.
|
||||
5. **Slow lane: Docker engine.**
|
||||
5. **Slow lane: Docker engine.** **BUILT 2026-10-04** — agent v0.142.0, hub v0.132.0; §5.8.
|
||||
6. **Slow lane: host kernel and Proxmox packages, with the reboot.**
|
||||
7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user