Files
felhom.eu/documentation/audits/day-2026-10-08/design-R-30.md
T

73 lines
5.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# R-30 — box presence from the live connection, not the report clock: a one-page design (2026-10-08)
Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`, felhom-agent `b228b44`. Architecture:
`05-hub-architecture.md` §4 (liveness / dead-man's-switch); `03-host-agent.md` §4 (the host-delete guard's reason).
**Status: design only. Nothing is built.**
## 1. The problem, with today's numbers
- **Measured 2026-07-21 (the row):** a powered-off box stayed „healthy" on the hub until `host_stale` fired after the
threshold („no report for 30m"). **Today the wait is longer:** the threshold is 45 minutes (live ConfigMap
`felhom-system/hub-config`, `alerting.stale_threshold: "45m"`, operator ruling A on R-549). The row's „30 min" is stale.
- The agent reports every 900 s (hub `internal/api/handler.go:653-655`, `defaultHostPollSeconds = 900`; agent
`internal/config/config.go:849`). „Online" on the host page is report age under the threshold
(`internal/web/hosts.go:31-44`). **Host delete refuses while „online"** (`hosts.go:935-937`, and since R-599 it says when
the refusal ends). **RESET refuses while any host row exists** (`internal/web/customer_reset.go:106-110`). So a forced
teardown of a box that is already off waits up to 45 minutes after its last report.
- **Measured today (read-only, the ingress access log on DooPlex, 07:0x–07:50 UTC):** the controller's wait channel
(`GET /api/v1/wait`) completes every **241–243 s** per box — three sources, 11–12 holds each, every hold 240.00x s.
So a healthy box starts a new wait at most ~243 s after the previous one started.
- **Not measured:** what the hub sees when a box loses power in the middle of a hold. Read in source: the hub writes a
newline every 25 s into nginx (`internal/api/wait.go:10-23`) and ends the hold itself at 240 s, so the hold most likely
ends normally and then **no new wait arrives**. That gap is the signal. Slice 1 measures it.
## 2. The three signals
| Signal | Cadence | What it proves | Blind spot |
|---|---|---|---|
| Agent host report | 900 s | the host and its agent | slow; the alarm needs 45 min hysteresis (R-549) |
| Controller wait channel | ~241 s | the guest's controller is running and reaching the hub | a crashed or updating controller looks like a dead box; per customer, not per host |
| WireGuard handshake age on ep0 | ~2 min (protocol rekey; not measured here) | the host's kernel, independent of both processes | needs a new forced command on ep0 and a hub poller; ep0 is protected |
## 3. Options
**A. Keep the report clock.** Nothing to build. A teardown waits ≤45 min; alarms stay right.
**B. The hub records wait-channel presence; the host page shows it; the delete guard may use it.** In `handleWait`
(once per request — not in `intent.Hub.Wait`, which runs once per 25-s window, `internal/intent/hub.go:77-102`) record
per customer: last wait start, holds open now. Presence = **connected** (a hold is open, or one started < 243 s + 90 s
grace ago), **not connected since T**, or **unknown** (hub restarted < 333 s ago — in memory only, so it falls back to
the report clock; the safe direction). Alerting is not changed. Cost: hub only, ~½ session. Wrong case: a crashed
controller on a live host reads „not connected"; if the guard trusts it, the operator deletes a live host and its agent
is locked out until re-enrolled — no household data is touched.
**C. B + the WireGuard handshake from ep0 as host-level presence.** The best signal (the host itself, a third channel).
Cost: an ep0 change, a new key and forced command, a hub poller; ep0 work needs an operator task.
**Pick: B, in two slices.** Slice 1 is display only and needs no decision. Slice 2 (the guard) needs the operator's word.
C only if slice 1's measurement shows the wait channel does not go quiet when a box dies.
## 4. First slice and its red test
- **Hub:** `intent.Hub` gains `MarkWaitStart / MarkWaitEnd / Presence(customerID, now)` behind a clock seam;
`handleWait` calls them; the host page shows „Box connection: connected now / last connected <time>" from the host's
customer. Debug log on each change of presence.
- **Red tests** (fail today — no such state): (1) a wait starts at t0 and ends at t0+240 s, and none follows → at t0+300 s presence
is `connected`, at t0+334 s it is `not connected since t0+240 s`; (2) a new start at t0+242 s keeps it `connected`; (3) a new
`Hub` (a hub restart) answers `unknown`, never `not connected`; (4) a render test per branch — the „not connected"
line appears only when absent (the seam-built-but-never-wired trap).
- **Live measurement** (after a hub release; scratch 9202 only): stop guest 9202, then read the time until the host
page says „not connected". Control from a different channel: the ingress access log's last `/api/v1/wait` line from
that box's address. Evidence copied off before 9202 is started again.
- **Slice 2 (after the operator's answer):** when report-age says „online" but presence says „not connected" for longer
than the grace, the delete refusal offers a tick box „I checked: the box is off"; the delete is logged and raises one
operator event.
## 5. One question for the operator
**When the hub has had no connection from a box for about six minutes, may „delete host" let you delete it at once,
after you tick „I checked: the box is off"?** My pick: yes (slice 2). It costs: if the box was really on (only its
controller had crashed), its host agent is locked out until it is re-enrolled; no household data is touched. *If you do
nothing:* slice 1 still shows „last connected", but deleting a powered-off host, and so a RESET, still waits up to
45 minutes after its last report.