Files
felhom.eu/documentation/audits/day-2026-10-08/design-R-30.md
T

5.6 KiB
Raw Blame History

R-30 — box presence from the live connection, not the report clock: a one-page design (2026-10-08)

Baselines read: felhom.eu b2dce901, felhom-controller a0370b4, felhom-agent b228b44. Architecture: 05-hub-architecture.md §4 (liveness / dead-man's-switch); 03-host-agent.md §4 (the host-delete guard's reason). Status: design only. Nothing is built.

1. The problem, with today's numbers

  • Measured 2026-07-21 (the row): a powered-off box stayed „healthy" on the hub until host_stale fired after the threshold („no report for 30m"). Today the wait is longer: the threshold is 45 minutes (live ConfigMap felhom-system/hub-config, alerting.stale_threshold: "45m", operator ruling A on R-549). The row's „30 min" is stale.
  • The agent reports every 900 s (hub internal/api/handler.go:653-655, defaultHostPollSeconds = 900; agent internal/config/config.go:849). „Online" on the host page is report age under the threshold (internal/web/hosts.go:31-44). Host delete refuses while „online" (hosts.go:935-937, and since R-599 it says when the refusal ends). RESET refuses while any host row exists (internal/web/customer_reset.go:106-110). So a forced teardown of a box that is already off waits up to 45 minutes after its last report.
  • Measured today (read-only, the ingress access log on DooPlex, 07:0x–07:50 UTC): the controller's wait channel (GET /api/v1/wait) completes every 241–243 s per box — three sources, 11–12 holds each, every hold 240.00x s. So a healthy box starts a new wait at most ~243 s after the previous one started.
  • Not measured: what the hub sees when a box loses power in the middle of a hold. Read in source: the hub writes a newline every 25 s into nginx (internal/api/wait.go:10-23) and ends the hold itself at 240 s, so the hold most likely ends normally and then no new wait arrives. That gap is the signal. Slice 1 measures it.

2. The three signals

Signal Cadence What it proves Blind spot
Agent host report 900 s the host and its agent slow; the alarm needs 45 min hysteresis (R-549)
Controller wait channel ~241 s the guest's controller is running and reaching the hub a crashed or updating controller looks like a dead box; per customer, not per host
WireGuard handshake age on ep0 ~2 min (protocol rekey; not measured here) the host's kernel, independent of both processes needs a new forced command on ep0 and a hub poller; ep0 is protected

3. Options

A. Keep the report clock. Nothing to build. A teardown waits ≤45 min; alarms stay right.

B. The hub records wait-channel presence; the host page shows it; the delete guard may use it. In handleWait (once per request — not in intent.Hub.Wait, which runs once per 25-s window, internal/intent/hub.go:77-102) record per customer: last wait start, holds open now. Presence = connected (a hold is open, or one started < 243 s + 90 s grace ago), not connected since T, or unknown (hub restarted < 333 s ago — in memory only, so it falls back to the report clock; the safe direction). Alerting is not changed. Cost: hub only, ~½ session. Wrong case: a crashed controller on a live host reads „not connected"; if the guard trusts it, the operator deletes a live host and its agent is locked out until re-enrolled — no household data is touched.

C. B + the WireGuard handshake from ep0 as host-level presence. The best signal (the host itself, a third channel). Cost: an ep0 change, a new key and forced command, a hub poller; ep0 work needs an operator task.

Pick: B, in two slices. Slice 1 is display only and needs no decision. Slice 2 (the guard) needs the operator's word. C only if slice 1's measurement shows the wait channel does not go quiet when a box dies.

4. First slice and its red test

  • Hub: intent.Hub gains MarkWaitStart / MarkWaitEnd / Presence(customerID, now) behind a clock seam; handleWait calls them; the host page shows „Box connection: connected now / last connected
  • Red tests (fail today — no such state): (1) a wait starts at t0 and ends at t0+240 s, and none follows → at t0+300 s presence is connected, at t0+334 s it is not connected since t0+240 s; (2) a new start at t0+242 s keeps it connected; (3) a new Hub (a hub restart) answers unknown, never not connected; (4) a render test per branch — the „not connected" line appears only when absent (the seam-built-but-never-wired trap).
  • Live measurement (after a hub release; scratch 9202 only): stop guest 9202, then read the time until the host page says „not connected". Control from a different channel: the ingress access log's last /api/v1/wait line from that box's address. Evidence copied off before 9202 is started again.
  • Slice 2 (after the operator's answer): when report-age says „online" but presence says „not connected" for longer than the grace, the delete refusal offers a tick box „I checked: the box is off"; the delete is logged and raises one operator event.

5. One question for the operator

When the hub has had no connection from a box for about six minutes, may „delete host" let you delete it at once, after you tick „I checked: the box is off"? My pick: yes (slice 2). It costs: if the box was really on (only its controller had crashed), its host agent is locked out until it is re-enrolled; no household data is touched. If you do nothing: slice 1 still shows „last connected", but deleting a powered-off host, and so a RESET, still waits up to 45 minutes after its last report.