# R-30 — box presence from the live connection, not the report clock: a one-page design (2026-10-08) Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`, felhom-agent `b228b44`. Architecture: `05-hub-architecture.md` §4 (liveness / dead-man's-switch); `03-host-agent.md` §4 (the host-delete guard's reason). **Status: design only. Nothing is built.** ## 1. The problem, with today's numbers - **Measured 2026-07-21 (the row):** a powered-off box stayed „healthy" on the hub until `host_stale` fired after the threshold („no report for 30m"). **Today the wait is longer:** the threshold is 45 minutes (live ConfigMap `felhom-system/hub-config`, `alerting.stale_threshold: "45m"`, operator ruling A on R-549). The row's „30 min" is stale. - The agent reports every 900 s (hub `internal/api/handler.go:653-655`, `defaultHostPollSeconds = 900`; agent `internal/config/config.go:849`). „Online" on the host page is report age under the threshold (`internal/web/hosts.go:31-44`). **Host delete refuses while „online"** (`hosts.go:935-937`, and since R-599 it says when the refusal ends). **RESET refuses while any host row exists** (`internal/web/customer_reset.go:106-110`). So a forced teardown of a box that is already off waits up to 45 minutes after its last report. - **Measured today (read-only, the ingress access log on DooPlex, 07:0x–07:50 UTC):** the controller's wait channel (`GET /api/v1/wait`) completes every **241–243 s** per box — three sources, 11–12 holds each, every hold 240.00x s. So a healthy box starts a new wait at most ~243 s after the previous one started. - **Not measured:** what the hub sees when a box loses power in the middle of a hold. Read in source: the hub writes a newline every 25 s into nginx (`internal/api/wait.go:10-23`) and ends the hold itself at 240 s, so the hold most likely ends normally and then **no new wait arrives**. That gap is the signal. Slice 1 measures it. ## 2. The three signals | Signal | Cadence | What it proves | Blind spot | |---|---|---|---| | Agent host report | 900 s | the host and its agent | slow; the alarm needs 45 min hysteresis (R-549) | | Controller wait channel | ~241 s | the guest's controller is running and reaching the hub | a crashed or updating controller looks like a dead box; per customer, not per host | | WireGuard handshake age on ep0 | ~2 min (protocol rekey; not measured here) | the host's kernel, independent of both processes | needs a new forced command on ep0 and a hub poller; ep0 is protected | ## 3. Options **A. Keep the report clock.** Nothing to build. A teardown waits ≤45 min; alarms stay right. **B. The hub records wait-channel presence; the host page shows it; the delete guard may use it.** In `handleWait` (once per request — not in `intent.Hub.Wait`, which runs once per 25-s window, `internal/intent/hub.go:77-102`) record per customer: last wait start, holds open now. Presence = **connected** (a hold is open, or one started < 243 s + 90 s grace ago), **not connected since T**, or **unknown** (hub restarted < 333 s ago — in memory only, so it falls back to the report clock; the safe direction). Alerting is not changed. Cost: hub only, ~½ session. Wrong case: a crashed controller on a live host reads „not connected"; if the guard trusts it, the operator deletes a live host and its agent is locked out until re-enrolled — no household data is touched. **C. B + the WireGuard handshake from ep0 as host-level presence.** The best signal (the host itself, a third channel). Cost: an ep0 change, a new key and forced command, a hub poller; ep0 work needs an operator task. **Pick: B, in two slices.** Slice 1 is display only and needs no decision. Slice 2 (the guard) needs the operator's word. C only if slice 1's measurement shows the wait channel does not go quiet when a box dies. ## 4. First slice and its red test - **Hub:** `intent.Hub` gains `MarkWaitStart / MarkWaitEnd / Presence(customerID, now)` behind a clock seam; `handleWait` calls them; the host page shows „Box connection: connected now / last connected