Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
5.6 KiB
R-30 — box presence from the live connection, not the report clock: a one-page design (2026-10-08)
Baselines read: felhom.eu b2dce901, felhom-controller a0370b4, felhom-agent b228b44. Architecture:
05-hub-architecture.md §4 (liveness / dead-man's-switch); 03-host-agent.md §4 (the host-delete guard's reason).
Status: design only. Nothing is built.
1. The problem, with today's numbers
- Measured 2026-07-21 (the row): a powered-off box stayed „healthy" on the hub until
host_stalefired after the threshold („no report for 30m"). Today the wait is longer: the threshold is 45 minutes (live ConfigMapfelhom-system/hub-config,alerting.stale_threshold: "45m", operator ruling A on R-549). The row's „30 min" is stale. - The agent reports every 900 s (hub
internal/api/handler.go:653-655,defaultHostPollSeconds = 900; agentinternal/config/config.go:849). „Online" on the host page is report age under the threshold (internal/web/hosts.go:31-44). Host delete refuses while „online" (hosts.go:935-937, and since R-599 it says when the refusal ends). RESET refuses while any host row exists (internal/web/customer_reset.go:106-110). So a forced teardown of a box that is already off waits up to 45 minutes after its last report. - Measured today (read-only, the ingress access log on DooPlex, 07:0x–07:50 UTC): the controller's wait channel
(
GET /api/v1/wait) completes every 241–243 s per box — three sources, 11–12 holds each, every hold 240.00x s. So a healthy box starts a new wait at most ~243 s after the previous one started. - Not measured: what the hub sees when a box loses power in the middle of a hold. Read in source: the hub writes a
newline every 25 s into nginx (
internal/api/wait.go:10-23) and ends the hold itself at 240 s, so the hold most likely ends normally and then no new wait arrives. That gap is the signal. Slice 1 measures it.
2. The three signals
| Signal | Cadence | What it proves | Blind spot |
|---|---|---|---|
| Agent host report | 900 s | the host and its agent | slow; the alarm needs 45 min hysteresis (R-549) |
| Controller wait channel | ~241 s | the guest's controller is running and reaching the hub | a crashed or updating controller looks like a dead box; per customer, not per host |
| WireGuard handshake age on ep0 | ~2 min (protocol rekey; not measured here) | the host's kernel, independent of both processes | needs a new forced command on ep0 and a hub poller; ep0 is protected |
3. Options
A. Keep the report clock. Nothing to build. A teardown waits ≤45 min; alarms stay right.
B. The hub records wait-channel presence; the host page shows it; the delete guard may use it. In handleWait
(once per request — not in intent.Hub.Wait, which runs once per 25-s window, internal/intent/hub.go:77-102) record
per customer: last wait start, holds open now. Presence = connected (a hold is open, or one started < 243 s + 90 s
grace ago), not connected since T, or unknown (hub restarted < 333 s ago — in memory only, so it falls back to
the report clock; the safe direction). Alerting is not changed. Cost: hub only, ~½ session. Wrong case: a crashed
controller on a live host reads „not connected"; if the guard trusts it, the operator deletes a live host and its agent
is locked out until re-enrolled — no household data is touched.
C. B + the WireGuard handshake from ep0 as host-level presence. The best signal (the host itself, a third channel). Cost: an ep0 change, a new key and forced command, a hub poller; ep0 work needs an operator task.
Pick: B, in two slices. Slice 1 is display only and needs no decision. Slice 2 (the guard) needs the operator's word. C only if slice 1's measurement shows the wait channel does not go quiet when a box dies.
4. First slice and its red test
- Hub:
intent.HubgainsMarkWaitStart / MarkWaitEnd / Presence(customerID, now)behind a clock seam;handleWaitcalls them; the host page shows „Box connection: connected now / last connected - Red tests (fail today — no such state): (1) a wait starts at t0 and ends at t0+240 s, and none follows → at t0+300 s presence
is
connected, at t0+334 s it isnot connected since t0+240 s; (2) a new start at t0+242 s keeps itconnected; (3) a newHub(a hub restart) answersunknown, nevernot connected; (4) a render test per branch — the „not connected" line appears only when absent (the seam-built-but-never-wired trap). - Live measurement (after a hub release; scratch 9202 only): stop guest 9202, then read the time until the host
page says „not connected". Control from a different channel: the ingress access log's last
/api/v1/waitline from that box's address. Evidence copied off before 9202 is started again. - Slice 2 (after the operator's answer): when report-age says „online" but presence says „not connected" for longer than the grace, the delete refusal offers a tick box „I checked: the box is off"; the delete is logged and raises one operator event.
5. One question for the operator
When the hub has had no connection from a box for about six minutes, may „delete host" let you delete it at once, after you tick „I checked: the box is off"? My pick: yes (slice 2). It costs: if the box was really on (only its controller had crashed), its host agent is locked out until it is re-enrolled; no household data is touched. If you do nothing: slice 1 still shows „last connected", but deleting a powered-off host, and so a RESET, still waits up to 45 minutes after its last report.