78121aa475
gates / gates (push) Successful in 3m57s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
73 lines
5.6 KiB
Markdown
73 lines
5.6 KiB
Markdown
# R-30 — box presence from the live connection, not the report clock: a one-page design (2026-10-08)
|
||
|
||
Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`, felhom-agent `b228b44`. Architecture:
|
||
`05-hub-architecture.md` §4 (liveness / dead-man's-switch); `03-host-agent.md` §4 (the host-delete guard's reason).
|
||
**Status: design only. Nothing is built.**
|
||
|
||
## 1. The problem, with today's numbers
|
||
|
||
- **Measured 2026-07-21 (the row):** a powered-off box stayed „healthy" on the hub until `host_stale` fired after the
|
||
threshold („no report for 30m"). **Today the wait is longer:** the threshold is 45 minutes (live ConfigMap
|
||
`felhom-system/hub-config`, `alerting.stale_threshold: "45m"`, operator ruling A on R-549). The row's „30 min" is stale.
|
||
- The agent reports every 900 s (hub `internal/api/handler.go:653-655`, `defaultHostPollSeconds = 900`; agent
|
||
`internal/config/config.go:849`). „Online" on the host page is report age under the threshold
|
||
(`internal/web/hosts.go:31-44`). **Host delete refuses while „online"** (`hosts.go:935-937`, and since R-599 it says when
|
||
the refusal ends). **RESET refuses while any host row exists** (`internal/web/customer_reset.go:106-110`). So a forced
|
||
teardown of a box that is already off waits up to 45 minutes after its last report.
|
||
- **Measured today (read-only, the ingress access log on DooPlex, 07:0x–07:50 UTC):** the controller's wait channel
|
||
(`GET /api/v1/wait`) completes every **241–243 s** per box — three sources, 11–12 holds each, every hold 240.00x s.
|
||
So a healthy box starts a new wait at most ~243 s after the previous one started.
|
||
- **Not measured:** what the hub sees when a box loses power in the middle of a hold. Read in source: the hub writes a
|
||
newline every 25 s into nginx (`internal/api/wait.go:10-23`) and ends the hold itself at 240 s, so the hold most likely
|
||
ends normally and then **no new wait arrives**. That gap is the signal. Slice 1 measures it.
|
||
|
||
## 2. The three signals
|
||
|
||
| Signal | Cadence | What it proves | Blind spot |
|
||
|---|---|---|---|
|
||
| Agent host report | 900 s | the host and its agent | slow; the alarm needs 45 min hysteresis (R-549) |
|
||
| Controller wait channel | ~241 s | the guest's controller is running and reaching the hub | a crashed or updating controller looks like a dead box; per customer, not per host |
|
||
| WireGuard handshake age on ep0 | ~2 min (protocol rekey; not measured here) | the host's kernel, independent of both processes | needs a new forced command on ep0 and a hub poller; ep0 is protected |
|
||
|
||
## 3. Options
|
||
|
||
**A. Keep the report clock.** Nothing to build. A teardown waits ≤45 min; alarms stay right.
|
||
|
||
**B. The hub records wait-channel presence; the host page shows it; the delete guard may use it.** In `handleWait`
|
||
(once per request — not in `intent.Hub.Wait`, which runs once per 25-s window, `internal/intent/hub.go:77-102`) record
|
||
per customer: last wait start, holds open now. Presence = **connected** (a hold is open, or one started < 243 s + 90 s
|
||
grace ago), **not connected since T**, or **unknown** (hub restarted < 333 s ago — in memory only, so it falls back to
|
||
the report clock; the safe direction). Alerting is not changed. Cost: hub only, ~½ session. Wrong case: a crashed
|
||
controller on a live host reads „not connected"; if the guard trusts it, the operator deletes a live host and its agent
|
||
is locked out until re-enrolled — no household data is touched.
|
||
|
||
**C. B + the WireGuard handshake from ep0 as host-level presence.** The best signal (the host itself, a third channel).
|
||
Cost: an ep0 change, a new key and forced command, a hub poller; ep0 work needs an operator task.
|
||
|
||
**Pick: B, in two slices.** Slice 1 is display only and needs no decision. Slice 2 (the guard) needs the operator's word.
|
||
C only if slice 1's measurement shows the wait channel does not go quiet when a box dies.
|
||
|
||
## 4. First slice and its red test
|
||
|
||
- **Hub:** `intent.Hub` gains `MarkWaitStart / MarkWaitEnd / Presence(customerID, now)` behind a clock seam;
|
||
`handleWait` calls them; the host page shows „Box connection: connected now / last connected <time>" from the host's
|
||
customer. Debug log on each change of presence.
|
||
- **Red tests** (fail today — no such state): (1) a wait starts at t0 and ends at t0+240 s, and none follows → at t0+300 s presence
|
||
is `connected`, at t0+334 s it is `not connected since t0+240 s`; (2) a new start at t0+242 s keeps it `connected`; (3) a new
|
||
`Hub` (a hub restart) answers `unknown`, never `not connected`; (4) a render test per branch — the „not connected"
|
||
line appears only when absent (the seam-built-but-never-wired trap).
|
||
- **Live measurement** (after a hub release; scratch 9202 only): stop guest 9202, then read the time until the host
|
||
page says „not connected". Control from a different channel: the ingress access log's last `/api/v1/wait` line from
|
||
that box's address. Evidence copied off before 9202 is started again.
|
||
- **Slice 2 (after the operator's answer):** when report-age says „online" but presence says „not connected" for longer
|
||
than the grace, the delete refusal offers a tick box „I checked: the box is off"; the delete is logged and raises one
|
||
operator event.
|
||
|
||
## 5. One question for the operator
|
||
|
||
**When the hub has had no connection from a box for about six minutes, may „delete host" let you delete it at once,
|
||
after you tick „I checked: the box is off"?** My pick: yes (slice 2). It costs: if the box was really on (only its
|
||
controller had crashed), its host agent is locked out until it is re-enrolled; no household data is touched. *If you do
|
||
nothing:* slice 1 still shows „last connected", but deleting a powered-off host, and so a RESET, still waits up to
|
||
45 minutes after its last report.
|