Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
7.0 KiB
R-314, R-279, R-177 — the operator's door into a running controller: a one-page design (2026-10-08)
Baselines read: felhom-controller a0370b4, felhom-agent b228b44, felhom.eu b2dce901. Architecture:
03-host-agent.md §4 (signature only for destroying or overwriting the only copy of customer data),
04-control-plane-authorization.md §6 („routine work: no signing"), 07-backup-architecture.md (decision 74, R-241/R-245).
Status: design only. Nothing is built.
1. What is still true (read in source today, not from the rows)
- R-177 is no longer true — close it. The row says fill-watch runs only at 03:30 and at start. Since controller
v0.297.0 (R-363, 2026-10-05) it also runs every 10 minutes:
cmd/controller/main.go:1579(sched.Every("fill-watch-interval", …)),main.go:2478(fillWatchInterval = 10 * time.Minute), pinned bycmd/controller/r363_fillwatch_interval_test.go. Every run logs a positive line (internal/fillwatch/fillwatch.go:267, „checked N filesystem(s) …"), which the operator reads through the hub's existing controller-log pull. „Did the warning clear after the household freed space?" is now a ≤10-minute wait, not a restart. Its general ask („run a named job now") is folded into option A below at no extra cost. - R-279 is still true. The only way to start an off-site run on demand is the household's dashboard
(
internal/web/server.go:752→offbox_handlers.go:264). A search of hub, controller report code and agent for any run-now path finds only the nightly call (main.go:1417). - R-314 is half true now. The CLI levers are unchanged (
main.go:91-93,229-270) and run as a SECOND process — the lost-update problem is written down atmain.go:208-219. What changed: decision 74 — on the pinned off-site tier the deletion is the HUB's, 7 days after the box asks for it, and the operator can cancel it at the hub (hub/internal/web/server.go:685-704→offsitekeys/service.go:474; a route, no button). The box follows that cancel (internal/backup/offbox_abandon.go:240). The gap that remains: the box's own 14-day countdown (offbox_abandon.go:31) runs first, and a hub cancel during it does nothing —CancelOffsiteAbandononly touches pending rows (hub/internal/store/offsite_keys.go:284-291), and on day 14 the box asks again (offbox_abandon.go:266). On a non-pinned target the box deletes by itself on day 14 (offbox_abandon.go:287) with no hub window at all. So a household that telephones in the first 14 days still needs a shell.
2. The door already exists — it is the report reply, not the controller's web surface
The 2026-10-05 note looked at the controller's HTTP surface (the household's session). The running process already
takes operator requests another way: operator presses a button → the hub stores a pending row and bumps the box's
intent generation → the controller's wait channel wakes (internal/report/waiter.go:131-138) → an out-of-cycle report
→ the reply carries the request (internal/report/pusher.go:43-50; hub internal/api/handler.go:600-610) → the running
controller acts. Proven for log pulls (hub internal/web/logbundle.go:43-67, controller internal/report/selftail.go).
Latency: seconds; worst case the 15-minute cycle. The hub never connects into the box.
Trust. 03 §4 asks for a signature only to destroy or overwrite the only copy. None of these does: an off-site
run adds a snapshot (the pinned tier cannot delete, decision 69); a check only reads; stop/extend keep data. The hub can
already cancel the hub-phase deletion (decision 74), so „stop" gives a compromised hub no new power. Not on the list:
anything that deletes or starts a countdown, and clearing a restore hold (R-379 — it lets an app start on a possibly
broken database; it stays on the CLI).
3. Options
| What | Cost | Risk | |
|---|---|---|---|
| A | Operator actions in the report reply. Hub table operator_actions(id, customer_id, action, arg, requested_at, done_at, outcome, message), buttons on the host page, reply field operator_actions:[{id,action,arg}] until a result arrives. Controller: a CLOSED switch in the running process; result in the next report operator_action_results:[{id,outcome,message}]; hub marks it done and saves a hub-minted operator event. |
~1 session, two repos, additive wire fields (wire-contract gate). | A controller that crashes mid-action gets it again — every listed action is safe to repeat (single-flight, refuse-when-nothing-running). |
| B | Signed job through the agent, which runs docker exec … --abandon-stop in the guest. |
New signed-op class + an agent exec path; a signing ceremony for a non-destructive act (against 04 §6). |
Still a second process: the lost-update at main.go:208-219 stays. |
| C | A unix socket in the container served by the running controller; the CLI flags become its clients. | ~½ session, controller only. | Fixes the second-process problem but the telephone path is still a shell. Good later for R-379. |
Pick: A. It reuses a proven channel, needs no key, and the household-visible state changes inside the process that owns it.
4. First slice and its red tests
- Controller
internal/report/opactions.go(theselftail.goshape) + a closed map inmain.go:offsite_backup_now(the same four checks asoffboxRunHandler, thenRunOffboxBackupin a goroutine);abandon_stop→StopAbandon(which also cancels a hub-phase request);abandon_extend(arg 1–30 days) →ExtendAbandon;run_job(arg ∈ fill-watch, offsite-integrity, offsite-proof, disk-health-check) → a newScheduler.RunNow(name)that refuses unknown or running jobs. An unknown action answersrefused, never silence. - Red tests (each fails today — the reply field does not exist): (1) a reply with
abandon_stopduring a countdown →AbandonStatus().Activeis false in the SAME manager ANDsettings.jsonon disk agrees; (2) the same id delivered twice →StopAbandonruns once; (3) an unknown action → resultrefused, nothing called; (4) hub: a host-page POST stores a row and bumps intent; the reply lists it until a result arrives, then not; a result naming another customer's id is ignored. - Live proof (scratch 9202 only, after the releases):
run_job fill-watchandoffsite_backup_now. Positive observables: the hub's result event; control from another channel: the controller log pull (the „checked N" line) and the off-site snapshot list on the box's backup page.
5. One question for the operator
May the hub, with only your hub password and no signing key, ask a box to run an off-site backup now, run a named check now, and stop or extend a household's deletion countdown? My pick: yes (option A). It costs one session in the controller and the hub. If you do nothing: nothing is built; a household that telephones in the first 14 days of a countdown is served by a shell on its box, and an off-site run can only be started by the household.