78121aa475
gates / gates (push) Successful in 3m57s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
78 lines
7.0 KiB
Markdown
78 lines
7.0 KiB
Markdown
# R-314, R-279, R-177 — the operator's door into a running controller: a one-page design (2026-10-08)
|
||
|
||
Baselines read: felhom-controller `a0370b4`, felhom-agent `b228b44`, felhom.eu `b2dce901`. Architecture:
|
||
`03-host-agent.md` §4 (signature only for destroying or overwriting the only copy of customer data),
|
||
`04-control-plane-authorization.md` §6 („routine work: no signing"), `07-backup-architecture.md` (decision 74, R-241/R-245).
|
||
**Status: design only. Nothing is built.**
|
||
|
||
## 1. What is still true (read in source today, not from the rows)
|
||
|
||
- **R-177 is no longer true — close it.** The row says fill-watch runs only at 03:30 and at start. Since controller
|
||
v0.297.0 (R-363, 2026-10-05) it also runs every 10 minutes: `cmd/controller/main.go:1579`
|
||
(`sched.Every("fill-watch-interval", …)`), `main.go:2478` (`fillWatchInterval = 10 * time.Minute`), pinned by
|
||
`cmd/controller/r363_fillwatch_interval_test.go`. Every run logs a positive line
|
||
(`internal/fillwatch/fillwatch.go:267`, „checked N filesystem(s) …"), which the operator reads through the hub's
|
||
existing controller-log pull. „Did the warning clear after the household freed space?" is now a ≤10-minute wait,
|
||
not a restart. Its general ask („run a named job now") is folded into option A below at no extra cost.
|
||
- **R-279 is still true.** The only way to start an off-site run on demand is the household's dashboard
|
||
(`internal/web/server.go:752` → `offbox_handlers.go:264`). A search of hub, controller report code and agent for any
|
||
run-now path finds only the nightly call (`main.go:1417`).
|
||
- **R-314 is half true now.** The CLI levers are unchanged (`main.go:91-93`, `229-270`) and run as a SECOND process —
|
||
the lost-update problem is written down at `main.go:208-219`. What changed: **decision 74** — on the pinned
|
||
off-site tier the deletion is the HUB's, 7 days after the box asks for it, and the operator can cancel it at the hub
|
||
(`hub/internal/web/server.go:685-704` → `offsitekeys/service.go:474`; a route, no button). The box follows that cancel
|
||
(`internal/backup/offbox_abandon.go:240`). **The gap that remains:** the box's own 14-day countdown
|
||
(`offbox_abandon.go:31`) runs first, and a hub cancel during it does nothing — `CancelOffsiteAbandon` only touches
|
||
pending rows (`hub/internal/store/offsite_keys.go:284-291`), and on day 14 the box asks again
|
||
(`offbox_abandon.go:266`). On a non-pinned target the box deletes by itself on day 14 (`offbox_abandon.go:287`) with
|
||
no hub window at all. So a household that telephones in the first 14 days still needs a shell.
|
||
|
||
## 2. The door already exists — it is the report reply, not the controller's web surface
|
||
|
||
The 2026-10-05 note looked at the controller's HTTP surface (the household's session). The running process already
|
||
takes operator requests another way: operator presses a button → the hub stores a pending row and bumps the box's
|
||
intent generation → the controller's wait channel wakes (`internal/report/waiter.go:131-138`) → an out-of-cycle report
|
||
→ the reply carries the request (`internal/report/pusher.go:43-50`; hub `internal/api/handler.go:600-610`) → the running
|
||
controller acts. Proven for log pulls (hub `internal/web/logbundle.go:43-67`, controller `internal/report/selftail.go`).
|
||
Latency: seconds; worst case the 15-minute cycle. The hub never connects into the box.
|
||
|
||
**Trust.** `03` §4 asks for a signature only to destroy or overwrite the only copy. None of these does: an off-site
|
||
run adds a snapshot (the pinned tier cannot delete, decision 69); a check only reads; stop/extend keep data. The hub can
|
||
already cancel the hub-phase deletion (decision 74), so „stop" gives a compromised hub no new power. **Not on the list:**
|
||
anything that deletes or starts a countdown, and clearing a restore hold (R-379 — it lets an app start on a possibly
|
||
broken database; it stays on the CLI).
|
||
|
||
## 3. Options
|
||
|
||
| | What | Cost | Risk |
|
||
|---|---|---|---|
|
||
| **A** | **Operator actions in the report reply.** Hub table `operator_actions(id, customer_id, action, arg, requested_at, done_at, outcome, message)`, buttons on the host page, reply field `operator_actions:[{id,action,arg}]` until a result arrives. Controller: a CLOSED switch in the running process; result in the next report `operator_action_results:[{id,outcome,message}]`; hub marks it done and saves a hub-minted operator event. | ~1 session, two repos, additive wire fields (wire-contract gate). | A controller that crashes mid-action gets it again — every listed action is safe to repeat (single-flight, refuse-when-nothing-running). |
|
||
| **B** | Signed job through the agent, which runs `docker exec … --abandon-stop` in the guest. | New signed-op class + an agent exec path; a signing ceremony for a non-destructive act (against `04` §6). | Still a second process: the lost-update at `main.go:208-219` stays. |
|
||
| **C** | A unix socket in the container served by the running controller; the CLI flags become its clients. | ~½ session, controller only. | Fixes the second-process problem but the telephone path is still a shell. Good later for R-379. |
|
||
|
||
**Pick: A.** It reuses a proven channel, needs no key, and the household-visible state changes inside the process
|
||
that owns it.
|
||
|
||
## 4. First slice and its red tests
|
||
|
||
- **Controller** `internal/report/opactions.go` (the `selftail.go` shape) + a closed map in `main.go`:
|
||
`offsite_backup_now` (the same four checks as `offboxRunHandler`, then `RunOffboxBackup` in a goroutine);
|
||
`abandon_stop` → `StopAbandon` (which also cancels a hub-phase request); `abandon_extend` (arg 1–30 days) →
|
||
`ExtendAbandon`; `run_job` (arg ∈ fill-watch, offsite-integrity, offsite-proof, disk-health-check) → a new
|
||
`Scheduler.RunNow(name)` that refuses unknown or running jobs. An unknown action answers `refused`, never silence.
|
||
- **Red tests** (each fails today — the reply field does not exist): (1) a reply with `abandon_stop` during a countdown
|
||
→ `AbandonStatus().Active` is false in the SAME manager AND `settings.json` on disk agrees; (2) the same id delivered
|
||
twice → `StopAbandon` runs once; (3) an unknown action → result `refused`, nothing called; (4) hub: a host-page POST
|
||
stores a row and bumps intent; the reply lists it until a result arrives, then not; a result naming another
|
||
customer's id is ignored.
|
||
- **Live proof** (scratch 9202 only, after the releases): `run_job fill-watch` and `offsite_backup_now`. Positive
|
||
observables: the hub's result event; control from another channel: the controller log pull (the „checked N" line)
|
||
and the off-site snapshot list on the box's backup page.
|
||
|
||
## 5. One question for the operator
|
||
|
||
**May the hub, with only your hub password and no signing key, ask a box to run an off-site backup now, run a named
|
||
check now, and stop or extend a household's deletion countdown?** My pick: yes (option A). It costs one session in the
|
||
controller and the hub. *If you do nothing:* nothing is built; a household that telephones in the first 14 days of a
|
||
countdown is served by a shell on its box, and an off-site run can only be started by the household.
|