Files
felhom.eu/documentation/audits/day-2026-10-08/design-R-314-279-177.md
T

78 lines
7.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# R-314, R-279, R-177 — the operator's door into a running controller: a one-page design (2026-10-08)
Baselines read: felhom-controller `a0370b4`, felhom-agent `b228b44`, felhom.eu `b2dce901`. Architecture:
`03-host-agent.md` §4 (signature only for destroying or overwriting the only copy of customer data),
`04-control-plane-authorization.md` §6 („routine work: no signing"), `07-backup-architecture.md` (decision 74, R-241/R-245).
**Status: design only. Nothing is built.**
## 1. What is still true (read in source today, not from the rows)
- **R-177 is no longer true — close it.** The row says fill-watch runs only at 03:30 and at start. Since controller
v0.297.0 (R-363, 2026-10-05) it also runs every 10 minutes: `cmd/controller/main.go:1579`
(`sched.Every("fill-watch-interval", …)`), `main.go:2478` (`fillWatchInterval = 10 * time.Minute`), pinned by
`cmd/controller/r363_fillwatch_interval_test.go`. Every run logs a positive line
(`internal/fillwatch/fillwatch.go:267`, „checked N filesystem(s) …"), which the operator reads through the hub's
existing controller-log pull. „Did the warning clear after the household freed space?" is now a ≤10-minute wait,
not a restart. Its general ask („run a named job now") is folded into option A below at no extra cost.
- **R-279 is still true.** The only way to start an off-site run on demand is the household's dashboard
(`internal/web/server.go:752` → `offbox_handlers.go:264`). A search of hub, controller report code and agent for any
run-now path finds only the nightly call (`main.go:1417`).
- **R-314 is half true now.** The CLI levers are unchanged (`main.go:91-93`, `229-270`) and run as a SECOND process —
the lost-update problem is written down at `main.go:208-219`. What changed: **decision 74** — on the pinned
off-site tier the deletion is the HUB's, 7 days after the box asks for it, and the operator can cancel it at the hub
(`hub/internal/web/server.go:685-704` → `offsitekeys/service.go:474`; a route, no button). The box follows that cancel
(`internal/backup/offbox_abandon.go:240`). **The gap that remains:** the box's own 14-day countdown
(`offbox_abandon.go:31`) runs first, and a hub cancel during it does nothing — `CancelOffsiteAbandon` only touches
pending rows (`hub/internal/store/offsite_keys.go:284-291`), and on day 14 the box asks again
(`offbox_abandon.go:266`). On a non-pinned target the box deletes by itself on day 14 (`offbox_abandon.go:287`) with
no hub window at all. So a household that telephones in the first 14 days still needs a shell.
## 2. The door already exists — it is the report reply, not the controller's web surface
The 2026-10-05 note looked at the controller's HTTP surface (the household's session). The running process already
takes operator requests another way: operator presses a button → the hub stores a pending row and bumps the box's
intent generation → the controller's wait channel wakes (`internal/report/waiter.go:131-138`) → an out-of-cycle report
→ the reply carries the request (`internal/report/pusher.go:43-50`; hub `internal/api/handler.go:600-610`) → the running
controller acts. Proven for log pulls (hub `internal/web/logbundle.go:43-67`, controller `internal/report/selftail.go`).
Latency: seconds; worst case the 15-minute cycle. The hub never connects into the box.
**Trust.** `03` §4 asks for a signature only to destroy or overwrite the only copy. None of these does: an off-site
run adds a snapshot (the pinned tier cannot delete, decision 69); a check only reads; stop/extend keep data. The hub can
already cancel the hub-phase deletion (decision 74), so „stop" gives a compromised hub no new power. **Not on the list:**
anything that deletes or starts a countdown, and clearing a restore hold (R-379 — it lets an app start on a possibly
broken database; it stays on the CLI).
## 3. Options
| | What | Cost | Risk |
|---|---|---|---|
| **A** | **Operator actions in the report reply.** Hub table `operator_actions(id, customer_id, action, arg, requested_at, done_at, outcome, message)`, buttons on the host page, reply field `operator_actions:[{id,action,arg}]` until a result arrives. Controller: a CLOSED switch in the running process; result in the next report `operator_action_results:[{id,outcome,message}]`; hub marks it done and saves a hub-minted operator event. | ~1 session, two repos, additive wire fields (wire-contract gate). | A controller that crashes mid-action gets it again — every listed action is safe to repeat (single-flight, refuse-when-nothing-running). |
| **B** | Signed job through the agent, which runs `docker exec … --abandon-stop` in the guest. | New signed-op class + an agent exec path; a signing ceremony for a non-destructive act (against `04` §6). | Still a second process: the lost-update at `main.go:208-219` stays. |
| **C** | A unix socket in the container served by the running controller; the CLI flags become its clients. | ~½ session, controller only. | Fixes the second-process problem but the telephone path is still a shell. Good later for R-379. |
**Pick: A.** It reuses a proven channel, needs no key, and the household-visible state changes inside the process
that owns it.
## 4. First slice and its red tests
- **Controller** `internal/report/opactions.go` (the `selftail.go` shape) + a closed map in `main.go`:
`offsite_backup_now` (the same four checks as `offboxRunHandler`, then `RunOffboxBackup` in a goroutine);
`abandon_stop` → `StopAbandon` (which also cancels a hub-phase request); `abandon_extend` (arg 1–30 days) →
`ExtendAbandon`; `run_job` (arg ∈ fill-watch, offsite-integrity, offsite-proof, disk-health-check) → a new
`Scheduler.RunNow(name)` that refuses unknown or running jobs. An unknown action answers `refused`, never silence.
- **Red tests** (each fails today — the reply field does not exist): (1) a reply with `abandon_stop` during a countdown
→ `AbandonStatus().Active` is false in the SAME manager AND `settings.json` on disk agrees; (2) the same id delivered
twice → `StopAbandon` runs once; (3) an unknown action → result `refused`, nothing called; (4) hub: a host-page POST
stores a row and bumps intent; the reply lists it until a result arrives, then not; a result naming another
customer's id is ignored.
- **Live proof** (scratch 9202 only, after the releases): `run_job fill-watch` and `offsite_backup_now`. Positive
observables: the hub's result event; control from another channel: the controller log pull (the „checked N" line)
and the off-site snapshot list on the box's backup page.
## 5. One question for the operator
**May the hub, with only your hub password and no signing key, ask a box to run an off-site backup now, run a named
check now, and stop or extend a household's deletion countdown?** My pick: yes (option A). It costs one session in the
controller and the hub. *If you do nothing:* nothing is built; a household that telephones in the first 14 days of a
countdown is served by a shell on its box, and an off-site run can only be started by the household.