Files
felhom-controller/REPORT.md
T
2026-09-17 10:59:21 +02:00

97 lines
6.8 KiB
Markdown

# REPORT — controller v0.246.0: an interrupted restore is told; the recovery-code reminder waits
**2026-09-17.** R-550 (operator ruling „fix") and R-546. Architecture read first:
`felhom.eu/documentation/architecture/07-backup-architecture.md` §3 (Lane 1) and
`08-alarm-ladder.md` §6.2.
## Claims in the brief that turned out wrong — named first
1. **„Readiness = `escrow.pbs_storage_id` set."** Readiness is the agent's own preflight `ok` over
**five** blocking items (`pbs_storage_id`, `dr_tier`, `age_binary`, `hub_upload`, `sudo_grant`;
`felhom-agent/internal/localapi/escrow_ceremony.go` `handleEscrowPreflight`). `pbs_storage_id` was the
one item R-546's box happened to show. The controller now reads the combined `ok`, never a copy.
2. **„The failure shows a raw error"** — not in a browser. The escrow page already listed the checklist
and showed its start form only when the preflight was `ok`. The raw `-storage` stderr came from the
chaos-night harness calling `POST /api/escrow/start` directly. What a browser user DID meet: red
crosses, the agent's English detail in muted brackets, no start button and no word about waiting.
Both paths are fixed; the direct one is now refused server-side.
3. **„The backup tiers persist their records atomically"** (read from `07`, not re-checked) — **checked
and TRUE.** The tier run records live in `settings.json`, and `Settings.save()`
(`internal/settings/settings.go:727-752`) writes a `.tmp` and renames it, with `.bak` recovery on a
corrupt primary. The restore record follows the same shape through the backup package's own
`atomicWrite` (tmp + rename; like `save()`, no fsync).
## What shipped
**B.1 — the restore record survives a restart (R-550).** A design REVERSED by ruling, recorded in the
file header, in `07` §3 and in `CONTEXT.md`: `restore-status.json` in `DataDir`, written at both ends of an
op (`internal/backup/restore_record.go`). A record still marked running at startup becomes a failed,
`Interrupted` result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — kept per
app until that app's next restore; `restore_interrupted` (warning, household) pushed once. Wired from
`main()` (`cmd/controller/restore_record_wiring.go`). Shown as a card on `/backups/restore` and in the
off-site wizard's outcome card. Notification cooldowns stay in memory.
**B.2 — the reminder waits (R-546).** `internal/web/escrow_readiness.go`: the agent's preflight `ok`,
cached 60 s, probed only while paused. The bar is held back while not ready; `/backup/escrow` shows a
waiting card that polls and reloads; `POST /api/escrow/start` refuses 409 before staging. Unknown
readiness keeps the bar (fail loud). Transitions logged at INFO.
**B.3 — the guide.** `VOLUNTEER-first-hour.md`: the recovery code is now step 7, after the first apps,
and begins „Amikor a sárga sáv megjelenik (a beállítás után néhány perccel)".
**MinAgent: 0.131.0, unchanged** — the preflight has been served since agent 0.88.0.
## Red-proofs — each seen failing, then passing
| test | break | failure seen |
|---|---|---|
| `TestRestoreRecord_InterruptedRestoreSurvivesRestart` | no persistence | `StartedAt:0001-01-01 … Last:<nil>` — the chaos-night value |
| `TestRestoreRecordAtStartup_RaisesInterruptedOnce` | helper skips loading | `raised no restore_interrupted: []` |
| `TestMainWiresRestoreRecord` | main() not wired | `does not call both … inert seam` |
| `TestR550_RestorePageShowsInterruptedRestore` | handler line removed | `the restore page does not show the interrupted restore` |
| `TestR546_NotReadyBoxHoldsTheBar` | readiness not consulted | `the bar urges a ceremony the box cannot run yet` |
| `TestR546_EscrowPageWaitsWhenNotReady` | flag never set | `does not tell the household the box is still preparing` |
| `TestR546_StartRefusedWhenNotReady` | gate removed | `= 200 … job_id escrow-1` |
Controls: finished restore stays finished; the next restore of that app clears the notice (another app's
does not); a ready box shows the bar; unknown readiness shows the bar; no interruption → no card.
Two of my own test mistakes, corrected: a `-storage` needle that matched the menu's `nav-group-storage`
id (narrowed to `requires -storage`); the three start-order tests updated to `preflight,…` — their
load-bearing assertion (stage BEFORE start) unchanged.
## Gates
`go build ./... && go vet ./... && go test ./...` — green, 28 packages. `controller_gates.py --fast` —
all OK (`golden-notice` advisory; golden stays 0.245.0 under the waiver to 2026-09-27).
## Live validation (endpoint-level; no browser on DooPlex)
**B.4(a) — PASS, on demo-hp guest 9201 with a throwaway `homebox`.** Restore accepted 08:51:51Z; 2 s in
the status and the file on disk both said `running … homebox`; `docker kill felhom-controller`; the
agent's supervisor restarted it in 42 s (not by hand). After restart the status read
`ok:false … "A visszaállítás megszakadt …" interrupted:true`; `/backups/restore` showed the card
(negative control 0); the controller logged `restore_interrupted pushed OK (HTTP 200)`; the hub stores the
event under demo-hp. A second restore of homebox completed `ok:true` and the card was gone (0).
Evidence: `felhom.eu/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt`.
**B.4(b) — NOT live-validated, and why (R-551).** No Tier-0 box is both paused and agent-connected:
9201 is escrowed (bar off by design; making it paused would be hand-set state on the standing demo box,
and a real ceremony would supersede its escrow — the hub keeps one `host_escrow` row per host); 9202 has
no local-API token. Proven by the three tests above through `ServeHTTP` with a fake agent; chaos night
measured live that the agent preflight is red ~17 min after a bind and turns green by itself.
**The ceremony was deliberately NOT run on demo-hp.**
## Teardown
- **Machine:** homebox removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up.
My first removal was refused 409 („still running — stop it first") — the product was right; stopped,
then removed. Password and scripts shredded in the guest (0 left).
- **Host:** nothing provisioned. 9201 stays on controller 0.246.0 (the release under validation).
- **Hub:** the `restore_interrupted` and `controller_restarted_by_agent` events stay as history.
## Observations
1. An interrupted-restore notice for an app that is then removed stays for ever. **FILED: R-552**
2. No Tier-0 box can exercise the paused + agent-connected escrow state. **FILED: R-551**
3. `PushEvent` is best-effort (3 attempts, 3 s apart); `restore_interrupted` inherits that. **NOT-A-FINDING: the page notice comes from the persisted file and does not depend on the event; the event-drop path is the existing, recorded behaviour of every controller event.**