REPORT: controller v0.246.0 - restore record, recovery-code readiness, live proof
gates / gates (push) Failing after 16s

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 10:59:02 +02:00
parent 0fe315b759
commit 29e2acb5f8
+80 -47
View File
@@ -1,63 +1,96 @@
# REPORT — controller v0.245.0: the household is asked for the recovery code # REPORT — controller v0.246.0: an interrupted restore is told; the recovery-code reminder waits
**2026-09-16.** R-543. Off-site backup is ON by default and does not RUN until the household creates **2026-09-17.** R-550 (operator ruling „fix") and R-546. Architecture read first:
its recovery code. **The pause is the design and is untouched here** — the escrow is zero-knowledge, `felhom.eu/documentation/architecture/07-backup-architecture.md` §3 (Lane 1) and
their code is the only key, and a run without one would write a copy nobody could ever open. What `08-alarm-ladder.md` §6.2.
was missing is that nothing asked them, while the page promised the very copy that had never run.
## Claims in the brief that turned out wrong — named first
1. **„Readiness = `escrow.pbs_storage_id` set."** Readiness is the agent's own preflight `ok` over
**five** blocking items (`pbs_storage_id`, `dr_tier`, `age_binary`, `hub_upload`, `sudo_grant`;
`felhom-agent/internal/localapi/escrow_ceremony.go` `handleEscrowPreflight`). `pbs_storage_id` was the
one item R-546's box happened to show. The controller now reads the combined `ok`, never a copy.
2. **„The failure shows a raw error"** — not in a browser. The escrow page already listed the checklist
and showed its start form only when the preflight was `ok`. The raw `-storage` stderr came from the
chaos-night harness calling `POST /api/escrow/start` directly. What a browser user DID meet: red
crosses, the agent's English detail in muted brackets, no start button and no word about waiting.
Both paths are fixed; the direct one is now refused server-side.
3. **„The backup tiers persist their records atomically"** (read from `07`, not re-checked) — **checked
and TRUE.** The tier run records live in `settings.json`, and `Settings.save()`
(`internal/settings/settings.go:727-752`) writes a `.tmp` and renames it, with `.bak` recovery on a
corrupt primary. The restore record follows the same shape through the backup package's own
`atomicWrite` (tmp + rename; like `save()`, no fsync).
## What shipped ## What shipped
**The reminder bar.** While the off-site tier is configured and its escrow state is not `escrowed`, **B.1 — the restore record survives a restart (R-550).** A design REVERSED by ruling, recorded in the
every authenticated page carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási file header, in `07` §3 and in `CONTEXT.md`: `restore-status.json` in `DataDir`, written at both ends of an
kódot." linking `/backup/escrow`. It is the **R-241 bar, second instance** — same session-cookie op (`internal/backup/restore_record.go`). A record still marked running at startup becomes a failed,
dismissal („Most nem"), back at the next visit, gone for good when escrowed. No second banner system `Interrupted` result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — kept per
was built. app until that app's next restore; `restore_interrupted` (warning, household) pushed once. Wired from
`main()` (`cmd/controller/restore_record_wiring.go`). Shown as a card on `/backups/restore` and in the
off-site wizard's outcome card. Notification cooldowns stay in memory.
*Where it hangs, and why it matters:* `executeTemplate` (server.go), the single render choke point, **B.2 — the reminder waits (R-546).** `internal/web/escrow_readiness.go`: the agent's preflight `ok`,
not the three `addRecoveryBanner` call sites. A per-handler helper would have covered the pages cached 60 s, probed only while paused. The bar is held back while not ready; `/backup/escrow` shows a
someone remembered — the seam-built-but-never-wired shape this repo has shipped four times. The waiting card that polls and reloads; `POST /api/escrow/start` refuses 409 before staging. Unknown
login and claim pages render through `ExecuteTemplate` directly and never pass through it; a session readiness keeps the bar (fail loud). Transitions logged at INFO.
check (`hasAdminSession`) additionally keeps it off the public guest share page.
**The tier-1 sentence renders by state.** v0.244.0 fixed the label (R-537) and put a new promise in **B.3 — the guide.** `VOLUNTEER-first-hour.md`: the recovery code is now step 7, after the first apps,
its place: „a távoli másolat … védi", printed from the app's shape alone. `driveFilesNoteFor` now and begins „Amikor a sárga sáv megjelenik (a beállítás után néhány perccel)".
takes `tier3State`'s own vocabulary — `active` → „védi"; `escrow_pending` → „**védené** — a távoli
mentés a helyreállítási kód létrehozásáig szünetel" + the route; no off-site and no second drive → **MinAgent: 0.131.0, unchanged** — the preflight has been served since agent 0.88.0.
„nincs másolat" + both ways out. One source of state, so the row and the sentence cannot disagree.
## Red-proofs — each seen failing, then passing ## Red-proofs — each seen failing, then passing
| fix | break | what failed | | test | break | failure seen |
|---|---|---| |---|---|---|
| the bar | delete the one `addEscrowBanner` line from `executeTemplate` | „/dashboard does not tell the household the off-site copy is PAUSED" — **and the same on /launcher** | | `TestRestoreRecord_InterruptedRestoreSurvivesRestart` | no persistence | `StartedAt:0001-01-01 … Last:<nil>` — the chaos-night value |
| the sentence | compute it from the app's shape, as v0.244.0 did | „the sentence claims the files ARE protected while the copy is paused", quoting the exact v0.244.0 wording | | `TestRestoreRecordAtStartup_RaisesInterruptedOnce` | helper skips loading | `raised no restore_interrupted: []` |
| `TestMainWiresRestoreRecord` | main() not wired | `does not call both … inert seam` |
| `TestR550_RestorePageShowsInterruptedRestore` | handler line removed | `the restore page does not show the interrupted restore` |
| `TestR546_NotReadyBoxHoldsTheBar` | readiness not consulted | `the bar urges a ceremony the box cannot run yet` |
| `TestR546_EscrowPageWaitsWhenNotReady` | flag never set | `does not tell the household the box is still preparing` |
| `TestR546_StartRefusedWhenNotReady` | gate removed | `= 200 … job_id escrow-1` |
Negative controls: an escrowed box is not nagged; an unconfigured box is not nagged; an app with no Controls: finished restore stays finished; the next restore of that app clears the notice (another app's
drive-side files gets no sentence in any state. The banner fixture asserts `OffboxConfigured()` does not); a ready box shows the bar; unknown readiness shows the bar; no interruption → no card.
itself, so it cannot pass by silently losing its precondition. Two of my own test mistakes, corrected: a `-storage` needle that matched the menu's `nav-group-storage`
id (narrowed to `requires -storage`); the three start-order tests updated to `preflight,…` — their
## Live validation (endpoint-level; no browser on DooPlex) load-bearing assertion (stage BEFORE start) unchanged.
Both boxes ran 0.245.0. Requests ran **inside** the guests (the controller answers on the container
address with the mandatory `Host` header; neither guest is reachable from DooPlex).
- **Escrowed (9201):** bar absent on /dashboard, /launcher, /backups/apps; the row reads „védi" (×2).
- **Paused (9202):** off-site configured through `POST /backup/offbox/config` → `escrow_state=pending`
with both secret files on disk. Bar present on /dashboard, /launcher, /backups/apps and /settings.
`POST /backup/offbox/run` → „A távoli mentés a kulcs letétbe helyezésére vár.", `last_run=None`,
`snapshot_count=None` — refused by the fork-4 gate, **not** by an unreachable target. A throwaway
class-A app's row rendered „…védené … szünetel" with „védi"=0. Dismissal sets a cookie with no
Max-Age and no Expires; same visit 0 hits, next visit 1 hit.
- **Teardown, three layers:** app removed with data and backups (0 containers, 0 drive folders); the
off-site target cleared and its three secrets `shred -u`'d (verified by re-reading); temp files
gone on guest and host; the hub provisioned nothing (9202 does not report).
## Gates ## Gates
`go build ./... && go vet ./... && go test ./...` — green. `controller_gates.py --fast` — 15/15 OK. `go build ./... && go vet ./... && go test ./...` — green, 28 packages. `controller_gates.py --fast` —
CI run 671 for the previous push (`ad398b60`) — **success**, matched by head_sha. all OK (`golden-notice` advisory; golden stays 0.245.0 under the waiver to 2026-09-27).
## Requires ## Live validation (endpoint-level; no browser on DooPlex)
Hub **v0.116.0** (unchanged from v0.244.0). MinAgent unchanged at **0.131.0**. No agent change, no **B.4(a) — PASS, on demo-hp guest 9201 with a throwaway `homebox`.** Restore accepted 08:51:51Z; 2 s in
hub change, no installer change, no golden. the status and the file on disk both said `running … homebox`; `docker kill felhom-controller`; the
agent's supervisor restarted it in 42 s (not by hand). After restart the status read
`ok:false … "A visszaállítás megszakadt …" interrupted:true`; `/backups/restore` showed the card
(negative control 0); the controller logged `restore_interrupted pushed OK (HTTP 200)`; the hub stores the
event under demo-hp. A second restore of homebox completed `ok:true` and the card was gone (0).
Evidence: `felhom.eu/documentation/audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt`.
**B.4(b) — NOT live-validated, and why (R-551).** No Tier-0 box is both paused and agent-connected:
9201 is escrowed (bar off by design; making it paused would be hand-set state on the standing demo box,
and a real ceremony would supersede its escrow — the hub keeps one `host_escrow` row per host); 9202 has
no local-API token. Proven by the three tests above through `ServeHTTP` with a fake agent; chaos night
measured live that the agent preflight is red ~17 min after a bind and turns green by itself.
**The ceremony was deliberately NOT run on demo-hp.**
## Teardown
- **Machine:** homebox removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up.
My first removal was refused 409 („still running — stop it first") — the product was right; stopped,
then removed. Password and scripts shredded in the guest (0 left).
- **Host:** nothing provisioned. 9201 stays on controller 0.246.0 (the release under validation).
- **Hub:** the `restore_interrupted` and `controller_restarted_by_agent` events stay as history.
## Observations
1. An interrupted-restore notice for an app that is then removed stays for ever. **FILED: R-552**
2. No Tier-0 box can exercise the paused + agent-connected escrow state. **FILED: R-551**
3. `PushEvent` is best-effort (3 attempts, 3 s apart); `restore_interrupted` inherits that. **NOT-A-FINDING: the page notice comes from the persisted file and does not depend on the event; the event-drop path is the existing, recorded behaviour of every controller event.**