R-543 closed: the household is asked for the recovery code (controller v0.245.0)
gates / gates (push) Successful in 21s

The tier-3 pause is the zero-knowledge escrow design and is untouched. What was
missing was the ASK, while the backup page promised the copy that had never run.

- VOLUNTEER-first-hour.md: a new step 6, right after the dashboard password and
  before the first app - what the code is, where, write it on PAPER, and that
  Felhom cannot get it back for them. Sections 6..12 renumbered to 7..13.
- day0-install.md A.2b: the operator step for a REBUILT box, which was missing.
  Acknowledged delete -> the hub re-issues by itself; otherwise ONE press of
  "Re-issue PBS credentials" (F-14 ruling 2026-07-13, hub/internal/web/pbsdr.go).
  This is the correction to last night's "zero presses" note.
- 07-backup-architecture.md: 6.1 records tier-3's paused state as a DESIGN, and
  2 records that the household is asked from first login.
- capability map: the first-hour row's last gap closed, with what it still does
  not claim (no volunteer has walked the ask from the written guide).
- register: R-543 CLOSED with the live measurements; R-545 filed (nothing
  un-configures an off-site target). R-511 was already closed yesterday.
- STATUS: the answered publish question removed (1.28.0 is live), readiness yes.
- evidence: red-proofs, the two-box live validation, teardown on three layers,
  and both of my own mistakes in this session.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 21:20:51 +02:00
parent 1acd693854
commit d124c77e17
9 changed files with 309 additions and 33 deletions
File diff suppressed because one or more lines are too long
@@ -77,6 +77,13 @@ So the honest statement of the property is:
> **The hub cannot open the escrow. The operator can reach a live box. Neither fact substitutes for
> the other, and R is what keeps the first true.**
**[DESIGN] Because R is the only key, the household is ASKED for it from the first login** (controller
v0.245.0, R-543). Off-site backup is enabled by default but does not RUN until the ceremony is done,
so the ask is not a nicety — it is the step that turns the default-on tier into an actual copy. The
volunteer guide asks for it immediately after the dashboard password and before the first app
(`runbooks/VOLUNTEER-first-hour.md` §6), and the product repeats the ask on every page until it is
done (§6.1).
---
## 3. The two lanes (D1)
@@ -299,6 +306,17 @@ never touches off-site snapshots (R-474). A removed app whose unit was kept is l
| **Plane-2** whole-guest, local | `local:` → `/var/lib/vz/dump` on the host | rootfs + `mp0 /var/lib/docker` + `mp1 /mnt/sys_drive` | **24 h** (`backup_cadence_seconds: 0`) | `local_backup_retention: 3` | **no** |
| **Plane-2** whole-guest, offsite | `felhom-pbs:` → ep0 datastore `felhom-offsite`, per-customer namespace, over WireGuard | same contents | **7 days** (`604800`) | server-side prune on ep0, `keep-last 2` at `03:30` — **and the box asks for none** (R-191) | **yes** (per-customer key) |
> **[FACT] Tier-3 has a fifth state the table above does not show: PAUSED (R-543, controller
> v0.245.0).** Off-site is ON by default from hub v0.116.0, and a run does not start until the
> household has performed the escrow ceremony — `tier3State` calls this `escrow_pending` and the page
> says „Kulcsletétre vár". **This is the design, not a defect:** the escrow is zero-knowledge (§2),
> the household's recovery code is the only key, and a run started without one would write a copy
> nobody could ever open. What was wrong until v0.245.0 is that **nothing asked the household for the
> code**, so a fresh box could sit paused indefinitely while its Tier-1 row promised that the off-site
> copy protected the app's files. Since v0.245.0 every dashboard page carries the reminder (the R-241
> bar, second instance) and the Tier-1 sentence renders by state — „védené … szünetel" while paused.
> Measured on a fresh box 2026-09-16: zero snapshots, and the page said the files were protected.
> **R-191 (2026-08-04) — this row was RIGHT and the configuration disagreed with it, weekly, for as long as R-89 has been in force.** The contract has not changed: offsite retention is ep0's, the box's token is write-only, and the box cannot delete its own history. What had not followed was the installer's `keep_last: 2` on the offsite tier, so every weekly run uploaded its snapshot successfully and then failed the whole JOB on a prune the token is refused — `whole_guest_backup_failed` in the operator's inbox about a backup that had already succeeded. Fixed in installer **1.25.0** (`keep_last: 0`) and on both live boxes; a gate now asserts it. **Verified before changing it:** ep0's two prune jobs have run every day since 2026-07-27, 18 tasks, all OK. A doc that states the contract does not enforce it — the gate does.