9.1 KiB
RUNBOOK — Escrow ceremony (customer recovery-code R)
When a ceremony runs
- Onboarding (with the PBS tier): seals K + the identity secrets.
- Offsite enable (fork-4 rule): the restic repo password MUST enter a CURRENT R-blob before any
offsite backup runs — the controller stages the password to the agent at enable and the run-gate holds
pendinguntil the hub-verified auto-confirm. - Any new unrecoverable secret / suspected R exposure: re-ceremony (R cannot be retro-fitted — it is never retained).
- Stale-blob warning (v0.127.0, Scenario F): the controller re-checks the hub blob's password
hash on every report ACK even in the
escrowedstate. If the CURRENT blob does not cover the current offsite password (hash mismatch, or a hash-less superseding blob — e.g. a CLI ceremony run without the staged secret), the Távoli mentés card shows "A letétben lévő helyreállítási csomag nem fedi a jelenlegi távoli mentési jelszót" and offers a re-ceremony. Runs are NOT blocked (offsite copies still protect against non-total loss); the fix is one wizard run.
PRIMARY PATH — the controller wizard (customer-driveable, agent ≥ v0.88.0, controller ≥ v0.127.0)
The customer (or the operator in the customer's session) opens Távoli mentés →
"Helyreállítási kód létrehozása" (/backup/escrow) and clicks through:
- Preflight — the wizard checks: PBS storage id configured, DR tier applied,
agepresent, hub upload target configured, the FELHOM_ESCROW sudo grant listed (list-mode probe), staged secret (informational). Red rows block the start. - Warnings — the R-is-final copy; a re-ceremony additionally shows the supersede warning ("a korábbi kód érvényét veszti…").
- Re-auth — the dashboard password, typed again (rides the login rate-limiter).
- Run — the controller RE-STAGES the current offsite password first (when offsite is
configured; a staging failure ABORTS the start — a ceremony without the staged secret would
mint a hash-less blob), then triggers the agent's local-API job. The agent runs
sudo -n felhom-agent --config /etc/felhom-agent/agent.json --selftest=escrow-create --upload --output=json— ONE fixed argv, byte-matched by sudoers (spike §2.2). - Reveal — R is delivered by a ONE-SHOT claim: displayed exactly once in the browser,
Cache-Control: no-store, never rendered server-side, never logged, wiped agent-side the moment it is claimed. Unclaimed R expires after 10 minutes → the ceremony is VOID and a re-run supersedes it (nothing is lost — the blob is already on the hub; the old R simply never existed for it). - Typed-back — the customer re-types two random words of R (client-side check), ticks the "felírtam" confirmation, finishes.
Then the existing auto-confirm flips pending → escrowed hands-free within one report cycle
(hash match in the ACK) and wipes the staged copy — same as ever.
Supersede/void semantics: every new ceremony's blob REPLACES the hub blob; the previous R keeps opening only the previous blob (pre-existing offsite history). An unclaimed/void R requires nothing but a re-run. An agent restart mid-ceremony or pre-claim loses R safely (in-memory by design) — the wizard reports it honestly; re-run.
F1 threat model (operator ruling 2026-07-13 — carried verbatim into the product)
R is displayed in the customer's browser, which means it transits the Cloudflare tunnel ONCE at
the reveal moment (TLS end-to-end, no-store, one-shot). This is an accepted, documented risk
— the same trust class as the claim code and the login password, which already ride the same
channel: an adversary positioned to read this leg already owns the dashboard session and the box.
The agent→controller leg (local API, pinned TLS) never leaves the box. A LAN-direct delivery
variant (à la plex.direct) is PARKED, not forgotten — revisit only if the trust class of the
tunnel changes.
OPERATOR FALLBACK — the CLI (text mode, unchanged)
felhom-agent --selftest=escrow-create --upload
- On the host, over SSH, as before; output byte-identical to the historical text mode. R is printed once inside the write-it-down banner.
--storage <pbs-storage-id>only ifescrow.pbs_storage_idisn't configured.- ⚠ Staged-secret rule: when offsite is configured, ensure the staged secret is present (re-save the offsite config to re-push it) BEFORE a CLI run — a CLI ceremony without it uploads a hash-less blob, which trips the Scenario-F stale warning and can never auto-confirm.
--output=jsonexists for the wizard's job runner (machine mode: ONE JSON object with R on stdout, humans to stderr). Never run json mode by hand into a terminal you are screen-sharing.
Prerequisites (the wizard preflight checks these for you)
- Agent ≥ v0.88.0 (wizard endpoints +
--output=json); ≥ v0.79.0 forrestic_pw_sha256. - ⚠ No-PBS hosts: the ceremony CANNOT run — POLICY, not a gap. Identity-only escrow is formally PARKED (operator decision 2026-07-12, DR-tier-by-default decision 4); every install ships the full DR plumbing (host-install ≥ v1.15.0), activation is the hub's per-customer DR-tier flag (default ON), and offsite backup REQUIRES that flag. A box always reaches the ceremony THROUGH the tier: WG peer → PBS-DR descriptor applied → ceremony.
- Host packages:
age(installed by host-install since v1.15.0). - K gate (PBS hosts):
escrow.pbs_storage_idset and the key file present. - Sudo grant: the
FELHOM_ESCROWalias in/etc/sudoers.d/felhom-agent(ships with host-install ≥ v1.16.0 via the canonical sudoers fetch; on older installs re-ship the sudoers with the agent deploy). - The customer is present: R is displayed ONCE and must land in their hands.
R handling — the rules to say out loud to the customer
- R is displayed once; the system keeps no copy anywhere. Write it down / password manager / safe — NOT on the server it protects.
- A new ceremony's R SUPERSEDES the old one. Replace the stored code; the old blob remains valid only for pre-existing history.
- R + the hub blob = full recovery after total loss (see Recovery below). Losing R does not lose data while the box lives — but after a disaster, an offsite copy without R is permanently sealed. That is the zero-knowledge trade, stated honestly.
After the ceremony — hands off (auto-confirm)
Within one report cycle the hub serves the blob's presence + hash in the ACK; the controller matches
sha256(local repo password) and flips pending → escrowed on its own, wiping the agent-staged copy.
Verify remotely: the customer page / report shows escrow_state: "escrowed" with zero manual clicks;
the next scheduled offsite run proceeds and snapshots appear. Hash-mismatch warning instead → the blob
does not cover the CURRENT password (stale blob) → re-run the ceremony (wizard); do NOT manual-confirm
around it. (The manual confirm ENDPOINT exists solely as a deprecated fallback for pre-hash legacy
blobs; its button is gone from the card since v0.127.0.)
Recovery (the other half — operator-assisted, sketch)
Total loss → the customer's R + the hub blob:
felhom-agent --selftest=identity-consume --blob <blob> --keydest <0600 scratch> with R → recovers the
bundle (restic password, tokens; K restored via its own path on PBS hosts). Offsite data: hub re-issues
SFTP credentials (Re-issue button — access is regenerable; the ENCRYPTION key is what R guards) → inject
the recovered password on the fresh controller (POST /backup/offbox/inject-password, honored by the
secrets writer) → restore through the real endpoints. Live-proven end-to-end 2026-07-09 (recovered
password opened the production repo).
Do NOT
- Pipe the CLI ceremony's output to a log or run it unattended (R would land in it). The wizard's machine mode is exempt BY DESIGN: its pipe terminates in the agent's in-memory one-shot holder, nowhere else.
- Manual-confirm a hash-mismatch; skip the self-verify; retain or transmit R on Felhom systems.
- Treat the SFTP credential as escrow-worthy (it's regenerable access; only the ENCRYPTION secrets ride R).
- Run a CLI ceremony on an offsite-configured box without the staged secret (hash-less blob → stale warning; the wizard's re-stage-first makes this impossible from the UI).