## THE RECOVERY-CODE STEP FAILED ON A FRESH BOX — recorded verbatim, before any retry
## 2026-09-16, box `tester-1-022354` / guest 9201, controller 0.245.0, agent 0.131.0

Context: this is the step controller v0.245.0 added to `VOLUNTEER-first-hour.md` a few hours ago —
„A helyreállítási kód (~2 perc) — ezt ne hagyd ki", placed deliberately right after the dashboard
password and BEFORE the first app, because until it is done the off-site backup does not run.

The box was at exactly that point of the guide: installed, bound with no press, claimed, data drive
initialised, no apps yet. The escrow reminder bar was on every page, telling the household to do
precisely this.

  GET  /api/escrow/preflight  -> (empty response body)
  POST /api/escrow/start      -> 200  {"job_id":"escrow-1789590499667361664","phase":"running"}
  GET  /api/escrow/status     -> claimable:false, claimed:false, and:

      detail: "exit 2: exit status 2 | stderr: selftest=escrow-create requires -storage
               <pbs-storage-id> (or escrow.pbs_storage…"

  POST /api/escrow/claim      -> **409**
      „A folyamat jelenlegi állapotában a kód nem kérhető le."

So: the ceremony starts, the agent's `escrow-create` selftest refuses for a missing PBS storage id,
and the one-shot claim then correctly declines. The refusal is fail-closed and the wording is honest
— nothing pretended to succeed. What is wrong is that the household is TOLD to do this now, on every
page, and at this moment it cannot be done.

Nothing was retried before this file was written, so the state above is the state the box was in.

## WHY it refused — measured on both sides, not guessed

**The hub's own Backup & DR panel says it in plain words:**
    host enrolled (tester-1-022354)                                          done
    WG tunnel peer registered                                                done
    descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1)   **waiting**
    „ceremony possible once the descriptor is applied on the box"

**The box agrees:**
    pvesm status  ->  only `local` (dir) and `local-lvm` (lvmthin). **No PBS storage exists yet.**
    /etc/pve/storage.cfg  ->  no `pbs:` entry
    /etc/felhom-agent/agent.json  ->  top-level keys are
        [authz, backup, deployment_mode, hub, lan_resolver, local_api, log_level, oob, privileged,
         proxmox, storage, wg_tunnel]
      — there is **no `escrow` section at all**, so `escrow.pbs_storage_id` is unset, which is
      precisely what the agent's selftest complained about.

**The controller, meanwhile, already has the off-site target:**
    offbox present=True, enabled=True, host=u629488-sub4.your-storagebox.de, escrow_state=**pending**

So the chain is: the hub provisioned the DR descriptor automatically at 20:19 (no press), the
controller already knows its off-site destination, but the AGENT has not yet applied the descriptor
on the box — and the escrow ceremony depends on that. The hub documents the dependency in the very
panel that shows it as „waiting".

## The question this does NOT yet answer, and how it is being measured
Whether the box applies the descriptor **by itself**, and how long that takes. That decides
everything about severity: a few minutes of convergence makes the new guide step slightly too early
in the journey; never converging without an operator press makes it a broken promise on every fresh
box. A watcher is now polling the box for `pvesm` gaining a PBS storage and the agent config gaining
`escrow.pbs_storage_id`, and the ceremony will be retried when it does. Nothing was pressed, and the
box is being left to do it alone.
