From 228dac4c06a4a229c1368706946e71082a923931 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 16 Jun 2026 14:18:39 +0200 Subject: [PATCH] =?UTF-8?q?docs(audit):=20SPIKE=20=E2=80=94=20Phase=202=20?= =?UTF-8?q?DR=20recipe=20+=20storage=20diagnosis=20(report-only)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Part 0 (live): flash apps on 9201 were down due to an operator pct reboot at 10:26 UTC + a boot-ordering race — dockerd auto-starts unless-stopped flash apps ~18s before the agent re-binds felhom-flash, so the create-time bind mkdir fails (permission denied) and RestartCount=0 never retries. Drive healthy, data intact, no USB drop, durable-id fine, drive-gate uninvolved. v0.70.0 self-restart RULED OUT (container restart, not a guest reboot; +38min after exits). Fix: restarted the 7 apps via the controller (drive present) — all Up. Flagged the intermediary mount app-start race as an architectural gap. Parts 1-3 (cited): characterized escrow (K + identity under recovery code R, fingerprint-gated, hub zero-knowledge) + PBS whole-CT contents (rootfs/secrets in, external drives out) + capstone DONE vs PENDING (agent-side recovery orchestration not wired, syncer.go:92). Defined the secret-free DR recipe (guest sizing + drive durable-id/role inventory + PVE storage + app bindings + PBS coords), sourced from facts the agent/controller already hold, landing in the reserved WireDesiredState.storage_manifest placeholder. Field-by-field boundary proof + a no-secrets test spec. Fork list + recommendation: spec/emit/store the recipe now, defer re-enrollment auth to slice 10D, never touch the escrow/PBS secret path. No code changes, no version bump. Co-Authored-By: Claude Opus 4.8 (1M context) --- .../audits/SPIKE-dr-recipe-2026-06-16.md | 188 ++++++++++++++++++ 1 file changed, 188 insertions(+) create mode 100644 documentation/audits/SPIKE-dr-recipe-2026-06-16.md diff --git a/documentation/audits/SPIKE-dr-recipe-2026-06-16.md b/documentation/audits/SPIKE-dr-recipe-2026-06-16.md new file mode 100644 index 0000000..c01c500 --- /dev/null +++ b/documentation/audits/SPIKE-dr-recipe-2026-06-16.md @@ -0,0 +1,188 @@ +# SPIKE — Phase 2 DR recipe: characterize existing DR, define the secret-free recipe, lock the boundary (+ storage diagnosis) + +**Date:** 2026-06-16 +**Type:** Report-only spike. **No code changes, no version bump.** Every claim cited against pushed source at `file:line`. Part 0 is a **live** diagnosis on guest 9201 / `felhom-pve` (one non-destructive fix applied — restarting already-present apps — and noted). +**Grounding:** `SPIKE-infra-backup-2026-06-15.md` line 24 — the revival is "a slim, **secret-free** DR recipe (customer → apps → drive/durable-id intents → guest sizing) that complements PBS's ingredients. Secrets are never in it; they are recovered from the PBS whole-CT snapshot, never regenerated." + +--- + +## One-screen summary + +- **PART 0 (storage diagnosis):** the flash-backed apps did **not** crash and felhom-flash did **not** drop. **Root cause: an operator `pct reboot` (PVE `vzreboot`) of guest 9201 at 10:26 UTC + a boot-ordering race.** On guest boot, dockerd auto-starts the `unless-stopped` flash apps **~18 s before** the agent re-binds felhom-flash into the guest; docker fails to create the not-yet-present bind source (`mkdir /mnt/felhom-drives/felhom-flash/userdata: permission denied`), and because that is a **create-time** mount failure (`RestartCount=0`) the restart policy never retries — the apps stay `Exited` forever even after the bind lands. The drive is healthy, mounted, bound (host + guest), data intact; no USB drop, no durable-id failure, the drive-gate was uninvolved. **The v0.70.0 self-restart is RULED OUT** (it restarts the felhom-controller *Docker container* via `os.Exit(0)`, never the LXC guest, and ran 11:04–11:17 UTC — ~38 min *after* the exits). Fix applied: started the 7 apps via the controller (drive present now) — all back `Up`. **Architectural gap flagged:** the intermediary mount is C1-immune (guest boots) but does not prevent this app-start race. +- **PART 1 (existing DR):** **DONE** — PBS whole-CT escrow + consume (`internal/escrow/*`, fingerprint-gated, real-data drilled), identity-escrow (`{tunnel_token, pbs_token}` age-wrapped under the same recovery code R), tunnel re-establishment, and the hub's **zero-knowledge** recovery-mode/re-enroll orchestration (`hub/internal/api/dr.go`). **PENDING** — the agent-side recovery orchestration that *consumes* the directive (the agent carries `restore_directive` forward-compat but explicitly does **not** act on it — `internal/desired/syncer.go:92`), and the secret-free recipe itself. +- **PART 2 (recipe):** the gap is the **non-secret scaffolding PBS does not capture** — guest sizing, drive inventory (durable-id → role → mount → enroll/decommission intent), PVE storage defs, app inventory + per-app storage bindings. The agent **already holds every fact** (`StorageTarget{durable_id, role, mount_path, total/used/avail}` at `report.go:86`; guest specs; PBS coords) — it can emit the recipe. The hub→agent wire **already reserves** `storage_manifest`/`backup_policy`/`pbs_namespace` placeholders (`report.go:271-273`). **Boundary:** every recipe field is an identifier/intent/size — no key, password, token, or hash. Secrets stay in PBS + escrow. +- **PART 3:** fork list + recommendation: emit the recipe **agent-side as an additive host-report section** (it owns the storage truth), store it **plaintext** on the hub next to the host-report, version the wire shape, and **defer** the re-enrollment auth (recovery-mode consumption) to the existing slice-10D track — the recipe complements escrow+PBS, it does not touch them. + +--- + +# PART 0 — Storage diagnosis: felhom-flash plugged in, yet flash-backed apps exited + +## What I observed (live, 12:00–12:10 UTC) + +Seven flash-backed apps `Exited`: audiobookshelf, calibre-web, immich-server, jellyfin, komga, radarr, romm. Exit codes **143** (audiobookshelf/immich-server/komga) and **128** (calibre-web/jellyfin/radarr/romm). felhom-flash physically connected. + +**Timezone note (this was load-bearing):** docker reports UTC (`…Z`); the host journal is `+02:00` (CEST). All times below normalized to **UTC** unless marked "local". + +## End-to-end walk (not "is the device present") + +1. **The drive is healthy and fully bound — host side.** `lsblk`: felhom-flash = `/dev/sdb1`, ext4, `LABEL=…`, `UUID=81a26531-62d8-408d-812f-a178b1d35310`, mounted at `/mnt/felhom-flash`. The intermediary chain is intact: + - stable parent: `/dev/mapper/pve-root on /mnt/felhom-drives` (on rootfs, always present), + - per-drive: `/dev/sdb1 on /mnt/felhom-drives/felhom-flash`. +2. **Bound through to the guest — guest side.** Inside 9201, `/proc/mounts` shows `/dev/sdb1 /mnt/felhom-drives/felhom-flash ext4 rw,relatime`, and `ls /mnt/felhom-drives/felhom-flash` → `appdata backups media userdata`. **The flash data is reachable in the guest right now.** +3. **The agent never lost it.** The reconcile loop logs every ~20 s, continuously: `reconcile: enrolled drive bound under shared parent (live, no reboot) … where=/mnt/felhom-flash guest_path=/mnt/felhom-drives/felhom-flash durable_id=uuid:81a26531-…` — including across the incident window. (`internal/reconcile` + `internal/localapi` guest-bind; durable-id matched.) +4. **No kernel-level USB event.** `dmesg -T` and `journalctl -k --since 10:20 --until 10:35` (UTC window) show **zero** USB disconnect / re-enumeration / `sd`/`ext4 I/O error` lines. **The task's "transiently dropped + re-enumerated under a new node" hypothesis is NOT supported** — the kernel never saw the device leave, and the agent re-bound it by the same `uuid:81a26531` throughout. Durable-id tracking did **not** fail. + +## What actually happened (root cause, with evidence) + +5. **Guest 9201 was rebooted by the operator at 10:26 UTC.** `uptime -s` = boot 10:26 (12:26 local); `/proc/uptime` ≈ 1.7 h. Host journal at 12:26:08 local: `pvedaemon[…]: end task UPID:demo-felhom:…:vzreboot:9201:root@pam: OK`. This is a **PVE-level `pct reboot`** of the LXC guest, operator-initiated. +6. **The "exits" are the reboot shutdown, not a fault.** romm's own log: `[2026-06-16 12:26:02] Stopping nginx … Stopping gunicorn … Worker … was sent SIGTERM!`; jellyfin: `[12:25:59] Disposing CoreAppHost`. Exit 143 = 128+SIGTERM; 128 = the image's PID1 exit on SIGTERM. The FinishedAt timestamps (10:25:58–10:26:02 UTC) are the guest-shutdown moment. +7. **All seven are `restart: unless-stopped`** — same policy as the DB/redis containers (`romm-db`, `immich-postgres`, …), which **did** come back (`Up ~1.7 h`). So policy is not the differentiator. +8. **THE SMOKING GUN — a create-time bind-mount failure at boot:** + ``` + romm State.Error = "error while creating mount source path + '/mnt/felhom-drives/felhom-flash/userdata/roms': + mkdir /mnt/felhom-drives/felhom-flash/userdata: permission denied" RestartCount=0 + jellyfin State.Error = "...mkdir /mnt/felhom-drives/felhom-flash/userdata: permission denied" RestartCount=0 + ``` + On boot, dockerd tries to restart the flash apps. Their binds point under `/mnt/felhom-drives/felhom-flash/…`, but the **real felhom-flash is not bound into the guest yet** — the agent re-binds it ~18 s after boot (host journal: `guest-attach: drive bound under shared parent (normalized to one bind, live) … where=/mnt/felhom-flash … 12:26:26` local = 10:26:26 UTC, vs boot 10:26:08). At that instant `/mnt/felhom-drives/felhom-flash` is the **empty, `nobody`-owned** stable-parent placeholder (the parent is `/dev/mapper/pve-root`, `drwxr-xr-x nobody nogroup` under the unprivileged-LXC idmap), so the guest-root docker daemon's attempt to `mkdir` the missing bind source is **permission-denied**. +9. **`RestartCount=0` → the restart policy never retries.** A *create-time* mount failure is not an exit-after-run, so `unless-stopped` does not apply — the container is left `Exited` permanently even though the bind lands 18 s later. The DB containers have no such bind (named volumes on rootfs) → they start immediately and stay up. That is the whole asymmetry. + +## Ruling out the v0.70.0 self-restart (explicitly, as requested) + +- The self-restart shipped this session is `os.Exit(0)` of the **felhom-controller Docker container** (`controller/internal/api/selfrestart.go`); the container is `restart: unless-stopped`, so Docker restarts *that one container*. It **never** issues a guest reboot. +- The incident was a **PVE `vzreboot` of the LXC guest** (host journal, step 5) — a different layer entirely. +- Timing: the exits were 10:25:58–10:26:02 UTC; my self-restart testing ran 11:04–11:17 UTC (controller `StartedAt` 11:17:51 UTC, "Up 46 min" at 12:04). The exits **precede** my testing by ~38 min. (The task's cited "~11:54–11:57" window matches neither the exits nor any controller restart I made; the controller's only live restart is 11:17:51 UTC.) **Not implicated.** + +## Fix applied (non-destructive, and it validates the diagnosis) + +The drive is present and the data intact, so the only thing wrong is that the apps never started. I started them through the **controller's own pipeline** (`POST /api/stacks/{name}/start`, the same path the UI uses) — romm first as a canary (`Up`, accessing `/romm/library` on flash), then the rest. Result: `audiobookshelf, calibre-web, jellyfin, radarr` healthy; `immich-server, komga` health-starting (normal); romm up. **That they start cleanly now confirms the root cause was the boot-ordering race, not any drive fault.** No remount/rebind was needed (the mount was already correct); no destructive action taken. + +## Architectural gap (flag — out of this spike's build scope, worth a follow-up) + +The intermediary-mount model (`SPIKE-intermediary-mount-2026-06-15.md`) is **C1-immune** (the guest boots clean because the bind *source* is the always-present stable parent). But it does **not** prevent an **app-start race**: the per-drive bind is (re)asserted by the agent's reconcile tick **seconds after** the guest's dockerd has already tried (and permanently failed, `RestartCount=0`) to auto-start drive-backed apps. Any operator `pct reboot`, host reboot, or guest auto-restart reproduces this. Candidate fixes (for a future slice, not here): +- **(A) order the bind before docker:** have the agent assert guest binds as part of guest-start (or gate the in-guest `docker.service` start until `/mnt/felhom-drives/` is populated), so the source exists before dockerd's auto-start. +- **(B) controller startup reconcile:** after confirming a drive bind is live, the controller starts any drive-backed app that is `Exited` with a create-time mount error (a targeted, idempotent "drive-present → start gate-eligible apps" pass). *Note: today's drive-gate stops apps when a drive goes missing but does not appear to re-start apps that failed to create at boot — that asymmetry is the bug surface.* +- **(Not C)** putting the drive back as a `pct mpN` would make it present at boot but reintroduces C1 (drive absent at boot → guest bricks) — the very thing the intermediary model removed. + +--- + +# PART 1 — The existing DR mechanism (cited) + +## 1A. Escrow + recovery — what is escrowed, how it's wrapped, what it unlocks + +**Two secrets, one recovery code `R`.** `R` = 10 EFF-large-wordlist words (`internal/escrow/wordlist.go:47`, `RecoveryCodeWords=10`, ~129.2 bits), hyphen-joined, surfaced to the customer once, never logged or hub-stored. + +- **K-escrow (the PBS encryption key).** The live PBS client key `K` (e.g. `/etc/pve/priv/storage/felhom-pbs.enc`) is copied and re-keyed to `kdf=scrypt` under `R` via PBS-native `proxmox-backup-client key change-passphrase` (`internal/escrow/escrow.go` `Wrap()` ~143). Result: an opaque blob PBS itself cannot open without `R`. `CreateResult.Blob` (`escrow.go:49`). +- **Identity-escrow (10D.1).** `IdentityBundle{TunnelToken, PBSToken}` (`internal/escrow/identity.go:22-27`) — "the secrets a re-enrolling box needs to come back 'as host X'" — wrapped under the **same `R`** via `age -p` (scrypt + ChaCha20-Poly1305) → `CreateResult.IdentityBlob` (`escrow.go:56`; `identity.go` `WrapIdentity()` ~29). Additive: the K-escrow path is unchanged when nil (`escrow.go:43`). +- **Self-verify before shipping:** Create unwraps a copy with `R` and checks the recovered fingerprint matches, "Never ship a blob we can't recover" (`escrow.go` ~96). + +**The hub is zero-knowledge.** `escrowUploadRequest{BlobB64, KeyFingerprint, Posture, IdentityBlobB64, DirectiveJSON}` (`hub/internal/api/handler.go:619-631`); the handler "Store the OPAQUE bytes. No decrypt path exists — the hub cannot open this" (`handler.go:678` area), persisted via `SaveHostEscrow` / `SaveHostDRBundle` (`hub/internal/store/store.go`, `GetHostDRBundle` returns opaque `KEscrowBlob`/`IdentityBlob`/`DirectiveJSON`). `dr.go:14-18`: "the hub ORCHESTRATES recovery … but holds **no usable secret** … the escrow blobs it serves are opaque (need R, which the hub never has)." + +**Consume + fingerprint gate (the unlock path).** `internal/escrow/consume.go` — four inputs: `(blob, R, expectedFingerprint, keyDest)` (`consume.go:13-32`). Flow: (1) `Unwrap` with `R` (`key change-passphrase --kdf none`); a wrong `R` "fails closed at the scrypt KDF: nonzero exit, no key emitted" (`consume.go:63`). (2) **Fingerprint gate** — recompute the recovered key's fingerprint and compare to the hub-served `expectedFingerprint`; mismatch → fail fast, no install (`consume.go:69-79`). (3) Atomic install at `keyDest` 0600 (`consume.go:81`). Tested: right-R+right-FP installs; wrong-R no install; FP-mismatch no install (`internal/escrow/consume_test.go`). + +**What `R` ultimately unlocks → how it decrypts the snapshot.** `R` → unwrap K-escrow → plaintext **`K` (PBS client encryption key)** → installed at `keyDest` (`$XDG_CONFIG_HOME/proxmox-backup/encryption-key.json` or `--keyfile`). PBS restore reads `K` from there and decrypts the client-side-encrypted snapshot chunks. `R` is the only out-of-band secret; it also unlocks the identity bundle for tunnel + PBS re-auth. (See 1B for the encryption side.) + +## 1B. What the PBS whole-CT snapshot contains vs not + +- **CONTAINS:** the guest **rootfs + CT config** via vzdump → PBS. "A Docker NAMED volume lives in the LXC rootfs (`/var/lib/docker/volumes/…/_data`) and is ALWAYS captured by vzdump" (`internal/proxmox/doc.go:52`). So the in-guest `app.yaml` (encrypted), the controller's encryption key, DB dumps, configs, and named-volume app data are all in the snapshot. Backup is `CrashConsistent: true` (`internal/backup/runner.go:83`). +- **DOES NOT contain — the bulk-volume gap:** external user-data drives (`felhom-usb`, `felhom-flash`) are bind-mounted `mpN` with `backup=0`/unset, which vzdump **excludes**. `uncoveredMountpoints` lists them (`internal/backup/runner.go:91-98, 228-252`: "covered ONLY when it carries an explicit `backup=1`"). These are restic cross-drive territory, **not** in the PBS whole-CT snapshot — and notably, **`felhom-flash`'s `userdata/media`, `appdata/romm/resources`, etc. (Part 0) are NOT in PBS.** Surfaced as `Backup.UncoveredVolumes` (`internal/hub/report.go:178-194`). +- **Encrypted?** Yes, client-side, by `K`. `PBSSnapshot.Encrypted` is derived from any data file with `crypt-mode == "encrypt"` (`internal/pbs/report.go:44-53`, `client.go:94`). The PBS server is zero-knowledge: "the PBS server has no client key" (`internal/pbs/doc.go:20`). Verify is key-free / ciphertext-level (`internal/pbs/client.go:60`). +- **Host-report fields:** `Backup{TargetID, VMID, Archive, Mode, CrashConsistent, SizeBytes, Success, Error, StartedAt, DurationSeconds, UncoveredVolumes}` (`report.go:182-194`); `PBSSnapshot{Namespace, BackupType, BackupID, BackupTime, SizeBytes, Owner, Protected, Encrypted, VerifyState, VerifyUPID}` (`report.go:223-234`); `VerifyState` ∈ `ok|failed|none` is "the load-bearing field" (`report.go:219`). + +## 1C. The capstone drill + DONE vs PENDING + +Drilled in `documentation/tests/slice10-escrow-consumption-spike-findings.md` and `slice10d-identity-restore-spike-findings.md`: +- **S1 identity round-trip (DONE):** `{tunnel_token, pbs_token}` recovered byte-identical on a fresh box from blob + `R`; wrong `R` fails closed. +- **S2 tunnel re-establishment (DONE):** the new connector with the recovered tunnel token re-registers; Cloudflare routes to it (no DNS change). **Caveat flagged in the drill:** the old connector/token stay valid → 10D must rotate them or delete the stale connector, and the hub's CF token is only zone-scoped WAF-Edit, not account-scoped Tunnel:Edit — a production blocker. +- **S3 PBS whole-CT restore (DONE):** real encrypted `spike-lxc` (~2.5 GB, `crypt-mode=encrypt`) restored on a keyless box using only the recovered `K`, fingerprint-gated. + +**Manual / operator-in-the-loop steps:** customer provides `R`; operator arms recovery-mode on the hub (the out-of-band "the old box is truly lost" gate); operator performs the external credential rotation (tunnel/PBS); operator signs a destructive restore-overwrite. + +**Hub side — DONE.** `hub/internal/api/dr.go:34-183`: `handleSetRecoveryMode`/`handleClearRecoveryMode` (TTL-bounded, auto-expires), `handleReEnroll` (revokes old host key, mints new, serves opaque blobs + directive), `handleGetRestoreDirective` (gated on recovery mode). `store.go`: `InRecoveryMode`, `SetRecoveryMode`/`ClearRecoveryMode`, `RotateHostAPIKey`, `GetHostDRBundle`. Desired-state serving (`SetHostDesired`, `GET /hosts/{id}/desired-state`) is live (slice 10A). + +**Agent side — PENDING (the gap).** The agent **carries but does not consume** the recovery path: `internal/desired/syncer.go:92-94` — "`restore_directive` present (consumed in slice 10D — ignored in 10A)"; the wire struct `WireDesiredState.RestoreDirective` is "forward-compat … 10A carries it through to the cache but does NOT consume it" (`internal/hub/report.go:266, 289-296`). There is **no** agent code that detects recovery mode, fetches re-enroll blobs, unwraps with `R`, or executes a restore. `internal/escrow/consume.go:17`: "The DR orchestration around it (re-enroll in restore mode, source the directive …)" is the missing wrapper. + +**The current end-to-end DR gap, precisely:** the crypto and real-data recovery are proven and the hub can serve the opaque blobs + directive, but **(a)** the agent has no recovery-orchestration executor, **(b)** the out-of-band `R` handshake (how the fresh box obtains `R`) is unspecified, **(c)** external credential rotation (CF tunnel + PBS token) is manual/blocked on token scope, and **(d)** there is **no secret-free recipe** telling the operator what host scaffolding to rebuild before the PBS bytes can land. This spike specs **(d)**. + +--- + +# PART 2 — The secret-free DR recipe (the gap it fills) + +## 2A. THE GAP — what PBS does *not* capture (the scaffolding the operator must rebuild) + +To put the PBS bytes back, the operator must first **rebuild the host + guest + storage scaffolding on new hardware**. None of this is in the PBS whole-CT snapshot (which is *inside-the-guest* bytes) or in the escrow (which is *secrets*): + +1. **Guest sizing** — cores / memory_bytes / disk_bytes to recreate the LXC at the right size (matches `GuestSpec` already on the wire, `desired-state.golden.json:8`). +2. **Drive inventory** — for each user-data drive: **durable-id** (`uuid:…`, wipe-binding `byid:`/`byuuid:`), **role** (primary | bulk-data | vzdump-target | pbs-offsite), **mount path** (`/mnt/` → intermediary `/mnt/felhom-drives/`), capacity, and **enroll vs decommissioned intent**. This is exactly what `StorageTarget` already reports (2B). +3. **PVE storage definitions** — which Proxmox storages exist (id, type ∈ `local-dir|lvmthin|usb|nfs|cifs|pbs|local`, content), so the new host's `/etc/pve/storage.cfg` can be rebuilt (`StorageTarget.{Name,Type,Content}`). +4. **App inventory + per-app storage bindings** — which apps are deployed and which drive/path each binds (e.g. romm.library → `felhom-flash:userdata/roms`) — so apps come up against the right (restic-restored) bulk data. The **controller** owns this (catalog refs + non-secret deploy fields + storage-role intents). +5. **PBS coordinates** — datastore/namespace/snapshot id to restore from (`PBSSnapshot.{Namespace,BackupID}`, the desired-state `pbs_namespace`). + +**Recipe contents (proposed, all non-secret):** +``` +recipe_version: 1 +customer_id, domain, tier +guest: { vmid, cores, memory_bytes, disk_bytes } +pbs: { repo_id, namespace, latest_snapshot_id } # coordinates, NOT the key +drives: [ { durable_id, wipe_durable_id, role, mount_path, fs_type, + total_bytes, intent: enrolled|decommissioned, userdata_layout } ] +pve_storage: [ { name, type, content } ] # to rebuild storage.cfg +apps: [ { name, catalog_ref, enabled, + storage_bindings: [ { container_path, drive_durable_id, subpath } ], + non_secret_deploy_fields: { ... } } ] # NO env secrets +``` +The recipe is the **re-provision plan**; PBS is the **ingredients**; escrow holds the **keys**. Three jobs, three stores. + +## 2B. SOURCE / FLOW — reuse, don't duplicate + +- **The agent already holds the storage/guest facts.** `StorageTarget` (`internal/hub/report.go:86-121`) carries `Name, Type, DurableID, State, Reachable, Total/Used/AvailBytes, Content, MountPath, BackingDevice, ClassHint, Role`. Guest specs and PBS snapshots are already in the host-report (`Backups[]`, `PBSSnapshots[]`, guest list). **The agent can emit the storage+guest+PBS half of the recipe from data it already collects** — additively, as a new host-report section (or a dedicated push), no new privileged reads. +- **The controller owns the app half** — app inventory, catalog refs, and per-app storage bindings (it generates the compose and knows the drive bindings; Part 0 showed romm's binds resolve to `felhom-flash/userdata/roms`). It already pushes a hub report; the app recipe is an additive, secret-free section there. +- **Where the hub stores it (plaintext-safe, non-secret).** Next to the host-report (e.g. a `dr_recipe` column/table keyed by host/customer + `recipe_version`), **not** encrypted (it has no secrets) and **separate** from `host_escrow` (which stays opaque). This is the clean inverse of the retired infra-backup: same "re-provision plan" goal, zero secrets, so plaintext-at-rest is correct rather than a violation. +- **How the operator consumes it during DR.** The hub→agent wire **already reserves the landing slot**: `WireDesiredState.StorageManifest` / `BackupPolicy` / `PBSNamespace` are forward-compat opaque placeholders "kept opaque so the wire is stable as those land" (`internal/hub/report.go:266-273`). DR flow: operator reads the stored recipe → authors the new host's **desired-state** (guest specs + `storage_manifest` from the recipe) → the agent provisions the guest at the right size and re-enrolls drives by durable-id with their roles → **then** escrow recovers `K`/identity and PBS restores the rootfs bytes (and restic restores the bulk drives) into the rebuilt scaffolding. The recipe feeds provisioning; escrow+PBS feed the data. No overlap. + +## 2C. BOUNDARY — the Phase-1 lesson, proven field-by-field + +**The rule (non-negotiable):** the recipe contains **only identifiers, intents, sizes, and coordinates** — never a key, password, token, hash, or any `ENC:`/secret value. Secrets live in the PBS whole-CT snapshot and the escrow blobs, recovered with `R`, **never regenerated** (consistent with the controller's fail-closed `data_key` gate). This is exactly where the retired infra-backup failed: it shipped `encryption_key_b64`, `restic_password`, and a `controller_config_b64` with `cf_api_token`/`session_secret`/`password_hash` (`SPIKE-infra-backup-2026-06-15.md:100-115`) — a zero-knowledge violation. The recipe must not. + +| Recipe field | Sensitive? | Why not | +|---|---|---| +| `customer_id`, `domain`, `tier` | No | Public identifiers | +| `guest.{cores,memory_bytes,disk_bytes,vmid}` | No | Sizing integers (already on the wire, `GuestSpec`) | +| `drives[].durable_id` / `wipe_durable_id` | No | FS-UUID / `byid` WWN — a hardware identifier, not a credential | +| `drives[].{role,mount_path,fs_type,total_bytes,intent,userdata_layout}` | No | Topology + intent; identical class to today's `StorageTarget` (already reported plaintext) | +| `pve_storage[].{name,type,content}` | No | Storage definitions, no auth (the PBS *token* is escrowed, not here) | +| `pbs.{repo_id,namespace,latest_snapshot_id}` | No | Coordinates; the encryption **key** is escrow-only, the access **token** is identity-escrow-only | +| `apps[].{name,catalog_ref,enabled,storage_bindings}` | No | Catalog reference + path bindings; compose comes from the catalog/PBS | +| `apps[].non_secret_deploy_fields` | **Must be filtered** | Only non-secret deploy fields; **every secret-typed field excluded at the source** | + +**The one place a leak could enter** is `apps[].non_secret_deploy_fields` (free-form deploy config). The boundary is enforced at the **emitter** (controller), which already distinguishes secret vs non-secret deploy fields (it encrypts `ENC:` secrets in `app.yaml` and the retired infra-backup work confirmed the controller knows which env fields are secret-typed). + +**A test that would catch a secret leaking in** (spec for the build slice): a `recipe_no_secrets_test` that (1) builds a recipe from a fixture customer whose apps include secret-typed deploy fields (passwords/tokens/keys) and asserts the serialized recipe JSON contains **none** of those values and **no** key whose name matches `(?i)(password|secret|token|key|hash|passphrase|api[_-]?key|ENC:)`; (2) a golden-shape test pinning the recipe's allowed top-level keys so a new field can't silently land without review; (3) a property check that every `apps[].*` field passing through is on an explicit **allowlist**, not a denylist (allowlist = a new sensitive field is excluded by default). Mirror it on both emitters (controller app-half, agent storage-half) so drift is caught on either side — the same cross-repo golden discipline already used for the host-report contract. + +--- + +# PART 3 — Fork list + recommendation + +**Recommendation (spec the recipe; defer the auth; never touch the secrets path):** + +1. **Recipe schema + fields** — adopt the 2A shape (`recipe_version` + customer/guest/pbs/drives/pve_storage/apps). **Fork:** exact field set for `apps[].non_secret_deploy_fields` (allowlist contents) — needs the catalog/controller owner to enumerate the non-secret deploy fields per app. *Recommend: start with `{catalog_ref, enabled, storage_bindings}` only; add non-secret deploy fields incrementally behind the allowlist test.* +2. **Agent-emit vs hub-derive** — **emit from the agent** (storage/guest/PBS half) as an **additive host-report section**, and from the **controller** (app half) as an additive report section; the hub **assembles + stores**, it does not derive facts it doesn't own. *Rationale: the agent is the storage source of truth (`StorageTarget`), the controller owns the app set; the hub deriving either would re-introduce drift.* **Fork:** one combined recipe record vs two half-records stitched at read time. +3. **Storage location/format** — **plaintext** JSON on the hub, keyed by host/customer + `recipe_version`, in a **dedicated `dr_recipe` store separate from `host_escrow`** (opaque) and from the retired `infra_backup*` tables (dropped Phase 1). Hub-manageable: download / prune / pin / "drive a DR re-provision". *Recommend: new table, not a reuse of `infra_backup_versions`.* +4. **Tie-in to escrow / PBS / capstone** — the recipe is the **re-provision plan only**: it feeds **provisioning + desired-state** (the reserved `WireDesiredState.StorageManifest` placeholder, `report.go:271`), then the **existing** escrow→`K` and PBS-restore (Part 1) put secrets+bytes into the rebuilt scaffolding, and **restic** restores the bulk drives the recipe enumerated (the `UncoveredVolumes` gap). The recipe **must not** carry keys/tokens (those stay in escrow) — it only names the `pbs.namespace`/`snapshot_id` and `drives[].durable_id` the restore targets. +5. **Re-enrollment auth (recovery-mode toggle + out-of-band `R`)** — **DEFER.** It is the separate PENDING slice-10D track (agent-side consumption of `restore_directive` + recovery mode, `syncer.go:92`). This slice **specs and emits/stores the recipe**; it does not implement the restore orchestration or the `R` handshake. *Rationale: the recipe is valuable on its own (operator can rebuild scaffolding by hand today) and is independently testable; coupling it to the auth slice delays both.* +6. **Recipe wire-shape versioning** — carry an explicit `recipe_version` (start at 1) and keep the hub store and both emitters version-aware, with a cross-repo **golden** byte-pinned like the existing `desired-state.golden.json` / host-report contracts, so the agent and controller halves can evolve without silent drift. **Fork:** strict-reject vs ignore-unknown-fields on version skew. *Recommend: ignore-unknown on read (forward-compat, as the wire already does for `storage_manifest`), pin known keys in the golden.* + +**Bottom line:** every fact the recipe needs is already collected (agent `StorageTarget` + guests + PBS; controller app set + bindings), the landing slot already exists (`WireDesiredState.StorageManifest`), and the boundary is enforceable at the emitter with an allowlist + no-secrets test. The recipe complements escrow (keys) and PBS/restic (bytes) without duplicating or weakening either — turning "I have PBS bytes" into "I can re-provision *this* customer." + +--- + +## Appendix — Part 0 evidence index (live, read-only except the app-start fix) + +- App states / exit codes / `State.Error` / `RestartCount`: `docker inspect` on guest 9201. +- Guest reboot: host `journalctl` `vzreboot:9201:root@pam OK` @ 12:26:08 local; guest `uptime -s` 10:26 UTC. +- Mount chain: host `lsblk`/`mount`; guest `/proc/mounts` + `ls /mnt/felhom-drives/felhom-flash`. +- No USB event: host `dmesg -T` + `journalctl -k` (10:20–10:35 UTC) empty for usb/sd/ext4. +- Agent rebind: `journalctl -u felhom-agent` `guest-attach: drive bound … 12:26:26` local + continuous `reconcile: enrolled drive bound`. +- Fix: `POST /api/stacks/{name}/start` ×7 → all `Up`; romm canary reads `/romm/library` on flash. + +*Report-only. No source modified. The only production action was restarting seven already-deployed apps whose drive + data were present — service restoration, non-destructive, and itself the confirmation of the root cause.*