Files
felhom.eu/documentation/audits/SPIKE-dr-recipe-2026-06-16.md
T
admin 228dac4c06 docs(audit): SPIKE — Phase 2 DR recipe + storage diagnosis (report-only)
Part 0 (live): flash apps on 9201 were down due to an operator pct reboot at
10:26 UTC + a boot-ordering race — dockerd auto-starts unless-stopped flash apps
~18s before the agent re-binds felhom-flash, so the create-time bind mkdir fails
(permission denied) and RestartCount=0 never retries. Drive healthy, data intact,
no USB drop, durable-id fine, drive-gate uninvolved. v0.70.0 self-restart RULED
OUT (container restart, not a guest reboot; +38min after exits). Fix: restarted
the 7 apps via the controller (drive present) — all Up. Flagged the intermediary
mount app-start race as an architectural gap.

Parts 1-3 (cited): characterized escrow (K + identity under recovery code R,
fingerprint-gated, hub zero-knowledge) + PBS whole-CT contents (rootfs/secrets in,
external drives out) + capstone DONE vs PENDING (agent-side recovery orchestration
not wired, syncer.go:92). Defined the secret-free DR recipe (guest sizing + drive
durable-id/role inventory + PVE storage + app bindings + PBS coords), sourced from
facts the agent/controller already hold, landing in the reserved
WireDesiredState.storage_manifest placeholder. Field-by-field boundary proof + a
no-secrets test spec. Fork list + recommendation: spec/emit/store the recipe now,
defer re-enrollment auth to slice 10D, never touch the escrow/PBS secret path.

No code changes, no version bump.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 14:18:39 +02:00

30 KiB
Raw Blame History

SPIKE — Phase 2 DR recipe: characterize existing DR, define the secret-free recipe, lock the boundary (+ storage diagnosis)

Date: 2026-06-16 Type: Report-only spike. No code changes, no version bump. Every claim cited against pushed source at file:line. Part 0 is a live diagnosis on guest 9201 / felhom-pve (one non-destructive fix applied — restarting already-present apps — and noted). Grounding: SPIKE-infra-backup-2026-06-15.md line 24 — the revival is "a slim, secret-free DR recipe (customer → apps → drive/durable-id intents → guest sizing) that complements PBS's ingredients. Secrets are never in it; they are recovered from the PBS whole-CT snapshot, never regenerated."


One-screen summary

  • PART 0 (storage diagnosis): the flash-backed apps did not crash and felhom-flash did not drop. Root cause: an operator pct reboot (PVE vzreboot) of guest 9201 at 10:26 UTC + a boot-ordering race. On guest boot, dockerd auto-starts the unless-stopped flash apps ~18 s before the agent re-binds felhom-flash into the guest; docker fails to create the not-yet-present bind source (mkdir /mnt/felhom-drives/felhom-flash/userdata: permission denied), and because that is a create-time mount failure (RestartCount=0) the restart policy never retries — the apps stay Exited forever even after the bind lands. The drive is healthy, mounted, bound (host + guest), data intact; no USB drop, no durable-id failure, the drive-gate was uninvolved. The v0.70.0 self-restart is RULED OUT (it restarts the felhom-controller Docker container via os.Exit(0), never the LXC guest, and ran 11:0411:17 UTC — ~38 min after the exits). Fix applied: started the 7 apps via the controller (drive present now) — all back Up. Architectural gap flagged: the intermediary mount is C1-immune (guest boots) but does not prevent this app-start race.
  • PART 1 (existing DR): DONE — PBS whole-CT escrow + consume (internal/escrow/*, fingerprint-gated, real-data drilled), identity-escrow ({tunnel_token, pbs_token} age-wrapped under the same recovery code R), tunnel re-establishment, and the hub's zero-knowledge recovery-mode/re-enroll orchestration (hub/internal/api/dr.go). PENDING — the agent-side recovery orchestration that consumes the directive (the agent carries restore_directive forward-compat but explicitly does not act on it — internal/desired/syncer.go:92), and the secret-free recipe itself.
  • PART 2 (recipe): the gap is the non-secret scaffolding PBS does not capture — guest sizing, drive inventory (durable-id → role → mount → enroll/decommission intent), PVE storage defs, app inventory + per-app storage bindings. The agent already holds every fact (StorageTarget{durable_id, role, mount_path, total/used/avail} at report.go:86; guest specs; PBS coords) — it can emit the recipe. The hub→agent wire already reserves storage_manifest/backup_policy/pbs_namespace placeholders (report.go:271-273). Boundary: every recipe field is an identifier/intent/size — no key, password, token, or hash. Secrets stay in PBS + escrow.
  • PART 3: fork list + recommendation: emit the recipe agent-side as an additive host-report section (it owns the storage truth), store it plaintext on the hub next to the host-report, version the wire shape, and defer the re-enrollment auth (recovery-mode consumption) to the existing slice-10D track — the recipe complements escrow+PBS, it does not touch them.

PART 0 — Storage diagnosis: felhom-flash plugged in, yet flash-backed apps exited

What I observed (live, 12:0012:10 UTC)

Seven flash-backed apps Exited: audiobookshelf, calibre-web, immich-server, jellyfin, komga, radarr, romm. Exit codes 143 (audiobookshelf/immich-server/komga) and 128 (calibre-web/jellyfin/radarr/romm). felhom-flash physically connected.

Timezone note (this was load-bearing): docker reports UTC (…Z); the host journal is +02:00 (CEST). All times below normalized to UTC unless marked "local".

End-to-end walk (not "is the device present")

  1. The drive is healthy and fully bound — host side. lsblk: felhom-flash = /dev/sdb1, ext4, LABEL=…, UUID=81a26531-62d8-408d-812f-a178b1d35310, mounted at /mnt/felhom-flash. The intermediary chain is intact:
    • stable parent: /dev/mapper/pve-root on /mnt/felhom-drives (on rootfs, always present),
    • per-drive: /dev/sdb1 on /mnt/felhom-drives/felhom-flash.
  2. Bound through to the guest — guest side. Inside 9201, /proc/mounts shows /dev/sdb1 /mnt/felhom-drives/felhom-flash ext4 rw,relatime, and ls /mnt/felhom-drives/felhom-flashappdata backups media userdata. The flash data is reachable in the guest right now.
  3. The agent never lost it. The reconcile loop logs every ~20 s, continuously: reconcile: enrolled drive bound under shared parent (live, no reboot) … where=/mnt/felhom-flash guest_path=/mnt/felhom-drives/felhom-flash durable_id=uuid:81a26531-… — including across the incident window. (internal/reconcile + internal/localapi guest-bind; durable-id matched.)
  4. No kernel-level USB event. dmesg -T and journalctl -k --since 10:20 --until 10:35 (UTC window) show zero USB disconnect / re-enumeration / sd/ext4 I/O error lines. The task's "transiently dropped + re-enumerated under a new node" hypothesis is NOT supported — the kernel never saw the device leave, and the agent re-bound it by the same uuid:81a26531 throughout. Durable-id tracking did not fail.

What actually happened (root cause, with evidence)

  1. Guest 9201 was rebooted by the operator at 10:26 UTC. uptime -s = boot 10:26 (12:26 local); /proc/uptime ≈ 1.7 h. Host journal at 12:26:08 local: pvedaemon[…]: <root@pam> end task UPID:demo-felhom:…:vzreboot:9201:root@pam: OK. This is a PVE-level pct reboot of the LXC guest, operator-initiated.
  2. The "exits" are the reboot shutdown, not a fault. romm's own log: [2026-06-16 12:26:02] Stopping nginx … Stopping gunicorn … Worker … was sent SIGTERM!; jellyfin: [12:25:59] Disposing CoreAppHost. Exit 143 = 128+SIGTERM; 128 = the image's PID1 exit on SIGTERM. The FinishedAt timestamps (10:25:5810:26:02 UTC) are the guest-shutdown moment.
  3. All seven are restart: unless-stopped — same policy as the DB/redis containers (romm-db, immich-postgres, …), which did come back (Up ~1.7 h). So policy is not the differentiator.
  4. THE SMOKING GUN — a create-time bind-mount failure at boot:
    romm     State.Error = "error while creating mount source path
      '/mnt/felhom-drives/felhom-flash/userdata/roms':
       mkdir /mnt/felhom-drives/felhom-flash/userdata: permission denied"   RestartCount=0
    jellyfin State.Error = "...mkdir /mnt/felhom-drives/felhom-flash/userdata: permission denied"  RestartCount=0
    
    On boot, dockerd tries to restart the flash apps. Their binds point under /mnt/felhom-drives/felhom-flash/…, but the real felhom-flash is not bound into the guest yet — the agent re-binds it ~18 s after boot (host journal: guest-attach: drive bound under shared parent (normalized to one bind, live) … where=/mnt/felhom-flash … 12:26:26 local = 10:26:26 UTC, vs boot 10:26:08). At that instant /mnt/felhom-drives/felhom-flash is the empty, nobody-owned stable-parent placeholder (the parent is /dev/mapper/pve-root, drwxr-xr-x nobody nogroup under the unprivileged-LXC idmap), so the guest-root docker daemon's attempt to mkdir the missing bind source is permission-denied.
  5. RestartCount=0 → the restart policy never retries. A create-time mount failure is not an exit-after-run, so unless-stopped does not apply — the container is left Exited permanently even though the bind lands 18 s later. The DB containers have no such bind (named volumes on rootfs) → they start immediately and stay up. That is the whole asymmetry.

Ruling out the v0.70.0 self-restart (explicitly, as requested)

  • The self-restart shipped this session is os.Exit(0) of the felhom-controller Docker container (controller/internal/api/selfrestart.go); the container is restart: unless-stopped, so Docker restarts that one container. It never issues a guest reboot.
  • The incident was a PVE vzreboot of the LXC guest (host journal, step 5) — a different layer entirely.
  • Timing: the exits were 10:25:5810:26:02 UTC; my self-restart testing ran 11:0411:17 UTC (controller StartedAt 11:17:51 UTC, "Up 46 min" at 12:04). The exits precede my testing by ~38 min. (The task's cited "~11:5411:57" window matches neither the exits nor any controller restart I made; the controller's only live restart is 11:17:51 UTC.) Not implicated.

Fix applied (non-destructive, and it validates the diagnosis)

The drive is present and the data intact, so the only thing wrong is that the apps never started. I started them through the controller's own pipeline (POST /api/stacks/{name}/start, the same path the UI uses) — romm first as a canary (Up, accessing /romm/library on flash), then the rest. Result: audiobookshelf, calibre-web, jellyfin, radarr healthy; immich-server, komga health-starting (normal); romm up. That they start cleanly now confirms the root cause was the boot-ordering race, not any drive fault. No remount/rebind was needed (the mount was already correct); no destructive action taken.

Architectural gap (flag — out of this spike's build scope, worth a follow-up)

The intermediary-mount model (SPIKE-intermediary-mount-2026-06-15.md) is C1-immune (the guest boots clean because the bind source is the always-present stable parent). But it does not prevent an app-start race: the per-drive bind is (re)asserted by the agent's reconcile tick seconds after the guest's dockerd has already tried (and permanently failed, RestartCount=0) to auto-start drive-backed apps. Any operator pct reboot, host reboot, or guest auto-restart reproduces this. Candidate fixes (for a future slice, not here):

  • (A) order the bind before docker: have the agent assert guest binds as part of guest-start (or gate the in-guest docker.service start until /mnt/felhom-drives/<drive> is populated), so the source exists before dockerd's auto-start.
  • (B) controller startup reconcile: after confirming a drive bind is live, the controller starts any drive-backed app that is Exited with a create-time mount error (a targeted, idempotent "drive-present → start gate-eligible apps" pass). Note: today's drive-gate stops apps when a drive goes missing but does not appear to re-start apps that failed to create at boot — that asymmetry is the bug surface.
  • (Not C) putting the drive back as a pct mpN would make it present at boot but reintroduces C1 (drive absent at boot → guest bricks) — the very thing the intermediary model removed.

PART 1 — The existing DR mechanism (cited)

1A. Escrow + recovery — what is escrowed, how it's wrapped, what it unlocks

Two secrets, one recovery code R. R = 10 EFF-large-wordlist words (internal/escrow/wordlist.go:47, RecoveryCodeWords=10, ~129.2 bits), hyphen-joined, surfaced to the customer once, never logged or hub-stored.

  • K-escrow (the PBS encryption key). The live PBS client key K (e.g. /etc/pve/priv/storage/felhom-pbs.enc) is copied and re-keyed to kdf=scrypt under R via PBS-native proxmox-backup-client key change-passphrase (internal/escrow/escrow.go Wrap() ~143). Result: an opaque blob PBS itself cannot open without R. CreateResult.Blob (escrow.go:49).
  • Identity-escrow (10D.1). IdentityBundle{TunnelToken, PBSToken} (internal/escrow/identity.go:22-27) — "the secrets a re-enrolling box needs to come back 'as host X'" — wrapped under the same R via age -p (scrypt + ChaCha20-Poly1305) → CreateResult.IdentityBlob (escrow.go:56; identity.go WrapIdentity() ~29). Additive: the K-escrow path is unchanged when nil (escrow.go:43).
  • Self-verify before shipping: Create unwraps a copy with R and checks the recovered fingerprint matches, "Never ship a blob we can't recover" (escrow.go ~96).

The hub is zero-knowledge. escrowUploadRequest{BlobB64, KeyFingerprint, Posture, IdentityBlobB64, DirectiveJSON} (hub/internal/api/handler.go:619-631); the handler "Store the OPAQUE bytes. No decrypt path exists — the hub cannot open this" (handler.go:678 area), persisted via SaveHostEscrow / SaveHostDRBundle (hub/internal/store/store.go, GetHostDRBundle returns opaque KEscrowBlob/IdentityBlob/DirectiveJSON). dr.go:14-18: "the hub ORCHESTRATES recovery … but holds no usable secret … the escrow blobs it serves are opaque (need R, which the hub never has)."

Consume + fingerprint gate (the unlock path). internal/escrow/consume.go — four inputs: (blob, R, expectedFingerprint, keyDest) (consume.go:13-32). Flow: (1) Unwrap with R (key change-passphrase --kdf none); a wrong R "fails closed at the scrypt KDF: nonzero exit, no key emitted" (consume.go:63). (2) Fingerprint gate — recompute the recovered key's fingerprint and compare to the hub-served expectedFingerprint; mismatch → fail fast, no install (consume.go:69-79). (3) Atomic install at keyDest 0600 (consume.go:81). Tested: right-R+right-FP installs; wrong-R no install; FP-mismatch no install (internal/escrow/consume_test.go).

What R ultimately unlocks → how it decrypts the snapshot. R → unwrap K-escrow → plaintext K (PBS client encryption key) → installed at keyDest ($XDG_CONFIG_HOME/proxmox-backup/encryption-key.json or --keyfile). PBS restore reads K from there and decrypts the client-side-encrypted snapshot chunks. R is the only out-of-band secret; it also unlocks the identity bundle for tunnel + PBS re-auth. (See 1B for the encryption side.)

1B. What the PBS whole-CT snapshot contains vs not

  • CONTAINS: the guest rootfs + CT config via vzdump → PBS. "A Docker NAMED volume lives in the LXC rootfs (/var/lib/docker/volumes/…/_data) and is ALWAYS captured by vzdump" (internal/proxmox/doc.go:52). So the in-guest app.yaml (encrypted), the controller's encryption key, DB dumps, configs, and named-volume app data are all in the snapshot. Backup is CrashConsistent: true (internal/backup/runner.go:83).
  • DOES NOT contain — the bulk-volume gap: external user-data drives (felhom-usb, felhom-flash) are bind-mounted mpN with backup=0/unset, which vzdump excludes. uncoveredMountpoints lists them (internal/backup/runner.go:91-98, 228-252: "covered ONLY when it carries an explicit backup=1"). These are restic cross-drive territory, not in the PBS whole-CT snapshot — and notably, felhom-flash's userdata/media, appdata/romm/resources, etc. (Part 0) are NOT in PBS. Surfaced as Backup.UncoveredVolumes (internal/hub/report.go:178-194).
  • Encrypted? Yes, client-side, by K. PBSSnapshot.Encrypted is derived from any data file with crypt-mode == "encrypt" (internal/pbs/report.go:44-53, client.go:94). The PBS server is zero-knowledge: "the PBS server has no client key" (internal/pbs/doc.go:20). Verify is key-free / ciphertext-level (internal/pbs/client.go:60).
  • Host-report fields: Backup{TargetID, VMID, Archive, Mode, CrashConsistent, SizeBytes, Success, Error, StartedAt, DurationSeconds, UncoveredVolumes} (report.go:182-194); PBSSnapshot{Namespace, BackupType, BackupID, BackupTime, SizeBytes, Owner, Protected, Encrypted, VerifyState, VerifyUPID} (report.go:223-234); VerifyStateok|failed|none is "the load-bearing field" (report.go:219).

1C. The capstone drill + DONE vs PENDING

Drilled in documentation/tests/slice10-escrow-consumption-spike-findings.md and slice10d-identity-restore-spike-findings.md:

  • S1 identity round-trip (DONE): {tunnel_token, pbs_token} recovered byte-identical on a fresh box from blob + R; wrong R fails closed.
  • S2 tunnel re-establishment (DONE): the new connector with the recovered tunnel token re-registers; Cloudflare routes to it (no DNS change). Caveat flagged in the drill: the old connector/token stay valid → 10D must rotate them or delete the stale connector, and the hub's CF token is only zone-scoped WAF-Edit, not account-scoped Tunnel:Edit — a production blocker.
  • S3 PBS whole-CT restore (DONE): real encrypted spike-lxc (~2.5 GB, crypt-mode=encrypt) restored on a keyless box using only the recovered K, fingerprint-gated.

Manual / operator-in-the-loop steps: customer provides R; operator arms recovery-mode on the hub (the out-of-band "the old box is truly lost" gate); operator performs the external credential rotation (tunnel/PBS); operator signs a destructive restore-overwrite.

Hub side — DONE. hub/internal/api/dr.go:34-183: handleSetRecoveryMode/handleClearRecoveryMode (TTL-bounded, auto-expires), handleReEnroll (revokes old host key, mints new, serves opaque blobs + directive), handleGetRestoreDirective (gated on recovery mode). store.go: InRecoveryMode, SetRecoveryMode/ClearRecoveryMode, RotateHostAPIKey, GetHostDRBundle. Desired-state serving (SetHostDesired, GET /hosts/{id}/desired-state) is live (slice 10A).

Agent side — PENDING (the gap). The agent carries but does not consume the recovery path: internal/desired/syncer.go:92-94 — "restore_directive present (consumed in slice 10D — ignored in 10A)"; the wire struct WireDesiredState.RestoreDirective is "forward-compat … 10A carries it through to the cache but does NOT consume it" (internal/hub/report.go:266, 289-296). There is no agent code that detects recovery mode, fetches re-enroll blobs, unwraps with R, or executes a restore. internal/escrow/consume.go:17: "The DR orchestration around it (re-enroll in restore mode, source the directive …)" is the missing wrapper.

The current end-to-end DR gap, precisely: the crypto and real-data recovery are proven and the hub can serve the opaque blobs + directive, but (a) the agent has no recovery-orchestration executor, (b) the out-of-band R handshake (how the fresh box obtains R) is unspecified, (c) external credential rotation (CF tunnel + PBS token) is manual/blocked on token scope, and (d) there is no secret-free recipe telling the operator what host scaffolding to rebuild before the PBS bytes can land. This spike specs (d).


PART 2 — The secret-free DR recipe (the gap it fills)

2A. THE GAP — what PBS does not capture (the scaffolding the operator must rebuild)

To put the PBS bytes back, the operator must first rebuild the host + guest + storage scaffolding on new hardware. None of this is in the PBS whole-CT snapshot (which is inside-the-guest bytes) or in the escrow (which is secrets):

  1. Guest sizing — cores / memory_bytes / disk_bytes to recreate the LXC at the right size (matches GuestSpec already on the wire, desired-state.golden.json:8).
  2. Drive inventory — for each user-data drive: durable-id (uuid:…, wipe-binding byid:<wwn>/byuuid:), role (primary | bulk-data | vzdump-target | pbs-offsite), mount path (/mnt/<name> → intermediary /mnt/felhom-drives/<name>), capacity, and enroll vs decommissioned intent. This is exactly what StorageTarget already reports (2B).
  3. PVE storage definitions — which Proxmox storages exist (id, type ∈ local-dir|lvmthin|usb|nfs|cifs|pbs|local, content), so the new host's /etc/pve/storage.cfg can be rebuilt (StorageTarget.{Name,Type,Content}).
  4. App inventory + per-app storage bindings — which apps are deployed and which drive/path each binds (e.g. romm.library → felhom-flash:userdata/roms) — so apps come up against the right (restic-restored) bulk data. The controller owns this (catalog refs + non-secret deploy fields + storage-role intents).
  5. PBS coordinates — datastore/namespace/snapshot id to restore from (PBSSnapshot.{Namespace,BackupID}, the desired-state pbs_namespace).

Recipe contents (proposed, all non-secret):

recipe_version: 1
customer_id, domain, tier
guest:        { vmid, cores, memory_bytes, disk_bytes }
pbs:          { repo_id, namespace, latest_snapshot_id }       # coordinates, NOT the key
drives:       [ { durable_id, wipe_durable_id, role, mount_path, fs_type,
                  total_bytes, intent: enrolled|decommissioned, userdata_layout } ]
pve_storage:  [ { name, type, content } ]                       # to rebuild storage.cfg
apps:         [ { name, catalog_ref, enabled,
                  storage_bindings: [ { container_path, drive_durable_id, subpath } ],
                  non_secret_deploy_fields: { ... } } ]         # NO env secrets

The recipe is the re-provision plan; PBS is the ingredients; escrow holds the keys. Three jobs, three stores.

2B. SOURCE / FLOW — reuse, don't duplicate

  • The agent already holds the storage/guest facts. StorageTarget (internal/hub/report.go:86-121) carries Name, Type, DurableID, State, Reachable, Total/Used/AvailBytes, Content, MountPath, BackingDevice, ClassHint, Role. Guest specs and PBS snapshots are already in the host-report (Backups[], PBSSnapshots[], guest list). The agent can emit the storage+guest+PBS half of the recipe from data it already collects — additively, as a new host-report section (or a dedicated push), no new privileged reads.
  • The controller owns the app half — app inventory, catalog refs, and per-app storage bindings (it generates the compose and knows the drive bindings; Part 0 showed romm's binds resolve to felhom-flash/userdata/roms). It already pushes a hub report; the app recipe is an additive, secret-free section there.
  • Where the hub stores it (plaintext-safe, non-secret). Next to the host-report (e.g. a dr_recipe column/table keyed by host/customer + recipe_version), not encrypted (it has no secrets) and separate from host_escrow (which stays opaque). This is the clean inverse of the retired infra-backup: same "re-provision plan" goal, zero secrets, so plaintext-at-rest is correct rather than a violation.
  • How the operator consumes it during DR. The hub→agent wire already reserves the landing slot: WireDesiredState.StorageManifest / BackupPolicy / PBSNamespace are forward-compat opaque placeholders "kept opaque so the wire is stable as those land" (internal/hub/report.go:266-273). DR flow: operator reads the stored recipe → authors the new host's desired-state (guest specs + storage_manifest from the recipe) → the agent provisions the guest at the right size and re-enrolls drives by durable-id with their roles → then escrow recovers K/identity and PBS restores the rootfs bytes (and restic restores the bulk drives) into the rebuilt scaffolding. The recipe feeds provisioning; escrow+PBS feed the data. No overlap.

2C. BOUNDARY — the Phase-1 lesson, proven field-by-field

The rule (non-negotiable): the recipe contains only identifiers, intents, sizes, and coordinates — never a key, password, token, hash, or any ENC:/secret value. Secrets live in the PBS whole-CT snapshot and the escrow blobs, recovered with R, never regenerated (consistent with the controller's fail-closed data_key gate). This is exactly where the retired infra-backup failed: it shipped encryption_key_b64, restic_password, and a controller_config_b64 with cf_api_token/session_secret/password_hash (SPIKE-infra-backup-2026-06-15.md:100-115) — a zero-knowledge violation. The recipe must not.

Recipe field Sensitive? Why not
customer_id, domain, tier No Public identifiers
guest.{cores,memory_bytes,disk_bytes,vmid} No Sizing integers (already on the wire, GuestSpec)
drives[].durable_id / wipe_durable_id No FS-UUID / byid WWN — a hardware identifier, not a credential
drives[].{role,mount_path,fs_type,total_bytes,intent,userdata_layout} No Topology + intent; identical class to today's StorageTarget (already reported plaintext)
pve_storage[].{name,type,content} No Storage definitions, no auth (the PBS token is escrowed, not here)
pbs.{repo_id,namespace,latest_snapshot_id} No Coordinates; the encryption key is escrow-only, the access token is identity-escrow-only
apps[].{name,catalog_ref,enabled,storage_bindings} No Catalog reference + path bindings; compose comes from the catalog/PBS
apps[].non_secret_deploy_fields Must be filtered Only non-secret deploy fields; every secret-typed field excluded at the source

The one place a leak could enter is apps[].non_secret_deploy_fields (free-form deploy config). The boundary is enforced at the emitter (controller), which already distinguishes secret vs non-secret deploy fields (it encrypts ENC: secrets in app.yaml and the retired infra-backup work confirmed the controller knows which env fields are secret-typed).

A test that would catch a secret leaking in (spec for the build slice): a recipe_no_secrets_test that (1) builds a recipe from a fixture customer whose apps include secret-typed deploy fields (passwords/tokens/keys) and asserts the serialized recipe JSON contains none of those values and no key whose name matches (?i)(password|secret|token|key|hash|passphrase|api[_-]?key|ENC:); (2) a golden-shape test pinning the recipe's allowed top-level keys so a new field can't silently land without review; (3) a property check that every apps[].* field passing through is on an explicit allowlist, not a denylist (allowlist = a new sensitive field is excluded by default). Mirror it on both emitters (controller app-half, agent storage-half) so drift is caught on either side — the same cross-repo golden discipline already used for the host-report contract.


PART 3 — Fork list + recommendation

Recommendation (spec the recipe; defer the auth; never touch the secrets path):

  1. Recipe schema + fields — adopt the 2A shape (recipe_version + customer/guest/pbs/drives/pve_storage/apps). Fork: exact field set for apps[].non_secret_deploy_fields (allowlist contents) — needs the catalog/controller owner to enumerate the non-secret deploy fields per app. Recommend: start with {catalog_ref, enabled, storage_bindings} only; add non-secret deploy fields incrementally behind the allowlist test.
  2. Agent-emit vs hub-deriveemit from the agent (storage/guest/PBS half) as an additive host-report section, and from the controller (app half) as an additive report section; the hub assembles + stores, it does not derive facts it doesn't own. Rationale: the agent is the storage source of truth (StorageTarget), the controller owns the app set; the hub deriving either would re-introduce drift. Fork: one combined recipe record vs two half-records stitched at read time.
  3. Storage location/formatplaintext JSON on the hub, keyed by host/customer + recipe_version, in a dedicated dr_recipe store separate from host_escrow (opaque) and from the retired infra_backup* tables (dropped Phase 1). Hub-manageable: download / prune / pin / "drive a DR re-provision". Recommend: new table, not a reuse of infra_backup_versions.
  4. Tie-in to escrow / PBS / capstone — the recipe is the re-provision plan only: it feeds provisioning + desired-state (the reserved WireDesiredState.StorageManifest placeholder, report.go:271), then the existing escrow→K and PBS-restore (Part 1) put secrets+bytes into the rebuilt scaffolding, and restic restores the bulk drives the recipe enumerated (the UncoveredVolumes gap). The recipe must not carry keys/tokens (those stay in escrow) — it only names the pbs.namespace/snapshot_id and drives[].durable_id the restore targets.
  5. Re-enrollment auth (recovery-mode toggle + out-of-band R)DEFER. It is the separate PENDING slice-10D track (agent-side consumption of restore_directive + recovery mode, syncer.go:92). This slice specs and emits/stores the recipe; it does not implement the restore orchestration or the R handshake. Rationale: the recipe is valuable on its own (operator can rebuild scaffolding by hand today) and is independently testable; coupling it to the auth slice delays both.
  6. Recipe wire-shape versioning — carry an explicit recipe_version (start at 1) and keep the hub store and both emitters version-aware, with a cross-repo golden byte-pinned like the existing desired-state.golden.json / host-report contracts, so the agent and controller halves can evolve without silent drift. Fork: strict-reject vs ignore-unknown-fields on version skew. Recommend: ignore-unknown on read (forward-compat, as the wire already does for storage_manifest), pin known keys in the golden.

Bottom line: every fact the recipe needs is already collected (agent StorageTarget + guests + PBS; controller app set + bindings), the landing slot already exists (WireDesiredState.StorageManifest), and the boundary is enforceable at the emitter with an allowlist + no-secrets test. The recipe complements escrow (keys) and PBS/restic (bytes) without duplicating or weakening either — turning "I have PBS bytes" into "I can re-provision this customer."


Appendix — Part 0 evidence index (live, read-only except the app-start fix)

  • App states / exit codes / State.Error / RestartCount: docker inspect on guest 9201.
  • Guest reboot: host journalctl vzreboot:9201:root@pam OK @ 12:26:08 local; guest uptime -s 10:26 UTC.
  • Mount chain: host lsblk/mount; guest /proc/mounts + ls /mnt/felhom-drives/felhom-flash.
  • No USB event: host dmesg -T + journalctl -k (10:2010:35 UTC) empty for usb/sd/ext4.
  • Agent rebind: journalctl -u felhom-agent guest-attach: drive bound … 12:26:26 local + continuous reconcile: enrolled drive bound.
  • Fix: POST /api/stacks/{name}/start ×7 → all Up; romm canary reads /romm/library on flash.

Report-only. No source modified. The only production action was restarting seven already-deployed apps whose drive + data were present — service restoration, non-destructive, and itself the confirmation of the root cause.