228dac4c06
Part 0 (live): flash apps on 9201 were down due to an operator pct reboot at 10:26 UTC + a boot-ordering race — dockerd auto-starts unless-stopped flash apps ~18s before the agent re-binds felhom-flash, so the create-time bind mkdir fails (permission denied) and RestartCount=0 never retries. Drive healthy, data intact, no USB drop, durable-id fine, drive-gate uninvolved. v0.70.0 self-restart RULED OUT (container restart, not a guest reboot; +38min after exits). Fix: restarted the 7 apps via the controller (drive present) — all Up. Flagged the intermediary mount app-start race as an architectural gap. Parts 1-3 (cited): characterized escrow (K + identity under recovery code R, fingerprint-gated, hub zero-knowledge) + PBS whole-CT contents (rootfs/secrets in, external drives out) + capstone DONE vs PENDING (agent-side recovery orchestration not wired, syncer.go:92). Defined the secret-free DR recipe (guest sizing + drive durable-id/role inventory + PVE storage + app bindings + PBS coords), sourced from facts the agent/controller already hold, landing in the reserved WireDesiredState.storage_manifest placeholder. Field-by-field boundary proof + a no-secrets test spec. Fork list + recommendation: spec/emit/store the recipe now, defer re-enrollment auth to slice 10D, never touch the escrow/PBS secret path. No code changes, no version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>