diff --git a/documentation/audits/CAMPAIGN-6D-2026-07-15.md b/documentation/audits/CAMPAIGN-6D-2026-07-15.md index e9ed482..2e24101 100644 --- a/documentation/audits/CAMPAIGN-6D-2026-07-15.md +++ b/documentation/audits/CAMPAIGN-6D-2026-07-15.md @@ -52,10 +52,22 @@ phases below — none block the core promise.) | **P4-DEEP** deep tiers | **PARTIAL** | **Snapshot coherence PASS** (calibre-web + immich snapshot paths == capture set, excludes excluded/optional-ro). F7 timed-cut / restic self-heal-mid-run / restore-to-verify-compare not run (timing-sensitive). | | **P5-REST** real crash (⚠) | **PASS** | SIGKILL host-PID mid-offbox → auto-restart ~15 s; journal marked interrupted run failed (no false success); all apps stayed up; post-crash offbox run clean (no stale lock, no marker-without-legs). | | **P-DAY0** golden provision | **CORE PASS** | Fresh nested PVE (qm 300 ← pre-day0-clean) → installer v1.16.0 fetched + **sha-verified agent 0.88.0 (9c8daf94…) + golden 0.136.0 (e4b27fee…)** vs the hub manifest → "Day-0 provision SUCCESS vmid=9201 host_id=demo-vm-felhom-2f4b00" → **controller 0.136.0 Up (healthy)**, guest at-floor, scoped ACL, all installer fixes shipped (age/pbs-apply/guarded wrappers), root@pam rotated+vaulted. Snapshot `post_day0_golden136`. **Deferred** (agent-side, proven 07-12, not golden-specific): escrow auto-confirm (correctly PENDING — fresh repo password vs hub's stale blob; ceremony needs the DR/PBS/WG chain) + offsite round-trip (guest claimed → needs the dashboard password; the 0.136.0 offsite engine itself is already proven on the demo in P-IMMICH/P3-DELIVERY). | +| **P-DAY0-DEEP** escrow+offsite on drill guest | **BLOCKED** | Legs E (escrow ceremony) + O (offsite round-trip) could not run: the drill guest's PBS-DR chain never configured (see the MED/HIGH finding). Viktor's re-register unblock proved not cleanly achievable (reinforces the finding). Per §5 the friend-alpha provisioning gate reverts to "supervise alpha #1's escrow ceremony LIVE" (normal customer-holds-R flow). **Custody note:** the authorized drill deviation (CC-held R) was never exercised — no R minted/claimed; the real friend-alpha escrow runs the normal customer-holds-R supervised flow. This drill would have proved the *mechanism*, not a new custody model. | | **P-OBS** aggregate-state | **NOT REPRODUCED** | No Exited(0) container present on the box → the cosmetic condition is absent right now. Observation-only. | ## 3. Findings (ranked) +- **MED/HIGH — day-0 PBS-DR chain has no self-heal (P-DAY0-DEEP).** On a fresh appliance provision the + agent's PBS-DR consume fires at install *before* the WG tunnel is established → `fingerprint probe + dial 10.77.0.1:8007: i/o timeout`, consumes nothing. Once the tunnel is up (it is: handshake fresh, + :8007 OPEN) the agent does NOT re-request the `pbs_dr` descriptor — it idles "no-op until a descriptor + arrives" and the one-shot descriptor's window has closed. On a re-provision with a reused WG peer, the + hub's auto-provision-on-registration never re-fires, so the descriptor is never re-pushed at all. There + is no operational re-trigger (hub UI has no per-peer delete; agent doesn't self-re-register). Net: a + fresh customer whose WG handshake is slow at install lands with PBS-DR unconfigured → escrow ceremony + cannot run → auto-confirm stays pending → offsite never activates, with no self-heal. **Follow-up TASK:** + bridge-side re-request of the descriptor once the WG peer is confirmed up (+ re-issue on reused-peer + re-provision). Recorded, not fixed inline (§13). - **LOW — duplicate recovery units.** offbox logs: "immich: multiple recovery units found, using newest (…felhom-usb…); ignoring: …felhom-flash…" (same for audiobookshelf). Controller handles it (uses newest), but a stale duplicate unit lingers on a second drive. Track: cleanup of superseded per-app recovery units. @@ -72,15 +84,18 @@ Core promise (SQ3 immich + Accept #1) **GREEN**, and the customer-facing **enlar + customer-notification gates are satisfied on real data, and the **P3-BROWSER** planes (escrow typed-back ceremony completed, hub 8-tab ring, cross-tab session) now PASS. **P-DAY0 CORE also PASSES** — golden 0.136.0 + agent 0.88.0 provision a fresh box to a healthy controller 0.136.0 (the friend-alpha -provisioning dress-rehearsal). Remaining (nice-to-have, not blocking): the P-DAY0 escrow/offsite deep -legs (DR/PBS/WG chain), plus P-TIER2-deep / P4-timing coverage. +provisioning dress-rehearsal). **One caveat surfaced by P-DAY0-DEEP:** the fresh-provision PBS-DR chain +did not self-complete (MED/HIGH finding — install-race + no self-heal), so alpha #1's escrow ceremony +must be **supervised live** until the self-heal follow-up lands. Also remaining (nice-to-have): P-TIER2-deep +/ P4-timing coverage. ## 5. Evidence index → `180:~/campaign6/6D/evidence/` `00-preflight-gate.md`, `01-P-REG.md`, `02-box-state-reality.md`, `03-immich-prestate.manifest`, `04-P-FAB-export.md`, `05-P-FAB-byteidentical-diff.txt`, `06-P-FAB-import.md`, `07-P-IMMICH-offsite.md`, `08-P-IMMICH-restore-diff.txt`, `09-P-IMMICH-restore.md`, `10-P-PLACE.md`, `11-P5-REST.md`, -`12-P4-DEEP-coherence.md`, `13-P3-DELIVERY.md`, `14-P3-BROWSER.md`, `15-P-DAY0-install.md`. +`12-P4-DEEP-coherence.md`, `13-P3-DELIVERY.md`, `14-P3-BROWSER.md`, `15-P-DAY0-install.md`, +`16-P-DAY0-escrow-BLOCKED.md`. ## 6. Box / morning-recovery state @@ -100,9 +115,11 @@ legs (DR/PBS/WG chain), plus P-TIER2-deep / P4-timing coverage. ## 7. Outstanding queue - **P-TIER2** deep 4 scenarios + **P4-DEEP** timing sub-tests (F7 cut / restic self-heal / restore-verify). -- **P-DAY0** remainder: escrow auto-confirm + offsite round-trip on the drill guest (needs the DR/PBS/WG - chain + the drill dashboard password). Drill VM qm 300 left provisioned (snapshot post_day0_golden136); - teardown items: stale/new hub host records for demo-vm-felhom, rotated root@pam. +- **P-DAY0-DEEP (BLOCKED → follow-up TASK):** PBS-DR self-heal (bridge re-requests the `pbs_dr` descriptor + once the WG peer is up; re-issue on reused-peer re-provision) — see the MED/HIGH finding. Until then, + friend-alpha alpha #1's escrow ceremony must be **supervised live** (normal customer-holds-R flow). +- **P-DAY0** teardown: drill VM qm 300 left provisioned + healthy (snapshot post_day0_golden136); + stale + new hub host records for demo-vm-felhom to delete; root@pam rotated+vaulted. - Minor P3-BROWSER remainder: session-expiry/logout re-auth pass. - Signed-update train (publish agent/golden to production customers) + Peti's box — out of 6D scope. - LOW: duplicate recovery-unit cleanup; campaign6 autofs (host reboot).