From 292a23bf39a1483f23ec20e1eaf785427c7697a1 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 10 Jul 2026 04:32:32 +0200 Subject: [PATCH] =?UTF-8?q?campaign:=20offsite=20overnight=20soak=202026-0?= =?UTF-8?q?7-10=20=E2=80=94=20PASS=20+=201=20HIGH=20finding=20(crash-mid-p?= =?UTF-8?q?rune=20stale=20lock)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Unattended live soak of the build-complete offsite arc. F4 re-issue + restore round-trip proved live (9/9 byte-identical, immich db-dump exact). Crash recovery surfaced a HIGH resilience gap: crash mid-prune leaves a stale restic exclusive lock the controller can't self-heal → subsequent runs fail until manual unlock. Demo restored to start-state. Track B (quota) + hub-orchestrated A blocked by session-gating → Peti runbook. Co-Authored-By: Claude Opus 4.8 Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6 --- .../CAMPAIGN-offsite-overnight-2026-07-10.md | 83 +++++++++++++++++++ 1 file changed, 83 insertions(+) create mode 100644 documentation/audits/CAMPAIGN-offsite-overnight-2026-07-10.md diff --git a/documentation/audits/CAMPAIGN-offsite-overnight-2026-07-10.md b/documentation/audits/CAMPAIGN-offsite-overnight-2026-07-10.md new file mode 100644 index 0000000..fc13748 --- /dev/null +++ b/documentation/audits/CAMPAIGN-offsite-overnight-2026-07-10.md @@ -0,0 +1,83 @@ +# CAMPAIGN — offsite arc overnight soak — 2026-07-10 (unattended) + +**Class:** unattended live test campaign. **Baselines:** controller v0.109.1 (`0bd4cd0`), hub v0.41.0 +(`09bd719`), agent v0.79.0 (`a6e8bcb`) — all confirmed live at start. **Operator action required:** none +overnight; two morning items below. **Demo end-state:** restored to start-state + newer clean snapshots. + +## VERDICT: PASS (arc resilient) with 1 HIGH finding + 2 notes; 2 tracks blocked by session-gating + +The build-complete offsite arc held up against reality. F4 credential re-issue and the restore round-trip +proved out live for the first time. One genuine resilience gap surfaced (crash-mid-prune stale lock). Two +tracks (the live quota-crossing and the hub-orchestrated provisioning) could not run: the hub operator +config surface is session/CSRF-gated and this run was unattended (no browser, no operator password) — they +join the auto-confirm happy chain in the "needs a supervised/Peti session" bucket. + +## Per-track results + +| Track | Scenario | Result | +|---|---|---| +| Pre-flight | baselines, demo snapshot, fence census | PASS (census: only demo 269008) | +| **A** | Hetzner API ops (create/reset/update-access/delete) | PASS — all API-valid against the real box | +| **A** | **F4 re-issue** (reset invalidates old pw, new works) | **PASS — first live proof** (throwaway 269090) | +| A | exactly-1 label guard data-condition | PASS (ListSubaccounts=2 → guard would refuse) | +| A | DNS-lag (F2) quantified | 3–48s clean; >30s / TIMEOUT under create/delete churn | +| A | Freeze readonly behavioral write-block | INCONCLUSIVE (unreliable write-probe; API op OK) | +| A | hub-orchestrated enable/consume/save | BLOCKED (session-gated) → supervised | +| **B** | live quota-crossing + descriptor churn | BLOCKED (quota change is session-gated) → supervised | +| **C** | single-flight double-trigger | PASS (run2 "skipped — another backup is running") | +| C | crash mid-run: escrow untouched | PASS (escrowed throughout) | +| C | crash restart: zero consume attempts | PASS (applied-marker idempotency held) | +| C | crash mid-run: settings.json integrity | PASS (valid after crash) | +| **C** | **crash mid-prune: repo recovery** | **FINDING C2 (HIGH)** — stale lock blocks all runs | +| **D** | alert pipeline | 1 intentional backup_failed; no flood; cooldown-dedupe not fully exercisable | +| **E** | **restore round-trip both apps vs USB** | **PASS — 9/9 byte-identical** incl. immich db-dump | + +## Findings + +- **C2 (HIGH — resilience gap).** A controller crash **mid-prune** leaves a stale restic *exclusive* lock + (confirmed: `repository is already locked exclusively by PID 33174 on demo-felhom`). The controller has + **no auto-unlock path**, so every subsequent offbox run FAILS (`exit status 1`) until the lock is cleared + manually or ages to ~30-minute staleness — the demo's scheduled 04:15 run would have failed. Repo itself + uncorrupted. Recovered in-campaign via `restic unlock --remove-all` (needed `--remove-all` because the + lock's hostname `demo-felhom` ≠ the container hostname, so restic couldn't auto-detect the dead PID as + stale), then a clean run succeeded. **Recommended follow-up (supervised, code change):** have + `ensureOffboxRepo` (or a pre-run guard) detect + `unlock` a stale lock it owns before failing, and/or the + run handler surface "repo locked — crash recovery" distinctly from a generic backup failure. +- **C1 (LOW — cosmetic).** A crash mid-run leaves `LastStatus="running"` persisted; the in-memory + single-flight mutex recovers (next run is attempted), and the status self-corrects on the next completed + run — but the UI shows a stale "running" until then. +- **A/F2 (confirms existing OPEN).** Fresh sub-account DNS lag is real and highly variable (3–48s clean, + >30s/TIMEOUT under churn) — the hub's scan-retry ladder is load-bearing; the ~60s default may occasionally + be tight under churn. (Already tracked from SLICE 4; this quantifies it.) + +## Cannot-cover (needs a supervised/Peti session — honest list) +The live quota-crossing (Track B) and hub-orchestrated provisioning/consume-once (Track A) both require the +hub operator config surface, which is session/CSRF-gated (no unattended path). The Freeze behavioral +write-block needs a reliable SFTP write-probe. All join the previously-listed auto-confirm-happy-chain / +fresh-enable-full-cycle bucket for the Peti onboarding runbook. + +## Fence audit (start vs end) +- Pool box 611714 sub-accounts: **start = {269008}**, **end = {269008}**, campaign-labelled = **0** both. + (Transient throwaways 269079–269090 created + deleted; Hetzner DELETE is async → census must poll — no + persistent leaks; every track ended campaign=0.) +- Demo offbox: escrow_state escrowed (unchanged), quota_gb 50 (unchanged — never touched), host/user/repo + unchanged, toggled apps {audiobookshelf, immich} preserved (the closing run backed up exactly those 2), + USB's 10 recovery units untouched, controller 0.109.1 healthy. Newer-only diffs: snapshot_count 2→4, + repo 45→86 MB, last_run advanced (within-quota accumulated snapshots — retention prunes over time). +- No throwaway or dummy byte left behind; Track-E restore scratch removed. +- Budgets: Hetzner API ~40 calls (≤600), upload ~0 (Track B blocked — no urandom staged; ≤12 GB), + throwaways ≤2 concurrent (≤8), alerts 1 email (≤100/day). No secret printed (throwaway passwords + generated/used/discarded; hashes only). + +## MORNING — two items for Viktor +1. **One expected alert email** to tick off: **backup_failed** (~02:26 CEST) — the Track-C crash-recovery + test's locked-repo run. Already recovered; no action needed. (No quota/fill/staleness emails — those + tracks needed an operator session.) +2. **Old-box archive retirement decision** (Track E hashes in hand): the new box restores **byte-perfect** + (9/9 files, immich db-dump sha256 `d78466fcbf595b…4526b`). The old-project archive `u629193*` is safe to + retire whenever you're satisfied. + +## Recommended follow-up (next supervised session) +- Fix C2: auto-unlock a self-owned stale restic lock on the offbox run path (small, high-value resilience fix). +- The Peti onboarding runbook remains the vehicle for the session-gated tracks (auto-confirm happy chain, + fresh-enable full cycle, live quota-crossing, hub-orchestrated provisioning + F4/freeze end-to-end).