v0.110.0: offbox stale-lock self-heal (campaign C2) + crash-truthful status (C1)

resticStep escalates a restic lock error to `unlock --remove-all` + one
retry (safe: single-writer repo — sub-account isolation + single-flight
mutex); plain `unlock` is stale-only and can't clear a crash lock across a
container-hostname change. Pre-run stale unlock hygiene on run+restore.
C1: NewManager flips a persisted LastStatus=running to a truthful error.
Both red-proofed (A reproduces the exact campaign backup failure).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
This commit is contained in:
2026-07-10 07:02:23 +02:00
parent 0bd4cd02be
commit cb9e992c7a
5 changed files with 224 additions and 7 deletions
+8
View File
@@ -716,6 +716,14 @@ not just those with HDD data. Non-HDD apps can configure destination, method, an
> anywhere) is a **hard error** → `LastStatus="error"` + operator alert (was a misleading `ok`/0 snapshots).
> A *partial* run (some units missing) stays `ok` but sets a Hungarian **`LastWarning`** naming the skipped
> apps, shown on `/backups`.
> - **Crash-lock self-heal (v0.110.0).** A crash mid-prune leaves a restic EXCLUSIVE lock that plain
> `restic unlock` can't clear (the recreated container's new hostname stops restic proving the dead PID
> stale for ~30 min). Every backup/prune/restore step runs through `resticStep`, which on a lock error
> escalates to `unlock --remove-all` + one retry — safe because the repo has a SINGLE legitimate writer
> (per-customer sub-account isolation + the in-process single-flight mutex). **Boundary:** a DR-cloned
> SECOND controller writing the same repo would defeat this premise — operator-supervised territory, out of
> scope for the auto-heal. A crash mid-run also flips a persisted `LastStatus="running"` to a truthful
> error on the next startup (self-corrects on the next successful run).
> - **Secrets** (SSH key + auto-gen repo password) are **0600 files in the data dir** — never logged/committed.
> - **Password custody + atomicity (v0.105.0, fork-4; pairs with agent v0.77.0).** The repo password is the
> irreplaceable DATA key for the offsite tier, so it rides the **customer-recovery-code (R) escrow**