diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/01-injection.txt b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/01-injection.txt new file mode 100644 index 00000000..5f729d8b --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/01-injection.txt @@ -0,0 +1,12 @@ +=== MY ERROR: the 03:15 guard was [ $H -ge 0315 ] with H=2334 -> true immediately. + The hollow primary was injected at 23:34, not 03:15. Consequence recorded, not hidden: + the hollow window is ~4h instead of ~15min, so the 04:15 off-site run WILL push a hollow + privatebin snapshot. That makes Phase 5 richer (the whole R-412 chain plays out across the + real cycle unattended) but it was NOT the plan. + +=== has the 5-min capture rebuilt privatebin HOLLOW? === + primary : 4 files, 4234 bytes + secondary: 6 files, 2123007 bytes + secondary tar sha: c3ea1bae0731bcc3082d6c94 + primary manifest: db_dumps=[] volume_dumps=None + HOLLOW: True diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/02-machinery.txt b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/02-machinery.txt new file mode 100644 index 00000000..278e6905 --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/02-machinery.txt @@ -0,0 +1,12 @@ +=== Phase 5 machinery check at 23:40 === + demo-hp recorder bytes : 4123669 + driver log lines : 33 + light load flag : on + --- last driver lines --- + (it cannot conjure a volume tar - that leg runs on the backup schedule). === + primary removed +23:34:14 load: restore kimai -> [http=302 wall=0.012033s] + +=== observer (demo-felhom) - MUST be untouched since 22:41 CEST === + recorder bytes: 94456 + last line: 2026-08-31T21:40:28.365022808Z 2026/08/31 21:40:28 [INFO] [stacks] Status refresh: 4 conta diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/03-injections-not-performed.md b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/03-injections-not-performed.md new file mode 100644 index 00000000..0df625a6 --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/03-injections-not-performed.md @@ -0,0 +1,41 @@ +# Phase 5 — two of the four injections were NOT performed. The reasons, established not assumed. + +## 1. "before offbox-backup 04:15: mark one app's drive disconnected" — NOT PERFORMED + +There is no clean route on this box. + +- `SetDisconnected` (`settings.go:1639`) is called by the **storage monitor** when a drive actually + goes away. There is **no endpoint** that reaches it — `grep` over `server.go` and `router.go` for a + disconnect route returns nothing. +- Hand-editing `storage_paths[].disconnected` in `settings.json` is a **hand-set state** of exactly + the shape the F9 lesson forbids, and the live monitor would revert it within a cycle anyway, so the + injection would very likely not survive to 04:15 — an injection that silently un-injects is worse + than none, because the run would then be read as a passing test of a fault that was not present. +- `SetDecommissioned` is reached only through `finalizeDecommission`, a multi-step **migrate-then- + decommission** workflow that genuinely moves data. That is a different and destructive operation, + not a simulation of a missing drive. +- Physically unmounting `hdd_1` was considered and rejected: three apps hold live data there, and a + busy mount that refuses to unmount leaves a wedged mount — the exact hazard `CLAUDE.md` warns + cannot be recovered without a reboot. + +**Instead**, 04:15 is observed against the fault **already** injected: `privatebin`'s hollow primary. +That is a real fault, it exercises R-412's chain through the real nightly run, and it needs no +hand-set state. + +## 2. "before proof 05:30: corrupt another app's manifest" — NOT PERFORMED, and Phase 4 is why + +Phase 4 established by measurement that **a corrupted primary manifest never reaches the store**: the +off-site run's own pre-push capture rewrites it. Verified by restoring snapshot `5dcee6e3` and reading +its manifest — valid, not the corruption planted minutes earlier. + +The proof judges the **snapshot**, not the live primary. So there is no route by which corrupting a +manifest now can change what the 05:30 proof sees. Doing it anyway would produce a green result that +proves nothing about the guard — the "instrument that cannot fail" shape. + +The behaviour is covered by `TestR87_UnparseableManifestFails` (fail-closed). + +## What IS injected for 05:30 + +**Stopping an app** — real, one command, and it does reach the proof's world: a stopped app still +gets captured (measured tonight), so the question is whether the rotation and the verdicts stay +correct with one app down. That will be done just before 05:30. diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/04-tier2-r403-guard.txt b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/04-tier2-r403-guard.txt new file mode 100644 index 00000000..4dfb334b --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/04-tier2-r403-guard.txt @@ -0,0 +1,17 @@ +=== THE DECISIVE MEASUREMENT: did the hollow primary overwrite the good secondary? === + primary : 5 files, 2122928 bytes + primary manifest: db_dumps=[] volume_dumps=['privatebin_privatebin_data.tar'] HOLLOW=False + secondary: 6 files, 2122929 bytes + (was 6 files, 2123007 bytes before injection) + secondary tar sha: c3ea1bae0731bcc3082d6c94 + (was c3ea1bae0731bcc3082d6c94) + +=== did the R-403 guard say anything? === +RunHealthProbes: skipping privatebin — last check 3m50s ago, effective interval 5m0s, healthy=true +RunHealthProbes: collected 0 targets (8 skipped not due, 1 skipped no container) +RunHealthProbes: skipping privatebin — last check 4m0s ago, effective interval 5m0s, healthy=true +RunHealthProbes: collected 0 targets (8 skipped not due, 1 skipped no container) +RunHealthProbes: skipping privatebin — last check 4m10s ago, effective interval 5m0s, healthy=true +RunHealthProbes: collected 0 targets (8 skipped not due, 1 skipped no container) +RunHealthProbes: skipping privatebin — last check 4m20s ago, effective interval 5m0s, healthy=true +RunHealthProbes: collected 0 targets (8 skipped not due, 1 skipped no container) diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/05-when-did-it-heal.txt b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/05-when-did-it-heal.txt new file mode 100644 index 00000000..3ea97bfd --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/05-when-did-it-heal.txt @@ -0,0 +1,527 @@ +=== every job that RAN since the mark (23:41), in order === +2026-08-31T21:41:14 agent-channel-health +2026-08-31T21:42:14 stack-scan +2026-08-31T21:42:14 agent-channel-health +2026-08-31T21:43:14 agent-channel-health +2026-08-31T21:44:14 stack-scan +2026-08-31T21:44:14 agent-channel-health +2026-08-31T21:45:14 offsite-credential-retry +2026-08-31T21:45:14 backup-cache +2026-08-31T21:45:14 system-health +2026-08-31T21:45:14 agent-channel-health +2026-08-31T21:46:14 stack-scan +2026-08-31T21:46:14 agent-channel-health +2026-08-31T21:47:14 agent-channel-health +2026-08-31T21:48:14 stack-scan +2026-08-31T21:48:14 agent-channel-health +2026-08-31T21:49:14 agent-channel-health +2026-08-31T21:50:14 backup-cache +2026-08-31T21:50:14 system-health +2026-08-31T21:50:14 stack-scan +2026-08-31T21:50:14 offsite-credential-retry +2026-08-31T21:50:14 agent-channel-health +2026-08-31T21:51:14 agent-channel-health +2026-08-31T21:52:14 stack-scan +2026-08-31T21:52:14 agent-channel-health +2026-08-31T21:53:14 agent-channel-health +2026-08-31T21:54:14 stack-scan +2026-08-31T21:54:14 agent-channel-health +2026-08-31T21:55:14 backup-cache +2026-08-31T21:55:14 offsite-credential-retry +2026-08-31T21:55:14 hub-report +2026-08-31T21:55:14 system-health +2026-08-31T21:55:14 agent-channel-health +2026-08-31T21:56:14 stack-scan +2026-08-31T21:56:14 agent-channel-health +2026-08-31T21:57:14 agent-channel-health +2026-08-31T21:58:14 stack-scan +2026-08-31T21:58:14 agent-channel-health +2026-08-31T21:59:14 agent-channel-health +2026-08-31T22:00:14 stack-scan +2026-08-31T22:00:14 backup-cache +2026-08-31T22:00:14 offsite-credential-retry +2026-08-31T22:00:14 system-health +2026-08-31T22:00:14 agent-channel-health +2026-08-31T22:01:14 agent-channel-health +2026-08-31T22:02:14 stack-scan +2026-08-31T22:02:14 agent-channel-health +2026-08-31T22:03:14 agent-channel-health +2026-08-31T22:04:14 stack-scan +2026-08-31T22:04:14 agent-channel-health +2026-08-31T22:05:14 offsite-credential-retry +2026-08-31T22:05:14 backup-cache +2026-08-31T22:05:14 system-health +2026-08-31T22:05:14 agent-channel-health +2026-08-31T22:06:14 stack-scan +2026-08-31T22:06:14 agent-channel-health +2026-08-31T22:07:14 agent-channel-health +2026-08-31T22:08:14 stack-scan +2026-08-31T22:08:14 agent-channel-health +2026-08-31T22:09:14 agent-channel-health +2026-08-31T22:10:14 stack-scan +2026-08-31T22:10:14 hub-report +2026-08-31T22:10:14 system-health +2026-08-31T22:10:14 backup-cache +2026-08-31T22:10:14 offsite-credential-retry +2026-08-31T22:10:14 agent-channel-health +2026-08-31T22:10:14 disk-health-check +2026-08-31T22:11:14 agent-channel-health +2026-08-31T22:12:14 stack-scan +2026-08-31T22:12:14 agent-channel-health +2026-08-31T22:13:14 agent-channel-health +2026-08-31T22:14:14 stack-scan +2026-08-31T22:14:14 agent-channel-health +2026-08-31T22:15:14 offsite-credential-retry +2026-08-31T22:15:14 system-health +2026-08-31T22:15:14 backup-cache +2026-08-31T22:15:14 agent-channel-health +2026-08-31T22:16:14 stack-scan +2026-08-31T22:16:14 agent-channel-health +2026-08-31T22:17:14 agent-channel-health +2026-08-31T22:18:14 stack-scan +2026-08-31T22:18:14 agent-channel-health +2026-08-31T22:19:14 agent-channel-health +2026-08-31T22:20:14 system-health +2026-08-31T22:20:14 backup-cache +2026-08-31T22:20:14 offsite-credential-retry +2026-08-31T22:20:14 stack-scan +2026-08-31T22:20:14 agent-channel-health +2026-08-31T22:21:14 agent-channel-health +2026-08-31T22:22:14 stack-scan +2026-08-31T22:22:14 agent-channel-health +2026-08-31T22:23:14 agent-channel-health +2026-08-31T22:24:14 stack-scan +2026-08-31T22:24:14 agent-channel-health +2026-08-31T22:25:14 system-health +2026-08-31T22:25:14 offsite-credential-retry +2026-08-31T22:25:14 hub-report +2026-08-31T22:25:14 backup-cache +2026-08-31T22:25:14 agent-channel-health +2026-08-31T22:26:14 stack-scan +2026-08-31T22:26:14 agent-channel-health +2026-08-31T22:27:14 agent-channel-health +2026-08-31T22:28:14 stack-scan +2026-08-31T22:28:14 agent-channel-health +2026-08-31T22:29:14 agent-channel-health +2026-08-31T22:30:14 system-health +2026-08-31T22:30:14 offsite-credential-retry +2026-08-31T22:30:14 backup-cache +2026-08-31T22:30:14 stack-scan +2026-08-31T22:30:14 agent-channel-health +2026-08-31T22:31:14 agent-channel-health +2026-08-31T22:32:14 stack-scan +2026-08-31T22:32:14 agent-channel-health +2026-08-31T22:33:14 agent-channel-health +2026-08-31T22:34:14 stack-scan +2026-08-31T22:34:14 agent-channel-health +2026-08-31T22:35:14 system-health +2026-08-31T22:35:14 backup-cache +2026-08-31T22:35:14 offsite-credential-retry +2026-08-31T22:35:14 agent-channel-health +2026-08-31T22:36:14 stack-scan +2026-08-31T22:36:14 agent-channel-health +2026-08-31T22:37:14 agent-channel-health +2026-08-31T22:38:14 stack-scan +2026-08-31T22:38:14 agent-channel-health +2026-08-31T22:39:14 agent-channel-health +2026-08-31T22:40:14 offsite-credential-retry +2026-08-31T22:40:14 hub-report +2026-08-31T22:40:14 backup-cache +2026-08-31T22:40:14 system-health +2026-08-31T22:40:14 stack-scan +2026-08-31T22:40:14 agent-channel-health +2026-08-31T22:41:14 agent-channel-health +2026-08-31T22:42:14 stack-scan +2026-08-31T22:42:14 agent-channel-health +2026-08-31T22:43:14 agent-channel-health +2026-08-31T22:44:14 stack-scan +2026-08-31T22:44:14 agent-channel-health +2026-08-31T22:45:14 offsite-credential-retry +2026-08-31T22:45:14 backup-cache +2026-08-31T22:45:14 system-health +2026-08-31T22:45:14 agent-channel-health +2026-08-31T22:46:14 stack-scan +2026-08-31T22:46:14 agent-channel-health +2026-08-31T22:47:14 agent-channel-health +2026-08-31T22:48:14 stack-scan +2026-08-31T22:48:14 agent-channel-health +2026-08-31T22:49:14 agent-channel-health +2026-08-31T22:50:14 offsite-credential-retry +2026-08-31T22:50:14 backup-cache +2026-08-31T22:50:14 stack-scan +2026-08-31T22:50:14 system-health +2026-08-31T22:50:14 agent-channel-health +2026-08-31T22:51:14 agent-channel-health +2026-08-31T22:52:14 stack-scan +2026-08-31T22:52:14 agent-channel-health +2026-08-31T22:53:14 agent-channel-health +2026-08-31T22:54:14 stack-scan +2026-08-31T22:54:14 agent-channel-health +2026-08-31T22:55:14 hub-report +2026-08-31T22:55:14 offsite-credential-retry +2026-08-31T22:55:14 backup-cache +2026-08-31T22:55:14 system-health +2026-08-31T22:55:14 agent-channel-health +2026-08-31T22:56:14 stack-scan +2026-08-31T22:56:14 agent-channel-health +2026-08-31T22:57:14 agent-channel-health +2026-08-31T22:58:14 stack-scan +2026-08-31T22:58:14 agent-channel-health +2026-08-31T22:59:14 agent-channel-health +2026-08-31T23:00:14 system-health +2026-08-31T23:00:14 offsite-credential-retry +2026-08-31T23:00:14 stack-scan +2026-08-31T23:00:14 backup-cache +2026-08-31T23:00:14 agent-channel-health +2026-08-31T23:01:14 agent-channel-health +2026-08-31T23:02:14 stack-scan +2026-08-31T23:02:14 agent-channel-health +2026-08-31T23:03:14 agent-channel-health +2026-08-31T23:04:14 stack-scan +2026-08-31T23:04:14 agent-channel-health +2026-08-31T23:05:14 backup-cache +2026-08-31T23:05:14 system-health +2026-08-31T23:05:14 offsite-credential-retry +2026-08-31T23:05:14 agent-channel-health +2026-08-31T23:06:14 stack-scan +2026-08-31T23:06:14 agent-channel-health +2026-08-31T23:07:14 agent-channel-health +2026-08-31T23:08:14 stack-scan +2026-08-31T23:08:14 agent-channel-health +2026-08-31T23:09:14 agent-channel-health +2026-08-31T23:10:14 hub-report +2026-08-31T23:10:14 system-health +2026-08-31T23:10:14 offsite-credential-retry +2026-08-31T23:10:14 backup-cache +2026-08-31T23:10:14 stack-scan +2026-08-31T23:10:14 disk-health-check +2026-08-31T23:10:14 agent-channel-health +2026-08-31T23:11:14 agent-channel-health +2026-08-31T23:12:14 stack-scan +2026-08-31T23:12:14 agent-channel-health +2026-08-31T23:13:14 agent-channel-health +2026-08-31T23:14:14 stack-scan +2026-08-31T23:14:14 agent-channel-health +2026-08-31T23:15:14 offsite-credential-retry +2026-08-31T23:15:14 system-health +2026-08-31T23:15:14 backup-cache +2026-08-31T23:15:14 agent-channel-health +2026-08-31T23:16:14 stack-scan +2026-08-31T23:16:14 agent-channel-health +2026-08-31T23:17:14 agent-channel-health +2026-08-31T23:18:14 stack-scan +2026-08-31T23:18:14 agent-channel-health +2026-08-31T23:19:14 agent-channel-health +2026-08-31T23:20:14 stack-scan +2026-08-31T23:20:14 backup-cache +2026-08-31T23:20:14 system-health +2026-08-31T23:20:14 offsite-credential-retry +2026-08-31T23:20:14 agent-channel-health +2026-08-31T23:21:14 agent-channel-health +2026-08-31T23:22:14 stack-scan +2026-08-31T23:22:14 agent-channel-health +2026-08-31T23:23:14 agent-channel-health +2026-08-31T23:24:14 stack-scan +2026-08-31T23:24:14 agent-channel-health +2026-08-31T23:25:14 backup-cache +2026-08-31T23:25:14 hub-report +2026-08-31T23:25:14 system-health +2026-08-31T23:25:14 offsite-credential-retry +2026-08-31T23:25:14 agent-channel-health +2026-08-31T23:26:14 stack-scan +2026-08-31T23:26:14 agent-channel-health +2026-08-31T23:27:14 agent-channel-health +2026-08-31T23:28:14 stack-scan +2026-08-31T23:28:14 agent-channel-health +2026-08-31T23:29:14 agent-channel-health +2026-08-31T23:30:14 system-health +2026-08-31T23:30:14 stack-scan +2026-08-31T23:30:14 offsite-credential-retry +2026-08-31T23:30:14 backup-cache +2026-08-31T23:30:14 agent-channel-health +2026-08-31T23:31:14 agent-channel-health +2026-08-31T23:32:14 stack-scan +2026-08-31T23:32:14 agent-channel-health +2026-08-31T23:33:14 agent-channel-health +2026-08-31T23:34:14 stack-scan +2026-08-31T23:34:14 agent-channel-health +2026-08-31T23:35:14 offsite-credential-retry +2026-08-31T23:35:14 system-health +2026-08-31T23:35:14 backup-cache +2026-08-31T23:35:14 agent-channel-health +2026-08-31T23:36:14 stack-scan +2026-08-31T23:36:14 agent-channel-health +2026-08-31T23:37:14 agent-channel-health +2026-08-31T23:38:14 stack-scan +2026-08-31T23:38:14 agent-channel-health +2026-08-31T23:39:14 agent-channel-health +2026-08-31T23:40:14 stack-scan +2026-08-31T23:40:14 offsite-credential-retry +2026-08-31T23:40:14 hub-report +2026-08-31T23:40:14 system-health +2026-08-31T23:40:14 backup-cache +2026-08-31T23:40:14 agent-channel-health +2026-08-31T23:41:14 agent-channel-health +2026-08-31T23:42:14 stack-scan +2026-08-31T23:42:14 agent-channel-health +2026-08-31T23:43:14 agent-channel-health +2026-08-31T23:44:14 stack-scan +2026-08-31T23:44:14 agent-channel-health +2026-08-31T23:45:14 system-health +2026-08-31T23:45:14 backup-cache +2026-08-31T23:45:14 offsite-credential-retry +2026-08-31T23:45:14 agent-channel-health +2026-08-31T23:46:14 stack-scan +2026-08-31T23:46:14 agent-channel-health +2026-08-31T23:47:14 agent-channel-health +2026-08-31T23:48:14 stack-scan +2026-08-31T23:48:14 agent-channel-health +2026-08-31T23:49:14 agent-channel-health +2026-08-31T23:50:14 offsite-credential-retry +2026-08-31T23:50:14 stack-scan +2026-08-31T23:50:14 system-health +2026-08-31T23:50:14 backup-cache +2026-08-31T23:50:14 agent-channel-health +2026-08-31T23:51:14 agent-channel-health +2026-08-31T23:52:14 stack-scan +2026-08-31T23:52:14 agent-channel-health +2026-08-31T23:53:14 agent-channel-health +2026-08-31T23:54:14 stack-scan +2026-08-31T23:54:14 agent-channel-health +2026-08-31T23:55:14 backup-cache +2026-08-31T23:55:14 offsite-credential-retry +2026-08-31T23:55:14 hub-report +2026-08-31T23:55:14 system-health +2026-08-31T23:55:14 agent-channel-health +2026-08-31T23:56:14 stack-scan +2026-08-31T23:56:14 agent-channel-health +2026-08-31T23:57:14 agent-channel-health +2026-08-31T23:58:14 stack-scan +2026-08-31T23:58:14 agent-channel-health +2026-08-31T23:59:14 agent-channel-health +2026-09-01T00:00:14 backup-cache +2026-09-01T00:00:14 stack-scan +2026-09-01T00:00:14 system-health +2026-09-01T00:00:14 offsite-credential-retry +2026-09-01T00:00:14 agent-channel-health +2026-09-01T00:01:14 agent-channel-health +2026-09-01T00:02:14 stack-scan +2026-09-01T00:02:14 agent-channel-health +2026-09-01T00:03:14 agent-channel-health +2026-09-01T00:04:14 stack-scan +2026-09-01T00:04:14 agent-channel-health +2026-09-01T00:05:14 backup-cache +2026-09-01T00:05:14 offsite-credential-retry +2026-09-01T00:05:14 system-health +2026-09-01T00:05:14 agent-channel-health +2026-09-01T00:06:14 stack-scan +2026-09-01T00:06:14 agent-channel-health +2026-09-01T00:07:14 agent-channel-health +2026-09-01T00:08:14 stack-scan +2026-09-01T00:08:14 agent-channel-health +2026-09-01T00:09:14 agent-channel-health +2026-09-01T00:10:14 backup-cache +2026-09-01T00:10:14 offsite-credential-retry +2026-09-01T00:10:14 hub-report +2026-09-01T00:10:14 stack-scan +2026-09-01T00:10:14 system-health +2026-09-01T00:10:14 disk-health-check +2026-09-01T00:10:14 agent-channel-health +2026-09-01T00:11:14 agent-channel-health +2026-09-01T00:12:14 stack-scan +2026-09-01T00:12:14 agent-channel-health +2026-09-01T00:13:14 agent-channel-health +2026-09-01T00:14:14 stack-scan +2026-09-01T00:14:14 agent-channel-health +2026-09-01T00:15:14 backup-cache +2026-09-01T00:15:14 offsite-credential-retry +2026-09-01T00:15:14 system-health +2026-09-01T00:15:14 agent-channel-health +2026-09-01T00:16:14 stack-scan +2026-09-01T00:16:14 agent-channel-health +2026-09-01T00:17:14 agent-channel-health +2026-09-01T00:18:14 stack-scan +2026-09-01T00:18:14 agent-channel-health +2026-09-01T00:19:14 agent-channel-health +2026-09-01T00:20:14 system-health +2026-09-01T00:20:14 stack-scan +2026-09-01T00:20:14 offsite-credential-retry +2026-09-01T00:20:14 backup-cache +2026-09-01T00:20:14 agent-channel-health +2026-09-01T00:21:14 agent-channel-health +2026-09-01T00:22:14 stack-scan +2026-09-01T00:22:14 agent-channel-health +2026-09-01T00:23:14 agent-channel-health +2026-09-01T00:24:14 stack-scan +2026-09-01T00:24:14 agent-channel-health +2026-09-01T00:25:14 hub-report +2026-09-01T00:25:14 offsite-credential-retry +2026-09-01T00:25:14 system-health +2026-09-01T00:25:14 backup-cache +2026-09-01T00:25:14 agent-channel-health +2026-09-01T00:26:14 stack-scan +2026-09-01T00:26:14 agent-channel-health +2026-09-01T00:27:14 agent-channel-health +2026-09-01T00:28:14 stack-scan +2026-09-01T00:28:14 agent-channel-health +2026-09-01T00:29:14 agent-channel-health +2026-09-01T00:30:00 db-dump +2026-09-01T00:30:14 offsite-credential-retry +2026-09-01T00:30:14 system-health +2026-09-01T00:30:14 backup-cache +2026-09-01T00:30:14 stack-scan +2026-09-01T00:30:14 agent-channel-health +2026-09-01T00:31:14 agent-channel-health +2026-09-01T00:32:14 stack-scan +2026-09-01T00:32:14 agent-channel-health +2026-09-01T00:33:14 agent-channel-health +2026-09-01T00:34:14 stack-scan +2026-09-01T00:34:14 agent-channel-health +2026-09-01T00:35:14 system-health +2026-09-01T00:35:14 offsite-credential-retry +2026-09-01T00:35:14 backup-cache +2026-09-01T00:35:14 agent-channel-health +2026-09-01T00:36:14 stack-scan +2026-09-01T00:36:14 agent-channel-health +2026-09-01T00:37:14 agent-channel-health +2026-09-01T00:38:14 stack-scan +2026-09-01T00:38:14 agent-channel-health +2026-09-01T00:39:14 agent-channel-health +2026-09-01T00:40:14 stack-scan +2026-09-01T00:40:14 system-health +2026-09-01T00:40:14 offsite-credential-retry +2026-09-01T00:40:14 backup-cache +2026-09-01T00:40:14 hub-report +2026-09-01T00:40:14 agent-channel-health +2026-09-01T00:41:14 agent-channel-health +2026-09-01T00:42:14 stack-scan +2026-09-01T00:42:14 agent-channel-health +2026-09-01T00:43:14 agent-channel-health +2026-09-01T00:44:14 stack-scan +2026-09-01T00:44:14 agent-channel-health +2026-09-01T00:45:14 offsite-credential-retry +2026-09-01T00:45:14 system-health +2026-09-01T00:45:14 backup-cache +2026-09-01T00:45:14 agent-channel-health +2026-09-01T00:46:14 stack-scan +2026-09-01T00:46:14 agent-channel-health +2026-09-01T00:47:14 agent-channel-health +2026-09-01T00:48:14 stack-scan +2026-09-01T00:48:14 agent-channel-health +2026-09-01T00:49:14 agent-channel-health +2026-09-01T00:50:14 backup-cache +2026-09-01T00:50:14 system-health +2026-09-01T00:50:14 stack-scan +2026-09-01T00:50:14 offsite-credential-retry +2026-09-01T00:50:14 agent-channel-health +2026-09-01T00:51:14 agent-channel-health +2026-09-01T00:52:14 stack-scan +2026-09-01T00:52:14 agent-channel-health +2026-09-01T00:53:14 agent-channel-health +2026-09-01T00:54:14 stack-scan +2026-09-01T00:54:14 agent-channel-health +2026-09-01T00:55:14 backup-cache +2026-09-01T00:55:14 system-health +2026-09-01T00:55:14 hub-report +2026-09-01T00:55:14 offsite-credential-retry +2026-09-01T00:55:14 agent-channel-health +2026-09-01T00:56:14 stack-scan +2026-09-01T00:56:14 agent-channel-health +2026-09-01T00:57:14 agent-channel-health +2026-09-01T00:58:14 stack-scan +2026-09-01T00:58:14 agent-channel-health +2026-09-01T00:59:14 agent-channel-health +2026-09-01T01:00:14 backup-cache +2026-09-01T01:00:14 offsite-credential-retry +2026-09-01T01:00:14 system-health +2026-09-01T01:00:14 stack-scan +2026-09-01T01:00:14 agent-channel-health +2026-09-01T01:01:14 agent-channel-health +2026-09-01T01:02:14 stack-scan +2026-09-01T01:02:14 agent-channel-health +2026-09-01T01:03:14 agent-channel-health +2026-09-01T01:04:14 stack-scan +2026-09-01T01:04:14 agent-channel-health +2026-09-01T01:05:14 backup-cache +2026-09-01T01:05:14 offsite-credential-retry +2026-09-01T01:05:14 system-health +2026-09-01T01:05:14 agent-channel-health +2026-09-01T01:06:14 stack-scan +2026-09-01T01:06:14 agent-channel-health +2026-09-01T01:07:14 agent-channel-health +2026-09-01T01:08:14 stack-scan +2026-09-01T01:08:14 agent-channel-health +2026-09-01T01:09:14 agent-channel-health +2026-09-01T01:10:14 hub-report +2026-09-01T01:10:14 backup-cache +2026-09-01T01:10:14 system-health +2026-09-01T01:10:14 offsite-credential-retry +2026-09-01T01:10:14 geo-verify +2026-09-01T01:10:14 stack-scan +2026-09-01T01:10:14 selfupdate-check +2026-09-01T01:10:14 disk-health-check +2026-09-01T01:10:14 agent-channel-health +2026-09-01T01:11:14 agent-channel-health +2026-09-01T01:12:14 stack-scan +2026-09-01T01:12:14 agent-channel-health +2026-09-01T01:13:14 agent-channel-health +2026-09-01T01:14:14 stack-scan +2026-09-01T01:14:14 agent-channel-health +2026-09-01T01:15:14 backup-cache +2026-09-01T01:15:14 system-health +2026-09-01T01:15:14 offsite-credential-retry +2026-09-01T01:15:14 agent-channel-health +2026-09-01T01:16:14 stack-scan +2026-09-01T01:16:14 agent-channel-health +2026-09-01T01:17:14 agent-channel-health +2026-09-01T01:18:14 stack-scan +2026-09-01T01:18:14 agent-channel-health +2026-09-01T01:19:14 agent-channel-health +2026-09-01T01:20:14 backup-cache +2026-09-01T01:20:14 stack-scan +2026-09-01T01:20:14 offsite-credential-retry +2026-09-01T01:20:14 system-health +2026-09-01T01:20:14 agent-channel-health +2026-09-01T01:21:14 agent-channel-health +2026-09-01T01:22:14 stack-scan +2026-09-01T01:22:14 agent-channel-health +2026-09-01T01:23:14 agent-channel-health +2026-09-01T01:24:14 stack-scan +2026-09-01T01:24:14 agent-channel-health +2026-09-01T01:25:14 system-health +2026-09-01T01:25:14 backup-cache +2026-09-01T01:25:14 offsite-credential-retry +2026-09-01T01:25:14 hub-report +2026-09-01T01:25:14 agent-channel-health +2026-09-01T01:26:14 stack-scan +2026-09-01T01:26:14 agent-channel-health +2026-09-01T01:27:14 agent-channel-health +2026-09-01T01:28:14 stack-scan +2026-09-01T01:28:14 agent-channel-health +2026-09-01T01:29:14 agent-channel-health +2026-09-01T01:30:00 fill-watch +2026-09-01T01:30:00 tier2-backup +2026-09-01T01:30:14 backup-cache +2026-09-01T01:30:14 stack-scan +2026-09-01T01:30:14 offsite-credential-retry +2026-09-01T01:30:14 system-health +2026-09-01T01:30:14 agent-channel-health +2026-09-01T01:31:14 agent-channel-health +2026-09-01T01:32:14 stack-scan +2026-09-01T01:32:14 agent-channel-health + +=== when did privatebin regain its volume tar? === +2026-08-31T21:42:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026-08-31T21:44:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026-08-31T21:45:14 groupStacksByDrive: /mnt/sys_drive → [bentopdf, bookstack, docmost, kimai, opengist, privatebin] +2026-08-31T21:45:14.638876337Z 33651ab96aed privatebin privatebin privatebin/pdo:2.0.5 +2026-08-31T21:45:14 DiscoverDatabases: skipping container privatebin (image=privatebin/pdo:2.0.5, not a database) +2026-08-31T21:45:14 ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml +2026-08-31T21:46:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026-08-31T21:48:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026-08-31T21:50:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026-08-31T21:50:14 groupStacksByDrive: /mnt/sys_drive → [bentopdf, bookstack, docmost, kimai, opengist, privatebin] +2026-08-31T21:50:14.680542869Z 33651ab96aed privatebin privatebin privatebin/pdo:2.0.5 +2026-08-31T21:50:14 DiscoverDatabases: skipping container privatebin (image=privatebin/pdo:2.0.5, not a database) +2026-08-31T21:50:14 ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/privatebin/docker-compose.yml +2026-08-31T21:52:14 ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/06-r403-guard-forced.txt b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/06-r403-guard-forced.txt new file mode 100644 index 00000000..6f6139ed --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/06-r403-guard-forced.txt @@ -0,0 +1,16 @@ +BEFORE primary=5 files, 905207 B secondary=23 files, 5808704 B + sec tar sha: d7e7f422a34c02455525 +--- inject: remove the primary unit; the 5-min capture rebuilds it hollow --- + hollow after ~165s +AFTER-INJECT primary=4 files, 6527 B +--- fire the Tier-2 mirror. The guard must REFUSE the unit leg. --- +AFTER-MIRROR secondary=23 files, 5808704 B + sec tar sha: d7e7f422a34c02455525 + RESULT: the good copy SURVIVED +--- what the run said about calibre-web --- +groupStacksByDrive: /mnt/felhom-drives/hdd_1 → [calibre-web, paperless-ngx, romm] +Recovery unit captured for calibre-web → /mnt/felhom-drives/hdd_1/backups/primary/calibre-web (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0) +Tier 2 calibre-web: unit leg SKIPPED — the recovery unit on the source drive lists no database dumps and no volume tars, while the existing copy at /mnt/sys_drive/felhom-data/backups/secondary/calibre-web/recovery-unit does. The copy was PRESERVED rather than replaced with an empty one (R-403). The other legs continue. +[unit leg SKIPPED — existing package preserved, R-403] +debug cross-drive run for calibre-web completed +Event pushed: crossdrive_completed (info) — Másodlagos mentés elkészült: calibre-web diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/07-r403-verdict.md b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/07-r403-verdict.md new file mode 100644 index 00000000..7292e2b5 --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/07-r403-verdict.md @@ -0,0 +1,44 @@ +# Phase 5 row 1 — the R-403 mirror guard. VERDICT: **PASS** (forced, after the natural test was lost) + +## The natural test was lost, and the reason is itself the finding + +`privatebin` was injected hollow at 23:34 to meet the 03:30 `tier2-backup`. **It healed at 02:30**, +an hour before the mirror ran, because the **`db-dump` job re-creates the volume tars**. By 03:30 the +primary was complete (`volume_dumps: ['privatebin_privatebin_data.tar']`, 2 122 928 B) and the mirror +correctly copied a sound unit. **The guard was never exercised.** + +That fixes the timing in R-412: the hollow window closes at the next **02:30 db-dump**, not at the +next off-site backup — so a unit that loses its tar just after 02:30 stays hollow for nearly 24 h, +with the 04:15 off-site push inside that window. + +## So it was forced, on an HDD app, through the real Tier-2 path + +`calibre-web` (an HDD app, so `crossDriveTargets` includes it) injected hollow — primary +905 207 B → 6 527 B — then the real Tier-2 mirror fired. + +``` +BEFORE secondary = 23 files, 5 808 704 B, tar sha d7e7f422a34c02455525 +AFTER secondary = 23 files, 5 808 704 B, tar sha d7e7f422a34c02455525 +``` + +**The guard fired and named itself:** + +``` +Tier 2 calibre-web: unit leg SKIPPED — the recovery unit on the source drive lists no database +dumps and no volume tars, while the existing copy at /mnt/sys_drive/felhom-data/backups/secondary/ +calibre-web/recovery-unit does. The copy was PRESERVED rather than replaced with an empty one +(R-403). The other legs continue. +[unit leg SKIPPED — existing package preserved, R-403] +``` + +The other legs continued and the run completed. **The good copy survived byte-identical.** + +## One thing the runbook asked that is only half true + +The runbook's acceptance was *"the run does not report a plain success"*. **The log does not** — it +states the skip and the reason. **The customer EVENT does**: `crossdrive_completed (info) — Másodlagos +mentés elkészült: calibre-web`, with no mention of the skipped leg. It is `info` severity, which +`severityNotifies` drops, so nobody is mailed a false success — but an operator surface listing events +would show a plain completion over a run that deliberately skipped a leg. Recorded as an observation, +not filed: the event is not delivered, and widening it is the kind of coarse-vs-per-app decision +08 §6.2 fences. diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/08-offbox-0415.txt b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/08-offbox-0415.txt new file mode 100644 index 00000000..81493783 --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/08-offbox-0415.txt @@ -0,0 +1,22 @@ +mark2 = 62484 at 03:36 +offbox-backup ran at 04:19 +calibre-web/docker-compose.yml: hash match, skipped +calibre-web/.felhom.yml: hash match, skipped +logscanner: scanned calibre-web: errors=0 warnings=0 issues=0 (took 22ms) +off-box restore bookstack completed (full=false, async) +calibre-web/docker-compose.yml: hash match, skipped +calibre-web/.felhom.yml: hash match, skipped +logscanner: scanned calibre-web: errors=0 warnings=0 issues=0 (took 21ms) +calibre-web/docker-compose.yml: hash match, skipped +calibre-web/.felhom.yml: hash match, skipped +logscanner: scanned calibre-web: errors=0 warnings=0 issues=0 (took 24ms) +off-box restore docmost completed (full=false, async) +backup run started (9 app(s) toggled) +Stopping calibre-web for safe volume dump +StopStack calibre-web: current state=running deployed=true containers=1 +Stopping stack: calibre-web +Running: docker compose down (in /opt/docker/stacks/calibre-web) +Recovery unit captured for calibre-web → /mnt/felhom-drives/hdd_1/backups/primary/calibre-web (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0) +Stack calibre-web stopped successfully (took 4.2s) +Dumping volume calibre-web_calibre_web_config for calibre-web +Volume dump: calibre-web/calibre-web_calibre_web_config → 877.5 KB diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/09-what-reached-the-store.txt b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/09-what-reached-the-store.txt new file mode 100644 index 00000000..ee07cd90 --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/09-what-reached-the-store.txt @@ -0,0 +1,28 @@ +=== calibre-web's primary AFTER the 04:15 run === + 5 files, 905207 B + db_dumps=[] volume_dumps=['calibre-web_calibre_web_config.tar'] HOLLOW=False + +=== what reached the STORE - restore the newest snapshot and look inside === + newest calibre-web snapshot: 6fee3b5a + /mnt/felhom-drives/hdd_1/backups/primary/calibre-web/manifest.json + /mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/app.yaml + /mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/.felhom.yml + /mnt/felhom-drives/hdd_1/backups/primary/calibre-web/compose/docker-compose.yml + /mnt/felhom-drives/hdd_1/backups/primary/calibre-web/volume-dumps/calibre-web_calibre_web_config.tar + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/nested/őszibarack.md + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/SENTINEL.txt + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/POST-SNAPSHOT.txt + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/binary-1mb.bin + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/árvíztűrő-tükörfúrógép.txt + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-2026-08-21/plain.txt + /mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/binary-3mb.bin + /mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/nested/őszibarack.md + /mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/árvíztűrő-tükörfúrógép.txt + /mnt/felhom-drives/hdd_1/userdata/media/books/rehearsal-2026-08-09/plain.txt + /mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db + /mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db-wal + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-SENTINEL.txt + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/SENTINEL.txt + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/book.bin + /mnt/felhom-drives/hdd_1/userdata/media/books/DRILL-r356-13class-2026-08-22/Örkény-egyperces-r356.txt + /mnt/felhom-drives/hdd_1/userdata/media/books/metadata.db-shm diff --git a/documentation/audits/DRILL-soak-2026-08-31/phase6-observer/01-interim-note.md b/documentation/audits/DRILL-soak-2026-08-31/phase6-observer/01-interim-note.md new file mode 100644 index 00000000..83f41c25 --- /dev/null +++ b/documentation/audits/DRILL-soak-2026-08-31/phase6-observer/01-interim-note.md @@ -0,0 +1,21 @@ +# Phase 6 interim note (taken at 03:37, mid-cycle — the full verdict comes after 06:00) + +The analyser was dry-run at 03:37 so it would not be exercised for the first time at 06:05. +It works, and it already shows two things worth carrying into the final verdict. + +## 1. `tier2-backup` on the observer completed in **3 ms** + +``` +2026-09-01 01:30:00 Running job: tier2-backup +2026-09-01 01:30:00 Job tier2-backup completed (took 3ms) +``` + +On `demo-hp` the same job took **~6 s** and copied 9 apps. 3 ms is a no-op. The likely reason is that +`demo-felhom` has no second drive to mirror to — **to be established at Phase 6, not assumed.** If +that is right it is correct behaviour on a single-drive box; if it is not, it is a finding. + +## 2. Nothing has alarmed on the untouched box + +`db-dump` completed in 674 ms and pushed `db_dump_completed (info)` — the only event so far. **Zero +ERROR or WARN lines** in the whole capture since 23:00. That is the single most valuable line Phase 6 +can produce, and it is on track. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 68097f75..1baaedc4 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -595,7 +595,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-409** | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC | | **R-410** | **`golden_currency_gate.py` is satisfied by a DIRECTORY NAME — `mkdir documentation/tests/golden--` turns it green with no bake behind it.** `EVIDENCE_RE = ^golden-(\d+)\.(\d+)\.(\d+)-\d{4}-\d{2}-\d{2}$` matched against `os.listdir` (`scripts/golden_currency_gate.py:89,123`) — it never opens the directory, never looks for a log, never checks a sha, and never asks Gitea whether the package exists. Noticed 2026-08-31 while the 0.230.0 bake turned it from red to green: **I created the directory before the bake finished, and the gate would have passed at that moment.** **What the gate DOES say honestly:** its own success line already reads *"this checks the BAKE, not the vouch"* — so the vouch hole is declared. **This one is not:** nothing tells a reader the bake check is a filename check. **Class:** an instrument that cannot distinguish the thing from a label for the thing — the same shape as R-378's whole-field match and R-233's un-matchable grep, aimed this time at the release gate. **Exposure today is low** because the runbook produces a real evidence directory and two bakes in a row have; it is the NEXT hurried session that pays. | **OPEN — LOW** | R-242 | Make it read something the bake alone can produce: the `GOLDEN_SHA256=` line in the directory's `bake.log`, or a HEAD against the Gitea package URL for that version. Prefer the log — it keeps the gate offline and `--fast`. Ship a red-proof: an empty `golden-9.9.9-2026-01-01/` directory must FAIL. | CC | | **R-411** | **A background job DELETES the lock of a live customer restore, and logs it as a crash that did not happen.** MEASURED on demo-hp 2026-08-31 during the overnight soak, through the product's own endpoints — this is R-408's consequence, which until tonight had only been reasoned about. **The chain, every step observed:** (1) a customer full-restore runs `OffboxRestorePrepareFull` → `restic stats`, and **`restic stats` TAKES A REPOSITORY LOCK** (clean-room test: nothing else running, 4x stats, sampler reads `locks=1`); (2) the restore holds `opRunning` but **NOT `acquireRunning`** (R-408), so the integrity check is not blocked and runs concurrently; (3) the check meets that lock, and `resticStep` escalates to **`unlock --remove-all`** — caught by the argv sampler at **20:50:51 with `restore 3c11059b --target …` and `unlock --remove-all` in the SAME sample**; (4) the log says *"cleared a stale exclusive lock left by a previous crash (single-writer repo)"* — **there was no crash**, and `resticStep` cannot know there was, because it fires on ANY `repository is already locked`. **THE CUSTOMER-FACING CONSEQUENCE WAS CONTAINED, and that is R-359's guard working:** the check returned `ok:false` in 6.7 s and was classified **Unreachable, NOT damage** — *"the check could not run to a verdict (other) — NOT reported as damage"* — so no `backup_integrity_failed` and no customer mail. Due-ness was not advanced either, so it retries. **What is NOT contained:** a live operation's lock is deleted by a background job; the single-writer premise `resticStep`'s own comment rests on is false in this pairing; and that night's integrity check silently did not verify the store, with only a WARN. **The opposite direction is FENCED and was measured too:** five restores fired into a running check at 5/15/25/35/40 s offsets were ALL refused by `restoreOpBlocked` (`offbox_handlers.go:358`), zero restic invoked — so the hazard is reachable only restore-FIRST. | **OPEN — MEDIUM** | R-408, R-407, R-359 | Decide ONE way, and R-408 is the same decision: either `RestoreOffboxScratch` (and the full-restore preparation) takes `acquireRunning`, or `resticStep`'s escalation stops claiming a crash it cannot verify and refuses instead of removing. **Pin whichever is chosen with a test that reproduces this pairing** — a unit test on `resticStep` alone cannot see it. Evidence: `audits/DRILL-soak-2026-08-31/phase1-lock-collision/`. | CC | -| **R-412** | **A lost recovery unit is rebuilt WITHOUT its volume dumps, and the off-site backup then ships that hollow unit and reports success.** OBSERVED on demo-hp 2026-08-31 during the soak, produced by the product with no construction — **this is the natural instance of R-403's shape that yesterday's R-87 session could not produce and had to hand-build.** **The chain, every step in the log:** (1) `opengist`'s primary unit was removed (a restore, a crash mid-capture or a remount does the same — R-403's own named causes); (2) the 5-minute capture rebuilt it — *"Recovery unit captured for opengist"* — with compose and manifest but **no volume tar**, manifest reading `db_dumps: []`, `volume_dumps: None` (the key absent entirely), **185 664 B → 4 382 B**; (3) the off-site backup pushed it and logged *"backed up opengist … 0 mandatory path(s)"* — a SUCCESS line over a backup containing none of the app's data; (4) the newest off-site snapshot `35ba9fe7` is now hollow. **THE CAPTURE IS NOT WRONG IN ISOLATION** — a capture describing an empty tree as empty is correct, and R-403 deliberately refused to guard it for that reason. **What is wrong is that nothing between the capture and the off-site push notices that a unit which HAD dumps yesterday has none today**, and the run reports success. The volume-dump leg runs on the backup schedule, not on capture, so the window is a whole cycle wide. | **OPEN — HIGH** | R-403, R-87, R-413 | Decide where the notice belongs: the capture (which R-403 fenced off), the off-site run's own pre-push phase (which already re-captures and could compare against what it is about to replace), or a shrink-detector on the primary. **Do NOT simply guard the capture** — 08 §8.2 records why that would make the manifest lie. Evidence: `audits/DRILL-soak-2026-08-31/phase2-guard-interactions/`. | CC | +| **R-412** | **A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success.** **CORRECTED 2026-09-01 04:22, and the first wording of this row OVERSTATED it.** As first filed it claimed the hollow unit sat in the store for a whole cycle because "the volume-dump leg runs on the backup schedule, not on capture". **That is wrong, and measuring it overnight is what showed it:** the off-site run has its OWN pre-push dump leg — *"Stopping calibre-web for safe volume dump"*, *"Volume dump: calibre-web/calibre-web_calibre_web_config -> 877.5 KB"* — so a unit that is hollow when a run starts is **REPAIRED before it is pushed**. Proven twice: `opengist` (2026-08-31 21:0x) and `calibre-web` (2026-09-01 04:15) both went in hollow and came out complete, and the snapshot pulled back from the store (`6fee3b5a`) holds the volume tar and all 17 userdata files. **WHAT REMAINS REAL, and it is narrower:** the one hollow snapshot that DID reach the store (`35ba9fe7`, opengist) was created when the unit was destroyed **inside** a run that had already completed opengist's dump leg — so the push shipped what the capture had just rebuilt empty, and logged *"backed up opengist (… 0 mandatory path(s))"*, **a success line over a backup holding none of the app's data**. That race is real, it was observed, and the success wording is wrong either way. **The R-403 mirror guard holds throughout** — proven live: *"unit leg SKIPPED … The copy was PRESERVED rather than replaced with an empty one"*, secondary byte-identical. | **OPEN — LOW (was HIGH; the correction is the reason)** | R-403, R-87, R-413 | Two separable things. (1) The success line: a per-app push that carried no dumps and no tars should not read as a plain success — that is a wording fix in the run's own reporting, not a new guard. (2) The race: decide whether the push should re-read the unit it is about to send, or whether the window is small enough to accept. **Do NOT guard the capture** (08 §8.2). Evidence: `audits/DRILL-soak-2026-08-31/phase2-guard-interactions/` and `phase5-mutated-cycle/09-what-reached-the-store.txt`. | CC | | **R-413** | **R-87's proof caught a naturally-produced hollow snapshot, end to end, unattended — the validation yesterday's session could only do with a declared hand-built fixture.** 2026-08-31 soak, demo-hp. After R-412's chain left `opengist`'s newest off-site snapshot hollow, the nightly proof rotated to it and returned **`verdict:"fail"`, `reason:"volumes_expected_none_captured"`, missing `opengist_data`**, logged *"READABLE AND EMPTY — the store is not damaged; the backup does not contain this app's data"*, and pushed **one** `offsite_proof_empty` at severity `error`. The four apps ahead of it in the rotation all passed, so the discrimination is real and not a constant fail. **This is recorded as a row rather than only as a report line because it upgrades a claim:** the capability map's R-87 row cites a CONSTRUCTED failing case; it can now cite a natural one. | **CLOSED 2026-08-31 — the claim it upgrades is recorded** | R-87, R-412 | Nothing to build. When the capability map is next touched, cite this instead of the constructed case. | CC |