soak drill phases 0-4: R-408's hazard is REAL and reachable; R-357 finally live (R-411..R-413)
gates / gates (push) Failing after 17s
gates / gates (push) Failing after 17s
Overnight soak on demo-hp, phases 0-4 of 7. Evidence documentation/audits/DRILL-soak-2026-08-31/. No production code written, no golden, no version bump - findings only. PHASE 1 FAIL - R-411. The R-408 hazard was only ever reasoned about; tonight it was measured through the product's own endpoints. restic stats TAKES A LOCK (clean-room: nothing else running, 4x stats, sampler reads locks=1). A customer full-restore runs stats in its preparation while holding NO acquireRunning, so the integrity check is not blocked, runs, meets that lock, and resticStep escalates to unlock --remove-all - the sampler caught "restore ..." and "unlock --remove-all" in the SAME sample at 20:50:51. The log calls it "a stale exclusive lock left by a previous crash"; there was no crash. THE CUSTOMER-FACING CONSEQUENCE IS CONTAINED and that is R-359 working: the check was classified Unreachable, NOT damage, so no alarm and due-ness held. The opposite direction is FENCED - five restores fired into a running check at 5/15/25/35/40s were all refused by restoreOpBlocked with zero restic invoked. PHASE 2 FAIL - R-412, and it is the natural instance of R-403's shape that yesterday's session had to hand-build. A lost primary unit is rebuilt by the capture WITHOUT its volume dumps (185664 B -> 4382 B, volume_dumps: None), and the off-site backup then ships it and logs "backed up opengist ... 0 mandatory path(s)" - a success line over a backup holding none of the app's data. PHASE 2 also R-413: the R-87 proof CAUGHT that hollow snapshot unattended - verdict "fail", volumes_expected_none_captured: opengist_data, one offsite_proof_empty at severity error, while the four apps ahead of it in the rotation passed. 2.3 stale marker PASS, 2.4 two restores PASS, 2.5 corrected the runbook's premise (RunTier2 has zero acquireRunning and zero restic references - a local mirror cannot collide with the repo, so the proof correctly does not skip for it). PHASE 3 PASS - R-357 live-validated at last, six days owed. A real full filesystem (1900544 B free vs 5878421 B needed) on a separate 1 TB device that does not back Docker. Refused BEFORE StopStack: app "Up 6 minutes (healthy)" unchanged, 0 safety dumps, live userdata tree fingerprint 5a9b2db75dfd4ba4131adb3255670472 unchanged, message names both numbers in Hungarian. Ballast removed, the SAME restore then worked and the fingerprint is still identical. On the same full disk the Tier-2 run, the proof and the integrity check all behaved: the proof refused before any download through the shared unitOnlyHeadroom gate extracted today, reached no verdict and did not alarm. PHASE 4 PASS - the false-alarm control the whole R-87 design rests on now EXISTS. Found by reading all 53 catalogue composes: bentopdf is the only template with neither a database service nor a named volume. Deployed, backed up, proved: PASSED in 2.224s with zero alarms. Four of my own instrument errors were caught by their own controls before any result was believed: a hub log line used as a controller positive control, a grep pattern that missed a registered job, a heredoc that mangled a planted marker, and a catalogue scan that read 0 of 0 apps from the wrong path. Phases 5-7 to follow: the real nightly cycle with injected faults, the untouched observer, and teardown.
This commit is contained in:
@@ -594,6 +594,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-408** | **`RestoreOffboxScratch` takes NO single-writer flag, and the file that depends on that invariant states it as universal.** `offbox_integrity.go:28` reads *"Every off-site operation takes `acquireRunning` for exactly that reason"* — the reason being that `resticStep` escalates to `unlock --remove-all` and is safe only because the in-process mutex proves no sibling operation is live. **Measured 2026-08-31:** `grep -rn 'acquireRunning()'` finds nine non-test callers and `RestoreOffboxScratch` (`offbox_restore.go:234`) is not among them. `restore_wizard.go:174` records the same fact independently — *"`RestoreOffboxScratch` never acquires it at all"* — and the UI works around it with a separate display flag (`opRunning`), so the gap is known at the web layer and unknown at the one that reasons about repository safety. **The exposure today is small and the exposure tomorrow is not:** the web handler's `restoreOpBlocked()` fences the only caller that exists, and restic's restore takes no lock (R-407's sampler), so nothing currently collides. **An unattended restore-test — R-87 — would be the first caller with no web handler in front of it.** | **OPEN — MEDIUM** | R-87, R-359 | Decide ONE way: either `RestoreOffboxScratch` takes `acquireRunning` (and every existing caller is re-checked for the double-acquire refusal `tier2_restore.go:183` warns about), or `offbox_integrity.go:28`'s sentence is corrected to name the exception. Whichever is chosen, **pin it with a test** — this is a comment asserting an invariant with nothing holding it. Do this BEFORE R-87 ships anything. | CC |
|
||||
| **R-409** | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC |
|
||||
| **R-410** | **`golden_currency_gate.py` is satisfied by a DIRECTORY NAME — `mkdir documentation/tests/golden-<VER>-<DATE>` turns it green with no bake behind it.** `EVIDENCE_RE = ^golden-(\d+)\.(\d+)\.(\d+)-\d{4}-\d{2}-\d{2}$` matched against `os.listdir` (`scripts/golden_currency_gate.py:89,123`) — it never opens the directory, never looks for a log, never checks a sha, and never asks Gitea whether the package exists. Noticed 2026-08-31 while the 0.230.0 bake turned it from red to green: **I created the directory before the bake finished, and the gate would have passed at that moment.** **What the gate DOES say honestly:** its own success line already reads *"this checks the BAKE, not the vouch"* — so the vouch hole is declared. **This one is not:** nothing tells a reader the bake check is a filename check. **Class:** an instrument that cannot distinguish the thing from a label for the thing — the same shape as R-378's whole-field match and R-233's un-matchable grep, aimed this time at the release gate. **Exposure today is low** because the runbook produces a real evidence directory and two bakes in a row have; it is the NEXT hurried session that pays. | **OPEN — LOW** | R-242 | Make it read something the bake alone can produce: the `GOLDEN_SHA256=` line in the directory's `bake.log`, or a HEAD against the Gitea package URL for that version. Prefer the log — it keeps the gate offline and `--fast`. Ship a red-proof: an empty `golden-9.9.9-2026-01-01/` directory must FAIL. | CC |
|
||||
| **R-411** | **A background job DELETES the lock of a live customer restore, and logs it as a crash that did not happen.** MEASURED on demo-hp 2026-08-31 during the overnight soak, through the product's own endpoints — this is R-408's consequence, which until tonight had only been reasoned about. **The chain, every step observed:** (1) a customer full-restore runs `OffboxRestorePrepareFull` → `restic stats`, and **`restic stats` TAKES A REPOSITORY LOCK** (clean-room test: nothing else running, 4x stats, sampler reads `locks=1`); (2) the restore holds `opRunning` but **NOT `acquireRunning`** (R-408), so the integrity check is not blocked and runs concurrently; (3) the check meets that lock, and `resticStep` escalates to **`unlock --remove-all`** — caught by the argv sampler at **20:50:51 with `restore 3c11059b --target …` and `unlock --remove-all` in the SAME sample**; (4) the log says *"cleared a stale exclusive lock left by a previous crash (single-writer repo)"* — **there was no crash**, and `resticStep` cannot know there was, because it fires on ANY `repository is already locked`. **THE CUSTOMER-FACING CONSEQUENCE WAS CONTAINED, and that is R-359's guard working:** the check returned `ok:false` in 6.7 s and was classified **Unreachable, NOT damage** — *"the check could not run to a verdict (other) — NOT reported as damage"* — so no `backup_integrity_failed` and no customer mail. Due-ness was not advanced either, so it retries. **What is NOT contained:** a live operation's lock is deleted by a background job; the single-writer premise `resticStep`'s own comment rests on is false in this pairing; and that night's integrity check silently did not verify the store, with only a WARN. **The opposite direction is FENCED and was measured too:** five restores fired into a running check at 5/15/25/35/40 s offsets were ALL refused by `restoreOpBlocked` (`offbox_handlers.go:358`), zero restic invoked — so the hazard is reachable only restore-FIRST. | **OPEN — MEDIUM** | R-408, R-407, R-359 | Decide ONE way, and R-408 is the same decision: either `RestoreOffboxScratch` (and the full-restore preparation) takes `acquireRunning`, or `resticStep`'s escalation stops claiming a crash it cannot verify and refuses instead of removing. **Pin whichever is chosen with a test that reproduces this pairing** — a unit test on `resticStep` alone cannot see it. Evidence: `audits/DRILL-soak-2026-08-31/phase1-lock-collision/`. | CC |
|
||||
| **R-412** | **A lost recovery unit is rebuilt WITHOUT its volume dumps, and the off-site backup then ships that hollow unit and reports success.** OBSERVED on demo-hp 2026-08-31 during the soak, produced by the product with no construction — **this is the natural instance of R-403's shape that yesterday's R-87 session could not produce and had to hand-build.** **The chain, every step in the log:** (1) `opengist`'s primary unit was removed (a restore, a crash mid-capture or a remount does the same — R-403's own named causes); (2) the 5-minute capture rebuilt it — *"Recovery unit captured for opengist"* — with compose and manifest but **no volume tar**, manifest reading `db_dumps: []`, `volume_dumps: None` (the key absent entirely), **185 664 B → 4 382 B**; (3) the off-site backup pushed it and logged *"backed up opengist … 0 mandatory path(s)"* — a SUCCESS line over a backup containing none of the app's data; (4) the newest off-site snapshot `35ba9fe7` is now hollow. **THE CAPTURE IS NOT WRONG IN ISOLATION** — a capture describing an empty tree as empty is correct, and R-403 deliberately refused to guard it for that reason. **What is wrong is that nothing between the capture and the off-site push notices that a unit which HAD dumps yesterday has none today**, and the run reports success. The volume-dump leg runs on the backup schedule, not on capture, so the window is a whole cycle wide. | **OPEN — HIGH** | R-403, R-87, R-413 | Decide where the notice belongs: the capture (which R-403 fenced off), the off-site run's own pre-push phase (which already re-captures and could compare against what it is about to replace), or a shrink-detector on the primary. **Do NOT simply guard the capture** — 08 §8.2 records why that would make the manifest lie. Evidence: `audits/DRILL-soak-2026-08-31/phase2-guard-interactions/`. | CC |
|
||||
| **R-413** | **R-87's proof caught a naturally-produced hollow snapshot, end to end, unattended — the validation yesterday's session could only do with a declared hand-built fixture.** 2026-08-31 soak, demo-hp. After R-412's chain left `opengist`'s newest off-site snapshot hollow, the nightly proof rotated to it and returned **`verdict:"fail"`, `reason:"volumes_expected_none_captured"`, missing `opengist_data`**, logged *"READABLE AND EMPTY — the store is not damaged; the backup does not contain this app's data"*, and pushed **one** `offsite_proof_empty` at severity `error`. The four apps ahead of it in the rotation all passed, so the discrimination is real and not a constant fail. **This is recorded as a row rather than only as a report line because it upgrades a claim:** the capability map's R-87 row cites a CONSTRUCTED failing case; it can now cite a natural one. | **CLOSED 2026-08-31 — the claim it upgrades is recorded** | R-87, R-412 | Nothing to build. When the capability map is next touched, cite this instead of the constructed case. | CC |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user