585ed654b4
gates / gates (push) Failing after 17s
Overnight soak on demo-hp, phases 0-4 of 7. Evidence documentation/audits/DRILL-soak-2026-08-31/. No production code written, no golden, no version bump - findings only. PHASE 1 FAIL - R-411. The R-408 hazard was only ever reasoned about; tonight it was measured through the product's own endpoints. restic stats TAKES A LOCK (clean-room: nothing else running, 4x stats, sampler reads locks=1). A customer full-restore runs stats in its preparation while holding NO acquireRunning, so the integrity check is not blocked, runs, meets that lock, and resticStep escalates to unlock --remove-all - the sampler caught "restore ..." and "unlock --remove-all" in the SAME sample at 20:50:51. The log calls it "a stale exclusive lock left by a previous crash"; there was no crash. THE CUSTOMER-FACING CONSEQUENCE IS CONTAINED and that is R-359 working: the check was classified Unreachable, NOT damage, so no alarm and due-ness held. The opposite direction is FENCED - five restores fired into a running check at 5/15/25/35/40s were all refused by restoreOpBlocked with zero restic invoked. PHASE 2 FAIL - R-412, and it is the natural instance of R-403's shape that yesterday's session had to hand-build. A lost primary unit is rebuilt by the capture WITHOUT its volume dumps (185664 B -> 4382 B, volume_dumps: None), and the off-site backup then ships it and logs "backed up opengist ... 0 mandatory path(s)" - a success line over a backup holding none of the app's data. PHASE 2 also R-413: the R-87 proof CAUGHT that hollow snapshot unattended - verdict "fail", volumes_expected_none_captured: opengist_data, one offsite_proof_empty at severity error, while the four apps ahead of it in the rotation passed. 2.3 stale marker PASS, 2.4 two restores PASS, 2.5 corrected the runbook's premise (RunTier2 has zero acquireRunning and zero restic references - a local mirror cannot collide with the repo, so the proof correctly does not skip for it). PHASE 3 PASS - R-357 live-validated at last, six days owed. A real full filesystem (1900544 B free vs 5878421 B needed) on a separate 1 TB device that does not back Docker. Refused BEFORE StopStack: app "Up 6 minutes (healthy)" unchanged, 0 safety dumps, live userdata tree fingerprint 5a9b2db75dfd4ba4131adb3255670472 unchanged, message names both numbers in Hungarian. Ballast removed, the SAME restore then worked and the fingerprint is still identical. On the same full disk the Tier-2 run, the proof and the integrity check all behaved: the proof refused before any download through the shared unitOnlyHeadroom gate extracted today, reached no verdict and did not alarm. PHASE 4 PASS - the false-alarm control the whole R-87 design rests on now EXISTS. Found by reading all 53 catalogue composes: bentopdf is the only template with neither a database service nor a named volume. Deployed, backed up, proved: PASSED in 2.224s with zero alarms. Four of my own instrument errors were caught by their own controls before any result was believed: a hub log line used as a controller positive control, a grep pattern that missed a registered job, a heredoc that mangled a planted marker, and a catalogue scan that read 0 of 0 apps from the wrong path. Phases 5-7 to follow: the real nightly cycle with injected faults, the untouched observer, and teardown.
45 lines
3.0 KiB
Markdown
45 lines
3.0 KiB
Markdown
# Phase 2 — the new guards against each other. VERDICT: **FAIL (one real defect found)**
|
|
|
|
Five of six sub-tests ran; 2.6 needs Phase 4's app and is deferred there.
|
|
|
|
| # | test | verdict | one sentence |
|
|
|---|---|---|---|
|
|
| 2.1 | rehydrate vs capture | **FAIL — R-412** | the capture won, and rebuilt the primary **hollow**: 185 664 B → 4 382 B, `volume_dumps: None` |
|
|
| 2.2 | rehydrate then mirror | **PASS (partial)** | the secondary survived byte-identical (`3e26592f…`), but the mirror run did not target this app, so R-403's guard was **not** exercised — deferred to Phase 5's 03:30 job, which is the natural test |
|
|
| 2.3 | stale scratch marker | **PASS** | a planted `deadbeef / full:true / 2020` marker was fully replaced by `180c5933 / full:false / now` |
|
|
| 2.4 | two restores at once | **PASS** | exactly one refusal, exactly one completion |
|
|
| 2.5 | proof during other runs | **PASS, with the runbook's premise corrected** | see below |
|
|
| 2.6 | legitimately-empty unit | deferred to Phase 4 | no such app exists on the box yet |
|
|
|
|
## The defect — R-412, and it is the natural instance of R-403's shape
|
|
|
|
I removed `opengist`'s primary unit. The 5-minute capture rebuilt it with compose and manifest and
|
|
**no volume tar**; the off-site backup then pushed that hollow unit and logged
|
|
`backed up opengist (… 0 mandatory path(s))` — **a success line over a backup holding none of the
|
|
app's data**. The newest off-site snapshot `35ba9fe7` is hollow.
|
|
|
|
**The capture is not wrong in isolation** — R-403 deliberately refused to guard it, because a capture
|
|
describing an empty tree as empty is correct. What is wrong is that **nothing between the capture and
|
|
the push notices a unit that had dumps yesterday and has none today.**
|
|
|
|
## And then the thing that makes tonight worth it — R-413
|
|
|
|
The R-87 proof, rotating normally, reached `opengist` and returned **`verdict:"fail"`,
|
|
`volumes_expected_none_captured: opengist_data`**, with one `offsite_proof_empty` at severity `error`.
|
|
The four apps ahead of it in the rotation all passed. **Yesterday's session had to hand-build its
|
|
failing case and declare it; tonight the product produced one by itself and the guard caught it.**
|
|
|
|
## 2.5 — the runbook's premise was wrong, and the behaviour is right
|
|
|
|
The runbook expected the proof to skip during a Tier-2 run. **It did not skip** (`stack:privatebin`,
|
|
`verdict:pass`, due-ness advanced 6→7). Established rather than assumed: `RunTier2` contains **zero**
|
|
`acquireRunning` calls and **zero** references to restic — it is a purely local rsync mirror that
|
|
cannot collide with the repository. The proof correctly skipped during the **off-site backup**
|
|
(`skipped:true`, due-ness held at 7), which is the operation that does share the repo.
|
|
|
|
## Deliberately left in place
|
|
|
|
`opengist` stays hollow-primary / complete-secondary overnight. That is **exactly** Phase 5's first
|
|
injected fault, and letting the real 03:30 `tier2-backup` meet it is a better test than forcing one.
|
|
Restored in Phase 7. Its live app data is untouched and healthy throughout.
|