# Overnight soak — the week's five new guards, run against each other **22:39 → 06:10 CEST, 2026-08-31/09-01. `demo-hp` the victim, `demo-felhom` the untouched observer.** No production code written. No golden baked. No version bumped. | # | phase | verdict | the one sentence | |---|---|---|---| | 1 | lock collision | **FAIL** | a background job **deletes the lock of a live customer restore** and logs a crash that did not happen — customer-facing consequence contained | | 2 | guard interactions | **PASS with one defect** | 4 of 5 guards clean; a unit lost *inside* a run ships hollow and is reported a success | | 3 | R-357 full disk | **PASS** | six days owed, now proven against a real full filesystem — refused before the app went down, data byte-identical | | 4 | proof edges | **PASS** | the false-alarm control the whole R-87 design rests on now exists and stays silent | | 5 | mutated cycle | **PASS** | every job ran; the R-403 guard fired and named itself; the nightly proof fired unattended for the first time | | 6 | observer | **FAIL** | on a box with no registered drive the nightly proof **cannot run at all**, every night, with only a WARN | | 7 | teardown | **PASS** | all three layers clean on both boxes; both healthy on 0.231.0 | --- ## 1. What surprised me, worst first ### The proof is inert on a whole class of box (R-414) — and only the untouched observer could find it `demo-felhom` has **zero registered storage paths**. At 05:30 the new job fired for the first time unattended and refused: *"proof: opengist has nowhere to restore to — nincs regisztrált adatmeghajtó"*. It will do that **every night, forever**, and the only signal is a `WARN`. Worse than a failure: the error path reaches no verdict, so `last_proof_result` stays **absent** — and absent is also what a pre-0.231.0 controller sends. **The hub cannot tell "never ran" from "not deployed".** That is the exact `StatsKnown` trap the field was designed to avoid, reappearing one level up. The same absence explains that box's `tier2-backup` finishing in **3 ms**. The box is not unprotected — its off-site backup ran normally in 46.9 s. ### `unlock --remove-all` really does run against a live operation (R-411) Measured, not reasoned. `restic stats` **takes a repository lock** (clean-room test). A customer full-restore runs `stats` while holding **no** single-writer flag, so the integrity check is not blocked, runs, meets that lock, and escalates. The sampler caught `restore …` and `unlock --remove-all` **in the same sample**, and the log said *"a stale exclusive lock left by a previous crash"*. There was no crash. **Contained**: the check was classified *unreachable*, not damage — no false "your backups are damaged" mail, and due-ness held. The opposite direction is fenced: five restores fired into a running check were all refused, zero restic invoked. ### I over-claimed a finding and had to correct it (R-412) Filed **HIGH** on a mechanism I had not finished establishing. Overnight measurement showed the off-site run has its **own** pre-push dump leg, so a hollow unit is **repaired before it ships** — proven on two apps, and confirmed by pulling the snapshot back out of the store. **Corrected to LOW**, with the over-claim written into the row rather than quietly edited away. What survives is narrower and real: a unit destroyed *inside* a run, after that app's dump leg, ships hollow and the run logs `backed up opengist (… 0 mandatory path(s))` — a success line over a backup holding none of the app's data. ### A quiet one worth saying: three attempts to damage a backup were repaired by the product A stopped app's tar was re-made; a deleted unit was rebuilt; a corrupted manifest was rewritten — each by the run's own capture, before any push. That is reassuring, and it is why the hollow snapshot needed a race to happen at all. --- ## 2. Findings filed | id | severity | what | |---|---|---| | **R-411** | MEDIUM | a background job deletes a live restore's lock and calls it a crash | | **R-412** | LOW *(was HIGH — corrected)* | a unit lost inside a run ships hollow and reports success | | **R-413** | CLOSED | the R-87 proof caught a product-produced hollow snapshot unattended | | **R-414** | MEDIUM | the nightly proof is inert on a box with no registered drive | Register: `OPEN-ITEMS.md` **173 → 174** rows. Evidence: `documentation/audits/DRILL-soak-2026-08-31/`, seven phase directories. --- ## 3. What I could not test, and why - **"Mark a drive disconnected" (Phase 5).** No endpoint reaches `SetDisconnected`; a hand-set flag would be reverted by the live monitor before 04:15, so the run would have read as a passing test of a fault that was not present. Unmounting a live drive risks a wedged mount that needs a reboot. 04:15 was observed against a real injected fault instead. - **"Corrupt a manifest before the proof."** Phase 4 established the capture rewrites it before any push, so nothing I do to the live manifest can change what the proof sees. Covered by `TestR87_UnparseableManifestFails`. - **Two failing apps in one hour → one mail.** Only two alarms fired all night and they were 1 h 42 m apart, so the coarse cooldown was never exercised. **UNDETERMINED.** - **The 06:00 integrity check at full depth on `demo-hp`.** It completed in **0 s — not due**, because my own manual runs during Phases 1–3 had already advanced its due-ness. My doing, and correct behaviour. --- ## 4. Left changed on the boxes - **`demo-felhom` is hand-deployed to 0.231.0** (was 0.230.0). Reversible; the golden bake supersedes it. **Recorded here because a hand-deployed box nobody records is how a fleet drifts.** - **`bentopdf` is deployed on `demo-hp` and I recommend KEEPING it.** It is the **only** template of 53 with neither a database nor a named volume, so it is the only possible live subject for R-87's false-alarm control. Deleting it deletes the control. - A hollow `opengist` snapshot (`35ba9fe7`) remains in the off-site store as history. Its newest snapshot is sound and the proof passes it. - Nothing else. All ballast, scratches, markers, probe scripts and credentials removed from all three layers on both boxes; `/root` on both hosts holds only `.cache` and `.ssh`; both container `/tmp` directories are empty. --- ## 5. My own mistakes, by name — eight, all caught by their own controls 1. Used a **hub** log line as the positive control for a **controller** recorder. 2. A grep pattern that reported a registered job as missing when it was there. 3. A heredoc that mangled a planted marker, and an escaping error that reported `"full":false` as a failure when the file plainly read `"full":false`. 4. A catalogue scan that read **0 of 0** apps from the wrong path. 5. A wait that matched my **own earlier output** replayed by the log follower — a false "job observed". 6. **A time guard `[ 2334 -ge 0315 ]` that fired the Phase 5 injection four hours early**, changing the drill. 7. **Filing R-412 at HIGH on a mechanism I had not finished measuring** — the worst one, because it went into the register wrong. 8. A baseline grep whose control failed, nearly making a 10-day-old stale directory look like tonight's damage; mtime settled it. Every zero reported above was earned by a positive control first. That is the only reason the list is eight and not eight-plus-a-wrong-verdict. --- ## 6. Owed to Viktor in the morning 1. **A golden carrying 0.231.0** — still owed from yesterday; the fleet floor is 0.230.0. 2. **A decision on `bentopdf`** — keep it as the permanent false-alarm control (my recommendation), or remove it. 3. **R-414** is the one worth reading first: a feature that shipped yesterday does not run at all on a box shaped like `demo-felhom`.