Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.
VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.
R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.
R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.
R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.
R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.
R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.
R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.
Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.
EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.
Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
7.7 KiB
Overnight soak — the week's five new guards, run against each other
22:39 → 06:10 CEST, 2026-08-31/09-01. demo-hp the victim, demo-felhom the untouched observer.
No production code written. No golden baked. No version bumped.
| # | phase | verdict | the one sentence |
|---|---|---|---|
| 1 | lock collision | FAIL | a background job deletes the lock of a live customer restore and logs a crash that did not happen — customer-facing consequence contained |
| 2 | guard interactions | PASS with one defect | 4 of 5 guards clean; a unit lost inside a run ships hollow and is reported a success |
| 3 | R-357 full disk | PASS | six days owed, now proven against a real full filesystem — refused before the app went down, data byte-identical |
| 4 | proof edges | PASS | the false-alarm control the whole R-87 design rests on now exists and stays silent |
| 5 | mutated cycle | PASS | every job ran; the R-403 guard fired and named itself; the nightly proof fired unattended for the first time |
| 6 | observer | FAIL | on a box with no registered drive the nightly proof cannot run at all, every night, with only a WARN |
| 7 | teardown | PASS | all three layers clean on both boxes; both healthy on 0.231.0 |
1. What surprised me, worst first
The proof is inert on a whole class of box (R-414) — and only the untouched observer could find it
demo-felhom has zero registered storage paths. At 05:30 the new job fired for the first time
unattended and refused: "proof: opengist has nowhere to restore to — nincs regisztrált
adatmeghajtó". It will do that every night, forever, and the only signal is a WARN.
Worse than a failure: the error path reaches no verdict, so last_proof_result stays absent — and
absent is also what a pre-0.231.0 controller sends. The hub cannot tell "never ran" from "not
deployed". That is the exact StatsKnown trap the field was designed to avoid, reappearing one
level up. The same absence explains that box's tier2-backup finishing in 3 ms.
The box is not unprotected — its off-site backup ran normally in 46.9 s.
unlock --remove-all really does run against a live operation (R-411)
Measured, not reasoned. restic stats takes a repository lock (clean-room test). A customer
full-restore runs stats while holding no single-writer flag, so the integrity check is not
blocked, runs, meets that lock, and escalates. The sampler caught restore … and
unlock --remove-all in the same sample, and the log said "a stale exclusive lock left by a
previous crash". There was no crash.
Contained: the check was classified unreachable, not damage — no false "your backups are damaged" mail, and due-ness held. The opposite direction is fenced: five restores fired into a running check were all refused, zero restic invoked.
I over-claimed a finding and had to correct it (R-412)
Filed HIGH on a mechanism I had not finished establishing. Overnight measurement showed the off-site run has its own pre-push dump leg, so a hollow unit is repaired before it ships — proven on two apps, and confirmed by pulling the snapshot back out of the store. Corrected to LOW, with the over-claim written into the row rather than quietly edited away.
What survives is narrower and real: a unit destroyed inside a run, after that app's dump leg, ships
hollow and the run logs backed up opengist (… 0 mandatory path(s)) — a success line over a backup
holding none of the app's data.
A quiet one worth saying: three attempts to damage a backup were repaired by the product
A stopped app's tar was re-made; a deleted unit was rebuilt; a corrupted manifest was rewritten — each by the run's own capture, before any push. That is reassuring, and it is why the hollow snapshot needed a race to happen at all.
2. Findings filed
| id | severity | what |
|---|---|---|
| R-411 | MEDIUM | a background job deletes a live restore's lock and calls it a crash |
| R-412 | LOW (was HIGH — corrected) | a unit lost inside a run ships hollow and reports success |
| R-413 | CLOSED | the R-87 proof caught a product-produced hollow snapshot unattended |
| R-414 | MEDIUM | the nightly proof is inert on a box with no registered drive |
Register: OPEN-ITEMS.md 173 → 174 rows. Evidence:
documentation/audits/DRILL-soak-2026-08-31/, seven phase directories.
3. What I could not test, and why
- "Mark a drive disconnected" (Phase 5). No endpoint reaches
SetDisconnected; a hand-set flag would be reverted by the live monitor before 04:15, so the run would have read as a passing test of a fault that was not present. Unmounting a live drive risks a wedged mount that needs a reboot. 04:15 was observed against a real injected fault instead. - "Corrupt a manifest before the proof." Phase 4 established the capture rewrites it before any
push, so nothing I do to the live manifest can change what the proof sees. Covered by
TestR87_UnparseableManifestFails. - Two failing apps in one hour → one mail. Only two alarms fired all night and they were 1 h 42 m apart, so the coarse cooldown was never exercised. UNDETERMINED.
- The 06:00 integrity check at full depth on
demo-hp. It completed in 0 s — not due, because my own manual runs during Phases 1–3 had already advanced its due-ness. My doing, and correct behaviour.
4. Left changed on the boxes
demo-felhomis hand-deployed to 0.231.0 (was 0.230.0). Reversible; the golden bake supersedes it. Recorded here because a hand-deployed box nobody records is how a fleet drifts.bentopdfis deployed ondemo-hpand I recommend KEEPING it. It is the only template of 53 with neither a database nor a named volume, so it is the only possible live subject for R-87's false-alarm control. Deleting it deletes the control.- A hollow
opengistsnapshot (35ba9fe7) remains in the off-site store as history. Its newest snapshot is sound and the proof passes it. - Nothing else. All ballast, scratches, markers, probe scripts and credentials removed from all three
layers on both boxes;
/rooton both hosts holds only.cacheand.ssh; both container/tmpdirectories are empty.
5. My own mistakes, by name — eight, all caught by their own controls
- Used a hub log line as the positive control for a controller recorder.
- A grep pattern that reported a registered job as missing when it was there.
- A heredoc that mangled a planted marker, and an escaping error that reported
"full":falseas a failure when the file plainly read"full":false. - A catalogue scan that read 0 of 0 apps from the wrong path.
- A wait that matched my own earlier output replayed by the log follower — a false "job observed".
- A time guard
[ 2334 -ge 0315 ]that fired the Phase 5 injection four hours early, changing the drill. - Filing R-412 at HIGH on a mechanism I had not finished measuring — the worst one, because it went into the register wrong.
- A baseline grep whose control failed, nearly making a 10-day-old stale directory look like tonight's damage; mtime settled it.
Every zero reported above was earned by a positive control first. That is the only reason the list is eight and not eight-plus-a-wrong-verdict.
6. Owed to Viktor in the morning
- A golden carrying 0.231.0 — still owed from yesterday; the fleet floor is 0.230.0.
- A decision on
bentopdf— keep it as the permanent false-alarm control (my recommendation), or remove it. - R-414 is the one worth reading first: a feature that shipped yesterday does not run at all on a
box shaped like
demo-felhom.