Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.
VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.
R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.
R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.
R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.
R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.
R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.
R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.
Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.
EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.
Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
2.6 KiB
Phase 6 — the untouched observer's verdict. FAIL, and it is the night's most valuable finding.
demo-felhom was hand-deployed to 0.231.0 at 22:40 and not touched again. Its whole nightly cycle
was recorded by a follow stream on the PVE host.
Did every job run? Yes — all eight, on time.
| job | CEST | duration | outcome |
|---|---|---|---|
| db-dump | 02:30 | 674 ms | completed, db_dump_completed (info) |
| tier2-backup | 03:30 | 3 ms | completed — a no-op |
| fill-watch | 03:30 | 0 s | completed |
| metrics-prune | 04:00 | 3 ms | completed |
| offbox-backup | 04:15 | 46.9 s | completed |
| offsite-abandon-sweep | 05:10 | 0 s | completed |
| offsite-proof | 05:30 | 2.6 s | REFUSED — could not run |
| offsite-integrity | 06:00 | 0 s | not due (last success 24 h ago) — correct due-ness |
Did any two overlap, and did the flag hold?
No two scheduled jobs overlapped on this box — the schedule spaces them 45–90 min apart and the
longest took 47 s. The flag was never contended here, so this box proves the schedule, not the
lock. Contention was tested on demo-hp (Phase 1).
Did the proof pick one app and record a snapshot? No — and that is the finding.
03:30:02 [WARN] [offbox] proof: opengist has nowhere to restore to: nincs regisztralt
adatmeghajto, ezert nincs hova visszaallitani — a meghajtok megvannak, csak ujra kell...
Cause established, not inferred: storage_paths: [] — zero registered storage paths, so
offboxRestoreScratchDir has nowhere to write a scratch and returns R-252's refusal. The same absence
explains tier2-backup at 3 ms: no second drive to mirror to.
last_proof_result is ABSENT and proved_snapshots is ABSENT — the error path reaches no
verdict, so nothing is recorded, so the hub cannot tell this from a controller too old to have the
feature. Filed R-414.
The box is not unprotected: its off-site backup ran normally in 46.9 s, because recovery units live on the system data path, which needs no registration. It is the PROOF that cannot run.
Did the integrity check run at full depth? No — and correctly.
not due (last successful check 24h0m0s ago). The 7-day max age had not elapsed. Due-ness working as
designed; this box's last real full-depth check was 2026-08-31 04:00.
Did anything alarm on a healthy box?
No. One event in the whole night — db_dump_completed (info), which severityNotifies drops.
Zero ERROR lines. One WARN, and it is the R-414 refusal above. On a box left completely alone for
seven hours, the product was silent, and silence was the correct answer.