cee8f70e98
gates / gates (push) Failing after 17s
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.
VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.
R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.
R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.
R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.
R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.
R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.
R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.
Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.
EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.
Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
53 lines
2.6 KiB
Markdown
53 lines
2.6 KiB
Markdown
# Phase 6 — the untouched observer's verdict. **FAIL, and it is the night's most valuable finding.**
|
||
|
||
`demo-felhom` was hand-deployed to 0.231.0 at 22:40 and **not touched again**. Its whole nightly cycle
|
||
was recorded by a follow stream on the PVE host.
|
||
|
||
## Did every job run? Yes — all eight, on time.
|
||
|
||
| job | CEST | duration | outcome |
|
||
|---|---|---|---|
|
||
| db-dump | 02:30 | 674 ms | completed, `db_dump_completed (info)` |
|
||
| tier2-backup | 03:30 | **3 ms** | completed — **a no-op** |
|
||
| fill-watch | 03:30 | 0 s | completed |
|
||
| metrics-prune | 04:00 | 3 ms | completed |
|
||
| offbox-backup | 04:15 | **46.9 s** | completed |
|
||
| offsite-abandon-sweep | 05:10 | 0 s | completed |
|
||
| **offsite-proof** | **05:30** | 2.6 s | **REFUSED — could not run** |
|
||
| offsite-integrity | 06:00 | 0 s | **not due** (last success 24 h ago) — correct due-ness |
|
||
|
||
## Did any two overlap, and did the flag hold?
|
||
|
||
No two scheduled jobs overlapped on this box — the schedule spaces them 45–90 min apart and the
|
||
longest took 47 s. The flag was never contended here, so this box proves the **schedule**, not the
|
||
lock. Contention was tested on `demo-hp` (Phase 1).
|
||
|
||
## Did the proof pick one app and record a snapshot? **No — and that is the finding.**
|
||
|
||
```
|
||
03:30:02 [WARN] [offbox] proof: opengist has nowhere to restore to: nincs regisztralt
|
||
adatmeghajto, ezert nincs hova visszaallitani — a meghajtok megvannak, csak ujra kell...
|
||
```
|
||
|
||
**Cause established, not inferred:** `storage_paths: []` — zero registered storage paths, so
|
||
`offboxRestoreScratchDir` has nowhere to write a scratch and returns R-252's refusal. The same absence
|
||
explains `tier2-backup` at 3 ms: no second drive to mirror to.
|
||
|
||
`last_proof_result` is **ABSENT** and `proved_snapshots` is **ABSENT** — the error path reaches no
|
||
verdict, so nothing is recorded, so the hub cannot tell this from a controller too old to have the
|
||
feature. Filed **R-414**.
|
||
|
||
**The box is not unprotected**: its off-site backup ran normally in 46.9 s, because recovery units
|
||
live on the system data path, which needs no registration. It is the PROOF that cannot run.
|
||
|
||
## Did the integrity check run at full depth? No — and correctly.
|
||
|
||
`not due (last successful check 24h0m0s ago)`. The 7-day max age had not elapsed. Due-ness working as
|
||
designed; this box's last real full-depth check was 2026-08-31 04:00.
|
||
|
||
## **Did anything alarm on a healthy box?**
|
||
|
||
**No.** One event in the whole night — `db_dump_completed (info)`, which `severityNotifies` drops.
|
||
**Zero ERROR lines. One WARN**, and it is the R-414 refusal above. On a box left completely alone for
|
||
seven hours, the product was silent, and silence was the correct answer.
|