cee8f70e98
gates / gates (push) Failing after 17s
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.
VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.
R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.
R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.
R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.
R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.
R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.
R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.
Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.
EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.
Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
51 lines
2.7 KiB
Markdown
51 lines
2.7 KiB
Markdown
# Phase 5 — `demo-hp`'s own nightly cycle, mutated. **PASS**
|
||
|
||
Every scheduled job ran, on time, with a light load (a unit restore every 20 min) contending for the
|
||
single-writer flag across the whole window.
|
||
|
||
| job | CEST | duration | outcome |
|
||
|---|---|---|---|
|
||
| db-dump | 02:30 | 1 m 24.7 s | completed — **and it re-made the volume tars**, healing the injected fault |
|
||
| tier2-backup | 03:30 | 5.9 s | completed, 9 apps |
|
||
| fill-watch | 03:30 | 0 s | completed |
|
||
| metrics-prune | 04:00 | 0 s | completed |
|
||
| offbox-backup | 04:15 | 2 m 57.7 s | completed — **its own pre-push dump leg repaired a hollow unit** |
|
||
| offsite-abandon-sweep | 05:10 | 0 s | completed, clean beside the rest |
|
||
| **offsite-proof** | **05:30** | **5.0 s** | **completed — `bentopdf` PASSED on `9d002b38` in 2.315 s** |
|
||
| offsite-integrity | 06:00 | 0 s | **not due** — my own manual runs had advanced due-ness |
|
||
|
||
## The headline: the nightly proof fired unattended, for the first time
|
||
|
||
Yesterday's report listed *"the unattended nightly firing"* as **not yet live-validated**. It is now:
|
||
the job fired on its own schedule at 05:30, picked one app, proved it in 2.315 s, and scheduled the
|
||
next run for `2026-09-02 05:30 CEST`. **One app per night held.**
|
||
|
||
## The R-403 guard: PASS, forced after the natural test evaporated
|
||
|
||
`privatebin` was injected hollow at 23:34 to meet the 03:30 mirror. **The 02:30 db-dump re-made its
|
||
tar**, so by 03:30 the primary was sound and the guard had nothing to refuse. Forced instead on
|
||
`calibre-web` through the real Tier-2 path — and the guard fired and named itself:
|
||
|
||
```
|
||
Tier 2 calibre-web: unit leg SKIPPED — the recovery unit on the source drive lists no database dumps
|
||
and no volume tars, while the existing copy ... does. The copy was PRESERVED rather than replaced
|
||
with an empty one (R-403). The other legs continue.
|
||
```
|
||
|
||
Secondary byte-identical: 23 files, 5 808 704 B, tar sha `d7e7f422…`. **Negative control** at
|
||
teardown: with every unit sound, the same job skipped **0** unit legs.
|
||
|
||
## Alarms across the whole mutated night
|
||
|
||
**One WARN** — the R-403 skip above. **Zero ERRORs.** Every event pushed was `info`
|
||
(`db_dump_completed`, `crossdrive_completed` ×12), all dropped by `severityNotifies`. **No customer
|
||
was mailed anything.**
|
||
|
||
## Half-marks, stated
|
||
|
||
- The runbook wanted *"the run does not report a plain success"* when a leg is skipped. **The log does
|
||
not** — it states the skip and the reason. **The event does**: `crossdrive_completed (info) —
|
||
Másodlagos mentés elkészült: calibre-web`, no mention of the skipped leg. It is not delivered
|
||
(`info`), so nobody is misled by mail, but an event list would read as a clean completion.
|
||
- Two of four injections were not performed; reasons in `03-injections-not-performed.md`.
|