Files
felhom.eu/documentation/audits/DRILL-soak-2026-08-31/phase5-mutated-cycle/00-VERDICT.md
T
admin cee8f70e98
gates / gates (push) Failing after 17s
soak drill COMPLETE: 7 phases, 4 findings, the observer found the worst one (R-414)
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.

VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.

R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.

R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.

R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.

R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.

R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.

R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.

Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.

EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.

Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
2026-09-01 06:07:13 +02:00

2.7 KiB
Raw Blame History

Phase 5 — demo-hp's own nightly cycle, mutated. PASS

Every scheduled job ran, on time, with a light load (a unit restore every 20 min) contending for the single-writer flag across the whole window.

job CEST duration outcome
db-dump 02:30 1 m 24.7 s completed — and it re-made the volume tars, healing the injected fault
tier2-backup 03:30 5.9 s completed, 9 apps
fill-watch 03:30 0 s completed
metrics-prune 04:00 0 s completed
offbox-backup 04:15 2 m 57.7 s completed — its own pre-push dump leg repaired a hollow unit
offsite-abandon-sweep 05:10 0 s completed, clean beside the rest
offsite-proof 05:30 5.0 s completed — bentopdf PASSED on 9d002b38 in 2.315 s
offsite-integrity 06:00 0 s not due — my own manual runs had advanced due-ness

The headline: the nightly proof fired unattended, for the first time

Yesterday's report listed "the unattended nightly firing" as not yet live-validated. It is now: the job fired on its own schedule at 05:30, picked one app, proved it in 2.315 s, and scheduled the next run for 2026-09-02 05:30 CEST. One app per night held.

The R-403 guard: PASS, forced after the natural test evaporated

privatebin was injected hollow at 23:34 to meet the 03:30 mirror. The 02:30 db-dump re-made its tar, so by 03:30 the primary was sound and the guard had nothing to refuse. Forced instead on calibre-web through the real Tier-2 path — and the guard fired and named itself:

Tier 2 calibre-web: unit leg SKIPPED — the recovery unit on the source drive lists no database dumps
and no volume tars, while the existing copy ... does. The copy was PRESERVED rather than replaced
with an empty one (R-403). The other legs continue.

Secondary byte-identical: 23 files, 5 808 704 B, tar sha d7e7f422…. Negative control at teardown: with every unit sound, the same job skipped 0 unit legs.

Alarms across the whole mutated night

One WARN — the R-403 skip above. Zero ERRORs. Every event pushed was info (db_dump_completed, crossdrive_completed ×12), all dropped by severityNotifies. No customer was mailed anything.

Half-marks, stated

  • The runbook wanted "the run does not report a plain success" when a leg is skipped. The log does not — it states the skip and the reason. The event does: crossdrive_completed (info) — Másodlagos mentés elkészült: calibre-web, no mention of the skipped leg. It is not delivered (info), so nobody is misled by mail, but an event list would read as a clean completion.
  • Two of four injections were not performed; reasons in 03-injections-not-performed.md.