Files
felhom.eu/documentation/audits/DRILL-soak-2026-08-31/phase3-r357-full-disk/00-VERDICT.md
T
admin 585ed654b4
gates / gates (push) Failing after 17s
soak drill phases 0-4: R-408's hazard is REAL and reachable; R-357 finally live (R-411..R-413)
Overnight soak on demo-hp, phases 0-4 of 7. Evidence
documentation/audits/DRILL-soak-2026-08-31/. No production code written, no golden,
no version bump - findings only.

PHASE 1 FAIL - R-411. The R-408 hazard was only ever reasoned about; tonight it was
measured through the product's own endpoints. restic stats TAKES A LOCK (clean-room:
nothing else running, 4x stats, sampler reads locks=1). A customer full-restore runs
stats in its preparation while holding NO acquireRunning, so the integrity check is not
blocked, runs, meets that lock, and resticStep escalates to unlock --remove-all - the
sampler caught "restore ..." and "unlock --remove-all" in the SAME sample at 20:50:51.
The log calls it "a stale exclusive lock left by a previous crash"; there was no crash.
THE CUSTOMER-FACING CONSEQUENCE IS CONTAINED and that is R-359 working: the check was
classified Unreachable, NOT damage, so no alarm and due-ness held. The opposite direction
is FENCED - five restores fired into a running check at 5/15/25/35/40s were all refused by
restoreOpBlocked with zero restic invoked.

PHASE 2 FAIL - R-412, and it is the natural instance of R-403's shape that yesterday's
session had to hand-build. A lost primary unit is rebuilt by the capture WITHOUT its
volume dumps (185664 B -> 4382 B, volume_dumps: None), and the off-site backup then ships
it and logs "backed up opengist ... 0 mandatory path(s)" - a success line over a backup
holding none of the app's data.

PHASE 2 also R-413: the R-87 proof CAUGHT that hollow snapshot unattended -
verdict "fail", volumes_expected_none_captured: opengist_data, one offsite_proof_empty at
severity error, while the four apps ahead of it in the rotation passed.
2.3 stale marker PASS, 2.4 two restores PASS, 2.5 corrected the runbook's premise
(RunTier2 has zero acquireRunning and zero restic references - a local mirror cannot
collide with the repo, so the proof correctly does not skip for it).

PHASE 3 PASS - R-357 live-validated at last, six days owed. A real full filesystem
(1900544 B free vs 5878421 B needed) on a separate 1 TB device that does not back Docker.
Refused BEFORE StopStack: app "Up 6 minutes (healthy)" unchanged, 0 safety dumps, live
userdata tree fingerprint 5a9b2db75dfd4ba4131adb3255670472 unchanged, message names both
numbers in Hungarian. Ballast removed, the SAME restore then worked and the fingerprint is
still identical. On the same full disk the Tier-2 run, the proof and the integrity check
all behaved: the proof refused before any download through the shared unitOnlyHeadroom
gate extracted today, reached no verdict and did not alarm.

PHASE 4 PASS - the false-alarm control the whole R-87 design rests on now EXISTS. Found by
reading all 53 catalogue composes: bentopdf is the only template with neither a database
service nor a named volume. Deployed, backed up, proved: PASSED in 2.224s with zero alarms.

Four of my own instrument errors were caught by their own controls before any result was
believed: a hub log line used as a controller positive control, a grep pattern that missed
a registered job, a heredoc that mangled a planted marker, and a catalogue scan that read
0 of 0 apps from the wrong path.

Phases 5-7 to follow: the real nightly cycle with injected faults, the untouched observer,
and teardown.
2026-08-31 23:33:32 +02:00

2.6 KiB

Phase 3 — R-357 against a real full disk. VERDICT: PASS

Owed since 2026-08-25. Until tonight the destructive restore's free-space gate had only ever been proven by seam tests. It has now met a genuinely full filesystem.

Method

calibre-web lives on hdd_1 — a separate 1 TB NVMe device that does not back Docker, chosen so filling it could not take the box down. fallocate reduced free space to 1 900 544 B against a measured need of 5 878 421 B.

The non-effects, which are what matter

assertion result
refused before StopStack — the app never went down Up 6 minutes (healthy) before and after, unchanged; 0 stop-related log lines
no safety dump written 0 files
live data byte-identical 22 files, tree fingerprint 5a9b2db75dfd4ba4131adb3255670472, unchanged
the message names the need and the free space, in Hungarian, not raw rsync Nincs eleg szabad hely a visszaallitashoz (5.6 MB szukseges, 1.8 MB szabad).
the log names both numbers and the non-effect REFUSING the destructive restore — need 5878421 B, free 1900544 B on /mnt/felhom-drives/hdd_1; the app was NOT stopped

And it works when there IS space — the other half, without which the above proves nothing

Ballast removed → the same restore ran: reconstituted calibre-web from snapshot ece64c87: 0 file(s) placed, 1 volume(s) replayed, app back in 37 s, and the userdata tree fingerprint is 5a9b2db75dfd4ba4131adb3255670472 — identical to before the whole drill.

The other three guards on the same full disk — none failed open

job behaviour verdict
Tier-2 run copied 3 apps successfully — its destination is on the other filesystem (/var/lib/felhom, 55 GB free), so a full hdd_1 correctly did not stop it. Events crossdrive_completed (info) — mails nobody correct
R-87 proof refused before any download: Nincs eleg szabad hely (2.0 GB szukseges, 1.8 MB szabad). Reached no verdict, did not alarm, and said so: "A visszaallitas nem futott le — a mentesrol semmi nem derult ki" correct — and this is the shared unitOnlyHeadroom gate, extracted today so the proof and the customer restore refuse at one floor with one sentence, now live-validated
integrity check PASSED in 53.8 s at 100 % depth — restic's cache lives on the container filesystem, not hdd_1 correct

No customer alarm fired from any of them. The only events pushed were three crossdrive_completed at severity info, which severityNotifies drops.