Files
felhom.eu/documentation/audits/DRILL-soak-2026-08-31/phase4-proof-edges/00-VERDICT.md
T
admin 585ed654b4
gates / gates (push) Failing after 17s
soak drill phases 0-4: R-408's hazard is REAL and reachable; R-357 finally live (R-411..R-413)
Overnight soak on demo-hp, phases 0-4 of 7. Evidence
documentation/audits/DRILL-soak-2026-08-31/. No production code written, no golden,
no version bump - findings only.

PHASE 1 FAIL - R-411. The R-408 hazard was only ever reasoned about; tonight it was
measured through the product's own endpoints. restic stats TAKES A LOCK (clean-room:
nothing else running, 4x stats, sampler reads locks=1). A customer full-restore runs
stats in its preparation while holding NO acquireRunning, so the integrity check is not
blocked, runs, meets that lock, and resticStep escalates to unlock --remove-all - the
sampler caught "restore ..." and "unlock --remove-all" in the SAME sample at 20:50:51.
The log calls it "a stale exclusive lock left by a previous crash"; there was no crash.
THE CUSTOMER-FACING CONSEQUENCE IS CONTAINED and that is R-359 working: the check was
classified Unreachable, NOT damage, so no alarm and due-ness held. The opposite direction
is FENCED - five restores fired into a running check at 5/15/25/35/40s were all refused by
restoreOpBlocked with zero restic invoked.

PHASE 2 FAIL - R-412, and it is the natural instance of R-403's shape that yesterday's
session had to hand-build. A lost primary unit is rebuilt by the capture WITHOUT its
volume dumps (185664 B -> 4382 B, volume_dumps: None), and the off-site backup then ships
it and logs "backed up opengist ... 0 mandatory path(s)" - a success line over a backup
holding none of the app's data.

PHASE 2 also R-413: the R-87 proof CAUGHT that hollow snapshot unattended -
verdict "fail", volumes_expected_none_captured: opengist_data, one offsite_proof_empty at
severity error, while the four apps ahead of it in the rotation passed.
2.3 stale marker PASS, 2.4 two restores PASS, 2.5 corrected the runbook's premise
(RunTier2 has zero acquireRunning and zero restic references - a local mirror cannot
collide with the repo, so the proof correctly does not skip for it).

PHASE 3 PASS - R-357 live-validated at last, six days owed. A real full filesystem
(1900544 B free vs 5878421 B needed) on a separate 1 TB device that does not back Docker.
Refused BEFORE StopStack: app "Up 6 minutes (healthy)" unchanged, 0 safety dumps, live
userdata tree fingerprint 5a9b2db75dfd4ba4131adb3255670472 unchanged, message names both
numbers in Hungarian. Ballast removed, the SAME restore then worked and the fingerprint is
still identical. On the same full disk the Tier-2 run, the proof and the integrity check
all behaved: the proof refused before any download through the shared unitOnlyHeadroom
gate extracted today, reached no verdict and did not alarm.

PHASE 4 PASS - the false-alarm control the whole R-87 design rests on now EXISTS. Found by
reading all 53 catalogue composes: bentopdf is the only template with neither a database
service nor a named volume. Deployed, backed up, proved: PASSED in 2.224s with zero alarms.

Four of my own instrument errors were caught by their own controls before any result was
believed: a hub log line used as a controller positive control, a grep pattern that missed
a registered job, a heredoc that mangled a planted marker, and a catalogue scan that read
0 of 0 apps from the wrong path.

Phases 5-7 to follow: the real nightly cycle with injected faults, the untouched observer,
and teardown.
2026-08-31 23:33:32 +02:00

3.4 KiB

Phase 4 — the proof's missing control and its edges. VERDICT: PASS

The one that mattered — the false-alarm control now EXISTS and PASSES

Yesterday's report recorded that no app on demo-hp had neither a database nor a named volume, so R-87's "must not alarm on a legitimately empty app" case had no live subject. The whole design rests on that control.

Found by reading all 53 catalogue composes, not by guessing: exactly one template qualifies — bentopdf (single service ghcr.io/alam00000/bentopdf:v2.8.6, no database image, no top-level volumes: block, needs_hdd: false). (My first scan reported 0 of 0 apps — the path was wrong; the positive control caught it.)

Deployed, toggled into the off-site set, backed up (454e6f8b), and proved:

proof: bentopdf PASSED on snapshot 454e6f8b in 2.224s — the backup holds what this app should have
offsite_proof_empty events: 0 before, 0 after   ->  PASS: no alarm was raised

Its unit is genuinely empty: db_dumps: [], volume_dumps: None, compose declares no volumes. The guard does not fire on a healthy app that legitimately has nothing.

The other edges

edge expected actual verdict
3. corrupt manifest → cannot judge cannot judge could not be reached through the live path — the off-site run's own capture phase rewrote my corrupted manifest before the push, so the snapshot was healthy and the pass was correct. Verified by restoring 5dcee6e3 and reading its manifest UNDETERMINED live; covered by TestR87_UnparseableManifestFails (fail-closed, not cannot-judge — a documented divergence from the runbook's expectation, and R-403's rule)
4. no newest snapshot → not a failure not a failure observed naturally: while bentopdf was deployed but not yet toggled into the off-site set, the proof rotated all 8 apps and returned no_snapshot:true. It never picked bentopdf and never failed PASS
5. two failures in an hour → one mail one mail UNDETERMINED — only two offsite_proof_empty were pushed tonight and they are 1 h 42 m apart (19:21:45, 21:03:00), so they fall in different cooldown windows and do not exercise it. The controller-side half of the contract (no stack_name, so the hub keys coarsely) is pinned by TestR87_NoPerAppCooldownEntryWasAdded UNDETERMINED
6. DB declared, no dump → fail, "intact but empty" fail proven twice tonight — Phase 2's natural opengist case, verbatim: "READABLE AND EMPTY … the store is not damaged" PASS

A third instance of the same product behaviour, worth naming

Three separate attempts to get a damaged unit into the store through the front door were repaired by the off-site run's own pre-push capture: a stopped app's volume tar was re-made, a deleted unit was rebuilt, and a corrupted manifest was rewritten. That is a reassuring property, and it sharpens R-412: the capture repairs the manifest and compose, but it cannot conjure a volume tar, because the dump legs run on the backup schedule. That is the one gap through which R-412's hollow unit reached the store.

And a confirmation of R-402's shape for the new fields

The hub's host page mentions proof zero times. The controller publishes last_proof_* and no hub surface reads them — exactly what the wire-contract allowlist entries added today say, and why they say "delete this entry when a hub surface reads it".