Files
felhom.eu/documentation/audits/DRILL-soak-2026-08-31/REPORT.md
T
admin cee8f70e98
gates / gates (push) Failing after 17s
soak drill COMPLETE: 7 phases, 4 findings, the observer found the worst one (R-414)
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.

VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.

R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.

R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.

R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.

R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.

R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.

R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.

Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.

EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.

Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
2026-09-01 06:07:13 +02:00

7.7 KiB
Raw Blame History

Overnight soak — the week's five new guards, run against each other

22:39 → 06:10 CEST, 2026-08-31/09-01. demo-hp the victim, demo-felhom the untouched observer. No production code written. No golden baked. No version bumped.

# phase verdict the one sentence
1 lock collision FAIL a background job deletes the lock of a live customer restore and logs a crash that did not happen — customer-facing consequence contained
2 guard interactions PASS with one defect 4 of 5 guards clean; a unit lost inside a run ships hollow and is reported a success
3 R-357 full disk PASS six days owed, now proven against a real full filesystem — refused before the app went down, data byte-identical
4 proof edges PASS the false-alarm control the whole R-87 design rests on now exists and stays silent
5 mutated cycle PASS every job ran; the R-403 guard fired and named itself; the nightly proof fired unattended for the first time
6 observer FAIL on a box with no registered drive the nightly proof cannot run at all, every night, with only a WARN
7 teardown PASS all three layers clean on both boxes; both healthy on 0.231.0

1. What surprised me, worst first

The proof is inert on a whole class of box (R-414) — and only the untouched observer could find it

demo-felhom has zero registered storage paths. At 05:30 the new job fired for the first time unattended and refused: "proof: opengist has nowhere to restore to — nincs regisztrált adatmeghajtó". It will do that every night, forever, and the only signal is a WARN.

Worse than a failure: the error path reaches no verdict, so last_proof_result stays absent — and absent is also what a pre-0.231.0 controller sends. The hub cannot tell "never ran" from "not deployed". That is the exact StatsKnown trap the field was designed to avoid, reappearing one level up. The same absence explains that box's tier2-backup finishing in 3 ms.

The box is not unprotected — its off-site backup ran normally in 46.9 s.

unlock --remove-all really does run against a live operation (R-411)

Measured, not reasoned. restic stats takes a repository lock (clean-room test). A customer full-restore runs stats while holding no single-writer flag, so the integrity check is not blocked, runs, meets that lock, and escalates. The sampler caught restore … and unlock --remove-all in the same sample, and the log said "a stale exclusive lock left by a previous crash". There was no crash.

Contained: the check was classified unreachable, not damage — no false "your backups are damaged" mail, and due-ness held. The opposite direction is fenced: five restores fired into a running check were all refused, zero restic invoked.

I over-claimed a finding and had to correct it (R-412)

Filed HIGH on a mechanism I had not finished establishing. Overnight measurement showed the off-site run has its own pre-push dump leg, so a hollow unit is repaired before it ships — proven on two apps, and confirmed by pulling the snapshot back out of the store. Corrected to LOW, with the over-claim written into the row rather than quietly edited away.

What survives is narrower and real: a unit destroyed inside a run, after that app's dump leg, ships hollow and the run logs backed up opengist (… 0 mandatory path(s)) — a success line over a backup holding none of the app's data.

A quiet one worth saying: three attempts to damage a backup were repaired by the product

A stopped app's tar was re-made; a deleted unit was rebuilt; a corrupted manifest was rewritten — each by the run's own capture, before any push. That is reassuring, and it is why the hollow snapshot needed a race to happen at all.


2. Findings filed

id severity what
R-411 MEDIUM a background job deletes a live restore's lock and calls it a crash
R-412 LOW (was HIGH — corrected) a unit lost inside a run ships hollow and reports success
R-413 CLOSED the R-87 proof caught a product-produced hollow snapshot unattended
R-414 MEDIUM the nightly proof is inert on a box with no registered drive

Register: OPEN-ITEMS.md 173 → 174 rows. Evidence: documentation/audits/DRILL-soak-2026-08-31/, seven phase directories.


3. What I could not test, and why

  • "Mark a drive disconnected" (Phase 5). No endpoint reaches SetDisconnected; a hand-set flag would be reverted by the live monitor before 04:15, so the run would have read as a passing test of a fault that was not present. Unmounting a live drive risks a wedged mount that needs a reboot. 04:15 was observed against a real injected fault instead.
  • "Corrupt a manifest before the proof." Phase 4 established the capture rewrites it before any push, so nothing I do to the live manifest can change what the proof sees. Covered by TestR87_UnparseableManifestFails.
  • Two failing apps in one hour → one mail. Only two alarms fired all night and they were 1 h 42 m apart, so the coarse cooldown was never exercised. UNDETERMINED.
  • The 06:00 integrity check at full depth on demo-hp. It completed in 0 s — not due, because my own manual runs during Phases 1–3 had already advanced its due-ness. My doing, and correct behaviour.

4. Left changed on the boxes

  • demo-felhom is hand-deployed to 0.231.0 (was 0.230.0). Reversible; the golden bake supersedes it. Recorded here because a hand-deployed box nobody records is how a fleet drifts.
  • bentopdf is deployed on demo-hp and I recommend KEEPING it. It is the only template of 53 with neither a database nor a named volume, so it is the only possible live subject for R-87's false-alarm control. Deleting it deletes the control.
  • A hollow opengist snapshot (35ba9fe7) remains in the off-site store as history. Its newest snapshot is sound and the proof passes it.
  • Nothing else. All ballast, scratches, markers, probe scripts and credentials removed from all three layers on both boxes; /root on both hosts holds only .cache and .ssh; both container /tmp directories are empty.

5. My own mistakes, by name — eight, all caught by their own controls

  1. Used a hub log line as the positive control for a controller recorder.
  2. A grep pattern that reported a registered job as missing when it was there.
  3. A heredoc that mangled a planted marker, and an escaping error that reported "full":false as a failure when the file plainly read "full":false.
  4. A catalogue scan that read 0 of 0 apps from the wrong path.
  5. A wait that matched my own earlier output replayed by the log follower — a false "job observed".
  6. A time guard [ 2334 -ge 0315 ] that fired the Phase 5 injection four hours early, changing the drill.
  7. Filing R-412 at HIGH on a mechanism I had not finished measuring — the worst one, because it went into the register wrong.
  8. A baseline grep whose control failed, nearly making a 10-day-old stale directory look like tonight's damage; mtime settled it.

Every zero reported above was earned by a positive control first. That is the only reason the list is eight and not eight-plus-a-wrong-verdict.


6. Owed to Viktor in the morning

  1. A golden carrying 0.231.0 — still owed from yesterday; the fleet floor is 0.230.0.
  2. A decision on bentopdf — keep it as the permanent false-alarm control (my recommendation), or remove it.
  3. R-414 is the one worth reading first: a feature that shipped yesterday does not run at all on a box shaped like demo-felhom.