Files
felhom.eu/documentation/audits/DRILL-soak-2026-08-31/phase6-observer/02-observer-table.txt
T
admin cee8f70e98
gates / gates (push) Failing after 17s
soak drill COMPLETE: 7 phases, 4 findings, the observer found the worst one (R-414)
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.

VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.

R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.

R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.

R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.

R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.

R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.

R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.

Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.

EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.

Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
2026-09-01 06:07:13 +02:00

34 lines
2.0 KiB
Plaintext

=== scheduled-job activity after 2026-08-31T23:00:00Z ===
-- db-dump -- (2 line(s))
2026-09-01 00:30:00 [scheduler] Running job: db-dump
2026-09-01 00:30:00 [scheduler] Job db-dump completed (took 674ms)
-- tier2-backup -- (2 line(s))
2026-09-01 01:30:00 [scheduler] Running job: tier2-backup
2026-09-01 01:30:00 [scheduler] Job tier2-backup completed (took 3ms)
-- offbox-backup -- (2 line(s))
2026-09-01 02:15:00 [scheduler] Running job: offbox-backup
2026-09-01 02:15:46 [scheduler] Job offbox-backup completed (took 46.946s)
-- offsite-abandon-sweep -- (2 line(s))
2026-09-01 03:10:00 [scheduler] Running job: offsite-abandon-sweep
2026-09-01 03:10:00 [scheduler] Job offsite-abandon-sweep completed (took 0s)
-- offsite-proof -- (2 line(s))
2026-09-01 03:30:00 [scheduler] Running job: offsite-proof
2026-09-01 03:30:02 [scheduler] Job offsite-proof completed (took 2.565s)
-- offsite-integrity -- (2 line(s))
2026-09-01 04:00:00 [scheduler] Running job: offsite-integrity
2026-09-01 04:00:00 [scheduler] Job offsite-integrity completed (took 0s)
-- metrics-prune -- (2 line(s))
2026-09-01 02:00:00 [scheduler] Running job: metrics-prune
2026-09-01 02:00:00 [scheduler] Job metrics-prune completed (took 3ms)
-- fill-watch -- (2 line(s))
2026-09-01 01:30:00 [scheduler] Running job: fill-watch
2026-09-01 01:30:00 [scheduler] Job fill-watch completed (took 0s)
=== every EVENT pushed ===
2026-09-01 00:30:00 db_dump_completed (info) — Adatbázis mentés elkészült
=== every ERROR / WARN of interest ===
2026-09-01 03:30:02 [offbox] proof: opengist has nowhere to restore to: nincs regisztrált adatmeghajtó, ezért nincs hová visszaállítani — a meghajtók megvannak, csak újra
=== proof verdicts ===
2026-09-01 03:30:02 opengist has nowhere to restore to: nincs regisztrált adatmeghajtó, ezért nincs hová visszaállítani — a meghajtók megvannak, csak újra kell
=== integrity verdicts ===
2026-09-01 04:00:00 not due (last successful check 24h0m0s ago, at 2026-08-31 04:00) — nothing was run