3208cb208366120288bbff98273fba76718288a4
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cee8f70e98 |
soak drill COMPLETE: 7 phases, 4 findings, the observer found the worst one (R-414)
gates / gates (push) Failing after 17s
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.
VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.
R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.
R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.
R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.
R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.
R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.
R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.
Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.
EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.
Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
|
||
|
|
ab8b884763 |
soak phase 5: R-403 guard PROVEN live; R-412 CORRECTED down after measuring the mechanism
gates / gates (push) Failing after 18s
R-403 MIRROR GUARD - PASS, forced after the natural test evaporated. privatebin was injected hollow at 23:34 to meet the 03:30 mirror; the 02:30 db-dump re-made its tar, so by 03:30 the primary was complete and the guard had nothing to refuse. Forced instead on calibre-web through the real Tier-2 path: the guard fired and named itself - "unit leg SKIPPED ... The copy was PRESERVED rather than replaced with an empty one (R-403). The other legs continue." Secondary byte-identical, 23 files, tar sha d7e7f422. R-412 CORRECTED, AND I OVERSTATED IT WHEN I FILED IT. The first wording claimed the hollow unit sits in the store for a whole cycle because the volume-dump leg runs only on the backup schedule. That is WRONG. The 04:15 off-site run has its OWN pre-push dump leg - "Stopping calibre-web for safe volume dump", "Volume dump: ... -> 877.5 KB" - so a unit that is hollow when a run starts is REPAIRED before it is pushed. Measured twice: opengist and calibre-web both went in hollow and came out complete, and the snapshot pulled back from the store (6fee3b5a) holds the volume tar and all 17 userdata files. What remains real is narrower: the one hollow snapshot that DID reach the store was created when the unit was destroyed INSIDE a run that had already completed that app's dump leg. The race is real and was observed, and "backed up opengist (... 0 mandatory path(s))" is a success line over a backup holding none of the app's data either way. Severity HIGH -> LOW, with the correction stated in the row rather than quietly rewritten. Phase 6 interim: the observer is clean so far - db-dump 674ms, tier2-backup 3ms (a no-op, cause to be established not assumed), zero ERROR/WARN since 23:00. Also recorded: two of Phase 5's four injections were NOT performed, with the reasons established rather than asserted - there is no endpoint that reaches SetDisconnected and a hand-set flag would be reverted by the live monitor before 04:15; and a corrupted manifest provably never reaches the store because the capture rewrites it first. |