cee8f70e98
gates / gates (push) Failing after 17s
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.
VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.
R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.
R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.
R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.
R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.
R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.
R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.
Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.
EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.
Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
78 lines
2.7 KiB
Plaintext
78 lines
2.7 KiB
Plaintext
### demo-hp 04:02:36 UTC
|
|
--- controller ---
|
|
gitea.dooplex.hu/admin/felhom-controller:0.231.0
|
|
gitea.dooplex.hu/admin/felhom-controller:0.231.0 Up 9 hours (healthy)
|
|
--- every app: running and healthy? ---
|
|
bentopdf Up 51 minutes (healthy)
|
|
bookstack Up 51 minutes (healthy)
|
|
bookstack-db Up 51 minutes (healthy)
|
|
calibre-web Up 51 minutes (healthy)
|
|
docmost Up 51 minutes (healthy)
|
|
docmost-postgres Up 51 minutes (healthy)
|
|
docmost-redis Up 51 minutes (healthy)
|
|
filebrowser Up 10 days (healthy)
|
|
kimai Up 51 minutes (healthy)
|
|
kimai-db Up 51 minutes (healthy)
|
|
opengist Up 51 minutes (healthy)
|
|
paperless-postgres Up 51 minutes (healthy)
|
|
paperless-redis Up 51 minutes (healthy)
|
|
paperless-webserver Up 51 minutes (healthy)
|
|
privatebin Up 51 minutes (healthy)
|
|
romm Up 50 minutes (healthy)
|
|
romm-db Up 51 minutes (healthy)
|
|
romm-redis Up 51 minutes (healthy)
|
|
--- unhealthy or exited (must be empty) ---
|
|
--- primary units: is each a REAL package? ---
|
|
bentopdf dumps=0 tars=0 HOLLOW
|
|
bookstack dumps=1 tars=2
|
|
docmost dumps=1 tars=3
|
|
kimai dumps=1 tars=2
|
|
opengist dumps=0 tars=1
|
|
paperless MANIFEST UNREADABLE: [Errno 2] No such file or directory: '/mnt/sys_drive/felhom-data/backups/primary/paperless/manifest.json'
|
|
privatebin dumps=0 tars=1
|
|
calibre-web dumps=0 tars=1
|
|
paperless-ngx dumps=1 tars=3
|
|
romm dumps=1 tars=3
|
|
--- tier2 copies ---
|
|
bentopdf: 5 files, 4156 bytes
|
|
bookstack: 11 files, 167027817 bytes
|
|
docmost: 12 files, 123773625 bytes
|
|
kimai: 8 files, 213231243 bytes
|
|
opengist: 6 files, 185665 bytes
|
|
privatebin: 6 files, 2122929 bytes
|
|
calibre-web: 23 files, 5808704 bytes
|
|
paperless-ngx: 13 files, 84582358 bytes
|
|
romm: 10 files, 185703681 bytes
|
|
--- leftover probes / scratches / ballast ---
|
|
offsite-restore: [bookstack calibre-web docmost kimai paperless-ngx ]
|
|
offsite-proof : []
|
|
root leftover: .soak-bentopdf-manifest.bak
|
|
root leftover: .soak-cookie
|
|
root leftover: .soak-csrf
|
|
root leftover: .soak-hdr
|
|
root leftover: .soak-pw
|
|
root leftover: insp.sh
|
|
root leftover: p21.sh
|
|
root leftover: p21b.sh
|
|
root leftover: p21c.sh
|
|
root leftover: p25.sh
|
|
root leftover: p3a.sh
|
|
root leftover: p3b.sh
|
|
root leftover: p3c.sh
|
|
root leftover: p3d.sh
|
|
root leftover: p4a.sh
|
|
root leftover: p4b.sh
|
|
root leftover: p4c.sh
|
|
root leftover: soak-api.sh
|
|
root leftover: soak-baseline.sh
|
|
root leftover: soak-collide.sh
|
|
root leftover: soak-locksample.sh
|
|
root leftover: soak-login.sh
|
|
root leftover: soak-reverse.sh
|
|
root leftover: stale.sh
|
|
container /tmp: [insp.sh soak-env.sh soak-locks.log soak-locksample.sh ]
|
|
--- free space ---
|
|
Mounted on Avail
|
|
/mnt/felhom-drives/hdd_1 949284630528
|
|
/var/lib/felhom 57631019008
|