cee8f70e98
gates / gates (push) Failing after 17s
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.
VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.
R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.
R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.
R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.
R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.
R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.
R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.
Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.
EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.
Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
137 lines
7.7 KiB
Markdown
137 lines
7.7 KiB
Markdown
# Overnight soak — the week's five new guards, run against each other
|
||
|
||
**22:39 → 06:10 CEST, 2026-08-31/09-01. `demo-hp` the victim, `demo-felhom` the untouched observer.**
|
||
No production code written. No golden baked. No version bumped.
|
||
|
||
| # | phase | verdict | the one sentence |
|
||
|---|---|---|---|
|
||
| 1 | lock collision | **FAIL** | a background job **deletes the lock of a live customer restore** and logs a crash that did not happen — customer-facing consequence contained |
|
||
| 2 | guard interactions | **PASS with one defect** | 4 of 5 guards clean; a unit lost *inside* a run ships hollow and is reported a success |
|
||
| 3 | R-357 full disk | **PASS** | six days owed, now proven against a real full filesystem — refused before the app went down, data byte-identical |
|
||
| 4 | proof edges | **PASS** | the false-alarm control the whole R-87 design rests on now exists and stays silent |
|
||
| 5 | mutated cycle | **PASS** | every job ran; the R-403 guard fired and named itself; the nightly proof fired unattended for the first time |
|
||
| 6 | observer | **FAIL** | on a box with no registered drive the nightly proof **cannot run at all**, every night, with only a WARN |
|
||
| 7 | teardown | **PASS** | all three layers clean on both boxes; both healthy on 0.231.0 |
|
||
|
||
---
|
||
|
||
## 1. What surprised me, worst first
|
||
|
||
### The proof is inert on a whole class of box (R-414) — and only the untouched observer could find it
|
||
|
||
`demo-felhom` has **zero registered storage paths**. At 05:30 the new job fired for the first time
|
||
unattended and refused: *"proof: opengist has nowhere to restore to — nincs regisztrált
|
||
adatmeghajtó"*. It will do that **every night, forever**, and the only signal is a `WARN`.
|
||
|
||
Worse than a failure: the error path reaches no verdict, so `last_proof_result` stays **absent** — and
|
||
absent is also what a pre-0.231.0 controller sends. **The hub cannot tell "never ran" from "not
|
||
deployed".** That is the exact `StatsKnown` trap the field was designed to avoid, reappearing one
|
||
level up. The same absence explains that box's `tier2-backup` finishing in **3 ms**.
|
||
|
||
The box is not unprotected — its off-site backup ran normally in 46.9 s.
|
||
|
||
### `unlock --remove-all` really does run against a live operation (R-411)
|
||
|
||
Measured, not reasoned. `restic stats` **takes a repository lock** (clean-room test). A customer
|
||
full-restore runs `stats` while holding **no** single-writer flag, so the integrity check is not
|
||
blocked, runs, meets that lock, and escalates. The sampler caught `restore …` and
|
||
`unlock --remove-all` **in the same sample**, and the log said *"a stale exclusive lock left by a
|
||
previous crash"*. There was no crash.
|
||
|
||
**Contained**: the check was classified *unreachable*, not damage — no false "your backups are
|
||
damaged" mail, and due-ness held. The opposite direction is fenced: five restores fired into a running
|
||
check were all refused, zero restic invoked.
|
||
|
||
### I over-claimed a finding and had to correct it (R-412)
|
||
|
||
Filed **HIGH** on a mechanism I had not finished establishing. Overnight measurement showed the
|
||
off-site run has its **own** pre-push dump leg, so a hollow unit is **repaired before it ships** —
|
||
proven on two apps, and confirmed by pulling the snapshot back out of the store. **Corrected to LOW**,
|
||
with the over-claim written into the row rather than quietly edited away.
|
||
|
||
What survives is narrower and real: a unit destroyed *inside* a run, after that app's dump leg, ships
|
||
hollow and the run logs `backed up opengist (… 0 mandatory path(s))` — a success line over a backup
|
||
holding none of the app's data.
|
||
|
||
### A quiet one worth saying: three attempts to damage a backup were repaired by the product
|
||
|
||
A stopped app's tar was re-made; a deleted unit was rebuilt; a corrupted manifest was rewritten — each
|
||
by the run's own capture, before any push. That is reassuring, and it is why the hollow snapshot
|
||
needed a race to happen at all.
|
||
|
||
---
|
||
|
||
## 2. Findings filed
|
||
|
||
| id | severity | what |
|
||
|---|---|---|
|
||
| **R-411** | MEDIUM | a background job deletes a live restore's lock and calls it a crash |
|
||
| **R-412** | LOW *(was HIGH — corrected)* | a unit lost inside a run ships hollow and reports success |
|
||
| **R-413** | CLOSED | the R-87 proof caught a product-produced hollow snapshot unattended |
|
||
| **R-414** | MEDIUM | the nightly proof is inert on a box with no registered drive |
|
||
|
||
Register: `OPEN-ITEMS.md` **173 → 174** rows. Evidence:
|
||
`documentation/audits/DRILL-soak-2026-08-31/`, seven phase directories.
|
||
|
||
---
|
||
|
||
## 3. What I could not test, and why
|
||
|
||
- **"Mark a drive disconnected" (Phase 5).** No endpoint reaches `SetDisconnected`; a hand-set flag
|
||
would be reverted by the live monitor before 04:15, so the run would have read as a passing test of
|
||
a fault that was not present. Unmounting a live drive risks a wedged mount that needs a reboot.
|
||
04:15 was observed against a real injected fault instead.
|
||
- **"Corrupt a manifest before the proof."** Phase 4 established the capture rewrites it before any
|
||
push, so nothing I do to the live manifest can change what the proof sees. Covered by
|
||
`TestR87_UnparseableManifestFails`.
|
||
- **Two failing apps in one hour → one mail.** Only two alarms fired all night and they were 1 h 42 m
|
||
apart, so the coarse cooldown was never exercised. **UNDETERMINED.**
|
||
- **The 06:00 integrity check at full depth on `demo-hp`.** It completed in **0 s — not due**, because
|
||
my own manual runs during Phases 1–3 had already advanced its due-ness. My doing, and correct
|
||
behaviour.
|
||
|
||
---
|
||
|
||
## 4. Left changed on the boxes
|
||
|
||
- **`demo-felhom` is hand-deployed to 0.231.0** (was 0.230.0). Reversible; the golden bake supersedes
|
||
it. **Recorded here because a hand-deployed box nobody records is how a fleet drifts.**
|
||
- **`bentopdf` is deployed on `demo-hp` and I recommend KEEPING it.** It is the **only** template of
|
||
53 with neither a database nor a named volume, so it is the only possible live subject for R-87's
|
||
false-alarm control. Deleting it deletes the control.
|
||
- A hollow `opengist` snapshot (`35ba9fe7`) remains in the off-site store as history. Its newest
|
||
snapshot is sound and the proof passes it.
|
||
- Nothing else. All ballast, scratches, markers, probe scripts and credentials removed from all three
|
||
layers on both boxes; `/root` on both hosts holds only `.cache` and `.ssh`; both container `/tmp`
|
||
directories are empty.
|
||
|
||
---
|
||
|
||
## 5. My own mistakes, by name — eight, all caught by their own controls
|
||
|
||
1. Used a **hub** log line as the positive control for a **controller** recorder.
|
||
2. A grep pattern that reported a registered job as missing when it was there.
|
||
3. A heredoc that mangled a planted marker, and an escaping error that reported `"full":false` as a
|
||
failure when the file plainly read `"full":false`.
|
||
4. A catalogue scan that read **0 of 0** apps from the wrong path.
|
||
5. A wait that matched my **own earlier output** replayed by the log follower — a false "job observed".
|
||
6. **A time guard `[ 2334 -ge 0315 ]` that fired the Phase 5 injection four hours early**, changing
|
||
the drill.
|
||
7. **Filing R-412 at HIGH on a mechanism I had not finished measuring** — the worst one, because it
|
||
went into the register wrong.
|
||
8. A baseline grep whose control failed, nearly making a 10-day-old stale directory look like
|
||
tonight's damage; mtime settled it.
|
||
|
||
Every zero reported above was earned by a positive control first. That is the only reason the list is
|
||
eight and not eight-plus-a-wrong-verdict.
|
||
|
||
---
|
||
|
||
## 6. Owed to Viktor in the morning
|
||
|
||
1. **A golden carrying 0.231.0** — still owed from yesterday; the fleet floor is 0.230.0.
|
||
2. **A decision on `bentopdf`** — keep it as the permanent false-alarm control (my recommendation),
|
||
or remove it.
|
||
3. **R-414** is the one worth reading first: a feature that shipped yesterday does not run at all on a
|
||
box shaped like `demo-felhom`.
|