v0.193.0 — the reserve guards the write that fills the disk, and its promise is true (R-181)
gates / gates (push) Successful in 9s
gates / gates (push) Successful in 9s
B2's capture floor (v0.192.0) was consulted in exactly ONE place — captureAllRecoveryUnits, which writes a few KB. The two legs that write the BULK into the same backups/primary/<app> tree, the DB dump and the volume dump, ran FIRST and unguarded. Measured live on demo-hp 2026-08-03 06:40:03: opengist's volume dump wrote 2.0 GB with no check, free fell to 1.0 GB, and the floor then refused the cheap write it had already lost the argument to. Its refusal message claimed "the previous unit is untouched" — measured false: that app's tar had gone 182,272 B -> 2,147,666,432 B under a stale manifest. Sixth entry in CLAUDE.md's table of shipped guarantees the code did not provide. Fix: ONE admission verdict per app per run (internal/backup/admission.go), taken before that app's FIRST write and covering all three legs — they write under one per-app root, which is why one verdict can honestly cover them. - Lazy, at the app's first write, NOT once at run start: app A's dump can put app B under the reserve, so a run-start verdict reads a disk that no longer exists. - Remembered for the run, never re-decided between an app's own legs — that is the split this closes. Reset per run. - Placed ahead of DumpAppVolumesSafe, which stops the stack as its first act, so a refused app is never bounced. After the volume-less check, which has no write. - Exactly one operator alert per refused app per run. - Leg order unchanged: volume dumps still precede the capture. The floor is now SIZE-AWARE: it asks whether THIS app's write would cross the reserve, not only whether the filesystem is already below it — which is how an app was admitted at 96% and then allowed to write 2 GB. Estimate = the app's previous .sql + .tar on disk. No history -> headroom-only, deliberately, and the alert says so. A container-based du per volume was MEASURED and rejected: 66 timed runs on demo-hp guest 9201, median ~355 ms/volume (341-404) on volumes holding tens of KB — container start-up, not the walk. Decisive on top: docker run needs the writable layer, so it can fail under exactly the pressure the reserve handles. The message was NOT weakened; the behaviour was moved so the wording became true. It now also names which term bound. Every claim is checked against a sha256 fingerprint of the tree it describes, never against the log line. Still refuses and never deletes: nothing here is generational. 11 new tests through the production functions. The DB leg cannot run without Docker, so its gate is pinned by an AST walk of backup.go asserting admitApp precedes DumpOne (strings.Contains is insufficient — a commented-out call still contains the string). 4 red-proofs demonstrated failing then restored.
This commit is contained in:
@@ -1,5 +1,75 @@
|
||||
## Changelog
|
||||
|
||||
### v0.193.0 — the reserve guards the write that fills the disk, and its promise is true (2026-08-03, R-181) — MinAgent: none
|
||||
|
||||
**The defect, found on live hardware and not by review.** v0.192.0's capture floor (B2) shipped as the
|
||||
deliberate replacement for the bulkhead the `mp1` partition used to give, and it was consulted in
|
||||
**exactly one place** — `captureAllRecoveryUnits`, which writes a manifest and three compose files: a
|
||||
few KB. The two legs that write the **bulk** into the same `backups/primary/<app>` tree — the database
|
||||
dump and the volume dump — ran **first** and **unguarded**. Measured on demo-hp 2026-08-03 06:40:03:
|
||||
opengist's volume dump wrote **2.0 GB with no check**, free fell to 1.0 GB, and the floor then refused
|
||||
the cheap write it had already lost the argument to.
|
||||
|
||||
**Second limb: the refusal message asserted something the code did not provide.** It printed *"the
|
||||
previous unit is untouched and NOTHING was deleted"*. *Nothing was deleted* held. *Untouched* was
|
||||
**measured false** — that app's tar had gone 182,272 B → 2,147,666,432 B under a `manifest.json` whose
|
||||
`created_at` and `checksums` had not moved. **This is the sixth entry in `CLAUDE.md`'s table of
|
||||
shipped guarantees the code did not provide**, and the fourth of those found on live hardware.
|
||||
|
||||
**The fix: ONE admission verdict per app per run, taken before that app's FIRST write, covering all
|
||||
three legs** (`internal/backup/admission.go`). The three write under one per-app root, which is
|
||||
exactly why one verdict can honestly cover them — and why the message may now claim what it claims.
|
||||
|
||||
- **Decided lazily, at the app's first write — NOT once at the start of the run.** Space changes
|
||||
during a run: app A's 2 GB dump can put app B under the reserve, and a run-start verdict would wave
|
||||
B through on a reading that was true before the disk filled. That is the same class of mistake,
|
||||
moved one level up.
|
||||
- **Remembered for the run, never re-decided between an app's own legs.** Re-deciding reintroduces
|
||||
the split this closes (DB admitted → volume admitted → capture refused, with the bulk written).
|
||||
Reset per run: a set carried between runs answers tonight's question with last night's disk.
|
||||
- **Placed ahead of `DumpAppVolumesSafe`, which stops the stack as its first act** — a refusal decided
|
||||
inside it would already have bounced the app it is refusing to back up. It sits *after* the
|
||||
volume-less check, because an app with no named volumes has no first write in that leg to gate.
|
||||
- **Exactly one operator alert per refused app per run.** Three legs must not mean three emails.
|
||||
- **The leg order is unchanged** — volume dumps still precede the capture so the manifests enumerate
|
||||
the fresh tars (`backup.go`'s load-bearing comment).
|
||||
|
||||
**The floor is now SIZE-AWARE, not merely headroom-aware.** It asks *would this app's write leave the
|
||||
filesystem below the reserve?*, not only *is it below the reserve now* — which is how an app was
|
||||
admitted at 96% used and then allowed to write 2 GB. The estimate is the app's **previous** `.sql` and
|
||||
`.tar` already on disk: free to read, and the next write is usually close. **No history →
|
||||
headroom-only**, deliberately, or the first backup would be the one that can never happen; the alert
|
||||
says so when that applies. Both post-write terms are evaluated, because a large write crosses the
|
||||
percentage bound on a small volume and the free-byte bound on a large one.
|
||||
|
||||
**A container-based `du` per volume was measured and REJECTED, not assumed.** 66 timed runs on
|
||||
demo-hp's guest 9201: **median ~355 ms per volume** (341–404 ms) on volumes holding tens of KB — the
|
||||
cost is container start-up, not the walk, so it does not shrink for small apps and only grows for real
|
||||
ones. Decisive on top of that: `docker run` needs the writable layer, so the measurement mechanism can
|
||||
fail under exactly the disk pressure the reserve exists to handle. The previous-dump estimate also
|
||||
measures the **artifact** that will be written rather than the live volume, which is the truer
|
||||
predictor. The figure and the decision are recorded rather than left as a "we could do better".
|
||||
|
||||
**THE MESSAGE WAS NOT WEAKENED — the behaviour was moved so the wording became true.** It still says
|
||||
the previous unit is untouched and nothing was deleted, and now adds *which* term bound (headroom or
|
||||
size) and the estimate that produced a size refusal.
|
||||
`TestAdmission_EveryClaimInTheRefusalMessageHoldsAgainstTheTree` checks **every** claim against a
|
||||
sha256 fingerprint of the tree it describes — not against the log line, because a log line is exactly
|
||||
what lied here.
|
||||
|
||||
**IT STILL REFUSES AND NEVER DELETES.** Unchanged and load-bearing: nothing on this filesystem is
|
||||
generational, so "prune the oldest" could only mean destroying a **different** app's only local
|
||||
recovery unit. `pruneStalePrimaryDirs` is not a retention policy and must never be repurposed for
|
||||
headroom.
|
||||
|
||||
**Tests (11 new, all through the production functions; 4 red-proofs demonstrated failing then
|
||||
restored).** The refusal assertions are **tree fingerprints before and after**, never log lines. The
|
||||
DB leg cannot run without Docker, so its gate is pinned by an **AST walk** of `backup.go` asserting
|
||||
`admitApp` precedes `DumpOne` — `strings.Contains` is insufficient, a commented-out call still
|
||||
contains the string. Red-proofs: both dump-leg gates removed (= v0.192.0) → Scenario A red, tree shown
|
||||
changing; the size term removed → Scenario D red; a prune injected into the refusal path → Scenario F
|
||||
red; the floor moved above the warning band → Scenario G red.
|
||||
|
||||
### v0.192.0 — the capture floor replaces the bulkhead (2026-08-03, R-165 · decision B2) — MinAgent: none
|
||||
|
||||
**Ships BEFORE the disk-layout merge it exists for, and is harmless on a box that never gets it.**
|
||||
|
||||
Reference in New Issue
Block a user