8.1 KiB
REPORT — R-181: the reserve guards the write that fills the disk, and its promise is true
Date: 2026-08-03 · Repo: felhom-controller · v0.192.0 → v0.193.0 → v0.193.1
Commits on main: fef07c3 (v0.193.0), 6c43bf6 (v0.193.1) · Baseline on arrival: 4be6467b501b, v0.192.0 (matched)
The defect
B2's capture floor (v0.192.0, R-165) shipped as the deliberate replacement for the bulkhead the mp1
partition used to give, and it was consulted in exactly one place — captureAllRecoveryUnits,
which writes a manifest and three compose files: a few KB. The two legs that write the bulk into
the same backups/primary/<app> tree — the database dump and the volume dump — ran first and
unguarded.
Measured live on demo-hp 2026-08-03 06:40:03: opengist's volume dump wrote 2.0 GB with no check, free fell to 1.0 GB, and the floor then refused the cheap write it had already lost the argument to.
Second limb: the refusal printed "the previous unit is untouched and NOTHING was deleted".
Nothing was deleted held. Untouched was measured false — that app's tar had gone
182,272 B → 2,147,666,432 B under a manifest.json whose created_at and checksums had not moved.
Sixth entry in CLAUDE.md's table of shipped guarantees the code did not provide.
The fix
One admission verdict per app per run (internal/backup/admission.go), taken before that app's
first write and consulted by all three legs. They write under one per-app root
(appbackup.RecoveryUnitPath), which is what makes one verdict able to cover them honestly.
- Lazy, at the app's first write — not once at run start. App A's dump can put app B under the reserve; a run-start verdict reads a disk that no longer exists.
- Never re-decided between an app's own legs (that is the split being closed); reset per run.
- Ahead of
DumpAppVolumesSafe, which stops the stack as its first act. After the volume-less check, which has no write to gate. - Exactly one operator alert per refused app per run. Leg order unchanged.
- Size term: would this app's write cross the reserve?, estimated from its previous
.sql+.tar. No history → headroom-only, or the first backup is the one that can never happen.
Files
internal/backup/admission.go (new), admission_test.go (new, 11 tests), backup.go (run scope +
DB-leg and volume-leg gates), recovery_unit.go (floorVerdict size-aware; capture leg via
admitApp), capture_floor_test.go (3 call sites), controller/README.md, REUSE.md, CHANGELOG.md.
Tests — go build ./... && go vet ./... && go test ./... → 28 packages ok, rc=0
Read separately from any commit (standing rule 1). Refusal assertions are sha256 tree fingerprints before and after, never log lines — the defect is a log line the tree contradicted.
The DB leg cannot run without Docker (DiscoverDatabases shells out), so its gate is pinned by an
AST walk of backup.go asserting admitApp precedes DumpOne; strings.Contains is
insufficient, a commented-out call still contains the string.
Red-proofs — demonstrated failing, then restored
| Mutation | Result |
|---|---|
| Both dump-leg gates removed (= exactly v0.192.0) | Scenario A RED ("the VOLUME leg ran for a refused app"), and with the leg assertions temporarily non-fatal the tree fingerprint changed too. Also red: C, D, and the AST wiring test |
The whole size term removed from floorVerdict |
Scenario D RED — 0 alerts where 1 required |
| The reserve removed entirely | Scenario F PASSED — recorded honestly. It does not exercise the assertion: every app then writes, which overwrites and adds but deletes nothing, so a deletion-watching test correctly stays green |
| A prune injected into the refusal path | Scenario F RED — the mutation that actually proves it watches deletion |
| Floor moved above the warning band (90% / 6 GiB) | TestFloorSitsBelowTheCriticalWarningBand RED |
python3 controller/scripts/controller_gates.py → all 8 gates OK.
Live validation — demo-hp guest 9201 (Tier 0)
Method: endpoint-level — POST /api/debug/backup/dbdump, the exact endpoint the debug UI button
calls, running the production RunDBDumps. No browser on DooPlex.
Instrument re-proven first: demo-hp's thin pool is 53.93 GiB, so a real fill of the 70 G volume
would exhaust it. A 5 GiB fallocate moved guest df 1.2G → 6.2G while data_percent held
36.83 → 36.83 — zero blocks allocated. Re-checked at every fill step.
Headroom term, 08:59:46 (906 MB free / 99%): TREE_SHA=111d1760c18d3440f700634ab325f8b8
identical before and after (opengist's tar still 182,272 B); 0 volume dumps; no
Stopping <app> for safe volume dump line at all — evidence rather than absence, because that line
is present in the 08:58 baseline run; one alert per app, HTTP 200. Freed and re-run 09:01:33 → both
captured normally.
Size term, 09:03:00, proven separately: a real 2 GiB file in opengist's volume made its previous
tar 2,147,666,432 B (the exact live figure); filesystem set to 91% used / 2.9 GB free — both
headroom terms deliberately clear. opengist refused (size); privatebin was ADMITTED and
dumped normally, so the term is per-app and not a global halt.
One honest correction: both apps' StartedAt did move, 26 s after the refusal — the quiesce
loop for the whole-guest PBS backup, which the fill had broken, not the app-data path. The
app-data claim rests on the absent Stopping … for safe volume dump line.
The du measurement — measured, then rejected
66 timed runs, docker run --rm -v <vol>:/v alpine du -sb /v: median ~355 ms/volume
(341–404 ms) on volumes holding tens of KB — container start-up, not the walk. Rejected on two
further grounds: docker run needs the writable layer, so the instrument can fail under exactly the
pressure the reserve handles; and the previous-dump estimate measures the artifact to be written
rather than the live volume. The estimate stands.
v0.193.1 — found by the proof run itself
The estimate was rendered fixed to two-decimal GiB, so opengist's real 178 KB printed as
estimated 0.00 GiB write — which reads as no estimate was available, the opposite of what
happened. Shipped the same session because it is the same defect class. Arithmetic unchanged (still
GiB, the reserve's own unit); rendering moved to humanizeBytes. Re-verified live after redeploy:
estimated 178.0 KB write.
Deployed
gitea.dooplex.hu/admin/felhom-controller:0.193.1 on demo-hp guest 9201 — Up (healthy).
Not deployed to demo-felhom (which carries the PBS-DR/offsite tier and was out of scope).
§3's correction — confirmed, not chased
restore_points.go:57-59 takes the manifest mtime then newestArtifact over .sql and .tar, so
the newest of the three wins. The restore point does not show a stale timestamp.
Teardown
Fill removed; planted 2 GiB file removed; a final backup regenerated a correct 178 KB tar;
pct fstrim 9201 returned 67.5 GiB and the pool settled at 29.43%, below its 36.83% baseline; the
tree is byte-identical to the pre-test fingerprint. Guest helper scripts and the credential file
shred-ed. No --no-verify on either push; the pre-push hook ran and reported gates OK both times.
Observations — NOT acted on (filed in felhom.eu/documentation/backlog/OPEN-ITEMS.md)
- R-182 (filed).
GetFullStatus→captureAllRecoveryUnitsruns with no admission scope, so a refused app re-alerts on every status poll — measured: a second identical alert pair 13 s after the run's. Pre-existing in v0.192.0; this change touched neither caller. - A reserve refusal does not fail the run. Both dump legs record
SKIP, notFAIL, solastDBDump.Successstays true. Deliberate and consistent with v0.192.0, but "backup succeeded" and "every app was backed up" are then not the same statement. UnitSpace.UsedPercentanddfdisagree (94% vs 99%) —dfaccounts for ext4 reserved blocks and the floor's figure does not. Harmless here, but the operator cannot reproduce the percent term withdf.