docs(v0.230.0): R-403 — CHANGELOG, CONTEXT rulings, README, REPORT
gates / gates (push) Failing after 13s

CHANGELOG v0.230.0, leading with the measurement rather than the fix: 120 082 104 B -> 7 036 B on
the shipped v0.229.0, reproduced before anything was built.

CONTEXT records three rulings: hollowness is a MANIFEST question and never a size question; the
guard fences one shape and NOT shrinking, because the derived-copy rebuild is a design decision; and
the rehydrate happens inside the restore because a follow-up job races the 5-minute capture. Plus
the shape the live run taught: a warning that fires on everything costs the same as the comforting
lie it replaces.

README documents the refusal, what each surface says, and why the capture job is deliberately not
guarded. REPORT leads with Part 1's result, carries the six red-proofs, the per-row Scenario D table
with its seven-app control, and eight observations including R-404 filed-not-acted-on and three
mistakes of mine recorded rather than tidied away.
This commit is contained in:
2026-08-31 14:39:13 +02:00
parent b48a7fa326
commit 1cfdde968f
4 changed files with 371 additions and 250 deletions
+31
View File
@@ -1220,6 +1220,37 @@ DIRECTORY, so the same restore that always worked from the primary now works fro
of 1, 28.65 s, an accented filename byte-identical, and `secrets recovered=2/2` with the guest's
`app.yaml` also moved aside:
`felhom.eu/documentation/audits/DRILL-r102-tier2-unit-2026-08-31/`.
- **Since v0.230.0 it also REFILLS the app's own drive** before returning — see the R-403 note below.
**The nightly copy refuses to replace a complete package with an empty one (R-403, v0.230.0)** —
`internal/backup/r403_hollow.go` + the precondition in `RunTier2`.
**The measurement, because this was run before it was fixed.** On the shipped v0.229.0, on `demo-hp`:
an app's Tier-2 copy went from **120 082 104 B (4 database dumps + 3 volume tars) to 7 036 B (none of
either) in one nightly run**, reported as a success. `RunTier2` guarded the unit leg with `os.Stat`
alone, `rsyncMirror` is `rsync -a --delete`, and nothing compared the two sides — and an empty
recovery unit is a folder that exists. Evidence:
`felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/`.
- **The predicate asks the MANIFEST, never the byte size.** `unitCarriesData` is true when the unit's
manifest lists a database dump or a volume tar. Absent or unparseable manifest ⇒ hollow, fail closed.
- **The refusal is one shape only:** source hollow AND destination not. complete→complete,
complete→hollow and hollow→hollow all mirror as before. **`--delete` stays and the data legs are
untouched** — §8 row 5's derived-copy rebuild is a design decision and a copy that legitimately
shrinks still shrinks.
- **The other legs still run** and the run is not failed; a preserved package must not cost the
customer their file legs or raise a red alarm on a healthy box.
- **What the surfaces say.** A preserved package is older than the run that preserved it, so the
per-app card carries „A másolat adatcsomagja régebbi, mint a legutóbbi mentés…", and the
„Teljes visszaállítás a másolatból" confirm names the **package's own date** (from the mirrored
manifest's `created_at`) plus a `FIGYELEM` clause saying why. The restore **outcome** names the same
date. Only an app whose leg was actually preserved shows any of it — a warning that fires on
everything costs the same as the comforting lie it replaces.
- **The cause is closed too:** `RestoreTier2Unit` refills an absent or hollow primary unit from the
mirror **inside the call**, because the hollow manifest was written two seconds later by the
5-minute capture job. Never over a complete primary, never after a failed restore, and **the capture
itself is not guarded** — it describes reality, and with the primary refilled there is nothing hollow
left to describe.
**Per-app Tier-2 config panel (v0.57.0)** — `GET/POST /stacks/{name}/backup`
(`internal/web/tier2_config_handler.go` + `templates/tier2_config.html`). The "2. mentés" row's