docs(v0.232.0): CHANGELOG, three CONTEXT rulings, README (R-411/408/407, R-414, R-412a)
gates / gates (push) Successful in 14s
gates / gates (push) Successful in 14s
This commit is contained in:
+57
-1
@@ -7,7 +7,63 @@
|
||||
>
|
||||
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
|
||||
|
||||
Last updated: 2026-08-31 (v0.231.0 — R-87: the box proves its own off-site copy still holds something)
|
||||
Last updated: 2026-09-01 (v0.232.0 — R-411/R-408/R-407 the lock family, R-414 reachability, R-412a wording)
|
||||
|
||||
> **2026-09-01 — v0.232.0. THREE RULINGS.**
|
||||
>
|
||||
> **1. `restic stats` TAKES A REPOSITORY LOCK, and that is the fact the whole R-411 chain rested on.**
|
||||
> Nobody had it. Clean-room measured on demo-hp 2026-08-31: nothing else running, four invocations,
|
||||
> the sampler reads `locks=1`. A customer FULL restore shells `stats` in its size probe, so it holds a
|
||||
> lock — and until v0.232.0 it held no single-writer flag, so the integrity check was not blocked, ran,
|
||||
> met that lock, and `resticStep` removed it with `unlock --remove-all` while logging *"a stale
|
||||
> exclusive lock left by a previous crash"*. There was no crash. Also measured, and recorded so the
|
||||
> next reader does not re-derive it: `restic check` takes a lock; `restic snapshots` and `restic list`
|
||||
> do **not**.
|
||||
>
|
||||
> **2. THE INVARIANT IS PINNED BY A WALK, NOT BY A COMMENT — and the walk is the deliverable, not the
|
||||
> acquire.** `offbox_integrity.go:28` asserted *"Every off-site operation takes `acquireRunning`"* from
|
||||
> v0.227.0 and it was false for months, which is the ninth instance of this project's most-repeated
|
||||
> class. `TestR408_EveryOffsiteEntryPointTakesTheFlagOrIsRegistered` is an AST pass over
|
||||
> `internal/backup` — deliberately not `strings.Contains`, because a commented-out call still contains
|
||||
> the string. **On its first run it found three entry points nobody had named**:
|
||||
> `OffboxRestorePrepareFull` (the request the customer's UI reaches FIRST, and the one that shells
|
||||
> `stats`), `RestoreSharesScratch` (R-411's exact shape on the shares tier, with a live caller) and
|
||||
> `RestoreOffbox` (no caller today). All four now take the flag. `OffsiteInventoryList` is registered
|
||||
> EXEMPT with its reason — `snapshots` only, measured not to lock, and flagging it would make a page
|
||||
> refuse to load during a backup for no safety gain. **Adding a line to `offsiteExempt` is a deliberate
|
||||
> act and belongs in the commit that adds it.**
|
||||
>
|
||||
> **3. R-414's BRANCH, AND THE EVIDENCE FOR IT.** The question was whether `offboxRestoreScratchDir`
|
||||
> was MISSED by R-356's unification or EXCLUDED on purpose. **Neither label fits: it was consciously
|
||||
> OUT OF SCOPE.** R-356's own commit (`08eb1a6`) says so in its test comments — *"the prepared scratch
|
||||
> still resolves to the registered storage path … only the DESTINATION moves, which is precisely what
|
||||
> this change is about"* — and every one of its fixtures assumed a registered storage path exists.
|
||||
> `demo-felhom`, with `storage_paths: []`, is the case it never had. It was **never ruled out on
|
||||
> state-only grounds**: the one comment about a `systemDataPath` fallback belonged to
|
||||
> `PlaceOffsiteRestore`, concerned bulk **userdata**, and R-356 deleted it deliberately. This
|
||||
> function's own documented exclusion is `cfg.Paths.DataDir` — the **rootfs** — a different filesystem.
|
||||
>
|
||||
> **So §6.3's `[DESIGN]` rule applies and now has a FOURTH consumer.** But the fallback is **SCOPED**,
|
||||
> because the two callers ask different questions and one predicate answering both is the R-356 defect
|
||||
> itself: **unit-only** may fall back (§7 records as `[FACT]` that a driveless app's unit already lives
|
||||
> on `systemDataPath` indefinitely and that the same-device placement is *"intended, not a defect"*);
|
||||
> **full** keeps the R-252 refusal, because it pulls bulk userdata onto a state-only tier (§2.2).
|
||||
>
|
||||
> **AND THE SILENCE ENDS EITHER WAY.** `ProofResultCannotRun` is recorded through
|
||||
> `RecordProofVerdict`, so `last_proof_result` is never ABSENT — absent already means *"controller
|
||||
> older than v0.231.0"*, and giving one field two meanings is the `StatsKnown` trap one level up. It
|
||||
> does **not** advance per-snapshot due-ness: nothing was proved, and marking one proved would stop the
|
||||
> app being retried once a drive is finally registered.
|
||||
>
|
||||
> **A MISTAKE OF MINE, RECORDED BECAUSE LIVE VALIDATION IS WHAT CAUGHT IT.** The fallback resolved a
|
||||
> scratch that `removeProofScratch` then refused to delete — its accepted-roots list is built from
|
||||
> REGISTERED drives, and a driveless box has none. Observed on `demo-felhom`: *"refusing to remove …
|
||||
> it is not inside a proof root"*, with the copy still on disk. Every nightly proof would have left one
|
||||
> behind, on exactly the boxes the fallback exists for. **The unit tests all registered a drive, so
|
||||
> none of them could see it.** Fixed, and pinned by a pair — one that the copy IS removed on a
|
||||
> driveless box, one that a path outside every proof root is still REFUSED, so the fix is not a
|
||||
> widening into uselessness.
|
||||
|
||||
|
||||
> **2026-08-31 — v0.231.0. FOUR RULINGS, recorded so none is re-litigated.**
|
||||
>
|
||||
|
||||
Reference in New Issue
Block a user