docs(v0.229.0): R-102 + R-103 — CHANGELOG, CONTEXT rulings, README, REUSE, REPORT
gates / gates (push) Failing after 13s

CHANGELOG v0.229.0. CONTEXT records three rulings: the source moves and the destination does not;
two predicates and not one wider one (R-356's cost restated); and a destructive operation reached
from a non-destructive surface carries the difference in the CONFIRM, not the label. README documents
the new action and route and corrects the coverage note to the measured count. REUSE maps the four
unit-directory-relative primitives and the new manager methods, with the traps.

REPORT covers the live drill on demo-hp (docmost, class B, primary unit moved aside — 3 volumes of 3
and 1 database of 1 in 28.65 s, accented filename byte-identical verified as hex, the app reading its
own row over TCP; Scenario D with the guest app.yaml also aside, secrets recovered=2/2), the settled
count (A=7 B=45 C=1, and why the earlier 9/43/1 was wrong), the five named red-proofs, and seven
observations including R-403 and a process error of mine that changed the box and is now in memory.
This commit is contained in:
2026-08-31 12:21:28 +02:00
parent 4c8f0d2919
commit 8aa95b5831
5 changed files with 413 additions and 406 deletions
+101
View File
@@ -1,3 +1,104 @@
## v0.229.0 — the second drive's copy becomes a way back (2026-08-31, R-102 + R-103)
**MinAgent: 0.129.0** (unchanged)
### The Tier-2 unit mirror is restorable (R-102)
Tier-2 has written a full `recovery-unit/` mirror — the app's definition, its portable secrets, its
database dump and its named-volume tars — to `<dest>/backups/secondary/<app>/recovery-unit/` on every
run for months. **No code path read it.** Every reader of a recovery unit could only name a path under
`backups/primary/`, because `appbackup/paths.go` joined that segment literally.
**The failure that made it worth doing:** Tier-2 exists for the loss of the primary drive, and in
exactly that loss the primary unit is gone while the mirror survives — unreachable by any customer
action (`07-backup-architecture.md` §6.3, §7.2). For the **45** class-B apps in the catalogue (Part 3
below) that is the whole of their data.
- **`appbackup`** gains four unit-directory-relative primitives — `UnitComposeDir`,
`UnitManifestFile`, `UnitDBDumpDir`, `UnitVolumeDumpDir`. The four `(nsRoot, stackName)` helpers
become thin wrappers over them and return byte-identical strings; every existing caller compiles
untouched. Pinned by `TestR102_PathWrappersAreByteIdenticalToToday` against hand-written literals,
not against the helpers under test.
- **`Manager.RestoreFromRecoveryUnitAt(stack, unitDir)`** holds the whole body;
`RestoreFromRecoveryUnit(stack)` is the thin caller naming the primary unit. ONE implementation, two
callers — the rule `restoreDockerVolumesFrom` already states beside itself, and for the same reason.
- **THE SOURCE MOVES; THE DESTINATION DOES NOT.** `unitDir` changes only where the manifest, the
compose capture, the `.sql` and the tars are READ from. Data still lands in the live Docker volumes
and the live database container, and the definition still in the guest. A restore that also
relocated the app's data would be a migration.
- **Unchanged and pinned:** the R-47 mutation order (stop → volumes → recreate → DB-only start →
replay → start), the secret reconciliation with unit-over-guest precedence, the fail-closed data-key
gate, and the no-unit fallback to `RestoreApp` with its `CountsUnknown` handling.
- **`Manager.RestoreTier2Unit(stack)`** resolves the recorded copy, refuses **fail-closed** unless the
mirror carries a parseable `manifest.json` — *a directory that exists is not a package* — and
delegates. The single-writer flag is taken inside `RestoreFromRecoveryUnitAt`, not beside it.
- The 35-minute DB-replay bound is now named once (`dbReimportTimeout`), so the two bounded entry
points cannot drift in how long a wedged import may hang a restore.
- The unit DIRECTORY is now logged on every restore. Which copy a restore read from is a real question
with two answers, and an absent log line is not evidence.
### The refusal becomes an action (R-103)
An app with no file legs but a full mirror was told to press a button on a **different page**.
- **`POST /backup/tier2/unit-restore`** + `backupTier2UnitRestoreHandler`: same guards, same
`restoreOpBlocked()` refusal (R-351b), same async shape as the file restore beside it, plus the
fail-closed pre-flight so the app is **never stopped** for a mirror that could not be opened.
- **`Tier2Coverage` gains `UnitRestorable`, and `CanRestore()` is NOT widened.** It still answers only
*"can the additive file restore run?"*. One predicate answering two questions is **R-356**, which
refused 40 running apps for months. `HasUnit` also keeps its old meaning — a half-copied mirror is
still unread data the file restore must disclose, even though the unit restore refuses it.
- **The row offers the action where the refusal was**, in `btn-danger-outline`, as a SEPARATE button.
The two are not merged: one adds what is missing, the other overwrites. **The confirm carries that
difference in words** and names the copy's date — and says so differently when that date is only an
ATTEMPT (R-101). It is assembled from named Go constants rather than inside an HTML attribute, so a
test asserts it verbatim; `fmtTimeStr` now delegates to a package-level `fmtRFC3339Local` so the
confirm and the outcome cannot render one date two ways.
- **`tier2NoCoverageMsg` is NARROWED** to the case that remains — no legs *and* no openable unit — and
still names the route that works. **`tier2UnitNotCoveredMsg` is NOT deleted:** it is appended where
the FILE restore ran and is still exactly true of it.
- The outcome reuses `unitRestoreOutcomeMsg` unchanged and appends which copy overwrote the live data.
### The count, settled (Part 3)
`07-backup-architecture.md` §6.2 recorded **two** Tier-2 coverage counts that disagreed — 9/43/1 and
7/45/1 — both unresolved. Counted at catalogue `459766cb16395fd1d1a66282f5cc6da59ead5924` by running
the PRODUCTION rule (`LoadMetadata` → `ParseComposeClassifiableBinds` → `ClassifyBinds` →
`ComputeCaptureSet` at `TierSecondary`) over all 53 templates:
**A = 7 · B = 45 · C = 1.** The INV Part B.1 enumeration was right. The two apps the C9-F1 Phase-0
count put in A are **radarr and sonarr**: both bind `${USERDATA_PATH}` paths **writably**, so the
`:ro`-default rule Phase 0 says it applied to plex/jellyfin/emby/navidrome does not catch them — they
are excluded by an **explicit** `class: excluded` entry instead. Class C is **bentopdf**, which
declares no volumes and no namespace binds at all. No catalogue file was changed.
### Live drill — demo-hp, endpoint level
`documentation/audits/DRILL-r102-tier2-unit-2026-08-31/` (in `felhom.eu`). docmost, class B, its
Tier-2 run reporting **0 leg(s)**. With the **primary unit moved aside** the restore returned 3 volumes
of 3 and 1 database of 1 from the secondary mirror in **28.65 s**; an accented Hungarian filename came
back byte-for-byte (verified as hex, not as rendered text — R-364) and the app read its own row **over
TCP with its own credential**. The post-backup discriminator was **gone**, so the replay was real.
**Scenario D** repeated it with the guest's `app.yaml` moved aside: `secrets recovered=2/2` from the
mirrored unit — this closes `00-capability-map.md`'s open *"not exercised live"* clause for Tier-2's
own cross-drive copy of a secret-bearing unit.
**Filed, not fixed — R-403.** Two seconds after a restore that ran with the primary unit absent, the
5-minute status refresh (`captureAllRecoveryUnits`) rewrote the primary unit from a drive with no
dumps, producing a manifest carrying `"db_dumps": []` and `"volume_dumps": null`. The ordinary restore
then read it and honestly reported that the backup held only settings. The dangerous half — that the
next Tier-2 run would mirror that hollow unit over the good secondary copy, since `rsyncMirror` carries
`--delete` — **was not tested and is recorded as unverified.**
### Tests
26 new test functions (1606 → 1632). Red-proofs run and reverted: **A1** (change one wrapper's join →
fails on all three fixtures), **A5** (swap the volume replay and the recreate → fails on the sequence),
**B2** (point the Tier-2 reader back at the primary → fails, and fails again with `permission denied`
once the primary tree is unreadable), **C1** (widen `CanRestore` to include `HasUnit` → the unit-only
cases fail), **D6** (drop `EndRestoreOp` from the handler's goroutine → *"the restore never published a
result"*). The Tier-2 fixtures build their mirror with the production `RunTier2`, so the claim under
test is *the copy Tier-2 writes is the copy this restore reads*.
## v0.228.0 — the check reads the data, and the debug page stops lying (2026-08-31, R-399 + R-400)
**MinAgent: 0.129.0** (unchanged)