v0.217.0: prefill from the app's own backup, where-the-data-goes on deploy, bounded inventory fan-out
gates / gates (push) Successful in 10s

Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok,
go vet clean, -race clean on the changed package - all run and read BEFORE this commit.

PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN
backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT
installed in this case, so there is no own drive to ask) and reads manifest.json plus the
captured compose/app.yaml. Local file reads only: no network, no restic, no restore.
RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live
deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as
"what your backup says" would be a fabricated fact. The prefill is labelled as coming from the
backup and stays editable: a memory, not a lock.

PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before
the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage
field; the other 40 have none and their data goes to the system drive, which no screen said.
Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for
the 40-class a recorded placement is a fact to state, never a value written into a field that
does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification.

PART 4 - measured before theorising, on the live off-site target:
  snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags
  => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds.
The cause is the shape already on file, so the per-app size calls now run concurrently,
BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent
UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted.
OffsiteInventoryList had no test at all before this.

TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a
template doing index/eq against an undefined key errors at RENDER time: green build, green vet,
green suite, 500 on the page. Four render tests, one per branch, because the existing deploy
render test only renders AutoFields and never reaches these blocks.

RED-PROOFS, mutation asserted applied then reverted to 0:
  A   three template guards dropped (count asserted 3) -> the blank form returned
  P4  inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential

DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory +
what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md
overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first.

NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit
carries no db_dumps and no volume_dumps still reports a bare completion.
This commit is contained in:
2026-08-21 21:29:01 +02:00
parent 985388c6e9
commit f94543ee5c
14 changed files with 1508 additions and 396 deletions
+25
View File
@@ -718,6 +718,31 @@ Per-app export creates a self-contained `.fab` file (tar.gz, optionally encrypte
The backup system implements a **3-2-1 backup architecture**. Each tier is a **complete,
self-sufficient backup** — any single tier can fully restore an app.
**The restore carries the customer's own previous answers (v0.217.0, R-351).**
`internal/backup/offbox_placement.go`. Every recovery unit's `manifest.json` records `drive` and
`namespace_root`, and its `compose/app.yaml` records `SUBDOMAIN`/`DOMAIN`. Until v0.217.0 nothing read
them back, so a restore into a destination different from the recorded one **succeeded silently**.
- **`CheckPlacement`** compares the recorded drive against where the restore is about to write,
**before the safety dump and before the first byte**. A difference is **named — both values** — and
refused. The customer may proceed deliberately with `ack_placement`, a **separate** form field from
`confirm=1`: one click must not carry two decisions.
- **An UNKNOWN recording is never a mismatch.** A pre-field or unreadable manifest falls back to the
previous behaviour rather than blocking, and is never rendered as an empty value.
- **The not-installed refusal names where the data belonged**, read from the prepared scratch.
- **The deploy page prefills the address and data folder from the app's own backup**
(`RecordedUnitForStack`), labelled as coming from the backup and still editable — a memory, not a
lock. `RecordedAddress.Known()` requires **both** halves: an absent `SUBDOMAIN` would otherwise
surface the *catalog's* default as though the customer had chosen it.
- **Where the app's data will live is stated on the deploy page before the button is pressed**
(R-352). Measured 2026-08-21: 13 of 53 catalogue templates declare a storage field; the other 40
have none and their data goes to the system drive. **Visibility only — no placement changed.**
- **Starting a restore is gated by `restoreOpBlocked()`**, which reads the display flag as well as the
concurrency flag. Before v0.217.0 a second press started a second run and was told it had.
- **The off-site listing's per-app size calls run concurrently, bounded to 4.** Measured before the
change: 2605 ms + 5 × 2697 ms ≈ 16 s. The bound protects the Storage Box's session cap; a refused
size call returns 0, which under-reports rather than fails visibly.
**The reserve — per-app backup admission (v0.192.0 decision B2, widened by v0.193.0 / R-181).**
`internal/backup/admission.go`. Since the `mp1`→`mp0` merge (R-165) local backups and Docker's
data-root share one filesystem, so an unbounded backup write is a stopped box rather than a slow one.