v0.217.0: prefill from the app's own backup, where-the-data-goes on deploy, bounded inventory fan-out
gates / gates (push) Successful in 10s

Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok,
go vet clean, -race clean on the changed package - all run and read BEFORE this commit.

PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN
backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT
installed in this case, so there is no own drive to ask) and reads manifest.json plus the
captured compose/app.yaml. Local file reads only: no network, no restic, no restore.
RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live
deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as
"what your backup says" would be a fabricated fact. The prefill is labelled as coming from the
backup and stays editable: a memory, not a lock.

PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before
the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage
field; the other 40 have none and their data goes to the system drive, which no screen said.
Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for
the 40-class a recorded placement is a fact to state, never a value written into a field that
does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification.

PART 4 - measured before theorising, on the live off-site target:
  snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags
  => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds.
The cause is the shape already on file, so the per-app size calls now run concurrently,
BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent
UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted.
OffsiteInventoryList had no test at all before this.

TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a
template doing index/eq against an undefined key errors at RENDER time: green build, green vet,
green suite, 500 on the page. Four render tests, one per branch, because the existing deploy
render test only renders AutoFields and never reaches these blocks.

RED-PROOFS, mutation asserted applied then reverted to 0:
  A   three template guards dropped (count asserted 3) -> the blank form returned
  P4  inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential

DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory +
what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md
overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first.

NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit
carries no db_dumps and no volume_dumps still reports a bare completion.
This commit is contained in:
2026-08-21 21:29:01 +02:00
parent 985388c6e9
commit f94543ee5c
14 changed files with 1508 additions and 396 deletions
+63
View File
@@ -1,3 +1,66 @@
## v0.217.0 — the restore knows where the data lived (2026-08-21, R-351/R-352/R-353)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
**Found by walking the screens on a rebuilt `demo-hp`, not by reading them.** A person restored an
app onto a rebuilt machine and had to remember two things the backup already held: the web address
and the data folder. Neither can be changed after installation without deleting the app and its data.
**The blindness (R-351a).** Every recovery unit's `manifest.json` has carried `drive` and
`namespace_root` since schema 1 (`internal/backup/recovery_unit.go:48-49`), written at capture from
the app's own live placement. `grep -rE '\.Drive\b|\.NamespaceRoot\b' --include=*.go` found **no
non-test reader anywhere in the repository**. The reconstitution opened that very manifest
(`offbox_reconstitute.go:235`) purely for the coherence stamp, then resolved its destination from the
LIVE app instead. **A restore into a destination different from the one the backup recorded therefore
succeeded silently, under a green message.** One side wrote the fact; the other never received it.
Now: `internal/backup/offbox_placement.go` — `CheckPlacement` (pure, total), `RecordedUnitForStack`,
`PlacementMismatchMessage`. The comparison happens **before the safety dump and before the first
byte**. A mismatch is **named** — both values, never "a destination differs" — and refused; the
customer may proceed deliberately via `ack_placement`, a **separate** field from `confirm=1`, because
one click must not carry two decisions. An **unknown** recording is never a mismatch: refusing on an
absence would strand every pre-field unit. The not-installed refusal (R-253) now names the recorded
drive. The deploy page prefills the address and folder **from the app's own backup**, labelled as
such, and still editable — a memory, not a lock.
**The second press (R-351b).** All seven restore handlers gated on `backupMgr.IsRunning()` — the
CONCURRENCY flag, which the restore goroutine acquires *inside* itself (`offbox_reconstitute.go:180`)
**after** the handler has returned. Established with a test before any change: the reconstitute and
place handlers both answered „…elindult" and **overwrote the first restore's op and stack**. The
wizard had read the correct flag since v0.154.0 and explained why in a comment; the handlers were
never moved over. `Server.restoreOpBlocked()` now reads **both** flags — the display flag covers the
whole off-box restore, the concurrency flag is the only one the nightly backup holds.
**The invisible result (R-351c).** The page *does* refresh; the defect was the RESULT. The banner
gated its terminal state on a page-local `sawRunning`, so a restore that finished before the page was
opened — or inside one 3 s poll — was shown to **nobody**. The 2026-08-21 OpenGist restore took
**8.666 s** and no screen ever said it completed. `RestoreOpStatus.LastRecent` now carries the
server's verdict, and `RestoreResultWindow` moved to `internal/backup` with `internal/web`'s constant
as an alias: **one expression, two surfaces.** The wizard's self-contradiction — that the state
refreshes automatically *and* that you must refresh the page — is gone.
**Where the data goes, said out loud (R-352).** Measured: **40 of 53** catalogue templates declare no
data path, and for those the deploy page offered no storage field and the configured default was
never consulted — `GetDefaultStoragePath()` has three non-test callers and **none of them places
data**. The deploy page now **states where the app's data will live before the button is pressed**.
**Visibility only: no placement changed and nothing was migrated.** The specification for the rest is
filed at `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`.
**The 16-second page (R-351d).** Measured on the live off-site target before changing anything:
`snapshots --json` 2605 ms once, then `stats` **2697 ms per app, sequentially** — 2605 + 5×2697 ≈
**16.1 s**. The per-app size calls now run concurrently, **bounded to 4**: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns 0, which silently *under-reports* the
customer's data rather than failing visibly. The bound is asserted by test, not just the speed.
**Red-proofs, each mutation asserted applied and reverted.** B with **both** guards removed **was
seen starting a restore with no drive attached** — no error, full 3.00 s run, writing into
`/tmp/mutant-destination`. C returned the silent divergent restore; E returned the fabricated empty
prefill; A returned the blank form; D forced on broke **8** ordinary reconstitute tests, proving the
guard is reachable in both directions; Part 4 reverted to sequential failed the concurrency assertion.
**Known and NOT fixed here, filed as R-353 and named as the next session's first item:** a restore
whose unit carries no `db_dumps` and no `volume_dumps` reports a bare completion. The OpenGist unit
contained configuration and nothing else, and the outcome said only that it had finished.
## v0.216.0 — one physical disk, one verdict (2026-08-14, R-335)
**MinAgent: 0.129.0**