Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok, go vet clean, -race clean on the changed package - all run and read BEFORE this commit. PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT installed in this case, so there is no own drive to ask) and reads manifest.json plus the captured compose/app.yaml. Local file reads only: no network, no restic, no restore. RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as "what your backup says" would be a fabricated fact. The prefill is labelled as coming from the backup and stays editable: a memory, not a lock. PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage field; the other 40 have none and their data goes to the system drive, which no screen said. Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for the 40-class a recorded placement is a fact to state, never a value written into a field that does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification. PART 4 - measured before theorising, on the live off-site target: snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds. The cause is the shape already on file, so the per-app size calls now run concurrently, BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted. OffsiteInventoryList had no test at all before this. TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a template doing index/eq against an undefined key errors at RENDER time: green build, green vet, green suite, 500 on the page. Four render tests, one per branch, because the existing deploy render test only renders AutoFields and never reaches these blocks. RED-PROOFS, mutation asserted applied then reverted to 0: A three template guards dropped (count asserted 3) -> the blank form returned P4 inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory + what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first. NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit carries no db_dumps and no volume_dumps still reports a bare completion.
11 KiB
REPORT — controller v0.216.0 → v0.217.0: the restore knows where the data lived
Date: 2026-08-21 · Task class: Implementation (Establish-first) · Repos touched:
felhom-controller (code), felhom.eu (register + specification, documentation only)
Register: R-351 CLOSED, R-352 partly closed (visibility shipped, placement open), R-353 OPEN — next session's first item. Ceiling moved R-350 → R-353.
0. Baselines, re-established live (nothing in the prompt was trusted)
| Value | How | |
|---|---|---|
| controller | 0.216.0 | docker inspect felhom-controller on guest 9201 |
| agent | 0.130.0 | felhom-agent --version on the host |
| golden | 0.216.0 | hub artifact manifest |
| register ceiling | R-350 (prompt said R-327 — stale) | grep -rhoE '\bR-[0-9]{1,4}\b' --include=*.md |
demo-hp |
reinstalled 2026-08-21, PVE 9.2.2 | break-glass root from hub host_recovery |
1. Part 0 — what the backup can tell us
1. Are the drive and folder readable before restoring? YES. RecoveryManifest.Drive /
.NamespaceRoot (internal/backup/recovery_unit.go:48-49). Confirmed on real data:
/mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json
drive='/mnt/sys_drive' namespace_root='/mnt/sys_drive/felhom-data'
/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/manifest.json
drive='/mnt/felhom-drives/hdd_1' namespace_root='/mnt/felhom-drives/hdd_1'
Cost, two prices: drive attached → a plain local file read, free. Off-site only → one
snapshots latest --tag plus one unit-only restore --include per app, not one for all; the
unit-only restore already exists and is already the default (offbox_restore.go:264), needing a
registered non-network drive for scratch and 2 GB free (offbox_restore.go:242).
2. Is the web address recoverable? YES, recorded — not inferred. app.yaml is captured into
every unit. Real values: opengist SUBDOMAIN: gist / DOMAIN: enkisfelhom.hu, calibre-web books.
Caveat carried into the design: an absent SUBDOMAIN makes the live path fall back to the catalog
default (stacks/deploy.go:88-90) — a catalog guess, not the customer's answer. Treated as UNKNOWN.
3. What does the restore compare today? NOTHING. A mismatched restore succeeded silently.
grep -rnE '\.Drive\b|\.NamespaceRoot\b' --include=*.go | grep -v _test returned only the
appbackup.NamespaceRoot function and Tier2Target's unrelated field.
4. Did the 21 August OpenGist restore complete? Yes — on the second try, by a different route.
16:32:43 off-box full-restore prepared for opengist (182.3 KB)
16:33:34 restored opengist (9e38b84c, full=true) -> .../backups/offsite-restore/opengist
16:37:14 [ERROR] off-box reconstitute opengist: ...nincs telepitve...
16:39:16 [WARN] Restore requested (async): stack=opengist, snapshot=helyi
16:39:25 Restore-from-unit completed: opengist in 8.666896042s
No screen said so — the answer existed only in docker logs. That is R-351c, and the contents
of what came back is R-353.
2. §2's ruling revisited
The R-253 comment (templates/backups_restore.html:104-110) removed the promise to reinstall,
because reconstitution writes to the app's own data path, "which exists only once the customer has
chosen a drive during deploy — the restore has no answer to that question and must not invent one."
The reasoning was correct. Its premise no longer holds. The restore does have an answer, and it is
the customer's own previous answer: manifest.Drive + the captured app.yaml. Reversed, openly,
recorded in R-351. What is not reversed: the restore still does not deploy the app for you —
deploy-then-restore as one atomic act stays out of scope (a half-failure leaves a half-installed app,
and the catalogue may have moved on since the backup, a hazard that must stay visible).
3. Part 1 — established, then ruled
The premise changed twice. It is not a typed path, and not a customer failing to choose.
| # | Measured untruth | Evidence |
|---|---|---|
| 1 | 40 of 53 templates declare no data path | 53 total, 13 with env_var: HDD_PATH; negative control: grep -c HDD_PATH on opengist's .felhom.yml and compose = 0 |
| 2 | The configured default is never consulted when placing data | GetDefaultStoragePath() has 3 non-test callers: metrics (main.go:410), dashboard panel (server.go:733), .fab import (handler_export_upload.go:154). grep over stacks/deploy.go+manager.go → nothing. Field comment // new apps use this by default (settings.go:453) has never been true |
| 3 | Tier-1 backup follows the data onto the same disk | GetAppDrivePath → systemDataPath (backup.go:324-334). Tier 2 refuses that posture outright (tier2.go:329) |
| 4 | The Drives count cannot include most apps | countAppsUsingPath matches Env["HDD_PATH"] only (handlers.go:2118) — "1 alkalmazás használja" means "1 of the apps that CAN" |
Protection. Whole-machine tier: covered — df shows /mnt/sys_drive, /var/lib/docker and
/var/lib/felhom all on pve-vm-9201-disk-1 = mp0, backup=1 in 9201.conf. Off-site tier:
covered by code, NOT observed — runVolumeDumps (backup.go:607+) passes its gates for these
apps on paper, but every unit reported volume_dumps: None, including calibre-web on the data
drive, because no nightly run had happened on a one-hour-old box. Recorded as unknown, not fine.
Ruled: visibility only tonight. My own earlier recommendation — refuse deployment until a drive
is registered — is WITHDRAWN: it assumed the customer had failed to choose; they had no choice.
Specification filed at felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md.
4. Red-proofs — every mutation asserted applied, then reverted to 0
| Scenario | Mutation | Outcome |
|---|---|---|
| B | both guards removed (MUTANT-B1+B2, count asserted 2) |
A restore WAS seen starting with no drive attached — no error, full 3.00 s run, wrote into /tmp/mutant-destination |
| C | Mismatch = false && … |
FAIL — the silent divergent restore returned |
| E | Known() { return true } |
FAIL — the fabricated empty prefill appeared |
| A | 3 template guards dropped (count asserted 3) | FAIL — the blank form returned (default subdomain, default drive, no notice) |
| D | Mismatch = true + unknown short-circuit removed |
FAIL — 8 ordinary reconstitute tests broke: the guard is reachable in both directions |
| 3a | false && restoreOpInFlight(st) |
FAIL — both handlers reported a started restore and overwrote the first op |
| 3b | st.LastRecent = false |
FAIL — just_finished, inside_the_window |
| P4 | inventorySizeConcurrency = 1 |
FAIL — "peak in flight was 1", elapsed 282 ms = the sequential cost |
Explicitly: yes, a restore was seen starting with no drive attached — but only once both guards were removed. The first attempt removed one and the new mismatch guard caught it; that proved defence in depth, not the stated observable, so it was redone.
Honesty note on D: the existing reconstitute fixtures write a schema-1 manifest with no
Drive, so they are scenario-E shaped. The matching case is covered in the scenario table, not
by them.
5. Part 3 — and a correction
A second press really did start a second run. Not refused deeper. Both offboxReconstituteHandler
and offboxPlaceHandler answered „…elindult" and overwrote the first restore's op/stack.
I was wrong about the refresh, and correct it here. I first reported that the list page had
neither banner nor poll. Both are present (backups_restore.html:10 and the script block at 214), the
endpoint exists (api/router.go:277), and all three banner pages poll. My grep pattern missed
restore_banner_js. The page refreshes. The real defect was the result: sawRunning hid every
terminal state from anyone who was not already watching.
Status line as it now behaves: the banner shows a running op, and now also shows a terminal result
that finished within RestoreResultWindow (10 min) regardless of whether this page saw it start.
6. Part 4 — measured, then fixed
Measured on the live off-site target (u629488-sub3.your-storagebox.de:23) before any change:
LEG1_snapshots_json_ms=2605 rc=0
snapshots=18 distinct_app_tags=5
LEG2_stats_calls=4 total_ms=10790 per_call_ms=2697
=> 2605 + 5 * 2697 = ~16.1 s
The cause is the shape on file, and the fix is contained: the per-app stats calls now run
concurrently, bounded to 4 (offbox_inventory.go). The bound is the safety property — the target
is a Storage Box with a session cap, and a refused size call returns 0, which under-reports the
customer's data rather than failing visibly. Peak-in-flight is asserted by test, with -race clean.
OffsiteInventoryList had no test at all before this.
One measurement trap worth recording: the first attempt failed with set -e and no message
because restic $BASE was unquoted — sftp.command=ssh … contains spaces and word-split.
7. Hungarian as shipped, bytes confirmed
Verified as hex, valid UTF-8, no BOM, and checked against eleven double-encoding sentinels — BAD = 0.
korábban itt voltak 6b6f72c3a16262616e2069747420766f6c74616b
erősítsd meg alább 6572c59173c3ad747364206d656720616cc3a16262
már fut, ezért most… 6dc3a172206675742c20657ac3a97274206d6f7374206e656d20696e64c3ad74686174c3b3…
8. Gates, tests, commits
python3 controller/scripts/controller_gates.py→ 11/11 OK,GATES_RC=0(hook also ran on push; no--no-verify)go test ./...→ 28 packages ok,SUITE_RC=0;go vetclean;-raceclean on the changed package- Commit 1:
985388c— Part 3 + Part 2's engine (14 files), pushed2fa1efc..985388c
9. What was dropped, named plainly
- Deploy-then-restore as one atomic act — deliberately out of scope, stated in the task and here.
- Placement change and migration for the 40-class — specification filed, ruling is the operator's.
- R-353 — a restore that returns configuration and no data still reports a bare completion. Next session's first item.
- The three items the task listed as out of scope (the re-issue skip on reinstall, the two cosmetic hub bugs, the CI runs failing with no log) were not touched.
10. Observations — noticed, not acted on
offbox_handlers.go:498— the verify-copy delete guard is blind in exactly the same way as the restore guards were. Deleting a verification copy while an off-box restore writes into it is a real hazard. Left alone because it does not start work, and widening the diff unasked is its own risk.- The OpenGist instance was removed by someone between 16:57 and 17:02 UTC, not by this session
(
ScanStacks: found stack "opengist" deployed=falsefrom 17:02:58). The task asked for it to be left in place as evidence. Its unit and manifest survive on/mnt/sys_drive, andprivatebinis now a live specimen of the same class. - A second reconstitution ran at 16:31:48 for
calibre-web(8 files placed, 0 DB dumps replayed) that was not mentioned in the task.