# REPORT — controller v0.216.0 → v0.217.0: the restore knows where the data lived **Date:** 2026-08-21 · **Task class:** Implementation (Establish-first) · **Repos touched:** `felhom-controller` (code), `felhom.eu` (register + specification, documentation only) Register: **R-351 CLOSED**, **R-352 partly closed** (visibility shipped, placement open), **R-353 OPEN — next session's first item**. Ceiling moved **R-350 → R-353**. --- ## 0. Baselines, re-established live (nothing in the prompt was trusted) | | Value | How | |---|---|---| | controller | **0.216.0** | `docker inspect felhom-controller` on guest 9201 | | agent | **0.130.0** | `felhom-agent --version` on the host | | golden | **0.216.0** | hub artifact manifest | | register ceiling | **R-350** (prompt said R-327 — stale) | `grep -rhoE '\bR-[0-9]{1,4}\b' --include=*.md` | | `demo-hp` | reinstalled 2026-08-21, PVE 9.2.2 | break-glass root from hub `host_recovery` | --- ## 1. Part 0 — what the backup can tell us **1. Are the drive and folder readable before restoring? YES.** `RecoveryManifest.Drive` / `.NamespaceRoot` (`internal/backup/recovery_unit.go:48-49`). Confirmed on real data: ``` /mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json drive='/mnt/sys_drive' namespace_root='/mnt/sys_drive/felhom-data' /mnt/felhom-drives/hdd_1/backups/primary/calibre-web/manifest.json drive='/mnt/felhom-drives/hdd_1' namespace_root='/mnt/felhom-drives/hdd_1' ``` **Cost, two prices:** drive attached → a plain local file read, **free**. Off-site only → one `snapshots latest --tag` plus one unit-only `restore --include` **per app**, not one for all; the unit-only restore already exists and is already the default (`offbox_restore.go:264`), needing a registered non-network drive for scratch and 2 GB free (`offbox_restore.go:242`). **2. Is the web address recoverable? YES, recorded — not inferred.** `app.yaml` is captured into every unit. Real values: `opengist SUBDOMAIN: gist / DOMAIN: enkisfelhom.hu`, `calibre-web books`. **Caveat carried into the design:** an absent `SUBDOMAIN` makes the live path fall back to the catalog default (`stacks/deploy.go:88-90`) — a catalog guess, not the customer's answer. Treated as UNKNOWN. **3. What does the restore compare today? NOTHING. A mismatched restore succeeded silently.** `grep -rnE '\.Drive\b|\.NamespaceRoot\b' --include=*.go | grep -v _test` returned only the `appbackup.NamespaceRoot` *function* and `Tier2Target`'s unrelated field. **4. Did the 21 August OpenGist restore complete? Yes — on the second try, by a different route.** ``` 16:32:43 off-box full-restore prepared for opengist (182.3 KB) 16:33:34 restored opengist (9e38b84c, full=true) -> .../backups/offsite-restore/opengist 16:37:14 [ERROR] off-box reconstitute opengist: ...nincs telepitve... 16:39:16 [WARN] Restore requested (async): stack=opengist, snapshot=helyi 16:39:25 Restore-from-unit completed: opengist in 8.666896042s ``` **No screen said so** — the answer existed only in `docker logs`. That is R-351c, and the *contents* of what came back is R-353. --- ## 2. §2's ruling revisited The R-253 comment (`templates/backups_restore.html:104-110`) removed the promise to reinstall, because reconstitution writes to the app's own data path, *"which exists only once the customer has chosen a drive during deploy — the restore has no answer to that question and must not invent one."* **The reasoning was correct. Its premise no longer holds.** The restore does have an answer, and it is the customer's own previous answer: `manifest.Drive` + the captured `app.yaml`. **Reversed, openly, recorded in R-351.** What is *not* reversed: the restore still does not deploy the app for you — deploy-then-restore as one atomic act stays out of scope (a half-failure leaves a half-installed app, and the catalogue may have moved on since the backup, a hazard that must stay visible). --- ## 3. Part 1 — established, then ruled **The premise changed twice.** It is not a typed path, and not a customer failing to choose. | # | Measured untruth | Evidence | |---|---|---| | 1 | **40 of 53** templates declare no data path | 53 total, 13 with `env_var: HDD_PATH`; negative control: `grep -c HDD_PATH` on opengist's `.felhom.yml` **and** compose = 0 | | 2 | The configured default is **never consulted** when placing data | `GetDefaultStoragePath()` has 3 non-test callers: metrics (`main.go:410`), dashboard panel (`server.go:733`), `.fab` import (`handler_export_upload.go:154`). `grep` over `stacks/deploy.go`+`manager.go` → nothing. Field comment `// new apps use this by default` (`settings.go:453`) has never been true | | 3 | Tier-1 backup follows the data onto the **same disk** | `GetAppDrivePath` → `systemDataPath` (`backup.go:324-334`). Tier 2 refuses that posture outright (`tier2.go:329`) | | 4 | The Drives count **cannot** include most apps | `countAppsUsingPath` matches `Env["HDD_PATH"]` only (`handlers.go:2118`) — "1 alkalmazás használja" means *"1 of the apps that CAN"* | **Protection.** Whole-machine tier: **covered** — `df` shows `/mnt/sys_drive`, `/var/lib/docker` and `/var/lib/felhom` all on `pve-vm-9201-disk-1` = `mp0`, `backup=1` in `9201.conf`. Off-site tier: **covered by code, NOT observed** — `runVolumeDumps` (`backup.go:607+`) passes its gates for these apps on paper, but every unit reported `volume_dumps: None`, including `calibre-web` on the data drive, because no nightly run had happened on a one-hour-old box. **Recorded as unknown, not fine.** **Ruled:** visibility only tonight. **My own earlier recommendation — refuse deployment until a drive is registered — is WITHDRAWN**: it assumed the customer had failed to choose; they had no choice. Specification filed at `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`. --- ## 4. Red-proofs — every mutation asserted applied, then reverted to 0 | Scenario | Mutation | Outcome | |---|---|---| | **B** | **both** guards removed (`MUTANT-B1`+`B2`, count asserted **2**) | **A restore WAS seen starting with no drive attached** — no error, full 3.00 s run, wrote into `/tmp/mutant-destination` | | **C** | `Mismatch = false && …` | FAIL — the silent divergent restore returned | | **E** | `Known() { return true }` | FAIL — the fabricated empty prefill appeared | | **A** | 3 template guards dropped (count asserted **3**) | FAIL — the blank form returned (default subdomain, default drive, no notice) | | **D** | `Mismatch = true` + unknown short-circuit removed | FAIL — **8** ordinary reconstitute tests broke: the guard is reachable in both directions | | **3a** | `false && restoreOpInFlight(st)` | FAIL — both handlers reported a started restore and overwrote the first op | | **3b** | `st.LastRecent = false` | FAIL — `just_finished`, `inside_the_window` | | **P4** | `inventorySizeConcurrency = 1` | FAIL — "peak in flight was 1", elapsed 282 ms = the sequential cost | **Explicitly:** yes, a restore was seen starting with no drive attached — but only once **both** guards were removed. The first attempt removed one and the *new* mismatch guard caught it; that proved defence in depth, not the stated observable, so it was redone. **Honesty note on D:** the existing reconstitute fixtures write a schema-1 manifest with **no `Drive`**, so they are scenario-**E** shaped. The matching case is covered in the scenario table, not by them. --- ## 5. Part 3 — and a correction **A second press really did start a second run.** Not refused deeper. Both `offboxReconstituteHandler` and `offboxPlaceHandler` answered „…elindult" and overwrote the first restore's op/stack. **I was wrong about the refresh, and correct it here.** I first reported that the list page had neither banner nor poll. Both are present (`backups_restore.html:10` and the script block at 214), the endpoint exists (`api/router.go:277`), and all three banner pages poll. My grep pattern missed `restore_banner_js`. **The page refreshes.** The real defect was the *result*: `sawRunning` hid every terminal state from anyone who was not already watching. **Status line as it now behaves:** the banner shows a running op, and now also shows a terminal result that finished within `RestoreResultWindow` (10 min) regardless of whether this page saw it start. --- ## 6. Part 4 — measured, then fixed Measured on the live off-site target (`u629488-sub3.your-storagebox.de:23`) **before** any change: ``` LEG1_snapshots_json_ms=2605 rc=0 snapshots=18 distinct_app_tags=5 LEG2_stats_calls=4 total_ms=10790 per_call_ms=2697 => 2605 + 5 * 2697 = ~16.1 s ``` The cause **is** the shape on file, and the fix is contained: the per-app `stats` calls now run concurrently, **bounded to 4** (`offbox_inventory.go`). The bound is the safety property — the target is a Storage Box with a session cap, and a refused size call returns 0, which *under-reports the customer's data* rather than failing visibly. Peak-in-flight is asserted by test, with `-race` clean. `OffsiteInventoryList` had **no test at all** before this. **One measurement trap worth recording:** the first attempt failed with `set -e` and no message because `restic $BASE` was unquoted — `sftp.command=ssh …` contains spaces and word-split. --- ## 7. Hungarian as shipped, bytes confirmed Verified as hex, valid UTF-8, no BOM, and checked against eleven double-encoding sentinels — `BAD = 0`. ``` korábban itt voltak 6b6f72c3a16262616e2069747420766f6c74616b erősítsd meg alább 6572c59173c3ad747364206d656720616cc3a16262 már fut, ezért most… 6dc3a172206675742c20657ac3a97274206d6f7374206e656d20696e64c3ad74686174c3b3… ``` --- ## 8. Gates, tests, commits - `python3 controller/scripts/controller_gates.py` → **11/11 OK**, `GATES_RC=0` (hook also ran on push; **no `--no-verify`**) - `go test ./...` → **28 packages ok**, `SUITE_RC=0`; `go vet` clean; `-race` clean on the changed package - Commit 1: `985388c` — Part 3 + Part 2's engine (14 files), pushed `2fa1efc..985388c` --- ## 9. What was dropped, named plainly - **Deploy-then-restore as one atomic act** — deliberately out of scope, stated in the task and here. - **Placement change and migration for the 40-class** — specification filed, ruling is the operator's. - **R-353** — a restore that returns configuration and no data still reports a bare completion. **Next session's first item.** - The three items the task listed as out of scope (the re-issue skip on reinstall, the two cosmetic hub bugs, the CI runs failing with no log) were not touched. ## 10. Observations — noticed, not acted on - `offbox_handlers.go:498` — the **verify-copy delete** guard is blind in exactly the same way as the restore guards were. Deleting a verification copy while an off-box restore writes into it is a real hazard. Left alone because it does not *start* work, and widening the diff unasked is its own risk. - **The OpenGist instance was removed by someone between 16:57 and 17:02 UTC**, not by this session (`ScanStacks: found stack "opengist" deployed=false` from 17:02:58). The task asked for it to be left in place as evidence. **Its unit and manifest survive** on `/mnt/sys_drive`, and `privatebin` is now a live specimen of the same class. - A second reconstitution ran at 16:31:48 for `calibre-web` (8 files placed, 0 DB dumps replayed) that was not mentioned in the task.