Files
felhom-controller/REPORT.md
T
admin f94543ee5c
gates / gates (push) Successful in 10s
v0.217.0: prefill from the app's own backup, where-the-data-goes on deploy, bounded inventory fan-out
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok,
go vet clean, -race clean on the changed package - all run and read BEFORE this commit.

PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN
backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT
installed in this case, so there is no own drive to ask) and reads manifest.json plus the
captured compose/app.yaml. Local file reads only: no network, no restic, no restore.
RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live
deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as
"what your backup says" would be a fabricated fact. The prefill is labelled as coming from the
backup and stays editable: a memory, not a lock.

PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before
the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage
field; the other 40 have none and their data goes to the system drive, which no screen said.
Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for
the 40-class a recorded placement is a fact to state, never a value written into a field that
does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification.

PART 4 - measured before theorising, on the live off-site target:
  snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags
  => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds.
The cause is the shape already on file, so the per-app size calls now run concurrently,
BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent
UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted.
OffsiteInventoryList had no test at all before this.

TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a
template doing index/eq against an undefined key errors at RENDER time: green build, green vet,
green suite, 500 on the page. Four render tests, one per branch, because the existing deploy
render test only renders AutoFields and never reaches these blocks.

RED-PROOFS, mutation asserted applied then reverted to 0:
  A   three template guards dropped (count asserted 3) -> the blank form returned
  P4  inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential

DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory +
what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md
overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first.

NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit
carries no db_dumps and no volume_dumps still reports a bare completion.
2026-08-21 21:29:01 +02:00

202 lines
11 KiB
Markdown

# REPORT — controller v0.216.0 → v0.217.0: the restore knows where the data lived
**Date:** 2026-08-21 · **Task class:** Implementation (Establish-first) · **Repos touched:**
`felhom-controller` (code), `felhom.eu` (register + specification, documentation only)
Register: **R-351 CLOSED**, **R-352 partly closed** (visibility shipped, placement open),
**R-353 OPEN — next session's first item**. Ceiling moved **R-350 → R-353**.
---
## 0. Baselines, re-established live (nothing in the prompt was trusted)
| | Value | How |
|---|---|---|
| controller | **0.216.0** | `docker inspect felhom-controller` on guest 9201 |
| agent | **0.130.0** | `felhom-agent --version` on the host |
| golden | **0.216.0** | hub artifact manifest |
| register ceiling | **R-350** (prompt said R-327 — stale) | `grep -rhoE '\bR-[0-9]{1,4}\b' --include=*.md` |
| `demo-hp` | reinstalled 2026-08-21, PVE 9.2.2 | break-glass root from hub `host_recovery` |
---
## 1. Part 0 — what the backup can tell us
**1. Are the drive and folder readable before restoring? YES.** `RecoveryManifest.Drive` /
`.NamespaceRoot` (`internal/backup/recovery_unit.go:48-49`). Confirmed on real data:
```
/mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json
drive='/mnt/sys_drive' namespace_root='/mnt/sys_drive/felhom-data'
/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/manifest.json
drive='/mnt/felhom-drives/hdd_1' namespace_root='/mnt/felhom-drives/hdd_1'
```
**Cost, two prices:** drive attached → a plain local file read, **free**. Off-site only → one
`snapshots latest --tag` plus one unit-only `restore --include` **per app**, not one for all; the
unit-only restore already exists and is already the default (`offbox_restore.go:264`), needing a
registered non-network drive for scratch and 2 GB free (`offbox_restore.go:242`).
**2. Is the web address recoverable? YES, recorded — not inferred.** `app.yaml` is captured into
every unit. Real values: `opengist SUBDOMAIN: gist / DOMAIN: enkisfelhom.hu`, `calibre-web books`.
**Caveat carried into the design:** an absent `SUBDOMAIN` makes the live path fall back to the catalog
default (`stacks/deploy.go:88-90`) — a catalog guess, not the customer's answer. Treated as UNKNOWN.
**3. What does the restore compare today? NOTHING. A mismatched restore succeeded silently.**
`grep -rnE '\.Drive\b|\.NamespaceRoot\b' --include=*.go | grep -v _test` returned only the
`appbackup.NamespaceRoot` *function* and `Tier2Target`'s unrelated field.
**4. Did the 21 August OpenGist restore complete? Yes — on the second try, by a different route.**
```
16:32:43 off-box full-restore prepared for opengist (182.3 KB)
16:33:34 restored opengist (9e38b84c, full=true) -> .../backups/offsite-restore/opengist
16:37:14 [ERROR] off-box reconstitute opengist: ...nincs telepitve...
16:39:16 [WARN] Restore requested (async): stack=opengist, snapshot=helyi
16:39:25 Restore-from-unit completed: opengist in 8.666896042s
```
**No screen said so** — the answer existed only in `docker logs`. That is R-351c, and the *contents*
of what came back is R-353.
---
## 2. §2's ruling revisited
The R-253 comment (`templates/backups_restore.html:104-110`) removed the promise to reinstall,
because reconstitution writes to the app's own data path, *"which exists only once the customer has
chosen a drive during deploy — the restore has no answer to that question and must not invent one."*
**The reasoning was correct. Its premise no longer holds.** The restore does have an answer, and it is
the customer's own previous answer: `manifest.Drive` + the captured `app.yaml`. **Reversed, openly,
recorded in R-351.** What is *not* reversed: the restore still does not deploy the app for you —
deploy-then-restore as one atomic act stays out of scope (a half-failure leaves a half-installed app,
and the catalogue may have moved on since the backup, a hazard that must stay visible).
---
## 3. Part 1 — established, then ruled
**The premise changed twice.** It is not a typed path, and not a customer failing to choose.
| # | Measured untruth | Evidence |
|---|---|---|
| 1 | **40 of 53** templates declare no data path | 53 total, 13 with `env_var: HDD_PATH`; negative control: `grep -c HDD_PATH` on opengist's `.felhom.yml` **and** compose = 0 |
| 2 | The configured default is **never consulted** when placing data | `GetDefaultStoragePath()` has 3 non-test callers: metrics (`main.go:410`), dashboard panel (`server.go:733`), `.fab` import (`handler_export_upload.go:154`). `grep` over `stacks/deploy.go`+`manager.go` → nothing. Field comment `// new apps use this by default` (`settings.go:453`) has never been true |
| 3 | Tier-1 backup follows the data onto the **same disk** | `GetAppDrivePath` → `systemDataPath` (`backup.go:324-334`). Tier 2 refuses that posture outright (`tier2.go:329`) |
| 4 | The Drives count **cannot** include most apps | `countAppsUsingPath` matches `Env["HDD_PATH"]` only (`handlers.go:2118`) — "1 alkalmazás használja" means *"1 of the apps that CAN"* |
**Protection.** Whole-machine tier: **covered** — `df` shows `/mnt/sys_drive`, `/var/lib/docker` and
`/var/lib/felhom` all on `pve-vm-9201-disk-1` = `mp0`, `backup=1` in `9201.conf`. Off-site tier:
**covered by code, NOT observed** — `runVolumeDumps` (`backup.go:607+`) passes its gates for these
apps on paper, but every unit reported `volume_dumps: None`, including `calibre-web` on the data
drive, because no nightly run had happened on a one-hour-old box. **Recorded as unknown, not fine.**
**Ruled:** visibility only tonight. **My own earlier recommendation — refuse deployment until a drive
is registered — is WITHDRAWN**: it assumed the customer had failed to choose; they had no choice.
Specification filed at `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`.
---
## 4. Red-proofs — every mutation asserted applied, then reverted to 0
| Scenario | Mutation | Outcome |
|---|---|---|
| **B** | **both** guards removed (`MUTANT-B1`+`B2`, count asserted **2**) | **A restore WAS seen starting with no drive attached** — no error, full 3.00 s run, wrote into `/tmp/mutant-destination` |
| **C** | `Mismatch = false && …` | FAIL — the silent divergent restore returned |
| **E** | `Known() { return true }` | FAIL — the fabricated empty prefill appeared |
| **A** | 3 template guards dropped (count asserted **3**) | FAIL — the blank form returned (default subdomain, default drive, no notice) |
| **D** | `Mismatch = true` + unknown short-circuit removed | FAIL — **8** ordinary reconstitute tests broke: the guard is reachable in both directions |
| **3a** | `false && restoreOpInFlight(st)` | FAIL — both handlers reported a started restore and overwrote the first op |
| **3b** | `st.LastRecent = false` | FAIL — `just_finished`, `inside_the_window` |
| **P4** | `inventorySizeConcurrency = 1` | FAIL — "peak in flight was 1", elapsed 282 ms = the sequential cost |
**Explicitly:** yes, a restore was seen starting with no drive attached — but only once **both**
guards were removed. The first attempt removed one and the *new* mismatch guard caught it; that
proved defence in depth, not the stated observable, so it was redone.
**Honesty note on D:** the existing reconstitute fixtures write a schema-1 manifest with **no
`Drive`**, so they are scenario-**E** shaped. The matching case is covered in the scenario table, not
by them.
---
## 5. Part 3 — and a correction
**A second press really did start a second run.** Not refused deeper. Both `offboxReconstituteHandler`
and `offboxPlaceHandler` answered „…elindult" and overwrote the first restore's op/stack.
**I was wrong about the refresh, and correct it here.** I first reported that the list page had
neither banner nor poll. Both are present (`backups_restore.html:10` and the script block at 214), the
endpoint exists (`api/router.go:277`), and all three banner pages poll. My grep pattern missed
`restore_banner_js`. **The page refreshes.** The real defect was the *result*: `sawRunning` hid every
terminal state from anyone who was not already watching.
**Status line as it now behaves:** the banner shows a running op, and now also shows a terminal result
that finished within `RestoreResultWindow` (10 min) regardless of whether this page saw it start.
---
## 6. Part 4 — measured, then fixed
Measured on the live off-site target (`u629488-sub3.your-storagebox.de:23`) **before** any change:
```
LEG1_snapshots_json_ms=2605 rc=0
snapshots=18 distinct_app_tags=5
LEG2_stats_calls=4 total_ms=10790 per_call_ms=2697
=> 2605 + 5 * 2697 = ~16.1 s
```
The cause **is** the shape on file, and the fix is contained: the per-app `stats` calls now run
concurrently, **bounded to 4** (`offbox_inventory.go`). The bound is the safety property — the target
is a Storage Box with a session cap, and a refused size call returns 0, which *under-reports the
customer's data* rather than failing visibly. Peak-in-flight is asserted by test, with `-race` clean.
`OffsiteInventoryList` had **no test at all** before this.
**One measurement trap worth recording:** the first attempt failed with `set -e` and no message
because `restic $BASE` was unquoted — `sftp.command=ssh …` contains spaces and word-split.
---
## 7. Hungarian as shipped, bytes confirmed
Verified as hex, valid UTF-8, no BOM, and checked against eleven double-encoding sentinels — `BAD = 0`.
```
korábban itt voltak 6b6f72c3a16262616e2069747420766f6c74616b
erősítsd meg alább 6572c59173c3ad747364206d656720616cc3a16262
már fut, ezért most… 6dc3a172206675742c20657ac3a97274206d6f7374206e656d20696e64c3ad74686174c3b3…
```
---
## 8. Gates, tests, commits
- `python3 controller/scripts/controller_gates.py` → **11/11 OK**, `GATES_RC=0` (hook also ran on push; **no `--no-verify`**)
- `go test ./...` → **28 packages ok**, `SUITE_RC=0`; `go vet` clean; `-race` clean on the changed package
- Commit 1: `985388c` — Part 3 + Part 2's engine (14 files), pushed `2fa1efc..985388c`
---
## 9. What was dropped, named plainly
- **Deploy-then-restore as one atomic act** — deliberately out of scope, stated in the task and here.
- **Placement change and migration for the 40-class** — specification filed, ruling is the operator's.
- **R-353** — a restore that returns configuration and no data still reports a bare completion.
**Next session's first item.**
- The three items the task listed as out of scope (the re-issue skip on reinstall, the two cosmetic
hub bugs, the CI runs failing with no log) were not touched.
## 10. Observations — noticed, not acted on
- `offbox_handlers.go:498` — the **verify-copy delete** guard is blind in exactly the same way as the
restore guards were. Deleting a verification copy while an off-box restore writes into it is a real
hazard. Left alone because it does not *start* work, and widening the diff unasked is its own risk.
- **The OpenGist instance was removed by someone between 16:57 and 17:02 UTC**, not by this session
(`ScanStacks: found stack "opengist" deployed=false` from 17:02:58). The task asked for it to be
left in place as evidence. **Its unit and manifest survive** on `/mnt/sys_drive`, and `privatebin`
is now a live specimen of the same class.
- A second reconstitution ran at 16:31:48 for `calibre-web` (8 files placed, 0 DB dumps replayed) that
was not mentioned in the task.