docs: R-356 CONTEXT + REPORT (v0.219.0, proven live on demo-hp)
gates / gates (push) Successful in 11s

This commit is contained in:
2026-08-22 13:25:36 +02:00
parent 08eb1a6e3a
commit cbc3fa589c
+88 -184
View File
@@ -1,201 +1,105 @@
# REPORT — controller v0.216.0 → v0.217.0: the restore knows where the data lived # REPORT — R-356: the off-site restore refused every app that has no data drive
**Date:** 2026-08-21 · **Task class:** Implementation (Establish-first) · **Repos touched:** **Controller v0.219.0**, shipped and proven live on `demo-hp` 2026-08-22.
`felhom-controller` (code), `felhom.eu` (register + specification, documentation only)
Register: **R-351 CLOSED**, **R-352 partly closed** (visibility shipped, placement open), ## Baselines used (re-verified at session start)
**R-353 OPEN — next session's first item**. Ceiling moved **R-350 → R-353**.
--- | repo | `main` @ | version | → |
|---|---|---|---|
| felhom-controller | `2da259af38cc` | v0.218.0 | **v0.219.0** |
| felhom-agent | `40d857b52711` | v0.130.0 | unchanged |
| felhom.eu (hub) | `f2edf7e54595` | v0.106.0 | unchanged |
| app-catalog | `459766cb1639` | — | unchanged |
## 0. Baselines, re-established live (nothing in the prompt was trusted) **MinAgent stays 0.129.0.** All four hashes matched the prompt.
| | Value | How | **The catalogue count, measured in this session and not taken from the prompt:** 53 templates,
**13 declare `needs_hdd: true`**, **40 declare `false`**, and exactly 13 carry a `backup:` block.
Agrees with §1.
## The defect
`ReconstituteFromOffsite` and `PlaceOffsiteRestore` both resolved the restore destination with the
**raw** `stackProvider.GetStackHDDPath(stack)` and refused when it was empty, with a message saying
the app is not installed. For the 40 driveless apps that value is never set, so both refused
permanently for a running, healthy app — and the remedy they offered ("reinstall it in the same
place") cannot be carried out, because those apps offer no place to pick.
The capture side never had this: `CaptureRecoveryUnit` resolves via `GetAppDrivePath`, which falls
back to `systemDataPath`. That is why the 2026-08-21 refusal could print `/mnt/sys_drive` — a
destination the backup had recorded and the restore refused to use.
## The change
- `internal/backup/backup.go` — new `Manager.isStackDeployed`, built on `knownStackNames()` →
`ListDeployedStacks()`. **Nil provider ⇒ false**, documented with its reason: the caller's next act
is a write, so "cannot tell" must fail into the recoverable refusal, not into a copy.
- `internal/backup/offbox_reconstitute.go` — the refusal now asks *installed?* first (R-351 sentence
and the recorded-drive half intact), then resolves the destination with `GetAppDrivePath`, then has
a **third, distinct** refusal for installed-but-no-resolvable-data-root. R-253 and R-351 comments
kept and extended with what the refusal stopped covering and why.
- `internal/backup/offbox_restore.go` — the same four steps in `PlaceOffsiteRestore`. The headroom
gate below it then measures the resolved namespace, which is the correct disk in both cases.
**Not changed, deliberately:** `offboxCaptureSet`'s raw `GetStackHDDPath` (`offbox_capture.go:43`) and
`offboxRestoreScratchDir`.
**`ListDeployedStacks()` semantics were verified, not assumed.** Source: `stackAdapter` skips
`!s.Deployed` (`cmd/controller/main.go:2138`). Test: positive and negative controls in
`TestR356_IsStackDeployedIsExactAndFailsClosed`. Seam: an AST walk in
`cmd/controller/r356_deployed_seam_test.go` pins both the `SetStackProvider` wiring in `main()` and
the `!s.Deployed` guard itself.
## Tests
New: `internal/backup/r356_hot_only_restore_test.go` (10 tests, Scenarios A–E across both entry
points) and `cmd/controller/r356_deployed_seam_test.go` (2 AST tests). **Test count in
`internal/backup` 315 → 325, `cmd/controller` 42 → 44.**
**Every red-proof was SEEN failing.** Mutation → observed failure:
| # | mutation applied | observed |
|---|---|---| |---|---|---|
| controller | **0.216.0** | `docker inspect felhom-controller` on guest 9201 | | A | destination back to `stackProvider.GetStackHDDPath` | `ScenarioA` failed: *"a(z) immich telepítve van, de a vezérlő nem tudja megállapítani, hová tartoznak az adatai…"* |
| agent | **0.130.0** | `felhom-agent --version` on the host | | B | `GetAppDrivePath` always returns `systemDataPath` | 3 tests failed, incl. *"placement …/sys/felhom-data/… left the app's own drive"* — this is also the **positive control** for the "13-class unchanged" absence claim |
| golden | **0.216.0** | hub artifact manifest | | C | `isStackDeployed` always true | 3 tests failed; the not-installed refusal was replaced by a placement-mismatch prompt |
| register ceiling | **R-350** (prompt said R-327 — stale) | `grep -rhoE '\bR-[0-9]{1,4}\b' --include=*.md` | | D | widen `nincs telepítve` to cover the no-data-root case | both ScenarioD tests failed: *"a running app must not be told it is not installed"* |
| `demo-hp` | reinstalled 2026-08-21, PVE 9.2.2 | break-glass root from hub `host_recovery` | | E | revert `PlaceOffsiteRestore` only (fix one entry point) | `ScenarioE_PlaceDriveless` failed |
--- **No red-proof passed.** Green gate: `go build ./... && go vet ./... && go test ./...` — all green.
## 1. Part 0 — what the backup can tell us **Fixture correction, stated rather than hidden:** several existing fixtures marked an app "installed"
by giving it an HDD path — the exact conflation this change removes. 17 tests failed on that alone
after the change; they now state deployment as its own fact (`prov.deployed`). No assertion was
weakened, and `TestPlace_UndeployedRefused` was passing for the wrong reason before.
**1. Are the drive and folder readable before restoring? YES.** `RecoveryManifest.Drive` / ## Live walk on `demo-hp` (controller v0.219.0, healthy)
`.NamespaceRoot` (`internal/backup/recovery_unit.go:48-49`). Confirmed on real data:
``` Endpoint-level — `claude-in-chrome` is not available on DooPlex. Full record and every artefact:
/mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json `felhom.eu/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/`.
drive='/mnt/sys_drive' namespace_root='/mnt/sys_drive/felhom-data'
/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/manifest.json
drive='/mnt/felhom-drives/hdd_1' namespace_root='/mnt/felhom-drives/hdd_1'
```
**Cost, two prices:** drive attached → a plain local file read, **free**. Off-site only → one **Driveless (`privatebin`):** planted through the app's own JSON API + a direct file plant with a
`snapshots latest --tag` plus one unit-only `restore --include` **per app**, not one for all; the Hungarian accented name; findability proved by plant→find→remove→fail-to-find; off-sited; **11 files
unit-only restore already exists and is already the default (`offbox_restore.go:264`), needing a deleted outright**; scratch prepared (snapshot `306accff`, **unit-only** — the shape never driven to
registered non-network drive for scratch and 2 GB free (`offbox_restore.go:242`). completion before); `POST /backup/offbox/reconstitute` — **no refusal**; **15/15 files back byte for
byte**. Message (160 bytes): *„A(z) privatebin: 0 fájl és 1 adatkötet visszaállítva (mentés:
2026-08-22 13:14) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."* It **names
the volume**, so it is not the R-354 shape.
**2. Is the web address recoverable? YES, recorded — not inferred.** `app.yaml` is captured into **Drive app (`calibre-web`):** same walk, **3/3 byte-identical**, placed on `/mnt/felhom-drives/hdd_1`;
every unit. Real values: `opengist SUBDOMAIN: gist / DOMAIN: enkisfelhom.hu`, `calibre-web books`. `/mnt/sys_drive/felhom-data/userdata/media` does not exist, so nothing was misplaced onto the SSD.
**Caveat carried into the design:** an absent `SUBDOMAIN` makes the live path fall back to the catalog
default (`stacks/deploy.go:88-90`) — a catalog guess, not the customer's answer. Treated as UNKNOWN.
**3. What does the restore compare today? NOTHING. A mismatched restore succeeded silently.** **Every downstream leg held on first run.** No finding raised against them.
`grep -rnE '\.Drive\b|\.NamespaceRoot\b' --include=*.go | grep -v _test` returned only the
`appbackup.NamespaceRoot` *function* and `Tier2Target`'s unrelated field.
**4. Did the 21 August OpenGist restore complete? Yes — on the second try, by a different route.** ## Also in this session
``` CI run **387** (job 386) failed on `instructions_gate` — **a real gate bug, not mine**: `ef6ac6f` in
16:32:43 off-box full-restore prepared for opengist (182.3 KB) `felhom.eu` compressed closed rows into `CLOSED-ITEMS.md`, which the gate does not read, so every
16:33:34 restored opengist (9e38b84c, full=true) -> .../backups/offsite-restore/opengist citation of a compressed item became "a reference to nothing". Fixed in
16:37:14 [ERROR] off-box reconstitute opengist: ...nincs telepitve... `felhom.eu/scripts/instructions_gate.py` (commit `e18668f`); both controls still convict.
16:39:16 [WARN] Restore requested (async): stack=opengist, snapshot=helyi
16:39:25 Restore-from-unit completed: opengist in 8.666896042s
```
**No screen said so** — the answer existed only in `docker logs`. That is R-351c, and the *contents* ## Operator follow-up
of what came back is R-353.
--- 1. Bake and **vouch** a golden carrying controller 0.219.0. The `golden-currency` gate is red until
then, correctly: a machine installed now still receives 0.218.0.
## 2. §2's ruling revisited 2. **Then** raise the update floor to 0.219.0 — last, in a separate save.
The R-253 comment (`templates/backups_restore.html:104-110`) removed the promise to reinstall,
because reconstitution writes to the app's own data path, *"which exists only once the customer has
chosen a drive during deploy — the restore has no answer to that question and must not invent one."*
**The reasoning was correct. Its premise no longer holds.** The restore does have an answer, and it is
the customer's own previous answer: `manifest.Drive` + the captured `app.yaml`. **Reversed, openly,
recorded in R-351.** What is *not* reversed: the restore still does not deploy the app for you —
deploy-then-restore as one atomic act stays out of scope (a half-failure leaves a half-installed app,
and the catalogue may have moved on since the backup, a hazard that must stay visible).
---
## 3. Part 1 — established, then ruled
**The premise changed twice.** It is not a typed path, and not a customer failing to choose.
| # | Measured untruth | Evidence |
|---|---|---|
| 1 | **40 of 53** templates declare no data path | 53 total, 13 with `env_var: HDD_PATH`; negative control: `grep -c HDD_PATH` on opengist's `.felhom.yml` **and** compose = 0 |
| 2 | The configured default is **never consulted** when placing data | `GetDefaultStoragePath()` has 3 non-test callers: metrics (`main.go:410`), dashboard panel (`server.go:733`), `.fab` import (`handler_export_upload.go:154`). `grep` over `stacks/deploy.go`+`manager.go` → nothing. Field comment `// new apps use this by default` (`settings.go:453`) has never been true |
| 3 | Tier-1 backup follows the data onto the **same disk** | `GetAppDrivePath` → `systemDataPath` (`backup.go:324-334`). Tier 2 refuses that posture outright (`tier2.go:329`) |
| 4 | The Drives count **cannot** include most apps | `countAppsUsingPath` matches `Env["HDD_PATH"]` only (`handlers.go:2118`) — "1 alkalmazás használja" means *"1 of the apps that CAN"* |
**Protection.** Whole-machine tier: **covered** — `df` shows `/mnt/sys_drive`, `/var/lib/docker` and
`/var/lib/felhom` all on `pve-vm-9201-disk-1` = `mp0`, `backup=1` in `9201.conf`. Off-site tier:
**covered by code, NOT observed** — `runVolumeDumps` (`backup.go:607+`) passes its gates for these
apps on paper, but every unit reported `volume_dumps: None`, including `calibre-web` on the data
drive, because no nightly run had happened on a one-hour-old box. **Recorded as unknown, not fine.**
**Ruled:** visibility only tonight. **My own earlier recommendation — refuse deployment until a drive
is registered — is WITHDRAWN**: it assumed the customer had failed to choose; they had no choice.
Specification filed at `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`.
---
## 4. Red-proofs — every mutation asserted applied, then reverted to 0
| Scenario | Mutation | Outcome |
|---|---|---|
| **B** | **both** guards removed (`MUTANT-B1`+`B2`, count asserted **2**) | **A restore WAS seen starting with no drive attached** — no error, full 3.00 s run, wrote into `/tmp/mutant-destination` |
| **C** | `Mismatch = false && …` | FAIL — the silent divergent restore returned |
| **E** | `Known() { return true }` | FAIL — the fabricated empty prefill appeared |
| **A** | 3 template guards dropped (count asserted **3**) | FAIL — the blank form returned (default subdomain, default drive, no notice) |
| **D** | `Mismatch = true` + unknown short-circuit removed | FAIL — **8** ordinary reconstitute tests broke: the guard is reachable in both directions |
| **3a** | `false && restoreOpInFlight(st)` | FAIL — both handlers reported a started restore and overwrote the first op |
| **3b** | `st.LastRecent = false` | FAIL — `just_finished`, `inside_the_window` |
| **P4** | `inventorySizeConcurrency = 1` | FAIL — "peak in flight was 1", elapsed 282 ms = the sequential cost |
**Explicitly:** yes, a restore was seen starting with no drive attached — but only once **both**
guards were removed. The first attempt removed one and the *new* mismatch guard caught it; that
proved defence in depth, not the stated observable, so it was redone.
**Honesty note on D:** the existing reconstitute fixtures write a schema-1 manifest with **no
`Drive`**, so they are scenario-**E** shaped. The matching case is covered in the scenario table, not
by them.
---
## 5. Part 3 — and a correction
**A second press really did start a second run.** Not refused deeper. Both `offboxReconstituteHandler`
and `offboxPlaceHandler` answered „…elindult" and overwrote the first restore's op/stack.
**I was wrong about the refresh, and correct it here.** I first reported that the list page had
neither banner nor poll. Both are present (`backups_restore.html:10` and the script block at 214), the
endpoint exists (`api/router.go:277`), and all three banner pages poll. My grep pattern missed
`restore_banner_js`. **The page refreshes.** The real defect was the *result*: `sawRunning` hid every
terminal state from anyone who was not already watching.
**Status line as it now behaves:** the banner shows a running op, and now also shows a terminal result
that finished within `RestoreResultWindow` (10 min) regardless of whether this page saw it start.
---
## 6. Part 4 — measured, then fixed
Measured on the live off-site target (`u629488-sub3.your-storagebox.de:23`) **before** any change:
```
LEG1_snapshots_json_ms=2605 rc=0
snapshots=18 distinct_app_tags=5
LEG2_stats_calls=4 total_ms=10790 per_call_ms=2697
=> 2605 + 5 * 2697 = ~16.1 s
```
The cause **is** the shape on file, and the fix is contained: the per-app `stats` calls now run
concurrently, **bounded to 4** (`offbox_inventory.go`). The bound is the safety property — the target
is a Storage Box with a session cap, and a refused size call returns 0, which *under-reports the
customer's data* rather than failing visibly. Peak-in-flight is asserted by test, with `-race` clean.
`OffsiteInventoryList` had **no test at all** before this.
**One measurement trap worth recording:** the first attempt failed with `set -e` and no message
because `restic $BASE` was unquoted — `sftp.command=ssh …` contains spaces and word-split.
---
## 7. Hungarian as shipped, bytes confirmed
Verified as hex, valid UTF-8, no BOM, and checked against eleven double-encoding sentinels — `BAD = 0`.
```
korábban itt voltak 6b6f72c3a16262616e2069747420766f6c74616b
erősítsd meg alább 6572c59173c3ad747364206d656720616cc3a16262
már fut, ezért most… 6dc3a172206675742c20657ac3a97274206d6f7374206e656d20696e64c3ad74686174c3b3…
```
---
## 8. Gates, tests, commits
- `python3 controller/scripts/controller_gates.py` → **11/11 OK**, `GATES_RC=0` (hook also ran on push; **no `--no-verify`**)
- `go test ./...` → **28 packages ok**, `SUITE_RC=0`; `go vet` clean; `-race` clean on the changed package
- Commit 1: `985388c` — Part 3 + Part 2's engine (14 files), pushed `2fa1efc..985388c`
---
## 9. What was dropped, named plainly
- **Deploy-then-restore as one atomic act** — deliberately out of scope, stated in the task and here.
- **Placement change and migration for the 40-class** — specification filed, ruling is the operator's.
- **R-353** — a restore that returns configuration and no data still reports a bare completion.
**Next session's first item.**
- The three items the task listed as out of scope (the re-issue skip on reinstall, the two cosmetic
hub bugs, the CI runs failing with no log) were not touched.
## 10. Observations — noticed, not acted on
- `offbox_handlers.go:498` — the **verify-copy delete** guard is blind in exactly the same way as the
restore guards were. Deleting a verification copy while an off-box restore writes into it is a real
hazard. Left alone because it does not *start* work, and widening the diff unasked is its own risk.
- **The OpenGist instance was removed by someone between 16:57 and 17:02 UTC**, not by this session
(`ScanStacks: found stack "opengist" deployed=false` from 17:02:58). The task asked for it to be
left in place as evidence. **Its unit and manifest survive** on `/mnt/sys_drive`, and `privatebin`
is now a live specimen of the same class.
- A second reconstitution ran at 16:31:48 for `calibre-web` (8 files placed, 0 DB dumps replayed) that
was not mentioned in the task.