diff --git a/REPORT.md b/REPORT.md index e802d40..b1b7355 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,201 +1,105 @@ -# REPORT — controller v0.216.0 → v0.217.0: the restore knows where the data lived +# REPORT — R-356: the off-site restore refused every app that has no data drive -**Date:** 2026-08-21 · **Task class:** Implementation (Establish-first) · **Repos touched:** -`felhom-controller` (code), `felhom.eu` (register + specification, documentation only) +**Controller v0.219.0**, shipped and proven live on `demo-hp` 2026-08-22. -Register: **R-351 CLOSED**, **R-352 partly closed** (visibility shipped, placement open), -**R-353 OPEN — next session's first item**. Ceiling moved **R-350 → R-353**. +## Baselines used (re-verified at session start) ---- +| repo | `main` @ | version | → | +|---|---|---|---| +| felhom-controller | `2da259af38cc` | v0.218.0 | **v0.219.0** | +| felhom-agent | `40d857b52711` | v0.130.0 | unchanged | +| felhom.eu (hub) | `f2edf7e54595` | v0.106.0 | unchanged | +| app-catalog | `459766cb1639` | — | unchanged | -## 0. Baselines, re-established live (nothing in the prompt was trusted) +**MinAgent stays 0.129.0.** All four hashes matched the prompt. -| | Value | How | +**The catalogue count, measured in this session and not taken from the prompt:** 53 templates, +**13 declare `needs_hdd: true`**, **40 declare `false`**, and exactly 13 carry a `backup:` block. +Agrees with §1. + +## The defect + +`ReconstituteFromOffsite` and `PlaceOffsiteRestore` both resolved the restore destination with the +**raw** `stackProvider.GetStackHDDPath(stack)` and refused when it was empty, with a message saying +the app is not installed. For the 40 driveless apps that value is never set, so both refused +permanently for a running, healthy app — and the remedy they offered ("reinstall it in the same +place") cannot be carried out, because those apps offer no place to pick. + +The capture side never had this: `CaptureRecoveryUnit` resolves via `GetAppDrivePath`, which falls +back to `systemDataPath`. That is why the 2026-08-21 refusal could print `/mnt/sys_drive` — a +destination the backup had recorded and the restore refused to use. + +## The change + +- `internal/backup/backup.go` — new `Manager.isStackDeployed`, built on `knownStackNames()` → + `ListDeployedStacks()`. **Nil provider ⇒ false**, documented with its reason: the caller's next act + is a write, so "cannot tell" must fail into the recoverable refusal, not into a copy. +- `internal/backup/offbox_reconstitute.go` — the refusal now asks *installed?* first (R-351 sentence + and the recorded-drive half intact), then resolves the destination with `GetAppDrivePath`, then has + a **third, distinct** refusal for installed-but-no-resolvable-data-root. R-253 and R-351 comments + kept and extended with what the refusal stopped covering and why. +- `internal/backup/offbox_restore.go` — the same four steps in `PlaceOffsiteRestore`. The headroom + gate below it then measures the resolved namespace, which is the correct disk in both cases. + +**Not changed, deliberately:** `offboxCaptureSet`'s raw `GetStackHDDPath` (`offbox_capture.go:43`) and +`offboxRestoreScratchDir`. + +**`ListDeployedStacks()` semantics were verified, not assumed.** Source: `stackAdapter` skips +`!s.Deployed` (`cmd/controller/main.go:2138`). Test: positive and negative controls in +`TestR356_IsStackDeployedIsExactAndFailsClosed`. Seam: an AST walk in +`cmd/controller/r356_deployed_seam_test.go` pins both the `SetStackProvider` wiring in `main()` and +the `!s.Deployed` guard itself. + +## Tests + +New: `internal/backup/r356_hot_only_restore_test.go` (10 tests, Scenarios A–E across both entry +points) and `cmd/controller/r356_deployed_seam_test.go` (2 AST tests). **Test count in +`internal/backup` 315 → 325, `cmd/controller` 42 → 44.** + +**Every red-proof was SEEN failing.** Mutation → observed failure: + +| # | mutation applied | observed | |---|---|---| -| controller | **0.216.0** | `docker inspect felhom-controller` on guest 9201 | -| agent | **0.130.0** | `felhom-agent --version` on the host | -| golden | **0.216.0** | hub artifact manifest | -| register ceiling | **R-350** (prompt said R-327 — stale) | `grep -rhoE '\bR-[0-9]{1,4}\b' --include=*.md` | -| `demo-hp` | reinstalled 2026-08-21, PVE 9.2.2 | break-glass root from hub `host_recovery` | +| A | destination back to `stackProvider.GetStackHDDPath` | `ScenarioA` failed: *"a(z) immich telepítve van, de a vezérlő nem tudja megállapítani, hová tartoznak az adatai…"* | +| B | `GetAppDrivePath` always returns `systemDataPath` | 3 tests failed, incl. *"placement …/sys/felhom-data/… left the app's own drive"* — this is also the **positive control** for the "13-class unchanged" absence claim | +| C | `isStackDeployed` always true | 3 tests failed; the not-installed refusal was replaced by a placement-mismatch prompt | +| D | widen `nincs telepítve` to cover the no-data-root case | both ScenarioD tests failed: *"a running app must not be told it is not installed"* | +| E | revert `PlaceOffsiteRestore` only (fix one entry point) | `ScenarioE_PlaceDriveless` failed | ---- +**No red-proof passed.** Green gate: `go build ./... && go vet ./... && go test ./...` — all green. -## 1. Part 0 — what the backup can tell us +**Fixture correction, stated rather than hidden:** several existing fixtures marked an app "installed" +by giving it an HDD path — the exact conflation this change removes. 17 tests failed on that alone +after the change; they now state deployment as its own fact (`prov.deployed`). No assertion was +weakened, and `TestPlace_UndeployedRefused` was passing for the wrong reason before. -**1. Are the drive and folder readable before restoring? YES.** `RecoveryManifest.Drive` / -`.NamespaceRoot` (`internal/backup/recovery_unit.go:48-49`). Confirmed on real data: +## Live walk on `demo-hp` (controller v0.219.0, healthy) -``` -/mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json - drive='/mnt/sys_drive' namespace_root='/mnt/sys_drive/felhom-data' -/mnt/felhom-drives/hdd_1/backups/primary/calibre-web/manifest.json - drive='/mnt/felhom-drives/hdd_1' namespace_root='/mnt/felhom-drives/hdd_1' -``` +Endpoint-level — `claude-in-chrome` is not available on DooPlex. Full record and every artefact: +`felhom.eu/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/`. -**Cost, two prices:** drive attached → a plain local file read, **free**. Off-site only → one -`snapshots latest --tag` plus one unit-only `restore --include` **per app**, not one for all; the -unit-only restore already exists and is already the default (`offbox_restore.go:264`), needing a -registered non-network drive for scratch and 2 GB free (`offbox_restore.go:242`). +**Driveless (`privatebin`):** planted through the app's own JSON API + a direct file plant with a +Hungarian accented name; findability proved by plant→find→remove→fail-to-find; off-sited; **11 files +deleted outright**; scratch prepared (snapshot `306accff`, **unit-only** — the shape never driven to +completion before); `POST /backup/offbox/reconstitute` — **no refusal**; **15/15 files back byte for +byte**. Message (160 bytes): *„A(z) privatebin: 0 fájl és 1 adatkötet visszaállítva (mentés: +2026-08-22 13:14) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."* It **names +the volume**, so it is not the R-354 shape. -**2. Is the web address recoverable? YES, recorded — not inferred.** `app.yaml` is captured into -every unit. Real values: `opengist SUBDOMAIN: gist / DOMAIN: enkisfelhom.hu`, `calibre-web books`. -**Caveat carried into the design:** an absent `SUBDOMAIN` makes the live path fall back to the catalog -default (`stacks/deploy.go:88-90`) — a catalog guess, not the customer's answer. Treated as UNKNOWN. +**Drive app (`calibre-web`):** same walk, **3/3 byte-identical**, placed on `/mnt/felhom-drives/hdd_1`; +`/mnt/sys_drive/felhom-data/userdata/media` does not exist, so nothing was misplaced onto the SSD. -**3. What does the restore compare today? NOTHING. A mismatched restore succeeded silently.** -`grep -rnE '\.Drive\b|\.NamespaceRoot\b' --include=*.go | grep -v _test` returned only the -`appbackup.NamespaceRoot` *function* and `Tier2Target`'s unrelated field. +**Every downstream leg held on first run.** No finding raised against them. -**4. Did the 21 August OpenGist restore complete? Yes — on the second try, by a different route.** +## Also in this session -``` -16:32:43 off-box full-restore prepared for opengist (182.3 KB) -16:33:34 restored opengist (9e38b84c, full=true) -> .../backups/offsite-restore/opengist -16:37:14 [ERROR] off-box reconstitute opengist: ...nincs telepitve... -16:39:16 [WARN] Restore requested (async): stack=opengist, snapshot=helyi -16:39:25 Restore-from-unit completed: opengist in 8.666896042s -``` +CI run **387** (job 386) failed on `instructions_gate` — **a real gate bug, not mine**: `ef6ac6f` in +`felhom.eu` compressed closed rows into `CLOSED-ITEMS.md`, which the gate does not read, so every +citation of a compressed item became "a reference to nothing". Fixed in +`felhom.eu/scripts/instructions_gate.py` (commit `e18668f`); both controls still convict. -**No screen said so** — the answer existed only in `docker logs`. That is R-351c, and the *contents* -of what came back is R-353. +## Operator follow-up ---- - -## 2. §2's ruling revisited - -The R-253 comment (`templates/backups_restore.html:104-110`) removed the promise to reinstall, -because reconstitution writes to the app's own data path, *"which exists only once the customer has -chosen a drive during deploy — the restore has no answer to that question and must not invent one."* - -**The reasoning was correct. Its premise no longer holds.** The restore does have an answer, and it is -the customer's own previous answer: `manifest.Drive` + the captured `app.yaml`. **Reversed, openly, -recorded in R-351.** What is *not* reversed: the restore still does not deploy the app for you — -deploy-then-restore as one atomic act stays out of scope (a half-failure leaves a half-installed app, -and the catalogue may have moved on since the backup, a hazard that must stay visible). - ---- - -## 3. Part 1 — established, then ruled - -**The premise changed twice.** It is not a typed path, and not a customer failing to choose. - -| # | Measured untruth | Evidence | -|---|---|---| -| 1 | **40 of 53** templates declare no data path | 53 total, 13 with `env_var: HDD_PATH`; negative control: `grep -c HDD_PATH` on opengist's `.felhom.yml` **and** compose = 0 | -| 2 | The configured default is **never consulted** when placing data | `GetDefaultStoragePath()` has 3 non-test callers: metrics (`main.go:410`), dashboard panel (`server.go:733`), `.fab` import (`handler_export_upload.go:154`). `grep` over `stacks/deploy.go`+`manager.go` → nothing. Field comment `// new apps use this by default` (`settings.go:453`) has never been true | -| 3 | Tier-1 backup follows the data onto the **same disk** | `GetAppDrivePath` → `systemDataPath` (`backup.go:324-334`). Tier 2 refuses that posture outright (`tier2.go:329`) | -| 4 | The Drives count **cannot** include most apps | `countAppsUsingPath` matches `Env["HDD_PATH"]` only (`handlers.go:2118`) — "1 alkalmazás használja" means *"1 of the apps that CAN"* | - -**Protection.** Whole-machine tier: **covered** — `df` shows `/mnt/sys_drive`, `/var/lib/docker` and -`/var/lib/felhom` all on `pve-vm-9201-disk-1` = `mp0`, `backup=1` in `9201.conf`. Off-site tier: -**covered by code, NOT observed** — `runVolumeDumps` (`backup.go:607+`) passes its gates for these -apps on paper, but every unit reported `volume_dumps: None`, including `calibre-web` on the data -drive, because no nightly run had happened on a one-hour-old box. **Recorded as unknown, not fine.** - -**Ruled:** visibility only tonight. **My own earlier recommendation — refuse deployment until a drive -is registered — is WITHDRAWN**: it assumed the customer had failed to choose; they had no choice. -Specification filed at `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`. - ---- - -## 4. Red-proofs — every mutation asserted applied, then reverted to 0 - -| Scenario | Mutation | Outcome | -|---|---|---| -| **B** | **both** guards removed (`MUTANT-B1`+`B2`, count asserted **2**) | **A restore WAS seen starting with no drive attached** — no error, full 3.00 s run, wrote into `/tmp/mutant-destination` | -| **C** | `Mismatch = false && …` | FAIL — the silent divergent restore returned | -| **E** | `Known() { return true }` | FAIL — the fabricated empty prefill appeared | -| **A** | 3 template guards dropped (count asserted **3**) | FAIL — the blank form returned (default subdomain, default drive, no notice) | -| **D** | `Mismatch = true` + unknown short-circuit removed | FAIL — **8** ordinary reconstitute tests broke: the guard is reachable in both directions | -| **3a** | `false && restoreOpInFlight(st)` | FAIL — both handlers reported a started restore and overwrote the first op | -| **3b** | `st.LastRecent = false` | FAIL — `just_finished`, `inside_the_window` | -| **P4** | `inventorySizeConcurrency = 1` | FAIL — "peak in flight was 1", elapsed 282 ms = the sequential cost | - -**Explicitly:** yes, a restore was seen starting with no drive attached — but only once **both** -guards were removed. The first attempt removed one and the *new* mismatch guard caught it; that -proved defence in depth, not the stated observable, so it was redone. - -**Honesty note on D:** the existing reconstitute fixtures write a schema-1 manifest with **no -`Drive`**, so they are scenario-**E** shaped. The matching case is covered in the scenario table, not -by them. - ---- - -## 5. Part 3 — and a correction - -**A second press really did start a second run.** Not refused deeper. Both `offboxReconstituteHandler` -and `offboxPlaceHandler` answered „…elindult" and overwrote the first restore's op/stack. - -**I was wrong about the refresh, and correct it here.** I first reported that the list page had -neither banner nor poll. Both are present (`backups_restore.html:10` and the script block at 214), the -endpoint exists (`api/router.go:277`), and all three banner pages poll. My grep pattern missed -`restore_banner_js`. **The page refreshes.** The real defect was the *result*: `sawRunning` hid every -terminal state from anyone who was not already watching. - -**Status line as it now behaves:** the banner shows a running op, and now also shows a terminal result -that finished within `RestoreResultWindow` (10 min) regardless of whether this page saw it start. - ---- - -## 6. Part 4 — measured, then fixed - -Measured on the live off-site target (`u629488-sub3.your-storagebox.de:23`) **before** any change: - -``` -LEG1_snapshots_json_ms=2605 rc=0 -snapshots=18 distinct_app_tags=5 -LEG2_stats_calls=4 total_ms=10790 per_call_ms=2697 -=> 2605 + 5 * 2697 = ~16.1 s -``` - -The cause **is** the shape on file, and the fix is contained: the per-app `stats` calls now run -concurrently, **bounded to 4** (`offbox_inventory.go`). The bound is the safety property — the target -is a Storage Box with a session cap, and a refused size call returns 0, which *under-reports the -customer's data* rather than failing visibly. Peak-in-flight is asserted by test, with `-race` clean. -`OffsiteInventoryList` had **no test at all** before this. - -**One measurement trap worth recording:** the first attempt failed with `set -e` and no message -because `restic $BASE` was unquoted — `sftp.command=ssh …` contains spaces and word-split. - ---- - -## 7. Hungarian as shipped, bytes confirmed - -Verified as hex, valid UTF-8, no BOM, and checked against eleven double-encoding sentinels — `BAD = 0`. - -``` -korábban itt voltak 6b6f72c3a16262616e2069747420766f6c74616b -erősítsd meg alább 6572c59173c3ad747364206d656720616cc3a16262 -már fut, ezért most… 6dc3a172206675742c20657ac3a97274206d6f7374206e656d20696e64c3ad74686174c3b3… -``` - ---- - -## 8. Gates, tests, commits - -- `python3 controller/scripts/controller_gates.py` → **11/11 OK**, `GATES_RC=0` (hook also ran on push; **no `--no-verify`**) -- `go test ./...` → **28 packages ok**, `SUITE_RC=0`; `go vet` clean; `-race` clean on the changed package -- Commit 1: `985388c` — Part 3 + Part 2's engine (14 files), pushed `2fa1efc..985388c` - ---- - -## 9. What was dropped, named plainly - -- **Deploy-then-restore as one atomic act** — deliberately out of scope, stated in the task and here. -- **Placement change and migration for the 40-class** — specification filed, ruling is the operator's. -- **R-353** — a restore that returns configuration and no data still reports a bare completion. - **Next session's first item.** -- The three items the task listed as out of scope (the re-issue skip on reinstall, the two cosmetic - hub bugs, the CI runs failing with no log) were not touched. - -## 10. Observations — noticed, not acted on - -- `offbox_handlers.go:498` — the **verify-copy delete** guard is blind in exactly the same way as the - restore guards were. Deleting a verification copy while an off-box restore writes into it is a real - hazard. Left alone because it does not *start* work, and widening the diff unasked is its own risk. -- **The OpenGist instance was removed by someone between 16:57 and 17:02 UTC**, not by this session - (`ScanStacks: found stack "opengist" deployed=false` from 17:02:58). The task asked for it to be - left in place as evidence. **Its unit and manifest survive** on `/mnt/sys_drive`, and `privatebin` - is now a live specimen of the same class. -- A second reconstitution ran at 16:31:48 for `calibre-web` (8 files placed, 0 DB dumps replayed) that - was not mentioned in the task. +1. Bake and **vouch** a golden carrying controller 0.219.0. The `golden-currency` gate is red until + then, correctly: a machine installed now still receives 0.218.0. +2. **Then** raise the update floor to 0.219.0 — last, in a separate save.