# REPORT — R-353 / R-357 / R-358 / R-360: the restore tells the truth **Controller v0.226.0 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`** --- ## 1. Confirmed baselines used — BOTH HAD MOVED | repo | spec baseline | **actual at start** | version | |---|---|---|---| | `felhom-controller` | `f8c9390` / v0.223.0 → v0.224.0 | **`e5eee50` / v0.225.0** | → **v0.226.0** | | `felhom.eu` | `c2c1fb4` | **`ac6ac03`** | docs only | | `felhom-agent` | not touched | v0.130.0 | unchanged | **Both targets were consumed earlier the same day by my own work** — v0.224.0 (R-330) and v0.225.0 (R-331). The drift was re-confirmed against live Gitea before the first edit and the operator authorised proceeding. **Every symbol in the spec's §5 table was re-verified present at the real baseline** before editing; all 20 resolved, and almost every line landmark still matched. `MinAgent` stays **0.129.0**. Both trees were clean and equal to `origin/main` at the start. ## 2. Files created / modified **`felhom-controller`** (all paths under repo root): | file | change | |---|---| | `controller/internal/backup/restore.go` | `restoreDockerVolumes` → `(int, error)`; caller updated | | `controller/internal/backup/restore_unit.go` | `UnitRestoreResult`; `RestoreFromRecoveryUnit` → `(UnitRestoreResult, error)` on every return path | | `controller/internal/web/handlers.go` | `unitRestoreOutcomeMsg` + 3 message constants; handler publishes the outcome | | `controller/internal/backup/offbox_reconstitute.go` | the R-357 headroom gate | | `controller/internal/backup/offbox_restore.go` | scratch marker (write/clear/read), `OffboxFullScratchReady` rewritten, shared refusal constants, `SetOffboxLatestSnapshotFn`, `WriteScratchMarkerForTest` | | `controller/internal/backup/backup.go` | `offboxLatestSnapFn` field | | `controller/internal/web/offbox_handlers.go` | server-side scratch refusals in 2 handlers; R-360 delete guard; corrected doc comment | | `controller/internal/backup/r357_reconstitute_headroom_test.go` | **new** | | `controller/internal/backup/r358_scratch_marker_test.go` | **new** | | `controller/internal/web/r353_unit_outcome_test.go` | **new** | | `controller/internal/web/r358_r360_handlers_test.go` | **new** | | 3 existing `*_test.go` in `internal/backup` | mechanical `_, err :=` for the changed signature | | `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md` | docs | **`felhom.eu`**: `STATUS.md`, `documentation/backlog/OPEN-ITEMS.md`, `documentation/backlog/CLOSED-ITEMS.md`, `documentation/architecture/07-backup-architecture.md`, `documentation/architecture/00-capability-map.md`, `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` (**new**). ## 3. Commits pushed to `main` | repo | commit | contents | |---|---|---| | `felhom-controller` | **`b8af7276`** | all four fixes, all tests, controller docs | | `felhom.eu` | **`e027b5d9`** | register closures, R-395 fix, architecture, capability map, evidence | | `felhom.eu` | *(this report's commit)* | housekeeping compression + REPORT | ## 4. Tests and red-proofs **All pass.** New tests, by group: | test | result | |---|---| | `TestUnitRestoreOutcome_VolumesAndDatabaseNamed` (A1) | pass | | `TestUnitRestoreOutcome_BackupHeldOnlySettings` (A2) | pass | | `TestUnitRestoreOutcome_ManifestListedDataThatDidNotReturn` (A3) | pass | | `TestUnitRestoreOutcome_DatabaseOnly` (A4) | pass | | `TestR353_HandlerPublishesTheOutcome` (A5, the seam test) | pass | | `TestR357_DestructiveRestoreRefusesWithoutHeadroom` (B1) | pass | | `TestR357_UnknownSizeFailsClosed` (B2) | pass | | `TestR357_UnknownFreeSpaceFailsClosed` (added — the mirror hole) | pass | | `TestR357_AmpleSpaceIsUnchanged` (B3) | pass | | `TestR358_FailedRestoreLeavesNoUsableScratch` (C1) | pass | | `TestR358_UnitOnlyScratchIsNotFullReady` (C2) | pass | | `TestR358_StaleMarkerIsClearedBeforeTheRun` (C3) | pass | | `TestR358_UnreadableMarkerFailsClosed` (C4) | pass | | `TestR358_WrongSchemaFailsClosed` (added) | pass | | `TestR358_CompletedFullScratchStillReady` (C5) | pass | | `TestR358_MarkerIsWrittenAt0600AndAtomically` (added) | pass | | `TestR358_MarkerIsNeverPlaced` (D3) | pass | | `TestR358_MarkerIsClearedBeforeResticAndWrittenAfter` (added, AST) | pass | | `TestR358_PlaceHandlerRefusesIncompleteScratch` (D1) | pass | | `TestR358_ReconstituteHandlerRefusesIncompleteScratch` (D2) | pass | | `TestR358_UnitOnlyScratchClosesTheFullRestoreCard` (Scenario F, flow level) | pass | | `TestR360_VerifyCopyDeleteRefusedDuringRestore` (D4) | pass | | `TestR360_VerifyCopyDeleteStillWorksWhenIdle` (Scenario H) | pass | **E1 — no existing test was modified for its content.** Every `TestReconstituteOutcome_*` passes unmodified. The only test-file edits were mechanical call-site updates for the changed `RestoreFromRecoveryUnit` signature (`err :=` → `_, err :=`) in three files. ### Red-proofs — each mutated, observed failing, reverted, `git diff` clean | # | mutation | observed failure | |---|---|---| | **A5** | `EndRestoreOp(true, stackName+" visszaállítva ("+snapshotID+").")` restored | `THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)."` | | **B1** | the whole R-357 gate deleted | `THE APP WAS STOPPED (1 call(s)) for a restore with 300 KB free for a 1 MB copy … (err=)` — and the same on both fail-closed tests | | **C1** | the old `len(entries) > 0` check restored | `a part-copy was reported READY`, plus the unit-only and unreadable-marker cases | | **D4** | the delete guard reverted to `IsRunning()` | `THE VERIFICATION COPY WAS DELETED while a restore was writing into it` | | **D1/D2** | both server-side scratch refusals removed | both handlers redirected with „…elindult" over a part-copy | > **⚠ THE FIRST B1 RED-PROOF EXPOSED A HOLLOW TEST OF MY OWN, and it is recorded rather than quietly > fixed.** With the gate removed, the run refused *earlier* — at the placement stat pre-pass — so > `stops == 0` passed against the pre-fix code and the test proved nothing about the thing it exists > for. Two corrections: the fixture now populates the scratch the way a completed download leaves it > (the snapshot's own absolute paths mirrored under the scratch, plus `SetSafetyDumpFn` so no Docker > is needed), and the assertions are **reordered** so a removed gate reports the outage rather than > "no error returned". Only after that does the red-proof print `THE APP WAS STOPPED (1 call(s))`. > The lesson is the doctrine's own: a test that cannot fail on the pre-fix shape is decoration. **Test count:** 24 new tests added across 4 new files. **Green gate:** `go build ./... && go vet ./... && go test ./...` → **28 packages, rc 0**, no failures. **All 12 controller gates OK.** ## 5. Deployed version ``` $ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'" gitea.dooplex.hu/admin/felhom-controller:0.226.0 Up 13 seconds (healthy) ``` **`demo-hp` only.** `demo-felhom` stays on 0.225.0 and the rest of the fleet on 0.223.0 (the floor). ## 6. Live validation — endpoint level, on `demo-hp` **Method:** the exact endpoints the UI invokes, driven from inside guest 9201 (no browser on DooPlex; the residual is client-side rendering only). No state was hand-set. Evidence copied off the box at the end of the phase: `felhom.eu/documentation/audits/evidence-r353-r360-live-2026-08-30/`. ### The verbatim messages **R-353** — `POST /backup/restore` for `opengist`, then the sentence read off the customer's own wizard page (`GET /backups/restore/app?name=opengist`): ``` A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult. ``` and the controller's own lines: ``` [INFO] [backup] Restore-from-unit completed: opengist — 1 volume(s) of 1 listed, 0 database(s) of 0 listed [INFO] [web] Restore completed (async): stack=opengist in 9.112668936s (volumes 1/1, dbs 0/0) ``` **This is Scenario A, not Scenario B, and the substitution is stated rather than glossed.** The spec named `opengist` as the data-less app from 21 August. **It is not one any more** — checked before relying on it, as the spec instructed: every unit on `demo-hp` today lists at least one volume dump (`opengist 1/0`, `privatebin 1/0`, `calibre-web 1/0`, the rest 2–3 volumes plus a database). **No app on the box has the Scenario B shape**, and manufacturing one would mean falsifying a manifest — the hand-set-state shortcut this project forbids. Scenario B is carried by `TestUnitRestoreOutcome_BackupHeldOnlySettings` and the A5 seam test. What the live run *does* prove is the whole path: real counts, correct clause selection (no database clause for 0 databases), and the sentence reaching the customer's page. **R-360** — a `mode=unit` scratch restore started, and the verification-copy delete POSTed while it ran. The refusal, urldecoded from the `Location` header: ``` /backups/restore?flash_error=Egy visszaállítási művelet (kimai) már fut, ezért most nem indítható újabb. Az állapotát ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást. ``` ``` [WARN] [web] verification-copy delete refused for kimai: a backup/restore op is running ``` **And the consequence, which is the assertion that matters:** a canary file planted in the copy was still there afterwards — `-rw-r--r-- 1 root root 10 Aug 30 17:34 canary.txt`, contents `defend-me`. **R-358 / R-396** — after the `mode=unit` restore completed: ``` $ cat …/backups/offsite-restore/kimai/.felhom-restore-complete.json {"schema":1,"snapshot_id":"84542ec8","full":false,"finished_at":"2026-08-30T17:34:42Z"} mode=600 (no .tmp left behind) [INFO] [offbox] kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy, so place-to-live stays closed ``` **This is Scenario F proven live**, and it is the case that pre-fix would have unlocked the destructive restore. ## 7. NOT yet live-validated — awaiting a supervised drill 1. **R-357's disk-full behaviour.** Filling a real filesystem to prove it is a drill step, not a build step. Carried by the seam tests, whose central assertion is `StopStack` call count 0. 2. **R-353's Scenario B** (the "backup held only settings" sentence) — no app on `demo-hp` has that unit shape any more; see §6. 3. **R-353's Scenario C** (the unit lists dumps, none return) — the R-367 stranded-dump shape was not reproduced on live hardware; unit-tested only. 4. **R-358's failed-download branch** was not induced live. It was not needed: the unit-only branch exercises the same marker gate through a real restore, without pointing restic at a bad snapshot. ## 8. Teardown **This run provisioned nothing.** No machine, no guest, no hub record — all three layers: - **Layer 1 (host):** nothing created on `felhom-pve` or `demo-hp`. No VM, no CT, no storage entry. - **Layer 2 (guest):** two throwaway shell scripts were pushed into guest 9201 to drive the endpoints and **both were removed**; one canary directory was planted under `backups/offsite-restore/kimai` on the system drive to prove R-360 and **was removed**. The `mode=unit` scratch restore left a normal verification copy on the HDD, which is ordinary product state a customer can delete. - **Layer 3 (hub):** no customer, appliance or host record created, so none to discard. `demo-felhom`, `ep0`, DooPlex and Peti's box were not touched. ## 9. Register **Size: 165 rows before → 167 after filing → 161 after housekeeping.** - **Closed:** R-353, R-357, R-358, R-360 — each with shipping version and evidence path. - **Filed and closed in the same session:** **R-396** (Scenario F's answer — see §11) and **R-395** (`STATUS.md` contradicting itself; **fixed**, not merely recorded). - **Re-ranked:** none. - **Housekeeping:** the six closed rows were compressed into `CLOSED-ITEMS.md` keeping title, version, evidence and every sentence that states a rule; the full original is `git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md`. **No open row was touched.** - **`ROADMAP.md` was not edited:** none of these four ever had a ROADMAP row — they live in the register only. Stated rather than silently skipped. ## 10. Documentation coupling (§5.5) - **`00-capability-map.md`** — one new row, splitting the claim: R-353/R-358/R-360 **PROVEN-LIVE** with the evidence citation; **R-357 IMPLEMENTED only**, with the reason. - **`07-backup-architecture.md`** — §10.2 gained five rows (the four plus R-396). **§8 matrix row 3 KEEPS its PROVEN status**, with a note recording why: R-353 was a defect in the *message*, not the mechanism. - **`OPEN-ITEMS.md` / `CLOSED-ITEMS.md` / `STATUS.md`** — as §9. - **`ROADMAP.md`** — no applicable rows. - **Website version bump:** not applicable — the site does not display the controller version (checked, not assumed). ## 11. Observations — noticed, documented, not acted on 1. **Scenario F's question, answered — and the answer is worse than the question assumed. Filed as R-396.** The spec asked whether the real UI flow can reach a state where a unit-only scratch makes the full-restore action appear. **It can, by the safest-looking action on the page.** „Ellenőrző visszaállítás" (`mode=unit`, the default, advertised as non-destructive) calls `RestoreOffboxScratch(full=false)`; `offboxRestoreScratchDir` **ignores `full`**, so both modes write the same directory, and `--include` limits *what* restic extracts, never *where*; the wizard sets `ScratchReady` from `OffboxFullScratchReady`; `deriveWizardStep` derives **both** `PlaceEnabled` and `RestoreEnabled` from that one flag. R-358 as filed assumed the bad state needed a *failed* download. It needed only a successful safe one. **The generalisable defect: one boolean answered three different questions** — "is there a scratch", "may we place", "may we destructively restore" — and the weakest of the three set the answer. 2. **`NotifyIntegrityOK` / `NotifyIntegrityFailed` are dead code** (noticed during R-331 earlier today, restated here because it is restore-adjacent): they exist and are called from nowhere; the controller runs no integrity check at all. 3. **`resticStep` is not a seam**, which is why R-358's ordering property needed an AST test rather than an execution test. If a future task needs to drive restic paths under test, that is the seam to add — and it should be added deliberately, not improvised inside a bugfix. 4. **`RestoreApp`'s own `restoreDockerVolumes` count is still discarded.** Left alone deliberately: the spec scoped `RestoreApp` out, and it is the no-unit fallback whose result is already reported as a zero `UnitRestoreResult` — which is honest, because a box with no recovery unit really did return no data from one. Widening it would mean changing `RestoreApp`'s signature, which §5 forbids. 5. **The spec's own §13 step 1 named an app whose shape has since changed.** Not a defect in the spec — a reminder that fixture assumptions about live boxes decay, and the instruction to "confirm its manifest first rather than assuming" is what caught it.