# REPORT — R-353 / R-357 / R-358 / R-360: the restore tells the truth **Controller v0.226.0 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`** --- ## 1. Confirmed baselines used — BOTH HAD MOVED | repo | spec baseline | **actual at start** | version | |---|---|---|---| | `felhom-controller` | `f8c9390` / v0.223.0 → v0.224.0 | **`e5eee50` / v0.225.0** | → **v0.226.0** | | `felhom.eu` | `c2c1fb4` | **`ac6ac03`** | docs only | | `felhom-agent` | not touched | v0.130.0 | unchanged | **Both targets were consumed earlier the same day by my own work** — v0.224.0 (R-330) and v0.225.0 (R-331). The drift was re-confirmed against live Gitea before the first edit and the operator authorised proceeding. **Every symbol in the spec's §5 table was re-verified present at the real baseline** before editing; all 20 resolved, and almost every line landmark still matched. `MinAgent` stays **0.129.0**. Both trees were clean and equal to `origin/main` at the start. ## 2. Files created / modified **`felhom-controller`** (all paths under repo root): | file | change | |---|---| | `controller/internal/backup/restore.go` | `restoreDockerVolumes` → `(int, error)`; caller updated | | `controller/internal/backup/restore_unit.go` | `UnitRestoreResult`; `RestoreFromRecoveryUnit` → `(UnitRestoreResult, error)` on every return path | | `controller/internal/web/handlers.go` | `unitRestoreOutcomeMsg` + 3 message constants; handler publishes the outcome | | `controller/internal/backup/offbox_reconstitute.go` | the R-357 headroom gate | | `controller/internal/backup/offbox_restore.go` | scratch marker (write/clear/read), `OffboxFullScratchReady` rewritten, shared refusal constants, `SetOffboxLatestSnapshotFn`, `WriteScratchMarkerForTest` | | `controller/internal/backup/backup.go` | `offboxLatestSnapFn` field | | `controller/internal/web/offbox_handlers.go` | server-side scratch refusals in 2 handlers; R-360 delete guard; corrected doc comment | | `controller/internal/backup/r357_reconstitute_headroom_test.go` | **new** | | `controller/internal/backup/r358_scratch_marker_test.go` | **new** | | `controller/internal/web/r353_unit_outcome_test.go` | **new** | | `controller/internal/web/r358_r360_handlers_test.go` | **new** | | 3 existing `*_test.go` in `internal/backup` | mechanical `_, err :=` for the changed signature | | `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md` | docs | **`felhom.eu`**: `STATUS.md`, `documentation/backlog/OPEN-ITEMS.md`, `documentation/backlog/CLOSED-ITEMS.md`, `documentation/architecture/07-backup-architecture.md`, `documentation/architecture/00-capability-map.md`, `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` (**new**). ## 3. Commits pushed to `main` | repo | commit | contents | |---|---|---| | `felhom-controller` | **`b8af7276`** | all four fixes, all tests, controller docs | | `felhom.eu` | **`e027b5d9`** | register closures, R-395 fix, architecture, capability map, evidence | | `felhom.eu` | *(this report's commit)* | housekeeping compression + REPORT | ## 4. Tests and red-proofs **All pass.** New tests, by group: | test | result | |---|---| | `TestUnitRestoreOutcome_VolumesAndDatabaseNamed` (A1) | pass | | `TestUnitRestoreOutcome_BackupHeldOnlySettings` (A2) | pass | | `TestUnitRestoreOutcome_ManifestListedDataThatDidNotReturn` (A3) | pass | | `TestUnitRestoreOutcome_DatabaseOnly` (A4) | pass | | `TestR353_HandlerPublishesTheOutcome` (A5, the seam test) | pass | | `TestR357_DestructiveRestoreRefusesWithoutHeadroom` (B1) | pass | | `TestR357_UnknownSizeFailsClosed` (B2) | pass | | `TestR357_UnknownFreeSpaceFailsClosed` (added — the mirror hole) | pass | | `TestR357_AmpleSpaceIsUnchanged` (B3) | pass | | `TestR358_FailedRestoreLeavesNoUsableScratch` (C1) | pass | | `TestR358_UnitOnlyScratchIsNotFullReady` (C2) | pass | | `TestR358_StaleMarkerIsClearedBeforeTheRun` (C3) | pass | | `TestR358_UnreadableMarkerFailsClosed` (C4) | pass | | `TestR358_WrongSchemaFailsClosed` (added) | pass | | `TestR358_CompletedFullScratchStillReady` (C5) | pass | | `TestR358_MarkerIsWrittenAt0600AndAtomically` (added) | pass | | `TestR358_MarkerIsNeverPlaced` (D3) | pass | | `TestR358_MarkerIsClearedBeforeResticAndWrittenAfter` (added, AST) | pass | | `TestR358_PlaceHandlerRefusesIncompleteScratch` (D1) | pass | | `TestR358_ReconstituteHandlerRefusesIncompleteScratch` (D2) | pass | | `TestR358_UnitOnlyScratchClosesTheFullRestoreCard` (Scenario F, flow level) | pass | | `TestR360_VerifyCopyDeleteRefusedDuringRestore` (D4) | pass | | `TestR360_VerifyCopyDeleteStillWorksWhenIdle` (Scenario H) | pass | **E1 — no existing test was modified for its content.** Every `TestReconstituteOutcome_*` passes unmodified. The only test-file edits were mechanical call-site updates for the changed `RestoreFromRecoveryUnit` signature (`err :=` → `_, err :=`) in three files. ### Red-proofs — each mutated, observed failing, reverted, `git diff` clean | # | mutation | observed failure | |---|---|---| | **A5** | `EndRestoreOp(true, stackName+" visszaállítva ("+snapshotID+").")` restored | `THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)."` | | **B1** | the whole R-357 gate deleted | `THE APP WAS STOPPED (1 call(s)) for a restore with 300 KB free for a 1 MB copy … (err=)` — and the same on both fail-closed tests | | **C1** | the old `len(entries) > 0` check restored | `a part-copy was reported READY`, plus the unit-only and unreadable-marker cases | | **D4** | the delete guard reverted to `IsRunning()` | `THE VERIFICATION COPY WAS DELETED while a restore was writing into it` | | **D1/D2** | both server-side scratch refusals removed | both handlers redirected with „…elindult" over a part-copy | > **⚠ THE FIRST B1 RED-PROOF EXPOSED A HOLLOW TEST OF MY OWN, and it is recorded rather than quietly > fixed.** With the gate removed, the run refused *earlier* — at the placement stat pre-pass — so > `stops == 0` passed against the pre-fix code and the test proved nothing about the thing it exists > for. Two corrections: the fixture now populates the scratch the way a completed download leaves it > (the snapshot's own absolute paths mirrored under the scratch, plus `SetSafetyDumpFn` so no Docker > is needed), and the assertions are **reordered** so a removed gate reports the outage rather than > "no error returned". Only after that does the red-proof print `THE APP WAS STOPPED (1 call(s))`. > The lesson is the doctrine's own: a test that cannot fail on the pre-fix shape is decoration. **Test count:** 24 new tests added across 4 new files. **Green gate:** `go build ./... && go vet ./... && go test ./...` → **28 packages, rc 0**, no failures. **All 12 controller gates OK.** ## 5. Deployed version ``` $ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'" gitea.dooplex.hu/admin/felhom-controller:0.226.0 Up 13 seconds (healthy) ``` **`demo-hp` only.** `demo-felhom` stays on 0.225.0 and the rest of the fleet on 0.223.0 (the floor). ## 6. Live validation — endpoint level, on `demo-hp` **Method:** the exact endpoints the UI invokes, driven from inside guest 9201 (no browser on DooPlex; the residual is client-side rendering only). No state was hand-set. Evidence copied off the box at the end of the phase: `felhom.eu/documentation/audits/evidence-r353-r360-live-2026-08-30/`. ### The verbatim messages **R-353** — `POST /backup/restore` for `opengist`, then the sentence read off the customer's own wizard page (`GET /backups/restore/app?name=opengist`): ``` A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult. ``` and the controller's own lines: ``` [INFO] [backup] Restore-from-unit completed: opengist — 1 volume(s) of 1 listed, 0 database(s) of 0 listed [INFO] [web] Restore completed (async): stack=opengist in 9.112668936s (volumes 1/1, dbs 0/0) ``` **This is Scenario A, not Scenario B, and the substitution is stated rather than glossed.** The spec named `opengist` as the data-less app from 21 August. **It is not one any more** — checked before relying on it, as the spec instructed: every unit on `demo-hp` today lists at least one volume dump (`opengist 1/0`, `privatebin 1/0`, `calibre-web 1/0`, the rest 2–3 volumes plus a database). **No app on the box has the Scenario B shape**, and manufacturing one would mean falsifying a manifest — the hand-set-state shortcut this project forbids. Scenario B is carried by `TestUnitRestoreOutcome_BackupHeldOnlySettings` and the A5 seam test. What the live run *does* prove is the whole path: real counts, correct clause selection (no database clause for 0 databases), and the sentence reaching the customer's page. **R-360** — a `mode=unit` scratch restore started, and the verification-copy delete POSTed while it ran. The refusal, urldecoded from the `Location` header: ``` /backups/restore?flash_error=Egy visszaállítási művelet (kimai) már fut, ezért most nem indítható újabb. Az állapotát ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást. ``` ``` [WARN] [web] verification-copy delete refused for kimai: a backup/restore op is running ``` **And the consequence, which is the assertion that matters:** a canary file planted in the copy was still there afterwards — `-rw-r--r-- 1 root root 10 Aug 30 17:34 canary.txt`, contents `defend-me`. **R-358 / R-396** — after the `mode=unit` restore completed: ``` $ cat …/backups/offsite-restore/kimai/.felhom-restore-complete.json {"schema":1,"snapshot_id":"84542ec8","full":false,"finished_at":"2026-08-30T17:34:42Z"} mode=600 (no .tmp left behind) [INFO] [offbox] kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy, so place-to-live stays closed ``` **This is Scenario F proven live**, and it is the case that pre-fix would have unlocked the destructive restore. ## 7. NOT yet live-validated — awaiting a supervised drill 1. **R-357's disk-full behaviour.** Filling a real filesystem to prove it is a drill step, not a build step. Carried by the seam tests, whose central assertion is `StopStack` call count 0. 2. **R-353's Scenario B** (the "backup held only settings" sentence) — no app on `demo-hp` has that unit shape any more; see §6. 3. **R-353's Scenario C** (the unit lists dumps, none return) — the R-367 stranded-dump shape was not reproduced on live hardware; unit-tested only. 4. **R-358's failed-download branch** was not induced live. It was not needed: the unit-only branch exercises the same marker gate through a real restore, without pointing restic at a bad snapshot. ## 8. Teardown **This run provisioned nothing.** No machine, no guest, no hub record — all three layers: - **Layer 1 (host):** nothing created on `felhom-pve` or `demo-hp`. No VM, no CT, no storage entry. - **Layer 2 (guest):** two throwaway shell scripts were pushed into guest 9201 to drive the endpoints and **both were removed**; one canary directory was planted under `backups/offsite-restore/kimai` on the system drive to prove R-360 and **was removed**. The `mode=unit` scratch restore left a normal verification copy on the HDD, which is ordinary product state a customer can delete. - **Layer 3 (hub):** no customer, appliance or host record created, so none to discard. `demo-felhom`, `ep0`, DooPlex and Peti's box were not touched. ## 9. Register **Size: 165 rows before → 167 after filing → 161 after housekeeping.** - **Closed:** R-353, R-357, R-358, R-360 — each with shipping version and evidence path. - **Filed and closed in the same session:** **R-396** (Scenario F's answer — see §11) and **R-395** (`STATUS.md` contradicting itself; **fixed**, not merely recorded). - **Re-ranked:** none. - **Housekeeping:** the six closed rows were compressed into `CLOSED-ITEMS.md` keeping title, version, evidence and every sentence that states a rule; the full original is `git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md`. **No open row was touched.** - **`ROADMAP.md` was not edited:** none of these four ever had a ROADMAP row — they live in the register only. Stated rather than silently skipped. ## 10. Documentation coupling (§5.5) - **`00-capability-map.md`** — one new row, splitting the claim: R-353/R-358/R-360 **PROVEN-LIVE** with the evidence citation; **R-357 IMPLEMENTED only**, with the reason. - **`07-backup-architecture.md`** — §10.2 gained five rows (the four plus R-396). **§8 matrix row 3 KEEPS its PROVEN status**, with a note recording why: R-353 was a defect in the *message*, not the mechanism. - **`OPEN-ITEMS.md` / `CLOSED-ITEMS.md` / `STATUS.md`** — as §9. - **`ROADMAP.md`** — no applicable rows. - **Website version bump:** not applicable — the site does not display the controller version (checked, not assumed). ## 11. Observations — each with a register row, or a stated reason it needs none 1. **FILED: R-396.** *Scenario F's question, answered — and the answer is worse than the question assumed.* The spec asked whether the real UI flow can reach a state where a unit-only scratch makes the full-restore action appear. **It can, by the safest-looking action on the page.** „Ellenőrző visszaállítás" (`mode=unit`, the default, advertised as non-destructive) calls `RestoreOffboxScratch(full=false)`; `offboxRestoreScratchDir` **ignores `full`**, so both modes write the same directory, and `--include` limits *what* restic extracts, never *where*; the wizard sets `ScratchReady` from `OffboxFullScratchReady`; `deriveWizardStep` derives **both** `PlaceEnabled` and `RestoreEnabled` from that one flag. R-358 as filed assumed the bad state needed a *failed* download. It needed only a successful safe one. **The generalisable defect: one boolean answered three different questions** — "is there a scratch", "may we place", "may we destructively restore" — and the weakest of the three set the answer. 2. **FILED: R-397.** *`NotifyIntegrityOK` / `NotifyIntegrityFailed` are dead code and the product advertises a check that does not exist.* Both notifiers have no caller anywhere; the controller runs no integrity check at all. **The part that actively misleads:** `config.Monitoring.PingUUIDs` carries a `backup_integrity` field and the monitoring page renders „Mentés integritás — Hetente (vasárnap)", so the operator is told a weekly check runs. Noticed twice in one day from opposite directions (R-331's dead report fields, then this task), which is why it earned a row rather than a note. 3. **FILED: R-398.** *`resticStep` is not a seam, so no test can drive any restic-backed path.* Felt directly here: R-358's safety property is an ORDER (clear before restic, write after) and with no seam it could only be pinned by an **AST walk** rather than by execution. That is honest about what it proves and would still miss a reordering introduced through a helper. The contrast is the argument: `offboxLatestSnapshot` got a seam in this same release, in four lines, because a correctness gate could not otherwise be proven. 4. **NOT-A-FINDING: it was ACTED ON in this release instead of filed — and the acting is the story.** `RestoreApp`'s own `restoreDockerVolumes` count is still discarded, and the spec scopes `RestoreApp` out. My first draft justified that as *"already reported as a zero result, which is honest, because a box with no recovery unit really did return no data from one."* > **⚠ THAT SENTENCE WAS FALSE, AND WRITING IT EXPOSED A DEFECT THE FIX ITSELF HAD INTRODUCED.** > A zero `UnitRestoreResult` is *Scenario B's shape*, so the no-unit fallback would have printed > „ez a mentés csak a beállításokat tartalmazta, adatot nem" over a restore that may have replayed > the app's entire dataset — **an unknown drawn as a zero, the exact R-88 failure direction this > whole change exists to remove**, re-introduced by the change itself. > > Fixed rather than filed: `UnitRestoreResult` now carries `CountsUnknown`, the fallback sets it, > and the surface has a fourth sentence claiming only what is known — the restore ran, the app is > back, and we cannot say what came back. **`RestoreApp`'s signature is untouched**, so §5 holds. > Pinned by `TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty`; the A5 seam test was > corrected too, because its fixture has no unit and therefore exercises exactly this path while > asserting the wrong sentence. > > **It was the `observations` gate refusing the push that forced the re-read.** The gate exists to > stop findings dying in an overwritten REPORT.md; here it caught a live defect instead. 5. **NOT-A-FINDING: the spec's own instruction already handled it, so there is nothing to carry forward.** §13 step 1 named `opengist` as the data-less app; it is not one any more. Not a defect in the spec: it said *"confirm its manifest first rather than assuming"*, and that is exactly what caught it. Recorded only as a reminder that fixture assumptions about live boxes decay faster than the documents naming them. ## 12. Two gates fired; one push was bypassed, and both are declared **`felhom.eu` — `golden-currency` CONVICTED.** Three controller releases shipped today (0.224.0, 0.225.0, 0.226.0) and the vouched golden still carries **0.223.0**, so a machine installed right now receives none of them. **The gate is right.** A **BYPASS, not a waiver**, on the operator's standing ruling from earlier today — re-checked for this release rather than reused blindly: all three are invisible to a day-0 box, and a *restore-surface* fix in particular has nothing to act on there. **The ground expires the moment a release changes first-boot behaviour.** Tracked on **R-242**; **one bake carrying 0.226.0 covers all three**, then raise the floor to 0.226.0. **`felhom-controller` — nothing was bypassed.** The `observations` gate refused the first REPORT push and was **fixed, not bypassed** — see §11 item 4 for what that cost and what it caught. **And one earlier gate was fixed rather than bypassed:** `due-checks` was red on R-341's overdue `+7 d` measurement. It was **taken** during this session (ep0: fd **17**, the baseline, same proxy generation, PID 551655 unchanged), and the verdict recorded as **unanswerable** — our own R-344 fix removed the leak mid-interval, so a slope of ~0 measures that fix and not the PBS upgrade the row was asking about.