Writing the REPORT's observation "the no-unit fallback already reports a zero result, which is honest" exposed that the sentence was FALSE. A zero UnitRestoreResult is Scenario B's shape. So RestoreFromRecoveryUnit's fallback to RestoreApp -- which returns only an error, and whose signature is deliberately out of scope -- would have printed "ez a mentes csak a beallitasokat tartalmazta, adatot nem" over a restore that may have replayed the app's entire dataset. That is an unknown drawn as a zero: the exact R-88 failure direction this whole change exists to remove, re-introduced by the change. UnitRestoreResult now carries CountsUnknown, the fallback sets it, and there is a fourth sentence claiming only what is known -- the restore ran, the app is back, and we cannot say what came back. RestoreApp's signature is untouched. Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty. The A5 seam test was corrected too: its fixture has no recovery unit, so it exercises exactly this path and had been asserting the wrong sentence -- it now asserts the unknown, which is what pins the fallback to it. IT WAS THE observations GATE REFUSING THE PUSH THAT FORCED THE RE-READ. A gate written to stop findings dying in an overwritten REPORT.md caught a live defect instead. Also files R-397 (NotifyIntegrityOK/Failed are dead code AND the monitoring page advertises a weekly integrity check that does not exist) and R-398 (resticStep is not a seam, which is why R-358's ordering needed an AST test) rather than leaving them in a file that is overwritten every session. REPORT.md is the full run record: baselines re-confirmed, per-test results, the five red-proofs with their observed output, the live validation with verbatim Hungarian messages, what was NOT validated and why, teardown across three layers, and the register 165 -> 167 -> 161. Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
18 KiB
REPORT — R-353 / R-357 / R-358 / R-360: the restore tells the truth
Controller v0.226.0 · 2026-08-30 · implemented on DooPlex, validated live on demo-hp
1. Confirmed baselines used — BOTH HAD MOVED
| repo | spec baseline | actual at start | version |
|---|---|---|---|
felhom-controller |
f8c9390 / v0.223.0 → v0.224.0 |
e5eee50 / v0.225.0 |
→ v0.226.0 |
felhom.eu |
c2c1fb4 |
ac6ac03 |
docs only |
felhom-agent |
not touched | v0.130.0 | unchanged |
Both targets were consumed earlier the same day by my own work — v0.224.0 (R-330) and v0.225.0
(R-331). The drift was re-confirmed against live Gitea before the first edit and the operator
authorised proceeding. Every symbol in the spec's §5 table was re-verified present at the real
baseline before editing; all 20 resolved, and almost every line landmark still matched. MinAgent
stays 0.129.0. Both trees were clean and equal to origin/main at the start.
2. Files created / modified
felhom-controller (all paths under repo root):
| file | change |
|---|---|
controller/internal/backup/restore.go |
restoreDockerVolumes → (int, error); caller updated |
controller/internal/backup/restore_unit.go |
UnitRestoreResult; RestoreFromRecoveryUnit → (UnitRestoreResult, error) on every return path |
controller/internal/web/handlers.go |
unitRestoreOutcomeMsg + 3 message constants; handler publishes the outcome |
controller/internal/backup/offbox_reconstitute.go |
the R-357 headroom gate |
controller/internal/backup/offbox_restore.go |
scratch marker (write/clear/read), OffboxFullScratchReady rewritten, shared refusal constants, SetOffboxLatestSnapshotFn, WriteScratchMarkerForTest |
controller/internal/backup/backup.go |
offboxLatestSnapFn field |
controller/internal/web/offbox_handlers.go |
server-side scratch refusals in 2 handlers; R-360 delete guard; corrected doc comment |
controller/internal/backup/r357_reconstitute_headroom_test.go |
new |
controller/internal/backup/r358_scratch_marker_test.go |
new |
controller/internal/web/r353_unit_outcome_test.go |
new |
controller/internal/web/r358_r360_handlers_test.go |
new |
3 existing *_test.go in internal/backup |
mechanical _, err := for the changed signature |
CHANGELOG.md, CONTEXT.md, REUSE.md, controller/README.md, REPORT.md |
docs |
felhom.eu: STATUS.md, documentation/backlog/OPEN-ITEMS.md,
documentation/backlog/CLOSED-ITEMS.md, documentation/architecture/07-backup-architecture.md,
documentation/architecture/00-capability-map.md,
documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt (new).
3. Commits pushed to main
| repo | commit | contents |
|---|---|---|
felhom-controller |
b8af7276 |
all four fixes, all tests, controller docs |
felhom.eu |
e027b5d9 |
register closures, R-395 fix, architecture, capability map, evidence |
felhom.eu |
(this report's commit) | housekeeping compression + REPORT |
4. Tests and red-proofs
All pass. New tests, by group:
| test | result |
|---|---|
TestUnitRestoreOutcome_VolumesAndDatabaseNamed (A1) |
pass |
TestUnitRestoreOutcome_BackupHeldOnlySettings (A2) |
pass |
TestUnitRestoreOutcome_ManifestListedDataThatDidNotReturn (A3) |
pass |
TestUnitRestoreOutcome_DatabaseOnly (A4) |
pass |
TestR353_HandlerPublishesTheOutcome (A5, the seam test) |
pass |
TestR357_DestructiveRestoreRefusesWithoutHeadroom (B1) |
pass |
TestR357_UnknownSizeFailsClosed (B2) |
pass |
TestR357_UnknownFreeSpaceFailsClosed (added — the mirror hole) |
pass |
TestR357_AmpleSpaceIsUnchanged (B3) |
pass |
TestR358_FailedRestoreLeavesNoUsableScratch (C1) |
pass |
TestR358_UnitOnlyScratchIsNotFullReady (C2) |
pass |
TestR358_StaleMarkerIsClearedBeforeTheRun (C3) |
pass |
TestR358_UnreadableMarkerFailsClosed (C4) |
pass |
TestR358_WrongSchemaFailsClosed (added) |
pass |
TestR358_CompletedFullScratchStillReady (C5) |
pass |
TestR358_MarkerIsWrittenAt0600AndAtomically (added) |
pass |
TestR358_MarkerIsNeverPlaced (D3) |
pass |
TestR358_MarkerIsClearedBeforeResticAndWrittenAfter (added, AST) |
pass |
TestR358_PlaceHandlerRefusesIncompleteScratch (D1) |
pass |
TestR358_ReconstituteHandlerRefusesIncompleteScratch (D2) |
pass |
TestR358_UnitOnlyScratchClosesTheFullRestoreCard (Scenario F, flow level) |
pass |
TestR360_VerifyCopyDeleteRefusedDuringRestore (D4) |
pass |
TestR360_VerifyCopyDeleteStillWorksWhenIdle (Scenario H) |
pass |
E1 — no existing test was modified for its content. Every TestReconstituteOutcome_* passes
unmodified. The only test-file edits were mechanical call-site updates for the changed
RestoreFromRecoveryUnit signature (err := → _, err :=) in three files.
Red-proofs — each mutated, observed failing, reverted, git diff clean
| # | mutation | observed failure |
|---|---|---|
| A5 | EndRestoreOp(true, stackName+" visszaállítva ("+snapshotID+").") restored |
THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)." |
| B1 | the whole R-357 gate deleted | THE APP WAS STOPPED (1 call(s)) for a restore with 300 KB free for a 1 MB copy … (err=<nil>) — and the same on both fail-closed tests |
| C1 | the old len(entries) > 0 check restored |
a part-copy was reported READY, plus the unit-only and unreadable-marker cases |
| D4 | the delete guard reverted to IsRunning() |
THE VERIFICATION COPY WAS DELETED while a restore was writing into it |
| D1/D2 | both server-side scratch refusals removed | both handlers redirected with „…elindult" over a part-copy |
⚠ THE FIRST B1 RED-PROOF EXPOSED A HOLLOW TEST OF MY OWN, and it is recorded rather than quietly fixed. With the gate removed, the run refused earlier — at the placement stat pre-pass — so
stops == 0passed against the pre-fix code and the test proved nothing about the thing it exists for. Two corrections: the fixture now populates the scratch the way a completed download leaves it (the snapshot's own absolute paths mirrored under the scratch, plusSetSafetyDumpFnso no Docker is needed), and the assertions are reordered so a removed gate reports the outage rather than "no error returned". Only after that does the red-proof printTHE APP WAS STOPPED (1 call(s)). The lesson is the doctrine's own: a test that cannot fail on the pre-fix shape is decoration.
Test count: 24 new tests added across 4 new files. Green gate:
go build ./... && go vet ./... && go test ./... → 28 packages, rc 0, no failures.
All 12 controller gates OK.
5. Deployed version
$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.226.0 Up 13 seconds (healthy)
demo-hp only. demo-felhom stays on 0.225.0 and the rest of the fleet on 0.223.0 (the floor).
6. Live validation — endpoint level, on demo-hp
Method: the exact endpoints the UI invokes, driven from inside guest 9201 (no browser on
DooPlex; the residual is client-side rendering only). No state was hand-set. Evidence copied off the
box at the end of the phase: felhom.eu/documentation/audits/evidence-r353-r360-live-2026-08-30/.
The verbatim messages
R-353 — POST /backup/restore for opengist, then the sentence read off the customer's own
wizard page (GET /backups/restore/app?name=opengist):
A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.
and the controller's own lines:
[INFO] [backup] Restore-from-unit completed: opengist — 1 volume(s) of 1 listed, 0 database(s) of 0 listed
[INFO] [web] Restore completed (async): stack=opengist in 9.112668936s (volumes 1/1, dbs 0/0)
This is Scenario A, not Scenario B, and the substitution is stated rather than glossed. The spec
named opengist as the data-less app from 21 August. It is not one any more — checked before
relying on it, as the spec instructed: every unit on demo-hp today lists at least one volume dump
(opengist 1/0, privatebin 1/0, calibre-web 1/0, the rest 2–3 volumes plus a database). No app
on the box has the Scenario B shape, and manufacturing one would mean falsifying a manifest — the
hand-set-state shortcut this project forbids. Scenario B is carried by
TestUnitRestoreOutcome_BackupHeldOnlySettings and the A5 seam test. What the live run does prove
is the whole path: real counts, correct clause selection (no database clause for 0 databases), and
the sentence reaching the customer's page.
R-360 — a mode=unit scratch restore started, and the verification-copy delete POSTed while it
ran. The refusal, urldecoded from the Location header:
/backups/restore?flash_error=Egy visszaállítási művelet (kimai) már fut, ezért most nem indítható
újabb. Az állapotát ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást.
[WARN] [web] verification-copy delete refused for kimai: a backup/restore op is running
And the consequence, which is the assertion that matters: a canary file planted in the copy was
still there afterwards — -rw-r--r-- 1 root root 10 Aug 30 17:34 canary.txt, contents defend-me.
R-358 / R-396 — after the mode=unit restore completed:
$ cat …/backups/offsite-restore/kimai/.felhom-restore-complete.json
{"schema":1,"snapshot_id":"84542ec8","full":false,"finished_at":"2026-08-30T17:34:42Z"}
mode=600 (no .tmp left behind)
[INFO] [offbox] kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy,
so place-to-live stays closed
This is Scenario F proven live, and it is the case that pre-fix would have unlocked the destructive restore.
7. NOT yet live-validated — awaiting a supervised drill
- R-357's disk-full behaviour. Filling a real filesystem to prove it is a drill step, not a build
step. Carried by the seam tests, whose central assertion is
StopStackcall count 0. - R-353's Scenario B (the "backup held only settings" sentence) — no app on
demo-hphas that unit shape any more; see §6. - R-353's Scenario C (the unit lists dumps, none return) — the R-367 stranded-dump shape was not reproduced on live hardware; unit-tested only.
- R-358's failed-download branch was not induced live. It was not needed: the unit-only branch exercises the same marker gate through a real restore, without pointing restic at a bad snapshot.
8. Teardown
This run provisioned nothing. No machine, no guest, no hub record — all three layers:
- Layer 1 (host): nothing created on
felhom-pveordemo-hp. No VM, no CT, no storage entry. - Layer 2 (guest): two throwaway shell scripts were pushed into guest 9201 to drive the endpoints
and both were removed; one canary directory was planted under
backups/offsite-restore/kimaion the system drive to prove R-360 and was removed. Themode=unitscratch restore left a normal verification copy on the HDD, which is ordinary product state a customer can delete. - Layer 3 (hub): no customer, appliance or host record created, so none to discard.
demo-felhom, ep0, DooPlex and Peti's box were not touched.
9. Register
Size: 165 rows before → 167 after filing → 161 after housekeeping.
- Closed: R-353, R-357, R-358, R-360 — each with shipping version and evidence path.
- Filed and closed in the same session: R-396 (Scenario F's answer — see §11) and R-395
(
STATUS.mdcontradicting itself; fixed, not merely recorded). - Re-ranked: none.
- Housekeeping: the six closed rows were compressed into
CLOSED-ITEMS.mdkeeping title, version, evidence and every sentence that states a rule; the full original isgit show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md. No open row was touched. ROADMAP.mdwas not edited: none of these four ever had a ROADMAP row — they live in the register only. Stated rather than silently skipped.
10. Documentation coupling (§5.5)
00-capability-map.md— one new row, splitting the claim: R-353/R-358/R-360 PROVEN-LIVE with the evidence citation; R-357 IMPLEMENTED only, with the reason.07-backup-architecture.md— §10.2 gained five rows (the four plus R-396). §8 matrix row 3 KEEPS its PROVEN status, with a note recording why: R-353 was a defect in the message, not the mechanism.OPEN-ITEMS.md/CLOSED-ITEMS.md/STATUS.md— as §9.ROADMAP.md— no applicable rows.- Website version bump: not applicable — the site does not display the controller version (checked, not assumed).
11. Observations — each with a register row, or a stated reason it needs none
-
FILED: R-396. Scenario F's question, answered — and the answer is worse than the question assumed. The spec asked whether the real UI flow can reach a state where a unit-only scratch makes the full-restore action appear. It can, by the safest-looking action on the page. „Ellenőrző visszaállítás" (
mode=unit, the default, advertised as non-destructive) callsRestoreOffboxScratch(full=false);offboxRestoreScratchDirignoresfull, so both modes write the same directory, and--includelimits what restic extracts, never where; the wizard setsScratchReadyfromOffboxFullScratchReady;deriveWizardStepderives bothPlaceEnabledandRestoreEnabledfrom that one flag. R-358 as filed assumed the bad state needed a failed download. It needed only a successful safe one. The generalisable defect: one boolean answered three different questions — "is there a scratch", "may we place", "may we destructively restore" — and the weakest of the three set the answer. -
FILED: R-397.
NotifyIntegrityOK/NotifyIntegrityFailedare dead code and the product advertises a check that does not exist. Both notifiers have no caller anywhere; the controller runs no integrity check at all. The part that actively misleads:config.Monitoring.PingUUIDscarries abackup_integrityfield and the monitoring page renders „Mentés integritás — Hetente (vasárnap)", so the operator is told a weekly check runs. Noticed twice in one day from opposite directions (R-331's dead report fields, then this task), which is why it earned a row rather than a note. -
FILED: R-398.
resticStepis not a seam, so no test can drive any restic-backed path. Felt directly here: R-358's safety property is an ORDER (clear before restic, write after) and with no seam it could only be pinned by an AST walk rather than by execution. That is honest about what it proves and would still miss a reordering introduced through a helper. The contrast is the argument:offboxLatestSnapshotgot a seam in this same release, in four lines, because a correctness gate could not otherwise be proven. -
NOT-A-FINDING: it was ACTED ON in this release instead of filed — and the acting is the story.
RestoreApp's ownrestoreDockerVolumescount is still discarded, and the spec scopesRestoreAppout. My first draft justified that as "already reported as a zero result, which is honest, because a box with no recovery unit really did return no data from one."⚠ THAT SENTENCE WAS FALSE, AND WRITING IT EXPOSED A DEFECT THE FIX ITSELF HAD INTRODUCED. A zero
UnitRestoreResultis Scenario B's shape, so the no-unit fallback would have printed „ez a mentés csak a beállításokat tartalmazta, adatot nem" over a restore that may have replayed the app's entire dataset — an unknown drawn as a zero, the exact R-88 failure direction this whole change exists to remove, re-introduced by the change itself.Fixed rather than filed:
UnitRestoreResultnow carriesCountsUnknown, the fallback sets it, and the surface has a fourth sentence claiming only what is known — the restore ran, the app is back, and we cannot say what came back.RestoreApp's signature is untouched, so §5 holds. Pinned byTestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty; the A5 seam test was corrected too, because its fixture has no unit and therefore exercises exactly this path while asserting the wrong sentence.It was the
observationsgate refusing the push that forced the re-read. The gate exists to stop findings dying in an overwritten REPORT.md; here it caught a live defect instead. -
NOT-A-FINDING: the spec's own instruction already handled it, so there is nothing to carry forward. §13 step 1 named
opengistas the data-less app; it is not one any more. Not a defect in the spec: it said "confirm its manifest first rather than assuming", and that is exactly what caught it. Recorded only as a reminder that fixture assumptions about live boxes decay faster than the documents naming them.
12. Two gates fired; one push was bypassed, and both are declared
felhom.eu — golden-currency CONVICTED. Three controller releases shipped today (0.224.0,
0.225.0, 0.226.0) and the vouched golden still carries 0.223.0, so a machine installed right now
receives none of them. The gate is right. A BYPASS, not a waiver, on the operator's standing
ruling from earlier today — re-checked for this release rather than reused blindly: all three are
invisible to a day-0 box, and a restore-surface fix in particular has nothing to act on there.
The ground expires the moment a release changes first-boot behaviour. Tracked on R-242; one
bake carrying 0.226.0 covers all three, then raise the floor to 0.226.0.
felhom-controller — nothing was bypassed. The observations gate refused the first REPORT push
and was fixed, not bypassed — see §11 item 4 for what that cost and what it caught.
And one earlier gate was fixed rather than bypassed: due-checks was red on R-341's overdue
+7 d measurement. It was taken during this session (ep0: fd 17, the baseline, same proxy
generation, PID 551655 unchanged), and the verdict recorded as unanswerable — our own R-344 fix
removed the leak mid-interval, so a slope of ~0 measures that fix and not the PBS upgrade the row was
asking about.