Records what was validated and, in equal detail, what was not.
PROVEN LIVE (demo-hp, endpoints the UI invokes, evidence copied off the box):
R-353 "A(z) opengist: 1 adatkotet visszaallitva -- az alkalmazas ujraindult."
read off the customer's own wizard page, with real counts 1/1 volumes
and 0/0 databases and correctly no database clause.
R-360 refused in the exact flag state that produced the bug, and the planted
canary file survived -- the consequence, not the branch.
R-358 a mode=unit restore wrote {"schema":1,...,"full":false} at mode 0600
with no .tmp left, and the gate logged place-to-live closed.
NOT live-validated, and each says why rather than being omitted:
R-357 filling a real filesystem is a drill step, not a build step.
R-353 Scenario B NO app on demo-hp still has a data-less unit -- the spec
named opengist from 21 August and it has since been recaptured (now
1 volume dump). Manufacturing one means falsifying a manifest, which is
the hand-set-state shortcut this project forbids.
R-353 Scenario C and R-358's failed-download branch: unit-tested only.
Also recorded, because a near-miss that is quietly fixed teaches nobody: the
first B1 red-proof exposed a HOLLOW TEST OF MY OWN. With the gate removed the
run refused earlier, at the placement stat pre-pass, so `stops == 0` passed
against the pre-fix code. Fixture corrected and assertions reordered; only then
does the red-proof print THE APP WAS STOPPED (1 call(s)).
Register 165 -> 167 -> 161. Six rows compressed into CLOSED-ITEMS keeping title,
version, evidence and every sentence stating a rule; full original at
`git show e027b5d9`. No open row touched. ROADMAP not edited -- none of these
four ever had a row there, stated rather than silently skipped.
Teardown: this run provisioned nothing, across all three layers. Two throwaway
scripts and one canary directory were planted in guest 9201 and both removed.
15 KiB
REPORT — R-353 / R-357 / R-358 / R-360: the restore tells the truth
Controller v0.226.0 · 2026-08-30 · implemented on DooPlex, validated live on demo-hp
1. Confirmed baselines used — BOTH HAD MOVED
| repo | spec baseline | actual at start | version |
|---|---|---|---|
felhom-controller |
f8c9390 / v0.223.0 → v0.224.0 |
e5eee50 / v0.225.0 |
→ v0.226.0 |
felhom.eu |
c2c1fb4 |
ac6ac03 |
docs only |
felhom-agent |
not touched | v0.130.0 | unchanged |
Both targets were consumed earlier the same day by my own work — v0.224.0 (R-330) and v0.225.0
(R-331). The drift was re-confirmed against live Gitea before the first edit and the operator
authorised proceeding. Every symbol in the spec's §5 table was re-verified present at the real
baseline before editing; all 20 resolved, and almost every line landmark still matched. MinAgent
stays 0.129.0. Both trees were clean and equal to origin/main at the start.
2. Files created / modified
felhom-controller (all paths under repo root):
| file | change |
|---|---|
controller/internal/backup/restore.go |
restoreDockerVolumes → (int, error); caller updated |
controller/internal/backup/restore_unit.go |
UnitRestoreResult; RestoreFromRecoveryUnit → (UnitRestoreResult, error) on every return path |
controller/internal/web/handlers.go |
unitRestoreOutcomeMsg + 3 message constants; handler publishes the outcome |
controller/internal/backup/offbox_reconstitute.go |
the R-357 headroom gate |
controller/internal/backup/offbox_restore.go |
scratch marker (write/clear/read), OffboxFullScratchReady rewritten, shared refusal constants, SetOffboxLatestSnapshotFn, WriteScratchMarkerForTest |
controller/internal/backup/backup.go |
offboxLatestSnapFn field |
controller/internal/web/offbox_handlers.go |
server-side scratch refusals in 2 handlers; R-360 delete guard; corrected doc comment |
controller/internal/backup/r357_reconstitute_headroom_test.go |
new |
controller/internal/backup/r358_scratch_marker_test.go |
new |
controller/internal/web/r353_unit_outcome_test.go |
new |
controller/internal/web/r358_r360_handlers_test.go |
new |
3 existing *_test.go in internal/backup |
mechanical _, err := for the changed signature |
CHANGELOG.md, CONTEXT.md, REUSE.md, controller/README.md, REPORT.md |
docs |
felhom.eu: STATUS.md, documentation/backlog/OPEN-ITEMS.md,
documentation/backlog/CLOSED-ITEMS.md, documentation/architecture/07-backup-architecture.md,
documentation/architecture/00-capability-map.md,
documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt (new).
3. Commits pushed to main
| repo | commit | contents |
|---|---|---|
felhom-controller |
b8af7276 |
all four fixes, all tests, controller docs |
felhom.eu |
e027b5d9 |
register closures, R-395 fix, architecture, capability map, evidence |
felhom.eu |
(this report's commit) | housekeeping compression + REPORT |
4. Tests and red-proofs
All pass. New tests, by group:
| test | result |
|---|---|
TestUnitRestoreOutcome_VolumesAndDatabaseNamed (A1) |
pass |
TestUnitRestoreOutcome_BackupHeldOnlySettings (A2) |
pass |
TestUnitRestoreOutcome_ManifestListedDataThatDidNotReturn (A3) |
pass |
TestUnitRestoreOutcome_DatabaseOnly (A4) |
pass |
TestR353_HandlerPublishesTheOutcome (A5, the seam test) |
pass |
TestR357_DestructiveRestoreRefusesWithoutHeadroom (B1) |
pass |
TestR357_UnknownSizeFailsClosed (B2) |
pass |
TestR357_UnknownFreeSpaceFailsClosed (added — the mirror hole) |
pass |
TestR357_AmpleSpaceIsUnchanged (B3) |
pass |
TestR358_FailedRestoreLeavesNoUsableScratch (C1) |
pass |
TestR358_UnitOnlyScratchIsNotFullReady (C2) |
pass |
TestR358_StaleMarkerIsClearedBeforeTheRun (C3) |
pass |
TestR358_UnreadableMarkerFailsClosed (C4) |
pass |
TestR358_WrongSchemaFailsClosed (added) |
pass |
TestR358_CompletedFullScratchStillReady (C5) |
pass |
TestR358_MarkerIsWrittenAt0600AndAtomically (added) |
pass |
TestR358_MarkerIsNeverPlaced (D3) |
pass |
TestR358_MarkerIsClearedBeforeResticAndWrittenAfter (added, AST) |
pass |
TestR358_PlaceHandlerRefusesIncompleteScratch (D1) |
pass |
TestR358_ReconstituteHandlerRefusesIncompleteScratch (D2) |
pass |
TestR358_UnitOnlyScratchClosesTheFullRestoreCard (Scenario F, flow level) |
pass |
TestR360_VerifyCopyDeleteRefusedDuringRestore (D4) |
pass |
TestR360_VerifyCopyDeleteStillWorksWhenIdle (Scenario H) |
pass |
E1 — no existing test was modified for its content. Every TestReconstituteOutcome_* passes
unmodified. The only test-file edits were mechanical call-site updates for the changed
RestoreFromRecoveryUnit signature (err := → _, err :=) in three files.
Red-proofs — each mutated, observed failing, reverted, git diff clean
| # | mutation | observed failure |
|---|---|---|
| A5 | EndRestoreOp(true, stackName+" visszaállítva ("+snapshotID+").") restored |
THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)." |
| B1 | the whole R-357 gate deleted | THE APP WAS STOPPED (1 call(s)) for a restore with 300 KB free for a 1 MB copy … (err=<nil>) — and the same on both fail-closed tests |
| C1 | the old len(entries) > 0 check restored |
a part-copy was reported READY, plus the unit-only and unreadable-marker cases |
| D4 | the delete guard reverted to IsRunning() |
THE VERIFICATION COPY WAS DELETED while a restore was writing into it |
| D1/D2 | both server-side scratch refusals removed | both handlers redirected with „…elindult" over a part-copy |
⚠ THE FIRST B1 RED-PROOF EXPOSED A HOLLOW TEST OF MY OWN, and it is recorded rather than quietly fixed. With the gate removed, the run refused earlier — at the placement stat pre-pass — so
stops == 0passed against the pre-fix code and the test proved nothing about the thing it exists for. Two corrections: the fixture now populates the scratch the way a completed download leaves it (the snapshot's own absolute paths mirrored under the scratch, plusSetSafetyDumpFnso no Docker is needed), and the assertions are reordered so a removed gate reports the outage rather than "no error returned". Only after that does the red-proof printTHE APP WAS STOPPED (1 call(s)). The lesson is the doctrine's own: a test that cannot fail on the pre-fix shape is decoration.
Test count: 24 new tests added across 4 new files. Green gate:
go build ./... && go vet ./... && go test ./... → 28 packages, rc 0, no failures.
All 12 controller gates OK.
5. Deployed version
$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.226.0 Up 13 seconds (healthy)
demo-hp only. demo-felhom stays on 0.225.0 and the rest of the fleet on 0.223.0 (the floor).
6. Live validation — endpoint level, on demo-hp
Method: the exact endpoints the UI invokes, driven from inside guest 9201 (no browser on
DooPlex; the residual is client-side rendering only). No state was hand-set. Evidence copied off the
box at the end of the phase: felhom.eu/documentation/audits/evidence-r353-r360-live-2026-08-30/.
The verbatim messages
R-353 — POST /backup/restore for opengist, then the sentence read off the customer's own
wizard page (GET /backups/restore/app?name=opengist):
A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.
and the controller's own lines:
[INFO] [backup] Restore-from-unit completed: opengist — 1 volume(s) of 1 listed, 0 database(s) of 0 listed
[INFO] [web] Restore completed (async): stack=opengist in 9.112668936s (volumes 1/1, dbs 0/0)
This is Scenario A, not Scenario B, and the substitution is stated rather than glossed. The spec
named opengist as the data-less app from 21 August. It is not one any more — checked before
relying on it, as the spec instructed: every unit on demo-hp today lists at least one volume dump
(opengist 1/0, privatebin 1/0, calibre-web 1/0, the rest 2–3 volumes plus a database). No app
on the box has the Scenario B shape, and manufacturing one would mean falsifying a manifest — the
hand-set-state shortcut this project forbids. Scenario B is carried by
TestUnitRestoreOutcome_BackupHeldOnlySettings and the A5 seam test. What the live run does prove
is the whole path: real counts, correct clause selection (no database clause for 0 databases), and
the sentence reaching the customer's page.
R-360 — a mode=unit scratch restore started, and the verification-copy delete POSTed while it
ran. The refusal, urldecoded from the Location header:
/backups/restore?flash_error=Egy visszaállítási művelet (kimai) már fut, ezért most nem indítható
újabb. Az állapotát ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást.
[WARN] [web] verification-copy delete refused for kimai: a backup/restore op is running
And the consequence, which is the assertion that matters: a canary file planted in the copy was
still there afterwards — -rw-r--r-- 1 root root 10 Aug 30 17:34 canary.txt, contents defend-me.
R-358 / R-396 — after the mode=unit restore completed:
$ cat …/backups/offsite-restore/kimai/.felhom-restore-complete.json
{"schema":1,"snapshot_id":"84542ec8","full":false,"finished_at":"2026-08-30T17:34:42Z"}
mode=600 (no .tmp left behind)
[INFO] [offbox] kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy,
so place-to-live stays closed
This is Scenario F proven live, and it is the case that pre-fix would have unlocked the destructive restore.
7. NOT yet live-validated — awaiting a supervised drill
- R-357's disk-full behaviour. Filling a real filesystem to prove it is a drill step, not a build
step. Carried by the seam tests, whose central assertion is
StopStackcall count 0. - R-353's Scenario B (the "backup held only settings" sentence) — no app on
demo-hphas that unit shape any more; see §6. - R-353's Scenario C (the unit lists dumps, none return) — the R-367 stranded-dump shape was not reproduced on live hardware; unit-tested only.
- R-358's failed-download branch was not induced live. It was not needed: the unit-only branch exercises the same marker gate through a real restore, without pointing restic at a bad snapshot.
8. Teardown
This run provisioned nothing. No machine, no guest, no hub record — all three layers:
- Layer 1 (host): nothing created on
felhom-pveordemo-hp. No VM, no CT, no storage entry. - Layer 2 (guest): two throwaway shell scripts were pushed into guest 9201 to drive the endpoints
and both were removed; one canary directory was planted under
backups/offsite-restore/kimaion the system drive to prove R-360 and was removed. Themode=unitscratch restore left a normal verification copy on the HDD, which is ordinary product state a customer can delete. - Layer 3 (hub): no customer, appliance or host record created, so none to discard.
demo-felhom, ep0, DooPlex and Peti's box were not touched.
9. Register
Size: 165 rows before → 167 after filing → 161 after housekeeping.
- Closed: R-353, R-357, R-358, R-360 — each with shipping version and evidence path.
- Filed and closed in the same session: R-396 (Scenario F's answer — see §11) and R-395
(
STATUS.mdcontradicting itself; fixed, not merely recorded). - Re-ranked: none.
- Housekeeping: the six closed rows were compressed into
CLOSED-ITEMS.mdkeeping title, version, evidence and every sentence that states a rule; the full original isgit show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md. No open row was touched. ROADMAP.mdwas not edited: none of these four ever had a ROADMAP row — they live in the register only. Stated rather than silently skipped.
10. Documentation coupling (§5.5)
00-capability-map.md— one new row, splitting the claim: R-353/R-358/R-360 PROVEN-LIVE with the evidence citation; R-357 IMPLEMENTED only, with the reason.07-backup-architecture.md— §10.2 gained five rows (the four plus R-396). §8 matrix row 3 KEEPS its PROVEN status, with a note recording why: R-353 was a defect in the message, not the mechanism.OPEN-ITEMS.md/CLOSED-ITEMS.md/STATUS.md— as §9.ROADMAP.md— no applicable rows.- Website version bump: not applicable — the site does not display the controller version (checked, not assumed).
11. Observations — noticed, documented, not acted on
- Scenario F's question, answered — and the answer is worse than the question assumed. Filed as
R-396. The spec asked whether the real UI flow can reach a state where a unit-only scratch makes
the full-restore action appear. It can, by the safest-looking action on the page.
„Ellenőrző visszaállítás" (
mode=unit, the default, advertised as non-destructive) callsRestoreOffboxScratch(full=false);offboxRestoreScratchDirignoresfull, so both modes write the same directory, and--includelimits what restic extracts, never where; the wizard setsScratchReadyfromOffboxFullScratchReady;deriveWizardStepderives bothPlaceEnabledandRestoreEnabledfrom that one flag. R-358 as filed assumed the bad state needed a failed download. It needed only a successful safe one. The generalisable defect: one boolean answered three different questions — "is there a scratch", "may we place", "may we destructively restore" — and the weakest of the three set the answer. NotifyIntegrityOK/NotifyIntegrityFailedare dead code (noticed during R-331 earlier today, restated here because it is restore-adjacent): they exist and are called from nowhere; the controller runs no integrity check at all.resticStepis not a seam, which is why R-358's ordering property needed an AST test rather than an execution test. If a future task needs to drive restic paths under test, that is the seam to add — and it should be added deliberately, not improvised inside a bugfix.RestoreApp's ownrestoreDockerVolumescount is still discarded. Left alone deliberately: the spec scopedRestoreAppout, and it is the no-unit fallback whose result is already reported as a zeroUnitRestoreResult— which is honest, because a box with no recovery unit really did return no data from one. Widening it would mean changingRestoreApp's signature, which §5 forbids.- The spec's own §13 step 1 named an app whose shape has since changed. Not a defect in the spec — a reminder that fixture assumptions about live boxes decay, and the instruction to "confirm its manifest first rather than assuming" is what caught it.