Files
felhom-controller/REPORT.md
T
admin c0c8fe67bf
gates / gates (push) Failing after 13s
An unknown drawn as a zero: the defect v0.226.0's own fix introduced
Writing the REPORT's observation "the no-unit fallback already reports a zero
result, which is honest" exposed that the sentence was FALSE.

A zero UnitRestoreResult is Scenario B's shape. So RestoreFromRecoveryUnit's
fallback to RestoreApp -- which returns only an error, and whose signature is
deliberately out of scope -- would have printed "ez a mentes csak a
beallitasokat tartalmazta, adatot nem" over a restore that may have replayed the
app's entire dataset. That is an unknown drawn as a zero: the exact R-88 failure
direction this whole change exists to remove, re-introduced by the change.

UnitRestoreResult now carries CountsUnknown, the fallback sets it, and there is a
fourth sentence claiming only what is known -- the restore ran, the app is back,
and we cannot say what came back. RestoreApp's signature is untouched.

Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty. The A5 seam
test was corrected too: its fixture has no recovery unit, so it exercises exactly
this path and had been asserting the wrong sentence -- it now asserts the
unknown, which is what pins the fallback to it.

IT WAS THE observations GATE REFUSING THE PUSH THAT FORCED THE RE-READ. A gate
written to stop findings dying in an overwritten REPORT.md caught a live defect
instead. Also files R-397 (NotifyIntegrityOK/Failed are dead code AND the
monitoring page advertises a weekly integrity check that does not exist) and
R-398 (resticStep is not a seam, which is why R-358's ordering needed an AST
test) rather than leaving them in a file that is overwritten every session.

REPORT.md is the full run record: baselines re-confirmed, per-test results, the
five red-proofs with their observed output, the live validation with verbatim
Hungarian messages, what was NOT validated and why, teardown across three
layers, and the register 165 -> 167 -> 161.

Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
2026-08-30 19:51:16 +02:00

18 KiB
Raw Blame History

REPORT — R-353 / R-357 / R-358 / R-360: the restore tells the truth

Controller v0.226.0 · 2026-08-30 · implemented on DooPlex, validated live on demo-hp


1. Confirmed baselines used — BOTH HAD MOVED

repo spec baseline actual at start version
felhom-controller f8c9390 / v0.223.0 → v0.224.0 e5eee50 / v0.225.0 → v0.226.0
felhom.eu c2c1fb4 ac6ac03 docs only
felhom-agent not touched v0.130.0 unchanged

Both targets were consumed earlier the same day by my own work — v0.224.0 (R-330) and v0.225.0 (R-331). The drift was re-confirmed against live Gitea before the first edit and the operator authorised proceeding. Every symbol in the spec's §5 table was re-verified present at the real baseline before editing; all 20 resolved, and almost every line landmark still matched. MinAgent stays 0.129.0. Both trees were clean and equal to origin/main at the start.

2. Files created / modified

felhom-controller (all paths under repo root):

file change
controller/internal/backup/restore.go restoreDockerVolumes → (int, error); caller updated
controller/internal/backup/restore_unit.go UnitRestoreResult; RestoreFromRecoveryUnit → (UnitRestoreResult, error) on every return path
controller/internal/web/handlers.go unitRestoreOutcomeMsg + 3 message constants; handler publishes the outcome
controller/internal/backup/offbox_reconstitute.go the R-357 headroom gate
controller/internal/backup/offbox_restore.go scratch marker (write/clear/read), OffboxFullScratchReady rewritten, shared refusal constants, SetOffboxLatestSnapshotFn, WriteScratchMarkerForTest
controller/internal/backup/backup.go offboxLatestSnapFn field
controller/internal/web/offbox_handlers.go server-side scratch refusals in 2 handlers; R-360 delete guard; corrected doc comment
controller/internal/backup/r357_reconstitute_headroom_test.go new
controller/internal/backup/r358_scratch_marker_test.go new
controller/internal/web/r353_unit_outcome_test.go new
controller/internal/web/r358_r360_handlers_test.go new
3 existing *_test.go in internal/backup mechanical _, err := for the changed signature
CHANGELOG.md, CONTEXT.md, REUSE.md, controller/README.md, REPORT.md docs

felhom.eu: STATUS.md, documentation/backlog/OPEN-ITEMS.md, documentation/backlog/CLOSED-ITEMS.md, documentation/architecture/07-backup-architecture.md, documentation/architecture/00-capability-map.md, documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt (new).

3. Commits pushed to main

repo commit contents
felhom-controller b8af7276 all four fixes, all tests, controller docs
felhom.eu e027b5d9 register closures, R-395 fix, architecture, capability map, evidence
felhom.eu (this report's commit) housekeeping compression + REPORT

4. Tests and red-proofs

All pass. New tests, by group:

test result
TestUnitRestoreOutcome_VolumesAndDatabaseNamed (A1) pass
TestUnitRestoreOutcome_BackupHeldOnlySettings (A2) pass
TestUnitRestoreOutcome_ManifestListedDataThatDidNotReturn (A3) pass
TestUnitRestoreOutcome_DatabaseOnly (A4) pass
TestR353_HandlerPublishesTheOutcome (A5, the seam test) pass
TestR357_DestructiveRestoreRefusesWithoutHeadroom (B1) pass
TestR357_UnknownSizeFailsClosed (B2) pass
TestR357_UnknownFreeSpaceFailsClosed (added — the mirror hole) pass
TestR357_AmpleSpaceIsUnchanged (B3) pass
TestR358_FailedRestoreLeavesNoUsableScratch (C1) pass
TestR358_UnitOnlyScratchIsNotFullReady (C2) pass
TestR358_StaleMarkerIsClearedBeforeTheRun (C3) pass
TestR358_UnreadableMarkerFailsClosed (C4) pass
TestR358_WrongSchemaFailsClosed (added) pass
TestR358_CompletedFullScratchStillReady (C5) pass
TestR358_MarkerIsWrittenAt0600AndAtomically (added) pass
TestR358_MarkerIsNeverPlaced (D3) pass
TestR358_MarkerIsClearedBeforeResticAndWrittenAfter (added, AST) pass
TestR358_PlaceHandlerRefusesIncompleteScratch (D1) pass
TestR358_ReconstituteHandlerRefusesIncompleteScratch (D2) pass
TestR358_UnitOnlyScratchClosesTheFullRestoreCard (Scenario F, flow level) pass
TestR360_VerifyCopyDeleteRefusedDuringRestore (D4) pass
TestR360_VerifyCopyDeleteStillWorksWhenIdle (Scenario H) pass

E1 — no existing test was modified for its content. Every TestReconstituteOutcome_* passes unmodified. The only test-file edits were mechanical call-site updates for the changed RestoreFromRecoveryUnit signature (err := → _, err :=) in three files.

Red-proofs — each mutated, observed failing, reverted, git diff clean

# mutation observed failure
A5 EndRestoreOp(true, stackName+" visszaállítva ("+snapshotID+").") restored THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)."
B1 the whole R-357 gate deleted THE APP WAS STOPPED (1 call(s)) for a restore with 300 KB free for a 1 MB copy … (err=<nil>) — and the same on both fail-closed tests
C1 the old len(entries) > 0 check restored a part-copy was reported READY, plus the unit-only and unreadable-marker cases
D4 the delete guard reverted to IsRunning() THE VERIFICATION COPY WAS DELETED while a restore was writing into it
D1/D2 both server-side scratch refusals removed both handlers redirected with „…elindult" over a part-copy

⚠ THE FIRST B1 RED-PROOF EXPOSED A HOLLOW TEST OF MY OWN, and it is recorded rather than quietly fixed. With the gate removed, the run refused earlier — at the placement stat pre-pass — so stops == 0 passed against the pre-fix code and the test proved nothing about the thing it exists for. Two corrections: the fixture now populates the scratch the way a completed download leaves it (the snapshot's own absolute paths mirrored under the scratch, plus SetSafetyDumpFn so no Docker is needed), and the assertions are reordered so a removed gate reports the outage rather than "no error returned". Only after that does the red-proof print THE APP WAS STOPPED (1 call(s)). The lesson is the doctrine's own: a test that cannot fail on the pre-fix shape is decoration.

Test count: 24 new tests added across 4 new files. Green gate: go build ./... && go vet ./... && go test ./... → 28 packages, rc 0, no failures. All 12 controller gates OK.

5. Deployed version

$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.226.0 Up 13 seconds (healthy)

demo-hp only. demo-felhom stays on 0.225.0 and the rest of the fleet on 0.223.0 (the floor).

6. Live validation — endpoint level, on demo-hp

Method: the exact endpoints the UI invokes, driven from inside guest 9201 (no browser on DooPlex; the residual is client-side rendering only). No state was hand-set. Evidence copied off the box at the end of the phase: felhom.eu/documentation/audits/evidence-r353-r360-live-2026-08-30/.

The verbatim messages

R-353 — POST /backup/restore for opengist, then the sentence read off the customer's own wizard page (GET /backups/restore/app?name=opengist):

A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.

and the controller's own lines:

[INFO] [backup] Restore-from-unit completed: opengist — 1 volume(s) of 1 listed, 0 database(s) of 0 listed
[INFO] [web]    Restore completed (async): stack=opengist in 9.112668936s (volumes 1/1, dbs 0/0)

This is Scenario A, not Scenario B, and the substitution is stated rather than glossed. The spec named opengist as the data-less app from 21 August. It is not one any more — checked before relying on it, as the spec instructed: every unit on demo-hp today lists at least one volume dump (opengist 1/0, privatebin 1/0, calibre-web 1/0, the rest 2–3 volumes plus a database). No app on the box has the Scenario B shape, and manufacturing one would mean falsifying a manifest — the hand-set-state shortcut this project forbids. Scenario B is carried by TestUnitRestoreOutcome_BackupHeldOnlySettings and the A5 seam test. What the live run does prove is the whole path: real counts, correct clause selection (no database clause for 0 databases), and the sentence reaching the customer's page.

R-360 — a mode=unit scratch restore started, and the verification-copy delete POSTed while it ran. The refusal, urldecoded from the Location header:

/backups/restore?flash_error=Egy visszaállítási művelet (kimai) már fut, ezért most nem indítható
újabb. Az állapotát ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást.
[WARN] [web] verification-copy delete refused for kimai: a backup/restore op is running

And the consequence, which is the assertion that matters: a canary file planted in the copy was still there afterwards — -rw-r--r-- 1 root root 10 Aug 30 17:34 canary.txt, contents defend-me.

R-358 / R-396 — after the mode=unit restore completed:

$ cat …/backups/offsite-restore/kimai/.felhom-restore-complete.json
{"schema":1,"snapshot_id":"84542ec8","full":false,"finished_at":"2026-08-30T17:34:42Z"}
mode=600        (no .tmp left behind)

[INFO] [offbox] kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy,
                so place-to-live stays closed

This is Scenario F proven live, and it is the case that pre-fix would have unlocked the destructive restore.

7. NOT yet live-validated — awaiting a supervised drill

  1. R-357's disk-full behaviour. Filling a real filesystem to prove it is a drill step, not a build step. Carried by the seam tests, whose central assertion is StopStack call count 0.
  2. R-353's Scenario B (the "backup held only settings" sentence) — no app on demo-hp has that unit shape any more; see §6.
  3. R-353's Scenario C (the unit lists dumps, none return) — the R-367 stranded-dump shape was not reproduced on live hardware; unit-tested only.
  4. R-358's failed-download branch was not induced live. It was not needed: the unit-only branch exercises the same marker gate through a real restore, without pointing restic at a bad snapshot.

8. Teardown

This run provisioned nothing. No machine, no guest, no hub record — all three layers:

  • Layer 1 (host): nothing created on felhom-pve or demo-hp. No VM, no CT, no storage entry.
  • Layer 2 (guest): two throwaway shell scripts were pushed into guest 9201 to drive the endpoints and both were removed; one canary directory was planted under backups/offsite-restore/kimai on the system drive to prove R-360 and was removed. The mode=unit scratch restore left a normal verification copy on the HDD, which is ordinary product state a customer can delete.
  • Layer 3 (hub): no customer, appliance or host record created, so none to discard.

demo-felhom, ep0, DooPlex and Peti's box were not touched.

9. Register

Size: 165 rows before → 167 after filing → 161 after housekeeping.

  • Closed: R-353, R-357, R-358, R-360 — each with shipping version and evidence path.
  • Filed and closed in the same session: R-396 (Scenario F's answer — see §11) and R-395 (STATUS.md contradicting itself; fixed, not merely recorded).
  • Re-ranked: none.
  • Housekeeping: the six closed rows were compressed into CLOSED-ITEMS.md keeping title, version, evidence and every sentence that states a rule; the full original is git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md. No open row was touched.
  • ROADMAP.md was not edited: none of these four ever had a ROADMAP row — they live in the register only. Stated rather than silently skipped.

10. Documentation coupling (§5.5)

  • 00-capability-map.md — one new row, splitting the claim: R-353/R-358/R-360 PROVEN-LIVE with the evidence citation; R-357 IMPLEMENTED only, with the reason.
  • 07-backup-architecture.md — §10.2 gained five rows (the four plus R-396). §8 matrix row 3 KEEPS its PROVEN status, with a note recording why: R-353 was a defect in the message, not the mechanism.
  • OPEN-ITEMS.md / CLOSED-ITEMS.md / STATUS.md — as §9.
  • ROADMAP.md — no applicable rows.
  • Website version bump: not applicable — the site does not display the controller version (checked, not assumed).

11. Observations — each with a register row, or a stated reason it needs none

  1. FILED: R-396. Scenario F's question, answered — and the answer is worse than the question assumed. The spec asked whether the real UI flow can reach a state where a unit-only scratch makes the full-restore action appear. It can, by the safest-looking action on the page. „Ellenőrző visszaállítás" (mode=unit, the default, advertised as non-destructive) calls RestoreOffboxScratch(full=false); offboxRestoreScratchDir ignores full, so both modes write the same directory, and --include limits what restic extracts, never where; the wizard sets ScratchReady from OffboxFullScratchReady; deriveWizardStep derives both PlaceEnabled and RestoreEnabled from that one flag. R-358 as filed assumed the bad state needed a failed download. It needed only a successful safe one. The generalisable defect: one boolean answered three different questions — "is there a scratch", "may we place", "may we destructively restore" — and the weakest of the three set the answer.

  2. FILED: R-397. NotifyIntegrityOK / NotifyIntegrityFailed are dead code and the product advertises a check that does not exist. Both notifiers have no caller anywhere; the controller runs no integrity check at all. The part that actively misleads: config.Monitoring.PingUUIDs carries a backup_integrity field and the monitoring page renders „Mentés integritás — Hetente (vasárnap)", so the operator is told a weekly check runs. Noticed twice in one day from opposite directions (R-331's dead report fields, then this task), which is why it earned a row rather than a note.

  3. FILED: R-398. resticStep is not a seam, so no test can drive any restic-backed path. Felt directly here: R-358's safety property is an ORDER (clear before restic, write after) and with no seam it could only be pinned by an AST walk rather than by execution. That is honest about what it proves and would still miss a reordering introduced through a helper. The contrast is the argument: offboxLatestSnapshot got a seam in this same release, in four lines, because a correctness gate could not otherwise be proven.

  4. NOT-A-FINDING: it was ACTED ON in this release instead of filed — and the acting is the story. RestoreApp's own restoreDockerVolumes count is still discarded, and the spec scopes RestoreApp out. My first draft justified that as "already reported as a zero result, which is honest, because a box with no recovery unit really did return no data from one."

    ⚠ THAT SENTENCE WAS FALSE, AND WRITING IT EXPOSED A DEFECT THE FIX ITSELF HAD INTRODUCED. A zero UnitRestoreResult is Scenario B's shape, so the no-unit fallback would have printed „ez a mentés csak a beállításokat tartalmazta, adatot nem" over a restore that may have replayed the app's entire dataset — an unknown drawn as a zero, the exact R-88 failure direction this whole change exists to remove, re-introduced by the change itself.

    Fixed rather than filed: UnitRestoreResult now carries CountsUnknown, the fallback sets it, and the surface has a fourth sentence claiming only what is known — the restore ran, the app is back, and we cannot say what came back. RestoreApp's signature is untouched, so §5 holds. Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty; the A5 seam test was corrected too, because its fixture has no unit and therefore exercises exactly this path while asserting the wrong sentence.

    It was the observations gate refusing the push that forced the re-read. The gate exists to stop findings dying in an overwritten REPORT.md; here it caught a live defect instead.

  5. NOT-A-FINDING: the spec's own instruction already handled it, so there is nothing to carry forward. §13 step 1 named opengist as the data-less app; it is not one any more. Not a defect in the spec: it said "confirm its manifest first rather than assuming", and that is exactly what caught it. Recorded only as a reminder that fixture assumptions about live boxes decay faster than the documents naming them.

12. Two gates fired; one push was bypassed, and both are declared

felhom.eu — golden-currency CONVICTED. Three controller releases shipped today (0.224.0, 0.225.0, 0.226.0) and the vouched golden still carries 0.223.0, so a machine installed right now receives none of them. The gate is right. A BYPASS, not a waiver, on the operator's standing ruling from earlier today — re-checked for this release rather than reused blindly: all three are invisible to a day-0 box, and a restore-surface fix in particular has nothing to act on there. The ground expires the moment a release changes first-boot behaviour. Tracked on R-242; one bake carrying 0.226.0 covers all three, then raise the floor to 0.226.0.

felhom-controller — nothing was bypassed. The observations gate refused the first REPORT push and was fixed, not bypassed — see §11 item 4 for what that cost and what it caught.

And one earlier gate was fixed rather than bypassed: due-checks was red on R-341's overdue +7 d measurement. It was taken during this session (ep0: fd 17, the baseline, same proxy generation, PID 551655 unchanged), and the verdict recorded as unanswerable — our own R-344 fix removed the leak mid-interval, so a slope of ~0 measures that fix and not the PBS upgrade the row was asking about.