Files
felhom-controller/REPORT.md
T
admin e4e0aa8f46 REPORT: v0.226.0 shipped, three of four fixes proven live on demo-hp
Records what was validated and, in equal detail, what was not.

PROVEN LIVE (demo-hp, endpoints the UI invokes, evidence copied off the box):
  R-353  "A(z) opengist: 1 adatkotet visszaallitva -- az alkalmazas ujraindult."
         read off the customer's own wizard page, with real counts 1/1 volumes
         and 0/0 databases and correctly no database clause.
  R-360  refused in the exact flag state that produced the bug, and the planted
         canary file survived -- the consequence, not the branch.
  R-358  a mode=unit restore wrote {"schema":1,...,"full":false} at mode 0600
         with no .tmp left, and the gate logged place-to-live closed.

NOT live-validated, and each says why rather than being omitted:
  R-357  filling a real filesystem is a drill step, not a build step.
  R-353 Scenario B  NO app on demo-hp still has a data-less unit -- the spec
         named opengist from 21 August and it has since been recaptured (now
         1 volume dump). Manufacturing one means falsifying a manifest, which is
         the hand-set-state shortcut this project forbids.
  R-353 Scenario C and R-358's failed-download branch: unit-tested only.

Also recorded, because a near-miss that is quietly fixed teaches nobody: the
first B1 red-proof exposed a HOLLOW TEST OF MY OWN. With the gate removed the
run refused earlier, at the placement stat pre-pass, so `stops == 0` passed
against the pre-fix code. Fixture corrected and assertions reordered; only then
does the red-proof print THE APP WAS STOPPED (1 call(s)).

Register 165 -> 167 -> 161. Six rows compressed into CLOSED-ITEMS keeping title,
version, evidence and every sentence stating a rule; full original at
`git show e027b5d9`. No open row touched. ROADMAP not edited -- none of these
four ever had a row there, stated rather than silently skipped.

Teardown: this run provisioned nothing, across all three layers. Two throwaway
scripts and one canary directory were planted in guest 9201 and both removed.
2026-08-30 19:41:48 +02:00

15 KiB
Raw Blame History

REPORT — R-353 / R-357 / R-358 / R-360: the restore tells the truth

Controller v0.226.0 · 2026-08-30 · implemented on DooPlex, validated live on demo-hp


1. Confirmed baselines used — BOTH HAD MOVED

repo spec baseline actual at start version
felhom-controller f8c9390 / v0.223.0 → v0.224.0 e5eee50 / v0.225.0 → v0.226.0
felhom.eu c2c1fb4 ac6ac03 docs only
felhom-agent not touched v0.130.0 unchanged

Both targets were consumed earlier the same day by my own work — v0.224.0 (R-330) and v0.225.0 (R-331). The drift was re-confirmed against live Gitea before the first edit and the operator authorised proceeding. Every symbol in the spec's §5 table was re-verified present at the real baseline before editing; all 20 resolved, and almost every line landmark still matched. MinAgent stays 0.129.0. Both trees were clean and equal to origin/main at the start.

2. Files created / modified

felhom-controller (all paths under repo root):

file change
controller/internal/backup/restore.go restoreDockerVolumes → (int, error); caller updated
controller/internal/backup/restore_unit.go UnitRestoreResult; RestoreFromRecoveryUnit → (UnitRestoreResult, error) on every return path
controller/internal/web/handlers.go unitRestoreOutcomeMsg + 3 message constants; handler publishes the outcome
controller/internal/backup/offbox_reconstitute.go the R-357 headroom gate
controller/internal/backup/offbox_restore.go scratch marker (write/clear/read), OffboxFullScratchReady rewritten, shared refusal constants, SetOffboxLatestSnapshotFn, WriteScratchMarkerForTest
controller/internal/backup/backup.go offboxLatestSnapFn field
controller/internal/web/offbox_handlers.go server-side scratch refusals in 2 handlers; R-360 delete guard; corrected doc comment
controller/internal/backup/r357_reconstitute_headroom_test.go new
controller/internal/backup/r358_scratch_marker_test.go new
controller/internal/web/r353_unit_outcome_test.go new
controller/internal/web/r358_r360_handlers_test.go new
3 existing *_test.go in internal/backup mechanical _, err := for the changed signature
CHANGELOG.md, CONTEXT.md, REUSE.md, controller/README.md, REPORT.md docs

felhom.eu: STATUS.md, documentation/backlog/OPEN-ITEMS.md, documentation/backlog/CLOSED-ITEMS.md, documentation/architecture/07-backup-architecture.md, documentation/architecture/00-capability-map.md, documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt (new).

3. Commits pushed to main

repo commit contents
felhom-controller b8af7276 all four fixes, all tests, controller docs
felhom.eu e027b5d9 register closures, R-395 fix, architecture, capability map, evidence
felhom.eu (this report's commit) housekeeping compression + REPORT

4. Tests and red-proofs

All pass. New tests, by group:

test result
TestUnitRestoreOutcome_VolumesAndDatabaseNamed (A1) pass
TestUnitRestoreOutcome_BackupHeldOnlySettings (A2) pass
TestUnitRestoreOutcome_ManifestListedDataThatDidNotReturn (A3) pass
TestUnitRestoreOutcome_DatabaseOnly (A4) pass
TestR353_HandlerPublishesTheOutcome (A5, the seam test) pass
TestR357_DestructiveRestoreRefusesWithoutHeadroom (B1) pass
TestR357_UnknownSizeFailsClosed (B2) pass
TestR357_UnknownFreeSpaceFailsClosed (added — the mirror hole) pass
TestR357_AmpleSpaceIsUnchanged (B3) pass
TestR358_FailedRestoreLeavesNoUsableScratch (C1) pass
TestR358_UnitOnlyScratchIsNotFullReady (C2) pass
TestR358_StaleMarkerIsClearedBeforeTheRun (C3) pass
TestR358_UnreadableMarkerFailsClosed (C4) pass
TestR358_WrongSchemaFailsClosed (added) pass
TestR358_CompletedFullScratchStillReady (C5) pass
TestR358_MarkerIsWrittenAt0600AndAtomically (added) pass
TestR358_MarkerIsNeverPlaced (D3) pass
TestR358_MarkerIsClearedBeforeResticAndWrittenAfter (added, AST) pass
TestR358_PlaceHandlerRefusesIncompleteScratch (D1) pass
TestR358_ReconstituteHandlerRefusesIncompleteScratch (D2) pass
TestR358_UnitOnlyScratchClosesTheFullRestoreCard (Scenario F, flow level) pass
TestR360_VerifyCopyDeleteRefusedDuringRestore (D4) pass
TestR360_VerifyCopyDeleteStillWorksWhenIdle (Scenario H) pass

E1 — no existing test was modified for its content. Every TestReconstituteOutcome_* passes unmodified. The only test-file edits were mechanical call-site updates for the changed RestoreFromRecoveryUnit signature (err := → _, err :=) in three files.

Red-proofs — each mutated, observed failing, reverted, git diff clean

# mutation observed failure
A5 EndRestoreOp(true, stackName+" visszaállítva ("+snapshotID+").") restored THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)."
B1 the whole R-357 gate deleted THE APP WAS STOPPED (1 call(s)) for a restore with 300 KB free for a 1 MB copy … (err=<nil>) — and the same on both fail-closed tests
C1 the old len(entries) > 0 check restored a part-copy was reported READY, plus the unit-only and unreadable-marker cases
D4 the delete guard reverted to IsRunning() THE VERIFICATION COPY WAS DELETED while a restore was writing into it
D1/D2 both server-side scratch refusals removed both handlers redirected with „…elindult" over a part-copy

⚠ THE FIRST B1 RED-PROOF EXPOSED A HOLLOW TEST OF MY OWN, and it is recorded rather than quietly fixed. With the gate removed, the run refused earlier — at the placement stat pre-pass — so stops == 0 passed against the pre-fix code and the test proved nothing about the thing it exists for. Two corrections: the fixture now populates the scratch the way a completed download leaves it (the snapshot's own absolute paths mirrored under the scratch, plus SetSafetyDumpFn so no Docker is needed), and the assertions are reordered so a removed gate reports the outage rather than "no error returned". Only after that does the red-proof print THE APP WAS STOPPED (1 call(s)). The lesson is the doctrine's own: a test that cannot fail on the pre-fix shape is decoration.

Test count: 24 new tests added across 4 new files. Green gate: go build ./... && go vet ./... && go test ./... → 28 packages, rc 0, no failures. All 12 controller gates OK.

5. Deployed version

$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.226.0 Up 13 seconds (healthy)

demo-hp only. demo-felhom stays on 0.225.0 and the rest of the fleet on 0.223.0 (the floor).

6. Live validation — endpoint level, on demo-hp

Method: the exact endpoints the UI invokes, driven from inside guest 9201 (no browser on DooPlex; the residual is client-side rendering only). No state was hand-set. Evidence copied off the box at the end of the phase: felhom.eu/documentation/audits/evidence-r353-r360-live-2026-08-30/.

The verbatim messages

R-353 — POST /backup/restore for opengist, then the sentence read off the customer's own wizard page (GET /backups/restore/app?name=opengist):

A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.

and the controller's own lines:

[INFO] [backup] Restore-from-unit completed: opengist — 1 volume(s) of 1 listed, 0 database(s) of 0 listed
[INFO] [web]    Restore completed (async): stack=opengist in 9.112668936s (volumes 1/1, dbs 0/0)

This is Scenario A, not Scenario B, and the substitution is stated rather than glossed. The spec named opengist as the data-less app from 21 August. It is not one any more — checked before relying on it, as the spec instructed: every unit on demo-hp today lists at least one volume dump (opengist 1/0, privatebin 1/0, calibre-web 1/0, the rest 2–3 volumes plus a database). No app on the box has the Scenario B shape, and manufacturing one would mean falsifying a manifest — the hand-set-state shortcut this project forbids. Scenario B is carried by TestUnitRestoreOutcome_BackupHeldOnlySettings and the A5 seam test. What the live run does prove is the whole path: real counts, correct clause selection (no database clause for 0 databases), and the sentence reaching the customer's page.

R-360 — a mode=unit scratch restore started, and the verification-copy delete POSTed while it ran. The refusal, urldecoded from the Location header:

/backups/restore?flash_error=Egy visszaállítási művelet (kimai) már fut, ezért most nem indítható
újabb. Az állapotát ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást.
[WARN] [web] verification-copy delete refused for kimai: a backup/restore op is running

And the consequence, which is the assertion that matters: a canary file planted in the copy was still there afterwards — -rw-r--r-- 1 root root 10 Aug 30 17:34 canary.txt, contents defend-me.

R-358 / R-396 — after the mode=unit restore completed:

$ cat …/backups/offsite-restore/kimai/.felhom-restore-complete.json
{"schema":1,"snapshot_id":"84542ec8","full":false,"finished_at":"2026-08-30T17:34:42Z"}
mode=600        (no .tmp left behind)

[INFO] [offbox] kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy,
                so place-to-live stays closed

This is Scenario F proven live, and it is the case that pre-fix would have unlocked the destructive restore.

7. NOT yet live-validated — awaiting a supervised drill

  1. R-357's disk-full behaviour. Filling a real filesystem to prove it is a drill step, not a build step. Carried by the seam tests, whose central assertion is StopStack call count 0.
  2. R-353's Scenario B (the "backup held only settings" sentence) — no app on demo-hp has that unit shape any more; see §6.
  3. R-353's Scenario C (the unit lists dumps, none return) — the R-367 stranded-dump shape was not reproduced on live hardware; unit-tested only.
  4. R-358's failed-download branch was not induced live. It was not needed: the unit-only branch exercises the same marker gate through a real restore, without pointing restic at a bad snapshot.

8. Teardown

This run provisioned nothing. No machine, no guest, no hub record — all three layers:

  • Layer 1 (host): nothing created on felhom-pve or demo-hp. No VM, no CT, no storage entry.
  • Layer 2 (guest): two throwaway shell scripts were pushed into guest 9201 to drive the endpoints and both were removed; one canary directory was planted under backups/offsite-restore/kimai on the system drive to prove R-360 and was removed. The mode=unit scratch restore left a normal verification copy on the HDD, which is ordinary product state a customer can delete.
  • Layer 3 (hub): no customer, appliance or host record created, so none to discard.

demo-felhom, ep0, DooPlex and Peti's box were not touched.

9. Register

Size: 165 rows before → 167 after filing → 161 after housekeeping.

  • Closed: R-353, R-357, R-358, R-360 — each with shipping version and evidence path.
  • Filed and closed in the same session: R-396 (Scenario F's answer — see §11) and R-395 (STATUS.md contradicting itself; fixed, not merely recorded).
  • Re-ranked: none.
  • Housekeeping: the six closed rows were compressed into CLOSED-ITEMS.md keeping title, version, evidence and every sentence that states a rule; the full original is git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md. No open row was touched.
  • ROADMAP.md was not edited: none of these four ever had a ROADMAP row — they live in the register only. Stated rather than silently skipped.

10. Documentation coupling (§5.5)

  • 00-capability-map.md — one new row, splitting the claim: R-353/R-358/R-360 PROVEN-LIVE with the evidence citation; R-357 IMPLEMENTED only, with the reason.
  • 07-backup-architecture.md — §10.2 gained five rows (the four plus R-396). §8 matrix row 3 KEEPS its PROVEN status, with a note recording why: R-353 was a defect in the message, not the mechanism.
  • OPEN-ITEMS.md / CLOSED-ITEMS.md / STATUS.md — as §9.
  • ROADMAP.md — no applicable rows.
  • Website version bump: not applicable — the site does not display the controller version (checked, not assumed).

11. Observations — noticed, documented, not acted on

  1. Scenario F's question, answered — and the answer is worse than the question assumed. Filed as R-396. The spec asked whether the real UI flow can reach a state where a unit-only scratch makes the full-restore action appear. It can, by the safest-looking action on the page. „Ellenőrző visszaállítás" (mode=unit, the default, advertised as non-destructive) calls RestoreOffboxScratch(full=false); offboxRestoreScratchDir ignores full, so both modes write the same directory, and --include limits what restic extracts, never where; the wizard sets ScratchReady from OffboxFullScratchReady; deriveWizardStep derives both PlaceEnabled and RestoreEnabled from that one flag. R-358 as filed assumed the bad state needed a failed download. It needed only a successful safe one. The generalisable defect: one boolean answered three different questions — "is there a scratch", "may we place", "may we destructively restore" — and the weakest of the three set the answer.
  2. NotifyIntegrityOK / NotifyIntegrityFailed are dead code (noticed during R-331 earlier today, restated here because it is restore-adjacent): they exist and are called from nowhere; the controller runs no integrity check at all.
  3. resticStep is not a seam, which is why R-358's ordering property needed an AST test rather than an execution test. If a future task needs to drive restic paths under test, that is the seam to add — and it should be added deliberately, not improvised inside a bugfix.
  4. RestoreApp's own restoreDockerVolumes count is still discarded. Left alone deliberately: the spec scoped RestoreApp out, and it is the no-unit fallback whose result is already reported as a zero UnitRestoreResult — which is honest, because a box with no recovery unit really did return no data from one. Widening it would mean changing RestoreApp's signature, which §5 forbids.
  5. The spec's own §13 step 1 named an app whose shape has since changed. Not a defect in the spec — a reminder that fixture assumptions about live boxes decay, and the instruction to "confirm its manifest first rather than assuming" is what caught it.