REPORT: v0.226.0 shipped, three of four fixes proven live on demo-hp

Records what was validated and, in equal detail, what was not.

PROVEN LIVE (demo-hp, endpoints the UI invokes, evidence copied off the box):
  R-353  "A(z) opengist: 1 adatkotet visszaallitva -- az alkalmazas ujraindult."
         read off the customer's own wizard page, with real counts 1/1 volumes
         and 0/0 databases and correctly no database clause.
  R-360  refused in the exact flag state that produced the bug, and the planted
         canary file survived -- the consequence, not the branch.
  R-358  a mode=unit restore wrote {"schema":1,...,"full":false} at mode 0600
         with no .tmp left, and the gate logged place-to-live closed.

NOT live-validated, and each says why rather than being omitted:
  R-357  filling a real filesystem is a drill step, not a build step.
  R-353 Scenario B  NO app on demo-hp still has a data-less unit -- the spec
         named opengist from 21 August and it has since been recaptured (now
         1 volume dump). Manufacturing one means falsifying a manifest, which is
         the hand-set-state shortcut this project forbids.
  R-353 Scenario C and R-358's failed-download branch: unit-tested only.

Also recorded, because a near-miss that is quietly fixed teaches nobody: the
first B1 red-proof exposed a HOLLOW TEST OF MY OWN. With the gate removed the
run refused earlier, at the placement stat pre-pass, so `stops == 0` passed
against the pre-fix code. Fixture corrected and assertions reordered; only then
does the red-proof print THE APP WAS STOPPED (1 call(s)).

Register 165 -> 167 -> 161. Six rows compressed into CLOSED-ITEMS keeping title,
version, evidence and every sentence stating a rule; full original at
`git show e027b5d9`. No open row touched. ROADMAP not edited -- none of these
four ever had a row there, stated rather than silently skipped.

Teardown: this run provisioned nothing, across all three layers. Two throwaway
scripts and one canary directory were planted in guest 9201 and both removed.
This commit is contained in:
2026-08-30 19:41:48 +02:00
parent b8af72764d
commit e4e0aa8f46
+219 -95
View File
@@ -1,132 +1,256 @@
# REPORT — R-331: the operator Backup card said every customer had no backups
# REPORT — R-353 / R-357 / R-358 / R-360: the restore tells the truth
**Controller v0.225.0 + hub v0.109.0 · 2026-08-30**
This is Option B of a two-part request. Option A (the nightly false `app_start_failed` e-mails) shipped
earlier today as controller v0.224.0 (R-330) and is recorded in this repo's `CHANGELOG.md`.
**Controller v0.226.0 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`**
---
## 1. What was wrong
## 1. Confirmed baselines used — BOTH HAD MOVED
The hub customer page's **Backup** card read, for **every customer, indefinitely**:
| repo | spec baseline | **actual at start** | version |
|---|---|---|---|
| `felhom-controller` | `f8c9390` / v0.223.0 → v0.224.0 | **`e5eee50` / v0.225.0** | → **v0.226.0** |
| `felhom.eu` | `c2c1fb4` | **`ac6ac03`** | docs only |
| `felhom-agent` | not touched | v0.130.0 | unchanged |
**Both targets were consumed earlier the same day by my own work** — v0.224.0 (R-330) and v0.225.0
(R-331). The drift was re-confirmed against live Gitea before the first edit and the operator
authorised proceeding. **Every symbol in the spec's §5 table was re-verified present at the real
baseline** before editing; all 20 resolved, and almost every line landmark still matched. `MinAgent`
stays **0.129.0**. Both trees were clean and equal to `origin/main` at the start.
## 2. Files created / modified
**`felhom-controller`** (all paths under repo root):
| file | change |
|---|---|
| `controller/internal/backup/restore.go` | `restoreDockerVolumes` → `(int, error)`; caller updated |
| `controller/internal/backup/restore_unit.go` | `UnitRestoreResult`; `RestoreFromRecoveryUnit` → `(UnitRestoreResult, error)` on every return path |
| `controller/internal/web/handlers.go` | `unitRestoreOutcomeMsg` + 3 message constants; handler publishes the outcome |
| `controller/internal/backup/offbox_reconstitute.go` | the R-357 headroom gate |
| `controller/internal/backup/offbox_restore.go` | scratch marker (write/clear/read), `OffboxFullScratchReady` rewritten, shared refusal constants, `SetOffboxLatestSnapshotFn`, `WriteScratchMarkerForTest` |
| `controller/internal/backup/backup.go` | `offboxLatestSnapFn` field |
| `controller/internal/web/offbox_handlers.go` | server-side scratch refusals in 2 handlers; R-360 delete guard; corrected doc comment |
| `controller/internal/backup/r357_reconstitute_headroom_test.go` | **new** |
| `controller/internal/backup/r358_scratch_marker_test.go` | **new** |
| `controller/internal/web/r353_unit_outcome_test.go` | **new** |
| `controller/internal/web/r358_r360_handlers_test.go` | **new** |
| 3 existing `*_test.go` in `internal/backup` | mechanical `_, err :=` for the changed signature |
| `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md` | docs |
**`felhom.eu`**: `STATUS.md`, `documentation/backlog/OPEN-ITEMS.md`,
`documentation/backlog/CLOSED-ITEMS.md`, `documentation/architecture/07-backup-architecture.md`,
`documentation/architecture/00-capability-map.md`,
`documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` (**new**).
## 3. Commits pushed to `main`
| repo | commit | contents |
|---|---|---|
| `felhom-controller` | **`b8af7276`** | all four fixes, all tests, controller docs |
| `felhom.eu` | **`e027b5d9`** | register closures, R-395 fix, architecture, capability map, evidence |
| `felhom.eu` | *(this report's commit)* | housekeeping compression + REPORT |
## 4. Tests and red-proofs
**All pass.** New tests, by group:
| test | result |
|---|---|
| `TestUnitRestoreOutcome_VolumesAndDatabaseNamed` (A1) | pass |
| `TestUnitRestoreOutcome_BackupHeldOnlySettings` (A2) | pass |
| `TestUnitRestoreOutcome_ManifestListedDataThatDidNotReturn` (A3) | pass |
| `TestUnitRestoreOutcome_DatabaseOnly` (A4) | pass |
| `TestR353_HandlerPublishesTheOutcome` (A5, the seam test) | pass |
| `TestR357_DestructiveRestoreRefusesWithoutHeadroom` (B1) | pass |
| `TestR357_UnknownSizeFailsClosed` (B2) | pass |
| `TestR357_UnknownFreeSpaceFailsClosed` (added — the mirror hole) | pass |
| `TestR357_AmpleSpaceIsUnchanged` (B3) | pass |
| `TestR358_FailedRestoreLeavesNoUsableScratch` (C1) | pass |
| `TestR358_UnitOnlyScratchIsNotFullReady` (C2) | pass |
| `TestR358_StaleMarkerIsClearedBeforeTheRun` (C3) | pass |
| `TestR358_UnreadableMarkerFailsClosed` (C4) | pass |
| `TestR358_WrongSchemaFailsClosed` (added) | pass |
| `TestR358_CompletedFullScratchStillReady` (C5) | pass |
| `TestR358_MarkerIsWrittenAt0600AndAtomically` (added) | pass |
| `TestR358_MarkerIsNeverPlaced` (D3) | pass |
| `TestR358_MarkerIsClearedBeforeResticAndWrittenAfter` (added, AST) | pass |
| `TestR358_PlaceHandlerRefusesIncompleteScratch` (D1) | pass |
| `TestR358_ReconstituteHandlerRefusesIncompleteScratch` (D2) | pass |
| `TestR358_UnitOnlyScratchClosesTheFullRestoreCard` (Scenario F, flow level) | pass |
| `TestR360_VerifyCopyDeleteRefusedDuringRestore` (D4) | pass |
| `TestR360_VerifyCopyDeleteStillWorksWhenIdle` (Scenario H) | pass |
**E1 — no existing test was modified for its content.** Every `TestReconstituteOutcome_*` passes
unmodified. The only test-file edits were mechanical call-site updates for the changed
`RestoreFromRecoveryUnit` signature (`err :=` → `_, err :=`) in three files.
### Red-proofs — each mutated, observed failing, reverted, `git diff` clean
| # | mutation | observed failure |
|---|---|---|
| **A5** | `EndRestoreOp(true, stackName+" visszaállítva ("+snapshotID+").")` restored | `THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)."` |
| **B1** | the whole R-357 gate deleted | `THE APP WAS STOPPED (1 call(s)) for a restore with 300 KB free for a 1 MB copy … (err=<nil>)` — and the same on both fail-closed tests |
| **C1** | the old `len(entries) > 0` check restored | `a part-copy was reported READY`, plus the unit-only and unreadable-marker cases |
| **D4** | the delete guard reverted to `IsRunning()` | `THE VERIFICATION COPY WAS DELETED while a restore was writing into it` |
| **D1/D2** | both server-side scratch refusals removed | both handlers redirected with „…elindult" over a part-copy |
> **⚠ THE FIRST B1 RED-PROOF EXPOSED A HOLLOW TEST OF MY OWN, and it is recorded rather than quietly
> fixed.** With the gate removed, the run refused *earlier* — at the placement stat pre-pass — so
> `stops == 0` passed against the pre-fix code and the test proved nothing about the thing it exists
> for. Two corrections: the fixture now populates the scratch the way a completed download leaves it
> (the snapshot's own absolute paths mirrored under the scratch, plus `SetSafetyDumpFn` so no Docker
> is needed), and the assertions are **reordered** so a removed gate reports the outage rather than
> "no error returned". Only after that does the red-proof print `THE APP WAS STOPPED (1 call(s))`.
> The lesson is the doctrine's own: a test that cannot fail on the pre-fix shape is decoration.
**Test count:** 24 new tests added across 4 new files. **Green gate:**
`go build ./... && go vet ./... && go test ./...` → **28 packages, rc 0**, no failures.
**All 12 controller gates OK.**
## 5. Deployed version
```
Enabled Yes Snapshots 0
Repo Size 0 MB Integrity Unknown
$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.226.0 Up 13 seconds (healthy)
```
Measured on `demo-hp` 2026-08-30, at which moment the truth was:
**`demo-hp` only.** `demo-felhom` stays on 0.225.0 and the rest of the fleet on 0.223.0 (the floor).
| source | value |
|---|---|
| the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` |
| that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` |
| the hub's **Offsite** page | `0.1 GB` used of a `50 GB` quota — read from the stored report |
## 6. Live validation — endpoint level, on `demo-hp`
**A card that reads "no backups" over a working backup is worse than no card.** It is the R-88
direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an
operator consults to answer "is this customer protected?".
**Method:** the exact endpoints the UI invokes, driven from inside guest 9201 (no browser on
DooPlex; the residual is client-side rendering only). No state was hand-set. Evidence copied off the
box at the end of the phase: `felhom.eu/documentation/audits/evidence-r353-r360-live-2026-08-30/`.
## 2. Root cause
### The verbatim messages
The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and
`integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice
8C) — `buildBackupReport` leaves them zero *deliberately* and says so in a comment. So the zeros were
not a bug in the controller; they were correct values for dead fields, being rendered as if live.
**R-353** — `POST /backup/restore` for `opengist`, then the sentence read off the customer's own
wizard page (`GET /backups/restore/app?name=opengist`):
**The data was never missing.** The live numbers ride in the report's **`offsite`** object
(`backup.OffboxReportStatus`), which the hub *already* reads for the Offsite page
(`offsiteUsageBytes`) and which `monitor.OffsiteChecker` *already* drives fill and staleness alarms
from. That the Offsite page rendered demo-hp's real usage from the same stored report, at the same
moment the Backup card said `0 MB`, is the proof the bytes were arriving.
```
A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.
```
So the fix is a **render change over an existing feed**, not a new pipeline.
and the controller's own lines:
## 3. Why it was not a one-line template swap
```
[INFO] [backup] Restore-from-unit completed: opengist — 1 volume(s) of 1 listed, 0 database(s) of 0 listed
[INFO] [web] Restore completed (async): stack=opengist in 9.112668936s (volumes 1/1, dbs 0/0)
```
`snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever
measured this repository*.
**This is Scenario A, not Scenario B, and the substitution is stated rather than glossed.** The spec
named `opengist` as the data-less app from 21 August. **It is not one any more** — checked before
relying on it, as the spec instructed: every unit on `demo-hp` today lists at least one volume dump
(`opengist 1/0`, `privatebin 1/0`, `calibre-web 1/0`, the rest 2–3 volumes plus a database). **No app
on the box has the Scenario B shape**, and manufacturing one would mean falsifying a manifest — the
hand-set-state shortcut this project forbids. Scenario B is carried by
`TestUnitRestoreOutcome_BackupHeldOnlySettings` and the A5 seam test. What the live run *does* prove
is the whole path: real counts, correct clause selection (no database clause for 0 databases), and
the sentence reaching the customer's page.
**R-225 already measured that exact confusion one layer down**: a rebuilt box rendered
„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and
`settings.OffboxTarget.StatsKnown` is what fixed it for the controller's own UI. It was never put on
the wire, so the hub was free to make the identical mistake one layer up — and rendering the count
without consulting it would have **moved R-225 to the hub instead of fixing anything**.
**R-360** — a `mode=unit` scratch restore started, and the verification-copy delete POSTed while it
ran. The refusal, urldecoded from the `Location` header:
## 4. What changed
```
/backups/restore?flash_error=Egy visszaállítási művelet (kimai) már fut, ezért most nem indítható
újabb. Az állapotát ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást.
```
**Controller (v0.225.0) — one field, and the dead ones labelled.**
```
[WARN] [web] verification-copy delete refused for kimai: a backup/restore op is running
```
- `OffboxReportStatus.StatsKnown` now forwards `settings.OffboxTarget.StatsKnown`. It is `omitempty`,
so a controller below this version sends no key, a reader sees `false`, and **false means "cannot
answer", never "the answer is zero"**. Absence is ignorance, not emptiness.
- The four dead `BackupReport` fields (`SnapshotCount`, `RepoSizeMB`, `LastIntegrityCheck`,
`IntegrityOK`) are **kept on the wire** so historical reports in the hub's store keep parsing, but
now carry a warning naming R-331 and pointing at `Offsite`. `TestBackupReport_DeadFieldsStayZero`
fails the moment a producer appears for any of them — the prompt to update the hub card in the
**same** change rather than shipping a field nothing renders.
**And the consequence, which is the assertion that matters:** a canary file planted in the copy was
still there afterwards — `-rw-r--r-- 1 root root 10 Aug 30 17:34 canary.txt`, contents `defend-me`.
**Hub (v0.109.0) — the card reads `offsite`, resolved in Go.**
**R-358 / R-396** — after the `mode=unit` restore completed:
`backup_card.go` builds a typed view, because the card's whole subject is a distinction a template
`{{if}}` chain over `map[string]interface{}` float64s cannot keep:
```
$ cat …/backups/offsite-restore/kimai/.felhom-restore-complete.json
{"schema":1,"snapshot_id":"84542ec8","full":false,"finished_at":"2026-08-30T17:34:42Z"}
mode=600 (no .tmp left behind)
| report state | card shows |
|---|---|
| no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** |
| `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" |
| enabled, `stats_known:false` | **&mdash;**, plus "never been measured". Never `0` |
| enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge |
[INFO] [offbox] kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy,
so place-to-live stays closed
```
**The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no
integrity check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from
nowhere**. A row that can only ever read "Unknown" is not information, and one that could read "OK"
from an unwritten field would be a lie.
**This is Scenario F proven live**, and it is the case that pre-fix would have unlocked the
destructive restore.
`fmtBytesAuto` is new rather than reusing the Offsite page's `fmtBytesGB`: that one is fixed at GB
because it renders against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` —
which on a card whose entire defect was under-reporting a real backup reads as "nearly nothing".
## 7. NOT yet live-validated — awaiting a supervised drill
## 5. Tests and red-proofs
1. **R-357's disk-full behaviour.** Filling a real filesystem to prove it is a drill step, not a build
step. Carried by the seam tests, whose central assertion is `StopStack` call count 0.
2. **R-353's Scenario B** (the "backup held only settings" sentence) — no app on `demo-hp` has that
unit shape any more; see §6.
3. **R-353's Scenario C** (the unit lists dumps, none return) — the R-367 stranded-dump shape was not
reproduced on live hardware; unit-tested only.
4. **R-358's failed-download branch** was not induced live. It was not needed: the unit-only branch
exercises the same marker gate through a real restore, without pointing restic at a bad snapshot.
The hub tests assert the **rendered page**, using demo-hp's real reported values, so a regression fails
against the same numbers the defect was measured against. The defect lived in the template's choice of
source object, so a test one layer below it would have been green against the shipped bug.
## 8. Teardown
| test | pins |
|---|---|
| `TestBackupCard_RendersTheRealSnapshotCount` (hub) | 67 and 134.3 MB are on the page; no `0 MB`; no Integrity row |
| `TestBackupCard_ThreeWayRuling` (hub) | unmeasured ≠ measured-empty ≠ no-object, all four branches |
| `TestBackupCard_OldControllerDegradesToUnknownNotEmpty` (hub) | upgrading the hub ahead of the fleet must not report every customer as empty |
| `TestFmtBytesAuto_...` (hub) | the real byte value renders as 134.3 MB, not 0.1 GB |
| `TestOffboxReportStatus_CarriesStatsKnown` (controller) | the field is on the wire |
| `TestOffboxReportStatus_MeasuredEmptyIsNotUnmeasured` (controller) | asserted on the **JSON**, not the struct — the hub sees bytes |
| `TestOffboxReportStatus_AbsentStatsKnownParsesAsUnknown` (controller) | the fail-safe direction |
| `TestBackupReport_DeadFieldsStayZero` (controller) | the dead fields stay dead |
**This run provisioned nothing.** No machine, no guest, no hub record — all three layers:
**Red-proofs, each printing the pre-fix value:**
- **Layer 1 (host):** nothing created on `felhom-pve` or `demo-hp`. No VM, no CT, no storage entry.
- **Layer 2 (guest):** two throwaway shell scripts were pushed into guest 9201 to drive the endpoints
and **both were removed**; one canary directory was planted under `backups/offsite-restore/kimai`
on the system drive to prove R-360 and **was removed**. The `mode=unit` scratch restore left a
normal verification copy on the HDD, which is ordinary product state a customer can delete.
- **Layer 3 (hub):** no customer, appliance or host record created, so none to discard.
1. Restore the pre-fix card markup → **all four hub tests fail**, reporting `67` and `134.3 MB` absent
from the rendered page and the `Integrity` row present.
2. Drop `StatsKnown: t.StatsKnown` from `OffboxReportStatus()` → `a MEASURED empty repository reported
stats_known=<nil>`.
`demo-felhom`, `ep0`, DooPlex and Peti's box were not touched.
Both restored immediately; `git diff` clean afterwards.
## 9. Register
**Green gates:** controller `go build && go vet && go test ./...` — 28 packages, rc 0. Hub, same
command in `hub/` — 18 packages, rc 0.
**Size: 165 rows before → 167 after filing → 161 after housekeeping.**
## 6. Deployment and live verification
- **Closed:** R-353, R-357, R-358, R-360 — each with shipping version and evidence path.
- **Filed and closed in the same session:** **R-396** (Scenario F's answer — see §11) and **R-395**
(`STATUS.md` contradicting itself; **fixed**, not merely recorded).
- **Re-ranked:** none.
- **Housekeeping:** the six closed rows were compressed into `CLOSED-ITEMS.md` keeping title,
version, evidence and every sentence that states a rule; the full original is
`git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md`. **No open row was touched.**
- **`ROADMAP.md` was not edited:** none of these four ever had a ROADMAP row — they live in the
register only. Stated rather than silently skipped.
Recorded in the hub's own `REPORT.md`; the two halves ship together (the hub renders "unknown" for any
box still on 0.224.0, which is correct rather than wrong).
## 10. Documentation coupling (§5.5)
## 7. Not done, and why
- **`00-capability-map.md`** — one new row, splitting the claim: R-353/R-358/R-360 **PROVEN-LIVE**
with the evidence citation; **R-357 IMPLEMENTED only**, with the reason.
- **`07-backup-architecture.md`** — §10.2 gained five rows (the four plus R-396). **§8 matrix row 3
KEEPS its PROVEN status**, with a note recording why: R-353 was a defect in the *message*, not the
mechanism.
- **`OPEN-ITEMS.md` / `CLOSED-ITEMS.md` / `STATUS.md`** — as §9.
- **`ROADMAP.md`** — no applicable rows.
- **Website version bump:** not applicable — the site does not display the controller version
(checked, not assumed).
- **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it.
A second verdict on the same data is two things that can disagree, which this codebase has been bitten
by before (`LastRun` vs `LastSuccess`, R-100).
- **The dead `BackupReport` fields were not removed from the wire.** Removing them would stop
historical reports already in the hub's store from parsing, for no gain — nothing renders them now,
and a test fails if anything starts producing them.
## 11. Observations — noticed, documented, not acted on
1. **Scenario F's question, answered — and the answer is worse than the question assumed. Filed as
R-396.** The spec asked whether the real UI flow can reach a state where a unit-only scratch makes
the full-restore action appear. **It can, by the safest-looking action on the page.**
„Ellenőrző visszaállítás" (`mode=unit`, the default, advertised as non-destructive) calls
`RestoreOffboxScratch(full=false)`; `offboxRestoreScratchDir` **ignores `full`**, so both modes
write the same directory, and `--include` limits *what* restic extracts, never *where*; the wizard
sets `ScratchReady` from `OffboxFullScratchReady`; `deriveWizardStep` derives **both**
`PlaceEnabled` and `RestoreEnabled` from that one flag. R-358 as filed assumed the bad state needed
a *failed* download. It needed only a successful safe one. **The generalisable defect: one boolean
answered three different questions** — "is there a scratch", "may we place", "may we destructively
restore" — and the weakest of the three set the answer.
2. **`NotifyIntegrityOK` / `NotifyIntegrityFailed` are dead code** (noticed during R-331 earlier
today, restated here because it is restore-adjacent): they exist and are called from nowhere; the
controller runs no integrity check at all.
3. **`resticStep` is not a seam**, which is why R-358's ordering property needed an AST test rather
than an execution test. If a future task needs to drive restic paths under test, that is the seam
to add — and it should be added deliberately, not improvised inside a bugfix.
4. **`RestoreApp`'s own `restoreDockerVolumes` count is still discarded.** Left alone deliberately:
the spec scoped `RestoreApp` out, and it is the no-unit fallback whose result is already reported
as a zero `UnitRestoreResult` — which is honest, because a box with no recovery unit really did
return no data from one. Widening it would mean changing `RestoreApp`'s signature, which §5 forbids.
5. **The spec's own §13 step 1 named an app whose shape has since changed.** Not a defect in the
spec — a reminder that fixture assumptions about live boxes decay, and the instruction to
"confirm its manifest first rather than assuming" is what caught it.