From 300d7e87d7cfcbb6dd355594f5d8936c06184a27 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 30 Aug 2026 21:29:15 +0200 Subject: [PATCH] REPORT: R-359/R-397 shipped and validated; the measurement changed what it is worth The headline is not the feature, it is what measuring it revealed: THE STRUCTURE CHECK THAT SHIPS ON DOES NOT CATCH SILENT CORRUPTION. A pack corrupted without a size change returned `no errors were found`, exit 0. Only --read-data caught it. So R-399 is not merely a bandwidth question -- at the shipped default a class of damage is not checked at all. The three numbers R-399 needed are MEASURED, not estimated: store 134.3 MB / 67 snapshots; structure check 35.0 s; and the full curve 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. At this size re-reading everything costs four seconds more than reading none. Stated limit: they do not extrapolate. Records four things that went wrong and were caught rather than shipped: - R-398 was MY OWN mistaken row. The seam already existed, Part 0 was not built, and building it would have HIDDEN the unlock --remove-all escalation from the assertions that must see it. Corrected, not closed. - the damage classifier matched restic's ORDINARY progress output; the negative control caught it (v0.227.1). - an exit code I misread through a pipe, corrected by re-measuring. - wire-contract convicted a field the hub cannot decode; allowlisted WITH A REASON rather than skipped, because building the hub display is a decision R-331 already ruled belongs to the operator. And the sweep the task asked for: EIGHT debug buttons post to endpoints that do not exist, not one. 24 referenced, 17 dispatched. Filed as R-400. Not live-validated and each says why: the weekly firing (a week away), read-data on a large store, and the join between "restic catches it" and "my code classifies it" -- both proven, the join is not, and the seam is named. --- REPORT.md | 432 +++++++++++++++++++++--------------------------------- 1 file changed, 168 insertions(+), 264 deletions(-) diff --git a/REPORT.md b/REPORT.md index 2dcdbcb..c4396c5 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,323 +1,227 @@ -# REPORT — R-353 / R-357 / R-358 / R-360: the restore tells the truth +# REPORT — R-359 / R-397: the off-site store gets checked -**Controller v0.226.0 → v0.226.1 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`** +**Controller v0.227.0 → v0.227.1 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`** --- -## 1. Confirmed baselines used — BOTH HAD MOVED +## ⚠ READ THIS FIRST — the measurement changed what the feature is worth -| repo | spec baseline | **actual at start** | version | +Part 5 corrupted one pack of a throwaway repository **without changing its size** (64 zero bytes at +offset 1024 — the subtlest form of rot). Both depths were run against it: + +| depth | exit | verdict | +|---|---|---| +| `restic check` — **the depth that ships ON** | **0** | **`no errors were found`** | +| every `--read-data*` form — **ships OFF** | 1 | `Pack ID does not match, want 288afd3e…, got 4b6847bb…` | + +**The structure check reported a corrupted store as healthy.** It verifies the index, the pack +inventory and the snapshot graph — it catches missing packs, broken indexes and unreadable snapshots, +which are real — but it does **not** re-hash pack contents. **R-399 is therefore not only a bandwidth +question: at the shipped default a class of damage is not checked at all**, and it is the class that +silently eats a customer's photos. The task could not have known this; nothing had measured it. + +## 1. Confirmed baselines + +| repo | spec baseline | actual at start | version | |---|---|---|---| -| `felhom-controller` | `f8c9390` / v0.223.0 → v0.224.0 | **`e5eee50` / v0.225.0** | → **v0.226.0** | -| `felhom.eu` | `c2c1fb4` | **`ac6ac03`** | docs only | +| `felhom-controller` | `c476ae51` | **`e64c84ae`** (drift: one docs-only commit — my own golden-bake report) | v0.226.1 → **v0.227.1** | +| `felhom.eu` | `4f875174` | `4f875174` — **exact match** | docs only | | `felhom-agent` | not touched | v0.130.0 | unchanged | -**Both targets were consumed earlier the same day by my own work** — v0.224.0 (R-330) and v0.225.0 -(R-331). The drift was re-confirmed against live Gitea before the first edit and the operator -authorised proceeding. **Every symbol in the spec's §5 table was re-verified present at the real -baseline** before editing; all 20 resolved, and almost every line landmark still matched. `MinAgent` -stays **0.129.0**. Both trees were clean and equal to `origin/main` at the start. +Both trees clean and equal to `origin/main`. **Every symbol in §5 was re-verified present** (20/20), +and all three claimed absences confirmed: 0 occurrences of a restic `check` verb, 0 callers of +`NotifyIntegrity*` in the controller, and the debug button present with no dispatch case. +`MinAgent` stays **0.129.0**. ## 2. Files created / modified -**`felhom-controller`** (all paths under repo root): +**`felhom-controller`:** `controller/internal/backup/offbox_integrity.go` (**new**), +`offbox.go` (report fields), `offbox_restore.go`+`r358_scratch_marker_test.go` (AST→execution test), +`internal/settings/settings.go`, `internal/config/config.go`, +`cmd/controller/main.go` (job + the one caller + the debug callback), +`internal/web/handler_debug.go`, `internal/web/handlers.go`, +tests: `r359_integrity_test.go`, `r359_lock_safety_test.go`, `r359_dueness_test.go` (**new**), +`cmd/controller/r359_wiring_test.go`, `internal/web/r397_debug_integrity_test.go` (**new**); +docs: `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md`. -| file | change | -|---|---| -| `controller/internal/backup/restore.go` | `restoreDockerVolumes` → `(int, error)`; caller updated | -| `controller/internal/backup/restore_unit.go` | `UnitRestoreResult`; `RestoreFromRecoveryUnit` → `(UnitRestoreResult, error)` on every return path | -| `controller/internal/web/handlers.go` | `unitRestoreOutcomeMsg` + 3 message constants; handler publishes the outcome | -| `controller/internal/backup/offbox_reconstitute.go` | the R-357 headroom gate | -| `controller/internal/backup/offbox_restore.go` | scratch marker (write/clear/read), `OffboxFullScratchReady` rewritten, shared refusal constants, `SetOffboxLatestSnapshotFn`, `WriteScratchMarkerForTest` | -| `controller/internal/backup/backup.go` | `offboxLatestSnapFn` field | -| `controller/internal/web/offbox_handlers.go` | server-side scratch refusals in 2 handlers; R-360 delete guard; corrected doc comment | -| `controller/internal/backup/r357_reconstitute_headroom_test.go` | **new** | -| `controller/internal/backup/r358_scratch_marker_test.go` | **new** | -| `controller/internal/web/r353_unit_outcome_test.go` | **new** | -| `controller/internal/web/r358_r360_handlers_test.go` | **new** | -| 3 existing `*_test.go` in `internal/backup` | mechanical `_, err :=` for the changed signature | -| `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md` | docs | - -**`felhom.eu`**: `STATUS.md`, `documentation/backlog/OPEN-ITEMS.md`, -`documentation/backlog/CLOSED-ITEMS.md`, `documentation/architecture/07-backup-architecture.md`, -`documentation/architecture/00-capability-map.md`, -`documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` (**new**). +**`felhom.eu`:** `STATUS.md`, `backlog/OPEN-ITEMS.md`, `backlog/CLOSED-ITEMS.md`, +`architecture/07-backup-architecture.md`, `architecture/08-alarm-ladder.md`, +`architecture/00-capability-map.md`, `scripts/wire_contract_gate.py` (allowlist **with a reason**), +`documentation/tests/r359-integrity-2026-08-30/` (**new**, 10 files). ## 3. Commits pushed to `main` | repo | commit | contents | |---|---|---| -| `felhom-controller` | **`b8af7276`** | all four fixes, all tests, controller docs | -| `felhom.eu` | **`e027b5d9`** | register closures, R-395 fix, architecture, capability map, evidence | -| `felhom.eu` | *(this report's commit)* | housekeeping compression + REPORT | +| `felhom-controller` | **`0d52a42c`** | v0.227.0 — the check, the job, the notifiers, the route, all tests | +| `felhom-controller` | **`45770f22`** | v0.227.1 — the classifier fix + CONTEXT/README/REUSE | +| `felhom.eu` | *(this report's commit)* | register, architecture, capability map, STATUS, evidence | ## 4. Tests and red-proofs -**All pass.** New tests, by group: +**All pass.** 22 new tests across 5 new files. Groups A (the check), B (the lock hazard), C (due-ness), +D (wiring), plus two fixtures built from restic's **real** bytes. -| test | result | -|---|---| -| `TestUnitRestoreOutcome_VolumesAndDatabaseNamed` (A1) | pass | -| `TestUnitRestoreOutcome_BackupHeldOnlySettings` (A2) | pass | -| `TestUnitRestoreOutcome_ManifestListedDataThatDidNotReturn` (A3) | pass | -| `TestUnitRestoreOutcome_DatabaseOnly` (A4) | pass | -| `TestR353_HandlerPublishesTheOutcome` (A5, the seam test) | pass | -| `TestR357_DestructiveRestoreRefusesWithoutHeadroom` (B1) | pass | -| `TestR357_UnknownSizeFailsClosed` (B2) | pass | -| `TestR357_UnknownFreeSpaceFailsClosed` (added — the mirror hole) | pass | -| `TestR357_AmpleSpaceIsUnchanged` (B3) | pass | -| `TestR358_FailedRestoreLeavesNoUsableScratch` (C1) | pass | -| `TestR358_UnitOnlyScratchIsNotFullReady` (C2) | pass | -| `TestR358_StaleMarkerIsClearedBeforeTheRun` (C3) | pass | -| `TestR358_UnreadableMarkerFailsClosed` (C4) | pass | -| `TestR358_WrongSchemaFailsClosed` (added) | pass | -| `TestR358_CompletedFullScratchStillReady` (C5) | pass | -| `TestR358_MarkerIsWrittenAt0600AndAtomically` (added) | pass | -| `TestR358_MarkerIsNeverPlaced` (D3) | pass | -| `TestR358_MarkerIsClearedBeforeResticAndWrittenAfter` (added, AST) | pass | -| `TestR358_PlaceHandlerRefusesIncompleteScratch` (D1) | pass | -| `TestR358_ReconstituteHandlerRefusesIncompleteScratch` (D2) | pass | -| `TestR358_UnitOnlyScratchClosesTheFullRestoreCard` (Scenario F, flow level) | pass | -| `TestR360_VerifyCopyDeleteRefusedDuringRestore` (D4) | pass | -| `TestR360_VerifyCopyDeleteStillWorksWhenIdle` (Scenario H) | pass | - -**E1 — no existing test was modified for its content.** Every `TestReconstituteOutcome_*` passes -unmodified. The only test-file edits were mechanical call-site updates for the changed -`RestoreFromRecoveryUnit` signature (`err :=` → `_, err :=`) in three files. - -### Red-proofs — each mutated, observed failing, reverted, `git diff` clean - -| # | mutation | observed failure | +| red-proof | mutation | observed failure | |---|---|---| -| **A5** | `EndRestoreOp(true, stackName+" visszaállítva ("+snapshotID+").")` restored | `THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)."` | -| **B1** | the whole R-357 gate deleted | `THE APP WAS STOPPED (1 call(s)) for a restore with 300 KB free for a 1 MB copy … (err=)` — and the same on both fail-closed tests | -| **C1** | the old `len(entries) > 0` check restored | `a part-copy was reported READY`, plus the unit-only and unreadable-marker cases | -| **D4** | the delete guard reverted to `IsRunning()` | `THE VERIFICATION COPY WAS DELETED while a restore was writing into it` | -| **D1/D2** | both server-side scratch refusals removed | both handlers redirected with „…elindult" over a part-copy | +| **A2** | classifier treats any non-zero as OK | `a repository restic said contains errors was reported as OK` | +| **B1** | remove the `acquireRunning` guard | `RESTIC WAS INVOKED while another operation held the single-writer flag: [… cat config] [… check]` | +| **C2** | gate due-ness on `time.Sunday` | `a 21-day-old check was NOT due on Saturday` | +| **D2** | remove the debug dispatch case | `POST /api/debug/backup/integrity was not dispatched` | -> **⚠ THE FIRST B1 RED-PROOF EXPOSED A HOLLOW TEST OF MY OWN, and it is recorded rather than quietly -> fixed.** With the gate removed, the run refused *earlier* — at the placement stat pre-pass — so -> `stops == 0` passed against the pre-fix code and the test proved nothing about the thing it exists -> for. Two corrections: the fixture now populates the scratch the way a completed download leaves it -> (the snapshot's own absolute paths mirrored under the scratch, plus `SetSafetyDumpFn` so no Docker -> is needed), and the assertions are **reordered** so a removed gate reports the outage rather than -> "no error returned". Only after that does the red-proof print `THE APP WAS STOPPED (1 call(s))`. -> The lesson is the doctrine's own: a test that cannot fail on the pre-fix shape is decoration. +**E1 — no existing test was modified.** `TestBackupReport_DeadFieldsStayZero` (R-331, D4) passes +unmodified, which is the check that I did not write to the retired report fields. -**Test count:** 24 new tests added across 4 new files. **Green gate:** -`go build ./... && go vet ./... && go test ./...` → **28 packages, rc 0**, no failures. -**All 12 controller gates OK.** +**Test count:** 28 packages green before and after; +22 tests. Green gate +`go build ./... && go vet ./... && go test ./...` → **rc 0**. ## 5. Deployed version ``` -$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'" -gitea.dooplex.hu/admin/felhom-controller:0.226.1 Up 13 seconds (healthy) +gitea.dooplex.hu/admin/felhom-controller:0.227.1 Up 13 seconds (healthy) +[INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST ``` -**A patch, not a rebuilt 0.226.0.** The live validation in §6 ran against **0.226.0**, which was -already on the box when §11 item 4's defect was found. Re-pushing a changed image under a tag already -running somewhere is the `:latest` hazard with extra steps — two images, one name, no way for a box to -tell which it has — so the fix shipped as **0.226.1** and the box was redeployed. **The §6 evidence -therefore describes 0.226.0**, and that is stated rather than quietly re-attributed; 0.226.1 differs -from it only by `CountsUnknown` and its two tests, neither of which touches any path §6 exercised. +**`demo-hp` only.** `demo-felhom` is on 0.226.1; the fleet floor and golden are 0.226.1. -**Fleet state at the end of the session — the debt named in §12 was PAID the same day.** Golden -**0.226.1** was baked, published, round-trip verified, **vouched** and the floor **raised to 0.226.1**. -Both demo machines now run 0.226.1, and `demo-felhom` got there by **self-update**, not by hand: +**v0.227.1 is a patch, not a rebuilt 0.227.0** — 0.227.0 was already running on `demo-hp` when the +classifier defect was found, and re-pushing changed bytes under a live tag is the `:latest` hazard. +The live validation in §8 ran against **0.227.0**; 0.227.1 differs only by the damage signatures and +their test. -``` -[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1) -``` +## 6. The schedule I chose, and what I read to choose it -`golden_currency_gate.py` went **red → green** on the same command. Evidence: -`felhom.eu/documentation/tests/golden-0.226.1-2026-08-30/`. +Live on `demo-hp`: `db-dump 02:30` · `tier2-backup 03:30` · `fill-watch 03:30` · `metrics-prune 04:00` +· **`offbox-backup 04:15`** · `offsite-abandon-sweep 05:10` · whole-guest gate `[04:30, 08:30)`. -## 6. Live validation — endpoint level, on `demo-hp` +**Chose 06:00** — 1h45 clear of the off-site leg (which runs ~2m20s, measured) and 50 min clear of the +sweep. The backup window is customer-configurable, so **no fixed time is collision-proof for every +box**; a collision costs one skipped day, not a missed check, because due-ness makes tomorrow try +again. A slot that collided *every* night would be a real defect; this is not one. -**Method:** the exact endpoints the UI invokes, driven from inside guest 9201 (no browser on -DooPlex; the residual is client-side rendering only). No state was hand-set. Evidence copied off the -box at the end of the phase: `felhom.eu/documentation/audits/evidence-r353-r360-live-2026-08-30/`. +## 7. NOT yet live-validated -### The verbatim messages +1. **A real weekly firing.** It is a week away. The job is confirmed **registered** on the box, which + is a different and weaker claim, and the capability map says so. +2. **Every `--read-data-subset` claim about a LARGE store.** The curve in §9 is measured on 134.3 MB + and does not extrapolate. +3. **The failure branch end-to-end through my code.** restic's damage output was captured live and fed + to the classifier as a fixture (`TestR359_RealResticDamageOutputIsClassifiedAsDamage`), but + `CheckOffboxIntegrity` reads the *configured* target, so pointing it at the damaged scratch repo + would have meant repointing the live off-site target. **The seam is named rather than glossed:** + restic-catches-it and my-code-classifies-it are each proven; the join is not. +4. **The skip path via a real backup.** See §8 — the intended trigger does not exist (R-279). -**R-353** — `POST /backup/restore` for `opengist`, then the sentence read off the customer's own -wizard page (`GET /backups/restore/app?name=opengist`): +## 8. Part 5 and the live validation, in full -``` -A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult. -``` +Evidence: `felhom.eu/documentation/tests/r359-integrity-2026-08-30/` (10 files, copied off **before** +teardown). -and the controller's own lines: +- **Scratch repo:** `/mnt/felhom-drives/hdd_1/r359-scratch-repo`, 3 files (195.4 KiB), snapshot + `3611a338`, hashes recorded. +- **Negative control FIRST:** healthy repo → `no errors were found`, exit 0, **703 ms**. +- **Damage:** pack `288afd3e868dc6bd…217bf0cc`, 64 zero bytes at offset 1024, `conv=notrunc`. Size + **unchanged** at 200 333 B; sha256 → `4b6847bb5eec…dc496d2b`. +- **Positive control:** the table at the top of this report. +- **The debug button that had never done anything:** + `{"data":{"duration_ms":35550,"ok":true,…},"message":"Az ellenőrzés rendben lezajlott","ok":true}` +- **R-397's notifier, end to end:** + `Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)` + — pushed to the hub, HTTP 200, severity `info`, so it mailed nobody. +- **THE HAZARD CONTROL, live.** The intended demonstration could not be run: `POST + /api/backup/offbox/run` returns **404** — there is no operator-triggerable off-site backup, which is + **R-279 and stays open**. So the same flag was exercised by its other holder — two checks 6 s apart: + ``` + B (while A held the flag): {"skipped":true,"skip_reason":"a backup or restore is already running","duration_ms":0} + A (completed): {"skipped":false,"ok":true,"duration_ms":34953} + ``` + **`duration_ms: 0` is the observable that matters: B never ran restic at all.** -``` -[INFO] [backup] Restore-from-unit completed: opengist — 1 volume(s) of 1 listed, 0 database(s) of 0 listed -[INFO] [web] Restore completed (async): stack=opengist in 9.112668936s (volumes 1/1, dbs 0/0) -``` +## 9. The three numbers R-399 needs — all MEASURED -**This is Scenario A, not Scenario B, and the substitution is stated rather than glossed.** The spec -named `opengist` as the data-less app from 21 August. **It is not one any more** — checked before -relying on it, as the spec instructed: every unit on `demo-hp` today lists at least one volume dump -(`opengist 1/0`, `privatebin 1/0`, `calibre-web 1/0`, the rest 2–3 volumes plus a database). **No app -on the box has the Scenario B shape**, and manufacturing one would mean falsifying a manifest — the -hand-set-state shortcut this project forbids. Scenario B is carried by -`TestUnitRestoreOutcome_BackupHeldOnlySettings` and the A5 seam test. What the live run *does* prove -is the whole path: real counts, correct clause selection (no database clause for 0 databases), and -the sentence reaching the customer's page. +| | | +|---|---| +| live store size | **140 829 678 B (134.3 MB)**, 2 651 blobs, 67 snapshots (`restic stats --mode raw-data`) | +| structure check wall-clock | **35.0 s** | +| 10% read-data | **35.9 s** (+0.9 s, **+3%**) | +| 50% | 37.3 s (+6%) | +| **100%** | **39.2 s (+12%)** | -**R-360** — a `mode=unit` scratch restore started, and the verification-copy delete POSTed while it -ran. The refusal, urldecoded from the `Location` header: +**Not an estimate — measured against the live store, read-only.** At this size, re-reading *all* the +data costs about four seconds more than reading none, because the wall clock is SFTP round-trips over +the WireGuard tunnel rather than transfer. **Stated assumption and its limit:** these do not +extrapolate — the structure check's cost tracks the INDEX, a read-data run's tracks the DATA, so a +50 GB store is ~370× the data and this curve says nothing about it. -``` -/backups/restore?flash_error=Egy visszaállítási művelet (kimai) már fut, ezért most nem indítható -újabb. Az állapotát ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást. -``` +## 10. Teardown — all three layers -``` -[WARN] [web] verification-copy delete refused for kimai: a backup/restore op is running -``` - -**And the consequence, which is the assertion that matters:** a canary file planted in the copy was -still there afterwards — `-rw-r--r-- 1 root root 10 Aug 30 17:34 canary.txt`, contents `defend-me`. - -**R-358 / R-396** — after the `mode=unit` restore completed: - -``` -$ cat …/backups/offsite-restore/kimai/.felhom-restore-complete.json -{"schema":1,"snapshot_id":"84542ec8","full":false,"finished_at":"2026-08-30T17:34:42Z"} -mode=600 (no .tmp left behind) - -[INFO] [offbox] kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy, - so place-to-live stays closed -``` - -**This is Scenario F proven live**, and it is the case that pre-fix would have unlocked the -destructive restore. - -## 7. NOT yet live-validated — awaiting a supervised drill - -1. **R-357's disk-full behaviour.** Filling a real filesystem to prove it is a drill step, not a build - step. Carried by the seam tests, whose central assertion is `StopStack` call count 0. -2. **R-353's Scenario B** (the "backup held only settings" sentence) — no app on `demo-hp` has that - unit shape any more; see §6. -3. **R-353's Scenario C** (the unit lists dumps, none return) — the R-367 stranded-dump shape was not - reproduced on live hardware; unit-tested only. -4. **R-358's failed-download branch** was not induced live. It was not needed: the unit-only branch - exercises the same marker gate through a real restore, without pointing restic at a bad snapshot. - -## 8. Teardown - -**This run provisioned nothing.** No machine, no guest, no hub record — all three layers: - -- **Layer 1 (host):** nothing created on `felhom-pve` or `demo-hp`. No VM, no CT, no storage entry. -- **Layer 2 (guest):** two throwaway shell scripts were pushed into guest 9201 to drive the endpoints - and **both were removed**; one canary directory was planted under `backups/offsite-restore/kimai` - on the system drive to prove R-360 and **was removed**. The `mode=unit` scratch restore left a - normal verification copy on the HDD, which is ordinary product state a customer can delete. +- **Layer 1 (host):** nothing provisioned on `demo-hp`. No VM, no CT, no storage entry. +- **Layer 2 (guest):** the scratch restic repo was created and **removed** — **1 511 424 bytes** + returned to `/mnt/felhom-drives/hdd_1`; seven throwaway scripts in the guest and the container + removed; the `mode=unit` scratch from yesterday untouched. - **Layer 3 (hub):** no customer, appliance or host record created, so none to discard. `demo-felhom`, `ep0`, DooPlex and Peti's box were not touched. -## 9. Register +## 11. Register -**Size: 165 rows before → 167 after filing → 161 after housekeeping.** +**Size: 163 before → 165 after filing → 163 after housekeeping.** -- **Closed:** R-353, R-357, R-358, R-360 — each with shipping version and evidence path. -- **Filed and closed in the same session:** **R-396** (Scenario F's answer — see §11) and **R-395** - (`STATUS.md` contradicting itself; **fixed**, not merely recorded). -- **Re-ranked:** none. -- **Housekeeping:** the six closed rows were compressed into `CLOSED-ITEMS.md` keeping title, - version, evidence and every sentence that states a rule; the full original is - `git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md`. **No open row was touched.** -- **`ROADMAP.md` was not edited:** none of these four ever had a ROADMAP row — they live in the - register only. Stated rather than silently skipped. +- **Closed:** R-359, R-397 — with version and evidence path, then compressed into `CLOSED-ITEMS.md`. +- **CORRECTED, not closed: R-398** — see §12 item 1. +- **Filed:** **R-399** (the depth decision, now with three measured numbers and a re-framed premise) + and **R-400** (the debug buttons — see §12 item 2). +- **Explicitly left open and said so:** **R-87** (the restic tier is never restore-**tested** — a check + is not a restore-test, and the two rows are adjacent, which is exactly how they would get + conflated), **R-279** (no operator-triggerable off-site backup), **R-104**. -## 10. Documentation coupling (§5.5) +## 12. Observations -- **`00-capability-map.md`** — one new row, splitting the claim: R-353/R-358/R-360 **PROVEN-LIVE** - with the evidence citation; **R-357 IMPLEMENTED only**, with the reason. -- **`07-backup-architecture.md`** — §10.2 gained five rows (the four plus R-396). **§8 matrix row 3 - KEEPS its PROVEN status**, with a note recording why: R-353 was a defect in the *message*, not the - mechanism. -- **`OPEN-ITEMS.md` / `CLOSED-ITEMS.md` / `STATUS.md`** — as §9. -- **`ROADMAP.md`** — no applicable rows. -- **Website version bump:** not applicable — the site does not display the controller version - (checked, not assumed). +1. **NOT-A-FINDING: acted on instead — R-398 was my own mistake and is CORRECTED, not closed.** + The row said `resticStep` is not a seam so no test can drive a restic-backed path. **The first half + is true and the conclusion was false:** `offboxRunner` / `SetOffboxRunner` / `m.runner()` has been + injectable since the off-site tier shipped, and `offbox_3a_test.go` has been using it five times + over. I read one function and generalised. **Part 0 was therefore not built, and building it would + have been actively harmful:** a `resticStepFn` seam replaces `resticStep`, hiding its + `unlock --remove-all` escalation from the very assertions that must observe it. What the row asked + for that *was* real is done — R-358's AST ordering test is now an execution test, which immediately + surfaced something the AST walk could not: `unlockStale` legitimately runs before the restore. -## 11. Observations — each with a register row, or a stated reason it needs none +2. **FILED: R-400 — and the sweep found EIGHT dead buttons, not one.** Comparing every + `/api/debug/...` reference in `debug.html` against every `subpath ==` case: **24 referenced, 17 + dispatched.** The seven others are `backup/crossdrive`, `backup/infra`, `dr/infra-status`, + `hub/infra-push`, `storage/simulate-disconnect`, `storage/simulate-reconnect`, + `storage/watchdog-status`. Single dispatcher, exact-match switch, `default: http.NotFound` — so they + 404 rather than silently succeed. **`controller/README.md` documented four `backup/*` debug routes + when two existed**; corrected. A third of a debug page does nothing, on the surface an operator + reaches for when something is already wrong. -1. **FILED: R-396.** *Scenario F's question, answered — and the answer is worse than the question - assumed.* The spec asked whether the real UI flow can reach a state where a unit-only scratch makes - the full-restore action appear. **It can, by the safest-looking action on the page.** - „Ellenőrző visszaállítás" (`mode=unit`, the default, advertised as non-destructive) calls - `RestoreOffboxScratch(full=false)`; `offboxRestoreScratchDir` **ignores `full`**, so both modes - write the same directory, and `--include` limits *what* restic extracts, never *where*; the wizard - sets `ScratchReady` from `OffboxFullScratchReady`; `deriveWizardStep` derives **both** - `PlaceEnabled` and `RestoreEnabled` from that one flag. R-358 as filed assumed the bad state needed - a *failed* download. It needed only a successful safe one. **The generalisable defect: one boolean - answered three different questions** — "is there a scratch", "may we place", "may we destructively - restore" — and the weakest of the three set the answer. +3. **NOT-A-FINDING: noted for R-359's own row, not a separate one.** `offbox_progress.go:328` carries a + near-identical lock-retry block to `resticStep`'s. Noted, **not merged** — §12 forbade it and the + duplication is load-bearing until someone proves otherwise. -2. **FILED: R-397.** *`NotifyIntegrityOK` / `NotifyIntegrityFailed` are dead code and the product - advertises a check that does not exist.* Both notifiers have no caller anywhere; the controller - runs no integrity check at all. **The part that actively misleads:** `config.Monitoring.PingUUIDs` - carries a `backup_integrity` field and the monitoring page renders - „Mentés integritás — Hetente (vasárnap)", so the operator is told a weekly check runs. Noticed - twice in one day from opposite directions (R-331's dead report fields, then this task), which is - why it earned a row rather than a note. +4. **NOT-A-FINDING: caught and fixed in the same session.** The damage classifier matched restic's + ordinary progress output (`"pack "` vs `check all packs`), so a check failing for a non-damage + reason would have told the customer their backups were corrupt. **The negative control caught it** — + which is precisely why a control that has only seen the failing case is worth nothing. Shipped as + v0.227.1. -3. **FILED: R-398.** *`resticStep` is not a seam, so no test can drive any restic-backed path.* Felt - directly here: R-358's safety property is an ORDER (clear before restic, write after) and with no - seam it could only be pinned by an **AST walk** rather than by execution. That is honest about what - it proves and would still miss a reordering introduced through a helper. The contrast is the - argument: `offboxLatestSnapshot` got a seam in this same release, in four lines, because a - correctness gate could not otherwise be proven. +5. **NOT-A-FINDING: my own measurement error, corrected before it reached a conclusion.** An early run + showed `exit=0` on the read-data subset forms while they printed `Fatal: repository contains + errors`. That was not a restic defect — the commands were piped through `tail`, so `$?` was tail's. + Re-measured without pipes: every read-data form exits 1. This project's own "exit codes that lie" + trap, caught by re-measuring rather than reasoning. -4. **NOT-A-FINDING: it was ACTED ON in this release instead of filed — and the acting is the story.** - `RestoreApp`'s own `restoreDockerVolumes` count is still discarded, and the spec scopes - `RestoreApp` out. My first draft justified that as *"already reported as a zero result, which is - honest, because a box with no recovery unit really did return no data from one."* +## 13. The push bypassed one gate, deliberately - > **⚠ THAT SENTENCE WAS FALSE, AND WRITING IT EXPOSED A DEFECT THE FIX ITSELF HAD INTRODUCED.** - > A zero `UnitRestoreResult` is *Scenario B's shape*, so the no-unit fallback would have printed - > „ez a mentés csak a beállításokat tartalmazta, adatot nem" over a restore that may have replayed - > the app's entire dataset — **an unknown drawn as a zero, the exact R-88 failure direction this - > whole change exists to remove**, re-introduced by the change itself. - > - > Fixed rather than filed: `UnitRestoreResult` now carries `CountsUnknown`, the fallback sets it, - > and the surface has a fourth sentence claiming only what is known — the restore ran, the app is - > back, and we cannot say what came back. **`RestoreApp`'s signature is untouched**, so §5 holds. - > Pinned by `TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty`; the A5 seam test was - > corrected too, because its fixture has no unit and therefore exercises exactly this path while - > asserting the wrong sentence. - > - > **It was the `observations` gate refusing the push that forced the re-read.** The gate exists to - > stop findings dying in an overwritten REPORT.md; here it caught a live defect instead. +**`felhom.eu` — `golden-currency` CONVICTED.** v0.227.0/0.227.1 are released and the vouched golden +carries 0.226.1. **The gate is right.** A **BYPASS, not a waiver**, and the spec directs it: §13 says +golden and fleet delivery are Viktor's (R-242). **A golden carrying 0.227.1 is OWED**, and it is item 3 +under "Waiting on you" in `STATUS.md`. -5. **NOT-A-FINDING: the spec's own instruction already handled it, so there is nothing to carry forward.** §13 step 1 named `opengist` as - the data-less app; it is not one any more. Not a defect in the spec: it said *"confirm its manifest - first rather than assuming"*, and that is exactly what caught it. Recorded only as a reminder that - fixture assumptions about live boxes decay faster than the documents naming them. - -## 12. Two gates fired; one push was bypassed, and both are declared - -**`felhom.eu` — `golden-currency` CONVICTED, and the debt is now PAID.** Three controller releases -shipped today and the vouched golden still carried **0.223.0**, so a machine installed at that moment -would have received none of them. **The gate was right.** Those pushes used `git push --no-verify` — a -**BYPASS, not a waiver**, on the operator's standing ruling, re-checked for this release rather than -reused blindly. - -**Later the same day the operator asked for the bake, and it was done: golden 0.226.1 baked, published, -round-trip verified, vouched, and the floor raised.** The gate went **red → green**, so those bypasses -are **historical rather than standing**. See §5 for the fleet state and -`felhom.eu/documentation/tests/golden-0.226.1-2026-08-30/` for the evidence, including the round trip -and the proven token-leak grep. - -**`felhom-controller` — nothing was bypassed.** The `observations` gate refused the first REPORT push -and was **fixed, not bypassed** — see §11 item 4 for what that cost and what it caught. - -**And one earlier gate was fixed rather than bypassed:** `due-checks` was red on R-341's overdue -`+7 d` measurement. It was **taken** during this session (ep0: fd **17**, the baseline, same proxy -generation, PID 551655 unchanged), and the verdict recorded as **unanswerable** — our own R-344 fix -removed the leak mid-interval, so a slope of ~0 measures that fix and not the PBS upgrade the row was -asking about. +**`wire-contract` was NOT bypassed — it caught something real and was answered.** It convicted +`offsite.last_integrity_ok`: the controller emits it and no hub struct can decode it. That is the +"emitting into the void" shape. It is now **allowlisted with a written reason** rather than skipped: +building a hub display is a hub change, and R-331 ruled that class a decision for the operator — the +*previous* integrity display was removed precisely because it rendered a value nothing wrote. The entry +says to delete it when a surface is built.