REPORT: R-359/R-397 shipped and validated; the measurement changed what it is worth
gates / gates (push) Failing after 12s

The headline is not the feature, it is what measuring it revealed: THE STRUCTURE
CHECK THAT SHIPS ON DOES NOT CATCH SILENT CORRUPTION. A pack corrupted without a
size change returned `no errors were found`, exit 0. Only --read-data caught it.
So R-399 is not merely a bandwidth question -- at the shipped default a class of
damage is not checked at all.

The three numbers R-399 needed are MEASURED, not estimated: store 134.3 MB / 67
snapshots; structure check 35.0 s; and the full curve 10% 35.9 s, 50% 37.3 s,
100% 39.2 s. At this size re-reading everything costs four seconds more than
reading none. Stated limit: they do not extrapolate.

Records four things that went wrong and were caught rather than shipped:
  - R-398 was MY OWN mistaken row. The seam already existed, Part 0 was not
    built, and building it would have HIDDEN the unlock --remove-all escalation
    from the assertions that must see it. Corrected, not closed.
  - the damage classifier matched restic's ORDINARY progress output; the
    negative control caught it (v0.227.1).
  - an exit code I misread through a pipe, corrected by re-measuring.
  - wire-contract convicted a field the hub cannot decode; allowlisted WITH A
    REASON rather than skipped, because building the hub display is a decision
    R-331 already ruled belongs to the operator.

And the sweep the task asked for: EIGHT debug buttons post to endpoints that do
not exist, not one. 24 referenced, 17 dispatched. Filed as R-400.

Not live-validated and each says why: the weekly firing (a week away), read-data
on a large store, and the join between "restic catches it" and "my code
classifies it" -- both proven, the join is not, and the seam is named.
This commit is contained in:
2026-08-30 21:29:15 +02:00
parent 45770f2282
commit 300d7e87d7
+168 -264
View File
@@ -1,323 +1,227 @@
# REPORT — R-353 / R-357 / R-358 / R-360: the restore tells the truth
# REPORT — R-359 / R-397: the off-site store gets checked
**Controller v0.226.0 → v0.226.1 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`**
**Controller v0.227.0 → v0.227.1 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`**
---
## 1. Confirmed baselines used — BOTH HAD MOVED
## ⚠ READ THIS FIRST — the measurement changed what the feature is worth
| repo | spec baseline | **actual at start** | version |
Part 5 corrupted one pack of a throwaway repository **without changing its size** (64 zero bytes at
offset 1024 — the subtlest form of rot). Both depths were run against it:
| depth | exit | verdict |
|---|---|---|
| `restic check` — **the depth that ships ON** | **0** | **`no errors were found`** |
| every `--read-data*` form — **ships OFF** | 1 | `Pack ID does not match, want 288afd3e…, got 4b6847bb…` |
**The structure check reported a corrupted store as healthy.** It verifies the index, the pack
inventory and the snapshot graph — it catches missing packs, broken indexes and unreadable snapshots,
which are real — but it does **not** re-hash pack contents. **R-399 is therefore not only a bandwidth
question: at the shipped default a class of damage is not checked at all**, and it is the class that
silently eats a customer's photos. The task could not have known this; nothing had measured it.
## 1. Confirmed baselines
| repo | spec baseline | actual at start | version |
|---|---|---|---|
| `felhom-controller` | `f8c9390` / v0.223.0 → v0.224.0 | **`e5eee50` / v0.225.0** | → **v0.226.0** |
| `felhom.eu` | `c2c1fb4` | **`ac6ac03`** | docs only |
| `felhom-controller` | `c476ae51` | **`e64c84ae`** (drift: one docs-only commit — my own golden-bake report) | v0.226.1 → **v0.227.1** |
| `felhom.eu` | `4f875174` | `4f875174` — **exact match** | docs only |
| `felhom-agent` | not touched | v0.130.0 | unchanged |
**Both targets were consumed earlier the same day by my own work** — v0.224.0 (R-330) and v0.225.0
(R-331). The drift was re-confirmed against live Gitea before the first edit and the operator
authorised proceeding. **Every symbol in the spec's §5 table was re-verified present at the real
baseline** before editing; all 20 resolved, and almost every line landmark still matched. `MinAgent`
stays **0.129.0**. Both trees were clean and equal to `origin/main` at the start.
Both trees clean and equal to `origin/main`. **Every symbol in §5 was re-verified present** (20/20),
and all three claimed absences confirmed: 0 occurrences of a restic `check` verb, 0 callers of
`NotifyIntegrity*` in the controller, and the debug button present with no dispatch case.
`MinAgent` stays **0.129.0**.
## 2. Files created / modified
**`felhom-controller`** (all paths under repo root):
**`felhom-controller`:** `controller/internal/backup/offbox_integrity.go` (**new**),
`offbox.go` (report fields), `offbox_restore.go`+`r358_scratch_marker_test.go` (AST→execution test),
`internal/settings/settings.go`, `internal/config/config.go`,
`cmd/controller/main.go` (job + the one caller + the debug callback),
`internal/web/handler_debug.go`, `internal/web/handlers.go`,
tests: `r359_integrity_test.go`, `r359_lock_safety_test.go`, `r359_dueness_test.go` (**new**),
`cmd/controller/r359_wiring_test.go`, `internal/web/r397_debug_integrity_test.go` (**new**);
docs: `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md`.
| file | change |
|---|---|
| `controller/internal/backup/restore.go` | `restoreDockerVolumes` → `(int, error)`; caller updated |
| `controller/internal/backup/restore_unit.go` | `UnitRestoreResult`; `RestoreFromRecoveryUnit` → `(UnitRestoreResult, error)` on every return path |
| `controller/internal/web/handlers.go` | `unitRestoreOutcomeMsg` + 3 message constants; handler publishes the outcome |
| `controller/internal/backup/offbox_reconstitute.go` | the R-357 headroom gate |
| `controller/internal/backup/offbox_restore.go` | scratch marker (write/clear/read), `OffboxFullScratchReady` rewritten, shared refusal constants, `SetOffboxLatestSnapshotFn`, `WriteScratchMarkerForTest` |
| `controller/internal/backup/backup.go` | `offboxLatestSnapFn` field |
| `controller/internal/web/offbox_handlers.go` | server-side scratch refusals in 2 handlers; R-360 delete guard; corrected doc comment |
| `controller/internal/backup/r357_reconstitute_headroom_test.go` | **new** |
| `controller/internal/backup/r358_scratch_marker_test.go` | **new** |
| `controller/internal/web/r353_unit_outcome_test.go` | **new** |
| `controller/internal/web/r358_r360_handlers_test.go` | **new** |
| 3 existing `*_test.go` in `internal/backup` | mechanical `_, err :=` for the changed signature |
| `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md` | docs |
**`felhom.eu`**: `STATUS.md`, `documentation/backlog/OPEN-ITEMS.md`,
`documentation/backlog/CLOSED-ITEMS.md`, `documentation/architecture/07-backup-architecture.md`,
`documentation/architecture/00-capability-map.md`,
`documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` (**new**).
**`felhom.eu`:** `STATUS.md`, `backlog/OPEN-ITEMS.md`, `backlog/CLOSED-ITEMS.md`,
`architecture/07-backup-architecture.md`, `architecture/08-alarm-ladder.md`,
`architecture/00-capability-map.md`, `scripts/wire_contract_gate.py` (allowlist **with a reason**),
`documentation/tests/r359-integrity-2026-08-30/` (**new**, 10 files).
## 3. Commits pushed to `main`
| repo | commit | contents |
|---|---|---|
| `felhom-controller` | **`b8af7276`** | all four fixes, all tests, controller docs |
| `felhom.eu` | **`e027b5d9`** | register closures, R-395 fix, architecture, capability map, evidence |
| `felhom.eu` | *(this report's commit)* | housekeeping compression + REPORT |
| `felhom-controller` | **`0d52a42c`** | v0.227.0 — the check, the job, the notifiers, the route, all tests |
| `felhom-controller` | **`45770f22`** | v0.227.1 — the classifier fix + CONTEXT/README/REUSE |
| `felhom.eu` | *(this report's commit)* | register, architecture, capability map, STATUS, evidence |
## 4. Tests and red-proofs
**All pass.** New tests, by group:
**All pass.** 22 new tests across 5 new files. Groups A (the check), B (the lock hazard), C (due-ness),
D (wiring), plus two fixtures built from restic's **real** bytes.
| test | result |
|---|---|
| `TestUnitRestoreOutcome_VolumesAndDatabaseNamed` (A1) | pass |
| `TestUnitRestoreOutcome_BackupHeldOnlySettings` (A2) | pass |
| `TestUnitRestoreOutcome_ManifestListedDataThatDidNotReturn` (A3) | pass |
| `TestUnitRestoreOutcome_DatabaseOnly` (A4) | pass |
| `TestR353_HandlerPublishesTheOutcome` (A5, the seam test) | pass |
| `TestR357_DestructiveRestoreRefusesWithoutHeadroom` (B1) | pass |
| `TestR357_UnknownSizeFailsClosed` (B2) | pass |
| `TestR357_UnknownFreeSpaceFailsClosed` (added — the mirror hole) | pass |
| `TestR357_AmpleSpaceIsUnchanged` (B3) | pass |
| `TestR358_FailedRestoreLeavesNoUsableScratch` (C1) | pass |
| `TestR358_UnitOnlyScratchIsNotFullReady` (C2) | pass |
| `TestR358_StaleMarkerIsClearedBeforeTheRun` (C3) | pass |
| `TestR358_UnreadableMarkerFailsClosed` (C4) | pass |
| `TestR358_WrongSchemaFailsClosed` (added) | pass |
| `TestR358_CompletedFullScratchStillReady` (C5) | pass |
| `TestR358_MarkerIsWrittenAt0600AndAtomically` (added) | pass |
| `TestR358_MarkerIsNeverPlaced` (D3) | pass |
| `TestR358_MarkerIsClearedBeforeResticAndWrittenAfter` (added, AST) | pass |
| `TestR358_PlaceHandlerRefusesIncompleteScratch` (D1) | pass |
| `TestR358_ReconstituteHandlerRefusesIncompleteScratch` (D2) | pass |
| `TestR358_UnitOnlyScratchClosesTheFullRestoreCard` (Scenario F, flow level) | pass |
| `TestR360_VerifyCopyDeleteRefusedDuringRestore` (D4) | pass |
| `TestR360_VerifyCopyDeleteStillWorksWhenIdle` (Scenario H) | pass |
**E1 — no existing test was modified for its content.** Every `TestReconstituteOutcome_*` passes
unmodified. The only test-file edits were mechanical call-site updates for the changed
`RestoreFromRecoveryUnit` signature (`err :=` → `_, err :=`) in three files.
### Red-proofs — each mutated, observed failing, reverted, `git diff` clean
| # | mutation | observed failure |
| red-proof | mutation | observed failure |
|---|---|---|
| **A5** | `EndRestoreOp(true, stackName+" visszaállítva ("+snapshotID+").")` restored | `THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)."` |
| **B1** | the whole R-357 gate deleted | `THE APP WAS STOPPED (1 call(s)) for a restore with 300 KB free for a 1 MB copy … (err=<nil>)` — and the same on both fail-closed tests |
| **C1** | the old `len(entries) > 0` check restored | `a part-copy was reported READY`, plus the unit-only and unreadable-marker cases |
| **D4** | the delete guard reverted to `IsRunning()` | `THE VERIFICATION COPY WAS DELETED while a restore was writing into it` |
| **D1/D2** | both server-side scratch refusals removed | both handlers redirected with „…elindult" over a part-copy |
| **A2** | classifier treats any non-zero as OK | `a repository restic said contains errors was reported as OK` |
| **B1** | remove the `acquireRunning` guard | `RESTIC WAS INVOKED while another operation held the single-writer flag: [… cat config] [… check]` |
| **C2** | gate due-ness on `time.Sunday` | `a 21-day-old check was NOT due on Saturday` |
| **D2** | remove the debug dispatch case | `POST /api/debug/backup/integrity was not dispatched` |
> **⚠ THE FIRST B1 RED-PROOF EXPOSED A HOLLOW TEST OF MY OWN, and it is recorded rather than quietly
> fixed.** With the gate removed, the run refused *earlier* — at the placement stat pre-pass — so
> `stops == 0` passed against the pre-fix code and the test proved nothing about the thing it exists
> for. Two corrections: the fixture now populates the scratch the way a completed download leaves it
> (the snapshot's own absolute paths mirrored under the scratch, plus `SetSafetyDumpFn` so no Docker
> is needed), and the assertions are **reordered** so a removed gate reports the outage rather than
> "no error returned". Only after that does the red-proof print `THE APP WAS STOPPED (1 call(s))`.
> The lesson is the doctrine's own: a test that cannot fail on the pre-fix shape is decoration.
**E1 — no existing test was modified.** `TestBackupReport_DeadFieldsStayZero` (R-331, D4) passes
unmodified, which is the check that I did not write to the retired report fields.
**Test count:** 24 new tests added across 4 new files. **Green gate:**
`go build ./... && go vet ./... && go test ./...` → **28 packages, rc 0**, no failures.
**All 12 controller gates OK.**
**Test count:** 28 packages green before and after; +22 tests. Green gate
`go build ./... && go vet ./... && go test ./...` → **rc 0**.
## 5. Deployed version
```
$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.226.1 Up 13 seconds (healthy)
gitea.dooplex.hu/admin/felhom-controller:0.227.1 Up 13 seconds (healthy)
[INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST
```
**A patch, not a rebuilt 0.226.0.** The live validation in §6 ran against **0.226.0**, which was
already on the box when §11 item 4's defect was found. Re-pushing a changed image under a tag already
running somewhere is the `:latest` hazard with extra steps — two images, one name, no way for a box to
tell which it has — so the fix shipped as **0.226.1** and the box was redeployed. **The §6 evidence
therefore describes 0.226.0**, and that is stated rather than quietly re-attributed; 0.226.1 differs
from it only by `CountsUnknown` and its two tests, neither of which touches any path §6 exercised.
**`demo-hp` only.** `demo-felhom` is on 0.226.1; the fleet floor and golden are 0.226.1.
**Fleet state at the end of the session — the debt named in §12 was PAID the same day.** Golden
**0.226.1** was baked, published, round-trip verified, **vouched** and the floor **raised to 0.226.1**.
Both demo machines now run 0.226.1, and `demo-felhom` got there by **self-update**, not by hand:
**v0.227.1 is a patch, not a rebuilt 0.227.0** — 0.227.0 was already running on `demo-hp` when the
classifier defect was found, and re-pushing changed bytes under a live tag is the `:latest` hazard.
The live validation in §8 ran against **0.227.0**; 0.227.1 differs only by the damage signatures and
their test.
```
[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)
```
## 6. The schedule I chose, and what I read to choose it
`golden_currency_gate.py` went **red → green** on the same command. Evidence:
`felhom.eu/documentation/tests/golden-0.226.1-2026-08-30/`.
Live on `demo-hp`: `db-dump 02:30` · `tier2-backup 03:30` · `fill-watch 03:30` · `metrics-prune 04:00`
· **`offbox-backup 04:15`** · `offsite-abandon-sweep 05:10` · whole-guest gate `[04:30, 08:30)`.
## 6. Live validation — endpoint level, on `demo-hp`
**Chose 06:00** — 1h45 clear of the off-site leg (which runs ~2m20s, measured) and 50 min clear of the
sweep. The backup window is customer-configurable, so **no fixed time is collision-proof for every
box**; a collision costs one skipped day, not a missed check, because due-ness makes tomorrow try
again. A slot that collided *every* night would be a real defect; this is not one.
**Method:** the exact endpoints the UI invokes, driven from inside guest 9201 (no browser on
DooPlex; the residual is client-side rendering only). No state was hand-set. Evidence copied off the
box at the end of the phase: `felhom.eu/documentation/audits/evidence-r353-r360-live-2026-08-30/`.
## 7. NOT yet live-validated
### The verbatim messages
1. **A real weekly firing.** It is a week away. The job is confirmed **registered** on the box, which
is a different and weaker claim, and the capability map says so.
2. **Every `--read-data-subset` claim about a LARGE store.** The curve in §9 is measured on 134.3 MB
and does not extrapolate.
3. **The failure branch end-to-end through my code.** restic's damage output was captured live and fed
to the classifier as a fixture (`TestR359_RealResticDamageOutputIsClassifiedAsDamage`), but
`CheckOffboxIntegrity` reads the *configured* target, so pointing it at the damaged scratch repo
would have meant repointing the live off-site target. **The seam is named rather than glossed:**
restic-catches-it and my-code-classifies-it are each proven; the join is not.
4. **The skip path via a real backup.** See §8 — the intended trigger does not exist (R-279).
**R-353** — `POST /backup/restore` for `opengist`, then the sentence read off the customer's own
wizard page (`GET /backups/restore/app?name=opengist`):
## 8. Part 5 and the live validation, in full
```
A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.
```
Evidence: `felhom.eu/documentation/tests/r359-integrity-2026-08-30/` (10 files, copied off **before**
teardown).
and the controller's own lines:
- **Scratch repo:** `/mnt/felhom-drives/hdd_1/r359-scratch-repo`, 3 files (195.4 KiB), snapshot
`3611a338`, hashes recorded.
- **Negative control FIRST:** healthy repo → `no errors were found`, exit 0, **703 ms**.
- **Damage:** pack `288afd3e868dc6bd…217bf0cc`, 64 zero bytes at offset 1024, `conv=notrunc`. Size
**unchanged** at 200 333 B; sha256 → `4b6847bb5eec…dc496d2b`.
- **Positive control:** the table at the top of this report.
- **The debug button that had never done anything:**
`{"data":{"duration_ms":35550,"ok":true,…},"message":"Az ellenőrzés rendben lezajlott","ok":true}`
- **R-397's notifier, end to end:**
`Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)`
— pushed to the hub, HTTP 200, severity `info`, so it mailed nobody.
- **THE HAZARD CONTROL, live.** The intended demonstration could not be run: `POST
/api/backup/offbox/run` returns **404** — there is no operator-triggerable off-site backup, which is
**R-279 and stays open**. So the same flag was exercised by its other holder — two checks 6 s apart:
```
B (while A held the flag): {"skipped":true,"skip_reason":"a backup or restore is already running","duration_ms":0}
A (completed): {"skipped":false,"ok":true,"duration_ms":34953}
```
**`duration_ms: 0` is the observable that matters: B never ran restic at all.**
```
[INFO] [backup] Restore-from-unit completed: opengist — 1 volume(s) of 1 listed, 0 database(s) of 0 listed
[INFO] [web] Restore completed (async): stack=opengist in 9.112668936s (volumes 1/1, dbs 0/0)
```
## 9. The three numbers R-399 needs — all MEASURED
**This is Scenario A, not Scenario B, and the substitution is stated rather than glossed.** The spec
named `opengist` as the data-less app from 21 August. **It is not one any more** — checked before
relying on it, as the spec instructed: every unit on `demo-hp` today lists at least one volume dump
(`opengist 1/0`, `privatebin 1/0`, `calibre-web 1/0`, the rest 2–3 volumes plus a database). **No app
on the box has the Scenario B shape**, and manufacturing one would mean falsifying a manifest — the
hand-set-state shortcut this project forbids. Scenario B is carried by
`TestUnitRestoreOutcome_BackupHeldOnlySettings` and the A5 seam test. What the live run *does* prove
is the whole path: real counts, correct clause selection (no database clause for 0 databases), and
the sentence reaching the customer's page.
| | |
|---|---|
| live store size | **140 829 678 B (134.3 MB)**, 2 651 blobs, 67 snapshots (`restic stats --mode raw-data`) |
| structure check wall-clock | **35.0 s** |
| 10% read-data | **35.9 s** (+0.9 s, **+3%**) |
| 50% | 37.3 s (+6%) |
| **100%** | **39.2 s (+12%)** |
**R-360** — a `mode=unit` scratch restore started, and the verification-copy delete POSTed while it
ran. The refusal, urldecoded from the `Location` header:
**Not an estimate — measured against the live store, read-only.** At this size, re-reading *all* the
data costs about four seconds more than reading none, because the wall clock is SFTP round-trips over
the WireGuard tunnel rather than transfer. **Stated assumption and its limit:** these do not
extrapolate — the structure check's cost tracks the INDEX, a read-data run's tracks the DATA, so a
50 GB store is ~370× the data and this curve says nothing about it.
```
/backups/restore?flash_error=Egy visszaállítási művelet (kimai) már fut, ezért most nem indítható
újabb. Az állapotát ezen az oldalon követheted; amint befejeződik, újra indíthatsz visszaállítást.
```
## 10. Teardown — all three layers
```
[WARN] [web] verification-copy delete refused for kimai: a backup/restore op is running
```
**And the consequence, which is the assertion that matters:** a canary file planted in the copy was
still there afterwards — `-rw-r--r-- 1 root root 10 Aug 30 17:34 canary.txt`, contents `defend-me`.
**R-358 / R-396** — after the `mode=unit` restore completed:
```
$ cat …/backups/offsite-restore/kimai/.felhom-restore-complete.json
{"schema":1,"snapshot_id":"84542ec8","full":false,"finished_at":"2026-08-30T17:34:42Z"}
mode=600 (no .tmp left behind)
[INFO] [offbox] kimai: scratch holds a UNIT-ONLY restore (snapshot 84542ec8) — not a full copy,
so place-to-live stays closed
```
**This is Scenario F proven live**, and it is the case that pre-fix would have unlocked the
destructive restore.
## 7. NOT yet live-validated — awaiting a supervised drill
1. **R-357's disk-full behaviour.** Filling a real filesystem to prove it is a drill step, not a build
step. Carried by the seam tests, whose central assertion is `StopStack` call count 0.
2. **R-353's Scenario B** (the "backup held only settings" sentence) — no app on `demo-hp` has that
unit shape any more; see §6.
3. **R-353's Scenario C** (the unit lists dumps, none return) — the R-367 stranded-dump shape was not
reproduced on live hardware; unit-tested only.
4. **R-358's failed-download branch** was not induced live. It was not needed: the unit-only branch
exercises the same marker gate through a real restore, without pointing restic at a bad snapshot.
## 8. Teardown
**This run provisioned nothing.** No machine, no guest, no hub record — all three layers:
- **Layer 1 (host):** nothing created on `felhom-pve` or `demo-hp`. No VM, no CT, no storage entry.
- **Layer 2 (guest):** two throwaway shell scripts were pushed into guest 9201 to drive the endpoints
and **both were removed**; one canary directory was planted under `backups/offsite-restore/kimai`
on the system drive to prove R-360 and **was removed**. The `mode=unit` scratch restore left a
normal verification copy on the HDD, which is ordinary product state a customer can delete.
- **Layer 1 (host):** nothing provisioned on `demo-hp`. No VM, no CT, no storage entry.
- **Layer 2 (guest):** the scratch restic repo was created and **removed** — **1 511 424 bytes**
returned to `/mnt/felhom-drives/hdd_1`; seven throwaway scripts in the guest and the container
removed; the `mode=unit` scratch from yesterday untouched.
- **Layer 3 (hub):** no customer, appliance or host record created, so none to discard.
`demo-felhom`, `ep0`, DooPlex and Peti's box were not touched.
## 9. Register
## 11. Register
**Size: 165 rows before → 167 after filing → 161 after housekeeping.**
**Size: 163 before → 165 after filing → 163 after housekeeping.**
- **Closed:** R-353, R-357, R-358, R-360 — each with shipping version and evidence path.
- **Filed and closed in the same session:** **R-396** (Scenario F's answer — see §11) and **R-395**
(`STATUS.md` contradicting itself; **fixed**, not merely recorded).
- **Re-ranked:** none.
- **Housekeeping:** the six closed rows were compressed into `CLOSED-ITEMS.md` keeping title,
version, evidence and every sentence that states a rule; the full original is
`git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md`. **No open row was touched.**
- **`ROADMAP.md` was not edited:** none of these four ever had a ROADMAP row — they live in the
register only. Stated rather than silently skipped.
- **Closed:** R-359, R-397 — with version and evidence path, then compressed into `CLOSED-ITEMS.md`.
- **CORRECTED, not closed: R-398** — see §12 item 1.
- **Filed:** **R-399** (the depth decision, now with three measured numbers and a re-framed premise)
and **R-400** (the debug buttons — see §12 item 2).
- **Explicitly left open and said so:** **R-87** (the restic tier is never restore-**tested** — a check
is not a restore-test, and the two rows are adjacent, which is exactly how they would get
conflated), **R-279** (no operator-triggerable off-site backup), **R-104**.
## 10. Documentation coupling (§5.5)
## 12. Observations
- **`00-capability-map.md`** — one new row, splitting the claim: R-353/R-358/R-360 **PROVEN-LIVE**
with the evidence citation; **R-357 IMPLEMENTED only**, with the reason.
- **`07-backup-architecture.md`** — §10.2 gained five rows (the four plus R-396). **§8 matrix row 3
KEEPS its PROVEN status**, with a note recording why: R-353 was a defect in the *message*, not the
mechanism.
- **`OPEN-ITEMS.md` / `CLOSED-ITEMS.md` / `STATUS.md`** — as §9.
- **`ROADMAP.md`** — no applicable rows.
- **Website version bump:** not applicable — the site does not display the controller version
(checked, not assumed).
1. **NOT-A-FINDING: acted on instead — R-398 was my own mistake and is CORRECTED, not closed.**
The row said `resticStep` is not a seam so no test can drive a restic-backed path. **The first half
is true and the conclusion was false:** `offboxRunner` / `SetOffboxRunner` / `m.runner()` has been
injectable since the off-site tier shipped, and `offbox_3a_test.go` has been using it five times
over. I read one function and generalised. **Part 0 was therefore not built, and building it would
have been actively harmful:** a `resticStepFn` seam replaces `resticStep`, hiding its
`unlock --remove-all` escalation from the very assertions that must observe it. What the row asked
for that *was* real is done — R-358's AST ordering test is now an execution test, which immediately
surfaced something the AST walk could not: `unlockStale` legitimately runs before the restore.
## 11. Observations — each with a register row, or a stated reason it needs none
2. **FILED: R-400 — and the sweep found EIGHT dead buttons, not one.** Comparing every
`/api/debug/...` reference in `debug.html` against every `subpath ==` case: **24 referenced, 17
dispatched.** The seven others are `backup/crossdrive`, `backup/infra`, `dr/infra-status`,
`hub/infra-push`, `storage/simulate-disconnect`, `storage/simulate-reconnect`,
`storage/watchdog-status`. Single dispatcher, exact-match switch, `default: http.NotFound` — so they
404 rather than silently succeed. **`controller/README.md` documented four `backup/*` debug routes
when two existed**; corrected. A third of a debug page does nothing, on the surface an operator
reaches for when something is already wrong.
1. **FILED: R-396.** *Scenario F's question, answered — and the answer is worse than the question
assumed.* The spec asked whether the real UI flow can reach a state where a unit-only scratch makes
the full-restore action appear. **It can, by the safest-looking action on the page.**
„Ellenőrző visszaállítás" (`mode=unit`, the default, advertised as non-destructive) calls
`RestoreOffboxScratch(full=false)`; `offboxRestoreScratchDir` **ignores `full`**, so both modes
write the same directory, and `--include` limits *what* restic extracts, never *where*; the wizard
sets `ScratchReady` from `OffboxFullScratchReady`; `deriveWizardStep` derives **both**
`PlaceEnabled` and `RestoreEnabled` from that one flag. R-358 as filed assumed the bad state needed
a *failed* download. It needed only a successful safe one. **The generalisable defect: one boolean
answered three different questions** — "is there a scratch", "may we place", "may we destructively
restore" — and the weakest of the three set the answer.
3. **NOT-A-FINDING: noted for R-359's own row, not a separate one.** `offbox_progress.go:328` carries a
near-identical lock-retry block to `resticStep`'s. Noted, **not merged** — §12 forbade it and the
duplication is load-bearing until someone proves otherwise.
2. **FILED: R-397.** *`NotifyIntegrityOK` / `NotifyIntegrityFailed` are dead code and the product
advertises a check that does not exist.* Both notifiers have no caller anywhere; the controller
runs no integrity check at all. **The part that actively misleads:** `config.Monitoring.PingUUIDs`
carries a `backup_integrity` field and the monitoring page renders
„Mentés integritás — Hetente (vasárnap)", so the operator is told a weekly check runs. Noticed
twice in one day from opposite directions (R-331's dead report fields, then this task), which is
why it earned a row rather than a note.
4. **NOT-A-FINDING: caught and fixed in the same session.** The damage classifier matched restic's
ordinary progress output (`"pack "` vs `check all packs`), so a check failing for a non-damage
reason would have told the customer their backups were corrupt. **The negative control caught it** —
which is precisely why a control that has only seen the failing case is worth nothing. Shipped as
v0.227.1.
3. **FILED: R-398.** *`resticStep` is not a seam, so no test can drive any restic-backed path.* Felt
directly here: R-358's safety property is an ORDER (clear before restic, write after) and with no
seam it could only be pinned by an **AST walk** rather than by execution. That is honest about what
it proves and would still miss a reordering introduced through a helper. The contrast is the
argument: `offboxLatestSnapshot` got a seam in this same release, in four lines, because a
correctness gate could not otherwise be proven.
5. **NOT-A-FINDING: my own measurement error, corrected before it reached a conclusion.** An early run
showed `exit=0` on the read-data subset forms while they printed `Fatal: repository contains
errors`. That was not a restic defect — the commands were piped through `tail`, so `$?` was tail's.
Re-measured without pipes: every read-data form exits 1. This project's own "exit codes that lie"
trap, caught by re-measuring rather than reasoning.
4. **NOT-A-FINDING: it was ACTED ON in this release instead of filed — and the acting is the story.**
`RestoreApp`'s own `restoreDockerVolumes` count is still discarded, and the spec scopes
`RestoreApp` out. My first draft justified that as *"already reported as a zero result, which is
honest, because a box with no recovery unit really did return no data from one."*
## 13. The push bypassed one gate, deliberately
> **⚠ THAT SENTENCE WAS FALSE, AND WRITING IT EXPOSED A DEFECT THE FIX ITSELF HAD INTRODUCED.**
> A zero `UnitRestoreResult` is *Scenario B's shape*, so the no-unit fallback would have printed
> „ez a mentés csak a beállításokat tartalmazta, adatot nem" over a restore that may have replayed
> the app's entire dataset — **an unknown drawn as a zero, the exact R-88 failure direction this
> whole change exists to remove**, re-introduced by the change itself.
>
> Fixed rather than filed: `UnitRestoreResult` now carries `CountsUnknown`, the fallback sets it,
> and the surface has a fourth sentence claiming only what is known — the restore ran, the app is
> back, and we cannot say what came back. **`RestoreApp`'s signature is untouched**, so §5 holds.
> Pinned by `TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty`; the A5 seam test was
> corrected too, because its fixture has no unit and therefore exercises exactly this path while
> asserting the wrong sentence.
>
> **It was the `observations` gate refusing the push that forced the re-read.** The gate exists to
> stop findings dying in an overwritten REPORT.md; here it caught a live defect instead.
**`felhom.eu` — `golden-currency` CONVICTED.** v0.227.0/0.227.1 are released and the vouched golden
carries 0.226.1. **The gate is right.** A **BYPASS, not a waiver**, and the spec directs it: §13 says
golden and fleet delivery are Viktor's (R-242). **A golden carrying 0.227.1 is OWED**, and it is item 3
under "Waiting on you" in `STATUS.md`.
5. **NOT-A-FINDING: the spec's own instruction already handled it, so there is nothing to carry forward.** §13 step 1 named `opengist` as
the data-less app; it is not one any more. Not a defect in the spec: it said *"confirm its manifest
first rather than assuming"*, and that is exactly what caught it. Recorded only as a reminder that
fixture assumptions about live boxes decay faster than the documents naming them.
## 12. Two gates fired; one push was bypassed, and both are declared
**`felhom.eu` — `golden-currency` CONVICTED, and the debt is now PAID.** Three controller releases
shipped today and the vouched golden still carried **0.223.0**, so a machine installed at that moment
would have received none of them. **The gate was right.** Those pushes used `git push --no-verify` — a
**BYPASS, not a waiver**, on the operator's standing ruling, re-checked for this release rather than
reused blindly.
**Later the same day the operator asked for the bake, and it was done: golden 0.226.1 baked, published,
round-trip verified, vouched, and the floor raised.** The gate went **red → green**, so those bypasses
are **historical rather than standing**. See §5 for the fleet state and
`felhom.eu/documentation/tests/golden-0.226.1-2026-08-30/` for the evidence, including the round trip
and the proven token-leak grep.
**`felhom-controller` — nothing was bypassed.** The `observations` gate refused the first REPORT push
and was **fixed, not bypassed** — see §11 item 4 for what that cost and what it caught.
**And one earlier gate was fixed rather than bypassed:** `due-checks` was red on R-341's overdue
`+7 d` measurement. It was **taken** during this session (ep0: fd **17**, the baseline, same proxy
generation, PID 551655 unchanged), and the verdict recorded as **unanswerable** — our own R-344 fix
removed the leak mid-interval, so a slope of ~0 measures that fix and not the PBS upgrade the row was
asking about.
**`wire-contract` was NOT bypassed — it caught something real and was answered.** It convicted
`offsite.last_integrity_ok`: the controller emits it and no hub struct can decode it. That is the
"emitting into the void" shape. It is now **allowlisted with a written reason** rather than skipped:
building a hub display is a hub change, and R-331 ruled that class a decision for the operator — the
*previous* integrity display was removed precisely because it rendered a value nothing wrote. The entry
says to delete it when a surface is built.