# REPORT — R-359 / R-397: the off-site store gets checked **Controller v0.227.0 → v0.227.1 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`** --- ## ⚠ READ THIS FIRST — the measurement changed what the feature is worth Part 5 corrupted one pack of a throwaway repository **without changing its size** (64 zero bytes at offset 1024 — the subtlest form of rot). Both depths were run against it: | depth | exit | verdict | |---|---|---| | `restic check` — **the depth that ships ON** | **0** | **`no errors were found`** | | every `--read-data*` form — **ships OFF** | 1 | `Pack ID does not match, want 288afd3e…, got 4b6847bb…` | **The structure check reported a corrupted store as healthy.** It verifies the index, the pack inventory and the snapshot graph — it catches missing packs, broken indexes and unreadable snapshots, which are real — but it does **not** re-hash pack contents. **R-399 is therefore not only a bandwidth question: at the shipped default a class of damage is not checked at all**, and it is the class that silently eats a customer's photos. The task could not have known this; nothing had measured it. ## 1. Confirmed baselines | repo | spec baseline | actual at start | version | |---|---|---|---| | `felhom-controller` | `c476ae51` | **`e64c84ae`** (drift: one docs-only commit — my own golden-bake report) | v0.226.1 → **v0.227.1** | | `felhom.eu` | `4f875174` | `4f875174` — **exact match** | docs only | | `felhom-agent` | not touched | v0.130.0 | unchanged | Both trees clean and equal to `origin/main`. **Every symbol in §5 was re-verified present** (20/20), and all three claimed absences confirmed: 0 occurrences of a restic `check` verb, 0 callers of `NotifyIntegrity*` in the controller, and the debug button present with no dispatch case. `MinAgent` stays **0.129.0**. ## 2. Files created / modified **`felhom-controller`:** `controller/internal/backup/offbox_integrity.go` (**new**), `offbox.go` (report fields), `offbox_restore.go`+`r358_scratch_marker_test.go` (AST→execution test), `internal/settings/settings.go`, `internal/config/config.go`, `cmd/controller/main.go` (job + the one caller + the debug callback), `internal/web/handler_debug.go`, `internal/web/handlers.go`, tests: `r359_integrity_test.go`, `r359_lock_safety_test.go`, `r359_dueness_test.go` (**new**), `cmd/controller/r359_wiring_test.go`, `internal/web/r397_debug_integrity_test.go` (**new**); docs: `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md`. **`felhom.eu`:** `STATUS.md`, `backlog/OPEN-ITEMS.md`, `backlog/CLOSED-ITEMS.md`, `architecture/07-backup-architecture.md`, `architecture/08-alarm-ladder.md`, `architecture/00-capability-map.md`, `scripts/wire_contract_gate.py` (allowlist **with a reason**), `documentation/tests/r359-integrity-2026-08-30/` (**new**, 10 files). ## 3. Commits pushed to `main` | repo | commit | contents | |---|---|---| | `felhom-controller` | **`0d52a42c`** | v0.227.0 — the check, the job, the notifiers, the route, all tests | | `felhom-controller` | **`45770f22`** | v0.227.1 — the classifier fix + CONTEXT/README/REUSE | | `felhom.eu` | *(this report's commit)* | register, architecture, capability map, STATUS, evidence | ## 4. Tests and red-proofs **All pass.** 22 new tests across 5 new files. Groups A (the check), B (the lock hazard), C (due-ness), D (wiring), plus two fixtures built from restic's **real** bytes. | red-proof | mutation | observed failure | |---|---|---| | **A2** | classifier treats any non-zero as OK | `a repository restic said contains errors was reported as OK` | | **B1** | remove the `acquireRunning` guard | `RESTIC WAS INVOKED while another operation held the single-writer flag: [… cat config] [… check]` | | **C2** | gate due-ness on `time.Sunday` | `a 21-day-old check was NOT due on Saturday` | | **D2** | remove the debug dispatch case | `POST /api/debug/backup/integrity was not dispatched` | **E1 — no existing test was modified.** `TestBackupReport_DeadFieldsStayZero` (R-331, D4) passes unmodified, which is the check that I did not write to the retired report fields. **Test count:** 28 packages green before and after; +22 tests. Green gate `go build ./... && go vet ./... && go test ./...` → **rc 0**. ## 5. Deployed version ``` gitea.dooplex.hu/admin/felhom-controller:0.227.1 Up 13 seconds (healthy) [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST ``` **`demo-hp` only.** `demo-felhom` is on 0.226.1; the fleet floor and golden are 0.226.1. **v0.227.1 is a patch, not a rebuilt 0.227.0** — 0.227.0 was already running on `demo-hp` when the classifier defect was found, and re-pushing changed bytes under a live tag is the `:latest` hazard. The live validation in §8 ran against **0.227.0**; 0.227.1 differs only by the damage signatures and their test. ## 6. The schedule I chose, and what I read to choose it Live on `demo-hp`: `db-dump 02:30` · `tier2-backup 03:30` · `fill-watch 03:30` · `metrics-prune 04:00` · **`offbox-backup 04:15`** · `offsite-abandon-sweep 05:10` · whole-guest gate `[04:30, 08:30)`. **Chose 06:00** — 1h45 clear of the off-site leg (which runs ~2m20s, measured) and 50 min clear of the sweep. The backup window is customer-configurable, so **no fixed time is collision-proof for every box**; a collision costs one skipped day, not a missed check, because due-ness makes tomorrow try again. A slot that collided *every* night would be a real defect; this is not one. ## 7. NOT yet live-validated 1. **A real weekly firing.** It is a week away. The job is confirmed **registered** on the box, which is a different and weaker claim, and the capability map says so. 2. **Every `--read-data-subset` claim about a LARGE store.** The curve in §9 is measured on 134.3 MB and does not extrapolate. 3. **The failure branch end-to-end through my code.** restic's damage output was captured live and fed to the classifier as a fixture (`TestR359_RealResticDamageOutputIsClassifiedAsDamage`), but `CheckOffboxIntegrity` reads the *configured* target, so pointing it at the damaged scratch repo would have meant repointing the live off-site target. **The seam is named rather than glossed:** restic-catches-it and my-code-classifies-it are each proven; the join is not. 4. **The skip path via a real backup.** See §8 — the intended trigger does not exist (R-279). ## 8. Part 5 and the live validation, in full Evidence: `felhom.eu/documentation/tests/r359-integrity-2026-08-30/` (10 files, copied off **before** teardown). - **Scratch repo:** `/mnt/felhom-drives/hdd_1/r359-scratch-repo`, 3 files (195.4 KiB), snapshot `3611a338`, hashes recorded. - **Negative control FIRST:** healthy repo → `no errors were found`, exit 0, **703 ms**. - **Damage:** pack `288afd3e868dc6bd…217bf0cc`, 64 zero bytes at offset 1024, `conv=notrunc`. Size **unchanged** at 200 333 B; sha256 → `4b6847bb5eec…dc496d2b`. - **Positive control:** the table at the top of this report. - **The debug button that had never done anything:** `{"data":{"duration_ms":35550,"ok":true,…},"message":"Az ellenőrzés rendben lezajlott","ok":true}` - **R-397's notifier, end to end:** `Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)` — pushed to the hub, HTTP 200, severity `info`, so it mailed nobody. - **THE HAZARD CONTROL, live.** The intended demonstration could not be run: `POST /api/backup/offbox/run` returns **404** — there is no operator-triggerable off-site backup, which is **R-279 and stays open**. So the same flag was exercised by its other holder — two checks 6 s apart: ``` B (while A held the flag): {"skipped":true,"skip_reason":"a backup or restore is already running","duration_ms":0} A (completed): {"skipped":false,"ok":true,"duration_ms":34953} ``` **`duration_ms: 0` is the observable that matters: B never ran restic at all.** ## 9. The three numbers R-399 needs — all MEASURED | | | |---|---| | live store size | **140 829 678 B (134.3 MB)**, 2 651 blobs, 67 snapshots (`restic stats --mode raw-data`) | | structure check wall-clock | **35.0 s** | | 10% read-data | **35.9 s** (+0.9 s, **+3%**) | | 50% | 37.3 s (+6%) | | **100%** | **39.2 s (+12%)** | **Not an estimate — measured against the live store, read-only.** At this size, re-reading *all* the data costs about four seconds more than reading none, because the wall clock is SFTP round-trips over the WireGuard tunnel rather than transfer. **Stated assumption and its limit:** these do not extrapolate — the structure check's cost tracks the INDEX, a read-data run's tracks the DATA, so a 50 GB store is ~370× the data and this curve says nothing about it. ## 10. Teardown — all three layers - **Layer 1 (host):** nothing provisioned on `demo-hp`. No VM, no CT, no storage entry. - **Layer 2 (guest):** the scratch restic repo was created and **removed** — **1 511 424 bytes** returned to `/mnt/felhom-drives/hdd_1`; seven throwaway scripts in the guest and the container removed; the `mode=unit` scratch from yesterday untouched. - **Layer 3 (hub):** no customer, appliance or host record created, so none to discard. `demo-felhom`, `ep0`, DooPlex and Peti's box were not touched. ## 11. Register **Size: 163 before → 165 after filing → 163 after housekeeping.** - **Closed:** R-359, R-397 — with version and evidence path, then compressed into `CLOSED-ITEMS.md`. - **CORRECTED, not closed: R-398** — see §12 item 1. - **Filed:** **R-399** (the depth decision, now with three measured numbers and a re-framed premise) and **R-400** (the debug buttons — see §12 item 2). - **Explicitly left open and said so:** **R-87** (the restic tier is never restore-**tested** — a check is not a restore-test, and the two rows are adjacent, which is exactly how they would get conflated), **R-279** (no operator-triggerable off-site backup), **R-104**. ## 12. Observations 1. **NOT-A-FINDING: acted on instead — R-398 was my own mistake and is CORRECTED, not closed.** The row said `resticStep` is not a seam so no test can drive a restic-backed path. **The first half is true and the conclusion was false:** `offboxRunner` / `SetOffboxRunner` / `m.runner()` has been injectable since the off-site tier shipped, and `offbox_3a_test.go` has been using it five times over. I read one function and generalised. **Part 0 was therefore not built, and building it would have been actively harmful:** a `resticStepFn` seam replaces `resticStep`, hiding its `unlock --remove-all` escalation from the very assertions that must observe it. What the row asked for that *was* real is done — R-358's AST ordering test is now an execution test, which immediately surfaced something the AST walk could not: `unlockStale` legitimately runs before the restore. 2. **FILED: R-400 — and the sweep found EIGHT dead buttons, not one.** Comparing every `/api/debug/...` reference in `debug.html` against every `subpath ==` case: **24 referenced, 17 dispatched.** The seven others are `backup/crossdrive`, `backup/infra`, `dr/infra-status`, `hub/infra-push`, `storage/simulate-disconnect`, `storage/simulate-reconnect`, `storage/watchdog-status`. Single dispatcher, exact-match switch, `default: http.NotFound` — so they 404 rather than silently succeed. **`controller/README.md` documented four `backup/*` debug routes when two existed**; corrected. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. 3. **NOT-A-FINDING: noted for R-359's own row, not a separate one.** `offbox_progress.go:328` carries a near-identical lock-retry block to `resticStep`'s. Noted, **not merged** — §12 forbade it and the duplication is load-bearing until someone proves otherwise. 4. **NOT-A-FINDING: caught and fixed in the same session.** The damage classifier matched restic's ordinary progress output (`"pack "` vs `check all packs`), so a check failing for a non-damage reason would have told the customer their backups were corrupt. **The negative control caught it** — which is precisely why a control that has only seen the failing case is worth nothing. Shipped as v0.227.1. 5. **NOT-A-FINDING: my own measurement error, corrected before it reached a conclusion.** An early run showed `exit=0` on the read-data subset forms while they printed `Fatal: repository contains errors`. That was not a restic defect — the commands were piped through `tail`, so `$?` was tail's. Re-measured without pipes: every read-data form exits 1. This project's own "exit codes that lie" trap, caught by re-measuring rather than reasoning. ## 13. The push bypassed one gate, deliberately **`felhom.eu` — `golden-currency` CONVICTED.** v0.227.0/0.227.1 are released and the vouched golden carries 0.226.1. **The gate is right.** A **BYPASS, not a waiver**, and the spec directs it: §13 says golden and fleet delivery are Viktor's (R-242). **A golden carrying 0.227.1 is OWED**, and it is item 3 under "Waiting on you" in `STATUS.md`. **`wire-contract` was NOT bypassed — it caught something real and was answered.** It convicted `offsite.last_integrity_ok`: the controller emits it and no hub struct can decode it. That is the "emitting into the void" shape. It is now **allowlisted with a written reason** rather than skipped: building a hub display is a hub change, and R-331 ruled that class a decision for the operator — the *previous* integrity display was removed precisely because it rendered a value nothing wrote. The entry says to delete it when a surface is built.