The headline is not the feature, it is what measuring it revealed: THE STRUCTURE
CHECK THAT SHIPS ON DOES NOT CATCH SILENT CORRUPTION. A pack corrupted without a
size change returned `no errors were found`, exit 0. Only --read-data caught it.
So R-399 is not merely a bandwidth question -- at the shipped default a class of
damage is not checked at all.
The three numbers R-399 needed are MEASURED, not estimated: store 134.3 MB / 67
snapshots; structure check 35.0 s; and the full curve 10% 35.9 s, 50% 37.3 s,
100% 39.2 s. At this size re-reading everything costs four seconds more than
reading none. Stated limit: they do not extrapolate.
Records four things that went wrong and were caught rather than shipped:
- R-398 was MY OWN mistaken row. The seam already existed, Part 0 was not
built, and building it would have HIDDEN the unlock --remove-all escalation
from the assertions that must see it. Corrected, not closed.
- the damage classifier matched restic's ORDINARY progress output; the
negative control caught it (v0.227.1).
- an exit code I misread through a pipe, corrected by re-measuring.
- wire-contract convicted a field the hub cannot decode; allowlisted WITH A
REASON rather than skipped, because building the hub display is a decision
R-331 already ruled belongs to the operator.
And the sweep the task asked for: EIGHT debug buttons post to endpoints that do
not exist, not one. 24 referenced, 17 dispatched. Filed as R-400.
Not live-validated and each says why: the weekly firing (a week away), read-data
on a large store, and the join between "restic catches it" and "my code
classifies it" -- both proven, the join is not, and the seam is named.
14 KiB
REPORT — R-359 / R-397: the off-site store gets checked
Controller v0.227.0 → v0.227.1 · 2026-08-30 · implemented on DooPlex, validated live on demo-hp
⚠ READ THIS FIRST — the measurement changed what the feature is worth
Part 5 corrupted one pack of a throwaway repository without changing its size (64 zero bytes at offset 1024 — the subtlest form of rot). Both depths were run against it:
| depth | exit | verdict |
|---|---|---|
restic check — the depth that ships ON |
0 | no errors were found |
every --read-data* form — ships OFF |
1 | Pack ID does not match, want 288afd3e…, got 4b6847bb… |
The structure check reported a corrupted store as healthy. It verifies the index, the pack inventory and the snapshot graph — it catches missing packs, broken indexes and unreadable snapshots, which are real — but it does not re-hash pack contents. R-399 is therefore not only a bandwidth question: at the shipped default a class of damage is not checked at all, and it is the class that silently eats a customer's photos. The task could not have known this; nothing had measured it.
1. Confirmed baselines
| repo | spec baseline | actual at start | version |
|---|---|---|---|
felhom-controller |
c476ae51 |
e64c84ae (drift: one docs-only commit — my own golden-bake report) |
v0.226.1 → v0.227.1 |
felhom.eu |
4f875174 |
4f875174 — exact match |
docs only |
felhom-agent |
not touched | v0.130.0 | unchanged |
Both trees clean and equal to origin/main. Every symbol in §5 was re-verified present (20/20),
and all three claimed absences confirmed: 0 occurrences of a restic check verb, 0 callers of
NotifyIntegrity* in the controller, and the debug button present with no dispatch case.
MinAgent stays 0.129.0.
2. Files created / modified
felhom-controller: controller/internal/backup/offbox_integrity.go (new),
offbox.go (report fields), offbox_restore.go+r358_scratch_marker_test.go (AST→execution test),
internal/settings/settings.go, internal/config/config.go,
cmd/controller/main.go (job + the one caller + the debug callback),
internal/web/handler_debug.go, internal/web/handlers.go,
tests: r359_integrity_test.go, r359_lock_safety_test.go, r359_dueness_test.go (new),
cmd/controller/r359_wiring_test.go, internal/web/r397_debug_integrity_test.go (new);
docs: CHANGELOG.md, CONTEXT.md, REUSE.md, controller/README.md, REPORT.md.
felhom.eu: STATUS.md, backlog/OPEN-ITEMS.md, backlog/CLOSED-ITEMS.md,
architecture/07-backup-architecture.md, architecture/08-alarm-ladder.md,
architecture/00-capability-map.md, scripts/wire_contract_gate.py (allowlist with a reason),
documentation/tests/r359-integrity-2026-08-30/ (new, 10 files).
3. Commits pushed to main
| repo | commit | contents |
|---|---|---|
felhom-controller |
0d52a42c |
v0.227.0 — the check, the job, the notifiers, the route, all tests |
felhom-controller |
45770f22 |
v0.227.1 — the classifier fix + CONTEXT/README/REUSE |
felhom.eu |
(this report's commit) | register, architecture, capability map, STATUS, evidence |
4. Tests and red-proofs
All pass. 22 new tests across 5 new files. Groups A (the check), B (the lock hazard), C (due-ness), D (wiring), plus two fixtures built from restic's real bytes.
| red-proof | mutation | observed failure |
|---|---|---|
| A2 | classifier treats any non-zero as OK | a repository restic said contains errors was reported as OK |
| B1 | remove the acquireRunning guard |
RESTIC WAS INVOKED while another operation held the single-writer flag: [… cat config] [… check] |
| C2 | gate due-ness on time.Sunday |
a 21-day-old check was NOT due on Saturday |
| D2 | remove the debug dispatch case | POST /api/debug/backup/integrity was not dispatched |
E1 — no existing test was modified. TestBackupReport_DeadFieldsStayZero (R-331, D4) passes
unmodified, which is the check that I did not write to the retired report fields.
Test count: 28 packages green before and after; +22 tests. Green gate
go build ./... && go vet ./... && go test ./... → rc 0.
5. Deployed version
gitea.dooplex.hu/admin/felhom-controller:0.227.1 Up 13 seconds (healthy)
[INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST
demo-hp only. demo-felhom is on 0.226.1; the fleet floor and golden are 0.226.1.
v0.227.1 is a patch, not a rebuilt 0.227.0 — 0.227.0 was already running on demo-hp when the
classifier defect was found, and re-pushing changed bytes under a live tag is the :latest hazard.
The live validation in §8 ran against 0.227.0; 0.227.1 differs only by the damage signatures and
their test.
6. The schedule I chose, and what I read to choose it
Live on demo-hp: db-dump 02:30 · tier2-backup 03:30 · fill-watch 03:30 · metrics-prune 04:00
· offbox-backup 04:15 · offsite-abandon-sweep 05:10 · whole-guest gate [04:30, 08:30).
Chose 06:00 — 1h45 clear of the off-site leg (which runs ~2m20s, measured) and 50 min clear of the sweep. The backup window is customer-configurable, so no fixed time is collision-proof for every box; a collision costs one skipped day, not a missed check, because due-ness makes tomorrow try again. A slot that collided every night would be a real defect; this is not one.
7. NOT yet live-validated
- A real weekly firing. It is a week away. The job is confirmed registered on the box, which is a different and weaker claim, and the capability map says so.
- Every
--read-data-subsetclaim about a LARGE store. The curve in §9 is measured on 134.3 MB and does not extrapolate. - The failure branch end-to-end through my code. restic's damage output was captured live and fed
to the classifier as a fixture (
TestR359_RealResticDamageOutputIsClassifiedAsDamage), butCheckOffboxIntegrityreads the configured target, so pointing it at the damaged scratch repo would have meant repointing the live off-site target. The seam is named rather than glossed: restic-catches-it and my-code-classifies-it are each proven; the join is not. - The skip path via a real backup. See §8 — the intended trigger does not exist (R-279).
8. Part 5 and the live validation, in full
Evidence: felhom.eu/documentation/tests/r359-integrity-2026-08-30/ (10 files, copied off before
teardown).
- Scratch repo:
/mnt/felhom-drives/hdd_1/r359-scratch-repo, 3 files (195.4 KiB), snapshot3611a338, hashes recorded. - Negative control FIRST: healthy repo →
no errors were found, exit 0, 703 ms. - Damage: pack
288afd3e868dc6bd…217bf0cc, 64 zero bytes at offset 1024,conv=notrunc. Size unchanged at 200 333 B; sha256 →4b6847bb5eec…dc496d2b. - Positive control: the table at the top of this report.
- The debug button that had never done anything:
{"data":{"duration_ms":35550,"ok":true,…},"message":"Az ellenőrzés rendben lezajlott","ok":true} - R-397's notifier, end to end:
Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)— pushed to the hub, HTTP 200, severityinfo, so it mailed nobody. - THE HAZARD CONTROL, live. The intended demonstration could not be run:
POST /api/backup/offbox/runreturns 404 — there is no operator-triggerable off-site backup, which is R-279 and stays open. So the same flag was exercised by its other holder — two checks 6 s apart:B (while A held the flag): {"skipped":true,"skip_reason":"a backup or restore is already running","duration_ms":0} A (completed): {"skipped":false,"ok":true,"duration_ms":34953}duration_ms: 0is the observable that matters: B never ran restic at all.
9. The three numbers R-399 needs — all MEASURED
| live store size | 140 829 678 B (134.3 MB), 2 651 blobs, 67 snapshots (restic stats --mode raw-data) |
| structure check wall-clock | 35.0 s |
| 10% read-data | 35.9 s (+0.9 s, +3%) |
| 50% | 37.3 s (+6%) |
| 100% | 39.2 s (+12%) |
Not an estimate — measured against the live store, read-only. At this size, re-reading all the data costs about four seconds more than reading none, because the wall clock is SFTP round-trips over the WireGuard tunnel rather than transfer. Stated assumption and its limit: these do not extrapolate — the structure check's cost tracks the INDEX, a read-data run's tracks the DATA, so a 50 GB store is ~370× the data and this curve says nothing about it.
10. Teardown — all three layers
- Layer 1 (host): nothing provisioned on
demo-hp. No VM, no CT, no storage entry. - Layer 2 (guest): the scratch restic repo was created and removed — 1 511 424 bytes
returned to
/mnt/felhom-drives/hdd_1; seven throwaway scripts in the guest and the container removed; themode=unitscratch from yesterday untouched. - Layer 3 (hub): no customer, appliance or host record created, so none to discard.
demo-felhom, ep0, DooPlex and Peti's box were not touched.
11. Register
Size: 163 before → 165 after filing → 163 after housekeeping.
- Closed: R-359, R-397 — with version and evidence path, then compressed into
CLOSED-ITEMS.md. - CORRECTED, not closed: R-398 — see §12 item 1.
- Filed: R-399 (the depth decision, now with three measured numbers and a re-framed premise) and R-400 (the debug buttons — see §12 item 2).
- Explicitly left open and said so: R-87 (the restic tier is never restore-tested — a check is not a restore-test, and the two rows are adjacent, which is exactly how they would get conflated), R-279 (no operator-triggerable off-site backup), R-104.
12. Observations
-
NOT-A-FINDING: acted on instead — R-398 was my own mistake and is CORRECTED, not closed. The row said
resticStepis not a seam so no test can drive a restic-backed path. The first half is true and the conclusion was false:offboxRunner/SetOffboxRunner/m.runner()has been injectable since the off-site tier shipped, andoffbox_3a_test.gohas been using it five times over. I read one function and generalised. Part 0 was therefore not built, and building it would have been actively harmful: aresticStepFnseam replacesresticStep, hiding itsunlock --remove-allescalation from the very assertions that must observe it. What the row asked for that was real is done — R-358's AST ordering test is now an execution test, which immediately surfaced something the AST walk could not:unlockStalelegitimately runs before the restore. -
FILED: R-400 — and the sweep found EIGHT dead buttons, not one. Comparing every
/api/debug/...reference indebug.htmlagainst everysubpath ==case: 24 referenced, 17 dispatched. The seven others arebackup/crossdrive,backup/infra,dr/infra-status,hub/infra-push,storage/simulate-disconnect,storage/simulate-reconnect,storage/watchdog-status. Single dispatcher, exact-match switch,default: http.NotFound— so they 404 rather than silently succeed.controller/README.mddocumented fourbackup/*debug routes when two existed; corrected. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. -
NOT-A-FINDING: noted for R-359's own row, not a separate one.
offbox_progress.go:328carries a near-identical lock-retry block toresticStep's. Noted, not merged — §12 forbade it and the duplication is load-bearing until someone proves otherwise. -
NOT-A-FINDING: caught and fixed in the same session. The damage classifier matched restic's ordinary progress output (
"pack "vscheck all packs), so a check failing for a non-damage reason would have told the customer their backups were corrupt. The negative control caught it — which is precisely why a control that has only seen the failing case is worth nothing. Shipped as v0.227.1. -
NOT-A-FINDING: my own measurement error, corrected before it reached a conclusion. An early run showed
exit=0on the read-data subset forms while they printedFatal: repository contains errors. That was not a restic defect — the commands were piped throughtail, so$?was tail's. Re-measured without pipes: every read-data form exits 1. This project's own "exit codes that lie" trap, caught by re-measuring rather than reasoning.
13. The push bypassed one gate, deliberately
felhom.eu — golden-currency CONVICTED. v0.227.0/0.227.1 are released and the vouched golden
carries 0.226.1. The gate is right. A BYPASS, not a waiver, and the spec directs it: §13 says
golden and fleet delivery are Viktor's (R-242). A golden carrying 0.227.1 is OWED, and it is item 3
under "Waiting on you" in STATUS.md.
wire-contract was NOT bypassed — it caught something real and was answered. It convicted
offsite.last_integrity_ok: the controller emits it and no hub struct can decode it. That is the
"emitting into the void" shape. It is now allowlisted with a written reason rather than skipped:
building a hub display is a hub change, and R-331 ruled that class a decision for the operator — the
previous integrity display was removed precisely because it rendered a value nothing wrote. The entry
says to delete it when a surface is built.