# R-359 / R-397 — the off-site integrity check, validated live on `demo-hp` (2026-08-30) Controller **v0.227.0**. Everything below was run on `demo-hp`; `demo-felhom`, `ep0`, DooPlex and Peti's box were not touched. --- ## ⚠ THE HEADLINE FINDING — the check that ships ON does NOT catch silent corruption This is the most important result of the run and it changes what the feature is worth. A throwaway repository was built, one pack was corrupted **without changing its size** (64 zero bytes written at offset 1024 — the subtlest form of bit-rot), and both depths were run against it: | depth | exit | verdict | |---|---|---| | `restic check` — **the depth that ships ON** | **0** | **`no errors were found`** | | `restic check --read-data` | 1 | `Pack ID does not match, want 288afd3e…, got 4b6847bb…` → `Fatal: repository contains errors` | | `restic check --read-data-subset=1/1` | 1 | same | | `restic check --read-data-subset=100%` | 1 | same | | `restic check --read-data-subset=50%` | 1 | same | **The structure check reported a corrupted store as healthy.** It verifies the index, the pack inventory and the snapshot graph — real failure modes, and it catches missing packs, broken indexes and unreadable snapshots. It does **not** re-hash pack contents, so it cannot see rot inside a pack that is still the right size. **What this means for R-399, and it is not what the task assumed.** R-399 was framed as a bandwidth and cadence question. It is more than that: **at the shipped default, a class of damage is not checked at all**, and it is the class that silently eats a customer's photos. The numbers below make the decision much easier than expected. > **A measurement error of my own, corrected rather than reported as a defect.** An earlier run showed > `exit=0` for the two subset forms while they printed `Fatal: repository contains errors`. That was > not restic: the commands were piped through `tail`, so `$?` was **tail's** exit code. Re-measured > without pipes, every read-data form exits **1**. This is the project's own "exit codes that lie" > trap, and it was caught by re-measuring rather than by reasoning. ## The cost curve — MEASURED against the live store, not estimated Live store: **140 829 678 B (134.3 MB)**, 2 651 blobs, **67 snapshots** (`restic stats --mode raw-data`). | depth | wall-clock | over structure-only | |---|---|---| | structure only (**ships ON**) | **35.0 s** | — | | `--read-data-subset=10%` | 35.9 s | **+0.9 s (+3%)** | | `--read-data-subset=50%` | 37.3 s | +2.2 s (+6%) | | `--read-data-subset=100%` | **39.2 s** | **+4.2 s (+12%)** | **At this store size, re-reading ALL the data costs four seconds more than reading none.** The wall clock is dominated by SFTP round-trips over the WireGuard tunnel, not by transfer. **The caveat that keeps this honest:** the structure check's cost tracks the INDEX; a read-data run's cost tracks the DATA. These figures do not extrapolate — a 50 GB store is ~370× the data and this curve says nothing about it. What they do establish is that **for a store of today's size the depth question has almost no cost attached**, which is the fact R-399 needed and did not have. ## Part 5, step by step 1. **Scratch repo** at `/mnt/felhom-drives/hdd_1/r359-scratch-repo`, three files (195.4 KiB), snapshot `3611a338`. Hashes recorded in `step1`–`step2`. 2. **Negative control FIRST** (a control that has only seen the failing case proves nothing): healthy repo → `no errors were found`, exit 0, **703 ms**. 3. **Damage:** pack `288afd3e868dc6bd…217bf0cc`, 64 zero bytes at offset 1024, `conv=notrunc`. Size **unchanged** at 200 333 B; sha256 moved to `4b6847bb5eec…dc496d2b`. 4. **Positive control:** see the table above. 5. **Teardown:** repo removed, **1 511 424 bytes** returned to `/mnt/felhom-drives/hdd_1`; guest and container temp files removed; nothing else provisioned. ## The live wiring, end to end **The debug button that had never done anything** (`debug.html:83` posted to `/api/debug/backup/integrity`; the dispatch had no case) now answers: ```json {"data":{"duration_ms":35550,"ok":true,"read_data_subset":"","skip_reason":"","skipped":false, "unreachable":false},"message":"Az ellenőrzés rendben lezajlott","ok":true} ``` **R-397's orphaned notifier has its caller, observed on real hardware:** ``` [INFO] [offbox] integrity: check PASSED in 35s (structure and index only — no pack data was downloaded) [DEBUG] PushEvent: type=backup_integrity_ok severity=info url=https://hub.felhom.eu/api/v1/event [DEBUG] PushEvent: backup_integrity_ok pushed OK (HTTP 200) [INFO] Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s) ``` Severity `info`, which `severityNotifies` drops before either leg — so it **mails nobody, by design**. **The scheduled job is registered live:** ``` [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST daily job registered: name="offsite-integrity" schedule="06:00" nextRun=2026-08-31T06:00:00+02:00 ``` ## THE HAZARD CONTROL, observed live `resticStep` escalates to **`unlock --remove-all`** on a lock error, and that is only safe because every caller holds the single-writer flag. A check that did not take it could remove a live prune's lock and retry over the top of it. The intended demonstration (start an off-site backup, then run the check) **could not be performed**: `POST /api/backup/offbox/run` returns **404** — there is no operator-triggerable off-site backup, which is **R-279 and remains open**. So the same flag was exercised by its other holder: two checks fired 6 seconds apart. ``` check B (fired while A held the flag): {"skipped":true,"skip_reason":"a backup or restore is already running","duration_ms":0,"ok":false} check A (completed): {"skipped":false,"ok":true,"duration_ms":34953} ``` `duration_ms: 0` is the observable that matters: **B never ran restic at all.** It yielded, it did not queue, and it did not advance due-ness. ## Files | file | what | |---|---| | `step1-negative-control.txt` | healthy scratch repo passes | | `step2-damage.txt` | which pack, how, before/after hashes | | `step3-positive-control.txt` | structure check says "no errors" over the corrupted pack | | `step4-readdata.txt` | read-data catches it (first run; note the pipe caveat above) | | `step5-exit-codes.txt` | exit codes re-measured without pipes | | `step7-readdata-cost.txt` | store size + the four-depth cost curve | | `step6-live-debug-route.txt` | the live debug route against the real store | | `step8-skip-live.txt` | the hazard control, live | | `step9-teardown.txt` | scratch removed, space returned |