v0.228.0 — the off-site check reads the data; the debug page stops lying (R-399 + R-400)
gates / gates (push) Successful in 12s

R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged
without changing its size made plain `restic check` report "no errors were found"
on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store:
35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means
not-configured, therefore the default; a malformed value falls back to the DEFAULT,
never to structure. A completed check over 5 minutes logs a WARN naming the
duration, the depth and R-401 — operator log only, no hub event, no depth change.
The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT
RECORDED, never "structure").

R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on
page LOAD, so those panels were permanently blank. backup/crossdrive implemented;
backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both
storage/simulate-* deleted with their panels and JavaScript.
scripts/debug_route_gate.py fails in both directions and is registered after the
seven were resolved. 18 referenced, 18 dispatched, none orphaned.

Corrections: the dead-field warning in report/types.go said the controller runs no
integrity check and the notifiers are called from nowhere — both false since
v0.227.0. controller.yaml.example gains its missing integrity: block.
integrityCheckTimeout's "ships OFF" comment rewritten.
This commit is contained in:
2026-08-31 10:24:29 +02:00
parent 300d7e87d7
commit 3c49dc8ea4
23 changed files with 1144 additions and 175 deletions
@@ -83,6 +83,21 @@ monitoring:
# ping_uuids: (deprecated — monitoring is now handled by the Hub event system)
system_health_interval: "5m"
health_check_schedule: "06:00"
# Off-site (restic) store integrity check — R-359 / R-399.
integrity:
max_age_days: 7 # A MAX AGE, not a weekday: the job runs daily and asks "is the last
# successful check older than this?", so a box that was off on its check
# day is checked the next day it is on.
read_data_subset: "" # HOW DEEP the check looks. Empty or absent = the default, 100% — every
# stored byte is downloaded and re-hashed. Set "off" to check the
# structure and index only; set "10%", "1/7" or "50M" for a partial
# re-read. A value restic does not accept WARNs and falls back to the
# default, never to "off".
# The default is 100% because the structure check does NOT detect a
# size-preserving pack corruption: measured on a real store 2026-08-30,
# `restic check` reported "no errors were found" over a damaged pack that
# every read-data form caught. Cost on that 134 MB store: 35.0 s at
# structure depth, 39.2 s at 100%.
thresholds:
disk_warn_percent: 80
disk_crit_percent: 90