v0.228.0 — the off-site check reads the data; the debug page stops lying (R-399 + R-400)
gates / gates (push) Successful in 12s

R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged
without changing its size made plain `restic check` report "no errors were found"
on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store:
35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means
not-configured, therefore the default; a malformed value falls back to the DEFAULT,
never to structure. A completed check over 5 minutes logs a WARN naming the
duration, the depth and R-401 — operator log only, no hub event, no depth change.
The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT
RECORDED, never "structure").

R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on
page LOAD, so those panels were permanently blank. backup/crossdrive implemented;
backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both
storage/simulate-* deleted with their panels and JavaScript.
scripts/debug_route_gate.py fails in both directions and is registered after the
seven were resolved. 18 referenced, 18 dispatched, none orphaned.

Corrections: the dead-field warning in report/types.go said the controller runs no
integrity check and the notifiers are called from nowhere — both false since
v0.227.0. controller.yaml.example gains its missing integrity: block.
integrityCheckTimeout's "ships OFF" comment rewritten.
This commit is contained in:
2026-08-31 10:24:29 +02:00
parent 300d7e87d7
commit 3c49dc8ea4
23 changed files with 1144 additions and 175 deletions
+83
View File
@@ -1,3 +1,86 @@
## v0.228.0 — the check reads the data, and the debug page stops lying (2026-08-31, R-399 + R-400)
**MinAgent: 0.129.0** (unchanged)
### The off-site check now re-reads the stored data (R-399)
`monitoring.integrity.read_data_subset` **defaults to `100%`**. Every box whose `controller.yaml` has
no `integrity:` block — which is every box in the fleet — now downloads and re-hashes the whole
off-site store on its weekly check instead of reading only the catalogue.
**The fact that forced it:** on `demo-hp`, 2026-08-30, a pack was damaged **without changing its
size**. Plain `restic check` — the structure-and-index check every box ran — reported
`no errors were found` and exited clean. Every `--read-data*` form caught it. A structurally-verified
store is one whose rot is found at restore time, with a customer waiting.
**The cost**, same store and day (140 829 678 B / 2 651 blobs / 67 snapshots): structure 35.0 s,
10% 35.9 s, 50% 37.3 s, **100% 39.2 s**.
- **`off` (any case) is the new off token.** Absent or empty means *not configured*, therefore the
default; a setting with no off switch is not a setting, and emptiness cannot mean both things.
- **A malformed value falls back to the DEFAULT, never to structure.** Downgrading on a typo would
silently remove the protection this row adds — R-357's shape, a guard that opens quietly.
- **A completed check over 5 minutes logs a WARN** naming the duration, the depth and R-401. Operator
log only: no hub event, no customer alarm, and it never changes the depth by itself. A skip or an
unreachable store never warns — neither has a duration to judge. The threshold is deliberately
imprecise (≈7.6× the only full-depth number that exists) because a notice changes no behaviour,
while a precise number invented from one measurement on one 134 MB store would not.
- **The depth is now recorded with the verdict** — `settings.OffboxTarget.LastIntegrityDepth` and
`OffboxReportStatus.LastIntegrityDepth`. Empty means NOT RECORDED (a pre-0.228.0 controller), never
"structure": absence means the box cannot answer, following the `StatsKnown` precedent beside it.
- **R-401 filed with a TRIGGER, not a date:** revisit the depth when the timing WARN fires on any box.
No rotation schedule, size threshold or bandwidth budget is built here — every one of those would be
a number invented from a single data point.
**R-87 (the restic tier is never restore-tested) stays OPEN.** Reading the bytes back out of the store
is not a restore.
### The debug page stops lying (R-400)
The shipped debug page referenced **24** `/api/debug/...` addresses and the dispatcher answered **17**.
Seven controls did nothing — and three of those seven were not buttons at all: `dr/infra-status` and
both `storage/watchdog-status` calls fetch on page **load**, so whole panels had been permanently
blank and nobody had to click anything to be misled. This is the page an operator opens when something
is already wrong.
| control | disposition | why |
|---|---|---|
| `backup/crossdrive` | **IMPLEMENTED** | `Manager.RunTier2` is live; only the route was missing |
| `backup/infra` | **DELETED** | the disk-tier infra backup moved to the host agent in slice 8C |
| `hub/infra-push` | **DELETED** | `Pusher.PushInfraBackup` was removed 2026-06-16 (it pushed plaintext secrets) |
| `dr/infra-status` | **DELETED** | it rendered the two retired mechanisms above; fetched on page load |
| `storage/watchdog-status` | **DELETED** | the slice-8C watchdog is retired; the drive-gate reconcile replaced it and publishes no such status. Fetched on page load, twice |
| `storage/simulate-disconnect` | **DELETED** | no backing capability, and it WRITES storage state — a button that fakes a drive disconnect on a customer's machine is a foot-gun |
| `storage/simulate-reconnect` | **DELETED** | same |
Each deletion took its panel and its JavaScript with it; the "Tárhely teszt" section went entirely.
A panel left behind renders nothing forever, which is how this class hides.
**`controller/scripts/debug_route_gate.py` makes the class impossible.** Two lists and a difference: it
fails when the template references an address the dispatcher lacks, **and** when the dispatcher answers
one nothing references. Registered in `controller_gates.py` **after** the seven were resolved — a
registered-but-failing gate refuses every push. Ten lines on purpose. Red-proofed in both directions.
Tree after the change: **18 referenced addresses, 18 dispatched, none orphaned.**
### Three corrections
- `internal/report/types.go` — the dead-field warning block said *"the controller runs no integrity
check, and `NotifyIntegrityOK`/`NotifyIntegrityFailed` … are called from nowhere"*. **Both halves
became false in v0.227.0.** The fields stay dead and unrendered (`TestBackupReport_DeadFieldsStayZero`
still passes unmodified); only the *reason* changed.
- `configs/controller.yaml.example` had **no `integrity:` block at all** — an operator could not
discover the setting exists. Added, with both keys, the default, the off token and the measurement.
- `internal/backup/offbox_integrity.go` — `integrityCheckTimeout`'s comment said read-data "ships OFF
… whoever turns it on must revisit this number". Rewritten: it is now the number a large store meets
first, and the slow notice exists to say so long before it does.
### Superseded tests
Two R-359 tests asserted the ruling this release reverses, and are replaced rather than weakened.
`TestR359_StructureCheckPassesNoReadDataFlag` → `TestR399_AbsentConfigRunsFullDepth`.
`TestR359_MalformedReadDataSubsetIsTreatedAsOff` → `TestR399_MalformedFallsBackToTheDefault` (its NAME
was the defect: treating a typo as "off" is the quiet downgrade). Both are recorded in place, so a
later reader does not re-derive the old ruling from an absence.
## v0.227.1 — the damage classifier matched restic's ordinary progress output (2026-08-30, R-359 follow-on)
**MinAgent: 0.129.0** (unchanged)