v0.227.1: the damage classifier matched restic's ordinary progress output
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
A patch and not a rebuilt 0.227.0: that tag was already running on demo-hp, and re-pushing changed bytes under a live tag is the :latest hazard with extra steps. looksLikeRepositoryDamage matched bare "pack ", "tree ", "snapshot ", "blob ". A HEALTHY restic check prints "check all packs" and "check snapshots, trees and blobs" -- so any check that failed for a NON-damage reason, a connection dropped mid-run for instance, would have been classified as a corrupted repository and told the customer their backups may be damaged. That is the false alarm that teaches an operator to ignore the true one. Caught by the NEGATIVE control, built from the real bytes of a real passing check on demo-hp. The spec made the negative control mandatory and this is what it was for: a control that has only ever seen the failing case proves nothing. Signatures are now phrases from restic's own error wording. Also in this commit: CONTEXT.md records the three rulings (take the flag and skip, due-ness not a weekday, publish on OffboxReportStatus not the R-331 dead fields) plus the measurement a future session would otherwise assume wrongly -- THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION. README documents the job, the route and the config, and corrects a line that listed four debug backup routes when only two exist. REUSE gains three rows, including one that records R-398 was my own mistake so nobody re-files it.
This commit is contained in:
+139
@@ -1,3 +1,142 @@
|
||||
## v0.227.1 — the damage classifier matched restic's ordinary progress output (2026-08-30, R-359 follow-on)
|
||||
**MinAgent: 0.129.0** (unchanged)
|
||||
|
||||
**A patch rather than a rebuilt 0.227.0, deliberately** — 0.227.0 was already deployed to `demo-hp`
|
||||
when this was found, and re-pushing changed bytes under a tag that is already running somewhere is the
|
||||
`:latest` hazard with extra steps.
|
||||
|
||||
`looksLikeRepositoryDamage` matched bare `"pack "`, `"tree "`, `"snapshot "` and `"blob "`. **A healthy
|
||||
`restic check` prints `check all packs` and `check snapshots, trees and blobs`**, so any check that
|
||||
failed for a NON-damage reason — a connection dropped mid-run, say — would have been classified as a
|
||||
corrupted repository and told the customer their backups may be damaged. That is the false alarm that
|
||||
teaches an operator to ignore the true one.
|
||||
|
||||
The signatures are now phrases from restic's own error wording (`does not match`, `not found in index`,
|
||||
`ciphertext verification failed`, `repository contains errors`, …), and the negative control that
|
||||
caught it — `TestR359_HealthyRealOutputIsNotDamage`, built from the real bytes of a real passing
|
||||
check — is what keeps it caught.
|
||||
|
||||
## v0.227.0 — the off-site store gets checked, and the check that was advertised becomes real (2026-08-30, R-359 + R-397)
|
||||
**MinAgent: 0.129.0** (unchanged)
|
||||
|
||||
### R-359 — nothing ever verified that the off-site copies are still readable
|
||||
|
||||
The whole-guest tier has verify jobs. The tier holding the customer's documents and photos had none:
|
||||
the complete set of restic verbs this controller used was `restore, snapshots, backup, unlock, stats,
|
||||
init, forget, prune, cat, config` — **no `check`**. We would have found out at restore time, with a
|
||||
customer waiting.
|
||||
|
||||
### ⚠ AND THE MEASUREMENT CHANGED WHAT THE FEATURE IS WORTH — read this before the rest
|
||||
|
||||
Part 5's positive control corrupted one pack of a throwaway repository **without changing its size**
|
||||
(64 zero bytes at offset 1024). Both depths were run against it:
|
||||
|
||||
| depth | exit | verdict |
|
||||
|---|---|---|
|
||||
| `restic check` — **the depth that ships ON** | **0** | **`no errors were found`** |
|
||||
| any `--read-data*` form — **ships OFF** | 1 | `Pack ID does not match, want 288afd3e…, got 4b6847bb…` |
|
||||
|
||||
**The structure check reported a corrupted store as healthy.** It verifies the index, the pack
|
||||
inventory and the snapshot graph — real failure modes, and it catches missing packs, broken indexes and
|
||||
unreadable snapshots. It does **not** re-hash pack contents, so it cannot see rot inside a pack that is
|
||||
still the right size. **R-399 is therefore not only a bandwidth question: at the shipped default a
|
||||
class of damage is not checked at all.**
|
||||
|
||||
**The cost curve, measured against the live store (134.3 MB, 67 snapshots) rather than reasoned:**
|
||||
|
||||
| depth | wall | over structure-only |
|
||||
|---|---|---|
|
||||
| structure only | 35.0 s | — |
|
||||
| `10%` | 35.9 s | +0.9 s (+3%) |
|
||||
| `50%` | 37.3 s | +2.2 s (+6%) |
|
||||
| `100%` | 39.2 s | **+4.2 s (+12%)** |
|
||||
|
||||
At this size, re-reading **all** the data costs four seconds more than reading none — the wall clock is
|
||||
dominated by SFTP round-trips, not transfer. **These figures do not extrapolate**: the structure
|
||||
check's cost tracks the index, a read-data run's tracks the data. The default is still not chosen here;
|
||||
R-399 now has numbers instead of guesses.
|
||||
|
||||
### The hazard that shapes the whole design
|
||||
|
||||
`resticStep` self-heals a crash lock by running **`unlock --remove-all`** and retrying, and its own
|
||||
comment records why that is safe: *every caller holds the in-process single-flight mutex, so any lock
|
||||
it meets is stale.* **A check that did not take that flag could meet a LIVE `forget --prune`'s lock
|
||||
from this same box, remove it, and retry over the top of it** — on the tier holding customer data.
|
||||
|
||||
So the check **takes the flag and SKIPS rather than waits**. Waiting would pin the nightly backup
|
||||
behind it; a skip costs nothing because due-ness makes tomorrow try again.
|
||||
`TestR359_SkipsWhenRunningFlagHeld` asserts the **non-effects** — restic never invoked, `unlock` never
|
||||
in any argv — and its red-proof prints the real thing: restic running `check` with the flag held.
|
||||
|
||||
**Proven live too.** The intended demonstration could not be run (`POST /api/backup/offbox/run` is a
|
||||
404 — that is **R-279 and stays open**), so the same flag was exercised by its other holder: two checks
|
||||
6 s apart. The second returned `skipped: true`, `"a backup or restore is already running"`,
|
||||
**`duration_ms: 0`** — it never ran restic at all.
|
||||
|
||||
### Due-ness, not a weekday
|
||||
|
||||
A daily job asking *"is the last successful check older than 7 days?"*, not *"is it Sunday?"*. A box
|
||||
switched off on its check day is checked the next day it is on. **R-341 is exactly the other shape** —
|
||||
a dated check quietly missed and never caught up. No `Weekly` primitive was added; due-ness is smaller
|
||||
and is what R-86 already chose for restore-tests. Registered at **06:00**, chosen from the live
|
||||
schedule read off `demo-hp` (db-dump 02:30, tier2 + fill-watch 03:30, metrics-prune 04:00, offbox-backup
|
||||
04:15, abandon-sweep 05:10, whole-guest gate 04:30–08:30).
|
||||
|
||||
### Three outcomes, not two
|
||||
|
||||
`Skipped`, `Unreachable` and failed are different facts. **"I could not look" is not "I looked and it
|
||||
is broken"** — R-339 already owns reachability, and a second alarm for the same fact trains the
|
||||
operator to discount the one alarm that means the backups are damaged. A timeout is unreachable, never
|
||||
damage. A failure advances due-ness (a broken store must not be re-checked nightly); a skip and an
|
||||
unreachable store do not.
|
||||
|
||||
### R-397 — the notifiers get their caller
|
||||
|
||||
`NotifyIntegrityOK` / `NotifyIntegrityFailed` existed with **no caller**; the hub allowlists both event
|
||||
types and carries the Hungarian customer text for both; the settings checkbox exists; the debug button
|
||||
posts to `/api/debug/backup/integrity` and **the dispatch had no such case**. Everything was built
|
||||
except the part that runs. **Sixth instance of that shape in this project.**
|
||||
|
||||
Success is severity `info`, which `severityNotifies` drops before either leg — it **mails nobody, by
|
||||
design**. A weekly success e-mail is how people stop reading their alerts. Observed live:
|
||||
`Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)`.
|
||||
|
||||
The customer gets a **sentence**; restic's words go to the log, truncated (R-379: 615 bytes of raw
|
||||
database text reached a customer once). Published on `OffboxReportStatus`, **not** on
|
||||
`report.BackupReport`'s `IntegrityOK` — those were retired by R-331 the day before, and
|
||||
`TestBackupReport_DeadFieldsStayZero` still passes unmodified.
|
||||
|
||||
`backup_integrity_failed` was checked against `perAppCooldownEvents`, `operatorOnlyEvents` and
|
||||
`DefaultEnabledEvents` and **deliberately left out of all three** — it is already in the right shape.
|
||||
|
||||
### Two defects this work introduced and then caught
|
||||
|
||||
**The damage classifier matched restic's ordinary progress output.** The first draft looked for bare
|
||||
`"pack "`, `"tree "`, `"snapshot "` — and a healthy run prints `check all packs` and
|
||||
`check snapshots, trees and blobs`, so a check that failed for a non-damage reason would have alarmed
|
||||
the customer that their backups were corrupt. **Caught by the negative control**
|
||||
(`TestR359_HealthyRealOutputIsNotDamage`) using the real bytes of a real passing check. The signatures
|
||||
are now phrases from restic's own error wording.
|
||||
|
||||
**An exit code I misread.** An early run showed `exit=0` on the subset forms while they printed
|
||||
`Fatal: repository contains errors`. That was not restic — the commands were piped through `tail`, so
|
||||
`$?` was tail's. Re-measured without pipes, every read-data form exits 1. The project's own
|
||||
"exit codes that lie" trap, caught by re-measuring rather than reasoning.
|
||||
|
||||
### Part 0 was NOT built, and R-398 was my own mistake
|
||||
|
||||
R-398 (filed by me yesterday) said `resticStep` is not a seam so no test can drive a restic-backed
|
||||
path. **The first half is true and the conclusion was false:** `offboxRunner` / `SetOffboxRunner` /
|
||||
`m.runner()` has been injectable since the off-site tier shipped, and other tests already drive restic
|
||||
paths through it. A `resticStepFn` seam would have been **worse** here — it would replace the
|
||||
`unlock --remove-all` escalation and hide it from the assertions that must observe it. R-358's AST
|
||||
ordering test is converted to a real execution test instead, which immediately surfaced something the
|
||||
AST walk could not: `unlockStale` legitimately runs before the restore. R-398 is **corrected, not
|
||||
closed**.
|
||||
|
||||
**Green gate:** 28 packages, rc 0. Four red-proofs (A2, B1, C2, D2), each printing the pre-fix
|
||||
behaviour. Evidence: `felhom.eu/documentation/tests/r359-integrity-2026-08-30/`.
|
||||
|
||||
## v0.226.1 — the unknown that the v0.226.0 fix drew as a zero (2026-08-30, R-353 follow-on)
|
||||
**MinAgent: 0.129.0** (unchanged)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user