Files
felhom.eu/documentation/audits/day-2026-10-08/design-R-435.md
T

5.4 KiB
Raw Blame History

R-435 — the snapshot-drop detector does not see one app's off-site history vanish: a one-page design (2026-10-08)

Baselines read: felhom.eu b2dce901, felhom-controller a0370b4. Architecture: 07-backup-architecture.md §11 row 10 (ransomware / malicious deletion; the 2026-10-03..05 [FACT] entries on the append-only key and the clean-up window); 09 §3 decisions 68, 69 (the box prunes only in a hub-opened window; the hub-provisioned key is append-only). Status: design only, nothing built.

1. The problem, and what changed under it since the row was filed (2026-09-01)

  • Still true in code. hub/internal/monitor/offsite.go:254-257: an alarm needs a fall of more than half the last count AND at least 5 (snapshotDropped, :285-295). One app's tag (~9 of 69 on demo-hp when measured) is invisible. The limit is written in the comment above the constants (:242-253), as the row says.
  • The row's second complaint is gone. STATUS.md no longer says „noticed within a day" (grep -n -i "within a day\|noticed" STATUS.md returned nothing today). 07 row 10 says „only a fall of MORE than half (R-435)".
  • The threat the row describes is now mostly PREVENTED, not only undetected. Since decisions 68–69 (2026-10-03) the box's off-site key cannot delete (provider refusal measured, 07 row 10). The box deletes only inside a hub-opened window (controller/internal/backup/offbox_window.go:203-225, t.Pinned()), and inside it the fake-snapshot guard refuses any plan that removes a snapshot younger than keepDaily days that is not superseded the same day (offbox_window.go:24-36). A one-app wipe must remove that app's newest snapshots, so the guard refuses it.
  • What is left uncovered: a deletion by something that is NOT the box's key — anyone with a deleting credential for the sub-account or the provider account — or a box that lies about its plan. The count detector sees those only above one half.

2. The new fact that makes a precise detector cheap

On a pinned tier the only legitimate way the count can fall is a clean-up window, and the hub records every window with count_before / count_after (hub/internal/store/store.go:835-845, offsite_keys.go:84-101). So the hub can say „this fall is explained" exactly, with no threshold: between two trustworthy reports, the count may fall by at most what the windows closed in that interval removed. Any fall beyond that is unexplained — even one snapshot.

3. Options

What Costs Risk
A Keep as is: documented blind spot. Nothing. A single-app deletion from outside the box stays silent.
B The row's own suggestion: the controller reports a count per app tag; the hub keeps a per-tag baseline and threshold. Two repos, a new report field (wire-contract gate), a per-tag threshold nobody has measured; retention also moves per-tag counts, so it needs its own calibration. ~1.5 sessions. Noise from retention on boundary days — the reason the global threshold is high.
C Hub only: on a pinned tier (hub holds a confirmed append-only key for the customer), compare prev − cur with the snapshots removed by windows closed since the previous trustworthy report. Any unexplained fall ≥ 1 raises offsite_snapshots_dropped. Non-pinned tiers (the household's NAS) keep today's half-rule. One repo. One store query (windows closed in an interval), one branch in snapshotDropped, tests. ~½ session. A window that closes by timeout without a box result has no count_after; treat it as „explains anything" (no alarm, log one INFO line) — the safe side for noise, recorded as a limit.

4. The pick — C

C covers the shape B was meant to cover — and any size — without a new threshold, a new field or the second repo. It uses the only fact that is exact (the windows the hub itself opened). It keeps the half-rule for NAS tiers, where the box still prunes by itself. B stays a note: it would add per-app naming in the message, which C does not have (C says „N snapshots went, outside any clean-up window"; the operator reads which ones on the box).

5. First slice and its red test

  • Hub: store.RemovedByWindowsBetween(customerID, from, to) (removed int, unknown bool); OffsiteChecker keeps the time of the last trustworthy baseline beside lastCounts; on a pinned customer, alarm when prev − cur > removed and not unknown. The message adds „outside any clean-up window".
  • Red test (seen failing on today's code): pinned customer, baseline 69, next trustworthy report 60, no window → exactly one offsite_snapshots_dropped event. Today: none (9 < 34.5).
  • Controls: same fall with a closed window that removed 9 → no event; NAS (not pinned) customer 69 → 60 → no event (half-rule kept); untrustworthy report → no event and baseline unchanged (existing rule).
  • Live proof, when releases resume: on scratch 9202 only — a window granted by hand removes N and no alarm; then the hub's own row for that window as the control from another channel.

6. One question for the operator

Should the hub raise an error mail when even ONE off-site snapshot disappears outside a clean-up window it opened? My pick: yes (option C). If you do nothing: deletions by the box stay blocked as today, but a deletion through any other credential stays silent unless it removes more than half of a household's off-site history.