# R-435 — the snapshot-drop detector does not see one app's off-site history vanish: a one-page design (2026-10-08) Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`. Architecture: `07-backup-architecture.md` §11 row 10 (ransomware / malicious deletion; the 2026-10-03..05 `[FACT]` entries on the append-only key and the clean-up window); `09` §3 decisions 68, 69 (the box prunes only in a hub-opened window; the hub-provisioned key is append-only). **Status:** design only, nothing built. ## 1. The problem, and what changed under it since the row was filed (2026-09-01) - **Still true in code.** `hub/internal/monitor/offsite.go:254-257`: an alarm needs a fall of more than half the last count AND at least 5 (`snapshotDropped`, `:285-295`). One app's tag (~9 of 69 on demo-hp when measured) is invisible. The limit is written in the comment above the constants (`:242-253`), as the row says. - **The row's second complaint is gone.** `STATUS.md` no longer says „noticed within a day" (`grep -n -i "within a day\|noticed" STATUS.md` returned nothing today). `07` row 10 says „only a fall of MORE than half (R-435)". - **The threat the row describes is now mostly PREVENTED, not only undetected.** Since decisions 68–69 (2026-10-03) the box's off-site key cannot delete (provider refusal measured, `07` row 10). The box deletes only inside a hub-opened window (`controller/internal/backup/offbox_window.go:203-225`, `t.Pinned()`), and inside it the fake-snapshot guard refuses any plan that removes a snapshot younger than `keepDaily` days that is not superseded the same day (`offbox_window.go:24-36`). A one-app wipe must remove that app's newest snapshots, so the guard refuses it. - **What is left uncovered:** a deletion by something that is NOT the box's key — anyone with a deleting credential for the sub-account or the provider account — or a box that lies about its plan. The count detector sees those only above one half. ## 2. The new fact that makes a precise detector cheap On a pinned tier the **only legitimate way the count can fall is a clean-up window**, and the hub records every window with `count_before` / `count_after` (`hub/internal/store/store.go:835-845`, `offsite_keys.go:84-101`). So the hub can say „this fall is explained" exactly, with no threshold: between two trustworthy reports, the count may fall by at most what the windows closed in that interval removed. **Any fall beyond that is unexplained — even one snapshot.** ## 3. Options | | What | Costs | Risk | |---|---|---|---| | **A** | Keep as is: documented blind spot. | Nothing. | A single-app deletion from outside the box stays silent. | | **B** | The row's own suggestion: the controller reports a count per app tag; the hub keeps a per-tag baseline and threshold. | Two repos, a new report field (wire-contract gate), a per-tag threshold nobody has measured; retention also moves per-tag counts, so it needs its own calibration. ~1.5 sessions. | Noise from retention on boundary days — the reason the global threshold is high. | | **C** | Hub only: on a **pinned** tier (hub holds a confirmed append-only key for the customer), compare `prev − cur` with the snapshots removed by windows closed since the previous trustworthy report. Any unexplained fall ≥ 1 raises `offsite_snapshots_dropped`. Non-pinned tiers (the household's NAS) keep today's half-rule. | One repo. One store query (windows closed in an interval), one branch in `snapshotDropped`, tests. ~½ session. | A window that closes by timeout without a box result has no `count_after`; treat it as „explains anything" (no alarm, log one INFO line) — the safe side for noise, recorded as a limit. | ## 4. The pick — C C covers the shape B was meant to cover — and any size — without a new threshold, a new field or the second repo. It uses the only fact that is exact (the windows the hub itself opened). It keeps the half-rule for NAS tiers, where the box still prunes by itself. B stays a note: it would add per-app naming in the message, which C does not have (C says „N snapshots went, outside any clean-up window"; the operator reads which ones on the box). ## 5. First slice and its red test - Hub: `store.RemovedByWindowsBetween(customerID, from, to) (removed int, unknown bool)`; `OffsiteChecker` keeps the time of the last trustworthy baseline beside `lastCounts`; on a pinned customer, alarm when `prev − cur > removed` and not `unknown`. The message adds „outside any clean-up window". - **Red test (seen failing on today's code):** pinned customer, baseline 69, next trustworthy report 60, no window → exactly one `offsite_snapshots_dropped` event. Today: none (9 < 34.5). - **Controls:** same fall with a closed window that removed 9 → no event; NAS (not pinned) customer 69 → 60 → no event (half-rule kept); untrustworthy report → no event and baseline unchanged (existing rule). - Live proof, when releases resume: on scratch 9202 only — a window granted by hand removes N and no alarm; then the hub's own row for that window as the control from another channel. ## 6. One question for the operator **Should the hub raise an error mail when even ONE off-site snapshot disappears outside a clean-up window it opened?** My pick: yes (option C). *If you do nothing:* deletions by the box stay blocked as today, but a deletion through any other credential stays silent unless it removes more than half of a household's off-site history.