78121aa475
gates / gates (push) Successful in 3m57s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
57 lines
5.4 KiB
Markdown
57 lines
5.4 KiB
Markdown
# R-435 — the snapshot-drop detector does not see one app's off-site history vanish: a one-page design (2026-10-08)
|
||
|
||
Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`. Architecture: `07-backup-architecture.md` §11 row 10
|
||
(ransomware / malicious deletion; the 2026-10-03..05 `[FACT]` entries on the append-only key and the clean-up window);
|
||
`09` §3 decisions 68, 69 (the box prunes only in a hub-opened window; the hub-provisioned key is append-only).
|
||
**Status:** design only, nothing built.
|
||
|
||
## 1. The problem, and what changed under it since the row was filed (2026-09-01)
|
||
- **Still true in code.** `hub/internal/monitor/offsite.go:254-257`: an alarm needs a fall of more than half the last
|
||
count AND at least 5 (`snapshotDropped`, `:285-295`). One app's tag (~9 of 69 on demo-hp when measured) is invisible.
|
||
The limit is written in the comment above the constants (`:242-253`), as the row says.
|
||
- **The row's second complaint is gone.** `STATUS.md` no longer says „noticed within a day" (`grep -n -i "within a day\|noticed" STATUS.md`
|
||
returned nothing today). `07` row 10 says „only a fall of MORE than half (R-435)".
|
||
- **The threat the row describes is now mostly PREVENTED, not only undetected.** Since decisions 68–69 (2026-10-03) the
|
||
box's off-site key cannot delete (provider refusal measured, `07` row 10). The box deletes only inside a hub-opened
|
||
window (`controller/internal/backup/offbox_window.go:203-225`, `t.Pinned()`), and inside it the fake-snapshot guard
|
||
refuses any plan that removes a snapshot younger than `keepDaily` days that is not superseded the same day
|
||
(`offbox_window.go:24-36`). A one-app wipe must remove that app's newest snapshots, so the guard refuses it.
|
||
- **What is left uncovered:** a deletion by something that is NOT the box's key — anyone with a deleting credential
|
||
for the sub-account or the provider account — or a box that lies about its plan. The count detector sees those only
|
||
above one half.
|
||
|
||
## 2. The new fact that makes a precise detector cheap
|
||
On a pinned tier the **only legitimate way the count can fall is a clean-up window**, and the hub records every window
|
||
with `count_before` / `count_after` (`hub/internal/store/store.go:835-845`, `offsite_keys.go:84-101`). So the hub can
|
||
say „this fall is explained" exactly, with no threshold: between two trustworthy reports, the count may fall by at
|
||
most what the windows closed in that interval removed. **Any fall beyond that is unexplained — even one snapshot.**
|
||
|
||
## 3. Options
|
||
| | What | Costs | Risk |
|
||
|---|---|---|---|
|
||
| **A** | Keep as is: documented blind spot. | Nothing. | A single-app deletion from outside the box stays silent. |
|
||
| **B** | The row's own suggestion: the controller reports a count per app tag; the hub keeps a per-tag baseline and threshold. | Two repos, a new report field (wire-contract gate), a per-tag threshold nobody has measured; retention also moves per-tag counts, so it needs its own calibration. ~1.5 sessions. | Noise from retention on boundary days — the reason the global threshold is high. |
|
||
| **C** | Hub only: on a **pinned** tier (hub holds a confirmed append-only key for the customer), compare `prev − cur` with the snapshots removed by windows closed since the previous trustworthy report. Any unexplained fall ≥ 1 raises `offsite_snapshots_dropped`. Non-pinned tiers (the household's NAS) keep today's half-rule. | One repo. One store query (windows closed in an interval), one branch in `snapshotDropped`, tests. ~½ session. | A window that closes by timeout without a box result has no `count_after`; treat it as „explains anything" (no alarm, log one INFO line) — the safe side for noise, recorded as a limit. |
|
||
|
||
## 4. The pick — C
|
||
C covers the shape B was meant to cover — and any size — without a new threshold, a new field or the second repo. It
|
||
uses the only fact that is exact (the windows the hub itself opened). It keeps the half-rule for NAS tiers, where the
|
||
box still prunes by itself. B stays a note: it would add per-app naming in the message, which C does not have (C says
|
||
„N snapshots went, outside any clean-up window"; the operator reads which ones on the box).
|
||
|
||
## 5. First slice and its red test
|
||
- Hub: `store.RemovedByWindowsBetween(customerID, from, to) (removed int, unknown bool)`; `OffsiteChecker` keeps the time
|
||
of the last trustworthy baseline beside `lastCounts`; on a pinned customer, alarm when `prev − cur > removed` and not
|
||
`unknown`. The message adds „outside any clean-up window".
|
||
- **Red test (seen failing on today's code):** pinned customer, baseline 69, next trustworthy report 60, no window →
|
||
exactly one `offsite_snapshots_dropped` event. Today: none (9 < 34.5).
|
||
- **Controls:** same fall with a closed window that removed 9 → no event; NAS (not pinned) customer 69 → 60 → no
|
||
event (half-rule kept); untrustworthy report → no event and baseline unchanged (existing rule).
|
||
- Live proof, when releases resume: on scratch 9202 only — a window granted by hand removes N and no alarm; then the
|
||
hub's own row for that window as the control from another channel.
|
||
|
||
## 6. One question for the operator
|
||
**Should the hub raise an error mail when even ONE off-site snapshot disappears outside a clean-up window it opened?**
|
||
My pick: yes (option C). *If you do nothing:* deletions by the box stay blocked as today, but a deletion through any
|
||
other credential stays silent unless it removes more than half of a household's off-site history.
|