Files
felhom.eu/documentation/audits/day-2026-10-08/design-R-435.md
T

57 lines
5.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# R-435 — the snapshot-drop detector does not see one app's off-site history vanish: a one-page design (2026-10-08)
Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`. Architecture: `07-backup-architecture.md` §11 row 10
(ransomware / malicious deletion; the 2026-10-03..05 `[FACT]` entries on the append-only key and the clean-up window);
`09` §3 decisions 68, 69 (the box prunes only in a hub-opened window; the hub-provisioned key is append-only).
**Status:** design only, nothing built.
## 1. The problem, and what changed under it since the row was filed (2026-09-01)
- **Still true in code.** `hub/internal/monitor/offsite.go:254-257`: an alarm needs a fall of more than half the last
count AND at least 5 (`snapshotDropped`, `:285-295`). One app's tag (~9 of 69 on demo-hp when measured) is invisible.
The limit is written in the comment above the constants (`:242-253`), as the row says.
- **The row's second complaint is gone.** `STATUS.md` no longer says „noticed within a day" (`grep -n -i "within a day\|noticed" STATUS.md`
returned nothing today). `07` row 10 says „only a fall of MORE than half (R-435)".
- **The threat the row describes is now mostly PREVENTED, not only undetected.** Since decisions 68–69 (2026-10-03) the
box's off-site key cannot delete (provider refusal measured, `07` row 10). The box deletes only inside a hub-opened
window (`controller/internal/backup/offbox_window.go:203-225`, `t.Pinned()`), and inside it the fake-snapshot guard
refuses any plan that removes a snapshot younger than `keepDaily` days that is not superseded the same day
(`offbox_window.go:24-36`). A one-app wipe must remove that app's newest snapshots, so the guard refuses it.
- **What is left uncovered:** a deletion by something that is NOT the box's key — anyone with a deleting credential
for the sub-account or the provider account — or a box that lies about its plan. The count detector sees those only
above one half.
## 2. The new fact that makes a precise detector cheap
On a pinned tier the **only legitimate way the count can fall is a clean-up window**, and the hub records every window
with `count_before` / `count_after` (`hub/internal/store/store.go:835-845`, `offsite_keys.go:84-101`). So the hub can
say „this fall is explained" exactly, with no threshold: between two trustworthy reports, the count may fall by at
most what the windows closed in that interval removed. **Any fall beyond that is unexplained — even one snapshot.**
## 3. Options
| | What | Costs | Risk |
|---|---|---|---|
| **A** | Keep as is: documented blind spot. | Nothing. | A single-app deletion from outside the box stays silent. |
| **B** | The row's own suggestion: the controller reports a count per app tag; the hub keeps a per-tag baseline and threshold. | Two repos, a new report field (wire-contract gate), a per-tag threshold nobody has measured; retention also moves per-tag counts, so it needs its own calibration. ~1.5 sessions. | Noise from retention on boundary days — the reason the global threshold is high. |
| **C** | Hub only: on a **pinned** tier (hub holds a confirmed append-only key for the customer), compare `prev − cur` with the snapshots removed by windows closed since the previous trustworthy report. Any unexplained fall ≥ 1 raises `offsite_snapshots_dropped`. Non-pinned tiers (the household's NAS) keep today's half-rule. | One repo. One store query (windows closed in an interval), one branch in `snapshotDropped`, tests. ~½ session. | A window that closes by timeout without a box result has no `count_after`; treat it as „explains anything" (no alarm, log one INFO line) — the safe side for noise, recorded as a limit. |
## 4. The pick — C
C covers the shape B was meant to cover — and any size — without a new threshold, a new field or the second repo. It
uses the only fact that is exact (the windows the hub itself opened). It keeps the half-rule for NAS tiers, where the
box still prunes by itself. B stays a note: it would add per-app naming in the message, which C does not have (C says
„N snapshots went, outside any clean-up window"; the operator reads which ones on the box).
## 5. First slice and its red test
- Hub: `store.RemovedByWindowsBetween(customerID, from, to) (removed int, unknown bool)`; `OffsiteChecker` keeps the time
of the last trustworthy baseline beside `lastCounts`; on a pinned customer, alarm when `prev − cur > removed` and not
`unknown`. The message adds „outside any clean-up window".
- **Red test (seen failing on today's code):** pinned customer, baseline 69, next trustworthy report 60, no window →
exactly one `offsite_snapshots_dropped` event. Today: none (9 < 34.5).
- **Controls:** same fall with a closed window that removed 9 → no event; NAS (not pinned) customer 69 → 60 → no
event (half-rule kept); untrustworthy report → no event and baseline unchanged (existing rule).
- Live proof, when releases resume: on scratch 9202 only — a window granted by hand removes N and no alarm; then the
hub's own row for that window as the control from another channel.
## 6. One question for the operator
**Should the hub raise an error mail when even ONE off-site snapshot disappears outside a clean-up window it opened?**
My pick: yes (option C). *If you do nothing:* deletions by the box stay blocked as today, but a deletion through any
other credential stays silent unless it removes more than half of a household's off-site history.