Part F: nine one-page designs (R-314/279/177, R-30, R-79, R-35, R-717, R-138, R-435, R-893); R-177 and R-298 closed with evidence; R-415 re-filed (lost in the 2026-10-03 triage); STATUS decision sheet D1-D9; 131 -> 130
gates / gates (push) Successful in 3m57s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-08 10:12:46 +02:00
parent 5beedcce1a
commit 78121aa475
11 changed files with 505 additions and 14 deletions
@@ -0,0 +1,56 @@
# R-435 — the snapshot-drop detector does not see one app's off-site history vanish: a one-page design (2026-10-08)
Baselines read: felhom.eu `b2dce901`, felhom-controller `a0370b4`. Architecture: `07-backup-architecture.md` §11 row 10
(ransomware / malicious deletion; the 2026-10-03..05 `[FACT]` entries on the append-only key and the clean-up window);
`09` §3 decisions 68, 69 (the box prunes only in a hub-opened window; the hub-provisioned key is append-only).
**Status:** design only, nothing built.
## 1. The problem, and what changed under it since the row was filed (2026-09-01)
- **Still true in code.** `hub/internal/monitor/offsite.go:254-257`: an alarm needs a fall of more than half the last
count AND at least 5 (`snapshotDropped`, `:285-295`). One app's tag (~9 of 69 on demo-hp when measured) is invisible.
The limit is written in the comment above the constants (`:242-253`), as the row says.
- **The row's second complaint is gone.** `STATUS.md` no longer says „noticed within a day" (`grep -n -i "within a day\|noticed" STATUS.md`
returned nothing today). `07` row 10 says „only a fall of MORE than half (R-435)".
- **The threat the row describes is now mostly PREVENTED, not only undetected.** Since decisions 68–69 (2026-10-03) the
box's off-site key cannot delete (provider refusal measured, `07` row 10). The box deletes only inside a hub-opened
window (`controller/internal/backup/offbox_window.go:203-225`, `t.Pinned()`), and inside it the fake-snapshot guard
refuses any plan that removes a snapshot younger than `keepDaily` days that is not superseded the same day
(`offbox_window.go:24-36`). A one-app wipe must remove that app's newest snapshots, so the guard refuses it.
- **What is left uncovered:** a deletion by something that is NOT the box's key — anyone with a deleting credential
for the sub-account or the provider account — or a box that lies about its plan. The count detector sees those only
above one half.
## 2. The new fact that makes a precise detector cheap
On a pinned tier the **only legitimate way the count can fall is a clean-up window**, and the hub records every window
with `count_before` / `count_after` (`hub/internal/store/store.go:835-845`, `offsite_keys.go:84-101`). So the hub can
say „this fall is explained" exactly, with no threshold: between two trustworthy reports, the count may fall by at
most what the windows closed in that interval removed. **Any fall beyond that is unexplained — even one snapshot.**
## 3. Options
| | What | Costs | Risk |
|---|---|---|---|
| **A** | Keep as is: documented blind spot. | Nothing. | A single-app deletion from outside the box stays silent. |
| **B** | The row's own suggestion: the controller reports a count per app tag; the hub keeps a per-tag baseline and threshold. | Two repos, a new report field (wire-contract gate), a per-tag threshold nobody has measured; retention also moves per-tag counts, so it needs its own calibration. ~1.5 sessions. | Noise from retention on boundary days — the reason the global threshold is high. |
| **C** | Hub only: on a **pinned** tier (hub holds a confirmed append-only key for the customer), compare `prev − cur` with the snapshots removed by windows closed since the previous trustworthy report. Any unexplained fall ≥ 1 raises `offsite_snapshots_dropped`. Non-pinned tiers (the household's NAS) keep today's half-rule. | One repo. One store query (windows closed in an interval), one branch in `snapshotDropped`, tests. ~½ session. | A window that closes by timeout without a box result has no `count_after`; treat it as „explains anything" (no alarm, log one INFO line) — the safe side for noise, recorded as a limit. |
## 4. The pick — C
C covers the shape B was meant to cover — and any size — without a new threshold, a new field or the second repo. It
uses the only fact that is exact (the windows the hub itself opened). It keeps the half-rule for NAS tiers, where the
box still prunes by itself. B stays a note: it would add per-app naming in the message, which C does not have (C says
„N snapshots went, outside any clean-up window"; the operator reads which ones on the box).
## 5. First slice and its red test
- Hub: `store.RemovedByWindowsBetween(customerID, from, to) (removed int, unknown bool)`; `OffsiteChecker` keeps the time
of the last trustworthy baseline beside `lastCounts`; on a pinned customer, alarm when `prev − cur > removed` and not
`unknown`. The message adds „outside any clean-up window".
- **Red test (seen failing on today's code):** pinned customer, baseline 69, next trustworthy report 60, no window →
exactly one `offsite_snapshots_dropped` event. Today: none (9 < 34.5).
- **Controls:** same fall with a closed window that removed 9 → no event; NAS (not pinned) customer 69 → 60 → no
event (half-rule kept); untrustworthy report → no event and baseline unchanged (existing rule).
- Live proof, when releases resume: on scratch 9202 only — a window granted by hand removes N and no alarm; then the
hub's own row for that window as the control from another channel.
## 6. One question for the operator
**Should the hub raise an error mail when even ONE off-site snapshot disappears outside a clean-up window it opened?**
My pick: yes (option C). *If you do nothing:* deletions by the box stay blocked as today, but a deletion through any
other credential stays silent unless it removes more than half of a household's off-site history.