# REPORT — R-429 corrected, R-95 re-scoped, R-431 shipped (2026-09-01) ## Part 1's answer, first, because everything reads differently after it **Can a customer's own account reach the snapshot tree? NO — it can reach the DOOR, and is refused writes to it, but the tree lists EMPTY.** Measured on **both** live boxes, over the credential each already holds, with controls in the same run. ``` === POSITIVE CONTROL: the account home (must list) === drwxr-xr-x 1058 4 Aug 22 03:10 ./. drwx------ 1058 3 Jul 23 09:53 ./.ssh drwxrwxr-x 1058 8 Aug 4 12:38 ./ === NEGATIVE CONTROL: a name that cannot exist (must fail) === Can't ls: "/home/./zzz-no-such-r429" not found === CANDIDATE A: /.zfs/snapshot === drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/. drwxrwxrwx root root 0 Jul 21 16:01 /.zfs/snapshot/.. === CANDIDATE B: /home/.zfs/snapshot === Can't ls: "/home/.zfs/snapshot" not found === CANDIDATE D: is the .zfs door itself visible? === drwxrwxrwx root root 2 Sep 1 12:12 /.zfs/shares drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot ``` **The write-refusal test — the load-bearing sentence of the whole re-scope, now proven not cited:** ``` --- control: the same write in the account home MUST succeed sftp> put … ./r429-write-control.txt Uploading … to /home/./r429-write-control.txt -rw-r--r-- 1058 11 Sep 1 12:12 ./r429-write-control.txt --- cleanup of the control file Removing /home/./r429-write-control.txt Can't ls: "/home/./r429-write-control.txt" not found --- THE TEST: write into /.zfs/snapshot (must be REFUSED) Uploading … to /.zfs/snapshot/r429-write-attempt.txt dest open "/.zfs/snapshot/r429-write-attempt.txt": Failure --- confirm nothing was left behind: drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/. ``` **Identical on demo-felhom** (gid 1019). `storage-box-pool-1` **is** `u629488` (`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`) — the same box that demonstrably holds seven snapshots — so the emptiness is **per-sub-account filtering**, not absence. → **R-432**. ## What changed in the record | where | from | to | |---|---|---| | **R-429** | *"the mitigation has never been confirmed"* | **CLOSED — confirmed working.** My probe used `.snapshots`; the vendor documents `/.zfs/snapshot`. The controls were sound, the subject was wrong. What remains true is the actual finding: the row had no id, its "confirm tomorrow" went 36 days unanswered, and `DUE-CHECKS` was empty. **The finding was never the snapshots — it was that nobody could tell.** | | **R-95** | *"can delete, exposure open-ended"* — #1 since July | **RE-SCOPED:** deletes the live repo, **cannot write to the snapshots of it**; costs ≤1 day plus per-file recovery. **Ranking left to Viktor.** | | **07 §8 row 10** | *"the restic repo is NOT protected the same way"* | **both halves stated:** (a) the live repository is deletable — R-95 stands; (b) the snapshots are not writable by anything — proven. **Status NOT moved** — the recovery route has never been walked, which is what PARTIAL means. | | **07 §10.2** | R-95's line | gains the re-scope + citation | | **§11-D** | — | **untouched.** Nothing here answers whether the two Hetzner services share an account. | ## Part 3 — what normal looks like, in numbers Hub's own `reports` table: **12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.** - **Nine decreases in the entire history, and every one lands exactly on ZERO** — 36→0, 18→0 ×2, 15→0, 12→0, 8→0, 3→0. **Not one gradual retention decrease anywhere.** - **Every one predates `stats_known`** — the R-331 shape, a zero meaning *unmeasured*. Several carry a declared `State` (`needs_credential`, `awaiting_recovery_key`) saying so outright. - **In the `stats_known`-true window (380 reports) there are ZERO decreases**: demo-felhom flat at 10; demo-hp 67→68→69, rises only. **So observed churn gave nothing to calibrate against, and I say so rather than inventing a number.** ## The threshold, and where it came from **A fall of more than HALF the previous count, AND at least 5.** Reasoned from what retention *can* do, since it was never seen to do anything: `--keep-daily 7 --keep-weekly 4 --keep-monthly 6 --group-by host,tags` over ~8 apps **cannot halve a total** — those floors are per group — while a mass deletion goes to ~0. The floor of 5 stops a small-count box twitching. Deliberately not sensitive: a detector that cries wolf is switched off within a fortnight. ## Files, commits, deployment | file | what | |---|---| | `hub/internal/monitor/offsite.go` | +170 lines: third signal, three guards, escalation latch | | `hub/internal/monitor/offsite_r431_test.go` | NEW — 5 tests | | `hub/internal/monitor/testdata/r431_real_history.json` | NEW — 9 009 real points, committed | | `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go` | allowlist + operator-only, same commit | | `hub/CHANGELOG.md`, `CONTEXT.md`, `STATUS.md`, `07`, `00-capability-map.md`, both registers | the record | Commits: **`3068176`** (code + record), **`65c82c4`** (manifest bump). **Deployed hub `v0.111.0`.** ArgoCD: `sync=Synced health=Healthy`, `rev=65c82c4aa0949510670c42eb5c25515c25e12040` — **exactly HEAD**; `deployment "hub" successfully rolled out`; live image `gitea.dooplex.hu/admin/felhom-hub:0.111.0`; startup line `Offsite checker initialized: … 3 ok-seeded`. Image presence verified in the registry before syncing. ## Tests and red-proofs | test | result | |---|---| | `TestR431_FiresOnAMassDeletion` | 69→4 fires exactly once, severity in the vocabulary, message carries both numbers and does not claim loss | | `TestR431_SilentWhenNotTrustworthy` | silent on all four shapes (no `stats_known`, declared state, failed run, incomplete run) **and the baseline is left untouched** | | `TestR431_OrdinaryRetentionIsSilent` | 69→60 silent · 10→7 silent · 10→5 silent (exactly half is not *more than* half) · 69→34 fires | | `TestR431_EscalationOnlyLatch` | a continuing deletion pages once; a clean sweep re-arms; a later deletion fires again | | **`TestR431_RealHistoryProducesZeroAlarms`** | **9 009 real points, 2 customers → ZERO alarms.** The acceptance step. | Full hub suite green (`go build ./... && go vet ./... && go test ./...`). **Red-proofs, by name:** 1. **Threshold** — `snapshotDropFraction` 0.5 → 0.99: `TestR431_FiresOnAMassDeletion` **FAILS** (`want exactly 1 …, got 0`). 2. **`StatsKnown`** — guard removed: `TestR431_SilentWhenNotTrustworthy` **FAILS** (`stats_known absent …: must NOT alarm; got 1`) **and the real-history replay FAILS too**. 3. **Escalation-only** — latch removed: `TestR431_EscalationOnlyLatch` **FAILS** (`a CONTINUING deletion must not re-page …; got 2 alarms`). ## The live firing **Negative control first**, because an allowlist that accepts everything proves nothing: ``` POST /api/v1/event event_type=zzz_not_allowlisted_r431 → HTTP 400 "Invalid event_type: zzz_not_allowlisted_r431" POST /api/v1/event event_type=offsite_snapshots_dropped → HTTP 200 {"ok":true} ``` Hub log: `[INFO] Event from demo-hp: offsite_snapshots_dropped (error) — …` then `[INFO] Operator email sent for demo-hp/offsite_snapshots_dropped`. **Routing, from the live DB — this is the operator-only claim, with a control:** ``` 788|customer|skipped|operator_only 787|operator|sent| ``` Other types do reach customers (`escrow_blob_served|customer|41`, `whole_guest_backup_failed|customer|28`), so "operator-only" is a real distinction here rather than everything being operator. **The mail, rendered from the hub's own `FormatOperatorEmail`:** ``` SUBJECT: [Felhom] 🔴 demo-hp: offsite_snapshots_dropped Customer: demo-hp Event: offsite_snapshots_dropped Severity: error Time: 2026-09-01 14:30 CEST Message: Customer demo-hp: off-site backup count fell from 69 to 4 snapshot(s) in one report - more than retention can explain. The daily Storage Box snapshots are read-only and still hold the older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check whether a deletion ran on the box before restoring anything. Details: {"previous_count":69,"current_count":4,"drop":65} Dashboard: https://hub.felhom.eu/customers/demo-hp ``` **⚠ THAT MAIL IS REAL AND VIKTOR RECEIVED IT. Nothing was deleted** — it is the required live firing, flagged as item 5 in `STATUS.md` so it is not acted on. ## Not validated - **PROVEN-LIVE for the capability row.** The firing proves *delivery*, not that a genuine deletion is caught on a live box. The capability row says **IMPLEMENTED**, with that gap written into it. - **Whether a NAMED snapshot can be entered** even though the directory does not list (ZFS permits exactly that). One panel read from Viktor settles it — R-432. - **Whether a snapshot can be restored from.** Deliberately not attempted, in this session or the last. ## No controller release, no golden **No controller change. No version bump there. No image. No golden owed.** `golden_currency_gate.py` exits 0; golden and fleet floor remain **0.232.0**. ## Still open, named **R-430** (`unlock` reports success while the lock survives — **latent**: it bites only when the box loses delete, which is not happening here), **R-242's vouch half**, **R-402**, **R-409**, **R-401**, **R-412 leg 2**, **§11-D** (whether the two Hetzner services share an account), **R-427** (twelve open rows carrying a closed verdict), **R-432**. ## Register **Before:** OPEN 181 · CLOSED 161. **After:** OPEN 181 · CLOSED 163. Corrected and closed **R-429**; shipped and closed **R-431**; filed **R-432**; re-scoped **R-95** (stays open, ranking to Viktor). Both closed rows were **moved into `CLOSED-ITEMS.md`** rather than left in the open register with a closed verdict — that pile is R-427 and I did not add to it. ## Observations, and my own mistakes by name 1. **A correct instrument aimed at the wrong subject produces a confident wrong answer, and controls cannot save you from it.** My `.snapshots` probe had a good positive and negative control and was still worthless. **FILED: R-429** — corrected there, with the cause named. 2. **My mistake — I asserted a mechanism from a task brief without checking the vendor documentation**, and reported "the safety net cannot be seen" to Viktor on that basis. **NOT-A-FINDING: this is item 1 seen from the other side, already recorded in R-429 and as a ruling in `CONTEXT.md`, so a second row would duplicate rather than add anything.** 3. **My mistake — my first escalation-only test was HOLLOW and its red-proof passed.** It re-swept the same report, so the baseline had already moved to the new count and the latch was never consulted. Caught because red-proof 3 did **not** fail. Rewritten to drive a continuously falling count, which is the only shape where the latch is load-bearing; the red-proof then failed correctly. **NOT-A-FINDING: caught inside the session by the red-proof discipline doing exactly its job — a test whose red-proof passes is not a test, and that is why they are run.** 4. **My mistake — I queried the wrong table for the history.** The brief pointed at `host_reports`; that is the *agent's* host report and carries no `offsite` object at all (8 434 rows, zero hits). The controller's report lives in `reports`. **NOT-A-FINDING: found in one query by checking the parse count instead of trusting the pointer; the measurement in the changelog and the fixture both come from the right table.** 5. **My mistake — I posted the live firing to `/event` instead of `/api/v1/event`** and got HTTP 302 for *both* the control and the test. **Two identical results are an instrument fault, not two findings** — the same shape as yesterday's `sftp -p`. **NOT-A-FINDING: corrected in one command once I read the router's `TrimPrefix`; no wrong conclusion was recorded.** 6. **A sub-account cannot see inside the snapshot tree, so recovery is operator-only today.** **FILED: R-432** — and it decides whether R-95's remedy can ever be product-driven.