Opens with Part 1's answer because everything reads differently after it: a customer's own account can reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY. The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather than cited - the control write to the account home succeeded and was cleaned up, the write into /.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes. storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence - which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's remedy can ever be product-driven. STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today - the live firing was required to prove delivery and nothing was deleted. Five of my own mistakes are named, including the one that matters most: my first escalation-only test was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved and the latch was never consulted. That is why red-proofs are run.
12 KiB
REPORT — R-429 corrected, R-95 re-scoped, R-431 shipped (2026-09-01)
Part 1's answer, first, because everything reads differently after it
Can a customer's own account reach the snapshot tree? NO — it can reach the DOOR, and is refused writes to it, but the tree lists EMPTY.
Measured on both live boxes, over the credential each already holds, with controls in the same run.
=== POSITIVE CONTROL: the account home (must list) ===
drwxr-xr-x <user> 1058 4 Aug 22 03:10 ./.
drwx------ <user> 1058 3 Jul 23 09:53 ./.ssh
drwxrwxr-x <user> 1058 8 Aug 4 12:38 ./<repo>
=== NEGATIVE CONTROL: a name that cannot exist (must fail) ===
Can't ls: "/home/./zzz-no-such-r429" not found
=== CANDIDATE A: /.zfs/snapshot ===
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
drwxrwxrwx root root 0 Jul 21 16:01 /.zfs/snapshot/..
=== CANDIDATE B: /home/.zfs/snapshot ===
Can't ls: "/home/.zfs/snapshot" not found
=== CANDIDATE D: is the .zfs door itself visible? ===
drwxrwxrwx root root 2 Sep 1 12:12 /.zfs/shares
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot
The write-refusal test — the load-bearing sentence of the whole re-scope, now proven not cited:
--- control: the same write in the account home MUST succeed
sftp> put … ./r429-write-control.txt
Uploading … to /home/./r429-write-control.txt
-rw-r--r-- <user> 1058 11 Sep 1 12:12 ./r429-write-control.txt
--- cleanup of the control file
Removing /home/./r429-write-control.txt
Can't ls: "/home/./r429-write-control.txt" not found
--- THE TEST: write into /.zfs/snapshot (must be REFUSED)
Uploading … to /.zfs/snapshot/r429-write-attempt.txt
dest open "/.zfs/snapshot/r429-write-attempt.txt": Failure
--- confirm nothing was left behind:
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
Identical on demo-felhom (gid 1019). storage-box-pool-1 is u629488
(RUNBOOK-ep0-datastore-volume-2026-07-27.md:386) — the same box that demonstrably holds seven
snapshots — so the emptiness is per-sub-account filtering, not absence. → R-432.
What changed in the record
| where | from | to |
|---|---|---|
| R-429 | "the mitigation has never been confirmed" | CLOSED — confirmed working. My probe used .snapshots; the vendor documents /.zfs/snapshot. The controls were sound, the subject was wrong. What remains true is the actual finding: the row had no id, its "confirm tomorrow" went 36 days unanswered, and DUE-CHECKS was empty. The finding was never the snapshots — it was that nobody could tell. |
| R-95 | "can delete, exposure open-ended" — #1 since July | RE-SCOPED: deletes the live repo, cannot write to the snapshots of it; costs ≤1 day plus per-file recovery. Ranking left to Viktor. |
| 07 §8 row 10 | "the restic repo is NOT protected the same way" | both halves stated: (a) the live repository is deletable — R-95 stands; (b) the snapshots are not writable by anything — proven. Status NOT moved — the recovery route has never been walked, which is what PARTIAL means. |
| 07 §10.2 | R-95's line | gains the re-scope + citation |
| §11-D | — | untouched. Nothing here answers whether the two Hetzner services share an account. |
Part 3 — what normal looks like, in numbers
Hub's own reports table: 12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.
- Nine decreases in the entire history, and every one lands exactly on ZERO — 36→0, 18→0 ×2, 15→0, 12→0, 8→0, 3→0. Not one gradual retention decrease anywhere.
- Every one predates
stats_known— the R-331 shape, a zero meaning unmeasured. Several carry a declaredState(needs_credential,awaiting_recovery_key) saying so outright. - In the
stats_known-true window (380 reports) there are ZERO decreases: demo-felhom flat at 10; demo-hp 67→68→69, rises only.
So observed churn gave nothing to calibrate against, and I say so rather than inventing a number.
The threshold, and where it came from
A fall of more than HALF the previous count, AND at least 5. Reasoned from what retention can
do, since it was never seen to do anything: --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --group-by host,tags over ~8 apps cannot halve a total — those floors are per group — while a
mass deletion goes to ~0. The floor of 5 stops a small-count box twitching. Deliberately not
sensitive: a detector that cries wolf is switched off within a fortnight.
Files, commits, deployment
| file | what |
|---|---|
hub/internal/monitor/offsite.go |
+170 lines: third signal, three guards, escalation latch |
hub/internal/monitor/offsite_r431_test.go |
NEW — 5 tests |
hub/internal/monitor/testdata/r431_real_history.json |
NEW — 9 009 real points, committed |
hub/internal/api/handler.go, hub/internal/notify/dispatcher.go |
allowlist + operator-only, same commit |
hub/CHANGELOG.md, CONTEXT.md, STATUS.md, 07, 00-capability-map.md, both registers |
the record |
Commits: 3068176 (code + record), 65c82c4 (manifest bump).
Deployed hub v0.111.0. ArgoCD: sync=Synced health=Healthy,
rev=65c82c4aa0949510670c42eb5c25515c25e12040 — exactly HEAD; deployment "hub" successfully rolled out; live image gitea.dooplex.hu/admin/felhom-hub:0.111.0; startup line
Offsite checker initialized: … 3 ok-seeded. Image presence verified in the registry before syncing.
Tests and red-proofs
| test | result |
|---|---|
TestR431_FiresOnAMassDeletion |
69→4 fires exactly once, severity in the vocabulary, message carries both numbers and does not claim loss |
TestR431_SilentWhenNotTrustworthy |
silent on all four shapes (no stats_known, declared state, failed run, incomplete run) and the baseline is left untouched |
TestR431_OrdinaryRetentionIsSilent |
69→60 silent · 10→7 silent · 10→5 silent (exactly half is not more than half) · 69→34 fires |
TestR431_EscalationOnlyLatch |
a continuing deletion pages once; a clean sweep re-arms; a later deletion fires again |
TestR431_RealHistoryProducesZeroAlarms |
9 009 real points, 2 customers → ZERO alarms. The acceptance step. |
Full hub suite green (go build ./... && go vet ./... && go test ./...).
Red-proofs, by name:
- Threshold —
snapshotDropFraction0.5 → 0.99:TestR431_FiresOnAMassDeletionFAILS (want exactly 1 …, got 0). StatsKnown— guard removed:TestR431_SilentWhenNotTrustworthyFAILS (stats_known absent …: must NOT alarm; got 1) and the real-history replay FAILS too.- Escalation-only — latch removed:
TestR431_EscalationOnlyLatchFAILS (a CONTINUING deletion must not re-page …; got 2 alarms).
The live firing
Negative control first, because an allowlist that accepts everything proves nothing:
POST /api/v1/event event_type=zzz_not_allowlisted_r431 → HTTP 400 "Invalid event_type: zzz_not_allowlisted_r431"
POST /api/v1/event event_type=offsite_snapshots_dropped → HTTP 200 {"ok":true}
Hub log: [INFO] Event from demo-hp: offsite_snapshots_dropped (error) — … then
[INFO] Operator email sent for demo-hp/offsite_snapshots_dropped.
Routing, from the live DB — this is the operator-only claim, with a control:
788|customer|skipped|operator_only
787|operator|sent|
Other types do reach customers (escrow_blob_served|customer|41, whole_guest_backup_failed|customer|28),
so "operator-only" is a real distinction here rather than everything being operator.
The mail, rendered from the hub's own FormatOperatorEmail:
SUBJECT: [Felhom] 🔴 demo-hp: offsite_snapshots_dropped
Customer: demo-hp
Event: offsite_snapshots_dropped
Severity: error
Time: 2026-09-01 14:30 CEST
Message: Customer demo-hp: off-site backup count fell from 69 to 4 snapshot(s) in one report - more
than retention can explain. The daily Storage Box snapshots are read-only and still hold the
older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check whether
a deletion ran on the box before restoring anything.
Details: {"previous_count":69,"current_count":4,"drop":65}
Dashboard: https://hub.felhom.eu/customers/demo-hp
⚠ THAT MAIL IS REAL AND VIKTOR RECEIVED IT. Nothing was deleted — it is the required live firing,
flagged as item 5 in STATUS.md so it is not acted on.
Not validated
- PROVEN-LIVE for the capability row. The firing proves delivery, not that a genuine deletion is caught on a live box. The capability row says IMPLEMENTED, with that gap written into it.
- Whether a NAMED snapshot can be entered even though the directory does not list (ZFS permits exactly that). One panel read from Viktor settles it — R-432.
- Whether a snapshot can be restored from. Deliberately not attempted, in this session or the last.
No controller release, no golden
No controller change. No version bump there. No image. No golden owed.
golden_currency_gate.py exits 0; golden and fleet floor remain 0.232.0.
Still open, named
R-430 (unlock reports success while the lock survives — latent: it bites only when the box
loses delete, which is not happening here), R-242's vouch half, R-402, R-409, R-401,
R-412 leg 2, §11-D (whether the two Hetzner services share an account), R-427 (twelve open
rows carrying a closed verdict), R-432.
Register
Before: OPEN 181 · CLOSED 161. After: OPEN 181 · CLOSED 163.
Corrected and closed R-429; shipped and closed R-431; filed R-432; re-scoped R-95
(stays open, ranking to Viktor). Both closed rows were moved into CLOSED-ITEMS.md rather than
left in the open register with a closed verdict — that pile is R-427 and I did not add to it.
Observations, and my own mistakes by name
- A correct instrument aimed at the wrong subject produces a confident wrong answer, and controls
cannot save you from it. My
.snapshotsprobe had a good positive and negative control and was still worthless. FILED: R-429 — corrected there, with the cause named. - My mistake — I asserted a mechanism from a task brief without checking the vendor documentation,
and reported "the safety net cannot be seen" to Viktor on that basis.
NOT-A-FINDING: this is item 1 seen from the other side, already recorded in R-429 and as a
ruling in
CONTEXT.md, so a second row would duplicate rather than add anything. - My mistake — my first escalation-only test was HOLLOW and its red-proof passed. It re-swept the same report, so the baseline had already moved to the new count and the latch was never consulted. Caught because red-proof 3 did not fail. Rewritten to drive a continuously falling count, which is the only shape where the latch is load-bearing; the red-proof then failed correctly. NOT-A-FINDING: caught inside the session by the red-proof discipline doing exactly its job — a test whose red-proof passes is not a test, and that is why they are run.
- My mistake — I queried the wrong table for the history. The brief pointed at
host_reports; that is the agent's host report and carries nooffsiteobject at all (8 434 rows, zero hits). The controller's report lives inreports. NOT-A-FINDING: found in one query by checking the parse count instead of trusting the pointer; the measurement in the changelog and the fixture both come from the right table. - My mistake — I posted the live firing to
/eventinstead of/api/v1/eventand got HTTP 302 for both the control and the test. Two identical results are an instrument fault, not two findings — the same shape as yesterday'ssftp -p. NOT-A-FINDING: corrected in one command once I read the router'sTrimPrefix; no wrong conclusion was recorded. - A sub-account cannot see inside the snapshot tree, so recovery is operator-only today. FILED: R-432 — and it decides whether R-95's remedy can ever be product-driven.