Files
felhom.eu/REPORT-r431-snapshots.md
admin 17d92e71a1
gates / gates (push) Successful in 18s
REPORT + CONTEXT + STATUS: the record corrected, and the thing worth building shipped
Opens with Part 1's answer because everything reads differently after it: a customer's own account can
reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY.

The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather
than cited - the control write to the account home succeeded and was cleaned up, the write into
/.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes.

storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence -
which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's
remedy can ever be product-driven.

STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two
minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today -
the live firing was required to prove delivery and nothing was deleted.

Five of my own mistakes are named, including the one that matters most: my first escalation-only test
was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved
and the latch was never consulted. That is why red-proofs are run.
2026-09-01 14:33:16 +02:00

12 KiB
Raw Permalink Blame History

REPORT — R-429 corrected, R-95 re-scoped, R-431 shipped (2026-09-01)

Part 1's answer, first, because everything reads differently after it

Can a customer's own account reach the snapshot tree? NO — it can reach the DOOR, and is refused writes to it, but the tree lists EMPTY.

Measured on both live boxes, over the credential each already holds, with controls in the same run.

=== POSITIVE CONTROL: the account home (must list) ===
    drwxr-xr-x  <user> 1058   4 Aug 22 03:10 ./.
    drwx------  <user> 1058   3 Jul 23 09:53 ./.ssh
    drwxrwxr-x  <user> 1058   8 Aug  4 12:38 ./<repo>
=== NEGATIVE CONTROL: a name that cannot exist (must fail) ===
    Can't ls: "/home/./zzz-no-such-r429" not found
=== CANDIDATE A: /.zfs/snapshot ===
    drwxrwxrwx  root root   2 Jan  1  1970 /.zfs/snapshot/.
    drwxrwxrwx  root root   0 Jul 21 16:01 /.zfs/snapshot/..
=== CANDIDATE B: /home/.zfs/snapshot ===
    Can't ls: "/home/.zfs/snapshot" not found
=== CANDIDATE D: is the .zfs door itself visible? ===
    drwxrwxrwx  root root   2 Sep  1 12:12 /.zfs/shares
    drwxrwxrwx  root root   2 Jan  1  1970 /.zfs/snapshot

The write-refusal test — the load-bearing sentence of the whole re-scope, now proven not cited:

--- control: the same write in the account home MUST succeed
    sftp> put … ./r429-write-control.txt
    Uploading … to /home/./r429-write-control.txt
    -rw-r--r--  <user> 1058  11 Sep  1 12:12 ./r429-write-control.txt
--- cleanup of the control file
    Removing /home/./r429-write-control.txt
    Can't ls: "/home/./r429-write-control.txt" not found
--- THE TEST: write into /.zfs/snapshot (must be REFUSED)
    Uploading … to /.zfs/snapshot/r429-write-attempt.txt
    dest open "/.zfs/snapshot/r429-write-attempt.txt": Failure
--- confirm nothing was left behind:
    drwxrwxrwx  root root  2 Jan  1  1970 /.zfs/snapshot/.

Identical on demo-felhom (gid 1019). storage-box-pool-1 is u629488 (RUNBOOK-ep0-datastore-volume-2026-07-27.md:386) — the same box that demonstrably holds seven snapshots — so the emptiness is per-sub-account filtering, not absence. → R-432.

What changed in the record

where from to
R-429 "the mitigation has never been confirmed" CLOSED — confirmed working. My probe used .snapshots; the vendor documents /.zfs/snapshot. The controls were sound, the subject was wrong. What remains true is the actual finding: the row had no id, its "confirm tomorrow" went 36 days unanswered, and DUE-CHECKS was empty. The finding was never the snapshots — it was that nobody could tell.
R-95 "can delete, exposure open-ended" — #1 since July RE-SCOPED: deletes the live repo, cannot write to the snapshots of it; costs ≤1 day plus per-file recovery. Ranking left to Viktor.
07 §8 row 10 "the restic repo is NOT protected the same way" both halves stated: (a) the live repository is deletable — R-95 stands; (b) the snapshots are not writable by anything — proven. Status NOT moved — the recovery route has never been walked, which is what PARTIAL means.
07 §10.2 R-95's line gains the re-scope + citation
§11-D — untouched. Nothing here answers whether the two Hetzner services share an account.

Part 3 — what normal looks like, in numbers

Hub's own reports table: 12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.

  • Nine decreases in the entire history, and every one lands exactly on ZERO — 36→0, 18→0 ×2, 15→0, 12→0, 8→0, 3→0. Not one gradual retention decrease anywhere.
  • Every one predates stats_known — the R-331 shape, a zero meaning unmeasured. Several carry a declared State (needs_credential, awaiting_recovery_key) saying so outright.
  • In the stats_known-true window (380 reports) there are ZERO decreases: demo-felhom flat at 10; demo-hp 67→68→69, rises only.

So observed churn gave nothing to calibrate against, and I say so rather than inventing a number.

The threshold, and where it came from

A fall of more than HALF the previous count, AND at least 5. Reasoned from what retention can do, since it was never seen to do anything: --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --group-by host,tags over ~8 apps cannot halve a total — those floors are per group — while a mass deletion goes to ~0. The floor of 5 stops a small-count box twitching. Deliberately not sensitive: a detector that cries wolf is switched off within a fortnight.

Files, commits, deployment

file what
hub/internal/monitor/offsite.go +170 lines: third signal, three guards, escalation latch
hub/internal/monitor/offsite_r431_test.go NEW — 5 tests
hub/internal/monitor/testdata/r431_real_history.json NEW — 9 009 real points, committed
hub/internal/api/handler.go, hub/internal/notify/dispatcher.go allowlist + operator-only, same commit
hub/CHANGELOG.md, CONTEXT.md, STATUS.md, 07, 00-capability-map.md, both registers the record

Commits: 3068176 (code + record), 65c82c4 (manifest bump). Deployed hub v0.111.0. ArgoCD: sync=Synced health=Healthy, rev=65c82c4aa0949510670c42eb5c25515c25e12040 — exactly HEAD; deployment "hub" successfully rolled out; live image gitea.dooplex.hu/admin/felhom-hub:0.111.0; startup line Offsite checker initialized: … 3 ok-seeded. Image presence verified in the registry before syncing.

Tests and red-proofs

test result
TestR431_FiresOnAMassDeletion 69→4 fires exactly once, severity in the vocabulary, message carries both numbers and does not claim loss
TestR431_SilentWhenNotTrustworthy silent on all four shapes (no stats_known, declared state, failed run, incomplete run) and the baseline is left untouched
TestR431_OrdinaryRetentionIsSilent 69→60 silent · 10→7 silent · 10→5 silent (exactly half is not more than half) · 69→34 fires
TestR431_EscalationOnlyLatch a continuing deletion pages once; a clean sweep re-arms; a later deletion fires again
TestR431_RealHistoryProducesZeroAlarms 9 009 real points, 2 customers → ZERO alarms. The acceptance step.

Full hub suite green (go build ./... && go vet ./... && go test ./...).

Red-proofs, by name:

  1. Threshold — snapshotDropFraction 0.5 → 0.99: TestR431_FiresOnAMassDeletion FAILS (want exactly 1 …, got 0).
  2. StatsKnown — guard removed: TestR431_SilentWhenNotTrustworthy FAILS (stats_known absent …: must NOT alarm; got 1) and the real-history replay FAILS too.
  3. Escalation-only — latch removed: TestR431_EscalationOnlyLatch FAILS (a CONTINUING deletion must not re-page …; got 2 alarms).

The live firing

Negative control first, because an allowlist that accepts everything proves nothing:

POST /api/v1/event  event_type=zzz_not_allowlisted_r431  → HTTP 400  "Invalid event_type: zzz_not_allowlisted_r431"
POST /api/v1/event  event_type=offsite_snapshots_dropped → HTTP 200  {"ok":true}

Hub log: [INFO] Event from demo-hp: offsite_snapshots_dropped (error) — … then [INFO] Operator email sent for demo-hp/offsite_snapshots_dropped.

Routing, from the live DB — this is the operator-only claim, with a control:

788|customer|skipped|operator_only
787|operator|sent|

Other types do reach customers (escrow_blob_served|customer|41, whole_guest_backup_failed|customer|28), so "operator-only" is a real distinction here rather than everything being operator.

The mail, rendered from the hub's own FormatOperatorEmail:

SUBJECT: [Felhom] 🔴 demo-hp: offsite_snapshots_dropped
Customer: demo-hp
Event:    offsite_snapshots_dropped
Severity: error
Time:     2026-09-01 14:30 CEST
Message:  Customer demo-hp: off-site backup count fell from 69 to 4 snapshot(s) in one report - more
          than retention can explain. The daily Storage Box snapshots are read-only and still hold the
          older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check whether
          a deletion ran on the box before restoring anything.
Details:  {"previous_count":69,"current_count":4,"drop":65}
Dashboard: https://hub.felhom.eu/customers/demo-hp

⚠ THAT MAIL IS REAL AND VIKTOR RECEIVED IT. Nothing was deleted — it is the required live firing, flagged as item 5 in STATUS.md so it is not acted on.

Not validated

  • PROVEN-LIVE for the capability row. The firing proves delivery, not that a genuine deletion is caught on a live box. The capability row says IMPLEMENTED, with that gap written into it.
  • Whether a NAMED snapshot can be entered even though the directory does not list (ZFS permits exactly that). One panel read from Viktor settles it — R-432.
  • Whether a snapshot can be restored from. Deliberately not attempted, in this session or the last.

No controller release, no golden

No controller change. No version bump there. No image. No golden owed. golden_currency_gate.py exits 0; golden and fleet floor remain 0.232.0.

Still open, named

R-430 (unlock reports success while the lock survives — latent: it bites only when the box loses delete, which is not happening here), R-242's vouch half, R-402, R-409, R-401, R-412 leg 2, §11-D (whether the two Hetzner services share an account), R-427 (twelve open rows carrying a closed verdict), R-432.

Register

Before: OPEN 181 · CLOSED 161. After: OPEN 181 · CLOSED 163. Corrected and closed R-429; shipped and closed R-431; filed R-432; re-scoped R-95 (stays open, ranking to Viktor). Both closed rows were moved into CLOSED-ITEMS.md rather than left in the open register with a closed verdict — that pile is R-427 and I did not add to it.

Observations, and my own mistakes by name

  1. A correct instrument aimed at the wrong subject produces a confident wrong answer, and controls cannot save you from it. My .snapshots probe had a good positive and negative control and was still worthless. FILED: R-429 — corrected there, with the cause named.
  2. My mistake — I asserted a mechanism from a task brief without checking the vendor documentation, and reported "the safety net cannot be seen" to Viktor on that basis. NOT-A-FINDING: this is item 1 seen from the other side, already recorded in R-429 and as a ruling in CONTEXT.md, so a second row would duplicate rather than add anything.
  3. My mistake — my first escalation-only test was HOLLOW and its red-proof passed. It re-swept the same report, so the baseline had already moved to the new count and the latch was never consulted. Caught because red-proof 3 did not fail. Rewritten to drive a continuously falling count, which is the only shape where the latch is load-bearing; the red-proof then failed correctly. NOT-A-FINDING: caught inside the session by the red-proof discipline doing exactly its job — a test whose red-proof passes is not a test, and that is why they are run.
  4. My mistake — I queried the wrong table for the history. The brief pointed at host_reports; that is the agent's host report and carries no offsite object at all (8 434 rows, zero hits). The controller's report lives in reports. NOT-A-FINDING: found in one query by checking the parse count instead of trusting the pointer; the measurement in the changelog and the fixture both come from the right table.
  5. My mistake — I posted the live firing to /event instead of /api/v1/event and got HTTP 302 for both the control and the test. Two identical results are an instrument fault, not two findings — the same shape as yesterday's sftp -p. NOT-A-FINDING: corrected in one command once I read the router's TrimPrefix; no wrong conclusion was recorded.
  6. A sub-account cannot see inside the snapshot tree, so recovery is operator-only today. FILED: R-432 — and it decides whether R-95's remedy can ever be product-driven.