Files
felhom.eu/REPORT-r431-snapshots.md
admin 17d92e71a1
gates / gates (push) Successful in 18s
REPORT + CONTEXT + STATUS: the record corrected, and the thing worth building shipped
Opens with Part 1's answer because everything reads differently after it: a customer's own account can
reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY.

The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather
than cited - the control write to the account home succeeded and was cleaned up, the write into
/.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes.

storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence -
which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's
remedy can ever be product-driven.

STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two
minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today -
the live firing was required to prove delivery and nothing was deleted.

Five of my own mistakes are named, including the one that matters most: my first escalation-only test
was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved
and the latch was never consulted. That is why red-proofs are run.
2026-09-01 14:33:16 +02:00

210 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — R-429 corrected, R-95 re-scoped, R-431 shipped (2026-09-01)
## Part 1's answer, first, because everything reads differently after it
**Can a customer's own account reach the snapshot tree? NO — it can reach the DOOR, and is refused
writes to it, but the tree lists EMPTY.**
Measured on **both** live boxes, over the credential each already holds, with controls in the same run.
```
=== POSITIVE CONTROL: the account home (must list) ===
drwxr-xr-x <user> 1058 4 Aug 22 03:10 ./.
drwx------ <user> 1058 3 Jul 23 09:53 ./.ssh
drwxrwxr-x <user> 1058 8 Aug 4 12:38 ./<repo>
=== NEGATIVE CONTROL: a name that cannot exist (must fail) ===
Can't ls: "/home/./zzz-no-such-r429" not found
=== CANDIDATE A: /.zfs/snapshot ===
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
drwxrwxrwx root root 0 Jul 21 16:01 /.zfs/snapshot/..
=== CANDIDATE B: /home/.zfs/snapshot ===
Can't ls: "/home/.zfs/snapshot" not found
=== CANDIDATE D: is the .zfs door itself visible? ===
drwxrwxrwx root root 2 Sep 1 12:12 /.zfs/shares
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot
```
**The write-refusal test — the load-bearing sentence of the whole re-scope, now proven not cited:**
```
--- control: the same write in the account home MUST succeed
sftp> put … ./r429-write-control.txt
Uploading … to /home/./r429-write-control.txt
-rw-r--r-- <user> 1058 11 Sep 1 12:12 ./r429-write-control.txt
--- cleanup of the control file
Removing /home/./r429-write-control.txt
Can't ls: "/home/./r429-write-control.txt" not found
--- THE TEST: write into /.zfs/snapshot (must be REFUSED)
Uploading … to /.zfs/snapshot/r429-write-attempt.txt
dest open "/.zfs/snapshot/r429-write-attempt.txt": Failure
--- confirm nothing was left behind:
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
```
**Identical on demo-felhom** (gid 1019). `storage-box-pool-1` **is** `u629488`
(`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`) — the same box that demonstrably holds seven
snapshots — so the emptiness is **per-sub-account filtering**, not absence. → **R-432**.
## What changed in the record
| where | from | to |
|---|---|---|
| **R-429** | *"the mitigation has never been confirmed"* | **CLOSED — confirmed working.** My probe used `.snapshots`; the vendor documents `/.zfs/snapshot`. The controls were sound, the subject was wrong. What remains true is the actual finding: the row had no id, its "confirm tomorrow" went 36 days unanswered, and `DUE-CHECKS` was empty. **The finding was never the snapshots — it was that nobody could tell.** |
| **R-95** | *"can delete, exposure open-ended"* — #1 since July | **RE-SCOPED:** deletes the live repo, **cannot write to the snapshots of it**; costs ≤1 day plus per-file recovery. **Ranking left to Viktor.** |
| **07 §8 row 10** | *"the restic repo is NOT protected the same way"* | **both halves stated:** (a) the live repository is deletable — R-95 stands; (b) the snapshots are not writable by anything — proven. **Status NOT moved** — the recovery route has never been walked, which is what PARTIAL means. |
| **07 §10.2** | R-95's line | gains the re-scope + citation |
| **§11-D** | — | **untouched.** Nothing here answers whether the two Hetzner services share an account. |
## Part 3 — what normal looks like, in numbers
Hub's own `reports` table: **12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.**
- **Nine decreases in the entire history, and every one lands exactly on ZERO** — 36→0, 18→0 ×2,
15→0, 12→0, 8→0, 3→0. **Not one gradual retention decrease anywhere.**
- **Every one predates `stats_known`** — the R-331 shape, a zero meaning *unmeasured*. Several carry a
declared `State` (`needs_credential`, `awaiting_recovery_key`) saying so outright.
- **In the `stats_known`-true window (380 reports) there are ZERO decreases**: demo-felhom flat at 10;
demo-hp 67→68→69, rises only.
**So observed churn gave nothing to calibrate against, and I say so rather than inventing a number.**
## The threshold, and where it came from
**A fall of more than HALF the previous count, AND at least 5.** Reasoned from what retention *can*
do, since it was never seen to do anything: `--keep-daily 7 --keep-weekly 4 --keep-monthly 6
--group-by host,tags` over ~8 apps **cannot halve a total** — those floors are per group — while a
mass deletion goes to ~0. The floor of 5 stops a small-count box twitching. Deliberately not
sensitive: a detector that cries wolf is switched off within a fortnight.
## Files, commits, deployment
| file | what |
|---|---|
| `hub/internal/monitor/offsite.go` | +170 lines: third signal, three guards, escalation latch |
| `hub/internal/monitor/offsite_r431_test.go` | NEW — 5 tests |
| `hub/internal/monitor/testdata/r431_real_history.json` | NEW — 9 009 real points, committed |
| `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go` | allowlist + operator-only, same commit |
| `hub/CHANGELOG.md`, `CONTEXT.md`, `STATUS.md`, `07`, `00-capability-map.md`, both registers | the record |
Commits: **`3068176`** (code + record), **`65c82c4`** (manifest bump).
**Deployed hub `v0.111.0`.** ArgoCD: `sync=Synced health=Healthy`,
`rev=65c82c4aa0949510670c42eb5c25515c25e12040` — **exactly HEAD**; `deployment "hub" successfully
rolled out`; live image `gitea.dooplex.hu/admin/felhom-hub:0.111.0`; startup line
`Offsite checker initialized: … 3 ok-seeded`. Image presence verified in the registry before syncing.
## Tests and red-proofs
| test | result |
|---|---|
| `TestR431_FiresOnAMassDeletion` | 69→4 fires exactly once, severity in the vocabulary, message carries both numbers and does not claim loss |
| `TestR431_SilentWhenNotTrustworthy` | silent on all four shapes (no `stats_known`, declared state, failed run, incomplete run) **and the baseline is left untouched** |
| `TestR431_OrdinaryRetentionIsSilent` | 69→60 silent · 10→7 silent · 10→5 silent (exactly half is not *more than* half) · 69→34 fires |
| `TestR431_EscalationOnlyLatch` | a continuing deletion pages once; a clean sweep re-arms; a later deletion fires again |
| **`TestR431_RealHistoryProducesZeroAlarms`** | **9 009 real points, 2 customers → ZERO alarms.** The acceptance step. |
Full hub suite green (`go build ./... && go vet ./... && go test ./...`).
**Red-proofs, by name:**
1. **Threshold** — `snapshotDropFraction` 0.5 → 0.99: `TestR431_FiresOnAMassDeletion` **FAILS**
(`want exactly 1 …, got 0`).
2. **`StatsKnown`** — guard removed: `TestR431_SilentWhenNotTrustworthy` **FAILS**
(`stats_known absent …: must NOT alarm; got 1`) **and the real-history replay FAILS too**.
3. **Escalation-only** — latch removed: `TestR431_EscalationOnlyLatch` **FAILS**
(`a CONTINUING deletion must not re-page …; got 2 alarms`).
## The live firing
**Negative control first**, because an allowlist that accepts everything proves nothing:
```
POST /api/v1/event event_type=zzz_not_allowlisted_r431 → HTTP 400 "Invalid event_type: zzz_not_allowlisted_r431"
POST /api/v1/event event_type=offsite_snapshots_dropped → HTTP 200 {"ok":true}
```
Hub log: `[INFO] Event from demo-hp: offsite_snapshots_dropped (error) — …` then
`[INFO] Operator email sent for demo-hp/offsite_snapshots_dropped`.
**Routing, from the live DB — this is the operator-only claim, with a control:**
```
788|customer|skipped|operator_only
787|operator|sent|
```
Other types do reach customers (`escrow_blob_served|customer|41`, `whole_guest_backup_failed|customer|28`),
so "operator-only" is a real distinction here rather than everything being operator.
**The mail, rendered from the hub's own `FormatOperatorEmail`:**
```
SUBJECT: [Felhom] 🔴 demo-hp: offsite_snapshots_dropped
Customer: demo-hp
Event: offsite_snapshots_dropped
Severity: error
Time: 2026-09-01 14:30 CEST
Message: Customer demo-hp: off-site backup count fell from 69 to 4 snapshot(s) in one report - more
than retention can explain. The daily Storage Box snapshots are read-only and still hold the
older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check whether
a deletion ran on the box before restoring anything.
Details: {"previous_count":69,"current_count":4,"drop":65}
Dashboard: https://hub.felhom.eu/customers/demo-hp
```
**⚠ THAT MAIL IS REAL AND VIKTOR RECEIVED IT. Nothing was deleted** — it is the required live firing,
flagged as item 5 in `STATUS.md` so it is not acted on.
## Not validated
- **PROVEN-LIVE for the capability row.** The firing proves *delivery*, not that a genuine deletion is
caught on a live box. The capability row says **IMPLEMENTED**, with that gap written into it.
- **Whether a NAMED snapshot can be entered** even though the directory does not list (ZFS permits
exactly that). One panel read from Viktor settles it — R-432.
- **Whether a snapshot can be restored from.** Deliberately not attempted, in this session or the last.
## No controller release, no golden
**No controller change. No version bump there. No image. No golden owed.**
`golden_currency_gate.py` exits 0; golden and fleet floor remain **0.232.0**.
## Still open, named
**R-430** (`unlock` reports success while the lock survives — **latent**: it bites only when the box
loses delete, which is not happening here), **R-242's vouch half**, **R-402**, **R-409**, **R-401**,
**R-412 leg 2**, **§11-D** (whether the two Hetzner services share an account), **R-427** (twelve open
rows carrying a closed verdict), **R-432**.
## Register
**Before:** OPEN 181 · CLOSED 161. **After:** OPEN 181 · CLOSED 163.
Corrected and closed **R-429**; shipped and closed **R-431**; filed **R-432**; re-scoped **R-95**
(stays open, ranking to Viktor). Both closed rows were **moved into `CLOSED-ITEMS.md`** rather than
left in the open register with a closed verdict — that pile is R-427 and I did not add to it.
## Observations, and my own mistakes by name
1. **A correct instrument aimed at the wrong subject produces a confident wrong answer, and controls
cannot save you from it.** My `.snapshots` probe had a good positive and negative control and was
still worthless. **FILED: R-429** — corrected there, with the cause named.
2. **My mistake — I asserted a mechanism from a task brief without checking the vendor documentation**,
and reported "the safety net cannot be seen" to Viktor on that basis.
**NOT-A-FINDING: this is item 1 seen from the other side, already recorded in R-429 and as a
ruling in `CONTEXT.md`, so a second row would duplicate rather than add anything.**
3. **My mistake — my first escalation-only test was HOLLOW and its red-proof passed.** It re-swept the
same report, so the baseline had already moved to the new count and the latch was never consulted.
Caught because red-proof 3 did **not** fail. Rewritten to drive a continuously falling count, which
is the only shape where the latch is load-bearing; the red-proof then failed correctly.
**NOT-A-FINDING: caught inside the session by the red-proof discipline doing exactly its job — a
test whose red-proof passes is not a test, and that is why they are run.**
4. **My mistake — I queried the wrong table for the history.** The brief pointed at `host_reports`;
that is the *agent's* host report and carries no `offsite` object at all (8 434 rows, zero hits).
The controller's report lives in `reports`. **NOT-A-FINDING: found in one query by checking the
parse count instead of trusting the pointer; the measurement in the changelog and the fixture both
come from the right table.**
5. **My mistake — I posted the live firing to `/event` instead of `/api/v1/event`** and got HTTP 302
for *both* the control and the test. **Two identical results are an instrument fault, not two
findings** — the same shape as yesterday's `sftp -p`. **NOT-A-FINDING: corrected in one command
once I read the router's `TrimPrefix`; no wrong conclusion was recorded.**
6. **A sub-account cannot see inside the snapshot tree, so recovery is operator-only today.**
**FILED: R-432** — and it decides whether R-95's remedy can ever be product-driven.