17d92e71a1
gates / gates (push) Successful in 18s
Opens with Part 1's answer because everything reads differently after it: a customer's own account can reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY. The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather than cited - the control write to the account home succeeded and was cleaned up, the write into /.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes. storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence - which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's remedy can ever be product-driven. STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today - the live firing was required to prove delivery and nothing was deleted. Five of my own mistakes are named, including the one that matters most: my first escalation-only test was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved and the latch was never consulted. That is why red-proofs are run.
210 lines
12 KiB
Markdown
210 lines
12 KiB
Markdown
# REPORT — R-429 corrected, R-95 re-scoped, R-431 shipped (2026-09-01)
|
||
|
||
## Part 1's answer, first, because everything reads differently after it
|
||
|
||
**Can a customer's own account reach the snapshot tree? NO — it can reach the DOOR, and is refused
|
||
writes to it, but the tree lists EMPTY.**
|
||
|
||
Measured on **both** live boxes, over the credential each already holds, with controls in the same run.
|
||
|
||
```
|
||
=== POSITIVE CONTROL: the account home (must list) ===
|
||
drwxr-xr-x <user> 1058 4 Aug 22 03:10 ./.
|
||
drwx------ <user> 1058 3 Jul 23 09:53 ./.ssh
|
||
drwxrwxr-x <user> 1058 8 Aug 4 12:38 ./<repo>
|
||
=== NEGATIVE CONTROL: a name that cannot exist (must fail) ===
|
||
Can't ls: "/home/./zzz-no-such-r429" not found
|
||
=== CANDIDATE A: /.zfs/snapshot ===
|
||
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
|
||
drwxrwxrwx root root 0 Jul 21 16:01 /.zfs/snapshot/..
|
||
=== CANDIDATE B: /home/.zfs/snapshot ===
|
||
Can't ls: "/home/.zfs/snapshot" not found
|
||
=== CANDIDATE D: is the .zfs door itself visible? ===
|
||
drwxrwxrwx root root 2 Sep 1 12:12 /.zfs/shares
|
||
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot
|
||
```
|
||
|
||
**The write-refusal test — the load-bearing sentence of the whole re-scope, now proven not cited:**
|
||
|
||
```
|
||
--- control: the same write in the account home MUST succeed
|
||
sftp> put … ./r429-write-control.txt
|
||
Uploading … to /home/./r429-write-control.txt
|
||
-rw-r--r-- <user> 1058 11 Sep 1 12:12 ./r429-write-control.txt
|
||
--- cleanup of the control file
|
||
Removing /home/./r429-write-control.txt
|
||
Can't ls: "/home/./r429-write-control.txt" not found
|
||
--- THE TEST: write into /.zfs/snapshot (must be REFUSED)
|
||
Uploading … to /.zfs/snapshot/r429-write-attempt.txt
|
||
dest open "/.zfs/snapshot/r429-write-attempt.txt": Failure
|
||
--- confirm nothing was left behind:
|
||
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
|
||
```
|
||
|
||
**Identical on demo-felhom** (gid 1019). `storage-box-pool-1` **is** `u629488`
|
||
(`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`) — the same box that demonstrably holds seven
|
||
snapshots — so the emptiness is **per-sub-account filtering**, not absence. → **R-432**.
|
||
|
||
## What changed in the record
|
||
|
||
| where | from | to |
|
||
|---|---|---|
|
||
| **R-429** | *"the mitigation has never been confirmed"* | **CLOSED — confirmed working.** My probe used `.snapshots`; the vendor documents `/.zfs/snapshot`. The controls were sound, the subject was wrong. What remains true is the actual finding: the row had no id, its "confirm tomorrow" went 36 days unanswered, and `DUE-CHECKS` was empty. **The finding was never the snapshots — it was that nobody could tell.** |
|
||
| **R-95** | *"can delete, exposure open-ended"* — #1 since July | **RE-SCOPED:** deletes the live repo, **cannot write to the snapshots of it**; costs ≤1 day plus per-file recovery. **Ranking left to Viktor.** |
|
||
| **07 §8 row 10** | *"the restic repo is NOT protected the same way"* | **both halves stated:** (a) the live repository is deletable — R-95 stands; (b) the snapshots are not writable by anything — proven. **Status NOT moved** — the recovery route has never been walked, which is what PARTIAL means. |
|
||
| **07 §10.2** | R-95's line | gains the re-scope + citation |
|
||
| **§11-D** | — | **untouched.** Nothing here answers whether the two Hetzner services share an account. |
|
||
|
||
## Part 3 — what normal looks like, in numbers
|
||
|
||
Hub's own `reports` table: **12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.**
|
||
|
||
- **Nine decreases in the entire history, and every one lands exactly on ZERO** — 36→0, 18→0 ×2,
|
||
15→0, 12→0, 8→0, 3→0. **Not one gradual retention decrease anywhere.**
|
||
- **Every one predates `stats_known`** — the R-331 shape, a zero meaning *unmeasured*. Several carry a
|
||
declared `State` (`needs_credential`, `awaiting_recovery_key`) saying so outright.
|
||
- **In the `stats_known`-true window (380 reports) there are ZERO decreases**: demo-felhom flat at 10;
|
||
demo-hp 67→68→69, rises only.
|
||
|
||
**So observed churn gave nothing to calibrate against, and I say so rather than inventing a number.**
|
||
|
||
## The threshold, and where it came from
|
||
|
||
**A fall of more than HALF the previous count, AND at least 5.** Reasoned from what retention *can*
|
||
do, since it was never seen to do anything: `--keep-daily 7 --keep-weekly 4 --keep-monthly 6
|
||
--group-by host,tags` over ~8 apps **cannot halve a total** — those floors are per group — while a
|
||
mass deletion goes to ~0. The floor of 5 stops a small-count box twitching. Deliberately not
|
||
sensitive: a detector that cries wolf is switched off within a fortnight.
|
||
|
||
## Files, commits, deployment
|
||
|
||
| file | what |
|
||
|---|---|
|
||
| `hub/internal/monitor/offsite.go` | +170 lines: third signal, three guards, escalation latch |
|
||
| `hub/internal/monitor/offsite_r431_test.go` | NEW — 5 tests |
|
||
| `hub/internal/monitor/testdata/r431_real_history.json` | NEW — 9 009 real points, committed |
|
||
| `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go` | allowlist + operator-only, same commit |
|
||
| `hub/CHANGELOG.md`, `CONTEXT.md`, `STATUS.md`, `07`, `00-capability-map.md`, both registers | the record |
|
||
|
||
Commits: **`3068176`** (code + record), **`65c82c4`** (manifest bump).
|
||
**Deployed hub `v0.111.0`.** ArgoCD: `sync=Synced health=Healthy`,
|
||
`rev=65c82c4aa0949510670c42eb5c25515c25e12040` — **exactly HEAD**; `deployment "hub" successfully
|
||
rolled out`; live image `gitea.dooplex.hu/admin/felhom-hub:0.111.0`; startup line
|
||
`Offsite checker initialized: … 3 ok-seeded`. Image presence verified in the registry before syncing.
|
||
|
||
## Tests and red-proofs
|
||
|
||
| test | result |
|
||
|---|---|
|
||
| `TestR431_FiresOnAMassDeletion` | 69→4 fires exactly once, severity in the vocabulary, message carries both numbers and does not claim loss |
|
||
| `TestR431_SilentWhenNotTrustworthy` | silent on all four shapes (no `stats_known`, declared state, failed run, incomplete run) **and the baseline is left untouched** |
|
||
| `TestR431_OrdinaryRetentionIsSilent` | 69→60 silent · 10→7 silent · 10→5 silent (exactly half is not *more than* half) · 69→34 fires |
|
||
| `TestR431_EscalationOnlyLatch` | a continuing deletion pages once; a clean sweep re-arms; a later deletion fires again |
|
||
| **`TestR431_RealHistoryProducesZeroAlarms`** | **9 009 real points, 2 customers → ZERO alarms.** The acceptance step. |
|
||
|
||
Full hub suite green (`go build ./... && go vet ./... && go test ./...`).
|
||
|
||
**Red-proofs, by name:**
|
||
|
||
1. **Threshold** — `snapshotDropFraction` 0.5 → 0.99: `TestR431_FiresOnAMassDeletion` **FAILS**
|
||
(`want exactly 1 …, got 0`).
|
||
2. **`StatsKnown`** — guard removed: `TestR431_SilentWhenNotTrustworthy` **FAILS**
|
||
(`stats_known absent …: must NOT alarm; got 1`) **and the real-history replay FAILS too**.
|
||
3. **Escalation-only** — latch removed: `TestR431_EscalationOnlyLatch` **FAILS**
|
||
(`a CONTINUING deletion must not re-page …; got 2 alarms`).
|
||
|
||
## The live firing
|
||
|
||
**Negative control first**, because an allowlist that accepts everything proves nothing:
|
||
|
||
```
|
||
POST /api/v1/event event_type=zzz_not_allowlisted_r431 → HTTP 400 "Invalid event_type: zzz_not_allowlisted_r431"
|
||
POST /api/v1/event event_type=offsite_snapshots_dropped → HTTP 200 {"ok":true}
|
||
```
|
||
|
||
Hub log: `[INFO] Event from demo-hp: offsite_snapshots_dropped (error) — …` then
|
||
`[INFO] Operator email sent for demo-hp/offsite_snapshots_dropped`.
|
||
|
||
**Routing, from the live DB — this is the operator-only claim, with a control:**
|
||
|
||
```
|
||
788|customer|skipped|operator_only
|
||
787|operator|sent|
|
||
```
|
||
|
||
Other types do reach customers (`escrow_blob_served|customer|41`, `whole_guest_backup_failed|customer|28`),
|
||
so "operator-only" is a real distinction here rather than everything being operator.
|
||
|
||
**The mail, rendered from the hub's own `FormatOperatorEmail`:**
|
||
|
||
```
|
||
SUBJECT: [Felhom] 🔴 demo-hp: offsite_snapshots_dropped
|
||
Customer: demo-hp
|
||
Event: offsite_snapshots_dropped
|
||
Severity: error
|
||
Time: 2026-09-01 14:30 CEST
|
||
Message: Customer demo-hp: off-site backup count fell from 69 to 4 snapshot(s) in one report - more
|
||
than retention can explain. The daily Storage Box snapshots are read-only and still hold the
|
||
older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check whether
|
||
a deletion ran on the box before restoring anything.
|
||
Details: {"previous_count":69,"current_count":4,"drop":65}
|
||
Dashboard: https://hub.felhom.eu/customers/demo-hp
|
||
```
|
||
|
||
**⚠ THAT MAIL IS REAL AND VIKTOR RECEIVED IT. Nothing was deleted** — it is the required live firing,
|
||
flagged as item 5 in `STATUS.md` so it is not acted on.
|
||
|
||
## Not validated
|
||
|
||
- **PROVEN-LIVE for the capability row.** The firing proves *delivery*, not that a genuine deletion is
|
||
caught on a live box. The capability row says **IMPLEMENTED**, with that gap written into it.
|
||
- **Whether a NAMED snapshot can be entered** even though the directory does not list (ZFS permits
|
||
exactly that). One panel read from Viktor settles it — R-432.
|
||
- **Whether a snapshot can be restored from.** Deliberately not attempted, in this session or the last.
|
||
|
||
## No controller release, no golden
|
||
|
||
**No controller change. No version bump there. No image. No golden owed.**
|
||
`golden_currency_gate.py` exits 0; golden and fleet floor remain **0.232.0**.
|
||
|
||
## Still open, named
|
||
|
||
**R-430** (`unlock` reports success while the lock survives — **latent**: it bites only when the box
|
||
loses delete, which is not happening here), **R-242's vouch half**, **R-402**, **R-409**, **R-401**,
|
||
**R-412 leg 2**, **§11-D** (whether the two Hetzner services share an account), **R-427** (twelve open
|
||
rows carrying a closed verdict), **R-432**.
|
||
|
||
## Register
|
||
|
||
**Before:** OPEN 181 · CLOSED 161. **After:** OPEN 181 · CLOSED 163.
|
||
Corrected and closed **R-429**; shipped and closed **R-431**; filed **R-432**; re-scoped **R-95**
|
||
(stays open, ranking to Viktor). Both closed rows were **moved into `CLOSED-ITEMS.md`** rather than
|
||
left in the open register with a closed verdict — that pile is R-427 and I did not add to it.
|
||
|
||
## Observations, and my own mistakes by name
|
||
|
||
1. **A correct instrument aimed at the wrong subject produces a confident wrong answer, and controls
|
||
cannot save you from it.** My `.snapshots` probe had a good positive and negative control and was
|
||
still worthless. **FILED: R-429** — corrected there, with the cause named.
|
||
2. **My mistake — I asserted a mechanism from a task brief without checking the vendor documentation**,
|
||
and reported "the safety net cannot be seen" to Viktor on that basis.
|
||
**NOT-A-FINDING: this is item 1 seen from the other side, already recorded in R-429 and as a
|
||
ruling in `CONTEXT.md`, so a second row would duplicate rather than add anything.**
|
||
3. **My mistake — my first escalation-only test was HOLLOW and its red-proof passed.** It re-swept the
|
||
same report, so the baseline had already moved to the new count and the latch was never consulted.
|
||
Caught because red-proof 3 did **not** fail. Rewritten to drive a continuously falling count, which
|
||
is the only shape where the latch is load-bearing; the red-proof then failed correctly.
|
||
**NOT-A-FINDING: caught inside the session by the red-proof discipline doing exactly its job — a
|
||
test whose red-proof passes is not a test, and that is why they are run.**
|
||
4. **My mistake — I queried the wrong table for the history.** The brief pointed at `host_reports`;
|
||
that is the *agent's* host report and carries no `offsite` object at all (8 434 rows, zero hits).
|
||
The controller's report lives in `reports`. **NOT-A-FINDING: found in one query by checking the
|
||
parse count instead of trusting the pointer; the measurement in the changelog and the fixture both
|
||
come from the right table.**
|
||
5. **My mistake — I posted the live firing to `/event` instead of `/api/v1/event`** and got HTTP 302
|
||
for *both* the control and the test. **Two identical results are an instrument fault, not two
|
||
findings** — the same shape as yesterday's `sftp -p`. **NOT-A-FINDING: corrected in one command
|
||
once I read the router's `TrimPrefix`; no wrong conclusion was recorded.**
|
||
6. **A sub-account cannot see inside the snapshot tree, so recovery is operator-only today.**
|
||
**FILED: R-432** — and it decides whether R-95's remedy can ever be product-driven.
|