REPORT + CONTEXT + STATUS: the record corrected, and the thing worth building shipped
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
Opens with Part 1's answer because everything reads differently after it: a customer's own account can reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY. The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather than cited - the control write to the account home succeeded and was cleaned up, the write into /.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes. storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence - which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's remedy can ever be product-driven. STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today - the live firing was required to prove delivery and nothing was deleted. Five of my own mistakes are named, including the one that matters most: my first escalation-only test was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved and the latch was never consulted. That is why red-proofs are run.
This commit is contained in:
@@ -0,0 +1,209 @@
|
||||
# REPORT — R-429 corrected, R-95 re-scoped, R-431 shipped (2026-09-01)
|
||||
|
||||
## Part 1's answer, first, because everything reads differently after it
|
||||
|
||||
**Can a customer's own account reach the snapshot tree? NO — it can reach the DOOR, and is refused
|
||||
writes to it, but the tree lists EMPTY.**
|
||||
|
||||
Measured on **both** live boxes, over the credential each already holds, with controls in the same run.
|
||||
|
||||
```
|
||||
=== POSITIVE CONTROL: the account home (must list) ===
|
||||
drwxr-xr-x <user> 1058 4 Aug 22 03:10 ./.
|
||||
drwx------ <user> 1058 3 Jul 23 09:53 ./.ssh
|
||||
drwxrwxr-x <user> 1058 8 Aug 4 12:38 ./<repo>
|
||||
=== NEGATIVE CONTROL: a name that cannot exist (must fail) ===
|
||||
Can't ls: "/home/./zzz-no-such-r429" not found
|
||||
=== CANDIDATE A: /.zfs/snapshot ===
|
||||
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
|
||||
drwxrwxrwx root root 0 Jul 21 16:01 /.zfs/snapshot/..
|
||||
=== CANDIDATE B: /home/.zfs/snapshot ===
|
||||
Can't ls: "/home/.zfs/snapshot" not found
|
||||
=== CANDIDATE D: is the .zfs door itself visible? ===
|
||||
drwxrwxrwx root root 2 Sep 1 12:12 /.zfs/shares
|
||||
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot
|
||||
```
|
||||
|
||||
**The write-refusal test — the load-bearing sentence of the whole re-scope, now proven not cited:**
|
||||
|
||||
```
|
||||
--- control: the same write in the account home MUST succeed
|
||||
sftp> put … ./r429-write-control.txt
|
||||
Uploading … to /home/./r429-write-control.txt
|
||||
-rw-r--r-- <user> 1058 11 Sep 1 12:12 ./r429-write-control.txt
|
||||
--- cleanup of the control file
|
||||
Removing /home/./r429-write-control.txt
|
||||
Can't ls: "/home/./r429-write-control.txt" not found
|
||||
--- THE TEST: write into /.zfs/snapshot (must be REFUSED)
|
||||
Uploading … to /.zfs/snapshot/r429-write-attempt.txt
|
||||
dest open "/.zfs/snapshot/r429-write-attempt.txt": Failure
|
||||
--- confirm nothing was left behind:
|
||||
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
|
||||
```
|
||||
|
||||
**Identical on demo-felhom** (gid 1019). `storage-box-pool-1` **is** `u629488`
|
||||
(`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`) — the same box that demonstrably holds seven
|
||||
snapshots — so the emptiness is **per-sub-account filtering**, not absence. → **R-432**.
|
||||
|
||||
## What changed in the record
|
||||
|
||||
| where | from | to |
|
||||
|---|---|---|
|
||||
| **R-429** | *"the mitigation has never been confirmed"* | **CLOSED — confirmed working.** My probe used `.snapshots`; the vendor documents `/.zfs/snapshot`. The controls were sound, the subject was wrong. What remains true is the actual finding: the row had no id, its "confirm tomorrow" went 36 days unanswered, and `DUE-CHECKS` was empty. **The finding was never the snapshots — it was that nobody could tell.** |
|
||||
| **R-95** | *"can delete, exposure open-ended"* — #1 since July | **RE-SCOPED:** deletes the live repo, **cannot write to the snapshots of it**; costs ≤1 day plus per-file recovery. **Ranking left to Viktor.** |
|
||||
| **07 §8 row 10** | *"the restic repo is NOT protected the same way"* | **both halves stated:** (a) the live repository is deletable — R-95 stands; (b) the snapshots are not writable by anything — proven. **Status NOT moved** — the recovery route has never been walked, which is what PARTIAL means. |
|
||||
| **07 §10.2** | R-95's line | gains the re-scope + citation |
|
||||
| **§11-D** | — | **untouched.** Nothing here answers whether the two Hetzner services share an account. |
|
||||
|
||||
## Part 3 — what normal looks like, in numbers
|
||||
|
||||
Hub's own `reports` table: **12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.**
|
||||
|
||||
- **Nine decreases in the entire history, and every one lands exactly on ZERO** — 36→0, 18→0 ×2,
|
||||
15→0, 12→0, 8→0, 3→0. **Not one gradual retention decrease anywhere.**
|
||||
- **Every one predates `stats_known`** — the R-331 shape, a zero meaning *unmeasured*. Several carry a
|
||||
declared `State` (`needs_credential`, `awaiting_recovery_key`) saying so outright.
|
||||
- **In the `stats_known`-true window (380 reports) there are ZERO decreases**: demo-felhom flat at 10;
|
||||
demo-hp 67→68→69, rises only.
|
||||
|
||||
**So observed churn gave nothing to calibrate against, and I say so rather than inventing a number.**
|
||||
|
||||
## The threshold, and where it came from
|
||||
|
||||
**A fall of more than HALF the previous count, AND at least 5.** Reasoned from what retention *can*
|
||||
do, since it was never seen to do anything: `--keep-daily 7 --keep-weekly 4 --keep-monthly 6
|
||||
--group-by host,tags` over ~8 apps **cannot halve a total** — those floors are per group — while a
|
||||
mass deletion goes to ~0. The floor of 5 stops a small-count box twitching. Deliberately not
|
||||
sensitive: a detector that cries wolf is switched off within a fortnight.
|
||||
|
||||
## Files, commits, deployment
|
||||
|
||||
| file | what |
|
||||
|---|---|
|
||||
| `hub/internal/monitor/offsite.go` | +170 lines: third signal, three guards, escalation latch |
|
||||
| `hub/internal/monitor/offsite_r431_test.go` | NEW — 5 tests |
|
||||
| `hub/internal/monitor/testdata/r431_real_history.json` | NEW — 9 009 real points, committed |
|
||||
| `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go` | allowlist + operator-only, same commit |
|
||||
| `hub/CHANGELOG.md`, `CONTEXT.md`, `STATUS.md`, `07`, `00-capability-map.md`, both registers | the record |
|
||||
|
||||
Commits: **`3068176`** (code + record), **`65c82c4`** (manifest bump).
|
||||
**Deployed hub `v0.111.0`.** ArgoCD: `sync=Synced health=Healthy`,
|
||||
`rev=65c82c4aa0949510670c42eb5c25515c25e12040` — **exactly HEAD**; `deployment "hub" successfully
|
||||
rolled out`; live image `gitea.dooplex.hu/admin/felhom-hub:0.111.0`; startup line
|
||||
`Offsite checker initialized: … 3 ok-seeded`. Image presence verified in the registry before syncing.
|
||||
|
||||
## Tests and red-proofs
|
||||
|
||||
| test | result |
|
||||
|---|---|
|
||||
| `TestR431_FiresOnAMassDeletion` | 69→4 fires exactly once, severity in the vocabulary, message carries both numbers and does not claim loss |
|
||||
| `TestR431_SilentWhenNotTrustworthy` | silent on all four shapes (no `stats_known`, declared state, failed run, incomplete run) **and the baseline is left untouched** |
|
||||
| `TestR431_OrdinaryRetentionIsSilent` | 69→60 silent · 10→7 silent · 10→5 silent (exactly half is not *more than* half) · 69→34 fires |
|
||||
| `TestR431_EscalationOnlyLatch` | a continuing deletion pages once; a clean sweep re-arms; a later deletion fires again |
|
||||
| **`TestR431_RealHistoryProducesZeroAlarms`** | **9 009 real points, 2 customers → ZERO alarms.** The acceptance step. |
|
||||
|
||||
Full hub suite green (`go build ./... && go vet ./... && go test ./...`).
|
||||
|
||||
**Red-proofs, by name:**
|
||||
|
||||
1. **Threshold** — `snapshotDropFraction` 0.5 → 0.99: `TestR431_FiresOnAMassDeletion` **FAILS**
|
||||
(`want exactly 1 …, got 0`).
|
||||
2. **`StatsKnown`** — guard removed: `TestR431_SilentWhenNotTrustworthy` **FAILS**
|
||||
(`stats_known absent …: must NOT alarm; got 1`) **and the real-history replay FAILS too**.
|
||||
3. **Escalation-only** — latch removed: `TestR431_EscalationOnlyLatch` **FAILS**
|
||||
(`a CONTINUING deletion must not re-page …; got 2 alarms`).
|
||||
|
||||
## The live firing
|
||||
|
||||
**Negative control first**, because an allowlist that accepts everything proves nothing:
|
||||
|
||||
```
|
||||
POST /api/v1/event event_type=zzz_not_allowlisted_r431 → HTTP 400 "Invalid event_type: zzz_not_allowlisted_r431"
|
||||
POST /api/v1/event event_type=offsite_snapshots_dropped → HTTP 200 {"ok":true}
|
||||
```
|
||||
|
||||
Hub log: `[INFO] Event from demo-hp: offsite_snapshots_dropped (error) — …` then
|
||||
`[INFO] Operator email sent for demo-hp/offsite_snapshots_dropped`.
|
||||
|
||||
**Routing, from the live DB — this is the operator-only claim, with a control:**
|
||||
|
||||
```
|
||||
788|customer|skipped|operator_only
|
||||
787|operator|sent|
|
||||
```
|
||||
|
||||
Other types do reach customers (`escrow_blob_served|customer|41`, `whole_guest_backup_failed|customer|28`),
|
||||
so "operator-only" is a real distinction here rather than everything being operator.
|
||||
|
||||
**The mail, rendered from the hub's own `FormatOperatorEmail`:**
|
||||
|
||||
```
|
||||
SUBJECT: [Felhom] 🔴 demo-hp: offsite_snapshots_dropped
|
||||
Customer: demo-hp
|
||||
Event: offsite_snapshots_dropped
|
||||
Severity: error
|
||||
Time: 2026-09-01 14:30 CEST
|
||||
Message: Customer demo-hp: off-site backup count fell from 69 to 4 snapshot(s) in one report - more
|
||||
than retention can explain. The daily Storage Box snapshots are read-only and still hold the
|
||||
older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check whether
|
||||
a deletion ran on the box before restoring anything.
|
||||
Details: {"previous_count":69,"current_count":4,"drop":65}
|
||||
Dashboard: https://hub.felhom.eu/customers/demo-hp
|
||||
```
|
||||
|
||||
**⚠ THAT MAIL IS REAL AND VIKTOR RECEIVED IT. Nothing was deleted** — it is the required live firing,
|
||||
flagged as item 5 in `STATUS.md` so it is not acted on.
|
||||
|
||||
## Not validated
|
||||
|
||||
- **PROVEN-LIVE for the capability row.** The firing proves *delivery*, not that a genuine deletion is
|
||||
caught on a live box. The capability row says **IMPLEMENTED**, with that gap written into it.
|
||||
- **Whether a NAMED snapshot can be entered** even though the directory does not list (ZFS permits
|
||||
exactly that). One panel read from Viktor settles it — R-432.
|
||||
- **Whether a snapshot can be restored from.** Deliberately not attempted, in this session or the last.
|
||||
|
||||
## No controller release, no golden
|
||||
|
||||
**No controller change. No version bump there. No image. No golden owed.**
|
||||
`golden_currency_gate.py` exits 0; golden and fleet floor remain **0.232.0**.
|
||||
|
||||
## Still open, named
|
||||
|
||||
**R-430** (`unlock` reports success while the lock survives — **latent**: it bites only when the box
|
||||
loses delete, which is not happening here), **R-242's vouch half**, **R-402**, **R-409**, **R-401**,
|
||||
**R-412 leg 2**, **§11-D** (whether the two Hetzner services share an account), **R-427** (twelve open
|
||||
rows carrying a closed verdict), **R-432**.
|
||||
|
||||
## Register
|
||||
|
||||
**Before:** OPEN 181 · CLOSED 161. **After:** OPEN 181 · CLOSED 163.
|
||||
Corrected and closed **R-429**; shipped and closed **R-431**; filed **R-432**; re-scoped **R-95**
|
||||
(stays open, ranking to Viktor). Both closed rows were **moved into `CLOSED-ITEMS.md`** rather than
|
||||
left in the open register with a closed verdict — that pile is R-427 and I did not add to it.
|
||||
|
||||
## Observations, and my own mistakes by name
|
||||
|
||||
1. **A correct instrument aimed at the wrong subject produces a confident wrong answer, and controls
|
||||
cannot save you from it.** My `.snapshots` probe had a good positive and negative control and was
|
||||
still worthless. **FILED: R-429** — corrected there, with the cause named.
|
||||
2. **My mistake — I asserted a mechanism from a task brief without checking the vendor documentation**,
|
||||
and reported "the safety net cannot be seen" to Viktor on that basis.
|
||||
**NOT-A-FINDING: this is item 1 seen from the other side, already recorded in R-429 and as a
|
||||
ruling in `CONTEXT.md`, so a second row would duplicate rather than add anything.**
|
||||
3. **My mistake — my first escalation-only test was HOLLOW and its red-proof passed.** It re-swept the
|
||||
same report, so the baseline had already moved to the new count and the latch was never consulted.
|
||||
Caught because red-proof 3 did **not** fail. Rewritten to drive a continuously falling count, which
|
||||
is the only shape where the latch is load-bearing; the red-proof then failed correctly.
|
||||
**NOT-A-FINDING: caught inside the session by the red-proof discipline doing exactly its job — a
|
||||
test whose red-proof passes is not a test, and that is why they are run.**
|
||||
4. **My mistake — I queried the wrong table for the history.** The brief pointed at `host_reports`;
|
||||
that is the *agent's* host report and carries no `offsite` object at all (8 434 rows, zero hits).
|
||||
The controller's report lives in `reports`. **NOT-A-FINDING: found in one query by checking the
|
||||
parse count instead of trusting the pointer; the measurement in the changelog and the fixture both
|
||||
come from the right table.**
|
||||
5. **My mistake — I posted the live firing to `/event` instead of `/api/v1/event`** and got HTTP 302
|
||||
for *both* the control and the test. **Two identical results are an instrument fault, not two
|
||||
findings** — the same shape as yesterday's `sftp -p`. **NOT-A-FINDING: corrected in one command
|
||||
once I read the router's `TrimPrefix`; no wrong conclusion was recorded.**
|
||||
6. **A sub-account cannot see inside the snapshot tree, so recovery is operator-only today.**
|
||||
**FILED: R-432** — and it decides whether R-95's remedy can ever be product-driven.
|
||||
Reference in New Issue
Block a user