REPORT + CONTEXT + STATUS: the record corrected, and the thing worth building shipped
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
Opens with Part 1's answer because everything reads differently after it: a customer's own account can reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY. The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather than cited - the control write to the account home succeeded and was cleaned up, the write into /.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes. storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence - which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's remedy can ever be product-driven. STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today - the live firing was required to prove delivery and nothing was deleted. Five of my own mistakes are named, including the one that matters most: my first escalation-only test was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved and the latch was never consulted. That is why red-proofs are run.
This commit is contained in:
+36
@@ -14,6 +14,42 @@
|
||||
> language, one screen, no identifiers in the prose. Same subjects, different readers; merging them
|
||||
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
## The record told a worse story than the truth for two months, and my own probe is why (2026-09-01, R-429 / R-95 / R-431)
|
||||
|
||||
**THREE RULINGS.**
|
||||
|
||||
**1. A correct instrument pointed at the wrong subject produces a confident wrong answer, and the
|
||||
controls cannot save you.** The 2026-08-31/09-01 spike probed for a directory called `.snapshots` and
|
||||
reported *"the safety net cannot be seen from the box"*. Its positive and negative controls were
|
||||
sound. **The vendor documents the path as `/.zfs/snapshot`.** So `not found` was true and meant
|
||||
nothing, and the register carried "the exposure is unbounded" on the strength of it. Re-probed at the
|
||||
documented path: `/.zfs` lists from inside the jail, and **a write into `/.zfs/snapshot` is REFUSED**
|
||||
(`dest open …: Failure`) while the identical write to the account home succeeds. **Controls prove the
|
||||
instrument works; they say nothing about whether it is aimed at the question.** Cite the vendor before
|
||||
asserting a mechanism — the task that set the spike asserted it without a citation, and I did not check.
|
||||
|
||||
**2. The real R-95 is much smaller than two months of register text.** The box can delete its own
|
||||
**live** repository, but it **cannot write to the daily snapshots of it** — measured, both boxes. So a
|
||||
deletion costs at most the data written since the last daily snapshot, and the rest comes back **file
|
||||
by file**, one customer at a time, with no effect on anyone else. **Not open-ended loss.** Two limits
|
||||
kept honest: a sub-account sees the snapshot directory EMPTY, so per-file recovery is **operator-only
|
||||
today** (R-432), and a panel-driven restore rolls back the whole Storage Box. **R-95's ranking is
|
||||
Viktor's** — it has been #1 since July on the old story.
|
||||
|
||||
**3. Detection belongs on the HUB, not the box, and the threshold is reasoned rather than invented.**
|
||||
The event being detected is a box deleting its own backups, so a detector living on that box is one
|
||||
the same event can silence. The hub already receives the count and keeps the history.
|
||||
**The measurement is the interesting part:** across 12 898 reports (2026-06-05 → 2026-09-01) every one
|
||||
of the nine decreases lands exactly on **zero**, and every one predates `stats_known` — they are the
|
||||
R-331 shape, a zero meaning *unmeasured*, several carrying a declared `State` that says so outright.
|
||||
In the 380-report window where `stats_known` is true there are **zero** decreases. **So observed churn
|
||||
gave nothing to calibrate against, and saying so beats inventing a number (R-401).** The threshold
|
||||
therefore comes from what retention *can* do: keeping 7 daily + 4 weekly + 6 monthly per group means
|
||||
it **cannot halve a total**, while a mass deletion goes to ~0 — hence *more than half, and at least
|
||||
five*. Three pre-conditions guard it, each with a scar: `StatsKnown` (R-331), the declared `State`
|
||||
(R-204), and run success (R-100, whose lesson lives in that very file). **Acceptance: 9 009 real
|
||||
report points replayed → zero alarms.**
|
||||
|
||||
## The checks were the one thing nothing checked — 16 of 29 could be fooled by a label (2026-09-01, R-421)
|
||||
|
||||
**THE CLASS: an instrument that matches a LABEL rather than the fact it names.** Five instances, and
|
||||
|
||||
@@ -0,0 +1,209 @@
|
||||
# REPORT — R-429 corrected, R-95 re-scoped, R-431 shipped (2026-09-01)
|
||||
|
||||
## Part 1's answer, first, because everything reads differently after it
|
||||
|
||||
**Can a customer's own account reach the snapshot tree? NO — it can reach the DOOR, and is refused
|
||||
writes to it, but the tree lists EMPTY.**
|
||||
|
||||
Measured on **both** live boxes, over the credential each already holds, with controls in the same run.
|
||||
|
||||
```
|
||||
=== POSITIVE CONTROL: the account home (must list) ===
|
||||
drwxr-xr-x <user> 1058 4 Aug 22 03:10 ./.
|
||||
drwx------ <user> 1058 3 Jul 23 09:53 ./.ssh
|
||||
drwxrwxr-x <user> 1058 8 Aug 4 12:38 ./<repo>
|
||||
=== NEGATIVE CONTROL: a name that cannot exist (must fail) ===
|
||||
Can't ls: "/home/./zzz-no-such-r429" not found
|
||||
=== CANDIDATE A: /.zfs/snapshot ===
|
||||
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
|
||||
drwxrwxrwx root root 0 Jul 21 16:01 /.zfs/snapshot/..
|
||||
=== CANDIDATE B: /home/.zfs/snapshot ===
|
||||
Can't ls: "/home/.zfs/snapshot" not found
|
||||
=== CANDIDATE D: is the .zfs door itself visible? ===
|
||||
drwxrwxrwx root root 2 Sep 1 12:12 /.zfs/shares
|
||||
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot
|
||||
```
|
||||
|
||||
**The write-refusal test — the load-bearing sentence of the whole re-scope, now proven not cited:**
|
||||
|
||||
```
|
||||
--- control: the same write in the account home MUST succeed
|
||||
sftp> put … ./r429-write-control.txt
|
||||
Uploading … to /home/./r429-write-control.txt
|
||||
-rw-r--r-- <user> 1058 11 Sep 1 12:12 ./r429-write-control.txt
|
||||
--- cleanup of the control file
|
||||
Removing /home/./r429-write-control.txt
|
||||
Can't ls: "/home/./r429-write-control.txt" not found
|
||||
--- THE TEST: write into /.zfs/snapshot (must be REFUSED)
|
||||
Uploading … to /.zfs/snapshot/r429-write-attempt.txt
|
||||
dest open "/.zfs/snapshot/r429-write-attempt.txt": Failure
|
||||
--- confirm nothing was left behind:
|
||||
drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/.
|
||||
```
|
||||
|
||||
**Identical on demo-felhom** (gid 1019). `storage-box-pool-1` **is** `u629488`
|
||||
(`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`) — the same box that demonstrably holds seven
|
||||
snapshots — so the emptiness is **per-sub-account filtering**, not absence. → **R-432**.
|
||||
|
||||
## What changed in the record
|
||||
|
||||
| where | from | to |
|
||||
|---|---|---|
|
||||
| **R-429** | *"the mitigation has never been confirmed"* | **CLOSED — confirmed working.** My probe used `.snapshots`; the vendor documents `/.zfs/snapshot`. The controls were sound, the subject was wrong. What remains true is the actual finding: the row had no id, its "confirm tomorrow" went 36 days unanswered, and `DUE-CHECKS` was empty. **The finding was never the snapshots — it was that nobody could tell.** |
|
||||
| **R-95** | *"can delete, exposure open-ended"* — #1 since July | **RE-SCOPED:** deletes the live repo, **cannot write to the snapshots of it**; costs ≤1 day plus per-file recovery. **Ranking left to Viktor.** |
|
||||
| **07 §8 row 10** | *"the restic repo is NOT protected the same way"* | **both halves stated:** (a) the live repository is deletable — R-95 stands; (b) the snapshots are not writable by anything — proven. **Status NOT moved** — the recovery route has never been walked, which is what PARTIAL means. |
|
||||
| **07 §10.2** | R-95's line | gains the re-scope + citation |
|
||||
| **§11-D** | — | **untouched.** Nothing here answers whether the two Hetzner services share an account. |
|
||||
|
||||
## Part 3 — what normal looks like, in numbers
|
||||
|
||||
Hub's own `reports` table: **12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.**
|
||||
|
||||
- **Nine decreases in the entire history, and every one lands exactly on ZERO** — 36→0, 18→0 ×2,
|
||||
15→0, 12→0, 8→0, 3→0. **Not one gradual retention decrease anywhere.**
|
||||
- **Every one predates `stats_known`** — the R-331 shape, a zero meaning *unmeasured*. Several carry a
|
||||
declared `State` (`needs_credential`, `awaiting_recovery_key`) saying so outright.
|
||||
- **In the `stats_known`-true window (380 reports) there are ZERO decreases**: demo-felhom flat at 10;
|
||||
demo-hp 67→68→69, rises only.
|
||||
|
||||
**So observed churn gave nothing to calibrate against, and I say so rather than inventing a number.**
|
||||
|
||||
## The threshold, and where it came from
|
||||
|
||||
**A fall of more than HALF the previous count, AND at least 5.** Reasoned from what retention *can*
|
||||
do, since it was never seen to do anything: `--keep-daily 7 --keep-weekly 4 --keep-monthly 6
|
||||
--group-by host,tags` over ~8 apps **cannot halve a total** — those floors are per group — while a
|
||||
mass deletion goes to ~0. The floor of 5 stops a small-count box twitching. Deliberately not
|
||||
sensitive: a detector that cries wolf is switched off within a fortnight.
|
||||
|
||||
## Files, commits, deployment
|
||||
|
||||
| file | what |
|
||||
|---|---|
|
||||
| `hub/internal/monitor/offsite.go` | +170 lines: third signal, three guards, escalation latch |
|
||||
| `hub/internal/monitor/offsite_r431_test.go` | NEW — 5 tests |
|
||||
| `hub/internal/monitor/testdata/r431_real_history.json` | NEW — 9 009 real points, committed |
|
||||
| `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go` | allowlist + operator-only, same commit |
|
||||
| `hub/CHANGELOG.md`, `CONTEXT.md`, `STATUS.md`, `07`, `00-capability-map.md`, both registers | the record |
|
||||
|
||||
Commits: **`3068176`** (code + record), **`65c82c4`** (manifest bump).
|
||||
**Deployed hub `v0.111.0`.** ArgoCD: `sync=Synced health=Healthy`,
|
||||
`rev=65c82c4aa0949510670c42eb5c25515c25e12040` — **exactly HEAD**; `deployment "hub" successfully
|
||||
rolled out`; live image `gitea.dooplex.hu/admin/felhom-hub:0.111.0`; startup line
|
||||
`Offsite checker initialized: … 3 ok-seeded`. Image presence verified in the registry before syncing.
|
||||
|
||||
## Tests and red-proofs
|
||||
|
||||
| test | result |
|
||||
|---|---|
|
||||
| `TestR431_FiresOnAMassDeletion` | 69→4 fires exactly once, severity in the vocabulary, message carries both numbers and does not claim loss |
|
||||
| `TestR431_SilentWhenNotTrustworthy` | silent on all four shapes (no `stats_known`, declared state, failed run, incomplete run) **and the baseline is left untouched** |
|
||||
| `TestR431_OrdinaryRetentionIsSilent` | 69→60 silent · 10→7 silent · 10→5 silent (exactly half is not *more than* half) · 69→34 fires |
|
||||
| `TestR431_EscalationOnlyLatch` | a continuing deletion pages once; a clean sweep re-arms; a later deletion fires again |
|
||||
| **`TestR431_RealHistoryProducesZeroAlarms`** | **9 009 real points, 2 customers → ZERO alarms.** The acceptance step. |
|
||||
|
||||
Full hub suite green (`go build ./... && go vet ./... && go test ./...`).
|
||||
|
||||
**Red-proofs, by name:**
|
||||
|
||||
1. **Threshold** — `snapshotDropFraction` 0.5 → 0.99: `TestR431_FiresOnAMassDeletion` **FAILS**
|
||||
(`want exactly 1 …, got 0`).
|
||||
2. **`StatsKnown`** — guard removed: `TestR431_SilentWhenNotTrustworthy` **FAILS**
|
||||
(`stats_known absent …: must NOT alarm; got 1`) **and the real-history replay FAILS too**.
|
||||
3. **Escalation-only** — latch removed: `TestR431_EscalationOnlyLatch` **FAILS**
|
||||
(`a CONTINUING deletion must not re-page …; got 2 alarms`).
|
||||
|
||||
## The live firing
|
||||
|
||||
**Negative control first**, because an allowlist that accepts everything proves nothing:
|
||||
|
||||
```
|
||||
POST /api/v1/event event_type=zzz_not_allowlisted_r431 → HTTP 400 "Invalid event_type: zzz_not_allowlisted_r431"
|
||||
POST /api/v1/event event_type=offsite_snapshots_dropped → HTTP 200 {"ok":true}
|
||||
```
|
||||
|
||||
Hub log: `[INFO] Event from demo-hp: offsite_snapshots_dropped (error) — …` then
|
||||
`[INFO] Operator email sent for demo-hp/offsite_snapshots_dropped`.
|
||||
|
||||
**Routing, from the live DB — this is the operator-only claim, with a control:**
|
||||
|
||||
```
|
||||
788|customer|skipped|operator_only
|
||||
787|operator|sent|
|
||||
```
|
||||
|
||||
Other types do reach customers (`escrow_blob_served|customer|41`, `whole_guest_backup_failed|customer|28`),
|
||||
so "operator-only" is a real distinction here rather than everything being operator.
|
||||
|
||||
**The mail, rendered from the hub's own `FormatOperatorEmail`:**
|
||||
|
||||
```
|
||||
SUBJECT: [Felhom] 🔴 demo-hp: offsite_snapshots_dropped
|
||||
Customer: demo-hp
|
||||
Event: offsite_snapshots_dropped
|
||||
Severity: error
|
||||
Time: 2026-09-01 14:30 CEST
|
||||
Message: Customer demo-hp: off-site backup count fell from 69 to 4 snapshot(s) in one report - more
|
||||
than retention can explain. The daily Storage Box snapshots are read-only and still hold the
|
||||
older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check whether
|
||||
a deletion ran on the box before restoring anything.
|
||||
Details: {"previous_count":69,"current_count":4,"drop":65}
|
||||
Dashboard: https://hub.felhom.eu/customers/demo-hp
|
||||
```
|
||||
|
||||
**⚠ THAT MAIL IS REAL AND VIKTOR RECEIVED IT. Nothing was deleted** — it is the required live firing,
|
||||
flagged as item 5 in `STATUS.md` so it is not acted on.
|
||||
|
||||
## Not validated
|
||||
|
||||
- **PROVEN-LIVE for the capability row.** The firing proves *delivery*, not that a genuine deletion is
|
||||
caught on a live box. The capability row says **IMPLEMENTED**, with that gap written into it.
|
||||
- **Whether a NAMED snapshot can be entered** even though the directory does not list (ZFS permits
|
||||
exactly that). One panel read from Viktor settles it — R-432.
|
||||
- **Whether a snapshot can be restored from.** Deliberately not attempted, in this session or the last.
|
||||
|
||||
## No controller release, no golden
|
||||
|
||||
**No controller change. No version bump there. No image. No golden owed.**
|
||||
`golden_currency_gate.py` exits 0; golden and fleet floor remain **0.232.0**.
|
||||
|
||||
## Still open, named
|
||||
|
||||
**R-430** (`unlock` reports success while the lock survives — **latent**: it bites only when the box
|
||||
loses delete, which is not happening here), **R-242's vouch half**, **R-402**, **R-409**, **R-401**,
|
||||
**R-412 leg 2**, **§11-D** (whether the two Hetzner services share an account), **R-427** (twelve open
|
||||
rows carrying a closed verdict), **R-432**.
|
||||
|
||||
## Register
|
||||
|
||||
**Before:** OPEN 181 · CLOSED 161. **After:** OPEN 181 · CLOSED 163.
|
||||
Corrected and closed **R-429**; shipped and closed **R-431**; filed **R-432**; re-scoped **R-95**
|
||||
(stays open, ranking to Viktor). Both closed rows were **moved into `CLOSED-ITEMS.md`** rather than
|
||||
left in the open register with a closed verdict — that pile is R-427 and I did not add to it.
|
||||
|
||||
## Observations, and my own mistakes by name
|
||||
|
||||
1. **A correct instrument aimed at the wrong subject produces a confident wrong answer, and controls
|
||||
cannot save you from it.** My `.snapshots` probe had a good positive and negative control and was
|
||||
still worthless. **FILED: R-429** — corrected there, with the cause named.
|
||||
2. **My mistake — I asserted a mechanism from a task brief without checking the vendor documentation**,
|
||||
and reported "the safety net cannot be seen" to Viktor on that basis.
|
||||
**NOT-A-FINDING: this is item 1 seen from the other side, already recorded in R-429 and as a
|
||||
ruling in `CONTEXT.md`, so a second row would duplicate rather than add anything.**
|
||||
3. **My mistake — my first escalation-only test was HOLLOW and its red-proof passed.** It re-swept the
|
||||
same report, so the baseline had already moved to the new count and the latch was never consulted.
|
||||
Caught because red-proof 3 did **not** fail. Rewritten to drive a continuously falling count, which
|
||||
is the only shape where the latch is load-bearing; the red-proof then failed correctly.
|
||||
**NOT-A-FINDING: caught inside the session by the red-proof discipline doing exactly its job — a
|
||||
test whose red-proof passes is not a test, and that is why they are run.**
|
||||
4. **My mistake — I queried the wrong table for the history.** The brief pointed at `host_reports`;
|
||||
that is the *agent's* host report and carries no `offsite` object at all (8 434 rows, zero hits).
|
||||
The controller's report lives in `reports`. **NOT-A-FINDING: found in one query by checking the
|
||||
parse count instead of trusting the pointer; the measurement in the changelog and the fixture both
|
||||
come from the right table.**
|
||||
5. **My mistake — I posted the live firing to `/event` instead of `/api/v1/event`** and got HTTP 302
|
||||
for *both* the control and the test. **Two identical results are an instrument fault, not two
|
||||
findings** — the same shape as yesterday's `sftp -p`. **NOT-A-FINDING: corrected in one command
|
||||
once I read the router's `TrimPrefix`; no wrong conclusion was recorded.**
|
||||
6. **A sub-account cannot see inside the snapshot tree, so recovery is operator-only today.**
|
||||
**FILED: R-432** — and it decides whether R-95's remedy can ever be product-driven.
|
||||
@@ -18,7 +18,7 @@ not an evening's work.**
|
||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||
nothing.*
|
||||
|
||||
1. **One thing is waiting on you — item 4, and it takes ten minutes.** Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
1. **Two things are waiting on you — item 4 (a snapshot name, two minutes) and item 5 (ignore one alarm mail).** Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
the real machines:
|
||||
- the background job that could delete a live restore's lock now waits its turn — and the check
|
||||
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
|
||||
@@ -37,24 +37,30 @@ nothing.*
|
||||
two register lines in the hub (already live). No customer action, no data migration, no
|
||||
credential change.
|
||||
|
||||
4. **The copy that holds the customers' documents and photos can still be deleted by the box that
|
||||
made it** (R-95 — first on the list since July, and this is the first time it has reached this
|
||||
page). I studied it today and did not change anything. **One thing is yours and it takes ten
|
||||
minutes.** Our notes say a safety net is switched on: the storage provider takes a snapshot every
|
||||
night and keeps seven. **I could not find a single one.** I looked from both machines, using the
|
||||
credentials they already have, and checked that my method could see other things and could
|
||||
correctly fail to see a made-up name. Either the snapshots are not being taken, or they are
|
||||
invisible to the machines — and if they are invisible, they are also useless to them: getting one
|
||||
back would be you, in the provider's control panel. **I did not log in to check, because your own
|
||||
notes say that question is yours.** **If you do nothing:** the register keeps saying the net is
|
||||
armed, and nobody knows whether it is. **Please open the Storage Box panel and tell me whether
|
||||
snapshots exist.** The answer changes which fix is worth building.
|
||||
4. **Good news, and I was wrong yesterday.** I told you I could not find the nightly safety net on
|
||||
the storage. **I was looking for the wrong name.** The snapshots are there, exactly as you saw in
|
||||
the control panel: seven of them, one a day. **And they are better than we thought** — I tried to
|
||||
write into the snapshot area from a customer's machine and the storage **refused**, while the same
|
||||
write to its normal folder worked. So a machine that wipes its own backup **cannot touch the
|
||||
snapshots of it**. The worst case is losing about a day, then copying the rest back file by file.
|
||||
**That is much smaller than what the notes have said since July.** I have corrected the notes.
|
||||
**What I still need from you:** the machines can see the snapshot *door* but not what is inside —
|
||||
only the main account can. So getting data back is you, in a browser, for now. **If you read one
|
||||
snapshot's name off the panel and send it to me, one command settles whether the machines can
|
||||
reach them directly** — and if they can, recovery becomes something the product does by itself.
|
||||
**If you do nothing:** it stays a manual job for you, which is workable but slow.
|
||||
|
||||
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
5. **You will have received an alarm email from me today about `demo-hp` losing 65 backups. It is a
|
||||
test and nothing is wrong.** I built the new "someone deleted the backups" alarm and had to fire it
|
||||
once for real to prove it reaches you. Subject: `[Felhom] 🔴 demo-hp: offsite_snapshots_dropped`.
|
||||
**No backups were deleted.** **If you do nothing:** nothing — but please do not act on that one
|
||||
mail.
|
||||
|
||||
6. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||
|
||||
6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
7. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||
as it misled one by an hour.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user