diff --git a/CONTEXT.md b/CONTEXT.md index 0935f633..8f2466a7 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,42 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## The record told a worse story than the truth for two months, and my own probe is why (2026-09-01, R-429 / R-95 / R-431) + +**THREE RULINGS.** + +**1. A correct instrument pointed at the wrong subject produces a confident wrong answer, and the +controls cannot save you.** The 2026-08-31/09-01 spike probed for a directory called `.snapshots` and +reported *"the safety net cannot be seen from the box"*. Its positive and negative controls were +sound. **The vendor documents the path as `/.zfs/snapshot`.** So `not found` was true and meant +nothing, and the register carried "the exposure is unbounded" on the strength of it. Re-probed at the +documented path: `/.zfs` lists from inside the jail, and **a write into `/.zfs/snapshot` is REFUSED** +(`dest open …: Failure`) while the identical write to the account home succeeds. **Controls prove the +instrument works; they say nothing about whether it is aimed at the question.** Cite the vendor before +asserting a mechanism — the task that set the spike asserted it without a citation, and I did not check. + +**2. The real R-95 is much smaller than two months of register text.** The box can delete its own +**live** repository, but it **cannot write to the daily snapshots of it** — measured, both boxes. So a +deletion costs at most the data written since the last daily snapshot, and the rest comes back **file +by file**, one customer at a time, with no effect on anyone else. **Not open-ended loss.** Two limits +kept honest: a sub-account sees the snapshot directory EMPTY, so per-file recovery is **operator-only +today** (R-432), and a panel-driven restore rolls back the whole Storage Box. **R-95's ranking is +Viktor's** — it has been #1 since July on the old story. + +**3. Detection belongs on the HUB, not the box, and the threshold is reasoned rather than invented.** +The event being detected is a box deleting its own backups, so a detector living on that box is one +the same event can silence. The hub already receives the count and keeps the history. +**The measurement is the interesting part:** across 12 898 reports (2026-06-05 → 2026-09-01) every one +of the nine decreases lands exactly on **zero**, and every one predates `stats_known` — they are the +R-331 shape, a zero meaning *unmeasured*, several carrying a declared `State` that says so outright. +In the 380-report window where `stats_known` is true there are **zero** decreases. **So observed churn +gave nothing to calibrate against, and saying so beats inventing a number (R-401).** The threshold +therefore comes from what retention *can* do: keeping 7 daily + 4 weekly + 6 monthly per group means +it **cannot halve a total**, while a mass deletion goes to ~0 — hence *more than half, and at least +five*. Three pre-conditions guard it, each with a scar: `StatsKnown` (R-331), the declared `State` +(R-204), and run success (R-100, whose lesson lives in that very file). **Acceptance: 9 009 real +report points replayed → zero alarms.** + ## The checks were the one thing nothing checked — 16 of 29 could be fooled by a label (2026-09-01, R-421) **THE CLASS: an instrument that matches a LABEL rather than the fact it names.** Five instances, and diff --git a/REPORT-r431-snapshots.md b/REPORT-r431-snapshots.md new file mode 100644 index 00000000..819025ac --- /dev/null +++ b/REPORT-r431-snapshots.md @@ -0,0 +1,209 @@ +# REPORT — R-429 corrected, R-95 re-scoped, R-431 shipped (2026-09-01) + +## Part 1's answer, first, because everything reads differently after it + +**Can a customer's own account reach the snapshot tree? NO — it can reach the DOOR, and is refused +writes to it, but the tree lists EMPTY.** + +Measured on **both** live boxes, over the credential each already holds, with controls in the same run. + +``` +=== POSITIVE CONTROL: the account home (must list) === + drwxr-xr-x 1058 4 Aug 22 03:10 ./. + drwx------ 1058 3 Jul 23 09:53 ./.ssh + drwxrwxr-x 1058 8 Aug 4 12:38 ./ +=== NEGATIVE CONTROL: a name that cannot exist (must fail) === + Can't ls: "/home/./zzz-no-such-r429" not found +=== CANDIDATE A: /.zfs/snapshot === + drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/. + drwxrwxrwx root root 0 Jul 21 16:01 /.zfs/snapshot/.. +=== CANDIDATE B: /home/.zfs/snapshot === + Can't ls: "/home/.zfs/snapshot" not found +=== CANDIDATE D: is the .zfs door itself visible? === + drwxrwxrwx root root 2 Sep 1 12:12 /.zfs/shares + drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot +``` + +**The write-refusal test — the load-bearing sentence of the whole re-scope, now proven not cited:** + +``` +--- control: the same write in the account home MUST succeed + sftp> put … ./r429-write-control.txt + Uploading … to /home/./r429-write-control.txt + -rw-r--r-- 1058 11 Sep 1 12:12 ./r429-write-control.txt +--- cleanup of the control file + Removing /home/./r429-write-control.txt + Can't ls: "/home/./r429-write-control.txt" not found +--- THE TEST: write into /.zfs/snapshot (must be REFUSED) + Uploading … to /.zfs/snapshot/r429-write-attempt.txt + dest open "/.zfs/snapshot/r429-write-attempt.txt": Failure +--- confirm nothing was left behind: + drwxrwxrwx root root 2 Jan 1 1970 /.zfs/snapshot/. +``` + +**Identical on demo-felhom** (gid 1019). `storage-box-pool-1` **is** `u629488` +(`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`) — the same box that demonstrably holds seven +snapshots — so the emptiness is **per-sub-account filtering**, not absence. → **R-432**. + +## What changed in the record + +| where | from | to | +|---|---|---| +| **R-429** | *"the mitigation has never been confirmed"* | **CLOSED — confirmed working.** My probe used `.snapshots`; the vendor documents `/.zfs/snapshot`. The controls were sound, the subject was wrong. What remains true is the actual finding: the row had no id, its "confirm tomorrow" went 36 days unanswered, and `DUE-CHECKS` was empty. **The finding was never the snapshots — it was that nobody could tell.** | +| **R-95** | *"can delete, exposure open-ended"* — #1 since July | **RE-SCOPED:** deletes the live repo, **cannot write to the snapshots of it**; costs ≤1 day plus per-file recovery. **Ranking left to Viktor.** | +| **07 §8 row 10** | *"the restic repo is NOT protected the same way"* | **both halves stated:** (a) the live repository is deletable — R-95 stands; (b) the snapshots are not writable by anything — proven. **Status NOT moved** — the recovery route has never been walked, which is what PARTIAL means. | +| **07 §10.2** | R-95's line | gains the re-scope + citation | +| **§11-D** | — | **untouched.** Nothing here answers whether the two Hetzner services share an account. | + +## Part 3 — what normal looks like, in numbers + +Hub's own `reports` table: **12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.** + +- **Nine decreases in the entire history, and every one lands exactly on ZERO** — 36→0, 18→0 ×2, + 15→0, 12→0, 8→0, 3→0. **Not one gradual retention decrease anywhere.** +- **Every one predates `stats_known`** — the R-331 shape, a zero meaning *unmeasured*. Several carry a + declared `State` (`needs_credential`, `awaiting_recovery_key`) saying so outright. +- **In the `stats_known`-true window (380 reports) there are ZERO decreases**: demo-felhom flat at 10; + demo-hp 67→68→69, rises only. + +**So observed churn gave nothing to calibrate against, and I say so rather than inventing a number.** + +## The threshold, and where it came from + +**A fall of more than HALF the previous count, AND at least 5.** Reasoned from what retention *can* +do, since it was never seen to do anything: `--keep-daily 7 --keep-weekly 4 --keep-monthly 6 +--group-by host,tags` over ~8 apps **cannot halve a total** — those floors are per group — while a +mass deletion goes to ~0. The floor of 5 stops a small-count box twitching. Deliberately not +sensitive: a detector that cries wolf is switched off within a fortnight. + +## Files, commits, deployment + +| file | what | +|---|---| +| `hub/internal/monitor/offsite.go` | +170 lines: third signal, three guards, escalation latch | +| `hub/internal/monitor/offsite_r431_test.go` | NEW — 5 tests | +| `hub/internal/monitor/testdata/r431_real_history.json` | NEW — 9 009 real points, committed | +| `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go` | allowlist + operator-only, same commit | +| `hub/CHANGELOG.md`, `CONTEXT.md`, `STATUS.md`, `07`, `00-capability-map.md`, both registers | the record | + +Commits: **`3068176`** (code + record), **`65c82c4`** (manifest bump). +**Deployed hub `v0.111.0`.** ArgoCD: `sync=Synced health=Healthy`, +`rev=65c82c4aa0949510670c42eb5c25515c25e12040` — **exactly HEAD**; `deployment "hub" successfully +rolled out`; live image `gitea.dooplex.hu/admin/felhom-hub:0.111.0`; startup line +`Offsite checker initialized: … 3 ok-seeded`. Image presence verified in the registry before syncing. + +## Tests and red-proofs + +| test | result | +|---|---| +| `TestR431_FiresOnAMassDeletion` | 69→4 fires exactly once, severity in the vocabulary, message carries both numbers and does not claim loss | +| `TestR431_SilentWhenNotTrustworthy` | silent on all four shapes (no `stats_known`, declared state, failed run, incomplete run) **and the baseline is left untouched** | +| `TestR431_OrdinaryRetentionIsSilent` | 69→60 silent · 10→7 silent · 10→5 silent (exactly half is not *more than* half) · 69→34 fires | +| `TestR431_EscalationOnlyLatch` | a continuing deletion pages once; a clean sweep re-arms; a later deletion fires again | +| **`TestR431_RealHistoryProducesZeroAlarms`** | **9 009 real points, 2 customers → ZERO alarms.** The acceptance step. | + +Full hub suite green (`go build ./... && go vet ./... && go test ./...`). + +**Red-proofs, by name:** + +1. **Threshold** — `snapshotDropFraction` 0.5 → 0.99: `TestR431_FiresOnAMassDeletion` **FAILS** + (`want exactly 1 …, got 0`). +2. **`StatsKnown`** — guard removed: `TestR431_SilentWhenNotTrustworthy` **FAILS** + (`stats_known absent …: must NOT alarm; got 1`) **and the real-history replay FAILS too**. +3. **Escalation-only** — latch removed: `TestR431_EscalationOnlyLatch` **FAILS** + (`a CONTINUING deletion must not re-page …; got 2 alarms`). + +## The live firing + +**Negative control first**, because an allowlist that accepts everything proves nothing: + +``` +POST /api/v1/event event_type=zzz_not_allowlisted_r431 → HTTP 400 "Invalid event_type: zzz_not_allowlisted_r431" +POST /api/v1/event event_type=offsite_snapshots_dropped → HTTP 200 {"ok":true} +``` + +Hub log: `[INFO] Event from demo-hp: offsite_snapshots_dropped (error) — …` then +`[INFO] Operator email sent for demo-hp/offsite_snapshots_dropped`. + +**Routing, from the live DB — this is the operator-only claim, with a control:** + +``` +788|customer|skipped|operator_only +787|operator|sent| +``` + +Other types do reach customers (`escrow_blob_served|customer|41`, `whole_guest_backup_failed|customer|28`), +so "operator-only" is a real distinction here rather than everything being operator. + +**The mail, rendered from the hub's own `FormatOperatorEmail`:** + +``` +SUBJECT: [Felhom] 🔴 demo-hp: offsite_snapshots_dropped +Customer: demo-hp +Event: offsite_snapshots_dropped +Severity: error +Time: 2026-09-01 14:30 CEST +Message: Customer demo-hp: off-site backup count fell from 69 to 4 snapshot(s) in one report - more + than retention can explain. The daily Storage Box snapshots are read-only and still hold the + older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check whether + a deletion ran on the box before restoring anything. +Details: {"previous_count":69,"current_count":4,"drop":65} +Dashboard: https://hub.felhom.eu/customers/demo-hp +``` + +**⚠ THAT MAIL IS REAL AND VIKTOR RECEIVED IT. Nothing was deleted** — it is the required live firing, +flagged as item 5 in `STATUS.md` so it is not acted on. + +## Not validated + +- **PROVEN-LIVE for the capability row.** The firing proves *delivery*, not that a genuine deletion is + caught on a live box. The capability row says **IMPLEMENTED**, with that gap written into it. +- **Whether a NAMED snapshot can be entered** even though the directory does not list (ZFS permits + exactly that). One panel read from Viktor settles it — R-432. +- **Whether a snapshot can be restored from.** Deliberately not attempted, in this session or the last. + +## No controller release, no golden + +**No controller change. No version bump there. No image. No golden owed.** +`golden_currency_gate.py` exits 0; golden and fleet floor remain **0.232.0**. + +## Still open, named + +**R-430** (`unlock` reports success while the lock survives — **latent**: it bites only when the box +loses delete, which is not happening here), **R-242's vouch half**, **R-402**, **R-409**, **R-401**, +**R-412 leg 2**, **§11-D** (whether the two Hetzner services share an account), **R-427** (twelve open +rows carrying a closed verdict), **R-432**. + +## Register + +**Before:** OPEN 181 · CLOSED 161. **After:** OPEN 181 · CLOSED 163. +Corrected and closed **R-429**; shipped and closed **R-431**; filed **R-432**; re-scoped **R-95** +(stays open, ranking to Viktor). Both closed rows were **moved into `CLOSED-ITEMS.md`** rather than +left in the open register with a closed verdict — that pile is R-427 and I did not add to it. + +## Observations, and my own mistakes by name + +1. **A correct instrument aimed at the wrong subject produces a confident wrong answer, and controls + cannot save you from it.** My `.snapshots` probe had a good positive and negative control and was + still worthless. **FILED: R-429** — corrected there, with the cause named. +2. **My mistake — I asserted a mechanism from a task brief without checking the vendor documentation**, + and reported "the safety net cannot be seen" to Viktor on that basis. + **NOT-A-FINDING: this is item 1 seen from the other side, already recorded in R-429 and as a + ruling in `CONTEXT.md`, so a second row would duplicate rather than add anything.** +3. **My mistake — my first escalation-only test was HOLLOW and its red-proof passed.** It re-swept the + same report, so the baseline had already moved to the new count and the latch was never consulted. + Caught because red-proof 3 did **not** fail. Rewritten to drive a continuously falling count, which + is the only shape where the latch is load-bearing; the red-proof then failed correctly. + **NOT-A-FINDING: caught inside the session by the red-proof discipline doing exactly its job — a + test whose red-proof passes is not a test, and that is why they are run.** +4. **My mistake — I queried the wrong table for the history.** The brief pointed at `host_reports`; + that is the *agent's* host report and carries no `offsite` object at all (8 434 rows, zero hits). + The controller's report lives in `reports`. **NOT-A-FINDING: found in one query by checking the + parse count instead of trusting the pointer; the measurement in the changelog and the fixture both + come from the right table.** +5. **My mistake — I posted the live firing to `/event` instead of `/api/v1/event`** and got HTTP 302 + for *both* the control and the test. **Two identical results are an instrument fault, not two + findings** — the same shape as yesterday's `sftp -p`. **NOT-A-FINDING: corrected in one command + once I read the router's `TrimPrefix`; no wrong conclusion was recorded.** +6. **A sub-account cannot see inside the snapshot tree, so recovery is operator-only today.** + **FILED: R-432** — and it decides whether R-95's remedy can ever be product-driven. diff --git a/STATUS.md b/STATUS.md index 7b0ddcb4..cf86a82f 100644 --- a/STATUS.md +++ b/STATUS.md @@ -18,7 +18,7 @@ not an evening's work.** *This section is allowed to be longer than one screen, and each item says what happens if you do nothing.* -1. **One thing is waiting on you — item 4, and it takes ten minutes.** Otherwise: Both problems the overnight test found are fixed and proven on +1. **Two things are waiting on you — item 4 (a snapshot name, two minutes) and item 5 (ignore one alarm mail).** Otherwise: Both problems the overnight test found are fixed and proven on the real machines: - the background job that could delete a live restore's lock now waits its turn — and the check that finds the next one like it is a test, not a comment, so it cannot come back quietly; @@ -37,24 +37,30 @@ nothing.* two register lines in the hub (already live). No customer action, no data migration, no credential change. -4. **The copy that holds the customers' documents and photos can still be deleted by the box that - made it** (R-95 — first on the list since July, and this is the first time it has reached this - page). I studied it today and did not change anything. **One thing is yours and it takes ten - minutes.** Our notes say a safety net is switched on: the storage provider takes a snapshot every - night and keeps seven. **I could not find a single one.** I looked from both machines, using the - credentials they already have, and checked that my method could see other things and could - correctly fail to see a made-up name. Either the snapshots are not being taken, or they are - invisible to the machines — and if they are invisible, they are also useless to them: getting one - back would be you, in the provider's control panel. **I did not log in to check, because your own - notes say that question is yours.** **If you do nothing:** the register keeps saying the net is - armed, and nobody knows whether it is. **Please open the Storage Box panel and tell me whether - snapshots exist.** The answer changes which fix is worth building. +4. **Good news, and I was wrong yesterday.** I told you I could not find the nightly safety net on + the storage. **I was looking for the wrong name.** The snapshots are there, exactly as you saw in + the control panel: seven of them, one a day. **And they are better than we thought** — I tried to + write into the snapshot area from a customer's machine and the storage **refused**, while the same + write to its normal folder worked. So a machine that wipes its own backup **cannot touch the + snapshots of it**. The worst case is losing about a day, then copying the rest back file by file. + **That is much smaller than what the notes have said since July.** I have corrected the notes. + **What I still need from you:** the machines can see the snapshot *door* but not what is inside — + only the main account can. So getting data back is you, in a browser, for now. **If you read one + snapshot's name off the panel and send it to me, one command settles whether the machines can + reach them directly** — and if they can, recovery becomes something the product does by itself. + **If you do nothing:** it stays a manual job for you, which is workable but slow. -5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August. +5. **You will have received an alarm email from me today about `demo-hp` losing 65 backups. It is a + test and nothing is wrong.** I built the new "someone deleted the backups" alarm and had to fire it + once for real to prove it reaches you. Subject: `[Felhom] 🔴 demo-hp: offsite_snapshots_dropped`. + **No backups were deleted.** **If you do nothing:** nothing — but please do not act on that one + mail. + +6. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August. Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as it is, at the risk you accept by leaving it. I can change it without ever showing you the new one. -6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is +7. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, as it misled one by an hour.