From db38f4c800fecfdafbb2b2f4aa6b9751426bfe59 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 1 Sep 2026 18:34:52 +0200 Subject: [PATCH] hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point. The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a new one, and the sentence is true under every possible answer to the provider questions, so it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not. was: "...still hold the older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check whether a deletion ran on the box before restoring." now: "...still hold the older copy. The route back out of them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a deletion ran on the box." It must not swing the other way either: "your backups are gone" is still usually false. Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed. Tests: offsite_r434_test.go, three, all driving the production path so they assert the sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls. RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated, so the wording keeps ONE home. R-435 written into the detector's own documentation, no threshold changed: it sees a mass deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and forget --prune groups by host,tags). Says explicitly not to lower the numbers. THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md. Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12, each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place: six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104) that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today. Two provider questions drafted, not sent, no API called (11-D stands): documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed for 36 days is the scar that block exists for. R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed", "zero snapshots") is corrected in place, order unchanged. Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169. No controller or agent change. No golden owed, no floor change. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM --- REPORT-drill-r95-recovery-2026-09-01.md | 166 ++++++++++++++++++ STATUS.md | 64 +++++-- .../architecture/07-backup-architecture.md | 32 +++- documentation/backlog/OPEN-ITEMS.md | 83 ++++++++- .../runbooks/provider-questions-2026-09-01.md | 101 +++++++++++ hub/CHANGELOG.md | 67 +++++++ hub/internal/monitor/offsite.go | 30 +++- hub/internal/monitor/offsite_r431_test.go | 10 +- hub/internal/monitor/offsite_r434_test.go | 155 ++++++++++++++++ 9 files changed, 673 insertions(+), 35 deletions(-) create mode 100644 REPORT-drill-r95-recovery-2026-09-01.md create mode 100644 documentation/runbooks/provider-questions-2026-09-01.md create mode 100644 hub/internal/monitor/offsite_r434_test.go diff --git a/REPORT-drill-r95-recovery-2026-09-01.md b/REPORT-drill-r95-recovery-2026-09-01.md new file mode 100644 index 00000000..a9755c33 --- /dev/null +++ b/REPORT-drill-r95-recovery-2026-09-01.md @@ -0,0 +1,166 @@ +# REPORT — the deletion we said is survivable: the recovery route does not exist (R-95 drill, 2026-09-01) + +**RUNBOOK, destructive class, `demo-hp` only. STOPPED at the end of Phase 1 on the operator's ruling, +before any destructive step. No delete verb was issued against any live store; no byte on either +Storage Box sub-account was written, moved or removed.** No production code, no version bump, no +image, no golden. Evidence: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`. + +| # | phase | verdict | one sentence | +|---|---|---|---| +| 1 | snapshot reachable, and its name | **NO — and it has no reachable name** | 777,600 exact names across nine days in the vendor's own format, plus 126 alternative shapes; zero resolve, with a control proving the sweep detects a path that exists. | +| 2 | the deletion | **NOT RUN — operator ruling** | With no recovery route, the deletion would have destroyed real history to buy nothing; put as a two-option decision, the ruling was stop. | +| 3 | the alarm fired | **NOT RUN — and it could not have fired at the specified size** | The shipped threshold needs a fall of more than half; one app's tag is ~9 of 69. → **R-435** | +| 4 | **the recovery** | **NOT RUN — no route exists that is not fenced** | Box-side: proven impossible. Panel: fenced and browserless. Hetzner API: fenced (§11-D), and the hub's client has no snapshot method at all. | +| — | **RTO from T₀** | **STILL BLANK** | Row 10's RTO cell is unchanged and remains a finding. | +| — | **data lost, quantified** | **NOT MEASURABLE THIS WAY** | The quantity only has meaning if the rest is recoverable, and the route that would recover it is not reachable. | +| 5 | re-arm | **NOT RUN** | Depended on Phase 4. | +| 6 | teardown | **PASS** | Store untouched at 69 snapshots; all three scratch layers removed; hub DB copies shredded; both boxes healthy. | + +--- + +## 1. Did the recovery work — and does yesterday's re-scope survive? + +**The recovery was never reachable, and the re-scope does not survive intact. Its first half stands; +its second half does not.** + +Yesterday's re-scope has two clauses. They must now be separated: + +* **(a) "The box can delete its live repository, but cannot write to the daily snapshots of it."** + **STANDS.** Re-confirmed here: `/.zfs/snapshot` is reachable and the write-refusal measurement is + unchanged. Nothing in this drill weakens it. +* **(b) "…so the rest is recoverable — file by file, one customer at a time."** **NOT SUPPORTED.** + A snapshot that cannot be opened cannot be copied out of. R-432 recorded the directory listing + empty and named the cheapest next step: *"a single `ls /.zfs/snapshot/` from a box then + settles whether a named snapshot can be entered even though the directory does not list (ZFS + allows exactly that)."* **That step is now done, exhaustively, and the answer is no.** + +**What was measured.** The port-23 restricted shell accepts a batched `stat`, which makes a cheap +existence oracle: 500–600 paths per round trip, stdout carrying only paths that exist. + +| sweep | candidates | hits | +|---|---|---| +| `/.zfs/snapshot/YYYY-MM-DDTHH-MM-SS`, nine full days, second granularity | **777,600** | **0** | +| 126 alternative name shapes and snapshot paths (`daily`, `snapshot-1`, colon and compact time forms, `/home/.snapshot`, …) | 126 | 0 | +| **control — the identical 600-name batch shape with one real path appended** | 6 batches | **6/6 returned it** | + +**And there is a structural reason, which is why I stopped sweeping.** The customer's data and the +snapshot door are on **different filesystems**: + +``` +df → u629488-sub3 mounted on /home +stat /home → Device 0,82 +stat /.zfs/snapshot → Device 0,276 ← a different device +stat /home/.zfs → cannot statx: No such file or directory +``` + +A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset that owns that `.zfs` — not to the +child mounted at `/home`. **So even a correctly named snapshot there could not contain +`felhom-repo`,** and the dataset that does hold it exposes no `.zfs` at all to this account. The +empty listing is not a display toggle hiding a reachable tree; from a sub-account there is no tree. + +**Three tools agree, each with controls in the same run:** SFTP, the port-23 shell, and +`rsync --list-only`. + +**What that does to R-95.** Its *exposure* is unchanged and its *remedy* is not. Yesterday the row +could say a deletion costs about a day because the rest comes back per-file. Today the only routes +to "the rest" are a whole-box panel rollback (which deletes newer snapshots and hits every customer +on the box) and the provider API (fenced, and unimplemented in the hub's client). **The re-scope's +comfort was resting on a route nobody had walked — which is precisely the standard this project +applies, and it is the reason this drill was called.** + +**The ranking is Viktor's and I am not re-ranking it.** What I will say plainly: the argument that +moved R-95 down yesterday is the argument this drill removed. On these facts I would put it back +where it was. + +## 2. The RTO + +**Still blank, and it stays a finding.** `07` §8 row 10's RTO cell has been empty since July and this +drill did not fill it. Nothing was recovered, so nothing was timed. The summary line at +`07-backup-architecture.md:948` — *"no ransomware-shaped recovery has ever been run"* — is still +true, and is now true for a sharper reason: **not "nobody has run it" but "from the box, it cannot +be run."** + +## 3. R-432's answer, and the naming scheme + +**R-432 is ANSWERED, and negatively. It did not need the panel read it was waiting on.** + +* **The naming scheme is `YYYY-MM-DDTHH-MM-SS`** — vendor-documented examples `2025-12-03T13-47-47`, + `2025-02-12T11-35-19`. Recorded so nobody hunts a console again. +* **Knowing it does not help.** Every name in that format for nine days is refused, and the st_dev + split above says why. **Per-file recovery is not operator-only — from the box it is nobody's,** and + for the operator it is a browser act against the main account that no credential in this project + can perform. +* **The panel cannot supply the missing piece either.** It offers Restore and Delete on a row and + does not show names; and the one name-shaped thing it could give would be tried against a door + that leads to the wrong dataset. + +## 4. The alarm's first real firing + +**It did not happen, and the drill as written could not have produced it.** The detector fires on a +fall of **more than half** the previous count **and at least 5** (`hub/internal/monitor/offsite.go`, +`snapshotDropFraction = 0.5`, `snapshotDropFloor = 5`). demo-hp's baseline is **69**. Phase 2 deletes +**one app's** history — about **9** snapshots. 9 is over the floor and nowhere near half, so the +alarm stays silent, **correctly and by design**. Firing it for real needs ~35+ snapshots destroyed, +i.e. most of demo-hp's off-site history. That trade is what the operator was asked to rule on. → **R-435** + +**One thing the alarm says is now wrong.** Its message, live in hub 0.111.0, reads: + +> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is +> recoverable file-by-file**; it is NOT confirmed data loss." + +The first clause is true; **the second promises a recovery the product cannot perform and the +operator cannot perform without a browser and the main account.** This is this project's own +corollary — *when a verdict changes which field it counts from, the alarm text has to change with +it* — landing on the alarm shipped the same day. → **R-434** + +## 5. Findings, as register rows + +All four filed in `documentation/backlog/OPEN-ITEMS.md`. + +| row | finding | +|---|---| +| **R-433** | A sub-account cannot reach any Storage Box snapshot **by any name**; `/home` and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Answers R-432 negatively and removes clause (b) of the R-95 re-scope. | +| **R-434** | `emitSnapshotDrop`'s message promises file-by-file recovery that is not reachable. Live in hub 0.111.0. | +| **R-435** | The drop detector cannot see a single-app deletion (>50% of 69 ⇒ ~35 needed). `offbox.go:1388` forgets **by tag**, so a single-tag wipe is exactly the shape the detector is blind to. Deliberate insensitivity, but the blind spot should be stated where the operator reads it. | +| **R-436** | **LEAD, not a defect.** Hetzner's port-23 shell offers `rclone serve restic --stdio` as a server-side backend, and restic 0.14.0 recognises the `rclone:` backend (measured; control `banana:` → invalid backend; rclone is absent from the controller image). `rclone serve restic` carries `--append-only`. **This could make R-95's real prevention far cheaper than the spike concluded — no new always-on machine, no data migration.** Caveat stated up front: the **client** supplies the server command line, so a compromised box could omit the flag unless the provider pins it. Settling that is a vendor question, not a code change. | + +**R-432 is marked ANSWERED**; its "one panel read settles it" next step is withdrawn as unnecessary. + +## 6. Does `07` §8 row 10 move? + +**No. It stays `PARTIAL`, and its RTO stays blank.** The status was already correct for the right +reason — *"the recovery ROUTE has never been walked, which is what PARTIAL means"* — and this drill +found the route is not walkable from the box at all. **What the row needs is a text correction, not a +status change:** its clause *"recoverable per-file (vendor)"* and its limit *"per-file recovery is +operator-only today (R-432)"* both overstate what exists. Updated in place with the citation. Moving +it only as far as the evidence goes means not moving it. + +## 7. What could not be tested, and why + +* **Whether the main account can see the snapshots.** No main-account credential exists in this + project — the hub holds only per-customer sub-accounts. This is the one question that would decide + whether per-file recovery exists *at all*, for anyone. +* **Whether the Hetzner API can list or read a snapshot.** Fenced by the runbook (§11-D). Separately, + `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method** — so this route needs new code + regardless of the fence. +* **The deletion, the alarm's first real firing, the recovery, the RTO, the re-arm.** Phases 2–5, not + run, on the operator's ruling. +* **Whether `rclone serve restic --stdio` is pinned server-side with `--append-only`** (R-436). + +## 8. My own mistakes + +* **I ran Phase 1 before finishing Phase 0's subject-app choice, and did not say so up front.** The + choice depends on the snapshot's timestamp, so the order was right, but the runbook's order is the + runbook's and a silent reordering is the thing this project keeps getting caught by. Stated at the + time in the session, recorded here. +* **My first sweep guessed the schedule instead of establishing it.** I probed 00:00 UTC and 22:00 + UTC — 600 names — on the strength of a register line reading *"daily 00:00"*, got nothing, and only + then widened to whole days. The narrow sweep was worth nothing on its own: a zero over a guessed + window is not evidence, and I should have gone to full days first or not run it at all. +* **I nearly reported the empty listing as "the display toggle is hiding it".** The vendor documents + exactly such a toggle and it fitted. The st_dev comparison — which I only ran because `df` printed + a filesystem name I did not expect — says the tree is on another dataset entirely. **A plausible + cause that fits the symptom is not a measured one**, and I had the wrong one for about ten minutes. +* **`REPORT.md` held the only copy of the R-331 report** (hub v0.109.0, 2026-08-30) — durable content + living only in the overwritten file, which `CLAUDE.md:82-87` forbids. Preserved as + `REPORT-r331-backup-card.md` before this report replaced it. diff --git a/STATUS.md b/STATUS.md index cf86a82f..5b4c509f 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,8 +1,8 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-01 (second pass) — both faults from the overnight test are fixed, and the golden -is baked, vouched and delivered. Both machines are on 0.232.0 and moved themselves. NOTHING is -waiting on you.** +**Updated 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down +where it stops. One alarm that was telling you something untrue is fixed and live (hub 0.111.1). +ONE THING is waiting on you: two short e-mails to Hetzner, drafted and ready to paste.** **Earlier 2026-08-31 — I measured whether the box could test its own remote restore without you, then built the narrow version you picked. The measurement is why it is 3 seconds a night and @@ -18,7 +18,7 @@ not an evening's work.** *This section is allowed to be longer than one screen, and each item says what happens if you do nothing.* -1. **Two things are waiting on you — item 4 (a snapshot name, two minutes) and item 5 (ignore one alarm mail).** Otherwise: Both problems the overnight test found are fixed and proven on +1. **One thing is waiting on you — item 4 (send two e-mails).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on the real machines: - the background job that could delete a live restore's lock now waits its turn — and the check that finds the next one like it is a test, not a comment, so it cannot come back quietly; @@ -37,24 +37,34 @@ nothing.* two register lines in the hub (already live). No customer action, no data migration, no credential change. -4. **Good news, and I was wrong yesterday.** I told you I could not find the nightly safety net on - the storage. **I was looking for the wrong name.** The snapshots are there, exactly as you saw in - the control panel: seven of them, one a day. **And they are better than we thought** — I tried to - write into the snapshot area from a customer's machine and the storage **refused**, while the same - write to its normal folder worked. So a machine that wipes its own backup **cannot touch the - snapshots of it**. The worst case is losing about a day, then copying the rest back file by file. - **That is much smaller than what the notes have said since July.** I have corrected the notes. - **What I still need from you:** the machines can see the snapshot *door* but not what is inside — - only the main account can. So getting data back is you, in a browser, for now. **If you read one - snapshot's name off the panel and send it to me, one command settles whether the machines can - reach them directly** — and if they can, recovery becomes something the product does by itself. - **If you do nothing:** it stays a manual job for you, which is workable but slow. +4. **Please send two short e-mails to Hetzner. They are written for you.** + `felhom.eu/documentation/runbooks/provider-questions-2026-09-01.md` — open it, copy, send. No + password or key is in that file, and none should be added. + + **Why.** Yesterday I told you the snapshots make a wiped backup survivable: lose about a day, copy + the rest back file by file. **The first half is still true. The second half is not, and I found + that out by trying it.** I tried **777,600** snapshot names on the storage, over nine days, in + Hetzner's own naming style. **None of them opened.** Then I found why: your data and the snapshot + door sit on **two different drives** inside the storage, and the door for your data **does not + exist at all**. So there is no way in from the machines. + + **What is still true, and it matters:** a machine that wipes its own backup **still cannot touch + the snapshots of it**. The older copy is there. What we do not have is a way to reach it. + + **The two questions.** One: can the **main** account pull single files out of a snapshot? Two: on + one of Hetzner's own tools, is a "cannot delete" switch forced by them, or chosen by the machine? + **The second one could remove the whole problem** — no new hardware, no moving anyone's data. + + **If you do nothing:** we cannot finish this. The backups keep working and keep being checked; we + simply cannot say what a wiped backup costs, and I would then put this risk back near the top of + your list. **My pick: send both. It is five minutes and it decides an evening's work.** 5. **You will have received an alarm email from me today about `demo-hp` losing 65 backups. It is a test and nothing is wrong.** I built the new "someone deleted the backups" alarm and had to fire it once for real to prove it reaches you. Subject: `[Felhom] 🔴 demo-hp: offsite_snapshots_dropped`. **No backups were deleted.** **If you do nothing:** nothing — but please do not act on that one - mail. + mail. **Closed now:** that was the only such mail, it was a test, and the alarm's wording has since + been corrected (item under *Decided* below). Nothing further is needed from you here. 6. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August. Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as @@ -66,6 +76,26 @@ nothing.* ## Decided — and what would reopen each +- **THE BACKUP AND RESTORE WORK IS FINISHED FOR BETA. DECIDED 2026-09-01.** + **What is done, and proven on the real machines:** everything **you or a customer** does alone — + getting deleted files back, getting an app's data back, getting a whole app back, and losing a + drive. The restore tells you what it put back, refuses if there is no room, will not accept a + half-copy, and puts your own data back if it fails. + **What is parked until after beta:** everything **only I do, with you** — rebuilding a machine as + itself, losing a whole box, recovering from ransomware, restoring the hub, and losing Hetzner. + **Six of these have never been timed, and the hub has never been restored.** They are written down, + they are real, and **none of them stops a beta customer.** + **Reopens if:** something a customer does for themselves turns out to be broken; **or** Hetzner's + two answers change what the snapshots are worth; **or** a real customer's data is at stake in one + of the parked items. + +- **The alarm that promised too much: FIXED and live (hub 0.111.1). DECIDED 2026-09-01.** + Yesterday's alarm mail said a deleted backup was *"recoverable file-by-file"*. **We now know it is + not.** I removed the promise rather than writing a new one, so the sentence stays true whatever + Hetzner answers. It no longer says the data is lost either — that is still usually untrue. + **Reopens if:** Hetzner's answers give us a real route back; then the alarm can name it. + + - **The missing-golden warning is now pointed at the people who can act on it (R-404). DECIDED 2026-09-01, and built the same day.** The warning was aimed at the wrong repository: the one where a release actually happens never checked at all, while the one that only holds documents was diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 727b003a..7e832174 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -859,6 +859,24 @@ crosses the line — **R-158**. ## 8. The failure → recovery matrix +> ### THE ARC IS CLOSED FOR BETA — read this before the blanks below (2026-09-01) +> +> **Closed at controller v0.232.0 / hub v0.111.1.** Everything a **customer** does for themselves is +> finished and proven live — rows **1, 2, 3, 3b, 3c, 6, 7, 14**. Everything only an **operator** does +> is **deliberately deferred until after beta** — rows **4, 8, 9, 10, 11 (+11b), 12**, each carrying +> the marker **`[BETA-DEFERRED]`** in its status cell so the set greps as a group. +> +> **The blanks in the RTO column below are now blank ON PURPOSE, and that is the whole difference.** +> Eleven rows carry a blank RTO — 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15 — and only six of those are +> deferred work. **Row 5** has no recovery to time (a derived copy), **row 13** has no route by +> design, **row 14** is proven and merely never stopwatched, **row 11b** is a note not a row, and +> **row 15 is an open DEFECT (R-104) that this line does NOT cover.** +> +> **NO STATUS MOVED ON THE DAY THIS WAS WRITTEN, because nothing was proven that day.** A stopping +> line that promotes a row is a stopping line that lies. Full text, and the conditions that reopen +> this: `documentation/backlog/OPEN-ITEMS.md`, *"DECIDED — the backup and restore arc is CLOSED FOR +> BETA"*. + **This is the core artifact.** One row per failure. It is authoritative for recovery routes; `00-capability-map.md` stays authoritative for per-capability status. @@ -879,16 +897,16 @@ crosses the line — **R-158**. | 3 | **An app's DB and named volumes are lost** | the guest, the unit | „Visszaállítás indítása" — Tier-1 unit restore (the **only** path that unpacks volume tars) | **customer** | **18.25 s** (path execution) · **27.6 s** (D5 drill, guest `app.yaml` absent) | 24 h | **PROVEN** (content recovery proven 2026-07-30) | CAMPAIGN-9 A2 proved the path executes; **D5 v0.188.0 closed the content gap** — after a restore with the guest's `app.yaml` moved aside, the app read the seeded row **over TCP with its own credential**, the pre-backup row returned and a post-backup row was gone (so the tar was really restored). No `.sql` dump in the unit ⇒ the DB came back from the volume tar. **2026-08-30 — this row KEEPS its PROVEN status and the reason is worth stating: R-353 was a defect in the MESSAGE, not in the mechanism.** The restore really did return what the unit held, every time; what it could not do was say so, because the count was discarded one call deep. Controller v0.226.0 fixed the sentence and changed nothing about the recovery path. A status that measures whether data comes back must not move because a status line was wrong | | 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery | | 3b | *same, for a class-B app via Tier-2* | the Tier-2 copy on the second drive | „Teljes visszaállítás a másolatból" — **Tier-2 unit restore** (`POST /backup/tier2/unit-restore` → `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`), controller **v0.229.0** | **customer** | **28.65 s** (3 volumes, 1 database, 114.5 MB unit) | 24 h | **PROVEN** (2026-08-31) | `audits/DRILL-r102-tier2-unit-2026-08-31/`. docmost — a class-B app whose Tier-2 run reports **0 leg(s)** — restored **with the primary unit moved aside**: 3 volumes of 3 and 1 database of 1, from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`. **The observable is the DATA:** an accented Hungarian filename returned byte-for-byte (verified as hex, R-364) and the app read its own row **over TCP with its own credential**; the post-backup discriminator was **gone**, so the tar was genuinely replayed. Repeated with the guest's `app.yaml` also aside → `secrets recovered=2/2` from the mirrored unit. **R-102 CLOSED** | -| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 of 53 apps) **and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own**; Tier-3 reconstitute for files + DB **+ the named-volume tars since v0.218.0** (`volReplay`) | **customer** (all) | | 24 h | **PARTIAL** | §7.2. **Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0).** This row stays **PARTIAL** deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. **No drive has ever actually died or been replaced under this recovery** — the drill removed a unit directory, not a disk, so drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore | +| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 of 53 apps) **and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own**; Tier-3 reconstitute for files + DB **+ the named-volume tars since v0.218.0** (`volReplay`) | **customer** (all) | | 24 h | **PARTIAL** **`[BETA-DEFERRED]`** | §7.2. **Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0).** This row stays **PARTIAL** deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. **No drive has ever actually died or been replaced under this recovery** — the drill removed a unit directory, not a disk, so drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore | | 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go`. **The derived-copy rule is UNCHANGED by R-403 (controller v0.230.0) and the single exception is stated in §8.2 below — read it before "fixing" a skip you find in the code** | | 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session | | 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | -| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) | -| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) | -| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day — **but only a fall of MORE than half (R-435).** **2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED.** The re-scope said a deletion is *recoverable per-file*. **It is not, from the box: no snapshot is reachable by ANY name.** MEASURED on demo-hp, read-only, no delete verb issued: **777,600** exact names in the vendor form `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes — **zero hits**, against a control where the identical 600-name batch returns a path that exists (6/6). **Structural cause:** `/home` (`u629488-sub3`) is **st_dev 0,82**, `/.zfs/snapshot` is **st_dev 0,276**, and `/home/.zfs` does not exist — a snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding the repository. **Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed.** So this row's *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). **R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed.** `audits/evidence-drill-r95-recovery-2026-09-01/`. | -| 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) | -| 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) | -| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D | +| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** **`[BETA-DEFERRED]`** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) | +| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** **`[BETA-DEFERRED]`** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) | +| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** **`[BETA-DEFERRED]`** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day — **but only a fall of MORE than half (R-435).** **2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED.** The re-scope said a deletion is *recoverable per-file*. **It is not, from the box: no snapshot is reachable by ANY name.** MEASURED on demo-hp, read-only, no delete verb issued: **777,600** exact names in the vendor form `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes — **zero hits**, against a control where the identical 600-name batch returns a path that exists (6/6). **Structural cause:** `/home` (`u629488-sub3`) is **st_dev 0,82**, `/.zfs/snapshot` is **st_dev 0,276**, and `/home/.zfs` does not exist — a snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding the repository. **Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed.** So this row's *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). **R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed.** `audits/evidence-drill-r95-recovery-2026-09-01/`. | +| 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** **`[BETA-DEFERRED]`** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) | +| 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** **`[BETA-DEFERRED]`** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) | +| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** **`[BETA-DEFERRED]`** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D | | 13 | **Customer loses R** | every byte, every tier, the Recipe | **on-premises recovery is unaffected** — Tier-1/2/3 and the whole-guest tiers all work while the guest and host live. What is lost: the ability to re-establish **host identity** and to open the escrowed offsite key material after a host loss | — | | | **NONE for host-loss** | R exists in **zero** system copies by design (INV Part C, row 24). Whether the operator should be able to recover is **open** → §11-A, §11-B. **D5 narrows this further:** Tier-1/2 app recovery needs only the drive, so losing R now costs the offsite route and host identity, never local app recovery (§7.4) | | 14 | **Host SSH management plane dies** | everything | 3-layer break-glass: tmpfiles → 60 s watchdog → hub-vaulted `root@pam` on the PVE console | operator | | | **PROVEN** | `runbooks/break-glass.md`; the original incident and its fix | | 15 | **Interrupted offsite run leaves a stale lock** | everything | manual `restic unlock --remove-all` | **operator** (SSH) | | | **DEFECT** | the self-heal exists (`offbox.go:634-648`) but the probe fails first and `classifyResticProbe` has no lock case → fail-fast, operator told *„ismeretlen okból"* → **R-104** | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 43348c12..0808e501 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -7,6 +7,66 @@ a non-overwritten `REPORT-.md` sibling instead (`CLAUDE.md:82-87`), of wh State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row has an owner. +## DECIDED — the backup and restore arc is CLOSED FOR BETA (2026-09-01) + +**Written in the voice this file uses for a settled decision: what was decided, and the condition +that reopens it.** It exists because nobody ever said the arc was finished, and an arc nobody closed +gets picked up again in a month by someone who reads the blanks in `07` §8 as unfinished work. + +**CLOSED FOR BETA at controller v0.232.0 / hub v0.111.1 (2026-09-01).** + +**What is finished, and proven live:** everything a customer does for themselves — losing files, +losing an app's data, losing a whole app, losing a drive. The restore states what it returned, +refuses without room, cannot be fed a part-copy, and puts the customer's own data back if it fails. +The second drive's copy is a route. A poorer copy cannot delete a richer one. The off-site store is +verified weekly at full depth and proved nightly to still contain something. +**Rows 1, 2, 3, 3b, 3c, 6, 7 and 14 of `07` §8 — every one PROVEN.** + +**What is deliberately deferred until after beta — `[BETA-DEFERRED]`, and these are the row numbers, +not a description:** + +| `07` §8 row | the failure | status today | +|---|---|---| +| **4** | primary drive dies — the drive-loss **journey** (the route is proven; no disk has ever died under it) | `PARTIAL` | +| **8** | host dies, drives intact — a host rebuilt as itself | `IMPLEMENTED`, never executed | +| **9** | whole box lost (fire/theft) | `IMPLEMENTED / UNPROVEN` | +| **10** | ransomware / malicious deletion — a ransomware-shaped recovery | `PARTIAL` | +| **11** (and **11b**) | hub lost — a hub restore | `UNPROVEN`; it has never been performed | +| **12** | off-site provider lost (Hetzner) | `[FACT]` only | + +**These are real, they are recorded, and none of them is a beta blocker.** Every one is invocable by +the operator, not the customer; every one needs hardware, a provider, or a destructive rehearsal that +beta does not. + +**A NUMBER IN THE BRIEF FOR THIS SECTION WAS WRONG AND IS CORRECTED HERE.** It said *"six rows of §8 +still have no measured time"*. **Six rows are DEFERRED; ELEVEN rows carry a blank RTO** — counted, not +estimated: 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15. The other five are blank for reasons that are not +deferred work, and collapsing them into one number is how a blank stops meaning anything: + +* **row 5** — `PROVEN` by construction; the secondary is a derived copy, so there is no recovery to time. +* **row 13** — `NONE for host-loss` **by design**; R exists in zero system copies, so there is no route. +* **row 14** — `PROVEN`; the break-glass route works and has simply never been stopwatched. +* **row 15** — an **open DEFECT** (R-104, the stale-lock path), not a deferred recovery. **It is NOT + inside this stopping line** and must not be read as parked by it. +* **row 11b** — a consequences note attached to row 11, not a recovery row of its own. + +**What is NOT deferred and stays open:** **R-95** and **R-433**, both `BLOCKED-ON-PROVIDER` behind +`documentation/runbooks/provider-questions-2026-09-01.md`; **R-435**, documentation only; and +**R-104**, row 15's defect. The alarm text (**R-434**) was fixed the same day and is closed. + +**THIS REOPENS IF:** a customer-facing recovery path is found broken; **or** Hetzner's answers change +what the snapshots are worth (either answer in Part 3's file can do it — "full restore only" makes +row 10 urgent, "append-only enforced" makes R-95's root cause cheap to remove); **or** a real +customer's data is at stake in one of the deferred rows. + +**THE MARKER.** Every deferred row carries the literal string **`[BETA-DEFERRED]`** in `07` §8, so +`grep -n '\[BETA-DEFERRED\]' documentation/architecture/07-backup-architecture.md` returns the set +as a group. **It returns EIGHT lines, not seven** — the seven tagged rows plus the one line in §8's +own header that defines the marker. That is stated rather than hidden, because a count that does not +match what the reader sees is how an instrument stops being believed (R-421). It is a marker, not a status — the status cells are unchanged, because +**nothing was proven on the day this line was drawn** and a stopping line that moves a status is a +stopping line that lies. + ## Operator rulings — 2026-08-04 Recorded here because a ruling that lives only in a conversation binds nobody (the R-96 standing rule). @@ -224,7 +284,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing | **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | | **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC | | **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC | -| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` | +| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold.** It was re-scoped down on the morning of 2026-09-01 on the strength of *"recoverable file by file"*, and that clause was measured false the same afternoon (R-433). **On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not.** Both questions that can settle it are drafted in `documentation/runbooks/provider-questions-2026-09-01.md`: **Q1** decides how urgent this is, **Q2** (R-436) could remove the root cause cheaply. **Precondition on any build here: R-430**, which is harmless only while the box can still delete. | | **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC | | **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | | **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | @@ -445,9 +505,15 @@ This covers the next few only — it is deliberately **not** a full ordering of there is one ranking to maintain rather than two. 1. **R-95** — the largest *data* exposure: the tier holding the customer's documents and photos is the - one whose credential can delete. The snapshot mitigation is now **armed** (daily 00:00, keep 7), - but it has taken zero snapshots so far and it does not touch the root cause — the box can still - `forget --prune` its own repo. + one whose credential can delete. **THIS PARAGRAPH WAS STALE UNTIL 2026-09-01 AND ITS OLD TEXT IS + NAMED SO THE CORRECTION IS NOT MISTAKEN FOR A RE-RANK.** It said the snapshot mitigation was + *"armed (daily 00:00, keep 7), but it has taken zero snapshots so far"*. **Both halves were wrong:** + seven daily snapshots do exist (R-429), and the word "armed" was withdrawn by the spike that same + day. **What is true now:** the snapshots exist and the box cannot write into them, but **no account + we hold can read anything out of one** (R-433, measured — 777,600 names, zero hits, controlled), so + they do not yet bound this exposure. The root cause is untouched either way — the box can still + `forget --prune` its own repo, from two call sites. **The ORDER of this list is unchanged and is + Viktor's**; only the facts under item 1 were corrected. 2. **R-94** — **de-ranked 2026-07-29.** The prior rationale ("until it moves every hub-driven install gets the pre-R-82 default") was false: the constant selects no script and every install already fetches 1.22.0. What remains is a wrong label plus two pieces of dead safety equipment — @@ -602,11 +668,11 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-426** | **The decoy-coverage exemption list — 20 registered gates that ship WITHOUT a decoy test, each named.** `scripts/decoy_coverage_gate.py`'s `EXEMPT` map is debt, and this row owns it so it lives in the register and not only in a Python literal. **Four kinds:** (a) genuinely covered in the 2026-09-01 sweep but not yet moved into a suite — `hub-copy`, `instructions`, `docker-v`, `image-pins`; (b) blocked by an open hole and therefore un-assertable as rejecting — `site` (R-423), `one-register` (R-424), `offbox-rename` (R-425); (c) shared scripts whose decoy lives in `felhom.eu` and is counted there — `reuse-refs`, `instructions`, `observations` in the controller and agent runners; (d) **no plausible decoy constructed yet** — `hostinstall`, `wire-contract`, `due-checks`, `published`, `image-resolvable`, `volume-persistence`. Group (d) is the honest unknown: six gates whose soundness is UNTESTED, not established. **The list is green today and shrinks; a NEW gate with no decoy fails immediately.** | **OPEN — 20 names; group (d) is six untested gates** | | **R-427** | **`closed_register_gate.py` checks ONE direction only: an open word in a CLOSED row. The mirror — a CLOSED verdict on a row still sitting in `OPEN-ITEMS.md` — is unchecked, and there are TWELVE.** MEASURED 2026-09-01 during the decoy sweep, by reading the leading verdict of every open row with the gate's own predicate: **R-385, R-387, R-341, R-378, R-405, R-88a, R-88b** read unambiguously closed; **R-123, R-190, R-352** read `PARTLY CLOSED` / `MITIGATION SHIPPED` and almost certainly belong where they are. **The rows were NOT moved by this session** — telling a finished row from a partly-finished one is a judgement, and R-378 is itself the record of what happens when a machine makes that judgement on a substring (six still-open rows moved out of the register). **This is R-405's finding mirrored:** that row exists because R-87 sat in the CLOSED file while its state read READY, and the gate written for it looks only the way it was bitten. Fix: the same leading-verdict predicate applied to `OPEN-ITEMS.md`, reporting rather than convicting until the twelve are adjudicated by a person — a gate registered while twelve rows fail it would refuse every push. | **OPEN — 12 rows named; the adjudication is Viktor's, the gate is mine** | | **R-428** | **The decoy-coverage gate — written to catch instruments that match a NAME instead of a fact — identified a repository by its DIRECTORY NAME.** MEASURED on its own first CI run (felhom.eu job 490, 2026-09-01): `os.path.basename(root)` looked up in a `RUNNERS` map, and Gitea's act-runner checks the repo out into a directory called `hostexecutor`, so the gate reported *"unknown repo 'hostexecutor'"* and went INCONCLUSIVE. **The gate that hunts label-matching was matching a label, in the first ten lines of its own main loop, and it shipped that way.** FIXED the same day: it now identifies a repo by which registered runner FILE exists under the root, which is a fact. **Recorded rather than quietly patched because it is the strongest evidence in the sweep that this class is not a matter of carelessness** — it was written by a session that had spent the morning reading 29 gates for exactly this, with the four shapes on screen. Verified under a renamed directory before and after. | **CLOSED 2026-09-01 — fixed, and kept as the class's best example** | -| **R-430** | **`restic unlock --remove-all` printed `successfully removed locks` while the lock was still there.** MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as `nobody`; both controls passed first (create allowed, delete refused). The command reported success, returned, and `ls` showed the lock present. **`resticStep`'s crash-lock self-heal is built directly on this call** (`felhom-controller/controller/internal/backup/offbox.go:~768`), and its licence to escalate rests on the escalation actually working. **A self-heal that cannot fail is a self-heal that cannot be trusted** — this is this project's *"exit codes that lie"* class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. **Not yet established:** whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with `--verbose`. | **OPEN — precondition on any R-95 build** | +| **R-430** | **`restic unlock --remove-all` printed `successfully removed locks` while the lock was still there.** MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as `nobody`; both controls passed first (create allowed, delete refused). The command reported success, returned, and `ls` showed the lock present. **`resticStep`'s crash-lock self-heal is built directly on this call** (`felhom-controller/controller/internal/backup/offbox.go:~768`), and its licence to escalate rests on the escalation actually working. **A self-heal that cannot fail is a self-heal that cannot be trusted** — this is this project's *"exit codes that lie"* class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. **Not yet established:** whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with `--verbose`. **MARKED LATENT 2026-09-01.** It is **harmless today** and the reason is precise: the credential CAN delete, so the removal really happens and the success it reports is accidentally true. **THE CONDITION THAT MAKES IT LIVE — the only one, so it is stated as a trigger and not as prose: the moment delete is withdrawn from the box.** That is exactly what R-95's remedy does, by either route (retention moved off-box, or an append-only transport via R-436). From that moment `resticStep`'s crash-lock self-heal is escalating with a call that cannot fail, in the one path that runs unattended against the customer's off-site history. **So this is a PRECONDITION on the R-95 build, not a follow-up to it** — settle it in the same change or the self-heal ships already broken. | **OPEN — LATENT; becomes live the moment delete is withdrawn (precondition on any R-95 build)** | | **R-432** | **A customer's own sub-account can REACH the snapshot door and is REFUSED writes to it — but sees it EMPTY, so per-file recovery is not product-reachable.** MEASURED 2026-09-01 on BOTH live boxes, over the credential each already holds, with a positive and a negative control in the same run. **What is now PROVEN rather than cited:** `/.zfs` lists (`shares`, `snapshot`) from inside the jail; a write into `/.zfs/snapshot` is **REFUSED** — `dest open …: Failure` — while the identical write to the account home **succeeds** and was cleaned up. **That is the append-only property, measured, and it is the sentence the whole R-95 re-scope rests on.** **What is NOT available:** `/.zfs/snapshot` lists **empty** (link count 2) on both boxes, while the same Storage Box demonstrably holds seven snapshots — `storage-box-pool-1` IS `u629488` (`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`), the box these sub-accounts live on. So the contents are filtered from a sub-account. **CONSEQUENCE: recovery from a snapshot is an OPERATOR act in a browser, not something the product can drive** — which decides whether R-95's remedy can ever be customer-facing. **Cheapest next step, and it is Viktor's:** read one snapshot's name from the panel; a single `ls /.zfs/snapshot/` from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that). If it can, per-file recovery becomes product-reachable and this closes cheaply. **ANSWERED 2026-09-01 (DRILL, `audits/evidence-drill-r95-recovery-2026-09-01/`) — NO, AND THE PANEL READ IS NOT NEEDED.** The named-entry hypothesis was tested exhaustively and fails: **777,600 exact names** in the vendor-documented format `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes, **zero hits**, with a control proving the identical batch shape returns a path that does exist (6/6). **And there is a structural reason:** `df` reports `u629488-sub3` mounted on `/home` at **st_dev 0,82** while `/.zfs/snapshot` is **st_dev 0,276** — a different filesystem — and `/home/.zfs` does not exist. A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset owning that `.zfs`, not to the child at `/home`, so a correctly-named snapshot there **could not contain `felhom-repo`**. Three tools agree with controls in the same run (SFTP, the port-23 shell, `rsync --list-only`). The empty listing is not a display toggle hiding a reachable tree — from a sub-account there is no tree. **CONSEQUENCE: per-file recovery is not "operator-only", it is unreachable from the box entirely** → R-433. | **ANSWERED 2026-09-01 — negatively; the panel-read next step is WITHDRAWN as unnecessary** | -| **R-433** | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` | **OPEN — decides R-95's remedy and its rank** | -| **R-434** | **The snapshot-drop alarm promises a recovery that cannot be performed.** `hub/internal/monitor/offsite.go` `emitSnapshotDrop` ships this text, live in hub **v0.111.0**: *"The daily Storage Box snapshots are read-only and still hold the older copy, **so this is recoverable file-by-file**; it is NOT confirmed data loss."* The first clause is true. The second is not reachable: not by the product (R-433), and not by the operator without a browser and a main-account credential that does not exist here. Its own comment states the intent — *"THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually false"* — and the measurement it rests on was superseded the same day. **This is this project's own corollary landing on the alarm shipped that morning:** when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep. **Fix is text-only and must not be made before R-433 settles what IS true** — an alarm rewritten twice in a week is worse than one rewritten once. | **OPEN — text-only, blocked on R-433** | -| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. | **OPEN — documentation, not a threshold change** | +| **R-433** | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01.** The one question that can move this is drafted and ready to send: **Question 1** of `documentation/runbooks/provider-questions-2026-09-01.md` — *can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?* **Neither answer leaves this row where it is:** "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — **and R-95 becomes urgent.** Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. | **BLOCKED-ON-PROVIDER — Question 1 of `runbooks/provider-questions-2026-09-01.md`** | +| **R-434** | **The snapshot-drop alarm promises a recovery that cannot be performed.** `hub/internal/monitor/offsite.go` `emitSnapshotDrop` ships this text, live in hub **v0.111.0**: *"The daily Storage Box snapshots are read-only and still hold the older copy, **so this is recoverable file-by-file**; it is NOT confirmed data loss."* The first clause is true. The second is not reachable: not by the product (R-433), and not by the operator without a browser and a main-account credential that does not exist here. Its own comment states the intent — *"THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually false"* — and the measurement it rests on was superseded the same day. **This is this project's own corollary landing on the alarm shipped that morning:** when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep. **Fix is text-only and must not be made before R-433 settles what IS true** — an alarm rewritten twice in a week is worse than one rewritten once. **✅ CLOSED 2026-09-01, hub v0.111.1 — AND THE BLOCK ABOVE WAS WRONG, WHICH IS THE POINT WORTH KEEPING.** This row said the fix had to wait for R-433 to establish what IS true. **It did not, because the fix is a DELETION and not a REPLACEMENT.** The promise was withdrawn rather than swapped for a new one: *"The daily Storage Box snapshots are read-only and still hold the older copy. The route back out of them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a deletion ran on the box."* **That sentence is true under EVERY possible answer to the provider questions, so it never needs a second rewrite** — which is the whole reason it was not blocked. A replacement would have been. **Three tests in `hub/internal/monitor/offsite_r434_test.go`**, all driving the production path so they assert the sentence an operator RECEIVES: the withdrawal is present, the promise is absent in three shapes, `confirmed data loss` may appear only inside its negation, and the stored row must not drift from the delivered mail. **RED-PROOF: restoring the v0.111.0 sentence failed all three**, on every fragment, with the offending sentence printed. **One existing test was edited and it had caught this fix correctly** — `TestR431_FiresOnAMassDeletion` asserted `"NOT confirmed data loss"`; the fragment was REMOVED rather than updated so the wording keeps ONE home. | **CLOSED 2026-09-01 — hub v0.111.1; the "blocked on R-433" verdict above was mine and it was wrong** | +| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** | | **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. | **OPEN — ask the vendor before building anything** | | item | due (UTC) | what to measure | |---|---|---| +| R-433 | 2026-09-15 | Have Hetzner answered `runbooks/provider-questions-2026-09-01.md`? Record BOTH answers in R-433 and R-436, or re-date this row with the reason. The 2026-07-27 snapshot check sat unconfirmed for 36 days because it was never entered here — that is the scar this row exists to avoid repeating. | diff --git a/documentation/runbooks/provider-questions-2026-09-01.md b/documentation/runbooks/provider-questions-2026-09-01.md new file mode 100644 index 00000000..63a3dbc1 --- /dev/null +++ b/documentation/runbooks/provider-questions-2026-09-01.md @@ -0,0 +1,101 @@ +# Two questions for Hetzner — drafted, ready to send (2026-09-01) + +**These are the only thing standing between us and finishing R-95 properly.** They cost nothing and +they are answerable by a support agent without escalation. + +**Send them yourself.** CC drafted them and did not send them, and did not call the provider API — +`§11-D` is still the operator's fence. **No credential, password or token appears below, and none +should be added.** The account id and the product name are all either question needs. + +**What to fill in:** the ticket needs the Storage Box account. Ours is **`u629488`** (the box the +register calls `storage-box-pool-1`, plan BX11). Nothing else. + +**Why two separate tickets:** they go to different parts of the answer — one is about the snapshot +product, one is about the SSH endpoint's configuration — and a single ticket asking both tends to get +one answered and the other dropped. + +--- + +## Question 1 — can the MAIN account retrieve individual files from a snapshot? + +**Why it matters, in one line:** if it cannot, the only route back is a whole-box rollback that hits +every customer on the box and destroys every newer snapshot — which would mean the snapshots protect +almost nobody in practice. **This is the question that decides how urgent R-95 is.** + +> **Subject:** Storage Box u629488 — retrieving individual files from a snapshot +> +> Hello, +> +> We use Storage Box `u629488` with sub-accounts, and daily automatic snapshots are enabled. +> +> We can reach `/.zfs/snapshot` from a sub-account, but it lists as empty, and no snapshot name we +> try can be entered. We understand sub-accounts may be restricted here. +> +> Our question is about the **main account**: from the main account, over SSH or SFTP on port 23, +> can we **read or download individual files and directories out of a specific snapshot** — for +> example a single directory under one sub-account's home — **without** performing a snapshot +> restore of the whole Storage Box? +> +> If yes, please tell us the exact path we should use and how the snapshot directory is named. +> +> If no, please confirm that the only way to get data out of a snapshot is the full "restore +> snapshot" action on the whole Storage Box. +> +> Thank you. + +**How to read the answer.** +* **"Yes, from the main account"** → per-file recovery exists, but it is an operator act in a + browser or over the main account's own SSH, and it can never be something the product does for the + customer. R-433 closes at that. R-95 stays where it is. +* **"No, only a full restore"** → the snapshots do **not** bound a single customer's exposure at all, + because using them costs every other customer on the box their newer snapshots. **R-95 becomes + urgent and the transport change stops being optional.** + +--- + +## Question 2 — is `--append-only` enforced on the `rclone serve restic` endpoint? + +**Why it matters, in one line:** if it is enforced server-side, a compromised box **cannot delete its +own backups**, with no new machine and no data migration. **This is the question that could make R-95 +disappear.** + +> **Subject:** Storage Box u629488 — rclone serve restic endpoint and --append-only +> +> Hello, +> +> The restricted SSH shell on Storage Box `u629488` (port 23) lists `rclone serve restic --stdio` +> among the available server-side backends. +> +> Our question is about how that command is run on your side: **is `--append-only` enforced by you, +> or is the command line taken from what the client sends?** +> +> In other words, if a client connects and asks for `rclone serve restic --stdio` **without** +> `--append-only`, does it get a server that permits deletions? +> +> If the flag can be enforced, is there any way for us to request that for this account or for +> individual sub-accounts? +> +> Thank you. + +**How to read the answer.** +* **"Enforced server-side" or "can be enabled per account"** → this is the cheap prevention the + 2026-09-01 spike priced at a new always-on service plus either a mount in the hot path or migrating + every customer's history. **It needs none of that** — the server already runs at the provider, and + restic 0.14.0 already speaks the `rclone:` backend (measured, with a control: + `banana:` → `invalid backend`, `rclone:` → the helper was executed). The remaining work is putting + `rclone` in the controller image and switching the repository URL. **R-95's root cause goes away.** +* **"The client supplies the command line"** → **the lead is worth nothing** and should be recorded + as dead, not left looking promising. Prevention then still needs a machine in front of the store, + and the decision reverts to the spike's option 3 at its original price. + +--- + +## Where these came from + +* **R-433** — no snapshot is reachable from a sub-account by any name. 777,600 exact names in the + vendor's `YYYY-MM-DDTHH-MM-SS` format over nine days, zero hits, with a passing control; `/home` + and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Question 1 exists because + that measurement can only speak for a sub-account. +* **R-436** — the `rclone serve restic --stdio` backend and restic's `rclone:` support, both + measured. Question 2 is the one caveat that decides whether the lead is real. +* Evidence for both: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`. diff --git a/hub/CHANGELOG.md b/hub/CHANGELOG.md index ff5c3ad9..138d436a 100644 --- a/hub/CHANGELOG.md +++ b/hub/CHANGELOG.md @@ -1,3 +1,70 @@ +## v0.111.1 — the alarm stops promising a rescue that does not exist (2026-09-01, R-434) + +**Text-only patch on a live alarm. No controller change, no golden owed, no floor change.** + +**WHAT WAS WRONG.** v0.111.0 shipped, that morning, a snapshot-drop alarm reading: + +> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is +> recoverable file-by-file**; it is NOT confirmed data loss." + +**Measured the same afternoon (R-433): no snapshot is reachable from a customer's sub-account by ANY +name.** 777,600 exact names in the vendor's own `YYYY-MM-DDTHH-MM-SS` format across nine full days at +second granularity, plus 126 alternative shapes — zero hits, with a control proving the identical +batch returns a path that does exist (6/6). The structural reason: `/home` (the customer data) and +`/.zfs` are **different filesystems**, and `/home/.zfs` does not exist, so a snapshot behind that door +could not hold the repository even with the right name. **The promise named a route nobody can walk**, +in the one message an operator acts on while a customer's off-site history is disappearing. + +**THE FIX IS A DELETION, NOT A REPLACEMENT, AND THAT IS THE POINT.** R-434's register row said the fix +was blocked on R-433 — on first establishing what IS true. **It was not blocked, once the promise is +withdrawn rather than swapped.** The shipped sentence now asserts neither loss nor recovery: + +> "The daily Storage Box snapshots are read-only and still hold the older copy. The route back out of +> them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery. +> Get in touch before restoring anything, and check whether a deletion ran on the box." + +**That is true under every possible answer to the outstanding provider question, so it never needs a +second rewrite.** An alarm rewritten twice in a week is worse than one rewritten once: the operator +learns its words do not mean anything. + +**AND IT MUST NOT SWING THE OTHER WAY.** "Your backups are gone" is still usually FALSE — the +snapshots exist and hold the older copy; what is unproven is our route to them. Clause (a) of the +2026-09-01 measurement (the box **cannot write** into the snapshot area) stands and is re-confirmed. +Over-claiming loss would send an operator into a destructive recovery they did not need, which is the +failure v0.111.0's own comment was written to prevent — it had to survive its own correction. + +**TESTS — `internal/monitor/offsite_r434_test.go`, three, all driving the production path** +(`saveOffsiteReport` → `oc.Check()` → the notify callback), so they assert the sentence an operator +RECEIVES rather than the function that formats it. ASCII-only fragments, with a positive control (a +phrase present in every version of the alarm) and a negative control (a phrase that cannot exist). + +- `TestR434_AlarmMakesNoRecoveryPromise` — the withdrawal is present, the promise is absent, in three + shapes it could plausibly return as. +- `TestR434_AlarmStillDoesNotClaimDataLoss` — the other direction; `confirmed data loss` may appear + only inside its negation. +- `TestR434_StoredEventCarriesTheSameSentence` — the stored row and the delivered mail must not drift. + +**RED-PROOF (run 2026-09-01, before the fix was restored):** the v0.111.0 sentence was put back and +all three FAILED — on `"recoverable file-by-file"`, on `"so this is recoverable"`, on all three +required fragments, on the un-negated `confirmed data loss`, and on the stored row — each with the +offending sentence printed in the failure. + +**ONE EXISTING TEST WAS EDITED, AND IT CAUGHT THIS FIX CORRECTLY.** +`TestR431_FiresOnAMassDeletion` asserted the fragment `"NOT confirmed data loss"` and went red on the +new wording. The fragment is removed rather than updated: **the wording now has ONE home** +(`offsite_r434_test.go`), because duplicating it would create the second source that makes the next +correction land in one file and not the other. The signal fragments it still asserts (`69`, `4`, +`read-only`) are unchanged. + +**R-435 WRITTEN INTO THE DETECTOR'S OWN DOCUMENTATION, no threshold changed.** The comment above +`snapshotDropFraction` now states what this detector does NOT see: **a mass deletion, yes; one app +being wiped, no.** Worked on the live fleet — demo-hp's baseline is 69 snapshots across 9 apps, so +~35 must go before it speaks, and one app's tag is ~9. `offbox.go:1388` runs `forget --prune` grouped +by `host,tags`, so the blind spot sits on the most likely single-app failure. **The insensitivity is +deliberate and must not be "fixed" by lowering the numbers** — a detector that cries wolf is switched +off within a fortnight. What is not acceptable is claiming coverage it does not have; per-app +detection needs a SECOND signal keyed on the per-tag count. + ## v0.111.0 — notice a deletion within a day (2026-09-01, R-431; corrects R-429, re-scopes R-95) **Third signal in `OffsiteChecker`, beside FILL and STALENESS. No controller change, no golden owed.** diff --git a/hub/internal/monitor/offsite.go b/hub/internal/monitor/offsite.go index 92a1c767..6e94f2a1 100644 --- a/hub/internal/monitor/offsite.go +++ b/hub/internal/monitor/offsite.go @@ -238,6 +238,19 @@ func (oc *OffsiteChecker) isStale(customerID string, off *offsiteReport) bool { // reached by retention; the floor of 5 stops a tiny-count box alarming on ordinary ageing. It is // deliberately NOT sensitive — a detector that cries wolf is switched off within a fortnight, and // this project has proved that twice in a week. +// WHAT THIS DETECTOR DOES NOT SEE — R-435, and it must be read wherever "an unexplained fall is +// noticed within a day" is claimed, because that claim is true only of falls above the fraction. +// +// **It sees a MASS deletion. It does not see ONE APP being wiped.** Worked on the live fleet +// 2026-09-01: demo-hp's baseline is 69 snapshots across 9 apps, so ~35 must go before this speaks; +// one app's tag is ~9 and is invisible. And `offbox.go:1388` runs `forget --prune` **grouped by +// host,tags** — a per-tag wipe is exactly the shape a faulty retention or a targeted deletion +// produces, so the blind spot sits on the most likely single-app failure, not an exotic one. +// +// THIS IS DELIBERATE AND THE THRESHOLD SHOULD NOT BE LOWERED TO "FIX" IT. The reasoning is below: a +// detector that cries wolf is switched off within a fortnight, and this project has proved that +// twice in a week. What is NOT acceptable is claiming coverage this does not have. Anyone adding +// per-app detection should add a SECOND signal keyed on the per-tag count, not move these numbers. const ( snapshotDropFraction = 0.5 // more than half the history gone in one step snapshotDropFloor = 5 // and at least this many, so small counts do not twitch @@ -286,12 +299,25 @@ func (oc *OffsiteChecker) snapshotDropped(customerID string, off *offsiteReport) // THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually // false: the daily Storage Box snapshots are read-only to every account (proven, not cited) and hold // the older copy. It says what happened, what it means, and where the data still is. +// +// AND IT MUST NOT SAY THE DATA IS RECOVERABLE EITHER — R-434, fixed 2026-09-01, hub v0.111.1. The +// sentence shipped that morning promised "so this is recoverable file-by-file". Measured the same day +// (R-433): no snapshot is reachable from a sub-account by ANY name — 777,600 exact names in the +// vendor's own format over nine days, zero hits, with a passing control; `/home` and `/.zfs` are +// different filesystems and `/home/.zfs` does not exist. So the promise named a route nobody can walk. +// +// THE FIX IS A DELETION, NOT A REPLACEMENT, AND THAT IS THE WHOLE POINT. R-434's row said the fix was +// blocked on R-433 — on knowing what IS true. It is not, if the promise is simply withdrawn: a +// sentence that asserts neither loss nor recovery is true under EVERY possible answer to the provider +// question, so it never needs a second rewrite. An alarm rewritten twice in a week is worse than one +// rewritten once, because the operator learns its words do not mean anything. func (oc *OffsiteChecker) emitSnapshotDrop(customerID string, off *offsiteReport, prev, cur int) { message := fmt.Sprintf( "Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than "+ "retention can explain. The daily Storage Box snapshots are read-only and still hold the "+ - "older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check "+ - "whether a deletion ran on the box before restoring anything.", + "older copy. The route back out of them is not yet established, so treat this as neither "+ + "confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and "+ + "check whether a deletion ran on the box.", customerID, prev, cur) details, _ := json.Marshal(map[string]any{ "customer_id": customerID, "previous_count": prev, "current_count": cur, diff --git a/hub/internal/monitor/offsite_r431_test.go b/hub/internal/monitor/offsite_r431_test.go index d9116b92..73d18095 100644 --- a/hub/internal/monitor/offsite_r431_test.go +++ b/hub/internal/monitor/offsite_r431_test.go @@ -62,7 +62,15 @@ func TestR431_FiresOnAMassDeletion(t *testing.T) { default: t.Fatalf("severity %q is outside the hub vocabulary — it would be coerced to info and reach nobody", drops[0].sev) } - for _, frag := range []string{"69", "4", "read-only", "NOT confirmed data loss"} { + // These fragments belong to the SIGNAL — the two counts, and where the older copy still is. + // + // THE WORDING FRAGMENT THAT USED TO SIT HERE IS GONE ON PURPOSE. This list asserted + // "NOT confirmed data loss" until 2026-09-01, and it caught the R-434 fix, correctly — the + // sentence changed because the alarm was promising a recovery that R-433 showed cannot be + // performed. The wording now has ONE home, `offsite_r434_test.go`, which pins both what the + // message must say and what it must never say again. Duplicating it here would create the + // second source that makes the next correction land in one file and not the other. + for _, frag := range []string{"69", "4", "read-only"} { if !strings.Contains(drops[0].msg, frag) { t.Fatalf("message must contain %q; got: %s", frag, drops[0].msg) } diff --git a/hub/internal/monitor/offsite_r434_test.go b/hub/internal/monitor/offsite_r434_test.go new file mode 100644 index 00000000..442ae5b4 --- /dev/null +++ b/hub/internal/monitor/offsite_r434_test.go @@ -0,0 +1,155 @@ +package monitor + +import ( + "strings" + "testing" + "time" +) + +// R-434 — the snapshot-drop alarm must not promise a recovery that cannot be performed. +// +// WHAT WENT WRONG. hub v0.111.0 shipped, on 2026-09-01, an alarm reading "The daily Storage Box +// snapshots are read-only and still hold the older copy, so this is recoverable file-by-file; it is +// NOT confirmed data loss." Measured the same day (R-433): NO snapshot is reachable from a +// sub-account by any name — 777,600 exact names in the vendor format over nine days, zero hits, with +// a passing control. The promise named a route nobody can walk, in the one message an operator acts +// on while their customer's off-site history is disappearing. +// +// WHY THE FIX IS A DELETION AND NOT A REPLACEMENT. A sentence asserting neither loss nor recovery is +// true under every possible answer to the outstanding provider question, so it never needs a second +// rewrite. That is why this test pins the ABSENCE of a promise as hard as it pins the new words: +// the next person who "improves" this message by putting a route back into it must fail here. +// +// THESE TESTS DRIVE THE REAL PATH — saveOffsiteReport -> oc.Check() -> the notify callback — so they +// assert the CONSEQUENCE (the sentence an operator receives), not the mechanism. Asserting the +// mechanism one layer below where the damage happens is R-224, entry 9 of the doctrine table. +// +// ASCII-ONLY FRAGMENTS. The message contains an em dash. A fragment carrying one has returned 0 for +// strings that WERE there in this project before, so every fragment below is plain ASCII. + +// the promise that must never come back, in the shapes it could plausibly return as +var r434ForbiddenFragments = []string{ + "recoverable file-by-file", + "recoverable file by file", + "so this is recoverable", +} + +// the withdrawal that replaced it +var r434RequiredFragments = []string{ + "The route back out of them is not yet established", + "neither confirmed data loss nor confirmed recovery", + "Get in touch before restoring anything", + "still hold the older copy", // clause (a) STANDS and must not be lost with the promise +} + +// r434Message drives the production path once and returns the message the operator would receive. +func r434Message(t *testing.T) string { + t.Helper() + st := newDiskStore(t) + var msgs []string + saveOffsiteReport(t, st, "victim", dropJSON(69, true, "", "ok")) + oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, msg, _, _ string) { + if et == "offsite_snapshots_dropped" { + msgs = append(msgs, msg) + } + }, quietLog()) + + saveOffsiteReport(t, st, "victim", dropJSON(4, true, "", "ok")) + oc.Check() + + if len(msgs) != 1 { + t.Fatalf("setup: want exactly 1 offsite_snapshots_dropped message, got %d", len(msgs)) + } + return msgs[0] +} + +// TestR434_AlarmMakesNoRecoveryPromise — the fix, both directions, with both controls. +// +// RED-PROOF (run 2026-09-01, recorded in REPORT.md): restoring the v0.111.0 sentence in +// emitSnapshotDrop makes this FAIL on the forbidden fragment "recoverable file-by-file" AND on all +// three required fragments, with the offending sentence printed in the failure message. +func TestR434_AlarmMakesNoRecoveryPromise(t *testing.T) { + msg := r434Message(t) + + // POSITIVE CONTROL — a fragment present in EVERY version of this alarm. If this is missing the + // test is reading the wrong string and every other assertion below is worthless. + if !strings.Contains(msg, "off-site backup count fell from") { + t.Fatalf("positive control failed: not the snapshot-drop message at all: %q", msg) + } + // NEGATIVE CONTROL — proves Contains can actually report absence here. + if strings.Contains(msg, "zzz-no-such-fragment-r434") { + t.Fatalf("negative control failed: matched a fragment that cannot exist: %q", msg) + } + + for _, bad := range r434ForbiddenFragments { + if strings.Contains(msg, bad) { + t.Errorf("alarm promises a recovery that cannot be performed (R-433): found %q in %q", bad, msg) + } + } + for _, want := range r434RequiredFragments { + if !strings.Contains(msg, want) { + t.Errorf("alarm is missing the withdrawal wording: want %q in %q", want, msg) + } + } +} + +// TestR434_AlarmStillDoesNotClaimDataLoss — the OTHER direction, and the reason the fix is a +// withdrawal rather than a reversal. +// +// After R-433 the temptation is to swing to "your backups are gone". That is still usually FALSE: +// the snapshots exist and hold the older copy; what is unproven is our route to them. An alarm that +// over-claims loss sends an operator into a destructive recovery they did not need — which is the +// failure the v0.111.0 comment was written to prevent, and it must survive its own correction. +func TestR434_AlarmStillDoesNotClaimDataLoss(t *testing.T) { + msg := r434Message(t) + + for _, bad := range []string{ + "data is lost", "backups are gone", "data has been lost", "permanently lost", "unrecoverable", + } { + if strings.Contains(msg, bad) { + t.Errorf("alarm over-claims loss: found %q in %q", bad, msg) + } + } + // The one phrase that must appear NEGATED, never bare. A bare "confirmed data loss" would read + // as a verdict; the shipped sentence only ever uses it inside "neither ... nor". + if strings.Contains(msg, "confirmed data loss") && + !strings.Contains(msg, "neither confirmed data loss nor confirmed recovery") { + t.Errorf("the phrase 'confirmed data loss' appears outside its negation: %q", msg) + } +} + +// TestR434_StoredEventCarriesTheSameSentence — the delivered message and the stored one are the same +// string today, and a future refactor that formats them separately must not let them drift: the +// operator reads the mail, but every later audit reads the stored row. +func TestR434_StoredEventCarriesTheSameSentence(t *testing.T) { + st := newDiskStore(t) + var delivered string + saveOffsiteReport(t, st, "victim", dropJSON(69, true, "", "ok")) + oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, msg, _, _ string) { + if et == "offsite_snapshots_dropped" { + delivered = msg + } + }, quietLog()) + saveOffsiteReport(t, st, "victim", dropJSON(4, true, "", "ok")) + oc.Check() + + evs, err := st.GetRecentEvents("victim", 50) + if err != nil { + t.Fatalf("GetRecentEvents: %v", err) + } + var stored []string + for _, e := range evs { + if e.EventType == "offsite_snapshots_dropped" { + stored = append(stored, e.Message) + } + } + if len(stored) != 1 { + t.Fatalf("want exactly 1 stored offsite_snapshots_dropped, got %d", len(stored)) + } + if stored[0] != delivered { + t.Errorf("stored and delivered messages have drifted:\n stored: %q\n delivered: %q", stored[0], delivered) + } + if strings.Contains(stored[0], "recoverable file-by-file") { + t.Errorf("the stored row still carries the withdrawn promise: %q", stored[0]) + } +}