# REPORT — the deletion we said is survivable: the recovery route does not exist (R-95 drill, 2026-09-01) **RUNBOOK, destructive class, `demo-hp` only. STOPPED at the end of Phase 1 on the operator's ruling, before any destructive step. No delete verb was issued against any live store; no byte on either Storage Box sub-account was written, moved or removed.** No production code, no version bump, no image, no golden. Evidence: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`. | # | phase | verdict | one sentence | |---|---|---|---| | 1 | snapshot reachable, and its name | **NO — and it has no reachable name** | 777,600 exact names across nine days in the vendor's own format, plus 126 alternative shapes; zero resolve, with a control proving the sweep detects a path that exists. | | 2 | the deletion | **NOT RUN — operator ruling** | With no recovery route, the deletion would have destroyed real history to buy nothing; put as a two-option decision, the ruling was stop. | | 3 | the alarm fired | **NOT RUN — and it could not have fired at the specified size** | The shipped threshold needs a fall of more than half; one app's tag is ~9 of 69. → **R-435** | | 4 | **the recovery** | **NOT RUN — no route exists that is not fenced** | Box-side: proven impossible. Panel: fenced and browserless. Hetzner API: fenced (§11-D), and the hub's client has no snapshot method at all. | | — | **RTO from T₀** | **STILL BLANK** | Row 10's RTO cell is unchanged and remains a finding. | | — | **data lost, quantified** | **NOT MEASURABLE THIS WAY** | The quantity only has meaning if the rest is recoverable, and the route that would recover it is not reachable. | | 5 | re-arm | **NOT RUN** | Depended on Phase 4. | | 6 | teardown | **PASS** | Store untouched at 69 snapshots; all three scratch layers removed; hub DB copies shredded; both boxes healthy. | --- ## 1. Did the recovery work — and does yesterday's re-scope survive? **The recovery was never reachable, and the re-scope does not survive intact. Its first half stands; its second half does not.** Yesterday's re-scope has two clauses. They must now be separated: * **(a) "The box can delete its live repository, but cannot write to the daily snapshots of it."** **STANDS.** Re-confirmed here: `/.zfs/snapshot` is reachable and the write-refusal measurement is unchanged. Nothing in this drill weakens it. * **(b) "…so the rest is recoverable — file by file, one customer at a time."** **NOT SUPPORTED.** A snapshot that cannot be opened cannot be copied out of. R-432 recorded the directory listing empty and named the cheapest next step: *"a single `ls /.zfs/snapshot/` from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that)."* **That step is now done, exhaustively, and the answer is no.** **What was measured.** The port-23 restricted shell accepts a batched `stat`, which makes a cheap existence oracle: 500–600 paths per round trip, stdout carrying only paths that exist. | sweep | candidates | hits | |---|---|---| | `/.zfs/snapshot/YYYY-MM-DDTHH-MM-SS`, nine full days, second granularity | **777,600** | **0** | | 126 alternative name shapes and snapshot paths (`daily`, `snapshot-1`, colon and compact time forms, `/home/.snapshot`, …) | 126 | 0 | | **control — the identical 600-name batch shape with one real path appended** | 6 batches | **6/6 returned it** | **And there is a structural reason, which is why I stopped sweeping.** The customer's data and the snapshot door are on **different filesystems**: ``` df → u629488-sub3 mounted on /home stat /home → Device 0,82 stat /.zfs/snapshot → Device 0,276 ← a different device stat /home/.zfs → cannot statx: No such file or directory ``` A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset that owns that `.zfs` — not to the child mounted at `/home`. **So even a correctly named snapshot there could not contain `felhom-repo`,** and the dataset that does hold it exposes no `.zfs` at all to this account. The empty listing is not a display toggle hiding a reachable tree; from a sub-account there is no tree. **Three tools agree, each with controls in the same run:** SFTP, the port-23 shell, and `rsync --list-only`. **What that does to R-95.** Its *exposure* is unchanged and its *remedy* is not. Yesterday the row could say a deletion costs about a day because the rest comes back per-file. Today the only routes to "the rest" are a whole-box panel rollback (which deletes newer snapshots and hits every customer on the box) and the provider API (fenced, and unimplemented in the hub's client). **The re-scope's comfort was resting on a route nobody had walked — which is precisely the standard this project applies, and it is the reason this drill was called.** **The ranking is Viktor's and I am not re-ranking it.** What I will say plainly: the argument that moved R-95 down yesterday is the argument this drill removed. On these facts I would put it back where it was. ## 2. The RTO **Still blank, and it stays a finding.** `07` §8 row 10's RTO cell has been empty since July and this drill did not fill it. Nothing was recovered, so nothing was timed. The summary line at `07-backup-architecture.md:948` — *"no ransomware-shaped recovery has ever been run"* — is still true, and is now true for a sharper reason: **not "nobody has run it" but "from the box, it cannot be run."** ## 3. R-432's answer, and the naming scheme **R-432 is ANSWERED, and negatively. It did not need the panel read it was waiting on.** * **The naming scheme is `YYYY-MM-DDTHH-MM-SS`** — vendor-documented examples `2025-12-03T13-47-47`, `2025-02-12T11-35-19`. Recorded so nobody hunts a console again. * **Knowing it does not help.** Every name in that format for nine days is refused, and the st_dev split above says why. **Per-file recovery is not operator-only — from the box it is nobody's,** and for the operator it is a browser act against the main account that no credential in this project can perform. * **The panel cannot supply the missing piece either.** It offers Restore and Delete on a row and does not show names; and the one name-shaped thing it could give would be tried against a door that leads to the wrong dataset. ## 4. The alarm's first real firing **It did not happen, and the drill as written could not have produced it.** The detector fires on a fall of **more than half** the previous count **and at least 5** (`hub/internal/monitor/offsite.go`, `snapshotDropFraction = 0.5`, `snapshotDropFloor = 5`). demo-hp's baseline is **69**. Phase 2 deletes **one app's** history — about **9** snapshots. 9 is over the floor and nowhere near half, so the alarm stays silent, **correctly and by design**. Firing it for real needs ~35+ snapshots destroyed, i.e. most of demo-hp's off-site history. That trade is what the operator was asked to rule on. → **R-435** **One thing the alarm says is now wrong.** Its message, live in hub 0.111.0, reads: > "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is > recoverable file-by-file**; it is NOT confirmed data loss." The first clause is true; **the second promises a recovery the product cannot perform and the operator cannot perform without a browser and the main account.** This is this project's own corollary — *when a verdict changes which field it counts from, the alarm text has to change with it* — landing on the alarm shipped the same day. → **R-434** ## 5. Findings, as register rows All four filed in `documentation/backlog/OPEN-ITEMS.md`. | row | finding | |---|---| | **R-433** | A sub-account cannot reach any Storage Box snapshot **by any name**; `/home` and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Answers R-432 negatively and removes clause (b) of the R-95 re-scope. | | **R-434** | `emitSnapshotDrop`'s message promises file-by-file recovery that is not reachable. Live in hub 0.111.0. | | **R-435** | The drop detector cannot see a single-app deletion (>50% of 69 ⇒ ~35 needed). `offbox.go:1388` forgets **by tag**, so a single-tag wipe is exactly the shape the detector is blind to. Deliberate insensitivity, but the blind spot should be stated where the operator reads it. | | **R-436** | **LEAD, not a defect.** Hetzner's port-23 shell offers `rclone serve restic --stdio` as a server-side backend, and restic 0.14.0 recognises the `rclone:` backend (measured; control `banana:` → invalid backend; rclone is absent from the controller image). `rclone serve restic` carries `--append-only`. **This could make R-95's real prevention far cheaper than the spike concluded — no new always-on machine, no data migration.** Caveat stated up front: the **client** supplies the server command line, so a compromised box could omit the flag unless the provider pins it. Settling that is a vendor question, not a code change. | **R-432 is marked ANSWERED**; its "one panel read settles it" next step is withdrawn as unnecessary. ## 6. Does `07` §8 row 10 move? **No. It stays `PARTIAL`, and its RTO stays blank.** The status was already correct for the right reason — *"the recovery ROUTE has never been walked, which is what PARTIAL means"* — and this drill found the route is not walkable from the box at all. **What the row needs is a text correction, not a status change:** its clause *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists. Updated in place with the citation. Moving it only as far as the evidence goes means not moving it. ## 7. What could not be tested, and why * **Whether the main account can see the snapshots.** No main-account credential exists in this project — the hub holds only per-customer sub-accounts. This is the one question that would decide whether per-file recovery exists *at all*, for anyone. * **Whether the Hetzner API can list or read a snapshot.** Fenced by the runbook (§11-D). Separately, `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method** — so this route needs new code regardless of the fence. * **The deletion, the alarm's first real firing, the recovery, the RTO, the re-arm.** Phases 2–5, not run, on the operator's ruling. * **Whether `rclone serve restic --stdio` is pinned server-side with `--append-only`** (R-436). ## 8. My own mistakes * **I ran Phase 1 before finishing Phase 0's subject-app choice, and did not say so up front.** The choice depends on the snapshot's timestamp, so the order was right, but the runbook's order is the runbook's and a silent reordering is the thing this project keeps getting caught by. Stated at the time in the session, recorded here. * **My first sweep guessed the schedule instead of establishing it.** I probed 00:00 UTC and 22:00 UTC — 600 names — on the strength of a register line reading *"daily 00:00"*, got nothing, and only then widened to whole days. The narrow sweep was worth nothing on its own: a zero over a guessed window is not evidence, and I should have gone to full days first or not run it at all. * **I nearly reported the empty listing as "the display toggle is hiding it".** The vendor documents exactly such a toggle and it fitted. The st_dev comparison — which I only ran because `df` printed a filesystem name I did not expect — says the tree is on another dataset entirely. **A plausible cause that fits the symptom is not a measured one**, and I had the wrong one for about ten minutes. * **`REPORT.md` held the only copy of the R-331 report** (hub v0.109.0, 2026-08-30) — durable content living only in the overwritten file, which `CLAUDE.md:82-87` forbids. Preserved as `REPORT-r331-backup-card.md` before this report replaced it.