Files
felhom.eu/REPORT.md
T
admin 10c223bdfe
gates / gates (push) Successful in 19s
DRILL R-95: the recovery route does not exist — stopped before the destructive phase
The drill was to delete demo-hp's off-site history and get it back out of a Storage
Box snapshot, filling row 10's blank RTO. Phase 1 found there is nothing to get it
back from: no snapshot is reachable from a sub-account BY ANY NAME.

Measured, read-only, no delete verb issued against any live store:
- 777,600 exact names in the vendor form YYYY-MM-DDTHH-MM-SS, nine full days at
  second granularity, plus 126 alternative shapes -> ZERO hits.
- The control is what makes that mean anything: the identical 600-name batch shape
  with one real path appended returned it, 6 of 6.
- Structural cause: /home (u629488-sub3) is st_dev 0,82; /.zfs/snapshot is st_dev
  0,276; /home/.zfs does not exist. A snapshot under /.zfs/snapshot belongs to a
  different dataset than the one holding felhom-repo.
- Three tools agree with controls in the same run: SFTP, the port-23 shell,
  rsync --list-only.

So yesterday's re-scope splits: clause (a) "the box cannot write into the snapshot
area" STANDS and is re-confirmed; clause (b) "the rest is recoverable file by file"
is NOT SUPPORTED. STOPPED before Phase 2 on the operator's ruling — with no recovery
leg the deletion would have destroyed real history to buy only an alarm test that
could not fire at the specified size. Store verified untouched at 69 snapshots.

R-432 ANSWERED (negatively; its panel-read next step withdrawn as unnecessary).
R-433 no snapshot reachable by any name — decides R-95's remedy and its rank.
R-434 the drop alarm's text promises a file-by-file recovery that cannot be performed.
R-435 the drop detector is blind to a single-app deletion (>50% of 69 needed, ~9 given).
R-436 LEAD: the provider offers `rclone serve restic --stdio` and restic 0.14.0 speaks
      `rclone:` (measured, controlled) — real prevention may need no new machine, IF
      the vendor pins --append-only. Ask before building.

07 §8 row 10: text corrected, status NOT moved, RTO still blank.
No code, no version bump, no image, no golden. REPORT.md's only copy of the R-331
report preserved as REPORT-r331-backup-card.md before overwrite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 16:55:50 +02:00

12 KiB
Raw Blame History

REPORT — the deletion we said is survivable: the recovery route does not exist (R-95 drill, 2026-09-01)

RUNBOOK, destructive class, demo-hp only. STOPPED at the end of Phase 1 on the operator's ruling, before any destructive step. No delete verb was issued against any live store; no byte on either Storage Box sub-account was written, moved or removed. No production code, no version bump, no image, no golden. Evidence: documentation/audits/evidence-drill-r95-recovery-2026-09-01/.

# phase verdict one sentence
1 snapshot reachable, and its name NO — and it has no reachable name 777,600 exact names across nine days in the vendor's own format, plus 126 alternative shapes; zero resolve, with a control proving the sweep detects a path that exists.
2 the deletion NOT RUN — operator ruling With no recovery route, the deletion would have destroyed real history to buy nothing; put as a two-option decision, the ruling was stop.
3 the alarm fired NOT RUN — and it could not have fired at the specified size The shipped threshold needs a fall of more than half; one app's tag is ~9 of 69. → R-435
4 the recovery NOT RUN — no route exists that is not fenced Box-side: proven impossible. Panel: fenced and browserless. Hetzner API: fenced (§11-D), and the hub's client has no snapshot method at all.
— RTO from T₀ STILL BLANK Row 10's RTO cell is unchanged and remains a finding.
— data lost, quantified NOT MEASURABLE THIS WAY The quantity only has meaning if the rest is recoverable, and the route that would recover it is not reachable.
5 re-arm NOT RUN Depended on Phase 4.
6 teardown PASS Store untouched at 69 snapshots; all three scratch layers removed; hub DB copies shredded; both boxes healthy.

1. Did the recovery work — and does yesterday's re-scope survive?

The recovery was never reachable, and the re-scope does not survive intact. Its first half stands; its second half does not.

Yesterday's re-scope has two clauses. They must now be separated:

  • (a) "The box can delete its live repository, but cannot write to the daily snapshots of it." STANDS. Re-confirmed here: /.zfs/snapshot is reachable and the write-refusal measurement is unchanged. Nothing in this drill weakens it.
  • (b) "…so the rest is recoverable — file by file, one customer at a time." NOT SUPPORTED. A snapshot that cannot be opened cannot be copied out of. R-432 recorded the directory listing empty and named the cheapest next step: "a single ls /.zfs/snapshot/<name> from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that)." That step is now done, exhaustively, and the answer is no.

What was measured. The port-23 restricted shell accepts a batched stat, which makes a cheap existence oracle: 500–600 paths per round trip, stdout carrying only paths that exist.

sweep candidates hits
/.zfs/snapshot/YYYY-MM-DDTHH-MM-SS, nine full days, second granularity 777,600 0
126 alternative name shapes and snapshot paths (daily, snapshot-1, colon and compact time forms, /home/.snapshot, …) 126 0
control — the identical 600-name batch shape with one real path appended 6 batches 6/6 returned it

And there is a structural reason, which is why I stopped sweeping. The customer's data and the snapshot door are on different filesystems:

df           →  u629488-sub3 mounted on /home
stat /home            →  Device 0,82
stat /.zfs/snapshot   →  Device 0,276      ← a different device
stat /home/.zfs       →  cannot statx: No such file or directory

A ZFS snapshot under /.zfs/snapshot belongs to the dataset that owns that .zfs — not to the child mounted at /home. So even a correctly named snapshot there could not contain felhom-repo, and the dataset that does hold it exposes no .zfs at all to this account. The empty listing is not a display toggle hiding a reachable tree; from a sub-account there is no tree.

Three tools agree, each with controls in the same run: SFTP, the port-23 shell, and rsync --list-only.

What that does to R-95. Its exposure is unchanged and its remedy is not. Yesterday the row could say a deletion costs about a day because the rest comes back per-file. Today the only routes to "the rest" are a whole-box panel rollback (which deletes newer snapshots and hits every customer on the box) and the provider API (fenced, and unimplemented in the hub's client). The re-scope's comfort was resting on a route nobody had walked — which is precisely the standard this project applies, and it is the reason this drill was called.

The ranking is Viktor's and I am not re-ranking it. What I will say plainly: the argument that moved R-95 down yesterday is the argument this drill removed. On these facts I would put it back where it was.

2. The RTO

Still blank, and it stays a finding. 07 §8 row 10's RTO cell has been empty since July and this drill did not fill it. Nothing was recovered, so nothing was timed. The summary line at 07-backup-architecture.md:948 — "no ransomware-shaped recovery has ever been run" — is still true, and is now true for a sharper reason: not "nobody has run it" but "from the box, it cannot be run."

3. R-432's answer, and the naming scheme

R-432 is ANSWERED, and negatively. It did not need the panel read it was waiting on.

  • The naming scheme is YYYY-MM-DDTHH-MM-SS — vendor-documented examples 2025-12-03T13-47-47, 2025-02-12T11-35-19. Recorded so nobody hunts a console again.
  • Knowing it does not help. Every name in that format for nine days is refused, and the st_dev split above says why. Per-file recovery is not operator-only — from the box it is nobody's, and for the operator it is a browser act against the main account that no credential in this project can perform.
  • The panel cannot supply the missing piece either. It offers Restore and Delete on a row and does not show names; and the one name-shaped thing it could give would be tried against a door that leads to the wrong dataset.

4. The alarm's first real firing

It did not happen, and the drill as written could not have produced it. The detector fires on a fall of more than half the previous count and at least 5 (hub/internal/monitor/offsite.go, snapshotDropFraction = 0.5, snapshotDropFloor = 5). demo-hp's baseline is 69. Phase 2 deletes one app's history — about 9 snapshots. 9 is over the floor and nowhere near half, so the alarm stays silent, correctly and by design. Firing it for real needs ~35+ snapshots destroyed, i.e. most of demo-hp's off-site history. That trade is what the operator was asked to rule on. → R-435

One thing the alarm says is now wrong. Its message, live in hub 0.111.0, reads:

"The daily Storage Box snapshots are read-only and still hold the older copy, so this is recoverable file-by-file; it is NOT confirmed data loss."

The first clause is true; the second promises a recovery the product cannot perform and the operator cannot perform without a browser and the main account. This is this project's own corollary — when a verdict changes which field it counts from, the alarm text has to change with it — landing on the alarm shipped the same day. → R-434

5. Findings, as register rows

All four filed in documentation/backlog/OPEN-ITEMS.md.

row finding
R-433 A sub-account cannot reach any Storage Box snapshot by any name; /home and /.zfs are different filesystems and /home/.zfs does not exist. Answers R-432 negatively and removes clause (b) of the R-95 re-scope.
R-434 emitSnapshotDrop's message promises file-by-file recovery that is not reachable. Live in hub 0.111.0.
R-435 The drop detector cannot see a single-app deletion (>50% of 69 ⇒ ~35 needed). offbox.go:1388 forgets by tag, so a single-tag wipe is exactly the shape the detector is blind to. Deliberate insensitivity, but the blind spot should be stated where the operator reads it.
R-436 LEAD, not a defect. Hetzner's port-23 shell offers rclone serve restic --stdio as a server-side backend, and restic 0.14.0 recognises the rclone: backend (measured; control banana: → invalid backend; rclone is absent from the controller image). rclone serve restic carries --append-only. This could make R-95's real prevention far cheaper than the spike concluded — no new always-on machine, no data migration. Caveat stated up front: the client supplies the server command line, so a compromised box could omit the flag unless the provider pins it. Settling that is a vendor question, not a code change.

R-432 is marked ANSWERED; its "one panel read settles it" next step is withdrawn as unnecessary.

6. Does 07 §8 row 10 move?

No. It stays PARTIAL, and its RTO stays blank. The status was already correct for the right reason — "the recovery ROUTE has never been walked, which is what PARTIAL means" — and this drill found the route is not walkable from the box at all. What the row needs is a text correction, not a status change: its clause "recoverable per-file (vendor)" and its limit "per-file recovery is operator-only today (R-432)" both overstate what exists. Updated in place with the citation. Moving it only as far as the evidence goes means not moving it.

7. What could not be tested, and why

  • Whether the main account can see the snapshots. No main-account credential exists in this project — the hub holds only per-customer sub-accounts. This is the one question that would decide whether per-file recovery exists at all, for anyone.
  • Whether the Hetzner API can list or read a snapshot. Fenced by the runbook (§11-D). Separately, hub/internal/hetznerapi/hetznerapi.go has no snapshot method — so this route needs new code regardless of the fence.
  • The deletion, the alarm's first real firing, the recovery, the RTO, the re-arm. Phases 2–5, not run, on the operator's ruling.
  • Whether rclone serve restic --stdio is pinned server-side with --append-only (R-436).

8. My own mistakes

  • I ran Phase 1 before finishing Phase 0's subject-app choice, and did not say so up front. The choice depends on the snapshot's timestamp, so the order was right, but the runbook's order is the runbook's and a silent reordering is the thing this project keeps getting caught by. Stated at the time in the session, recorded here.
  • My first sweep guessed the schedule instead of establishing it. I probed 00:00 UTC and 22:00 UTC — 600 names — on the strength of a register line reading "daily 00:00", got nothing, and only then widened to whole days. The narrow sweep was worth nothing on its own: a zero over a guessed window is not evidence, and I should have gone to full days first or not run it at all.
  • I nearly reported the empty listing as "the display toggle is hiding it". The vendor documents exactly such a toggle and it fitted. The st_dev comparison — which I only ran because df printed a filesystem name I did not expect — says the tree is on another dataset entirely. A plausible cause that fits the symptom is not a measured one, and I had the wrong one for about ten minutes.
  • REPORT.md held the only copy of the R-331 report (hub v0.109.0, 2026-08-30) — durable content living only in the overwritten file, which CLAUDE.md:82-87 forbids. Preserved as REPORT-r331-backup-card.md before this report replaced it.