Files
felhom.eu/documentation/audits/evidence-drill-r95-recovery-2026-09-01
admin 10c223bdfe
gates / gates (push) Successful in 19s
DRILL R-95: the recovery route does not exist — stopped before the destructive phase
The drill was to delete demo-hp's off-site history and get it back out of a Storage
Box snapshot, filling row 10's blank RTO. Phase 1 found there is nothing to get it
back from: no snapshot is reachable from a sub-account BY ANY NAME.

Measured, read-only, no delete verb issued against any live store:
- 777,600 exact names in the vendor form YYYY-MM-DDTHH-MM-SS, nine full days at
  second granularity, plus 126 alternative shapes -> ZERO hits.
- The control is what makes that mean anything: the identical 600-name batch shape
  with one real path appended returned it, 6 of 6.
- Structural cause: /home (u629488-sub3) is st_dev 0,82; /.zfs/snapshot is st_dev
  0,276; /home/.zfs does not exist. A snapshot under /.zfs/snapshot belongs to a
  different dataset than the one holding felhom-repo.
- Three tools agree with controls in the same run: SFTP, the port-23 shell,
  rsync --list-only.

So yesterday's re-scope splits: clause (a) "the box cannot write into the snapshot
area" STANDS and is re-confirmed; clause (b) "the rest is recoverable file by file"
is NOT SUPPORTED. STOPPED before Phase 2 on the operator's ruling — with no recovery
leg the deletion would have destroyed real history to buy only an alarm test that
could not fire at the specified size. Store verified untouched at 69 snapshots.

R-432 ANSWERED (negatively; its panel-read next step withdrawn as unnecessary).
R-433 no snapshot reachable by any name — decides R-95's remedy and its rank.
R-434 the drop alarm's text promises a file-by-file recovery that cannot be performed.
R-435 the drop detector is blind to a single-app deletion (>50% of 69 needed, ~9 given).
R-436 LEAD: the provider offers `rclone serve restic --stdio` and restic 0.14.0 speaks
      `rclone:` (measured, controlled) — real prevention may need no new machine, IF
      the vendor pins --append-only. Ask before building.

07 §8 row 10: text corrected, status NOT moved, RTO still blank.
No code, no version bump, no image, no golden. REPORT.md's only copy of the R-331
report preserved as REPORT-r331-backup-card.md before overwrite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 16:55:50 +02:00
..

Evidence — DRILL: the deletion we said is survivable (R-95 recovery drill, 2026-09-01)

The drill STOPPED at the end of Phase 1, on the operator's ruling, before any destructive step. No delete verb was issued against any live store. No byte on either Storage Box sub-account was written, moved or removed. Every probe in this directory is a read: ls, stat, tree, du, df, rsync --list-only, restic snapshots.

Why it stopped

Phase 1 was supposed to find a Storage Box snapshot so Phase 4 could recover from it. It found that no snapshot is reachable from the box by any name, and that no unfenced route to one exists for CC at all. The recovery leg therefore could not run, and deleting a real app's off-site history would have bought only an alarm test that — measured against the shipped threshold — could not have fired at the size the runbook specifies. Put to the operator as a two-option decision; the ruling was stop and report.

Files

file what it is
phase0-ground-truth/01-offsite-inventory.txt demo-hp's full off-site inventory — 69 restic snapshots, 9 apps, ids + tags + paths. The baseline.
phase0-ground-truth/02-hub-view.txt the hub's latest offsite object for demo-hp (snapshot_count: 69, stats_known: true, last_status: ok). Cross-checks the restic count exactly.
phase1-find-the-snapshot/01-probe-tree.txt SFTP probes of /.zfs, /.zfs/snapshot, /home/.zfs, with a positive and a negative control. Reproduces R-432.
phase1-find-the-snapshot/02-shell-and-rsync-doors.txt two independent tools (a real shell on port 23, and rsync --list-only) agree the tree lists empty. Both controlled.
phase1-find-the-snapshot/03-shell-capabilities.txt the port-23 restricted shell's full help — the command set the credential actually reaches.
phase1-find-the-snapshot/04-enumeration-attempts.txt df/stat/tree/du on the snapshot door. Where the st_dev split was found.
phase1-find-the-snapshot/05-home-zfs-door.txt /home/.zfs does not exist — three tools, negative control in the same run.
phase1-find-the-snapshot/06-name-sweep.txt the name sweep: nine full days, second granularity, 777,600 candidate names, zero hits.
phase1-find-the-snapshot/07-sweep-control.txt the control for that sweep — the identical 600-name batch shape with one real path appended; 6/6 batches returned it. Without this the zero result would prove nothing.
phase1-find-the-snapshot/08-alt-formats.txt 126 non-timestamp name shapes and alternative snapshot paths. Only the controls resolved.
phase1-find-the-snapshot/09-rclone-restic-lead.txt the rclone: backend measurement behind R-436, with the banana: invalid-backend control.
phase6-teardown/01-verify-then-clean.txt the store is untouched: still 69 snapshots, home is exactly .ssh + felhom-repo, repo top level intact.
phase6-teardown/02-guest-and-host-clean.txt scratch removed from the guest and the PVE host; controller 0.232.0 healthy, 19 healthy containers.
phase6-teardown/03-dooplex-clean-and-final-state.txt hub DB copies shredded (both, -wal included); both customers' tiers still ok; no alarm event raised by this session.

The method that makes the negative result trustworthy

A sweep that finds nothing is worthless unless it is shown it could have found something. The oracle is a batched stat -c %n over the port-23 shell: 500–600 paths per round trip, and stdout carries only the paths that exist. 07-sweep-control.txt runs the identical batch shape with /home appended and gets /home back, 6 times out of 6. So the zero in 06-name-sweep.txt is a measurement, not a silence.

Not done, and why

  • No Hetzner API call. Fenced by the runbook (§11-D). The hub's own API client has no snapshot method at all, so even unfenced it would have needed new code.
  • No panel action. No browser on DooPlex, and "Restore snapshot" is fenced in any case — it rolls back the whole Storage Box and deletes newer snapshots.
  • demo-felhom never addressed. Every storage-box connection in this session authenticated as u629488-sub3, demo-hp's own sub-account. Its tier is verified still reporting ok in phase6-teardown/03-*.