db38f4c800
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.
was: "...still hold the older copy, so this is recoverable file-by-file; it is NOT
confirmed data loss. Check whether a deletion ran on the box before restoring."
now: "...still hold the older copy. The route back out of them is not yet established,
so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
before restoring anything, and check whether a deletion ran on the box."
It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.
Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.
R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.
THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.
Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.
R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.
Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
167 lines
12 KiB
Markdown
167 lines
12 KiB
Markdown
# REPORT — the deletion we said is survivable: the recovery route does not exist (R-95 drill, 2026-09-01)
|
||
|
||
**RUNBOOK, destructive class, `demo-hp` only. STOPPED at the end of Phase 1 on the operator's ruling,
|
||
before any destructive step. No delete verb was issued against any live store; no byte on either
|
||
Storage Box sub-account was written, moved or removed.** No production code, no version bump, no
|
||
image, no golden. Evidence: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`.
|
||
|
||
| # | phase | verdict | one sentence |
|
||
|---|---|---|---|
|
||
| 1 | snapshot reachable, and its name | **NO — and it has no reachable name** | 777,600 exact names across nine days in the vendor's own format, plus 126 alternative shapes; zero resolve, with a control proving the sweep detects a path that exists. |
|
||
| 2 | the deletion | **NOT RUN — operator ruling** | With no recovery route, the deletion would have destroyed real history to buy nothing; put as a two-option decision, the ruling was stop. |
|
||
| 3 | the alarm fired | **NOT RUN — and it could not have fired at the specified size** | The shipped threshold needs a fall of more than half; one app's tag is ~9 of 69. → **R-435** |
|
||
| 4 | **the recovery** | **NOT RUN — no route exists that is not fenced** | Box-side: proven impossible. Panel: fenced and browserless. Hetzner API: fenced (§11-D), and the hub's client has no snapshot method at all. |
|
||
| — | **RTO from T₀** | **STILL BLANK** | Row 10's RTO cell is unchanged and remains a finding. |
|
||
| — | **data lost, quantified** | **NOT MEASURABLE THIS WAY** | The quantity only has meaning if the rest is recoverable, and the route that would recover it is not reachable. |
|
||
| 5 | re-arm | **NOT RUN** | Depended on Phase 4. |
|
||
| 6 | teardown | **PASS** | Store untouched at 69 snapshots; all three scratch layers removed; hub DB copies shredded; both boxes healthy. |
|
||
|
||
---
|
||
|
||
## 1. Did the recovery work — and does yesterday's re-scope survive?
|
||
|
||
**The recovery was never reachable, and the re-scope does not survive intact. Its first half stands;
|
||
its second half does not.**
|
||
|
||
Yesterday's re-scope has two clauses. They must now be separated:
|
||
|
||
* **(a) "The box can delete its live repository, but cannot write to the daily snapshots of it."**
|
||
**STANDS.** Re-confirmed here: `/.zfs/snapshot` is reachable and the write-refusal measurement is
|
||
unchanged. Nothing in this drill weakens it.
|
||
* **(b) "…so the rest is recoverable — file by file, one customer at a time."** **NOT SUPPORTED.**
|
||
A snapshot that cannot be opened cannot be copied out of. R-432 recorded the directory listing
|
||
empty and named the cheapest next step: *"a single `ls /.zfs/snapshot/<name>` from a box then
|
||
settles whether a named snapshot can be entered even though the directory does not list (ZFS
|
||
allows exactly that)."* **That step is now done, exhaustively, and the answer is no.**
|
||
|
||
**What was measured.** The port-23 restricted shell accepts a batched `stat`, which makes a cheap
|
||
existence oracle: 500–600 paths per round trip, stdout carrying only paths that exist.
|
||
|
||
| sweep | candidates | hits |
|
||
|---|---|---|
|
||
| `/.zfs/snapshot/YYYY-MM-DDTHH-MM-SS`, nine full days, second granularity | **777,600** | **0** |
|
||
| 126 alternative name shapes and snapshot paths (`daily`, `snapshot-1`, colon and compact time forms, `/home/.snapshot`, …) | 126 | 0 |
|
||
| **control — the identical 600-name batch shape with one real path appended** | 6 batches | **6/6 returned it** |
|
||
|
||
**And there is a structural reason, which is why I stopped sweeping.** The customer's data and the
|
||
snapshot door are on **different filesystems**:
|
||
|
||
```
|
||
df → u629488-sub3 mounted on /home
|
||
stat /home → Device 0,82
|
||
stat /.zfs/snapshot → Device 0,276 ← a different device
|
||
stat /home/.zfs → cannot statx: No such file or directory
|
||
```
|
||
|
||
A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset that owns that `.zfs` — not to the
|
||
child mounted at `/home`. **So even a correctly named snapshot there could not contain
|
||
`felhom-repo`,** and the dataset that does hold it exposes no `.zfs` at all to this account. The
|
||
empty listing is not a display toggle hiding a reachable tree; from a sub-account there is no tree.
|
||
|
||
**Three tools agree, each with controls in the same run:** SFTP, the port-23 shell, and
|
||
`rsync --list-only`.
|
||
|
||
**What that does to R-95.** Its *exposure* is unchanged and its *remedy* is not. Yesterday the row
|
||
could say a deletion costs about a day because the rest comes back per-file. Today the only routes
|
||
to "the rest" are a whole-box panel rollback (which deletes newer snapshots and hits every customer
|
||
on the box) and the provider API (fenced, and unimplemented in the hub's client). **The re-scope's
|
||
comfort was resting on a route nobody had walked — which is precisely the standard this project
|
||
applies, and it is the reason this drill was called.**
|
||
|
||
**The ranking is Viktor's and I am not re-ranking it.** What I will say plainly: the argument that
|
||
moved R-95 down yesterday is the argument this drill removed. On these facts I would put it back
|
||
where it was.
|
||
|
||
## 2. The RTO
|
||
|
||
**Still blank, and it stays a finding.** `07` §8 row 10's RTO cell has been empty since July and this
|
||
drill did not fill it. Nothing was recovered, so nothing was timed. The summary line at
|
||
`07-backup-architecture.md:948` — *"no ransomware-shaped recovery has ever been run"* — is still
|
||
true, and is now true for a sharper reason: **not "nobody has run it" but "from the box, it cannot
|
||
be run."**
|
||
|
||
## 3. R-432's answer, and the naming scheme
|
||
|
||
**R-432 is ANSWERED, and negatively. It did not need the panel read it was waiting on.**
|
||
|
||
* **The naming scheme is `YYYY-MM-DDTHH-MM-SS`** — vendor-documented examples `2025-12-03T13-47-47`,
|
||
`2025-02-12T11-35-19`. Recorded so nobody hunts a console again.
|
||
* **Knowing it does not help.** Every name in that format for nine days is refused, and the st_dev
|
||
split above says why. **Per-file recovery is not operator-only — from the box it is nobody's,** and
|
||
for the operator it is a browser act against the main account that no credential in this project
|
||
can perform.
|
||
* **The panel cannot supply the missing piece either.** It offers Restore and Delete on a row and
|
||
does not show names; and the one name-shaped thing it could give would be tried against a door
|
||
that leads to the wrong dataset.
|
||
|
||
## 4. The alarm's first real firing
|
||
|
||
**It did not happen, and the drill as written could not have produced it.** The detector fires on a
|
||
fall of **more than half** the previous count **and at least 5** (`hub/internal/monitor/offsite.go`,
|
||
`snapshotDropFraction = 0.5`, `snapshotDropFloor = 5`). demo-hp's baseline is **69**. Phase 2 deletes
|
||
**one app's** history — about **9** snapshots. 9 is over the floor and nowhere near half, so the
|
||
alarm stays silent, **correctly and by design**. Firing it for real needs ~35+ snapshots destroyed,
|
||
i.e. most of demo-hp's off-site history. That trade is what the operator was asked to rule on. → **R-435**
|
||
|
||
**One thing the alarm says is now wrong.** Its message, live in hub 0.111.0, reads:
|
||
|
||
> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is
|
||
> recoverable file-by-file**; it is NOT confirmed data loss."
|
||
|
||
The first clause is true; **the second promises a recovery the product cannot perform and the
|
||
operator cannot perform without a browser and the main account.** This is this project's own
|
||
corollary — *when a verdict changes which field it counts from, the alarm text has to change with
|
||
it* — landing on the alarm shipped the same day. → **R-434**
|
||
|
||
## 5. Findings, as register rows
|
||
|
||
All four filed in `documentation/backlog/OPEN-ITEMS.md`.
|
||
|
||
| row | finding |
|
||
|---|---|
|
||
| **R-433** | A sub-account cannot reach any Storage Box snapshot **by any name**; `/home` and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Answers R-432 negatively and removes clause (b) of the R-95 re-scope. |
|
||
| **R-434** | `emitSnapshotDrop`'s message promises file-by-file recovery that is not reachable. Live in hub 0.111.0. |
|
||
| **R-435** | The drop detector cannot see a single-app deletion (>50% of 69 ⇒ ~35 needed). `offbox.go:1388` forgets **by tag**, so a single-tag wipe is exactly the shape the detector is blind to. Deliberate insensitivity, but the blind spot should be stated where the operator reads it. |
|
||
| **R-436** | **LEAD, not a defect.** Hetzner's port-23 shell offers `rclone serve restic --stdio` as a server-side backend, and restic 0.14.0 recognises the `rclone:` backend (measured; control `banana:` → invalid backend; rclone is absent from the controller image). `rclone serve restic` carries `--append-only`. **This could make R-95's real prevention far cheaper than the spike concluded — no new always-on machine, no data migration.** Caveat stated up front: the **client** supplies the server command line, so a compromised box could omit the flag unless the provider pins it. Settling that is a vendor question, not a code change. |
|
||
|
||
**R-432 is marked ANSWERED**; its "one panel read settles it" next step is withdrawn as unnecessary.
|
||
|
||
## 6. Does `07` §8 row 10 move?
|
||
|
||
**No. It stays `PARTIAL`, and its RTO stays blank.** The status was already correct for the right
|
||
reason — *"the recovery ROUTE has never been walked, which is what PARTIAL means"* — and this drill
|
||
found the route is not walkable from the box at all. **What the row needs is a text correction, not a
|
||
status change:** its clause *"recoverable per-file (vendor)"* and its limit *"per-file recovery is
|
||
operator-only today (R-432)"* both overstate what exists. Updated in place with the citation. Moving
|
||
it only as far as the evidence goes means not moving it.
|
||
|
||
## 7. What could not be tested, and why
|
||
|
||
* **Whether the main account can see the snapshots.** No main-account credential exists in this
|
||
project — the hub holds only per-customer sub-accounts. This is the one question that would decide
|
||
whether per-file recovery exists *at all*, for anyone.
|
||
* **Whether the Hetzner API can list or read a snapshot.** Fenced by the runbook (§11-D). Separately,
|
||
`hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method** — so this route needs new code
|
||
regardless of the fence.
|
||
* **The deletion, the alarm's first real firing, the recovery, the RTO, the re-arm.** Phases 2–5, not
|
||
run, on the operator's ruling.
|
||
* **Whether `rclone serve restic --stdio` is pinned server-side with `--append-only`** (R-436).
|
||
|
||
## 8. My own mistakes
|
||
|
||
* **I ran Phase 1 before finishing Phase 0's subject-app choice, and did not say so up front.** The
|
||
choice depends on the snapshot's timestamp, so the order was right, but the runbook's order is the
|
||
runbook's and a silent reordering is the thing this project keeps getting caught by. Stated at the
|
||
time in the session, recorded here.
|
||
* **My first sweep guessed the schedule instead of establishing it.** I probed 00:00 UTC and 22:00
|
||
UTC — 600 names — on the strength of a register line reading *"daily 00:00"*, got nothing, and only
|
||
then widened to whole days. The narrow sweep was worth nothing on its own: a zero over a guessed
|
||
window is not evidence, and I should have gone to full days first or not run it at all.
|
||
* **I nearly reported the empty listing as "the display toggle is hiding it".** The vendor documents
|
||
exactly such a toggle and it fitted. The st_dev comparison — which I only ran because `df` printed
|
||
a filesystem name I did not expect — says the tree is on another dataset entirely. **A plausible
|
||
cause that fits the symptom is not a measured one**, and I had the wrong one for about ten minutes.
|
||
* **`REPORT.md` held the only copy of the R-331 report** (hub v0.109.0, 2026-08-30) — durable content
|
||
living only in the overwritten file, which `CLAUDE.md:82-87` forbids. Preserved as
|
||
`REPORT-r331-backup-card.md` before this report replaced it.
|