hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.
was: "...still hold the older copy, so this is recoverable file-by-file; it is NOT
confirmed data loss. Check whether a deletion ran on the box before restoring."
now: "...still hold the older copy. The route back out of them is not yet established,
so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
before restoring anything, and check whether a deletion ran on the box."
It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.
Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.
R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.
THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.
Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.
R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.
Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
@@ -0,0 +1,166 @@
|
||||
# REPORT — the deletion we said is survivable: the recovery route does not exist (R-95 drill, 2026-09-01)
|
||||
|
||||
**RUNBOOK, destructive class, `demo-hp` only. STOPPED at the end of Phase 1 on the operator's ruling,
|
||||
before any destructive step. No delete verb was issued against any live store; no byte on either
|
||||
Storage Box sub-account was written, moved or removed.** No production code, no version bump, no
|
||||
image, no golden. Evidence: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`.
|
||||
|
||||
| # | phase | verdict | one sentence |
|
||||
|---|---|---|---|
|
||||
| 1 | snapshot reachable, and its name | **NO — and it has no reachable name** | 777,600 exact names across nine days in the vendor's own format, plus 126 alternative shapes; zero resolve, with a control proving the sweep detects a path that exists. |
|
||||
| 2 | the deletion | **NOT RUN — operator ruling** | With no recovery route, the deletion would have destroyed real history to buy nothing; put as a two-option decision, the ruling was stop. |
|
||||
| 3 | the alarm fired | **NOT RUN — and it could not have fired at the specified size** | The shipped threshold needs a fall of more than half; one app's tag is ~9 of 69. → **R-435** |
|
||||
| 4 | **the recovery** | **NOT RUN — no route exists that is not fenced** | Box-side: proven impossible. Panel: fenced and browserless. Hetzner API: fenced (§11-D), and the hub's client has no snapshot method at all. |
|
||||
| — | **RTO from T₀** | **STILL BLANK** | Row 10's RTO cell is unchanged and remains a finding. |
|
||||
| — | **data lost, quantified** | **NOT MEASURABLE THIS WAY** | The quantity only has meaning if the rest is recoverable, and the route that would recover it is not reachable. |
|
||||
| 5 | re-arm | **NOT RUN** | Depended on Phase 4. |
|
||||
| 6 | teardown | **PASS** | Store untouched at 69 snapshots; all three scratch layers removed; hub DB copies shredded; both boxes healthy. |
|
||||
|
||||
---
|
||||
|
||||
## 1. Did the recovery work — and does yesterday's re-scope survive?
|
||||
|
||||
**The recovery was never reachable, and the re-scope does not survive intact. Its first half stands;
|
||||
its second half does not.**
|
||||
|
||||
Yesterday's re-scope has two clauses. They must now be separated:
|
||||
|
||||
* **(a) "The box can delete its live repository, but cannot write to the daily snapshots of it."**
|
||||
**STANDS.** Re-confirmed here: `/.zfs/snapshot` is reachable and the write-refusal measurement is
|
||||
unchanged. Nothing in this drill weakens it.
|
||||
* **(b) "…so the rest is recoverable — file by file, one customer at a time."** **NOT SUPPORTED.**
|
||||
A snapshot that cannot be opened cannot be copied out of. R-432 recorded the directory listing
|
||||
empty and named the cheapest next step: *"a single `ls /.zfs/snapshot/<name>` from a box then
|
||||
settles whether a named snapshot can be entered even though the directory does not list (ZFS
|
||||
allows exactly that)."* **That step is now done, exhaustively, and the answer is no.**
|
||||
|
||||
**What was measured.** The port-23 restricted shell accepts a batched `stat`, which makes a cheap
|
||||
existence oracle: 500–600 paths per round trip, stdout carrying only paths that exist.
|
||||
|
||||
| sweep | candidates | hits |
|
||||
|---|---|---|
|
||||
| `/.zfs/snapshot/YYYY-MM-DDTHH-MM-SS`, nine full days, second granularity | **777,600** | **0** |
|
||||
| 126 alternative name shapes and snapshot paths (`daily`, `snapshot-1`, colon and compact time forms, `/home/.snapshot`, …) | 126 | 0 |
|
||||
| **control — the identical 600-name batch shape with one real path appended** | 6 batches | **6/6 returned it** |
|
||||
|
||||
**And there is a structural reason, which is why I stopped sweeping.** The customer's data and the
|
||||
snapshot door are on **different filesystems**:
|
||||
|
||||
```
|
||||
df → u629488-sub3 mounted on /home
|
||||
stat /home → Device 0,82
|
||||
stat /.zfs/snapshot → Device 0,276 ← a different device
|
||||
stat /home/.zfs → cannot statx: No such file or directory
|
||||
```
|
||||
|
||||
A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset that owns that `.zfs` — not to the
|
||||
child mounted at `/home`. **So even a correctly named snapshot there could not contain
|
||||
`felhom-repo`,** and the dataset that does hold it exposes no `.zfs` at all to this account. The
|
||||
empty listing is not a display toggle hiding a reachable tree; from a sub-account there is no tree.
|
||||
|
||||
**Three tools agree, each with controls in the same run:** SFTP, the port-23 shell, and
|
||||
`rsync --list-only`.
|
||||
|
||||
**What that does to R-95.** Its *exposure* is unchanged and its *remedy* is not. Yesterday the row
|
||||
could say a deletion costs about a day because the rest comes back per-file. Today the only routes
|
||||
to "the rest" are a whole-box panel rollback (which deletes newer snapshots and hits every customer
|
||||
on the box) and the provider API (fenced, and unimplemented in the hub's client). **The re-scope's
|
||||
comfort was resting on a route nobody had walked — which is precisely the standard this project
|
||||
applies, and it is the reason this drill was called.**
|
||||
|
||||
**The ranking is Viktor's and I am not re-ranking it.** What I will say plainly: the argument that
|
||||
moved R-95 down yesterday is the argument this drill removed. On these facts I would put it back
|
||||
where it was.
|
||||
|
||||
## 2. The RTO
|
||||
|
||||
**Still blank, and it stays a finding.** `07` §8 row 10's RTO cell has been empty since July and this
|
||||
drill did not fill it. Nothing was recovered, so nothing was timed. The summary line at
|
||||
`07-backup-architecture.md:948` — *"no ransomware-shaped recovery has ever been run"* — is still
|
||||
true, and is now true for a sharper reason: **not "nobody has run it" but "from the box, it cannot
|
||||
be run."**
|
||||
|
||||
## 3. R-432's answer, and the naming scheme
|
||||
|
||||
**R-432 is ANSWERED, and negatively. It did not need the panel read it was waiting on.**
|
||||
|
||||
* **The naming scheme is `YYYY-MM-DDTHH-MM-SS`** — vendor-documented examples `2025-12-03T13-47-47`,
|
||||
`2025-02-12T11-35-19`. Recorded so nobody hunts a console again.
|
||||
* **Knowing it does not help.** Every name in that format for nine days is refused, and the st_dev
|
||||
split above says why. **Per-file recovery is not operator-only — from the box it is nobody's,** and
|
||||
for the operator it is a browser act against the main account that no credential in this project
|
||||
can perform.
|
||||
* **The panel cannot supply the missing piece either.** It offers Restore and Delete on a row and
|
||||
does not show names; and the one name-shaped thing it could give would be tried against a door
|
||||
that leads to the wrong dataset.
|
||||
|
||||
## 4. The alarm's first real firing
|
||||
|
||||
**It did not happen, and the drill as written could not have produced it.** The detector fires on a
|
||||
fall of **more than half** the previous count **and at least 5** (`hub/internal/monitor/offsite.go`,
|
||||
`snapshotDropFraction = 0.5`, `snapshotDropFloor = 5`). demo-hp's baseline is **69**. Phase 2 deletes
|
||||
**one app's** history — about **9** snapshots. 9 is over the floor and nowhere near half, so the
|
||||
alarm stays silent, **correctly and by design**. Firing it for real needs ~35+ snapshots destroyed,
|
||||
i.e. most of demo-hp's off-site history. That trade is what the operator was asked to rule on. → **R-435**
|
||||
|
||||
**One thing the alarm says is now wrong.** Its message, live in hub 0.111.0, reads:
|
||||
|
||||
> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is
|
||||
> recoverable file-by-file**; it is NOT confirmed data loss."
|
||||
|
||||
The first clause is true; **the second promises a recovery the product cannot perform and the
|
||||
operator cannot perform without a browser and the main account.** This is this project's own
|
||||
corollary — *when a verdict changes which field it counts from, the alarm text has to change with
|
||||
it* — landing on the alarm shipped the same day. → **R-434**
|
||||
|
||||
## 5. Findings, as register rows
|
||||
|
||||
All four filed in `documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
| row | finding |
|
||||
|---|---|
|
||||
| **R-433** | A sub-account cannot reach any Storage Box snapshot **by any name**; `/home` and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Answers R-432 negatively and removes clause (b) of the R-95 re-scope. |
|
||||
| **R-434** | `emitSnapshotDrop`'s message promises file-by-file recovery that is not reachable. Live in hub 0.111.0. |
|
||||
| **R-435** | The drop detector cannot see a single-app deletion (>50% of 69 ⇒ ~35 needed). `offbox.go:1388` forgets **by tag**, so a single-tag wipe is exactly the shape the detector is blind to. Deliberate insensitivity, but the blind spot should be stated where the operator reads it. |
|
||||
| **R-436** | **LEAD, not a defect.** Hetzner's port-23 shell offers `rclone serve restic --stdio` as a server-side backend, and restic 0.14.0 recognises the `rclone:` backend (measured; control `banana:` → invalid backend; rclone is absent from the controller image). `rclone serve restic` carries `--append-only`. **This could make R-95's real prevention far cheaper than the spike concluded — no new always-on machine, no data migration.** Caveat stated up front: the **client** supplies the server command line, so a compromised box could omit the flag unless the provider pins it. Settling that is a vendor question, not a code change. |
|
||||
|
||||
**R-432 is marked ANSWERED**; its "one panel read settles it" next step is withdrawn as unnecessary.
|
||||
|
||||
## 6. Does `07` §8 row 10 move?
|
||||
|
||||
**No. It stays `PARTIAL`, and its RTO stays blank.** The status was already correct for the right
|
||||
reason — *"the recovery ROUTE has never been walked, which is what PARTIAL means"* — and this drill
|
||||
found the route is not walkable from the box at all. **What the row needs is a text correction, not a
|
||||
status change:** its clause *"recoverable per-file (vendor)"* and its limit *"per-file recovery is
|
||||
operator-only today (R-432)"* both overstate what exists. Updated in place with the citation. Moving
|
||||
it only as far as the evidence goes means not moving it.
|
||||
|
||||
## 7. What could not be tested, and why
|
||||
|
||||
* **Whether the main account can see the snapshots.** No main-account credential exists in this
|
||||
project — the hub holds only per-customer sub-accounts. This is the one question that would decide
|
||||
whether per-file recovery exists *at all*, for anyone.
|
||||
* **Whether the Hetzner API can list or read a snapshot.** Fenced by the runbook (§11-D). Separately,
|
||||
`hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method** — so this route needs new code
|
||||
regardless of the fence.
|
||||
* **The deletion, the alarm's first real firing, the recovery, the RTO, the re-arm.** Phases 2–5, not
|
||||
run, on the operator's ruling.
|
||||
* **Whether `rclone serve restic --stdio` is pinned server-side with `--append-only`** (R-436).
|
||||
|
||||
## 8. My own mistakes
|
||||
|
||||
* **I ran Phase 1 before finishing Phase 0's subject-app choice, and did not say so up front.** The
|
||||
choice depends on the snapshot's timestamp, so the order was right, but the runbook's order is the
|
||||
runbook's and a silent reordering is the thing this project keeps getting caught by. Stated at the
|
||||
time in the session, recorded here.
|
||||
* **My first sweep guessed the schedule instead of establishing it.** I probed 00:00 UTC and 22:00
|
||||
UTC — 600 names — on the strength of a register line reading *"daily 00:00"*, got nothing, and only
|
||||
then widened to whole days. The narrow sweep was worth nothing on its own: a zero over a guessed
|
||||
window is not evidence, and I should have gone to full days first or not run it at all.
|
||||
* **I nearly reported the empty listing as "the display toggle is hiding it".** The vendor documents
|
||||
exactly such a toggle and it fitted. The st_dev comparison — which I only ran because `df` printed
|
||||
a filesystem name I did not expect — says the tree is on another dataset entirely. **A plausible
|
||||
cause that fits the symptom is not a measured one**, and I had the wrong one for about ten minutes.
|
||||
* **`REPORT.md` held the only copy of the R-331 report** (hub v0.109.0, 2026-08-30) — durable content
|
||||
living only in the overwritten file, which `CLAUDE.md:82-87` forbids. Preserved as
|
||||
`REPORT-r331-backup-card.md` before this report replaced it.
|
||||
@@ -1,8 +1,8 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-01 (second pass) — both faults from the overnight test are fixed, and the golden
|
||||
is baked, vouched and delivered. Both machines are on 0.232.0 and moved themselves. NOTHING is
|
||||
waiting on you.**
|
||||
**Updated 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down
|
||||
where it stops. One alarm that was telling you something untrue is fixed and live (hub 0.111.1).
|
||||
ONE THING is waiting on you: two short e-mails to Hetzner, drafted and ready to paste.**
|
||||
|
||||
**Earlier 2026-08-31 — I measured whether the box could test its own remote restore without you,
|
||||
then built the narrow version you picked. The measurement is why it is 3 seconds a night and
|
||||
@@ -18,7 +18,7 @@ not an evening's work.**
|
||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||
nothing.*
|
||||
|
||||
1. **Two things are waiting on you — item 4 (a snapshot name, two minutes) and item 5 (ignore one alarm mail).** Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
1. **One thing is waiting on you — item 4 (send two e-mails).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
the real machines:
|
||||
- the background job that could delete a live restore's lock now waits its turn — and the check
|
||||
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
|
||||
@@ -37,24 +37,34 @@ nothing.*
|
||||
two register lines in the hub (already live). No customer action, no data migration, no
|
||||
credential change.
|
||||
|
||||
4. **Good news, and I was wrong yesterday.** I told you I could not find the nightly safety net on
|
||||
the storage. **I was looking for the wrong name.** The snapshots are there, exactly as you saw in
|
||||
the control panel: seven of them, one a day. **And they are better than we thought** — I tried to
|
||||
write into the snapshot area from a customer's machine and the storage **refused**, while the same
|
||||
write to its normal folder worked. So a machine that wipes its own backup **cannot touch the
|
||||
snapshots of it**. The worst case is losing about a day, then copying the rest back file by file.
|
||||
**That is much smaller than what the notes have said since July.** I have corrected the notes.
|
||||
**What I still need from you:** the machines can see the snapshot *door* but not what is inside —
|
||||
only the main account can. So getting data back is you, in a browser, for now. **If you read one
|
||||
snapshot's name off the panel and send it to me, one command settles whether the machines can
|
||||
reach them directly** — and if they can, recovery becomes something the product does by itself.
|
||||
**If you do nothing:** it stays a manual job for you, which is workable but slow.
|
||||
4. **Please send two short e-mails to Hetzner. They are written for you.**
|
||||
`felhom.eu/documentation/runbooks/provider-questions-2026-09-01.md` — open it, copy, send. No
|
||||
password or key is in that file, and none should be added.
|
||||
|
||||
**Why.** Yesterday I told you the snapshots make a wiped backup survivable: lose about a day, copy
|
||||
the rest back file by file. **The first half is still true. The second half is not, and I found
|
||||
that out by trying it.** I tried **777,600** snapshot names on the storage, over nine days, in
|
||||
Hetzner's own naming style. **None of them opened.** Then I found why: your data and the snapshot
|
||||
door sit on **two different drives** inside the storage, and the door for your data **does not
|
||||
exist at all**. So there is no way in from the machines.
|
||||
|
||||
**What is still true, and it matters:** a machine that wipes its own backup **still cannot touch
|
||||
the snapshots of it**. The older copy is there. What we do not have is a way to reach it.
|
||||
|
||||
**The two questions.** One: can the **main** account pull single files out of a snapshot? Two: on
|
||||
one of Hetzner's own tools, is a "cannot delete" switch forced by them, or chosen by the machine?
|
||||
**The second one could remove the whole problem** — no new hardware, no moving anyone's data.
|
||||
|
||||
**If you do nothing:** we cannot finish this. The backups keep working and keep being checked; we
|
||||
simply cannot say what a wiped backup costs, and I would then put this risk back near the top of
|
||||
your list. **My pick: send both. It is five minutes and it decides an evening's work.**
|
||||
|
||||
5. **You will have received an alarm email from me today about `demo-hp` losing 65 backups. It is a
|
||||
test and nothing is wrong.** I built the new "someone deleted the backups" alarm and had to fire it
|
||||
once for real to prove it reaches you. Subject: `[Felhom] 🔴 demo-hp: offsite_snapshots_dropped`.
|
||||
**No backups were deleted.** **If you do nothing:** nothing — but please do not act on that one
|
||||
mail.
|
||||
mail. **Closed now:** that was the only such mail, it was a test, and the alarm's wording has since
|
||||
been corrected (item under *Decided* below). Nothing further is needed from you here.
|
||||
|
||||
6. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
@@ -66,6 +76,26 @@ nothing.*
|
||||
|
||||
## Decided — and what would reopen each
|
||||
|
||||
- **THE BACKUP AND RESTORE WORK IS FINISHED FOR BETA. DECIDED 2026-09-01.**
|
||||
**What is done, and proven on the real machines:** everything **you or a customer** does alone —
|
||||
getting deleted files back, getting an app's data back, getting a whole app back, and losing a
|
||||
drive. The restore tells you what it put back, refuses if there is no room, will not accept a
|
||||
half-copy, and puts your own data back if it fails.
|
||||
**What is parked until after beta:** everything **only I do, with you** — rebuilding a machine as
|
||||
itself, losing a whole box, recovering from ransomware, restoring the hub, and losing Hetzner.
|
||||
**Six of these have never been timed, and the hub has never been restored.** They are written down,
|
||||
they are real, and **none of them stops a beta customer.**
|
||||
**Reopens if:** something a customer does for themselves turns out to be broken; **or** Hetzner's
|
||||
two answers change what the snapshots are worth; **or** a real customer's data is at stake in one
|
||||
of the parked items.
|
||||
|
||||
- **The alarm that promised too much: FIXED and live (hub 0.111.1). DECIDED 2026-09-01.**
|
||||
Yesterday's alarm mail said a deleted backup was *"recoverable file-by-file"*. **We now know it is
|
||||
not.** I removed the promise rather than writing a new one, so the sentence stays true whatever
|
||||
Hetzner answers. It no longer says the data is lost either — that is still usually untrue.
|
||||
**Reopens if:** Hetzner's answers give us a real route back; then the alarm can name it.
|
||||
|
||||
|
||||
- **The missing-golden warning is now pointed at the people who can act on it (R-404). DECIDED
|
||||
2026-09-01, and built the same day.** The warning was aimed at the wrong repository: the one where
|
||||
a release actually happens never checked at all, while the one that only holds documents was
|
||||
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,101 @@
|
||||
# Two questions for Hetzner — drafted, ready to send (2026-09-01)
|
||||
|
||||
**These are the only thing standing between us and finishing R-95 properly.** They cost nothing and
|
||||
they are answerable by a support agent without escalation.
|
||||
|
||||
**Send them yourself.** CC drafted them and did not send them, and did not call the provider API —
|
||||
`§11-D` is still the operator's fence. **No credential, password or token appears below, and none
|
||||
should be added.** The account id and the product name are all either question needs.
|
||||
|
||||
**What to fill in:** the ticket needs the Storage Box account. Ours is **`u629488`** (the box the
|
||||
register calls `storage-box-pool-1`, plan BX11). Nothing else.
|
||||
|
||||
**Why two separate tickets:** they go to different parts of the answer — one is about the snapshot
|
||||
product, one is about the SSH endpoint's configuration — and a single ticket asking both tends to get
|
||||
one answered and the other dropped.
|
||||
|
||||
---
|
||||
|
||||
## Question 1 — can the MAIN account retrieve individual files from a snapshot?
|
||||
|
||||
**Why it matters, in one line:** if it cannot, the only route back is a whole-box rollback that hits
|
||||
every customer on the box and destroys every newer snapshot — which would mean the snapshots protect
|
||||
almost nobody in practice. **This is the question that decides how urgent R-95 is.**
|
||||
|
||||
> **Subject:** Storage Box u629488 — retrieving individual files from a snapshot
|
||||
>
|
||||
> Hello,
|
||||
>
|
||||
> We use Storage Box `u629488` with sub-accounts, and daily automatic snapshots are enabled.
|
||||
>
|
||||
> We can reach `/.zfs/snapshot` from a sub-account, but it lists as empty, and no snapshot name we
|
||||
> try can be entered. We understand sub-accounts may be restricted here.
|
||||
>
|
||||
> Our question is about the **main account**: from the main account, over SSH or SFTP on port 23,
|
||||
> can we **read or download individual files and directories out of a specific snapshot** — for
|
||||
> example a single directory under one sub-account's home — **without** performing a snapshot
|
||||
> restore of the whole Storage Box?
|
||||
>
|
||||
> If yes, please tell us the exact path we should use and how the snapshot directory is named.
|
||||
>
|
||||
> If no, please confirm that the only way to get data out of a snapshot is the full "restore
|
||||
> snapshot" action on the whole Storage Box.
|
||||
>
|
||||
> Thank you.
|
||||
|
||||
**How to read the answer.**
|
||||
* **"Yes, from the main account"** → per-file recovery exists, but it is an operator act in a
|
||||
browser or over the main account's own SSH, and it can never be something the product does for the
|
||||
customer. R-433 closes at that. R-95 stays where it is.
|
||||
* **"No, only a full restore"** → the snapshots do **not** bound a single customer's exposure at all,
|
||||
because using them costs every other customer on the box their newer snapshots. **R-95 becomes
|
||||
urgent and the transport change stops being optional.**
|
||||
|
||||
---
|
||||
|
||||
## Question 2 — is `--append-only` enforced on the `rclone serve restic` endpoint?
|
||||
|
||||
**Why it matters, in one line:** if it is enforced server-side, a compromised box **cannot delete its
|
||||
own backups**, with no new machine and no data migration. **This is the question that could make R-95
|
||||
disappear.**
|
||||
|
||||
> **Subject:** Storage Box u629488 — rclone serve restic endpoint and --append-only
|
||||
>
|
||||
> Hello,
|
||||
>
|
||||
> The restricted SSH shell on Storage Box `u629488` (port 23) lists `rclone serve restic --stdio`
|
||||
> among the available server-side backends.
|
||||
>
|
||||
> Our question is about how that command is run on your side: **is `--append-only` enforced by you,
|
||||
> or is the command line taken from what the client sends?**
|
||||
>
|
||||
> In other words, if a client connects and asks for `rclone serve restic --stdio` **without**
|
||||
> `--append-only`, does it get a server that permits deletions?
|
||||
>
|
||||
> If the flag can be enforced, is there any way for us to request that for this account or for
|
||||
> individual sub-accounts?
|
||||
>
|
||||
> Thank you.
|
||||
|
||||
**How to read the answer.**
|
||||
* **"Enforced server-side" or "can be enabled per account"** → this is the cheap prevention the
|
||||
2026-09-01 spike priced at a new always-on service plus either a mount in the hot path or migrating
|
||||
every customer's history. **It needs none of that** — the server already runs at the provider, and
|
||||
restic 0.14.0 already speaks the `rclone:` backend (measured, with a control:
|
||||
`banana:` → `invalid backend`, `rclone:` → the helper was executed). The remaining work is putting
|
||||
`rclone` in the controller image and switching the repository URL. **R-95's root cause goes away.**
|
||||
* **"The client supplies the command line"** → **the lead is worth nothing** and should be recorded
|
||||
as dead, not left looking promising. Prevention then still needs a machine in front of the store,
|
||||
and the decision reverts to the spike's option 3 at its original price.
|
||||
|
||||
---
|
||||
|
||||
## Where these came from
|
||||
|
||||
* **R-433** — no snapshot is reachable from a sub-account by any name. 777,600 exact names in the
|
||||
vendor's `YYYY-MM-DDTHH-MM-SS` format over nine days, zero hits, with a passing control; `/home`
|
||||
and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Question 1 exists because
|
||||
that measurement can only speak for a sub-account.
|
||||
* **R-436** — the `rclone serve restic --stdio` backend and restic's `rclone:` support, both
|
||||
measured. Question 2 is the one caveat that decides whether the lead is real.
|
||||
* Evidence for both: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`.
|
||||
@@ -1,3 +1,70 @@
|
||||
## v0.111.1 — the alarm stops promising a rescue that does not exist (2026-09-01, R-434)
|
||||
|
||||
**Text-only patch on a live alarm. No controller change, no golden owed, no floor change.**
|
||||
|
||||
**WHAT WAS WRONG.** v0.111.0 shipped, that morning, a snapshot-drop alarm reading:
|
||||
|
||||
> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is
|
||||
> recoverable file-by-file**; it is NOT confirmed data loss."
|
||||
|
||||
**Measured the same afternoon (R-433): no snapshot is reachable from a customer's sub-account by ANY
|
||||
name.** 777,600 exact names in the vendor's own `YYYY-MM-DDTHH-MM-SS` format across nine full days at
|
||||
second granularity, plus 126 alternative shapes — zero hits, with a control proving the identical
|
||||
batch returns a path that does exist (6/6). The structural reason: `/home` (the customer data) and
|
||||
`/.zfs` are **different filesystems**, and `/home/.zfs` does not exist, so a snapshot behind that door
|
||||
could not hold the repository even with the right name. **The promise named a route nobody can walk**,
|
||||
in the one message an operator acts on while a customer's off-site history is disappearing.
|
||||
|
||||
**THE FIX IS A DELETION, NOT A REPLACEMENT, AND THAT IS THE POINT.** R-434's register row said the fix
|
||||
was blocked on R-433 — on first establishing what IS true. **It was not blocked, once the promise is
|
||||
withdrawn rather than swapped.** The shipped sentence now asserts neither loss nor recovery:
|
||||
|
||||
> "The daily Storage Box snapshots are read-only and still hold the older copy. The route back out of
|
||||
> them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery.
|
||||
> Get in touch before restoring anything, and check whether a deletion ran on the box."
|
||||
|
||||
**That is true under every possible answer to the outstanding provider question, so it never needs a
|
||||
second rewrite.** An alarm rewritten twice in a week is worse than one rewritten once: the operator
|
||||
learns its words do not mean anything.
|
||||
|
||||
**AND IT MUST NOT SWING THE OTHER WAY.** "Your backups are gone" is still usually FALSE — the
|
||||
snapshots exist and hold the older copy; what is unproven is our route to them. Clause (a) of the
|
||||
2026-09-01 measurement (the box **cannot write** into the snapshot area) stands and is re-confirmed.
|
||||
Over-claiming loss would send an operator into a destructive recovery they did not need, which is the
|
||||
failure v0.111.0's own comment was written to prevent — it had to survive its own correction.
|
||||
|
||||
**TESTS — `internal/monitor/offsite_r434_test.go`, three, all driving the production path**
|
||||
(`saveOffsiteReport` → `oc.Check()` → the notify callback), so they assert the sentence an operator
|
||||
RECEIVES rather than the function that formats it. ASCII-only fragments, with a positive control (a
|
||||
phrase present in every version of the alarm) and a negative control (a phrase that cannot exist).
|
||||
|
||||
- `TestR434_AlarmMakesNoRecoveryPromise` — the withdrawal is present, the promise is absent, in three
|
||||
shapes it could plausibly return as.
|
||||
- `TestR434_AlarmStillDoesNotClaimDataLoss` — the other direction; `confirmed data loss` may appear
|
||||
only inside its negation.
|
||||
- `TestR434_StoredEventCarriesTheSameSentence` — the stored row and the delivered mail must not drift.
|
||||
|
||||
**RED-PROOF (run 2026-09-01, before the fix was restored):** the v0.111.0 sentence was put back and
|
||||
all three FAILED — on `"recoverable file-by-file"`, on `"so this is recoverable"`, on all three
|
||||
required fragments, on the un-negated `confirmed data loss`, and on the stored row — each with the
|
||||
offending sentence printed in the failure.
|
||||
|
||||
**ONE EXISTING TEST WAS EDITED, AND IT CAUGHT THIS FIX CORRECTLY.**
|
||||
`TestR431_FiresOnAMassDeletion` asserted the fragment `"NOT confirmed data loss"` and went red on the
|
||||
new wording. The fragment is removed rather than updated: **the wording now has ONE home**
|
||||
(`offsite_r434_test.go`), because duplicating it would create the second source that makes the next
|
||||
correction land in one file and not the other. The signal fragments it still asserts (`69`, `4`,
|
||||
`read-only`) are unchanged.
|
||||
|
||||
**R-435 WRITTEN INTO THE DETECTOR'S OWN DOCUMENTATION, no threshold changed.** The comment above
|
||||
`snapshotDropFraction` now states what this detector does NOT see: **a mass deletion, yes; one app
|
||||
being wiped, no.** Worked on the live fleet — demo-hp's baseline is 69 snapshots across 9 apps, so
|
||||
~35 must go before it speaks, and one app's tag is ~9. `offbox.go:1388` runs `forget --prune` grouped
|
||||
by `host,tags`, so the blind spot sits on the most likely single-app failure. **The insensitivity is
|
||||
deliberate and must not be "fixed" by lowering the numbers** — a detector that cries wolf is switched
|
||||
off within a fortnight. What is not acceptable is claiming coverage it does not have; per-app
|
||||
detection needs a SECOND signal keyed on the per-tag count.
|
||||
|
||||
## v0.111.0 — notice a deletion within a day (2026-09-01, R-431; corrects R-429, re-scopes R-95)
|
||||
|
||||
**Third signal in `OffsiteChecker`, beside FILL and STALENESS. No controller change, no golden owed.**
|
||||
|
||||
@@ -238,6 +238,19 @@ func (oc *OffsiteChecker) isStale(customerID string, off *offsiteReport) bool {
|
||||
// reached by retention; the floor of 5 stops a tiny-count box alarming on ordinary ageing. It is
|
||||
// deliberately NOT sensitive — a detector that cries wolf is switched off within a fortnight, and
|
||||
// this project has proved that twice in a week.
|
||||
// WHAT THIS DETECTOR DOES NOT SEE — R-435, and it must be read wherever "an unexplained fall is
|
||||
// noticed within a day" is claimed, because that claim is true only of falls above the fraction.
|
||||
//
|
||||
// **It sees a MASS deletion. It does not see ONE APP being wiped.** Worked on the live fleet
|
||||
// 2026-09-01: demo-hp's baseline is 69 snapshots across 9 apps, so ~35 must go before this speaks;
|
||||
// one app's tag is ~9 and is invisible. And `offbox.go:1388` runs `forget --prune` **grouped by
|
||||
// host,tags** — a per-tag wipe is exactly the shape a faulty retention or a targeted deletion
|
||||
// produces, so the blind spot sits on the most likely single-app failure, not an exotic one.
|
||||
//
|
||||
// THIS IS DELIBERATE AND THE THRESHOLD SHOULD NOT BE LOWERED TO "FIX" IT. The reasoning is below: a
|
||||
// detector that cries wolf is switched off within a fortnight, and this project has proved that
|
||||
// twice in a week. What is NOT acceptable is claiming coverage this does not have. Anyone adding
|
||||
// per-app detection should add a SECOND signal keyed on the per-tag count, not move these numbers.
|
||||
const (
|
||||
snapshotDropFraction = 0.5 // more than half the history gone in one step
|
||||
snapshotDropFloor = 5 // and at least this many, so small counts do not twitch
|
||||
@@ -286,12 +299,25 @@ func (oc *OffsiteChecker) snapshotDropped(customerID string, off *offsiteReport)
|
||||
// THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually
|
||||
// false: the daily Storage Box snapshots are read-only to every account (proven, not cited) and hold
|
||||
// the older copy. It says what happened, what it means, and where the data still is.
|
||||
//
|
||||
// AND IT MUST NOT SAY THE DATA IS RECOVERABLE EITHER — R-434, fixed 2026-09-01, hub v0.111.1. The
|
||||
// sentence shipped that morning promised "so this is recoverable file-by-file". Measured the same day
|
||||
// (R-433): no snapshot is reachable from a sub-account by ANY name — 777,600 exact names in the
|
||||
// vendor's own format over nine days, zero hits, with a passing control; `/home` and `/.zfs` are
|
||||
// different filesystems and `/home/.zfs` does not exist. So the promise named a route nobody can walk.
|
||||
//
|
||||
// THE FIX IS A DELETION, NOT A REPLACEMENT, AND THAT IS THE WHOLE POINT. R-434's row said the fix was
|
||||
// blocked on R-433 — on knowing what IS true. It is not, if the promise is simply withdrawn: a
|
||||
// sentence that asserts neither loss nor recovery is true under EVERY possible answer to the provider
|
||||
// question, so it never needs a second rewrite. An alarm rewritten twice in a week is worse than one
|
||||
// rewritten once, because the operator learns its words do not mean anything.
|
||||
func (oc *OffsiteChecker) emitSnapshotDrop(customerID string, off *offsiteReport, prev, cur int) {
|
||||
message := fmt.Sprintf(
|
||||
"Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than "+
|
||||
"retention can explain. The daily Storage Box snapshots are read-only and still hold the "+
|
||||
"older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check "+
|
||||
"whether a deletion ran on the box before restoring anything.",
|
||||
"older copy. The route back out of them is not yet established, so treat this as neither "+
|
||||
"confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and "+
|
||||
"check whether a deletion ran on the box.",
|
||||
customerID, prev, cur)
|
||||
details, _ := json.Marshal(map[string]any{
|
||||
"customer_id": customerID, "previous_count": prev, "current_count": cur,
|
||||
|
||||
@@ -62,7 +62,15 @@ func TestR431_FiresOnAMassDeletion(t *testing.T) {
|
||||
default:
|
||||
t.Fatalf("severity %q is outside the hub vocabulary — it would be coerced to info and reach nobody", drops[0].sev)
|
||||
}
|
||||
for _, frag := range []string{"69", "4", "read-only", "NOT confirmed data loss"} {
|
||||
// These fragments belong to the SIGNAL — the two counts, and where the older copy still is.
|
||||
//
|
||||
// THE WORDING FRAGMENT THAT USED TO SIT HERE IS GONE ON PURPOSE. This list asserted
|
||||
// "NOT confirmed data loss" until 2026-09-01, and it caught the R-434 fix, correctly — the
|
||||
// sentence changed because the alarm was promising a recovery that R-433 showed cannot be
|
||||
// performed. The wording now has ONE home, `offsite_r434_test.go`, which pins both what the
|
||||
// message must say and what it must never say again. Duplicating it here would create the
|
||||
// second source that makes the next correction land in one file and not the other.
|
||||
for _, frag := range []string{"69", "4", "read-only"} {
|
||||
if !strings.Contains(drops[0].msg, frag) {
|
||||
t.Fatalf("message must contain %q; got: %s", frag, drops[0].msg)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
package monitor
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// R-434 — the snapshot-drop alarm must not promise a recovery that cannot be performed.
|
||||
//
|
||||
// WHAT WENT WRONG. hub v0.111.0 shipped, on 2026-09-01, an alarm reading "The daily Storage Box
|
||||
// snapshots are read-only and still hold the older copy, so this is recoverable file-by-file; it is
|
||||
// NOT confirmed data loss." Measured the same day (R-433): NO snapshot is reachable from a
|
||||
// sub-account by any name — 777,600 exact names in the vendor format over nine days, zero hits, with
|
||||
// a passing control. The promise named a route nobody can walk, in the one message an operator acts
|
||||
// on while their customer's off-site history is disappearing.
|
||||
//
|
||||
// WHY THE FIX IS A DELETION AND NOT A REPLACEMENT. A sentence asserting neither loss nor recovery is
|
||||
// true under every possible answer to the outstanding provider question, so it never needs a second
|
||||
// rewrite. That is why this test pins the ABSENCE of a promise as hard as it pins the new words:
|
||||
// the next person who "improves" this message by putting a route back into it must fail here.
|
||||
//
|
||||
// THESE TESTS DRIVE THE REAL PATH — saveOffsiteReport -> oc.Check() -> the notify callback — so they
|
||||
// assert the CONSEQUENCE (the sentence an operator receives), not the mechanism. Asserting the
|
||||
// mechanism one layer below where the damage happens is R-224, entry 9 of the doctrine table.
|
||||
//
|
||||
// ASCII-ONLY FRAGMENTS. The message contains an em dash. A fragment carrying one has returned 0 for
|
||||
// strings that WERE there in this project before, so every fragment below is plain ASCII.
|
||||
|
||||
// the promise that must never come back, in the shapes it could plausibly return as
|
||||
var r434ForbiddenFragments = []string{
|
||||
"recoverable file-by-file",
|
||||
"recoverable file by file",
|
||||
"so this is recoverable",
|
||||
}
|
||||
|
||||
// the withdrawal that replaced it
|
||||
var r434RequiredFragments = []string{
|
||||
"The route back out of them is not yet established",
|
||||
"neither confirmed data loss nor confirmed recovery",
|
||||
"Get in touch before restoring anything",
|
||||
"still hold the older copy", // clause (a) STANDS and must not be lost with the promise
|
||||
}
|
||||
|
||||
// r434Message drives the production path once and returns the message the operator would receive.
|
||||
func r434Message(t *testing.T) string {
|
||||
t.Helper()
|
||||
st := newDiskStore(t)
|
||||
var msgs []string
|
||||
saveOffsiteReport(t, st, "victim", dropJSON(69, true, "", "ok"))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, msg, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
msgs = append(msgs, msg)
|
||||
}
|
||||
}, quietLog())
|
||||
|
||||
saveOffsiteReport(t, st, "victim", dropJSON(4, true, "", "ok"))
|
||||
oc.Check()
|
||||
|
||||
if len(msgs) != 1 {
|
||||
t.Fatalf("setup: want exactly 1 offsite_snapshots_dropped message, got %d", len(msgs))
|
||||
}
|
||||
return msgs[0]
|
||||
}
|
||||
|
||||
// TestR434_AlarmMakesNoRecoveryPromise — the fix, both directions, with both controls.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-01, recorded in REPORT.md): restoring the v0.111.0 sentence in
|
||||
// emitSnapshotDrop makes this FAIL on the forbidden fragment "recoverable file-by-file" AND on all
|
||||
// three required fragments, with the offending sentence printed in the failure message.
|
||||
func TestR434_AlarmMakesNoRecoveryPromise(t *testing.T) {
|
||||
msg := r434Message(t)
|
||||
|
||||
// POSITIVE CONTROL — a fragment present in EVERY version of this alarm. If this is missing the
|
||||
// test is reading the wrong string and every other assertion below is worthless.
|
||||
if !strings.Contains(msg, "off-site backup count fell from") {
|
||||
t.Fatalf("positive control failed: not the snapshot-drop message at all: %q", msg)
|
||||
}
|
||||
// NEGATIVE CONTROL — proves Contains can actually report absence here.
|
||||
if strings.Contains(msg, "zzz-no-such-fragment-r434") {
|
||||
t.Fatalf("negative control failed: matched a fragment that cannot exist: %q", msg)
|
||||
}
|
||||
|
||||
for _, bad := range r434ForbiddenFragments {
|
||||
if strings.Contains(msg, bad) {
|
||||
t.Errorf("alarm promises a recovery that cannot be performed (R-433): found %q in %q", bad, msg)
|
||||
}
|
||||
}
|
||||
for _, want := range r434RequiredFragments {
|
||||
if !strings.Contains(msg, want) {
|
||||
t.Errorf("alarm is missing the withdrawal wording: want %q in %q", want, msg)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestR434_AlarmStillDoesNotClaimDataLoss — the OTHER direction, and the reason the fix is a
|
||||
// withdrawal rather than a reversal.
|
||||
//
|
||||
// After R-433 the temptation is to swing to "your backups are gone". That is still usually FALSE:
|
||||
// the snapshots exist and hold the older copy; what is unproven is our route to them. An alarm that
|
||||
// over-claims loss sends an operator into a destructive recovery they did not need — which is the
|
||||
// failure the v0.111.0 comment was written to prevent, and it must survive its own correction.
|
||||
func TestR434_AlarmStillDoesNotClaimDataLoss(t *testing.T) {
|
||||
msg := r434Message(t)
|
||||
|
||||
for _, bad := range []string{
|
||||
"data is lost", "backups are gone", "data has been lost", "permanently lost", "unrecoverable",
|
||||
} {
|
||||
if strings.Contains(msg, bad) {
|
||||
t.Errorf("alarm over-claims loss: found %q in %q", bad, msg)
|
||||
}
|
||||
}
|
||||
// The one phrase that must appear NEGATED, never bare. A bare "confirmed data loss" would read
|
||||
// as a verdict; the shipped sentence only ever uses it inside "neither ... nor".
|
||||
if strings.Contains(msg, "confirmed data loss") &&
|
||||
!strings.Contains(msg, "neither confirmed data loss nor confirmed recovery") {
|
||||
t.Errorf("the phrase 'confirmed data loss' appears outside its negation: %q", msg)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR434_StoredEventCarriesTheSameSentence — the delivered message and the stored one are the same
|
||||
// string today, and a future refactor that formats them separately must not let them drift: the
|
||||
// operator reads the mail, but every later audit reads the stored row.
|
||||
func TestR434_StoredEventCarriesTheSameSentence(t *testing.T) {
|
||||
st := newDiskStore(t)
|
||||
var delivered string
|
||||
saveOffsiteReport(t, st, "victim", dropJSON(69, true, "", "ok"))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, msg, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
delivered = msg
|
||||
}
|
||||
}, quietLog())
|
||||
saveOffsiteReport(t, st, "victim", dropJSON(4, true, "", "ok"))
|
||||
oc.Check()
|
||||
|
||||
evs, err := st.GetRecentEvents("victim", 50)
|
||||
if err != nil {
|
||||
t.Fatalf("GetRecentEvents: %v", err)
|
||||
}
|
||||
var stored []string
|
||||
for _, e := range evs {
|
||||
if e.EventType == "offsite_snapshots_dropped" {
|
||||
stored = append(stored, e.Message)
|
||||
}
|
||||
}
|
||||
if len(stored) != 1 {
|
||||
t.Fatalf("want exactly 1 stored offsite_snapshots_dropped, got %d", len(stored))
|
||||
}
|
||||
if stored[0] != delivered {
|
||||
t.Errorf("stored and delivered messages have drifted:\n stored: %q\n delivered: %q", stored[0], delivered)
|
||||
}
|
||||
if strings.Contains(stored[0], "recoverable file-by-file") {
|
||||
t.Errorf("the stored row still carries the withdrawn promise: %q", stored[0])
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user