hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.
was: "...still hold the older copy, so this is recoverable file-by-file; it is NOT
confirmed data loss. Check whether a deletion ran on the box before restoring."
now: "...still hold the older copy. The route back out of them is not yet established,
so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
before restoring anything, and check whether a deletion ran on the box."
It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.
Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.
R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.
THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.
Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.
R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.
Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,101 @@
|
||||
# Two questions for Hetzner — drafted, ready to send (2026-09-01)
|
||||
|
||||
**These are the only thing standing between us and finishing R-95 properly.** They cost nothing and
|
||||
they are answerable by a support agent without escalation.
|
||||
|
||||
**Send them yourself.** CC drafted them and did not send them, and did not call the provider API —
|
||||
`§11-D` is still the operator's fence. **No credential, password or token appears below, and none
|
||||
should be added.** The account id and the product name are all either question needs.
|
||||
|
||||
**What to fill in:** the ticket needs the Storage Box account. Ours is **`u629488`** (the box the
|
||||
register calls `storage-box-pool-1`, plan BX11). Nothing else.
|
||||
|
||||
**Why two separate tickets:** they go to different parts of the answer — one is about the snapshot
|
||||
product, one is about the SSH endpoint's configuration — and a single ticket asking both tends to get
|
||||
one answered and the other dropped.
|
||||
|
||||
---
|
||||
|
||||
## Question 1 — can the MAIN account retrieve individual files from a snapshot?
|
||||
|
||||
**Why it matters, in one line:** if it cannot, the only route back is a whole-box rollback that hits
|
||||
every customer on the box and destroys every newer snapshot — which would mean the snapshots protect
|
||||
almost nobody in practice. **This is the question that decides how urgent R-95 is.**
|
||||
|
||||
> **Subject:** Storage Box u629488 — retrieving individual files from a snapshot
|
||||
>
|
||||
> Hello,
|
||||
>
|
||||
> We use Storage Box `u629488` with sub-accounts, and daily automatic snapshots are enabled.
|
||||
>
|
||||
> We can reach `/.zfs/snapshot` from a sub-account, but it lists as empty, and no snapshot name we
|
||||
> try can be entered. We understand sub-accounts may be restricted here.
|
||||
>
|
||||
> Our question is about the **main account**: from the main account, over SSH or SFTP on port 23,
|
||||
> can we **read or download individual files and directories out of a specific snapshot** — for
|
||||
> example a single directory under one sub-account's home — **without** performing a snapshot
|
||||
> restore of the whole Storage Box?
|
||||
>
|
||||
> If yes, please tell us the exact path we should use and how the snapshot directory is named.
|
||||
>
|
||||
> If no, please confirm that the only way to get data out of a snapshot is the full "restore
|
||||
> snapshot" action on the whole Storage Box.
|
||||
>
|
||||
> Thank you.
|
||||
|
||||
**How to read the answer.**
|
||||
* **"Yes, from the main account"** → per-file recovery exists, but it is an operator act in a
|
||||
browser or over the main account's own SSH, and it can never be something the product does for the
|
||||
customer. R-433 closes at that. R-95 stays where it is.
|
||||
* **"No, only a full restore"** → the snapshots do **not** bound a single customer's exposure at all,
|
||||
because using them costs every other customer on the box their newer snapshots. **R-95 becomes
|
||||
urgent and the transport change stops being optional.**
|
||||
|
||||
---
|
||||
|
||||
## Question 2 — is `--append-only` enforced on the `rclone serve restic` endpoint?
|
||||
|
||||
**Why it matters, in one line:** if it is enforced server-side, a compromised box **cannot delete its
|
||||
own backups**, with no new machine and no data migration. **This is the question that could make R-95
|
||||
disappear.**
|
||||
|
||||
> **Subject:** Storage Box u629488 — rclone serve restic endpoint and --append-only
|
||||
>
|
||||
> Hello,
|
||||
>
|
||||
> The restricted SSH shell on Storage Box `u629488` (port 23) lists `rclone serve restic --stdio`
|
||||
> among the available server-side backends.
|
||||
>
|
||||
> Our question is about how that command is run on your side: **is `--append-only` enforced by you,
|
||||
> or is the command line taken from what the client sends?**
|
||||
>
|
||||
> In other words, if a client connects and asks for `rclone serve restic --stdio` **without**
|
||||
> `--append-only`, does it get a server that permits deletions?
|
||||
>
|
||||
> If the flag can be enforced, is there any way for us to request that for this account or for
|
||||
> individual sub-accounts?
|
||||
>
|
||||
> Thank you.
|
||||
|
||||
**How to read the answer.**
|
||||
* **"Enforced server-side" or "can be enabled per account"** → this is the cheap prevention the
|
||||
2026-09-01 spike priced at a new always-on service plus either a mount in the hot path or migrating
|
||||
every customer's history. **It needs none of that** — the server already runs at the provider, and
|
||||
restic 0.14.0 already speaks the `rclone:` backend (measured, with a control:
|
||||
`banana:` → `invalid backend`, `rclone:` → the helper was executed). The remaining work is putting
|
||||
`rclone` in the controller image and switching the repository URL. **R-95's root cause goes away.**
|
||||
* **"The client supplies the command line"** → **the lead is worth nothing** and should be recorded
|
||||
as dead, not left looking promising. Prevention then still needs a machine in front of the store,
|
||||
and the decision reverts to the spike's option 3 at its original price.
|
||||
|
||||
---
|
||||
|
||||
## Where these came from
|
||||
|
||||
* **R-433** — no snapshot is reachable from a sub-account by any name. 777,600 exact names in the
|
||||
vendor's `YYYY-MM-DDTHH-MM-SS` format over nine days, zero hits, with a passing control; `/home`
|
||||
and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Question 1 exists because
|
||||
that measurement can only speak for a sub-account.
|
||||
* **R-436** — the `rclone serve restic --stdio` backend and restic's `rclone:` support, both
|
||||
measured. Question 2 is the one caveat that decides whether the lead is real.
|
||||
* Evidence for both: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`.
|
||||
Reference in New Issue
Block a user