hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s

R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.

  was:  "...still hold the older copy, so this is recoverable file-by-file; it is NOT
         confirmed data loss. Check whether a deletion ran on the box before restoring."
  now:  "...still hold the older copy. The route back out of them is not yet established,
         so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
         before restoring anything, and check whether a deletion ran on the box."

It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.

Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.

R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.

THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.

Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.

R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.

Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
2026-09-01 18:34:52 +02:00
parent 10c223bdfe
commit db38f4c800
9 changed files with 673 additions and 35 deletions
+47 -17
View File
@@ -1,8 +1,8 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-01 (second pass) — both faults from the overnight test are fixed, and the golden
is baked, vouched and delivered. Both machines are on 0.232.0 and moved themselves. NOTHING is
waiting on you.**
**Updated 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down
where it stops. One alarm that was telling you something untrue is fixed and live (hub 0.111.1).
ONE THING is waiting on you: two short e-mails to Hetzner, drafted and ready to paste.**
**Earlier 2026-08-31 — I measured whether the box could test its own remote restore without you,
then built the narrow version you picked. The measurement is why it is 3 seconds a night and
@@ -18,7 +18,7 @@ not an evening's work.**
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Two things are waiting on you — item 4 (a snapshot name, two minutes) and item 5 (ignore one alarm mail).** Otherwise: Both problems the overnight test found are fixed and proven on
1. **One thing is waiting on you — item 4 (send two e-mails).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
the real machines:
- the background job that could delete a live restore's lock now waits its turn — and the check
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
@@ -37,24 +37,34 @@ nothing.*
two register lines in the hub (already live). No customer action, no data migration, no
credential change.
4. **Good news, and I was wrong yesterday.** I told you I could not find the nightly safety net on
the storage. **I was looking for the wrong name.** The snapshots are there, exactly as you saw in
the control panel: seven of them, one a day. **And they are better than we thought** — I tried to
write into the snapshot area from a customer's machine and the storage **refused**, while the same
write to its normal folder worked. So a machine that wipes its own backup **cannot touch the
snapshots of it**. The worst case is losing about a day, then copying the rest back file by file.
**That is much smaller than what the notes have said since July.** I have corrected the notes.
**What I still need from you:** the machines can see the snapshot *door* but not what is inside —
only the main account can. So getting data back is you, in a browser, for now. **If you read one
snapshot's name off the panel and send it to me, one command settles whether the machines can
reach them directly** — and if they can, recovery becomes something the product does by itself.
**If you do nothing:** it stays a manual job for you, which is workable but slow.
4. **Please send two short e-mails to Hetzner. They are written for you.**
`felhom.eu/documentation/runbooks/provider-questions-2026-09-01.md` — open it, copy, send. No
password or key is in that file, and none should be added.
**Why.** Yesterday I told you the snapshots make a wiped backup survivable: lose about a day, copy
the rest back file by file. **The first half is still true. The second half is not, and I found
that out by trying it.** I tried **777,600** snapshot names on the storage, over nine days, in
Hetzner's own naming style. **None of them opened.** Then I found why: your data and the snapshot
door sit on **two different drives** inside the storage, and the door for your data **does not
exist at all**. So there is no way in from the machines.
**What is still true, and it matters:** a machine that wipes its own backup **still cannot touch
the snapshots of it**. The older copy is there. What we do not have is a way to reach it.
**The two questions.** One: can the **main** account pull single files out of a snapshot? Two: on
one of Hetzner's own tools, is a "cannot delete" switch forced by them, or chosen by the machine?
**The second one could remove the whole problem** — no new hardware, no moving anyone's data.
**If you do nothing:** we cannot finish this. The backups keep working and keep being checked; we
simply cannot say what a wiped backup costs, and I would then put this risk back near the top of
your list. **My pick: send both. It is five minutes and it decides an evening's work.**
5. **You will have received an alarm email from me today about `demo-hp` losing 65 backups. It is a
test and nothing is wrong.** I built the new "someone deleted the backups" alarm and had to fire it
once for real to prove it reaches you. Subject: `[Felhom] 🔴 demo-hp: offsite_snapshots_dropped`.
**No backups were deleted.** **If you do nothing:** nothing — but please do not act on that one
mail.
mail. **Closed now:** that was the only such mail, it was a test, and the alarm's wording has since
been corrected (item under *Decided* below). Nothing further is needed from you here.
6. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
@@ -66,6 +76,26 @@ nothing.*
## Decided — and what would reopen each
- **THE BACKUP AND RESTORE WORK IS FINISHED FOR BETA. DECIDED 2026-09-01.**
**What is done, and proven on the real machines:** everything **you or a customer** does alone —
getting deleted files back, getting an app's data back, getting a whole app back, and losing a
drive. The restore tells you what it put back, refuses if there is no room, will not accept a
half-copy, and puts your own data back if it fails.
**What is parked until after beta:** everything **only I do, with you** — rebuilding a machine as
itself, losing a whole box, recovering from ransomware, restoring the hub, and losing Hetzner.
**Six of these have never been timed, and the hub has never been restored.** They are written down,
they are real, and **none of them stops a beta customer.**
**Reopens if:** something a customer does for themselves turns out to be broken; **or** Hetzner's
two answers change what the snapshots are worth; **or** a real customer's data is at stake in one
of the parked items.
- **The alarm that promised too much: FIXED and live (hub 0.111.1). DECIDED 2026-09-01.**
Yesterday's alarm mail said a deleted backup was *"recoverable file-by-file"*. **We now know it is
not.** I removed the promise rather than writing a new one, so the sentence stays true whatever
Hetzner answers. It no longer says the data is lost either — that is still usually untrue.
**Reopens if:** Hetzner's answers give us a real route back; then the alarm can name it.
- **The missing-golden warning is now pointed at the people who can act on it (R-404). DECIDED
2026-09-01, and built the same day.** The warning was aimed at the wrong repository: the one where
a release actually happens never checked at all, while the one that only holds documents was