hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s

R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.

  was:  "...still hold the older copy, so this is recoverable file-by-file; it is NOT
         confirmed data loss. Check whether a deletion ran on the box before restoring."
  now:  "...still hold the older copy. The route back out of them is not yet established,
         so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
         before restoring anything, and check whether a deletion ran on the box."

It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.

Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.

R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.

THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.

Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.

R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.

Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
2026-09-01 18:34:52 +02:00
parent 10c223bdfe
commit db38f4c800
9 changed files with 673 additions and 35 deletions
+67
View File
@@ -1,3 +1,70 @@
## v0.111.1 — the alarm stops promising a rescue that does not exist (2026-09-01, R-434)
**Text-only patch on a live alarm. No controller change, no golden owed, no floor change.**
**WHAT WAS WRONG.** v0.111.0 shipped, that morning, a snapshot-drop alarm reading:
> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is
> recoverable file-by-file**; it is NOT confirmed data loss."
**Measured the same afternoon (R-433): no snapshot is reachable from a customer's sub-account by ANY
name.** 777,600 exact names in the vendor's own `YYYY-MM-DDTHH-MM-SS` format across nine full days at
second granularity, plus 126 alternative shapes — zero hits, with a control proving the identical
batch returns a path that does exist (6/6). The structural reason: `/home` (the customer data) and
`/.zfs` are **different filesystems**, and `/home/.zfs` does not exist, so a snapshot behind that door
could not hold the repository even with the right name. **The promise named a route nobody can walk**,
in the one message an operator acts on while a customer's off-site history is disappearing.
**THE FIX IS A DELETION, NOT A REPLACEMENT, AND THAT IS THE POINT.** R-434's register row said the fix
was blocked on R-433 — on first establishing what IS true. **It was not blocked, once the promise is
withdrawn rather than swapped.** The shipped sentence now asserts neither loss nor recovery:
> "The daily Storage Box snapshots are read-only and still hold the older copy. The route back out of
> them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery.
> Get in touch before restoring anything, and check whether a deletion ran on the box."
**That is true under every possible answer to the outstanding provider question, so it never needs a
second rewrite.** An alarm rewritten twice in a week is worse than one rewritten once: the operator
learns its words do not mean anything.
**AND IT MUST NOT SWING THE OTHER WAY.** "Your backups are gone" is still usually FALSE — the
snapshots exist and hold the older copy; what is unproven is our route to them. Clause (a) of the
2026-09-01 measurement (the box **cannot write** into the snapshot area) stands and is re-confirmed.
Over-claiming loss would send an operator into a destructive recovery they did not need, which is the
failure v0.111.0's own comment was written to prevent — it had to survive its own correction.
**TESTS — `internal/monitor/offsite_r434_test.go`, three, all driving the production path**
(`saveOffsiteReport` → `oc.Check()` → the notify callback), so they assert the sentence an operator
RECEIVES rather than the function that formats it. ASCII-only fragments, with a positive control (a
phrase present in every version of the alarm) and a negative control (a phrase that cannot exist).
- `TestR434_AlarmMakesNoRecoveryPromise` — the withdrawal is present, the promise is absent, in three
shapes it could plausibly return as.
- `TestR434_AlarmStillDoesNotClaimDataLoss` — the other direction; `confirmed data loss` may appear
only inside its negation.
- `TestR434_StoredEventCarriesTheSameSentence` — the stored row and the delivered mail must not drift.
**RED-PROOF (run 2026-09-01, before the fix was restored):** the v0.111.0 sentence was put back and
all three FAILED — on `"recoverable file-by-file"`, on `"so this is recoverable"`, on all three
required fragments, on the un-negated `confirmed data loss`, and on the stored row — each with the
offending sentence printed in the failure.
**ONE EXISTING TEST WAS EDITED, AND IT CAUGHT THIS FIX CORRECTLY.**
`TestR431_FiresOnAMassDeletion` asserted the fragment `"NOT confirmed data loss"` and went red on the
new wording. The fragment is removed rather than updated: **the wording now has ONE home**
(`offsite_r434_test.go`), because duplicating it would create the second source that makes the next
correction land in one file and not the other. The signal fragments it still asserts (`69`, `4`,
`read-only`) are unchanged.
**R-435 WRITTEN INTO THE DETECTOR'S OWN DOCUMENTATION, no threshold changed.** The comment above
`snapshotDropFraction` now states what this detector does NOT see: **a mass deletion, yes; one app
being wiped, no.** Worked on the live fleet — demo-hp's baseline is 69 snapshots across 9 apps, so
~35 must go before it speaks, and one app's tag is ~9. `offbox.go:1388` runs `forget --prune` grouped
by `host,tags`, so the blind spot sits on the most likely single-app failure. **The insensitivity is
deliberate and must not be "fixed" by lowering the numbers** — a detector that cries wolf is switched
off within a fortnight. What is not acceptable is claiming coverage it does not have; per-app
detection needs a SECOND signal keyed on the per-tag count.
## v0.111.0 — notice a deletion within a day (2026-09-01, R-431; corrects R-429, re-scopes R-95)
**Third signal in `OffsiteChecker`, beside FILL and STALENESS. No controller change, no golden owed.**