hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s

R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.

  was:  "...still hold the older copy, so this is recoverable file-by-file; it is NOT
         confirmed data loss. Check whether a deletion ran on the box before restoring."
  now:  "...still hold the older copy. The route back out of them is not yet established,
         so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
         before restoring anything, and check whether a deletion ran on the box."

It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.

Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.

R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.

THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.

Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.

R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.

Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
2026-09-01 18:34:52 +02:00
parent 10c223bdfe
commit db38f4c800
9 changed files with 673 additions and 35 deletions
+28 -2
View File
@@ -238,6 +238,19 @@ func (oc *OffsiteChecker) isStale(customerID string, off *offsiteReport) bool {
// reached by retention; the floor of 5 stops a tiny-count box alarming on ordinary ageing. It is
// deliberately NOT sensitive — a detector that cries wolf is switched off within a fortnight, and
// this project has proved that twice in a week.
// WHAT THIS DETECTOR DOES NOT SEE — R-435, and it must be read wherever "an unexplained fall is
// noticed within a day" is claimed, because that claim is true only of falls above the fraction.
//
// **It sees a MASS deletion. It does not see ONE APP being wiped.** Worked on the live fleet
// 2026-09-01: demo-hp's baseline is 69 snapshots across 9 apps, so ~35 must go before this speaks;
// one app's tag is ~9 and is invisible. And `offbox.go:1388` runs `forget --prune` **grouped by
// host,tags** — a per-tag wipe is exactly the shape a faulty retention or a targeted deletion
// produces, so the blind spot sits on the most likely single-app failure, not an exotic one.
//
// THIS IS DELIBERATE AND THE THRESHOLD SHOULD NOT BE LOWERED TO "FIX" IT. The reasoning is below: a
// detector that cries wolf is switched off within a fortnight, and this project has proved that
// twice in a week. What is NOT acceptable is claiming coverage this does not have. Anyone adding
// per-app detection should add a SECOND signal keyed on the per-tag count, not move these numbers.
const (
snapshotDropFraction = 0.5 // more than half the history gone in one step
snapshotDropFloor = 5 // and at least this many, so small counts do not twitch
@@ -286,12 +299,25 @@ func (oc *OffsiteChecker) snapshotDropped(customerID string, off *offsiteReport)
// THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually
// false: the daily Storage Box snapshots are read-only to every account (proven, not cited) and hold
// the older copy. It says what happened, what it means, and where the data still is.
//
// AND IT MUST NOT SAY THE DATA IS RECOVERABLE EITHER — R-434, fixed 2026-09-01, hub v0.111.1. The
// sentence shipped that morning promised "so this is recoverable file-by-file". Measured the same day
// (R-433): no snapshot is reachable from a sub-account by ANY name — 777,600 exact names in the
// vendor's own format over nine days, zero hits, with a passing control; `/home` and `/.zfs` are
// different filesystems and `/home/.zfs` does not exist. So the promise named a route nobody can walk.
//
// THE FIX IS A DELETION, NOT A REPLACEMENT, AND THAT IS THE WHOLE POINT. R-434's row said the fix was
// blocked on R-433 — on knowing what IS true. It is not, if the promise is simply withdrawn: a
// sentence that asserts neither loss nor recovery is true under EVERY possible answer to the provider
// question, so it never needs a second rewrite. An alarm rewritten twice in a week is worse than one
// rewritten once, because the operator learns its words do not mean anything.
func (oc *OffsiteChecker) emitSnapshotDrop(customerID string, off *offsiteReport, prev, cur int) {
message := fmt.Sprintf(
"Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than "+
"retention can explain. The daily Storage Box snapshots are read-only and still hold the "+
"older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check "+
"whether a deletion ran on the box before restoring anything.",
"older copy. The route back out of them is not yet established, so treat this as neither "+
"confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and "+
"check whether a deletion ran on the box.",
customerID, prev, cur)
details, _ := json.Marshal(map[string]any{
"customer_id": customerID, "previous_count": prev, "current_count": cur,
+9 -1
View File
@@ -62,7 +62,15 @@ func TestR431_FiresOnAMassDeletion(t *testing.T) {
default:
t.Fatalf("severity %q is outside the hub vocabulary — it would be coerced to info and reach nobody", drops[0].sev)
}
for _, frag := range []string{"69", "4", "read-only", "NOT confirmed data loss"} {
// These fragments belong to the SIGNAL — the two counts, and where the older copy still is.
//
// THE WORDING FRAGMENT THAT USED TO SIT HERE IS GONE ON PURPOSE. This list asserted
// "NOT confirmed data loss" until 2026-09-01, and it caught the R-434 fix, correctly — the
// sentence changed because the alarm was promising a recovery that R-433 showed cannot be
// performed. The wording now has ONE home, `offsite_r434_test.go`, which pins both what the
// message must say and what it must never say again. Duplicating it here would create the
// second source that makes the next correction land in one file and not the other.
for _, frag := range []string{"69", "4", "read-only"} {
if !strings.Contains(drops[0].msg, frag) {
t.Fatalf("message must contain %q; got: %s", frag, drops[0].msg)
}
+155
View File
@@ -0,0 +1,155 @@
package monitor
import (
"strings"
"testing"
"time"
)
// R-434 — the snapshot-drop alarm must not promise a recovery that cannot be performed.
//
// WHAT WENT WRONG. hub v0.111.0 shipped, on 2026-09-01, an alarm reading "The daily Storage Box
// snapshots are read-only and still hold the older copy, so this is recoverable file-by-file; it is
// NOT confirmed data loss." Measured the same day (R-433): NO snapshot is reachable from a
// sub-account by any name — 777,600 exact names in the vendor format over nine days, zero hits, with
// a passing control. The promise named a route nobody can walk, in the one message an operator acts
// on while their customer's off-site history is disappearing.
//
// WHY THE FIX IS A DELETION AND NOT A REPLACEMENT. A sentence asserting neither loss nor recovery is
// true under every possible answer to the outstanding provider question, so it never needs a second
// rewrite. That is why this test pins the ABSENCE of a promise as hard as it pins the new words:
// the next person who "improves" this message by putting a route back into it must fail here.
//
// THESE TESTS DRIVE THE REAL PATH — saveOffsiteReport -> oc.Check() -> the notify callback — so they
// assert the CONSEQUENCE (the sentence an operator receives), not the mechanism. Asserting the
// mechanism one layer below where the damage happens is R-224, entry 9 of the doctrine table.
//
// ASCII-ONLY FRAGMENTS. The message contains an em dash. A fragment carrying one has returned 0 for
// strings that WERE there in this project before, so every fragment below is plain ASCII.
// the promise that must never come back, in the shapes it could plausibly return as
var r434ForbiddenFragments = []string{
"recoverable file-by-file",
"recoverable file by file",
"so this is recoverable",
}
// the withdrawal that replaced it
var r434RequiredFragments = []string{
"The route back out of them is not yet established",
"neither confirmed data loss nor confirmed recovery",
"Get in touch before restoring anything",
"still hold the older copy", // clause (a) STANDS and must not be lost with the promise
}
// r434Message drives the production path once and returns the message the operator would receive.
func r434Message(t *testing.T) string {
t.Helper()
st := newDiskStore(t)
var msgs []string
saveOffsiteReport(t, st, "victim", dropJSON(69, true, "", "ok"))
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, msg, _, _ string) {
if et == "offsite_snapshots_dropped" {
msgs = append(msgs, msg)
}
}, quietLog())
saveOffsiteReport(t, st, "victim", dropJSON(4, true, "", "ok"))
oc.Check()
if len(msgs) != 1 {
t.Fatalf("setup: want exactly 1 offsite_snapshots_dropped message, got %d", len(msgs))
}
return msgs[0]
}
// TestR434_AlarmMakesNoRecoveryPromise — the fix, both directions, with both controls.
//
// RED-PROOF (run 2026-09-01, recorded in REPORT.md): restoring the v0.111.0 sentence in
// emitSnapshotDrop makes this FAIL on the forbidden fragment "recoverable file-by-file" AND on all
// three required fragments, with the offending sentence printed in the failure message.
func TestR434_AlarmMakesNoRecoveryPromise(t *testing.T) {
msg := r434Message(t)
// POSITIVE CONTROL — a fragment present in EVERY version of this alarm. If this is missing the
// test is reading the wrong string and every other assertion below is worthless.
if !strings.Contains(msg, "off-site backup count fell from") {
t.Fatalf("positive control failed: not the snapshot-drop message at all: %q", msg)
}
// NEGATIVE CONTROL — proves Contains can actually report absence here.
if strings.Contains(msg, "zzz-no-such-fragment-r434") {
t.Fatalf("negative control failed: matched a fragment that cannot exist: %q", msg)
}
for _, bad := range r434ForbiddenFragments {
if strings.Contains(msg, bad) {
t.Errorf("alarm promises a recovery that cannot be performed (R-433): found %q in %q", bad, msg)
}
}
for _, want := range r434RequiredFragments {
if !strings.Contains(msg, want) {
t.Errorf("alarm is missing the withdrawal wording: want %q in %q", want, msg)
}
}
}
// TestR434_AlarmStillDoesNotClaimDataLoss — the OTHER direction, and the reason the fix is a
// withdrawal rather than a reversal.
//
// After R-433 the temptation is to swing to "your backups are gone". That is still usually FALSE:
// the snapshots exist and hold the older copy; what is unproven is our route to them. An alarm that
// over-claims loss sends an operator into a destructive recovery they did not need — which is the
// failure the v0.111.0 comment was written to prevent, and it must survive its own correction.
func TestR434_AlarmStillDoesNotClaimDataLoss(t *testing.T) {
msg := r434Message(t)
for _, bad := range []string{
"data is lost", "backups are gone", "data has been lost", "permanently lost", "unrecoverable",
} {
if strings.Contains(msg, bad) {
t.Errorf("alarm over-claims loss: found %q in %q", bad, msg)
}
}
// The one phrase that must appear NEGATED, never bare. A bare "confirmed data loss" would read
// as a verdict; the shipped sentence only ever uses it inside "neither ... nor".
if strings.Contains(msg, "confirmed data loss") &&
!strings.Contains(msg, "neither confirmed data loss nor confirmed recovery") {
t.Errorf("the phrase 'confirmed data loss' appears outside its negation: %q", msg)
}
}
// TestR434_StoredEventCarriesTheSameSentence — the delivered message and the stored one are the same
// string today, and a future refactor that formats them separately must not let them drift: the
// operator reads the mail, but every later audit reads the stored row.
func TestR434_StoredEventCarriesTheSameSentence(t *testing.T) {
st := newDiskStore(t)
var delivered string
saveOffsiteReport(t, st, "victim", dropJSON(69, true, "", "ok"))
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, msg, _, _ string) {
if et == "offsite_snapshots_dropped" {
delivered = msg
}
}, quietLog())
saveOffsiteReport(t, st, "victim", dropJSON(4, true, "", "ok"))
oc.Check()
evs, err := st.GetRecentEvents("victim", 50)
if err != nil {
t.Fatalf("GetRecentEvents: %v", err)
}
var stored []string
for _, e := range evs {
if e.EventType == "offsite_snapshots_dropped" {
stored = append(stored, e.Message)
}
}
if len(stored) != 1 {
t.Fatalf("want exactly 1 stored offsite_snapshots_dropped, got %d", len(stored))
}
if stored[0] != delivered {
t.Errorf("stored and delivered messages have drifted:\n stored: %q\n delivered: %q", stored[0], delivered)
}
if strings.Contains(stored[0], "recoverable file-by-file") {
t.Errorf("the stored row still carries the withdrawn promise: %q", stored[0])
}
}