hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.
was: "...still hold the older copy, so this is recoverable file-by-file; it is NOT
confirmed data loss. Check whether a deletion ran on the box before restoring."
now: "...still hold the older copy. The route back out of them is not yet established,
so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
before restoring anything, and check whether a deletion ran on the box."
It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.
Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.
R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.
THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.
Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.
R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.
Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
@@ -238,6 +238,19 @@ func (oc *OffsiteChecker) isStale(customerID string, off *offsiteReport) bool {
|
||||
// reached by retention; the floor of 5 stops a tiny-count box alarming on ordinary ageing. It is
|
||||
// deliberately NOT sensitive — a detector that cries wolf is switched off within a fortnight, and
|
||||
// this project has proved that twice in a week.
|
||||
// WHAT THIS DETECTOR DOES NOT SEE — R-435, and it must be read wherever "an unexplained fall is
|
||||
// noticed within a day" is claimed, because that claim is true only of falls above the fraction.
|
||||
//
|
||||
// **It sees a MASS deletion. It does not see ONE APP being wiped.** Worked on the live fleet
|
||||
// 2026-09-01: demo-hp's baseline is 69 snapshots across 9 apps, so ~35 must go before this speaks;
|
||||
// one app's tag is ~9 and is invisible. And `offbox.go:1388` runs `forget --prune` **grouped by
|
||||
// host,tags** — a per-tag wipe is exactly the shape a faulty retention or a targeted deletion
|
||||
// produces, so the blind spot sits on the most likely single-app failure, not an exotic one.
|
||||
//
|
||||
// THIS IS DELIBERATE AND THE THRESHOLD SHOULD NOT BE LOWERED TO "FIX" IT. The reasoning is below: a
|
||||
// detector that cries wolf is switched off within a fortnight, and this project has proved that
|
||||
// twice in a week. What is NOT acceptable is claiming coverage this does not have. Anyone adding
|
||||
// per-app detection should add a SECOND signal keyed on the per-tag count, not move these numbers.
|
||||
const (
|
||||
snapshotDropFraction = 0.5 // more than half the history gone in one step
|
||||
snapshotDropFloor = 5 // and at least this many, so small counts do not twitch
|
||||
@@ -286,12 +299,25 @@ func (oc *OffsiteChecker) snapshotDropped(customerID string, off *offsiteReport)
|
||||
// THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually
|
||||
// false: the daily Storage Box snapshots are read-only to every account (proven, not cited) and hold
|
||||
// the older copy. It says what happened, what it means, and where the data still is.
|
||||
//
|
||||
// AND IT MUST NOT SAY THE DATA IS RECOVERABLE EITHER — R-434, fixed 2026-09-01, hub v0.111.1. The
|
||||
// sentence shipped that morning promised "so this is recoverable file-by-file". Measured the same day
|
||||
// (R-433): no snapshot is reachable from a sub-account by ANY name — 777,600 exact names in the
|
||||
// vendor's own format over nine days, zero hits, with a passing control; `/home` and `/.zfs` are
|
||||
// different filesystems and `/home/.zfs` does not exist. So the promise named a route nobody can walk.
|
||||
//
|
||||
// THE FIX IS A DELETION, NOT A REPLACEMENT, AND THAT IS THE WHOLE POINT. R-434's row said the fix was
|
||||
// blocked on R-433 — on knowing what IS true. It is not, if the promise is simply withdrawn: a
|
||||
// sentence that asserts neither loss nor recovery is true under EVERY possible answer to the provider
|
||||
// question, so it never needs a second rewrite. An alarm rewritten twice in a week is worse than one
|
||||
// rewritten once, because the operator learns its words do not mean anything.
|
||||
func (oc *OffsiteChecker) emitSnapshotDrop(customerID string, off *offsiteReport, prev, cur int) {
|
||||
message := fmt.Sprintf(
|
||||
"Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than "+
|
||||
"retention can explain. The daily Storage Box snapshots are read-only and still hold the "+
|
||||
"older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check "+
|
||||
"whether a deletion ran on the box before restoring anything.",
|
||||
"older copy. The route back out of them is not yet established, so treat this as neither "+
|
||||
"confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and "+
|
||||
"check whether a deletion ran on the box.",
|
||||
customerID, prev, cur)
|
||||
details, _ := json.Marshal(map[string]any{
|
||||
"customer_id": customerID, "previous_count": prev, "current_count": cur,
|
||||
|
||||
Reference in New Issue
Block a user